Abstract
Cardiovascular disease prediction using structured clinical data is commonly formulated as a binary classification problem, while multi-class prediction is more challenging because of overlapping clinical characteristics and class imbalance. This study presents a hybrid ensemble and imbalance-aware machine-learning framework for binary and five-class heart-disease prediction using the combined UCI Heart Disease dataset comprising 920 records from Cleveland, Hungary, Switzerland, and VA Long Beach. Six classifiers—k-Nearest Neighbors, Decision Tree, Logistic Regression, Support Vector Machine, Random Forest, and XGBoost—are evaluated together with majority-voting and stacking ensembles. To address the highly imbalanced five-class distribution, SMOTE is applied exclusively to the training folds during stratified cross-validation. A confidence-aware hierarchical voting mechanism is also introduced to resolve classifier disagreements by retaining majority decisions when consensus exists and using posterior-confidence information in ambiguous cases. On the original imbalanced five-class data, XGBoost achieves the highest accuracy of 86.74%. After SMOTE-based balancing, stacking achieves the best overall performance, with 94.94% accuracy, 94.92% precision, 94.94% recall, and 94.92% F1-score, closely followed by majority voting. Feature-importance analysis identifies chest-pain type, maximum heart rate, ST depression, and number of major vessels as influential predictive attributes. The results demonstrate the potential of imbalance-aware heterogeneous ensembles for robust multi-class heart-disease prediction, while the proposed confidence-aware voting strategy provides an explicit mechanism for resolving classifier disagreement.
1. Introduction
Cardiovascular disease (CVD) remains a leading cause of mortality worldwide and a major challenge for healthcare systems [1,2]. Accurate identification of cardiovascular abnormalities is important for supporting clinical assessment and risk management. Conventional diagnostic procedures can provide valuable information but may require specialized resources and additional costs [1,3]. Consequently, machine learning (ML) has received considerable attention as a complementary approach for analyzing structured clinical data and identifying relationships among demographic, physiological, laboratory, and exercise-related cardiovascular variables [2,4].
Various supervised ML techniques have been investigated for heart-disease prediction, including Logistic Regression (LR), k-Nearest Neighbors (kNN), Decision Tree (DT), Support Vector Machine (SVM), Random Forest (RF), and XGBoost [2,5,6,7,8,9,10,11]. Recent studies have further incorporated feature engineering, class balancing, voting and stacking ensembles, and explainable artificial intelligence to improve predictive performance and interpretability [4,6,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27]. These developments demonstrate the continued effectiveness of ML for tabular clinical data when appropriate preprocessing and ensemble strategies are employed.
Despite these advances, several challenges remain. First, most heart-disease studies formulate prediction as a binary classification problem. Although useful for screening, binary prediction does not differentiate among multiple disease categories. Studies considering multi-class or severity-level classification have demonstrated that distinguishing several categories is substantially more difficult because their clinical characteristics may overlap [19,20].
Second, clinical datasets frequently exhibit substantial class imbalance, particularly for less frequent disease categories [13,14]. Models trained on imbalanced data may favor majority classes and inadequately recognize minority cases. Resampling techniques such as the Synthetic Minority Over-sampling Technique (SMOTE) can improve minority-class learning [13,14,19,20,21,22,23], but the problem becomes more challenging when several minority categories coexist in a multi-class setting.
Third, heterogeneous classifiers may produce conflicting predictions for the same observation. Ensemble methods, including majority voting and stacking, improve robustness by combining complementary models [4,10,11,12,15,22,24,25,26,27,28,29]. However, majority voting follows a fixed aggregation rule, whereas stacking relies on a learned meta-classifier. When classifiers strongly disagree, incorporating prediction confidence can provide additional information for resolving ambiguous decisions. Recent uncertainty-aware ensemble approaches further support the usefulness of considering prediction reliability in addition to class labels [26].
To address these challenges, this study presents a hybrid ensemble and imbalance-aware ML framework for binary and five-class heart-disease prediction using tabular clinical data. The combined UCI Heart Disease dataset contains 920 records from Cleveland, Hungary, Switzerland, and VA Long Beach [30] and exhibits substantial imbalance across its five target categories. The framework integrates data preprocessing, feature analysis, SMOTE-based imbalance handling, individual and ensemble classifiers, and a confidence-aware hierarchical voting mechanism.
The main contributions of this study are as follows:
- Six supervised ML algorithms—kNN, DT, LR, SVM, RF, and XGBoost—are systematically evaluated together with majority-voting and stacking ensembles for both binary and five-class classification.
- The effect of severe class imbalance on multi-class prediction is investigated, with SMOTE incorporated into the training process to improve minority-class representation.
- A confidence-aware hierarchical voting mechanism is introduced to resolve disagreements among heterogeneous classifiers. Majority decisions are retained when sufficient consensus exists, while posterior-confidence information is used for ambiguous predictions.
- The proposed framework provides a comprehensive comparison of individual classifiers, conventional ensembles, and confidence-aware decision-making for an imbalanced five-class heart-disease prediction problem.
Unlike many recent studies that primarily address binary heart-disease detection, this work jointly considers binary and five-class prediction while explicitly addressing class imbalance and classifier disagreement. Its contribution therefore lies not solely in numerical accuracy but in achieving competitive predictive performance for a more granular and imbalanced multi-class problem while providing an explicit confidence-aware strategy for resolving conflicting ensemble predictions.
2. Related Work
2.1. Machine Learning Approaches for Heart Disease Prediction
Machine learning (ML) has been widely investigated for heart-disease prediction using structured demographic, physiological, laboratory, and exercise-related variables [1,2]. Common approaches include Logistic Regression (LR), k-Nearest Neighbors (kNN), Decision Tree (DT), Support Vector Machine (SVM), Random Forest (RF), and boosting algorithms. LR offers interpretability and computational efficiency [5], while SVM supports nonlinear classification through kernel transformations [12,13]. DT provides interpretable rules but may be susceptible to overfitting [9,10,14], whereas RF and XGBoost improve robustness and capture nonlinear feature interactions through ensemble learning [6,31,32].
Recent studies emphasize the importance of feature selection and representation. Li et al. [18] combined feature-selection techniques with conventional classifiers for heart-disease identification. Gárate-Escamila et al. [19] investigated Chi-square feature selection and PCA using Cleveland, Hungarian, and combined Cleveland–Hungarian heart-disease data; their CHI-PCA approach with RF achieved strong classification performance. Dissanayake and Johar [20] compared multiple feature-selection strategies with several classifiers, while Pan et al. [21] examined the influence of categorical and numerical variables in ensemble frameworks. Together with entropy-based feature engineering [6], these studies demonstrate that conventional ML remains competitive for tabular heart-disease prediction when appropriate preprocessing and feature representation are employed.
2.2. Binary and Multi-Class Heart Disease Classification
Most heart-disease prediction studies formulate the task as binary classification, distinguishing disease from non-disease cases [2,4]. Although suitable for screening, binary classification does not differentiate among multiple disease categories and therefore provides a less granular prediction problem.
Kibria and Matin [22] directly investigated both binary and multi-class cardiovascular-disease prediction using several ML algorithms and three weighted score-based fusion models. Their study specifically addresses severity prediction and confirms that multi-class prediction is considerably more challenging than binary classification. Abdellatif et al. [23] similarly investigated heart-disease detection and severity-level classification using ML and hyperparameter optimization. These studies indicate that performance from binary and multi-class experiments should not be directly compared because multi-class prediction requires separation among several potentially overlapping categories. The present study therefore evaluates both settings using the combined UCI Heart Disease dataset [33], considering Class 0 versus Classes 1–4 for binary prediction and retaining all five original target categories for multi-class classification.
2.3. Class Imbalance and Imbalance-Aware Learning
Clinical datasets frequently exhibit class imbalance, which may cause classifiers to favor majority categories and inadequately recognize minority cases [16,30]. SMOTE is a widely used approach for addressing this problem by generating synthetic minority-class observations within the feature space.
Recent cardiovascular studies have incorporated different balancing strategies. Kibria and Matin [22] considered oversampling in multi-class cardiovascular prediction, while Abdellatif et al. [23] addressed imbalance in severity-level classification. Zhang et al. [24] evaluated class-balancing methods with LightGBM for coronary artery disease prediction, whereas El-Sofany et al. [26] incorporated balancing techniques within an ML and explainable-AI framework.
Class imbalance is particularly relevant to the present dataset, where the five original target categories contain 411, 265, 109, 107, and 28 observations, respectively. The proposed framework therefore compares learning from the original distribution with SMOTE-based training. SMOTE is applied only to training folds during cross-validation, preventing synthetic samples from entering the corresponding validation folds. Unlike many binary heart-disease studies, the present work investigates imbalance across five target categories.
2.4. Ensemble Learning, Voting, and Stacking
Ensemble learning combines multiple predictors to improve robustness and generalization. RF and XGBoost represent homogeneous bagging and boosting approaches [31,32], whereas heterogeneous ensembles integrate classifiers with different learning characteristics [29]. Majority voting selects the class receiving the greatest classifier support [15,17], while stacking learns how to combine base-model outputs through a meta-classifier.
Recent studies demonstrate the potential of these approaches for heart-disease prediction. Daza et al. [25] developed four stacking configurations with hyperparameter optimization using a dataset reported as containing 918 records and 12 attributes. Their best models achieved 88.24% test accuracy and 92.00% ROC-AUC, illustrating the potential benefit of combining stacking, oversampling, and model optimization. Ganie et al. [4] compared voting and stacking ensembles across multiple heart-disease datasets and incorporated explainable AI. Ghose et al. [27] combined feature engineering with stacked ensemble learning, while Ashika and Hannah Grace [34] integrated Rough Set Theory-based feature selection, nine ML classifiers, five stacking configurations, multi-criteria model ranking, hyperparameter optimization, and XAI. More recently, CardiaTics [35] combined feature engineering, multiple ML models, stacked ensemble learning, and explainable AI.
These studies demonstrate the effectiveness of heterogeneous ensembles. However, conventional voting generally uses a predetermined aggregation rule, while stacking learns a meta-level mapping from base-model outputs. This motivates the investigation of ensemble mechanisms that explicitly consider the agreement and disagreement patterns among classifiers when determining the final prediction.
2.5. Explainable Machine Learning for Heart Disease Prediction
Interpretability is increasingly important in clinical ML because predictive performance alone does not explain how input variables contribute to model decisions [3,4]. Tree-based feature-importance methods provide global measures of predictive relevance [10,11], while recent studies increasingly incorporate dedicated explainable-AI techniques.
El-Sofany et al. [26] combined heart-disease prediction with explainability analysis, while Ganie et al. [4] incorporated global and local explanations into voting and stacking ensembles. Ghose et al. [27] combined feature engineering, stacked ensembles, and XAI, and Ashika and Hannah Grace [34] also employed XAI to interpret their optimized RST-ML ensemble framework. CardiaTics [35] further integrates explainability with feature engineering and stacked ensemble learning.
In the present study, feature-importance analyses from DT, RF, and XGBoost are used to identify attributes contributing strongly to model predictions. Chest-pain type, maximum heart rate, ST depression, and number of major vessels emerge as influential predictive features. These results represent predictive associations rather than evidence of clinical causality.
2.6. Confidence- and Uncertainty-Aware Ensemble Decision Making
Recent ensemble research has extended conventional prediction aggregation by explicitly considering confidence and uncertainty. Majority voting generally treats classifier votes uniformly [15,17,29], which may be insufficient when heterogeneous classifiers strongly disagree.
Wang et al. [36] recently proposed an uncertainty-aware feature-weighted ensemble (UAFE) for tabular heart-disease prediction. Their framework combines feature stratification, disagreement-based uncertainty estimation, dynamic uncertainty weighting, and neighborhood refinement. It achieved an overall accuracy of 86.60%, while leave-one-center-out validation across four clinical sites yielded a mean accuracy of 83.32%. Importantly, UAFE dynamically adjusts model contributions according to uncertainty rather than relying solely on static voting.
The present study follows the broader confidence-aware direction but uses a different decision strategy. The proposed hierarchical voting mechanism first examines agreement among heterogeneous classifiers. A clear majority is accepted directly, while posterior-confidence information is invoked to resolve complete or partial disagreement. Thus, unlike uncertainty-based dynamic weighting [36], the proposed mechanism uses confidence selectively as a conflict-resolution criterion when conventional voting becomes ambiguous. This distinction provides the principal methodological positioning of the proposed voting strategy relative to recent uncertainty-aware ensemble research.
2.7. Research Gap and Positioning of the Proposed Framework
Existing studies demonstrate strong potential for conventional ML, feature engineering, and feature selection in tabular heart-disease prediction [5,6,7,8,9,10,11,18,19,20,21]. Imbalance-aware approaches address minority-class underrepresentation [16,22,23,24,26,30], while voting, stacking, XAI, and uncertainty-aware ensembles provide increasingly sophisticated prediction frameworks [4,15,17,25,26,27,29,34,35,36].
Nevertheless, much of the recent literature focuses on binary heart-disease prediction, while fewer studies explicitly address multiple disease categories [22,23]. Moreover, comparatively limited research jointly considers five-class prediction, severe class imbalance, heterogeneous ensemble learning, and explicit confidence-based resolution of classifier disagreement.
The present study addresses these issues using the combined 920-record UCI Heart Disease dataset from Cleveland, Hungary, Switzerland, and VA Long Beach [33]. It systematically evaluates six ML classifiers together with majority voting and stacking under binary and five-class settings, incorporates SMOTE for imbalance handling, and introduces a confidence-aware hierarchical voting mechanism for resolving ambiguous ensemble decisions.
Accordingly, the contribution of this work should not be interpreted solely through direct numerical comparison with previous studies, particularly because datasets, class definitions, validation procedures, and evaluation protocols differ considerably across the literature. Instead, the study provides competitive predictive performance for a granular and highly imbalanced five-class heart-disease problem while introducing an explicit confidence-aware mechanism for resolving conflicting predictions among heterogeneous classifiers.
3. Classification Models, Data Handling, and Evaluation
This study proposes a comprehensive machine learning framework for cardiovascular disease (CVD) severity prediction using both single-model and ensemble-based classification strategies. The framework integrates data preprocessing, imbalance-aware learning, feature engineering, cross-validation, and advanced ensemble decision mechanisms to improve predictive reliability across both binary and multi-class classification settings.
3.1. Data Description
This study utilized the Heart Disease dataset from the UCI Machine Learning Repository [33] using the version curated on Kaggle by Redwankarimsony. The dataset contains clinical and demographic information collected from multiple medical institutions located in Cleveland, Hungary, Switzerland, and the VA Long Beach medical centers. Due to its comprehensive cardiovascular attributes and established benchmark status, the dataset has been extensively used in cardiovascular disease prediction research.
The dataset consists of 920 patient records with 14 clinical features, as shown in Table 1 and Table 2, including numerical and categorical variables associated with cardiovascular risk assessment. Numerical features include age, resting blood pressure, serum cholesterol, maximum heart rate, ST depression, and the number of major vessels. Categorical features include sex, chest pain type, fasting blood sugar, resting electrocardiographic results, exercise-induced angina, ST slope, thalassemia status, and hospital source.
Table 1.
Numerical features of the input data.
Table 2.
Categorical features of the input data.
The target variable represents cardiovascular disease severity levels ranging from 0 to 4, where Class 0 indicates the absence of cardiovascular disease, and Classes 1–4 represent increasing severity levels of cardiovascular disease. To comprehensively evaluate the proposed framework, two classification scenarios were investigated:
- Binary classification, where the target variable was grouped into non-disease (Class 0), and disease (Classes 1–4).
- Multi-class classification, where all five severity levels were preserved to evaluate severity stratification capability.
The original dataset exhibits substantial class imbalance, particularly for severe disease classes (Classes 3 and 4), making it suitable for evaluating imbalance-aware learning techniques.
3.2. Single-Model Classification Approaches
Four supervised machine learning algorithms were initially investigated as baseline classifiers to evaluate the discriminative capability of conventional learning methods for cardiovascular disease prediction.
3.2.1. k-Nearest Neighbors (kNN)
The k-Nearest Neighbors (kNN) algorithm is a non-parametric distance-based classifier that predicts the target class based on the majority class among the k closest training samples in the feature space. In this study, distance similarity was measured using the Euclidean distance metric. The model assumes that samples belonging to similar classes are located in close proximity within the multidimensional feature space.
Due to its simplicity and effectiveness in low-dimensional classification problems, kNN serves as an important baseline model for comparison with more advanced classifiers. However, its performance is highly dependent on feature scaling and neighborhood selection.
3.2.2. Decision Tree (DT)
Decision Tree (DT) is a hierarchical rule-based classification model that recursively partitions the feature space into homogeneous subsets using impurity-based splitting criteria [9,10]. In this study, node partitioning was performed using the Gini impurity measure to maximize class separation at each decision node. Decision Tree iteratively selects the feature and threshold that minimize impurity, thereby generating interpretable classification rules that are particularly valuable in clinical applications. Despite its interpretability advantages, DT models are susceptible to overfitting, especially when trained on noisy or highly variable datasets. Figure 1 illustrates the general structure and terminology of a decision tree classifier.
Figure 1.
Structure and terminology of Decision Tree.
3.2.3. Logistic Regression (LR)
Logistic Regression (LR) is a probabilistic linear classification model that estimates posterior class probabilities through the logistic sigmoid function. The logistic response function [28] is defined as Equation (1). The model constructs linear decision boundaries by learning weighted relationships between input features and class labels.
Because of its statistical interpretability and computational efficiency, LR is widely adopted in medical prediction tasks. However, its linear assumptions may limit its ability to model complex nonlinear interactions commonly observed in cardiovascular disease data.
3.2.4. Support Vector Machine (SVM)
Support Vector Machine (SVM) is a supervised discriminative learning model that identifies an optimal separating hyperplane with maximum margin between classes [12]. Both linear and nonlinear kernel-based SVMs were investigated in this study.
For linearly separable data, the separating hyperplane is represented by Equation (2):
where denotes the weight vector, and b represents the bias term.
To address nonlinear relationships among cardiovascular risk factors, kernel functions were employed to map the original feature space into higher-dimensional representations [12,13]. In this work, both linear and Radial Basis Function (RBF) kernels were evaluated. Figure 2 presents the concept of data transformation in SVM, where nonlinear data distributions become linearly separable in a transformed feature space.
Figure 2.
Data transformation with the SVM kernel function.
3.3. Ensemble Learning Models
To improve predictive robustness and reduce model variance, multiple ensemble learning strategies were investigated.
3.3.1. Random Forest (RF)
Random Forest (RF) is a bagging-based ensemble learning method that combines multiple decision trees trained on randomly sampled subsets of the training data. Final predictions are obtained through majority voting across all trees.
The use of bootstrap aggregation reduces variance and improves model generalization compared with individual decision trees. In addition, random feature selection during node splitting enhances diversity among trees and mitigates overfitting.
Feature importance in RF was estimated using the Mean Decrease Impurity (MDI) approach, which quantifies the average reduction in Gini impurity contributed by each feature across all trees.
3.3.2. Extreme Gradient Boosting (XGBoost)
Extreme Gradient Boosting (XGBoost) is an optimized gradient boosting framework that sequentially trains weak learners to minimize classification error through gradient optimization and regularization.
Unlike bagging-based approaches, XGBoost iteratively constructs decision trees where each new tree focuses on correcting the residual errors produced by previous learners. The objective function incorporates both loss minimization and regularization terms to improve generalization and reduce overfitting [32].
Feature importance was computed based on information gain accumulated across all tree splits involving each feature.
3.3.3. Majority Voting Ensemble
Majority voting is a conventional ensemble strategy that aggregates predictions from multiple base classifiers and selects the class receiving the highest number of votes [15,17,29].
Although this method improves prediction stability relative to single models, conventional majority voting may fail when classifiers produce conflicting predictions, particularly in multi-class classification scenarios with overlapping feature distributions.
3.3.4. Hierarchical Voting Mechanism
Conventional majority (hard) voting often fails to reach a definitive consensus due to voting ties [15,29]. This limitation is manifested in two structural edge cases where the mode cannot be determined:
- Unresolvable ties in even-numbered binary ensembles (a 50/50 split);
- Multimodal ties in multi-class problems, where either several classes receive an identical maximum vote count or all classifiers predict different classes entirely.
Such ties are particularly common in multi-class classification tasks involving overlapping clinical characteristics. To address this limitation, this study proposes a hierarchical voting mechanism that integrates majority voting with confidence-aware decision rules to guarantee a deterministic outcome. The proposed framework operates through three hierarchical decision stages, as follows.
- If a clear majority class exists among the base learners, the majority class is selected directly.
- If all classifiers predict different classes, the class associated with the highest posterior probability is selected.
- In partially conflicting cases where no dominant majority exists, the final prediction is determined using the highest average confidence score among supporting classifiers.
By incorporating classifier confidence into the decision process, the proposed mechanism guarantees a deterministic prediction in every case by construction, thereby eliminating the undefined outcomes that conventional hard voting produces under tie conditions. A quantitative comparison against standard hard and soft voting strategies is left for future work.
The predicted class labels obtained from the majority vote are shown in Equation (3):
where is the predicted class for the observation from the classifier, and K represents the number of classifiers.
In multi-class classification analysis, there is a possibility that the normal majority voting method fails to produce a reliable prediction, such as the model predicting different classes, or when there is uniformity in many classes. Therefore, this research proposes a hierarchical voting mechanism. Figure 3 presents the operational workflow of the proposed hierarchical voting mechanism. The framework combines majority voting with confidence-based decision rules (posterior probabilities) to resolve conflicts among base classifiers and improve prediction stability in multi-class classification.
Figure 3.
Operation of the hierarchical voting mechanism.
When the prediction results of the sub-models are inconsistent, the algorithm starts by receiving the predicted target classes from K classifiers in the form of a vector.
and the model confidence (posterior probability) of the model set of each model is as follows.
where represents the predicted class for the observation from the classifier, and represents the model confidence (posterior probability) of the model for classifier.
The first step of the decision is based on the majority voting principle by checking which class is the mode of the set of prediction results. If a mode is found, that class is immediately assigned as the result, as in Equation (4):
However, if the mode value cannot be found, two cases are considered. In the first case, when the prediction values from the model are completely different, the target class is predicted by considering the class with the highest model confidence (posterior probability), as shown in Equation (5):
Another case is when some classes are predicted more than once but not enough to produce a mode. The algorithm selects the class with the highest mean of the model confidence (posterior probability) from the models predicting that class, as in Equation (6):
3.3.5. Stacking Classifier
A stacking classifier is a stacking ensemble technique that combines the outputs of multiple base learners and then feeds the result to a meta-learner to learn how to combine the data to produce the best results by using posterior probabilities from each model as features [15]. The new features obtained can be defined as in Equation (7):
where represents posterior probabilities that the label of target class is c given input data processed by the model.
The resulting new feature is used as an input feature for the meta-learner to predict the target class, which can be represented by Equation (8):
3.4. Synthetic Minority Over-Sampling Technique
The class imbalance problem typically occurs in classification tasks when the proportions of the target classes differ significantly. This disparity hinders supervised learning algorithms from adequately learning the decision boundaries and characteristics of minority class samples. In such scenarios, standard classifiers tend to be overwhelmed by the majority classes, leading to a biased optimization that ignores the minority instances [16,30]. This issue is highly prevalent in critical domains such as fraud detection, abnormal event detection, and clinical disease diagnosis, where the class of primary interest (e.g., presence or high severity of a disease) constitutes only a small fraction of the overall dataset.
To address these limitations, the Synthetic Minority Over-sampling Technique (SMOTE) is widely employed to rebalance the class distribution [30]. SMOTE mitigates class imbalance by generating synthetic examples for the minority class rather than relying on simple over-sampling with replacement. This is achieved through spatial interpolation: for each minority instance, synthetic samples are created along the line segments joining any or all of its k-nearest neighbors within the same class in the feature space. As illustrated in Figure 4, adjusting the dataset with SMOTE yields a balanced class representation, preventing the learning process from favoring the majority class and thereby enhancing classification performance for minority severity categories.
Figure 4.
(a) Before applying SMOTE. (b) After applying SMOTE.
3.5. Performance Metrics
Model performance was evaluated using multiple classification metrics derived from the confusion matrix.
3.5.1. Confusion Matrix
A confusion matrix in multi-class classification which has M classes is an M×M table that summarizes the number of samples classified into each class using frequency, showing the relationship between the actual and predicted values. In binary classification, the confusion matrix consists of four values: 1. True Positive (TP), predicted as positive and positive; 2. True Negative (TN), predicted as negative and negative; 3. False Positive (FP), predicted as positive but negative; and 4. False Negative (FN), predicted as negative but positive. These values are important in calculating the following evaluation metrics.
3.5.2. Accuracy
Accuracy measures the overall proportion of correctly classified samples. Accuracy can be calculated using Equation (9):
3.5.3. Precision
Precision measures the accuracy of the positive class prediction, considering the probability that the model correctly predicts that the instance is in the positive class. It is useful in applications where false positive predictions are costly, such as fraud detection and spam filtering. It can be calculated using Equation (10):
3.5.4. Recall
Recall (or sensitivity, True Positive Rate) measures the ability to detect all positives in real data. It is suitable for tasks where positives cannot be missed, such as disease screening and warning systems. It can be calculated using Equation (11):
3.5.5. F1-Score
The F1-Score is the harmonic mean of precision and recall, useful when precision and recall need to be balanced, especially for data where the target class has imbalance issues. It can be calculated using Equation (12):
In multi-class classification, Macro F1-Score, Weighted F1-Score, and Micro Average F1-Score are also evaluated to provide a comprehensive assessment across all severity levels.
The Macro F1-Score calculates the F1 value for each class and then averages it without weighting. The equation of the Macro F1-Score can be written as in Equation (13):
The Weighted F1-Score calculates the F1 value of each class and weights it by the proportion of samples in that class. Equation (14) represents the Weighted F1-Score:
where
The Micro Average F1-Score calculates the F1 value by first summing the TP, FP, and FN values of all classes, followed by calculating the precision and recall, as shown in Equations (15) and (16), respectively:
Then, the Micro Average F1-Score value is given by Equation (17):
The Micro Average F1-Score provides the same results as the accuracy in multi-class classification. This is suitable for imbalanced data and reflects performance based on the total number of samples.
3.6. Implementation Details
All experiments were implemented using Python 3.12.4 with the Scikit-learn and XGBoost libraries. Hyperparameter tuning was performed using a grid search combined with stratified k-fold cross-validation.
To ensure reproducibility, a fixed random seed was used across all experiments, and SMOTE was applied only to the training folds to prevent data leakage. In addition, feature scaling was performed using standardization prior to model training.
4. Experimental Design
4.1. Data Preprocessing and Feature Engineering
To improve data quality, reduce model bias, and enhance predictive performance, several preprocessing and feature engineering procedures were applied prior to model training.
4.1.1. Missing Value Imputation
The dataset contains missing values in both numerical and categorical variables. Numerical missing values were imputed using the class-wise mean value, while categorical missing values were replaced using the class-wise mode. This strategy preserved class-specific statistical characteristics and reduced potential bias caused by global imputation.
4.1.2. Categorical Feature Encoding
Categorical variables were transformed into numerical representations using one-hot encoding. This approach converts categorical attributes into binary indicator variables while preventing ordinal assumptions among categories. To avoid multicollinearity and redundant feature representation, one category from each encoded feature group was selected as the reference category and excluded from the transformed feature space. As a result of one-hot encoding, the dimensionality of the input feature space increased beyond the original 14 variables.
4.1.3. Feature Standardization
Since the dataset contains features measured on different numerical scales, feature standardization was applied before model training. Each feature was transformed into a zero-mean, unit-variance distribution using Equation (18):
where denotes the original feature value at the observation, is the feature mean, and represents the standard deviation of the feature, respectively.
Feature scaling is particularly important for distance-based and gradient-based algorithms such as kNN, SVM, and Logistic Regression, where differences in feature magnitude may significantly affect model behavior.
4.1.4. Feature Importance Analysis
Feature importance analysis was conducted using tree-based models, including Decision Tree, Random Forest, and XGBoost. These methods estimate feature relevance based on impurity reduction and information gain. The analysis was performed to
- Identify clinically significant cardiovascular risk factors;
- Improve interpretability of the prediction framework;
- Evaluate the relative contribution of individual variables to disease severity prediction.
Feature importance scores were estimated based on impurity reduction and information gain accumulated during the tree construction process.
4.2. Data Partitioning and Cross-Validation
To ensure reliable model evaluation and reduce performance variance caused by dataset partitioning, stratified k-fold cross-validation, with k set to five, was employed throughout all experiments.
Figure 5 illustrates the data partitioning procedure using stratified k-fold cross-validation. In this process, the dataset is partitioned into k-approximately equal subsets (folds). During each iteration, one fold is used as the validation set, while the remaining (k − 1) folds are used for model training. This process is repeated k times so that each fold serves as the validation set exactly once.
Figure 5.
Data partitioning using k-fold cross-validation.
Stratification was applied to preserve the original class distribution within each fold, which is particularly important for imbalanced multi-class classification problems.
Performance metrics from all folds were aggregated to obtain stable and unbiased evaluation results. Confusion matrix statistics were accumulated across all validation iterations to compute overall classification performance measures.
Several measures were employed to reduce the risk of overfitting. Stratified k-fold cross-validation was used to evaluate model performance across multiple data partitions rather than relying on a single training–test split. For the imbalanced multi-class experiments, SMOTE was applied exclusively to the training folds within the cross-validation procedure, while the corresponding validation folds remained unchanged, thereby preventing synthetic samples from leaking into model evaluation. In addition, regularized and ensemble classifiers, including Random Forest and XGBoost, were evaluated alongside heterogeneous voting and stacking approaches to improve generalization. Nevertheless, because no independent external test dataset was available, the reported cross-validation performance should be interpreted as internal validation, and further external validation is required to confirm model generalizability.
4.3. Experimental Design Process
The overall experimental workflow consisted of multiple sequential stages, including data preprocessing, feature engineering, imbalance handling, model training, and performance evaluation.
Initially, the dataset was imported and subjected to preprocessing procedures including missing value imputation, categorical encoding, and feature standardization.
Following preprocessing, the dataset was partitioned using stratified k-fold cross-validation. Two classification scenarios were investigated.
4.3.1. Binary Classification
In the binary classification setting, the target labels were grouped into:
- Non-disease class (Class 0);
- Disease class (Classes 1–4).
Since the binary dataset exhibits relatively balanced class proportions (411 non-disease samples versus 509 disease samples), imbalance handling was not applied in that experimental setting.
Table 3 presents the machine learning models evaluated in this study, along with their respective hyperparameter search spaces used for the grid search.
Table 3.
Hyperparameter search space for all the classification models.
4.3.2. Multi-Class Classification
In the multi-class setting, all five cardiovascular severity classes were preserved to evaluate severity stratification capability.
Two experimental conditions were investigated:
- Multi-class classification using the original imbalanced dataset.
- Multi-class classification after with SMOTE.
For the balanced-data experiment, SMOTE was applied exclusively to the training folds within each cross-validation iteration to prevent data leakage and ensure unbiased evaluation.
The same set of machine learning and ensemble learning models used in binary classification was also evaluated for the multi-class setting.
Performance comparison across all models was conducted using accuracy, precision, recall, F1-score, and additional class-wise evaluation metrics.
Figure 6 presents the complete flowchart of the proposed experimental framework, including preprocessing, feature engineering, imbalance-aware learning, model training, ensemble classification, and performance evaluation stages.
Figure 6.
Flowchart of the experimental design process.
4.4. Experimental Objectives
The experimental design was developed to investigate several key research objectives:
- Compare the effectiveness of conventional machine learning classifiers and ensemble learning approaches for cardiovascular disease prediction.
- Evaluate the impact of imbalance-aware learning using SMOTE on minority severity class detection.
- Assess the effectiveness of the proposed hierarchical voting mechanism in resolving prediction conflicts among ensemble classifiers.
- Investigate the ability of the proposed framework to perform clinically meaningful multi-class severity stratification.
- Analyze the influence of feature engineering and feature importance analysis on predictive performance and interpretability.
5. Experimental Results and Discussion
5.1. Binary Classification Results
This section presents the experimental results obtained from binary cardiovascular disease classification, where the target labels were grouped into two categories: non-disease (Class 0) and disease (Classes 1–4). The experiments aimed to evaluate the predictive performance of both conventional machine learning classifiers and ensemble learning approaches under relatively balanced class conditions.
Before model training, feature correlation and feature importance analyses were conducted using Decision Tree, Random Forest, and XGBoost models to identify the most influential clinical variables associated with cardiovascular disease prediction.
Figure 7 and Figure 8 illustrate the correlation analysis and feature importance rankings obtained from the tree-based models.
Figure 7.
Feature correlation with binary target class.
Figure 8.
(a) Feature importance score from Decision Tree. (b) Feature importance score from Random Forest. (c) Feature importance score from XGBoost.
The analysis consistently identified several clinically relevant variables as dominant predictive factors, including
- Chest pain type;
- Maximum heart rate achieved;
- ST depression induced by exercise;
- The number of major vessels.
These findings align with the established cardiovascular risk assessment literature and support the clinical validity of the proposed framework. The expanded feature space observed in the analysis results from the one-hot encoding of categorical variables.
Overall, the feature importance analysis demonstrates that physiologically meaningful cardiovascular indicators contribute most significantly to disease prediction, thereby improving both predictive performance and interpretability.
5.1.1. Hyperparameter Configuration
The hyperparameters of all binary classification models were optimized using grid-search-based tuning combined with stratified cross-validation. The final parameter configurations are summarized in Table 4. The evaluated classifiers included k-Nearest Neighbors (kNN), Decision Tree (DT), Random Forest (RF), Logistic Regression (LR), Support Vector Machine (SVM), XGBoost, majority voting ensemble, and stacking ensemble.
Table 4.
Parameter setting for all binary classification models.
5.1.2. Binary Classification Performance
The quantitative performance results obtained from all binary classification models are presented in Table 5.
Table 5.
Results obtained from all the binary classification models.
The experimental results indicate that all classifiers achieved strong predictive performance, with overall classification accuracies ranging from approximately 92.9% to 94.7%. The consistently high performance across all evaluation metrics demonstrates the suitability of machine learning methods for cardiovascular disease screening tasks.
Among the evaluated models, XGBoost achieved the highest classification accuracy and precision, while the majority voting ensemble achieved the highest recall and F1-score.
The superior performance of XGBoost can be attributed to its gradient boosting architecture, which effectively captures nonlinear feature interactions and difficult classification patterns. In contrast, the majority voting ensemble benefits from the complementary strengths of multiple base learners, improving sensitivity and overall classification stability.
The stacking ensemble did not outperform XGBoost or majority voting despite integrating multiple classifiers. This behavior may result from insufficient diversity among the base learners and limitations of the meta-classifier in modeling complex inter-model dependencies.
5.1.3. ROC Curve Analysis
In addition to accuracy-based metrics, the Area Under the Receiver Operating Characteristic Curve (AUC–ROC) was evaluated to assess the discriminative capability of each classifier across varying decision thresholds. Figure 9 presents the ROC curves obtained from all binary classification models.
Figure 9.
AUC-ROC results obtained from the binary classification models.
The results indicate that ensemble-based approaches, particularly XGBoost and the proposed hybrid framework, achieved superior AUC performance compared with single classifiers. High AUC values demonstrate strong separability between patients with and without cardiovascular disease and confirm the robustness of the proposed framework for clinical decision-support applications.
From a clinical perspective, strong AUC performance is particularly important because it reflects reliable disease discrimination across varying sensitivity and specificity requirements.
5.2. Multi-Class Cardiovascular Disease Severity Classification
To evaluate the ability of the proposed framework to perform clinically meaningful severity stratification, multi-class classification experiments were conducted using five cardiovascular disease severity levels represented in Figure 10. In this setting, Class 0 represents patients without cardiovascular disease, while Classes 1–4 represent progressively increasing disease severity levels.
Figure 10.
Data distribution for the 5 output classes.
Because the dataset exhibits substantial class imbalance, particularly for severe disease classes, two experimental conditions were investigated:
- Multi-class classification using the original imbalanced dataset.
- Multi-class classification after with SMOTE.
5.3. Multi-Class Classification Using the Original Imbalanced Dataset
The original class distribution of the five severity categories is presented in Figure 10. The dataset is strongly skewed toward lower severity levels, with Class 0 containing the largest number of samples and Class 4 containing the fewest observations. Such an imbalance introduces substantial classification difficulty because learning algorithms tend to favor majority classes during training.
5.3.1. Feature Analysis Under Imbalanced Conditions
Feature correlation and feature importance analyses were again performed for the multi-class classification setting using Decision Tree, Random Forest, and XGBoost models. Figure 11 and Figure 12 illustrate the feature correlation and feature importance results.
Figure 11.
Feature correlation with the original imbalanced multi-class classification.
Figure 12.
Comparison of feature importance of the input features obtained from different techniques (a) Feature importance score from Decision Tree for original imbalanced multi-class classification. (b) Feature importance score from Random Forest for original imbalanced multi-class classification. (c) Feature importance score from XGBoost for original imbalanced multi-class classification.
The analysis consistently identified the number of major vessels, chest pain type, maximum heart rate, ST depression, and age as the most influential variables for cardiovascular severity prediction.
These findings indicate that physiologically relevant features remain dominant predictors even under imbalanced multi-class conditions.
5.3.2. Multi-Class Hyperparameter Optimization
The optimized hyperparameter settings for all multi-class classification models are summarized in Table 6. Compared with the binary classification setting, several models required modified hyperparameter configurations to improve performance under increased classification complexity and class imbalance.
Table 6.
Parameter settings for the multi-class classification.
5.3.3. Multi-Class Performance Under Imbalanced Conditions
The quantitative results for all multi-class classifiers trained on the original imbalanced dataset are presented in Table 7.
Table 7.
Performance evaluation matrices obtained for the multi-class classification.
The experimental results demonstrate that multi-class classification is substantially more challenging than binary classification, with accuracy values ranging from approximately 73% to 86.7%.
Among the evaluated models, XGBoost achieved the highest classification accuracy, while the stacking ensemble achieved the best overall balance among precision, recall, and F1-score metrics.
The strong performance of XGBoost highlights the effectiveness of boosting-based methods for handling nonlinear multi-class medical prediction tasks. Similarly, the stacking framework benefits from integrating diverse model predictions through a meta-learning strategy.
However, models such as Logistic Regression and kNN exhibited significantly lower performance, particularly in recall and F1-score metrics. This suggests that simpler linear or distance-based decision boundaries are insufficient for capturing the complex relationships associated with cardiovascular disease severity progression.
5.3.4. Confusion Matrix Analysis
The confusion matrix obtained from the XGBoost classifier for the imbalanced dataset is shown in Table 8. The results are averaged across folds. They reveal several important observations:
Table 8.
Confusion matrix of XGBoost for the imbalanced dataset.
- Class 0 was classified with high accuracy because of its dominant representation.
- Intermediate severity classes (C1–C3) exhibited moderate confusion with neighboring severity levels.
- Severe cases (Class 4) demonstrated reduced recall because of insufficient training samples.
These results highlight the limitations of conventional machine learning models when applied to highly imbalanced multi-class medical datasets.
5.4. Multi-Class Classification After Applying SMOTE
To address the imbalance problem, the Synthetic Minority Over-sampling Technique (SMOTE) was applied to the training folds during cross-validation. Figure 13 illustrates the class distribution before and after applying SMOTE. After oversampling, all severity classes contain balanced sample distributions, enabling the classifiers to better learn minority class decision boundaries.
Figure 13.
Class distribution of the cardiovascular disease severity dataset before and after applying SMOTE: (a) original imbalanced dataset and (b) balanced dataset generated using SMOTE with .
Figure 14 and Figure 15 present the feature correlation and feature importance results after applying SMOTE. The analysis shows that the most important predictive variables remained consistent with those identified under imbalanced conditions, confirming the robustness and clinical relevance of the selected features.
Figure 14.
Feature correlation for multi-class classification with SMOTE.
Figure 15.
Comparison of feature importance of the input features obtained from different techniques: (a) Feature importance score from Decision Tree for multi-class classification with SMOTE. (b) Feature importance score from Random Forest for multi-class classification with SMOTE. (c) Feature importance score from XGBoost for multi-class classification with SMOTE.
The persistent dominance of variables such as number of major vessels, chest pain type, maximum heart rate, ST depression, and age indicates that the oversampling process preserved meaningful clinical relationships while improving class representation.
5.4.1. Hyperparameter Settings for Balanced Classification with SMOTE
The optimized hyperparameter configurations for the balanced-data experiments are summarized in Table 9. Several classifiers required different parameter settings after oversampling because balanced data distributions altered the underlying decision boundaries and model learning behavior.
Table 9.
Parameter settings for the multi-class classification with SMOTE.
5.4.2. Performance After Applying SMOTE
The quantitative classification results obtained after applying SMOTE are presented in Table 10. The experimental results demonstrate a substantial improvement across all evaluation metrics compared with the imbalanced-data experiments.
Table 10.
Results obtained for the multi-class classification of the balanced dataset.
Most classifiers achieved performance values within the 80–95% range, confirming the effectiveness of imbalance-aware learning in multi-class cardiovascular severity prediction.
Among all evaluated models, the stacking ensemble achieved the highest overall classification performance, closely followed by Majority Voting and Random Forest. The strong performance of ensemble methods indicates that combining diverse classifiers improves robustness and generalization under balanced multi-class conditions.
Interestingly, Logistic Regression continued to underperform despite the improved class distribution, suggesting that linear classification boundaries remained insufficient for modeling the complex nonlinear relationships associated with cardiovascular disease severity progression.
5.4.3. Confusion Matrix Analysis After SMOTE
The confusion matrix obtained from the best classifier, the stacking classifier, after applying SMOTE is shown in Table 11. The results are averaged across folds. Compared with the imbalanced-data experiment, the balanced-data confusion matrix demonstrates improved recall for minority severity classes, reduced confusion among adjacent severity levels, and more balanced classification performance across all classes.
Table 11.
Confusion matrix of stacking classifier after SMOTE.
In particular, the detection capability for severe cardiovascular disease classes improved substantially after oversampling, highlighting the importance of imbalance-aware learning for clinically critical cases.
The analysis of the best-performing five-class model with SMOTE is presented in the Table 12. Particular attention should be given to Classes 3 and 4 to assess the model’s ability to distinguish these higher disease categories. The results show that the classifier can identify the classes very well, achieving more than 95% precision, recall, and F1-score for Classes 2, 3, and 4.
Table 12.
Performance metrics by class.
5.5. Comparative Discussion of the Results
The experimental results clearly demonstrate that ensemble learning methods consistently outperform individual classifiers across both binary and multi-class cardiovascular disease prediction tasks.
Among the evaluated approaches, XGBoost achieved the strongest overall performance for imbalanced multi-class classification, while stacking and majority voting ensembles achieved superior performance after applying SMOTE.
The superior performance of ensemble models can be attributed to their ability to reduce model variance, capture nonlinear feature interactions, and integrate complementary predictive information from multiple learners.
The experiments further demonstrate that class imbalance significantly degrades the predictive performance of conventional classifiers, particularly for severe disease categories. Applying SMOTE substantially improved minority class detection and overall classification fairness.
5.5.1. Effectiveness of the Proposed Hierarchical Voting Mechanism
The proposed hierarchical voting mechanism further improved prediction consistency by resolving conflicts among base classifiers using confidence-aware decision rules.
Unlike conventional majority voting, which may fail under evenly distributed predictions, the proposed framework incorporates posterior probability information to prioritize high-confidence predictions, reduce ambiguity in borderline cases, and improve robustness across all severity levels.
This improvement is especially important in clinical decision-making environments, where failure to detect severe cardiovascular conditions may lead to delayed intervention and increased patient risk.
5.5.2. Clinical Implications and Significance
The results demonstrate the potential of the proposed framework to support heart-disease assessment using routinely available clinical attributes. The five-class analysis provides more granular predictive information than binary classification, while SMOTE-based imbalance handling improves the representation of less frequent categories. Feature-importance analysis identifies chest-pain type, maximum heart rate, ST depression, and number of major vessels as influential predictive attributes, although these represent predictive associations rather than clinical causality.
Classification errors also have important clinical implications. False negatives may delay further assessment of patients with heart disease, whereas false positives may lead to unnecessary follow-up examinations and additional healthcare-resource utilization. In multi-class prediction, reliable identification of the less-represented higher disease categories, particularly Classes 3 and 4, is therefore important.
The proposed framework is intended to support rather than replace clinical judgment. In practical settings, predictions and confidence information could complement clinical assessment, particularly when classifiers disagree. Real-world implementation would require integration with clinical workflows, appropriate handling of heterogeneous or missing data, external validation, and continued performance monitoring.
5.6. Class-Wise Performance Analysis
To further investigate the effectiveness of the proposed framework, a class-wise performance analysis was conducted for the multi-class classification experiments.
The analysis demonstrated that imbalanced datasets disproportionately degraded performance for severe disease classes, ensemble models improved classification consistency across all severity levels, and SMOTE substantially enhanced minority class recall.
5.6.1. Performance Under Imbalanced Conditions
Under the original imbalanced dataset,
- Class 0 achieved consistently high classification accuracy because of its dominant representation;
- Classes C1 and C2 exhibited moderate confusion with neighboring categories;
- Classes C3 and C4 experienced low recall and higher misclassification rates.
These findings indicate that conventional models tend to prioritize majority classes and under-detect clinically critical minority cases.
5.6.2. Performance After SMOTE
After applying SMOTE, the recall for severe classes improved substantially, confusion among adjacent severity levels decreased, and classification stability increased across all classes. The improvement was particularly evident in ensemble-based models such as Random Forest, XGBoost, majority voting, and stacking.
5.6.3. Related Work Comparison and Overall Findings
The class-wise analysis confirms that imbalance-aware learning is essential for reliable multi-class medical severity classification, particularly for underrepresented severity levels. Ensemble learning further improves robustness across classes, while the proposed hierarchical voting mechanism enhances prediction consistency and reliability. Overall, the findings demonstrate that integrating preprocessing, feature engineering, imbalance handling, ensemble learning, and confidence-aware voting provides a robust framework for cardiovascular disease severity prediction.
The proposed framework also demonstrates competitive performance compared with recent tabular heart-disease studies. Kibria and Matin [19] reported accuracies of 95% for binary classification and 75% for five-class classification using weighted-score fusion. In comparison, the present study achieved 86.74% accuracy on the original imbalanced five-class dataset and 94.94% after applying SMOTE with stacking. Daza et al. [22] achieved 88.24% accuracy using stacking and oversampling for binary classification, whereas Wang et al. [26] reported 86.60% accuracy and an F1-score of 88.26% using an uncertainty-aware ensemble. These findings indicate that the proposed framework achieves competitive performance despite the greater complexity of the five-class severity classification task. Nevertheless, direct numerical comparisons should be interpreted cautiously because the studies differ in datasets, class definitions, preprocessing strategies, imbalance-handling methods, and validation protocols.
6. Conclusions and Future Work
6.1. Conclusions
This study presented a hybrid ensemble and imbalance-aware machine-learning framework for binary and five-class heart-disease prediction using structured clinical data from the combined UCI Heart Disease dataset. The framework integrated preprocessing, feature analysis, SMOTE-based imbalance handling, six machine-learning classifiers, majority voting, stacking, and a confidence-aware hierarchical voting mechanism.
The results showed that five-class prediction was substantially more challenging than binary classification when using the original imbalanced data. XGBoost achieved the highest multi-class accuracy of 86.74%, while stacking provided strong overall performance across the evaluation metrics. After applying SMOTE to the training folds, ensemble performance improved considerably. Stacking achieved the best overall results, with 94.94% accuracy, 94.92% precision, 94.94% recall, and 94.92% F1-score, closely followed by majority voting. These findings demonstrate the importance of imbalance-aware learning and heterogeneous ensembles for multi-class heart-disease prediction.
Feature-importance analysis identified chest-pain type, maximum heart rate, ST depression, and number of major vessels as influential predictive attributes. These results represent predictive associations rather than evidence of clinical causality.
The proposed confidence-aware hierarchical voting mechanism provides an explicit strategy for resolving disagreements among heterogeneous classifiers. It retains majority decisions when sufficient consensus exists and uses posterior-confidence information for ambiguous predictions. Thus, the contribution of this study lies not only in competitive predictive performance for an imbalanced five-class problem, but also in providing a confidence-aware mechanism for ensemble conflict resolution. The findings demonstrate the potential of the proposed framework for multi-class heart-disease prediction using tabular clinical data, although further validation is required before considering clinical applications.
6.2. Limitations and Future Work
Although the study demonstrates promising predictive performance using the 920-record UCI Heart Disease dataset, several limitations remain. First, the original data are highly imbalanced, particularly in minority categories; while SMOTE was used during training, real-world generalizability may still be affected by variations in patient demographics, missing data, and differing clinical protocols. Furthermore, the use of tree-based feature importance metrics introduces potential bias toward variables with more split points or strong correlations. Consequently, these importance scores should be viewed merely as predictive indicators rather than measures of true clinical or causal significance.
To address these limitations and build robust clinical decision-support applications, future research must prioritize external validation using larger, independent, and contemporary datasets under heterogeneous conditions. The predictive framework can be further enhanced by incorporating complementary modalities like ECG signals, medical imaging, and wearable sensor data. Finally, integrating advanced techniques—such as uncertainty estimation, adaptive ensemble strategies, continuous model monitoring, and SHAP-based explainability—will be crucial for providing reliable, patient-level interpretations in real-world medical settings.
Author Contributions
Conceptualization, B.T. and N.W.; methodology, B.T. and P.K.; software, P.K.; validation, B.T., P.K. and N.W.; formal analysis, B.T. and P.K.; resources, S.V. and B.T.; data curation, B.T. and P.K.; writing—original draft preparation, B.T.; writing—review and editing, B.T., S.V., P.K., N.W. and S.V.; visualization, P.K.; supervision, N.W.; project administration, B.T.; funding acquisition, B.T. and S.V. All authors have read and agreed to the published version of the manuscript.
Funding
This research was funded by the Faculty of Science and Technology, Thammasat University (General Research Grant, Fiscal Year 2025), grant number SciGR5/2568. The same grant funded the APC.
Institutional Review Board Statement
Not applicable, as this study did not involve humans or animals. The data used in this study were obtained from publicly available web sources and do not contain personal identifiable information.
Informed Consent Statement
Not applicable.
Data Availability Statement
The data utilized in this study were derived from the Heart Disease dataset, which is publicly available at the UCI Machine Learning Repository. Detailed documentation and the curated version of the dataset can also be accessed via Kaggle (contributed by Redwankarimsony).
Acknowledgments
During the preparation of this manuscript, the authors used Google Gemini 3.7 Flash, OpenAI ChatGPT (GPT-5), and Grammarly Premium (Grammarly Inc., San Francisco, CA, USA) for language editing, grammar checking, and refinement. The authors have reviewed and edited the output and take full responsibility for the content of this publication.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- World Health Organization. Cardiovascular Diseases (CVDs). 2025. Available online: https://www.who.int/news-room/fact-sheets/detail/cardiovascular-diseases-(cvds) (accessed on 1 August 2026).
- Karna, V.V.R.; Karna, V.R.; Janamala, V.; Devana, V.N.K.R.; Ch, V.R.S.; Tummala, A.B. A Comprehensive Review on Heart Disease Risk Prediction using Machine Learning and Deep Learning Algorithms. Arch. Comput. Methods Eng. 2025, 32, 1763–1795. [Google Scholar] [CrossRef] [Scilit]
- E, B.; Rajakumari, K.; V., K. An Interpretable Machine Learning Approach for Heart Disease Detection Using Explainable AI Techniques. In Proceedings of the 2025 15th International Conference on Mathematics, Actuarial Science, Computer Science and Statistics (MACS), Karachi, Pakistan, 20–21 December 2025; pp. 1–7. [Google Scholar] [CrossRef] [Scilit]
- Ganie, S.M.; Pramanik, P.K.D.; Zhao, Z. Ensemble Learning with Explainable AI for Improved Heart Disease Prediction Based on Multiple Datasets. Sci. Rep. 2025, 15, 13912. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Anshori, M.; Haris, M.S. Predicting Heart Disease using Logistic Regression. Knowl. Eng. Data Sci. 2022, 5, 188–196. [Google Scholar] [CrossRef] [Scilit]
- Rajendran, R.; Karthi, A. Heart disease prediction using entropy based feature engineering and ensembling of machine learning classifiers. Expert Syst. Appl. 2022, 207, 117882. [Google Scholar] [CrossRef] [Scilit]
- Ozturk Kiyak, E.; Ghasemkhani, B.; Birant, D. High-Level K-Nearest Neighbors (HLKNN): A Supervised Machine Learning Model for Classification Analysis. Electronics 2023, 12, 3828. [Google Scholar] [CrossRef] [Scilit]
- Hu, L.Y.; Huang, M.W.; Ke, S.W.; Tsai, C.F. The Distance Function Effect on K-Nearest Neighbor Classification for Medical Datasets. SpringerPlus 2016, 5, 1304. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, G.; Gionis, A. Regularized Impurity Reduction: Accurate Decision Trees with Complexity Guarantees. Data Min. Knowl. Discov. 2023, 37, 434–475. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kim, Y.; Kwon, B.A.; Kang, K. Assessing the Impact of Data Impurity Measures on Decision Tree Algorithms in Spam Classification. J. Artif. Intell. Res. Appl. 2024, 1, 33–46. [Google Scholar]
- Tetteh, E.T.; Zielosko, B. Greedy Algorithm for Deriving Decision Rules from Decision Tree Ensembles. Entropy 2025, 27, 35. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cervantes, J.; Garcia-Lamont, F.; Rodríguez-Mazahua, L.; Lopez, A. A comprehensive survey on support vector machine classification: Applications, challenges and trends. Neurocomputing 2020, 408, 189–215. [Google Scholar] [CrossRef] [Scilit]
- Guido, R.; Ferrisi, S.; Lofaro, D.; Conforti, D. An Overview on the Advancements of Support Vector Machine Models in Healthcare Applications: A Review. Information 2024, 15, 235. [Google Scholar] [CrossRef] [Scilit]
- Halabaku, E.; Bytyçi, E. Overfitting in Machine Learning: A Comparative Analysis of Decision Trees and Random Forests. Intell. Autom. Soft Comput. 2024, 39, 987–1006. [Google Scholar] [CrossRef] [Scilit]
- Muniappan, M.; Paruvachi Subramanian, N.D. A Majority Voting Mechanism-Based Ensemble Learning Approach for Financial Distress Prediction in Indian Automobile Industry. J. Risk Financ. Manag. 2025, 18, 197. [Google Scholar] [CrossRef] [Scilit]
- He, H.; Garcia, E.A. Learning from Imbalanced Data. IEEE Trans. Knowl. Data Eng. 2009, 21, 1263–1284. [Google Scholar] [CrossRef] [Scilit]
- Azad, M.; Nehal, T.H.; Moshkov, M. A Novel Ensemble Learning Method Using Majority Based Voting of Multiple Selective Decision Trees. Computing 2025, 107, 42. [Google Scholar] [CrossRef] [Scilit]
- Li, J.P.; Haq, A.U.; Din, S.U.; Khan, J.; Khan, A.; Saboor, A. Heart Disease Identification Method Using Machine Learning Classification in E-Healthcare. IEEE Access 2020, 8, 107562–107582. [Google Scholar] [CrossRef] [Scilit]
- Gárate-Escamila, A.K.; Hajjam El Hassani, A.; Andrès, E. Classification Models for Heart Disease Prediction Using Feature Selection and PCA. Inform. Med. Unlocked 2020, 19, 100330. [Google Scholar] [CrossRef] [Scilit]
- Dissanayake, K.; Md Johar, M.G. Comparative Study on Heart Disease Prediction Using Feature Selection Techniques on Classification Algorithms. Appl. Comput. Intell. Soft Comput. 2021, 2021, 5581806. [Google Scholar] [CrossRef] [Scilit]
- Pan, C.; Poddar, A.; Mukherjee, R.; Ray, A.K. Impact of Categorical and Numerical Features in Ensemble Machine Learning Frameworks for Heart Disease Prediction. Biomed. Signal Process. Control 2022, 76, 103666. [Google Scholar] [CrossRef] [Scilit]
- Kibria, H.B.; Matin, A. The Severity Prediction of the Binary and Multi-Class Cardiovascular Disease—A Machine Learning-Based Fusion Approach. Comput. Biol. Chem. 2022, 98, 107672. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abdellatif, A.; Abdellatef, H.; Kanesan, J.; Chow, C.O.; Chuah, J.H.; Gheni, H.M. An Effective Heart Disease Detection and Severity Level Classification Model Using Machine Learning and Hyperparameter Optimization Methods. IEEE Access 2022, 10, 79974–79985. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Yuan, Y.; Yao, Z.; Yang, J.; Wang, X.; Tian, J. Coronary Artery Disease Detection Model Based on Class Balancing Methods and LightGBM Algorithm. Electronics 2022, 11, 1495. [Google Scholar] [CrossRef] [Scilit]
- Daza, A.; Bobadilla, J.; Herrera, J.C.; Medina, A.; Saboya, N.; Zavaleta, K.; Siguenas, S. Stacking Ensemble Based Hyperparameters to Diagnosing of Heart Disease: Future Works. Results Eng. 2024, 21, 101894. [Google Scholar] [CrossRef] [Scilit]
- El-Sofany, H.; Bouallegue, B.; Abd El-Latif, Y.M. A Proposed Technique for Predicting Heart Disease Using Machine Learning Algorithms and an Explainable AI Method. Sci. Rep. 2024, 14, 23277. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ghose, P.; Oliullah, K.; Mahbub, M.K.; Biswas, M.; Uddin, K.N.; Jamil, H.M. Explainable AI Assisted Heart Disease Diagnosis Through Effective Feature Engineering and Stacked Ensemble Learning. Expert Syst. Appl. 2025, 265, 125928. [Google Scholar] [CrossRef] [Scilit]
- Chatterjee, S.; Hadi, A.S. Regression Analysis by Example, 5th ed.; John Wiley & Sons, Inc.: Hoboken, NJ, USA, 2012; pp. 58–60. [Google Scholar]
- Du, K.L.; Zhang, R.; Jiang, B.; Zeng, J.; Lu, J. Foundations and Innovations in Data Fusion and Ensemble Learning for Effective Consensus. Mathematics 2025, 13, 587. [Google Scholar] [CrossRef] [Scilit]
- Chawla, N.V.; Bowyer, K.W.; Hall, L.O.; Kegelmeyer, W.P. SMOTE: Synthetic Minority Over-sampling Technique. J. Artif. Intell. Res. 2002, 16, 321–357. [Google Scholar] [CrossRef] [Scilit]
- Ibrahim, M. Evolution of Random Forest from Decision Tree and Bagging: A Bias-Variance Perspective. Dhaka Univ. J. Appl. Sci. Eng. 2022, 7, 66–71. [Google Scholar] [CrossRef] [Scilit]
- Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; Association for Computing Machinery: New York, NY, USA, 2016; pp. 785–794. [Google Scholar] [CrossRef] [Scilit]
- Janosi, A.; Steinbrunn, W.; Pfisterer, M.; Detrano, R. Heart Disease [Dataset]. UCI Machine Learning Repository. 1989. Available online: https://archive.ics.uci.edu/dataset/45/heart+disease (accessed on 13 August 2025).
- Ashika, T.; Hannah Grace, G. Enhancing Heart Disease Prediction with Stacked Ensemble and MCDM-Based Ranking: An Optimized RST-ML Approach. Front. Digit. Health 2025, 7, 1609308. [Google Scholar] [CrossRef] [Scilit]
- Ghose, P.; Jamil, H. CardiaTics: An Explainable AI Integrated Heart Disease Diagnosis Model with Feature Engineering and Stacked Ensemble Approach. J. Big Data 2026, 13, 59. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, X.; Fan, Y.; Yu, M.; Yuan, F. Uncertainty-Aware Feature-Weighted Ensemble Framework for Heart Disease Prediction. Sci. Rep. 2026, 16, 13321. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.














