Next Article in Journal
KIF18B Is Essential for Lung Adenocarcinoma Progression Through the E2F Transcriptional Network
Next Article in Special Issue
RAGE Axis in the Pathogenesis and Treatment of CNS Neurodegeneration in Long-Term Hyperglycemia
Previous Article in Journal
Antimicrobial Activity of Aurisin A Against Streptococcus suis and Its Protective Effect on Epithelial Cells
Previous Article in Special Issue
From Adipose to Action: Reprogramming Stem Cells for Functional Neural Progenitors for Neural Regenerative Therapy
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Interpretable Machine Learning with SHAP Identifies Key Biomarkers in a Multi-Factorial Spectrum of Age-Related Neurological and Metabolic Conditions

by
Daniil V. Artamonov
1,2,
Polina I. Popova
3,
Ekaterina A. Korf
4,
Natalia G. Voitenko
4,
Alisa A. Chernysheva
4,5,
Pavel V. Avdonin
6,
Richard O. Jenkins
7 and
Nikolay V. Goncharov
4,8,9,*
1
Group of Theoretical Chemistry, N.D. Zelinsky Institute of Organic Chemistry of Russian Academy of Sciences, Leninsky Prospect 47, Moscow 119991, Russia
2
Infochemistry Scientific Center, National Research University ITMO, Kronverksky Ave. 49, Building A, St. Petersburg 197101, Russia
3
St. Petersburg State Budgetary Healthcare Institution “City Polyclinic No. 107”, St. Petersburg 195030, Russia
4
Sechenov Institute of Evolutionary Physiology and Biochemistry, Russian Academy of Sciences, pr. Torez 44, St. Petersburg 194223, Russia
5
Department of Computational Biology, Sirius University of Science and Technology, Olympic Ave. 1, Sirius 354340, Russia
6
Koltsov Institute of Developmental Biology, Russian Academy of Sciences, 26 Vavilov St., Moscow 119334, Russia
7
Leicester School of Allied Health Sciences, De Montfort University, The Gateway, Leicester LE1 9BH, UK
8
Research Institute of Hygiene, Occupational Pathology and Human Ecology of the Federal Medical Biological Agency, p.o. Kuz’molovsky bld.93, St. Petersburg 188663, Russia
9
Department of Biological Chemistry, Petersburg State Pediatric Medical University, Ministry of Health of the Russian Federation, Litovskaya St. 2, St. Petersburg 194100, Russia
*
Author to whom correspondence should be addressed.
Int. J. Mol. Sci. 2026, 27(4), 1805; https://doi.org/10.3390/ijms27041805
Submission received: 6 January 2026 / Revised: 2 February 2026 / Accepted: 9 February 2026 / Published: 13 February 2026
(This article belongs to the Special Issue Challenges and Innovation in Neurodegenerative Diseases, 2nd Edition)

Abstract

Vascular and metabolic disorders in the elderly—including acute ischemic stroke (AIS), chronic cerebral circulation insufficiency (CCCI), type 2 diabetes mellitus (DM), and subcortical ischemic vascular dementia (SIVD)—pose a major diagnostic challenge due to their reliance on multi-parameter blood chemistry. In this study, 49 biochemical features were analyzed within a cohort of 120 patients. The application of variance-aware statistical testing revealed that several features (e.g., Fe, Transf, RDW%, LDL) exhibited statistically significant heterogeneity of variance (p < 0.05), which is known to distort standard ANOVA inference. While standard machine-learning (ML) classifiers demonstrated variable performance across clinical groups, a gradient boosting model with restricted tree depth (max depth = 3) achieved high discriminative accuracy, yielding F1-scores between 0.87 and 0.96 across all five clinical classes. Through the use of Shapley Additive Explanations (SHAP), key stable biomarkers including iron (Fe), transferrin, and glucose were identified as having synergistic interactions in model predictions. A comparative analysis of feature importance ranks indicated consistency between statistical significance and SHAP values, with Spearman correlation coefficients reaching 0.53 for groups 1–2 and 0.59 for groups 1–5. Conversely, unsupervised KMeans clustering (k = 5) revealed a poor correspondence with clinical labels, yielding an Adjusted Rand Index (ARI) of 0.198 and Normalized Mutual Information (NMI) of 0.286. These results underscore that statistical structures in biochemical data do not always map to meaningful clinical categories and advocate for the adoption of variance-aware workflows and interpretable ML to enhance diagnostic reliability in aging populations.

1. Introduction

Diagnostic error is commonly defined as a missed, delayed, or incorrect diagnosis and is considered one of the most serious threats to patient safety [1]. A large-scale study of patients transferred to intensive care units or who died in hospitals found that diagnostic errors were present in nearly one-quarter of cases. Approximately every fifteenth death was associated with a diagnostic error. These errors were most often linked to delays in diagnosis and incorrect selection of diagnostic tests, leading to worse outcomes and high mortality rates [2]. In 2010–2011, an estimated 40,000 adult patients died annually in the United States due to misdiagnosis in intensive care units and emergency departments, while 10 years later, this number had risen to 250,000 [3,4]. The leading causes of death are vascular disease and infections. The implementation of artificial intelligence (AI) systems in clinical diagnosis represents a promising direction for developing reliable support tools that can improve the accuracy of diagnostic decisions, optimize examination time, and reduce the risk of medical errors [5].
However, beyond the emergency departments, vascular and metabolic disorders represent a substantial clinical and public health burden in the aging population, with acute ischemic stroke (AIS), chronic cerebral circulation insufficiency (CCCI), prediabetes, type 2 diabetes (DM), and subcortical ischemic vascular dementia (SIVD) being among the most prevalent and debilitating conditions affecting elderly individuals [6]. AIS, a major cause of mortality and morbidity in the elderly, has become increasingly prevalent in recent years, with ischemic stroke accounting for approximately 70–80% of all stroke cases and demonstrating age-dependent increases in incidence, particularly among individuals aged 85 years and older [7]. Chronic cerebral vascular insufficiency and cerebral small vessel disease contribute substantially to cognitive decline and neurological dysfunction through progressive deterioration of small penetrating arteries, leading to white matter hyperintensities, lacunar infarctions, and impaired cerebral autoregulation [8]. Prediabetes and DM affect a large proportion of elderly individuals, with up to 70% of patients with prediabetes eventually progressing to diabetes [9]. In elderly patients with DM, macrovascular complications including coronary heart disease (42.4%), stroke (11.4%), and chronic obliterating disease of the lower limb arteries (10.1%) are highly prevalent, alongside microvascular complications such as retinopathy (64.9%) and polyneuropathy (90.7%) [10]. Subcortical ischemic vascular dementia, often referred to as Binswanger’s disease in its severe form, represents the most common subtype of vascular dementia and is characterized by confluent white matter hyperintensities, multiple lacunar infarcts, and progressive cognitive decline with prominent executive dysfunction, psychomotor slowing, and gait disturbances [11]. The clinical significance of these interconnected vascular and metabolic disorders extends beyond individual morbidity and mortality, as cerebral small vessel disease has been identified as a critical contributor to approximately 50% of dementia cases globally and serves as both an initiator and amplifier of neurodegenerative pathology [12,13]. Biomarkers serve as a fundamental basis for optimizing primary screening by providing enhanced criteria for early detection and risk stratification. They contribute to the development of personalized approaches to diagnosis and prevention in oncology, neurodegenerative, and metabolic disorders [14].
Classical analysis of variance (ANOVA) remains a widely used statistical method for comparing mean outcomes across multiple treatment or diagnostic groups in biomedical research. However, a fundamental assumption of ANOVA—homogeneity of variance (homoscedasticity)—is frequently violated in clinical and biochemical datasets, leading to potentially spurious results and compromised statistical inference. When the assumption of equal variances is violated, the standard ANOVA F-test becomes unreliable, with Type I error rates deviating substantially from the nominal significance level [15]. Heteroscedasticity in biomedical data is particularly problematic because it distorts estimated standard errors, rendering confidence intervals inaccurate and hypothesis tests unreliable, which can lead to incorrect conclusions about treatment efficacy or diagnostic group differences [16]. The presence of heteroscedasticity is common in clinical trials and multi-site studies, where inherent biological variability, differences in measurement precision across subgroups, and varying sample sizes combine to produce unequal variances that systematically reduce statistical power to detect true treatment effects [17].
Importantly, Welch’s ANOVA maintains comparable or even superior statistical power to classical ANOVA when variances are equal, meaning that it can be used routinely without concern for power loss, effectively eliminating the need to test for homogeneity of variance as a prerequisite for analysis [18]. In clinical trials and biomedical research settings, Welch’s ANOVA has been successfully applied to analyze biochemical markers and clinical endpoints where unequal variances are anticipated; for instance, in studies examining biomarker changes in knee synovial fluid from patients with anterior cruciate ligament injuries, Welch’s ANOVA enabled valid comparison of cytokine concentrations across injury groups despite significant heteroscedasticity [19]. Comparative studies have demonstrated that Welch’s ANOVA consistently provides more accurate Type I error control and better power performance than both the traditional ANOVA F-test and several other alternatives under conditions of heteroscedasticity, establishing it as a preferred method for analyzing clinical and experimental data with unequal variances [20].
Machine learning (ML) algorithms have demonstrated substantial efficacy in classification and diagnosis tasks based on biochemical and clinical data, with various supervised learning methods including linear discriminant analysis (LDA), random forest (RF), support vector machines (SVM), logistic regression (LR), k-nearest neighbors (kNN), and decision trees (DT) showing strong performance across diverse clinical applications [21]. Random forest, an ensemble learning method that constructs multiple decision trees and aggregates their predictions, has emerged as one of the most widely applied ML techniques in medical research with particular success in handling high-dimensional datasets, such as electronic health records, and demonstrating robustness to missing data and resistance to overfitting [22]. Investigations into oncological applications have shown that SVM and DT approaches can yield accuracies exceeding 97% when differentiating between benign and malignant lesions, as seen in skin cancer detection studies, underscoring the capacity of these algorithms to handle high-dimensional data while maintaining exceptional sensitivity and specificity [23]. Logistic Regression continues to play a vital role in clinical prediction models by providing a robust probabilistic framework that can integrate multiple covariates to predict disease risk, offering interpretable coefficients that directly translate into clinical risk estimates [24]. Support vector machines, employed in 32% of contemporary studies, excel in handling high-dimensional data and nonlinear relationships between features, demonstrating particular utility in genomic data analysis, medical imaging classification, and precision medicine applications where complex patterns must be identified from multi-modal datasets [25]. K-nearest neighbors algorithms have been successfully applied to medical diagnosis tasks, particularly in scenarios where local patterns in feature space are informative for classification, although their performance can be sensitive to the choice of distance metric and the number of neighbors [26]. Furthermore, the flexibility of these ML algorithms to adapt to high-dimensional and heterogeneous data environments underscores their pivotal role in modern clinical diagnostics, where they not only enhance predictive performance but also contribute to a deeper understanding of underlying disease mechanisms [27].
Shapley Additive Explanations (SHAP), a framework grounded in cooperative game theory for explaining machine learning predictions, has become a mainstream approach in data science for interpreting complex models and enhancing the clinical utility of ML-based diagnostic and prognostic systems. SHAP values assign each feature an importance value for a particular prediction by fairly distributing the ‘payout’ among the features, thus providing explainability locally for individual predictions as well as globally across datasets [28]. Recent studies integrating convolutional neural networks with attention mechanisms and SHAP for personalized health monitoring have achieved remarkable accuracies of 97.86%, with SHAP providing both population-level feature importance rankings and patient-specific explanations through force plots and dependence visualizations, which enable clinicians to understand why specific predictions were made for individual patients [29]. The XGBoost model combined with SHAP analysis accurately predicted lymph node metastasis in gastric cancer patients. SHAP values identified inflammatory indices and peripheral lymphocyte subpopulations as top contributors, with SHAP values ranging from 0.15 to 0.32 on average. This approach enhanced model interpretability and achieved high predictive accuracy (AUC > 0.85) [30]. In acute heart failure prediction, XGBoost models combined with SHAP analysis have demonstrated superior performance with AUC values of 0.817 for in-hospital mortality and 0.766 for worsening heart failure, while SHAP clustering has enabled risk stratification by forming interpretable patient clusters with similar mortality and complication rates [31].
This study focuses on comparing one-way ANOVA and Welch’s ANOVA for selecting biochemical features across patient groups, followed by the evaluation of ML classifiers (LDA, LR, kNN, DT, RF, SVM) and interpretation using SHAP. The objective is to develop ML models for a personalized diagnosis of cardiovascular diseases, providing enhanced accuracy and interpretability of results.

2. Results

2.1. Data Analysis

Statistical analysis of the dataset provides an essential foundation for understanding the distribution and variability of biochemical features across patient groups. Initial screening was performed with classical statistical methods to identify features that differ between groups and to inform subsequent feature selection for machine learning. In particular, assessment of variance homogeneity and between-group differences is important because these characteristics influence both the interpretability and predictive performance of supervised models.
To test variance homogeneity, Levene’s test was applied to each biochemical parameter. Figure 1 shows the Levene W statistics and corresponding p-values (right axis, log scale) for all parameters. Several features exhibited statistically significant heterogeneity of variance (p < 0.05, e.g., Fe, Transf, RDW%, LDL, LDH), indicating violation of the equal-variance assumption required by classical one-way ANOVA.
Due to the observed heteroscedasticity, the results of classical one-way ANOVA were compared with Welch’s ANOVA, which relaxes the equal-variance assumption by adjusting the F-statistic and degrees of freedom. Figure 2 presents p-values from both methods for all blood parameters (p-values plotted on a log scale and sorted by Welch’s ANOVA). The comparison highlights marked discrepancies in statistical significance for a subset of parameters.
Notably, the divergence between methods is pronounced for several parameters (e.g., Fe, Transferrin (Transf), RDW%, Glucose, Albumin, RDWa, von Willebrand Factor (vWF), LDL). For these features, the p-values produced by Welch’s ANOVA differ substantially from those of classical ANOVA, reflecting the influence of unequal variances across groups. Box plots of the top 5 features ranked by Welch’s ANOVA and by one-way ANOVA are shown in Figure 3 and Figure 4, respectively. For the boxplots, feature values were standardized (z-score) prior to plotting to facilitate visual comparison of group distributions.
This phenomenon—substantial differences in ranked lists of significant features between the two ANOVA variants—indicates that heteroskedasticity materially affects significance calls. Selecting an appropriate test (e.g., Welch’s ANOVA when variances are unequal) therefore improves the reliability of feature selection, which in turn supports construction of more robust and interpretable machine-learning models.
Multicollinearity Check: To ensure that features included in the models are not highly correlated with one another, Pearson correlation was used to examine the relationships between all features. A heatmap of the correlation matrix was generated to visualize multicollinearity, with highly correlated features (correlation coefficient > 0.8) flagged for potential exclusion from subsequent analysis to reduce collinearity and improve model performance. The correlation heatmap is shown in Figure 5.

2.2. Comparison of ANOVA Variants Under Heteroscedasticity

To quantitatively validate the impact of heteroscedasticity on classical ANOVA inference, five statistical tests were compared: classical ANOVA, Welch’s ANOVA, Levene’s test, Brown-Forsythe (Levene on medians), and Heteroskedasticity-Consistent Covariance Matrix Estimator (HC3) (Table S1). The comparative testing results reveal the fundamental limitations of ANOVA under violated homoscedasticity assumptions. As shown in Table S1, which presents the number of significant parameters (at α = 0.05 significance level), classical ANOVA identified 25 significant features, clearly illustrating the tendency toward Type I error inflation in the presence of unequal variances across patient groups. This inflation occurs because the standard ANOVA F-test is highly sensitive to deviations from variance homogeneity, leading to distorted statistical inference. In contrast, Welch’s ANOVA identified 26 significant parameters, demonstrating a more conservative and robust approach that maintains statistical power without compromising Type I error control even under heteroscedasticity conditions. Comparison with alternative robust methods—including Brown-Forsythe (15 significant), HC3 (27 significant), and Kruskal-Wallis (27 significant)—confirms consensus between Welch’s ANOVA and Kruskal-Wallis (Spearman ρ = 0.92), providing strong methodological justification for selecting Welch’s ANOVA as the primary feature selection method under the heteroscedastic conditions characteristic of biochemical biomarker datasets.

2.3. Machine Learning

A set of classical supervised classifiers was evaluated to quantify discriminative performance and the effect of different feature-selection strategies. The following models were considered: LDA, LR, kNN, DT, RF and SVM. These methods cover linear and non-linear paradigms and provide complementary perspectives on class separability and feature relevance.
Three feature sets were used in the experiments: (1) the full preprocessed feature set including all biochemical variables, (2) a subset of features selected by Welch’s ANOVA, and (3) a subset selected by classical one-way ANOVA. For both ANOVA-based selections, only features with p-values below 0.05 were retained. The one-way ANOVA subset included the following features: Transf, ALB, Glu, Chol, Ca, BChE, LDL, LDH, Fe, RBC, RDW%, GRA%, NEFA, LYM%, HGB, GRAN, HDL, HCT, ALP, vWF, CK-NAC, Ur.Acid, TP, a1-AGP, MCHC, and MCV. The Welch ANOVA subset consisted of: Fe, Transf, RDW%, LDL, LDH, vWF, Chol, BChE, Glu, ALB, RBC, MCHC, NEFA, GRAN, HDL, HGB, Ca, HCT, Ur.Acid, PDW, RDWa, GRA%, ALP, CK-NAC, LYM%, TP, WBC, and a1-AGP. All numerical predictors were standardized (z-score), and all preprocessing steps were performed within cross-validation folds to prevent information leakage.
Model selection and hyperparameter optimization were performed using stratified 3-fold cross-validation with an inner grid search for each classifier. Performance was summarized by per-class F1-score (reported in tables) together with macro-average F1, Precision, and Recall metrics. Detailed per-class results and aggregated metrics are provided in Table 1.
Figure 6 shows a radar (spider) plot of the mean per-class F1-score for six classifiers (LDA, kNN, LR, DT, RF, SVM) under three feature-selection strategies. Each axis represents a classifier, and the closed polygons correspond to the F1-score averaged across all classes for each feature set: the full set of features, features selected by one-way ANOVA, and features selected by Welch’s ANOVA. The plot provides a concise visual comparison of classifier performance across feature-selection strategies, highlighting relative differences rather than absolute F1 values.

2.4. Personalized and Interpretable Models

For the SHAP analysis, a Gradient Boosting model with a tree depth of 3 was used to prevent overfitting on the full dataset while maintaining interpretability. This shallow tree structure ensures that the computed SHAP values are stable and reliable. In fact, this choice is heuristic, so we obtained the learning curves for the gradient boosting classifier (Figure S1) and conducted a sensitivity analysis of the gradient boosting model (Figure S2). Tree SHAP allows for exact and computationally efficient calculation of Shapley values, capturing not only the individual contributions of each biochemical feature to the model’s predictions, but also interactions between features. As a result, the analysis provides a nuanced understanding of how combinations of variables influence patient classification and helps identify the most influential features beyond simple correlations. The model’s performance is presented in the confusion matrix (Figure 7) and summarized in Table 2.
Subsequently, SHAP values were computed for all features (blood test parameters) across all classes to evaluate their individual contributions to the model’s predictions. The results are presented for each class using paired plots: one showing feature interactions and the other summarizing individual feature impacts. Figure 8, Figure 9, Figure 10, Figure 11 and Figure 12 display these visualizations for classes 1 through 5, highlighting the most influential features and the strength and direction of their effects on the predicted probabilities. Features are ranked by global importance. Each dot represents a patient; red indicates a high biomarker value, while blue indicates a low value. Dots to the right of the midline increase the diagnostic probability for the respective class, whereas dots to the left decrease it.

2.5. Mutual Information for Feature Synergy (MI)

To further investigate feature relationships and interactions, two complementary visualizations were constructed. First, a feature interaction network was created where each feature was represented as a node. The top 50% of features were selected based on mutual information with the target variable. Edges were added between feature pairs exhibiting absolute Pearson correlations above 0.3, with edge weights proportional to correlation strength. Isolated nodes were removed, retaining only interacting features. Community detection was performed using the Louvain algorithm, and nodes were arranged in a circular layout, with cluster membership indicated by distinct pastel colors. Edge widths reflected correlation magnitude, providing an intuitive representation of clusters of features that may jointly contribute to class discrimination.
Second, to capture interaction effects identified by the predictive model, the absolute SHAP interaction values were computed and averaged across all classes. The resulting mean interaction matrix was visualized as a heatmap (Figure 13a). This visualization highlights pairs of features whose combined effects contribute substantially to model predictions, complementing the correlation-based network by revealing non-linear and class-dependent interactions. Together, these analyses provide a multi-level perspective on both linear dependencies and model-informed synergistic relationships among features. The combined results of the feature interaction network and the mean SHAP interaction values across classes are presented in Figure 13.

2.6. Unsupervised Pattern Discovery

Unsupervised clustering was performed on the full feature space using the KMeans algorithm. The optimal number of clusters was determined via the elbow method, which involves plotting the within-cluster sum of squares (WCSS) as a function of the number of clusters and identifying the point where the rate of decrease sharply diminishes. This approach suggested k = 5, coincidentally matching the number of predefined classes in the dataset (Figure 14a).
To visualize the clustering results, the high-dimensional feature space was projected into two dimensions using t-distributed Stochastic Neighbor Embedding (t-SNE), which preserves local neighborhood structures and allows for intuitive visualization of complex patterns. The resulting 2D embeddings of the clustered objects are shown in Figure 14b.
To assess the quality of the clustering and its correspondence to the ground-truth classes, the Silhouette score for the original class labels in the full feature space was first evaluated, yielding a value of −0.139. After applying KMeans clustering, the Silhouette score increased to 0.122, indicating a moderate improvement in the cohesion of the assigned groups. The agreement between the derived clusters and the true labels was further quantified using the Adjusted Rand Index (ARI = 0.198) and Normalized Mutual Information (NMI = 0.286). The ARI measures the similarity between two labels while correcting for chance, with values closer to 1 indicating strong agreement, whereas the NMI evaluates the mutual dependence between cluster assignments and true classes, also ranging from 0 (no correspondence) to 1 (perfect match). Despite the moderate improvements in these metrics, the results suggest that the clusters do not fully capture the original class structure. This discrepancy may be partly due to the high dimensionality of the feature space, which can complicate the separation of clusters and obscure intrinsic patterns in the data. The composition of each cluster with respect to the original classes is summarized in Table 3.
To further characterize the identified clusters, differences in mean blood parameter values were analyzed. Prior to the analysis, all features were standardized to ensure valid comparisons between variables with different scales. The resulting differences are visualized as a heatmap in Figure 15, illustrating the relative profiles of blood indicators for each cluster. This visualization highlights distinct trends in certain biochemical parameters, suggesting potential physiological distinctions among the identified groups.
To objectively assess the agreement between the model-estimated and statistical significance of features [6], a rank comparison was performed and the Spearman correlation was employed. The negative logarithm of the p-value from the nonparametric Kruskal-Wallis tests and Dunn’s post hoc test was used as a measure of the statistical significance of each feature: log 10 ( p ) . The model-estimated significance of features was expressed as normalized mean absolute values:
w j k = E i G φ i j k ,
w ~ j k = w j k max j   E i G φ i j k ,
where φ i j k are the SHAP values for feature (blood parameter) j, patient i, class k.
The Spearman correlation coefficient was calculated based on the ranks:
ρ = 1   6   i d i 2 n   n 2 1 ,
where di is the difference in ranks between the ranked sets of statistical significance (the negative logarithm of statistical significance) and the features selected as a result of the SHAP analysis; n is the number of features.
To determine the Spearman correlation coefficient, the sets of statistical significance of features and SHAP values were ranked—each element was assigned a rank R i i s t a t and R i S H A P according to the order in the set. The results of the comparative analysis are presented in Table 4 and Table 5.

2.7. External Validation on a Single-Class Cohort

To evaluate the generalization capability of the proposed framework and substantiate the clinical relevance of the identified biomarker signatures, external validation was performed on an independent, single-class cohort of patients diagnosed with SIVD. Given that the initial findings were considered exploratory due to the high-dimensional nature of the biochemical feature space (pn), this step aimed to determine whether the model-estimated feature importance remains stable when applied to a non-overlapping dataset.
The validation cohort comprised 28 patients, representing a clinical phenotype closely aligned with the SIVD group analyzed in the primary study. It should be noted that certain biomarkers, specifically {‘BilAc’, ‘LAC’, ‘Glyc’, ‘Transf’, ‘PON1’, ‘NEFA’, ‘Ur.Acid’, ‘a1-AGP’, ‘K’, ‘vWF’}, were unavailable in the validation dataset due to differences in clinical recording protocols. To maintain the integrity of the input feature vector and enable the deployment of the pre-trained classifier, these missing values were imputed using the mean values derived from the primary training dataset.
Furthermore, a rigorous comparative analysis was conducted between the primary and validation cohorts to assess potential out-of-distribution (OOD) effects and distributional shifts. This evaluation included mean values with 95% confidence intervals (CIs), medians, and interquartile ranges (IQRs) for each biochemical parameter. The Population Stability Index (PSI) was calculated for every feature; the average PSI was 2.111, indicating substantial distributional differences between the training and validation cohorts. These detailed comparative statistics and PSI values are provided in the Supplementary Materials (Tables S2 and S3).
To ensure methodological consistency and mitigate the impact of heteroscedasticity, the pre-trained Gradient Boosting model (maximum tree depth = 3) was deployed alongside the previously optimized variance-aware preprocessing pipeline with accuracy 71.43%. Performance was quantified not only by classification accuracy but also by the consistency of feature attribution, using SHAP (Figure 16) to assess whether key markers—such as iron (Fe), glucose, and transferrin—retain their synergistic predictive roles in an independent clinical setting. This external validation serves to confirm that the identified biomarker signatures are not artifacts of the original sample’s variance structure but reflect stable physiological patterns associated with vascular cognitive impairment.

3. Discussion

Classical classifiers (LDA, LR, kNN, DT, RF, SVM) were evaluated on three feature set variants: the full set, features selected via one-way ANOVA, and features selected via Welch’s ANOVA. The results revealed considerable variability across classes: for some classes, the models achieved acceptable accuracy, whereas for others, F1 scores were low (see Table 1). SHAP analysis and the interaction network identified a number of stable markers and inter-feature relationships (Figure 8, Figure 9, Figure 10, Figure 11, Figure 12 and Figure 13).
Several factors likely contribute to the observed limitations in classification performance. First, the dataset exhibits high dimensionality relative to the sample size, with approximately 49 numerical features for only 120 patients. This “wide” setup (pn) increases estimate variability, promotes overfitting, and destabilizes model parameters, particularly for linear models and SVMs with nonlinear kernels. Second, class imbalance shifts the optimization toward the more prevalent classes, reducing macro-F1 scores even when overall accuracy or dominant-class F1 appears satisfactory, as reflected in Table 1. Third, heterogeneity within clinical groups, arising from diverse comorbidities, increases within-class variance and diminishes between-class discriminability. This is consistent with high heteroscedasticity observed in several features (Levene test), explaining differences in feature rankings between Welch’s and standard ANOVA. Finally, strong correlations and redundancy among features (e.g., RBC/HGB/HCT; PLT/PCT; Chol/LDL) complicate coefficient interpretation and dilute the contribution of individual variables in linear models.
SHAP analysis further provided insights into feature importance. A consistent set of markers—including Fe, Transferrin, RDW%, LDL/Chol, LDH, vWF, Glu, ALB, and several hematological parameters—was repeatedly identified across classes. These features not only show high individual contributions but also participate in synergistic interactions, as evidenced in the SHAP interaction network (Figure 8, Figure 9, Figure 10, Figure 11, Figure 12 and Figure 13). Importantly, feature relevance is class-specific: certain markers are critical for distinguishing vascular-metabolic conditions, while others are more relevant for cognitive-vascular syndromes. This class-specificity explains why global metrics, such as macro-F1, may not fully capture the clinical utility of the models for particular diagnoses.
The results of the comparative analysis of ranks showed that there is a consistency between the results of the statistical analysis and the SHAP method—it is especially pronounced for combinations of groups 1–2 (0.5277) and 1–5 (0.5919) according to the Spearman coefficient. However, blood parameters that are important for predicting groups, but not for statistical analysis, were also identified: for group 2, they were Glu, Transf, LAC, WBC, LDL; for group 3, they were UREA, Ca, CREA, Fe, BilAc; for group 4, they were Chol, Fe, HDL, Ur. Acid, Ca; and for group 5, they were MCHC, Ca, K-RUV, BilAc, NEFA.
To further assess model robustness, all analyses were validated on an independent cohort consisting exclusively of patients with dementia. Statistical comparisons between the original and dementia-specific cohorts revealed a significant shift in data distributions, and several features present in the original dataset were unavailable in the new cohort. Despite these limitations, the gradient boosting model, combined with SHAP-based interpretability, maintained high classification accuracy. Notably, the same set of stable markers and class-specific feature interactions identified in the original cohort were largely preserved, indicating that the predictive patterns uncovered are robust to cohort heterogeneity and missing variables, at least within the context of dementia patients.

4. Materials and Methods

4.1. Chemicals

PBS (pH 7.4) was purchased from Biolot (St. Petersburg, Russia). Diagnostic kits were produced by Randox Laboratories (Crumlin, UK). All other reagents were from Sigma-Aldrich (Rockville, MD, USA).

4.2. Study Design and Setting

The study was conducted in accordance with the Declaration of Helsinki and approved by the Ethics Committee of the Research Institute of Hygiene, Occupational Pathology and Human Ecology of the Federal Medical Biological Agency (Approval No. 3, registration date: 2 June 2022). Informed consent was obtained from all subjects involved in this study. The period during which the patient groups were formed, and their neurological examination and blood sampling for subsequent analysis were performed was approximately one year—from mid-2022 to mid-2023. The dataset comprised 120 elderly patients (60 to 78 years old) living in St. Petersburg and registered in the geriatric departments of city clinics. Structured interviews were conducted to obtain socio-demographic data. Medical history and neurological examination data were collected. The patients were divided into five groups based on clinical diagnoses established according to standard criteria detailed in the prior study [6]. Exclusion criteria included data on progressive cancer and recent infectious diseases. As for the inclusion criteria, numerical labels (1–5) map to the following groups for consistency across analyses (Table 6). For validation purposes, a group of 28 patients aged 60–80 undergoing inpatient treatment at City Hospital No. 28 in St. Petersburg was formed in the fall of 2025.

4.3. Sample Preparation

Blood samples were collected, processed, and stored in accordance with international guidelines [32]. Blood was collected from the subjects on an empty stomach from the cubital vein into BD Vacutainer vacuum tubes, Becton Dickinson, Franklin Lakes, NJ, USA with anticoagulants (K3EDTA, heparin, and citrate). The plasma was stored at −70 °C until needed.

4.4. Biochemical and Hematological Analysis

Biochemical parameters were determined using a Sapphire 400 analyzer (Tokyo Boeki Medisys, Tokyo, Japan) and commercial RANDOX kits. Immunological and hematological studies of whole blood were performed on the day of blood collection. Baseline parameters and histograms of the distribution of erythrocytes, leukocytes, and platelets by volume were determined using a Medonic hematology analyzer (Boule Diagnostics, Spånga, Sweden).

4.5. Von Willebrand Factor and ADAMTS13

To quantify vWF in blood plasma using the standard method, the Technozym vWF:Ag ELISA kit (Technoclone GmbH, Vienna, Austria) was used. The activity of the von Willebrand factor in the plasma of patients was determined in the agglutination reaction of lyophilized platelets with ristocetin using a set of reagents from the Renam company (Moscow, Russia, No. AG-5). The determination of ADAMTS13 activity (VWF cleaving agent) in plasma was carried out using a Synergy 2 plate fluorometer (BioTech, Hudson, MA, USA) and the fluorescent substrate FRETS-VWF73.

4.6. Dataset Description

This study utilized a dataset comprising biochemical blood parameters from 120 patients divided into four groups: control (group 1), AIS (group 2), CCCI (group 3), DM (group 4) and SIVD (group 5). The dataset included 49 features such as RBC, MCV, PLT, and various enzymes (e.g., BChE, TP, vWF). In Table 7 below is the detail description of the dataset. Data preprocessing involved imputation of missing values using median filling, normalization via StandardScaler to ensure zero mean and unit variance, and outlier detection with a contamination level of 0.01 to reduce noise in heterogeneous biomedical data. All analyses were conducted using Python 3.10 with scikit-learn (v. 1.8.0), NumPy (v. 2.3.5), Pandas (v. 2.3.3), SHAP (v. 0.48.0), CatBoost (v. 1.2.8) and Matplotlib (v. 3.10.8) libraries.

4.6.1. Feature Selection Methods

To address heteroscedasticity and enhance model robustness, three feature selection methods were applied, with a specific focus on mitigating the risks associated with violated statistical assumptions.
  • Full Feature Set: All 49 blood parameters were used without filtering, serving as a baseline to assess the impact of the complete biomarker set on classification performance. This “wide” setup (pn) often increases estimate variability and promotes overfitting in high-dimensional biomedical data, necessitating comparative selection strategies.
  • One-Way ANOVA Method: Features were selected based on standard ANOVA, which evaluates differences between groups by comparing mean values. While this method assumes homogeneity of variances (homoscedasticity), it was employed as a comparative benchmark to identify biomarkers with significant variance across imbalanced groups. However, when this assumption is violated, the standard F-test becomes unreliable, leading to Type I error rates that deviate substantially from nominal significance levels and potentially producing spurious results.
  • Welch’s ANOVA Method (Variance-Aware): To account for the unequal variances typical in clinical biochemical data, Welch’s ANOVA was applied. This method is specifically designed to handle heteroscedasticity by adjusting the F-statistic and degrees of freedom
The justification for this approach is twofold:
(1)
Empirical Evidence: In our study, Levene’s test confirmed statistically significant heterogeneity of variance (p < 0.05) for several critical features, including iron (Fe), transferrin (Transf), RDW%, and LDL receptor (LDL). These findings directly demonstrated that the equal-variance assumption was untenable for our dataset.
(2)
Methodological Robustness: Welch’s ANOVA provides more accurate Type I error control and maintains comparable or superior statistical power to classical ANOVA even when variances are equal. It effectively eliminates the need for homogeneity as a prerequisite, ensuring that the identified biomarker signatures are not artifacts of unequal group variances.
Features were ranked according to their F-statistics, and the comparison of p-values between the two ANOVA variants revealed marked discrepancies in statistical significance for key parameters, underscoring the necessity of using variance-aware workflows to ensure the reliability of subsequent machine learning tasks.

4.6.2. ML Models

Six classical supervised classification algorithms were employed to predict patient group membership based on the selected blood biomarkers. Prior to model training, feature values were standardized to ensure consistent scaling across variables. The chosen models represent a complementary spectrum of analytical approaches: linear methods facilitate interpretability and identification of direct associations, whereas nonlinear models are capable of capturing intricate, nonlinearly separable patterns inherent in heterogeneous clinical datasets.
Linear Discriminant Analysis (LDA)
LDA is a classical statistical approach that projects data onto a lower-dimensional space to achieve maximum class separability. It constructs a linear combination of features that optimally distinguishes between predefined groups by maximizing the ratio of between-class to within-class variance [33].
Logistic Regression (LR)
LR was employed as a complementary linear classifier to benchmark model performance. It models the probability of class membership using the logistic (sigmoid) function and provides interpretable coefficients that quantify each feature’s contribution to the prediction. In this study, LR was extended to multiclass classification using the one-vs-rest strategy. To mitigate overfitting and improve generalization, both L1 (Lasso) and L2 (Ridge) regularization techniques were applied [34].
k-Nearest Neighbor (kNN)
KNN is a non-parametric, instance-based learning algorithm that classifies samples based on their similarity to previously observed data. Rather than building an explicit model during training, KNN stores all training examples and determines the class of a new sample by analyzing the labels of its closest neighbors in the feature space. kNN captured local patterns in the biochemical feature space [35].
Decision Tree (DT)
A non-linear, tree-based method that recursively partitions the feature space based on feature thresholds to minimize impurity. DT was included to model potential non-linear decision boundaries [36,37].
Random Forest (RF)
An ensemble of decision trees that aggregates predictions via majority voting, offering robustness to overfitting and high-dimensional data. RF was used here to handle feature redundancy and interactions among correlated biochemical markers [38]. Key parameters tuned during hyperparameter optimization include:
N estimators: This defines the number of decision trees in the RF. Values 50, 100 and 200 were tested to strike a balance between predictive performance and computational efficiency.
Max_depth: This controls the maximum number of splits per decision tree. A low value risks underfitting, while a high value may lead to overfitting. In this study, tree depths were adjusted to 5, 10 and 15 to capture meaningful patterns without excessive complexity.
These parameter adjustments resulted in a robust and accurate model, well-suited for handling the complexity and high dimensionality of biological data.
Support Vector Machines (SVM)
A kernel-based method that finds an optimal hyperplane, maximizing the margin between classes, particularly effective for high-dimensional data. SVM was applied to capture potentially non-linear separations in the biomarker space; tuned parameters included regularization parameter C, kernel type, and gamma parameter for RBF and polynomial kernels [39].

4.6.3. Cross-Validation

To ensure robust model evaluation and prevent overfitting, stratified k-fold cross-validation was employed with k = 3. Stratification preserved the class distribution within each fold, reflecting the multiclass nature of the dataset (four patient groups). The dataset was partitioned into non-overlapping training and validation subsets for each iteration, with models trained on the training set and evaluated on the test set. This process was repeated across all folds, and performance metrics were averaged to improve reliability and assess model sensitivity to training subsets.

4.6.4. Hyperparameter Optimization

Hyperparameters were tuned using GridSearchCV, an exhaustive search over a predefined parameter grid with internal 3-fold cross-validation, using the macro-averaged F1-score as the evaluation metric. The tuning parameters for each model included:
LDA: Solver type (‘svd’, ‘lsqr’, ‘eigen’) and shrinkage parameter (applicable for ‘lsqr’ and ‘eigen’ solvers). In this study, the default ‘svd’ solver was used without additional shrinkage optimization, as it is robust to multicollinearity and does not require regularization tuning.
Logistic Regression (LR): Regularization type (‘l1’, ‘l2’, ‘elasticnet’, ‘none’), regularization strength (C ∈ {0.01, 0.1, 1, 10, 100}), solver (‘saga’, compatible with all penalty types), and L1 ratio (∈{0, 0.5, 1}, applicable only for elastic-net regularization). Based on GridSearchCV results, the optimal LR configuration employed L1 regularization with C = 1, using the saga solver.
K-Nearest Neighbors (KNN): Number of neighbors (k ∈ {3, 5, 7, 9, 11}), weighting strategy (‘uniform’, ‘distance’), and distance metric (‘euclidean’, ‘manhattan’). Based on GridSearchCV results, the optimal KNN configuration employed k = 9 neighbors, distance-based weighting, and the Euclidean distance metric.
Decision Tree (DT): Maximum tree depth (None, 3, 5, 7, 9), minimum number of samples required to split an internal node (2, 5, 10), minimum number of samples required at a leaf node (1, 2, 4), and splitting criterion (‘gini’, ‘entropy’). Based on GridSearchCV results, the optimal DT configuration employed a maximum tree depth of 5, a minimum of 2 samples per leaf, a minimum of 2 samples to split an internal node, and the entropy-based splitting criterion.
Random Forest (RF): Number of trees (n_estimators ∈ {50, 100, 200}), maximum tree depth (None, 5, 10, 15), minimum number of samples required to split an internal node (2, 5, 10), minimum number of samples required at a leaf node (1, 2, 4), and splitting criterion (‘gini’, ‘entropy’). Based on GridSearchCV results, the optimal RF configuration employed 50 trees, no explicit limit on tree depth, a minimum of 5 samples to split an internal node, a minimum of 1 sample per leaf, and the entropy-based splitting criterion.
Support Vector Machine (SVM): Regularization parameter C (0.1, 1, 10, 100), kernel type (‘linear’, ‘rbf’, ‘poly’), and kernel coefficient gamma (‘scale’, ‘auto’; used only for ‘rbf’ and ‘poly’ kernels). Based on GridSearchCV results, the optimal SVM configuration employed an RBF kernel with C = 10 and γ = scale.
Optimal hyperparameters were selected based on the highest cross-validated F1-score to ensure strong generalization capability on unseen data.

4.6.5. Evaluation Metrics

Model performance was assessed using the macro-averaged F1-score, balancing precision and recall across all classes, appropriate for multiclass, imbalanced datasets.
Mean F1-scores and standard deviations were calculated across cross-validation folds for each pipeline (full feature set, one-way ANOVA, Welch’s ANOVA).
Ethical considerations complied with institutional guidelines, and data were anonymized to protect patient confidentiality [40].

5. Conclusions

In this study, a variance-robust and interpretable analytical framework was proposed and validated for biomarker discovery in multivariate clinical datasets characterized by small sample sizes, heteroscedasticity, and overlapping diagnostic categories. It was demonstrated that ignoring variance heterogeneity leads to unstable feature selection and potentially misleading interpretations, while the use of Welch’s ANOVA provides more robust results without sacrificing statistical power.
Integrating interpretable machine learning methods with SHAP analysis enables the identification of nonlinear, class-specific, and synergistic effects between biochemical markers that go beyond traditional single-variate statistics. The results demonstrate that interpretability is not simply an auxiliary tool for explaining patterns, but a critical component of robust clinical machine learning.
The proposed workflow can be directly applied to other biomedical fields where data heterogeneity and limited sample sizes make robust statistical inference difficult, providing a transparent link between statistical testing, predictive modeling, and clinical interpretation.

6. Limitations of the Research

Several limitations of the present study should be acknowledged. First, the analysis is based on a relatively small cohort (n = 120) while the feature space has high dimensionality (≈49 biochemical parameters), resulting in a pn configuration that increases estimator variability and the risk of overfitting despite cross-validation and model regularization. This limited sample size reduces the statistical power to detect small-to-moderate effects and constrains the generalizability of the findings.
Second, the dataset contains substantial inter-feature correlation and redundancy (e.g., RBC/HGB/HCT, PLT/PCT, Chol/LDL), which complicates model interpretability and can destabilize feature importance rankings; although SHAP provides local and global explainability, correlated predictors may still lead to ambiguous attribution of effects.
Third, several parameters exhibit pronounced heteroscedasticity (Levene’s test), and one of the groups shows marked separation in the clustering analysis. Unequal variances across groups violate assumptions of classical parametric tests and can bias significance rankings and cross-group comparisons; while Welch’s ANOVA and variance-aware procedures were employed to mitigate this issue, residual effects of heteroscedasticity and group-specific variance structure may still affect both statistical and machine-learning inferences.
Finally, class imbalance and the moderate agreement between unsupervised clusters and clinical labels (low ARI/NMI and modest Silhouette scores) limit the immediate clinical translatability of the classifiers. Together, these limitations mean that the identified biomarker signatures should be considered exploratory; future work must validate the results on larger, multi-center cohorts, apply more extensive external validation, and test robustness to alternative preprocessing, dimensionality-reduction, and class-balancing strategies.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/ijms27041805/s1.

Author Contributions

Conceptualization, N.V.G. and D.V.A.; data curation, P.I.P., E.A.K., N.G.V. and A.A.C.; investigation, E.A.K., P.I.P., E.A.K. and P.V.A.; supervision, N.V.G., D.V.A. and P.V.A.; writing—original draft, D.V.A., P.I.P. and A.A.C.; writing—review and editing, N.V.G. and R.O.J. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by the Russian Science Foundation (grant number 22-15-00155-∏).

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and approved by the Ethics Committee of the Research Institute of Hygiene, Occupational Pathology and Human Ecology of the Federal Medical Biological Agency (Approval No. 3, registration date: 2 June 2022).

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

a1-AGP Alpha-1-acid glycoprotein
ADAMTS13A Disintegrin And Metalloproteinase with a ThromboSpondin type 1 motif, member 13
AIArtificial Intelligence
AISAcute ischemic stroke
ALBAlbumin
ALPAlkaline phosphatase
ALT Alanine aminotransferase
ANCOVA Analysis of covariance
ANOVAAnalysis of variance
ARIAdjusted Rand Index
ASTAspartate aminotransferase
BilAcDirect bilirubin
BChE Butyrylcholinesterase
BMIBody mass index
CCCIChronic cerebral circulation insufficiency
CholTotal cholesterol
CIConfidence interval
CK-NAC Creatine kinase
CREA Creatinine
DMType 2 diabetes mellitus
DTDecision Tree
ERDEarly dementia risk
FeIron
GenderGender
GGTGamma-glutamyl transferase
GluGlucose
GRA%Granulocyte percentage
GRANGranulocyte absolute count
HC3Heteroskedasticity-Consistent Covariance Matrix Estimator
HCTHematocrit
HDHealthy donor
HDL High-density lipoprotein cholesterol
HGBHemoglobin
IQRInterquartile range
K Potassium
kNNk-Nearest Neighbors
LACLactate
LDALinear Discriminant Analysis
LDH Lactate dehydrogenase
LDLLow-density lipoprotein cholesterol
LPCRLarge platelet ratio
LRLogistic Regression
LYMLymphocyte absolute count
LYM%Lymphocyte percentage
MCHMean corpuscular hemoglobin
MCHCMean corpuscular hemoglobin concentration
MCVMean corpuscular volume
MID%/Mon% Monocyte percentage
MID/Mon Monocyte absolute count
MMSEMini-mental state examination
MLMachine Learning
MOCAMontreal cognitive assessment
MPVMean platelet volume
MVDMicrovascular dysfunction
NEFANon-esterified fatty acids
NMINormalized Mutual Information
PCT Plateletcrit
PDWPlatelet distribution width
PHOSInorganic phosphorus
PLTPlatelets
PON1Paraoxonase 1
PSIPopulation Stability Index
RBCRed blood cells
RDWRed cell distribution width
RDWaRed cell distribution width absolute
RFRandom Forest
Set CoverSet cover algorithm
SIVD Subcortical ischemic vascular dementia
SVMSupport Vector Machine
TIATransient ischemic attack
Tot. Prot., TPTotal protein
Transf Transferrin
TRIGSTriglycerides
Uric AcidUric acid
UREAUrea
VWFVon Willebrand factor
WBCWhite blood cells
WCSSWithin-cluster sum of squares

References

  1. Gunderson, C.G.; Bilan, V.P.; Holleck, J.L.; Nickerson, P.; Cherry, B.M.; Chui, P.; Bastian, L.A.; Grimshaw, A.A.; Rodwin, B.A. Prevalence of Harmful Diagnostic Errors in Hospitalised Adults: A Systematic Review and Meta-Analysis. BMJ Qual. Saf. 2020, 29, 1008–1018. [Google Scholar] [CrossRef] [Scilit]
  2. Auerbach, A.D.; Lee, T.M.; Hubbard, C.C.; Ranji, S.R.; Raffel, K.; Valdes, G.; Boscardin, J.; Dalal, A.K.; Harris, A.; Flynn, E.; et al. Diagnostic Errors in Hospitalized Adults Who Died or Were Transferred to Intensive Care. JAMA Intern. Med. 2024, 184, 164. [Google Scholar] [CrossRef] [Scilit]
  3. Newman-Toker, D.E.; Peterson, S.M.; Badihian, S.; Hassoon, A.; Nassery, N.; Parizadeh, D.; Wilson, L.M.; Jia, Y.; Omron, R.; Tharmarajah, S.; et al. Diagnostic Errors in the Emergency Department: A Systematic Review; AHRQ Comparative Effectiveness Reviews; Agency for Healthcare Research and Quality (US): Rockville, MD, USA, 2022.
  4. Winters, B.; Custer, J.; Galvagno, S.M.; Colantuoni, E.; Kapoor, S.G.; Lee, H.; Goode, V.; Robinson, K.; Nakhasi, A.; Pronovost, P.; et al. Diagnostic Errors in the Intensive Care Unit: A Systematic Review of Autopsy Studies. BMJ Qual. Saf. 2012, 21, 894–902. [Google Scholar] [CrossRef] [Scilit]
  5. Kai, Q. Artificial Intelligence Applications in Healthcare Diagnosis and Treatment in the Era of Big Data. Pac. Int. J. 2025, 8, 60–66. [Google Scholar] [CrossRef] [Scilit]
  6. Goncharov, N.V.; Popova, P.I.; Kudryavtsev, I.V.; Golovkin, A.S.; Savitskaya, I.V.; Avdonin, P.P.; Korf, E.A.; Voitenko, N.G.; Belinskaia, D.A.; Serebryakova, M.K.; et al. Immunological Profile and Markers of Endothelial Dysfunction in Elderly Patients with Cognitive Impairments. Int. J. Mol. Sci. 2024, 25, 1888. [Google Scholar] [CrossRef] [Scilit]
  7. Hua, S.; Dong, Z.; Wang, H.; Liu, T. Global, Regional, and National Burden of Ischemic Stroke in Older Adults (≥60 Years) from 1990 to 2021 and Projections to 2030. Front. Neurol. 2025, 16, 1567609. [Google Scholar] [CrossRef] [Scilit]
  8. Deshpande, A.; Zhang, L.Q.; Balu, R.; Yahyavi-Firouz-Abadi, N.; Badjatia, N.; Laksari, K.; Tahsili-Fahadan, P. Cerebrovascular Morphology: Insights into Normal Variations, Aging Effects and Disease Implications. J. Cereb. Blood Flow Metab. 2025, 45, 1249–1264. [Google Scholar] [CrossRef] [Scilit]
  9. Hu, S.; Ji, W.; Zhang, Y.; Zhu, W.; Sun, H.; Sun, Y. Risk Factors for Progression to Type 2 Diabetes in Prediabetes: A Systematic Review and Meta-Analysis. BMC Public Health 2025, 25, 1220. [Google Scholar] [CrossRef] [Scilit]
  10. Pervyshin, N.A. Clinical and Laboratory Characteristics of Elderly Patients with Type 2 Diabetes Mellitus in Samara, Russia: A Registry-Based Study. Russ. Open Med. J. 2025, 14, e0314. [Google Scholar] [CrossRef] [Scilit]
  11. Yang, H.-M. Vascular Dementia: From Pathophysiology to Therapeutic Frontiers. J. Clin. Med. 2025, 14, 6611. [Google Scholar] [CrossRef] [Scilit]
  12. Markus, H.S.; De Leeuw, F.E. Cerebral Small Vessel Disease: Recent Advances and Future Directions. Int. J. Stroke 2023, 18, 4–14. [Google Scholar] [CrossRef] [Scilit]
  13. Jin, C.; Bhaskar, S. Unifying Vascular Injury and Neurodegeneration: A Mechanistic Continuum in Cerebral Small Vessel Disease and Dementia. Eur. J. Neurosci. 2025, 62, e70246. [Google Scholar] [CrossRef] [Scilit]
  14. Vo, D.-K.; Trinh, K.T.L. Emerging Biomarkers in Metabolomics: Advancements in Precision Health and Disease Diagnosis. Int. J. Mol. Sci. 2024, 25, 13190. [Google Scholar] [CrossRef] [Scilit]
  15. Zhou, Y.; Zhu, Y.; Wong, W.K. Statistical Tests for Homogeneity of Variance for Clinical Trials and Recommendations. Contemp. Clin. Trials Commun. 2023, 33, 101119. [Google Scholar] [CrossRef] [Scilit]
  16. Rosopa, P.J.; Schaffer, M.M.; Schroeder, A.N. Managing Heteroscedasticity in General Linear Models. Psychol. Methods 2013, 18, 335–351. [Google Scholar] [CrossRef] [Scilit]
  17. Salsburg, D. Why Analysis of Variance Is Inappropriate for Multiclinic Trials. Control. Clin. Trials 1999, 20, 453–468. [Google Scholar] [CrossRef] [Scilit]
  18. Delacre, M.; Leys, C.; Mora, Y.L.; Lakens, D. Taking Parametric Assumptions Seriously: Arguments for the Use of Welch’s F-Test Instead of the Classical F-Test in One-Way ANOVA. Int. Rev. Soc. Psychol. 2019, 32, 13. [Google Scholar] [CrossRef] [Scilit]
  19. Strauss, E.J.; Kaplan, D.J.; Jazrawi, L.M. Biomarker Changes in ACL Deficient Knees Compared with Contralaterals. Orthop. J. Sports Med. 2016, 4, 2325967116S00139. [Google Scholar] [CrossRef] [Scilit]
  20. Jan, S.; Shieh, G. Sample Size Determinations for Welch’s Test in One-way Heteroscedastic ANOVA. Br. Math. Stat. Psychol. 2014, 67, 72–93. [Google Scholar] [CrossRef] [Scilit]
  21. Ahsan, M.M.; Luna, S.A.; Siddique, Z. Machine-Learning-Based Disease Diagnosis: A Comprehensive Review. Healthcare 2022, 10, 541. [Google Scholar] [CrossRef] [Scilit]
  22. Alhumaidi, N.H.; Dermawan, D.; Kamaruzaman, H.F.; Alotaiq, N. The Use of Machine Learning for Analyzing Real-World Data in Disease Prediction and Management: Systematic Review. JMIR Med. Inform. 2025, 13, e68898. [Google Scholar] [CrossRef] [Scilit]
  23. Ibrahim, I.; Abdulazeez, A. The Role of Machine Learning Algorithms for Diagnosing Diseases. J. Appl. Sci. Technol. Trends 2021, 2, 10–19. [Google Scholar] [CrossRef] [Scilit]
  24. Shipe, M.E.; Deppen, S.A.; Farjah, F.; Grogan, E.L. Developing Prediction Models for Clinical Use Using Logistic Regression: An Overview. J. Thorac. Dis. 2019, 11, S574–S584. [Google Scholar] [CrossRef] [Scilit]
  25. Wynants, L.; Van Calster, B.; Collins, G.S.; Riley, R.D.; Heinze, G.; Schuit, E.; Albu, E.; Arshi, B.; Bellou, V.; Bonten, M.M.J.; et al. Prediction Models for Diagnosis and Prognosis of Covid-19: Systematic Review and Critical Appraisal. BMJ 2020, 369, m1328, Erratum in BMJ 2020, 369, m2204. https://doi.org/10.1136/bmj.m1328.. [Google Scholar] [CrossRef] [Scilit]
  26. Starcke, J.; Spadafora, J.; Spadafora, J.; Spadafora, P.; Toma, M. The Effect of Data Leakage and Feature Selection on Machine Learning Performance for Early Parkinson’s Disease Detection. Bioengineering 2025, 12, 845. [Google Scholar] [CrossRef] [Scilit]
  27. Rajkomar, A.; Dean, J.; Kohane, I. Machine Learning in Medicine. N. Engl. J. Med. 2019, 380, 1347–1358. [Google Scholar] [CrossRef] [Scilit]
  28. Lundberg, S.M.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems; Guyon, I., Von Luxburg, U., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30, Available online: https://proceedings.neurips.cc/paper_files/paper/2017/file/8a20a8621978632d76c43dfd28b67767-Paper.pdf (accessed on 4 January 2026).
  29. Vani, M.S.; Sudhakar, R.V.; Mahendar, A.; Ledalla, S.; Radha, M.; Sunitha, M. Personalized Health Monitoring Using Explainable AI: Bridging Trust in Predictive Healthcare. Sci. Rep. 2025, 15, 31892. [Google Scholar] [CrossRef] [Scilit]
  30. Zhu, Z.; Wang, C.; Shi, L.; Li, M.; Li, J.; Liang, S.; Yin, Z.; Xue, Y. Integrating Machine Learning and the SHapley Additive exPlanations (SHAP) Framework to Predict Lymph Node Metastasis in Gastric Cancer Patients Based on Inflammation Indices and Peripheral Lymphocyte Subpopulations. J. Inflamm. Res. 2024, 17, 9551–9566. [Google Scholar] [CrossRef] [Scilit]
  31. Tanaka, M.; Kohjitani, H.; Yamamoto, E.; Morimoto, T.; Kato, T.; Yaku, H.; Inuzuka, Y.; Tamaki, Y.; Ozasa, N.; Seko, Y.; et al. Development of Interpretable Machine Learning Models to Predict In-hospital Prognosis of Acute Heart Failure Patients. ESC Heart Fail. 2024, 11, 2798–2812. [Google Scholar] [CrossRef] [Scilit]
  32. Barnes, P.W.; Mcfadden, S.L.; Machin, S.J.; Simson, E.; International Consensus Group for Hematology. The international consensus group for hematology review: Suggested criteria for action following automated CBC and WBC differential analysis. Lab. Hematol. 2005, 11, 83–90. [Google Scholar] [CrossRef] [Scilit]
  33. Tharwat, A.; Gaber, T.; Ibrahim, A.; Hassanien, A.E. Linear Discriminant Analysis: A Detailed Tutorial. AI Commun. 2017, 30, 169–190. [Google Scholar] [CrossRef] [Scilit]
  34. Sperandei, S. Understanding Logistic Regression Analysis. Biochem. Med. 2014, 24, 12–18. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Cover, T.; Hart, P. Nearest Neighbor Pattern Classification. IEEE Trans. Inform. Theory 1967, 13, 21–27. [Google Scholar] [CrossRef] [Scilit]
  36. Breiman, L.; Friedman, J.H.; Olshen, R.A.; Stone, C.J. Classification and Regression Trees, 1st ed.; Routledge: London, UK, 2017; ISBN 978-1-315-13947-0. [Google Scholar]
  37. Quinlan, J.R. Induction of Decision Trees. Mach. Learn. 1986, 1, 81–106. [Google Scholar] [CrossRef] [Scilit]
  38. Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
  39. Cortes, C.; Vapnik, V. Support-Vector Networks. Mach. Learn. 1995, 20, 273–297. [Google Scholar] [CrossRef] [Scilit]
  40. Hand, D.J.; Christen, P.; Kirielle, N. F*: An Interpretable Transformation of the F-Measure. Mach. Learn. 2021, 110, 451–456. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Levene test results for biochemical parameters across patient groups. The blue bars represent the Levene W statistic, while the red bars indicate p-values (logarithmic scale). The heavy dashed horizontal line denotes the p = 0.05 significance threshold. A p-value below this line (p < 0.05) indicates that the assumption of homogeneity of variances is statistically violated (heteroscedasticity). In this dataset, parameters to the left of PDW (e.g., Ca, ALB, ALP, Glyc, RDW%, LDL, LDH, Fe, UREA, BChE, vWF, HCT, RBC and HGB) demonstrate significant variance heterogeneity, necessitating the use of Welch’s ANOVA to ensure robust statistical inference.
Figure 1. Levene test results for biochemical parameters across patient groups. The blue bars represent the Levene W statistic, while the red bars indicate p-values (logarithmic scale). The heavy dashed horizontal line denotes the p = 0.05 significance threshold. A p-value below this line (p < 0.05) indicates that the assumption of homogeneity of variances is statistically violated (heteroscedasticity). In this dataset, parameters to the left of PDW (e.g., Ca, ALB, ALP, Glyc, RDW%, LDL, LDH, Fe, UREA, BChE, vWF, HCT, RBC and HGB) demonstrate significant variance heterogeneity, necessitating the use of Welch’s ANOVA to ensure robust statistical inference.
Ijms 27 01805 g001
Figure 2. Comparison of p-values obtained from classical one-way ANOVA and Welch ANOVA for all blood parameters. The bar chart displays Welch ANOVA results (blue bars, left group) and one-way ANOVA results (red bars, right group) for each parameter. p-values are plotted on a logarithmic scale to enhance visibility of small values. The dashed horizontal line at p = 0.05 indicates the threshold for statistical significance. Features are sorted according to Welch ANOVA p-values to facilitate comparison. This visualization highlights discrepancies between the two methods, illustrating the impact of variance heterogeneity on the assessment of statistical significance.
Figure 2. Comparison of p-values obtained from classical one-way ANOVA and Welch ANOVA for all blood parameters. The bar chart displays Welch ANOVA results (blue bars, left group) and one-way ANOVA results (red bars, right group) for each parameter. p-values are plotted on a logarithmic scale to enhance visibility of small values. The dashed horizontal line at p = 0.05 indicates the threshold for statistical significance. Features are sorted according to Welch ANOVA p-values to facilitate comparison. This visualization highlights discrepancies between the two methods, illustrating the impact of variance heterogeneity on the assessment of statistical significance.
Ijms 27 01805 g002
Figure 3. Top 5 blood parameters based on Welch’s ANOVA across patient groups.
Figure 3. Top 5 blood parameters based on Welch’s ANOVA across patient groups.
Ijms 27 01805 g003
Figure 4. Top 5 blood parameters based on One-ANOVA across patient groups.
Figure 4. Top 5 blood parameters based on One-ANOVA across patient groups.
Ijms 27 01805 g004
Figure 5. Reordered correlation heatmap of biochemical features. To enhance interpretability, variables were grouped using hierarchical clustering (Ward’s method) based on the distance metric d = 1 − ∣r∣. This approach automatically clusters biologically related markers—such as red cell indices (RBC, HGB, HCT), platelet parameters (PLT, PCT, LPCR), and lipid profiles (Chol, LDL)—revealing dense blocks of high multicollinearity (∣r∣ > 0.8). Such structural redundancy informs the subsequent use of SHAP interaction analysis to capture the synergistic roles of correlated predictors.
Figure 5. Reordered correlation heatmap of biochemical features. To enhance interpretability, variables were grouped using hierarchical clustering (Ward’s method) based on the distance metric d = 1 − ∣r∣. This approach automatically clusters biologically related markers—such as red cell indices (RBC, HGB, HCT), platelet parameters (PLT, PCT, LPCR), and lipid profiles (Chol, LDL)—revealing dense blocks of high multicollinearity (∣r∣ > 0.8). Such structural redundancy informs the subsequent use of SHAP interaction analysis to capture the synergistic roles of correlated predictors.
Ijms 27 01805 g005
Figure 6. Radar (spider) plot comparing classifiers by mean per-class F1-score under three feature-selection strategies. Each axis corresponds to a classifier (LDA, kNN, LR, DT, RF, SVM); closed polygons show average F1 across classes for a given feature set: all features, features selected by one-way ANOVA, and features selected by Welch’s ANOVA. Radial scale indicates mean F1 (higher values denote better class-level performance); plot range is the [0.2, 0.55] interval used for visualization.
Figure 6. Radar (spider) plot comparing classifiers by mean per-class F1-score under three feature-selection strategies. Each axis corresponds to a classifier (LDA, kNN, LR, DT, RF, SVM); closed polygons show average F1 across classes for a given feature set: all features, features selected by one-way ANOVA, and features selected by Welch’s ANOVA. Radial scale indicates mean F1 (higher values denote better class-level performance); plot range is the [0.2, 0.55] interval used for visualization.
Ijms 27 01805 g006
Figure 7. Confusion matrix of the Gradient Boosting model on the full dataset. Cell colors indicate the percentage of predictions: diagonal cells show correctly classified instances, while off-diagonal cells indicate misclassifications. Percentages are normalized per true class.
Figure 7. Confusion matrix of the Gradient Boosting model on the full dataset. Cell colors indicate the percentage of predictions: diagonal cells show correctly classified instances, while off-diagonal cells indicate misclassifications. Percentages are normalized per true class.
Ijms 27 01805 g007
Figure 8. Feature contributions for class 1 (HD): (a) Interactions between features; (b) Summary of individual feature impacts.
Figure 8. Feature contributions for class 1 (HD): (a) Interactions between features; (b) Summary of individual feature impacts.
Ijms 27 01805 g008
Figure 9. Feature contributions for class 2 (TIA): (a) Interactions between features; (b) Summary of individual feature impacts.
Figure 9. Feature contributions for class 2 (TIA): (a) Interactions between features; (b) Summary of individual feature impacts.
Ijms 27 01805 g009
Figure 10. Feature contributions for class 3 (CCCI): (a) Interactions between features; (b) Summary of individual feature impacts.
Figure 10. Feature contributions for class 3 (CCCI): (a) Interactions between features; (b) Summary of individual feature impacts.
Ijms 27 01805 g010
Figure 11. Feature contributions for class 4 (DM): (a) Interactions between features; (b) Summary of individual feature impacts.
Figure 11. Feature contributions for class 4 (DM): (a) Interactions between features; (b) Summary of individual feature impacts.
Ijms 27 01805 g011
Figure 12. Feature contributions for class 5 (SIVD): (a) Interactions between features; (b) Summary of individual feature impacts.
Figure 12. Feature contributions for class 5 (SIVD): (a) Interactions between features; (b) Summary of individual feature impacts.
Ijms 27 01805 g012
Figure 13. Mean SHAP Interaction Values (a) and Corresponding Feature Interaction Network (b) Across Classes.
Figure 13. Mean SHAP Interaction Values (a) and Corresponding Feature Interaction Network (b) Across Classes.
Ijms 27 01805 g013
Figure 14. Clustering analysis of the dataset. (a) Elbow method used to determine the optimal number of clusters (k = 5); (b) Two-dimensional t-SNE projection of the data points colored by their assigned KMeans cluster.
Figure 14. Clustering analysis of the dataset. (a) Elbow method used to determine the optimal number of clusters (k = 5); (b) Two-dimensional t-SNE projection of the data points colored by their assigned KMeans cluster.
Ijms 27 01805 g014
Figure 15. Heatmap of mean standardized blood parameters across clusters.
Figure 15. Heatmap of mean standardized blood parameters across clusters.
Ijms 27 01805 g015
Figure 16. Mean absolute SHAP values for each feature, broken down by class, illustrating their contribution to the model predictions in the validation cohort.
Figure 16. Mean absolute SHAP values for each feature, broken down by class, illustrating their contribution to the model predictions in the validation cohort.
Ijms 27 01805 g016
Table 1. F1 Metrics by Class for Various Classification Models.
Table 1. F1 Metrics by Class for Various Classification Models.
ModelData RepresentationClass 1 (HD)Class 2 (TIA)Class 3 (CCCI)Class 4 (DM)Class 5 (SIVD)Macro-F1
LDAFull0.37 ± 0.060.23 ± 0.090.18 ± 0.050.44 ± 0.230.74 ± 0.020.39
One-way0.36 ± 0.150.35 ± 0.080.30 ± 0.090.38 ± 0.040.86 ± 0.100.45
Welch0.30 ± 0.040.36 ± 0.090.38 ± 0.140.37 ± 0.030.78 ± 0.060.44
kNNFull0.48 ± 0.100.33 ± 0.080.19 ± 0.030.47 ± 0.140.80 ± 0.040.45
One-way0.55 ± 0.040.46 ± 0.110.18 ± 0.130.49 ± 0.130.89 ± 0.140.51
Welch0.54 ± 0.010.43 ± 0.030.18 ± 0.130.54 ± 0.130.91 ± 0.080.52
LRFull0.45 ± 0.070.43 ± 0.170.30 ± 0.090.39 ± 0.070.83 ± 0.020.48
One-way0.53 ± 0.140.24 ± 0.170.37 ± 0.150.42 ± 0.040.72 ± 0.160.46
Welch0.51 ± 0.140.38 ± 0.070.38 ± 0.160.41 ± 0.060.83 ± 0.040.50
DTFull0.36 ± 0.190.27 ± 0.060.26 ± 0.090.32 ± 0.080.58 ± 0.140.36
One-way0.43 ± 0.190.31 ± 0.090.30 ± 0.050.34 ± 0.190.58 ± 0.060.39
Welch0.37 ± 0.060.28 ± 0.070.36 ± 0.080.36 ± 0.130.67 ± 0.130.41
RFFull0.47 ± 0.120.47 ± 0.070.26 ± 0.120.54 ± 0.090.82 ± 0.10.51
One-way0.50 ± 0.120.44 ± 0.090.23 ± 0.160.49 ± 0.040.82 ± 0.150.50
Welch0.55 ± 0.130.43 ± 0.110.34 ± 0.180.52 ± 0.030.84 ± 0.080.54
SVMFull0.63 ± 0.030.37 ± 0.130.31 ± 0.170.48 ± 0.110.81 ± 0.110.52
One-way0.55 ± 0.140.29 ± 0.070.14 ± 0.090.37 ± 0.110.72 ± 0.040.41
Welch0.54 ± 0.170.35 ± 0.030.19 ± 0.130.41 ± 0.150.72 ± 0.110.44
Table 2. Performance metrics of the Gradient Boosting model for each class.
Table 2. Performance metrics of the Gradient Boosting model for each class.
ClassF1-ScorePrecisionRecall
Class 1 (HD)0.960.960.86
Class 2 (TIA)0.890.880.76
Class 3 (CCCI)0.871.000.91
Class 4 (DM)0.940.910.81
Class 5 (SIVD)0.950.910.81
Table 3. Composition of clusters with respect to ground-truth classes. The share of the dominant class in the entire cluster is shown in parentheses.
Table 3. Composition of clusters with respect to ground-truth classes. The share of the dominant class in the entire cluster is shown in parentheses.
ClassCluster Label
Cluster 1Cluster 2Cluster 3Cluster 4Cluster 5
Class 174561
Class 2171230
Class 367351
Class 41312511
Class 5100019
DominantDM—4 (0.46)DM—4 (0.40)TIA—2 (0.48)HD—1 (0.40)SIVD—5 (0.86)
Table 4. Significance of Differences Across Groups Relative to the Reference Group.
Table 4. Significance of Differences Across Groups Relative to the Reference Group.
Group Combination ρ
1–20.5277
1–30.4849
1–40.2492
1–50.5919
Table 5. Comparison of Feature Rankings Based on Spearman Correlation and SHAP Values.
Table 5. Comparison of Feature Rankings Based on Spearman Correlation and SHAP Values.
Feature Name d 12 R 12 s t a t R 2 S H A P d 13 R 13 s t a t R 3 S H A P d 14 R 14 s t a t R 4 S H A P d 15 R 15 s t a t R 5 S H A P
RBC3.018.015.09.013.04.0−5.05.010.013.041.028.0
MCV8.037.029.01.039.038.0−17.016.033.04.011.07.0
RDW%4.032.028.0−17.03.020.03.030.027.0−15.023.038.0
RDWa4.035.031.0−18.09.027.015.034.019.0−4.027.031.0
HCT2.029.027.0−1.023.024.0−1.014.015.08.034.026.0
PLT−14.011.025.09.037.028.09.040.031.013.030.017.0
MPV8.014.06.03.08.05.08.017.09.00.05.05.0
PDW7.016.09.0−5.02.07.00.08.08.07.022.015.0
PCT11.012.01.032.033.01.037.044.07.024.025.01.0
LPCR−10.08.018.0−4.07.011.0−12.09.021.0−10.019.029.0
WBC−19.017.036.012.031.019.0−1.038.039.016.037.021.0
HGB12.028.016.0−8.021.029.0−2.012.014.016.038.022.0
MCH−10.033.043.0−13.035.048.0−24.022.046.0−4.09.013.0
MCHC33.041.08.010.042.032.08.049.041.0−33.01.034.0
LYM3.05.02.09.011.02.01.02.01.08.010.02.0
GRAN−12.021.033.08.043.035.016.048.032.07.040.033.0
MID/Mon−2.02.04.012.015.03.027.029.02.015.018.03.0
LYM%−12.010.022.021.044.023.014.043.029.019.035.016.0
GRA%−17.013.030.08.045.037.03.045.042.09.036.027.0
MID%/Mon%16.023.07.020.026.06.016.020.04.03.013.010.0
BChE4.043.039.00.022.022.02.019.017.0−1.042.043.0
PON114.026.012.016.025.09.017.037.020.012.026.014.0
ALT−10.04.014.0−1.029.030.09.027.018.0−4.014.018.0
AST31.034.03.014.028.014.028.039.011.01.012.011.0
ALB2.019.017.0−5.016.021.07.031.024.03.045.042.0
Glu−33.09.042.00.049.049.0−5.042.047.0−1.046.047.0
GGT14.024.010.02.046.044.019.032.013.09.015.06.0
TP−4.036.040.02.019.017.04.041.037.0−3.033.036.0
PHOS−10.03.013.0−4.012.016.00.023.023.0−1.08.09.0
UREA 1.042.041.0−37.05.042.0−13.025.038.0−5.07.012.0
TRIGS2.07.05.024.034.010.018.021.03.02.06.04.0
CREA−2.030.032.0−30.04.034.0−21.015.036.0−15.017.032.0
Ca−2.045.047.0−31.010.041.0−27.01.028.0−26.04.030.0
ALP0.048.048.0−11.020.031.017.033.016.0−15.024.039.0
Glyc8.027.019.00.040.040.010.035.025.0−9.028.037.0
K−9.025.034.0−12.01.013.0−24.06.030.0−25.016.041.0
BilAc17.038.021.0−19.06.025.0−12.010.022.0−22.03.025.0
HDL 29.040.011.03.036.033.0−30.04.034.04.039.035.0
LDL −18.06.024.017.032.015.01.013.012.024.048.024.0
CK-NAC−15.031.046.0−2.024.026.02.028.026.0−13.032.045.0
Chol1.039.038.0−6.041.047.0−46.03.049.01.047.046.0
LDH 0.044.044.02.014.012.05.011.06.0−15.029.044.0
Ur.Acid1.046.045.01.047.046.0−27.018.045.011.031.020.0
LAC−22.01.023.09.017.08.01.036.035.013.021.08.0
NEFA12.047.035.012.030.018.019.024.05.0−17.02.019.0
Transf−27.022.049.0−16.027.043.0−14.026.040.00.049.049.0
Fe−5.015.020.0−27.018.045.0−41.07.048.0−4.044.048.0
a1-AGP−6.020.026.02.038.036.02.046.044.0−3.020.023.0
vWF12.049.037.09.048.039.04.047.043.03.043.040.0
Table 6. Patient groups and inclusion criteria.
Table 6. Patient groups and inclusion criteria.
LabelClassDiagnostic Criterian
1HD (healthy donors)No history of acute or chronic cerebrovascular accident, signs of diabetes or cognitive impairment.23
2AIS (acute ischemic stroke)At least one registered visit to the clinic in the last 3 years due to acute cerebrovascular accident, confirmed by instrumental analysis data.24
3CCCI (chronic cerebral circulation insufficiency)Neurological examination and instrumental analysis data indicate the presence of chronic cerebral circulation insufficiency.21
4DM (type 2 diabetes)Clinical and biochemical features of early and untreated type 2 diabetes.32
5SIVD (subcortical ischemic vascular dementia)Neurological examination and instrumental analysis data indicate the presence of subcortical ischemic vascular dementia.20
Table 7. Description of features in the dataset, including feature names, detailed descriptions, and units of measurement.
Table 7. Description of features in the dataset, including feature names, detailed descriptions, and units of measurement.
Feature NameDescriptionUnits
Patient groupDisease class (1 = HD, 2 = TIA, 3 = CCH, 4 = DM, 5 = SIVD)
Gender1= male, 0= female
RBCred blood cells106/μL
MCVmean corpuscular volumefL
RDW%red cell distribution width (relative value)%
RDWared cell distribution width (absolute value)fL
HCThematocrit%
PLTplatelets103/μL
MPVmean platelet volumefL
PDWplatelet distribution width by volume%
PCTplateletcrit%
LPCRlarge platelet coefficient%
WBCwhite blood cells103/μL
HGBhemoglobing/dL
MCHmean corpuscular hemoglobinpg
MCHCmean corpuscular hemoglobin concentrationg/dL
LYMlymphocytes (absolute)103/μL
GRANgranulocytes (absolute)103/μL
MID/Monmonocytes/medium cells (absolute)103/μL
LYM%lymphocytes%
GRA%granulocytes%
MID%/Mon%monocytes/medium cells%
BChEbutyrylcholinesteraseU/L
PON1paraoxonase 1 (enzyme)U/L
ALTalanine aminotransferaseU/L
ASTaspartate aminotransferaseU/L
ALBalbuming/L
Gluglucosemmol/L
GGTgamma-glutamyl transferaseU/L
Tot.Prottotal proteing/L
PHOSinorganic phosphorusmmol/L
UREAureammol/L
TRIGStriglyceridesmmol/L
CREAcreatinineμmol/L
ALPalkaline phosphataseU/L
Kpotassium (measurement method)mmol/L
BilAcdirect (conjugated) bilirubinμmol/L
HDLhigh-density lipoprotein cholesterolmmol/L
LDLlow-density lipoprotein cholesterolmmol/L
CK-NACcreatine kinaseU/L
Choltotal cholesterolmmol/L
LDHlactate dehydrogenaseU/L
Ur.Aciduric acidμmol/L
LAClactatemmol/L
NEFAnon-esterified fatty acidsmmol/L
Transftransferring/L
Feironμmol/L
a1-AGPalpha-1-acid glycoproteing/L
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Artamonov, D.V.; Popova, P.I.; Korf, E.A.; Voitenko, N.G.; Chernysheva, A.A.; Avdonin, P.V.; Jenkins, R.O.; Goncharov, N.V. Interpretable Machine Learning with SHAP Identifies Key Biomarkers in a Multi-Factorial Spectrum of Age-Related Neurological and Metabolic Conditions. Int. J. Mol. Sci. 2026, 27, 1805. https://doi.org/10.3390/ijms27041805

AMA Style

Artamonov DV, Popova PI, Korf EA, Voitenko NG, Chernysheva AA, Avdonin PV, Jenkins RO, Goncharov NV. Interpretable Machine Learning with SHAP Identifies Key Biomarkers in a Multi-Factorial Spectrum of Age-Related Neurological and Metabolic Conditions. International Journal of Molecular Sciences. 2026; 27(4):1805. https://doi.org/10.3390/ijms27041805

Chicago/Turabian Style

Artamonov, Daniil V., Polina I. Popova, Ekaterina A. Korf, Natalia G. Voitenko, Alisa A. Chernysheva, Pavel V. Avdonin, Richard O. Jenkins, and Nikolay V. Goncharov. 2026. "Interpretable Machine Learning with SHAP Identifies Key Biomarkers in a Multi-Factorial Spectrum of Age-Related Neurological and Metabolic Conditions" International Journal of Molecular Sciences 27, no. 4: 1805. https://doi.org/10.3390/ijms27041805

APA Style

Artamonov, D. V., Popova, P. I., Korf, E. A., Voitenko, N. G., Chernysheva, A. A., Avdonin, P. V., Jenkins, R. O., & Goncharov, N. V. (2026). Interpretable Machine Learning with SHAP Identifies Key Biomarkers in a Multi-Factorial Spectrum of Age-Related Neurological and Metabolic Conditions. International Journal of Molecular Sciences, 27(4), 1805. https://doi.org/10.3390/ijms27041805

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop