1. Introduction
Heart failure (HF) is a clinical syndrome caused by structural and/or functional abnormalities of the heart, leading to impaired ventricular filling and/or blood ejection [
1]. It affects approximately 64 million people globally, and its already high worldwide prevalence (over 10% in those over 70 years old) is expected to rise further due to an aging population, improved treatment and improved survival rates [
2], underpinning the need for timely diagnosis and risk management. Among the challenges in HF is the heterogeneity of the disease. Clinically relevant patient categories are defined based on the ejection fraction (EF), which is assessed via transthoracic echocardiography and calculated as the ratio of stroke volume to end-diastolic volume [
3]. Patients exhibiting preserved EF (HFpEF, EF ≥ 50%) receive different treatments than those with reduced EF (HFrEF, EF ≤ 40%) [
4]. However, there is also an EF zone between 41 and 49% called mildly reduced EF (HFmrEF), where patients present characteristics more similar to those of HFrEF and respond better to HFrEF than HFpEF therapies [
4]. Studies have also suggested that HFmrEF is a transitional state from HFpEF to HFrEF, and it is referred to as a gray zone [
4]. HFpEF comprises approximately 50% of heart failure cases [
2]. Despite the preserved ejection fraction, HFpEF is a condition related to high morbidity, mortality, and healthcare costs, making its timely identification clinically essential.
In clinical practice, ECG plays a supportive role in the diagnosis of HFpEF, helping to evaluate patients for referral to B-type natriuretic peptide (BNP)/NT-proBNP tests and further echocardiographic examination. According to guidelines, a normal or not particularly abnormal ECG and a low BNP/NT-proBNP value lower the suspicion of HFpEF and lead to underdiagnosed cases [
4]. In order to further investigate ECG value, previous work [
5] extensively studied traditional ECG with respect to structural heart abnormalities and low specificity for functional abnormalities, highlighting the need to develop novel methods for the evaluation of HF. Differences in ECG and HF subtypes were examined in [
6], suggesting that ventricular repolarization and ventricular depolarization markers can distinguish between patients with HFrEF and HFpEF, showcasing the role of ECG in the personalized risk assessment of HF subtypes. Differences do occur—not only among subtypes but also within the same subtype, resulting in different phenotypes that require precise medical interventions [
7,
8,
9].
Recent findings [
10,
11] suggest that HFpEF is often underdiagnosed using standard guidelines, indicating a pressing need for other non-invasive diagnostic devices and tools, including AI-DSS tools.
Previous studies have addressed this problem using different AI approaches to improve diagnosis, focusing mainly on the HFpEF vs. healthy classification. In [
12], an ML model (XGBoost) was developed by combining demographic, clinical, and echocardiography data to detect HFpEF vs. controls, demonstrating strong performance, with an AUROC of 0.89 (95% CI: 0.85–0.93) in an external validation dataset of 3349 samples. In [
13], an AI-ECG model was implemented based on a CNN solution, where patients were classified as HFpEF (HFA-PEFF score ≥ 5) or controls, achieving a good discriminative performance of AUROC = 0.81 (95% CI: 0.80–0.82) in an internal validation set of 2167 samples. Recently, in [
14], a ResNet-based deep learning model (ECG-AI) was extensively validated in 100.000 recordings for early identification of HFpEF based on ECG obtained up to six months prior to diagnosis, exhibiting impressive results, with AUROC = 0.79 (95% CI: 0.76–0.82). Similarly, in [
15], a CNN-based model demonstrated competitive results in external validation (samples = 203), with an AUROC of 0.80 (95% CI: 0.74–0.86).
Studies that address HFrEF vs. healthy or HF subtype stratification using ECG have received less attention. For HFrEF detection against healthy controls, a CNN model [
16] achieved an AUROC of 0.93 in identifying left ventricular (LV) systolic dysfunction in an unselected, large population setting. In [
17], a deep learning model (DeepECG-HFrEF) was validated in symptomatic HF patients with EF < 40% and achieved an AUROC of 0.844. Regarding HFrEF vs. HFpEF classification, ref. [
18] developed and evaluated traditional machine learning models; however evaluation was conducted on a single training/test split, limiting the assessment of the stability of the results. Another study [
19], employed interval-based features (QTc, QRS, ST-T, and TP) extracted from 24 h Holter recordings and classified HFrEF, HFmrEF, and HFpEF (303 patients) using machine learning models, achieving AUROCs above 0.98. A study with the same cohort and a group of authors from the latter approached HF stratification by encoding 24 h circadian HRV features and clinical information into multi-parameter polar images, achieving an AUROC of 0.833 [
20]. Additionally, using data from the MUSIC database and LightGBM, along with P-wave, QRS-complex and T-wave energies, ref. [
21] achieved an accuracy of 98.45% and AUROC of 0.9989 for HFrEF vs. HFmrEF vs. HFpEF classification.
Regarding heterogeneity beyond EF subtypes, in [
22], a transformer-based approach on large-scale electronic health records identified seven distinct HF subtypes, including unrecognized high-risk clusters, underscoring the heterogeneity of HF.
Beyond HF detection, wavelet-based ECG features combined with traditional machine learning models have also been explored for early-stage subclinical LV dysfunction [
23,
24,
25,
26], a condition related to HF pathophysiology, providing model interpretability through features suitable for specific physiological problems.
These works illustrate the variety of signal modalities and detection approaches applied to HF detection. Similarly, a meta-analysis of AI models for HF assessment showed that studies using CNN and larger datasets achieved consistently higher accuracy, while traditional machine learning models showed greater variability in their performance [
27]. However, a common limitation persists. Vectorcardiography (VCG) provides a three-dimensional representation of cardiac electrical activity, which may reveal more detailed information than the standard 12-lead ECG [
28]. In the cited studies, models based on deep learning methods and 12-lead ECG tend to offer limited interpretability. This deficiency does not allow clinicians to examine the features driving the classification, which is a critical requirement for individualized patient assessment and personalized care. Among the traditional machine learning approaches, ref. [
18] relied on a single bipolar limb lead (I) without a healthy control group, while [
19] used long-duration ambulatory recordings from a pseudo-orthogonal lead configuration, also without healthy individuals. All things considered, there is a motivation to develop interpretable, VCG-based approaches that incorporate features with physiological meaning, morphology and subtle rhythmic changes.
This study aimed to develop and internally evaluate a non-invasive ECG/VCG-based machine learning approach for HF assessment in individuals with sinus rhythm, with particular emphasis on HFpEF. The first endpoint was to identify patients with HFpEF from healthy individuals, and the second was to screen healthy individuals and patients with HFrEF and HFpEF. The model is intended as a potential screening tool to identify individuals who may benefit from further echocardiographic and biomarker-based evaluations rather than as a standalone diagnostic test.
4. Discussion
4.1. The Value of This Work in Personalized Medicine
Heart failure management needs differ among patients and require interventions that focus on each patient individually. To create representative clusters of phenotypes, each HF subtype must be precisely diagnosed and stratified. This study demonstrated that ECG, a widely available clinical examination, has the potential to aid in HF assessment, with some considerations regarding HF subtype classification. In detail, normal vs. HFpEF and normal vs. HFrEF models exhibited impressive and robust performance for multiple train/test splits. This is evident from the metrics: Recall = 0.880 (95% CI: 0.861–0.899), F1 score = 0.897 (95% CI: 0.885–0.909), and AUROC = 0.983 (95% CI: 0.979–0.987) for the first one and Recall = 0.932 (95% CI: 0.924–0.940), F1 score = 0.950 (95% CI: 0.945–0.955), and AUROC = 0.983 (95% CI: 0.980–0.986) for the latter one. Although the best models included SMOTE use,
Data S1 and S2 (see
Supplementary Materials) show that performance metrics without SMOTE differ to a very limited extent.
The HFrEF vs. HFpEF model showed that it could identify most of the true positives by its metrics—Recall = 0.749 (95% CI: 0.724–0.774) and AUROC = 0.766 (95% CI: 0.749–0.782)—but was susceptible to false positives. The difficulty of discriminating between HFrEF and HFpEF is also shown in the aggregated results from the n
2 classifiers.
Data S3 (see
Supplementary Materials) shows that without SMOTE, the models severely underperformed. It should be acknowledged that its use might produce a bias in the HFrEF vs. HFpEF model, and these findings should be interpreted cautiously due to the lack of an external validation set.
The findings suggest that the proposed approach may be more suitable as an HF screening tool, whereas EF subtype classification should currently be considered exploratory and requires validation in larger, well-phenotyped cohorts. However, most patients with HF are not undiagnosed, allowing clinicians to proceed with further tests to determine EF, identify the phenotype to which they belong and treat them accordingly.
4.2. Clinical Relevance of the Selected Features
HRV is consistently reduced in congestive HF, reflecting autonomic imbalance (increased sympathetic versus vagal tone), and it is generally related to cardiovascular risk [
75], constituting a potential biomarker for continuous monitoring, and prognostic value (especially in terms of longitudinal variations of HRV features). While HRV features are relevant to HF, they were not selected, which is consistent with the controversial findings reported in the review by Míková et al. [
76]. Although STD HR, SD1, SD2, NN50, pNN50, SDNN, RMSSD and Max HR were found to statistically significant between normal vs. HF and Max HR between HFrEF vs. HFpEF (see
Data S4 in Supplementary Materials), their importance was not higher than the ECG morphology features. The ultra-short term window for HRV feature calculation may limit diagnostic value.
HFrEF is characterized by more severe conduction abnormalities, such as a wide QRS, left bundle branch block (LBBB), pathological Q waves, and left ventricular hypertrophy (LVH) [
77], whereas HFpEF shows subtler changes in electrical activity, such as LVH with repolarization abnormalities, myocardial fibrosis, and diastolic dysfunction [
78]. These underlying pathologies leave a footprint in the frequency domain. Within the QRS complex, the low-frequency content is influenced by ventricular conduction time and depolarization such that slow conduction and widening of QRS distribute power towards the lowest frequencies, where normal QRS already concentrates most of its spectral energy [
79]. In contrast, QRS with high-frequency content is linked to fragmentation by fibrosis and microscarring [
80]. In the repolarization interval, structural remodeling (LVH and myocardial fibrosis) increases repolarization variability, resulting in secondary changes in the ST segment and alterations in T-wave morphology, which reflect heterogeneous recovery of the myocardium [
81,
82,
83].
These physiological associations were drawn from the broader literature on QRS and ST-T pathophysiology in heart disease. The CWT-based features identified in this study should be interpreted as hypothesis-generating electrical signatures rather than direct mechanistic biomarkers.
The feature selection process retained features concentrated in the lower portion of the spectrum. For the QRS complex, the selected features were located in the low band (10–50 Hz) with only a single high-band feature (VZ_QRS_High_RelEnergy_mean, 100–150 Hz) and only for the HFrEF vs. HFpEF model. Regarding the ST-T region, the selected features came from both the low (5–15 Hz) and mid (15–50 Hz) bands. The VX_ST_Mid_MeanAmp_std and VX_ST_Low_RelEnergy_mean repolarization features were higher in HF subtypes than in normals and higher in HFrEF than in HFpEF, indicating greater beat-to-beat variability of mid-band amplitude and a larger share of low-band energy concentrated at the lowest frequencies. This X-axis ST pattern reflects differences in left–right ventricular wall repolarization, potentially related to unstable and inconsistent repolarization due to structural substrate changes (such as fibrosis and ischemia). On the contrary, the beat-to-beat variability of the absolute low-band energy on both the X and Y axes (VX_ST_Low_Energy_std and VY_ST_Low_Energy_std) was reduced in HFrEF relative to normals. In the QRS complex, the low-band QRS energy (VX_QRS_Low_Energy_mean) was higher in HFpEF than in HFrEF. As this feature reflects the absolute band energy, it depends on the signal amplitude. Diffuse myocardial fibrosis is known to attenuate ECG voltage independently of left ventricular mass [
84]; therefore, the lower low-band energy in HFrEF may indicate voltage reduction rather than redistribution of the spectral content. The relative energy features might be better for assessing the low-frequency shift.
To interpret the predictive ability of the best n2 classifier, PFI was utilized. Using the MCC metric for permutation, ’sensitive’ behavior is expected to be observed, as it presents changes across all elements of the confusion matrix. Additionally, due to the small test-set sample, a high variance was anticipated. Indeed, the overall picture shows that the drops in the MCC value differ significantly for each feature. However, the IQRs and mean values show that all the selected features contribute positively to each classification pair. These findings were complimentarily validated by SHAP analysis on sample-level and confirmed the direction and ranking observed with PFI.
For the separation of normals from HF patients, the features of the X-axis repolarization region dominated the importance rankings, in concordance with the described ST-T characteristics and the role of repolarization in HF. The SHAP plots support this pattern, with X-axis features consistently ranking among the top model factors, showing a directional relationship between feature value and impact on model output.
Within the two HF subtypes, VX_QRS_Low_Energy_mean was the most important feature, in line with the expectation that conduction slowing in HFrEF redistributes QRS energy to the lowest band. This was similarly reflected in the SHAP plot, where VX_QRS_Low_Energy_mean ranked at the top, with lower feature values associated with HFrEF pathology, as expected. The single high-band QRS feature is consistent with fibrosis-related microstructural differences between subtypes; however, with only one surviving high-band feature and a comparatively smaller and less consistent SHAP contribution relative to the other features, this remains exploratory.
Across all three pairs of classifications, VX_ST_Mid_MeanAmp_std was assigned high importance, making it a biomarker on which the n
2 classifier was most consistently based, showing that average instability in high-frequency components during the repolarization phase might have diagnostic value in HF assessment. This was also reflected in the SHAP results, where it was ranked among the highest contributors, strengthening the evidence that it might be a highly influential biomarker. In
Appendix E, four representative scenarios are presented for HFpEF and HFrEF patients using SHAP to show the potential feature contributions in different clinical scenarios.
These features presented here are referred to as factors that might be associated with heart failure and the physiology behind the disease rather than as causal factors. In addition, the frequency-band ranges were not selected as the optimal ones to detect HF and should be considered exploratory. Further studies are needed to explore and evaluate different cutoffs for HF assessment.
4.3. Comparison with Relevant Published Models
Compared with existing approaches, the normal vs. HFpEF and normal vs. HFrEF models achieved an AUROC of 0.983, which is comparable to or exceeds previously reported results. For HFpEF detection, the AUROC range was 0.79–0.89 [
12,
13,
14,
15], and for HFrEF detection, the AUROC was 0.844–0.93 [
16,
17]. A direct comparison is limited by differences in population size, signal modalities and HF definition criteria. Different origins of healthy and HF patients and a lack of external validation should also be considered in this section.
Regarding the discrimination between HFrEF and HFpEF, the present model did not reach the performance reported by [
21], although this comparison should be interpreted with caution. The two studies used an identical source and relied on feature-based machine learning, but they differ in scope. Also, ref. [
21] did not include healthy subjects and stratified patients regardless of rhythm, in contrast with the current work, which was restricted to individuals in sinus rhythm. Additionally, in [
21], insightful findings about feature importance are not reported, hindering the interpretation of the results in general and the key differences during the classification task between the two studies. Notably, ref. [
21] was not externally validated—similarly to this study—and evaluation on independent datasets is required to establish generalizability. In a different cohort, ref. [
19] also addressed a three-classification problem within a population selected for prior myocardial infarction or ischemic heart disease, without a healthy control group, and used long-duration Holter recordings, while the present work addressed a binary distinction within a more heterogeneous HF cohort that also included healthy controls. However, convergence was observed at the feature level in the present study. Ref. [
19] identified QTc, QRS, ST-T and TP intervals as the most important features, relating this to ventricular remodeling and repolarization abnormalities. The PFI analysis in the present study similarly identified the ST-T- and QRS-region spectral features as important markers of HF. An interesting finding is that both [
19,
20] found that the HF subtype classification accuracy varies substantially by time of day. This has direct implications for the design of outpatient monitoring protocols and may explain some of the heterogeneity in the results between retrospective studies that do not take into account the time at which the ECG was recorded.
It is also worth mentioning related work regarding wavelet-based ECG feature approaches for LV dysfunction, as mentioned above. In [
23], wavelet ECG features and clinical variables were combined to estimate myocardial relaxation velocity and discriminate systolic and diastolic LV dysfunction in a multicenter cohort of 1202 subjects, achieving AUROCs of 0.75–0.94 across internal and external test sets. Similarly, ref. [
25] combined wavelet energy-based features (ewECG) with the ARIC HF clinical score to detect subclinical systolic and diastolic LV dysfunction in 398 individuals at risk of HF, achieving an AUROC of 0.83 in an independent test set of 111 participants. An earlier study from the same group utilized ewECG features and a RF classifier for the screening of 319 asymptomatic subjects. External validation showed a sensitivity of 93% and an AUROC of 0.76, with the authors estimating that this could reduce unnecessary echocardiography by 38%. In [
26], similar ewECG features were used in combination with clinical variables and NT-proBNP to detect subclinical LV dysfunction in patients with type 2 diabetes mellitus, achieving an AUROC of 0.81, significantly outperforming NT-proBNP and the ARIC HF score alone, with AUROC values of 0.56 and 0.67, respectively, in a validation cohort of 97 participants. These studies support the notion that traditional machine learning models based on wavelet-based ECG features can achieve performance comparable to that of more complex deep learning approaches while offering greater interpretability, consistent with the feature-based approach adopted in this work. However, these studies incorporated clinical data alongside ECG features, which the present study did not include, and they targeted subclinical LV dysfunction rather than the HF subtype distinction addressed here.
An important point is that the EF threshold definitions vary across studies, which limits direct comparability. In [
19], the ASE/EACVI criteria were adopted (HFrEF < 50%; HFmrEF, 50–55%; HFpEF > 55%), differing from the ESC/ACC-AHA thresholds used in this work. In addition, ref. [
85] showed that the AUROC for low-EF classification varied significantly with the chosen threshold (35–50%). This suggests that the performance metrics for EF classification tasks are threshold-dependent to some extent, and comparisons between studies using different EF thresholds should be evaluated circumspectly.
4.4. Dataset Heterogeneity
Normals and HF patients come from different databases, and differences in hardware, conditions and acquisition protocol might introduce bias to a model. First, in
Table 1, it is presented that normal and HF patients have an age gap, which is backed up by the HF epidemiology. In [
4], it was reported that HF affects approximately 1% of those under 55 years and over 10% of those aged 70 or older. This gap is difficult to avoid when relying only on publicly available databases with different origins.
Sex differences were not found to be significant, as normal and HF samples had nearly the same sex percentages, and HFpEF had a 10% greater female population than the other two classes. This difference is attributed to the difference in HF subtype epidemiology; HFpEF is more common in females, and HFrEF is more common in males [
4].
The drop in the HFrEF vs. HFpEF task might raise some flags, but HF is a complex syndrome characterized by comorbidity; thus, distinguishing HF subtypes from each other is more complicated than normals vs. HF. Also, the small number of dataset in this study is not ideal for fully investigating this kind of heterogeneity of HF.
Additionally, the most significant features selected for each task and the feature importance analysis are reasonable and in concordance with literature findings regarding the underlying pathophysiological mechanisms and what could really differ among the three classes. In addition, the findings of the
Appendix A.3 suggest that model predictions were mostly driven by physiological factors rather than dataset-specific differences; however, with such a small number of HF patients from PTB, further evaluation on larger cohorts is essential to ensure robustness and reliability.
4.5. Power Analysis Results
Post hoc power analysis showed that all three binary classifiers, achieved 0.84 power to detect AUROC ≥ 0.65 and ≥0.98 power at AUROC ≥ 0.70. The combined n
2 classifier had ≥0.99 power to confirm accuracy ≥ 0.65 against the NIR (0.531). These results, together with the 95% CIs from 100 repeated nested cross-validation iterations (
Table 6 and
Table 7), the metrics from the screening task (
Table 8) and the EPV ratios confirming model stability (
Appendix D) suggest that the sample size (and the HFpEF group) is sufficient for the reported analyses.
4.6. Limitations and Future Steps
Although the results suggest that this method can contribute to HF assessment, several limitations should be considered. First, the dataset was relatively small and imbalanced, particularly in the HFpEF group, increasing the risk of overfitting. Although SMOTE was deployed and assisted in the classification process, imbalance can still reduce the sensitivity of the minority class (HFpEF). Second, healthy controls and HF patients were derived from different open-access databases, introducing potential domain-shift bias related to the acquisition protocols, recording conditions, and population characteristics. Third, the study included only individuals with sinus rhythm in order to isolate morphological changes from rhythmic abnormalities. Thus, the selected features and models have not been validated for patients with arrhythmias, frequent ectopy or paced rhythm—which are common in real-world HF populations; applying them in practice would likely require a prior rhythm inspection step. Also, beyond subtype classification, a recent study using paired recordings showed that HF status changes over time [
86], suggesting that long-term assessment might provide diagnostic information and should not be limited to a single time point. Finally, this study uses cross-validation and tests with multiple splits, but the models were not validated in an external data source.
Future studies should include external validation of the models to assess their performance and robustness with current findings. Additionally, the development of new models with 12-lead ECGs is a reasonable extension, as this lead system is intrinsically linked to the actual clinical setting. Demographics and clinical characteristics of a patient, such as systolic/diastolic pressure, blood glucose levels, and natriuretic peptides, constitute a notable addition in the detection of HF [
87] and linking them with ECG markers.