Author Contributions
Conceptualization, G.K., G.N., G.D.B. and G.P.T.; Methodology, G.K., G.N., G.D.B. and G.P.T.; Software, G.K.; Formal Analysis, G.K., G.N., G.D.B. and G.P.T.; Investigation, G.K., G.N., G.D.B. and G.P.T.; Visualization, G.K.; Validation, G.N., G.D.B. and G.P.T.; Writing—original draft preparation, G.K.; Writing—review and editing, G.K., G.N., G.D.B. and G.P.T.; Supervision, G.N., G.D.B. and G.P.T.; Funding acquisition, G.P.T. All authors have read and agreed to the published version of the manuscript.
Figure 1.
Beat-annotation distribution in the MIT-BIH database prior to filtering. Normal beats are 82.9% of raw annotations; the unknown class (7.3%) was discarded. Among the four retained AAMI classes, Class is 89.4%, with the minority classes (, , and ) together only 10.6%—the primary motivation for generative augmentation.
Figure 1.
Beat-annotation distribution in the MIT-BIH database prior to filtering. Normal beats are 82.9% of raw annotations; the unknown class (7.3%) was discarded. Among the four retained AAMI classes, Class is 89.4%, with the minority classes (, , and ) together only 10.6%—the primary motivation for generative augmentation.
Figure 2.
Representative beat segments from the training split for each of the four AAMI classes described in
Section 2.1 (three randomly selected samples per row). Amplitude is in arbitrary units after global normalization and time is in samples at 360
.
Figure 2.
Representative beat segments from the training split for each of the four AAMI classes described in
Section 2.1 (three randomly selected samples per row). Amplitude is in arbitrary units after global normalization and time is in samples at 360
.
Figure 3.
Training class counts before and after synthetic augmentation (log scale). Grey bars are real beats; blue hatched layers show cumulative synthetic beats added at each ratio , where brings minority classes to parity with . Class has under 1% as many real beats as Class , making it most dependent on synthetic data.
Figure 3.
Training class counts before and after synthetic augmentation (log scale). Grey bars are real beats; blue hatched layers show cumulative synthetic beats added at each ratio , where brings minority classes to parity with . Class has under 1% as many real beats as Class , making it most dependent on synthetic data.
Figure 4.
Overview of the cVAE latent space and its role in the pipeline. The encoder maps ECG beats into the shared 32-dimensional latent space via the class-conditional posterior (steps 1–3). cVAE-only generation (step 4) samples unconditionally with decoder-only class conditioning. At inference, DDPM generates class-conditional latents (step 5), QLR applies a bounded correction (step 6), and the shared decoder reconstructs synthetic beats (step 7).
Figure 4.
Overview of the cVAE latent space and its role in the pipeline. The encoder maps ECG beats into the shared 32-dimensional latent space via the class-conditional posterior (steps 1–3). cVAE-only generation (step 4) samples unconditionally with decoder-only class conditioning. At inference, DDPM generates class-conditional latents (step 5), QLR applies a bounded correction (step 6), and the shared decoder reconstructs synthetic beats (step 7).
Figure 5.
Class-conditional latent DDPM. Noise added to
forms
; the denoiser
minimizes
. Sampling reverses this with classifier-free guidance (
) to recover
. The denoiser’s internal AdaLN-Zero block architecture is detailed in
Appendix A,
Figure A2.
Figure 5.
Class-conditional latent DDPM. Noise added to
forms
; the denoiser
minimizes
. Sampling reverses this with classifier-free guidance (
) to recover
. The denoiser’s internal AdaLN-Zero block architecture is detailed in
Appendix A,
Figure A2.
Figure 6.
QLR module. (
a) The PQC produces a correction direction, gated and scaled before being added to the normalized latent and mapped back to latent space (Equation (
3)). (
b) A DDPM-sampled latent is refined and compared to its source (stay loss) and the real class bank (MMD loss), jointly optimizing the PQC and gate (Equation (
4)).
Figure 6.
QLR module. (
a) The PQC produces a correction direction, gated and scaled before being added to the normalized latent and mapped back to latent space (Equation (
3)). (
b) A DDPM-sampled latent is refined and compared to its source (stay loss) and the real class bank (MMD loss), jointly optimizing the PQC and gate (Equation (
4)).
Figure 7.
Overview of the hybrid generative pipeline. ECG beats are encoded into a 32-dimensional cVAE latent space, which feeds cVAE sampling, class-conditional latent DDPM, and optional QLR. A shared decoder maps latents back to synthetic beats, combined with real beats to augment training. A 1D MobileNetV2 classifier is trained on the augmented set and evaluated via Macro F1, precision, and recall.
Figure 7.
Overview of the hybrid generative pipeline. ECG beats are encoded into a 32-dimensional cVAE latent space, which feeds cVAE sampling, class-conditional latent DDPM, and optional QLR. A shared decoder maps latents back to synthetic beats, combined with real beats to augment training. A 1D MobileNetV2 classifier is trained on the augmented set and evaluated via Macro F1, precision, and recall.
Figure 8.
Training and validation convergence curves for the cVAE (Panel (A)) and latent DDPM (Panel (B)). The KL annealing-induced rise in the ELBO loss (Panel (A), epochs 0–40) is characteristic of -VAE training; the loss descends monotonically after annealing completes. The EMA-smoothed DDPM loss (Panel (B)) decreases consistently from ∼0.42 to ∼0.16.
Figure 8.
Training and validation convergence curves for the cVAE (Panel (A)) and latent DDPM (Panel (B)). The KL annealing-induced rise in the ELBO loss (Panel (A), epochs 0–40) is characteristic of -VAE training; the loss descends monotonically after annealing completes. The EMA-smoothed DDPM loss (Panel (B)) decreases consistently from ∼0.42 to ∼0.16.
Figure 9.
Morphology comparison for Class . Real mean envelope overlaid with synthetic statistics for cVAE (A), DDPM (B), and DDPM + QLR (C). DDPM + QLR achieves the lowest RMSD (0.047), the smallest maximum absolute deviation (0.212), and ties DDPM for the highest cosine similarity (0.995)—the clearest and most consistent morphological improvement QLR provides across the three minority classes.
Figure 9.
Morphology comparison for Class . Real mean envelope overlaid with synthetic statistics for cVAE (A), DDPM (B), and DDPM + QLR (C). DDPM + QLR achieves the lowest RMSD (0.047), the smallest maximum absolute deviation (0.212), and ties DDPM for the highest cosine similarity (0.995)—the clearest and most consistent morphological improvement QLR provides across the three minority classes.
Figure 10.
Morphology comparison for Class . Real mean envelope (broadest of the three classes, reflecting high morphological variability) overlaid with synthetic statistics for (A) cVAE, a clear outlier (RMSD 0.140), reflecting the difficulty of modeling ’s multimodal structure without a diffusion stage; (B) DDPM, which achieves the best RMSD (0.070) and lowest max deviation (0.289); and (C) DDPM + QLR, which tracks closely (RMSD 0.070, max deviation 0.302) without further improvement.
Figure 10.
Morphology comparison for Class . Real mean envelope (broadest of the three classes, reflecting high morphological variability) overlaid with synthetic statistics for (A) cVAE, a clear outlier (RMSD 0.140), reflecting the difficulty of modeling ’s multimodal structure without a diffusion stage; (B) DDPM, which achieves the best RMSD (0.070) and lowest max deviation (0.289); and (C) DDPM + QLR, which tracks closely (RMSD 0.070, max deviation 0.302) without further improvement.
Figure 11.
Morphology comparison for Class . Real mean envelope overlaid with synthetic statistics for (A) cVAE, which achieves the best RMSD (0.120) and lowest maximum absolute deviation (0.448) for this class; (B) DDPM; and (C) DDPM + QLR, which attains the highest cosine similarity (0.999), marginally ahead of DDPM (0.999). All three methods capture the hybrid QRS morphology with similar fidelity.
Figure 11.
Morphology comparison for Class . Real mean envelope overlaid with synthetic statistics for (A) cVAE, which achieves the best RMSD (0.120) and lowest maximum absolute deviation (0.448) for this class; (B) DDPM; and (C) DDPM + QLR, which attains the highest cosine similarity (0.999), marginally ahead of DDPM (0.999). All three methods capture the hybrid QRS morphology with similar fidelity.
Figure 12.
PCA projections of the 32-dimensional cVAE latent space for minority classes (A), (B), and (C). Grey scatter points are DDPM-generated latent samples, and shaded regions show the KDE of the real latent bank. Distribution drift values are 0.06, 0.12, and 0.19 for classes , , and , respectively.
Figure 12.
PCA projections of the 32-dimensional cVAE latent space for minority classes (A), (B), and (C). Grey scatter points are DDPM-generated latent samples, and shaded regions show the KDE of the real latent bank. Distribution drift values are 0.06, 0.12, and 0.19 for classes , , and , respectively.
Figure 13.
Latent shift induced by the QLR module for Class . Hollow circles: DDPM latents before refinement; filled circles: quantum-refined latents, overlaid on the real latent KDE. Refined points shift toward higher-density regions without collapsing to a single mode.
Figure 13.
Latent shift induced by the QLR module for Class . Hollow circles: DDPM latents before refinement; filled circles: quantum-refined latents, overlaid on the real latent KDE. Refined points shift toward higher-density regions without collapsing to a single mode.
Figure 14.
Latent shift induced by the QLR module for Class . The QLR moves several outlier DDPM points toward the main density mass of the Ventricular latent distribution.
Figure 14.
Latent shift induced by the QLR module for Class . The QLR moves several outlier DDPM points toward the main density mass of the Ventricular latent distribution.
Figure 15.
Latent shift induced by the QLR module for Class . The multimodal Fusion latent distribution has two to three distinct clusters. The QLR realigns samples within each cluster without merging distinct modes, preserving intra-class morphological diversity.
Figure 15.
Latent shift induced by the QLR module for Class . The multimodal Fusion latent distribution has two to three distinct clusters. The QLR realigns samples within each cluster without merging distinct modes, preserving intra-class morphological diversity.
Figure 16.
Confusion matrix for the unaugmented baseline (seed 42, Macro F1 = 0.799). -class recall is high (96.3%) but precision is only 71.9%, generating a high false Ventricular alarm rate. -class recall is 57.4%, with 19.2% of Supraventricular beats misclassified as Ventricular.
Figure 16.
Confusion matrix for the unaugmented baseline (seed 42, Macro F1 = 0.799). -class recall is high (96.3%) but precision is only 71.9%, generating a high false Ventricular alarm rate. -class recall is 57.4%, with 19.2% of Supraventricular beats misclassified as Ventricular.
Figure 17.
SMOTE (
).
-class recall reaches 76.2%, but 29.2% of
-class beats are misclassified as Ventricular—higher than baseline (19.2%) and the deep generative methods (
Figure 18,
Figure 19 and
Figure 20)—suggesting SMOTE-interpolated
beats overlap with the Ventricular feature space.
Figure 17.
SMOTE (
).
-class recall reaches 76.2%, but 29.2% of
-class beats are misclassified as Ventricular—higher than baseline (19.2%) and the deep generative methods (
Figure 18,
Figure 19 and
Figure 20)—suggesting SMOTE-interpolated
beats overlap with the Ventricular feature space.
Figure 18.
cVAE (
).
-class recall improves to 98.9% and
-class precision improves markedly over baseline (71.9% to 89.8%).
-to-
confusion falls to just 0.9%, the lowest of any method shown in
Figure 16,
Figure 17,
Figure 18,
Figure 19 and
Figure 20.
Figure 18.
cVAE (
).
-class recall improves to 98.9% and
-class precision improves markedly over baseline (71.9% to 89.8%).
-to-
confusion falls to just 0.9%, the lowest of any method shown in
Figure 16,
Figure 17,
Figure 18,
Figure 19 and
Figure 20.
Figure 19.
DDPM (). -class recall rises to 97.7%, -class precision improves to 79.4%, and confusion falls to 3.5%.
Figure 19.
DDPM (). -class recall rises to 97.7%, -class precision improves to 79.4%, and confusion falls to 3.5%.
Figure 20.
DDPM + QLR (). -class precision reaches 88.5%. confusion (10.5%) is well below SMOTE (29.2%) and baseline (19.2%).
Figure 20.
DDPM + QLR (). -class precision reaches 88.5%. confusion (10.5%) is well below SMOTE (29.2%) and baseline (19.2%).
Figure 21.
Class
ablation, seed 42. (
Left): validation loss (Equation (
4)-equivalent objective) over training for QLR (689 parameters) and two classical MLP refiners (MLP-S: 292; MLP-M: 1072). (
Right): downstream
-class F1 and Macro F1 for each refiner plus the unrefined DDPM baseline.
Figure 21.
Class
ablation, seed 42. (
Left): validation loss (Equation (
4)-equivalent objective) over training for QLR (689 parameters) and two classical MLP refiners (MLP-S: 292; MLP-M: 1072). (
Right): downstream
-class F1 and Macro F1 for each refiner plus the unrefined DDPM baseline.
Table 1.
Real-beat counts per AAMI class and data partition after intra-patient temporal splitting across all 44 retained MIT-BIH records (80/10/10 split).
Table 1.
Real-beat counts per AAMI class and data partition after intra-patient temporal splitting across all 44 retained MIT-BIH records (80/10/10 split).
| Class | Train | Validation | Test |
|---|
| (Normal) | 72,136 | 9046 | 8943 |
| (Supraventricular) | 2186 | 252 | 343 |
| (Ventricular) | 5527 | 718 | 764 |
| (Fusion) | 708 | 53 | 42 |
| Total | 80,557 | 10,069 | 10,092 |
Table 2.
Morphological fidelity metrics comparing real and synthetic class-mean waveforms. RMSD: root-mean-square deviation. Max : maximum point-wise absolute deviation. CosSim: cosine similarity. Lower RMSD and Max and higher CosSim indicate better waveform alignment.
Table 2.
Morphological fidelity metrics comparing real and synthetic class-mean waveforms. RMSD: root-mean-square deviation. Max : maximum point-wise absolute deviation. CosSim: cosine similarity. Lower RMSD and Max and higher CosSim indicate better waveform alignment.
| Class | Method | RMSD | Max | CosSim |
|---|
| cVAE | 0.0518 | 0.3036 | 0.9935 |
| DDPM | 0.0474 | 0.2170 | 0.9946 |
| DDPM + QLR | 0.0471 | | |
| cVAE | 0.1400 | 0.6837 | 0.9838 |
| DDPM | | | |
| DDPM + QLR | 0.0700 | 0.3018 | 0.9845 |
| cVAE | | | 0.9951 |
| DDPM | 0.1269 | 0.6764 | 0.9988 |
| DDPM + QLR | 0.1232 | 0.6522 | |
Table 3.
Mean diversity and novelty on decoded-signal space over ten seeds.
Table 3.
Mean diversity and novelty on decoded-signal space over ten seeds.
| Class | Method | Diversity | Novelty |
|---|
| cVAE | 11.63 | 0.1023 |
| DDPM | 10.40 | 0.0391 |
| DDPM + QLR | 10.08 | 0.0389 |
| cVAE | 22.67 | 0.0876 |
| DDPM | 36.45 | 0.0280 |
| DDPM + QLR | 35.51 | 0.0281 |
| cVAE | 22.47 | 0.0620 |
| DDPM | 27.30 | 0.0117 |
| DDPM + QLR | 27.14 | 0.0131 |
Table 4.
Macro F1 (mean ± std, ten seeds) for each augmentation method across ratios . Unaugmented baseline: . Best result per ratio shown in bold.
Table 4.
Macro F1 (mean ± std, ten seeds) for each augmentation method across ratios . Unaugmented baseline: . Best result per ratio shown in bold.
| Method | | | | |
|---|
| SMOTE | | | | |
| cVAE | | | | |
| DDPM | | | | |
| DDPM + QLR | | | | |
Table 5.
Per-class precision (PPV) and recall (SEN) at each method’s best ratio (mean ± std, ten seeds): SMOTE, cVAE, DDPM at ; DDPM + QLR at . Best value per metric in bold. Top: macro-averaged PPV/SEN, Classes , . Bottom: Classes , .
Table 5.
Per-class precision (PPV) and recall (SEN) at each method’s best ratio (mean ± std, ten seeds): SMOTE, cVAE, DDPM at ; DDPM + QLR at . Best value per metric in bold. Top: macro-averaged PPV/SEN, Classes , . Bottom: Classes , .
| | All Classes | Class | Class |
| Method | PPV | SEN | PPV | SEN | PPV | SEN |
| Baseline | | | | | | |
| SMOTE | | | | | | |
| cVAE | | | | | | |
| DDPM | | | | | | |
| DDPM + QLR | | | | | | |
| | Class | Class | | |
| Method | PPV | SEN | PPV | SEN | | |
| Baseline | | | | | | |
| SMOTE | | | | | | |
| cVAE | | | | | | |
| DDPM | | | | | | |
| DDPM + QLR | | | | | | |