Next Article in Journal
Effects of Electrode Wear on Nugget Formation in Continuous Resistance Spot Welding of Aluminum Alloys
Previous Article in Journal
Investigation of Optimal Temperature Parameters in ABS Additive Printing
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Proceeding Paper

Impact of SMOTE Oversampling on Machine Learning Classifiers for Preeclampsia Prediction Under Severe Class Imbalance: Evidence from a Bulgarian Screening Cohort †

1
Department of Computer Systems and Technologies, Faculty of Electronics and Automation, Technical University of Sofia, 4000 Plovdiv, Bulgaria
2
Department of Obstetrics and Gynaecology, Medical University of Plovdiv, 4002 Plovdiv, Bulgaria
3
Center of Competence “Smart mechatronic, Eco-and Energy-Saving Systems and Technologies”, 4000 Plovdiv, Bulgaria
*
Author to whom correspondence should be addressed.
Presented at the 15th International Scientific Conference TechSys 2026—Engineering, Technologies and Systems, Plovdiv, Bulgaria, 14–16 May 2026.
Eng. Proc. 2026, 150(1), 84; https://doi.org/10.3390/engproc2026150084
Published: 27 July 2026

Abstract

This paper evaluates machine learning (ML) classifiers for predicting preeclampsia (PE) and pregnancy-induced hypertension (PIH) using first-trimester screening data from 1383 pregnant women in Plovdiv, Bulgaria (2018–2020). Three classifiers—Logistic Regression (LR), Extra Trees Classifier (ETC), and Voting Classifier (VC)—are compared across multiple prediction targets and feature configurations. The impact of SMOTE oversampling strategies on model performance in the context of a severe class imbalance (2.46% PE prevalence) is assessed. Logistic Regression achieves the highest AUC of 0.853 for preterm PE prediction without oversampling, while SMOTE significantly improves tree-based models (ETC: +0.058 AUC). A non-screened control cohort of 533 patients is evaluated separately using maternal characteristics alone (AUC 0.751). The ML model showed promising discrimination for preterm PE in this local cohort and warrants direct comparison with FMF-based risk stratification in future studies. These results support further validation of ML-based tools as potential components of future clinical decision support systems to extend systematic PE screening in resource-constrained settings.

1. Introduction

Preeclampsia is one of the most serious hypertensive complications of pregnancy, affecting 2–8% of pregnancies worldwide and contributing to approximately 76,000 maternal deaths annually [1,2]. The condition is defined by new-onset hypertension after the 20th gestational week, accompanied by end-organ manifestations including proteinuria, thrombocytopenia, impaired liver or renal function, pulmonary oedema, or central nervous system disturbances [3]. The global burden increased by 10.92% between 1990 and 2019 [4], and in resource-constrained settings such as Bulgaria and Eastern Europe, structural barriers—limited diagnostic infrastructure, uneven access to specialised obstetric care, and socioeconomic disparities—further limit the effectiveness of existing prevention strategies [5].
Early identification of high-risk patients during the first trimester enables aspirin prophylaxis to be initiated before 16 gestational weeks, which significantly reduces the incidence of preterm PE [6]. The Fetal Medicine Foundation (FMF) combined screening algorithm, incorporating maternal characteristics, mean arterial pressure (MAP), uterine artery pulsatility index (UtA-PI), and biochemical markers (PAPP-A, PlGF), achieves detection rates of 75–90% for early-onset PE at a 10% false-positive rate in large Western European cohorts [7]. Conventional risk stratification based on individual risk factors achieves detection of only 34–41% [8], highlighting the clinical value of multiparametric first-trimester screening.
Recent advances in machine learning have demonstrated further potential to improve predictive performance and adapt algorithms to local population characteristics. A 2025 systematic review of 11 studies encompassing 116,253 pregnancies showed that ML algorithms consistently outperform traditional statistical approaches [9]. Extreme Gradient Boosting (XGBoost) achieves an AUC of 0.955 for PE prediction. ML models have also demonstrated improved prediction of PE-associated adverse outcomes such as HELLP syndrome and maternal ICU admission [10]. At the same time, deep neural networks applied in the PREVAL study (n = 10,110) reached AUC 0.920 with an 84.4% detection rate and a 10% false-positive rate [11,12]. However, the majority of these models were trained on large, ethnically diverse cohorts and have not been validated in Eastern European populations. The deployment of such models in clinical practice also requires transparency and interpretability, particularly under the European AI Act (Regulation EU 2024/1689), which classifies AI systems for maternal and foetal health monitoring as high-risk [13].
Parallel developments in wearable sensor technologies are enabling continuous, non-invasive monitoring of physiological parameters throughout pregnancy. ECG-based AI models have demonstrated 85% AUC for PE prediction up to 90 days before clinical diagnosis [14], and multi-parameter wearable sensor platforms have been validated across both high-resource and low-resource settings [15,16]. The integration of real-time sensor data with ML-based risk stratification through Internet of Things (IoT) architectures creates an opportunity for continuous, adaptive monitoring that can compensate for the limitations of single-point first-trimester screening [17,18].
This paper presents a machine learning analysis of two complementary clinical datasets from the “Fetal Medicine” outpatient clinic in Plovdiv, Bulgaria: a first-trimester screening cohort (n = 1383; 2018–2020) and a non-screened control cohort (n = 533). Three classifier families are compared—Logistic Regression (LR), Extra Trees Classifier (ETC), and Voting Classifier (VC)—across binary and multiclass prediction targets, with systematic evaluation of SMOTE-based oversampling strategies. The study provides locally validated prediction models as a foundation for an AI-powered clinical decision support system targeting high-risk patients, geographically isolated communities, and socioeconomically disadvantaged groups within the Bulgarian healthcare system.

2. Materials and Methods

2.1. Study Population

This retrospective observational study analysed clinical data collected at the “Fetal Medicine” outpatient clinic in Plovdiv, Bulgaria. The dataset includes pregnant women examined between 30 January 2018 and 30 November 2020 as part of routine first-trimester screening for preeclampsia [19]. The study protocol was conducted in accordance with the Declaration of Helsinki and approved by the Ethics Committee of the Medical University of Plovdiv.
Two complementary cohorts were analysed:
  • The screening cohort consisted of 1383 pregnant women who underwent combined first-trimester screening between 11 + 0 and 13 + 6 weeks of gestation. Screening included maternal characteristics, mean arterial pressure measurements, uterine artery Doppler assessment, and biochemical markers.
  • A comparison cohort of 533 pregnant women who had not undergone first-trimester screening and presented later in pregnancy (approximately 22 weeks of gestation) was analysed separately. This cohort allowed for evaluation of predictive performance when only maternal characteristics were available.
Inclusion criteria for the study were:
  • Singleton pregnancy;
  • Gestational age between 11 + 0 and 13 + 6 weeks for the screening cohort;
  • Maternal age ≥18 years;
  • Availability of clinical follow-up data.
Pregnancies with major foetal anomalies or incomplete clinical data were excluded. Table 1 summarises the cohort characteristics.
The lower PE rate in the screened cohort (2.46% vs. 3.19%) is consistent with a protective effect of aspirin prophylaxis following high-risk classification.

2.2. Feature Sets and Prediction Targets

Features were organised into five incremental groups to assess the marginal contribution of each biomarker category (Table 2). Post-screening treatment variables (aspirin intake, antihypertensive medication, dosage) and post-delivery outcomes (birthweight, delivery mode) were excluded to prevent data leakage.
Three binary prediction targets were defined: pe_all (any PE, n = 34, prevalence 2.46%), preterm_pe (PE with delivery < 37 weeks, n = 19, 1.37%), and term_pe (PE with delivery ≥ 37 weeks, n = 15, 1.08%). A three-class target (pe_pih_multiclass: None/PE/PIH) was also evaluated (8.82% combined prevalence).

2.3. Machine Learning Models

Three classifiers were selected for their complementary strengths: Logistic Regression (LR, balanced class weights, linear, interpretable); Extra Trees Classifier (ETC, balanced class weights, ensemble of 100 randomised trees, captures nonlinear patterns); and Voting Classifier (VC, soft voting ensemble of Random Forest and Extra Trees). All models used balanced weight class (class_weight = “balanced”) to penalise minority-class misclassification proportionally to class imbalance.

2.4. Oversampling Strategies

Given the extreme class imbalance (2.46% PE), three SMOTE-based strategies were evaluated, applied exclusively within training folds: standard SMOTE (synthetic minority interpolation between k-nearest neighbours); BorderlineSMOTE (synthetic samples generated only near the decision boundary); and ADASYN (adaptive generation weighted toward harder-to-classify instances). Test data was never augmented, ensuring unbiased evaluation.

2.5. Validation Protocol

The 5 × 5-fold stratified cross-validation (25 evaluation folds) was used throughout, with stratification preserving class ratios in all folds. The primary metric was AUC-ROC (threshold-independent, appropriate for imbalanced data). Preprocessing used MinMaxScaler normalisation. All analyses were performed in Python 3.14 with scikit-learn 1.8.0 and imbalanced-learn 0.14.1. Because of the limited number of positive cases, model evaluation was performed using repeated stratified cross-validation to reduce variance in performance estimates and minimise overfitting.

3. Results

3.1. Model Comparison Without Oversampling

Table 3 presents the 5 × 5-fold cross-validated AUC-ROC for the three classifiers using full-screening features.
Logistic Regression significantly outperforms tree-based models on both prediction targets. Preterm PE is more predictable than all PE (AUC 0.853 vs. 0.766), which is consistent with the literature indicating that early-onset PE exhibits stronger first-trimester biomarker signatures. The limited number of positive cases (n = 34) is insufficient for tree-based models to learn robust nonlinear patterns without overfitting. This finding aligns with the general machine learning principle that simpler models generalise better when data is scarce.

3.2. Impact of SMOTE Oversampling

The effect of oversampling strategies on AUC-ROC is shown in Table 4, Figure 1 and Figure 2.
Standard SMOTE significantly improves tree-based models: ETC gains +0.058 AUC on preterm PE, nearly matching LR (0.842 vs. 0.853). However, SMOTE slightly degrades LR performance because LR’s balanced class weighting already handles imbalance internally, making oversampling redundant. BorderlineSMOTE underperforms standard SMOTE because accurate boundary estimation requires more positive samples than are available with only 19–34 cases. With standard SMOTE, all three models converge to AUCs of 0.84–0.85 for preterm PE.

3.3. Multiclass Classification (PE/PIH/None)

The results for the three-class prediction target are shown in Table 5 and Figure 3.
SMOTE degrades all models in the multiclass setting, unlike in binary PE prediction. With three classes and 122 total positive cases (34 PE and 88 PIH), class weighting alone is sufficient. BorderlineSMOTE slightly improves PIH detection (0.010 to 0.016 AUC) at the cost of significantly worse PE detection (−0.021 to −0.054 AUC)—a trade-off given that PE is the more clinically critical outcome. PE is consistently harder to predict than PIH across all models (AUC ~0.68–0.73 vs. ~0.71–0.77), reflecting its smaller sample size and greater clinical heterogeneity.

3.4. Control Group Results

The non-screened control cohort (n = 533) was evaluated using maternal characteristics alone (Table 6), as biochemical and full biophysical markers were not available at the mid-pregnancy visit.
Without biochemical and full biophysical biomarkers, prediction from maternal characteristics alone achieves AUC 0.710–0.751, approximately 0.05–0.10 lower than the screening group’s full model. Mean UtA-PI measured at ~22 weeks added negligible predictive value in this context. Preterm PE prediction was not feasible in the control group (only seven cases; AUC ~0.5), underscoring the need for first-trimester biomarkers for high-performance prediction. This is consistent with externally validated models using routine maternal characteristics alone, which achieve AUC 0.84 in large-scale temporal validation studies [20], though these were trained on substantially larger cohorts.

3.5. Comparison with FMF Screening Algorithm

The performance of the Logistic Regression model for predicting preterm PE (AUC 0.853) is within the range reported for first-trimester screening models in the literature. Although direct comparison with the FMF algorithm is limited by differences in model structure and operating thresholds, the results (Table 7) suggest that locally trained machine learning models may achieve clinically meaningful discrimination in regional populations.
Further studies using direct head-to-head comparisons with established screening algorithms are required to determine the potential role of machine learning models in clinical risk stratification.
Crucially, the LR model’s logistic coefficients are directly interpretable, fulfilling transparency requirements under the European AI Act for high-risk AI systems in maternal health monitoring [13].

4. Discussion

4.1. Clinical Relevance

The result of an AUC of 0.853 for preterm PE is a strong indicator of the discriminatory power of the model in the most severe forms of the condition. With a detection rate of 70–75% with 10% false positives, our algorithm is fully comparable to that of the Fetal Medicine Foundation (FMF). The key advantage here is that these indicators were achieved with data from only one Bulgarian centre, relying on routine measurements in the first trimester and without the need for closed commercial software.
The comparison between the two groups confirms the effectiveness of systematic screening and aspirin prophylaxis: the cases of premature PE decreased from 3.0% to 1.95%. However, the analysis also revealed two critical gaps: 24.1% of women at risk did not start aspirin therapy, and the detection of pregnancy-induced hypertension (PIH) remained low (nearly 45% of cases fell into the low-risk group). These data outline a clear need for AI systems to support decision-making. The integration of continuous monitoring through wearables and IoT architectures can compensate for the limitations of one-time screening, especially in remote regions of Bulgaria with difficult access to specialised help.

4.2. Model Selection and Oversampling

The dominance of Logistic Regression (LR) in this study confirms a basic statistical principle: with a small number of positive cases (only 34 in PE), simpler models are more robust. Complex ensemble methods (such as Extra Trees) are prone to overfitting when data are limited. The built-in weight balancing in LR proved to be more effective than the artificial generation of synthetic examples (SMOTE), which, in our case, introduced more noise than precision.
The observation that first-trimester biomarkers add about +0.05 AUC over maternal characteristics alone to the baseline model (0.766 vs. 0.71) is of direct relevance for health resource management. In hospitals where biochemical tests (PAPP-A, PlGF) are expensive or unavailable, the model relying only on maternal characteristics still offers meaningful risk stratification. This supports the idea of “stepped screening”: mass pre-selection based on history, followed by full testing only for patients at intermediate and high risk.

4.3. Strengths and Limitations

This study has several strengths. First, it uses real-world clinical data from a well-characterised first-trimester screening cohort. Second, it systematically evaluates multiple machine learning models and oversampling strategies under conditions of extreme class imbalance. Third, the use of repeated cross-validation provides a more stable estimate of model performance than a single train–test split.
Nevertheless, several limitations should be acknowledged. The number of preeclampsia cases was relatively small, which may limit the stability of performance estimates and increases the risk of optimistic bias. The study was conducted in a single centre, and external validation in independent populations is necessary before clinical implementation. In addition, the comparison cohort differed in timing of clinical assessment and available biomarkers, which limits direct comparisons between groups.

4.4. Future Research

Future research should focus on validating these models in larger multi-centre datasets and evaluating their performance in prospective clinical settings. Combining datasets from multiple centres could substantially increase the number of positive cases and allow for more robust evaluation of complex machine learning algorithms.
Integration of continuous physiological monitoring data—such as blood pressure measurements or wearable sensor data—may further improve predictive accuracy and enable dynamic risk stratification throughout pregnancy.

5. Conclusions

In this Bulgarian screening cohort, ML models demonstrated a promising ability to predict preterm preeclampsia using first-trimester clinical and biomarker data. Logistic Regression provided the most robust performance in the presence of severe class imbalance, while SMOTE oversampling improved the performance of tree-based models. These results support further development and external validation of interpretable ML approaches as potential complements to existing preeclampsia screening strategies.
This study presents an analysis of machine learning for predicting preeclampsia (PE) in a Bulgarian clinical setting. Based on data from 1383 screened women and a control group of 533 patients, we draw the following key conclusions:
  • The Logistic Regression model is the most effective for this dataset (for preterm PE). It outperforms more complex ensemble models that are prone to overfitting with a small number of positive cases.
  • SMOTE oversampling only helps tree models. While standard SMOTE improves the results of Extra Trees Classifier by up to +0.038–0.058 AUC, it is redundant and even slightly detrimental for LR, which performs better by internally balancing the weights.
  • First-trimester biomarkers are critical. Adding data from MAP, PlGF, PAPP-A, and UtA-PI increases accuracy by about +0.05 AUC compared to models relying only on maternal characteristics. This justifies their inclusion in the standard protocol.
  • ML models are a real alternative to the FMF algorithm. The achieved detection of 70–75% with 10% false positives is comparable to commercial software, but offers full transparency according to the EU AI Act requirements.
  • Prevention gap: Despite the effectiveness of aspirin prophylaxis, 24.1% of high-risk patients do not adhere to therapy. This is a serious niche for future monitoring and reminder systems.
Future work will focus on integration of continuous wearable sensor data (blood pressure, ECG, PPG) into a real-time AI clinical decision support prototype; threshold optimisation for clinical deployment; and a cost-effectiveness analysis comparing ML-based screening with the FMF algorithm in the Bulgarian healthcare context.

Author Contributions

Conceptualisation, M.S. and B.S.; methodology, B.S. and M.S.; software, V.D.; validation, V.D., M.S. and B.S.; formal analysis, V.D. and M.S.; investigation, V.D.; resources, V.D. and B.S.; data curation, V.D.; writing—original draft preparation, V.D.; writing—review and editing, M.S. and B.S.; visualisation, V.D.; supervision, M.S.; project administration, M.S.; funding acquisition, M.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the European Regional Development Fund within the OP “Research, Innovation and Digitalization Programme for Intelligent Transformation 2021–2027”, Project No. BG16RFPR002-1.014-0005, Center of Competence “Smart Mechatronics, Eco- and Energy Saving Systems and Technologies”.

Institutional Review Board Statement

This study was conducted in accordance with the Declaration of Helsinki and approved by the Ethical Committee of the Medical University of Plovdiv (Protocol No.6/07.07.2022). All participants provided written informed consent.

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

The clinical datasets supporting the findings of this study are available from the corresponding author upon reasonable request, subject to institutional ethical approval.

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

Abbreviations

The following abbreviations are used in this manuscript:
ADASYNAdaptive Synthetic Sampling
ACOGAmerican College of Obstetricians and Gynecologists
AIArtificial Intelligence
AUCArea Under the Curve
BMIBody Mass Index
ECGElectrocardiogram
ETCExtra Trees Classifier
FGRFoetal Growth Restriction
FMFFetal Medicine Foundation
FPRFalse-Positive Rate
GDMGestational Diabetes Mellitus
IoTInternet of Things
LRLogistic Regression
MAPMean Arterial Pressure
MLMachine Learning
MoMMultiple of the Median
NICENational Institute for Health and Care Excellence
NICUNeonatal Intensive Care Unit
PAPP-APregnancy-Associated Plasma Protein A
PEPreeclampsia
PIHPregnancy-Induced Hypertension
PlGFPlacental Growth Factor
PPGPhotoplethysmography
SMOTESynthetic Minority Oversampling Technique
UtA-PIUterine Artery Pulsatility Index
VCVoting Classifier

References

  1. Say, L.; Chou, D.; Gemmill, A.; Tunçalp, Ö.; Moller, A.-B.; Daniels, J.; Gülmezoglu, A.M.; Temmerman, M.; Alkema, L. Global causes of maternal death: A WHO systematic analysis. Lancet Glob. Health 2014, 2, E323–E333. [Google Scholar] [CrossRef]
  2. Duley, L. The Global Impact of Pre-eclampsia and Eclampsia. Semin. Perinatol. 2009, 33, 130–137. [Google Scholar] [CrossRef] [PubMed]
  3. Magee, L.A.; Nicolaides, K.H.; von Dadelszen, P. Preeclampsia. N. Engl. J. Med. 2022, 386, 1817–1832. [Google Scholar] [CrossRef] [PubMed]
  4. Wang, W.; Xie, X.; Yuan, T.; Wang, Y.; Zhao, F.; Zhou, Z.; Zhang, H. Epidemiological trends of maternal hypertensive disorders of pregnancy at the global, regional, and national levels: A population-based study. BMC Pregnancy Childbirth 2021, 21, 364. [Google Scholar] [CrossRef] [PubMed]
  5. European Commission. State of Health in the EU: Bulgaria Country Health Profile 2023; OECD Publishing: Paris, France, 2023. [Google Scholar]
  6. Rolnik, D.L.; Wright, D.; Poon, L.C.; O’Gorman, N.; Syngelaki, A.; de Paco Matallana, C.; Akolekar, R.; Cicero, S.; Janga, D.; Singh, M.; et al. Aspirin versus Placebo in Pregnancies at High Risk for Preterm Preeclampsia. N. Engl. J. Med. 2017, 377, 613–622. [Google Scholar] [CrossRef] [PubMed]
  7. Chaemsaithong, P.; Sahota, D.S.; Poon, L.C. First trimester preeclampsia screening and prediction. Am. J. Obstet. Gynecol. 2022, 226, S1071–S1097.e2. [Google Scholar] [CrossRef] [PubMed]
  8. O’ Gorman, N.; Wright, D.; Poon, L.C.; Rolnik, D.L.; Syngelaki, A.; De Alvarado, M.; Carbone, I.F.; Dutemeyer, V.; Fiolna, M.; Frick, A.; et al. Multicenter screening for pre-eclampsia by maternal factors and biomarkers at 11-13 weeks’ gestation: Comparison with NICE guidelines and ACOG recommendations. Ultrasound Obstet. Gynecol. 2017, 49, 756–760. [Google Scholar] [CrossRef] [PubMed]
  9. Dkeen, N.O.M.; Radwan, M.E.D.; Zumam, I.A.A.; Mohamed, N.A.A.E.; Abdelmahmoud, E.M.A.; Magboul, N.M.E. Artificial Intelligence Applications in Obstetric Risk Prediction: A Systematic Review of Machine Learning Models for Preeclampsia. Cureus 2025, 17, e83961. [Google Scholar] [CrossRef] [PubMed]
  10. Schmidt, L.J.; Rieger, O.; Neznansky, M.; Hackelöer, M.; Dröge, L.A.; Henrich, W.; Higgins, D.; Verlohren, S. A machine-learning–based algorithm improves prediction of preeclampsia-associated adverse outcomes. Am. J. Obstet. Gynecol. 2022, 227, 77.e1–77.e30. [Google Scholar] [CrossRef] [PubMed]
  11. Gil, M.M.; Cuenca-Gómez, D.; Rolle, V.; Pertegal, M.; Díaz, C.; Revello, R.; Adiego, B.; Mendoza, M.; Molina, F.S.; Santacruz, B.; et al. Validation of machine-learning model for first-trimester prediction of pre-eclampsia using cohort from PREVAL study. Ultrasound Obstet. Gynecol. 2024, 63, 68–74. [Google Scholar] [CrossRef] [PubMed]
  12. Ansbacher-Feldman, Z.; Syngelaki, A.; Meiri, H.; Cirkin, R.; Nicolaides, K.H.; Louzoun, Y. Machine-learning-based prediction of pre-eclampsia using first-trimester maternal characteristics and biomarkers. Ultrasound Obstet. Gynecol. 2022, 60, 739–745. [Google Scholar] [CrossRef] [PubMed]
  13. European Parliament and Council. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Available online: http://data.europa.eu/eli/reg/2024/1689/oj (accessed on 9 April 2026).
  14. Butler, L.; Gunturkun, F.; Chinthala, L.; Karabayir, I.; Tootooni, M.S.; Bakir-Batu, B.; Celik, T.; Akbilgic, O.; Davis, R.L. AI-based preeclampsia detection and prediction with electrocardiogram data. Front. Cardiovasc. Med. 2024, 11, 1360238. [Google Scholar] [CrossRef] [PubMed]
  15. Ryu, D.; Kim, D.H.; Price, J.T.; Lee, J.Y.; Chung, H.U.; Allen, E.; Walter, J.R.; Jeong, H.; Cao, J.; Kulikova, E.; et al. Comprehensive pregnancy monitoring with a network of wireless, soft, and flexible sensors in high- and low-resource health settings. Proc. Natl. Acad. Sci. USA 2021, 118, e2100466118. [Google Scholar] [CrossRef] [PubMed]
  16. Zhou, S.; Park, G.; Longardner, K.; Lin, M.; Qi, B.; Yang, X.; Gao, X.; Huang, H.; Chen, X.; Bian, Y.; et al. Clinical validation of a wearable ultrasound sensor of blood pressure. Nat. Biomed. Eng. 2025, 9, 865–881. [Google Scholar] [CrossRef] [PubMed]
  17. Munyao, M.M.; Maina, E.M.; Mambo, S.M.; Wanyoro, A. Real-time pre-eclampsia prediction model based on IoT and machine learning. Discov. Internet Things 2024, 4, 10. [Google Scholar] [CrossRef]
  18. Marques, J.A.L.; Han, T.; Wu, W.; Madeiro, J.P.D.V.; Neto, A.V.L.; Gravina, R.; Fortino, G.; de Albuquerque, V.H.C. IoT-Based Smart Health System for Ambulatory Maternal and Fetal Monitoring. IEEE Internet Things J. 2021, 8, 16814–16824. [Google Scholar] [CrossRef]
  19. Stoilov, B.T. Assessment of Preeclampsia Screening and Prevention with Acetylsalicylic Acid in a High-Risk Group of Women in the Bulgarian Population. Ph.D. Thesis, Medical University of Plovdiv, Plovdiv, Bulgaria, 2023. [Google Scholar]
  20. Tiruneh, S.A.; Rolnik, D.L.; Teede, H.; Enticott, J. Temporal validation of machine learning models for pre-eclampsia prediction using routinely collected maternal characteristics: A validation study. Comput. Biol. Med. 2025, 191, 110183. [Google Scholar] [CrossRef] [PubMed]
Figure 1. ROC curves for preterm PE prediction comparing no oversampling (left), standard SMOTE (centre), and BorderlineSMOTE (right). Standard SMOTE brings tree-based models to near-parity with LR.
Figure 1. ROC curves for preterm PE prediction comparing no oversampling (left), standard SMOTE (centre), and BorderlineSMOTE (right). Standard SMOTE brings tree-based models to near-parity with LR.
Engproc 150 00084 g001
Figure 2. ROC curves for all PE predictions with the same three oversampling strategies.
Figure 2. ROC curves for all PE predictions with the same three oversampling strategies.
Engproc 150 00084 g002
Figure 3. One-vs-rest ROC curves for the 3-class model (None/PE/PIH).
Figure 3. One-vs-rest ROC curves for the 3-class model (None/PE/PIH).
Engproc 150 00084 g003
Table 1. Cohort characteristics: screened vs. non-screened group.
Table 1. Cohort characteristics: screened vs. non-screened group.
ParameterScreened Group (n = 1383)Control Group (n = 533)
Visit timing11–13 + 6 weeks (1st trimester)~22 weeks (mid-pregnancy)
PE cases34 (2.46%)17 (3.19%)
PIH cases88 (6.36%)36 (6.75%)
Unaffected1261 (91.18%)480 (90.06%)
Mean age (years)30.0 ± 5.129.7 ± 5.2
Mean BMI24.2 ± 5.025.5 ± 4.8
Caesarean section rate63.50%63.00%
Biomarkers availablePlGF, PAPP-A, MAP MoM, UtA-PIUtA-PI only
Table 2. Feature sets used for model evaluation.
Table 2. Feature sets used for model evaluation.
Feature Setn FeaturesDescription
maternal_only10Age, height, weight, BMI, parity, conception, smoking, family history of PE, chronic HTN, SLE
maternal_utpi11Maternal + mean uterine artery PI
maternal_biophysical12Maternal + MAP MoM + mean UtA-PI
maternal_biochemical12Maternal + PlGF MoM + PAPP-A MoM
full_screening14All maternal + biophysical + biochemical markers
Table 3. AUC-ROC (95% CI) for binary PE prediction, full-screening features, n = 1383.
Table 3. AUC-ROC (95% CI) for binary PE prediction, full-screening features, n = 1383.
Modelpe_all (n = 34)preterm_pe (n = 19)
Logistic Regression0.766 (0.733–0.797)0.853 (0.820–0.885)
Extra Trees Classifier0.678 (0.647–0.709)0.784 (0.734–0.830)
Voting Classifier0.691 (0.658–0.724)0.802 (0.755–0.846)
Table 4. AUC-ROC comparison across oversampling strategies (full_screening features).
Table 4. AUC-ROC comparison across oversampling strategies (full_screening features).
Model/TargetNo SMOTEStandard SMOTEBorderline SMOTE
LR—pe_all0.7660.755 (−0.011)0.691 (−0.075)
ETC—pe_all0.6780.690 (+0.012)0.679 (+0.001)
VC—pe_all0.6910.703 (+0.012)0.698 (+0.007)
LR—preterm_pe0.8530.847 (−0.006)0.769 (−0.084)
ETC—preterm_pe0.7840.842 (+0.058)0.805 (+0.021)
VC—preterm_pe0.8020.840 (+0.038)0.816 (+0.014)
Table 5. Multiclass AUC-ROC (95% CI), full_screening features, 3-class target, n = 1383.
Table 5. Multiclass AUC-ROC (95% CI), full_screening features, 3-class target, n = 1383.
ModelAUC MacroAUC NoneAUC PEAUC PIH
Logistic Regression0.757 (0.734–0.777)0.7790.7290.762
Voting Classifier0.724 (0.703–0.744)0.7440.7020.727
Extra Trees Classifier0.711 (0.691–0.729)0.7360.6810.715
Table 6. Five-fold CV results on control group (n = 533), maternal-only features.
Table 6. Five-fold CV results on control group (n = 533), maternal-only features.
ModelTargetAUC
Logistic Regressionpe_pih_multiclass (OvR)0.71
Logistic Regressionpe_all0.751
Extra Trees Classifierpe_all0.613
Voting Classifierpe_all0.594
Table 7. Comparison of ML model vs. FMF algorithm performance.
Table 7. Comparison of ML model vs. FMF algorithm performance.
MethodTargetDetection RateFalse-Positive Rate
FMF algorithm [19]All PE70.60%36.60%
FMF algorithm [19]Preterm PE (<34 wks)81.25%36.60%
LR (this study)All PE~50% at similar FPR~25%
LR (this study)Preterm PE~70–75% at 10% FPRAUC 0.853
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Derimanov, V.; Stoilov, B.; Shopov, M. Impact of SMOTE Oversampling on Machine Learning Classifiers for Preeclampsia Prediction Under Severe Class Imbalance: Evidence from a Bulgarian Screening Cohort. Eng. Proc. 2026, 150, 84. https://doi.org/10.3390/engproc2026150084

AMA Style

Derimanov V, Stoilov B, Shopov M. Impact of SMOTE Oversampling on Machine Learning Classifiers for Preeclampsia Prediction Under Severe Class Imbalance: Evidence from a Bulgarian Screening Cohort. Engineering Proceedings. 2026; 150(1):84. https://doi.org/10.3390/engproc2026150084

Chicago/Turabian Style

Derimanov, Vasil, Boris Stoilov, and Mitko Shopov. 2026. "Impact of SMOTE Oversampling on Machine Learning Classifiers for Preeclampsia Prediction Under Severe Class Imbalance: Evidence from a Bulgarian Screening Cohort" Engineering Proceedings 150, no. 1: 84. https://doi.org/10.3390/engproc2026150084

APA Style

Derimanov, V., Stoilov, B., & Shopov, M. (2026). Impact of SMOTE Oversampling on Machine Learning Classifiers for Preeclampsia Prediction Under Severe Class Imbalance: Evidence from a Bulgarian Screening Cohort. Engineering Proceedings, 150(1), 84. https://doi.org/10.3390/engproc2026150084

Article Metrics

Back to TopTop