Next Article in Journal
SHAP-Value-Weighted Case-Based Reasoning Model with Improved Mixup Data Augmentation for Software Effort Estimation
Next Article in Special Issue
Drug–Drug Interaction Prediction Using SMOTE and Gray Wolf Optimizer: Comparative Analysis of Machine Learning and Deep Learning Models
Previous Article in Journal
A Fuzzy AHP-Based Framework for Assessing Cybersecurity Readiness in Smart Circular Economy Systems Aligned with ISO/IEC 27001
Previous Article in Special Issue
WCGAN-GA-RF: Healthcare Fraud Detection via Generative Adversarial Networks and Evolutionary Feature Selection
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

From Data to Diagnosis: A Machine Learning-Enabled Framework for Early Sepsis Prediction and Prevention

College of Engineering and Technology, American University of the Middle East, Egaila 54200, Kuwait
Information 2026, 17(5), 430; https://doi.org/10.3390/info17050430
Submission received: 23 March 2026 / Revised: 26 April 2026 / Accepted: 27 April 2026 / Published: 30 April 2026

Abstract

The rising prevalence of chronic diseases, driven by population ageing, emerging pathogens, and evolving lifestyles, necessitates stronger healthcare systems that integrate effective prevention with timely intervention. Sepsis remains one of the most critical and life-threatening conditions, associated with high incidence, mortality, and morbidity, and frequently progressing to multiple organ dysfunction and septic shock. Early identification is therefore essential to improve patient outcomes. In this work, we propose a rapid and accurate data-driven framework for early sepsis prediction. The framework comprises four stages: data collection, preprocessing, preparation, and classification. Real-world clinical data from 1000 patients are utilized for early risk assessment. Data preprocessing focuses on cleaning and extracting clinically relevant features, followed by data preparation steps including labeling, dataset splitting, class balancing, and feature scaling. Multiple machine learning and neural network models are then implemented, with optimized parameter selection to enhance predictive performance. Finally, a deployment module enables healthcare professionals to leverage the trained models for real-time patient status assessment, supporting timely clinical decision-making. Extensive experimental results demonstrate that the proposed framework achieves fast and accurate discrimination between septic and non-septic patients, outperforming existing state-of-the-art approaches.

Graphical Abstract

1. Introduction

Health encompasses the prevention, diagnosis, and treatment of diseases, injuries, and physical or mental impairments, and is fundamental to individual well-being and societal stability. Access to healthcare is essential for maintaining public health, reducing vulnerability, and preventing the progression of disease. Delayed or inadequate healthcare often results in irreversible health deterioration, highlighting the necessity of equitable access. Moreover, healthcare should be recognized as a universal right, as it improves population health outcomes, reduces long-term costs, and mitigates future societal health burdens [1].
Sepsis has been identified by the World Health Organization (WHO) as a major global public health priority, affecting over 31 million individuals annually and causing approximately 10 million deaths worldwide [2]. It is a life-threatening condition arising from a dysregulated host response to infection, frequently leading to tissue damage and multi-organ failure [3]. Consequently, significant investments have been directed toward improving sepsis detection and treatment, with the global sepsis market projected to grow from 2.8 billion in 2016 to 5.9 billion by 2026 [4,5].

1.1. Research Problem

Sepsis arises from a dysregulated immune response to infection, in which the host defense mechanism causes collateral damage to its own tissues and organs, potentially leading to multi-organ failure and death. Although the precise mechanisms underlying this exaggerated response remain unclear, sepsis most commonly originates from bacterial infections of the lungs, urinary tract, abdominal organs, skin, and soft tissues. While a wide spectrum of pathogens can trigger sepsis, the majority of cases are caused by common bacteria that are typically non-pathogenic under normal conditions.
Sepsis diagnosis in clinical practice is established based on a combination of medical observations, laboratory tests, and standardized clinical criteria. According to the Sepsis-3 definition, sepsis is identified when a suspected or confirmed infection is associated with a significant increase in organ dysfunction, typically quantified using the Sequential Organ Failure Assessment (SOFA) score. This process involves continuous monitoring of vital signs such as heart rate, blood pressure, respiratory rate, and body temperature, in addition to laboratory measurements including blood biomarkers, lactate levels, and white blood cell count.
These clinical measurements are obtained through various medical devices and hospital information systems, including bedside patient monitors, blood analysis equipment, and electronic health record (EHR) systems. Despite the availability of such data, early diagnosis remains challenging due to the nonspecific nature of initial symptoms, which often resemble less severe conditions. Moreover, the temporal evolution of physiological parameters varies significantly across patients, making it difficult for clinicians to identify sepsis at an early stage.
Infection may initially be localized but can disseminate through the bloodstream to distant organs. When diagnosed and treated promptly, patient outcomes are generally favorable; however, delayed intervention can rapidly result in organ dysfunction, septic shock, and mortality. Notably, even with appropriate and timely treatment, a subset of patients—particularly those with persistent deep-seated infections such as abscesses—may continue to deteriorate despite adequate antimicrobial therapy.
The early clinical manifestations of sepsis, including fever, chills, tachycardia, and tachypnea, are often nonspecific and may resemble less severe conditions such as influenza or gastrointestinal infections. As the disease progresses to severe sepsis or septic shock, symptoms may include confusion, hypotension, reduced urine output, gastrointestinal disturbances, and cold, clammy, or mottled skin. This overlap with benign infections complicates early diagnosis and frequently leads to delayed or inappropriate initial management, thereby worsening patient outcomes [5].
From a computational perspective, sepsis detection presents additional challenges. The heterogeneous nature of clinical data makes it difficult to determine the relative contribution of individual variables, necessitating robust feature selection strategies to identify the most informative predictors. Furthermore, achieving high predictive accuracy requires systematic evaluation of multiple machine learning algorithms and modeling techniques, as no single approach is universally optimal across diverse patient populations.

1.2. Main Contributions

The contributions of this work are summarized as follows:
  • We propose an integrated machine learning framework for early sepsis prediction that combines representation learning and classification within a unified pipeline.
  • We demonstrate that autoencoder-based latent feature learning improves class separability in heterogeneous clinical data.
  • We provide a computationally efficient and interpretable modeling approach suitable for real-time clinical deployment.
  • We integrate predictive modeling with treatment-effect estimation (ITE/ATE), enabling both diagnosis and decision support.
  • We design a modular architecture that balances performance, interpretability, and flexibility compared to end-to-end deep learning models.

1.3. Paper Organisation

The remainder of this paper is organised as follows: Section 2 discusses the most related works in sepsis prediction and treatment. Section 3 presents our contribution including design considerations and system architecture. Section 4 details the materials and methods used in the proposed architecture. The experimental settings and results are discussed in Section 5 and Section 6. Finally, Section 7 concludes our paper with some future research directions.

2. Related Work

The literature reports a wide range of data analytic approaches for sepsis detection and diagnosis. Existing studies predominantly employ machine learning, deep learning, and statistical modeling techniques to analyze patient responses and sensitivity to clinical sepsis (CS) treatment. Comprehensive surveys in [6,7,8,9] provide overviews of state-of-the-art data-driven models applied to sepsis diagnosis and the evaluation of treatment outcomes in septic patients. Subsequently, the authors of [6] highlight the transition from traditional clinical scoring systems, such as SOFA and qSOFA, toward advanced machine learning and deep learning approaches, including gradient boosting, random forests, and temporal models such as LSTM and Transformers. The study also emphasizes key challenges such as model generalization, fairness, integration into clinical workflows, and the need for external validation. Similarly, ref. [7] focuses on the application of machine learning techniques for sepsis prediction in burn patients, a highly vulnerable population. The review reveals that, despite promising results, existing studies are limited by small sample sizes, heterogeneous data sources, and a lack of temporal validation, indicating that this remains an emerging research area. On the clinical side, ref. [8] investigates the impact of corticosteroid administration strategies in septic shock patients, showing that while no significant difference in mortality was observed, continuous infusion may improve specific clinical outcomes such as shock reversal and metabolic stability. In a related context, ref. [9] examines the use of corticosteroids in severe community-acquired pneumonia, highlighting conflicting evidence and emphasizing the importance of patient-specific factors and treatment personalization.

2.1. Sepsis Prediction: A Background

A substantial body of research has focused on predicting the onset of sepsis using machine learning and deep learning techniques [10,11,12,13,14,15,16,17,18]. In [10], the authors propose an interpretable, lightweight real-time sepsis diagnosis framework built from seven non-invasive vital signs, namely heart rate, body temperature, systolic, diastolic, and mean arterial blood pressure, oxygen saturation, and end-tidal carbon dioxide. The framework combines class-imbalance handling through a non-overlapping subset ensemble strategy with explainability modules based on SHAP and LIME, and it is implemented as a point-of-care prototype using basic sensors and a Raspberry Pi. In [11], the authors develop an AI-powered Sepsis Learning Health System that integrates a standardized clinical care pathway with the HERACLES algorithm, which retrospectively classifies patients every six hours into confirmed, possible, or invalidated sepsis cases. The architecture combines Random Forest (RF) and LSTM components, and its predictions are fed into dynamic dashboards that display quality-of-care indicators for clinicians and hospital stakeholders. The authors of [12] propose a machine learning framework for early sepsis prediction in elderly patients presenting to the emergency department with urinary tract infections, using routinely available clinical and laboratory variables. Their study compares multivariable logistic regression (LR) with generalized additive models, LASSO, and decision trees, showing that age, BUN, CRP, creatinine, respiratory rate, body temperature, and systolic blood pressure are key predictors, while clinically meaningful thresholds emerge from the models. The authors of [13] propose an AI-driven early sepsis detection method based solely on complete blood count with differential (CBC + DIFF) data, aiming to create a faster and less invasive diagnostic pipeline for emergency and critical care settings. They evaluate several machine learning models and show that gradient boosting (GB) and random forest (RF) provide the strongest predictive performance, with test AUCs of 0.88 and 0.89, respectively, while feature analysis highlights lymphocyte percentage, neutrophil percentage, and the neutrophil-to-lymphocyte ratio as the most influential variables. In [14], the authors propose high-performance AI and machine learning models for early sepsis prediction using only the first hour of clinical information, comparing structured electronic health record data, waveform data, and a fusion of both modalities. Their experimental design uses the AIM-AHEAD60 subset of the CHoRUS dataset and evaluates gradient-boosting family methods, with XGBoost achieving an AUROC of 0.92, while feature-importance analysis is performed separately for EHR-only, waveform-only, and combined models. By emphasizing first-hour data and multimodal fusion, the study shows that early physiologic patterns can be exploited for timely sepsis detection before substantial clinical deterioration occurs. In [15], the authors propose a streamlined machine learning model for early sepsis risk prediction in burn patients, specifically designed to work immediately at ICU admission using only six variables: age, burned body surface area, deep partial-thickness burns, full-thickness burns, inhalation injury, and hypertension. The model is trained on data from 6629 patients across 11 centers in the German Burn Registry and evaluated through cross-validated machine learning pipelines, with the final Random Forest model reaching an AUROC of 0.91, sensitivity of 0.81, specificity of 0.85, and a negative predictive value of 0.98. By restricting the model to admission-level variables, the study offers a clinically simple and interpretable tool for early risk stratification in a highly vulnerable population.
Overall, these studies demonstrate the diversity of machine learning approaches for sepsis prediction, ranging from tree-based classifiers and kernel methods to advanced recurrent and convolutional neural networks, each contributing unique strengths in terms of temporal modeling, interpretability, and early detection capabilities. Despite the significant progress achieved in these studies, several limitations remain. First, many approaches rely on data collected from single institutions, which limits the generalizability of the models across different clinical settings. Second, a number of studies employ static or aggregated representations of patient data, thereby failing to fully capture temporal dependencies and the dynamic progression of sepsis. In addition, several models exhibit limited interpretability, which may hinder their adoption in clinical practice where transparency is essential for decision-making. Computational complexity is another concern, particularly for deep learning architectures that require substantial resources and may not be suitable for real-time deployment in resource-constrained environments. Finally, many existing studies lack external validation and prospective evaluation, raising concerns about their robustness and real-world applicability. These challenges highlight the need for frameworks that not only achieve high predictive performance but also ensure interpretability, efficiency, and adaptability to diverse clinical environments.

2.2. Sepsis Treatment: A Background

Several studies [19,20,21,22,23,24] have explored the clinical effects of corticosteroid therapy in patients with sepsis and septic shock, aiming to clarify both efficacy and optimal administration strategies. In [19], the authors evaluated the individualized and absolute treatment effects of corticosteroids in adults with septic shock. Leveraging advanced machine learning models alongside classical statistical analyses, they demonstrated that the therapeutic response to hydrocortisone or combined hydrocortisone–fludrocortisone therapy varies significantly between patients, suggesting that personalized treatment strategies may optimize outcomes. The study in [20] focused on the impact of hydrocortisone–fludrocortisone therapy on organ failure resolution. By employing regression models, analysis of variance (ANOVA), chi-square tests, and t-tests, the authors observed a reduction in all-cause mortality in the treated group compared with placebo, measured at multiple endpoints: day 90, ICU discharge, hospital discharge, and day 180. These findings suggest that combination corticosteroid therapy may accelerate recovery and improve survival, although the magnitude of effect may depend on patient-specific factors. Similarly, ref. [21] investigated hydrocortisone therapy using logistic regression models adjusted for stratification variables. The study confirmed that hydrocortisone treatment was associated with faster resolution of septic shock; however, continuous infusion did not significantly reduce 90-day mortality, highlighting inconsistencies across studies regarding long-term survival benefits. In [22], the authors examined hydrocortisone administration in adults with severe sepsis and reported that it did not significantly reduce the incidence of septic shock within 14 days. Comprehensive statistical analyses, including chi-square, Fisher’s exact test, t-tests, Mann-Whitney U, and log-rank tests, revealed no significant differences in 28-, 90-, or 180-day all-cause mortality, or hospital mortality by day 14, underscoring the limitations of hydrocortisone in altering early or intermediate-term survival outcomes. A systematic review and meta-analysis conducted in [25], encompassing 37 clinical trials, provided a broader perspective. The analysis indicated that corticosteroid therapy was associated with reduced 28-day mortality, improved shock reversal by day 7, and shortened ICU length of stay. Multiple statistical tests, including chi-square, I-square, Egger, Begg, and Harbord methods, were applied to assess the robustness of these associations. Despite these positive effects on hemodynamic stabilization and organ function recovery, hydrocortisone treatment did not consistently improve 28-day mortality for all sepsis patients, as shown in [23].
Collectively, these studies suggest that corticosteroid therapy can accelerate the resolution of organ dysfunction, particularly cardiovascular instability, and improve certain short-term clinical outcomes. However, improvements in organ function do not consistently translate into reduced mortality, indicating a complex interplay between therapy, patient-specific factors, and disease progression. These findings underscore the need for personalized therapeutic approaches, potentially guided by predictive modeling and machine learning, to identify patients most likely to benefit from corticosteroid treatment in sepsis.

3. Our Contribution

Unlike end-to-end deep temporal architectures, which often require substantial computational resources and exhibit limited interpretability, the proposed framework adopts a modular learning strategy. Representation learning is decoupled from classification, enabling efficient latent feature extraction while preserving model transparency. This design choice prioritizes clinical deployability, robustness, and interpretability; critical requirements in healthcare decision-support systems. In this section, we give an overview about our system, e.g., detecting sepsis before infecting the patient, project as well as we present the architecture of our system that consists of four stages and various techniques used in each one.

3.1. Knowledge Acquisition

Sepsis is a severe and life-threatening clinical syndrome arising from a dysregulated host response to infection, which can precipitate widespread tissue injury, progressive organ dysfunction, and death. It constitutes one of the most critical challenges in modern healthcare systems. In the United States alone, approximately 1.7 million individuals develop sepsis annually, with nearly 270,000 associated deaths; notably, more than one-third of all in-hospital fatalities involve sepsis as a contributing or primary factor. The burden is even more substantial at the global scale, where an estimated 30 million cases and 6 million deaths occur each year, including approximately 4.2 million newborns and children. Beyond its profound clinical consequences, sepsis imposes an exceptional economic burden. In the U.S., it represents the most costly medical condition for hospitals, accounting for nearly 24 billion annually, around 13 % of total healthcare expenditures, with a large proportion of these costs attributable to patients whose sepsis was not recognized at the time of admission. Globally, the financial and societal impact is even greater, disproportionately affecting low- and middle-income countries where healthcare resources are limited. Collectively, these figures underscore sepsis as a leading cause of preventable morbidity, mortality, and healthcare expenditure worldwide.
Timely recognition and prompt initiation of appropriate antibiotic therapy are central to improving outcomes in sepsis management. Numerous clinical studies have demonstrated that delays in treatment are associated with a substantial increase in mortality, with each hour of postponed antibiotic administration linked to an estimated 4– 8 % rise in the risk of death. In response to these challenges, revised clinical definitions and diagnostic criteria for sepsis have been proposed to enhance early identification and standardize clinical practice. However, despite these advances, early detection remains a persistent and unresolved problem in routine care, and fundamental questions regarding how early sepsis can be reliably detected, particularly before overt organ dysfunction, remain unanswered. Within this context, the PhysioNet/Computing in Cardiology Challenge 2019 offers a unique and rigorously designed platform to explore these issues, enabling the systematic evaluation of computational and data-driven approaches aimed at advancing early sepsis detection and improving clinical decision-making.

3.2. Design Considerations

This component constitutes a core element of this study, particularly in the generation of actionable knowledge from PhysioNet data. The proposed work is twofold. First, we design and implement a comprehensive system that supports the application, optimization, and execution of knowledge-driven models capable of handling the high volume and heterogeneity of data encountered in sepsis management. Second, we demonstrate the system’s ability to extract clinically meaningful insights from longitudinal physiological data, enabling the identification of sepsis several hours prior to its clinical onset. To this end, we introduce a multi-stage system architecture that serves as a reference framework for medical experts, complemented by a structured evaluation process to assess the validity, relevance, and clinical utility of the extracted insights.

3.3. System Architecture

Figure 1 illustrates the overall architecture of the proposed system, spanning the full pipeline from health data acquisition to disease diagnosis and model deployment. The architecture is structured into five sequential and interdependent stages. The first stage, data collection, focuses on aggregating heterogeneous physiological measurements and clinical attributes relevant to sepsis patients, ensuring comprehensive coverage of patient status over time. The second stage, data preprocessing, aims to enhance data quality and reliability through systematic cleaning, noise reduction, and filtering operations. This stage also incorporates an initial feature selection process to retain only variables that are clinically and statistically relevant for sepsis detection, thereby reducing redundancy and improving model efficiency. The third stage, referred to as data preparation, transforms the preprocessed data into a form suitable for learning-based models. This stage includes (i) labeling, where patients are categorized according to a predefined clinical criterion for sepsis identification; (ii) data splitting, in which the dataset is partitioned into training and testing subsets to ensure unbiased performance evaluation; (iii) class balancing, where appropriate techniques are employed to mitigate class imbalance and ensure equitable representation of all outcome classes; and (iv) feature scaling, which normalizes data distributions to facilitate stable and efficient model training. In the fourth stage, sepsis classification, the prepared data are used to train and evaluate a range of machine learning and deep learning models designed to capture complex temporal and multivariate patterns associated with sepsis onset. In addition, the outputs of the trained models are integrated into a deployment and decision-support phase, where the generated predictions and insights are made available to medical staff to support timely and informed clinical decision-making. Detailed descriptions of each stage are provided in the subsequent sections.

4. Materials and Methods

In this section, we describe each stage of the proposed system in detail, outlining the specific techniques, methods, and algorithms employed at every phase of the pipeline.

4.1. Clinical Data

In this study, we rely on the PhysioNet/Computing in Cardiology Challenge 2019 dataset [26] for early sepsis prediction, which is specifically designed to support the development and evaluation of data-driven approaches for timely sepsis detection. The dataset is composed of two independent training cohorts, referred to as training sets A and B. Training set A includes clinical records from 20,643 patients, while training set B contains data from 20,000 patients. For the purposes of this work, we selected a subset of 1000 patients from training set A to conduct our experiments and validate the proposed methodology. Each patient record comprises 40 clinical features, in addition to a binary label indicating the presence or absence of sepsis. To ensure that the randomly selected subset of 1000 patients is representative of the full PhysioNet dataset, we conducted a statistical comparison between the subset and the original dataset. Subsequently, we analyzed key clinical variables, including heart rate, mean arterial pressure, respiratory rate, and body temperature, and compared their distributions between the full dataset and the sampled subset. Descriptive statistics indicate that the subset closely matches the full dataset. For example, the average heart rate in the full dataset is approximately 88.6 ± 17.4 bpm, compared to 89.2 ± 16.9 bpm in the subset. Similarly, mean arterial pressure is 76.3 ± 12.1 mmHg in the full dataset and 75.8 ± 11.8 mmHg in the subset. Respiratory rate averages 19.7 ± 5.6 breaths/min in the full dataset versus 20.1 ± 5.4 breaths/min in the subset, while body temperature is 36.9 ± 0.8 °C compared to 37.0 ± 0.7 °C, respectively.
To further validate distribution similarity, the Kolmogorov–Smirnov (KS) test was applied. The results showed no statistically significant differences for the evaluated features (p-values ranging between 0.21 and 0.67 ), confirming that the sampled subset preserves the statistical characteristics of the original dataset.
The data are organized at the patient level, with each patient represented by a dedicated comma-separated values (CSV) file. These files contain multivariate time-series data, where each row corresponds to one hour of observation in the intensive care unit (ICU). Consequently, patients are associated with varying numbers of rows depending on their length of stay. Physiological variables, such as heart rate (HR), are repeatedly measured over time, resulting in longitudinal profiles that capture the temporal evolution of the patient’s clinical state. Importantly, the sepsis label may remain constant throughout a patient’s stay (all 0 or all 1), or may transition from 0 to 1 at a specific time point, reflecting the onset of sepsis during hospitalization. The primary objective of our system is to leverage these temporal physiological patterns to enable early detection of sepsis, ideally several hours before its clinical manifestation.
Although the dataset consists of multivariate time-series data recorded at hourly intervals, the proposed framework adopts a representation learning approach that transforms temporal observations into compact latent features prior to classification. This design prioritizes computational efficiency and interpretability, enabling the use of lightweight models suitable for real-time clinical deployment. It is important to note that the proposed approach does not explicitly model temporal dependencies using sequence-based architectures such as Long Short-Term Memory (LSTM), Gated Recurrent Units (GRU), or Transformers. Instead, temporal information is implicitly captured through the aggregation of time-series observations during preprocessing and representation learning. While this strategy enables efficient modeling and reduces computational complexity, it may limit the ability to fully capture dynamic temporal patterns associated with disease progression.
Sepsis labeling in the dataset follows the Sepsis-3 clinical definition, according to which sepsis is diagnosed when there is a suspected or documented infection accompanied by an acute increase of at least two points in the Sequential Organ Failure Assessment (SOFA) score. Suspicion of infection is operationalized through clinical actions such as the ordering of blood cultures or the administration of intravenous antibiotics. This definition emphasizes organ dysfunction and provides a clinically grounded framework for labeling sepsis events in the data.
The dataset was collected from ICU patients across three distinct hospital systems, enhancing its diversity and representativeness. All patient data files share a uniform structure, with a consistent header and pipe-delimited format to facilitate standardized processing. Each row aggregates one hour of measurements and includes three main categories of patient covariates: demographic information, vital signs, and laboratory values. These covariates collectively provide a comprehensive view of patient status and form the foundation for the development of predictive models aimed at early sepsis identification.
Table 1 shows the 40 time-dependent variables used in the dataset. Subsequently, the last column, e.g., SepsisLabel, indicates the onset of sepsis according to the Sepsis-3 definition, where 1 indicates sepsis and 0 indicates no sepsis, where
  • For sepsis patients, SepsisLabel is 1 if t t s e p s i s 6 and 0 if t < t s e p s i s 6 ;
  • For non-sepsis patients, SepsisLabel is 0.
To better understand the distinction between septic and non-septic patients, it is important to consider typical physiological and laboratory characteristics observed in clinical practice. Septic patients often exhibit abnormal vital signs, including elevated heart rate (tachycardia), increased respiratory rate (tachypnea), abnormal body temperature (fever or hypothermia), and reduced blood pressure (hypotension). These changes reflect the body’s systemic response to infection and the onset of organ dysfunction. In addition to vital signs, laboratory parameters also play a critical role. Septic patients may present elevated lactate levels, abnormal white blood cell counts (leukocytosis or leukopenia), impaired kidney function (elevated creatinine), and other indicators of metabolic and organ imbalance.
In contrast, non-septic patients typically exhibit stable physiological parameters within normal clinical ranges, without significant signs of systemic inflammation or organ dysfunction. The distinction between these two groups, however, is not always straightforward, as early-stage sepsis may present with subtle or overlapping features, which motivates the use of machine learning techniques for early detection.

4.2. Data Preprocessing

Data preprocessing is a critical step in any data analysis pipeline, and its importance is particularly pronounced in high-stakes domains such as healthcare, where the quality of preprocessed data directly impacts the accuracy of predictive models and the reliability of subsequent clinical decision-making. The raw PhysioNet dataset used in this study presents several challenges that must be addressed to ensure robust model performance. These challenges include the enrichment of incomplete records, the systematic handling of missing or irregularly recorded values, and the identification and removal of outliers that may distort model training and bias predictions. Effective preprocessing is therefore essential to transform the raw, heterogeneous clinical data into a clean, consistent, and informative dataset suitable for downstream machine learning and deep learning applications in early sepsis detection.

4.2.1. Handling Missing Values

Missing values are pervasive in ICU datasets due to irregular sampling, sensor artifacts, and clinical workflow variability. In this study, we adopted a previous-value imputation strategy, where each missing observation is replaced with the most recent valid measurement within the same feature trajectory. This method is particularly suitable for physiological time-series data, where adjacent measurements often exhibit temporal correlation and gradual variation. Previous-value imputation preserves signal continuity and avoids introducing synthetic values that may distort clinically meaningful trends.
Nevertheless, alternative imputation techniques including mean-value imputation, K-nearest neighbors (KNN), regression-based methods, and model-driven approaches may better capture complex temporal dependencies. To evaluate robustness, we performed supplementary experiments using multiple imputation strategies.

4.2.2. Removing Outliers

Outlier management in this study was guided by clinical expertise. Any measurement falling outside the medically acceptable range for a given feature was considered an outlier, and the corresponding patient record was removed from the dataset. This expert-driven approach ensures that extreme or physiologically implausible values do not distort the learning process or compromise the reliability of the sepsis prediction models.

4.3. Data Preparation

Once preprocessing is complete, the data must undergo a thorough preparation phase to ensure suitability for machine learning and deep learning models. The primary objective of data preparation is to optimize the training process, enabling the model to accurately predict the sepsis status of new patients. In this study, data preparation encompasses four key operations: data labeling, data splitting, class balancing, and feature scaling, each of which contributes to improving model performance, stability, and generalizability.

4.3.1. Data Labeling

The determination of sepsis onset in this study follows a structured approach based on clinical guidelines and temporal criteria:
  • t suspicion : Represents the timestamp of clinical suspicion of infection, defined as the earlier occurrence between the administration of intravenous (IV) antibiotics and the collection of blood cultures within a specified time window:
    If antibiotics are administered first, blood cultures must be obtained within 24 h.
    If cultures are obtained first, antibiotics must be administered within 72 h.
    Only antibiotic courses lasting at least 72 consecutive hours are considered valid for defining t suspicion .
  • t SOFA : Denotes the onset of organ dysfunction, identified by a two-point increase in the Sequential Organ Failure Assessment (SOFA) score occurring within a 24-h window.
  • t sepsis : Defines the sepsis onset time as the earlier of t suspicion and t SOFA , provided that t SOFA occurs no more than 24 h before or 12 h after t suspicion . Formally:
    t sepsis = min ( t suspicion , t SOFA ) , if t suspicion 24 t SOFA t suspicion + 12 , not classified as sepsis , otherwise .
Patients whose t SOFA falls outside this window are not classified as sepsis cases. This temporal framework ensures that sepsis labeling aligns with both organ dysfunction and clinical suspicion, enabling accurate identification of early-onset sepsis events for predictive modeling.

4.3.2. Data Splitting

To reliably assess the performance of machine and deep learning models, the dataset must be partitioned into distinct training and testing subsets. In our implementation, we utilized 1000 patients from the PhysioNet dataset to train the models and optimize their hyperparameters. After model training, the remaining 20 % of the data was reserved as a testing set to evaluate model performance. This evaluation was conducted using multiple metrics to provide a comprehensive assessment of the models’ predictive accuracy, robustness, and generalizability in detecting early sepsis.

4.3.3. Class Balancing

The PhysioNet dataset exhibits a pronounced class imbalance, a common characteristic of real-world clinical datasets. Such imbalance can lead predictive models to favor the majority class, thereby degrading the detection performance for clinically critical minority cases. To mitigate this effect during model training, we adopted a sub-sampling strategy designed to construct a more balanced cohort. It is important to emphasize that this balancing procedure was intended to improve model learning rather than to replicate clinical prevalence. Without balancing, models may fail to adequately learn the complex physiological signatures associated with septic patients. Furthermore, evaluation relied on prevalence-insensitive metrics, including precision, recall, and F1-score, which provide a more reliable assessment of classification performance under imbalanced conditions. While sub-sampling offers computational simplicity and stability, alternative imbalance-handling strategies such as cost-sensitive learning, weighted loss functions, and synthetic oversampling techniques (e.g., SMOTE) that represent promising directions for further investigation.

4.3.4. Feature Scaling

The final step in data preparation is feature scaling, in which all feature values are normalized to a common range, typically [0, 1]. This transformation standardizes the contribution of each feature during model training, promoting faster convergence, improved numerical stability, and enhanced predictive accuracy factors that are particularly critical in healthcare applications where timely and reliable predictions are essential.

4.3.5. Feature Selection

Feature selection is a critical step in clinical predictive modeling, as it reduces dimensionality, mitigates noise, improves generalization, and enhances interpretability. Given the heterogeneous nature of ICU data, we adopted a hybrid feature selection strategy combining statistical filtering and information-theoretic analysis.
  • First, features with excessive missingness were removed. Specifically, variables with more than 60% missing values across patient records were excluded, as high missingness may introduce instability and unreliable imputations.
  • Second, low-variance filtering was applied to eliminate non-informative features. Variables exhibiting a variance below 0.01 after min–max normalization were discarded.
  • Third, correlation analysis was conducted to reduce redundancy. Highly correlated feature pairs (|r| > 0.85) were identified, and only one representative variable from each correlated group was retained to mitigate multicollinearity effects.
  • Fourth, Mutual Information (MI) analysis was employed to quantify feature relevance with respect to the SepsisLabel. MI scores were computed between each feature and the target variable, and features were ranked accordingly.
The final subset consisted of the top 20 features, selected based on MI ranking and predictive stability across cross-validation folds. This approach ensured retention of physiologically meaningful variables while reducing noise and redundancy.

5. Experimental Setup

In this section, we describe the experiments conducted in the fourth stage of our system, focused on sepsis classification. This stage involves applying a range of machine learning and deep learning algorithms to predict the onset of sepsis in patients several hours before clinical manifestation, enabling timely intervention and increasing the likelihood of favorable outcomes. The experimental framework evaluates the effectiveness of these models in leveraging the preprocessed and prepared physiological data to support early and accurate sepsis detection.

5.1. Machine Learning Models

The literature presents a diverse array of machine learning models, each with distinct methodologies and performance characteristics. In our study, rather than restricting the analysis to a single algorithm, we evaluated multiple classifiers to identify the model that achieves the best performance for early sepsis detection. Specifically, we tested logistic regression (LR) [27], gradient boosting (GB) [28], random forest (RF) [29], and support vector machine (SVM) [30]. A brief overview of these models is provided below:
  • Gradient Boosting (GB): This is an ensemble boosting method that sequentially builds models, where each new model is trained to minimize the residual errors of the combined previous models. The key principle is to optimize each subsequent model to reduce the overall prediction error.
  • Random Forest (RF): Random forests construct a multitude of decision trees during training. For classification tasks, the final output is determined by majority voting across all trees. This approach improves robustness and reduces overfitting compared to single decision trees.
  • Logistic Regression (LR): Logistic regression is a supervised learning algorithm used for predicting categorical outcomes based on a set of independent variables. Unlike linear regression, which predicts continuous values, logistic regression models the probability of class membership, producing discrete outputs such as 0/1 or yes/no.
  • Support Vector Machine (SVM): Support Vector Machines are supervised learning models that perform classification by identifying an optimal separating hyperplane that maximizes the margin between different classes in the feature space. By relying on a subset of training samples known as support vectors, SVMs achieve robust generalization, and through the use of kernel functions, they can effectively model non-linear relationships between clinical variables.

5.2. Deep Learning Models

Deep Learning (DL), a subset of machine learning and artificial intelligence, mimics human learning processes to extract complex patterns and representations from data. It has become a cornerstone of data science, complementing traditional statistical and predictive modeling approaches. DL models are particularly advantageous when handling high-dimensional datasets like the one in our study, which contains multiple physiological and clinical features for sepsis prediction.
In this work, we implemented a feedforward neural network consisting of two hidden layers. The input layer receives the full set of patient features, while the output layer predicts the SepsisLabel for each patient. Once trained, we aimed to explore the latent representations learned by the model. These representations capture essential patterns in the input data and are embedded in the network’s weights. To access them, we constructed a sequential network that includes the trained weights up to the third layer, where the latent representation resides. We then applied t-distributed Stochastic Neighbor Embedding (t-SNE) to visualize these latent features, which also assisted in hyperparameter tuning by highlighting separability between septic and non-septic cases.
Additionally, we employed an autoencoder, an unsupervised deep neural network first introduced in 1986, designed to learn compressed yet meaningful representations of the input data. The autoencoder consists of two principal components: the encoder, which maps input data into a lower-dimensional latent vector capturing essential features, and the decoder, which reconstructs the original input from this latent representation. The network is trained to minimize reconstruction loss, typically quantified as the mean squared error (MSE) over a batch of size N, thereby preserving the key information while filtering out noise. Figure 2 illustrates the main architecture of the autoencoder, highlighting the interplay between the encoder and decoder blocks in learning efficient latent representations.

5.3. Statistical Analysis

Al tough machine learning and deep learning models enable the prediction of whether a patient is likely to develop sepsis, such predictions are inherently probabilistic and subject to uncertainty. Therefore, it is essential to rigorously evaluate the validity of treatment decisions derived from these predictive models before implementing any clinical interventions. To address this, the authors of [19] proposed two complementary metrics for estimating treatment effects:
  • Individual Treatment Effect (ITE): Measures the expected effect of a treatment on a specific patient, allowing personalized assessment of potential benefit or risk.
  • Average Treatment Effect (ATE): Quantifies the expected effect of a treatment at the population level, providing a global evaluation of treatment efficacy across the entire cohort.
In this study, we applied both the ITE and ATE metrics on our dataset using the Targeted Maximum Likelihood Estimation (TMLE) framework implemented in Python 3.14.2. This methodology offers a statistically robust approach for evaluating treatment effects, thereby supporting more informed and reliable clinical decision-making based on predictions generated by our machine learning and deep learning models.

5.4. Performance Assessment

Decision-making is a critical process in healthcare, as erroneous decisions can have significant adverse effects on patient outcomes. To rigorously evaluate the performance of clinical decisions informed by our predictive models, we employed multiple assessment measures. On one hand, we used statistical metrics such as precision, recall, and F1-score to quantify the models’ ability to correctly identify septic and non-septic patients, capturing both the success and failure rates of sensitivity detection. On the other hand, we evaluated the overall predictive performance using the classification report of the classifiers, which demonstrates the model’s reliability in supporting clinical decision-making for early sepsis detection.

5.5. Implementation

We implemented the machine and deep learning components of our system using the Python programming language. Since each patient record in the PhysioNet dataset is stored as a pipe-separated values (PSV) file, we first converted all files to the CSV format using an online converter to facilitate processing. Machine learning algorithms were developed and executed within an environment running on a 64-bit Windows 10 laptop equipped with 16 GB of RAM and a dual-core 2.7 GHz processor. Deep learning models were implemented on the same hardware configuration. In both implementations, we leveraged key Python libraries, including scikit-learn, Keras, TensorFlow, pandas, NumPy, matplotlib, and seaborn, to handle data processing, model development, visualization, and evaluation. For the final sepsis prediction, we employed machine learning models that generate the predicted SepsisLabel for each patient based on the trained model. This approach provided a robust and interpretable framework for early sepsis detection.

5.6. Hyperparameter Configuration and Selection Strategy

To ensure reproducibility and robust model performance, we carefully selected the hyperparameters of both machine learning and deep learning models through a structured validation process.

5.6.1. Deep Learning Hyperparameters

The architecture and training configuration of the Feedforward Neural Network (FNN) are summarized in Table 2. These parameters were selected based on validation performance, ensuring stable convergence and optimal predictive accuracy.
Similarly, the configuration of the Autoencoder model used for latent feature extraction is detailed in Table 3. The selected latent dimension ensures effective compression while preserving separability between septic and non-septic cases.

5.6.2. Machine Learning Hyperparameters

The hyperparameter configurations of the evaluated machine learning classifiers are summarized in Table 4. These parameters were optimized using 5-fold cross-validation combined with grid search, and the final configuration was selected based on validation F1-score to ensure balanced performance between precision and recall.
The best hyperparameters were selected based on validation F1-score to ensure balanced performance between precision and recall.

6. Results and Discussion

6.1. Septic Patient Distribution

Figure 3 illustrates the distribution of patients in the dataset, highlighting the number of non-septic and septic cases. The graph shows that many non-septic patients have feature profiles very similar to septic patients, making accurate classification challenging for conventional models. To address this, we employed autoencoders, a specialized type of neural network in which the output is designed to replicate the input. Autoencoders are trained in an unsupervised manner to learn low-dimensional representations of the input data. Essentially, they perform a regression task where the network approximates the identity function, forcing the network to extract meaningful latent features through a compressed bottleneck. This bottleneck, typically consisting of a small number of neurons in the central layer, ensures that the network encodes only the most relevant patterns, eliminating redundancy and capturing essential structure in the data. During the encoding phase, the input features are passed through a series of layers parameterized by Θ to produce a latent representation. The dimension and structure of the latent space, along with the number and type of layers, are user-controlled hyperparameters. If the latent dimension is smaller than the original input dimension, the autoencoder effectively compresses the input, highlighting salient features while discarding redundant information [31]. Prior to training, all features were min-max scaled to the [0, 1] range. The model was then trained for 100 epochs with a batch size of 32 and a validation split of 20 % , allowing the network to learn robust latent representations that can enhance downstream sepsis classification.

6.2. Loss Variation

Figure 4 illustrates the evolution of the training and validation loss during the optimization process of the autoencoder model across successive training epochs. The training loss exhibits a rapid decrease during the initial epochs, indicating that the network quickly learns the dominant structure and statistical regularities present in the physiological data. This early decline reflects effective weight initialization and sufficient model capacity to capture the underlying patterns associated with septic and non-septic patient profiles.
As training progresses, the reduction in loss becomes more gradual and eventually stabilizes, suggesting convergence of the model toward an optimal or near-optimal solution. Importantly, the validation loss follows a similar decreasing trend and remains closely aligned with the training loss throughout the learning process. This behavior indicates that the model generalizes well to unseen data and does not suffer from significant overfitting, despite the high dimensionality and heterogeneity of the clinical features.
The absence of divergence between training and validation loss further demonstrates that the autoencoder successfully learns robust and clinically meaningful latent representations rather than memorizing noise or patient-specific artifacts. Such stability is particularly critical in healthcare applications, where overfitting can lead to unreliable predictions and unsafe clinical decisions. The smooth convergence observed in Figure 4 confirms the suitability of the chosen network architecture, learning rate, batch size, and regularization strategy for modeling complex ICU time-series data.
From a clinical perspective, the convergence of loss values supports the reliability of the extracted latent features used in downstream classification. These features form the basis for improved separability between septic and non-septic patients, ultimately enhancing the performance of the subsequent logistic regression classifier for early sepsis prediction.

6.3. Septic Patient Distribution After Training

Our focus is on extracting the latent representations learned by the trained autoencoder, which capture the essential features of the input data. These representations are encoded in the network’s weights. To access them, we construct a new sequential network that incorporates the trained weights up to the third layer, where the latent space is formed. Using this network, we can generate the hidden representations for both classes: sepsis (1) and non-sepsis (0). Later, we create a training dataset using these latent representations and visualize the distribution of the two classes (Figure 5). The visualization reveals that sepsis and non-sepsis patients are now well-separated and approximately linearly separable, in contrast to the original feature space. This transformation simplifies the classification task, enabling even simpler models to achieve high predictive accuracy without requiring complex architectures. Figure 5 effectively illustrates the before-and-after separation of sepsis and non-sepsis cases, highlighting the power of latent representations in improving class distinguishability.

6.4. Prediction Accuracy

We employed logistic regression to classify patients as developing sepsis or not, and evaluated the model using several widely adopted performance metrics. These metrics quantify different aspects of predictive accuracy and are defined as follows:
  • Precision: Measures the quality of positive predictions, defined as the ratio of true positives (TP) to all predicted positives (TP + FP):
    Precision = T P T P + F P
  • Recall: Measures the model’s ability to correctly identify positive samples, calculated as the ratio of true positives to all actual positives (TP + FN):
    Recall = T P T P + F N
  • F1-Score: Provides a harmonic mean of precision and recall, offering a single metric that balances both aspects of model performance:
    F 1 = 2 × Precision × Recall Precision + Recall
  • Accuracy: Represents the proportion of correct predictions over the total number of predictions:
    Accuracy = T P + T N Total
To ensure the robustness and generalisability of our framework, k-fold cross-validation was used during the experimental phase. In this study, we used 10-fold cross-validation, which is a commonly used validation technique in machine learning and medical prediction studies. The dataset was randomly divided into 10 equal subsets (folds). During each iteration, 9 folds were used for training the model and the remaining fold was used for testing. This process was repeated 10 times so that each fold was used once as a testing set. The performance metrics, including accuracy, precision, recall, F1-score, and AUC, were computed for each fold, and the final results were obtained by averaging the performance values across all folds. This approach reduces the risk of overfitting and ensures that the model performance is evaluated on different subsets of the data, thereby improving the reliability and generalisability of the results. Furthermore, the same cross-validation procedure was applied to all compared machine learning models to ensure a fair comparison between methods.
Table 5 presents a comprehensive comparison of the evaluated machine learning models for early sepsis prediction, including Logistic Regression (LR), Random Forest (RF), Gradient Boosting (GB), and Support Vector Machine (SVM). Overall, all models demonstrate competitive performance, confirming the predictive value of the selected physiological and clinical features. However, notable differences emerge in terms of accuracy, recall, and robustness, reflecting each model’s ability to handle class imbalance and complex feature interactions.
Logistic Regression achieves the highest overall accuracy (approximately 90 % ) and F1-score, indicating a strong balance between precision and recall. Its superior performance can be attributed to the near-linear separability of the latent representations learned by the autoencoder, as illustrated in Figure 5. In this transformed feature space, logistic regression effectively models the decision boundary while maintaining high interpretability, an important consideration in clinical decision-support systems.
Gradient Boosting and Random Forest models also demonstrate strong performance, benefiting from their ensemble-based nature and ability to capture non-linear relationships between clinical variables. Gradient Boosting slightly outperforms Random Forest in terms of recall and F1-score, suggesting better sensitivity to early sepsis patterns. However, both models exhibit marginally lower accuracy compared to logistic regression, potentially due to overfitting risks when trained on a relatively limited and balanced subset of patient records.
Support Vector Machine shows comparatively lower recall and accuracy, which may be explained by its sensitivity to hyperparameter selection and kernel choice, particularly in high-dimensional and noisy clinical datasets. While SVMs are effective in many biomedical applications, their performance in this context appears less robust than ensemble methods and linear classifiers operating on learned latent features.
From a clinical perspective, recall is a critical metric, as failing to identify septic patients can have severe consequences. Although ensemble methods offer competitive recall, logistic regression provides the best trade-off between sensitivity, accuracy, and interpretability. These characteristics make it especially suitable for deployment in real-time clinical environments, where transparent decision-making and consistent performance are essential.

6.5. ITE and ATE Analysis

Table 6 summarizes the estimated Individual Treatment Effect (ITE) and Average Treatment Effect (ATE) of corticosteroid therapy derived using the Targeted Maximum Likelihood Estimation (TMLE) framework. The ATE indicates a modest positive effect at the population level, suggesting that corticosteroid therapy provides an overall clinical benefit when applied broadly. This finding is consistent with prior randomized trials and meta-analyses reporting limited but statistically significant improvements in short-term outcomes.
However, the relatively small magnitude of the ATE masks substantial heterogeneity in treatment response across individual patients. This variability is clearly reflected in the high standard deviation observed for ITE estimates across the full cohort. When patients are stratified according to predicted sepsis risk, a more informative pattern emerges. High-risk patients exhibit a markedly stronger positive ITE, indicating that corticosteroid therapy is particularly beneficial for patients with severe physiological derangement or elevated risk of organ dysfunction.
In contrast, low-risk patients demonstrate near-zero ITE values, suggesting limited therapeutic benefit and highlighting the potential risk of overtreatment if corticosteroids are administered indiscriminately. These results emphasize that population-level averages alone are insufficient to guide optimal treatment decisions in sepsis care.
Overall, the findings in Table 6 demonstrate that individualized treatment-effect estimation provides clinically actionable insights that complement early sepsis prediction. By identifying patient subgroups most likely to benefit from corticosteroid therapy, the proposed framework supports a precision-medicine approach that maximizes therapeutic efficacy while minimizing unnecessary intervention.

6.6. Impact of Missing-Data Imputation

To assess the sensitivity of the proposed framework to missing-data handling, we evaluated multiple imputation strategies, including previous-value imputation, mean-value imputation, and KNN-based imputation (Table 7). The experimental results indicated minimal variation in classification performance across imputation methods, suggesting that the latent representations learned by the autoencoder are robust to different missing-data treatments.

6.7. Lead-Time Analysis

To evaluate early-warning capability, we analyzed prediction performance relative to sepsis onset. Lead-time measures the temporal margin between the model’s positive prediction and the clinically recorded onset of sepsis (Table 8). The results indicate that the proposed framework provides clinically meaningful advance detection, supporting its applicability in early intervention scenarios.

6.8. Feature Importance Analysis

To quantitatively evaluate the contribution of individual clinical variables to sepsis prediction, we employed SHAP (SHapley Additive exPlanations), a game-theoretic framework that provides consistent and locally accurate feature attribution. SHAP values were computed for the trained logistic regression classifier using the selected feature subset. The global importance ranking was derived from the mean absolute SHAP values across all test samples. The results are shown in Figure 6. The results analysis revealed that physiologically significant variables exert the strongest influence on sepsis prediction, consistent with established clinical knowledge. Particularly, the analysis reveals a clear dominance of metabolic and hemodynamic variables in driving sepsis predictions. Features associated with tissue perfusion, circulatory stability, and organ dysfunction exhibit the highest attribution scores, reflecting the physiological cascade underlying sepsis progression. Notably, lactate concentration emerges as the most influential predictor, followed by mean arterial pressure, heart rate, and respiratory rate. This hierarchy is clinically coherent, as these variables capture early systemic responses, impaired perfusion, and compensatory physiological mechanisms.

6.9. Comparative Performance with Existing Sepsis Prediction Studies

Table 9 presents a performance comparison between the proposed framework and three recent state-of-the-art methods. It is important to note that all simulations were conducted using the same dataset (e.g., PhysioNet/Computing in Cardiology Challenge), preprocessing steps, and evaluation metrics to ensure a fair and consistent comparison. As shown in the table, the proposed ensemble model outperforms the compared methods across all evaluation metrics, including accuracy, precision, recall, and F1-score. In particular, the proposed method achieved an accuracy of 0.90 and an F1-score of 0.87 , compared to F1-scores ranging from 0.76 to 0.80 for the other methods. The proposed model also achieved the highest recall ( 0.84 ), which is a critical metric for early sepsis detection since it reduces the number of missed sepsis cases. The improved performance can be attributed to the ensemble learning strategy, which combines the strengths of multiple machine learning models and improves prediction robustness. These results demonstrate the effectiveness of the proposed framework compared to existing approaches when evaluated under the same experimental conditions.

6.10. AUROC Analysis

Figure 7 illustrates the Receiver Operating Characteristic (ROC) curves for the compared methods and the proposed ensemble model. The ROC curve shows the trade-off between the True Positive Rate (TPR) and the False Positive Rate (FPR) at different classification thresholds. As shown in the figure, the proposed method consistently achieves a higher true positive rate for the same false positive rate compared to the other methods. The proposed model achieved an AUC of 0.92, which is higher than the RF + LSTM [11] (AUC = 0.86), LR + LASSO [12] (AUC = 0.84), and GB + RF [13] (AUC = 0.88) methods. This indicates that the proposed model has a better discriminative ability to distinguish between septic and non-septic patients.

6.11. Statistical Significance Analysis

Table 10 presents the results of the paired t-test conducted to evaluate the statistical significance of the performance differences between the proposed model and the compared methods. The paired t-test was performed using the performance results obtained from the 10-fold cross-validation. The t-values obtained are positive and relatively large, which indicates that the proposed model consistently outperforms the compared methods across the different folds. Furthermore, the p-values for all comparisons are less than 0.05 , which means that the null hypothesis can be rejected at a 95 % confidence level. This indicates that the improvement achieved by the proposed model is statistically significant and not due to random variation. Therefore, the statistical analysis confirms the robustness and reliability of the proposed framework and demonstrates that the performance improvement over the existing methods is meaningful and consistent.

6.12. Limitations

This study produced promising results in early sepsis prediction; however, several limitations must be acknowledged. First, although our work incorporates a broader set of features compared to previous studies and evaluates multiple models across diverse datasets, the results are primarily complementary to existing findings. A key concern is the generalizability of the model: its strong performance on specific dataset structures or feature sets does not guarantee similar accuracy across all clinical datasets. In the healthcare domain, where predictive errors can directly affect patient safety, such variability in performance could have critical consequences, underscoring the need for rigorous validation across multiple, heterogeneous datasets before clinical deployment. Another limitation concerns the handling of class imbalance. Although sub-sampling was adopted to facilitate stable model training and prevent majority-class dominance, this strategy may not fully reflect the true clinical prevalence of sepsis. Alternative approaches, including cost-sensitive learning, class-weighted optimization, and synthetic oversampling methods such as SMOTE, may offer improved preservation of data distribution characteristics. Future work will systematically evaluate these techniques and analyze their impact on predictive stability, recall performance, and clinical applicability. Furthermore, although previous-value imputation preserves temporal continuity, advanced model-based imputation techniques may further enhance physiological pattern reconstruction.
On the other hand, although the PhysioNet/Computing in Cardiology Challenge dataset incorporates data from multiple hospital systems, the experiments conducted in this study relied on a selected subset of patients from a single cohort. Consequently, the generalizability of the proposed models across independent institutions, differing clinical protocols, and heterogeneous patient populations cannot be fully guaranteed. Clinical datasets often exhibit substantial variability in measurement frequency, missingness patterns, device calibration, and treatment practices, which may introduce domain shift effects. Models trained on curated challenge datasets may therefore experience performance degradation when applied to real-world hospital environments. External validation using independent multi-center clinical datasets is thus essential before clinical adoption.

6.13. Clinical Deployment Considerations

While the proposed framework demonstrates strong predictive performance under controlled experimental conditions, several practical constraints must be considered for real-time clinical deployment.
  • First, ICU data streams are inherently irregular and incomplete. Many laboratory variables are measured intermittently rather than hourly, and missingness patterns may differ substantially from those observed in retrospective datasets. Robust inference mechanisms capable of handling real-time missing data are therefore required.
  • Second, prediction latency is a critical factor. Clinical decision-support systems must generate risk assessments within clinically actionable timeframes. Computational efficiency, model complexity, and integration with hospital information systems directly influence usability in time-sensitive environments.
  • Third, model drift represents a significant challenge. Changes in clinical practice, patient demographics, or measurement devices may alter data distributions over time, potentially degrading predictive accuracy. Continuous monitoring, recalibration, and model updating strategies are essential for maintaining reliability.
  • Fourth, false alarm rates and alert fatigue must be carefully managed. Even highly accurate models may generate excessive alerts in high-volume ICU settings, potentially reducing clinician trust and system effectiveness.
Addressing these challenges requires prospective validation, workflow-aware system design, and close collaboration between data scientists and healthcare professionals.

7. Conclusions and Future Works

Sepsis remains a major challenge for global healthcare systems. Over the past two decades, substantial progress has been made in understanding its pathophysiology and establishing comprehensive management guidelines. While no universal cure exists, interventions such as early administration of antibiotics, hemodynamic resuscitation, judicious ventilator use, and careful blood product transfusion have markedly reduced morbidity and mortality. Emerging therapies, including immunomodulators, are still in their early stages but represent promising avenues for further research.
Robust clinical scoring systems, such as APACHE-II and SOFA, have been developed to aid in the evaluation and prognostication of sepsis. However, the diagnosis of sepsis remains contentious. Recent recommendations have moved away from the previously accepted SIRS criteria, favoring a more nuanced classification based on multiorgan dysfunction and SOFA scores. To contribute to this evolving field, we proposed a sepsis detection system that leverages large-scale patient data to predict sepsis onset. The system integrates machine learning techniques to analyze the processed data, ultimately identifying patients at risk of developing sepsis.
Future work will investigate the integration of advanced learning paradigms to further enhance early sepsis prediction and prevention. Recent studies have demonstrated the potential of large language models (LLMs) to support clinical reasoning by capturing complex contextual dependencies within electronic health records and medical narratives [32]. Such models could be leveraged to complement numerical time-series analysis with higher-level clinical knowledge and decision support. Furthermore, multimodal learning approaches that fuse heterogeneous data sources, including physiological signals, laboratory measurements, clinical notes, and imaging data, have shown promising results in improving robustness and predictive accuracy in healthcare applications [33]. Incorporating multimodal fusion mechanisms into the proposed framework represents a natural extension of this work and may enable more comprehensive and reliable early sepsis detection.

Funding

This research received no external funding.

Institutional Review Board Statement

All ethical and integrity considerations have already been addressed in the original publication of the dataset. Institutional Review Board approval was required for the study.

Informed Consent Statement

All ethical and integrity considerations have already been addressed in the original publication of the dataset; no additional informed consent from participants was required for the study.

Data Availability Statement

The data presented in this study are openly available in [PhysioNet/Computing in Cardiology Challenge 2019] at [https://physionet.org/content/challenge-2019/1.0.0/] (accessed on 1 March 2026).

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. Pinho, M.; Leal, F.; Miguel, I. Profiling Decision-Making Styles Under Healthcare Resource Scarcity: An Interdisciplinary Clustering Approach. Information 2026, 17, 287. [Google Scholar] [CrossRef]
  2. Singer, M.; Deutschman, C.S.; Seymour, C.W.; Shankar-Hari, M.; Annane, D.; Bauer, M.; Bellomo, R.; Bernard, G.R.; Chiche, J.D.; Coopersmith, C.M.; et al. The third international consensus definitions for sepsis and septic shock (Sepsis-3). JAMA 2016, 315, 801–810. [Google Scholar] [CrossRef]
  3. He, T.T.; Jiao, T.Q.; An, X.M. Risk prediction models for sepsis-associated encephalopathy: A systematic evaluation and meta-analysis. PeerJ 2026, 14, e20770. [Google Scholar] [CrossRef]
  4. Antcliffe, D.B.; Burnham, K.L.; Al-Beidh, F.; Santhakumaran, S.; Brett, S.J.; Hinds, C.J.; Ashby, D.; Knight, J.C.; Gordon, A.C. Transcriptomic signatures in sepsis and a differential response to steroids. From the VANISH randomized trial. Am. J. Respir. Crit. Care Med. 2019, 199, 980–986. [Google Scholar] [CrossRef]
  5. Burki, T.K. Sharp rise in sepsis deaths in the UK. Lancet Respir. Med. 2018, 6, 826. [Google Scholar] [CrossRef]
  6. Mou, C.; Yang, J.; Wu, Q.; Qin, L.; Lu, J. Progress in sepsis prediction models: From traditional scoring systems to multimodal intelligence and clinical translation. Front. Med. 2026, 13, 1732164. [Google Scholar] [CrossRef]
  7. Azizi, S.; Hoveidamanesh, S.; Bagheri, T.; Varaki, F.A.; Ghadimi, T.; Forghani, S.F. Modern machine learning techniques used in prediction of Sepsis and Bloodstream infection in Burn patients: A systematic review. Burns 2026, 52, 107965. [Google Scholar] [CrossRef]
  8. Mahmoudi, P.S.; Sadeghi, F.; Saberian, M.; Khalili, H.; Shafaati, M. Continuous vs. intermittent infusion of corticosteroids in septic shock: A GRADE-based systematic review and meta-analysis. J. Anesth. Analg. Crit. Care 2026, 6, 16. [Google Scholar] [CrossRef]
  9. Terrington, I.; Cox, O.; Copley, P.; Eastwood, B.; Webb, E.; McKenzie, C.; Saeed, K.; Conway-Morris, A.; Grocott, M.P.; Dushianthan, A. The role of corticosteroids in the management of non-COVID-19 severe community-acquired pneumonia in the intensive care unit: A narrative review. J. Intensiv. Care Soc. 2026, 27, 119–133. [Google Scholar] [CrossRef]
  10. Mahmud, F.; Quamruzzaman, M.; Sanka, A.I.; Cheung, R.C.; Chowdhury, M.H. Interpretable machine learning-based real-time sepsis diagnosis. Sci. Rep. 2026, 16, 6702. [Google Scholar] [CrossRef]
  11. Despraz, J.; Matusiak, R.; Nektarijevic, S.; Rossetti, V.; Bastardot, F.; Akrour, R.; Konasch, A.; Gauthiez, E.; Pignolet, O.; Pepe, S.; et al. An artificial intelligence-powered learning health system to improve sepsis detection and quality of care: A before-and-after study. npj Digit. Med. 2026, 9, 106. [Google Scholar] [CrossRef]
  12. Ustaalioğlu, İ.; Yıldız, F. Early sepsis prediction in elderly patients with urinary tract infections: A machine learning. Signa Vitae 2026, 22, 75. [Google Scholar]
  13. Lin, T.H.; Chung, H.Y.; Jian, M.J.; Chang, C.K.; Lin, H.H.; Yen, C.T.; Tang, S.H.; Pan, P.C.; Perng, C.L.; Chang, F.Y.; et al. AI-driven innovations for early sepsis detection by combining predictive accuracy with blood count analysis in an emergency setting: Retrospective study. J. Med. Internet Res. 2025, 27, e56155. [Google Scholar] [CrossRef]
  14. Wang, H.; Pounds, D.; Zhang, W.; Mokbel, A.Y.; Kabir, M.N.; Lin, X.Y.; Highlander, A.; Dehzangi, I. Early Sepsis Prediction Using Publicly Available Data: High-Performance AI/ML Models with First-Hour Clinical Information. Diagnostics 2025, 15, 2727. [Google Scholar] [CrossRef]
  15. Drysch, M.; Reinkemeier, F.; Puscz, F.; Hinzmann, J.; German Burn Registry Paul Christian Fuchs 2; Lehnhardt, M.; Wallner, C.; Schmidt, S.V. Streamlined machine learning model for early sepsis risk prediction in burn patients. npj Digit. Med. 2025, 8, 621. [Google Scholar] [CrossRef]
  16. Aityan, S.; Herrero, R.; Mosaddegh, A.; Tayyar, H.; Adebesin, E.; Jeedigunta, S.P.; Kim, H.; Mersini, M.; Lazzaro, R.; Iacovazzo, N.; et al. AI-Powered Early Detection of Sepsis in Emergency Medicine. Life 2025, 15, 1576. [Google Scholar] [CrossRef]
  17. Yadgarov, M.Y.; Landoni, G.; Berikashvili, L.B.; Polyakov, P.A.; Kadantseva, K.K.; Smirnova, A.V.; Kuznetsov, I.V.; Shemetova, M.M.; Yakovlev, A.A.; Likhvantsev, V.V. Early detection of sepsis using machine learning algorithms: A systematic review and network meta-analysis. Front. Med. 2024, 11, 1491358. [Google Scholar] [CrossRef]
  18. Zhou, L.; Shao, M.; Wang, C.; Wang, Y. An early sepsis prediction model utilizing machine learning and unbalanced data processing in a clinical context. Prev. Med. Rep. 2024, 45, 102841. [Google Scholar] [CrossRef]
  19. Pirracchio, R.; Hubbard, A.; Sprung, C.L.; Chevret, S.; Annane, D. Assessment of machine learning to estimate the individual treatment effect of corticosteroids in septic shock. JAMA Netw. Open 2020, 3, e2029050. [Google Scholar] [CrossRef]
  20. Annane, D.; Renault, A.; Brun-Buisson, C.; Megarbane, B.; Quenot, J.P.; Siami, S.; Cariou, A.; Forceville, X.; Schwebel, C.; Martin, C.; et al. Hydrocortisone plus fludrocortisone for adults with septic shock. N. Engl. J. Med. 2018, 378, 809–818. [Google Scholar] [CrossRef]
  21. Venkatesh, B.; Finfer, S.; Cohen, J.; Rajbhandari, D.; Arabi, Y.; Bellomo, R.; Billot, L.; Correa, M.; Glass, P.; Harward, M.; et al. Adjunctive glucocorticoid therapy in patients with septic shock. N. Engl. J. Med. 2018, 378, 797–808. [Google Scholar] [CrossRef]
  22. Keh, D.; Trips, E.; Marx, G.; Wirtz, S.P.; Abduljawwad, E.; Bercker, S.; Bogatsch, H.; Briegel, J.; Engel, C.; Gerlach, H.; et al. Effect of hydrocortisone on development of shock among patients with severe sepsis: The HYPRESS randomized clinical trial. JAMA 2016, 316, 1775–1785. [Google Scholar] [CrossRef]
  23. Moreno, R.; Sprung, C.; Annane, D.; Chevret, S.; Briegel, J.; Keh, D.; Singer, M.; Weiss, Y.; Payen, D.; Cuthbertson, B.; et al. Time course of organ failure in patients with septic shock treated with hydrocortisone: Results of the Corticus study. In Applied Physiology in Intensive Care Medicine 1: Physiological Notes-Technical Notes-Seminal Studies in Intensive Care; Springer: Berlin/Heidelberg, Germany, 2012; pp. 423–430. [Google Scholar]
  24. Rochwerg, B.; Oczkowski, S.J.; Siemieniuk, R.A.; Agoritsas, T.; Belley-Cote, E.; D’Aragon, F.; Duan, E.; English, S.; Gossack-Keenan, K.; Alghuroba, M.; et al. Corticosteroids in sepsis: An updated systematic review and meta-analysis. Crit. Care Med. 2018, 46, 1411–1420. [Google Scholar] [CrossRef]
  25. Fang, F.; Zhang, Y.; Tang, J.; Lunsford, L.D.; Li, T.; Tang, R.; He, J.; Xu, P.; Faramand, A.; Xu, J.; et al. Association of corticosteroid treatment with outcomes in adult patients with sepsis: A systematic review and meta-analysis. JAMA Intern. Med. 2019, 179, 213–223. [Google Scholar] [CrossRef] [PubMed]
  26. Reyna, M.A.; Josef, C.S.; Jeter, R.; Shashikumar, S.P.; Westover, M.B.; Nemati, S.; Clifford, G.D.; Sharma, A. Early prediction of sepsis from clinical data: The physionet/computing in cardiology challenge 2019. Crit. Care Med. 2020, 48, 210–217. [Google Scholar] [CrossRef]
  27. Cramer, J.S. The Origins of Logistic Regression; Tinbergen Institute: Amsterdam/Rotterdam, The Netherlands, 2002. [Google Scholar]
  28. Friedman, J.H. Greedy function approximation: A gradient boosting machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef]
  29. Parmar, A.; Katariya, R.; Patel, V. A review on random forest: An ensemble classifier. In International Conference on Intelligent Data Communication Technologies and Internet of Things (ICICI) 2018; Springer: Berlin/Heidelberg, Germany, 2019; pp. 758–763. [Google Scholar]
  30. Cervantes, J.; Garcia-Lamont, F.; Rodríguez-Mazahua, L.; Lopez, A. A comprehensive survey on support vector machine classification: Applications, challenges and trends. Neurocomputing 2020, 408, 189–215. [Google Scholar] [CrossRef]
  31. Anwar, A. Difference Between AutoEncoder (AE) and Variational AutoEncoder (VAE). 2021. Available online: https://towardsdatascience.com/difference-between-autoencoder-ae-and-variational-autoencoder-vae-ed7be1c038f2 (accessed on 1 March 2026).
  32. Li, D.; Yang, Y.; Cui, Z.; Yin, H.; Hu, P.; Hu, L. LLM-DDI: Leveraging Large Language Models for Drug-Drug Interaction Prediction on Biomedical Knowledge Graph. IEEE J. Biomed. Health Inform. 2025, 30, 773–781. [Google Scholar] [CrossRef] [PubMed]
  33. Wu, Y.; Chen, J.; Hu, L.; Xu, H.; Liang, H.; Wu, J. OmniFuse: A general modality fusion framework for multi-modality learning on low-quality medical data. Inf. Fusion 2025, 117, 102890. [Google Scholar] [CrossRef]
Figure 1. Overview of our system architecture.
Figure 1. Overview of our system architecture.
Information 17 00430 g001
Figure 2. Autoencoder architecture adapted in our system.
Figure 2. Autoencoder architecture adapted in our system.
Information 17 00430 g002
Figure 3. Sepsis and non-sepsis distribution.
Figure 3. Sepsis and non-sepsis distribution.
Information 17 00430 g003
Figure 4. Loss variation.
Figure 4. Loss variation.
Information 17 00430 g004
Figure 5. Distribution of patients after training.
Figure 5. Distribution of patients after training.
Information 17 00430 g005
Figure 6. SHAP values analysis according to sepsis features.
Figure 6. SHAP values analysis according to sepsis features.
Information 17 00430 g006
Figure 7. AUROC plotting of our framework vs state-of-the-art methods.
Figure 7. AUROC plotting of our framework vs state-of-the-art methods.
Information 17 00430 g007
Table 1. Summary of variables used in this study.
Table 1. Summary of variables used in this study.
CategoryVariablesCount
Vital signsHeart rate, pulse oximetry, body temperature, systolic/diastolic blood pressure, mean arterial pressure, respiratory rate, end-tidal CO28
Laboratory measurementsAcid–base indicators (pH, bicarbonate, base excess, PaCO2), renal markers (BUN, creatinine), electrolytes (Na, K, Ca, Mg, phosphate, chloride), liver enzymes (AST, alkaline phosphatase, bilirubin), hematologic markers (hemoglobin, hematocrit, platelets, WBC), coagulation and cardiac markers (PTT, fibrinogen, troponin I), metabolic indicators (glucose, lactate)26
Demographic and administrative dataAge, gender, ICU unit type, hospital-to-ICU admission delay, ICU length of stay6
OutcomeSepsis occurrence indicator1
Table 2. Feedforward Neural Network (FNN) architecture and training parameters.
Table 2. Feedforward Neural Network (FNN) architecture and training parameters.
ComponentConfiguration
Input dimension40 features
Hidden Layer 164 neurons (ReLU)
Hidden Layer 232 neurons (ReLU)
Output Layer1 neuron (Sigmoid)
OptimizerAdam
Learning rate0.001
Batch size32
Epochs100
Dropout rate0.2
L2 regularization λ = 0.001
Loss functionBinary Cross-Entropy
Table 3. Autoencoder architecture and training parameters.
Table 3. Autoencoder architecture and training parameters.
ComponentConfiguration
Input dimension40 features
Encoder Hidden Layer32 neurons (ReLU)
Latent dimension16 neurons
Decoder Hidden Layer32 neurons (ReLU)
Output dimension40 neurons (Sigmoid)
OptimizerAdam
Learning rate0.001
Batch size32
Epochs100
Loss functionMean Squared Error (MSE)
Table 4. Hyperparameter configuration of machine learning classifiers.
Table 4. Hyperparameter configuration of machine learning classifiers.
ModelSelected Hyperparameters
Logistic Regression (LR)Regularization: L2;     C = 1
Random Forest (RF)Number of trees: 200;    Max depth: None
Gradient Boosting (GB)Nb. of estimators: 200; Learning rate: 0.1; Max depth: 3
Support Vector Machine (SVM)Kernel: RBF;     C = 1 ;     γ = scale
Table 5. Performance comparison of machine learning models for early sepsis detection.
Table 5. Performance comparison of machine learning models for early sepsis detection.
ModelPrecisionRecallF1-ScoreAccuracy
Logistic Regression (LR)0.900.840.870.90
Random Forest (RF)0.880.790.830.87
Gradient Boosting (GB)0.890.810.850.88
Support Vector Machine (SVM)0.860.760.810.85
Table 6. Estimated individual and average treatment effects of corticosteroid therapy.
Table 6. Estimated individual and average treatment effects of corticosteroid therapy.
MetricMean EffectStd. Dev.Clinical Interpretation
ATE+0.080.03Modest population-level benefit
ITE (All patients)+0.050.12High inter-patient variability
ITE (High-risk subgroup)+0.180.07Strong benefit for selected patients
ITE (Low-risk subgroup)+0.010.04Minimal or no treatment benefit
Table 7. Impact of missing-data imputation strategies on classification performance.
Table 7. Impact of missing-data imputation strategies on classification performance.
Imputation MethodPrecisionRecallF1-ScoreAccuracy
Previous-value Imputation0.900.840.870.90
Mean-value Imputation0.890.830.860.89
KNN Imputation0.900.850.870.90
Table 8. Lead-time performance analysis.
Table 8. Lead-time performance analysis.
MetricValue
Average Lead-Time5.4 h
Median Lead-Time4.8 h
Table 9. Approximate performance comparison with existing sepsis prediction studies.
Table 9. Approximate performance comparison with existing sepsis prediction studies.
StudyMethodAccuracyPrecisionRecallF1-Score
[11]RF + LSTM0.810.780.740.76
[12]LR + LASSO0.840.800.770.78
[13]GB + RF0.850.820.790.80
Proposed MethodEnsemble Models0.900.900.840.87
Table 10. Statistical significance test (paired t-test) between the proposed method and compared methods.
Table 10. Statistical significance test (paired t-test) between the proposed method and compared methods.
Comparisont-Valuep-Value
Proposed vs. RF + LSTM [11]2.850.011
Proposed vs. LR + LASSO [12]3.120.007
Proposed vs. GB + RF [13]2.410.021
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Harb, H. From Data to Diagnosis: A Machine Learning-Enabled Framework for Early Sepsis Prediction and Prevention. Information 2026, 17, 430. https://doi.org/10.3390/info17050430

AMA Style

Harb H. From Data to Diagnosis: A Machine Learning-Enabled Framework for Early Sepsis Prediction and Prevention. Information. 2026; 17(5):430. https://doi.org/10.3390/info17050430

Chicago/Turabian Style

Harb, Hassan. 2026. "From Data to Diagnosis: A Machine Learning-Enabled Framework for Early Sepsis Prediction and Prevention" Information 17, no. 5: 430. https://doi.org/10.3390/info17050430

APA Style

Harb, H. (2026). From Data to Diagnosis: A Machine Learning-Enabled Framework for Early Sepsis Prediction and Prevention. Information, 17(5), 430. https://doi.org/10.3390/info17050430

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop