Next Article in Journal
Greedy Pursuit-Based Hierarchical Iterative Algorithm for Multi-Input Systems with Unknown Time Delays and Colored Noise
Previous Article in Journal
A Structure-Preserving Power-Adaptive Colored-Noise Kalman Filtering Algorithm for Four-Wire Pendulum Velocity Monitoring
Previous Article in Special Issue
Multimodal AI Algorithms for Risk Early Warning and Proactive Intervention in Community-Based Elderly Care: A Review
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Perspective

Machine Learning Applications in the ICU: Opportunities and Pitfalls

1
Division of Cardiology, Department of Medicine, University of Minnesota School of Medicine, Minneapolis, MN 55455, USA
2
Department of Electrical and Computer Engineering, Scott Engineering Center, University of Nebraska-Lincoln, Lincoln, NE 68588-0511, USA
*
Author to whom correspondence should be addressed.
Algorithms 2026, 19(10), 836; https://doi.org/10.3390/a19100836
Submission received: 9 June 2026 / Revised: 28 August 2026 / Accepted: 22 September 2026 / Published: 1 October 2026

Abstract

Intensive Care Units (ICUs) care for the sickest patients in any given hospital system and are associated with the worst outcomes and healthcare-associated costs. There is great interest in tailoring or streamlining ICU care to improve outcomes and reduce costs. With the growing use of machine learning (ML) in healthcare research, there has been increased utilization of ML-based approaches to address questions in ICU care. This necessitates an approach that considers the opportunities and pitfalls inherent to these tools. This paper aims to provide a brief overview of ML approaches to ICU datasets before outlining the opportunities for ML applications in this area and describing key pitfalls that arise when using ML tools to solve ICU and healthcare-related problems, namely that the current methods of addressing class imbalances and data missingness in healthcare datasets are at risk of causing deviations between model data and real-world signals, and that the resulting models are often difficult to apply to patient care. The contribution of this paper is to critically examine these pitfalls and provide guidance in the use of ML in ICU applications, such that solutions can be designed in a way that accounts for these potential weaknesses.

1. Introduction

Intensive care units (ICUs) typically care for the sickest patients within their hospital systems. Appropriate treatment and subsequent rehabilitation of critically ill patients are the primary goals of the various treatment teams involved in their care, and there is a significant incentive to streamline and tailor their treatment. From a systems perspective, appropriate treatment and minimization of long-term morbidity carry high societal and economic costs. Williams [1] found that in-hospital mortality for patients who spent less time (≤10 days) was significantly lower than those with longer stays, with a tendency toward lower long-term mortality. Longer ICU stays tend to consume more hospital resources—Rodriguez-Villar [2] found that the 6.5% of patients across 30 years with >20-day stays in an ICU consumed 29.5% of the ICU’s resources, totaling approximately $14 million. The effects and costs can extend beyond the walls of the ICU as well—a meta-analysis by Williams [3] found that the 5-year mortality of general ICU admissions ranged from 40–58%, suggesting significant attrition among patients even after surviving an ICU stay.
In the above examples, length of stay and cost of stay are metrics of ICU care. Logically, the longer a patient spends in an intensive care setting, the sicker they are, the more costs are accrued, and the more complex their care is, with “complexity” being a catch-all term for the number of medications, procedures, consultations, and disciplines involved in a patient’s care. These metrics can also be viewed as surrogates for the effectiveness of the care provided to the patients: more effective care should ideally translate to fewer unnecessary interventions, a shorter overall stay, and better outcomes. This characterization is rife with oversimplification but can still serve as the basis for a long-standing question: how can appropriate therapies be more rapidly identified and delivered to patients to improve their care?
Given the complexity of the human body, rapid identification (or phenotyping) of patients at risk of various health conditions, and, by extension, their candidacy for different treatments and interventions, is a skill honed by healthcare practitioners across every medical system at every level. Medical research is a complex and multifaceted endeavor, but one characterization of its goals is to produce knowledge that can be applied to the care of future patients [4]—however, the sheer volume of data that can be accessed for each individual patient and the difficulty in collating such data in a way that is easily parsed make this an increasingly complex endeavor. Beyond some basic (and often subjective) distinctions between ICU subtypes (such as cardiac, surgical, or medical units), patient phenotyping within intensive care frameworks often ends at the unit’s door. As a result, patient populations within ICUs tend to be very heterogeneous in terms of demographics, initial presentation, illness, and acuity. In addition, most ICU patients are at risk of poor outcomes, be it from mortality or long-term morbidity as a result of their presenting condition. All of these factors pose a significant challenge for both clinicians and hospital systems. Given this heterogeneity of data, as well as the high workload placed on ICU practitioners and high financial costs faced by healthcare systems, there is significant interest in using machine learning (ML) practices to sift through data in an attempt to more rapidly stratify patients, improve care delivery, and, in theory, reduce clinician workload. However, it is important to be mindful of the assumptions and adjustments to data that are made in the process of using these tools, especially when applied to critically ill and otherwise vulnerable populations in the healthcare sphere. The goal of this paper is to provide a brief overview of ML use in ICU populations, examine opportunities for ML applications in the ICU, and explore three significant pitfalls faced by projects seeking to use these tools in this setting.

2. Overview: ML Approaches to the ICU

While ML tools are increasingly applied to a wide range of problems across many fields, care should be taken when applying them in healthcare settings. Both healthcare problems and the human populations they affect tend to hide significant heterogeneity, which can hamstring underprepared algorithmic approaches. In this section, we provide a brief overview of how ML approaches are typically framed when applied to healthcare and ICU settings.
When defining a question in a healthcare setting, it is most helpful to frame it in terms of a population and an outcome, where ‘population’ refers to a specific set of patients of interest and ‘outcome’ to a characteristic, event, or result to be identified. Given the heterogeneity of healthcare populations, distilling the population of interest into a focused subset helps ensure a cleaner signal when looking for relationships between these populations and their outcomes. Most medical trials over the past 30 to 40 years have shown that a stringent set of inclusion and exclusion criteria is instrumental in improving the final outcome. This naturally comes at the cost of generalizability of the final conclusion, and it is important that this not be forgotten. Similarly, the ‘outcome’ that is under study should be defined as clearly as possible.
For example, say we wish to examine complication rates across different surgeries. To focus this query, it is important to ask a series of clarifying questions: Which surgeries are we examining? Are they done emergently or electively? Which patients are we interested in? Is there a specific age range, and are there any associated conditions? And what types of complications are we looking for? When do the complications occur, i.e., during the surgery or within a time period after? By doing this, we can more easily fit our desired question into a framework of inputs and outputs, or features and labels, thereby improving our chances of finding an answer.
The inputs to machine learning algorithms are generally referred to as features. We want to identify and select the inputs or features most helpful for the task at hand. Feature identification in medical populations is often complex. Electronic health records (EHRs) now enable access to large volumes of patient data, ranging from basic vitals like height, weight, heart rate, and blood pressure to lab values, diagnostic and procedure codes, and even clinician notes spanning nursing, nutrition, and physician care. This data is multiplied when patients are admitted to the hospital, as vitals assessments are repeated regularly throughout the day and multiple medical teams enter documentation on patients concurrently over the course of the stay. Once a patient is in an ICU, readings are obtained even more frequently, and continuous physiologic data, including heart rate, heart rhythm, blood oxygen levels, and cardiac function, can be added to the mix. Add to this imperfections in charting, errors in machine readings, and occasional missing data, and the task of feature selection becomes more complicated than simply querying all patient data and plugging it directly into an algorithm.
The primary challenge in feature selection is finding features that not only pertain to the outcome of interest but are also obtained frequently enough for analysis. The latter can be determined from EHR data, but the former is difficult to determine without substantial research or input from experienced clinicians. This input helps nonmedical research staff distinguish clinically or situationally relevant variables from redundant or extraneous data. Alternatively, maximizing feature count can improve the final model’s fit but risks including redundant data, since records may contain multiple correlated signals or even duplicate data. A higher feature count may boost model performance, but we should always exercise caution to avoid overfitting. To that end, dimension-reduction techniques such as principal component analysis and regularization (where applicable) can help avoid overfitting in the final model.
Once features are selected, as with any other ML application, missingness can be assessed. Imputation is commonly used in different healthcare ML studies to ‘fill in’ missing data, provided missingness is below an arbitrary threshold, but it should be noted that data in a healthcare population is rarely missing completely at random (MCAR), and making assumptions about missing data can negatively impact the overall ‘truth’ of the resulting model. Similarly, depending on the question being asked, the population of interest may be quite small. Different studies have used oversampling techniques such as bootstrapping and synthetic minority oversampling (SMOTE) [5,6,7,8,9]. While these are always helpful for increasing the amount of data for model development, they may again adversely affect the final model’s performance during validation, as we discuss later.
In supervised model construction, validation is often one of the last steps to determine whether the model holds up under scrutiny. This is often the best time to determine whether the model performs well across different datasets or has been overfit to the training data. Again, if there is a limited number of patient records available, this can prove difficult, as strict partitioning of data into training and test sets can reduce our ability to train the model. To that end, K-fold cross-validation is a common tool in medical ML studies [7,10,11], allowing researchers to train the model multiple times to assess its performance across varied data. If possible, validation on an external dataset is preferred, as it more cleanly avoids overfitting.
By comparison, label selection is somewhat easier, if it is done at all. Typically, ML in healthcare populations targets identification based on diagnostic criteria. In most EHRs, diagnoses are encoded using International Classification of Diseases (ICD) codes, which are publicly available. Procedures are similarly coded using Current Procedural Terminology (CPT) codes. Provided the desired diagnosis or procedure is found within these codes and the EHR is deemed to reliably document them when needed, label identification becomes trivial. If the label of interest is not captured by these codes, or if alternative diagnostic criteria are desired (for example, a more stringent set of criteria for a diagnosis than the current standard of care), then post-processing of the data will be needed in order to properly identify the patients with the desired label.
It should be noted that while labels are necessary for supervised machine learning, unsupervised machine learning can also be used in ICU studies. Specifically, it can be used to identify and phenotype patient populations that either lack an official diagnosis per the current standard of care or fall at the boundaries of certain diagnostic designations. In this case, it can be useful to subject a population to unsupervised model construction and then examine the labels to see whether the model identifies the same populations as the clinical teams that diagnosed them. Alternatively, an explainable supervised learning algorithm can be used to construct a model that identifies the weights of different features in classifying known labels. This also allows for characterization of a population, but limits it to known labels. An example of this can be found in Kresoja et al. [12], which analyzes patient survival in the previously conducted ECLS-SHOCK trial.
While discussing work involving patient health information, we would be remiss not to reinforce that, as with all research involving individual patient data, privacy and ethics guidelines should be rigorously followed and enforced. Given the growing awareness of ML and general artificial intelligence approaches, extra care should be taken to foster dialogue between research teams and institutional review boards to ensure that all institutional restrictions and safeguards against data misappropriation and patient safety breaches are respected. The growing off-site and outsourced nature of computational analysis may make this difficult and require transparent discussions about the feasibility of using on-site tools and data storage to minimize the risk of data breaches. It can also require simple elbow grease to anonymize large swaths of patient data to avoid both data breaches and the possibility of protected patient information skewing model construction. Additional methods for disclosing protocols to patients and for oversight specific to ML applications should be seriously considered and implemented where necessary [13]. And finally, any trials arising from ML-based research should be conducted to the same rigorous standards as other medical trials. The CONSORT-AI and SPIRIT-AI guidelines can act as starting points for the responsible design and implementation of such trials [14].

3. Opportunities for ML Use in the ICU

As we have noted, machine learning is not a new tool. Use of machine learning to analyze ICU populations has been on the rise since at least 2015, though attempts to use it predate that. Shillan et al. [15] conducted a systematic review of ML analysis of ICU data from 1991 to 2018 and found that nearly half of the studies were published after 2015. Until 2000, all studies used neural networks, but other modalities gained popularity thereafter, including support vector machines (SVMs), random forests, and classification trees. Hong et al. [16] conducted a broader literature review from 1980 to 2020 and found that the majority of studies used supervised learning methods, most of which focused on outcome prediction and prognosis. Syed et al.’s [17] focused review of ML- and Deep Learning (DL)-based applications using the MIMIC-III dataset (a large publicly available database of patients admitted to critical care units at Beth Israel Deaconess Medical Center) found that most studies focused on mortality prediction, followed by risk stratification, prediction diagnoses such as sepsis, cardiac episodes, or acute kidney injury, then resource management, and readmission risk. While ML applications in intensive care have broadened over the past three decades, approaches to each problem typically follow the same beats outlined in our overview above, and the breadth of work highlights opportunities for ML-based improvements in ICU work. To provide examples of these opportunities, we will look at four domains of ML use in the ICU: mortality prediction, disease state, clinical course, and signal analysis.

3.1. Mortality Prediction

Mortality is the easiest to understand among the three domains and is the most common application of supervised machine learning in intensive care datasets, as there is significant weight placed on quickly and accurately identifying patients at high risk of mortality, and thus with the most to gain from care in the ICU.
There are a handful of score-based clinical decision support tools that are used by clinicians to rapidly risk-stratify patients, such as the Sequential Organ Failure Assessment score (SOFA, with 8 variables addressing six organ systems [18,19]), the Simplified Acute Physiology Score (SAPS II, with 17 variables [20]), and the Acute Physiology and Chronic Health Evaluation (APACHE II, with 12 variables [21]). In particular, SOFA and APACHE II were built using expert consensus, while SAPS II was derived using logistic regression, making it an early example of the use of supervised machine learning in ICU care. These tools generally perform favorably in validation—SOFA has reported AUCs ranging from 0.6 to 0.9 across studies, SAPS II’s initial derivation noted a validation AUC of 0.86, and APACHE II performs similarly in the literature. They also have the advantage of being relatively portable and well-known, but it must be noted that they provide a single data point in a complex and fragile patient population. As such, there is strong interest in developing improved, more flexible prognostication tools.
In the existing literature surrounding ML approaches to mortality prediction, models vary widely in the number of features they use. Many publications go far beyond the 17 variables of SAPS II, including dozens [22,23] and sometimes hundreds [24,25] of available categorical and numeric features. Some publications try to narrow the target population in order to improve the performance of the model, focusing on specific groups with known diagnoses such as acute kidney failure [26,27], sepsis [28], or myocardial infarction [29,30,31], or patients in specific wards [32,33] or after specific procedures [34]. All of these represent promising avenues towards higher-fidelity mortality prediction, provided they can overcome the pitfalls we will outline later.

3.2. Predicting Disease States and Identifying Phenotypes

“Disease state” refers to whether a patient developed or received treatment for a specific diagnosis or complication during their hospitalization. As an example, consider sepsis. Sepsis (and its more severe form, septic shock) is a disease state broadly consisting of a maladaptive and detrimental physiologic response to an infection. In 2011, sepsis was estimated to account for more than $20 billion in total US hospital costs, and it is estimated to be a leading cause of mortality worldwide [35]. The mainstay of sepsis management is early recognition and treatment, and the U.S. Centers for Medicare & Medicaid Services has instituted sepsis care guidelines tied to hospital reimbursement. The current guidelines require lab testing and patient treatment within strict time periods from presentation [36], the latest of which is termed SEP-1. Prior to this, scoring systems such as SIRS [37] and qSOFA [38] were designed to help clinicians more easily identify patients with sepsis and guide treatment, though SIRS has since fallen out of favor.
Given how closely sepsis outcomes are monitored and tied to hospital remuneration, it is no surprise that a large body of data exists on the timely identification and intervention for patients with or at risk of sepsis (or, in some cases, infection in general [39]). Desautels et al. have published work that attempts to predict sepsis in patients using a proprietary algorithm called InSight [40,41]. And, as in our examples on mortality prediction, there is interest in applying machine learning to both cross-population comparisons of sepsis and other disease states [42] and in identifying sub-phenotypes of sepsis [43]. In this case, a successfully implemented ML tool that identifies sepsis could potentially save hospital systems (and, by extension, patients) significant sums in hospital costs.
Acute kidney injury (AKI) is another example of a common diagnosis in ICU populations with potentially outsized effects. AKI refers to a rapid decrease in renal function, leading to the accumulation of metabolic byproducts that would otherwise be removed by functional kidneys [44]. This accumulation can lead to significant toxicity, associated organ dysfunction, and, when severe, death. As with sepsis, there is significant interest in using machine learning to predict AKI. Given the multiple diagnostic criteria, efforts at early identification using ML can and have tried to use individual aspects of the diagnosis as the label of interest, such as urine output [45] and blood lab values [46]. However, in both sepsis and AKI prediction models, many publications default to using diagnostic codes as the label of interest, relying on clinician records over specific data points to identify the population of interest [47,48,49,50]. Acute kidney injury is commonly associated with a host of poor prognostic markers for patients, including delirium [51], need for chronic dialysis, and death [52]. An ML tool that could reliably identify patients either at risk for or in the process of developing an AKI could potentially either improve the lives of patients or save them outright by preventing the complication from occurring.
Special care should be taken to note the potential investigative opportunities in this domain. We had previously noted that unsupervised algorithms can be used to identify subsets of patient populations, such as in the work of Kresoja et al. [12]. Another example of this can be found in the work of Zweck et al. [53], which applied unsupervised clustering to a group of patients diagnosed with cardiogenic shock. In doing so, the group identified previously undefined phenotypes within the diagnosis, which were subsequently reproduced in separate datasets and later confirmed by biochemical analyses of similar populations [54]. This can be taken as an example of how machine learning applications can be extended to identify new diagnoses as well as existing ones with the right approach.

3.3. Predicting Clinical Course

Clinical course more broadly assesses whether a patient underwent specific events during their hospitalization, such as a procedure, surgery, admission to a specific ward or unit, or any identifiable clinical change such as deterioration, organ failure, death, or survival. In particular, repeated patient trips to the ICU during one hospitalization often signal poor prognosis and outcomes, making them an important signal to detect and reduce when determining appropriate patient care. In addition, CMS reimbursement is closely tied to length of stay, so repeat ICU trips or unanticipated clinical changes that lengthen stay are events administrative systems are keen to limit. ML techniques have been applied to determining patient candidacy for treatments or interventions, as well as the likelihood of change in clinical status. The more strenuous the intervention or deleterious the change, the more important it becomes to accurately identify patients beforehand. If we assume that changes in patient clinical status are always preceded by some signal embedded within their clinical data (that is, the process by which patient clinical status changes is deterministic), then if patient data is subjected to analysis to find these patterns, the deterioration can be predicted before it occurs, at least in theory. Multiple studies have explored this theory using EHR and biosignal data, with varying degrees of success. The assumptions underlying this result are subject to significant caveats, as we will discuss later.
Practically speaking, given adequate data, applying a label to individual patient records for retrospective model construction is straightforward. In this case, the sheer volume of data accumulated over the course of a patient’s hospitalization (vitals, labs, diagnostic and procedural codes, notes, medications, etc.) can become an obstacle with longer or more morbid stays, resulting in very high numbers of features if not carefully curated. Feature selection can be done using existing frameworks for patient care de-escalation [55], or regularization can be applied to cohorts to decrease model complexity [56]. Alternatively, approaches using deep learning algorithms, such as neural networks, tend to have higher feature counts [57,58,59], often at the cost of explainability. Neural networks of various types, such as feedforward neural networks [60,61] and deep neural networks [62], can still play a role in model construction, provided care is taken to account for their outputs and to explain the final model. The manuscript by Stephens et al. [62] should receive special attention for this. This group used a deep neural network approach to try to build a calculator to predict patient outcomes on ECMO. In their work, they use a tool called SHAP (SHapley Additive exPlanations), first described by Lundberg and Lee [63] as a way to explain the contributions of various features to complex models. While this will not explain exactly how the model was constructed, it provides significant benefit by showing how much each feature affects the outcome, offering more insight to clinicians who may use the model.
The same data can be used to try to predict patient deterioration. This requires defining what constitutes a deteriorating patient course to generate the appropriate label. No single vital sign or lab value can be used to determine whether a patient is ‘doing worse’ within a reasonable spectrum. Obviously, if a patient’s heart rate, blood pressure, or respiration were to go to 0, it is a sign of dire straits, but this would not happen in isolation in the overwhelming majority of patients (and depending on the vital sign in question, could not physiologically happen without affecting some other signal). Given that the criteria for escalation of care can vary from hospital to hospital, this definition remains mutable. Attempts range from limiting it to a protracted decline in one vital sign [10] to detecting deviation in multiple vital signs from pre-established normal ranges [64,65].
Finally, determining patient candidacy for interventions and procedures is a large part of clinical care and can be difficult when a quick decision is needed. In critical care, this can apply to a myriad of procedures, but for the sake of discussion, we will use extracorporeal membrane oxygenation (ECMO) as our example. ECMO is a form of percutaneous pulmonary or cardiopulmonary bypass and one of the most strenuous and resource-intensive therapies available outside of an operating theater. It is considered for patients at the extremes of pulmonary and cardiovascular collapse as a way to allow time for medical teams to treat diseases that would otherwise kill the patient quickly. The most notable example of this is severe shock from a heart attack, but it can also be considered in other situations that compromise cardiopulmonary function, such as severe trauma or pneumonia from infections such as COVID-19. Given the resources required to place and maintain a patient on ECMO, as well as the significant morbidity associated with its use even in successful cases, selection of patients for its use is a significant discussion point in intensive care, especially given that patients who could benefit from ECMO typically can’t tolerate significant delays in therapy. Large repositories of international data collected by the Extracorporeal Life Support Organization (ELSO) have been assembled to provide data support for understanding ECMO use and effects. This has been leveraged to identify the best patient candidates, with the most visible attempts (the SAVE and RESP scores for different ECMO modalities) using logistic regression [66,67,68].
The ultimate target of the above interventions can be boiled down to the early identification of patients who are at risk of deterioration or would benefit from early intervention. Successful implementation of such tools would mean not only better patient outcomes but also more tailored resource allocation and utilization, all of which represent significant opportunities for growth and improvement.

3.4. Signal Analysis

Hospital systems now universally use digital systems to monitor and display patient vitals, including ECGs and invasive hemodynamic monitoring of systemic and pulmonary arterial waveforms. This data is often stored in EHR systems for record-keeping and can thus be analyzed. Vital sign monitors include rudimentary alarm systems to support nursing care based on immediate signal analysis (for example, alarms can be triggered by low heart rate, low blood pressure, or certain changes in ECG waveforms). With increasing access to biosignal data repositories, some groups have attempted to mine these signals to develop methods for predicting changes in the clinical course. Biosignals are typically processed to extract time- and frequency-domain features for analysis, either in isolation [69,70] or in conjunction with information from the medical record [11]. If a tool is successfully developed, it could augment any of the fields above to improve patient care, whether through early prediction systems or patient identification, using tools already in the field.

4. Pitfalls of ML Applications in ICU Datasets

Thus far, we have provided a short overview of ML approaches to ICU care and attempted to explain how to frame this approach. However, implicit in this approach are pitfalls that may prove detrimental to the utility of ML in this particular setting. In this section, we will explain three of these pitfalls: class imbalance and the inherent difficulty of correcting it in this population; missing data and the dangers of imputation in healthcare data; and the inherent gap between model performance and actual clinical utility.

4.1. Reliance on Oversampling to Overcome Class Imbalance Alters the Data and Results in Synthetic Results That May Degrade Model Reliability

In investigations of ICU data, populations of interest often constitute a minority of the overall dataset, leading to significant class imbalance. This imbalance violates any prior assumptions ML algorithms may have about whether the different classes have similar prior probabilities and can result in classifiers being either overwhelmed by the majority class or lacking sufficient resolution to reliably distinguish the minority class. Class imbalance is a well-documented phenomenon, and within the ML space there are a number of remedies that have been described, such as oversampling and undersampling [71]. Within the medical sphere, this is most often remedied using techniques that involve oversampling of the minority class of interest in order to synthetically inflate the size of the class of interest, typically using an algorithm such as Synthetic Minority Oversampling Technique (SMOTE), originally described by Chawla et al. in 2002 [5]. While these techniques are adept at making class imbalances ‘go away’, the actual effect that this has on the dataset must be examined and acknowledged.
As an example, consider the work by Awad et al. on predicting early mortality in ICU patients [22]. Their investigation built a classifier using pre-collected ICU data. After filtering down to the patient population of interest, this data consisted of 11,712 records, of which 1488 documented patient deaths. Throughout the paper, the investigators run an ensemble learning algorithm on several different ’versions’ of this data, including the original, mean-imputed, algorithmically imputed, and oversampled data, as well as all combinations thereof. Detailing these findings in full is beyond the scope of this paper, but we note that SMOTE improved the classification performance in a majority of cases. This intuitively makes sense, but the true effect of oversampling on model fidelity is difficult to fully elucidate. Central to this is the fact that even though oversampling in theory uses existing data to create its synthetic data points, it is still synthetic data, and this synthetic data may have been injected into the entire dataset, not just the training set. This introduced ’new’ information can degrade the fidelity of the final model. This is supported by the literature, including Tarawneh et al. [72] and Alkhawaldeh et al. [73], which note that all validated oversampling approaches can suffer from errors in the synthesized data. If erroneous data is introduced into model training, it can have a knock-on effect on model performance when it is put into clinical use, and if we do not know when the data was introduced during model construction, it becomes very difficult to account for.
Adding to this is the fact that the synthetic data is not random. Oversampling of a specific class results in the injection of data that is, in effect, already trained and can result in data leakage if inappropriately applied. This, in turn, can result in overoptimistic performance metrics. Demircioglu [74] demonstrated that applying SMOTE to the entire dataset prior to cross-validation, rather than only to the training data, can lead to an overestimation of the ML algorithm’s performance. This inaccuracy, in turn, can have unpredictable effects on the real performance of the final model.
While this uncertainty may fall within operational tolerance for certain applications, healthcare is likely not one of them. Theoretically, a skewed model can overdiagnose, underdiagnose, or diagnose at random. All of these possibilities carry significant risks for individual patients and healthcare systems. If we take a maximalist view of healthcare, it can represent, at a minimum, overuse of procedures and overcommitment of limited resources, and at a maximum, needless patient mortality and morbidity from over- or underutilization of said resources. Thus, although the effect of synthetic data generated by oversampling entire data sets is uncertain, it cannot ultimately be ignored.

4.2. Missing Data Is a Significant Barrier That Cannot Be Reasonably Overcome with Imputation

Another common barrier to analysis in ICU and healthcare research is missingness. Nearly every healthcare dataset of substantial size has some degree of missing data, which is commonly addressed by either imputation or elimination of missing records prior to model construction. Imputation of missing data can take place in any number of ways, but it always operates under the assumption that data is either missing completely at random or is at least missing at random. Violating these assumptions can skew the data. Therefore, these can be dangerous assumptions to make, particularly when dealing with healthcare data.
As an example, consider patients with late- or end-stage diseases such as severe emphysema or cancer who are admitted to the ICU. If a patient with a known poor prognosis is admitted to the ICU and there is already momentum toward comfort-oriented care, vitals and labs may be obtained less frequently, which may be reflected as ’missing’ in a dataset. If the management of sicker patients follows this trend, then any missing data is, by definition, not missing completely at random. The hypothetical situation can extend in the other direction as well, with healthier or less unstable patients needing less frequent check-ins and thus having less data overall.
As it stands, we don’t need to treat variations in the density of patient data as theoretically non-random. There is one particular study that suggests that it is truly nonrandom. The CONCERN early warning system [75] attempts to predict patient deterioration by processing metadata from nurse-entered documentation, as well as the language used in their narrative notes, with particular attention to increased assessment frequency and assessments at ’uncommon times’, among other signals. This information is converted into a risk score and conveyed to a covering provider. A cluster-randomized controlled trial of this system found that use of the early warning system resulted in a 35.6% reduction in the instantaneous risk of death, decreased sepsis risk, decreased length of stay, and increased risk of unanticipated ICU transfer, compared with groups that didn’t use the system. Broadly speaking, we can describe the CONCERN system as monitoring for changes in data density during patient stays. If an increase in density can be interpreted as a signal portending a change in condition, loss of data density (like what is seen in EHR databases) cannot be treated as statistically neutral, thus leading to the breakdown of the assumptions driving imputation and implying that any imputation in these datasets will result in skewed data.

4.3. Strong Performance Metrics of Models in Retrospective Datasets Do Not Necessarily Confer a Benefit to Actual Frontline Medical Care

The ultimate goal of most medical investigations is to try to improve patient care in some way, shape, or form. To that end, we should always consider how the ultimate result of our work can be implemented to achieve our goal. Despite the large number of models that have been developed and refined in the literature, very few have been directly applied to patient care in this way. This dearth of implemented models highlights the fact that, while it’s possible to come up with a prediction for almost any outcome, it can be difficult to translate that into a change in care. This difficulty can arise from infrastructural barriers, such as the cost of implementation, or from the simple fact that there is no role for the model in the current standard of care. A review by Munõz et al. in 2026 [76] of RCTs focused on machine learning interventions in ICUs found only 10 randomized trials, and of those, only 2 showed a benefit in patient mortality. One trial [77] tested a proprietary sepsis warning algorithm, involving 142 patients across two units of one hospital in a non-blinded study. The other was the CONCERN trial, which we have discussed already.
Another trial that covered a broader footprint than just the ICU was conducted in 2022 to test the TREWS model. This is a proprietary algorithm also targeting sepsis [78]. This time, 6877 patients across all units in 5 hospitals were involved, but the trial was not randomized—the arms consisted of patients for whom a model-generated alert was acknowledged within a specific timeframe, and those for whom it wasn’t. This leaves open a host of questions centering on why the alert wasn’t acknowledged, and whether that led to the difference in outcomes.
We detail this not to discourage the development of ML-based tools in ICU care, but to highlight the need for a thoughtful approach when developing such tools. Again, we highlight the existence and availability of resources to guide the development of clinical trials that use ML- and AI-based tools, namely the CONSORT-AI extension of the 2010 CONSORT statement detailing RCT trial reporting, specifically covering AI and ML-based trial reporting [14], and the SPIRIT-AI guidelines on clinical trial protocols centered around AI and ML tools. These guidelines, combined with guidance from IRBs and research advisory boards, should be considered when constructing tools and the subsequent trials to test them.

5. Conclusions

ML-based approaches to medical research are becoming increasingly common. As such techniques appear more often in medical settings, it becomes necessary to understand not only these techniques but also the potential merits and shortcomings.
In this paper, we briefly outline how ML tools can be applied to ICU care, and present a perspective on opportunities and major pitfalls for anyone seeking to pursue them. In doing so, we highlight the challenge of working with class imbalance in data where the class imbalance is itself intrinsic to the populations of interest, and the difficulty in addressing missing data where the missingness itself can carry a signal. We also discuss the importance of understanding how the resulting model will be used to inform healthcare provision, and the need for active dialogue between data scientists and healthcare professionals to determine the overall usefulness of the research.
The breadth of ML-based healthcare research and the paucity of useful trials within it are notable features of the current landscape, highlighting the importance of a thoughtful approach to any project that uses ML to solve an ICU-related problem. With careful data acquisition and stewardship, thoughtful model construction, and responsible subsequent trial design, we may yet see the development of tools that ultimately change the face of intensive care and improve care and outcomes for critically ill patients.

Author Contributions

Conceptualization, S.S.S. and K.S.; writing—original draft preparation S.S.S.; writing—review and editing S.S.S. and K.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Acknowledgments

We would like to thank the reviewers for their careful reading and their very constructive comments. In particular, we would like to thank Reviewer 3, whose guidance led us to significantly change our submission.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Williams, T.A.; Ho, K.M.; Dobb, G.J.; Finn, J.C.; Knuiman, M.; Webb, S.A.R. Effect of length of stay in intensive care unit on hospital and long-term mortality of critically ill adult patients. Br. J. Anaesth. 2010, 104, 459–464. [Google Scholar] [CrossRef] [Scilit]
  2. Rodriguez-Villar, S.; Yuste, R.B. Long-term admission to the intensive care unit: A cost-benefit analysis. Rev. Esp. Anestesiol. Reanim. 2014, 61, 489–496. [Google Scholar] [CrossRef] [Scilit]
  3. Williams, T.A.; Dobb, G.J.; Finn, J.C.; Webb, S.A.R. Long-term survival from intensive care: A review. Intensive Care Med. 2005, 31, 1306–1315. [Google Scholar] [CrossRef] [Scilit]
  4. Litton, P.; Miller, F.G. A normative justification for distinguishing the ethics of clinical research from the ethics of medical care. J. Law Med. Ethics 2005, 33, 566–574. [Google Scholar] [CrossRef] [Scilit]
  5. Chawla, N.V.; Bowyer, K.W.; Hall, L.O.; Kegelmeyer, W.P. SMOTE: Synthetic Minority Over-sampling Technique. J. Artif. Intell. Res. 2002, 16, 321–357. [Google Scholar] [CrossRef] [Scilit]
  6. Mikhno, A.; Ennett, C.M. Prediction of extubation failure for neonates with respiratory distress syndrome using the MIMIC-II clinical database. In Proceedings of the 2012 Annual International Conference of the IEEE Engineering in Medicine and Biology Society; IEEE: New York, NY, USA, 2012; pp. 5094–5097. ISSN 1558-4615. [Google Scholar] [CrossRef] [Scilit]
  7. Huddar, V.; Desiraju, B.K.; Rajan, V.; Bhattacharya, S.; Roy, S.; Reddy, C.K. Predicting Complications in Critical Care Using Heterogeneous Clinical Data. IEEE Access 2016, 4, 7988–8001. [Google Scholar] [CrossRef] [Scilit]
  8. Dervishi, A. Fuzzy risk stratification and risk assessment model for clinical monitoring in the ICU. Comput. Biol. Med. 2017, 87, 169–178. [Google Scholar] [CrossRef] [Scilit]
  9. Nakatsu, R.T. Validation of machine learning ridge regression models using Monte Carlo, bootstrap, and variations in cross-validation. J. Intell. Syst. 2023, 32, 20220224. [Google Scholar] [CrossRef] [Scilit]
  10. Lee, J.; Mark, R.G. An investigation of patterns in hemodynamic data indicative of impending hypotension in intensive care. Biomed. Eng. Online 2010, 9, 62. [Google Scholar] [CrossRef] [Scilit]
  11. Cherifa, M.; Blet, A.; Chambaz, A.; Gayat, E.; Resche-Rigon, M.; Pirracchio, R. Prediction of an Acute Hypotensive Episode During an ICU Hospitalization With a Super Learner Machine-Learning Algorithm. Anesth. Analg. 2020, 130, 1157. [Google Scholar] [CrossRef] [Scilit]
  12. Kresoja, K.P.; Zeymer, U.; Thevathasan, T.; Rassaf, T.; Jung, C.; Pöss, J.; Schneider, S.; Desch, S.; Freund, A.; Thiele, H. The Influence of Extracorporeal Life Support on Patients in Cardiogenic Shock Assessed by Machine Learning. JACC Cardiovasc. Interv. 2025, 18, 273–275. [Google Scholar] [CrossRef] [Scilit]
  13. Tilala, M.H.; Chenchala, P.K.; Choppadandi, A.; Kaur, J.; Naguri, S.; Saoji, R.; Devaguptapu, B.; Tilala, M. Ethical considerations in the use of artificial intelligence and machine learning in health care: A comprehensive review. Cureus 2024, 16, e62443. [Google Scholar] [CrossRef] [Scilit]
  14. Liu, X.; Rivera, S.C.; Moher, D.; Calvert, M.J.; Denniston, A.K.; Ashrafian, H.; Beam, A.L.; Chan, A.W.; Collins, G.S.; Deeks, A.D.J.; et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. Lancet Digit. Health 2020, 2, e537–e548. [Google Scholar] [CrossRef] [Scilit]
  15. Shillan, D.; Sterne, J.A.C.; Champneys, A.; Gibbison, B. Use of machine learning to analyse routinely collected intensive care unit data: A systematic review. Crit. Care 2019, 23, 284. [Google Scholar] [CrossRef] [Scilit]
  16. Hong, N.; Liu, C.; Gao, J.; Han, L.; Chang, F.; Gong, M.; Su, L. State of the Art of Machine Learning–Enabled Clinical Decision Support in Intensive Care Units: Literature Review. JMIR Med. Inform. 2022, 10, e28781. [Google Scholar] [CrossRef] [Scilit]
  17. Syed, M.; Syed, S.; Sexton, K.; Syeda, H.B.; Garza, M.; Zozus, M.; Syed, F.; Begum, S.; Syed, A.U.; Sanford, J.; et al. Application of Machine Learning in Intensive Care Unit (ICU) Settings Using MIMIC Dataset: Systematic Review. Informatics 2021, 8, 16. [Google Scholar] [CrossRef] [Scilit]
  18. Vincent, J.L.; Moreno, R.; Takala, J.; Willatts, S.; De Mendonça, A.; Bruining, H.; Reinhart, C.K.; Suter, P.; Thijs, L.G. The SOFA (Sepsis-related Organ Failure Assessment) score to describe organ dysfunction/failure: On behalf of the Working Group on Sepsis-Related Problems of the European Society of Intensive Care Medicine (see contributors to the project in the appendix). Intensive Care Med. 1996, 22, 707–710. [Google Scholar]
  19. Lambden, S.; Laterre, P.F.; Levy, M.M.; Francois, B. The SOFA score—Development, utility and challenges of accurate assessment in clinical trials. Crit. Care 2019, 23, 374. [Google Scholar] [CrossRef] [Scilit]
  20. Le Gall, J.R.; Lemeshow, S.; Saulnier, F. A new simplified acute physiology score (SAPS II) based on a European/North American multicenter study. JAMA 1993, 270, 2957–2963. [Google Scholar] [CrossRef] [Scilit]
  21. Knaus, W.A.; Draper, E.A.; Wagner, D.P.; Zimmerman, J.E. APACHE II: A severity of disease classification system. Crit. Care Med. 1985, 13, 818–829. [Google Scholar] [CrossRef] [Scilit]
  22. Awad, A.; Bader-El-Den, M.; McNicholas, J.; Briggs, J. Early hospital mortality prediction of intensive care unit patients using an ensemble learning approach. Int. J. Med. Inform. 2017, 108, 185–195. [Google Scholar] [CrossRef] [Scilit]
  23. Du, H.; Ghassemi, M.M.; Feng, M. The effects of deep network topology on mortality prediction. In Proceedings of the 2016 38th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Orlando, FL, USA, 16–20 August 2016; pp. 2602–2605. [Google Scholar] [CrossRef] [Scilit]
  24. Hoogendoorn, M.; El Hassouni, A.; Mok, K.; Ghassemi, M.; Szolovits, P. Prediction using patient comparison vs. modeling: A case study for mortality prediction. In Proceedings of the 2016 38th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Orlando, FL, USA, 16–20 August 2016; pp. 2464–2467. [Google Scholar]
  25. García-Gallo, J.E.; Fonseca-Ruiz, N.J.; Celi, L.A.; Duitama-Muñoz, J.F. A machine learning-based model for 1-year mortality prediction in patients admitted to an Intensive Care Unit with a diagnosis of sepsis. Med. Intensiv. (Engl. Ed.) 2020, 44, 160–170. [Google Scholar] [CrossRef] [Scilit]
  26. Lin, K.; Hu, Y.; Kong, G. Predicting in-hospital mortality of patients with acute kidney injury in the ICU using random forest model. Int. J. Med. Inform. 2019, 125, 55–61. [Google Scholar] [CrossRef] [Scilit]
  27. Celi, L.A.; Galvin, S.; Davidzon, G.; Lee, J.; Scott, D.; Mark, R. A Database-driven Decision Support System: Customized Mortality Prediction. J. Pers. Med. 2012, 2, 138–148. [Google Scholar] [CrossRef] [Scilit]
  28. Kong, G.; Lin, K.; Hu, Y. Using machine learning methods to predict in-hospital mortality of sepsis patients in the ICU. BMC Med. Inform. Decis. Mak. 2020, 20, 251. [Google Scholar] [CrossRef] [Scilit]
  29. Xia, J.; Pan, S.; Zhu, M.; Cai, G.; Yan, M.; Su, Q.; Yan, J.; Ning, G. A Long Short-Term Memory Ensemble Approach for Improving the Outcome Prediction in Intensive Care Unit. Comput. Math. Methods Med. 2019, 2019, 8152713. [Google Scholar] [CrossRef] [Scilit]
  30. Rongali, S.; Rose, A.J.; McManus, D.D.; Bajracharya, A.S.; Kapoor, A.; Granillo, E.; Yu, H. Learning Latent Space Representations to Predict Patient Outcomes: Model Development and Validation. J. Med. Internet Res. 2020, 22, e16374. [Google Scholar] [CrossRef] [Scilit]
  31. Payrovnaziri, S.N.; Barrett, L.A.; Bis, D.; Bian, J.; He, Z. Enhancing Prediction Models for One-Year Mortality in Patients with Acute Myocardial Infarction and Post Myocardial Infarction Syndrome. In Studies in Health Technology and Informatics; IOS Press: Amsterdam, The Netherlands, 2019. [Google Scholar] [CrossRef] [Scilit]
  32. Ahmed, F.S.; Ali, L.; Joseph, B.A.; Ikram, A.; Ul Mustafa, R.; Bukhari, S.A.C. A statistically rigorous deep neural network approach to predict mortality in trauma patients admitted to the intensive care unit. J. Trauma Acute Care Surg. 2020, 89, 736. [Google Scholar] [CrossRef] [Scilit]
  33. Jentzer, J.C.; Anavekar, N.S.; Bennett, C.; Murphree, D.H.; Keegan, M.T.; Wiley, B.; Morrow, D.A.; Murphy, J.G.; Bell, M.R.; Barsness, G.W. Derivation and Validation of a Novel Cardiac Intensive Care Unit Admission Risk Score for Mortality. J. Am. Heart Assoc. 2019, 8, e013675. [Google Scholar] [CrossRef] [Scilit]
  34. Meyer, A.; Zverinski, D.; Pfahringer, B.; Kempfert, J.; Kuehne, T.; Sündermann, S.H.; Stamm, C.; Hofmann, T.; Falk, V.; Eickhoff, C. Machine learning for real-time prediction of complications in critical care: A retrospective study. Lancet Respir. Med. 2018, 6, 905–914. [Google Scholar] [CrossRef] [Scilit]
  35. Singer, M.; Deutschman, C.S.; Seymour, C.W.; Shankar-Hari, M.; Annane, D.; Bauer, M.; Bellomo, R.; Bernard, G.R.; Chiche, J.D.; Coopersmith, C.M.; et al. The third international consensus definitions for sepsis and septic shock (Sepsis-3). JAMA 2016, 315, 801–810. [Google Scholar] [CrossRef] [Scilit]
  36. Barbash, I.J.; Davis, B.S.; Yabes, J.G.; Seymour, C.W.; Angus, D.C.; Kahn, J.M. Treatment Patterns and Clinical Outcomes After the Introduction of the Medicare Sepsis Performance Measure (SEP-1). Ann. Intern. Med. 2021, 174, 927–935. [Google Scholar] [CrossRef] [Scilit]
  37. Bone, R.C.; Balk, R.A.; Cerra, F.B.; Dellinger, R.P.; Fein, A.M.; Knaus, W.A.; Schein, R.M.H.; Sibbald, W.J. Definitions for Sepsis and Organ Failure and Guidelines for the Use of Innovative Therapies in Sepsis. Chest 1992, 101, 1644–1655. [Google Scholar] [CrossRef] [Scilit]
  38. Seymour, C.W.; Liu, V.X.; Iwashyna, T.J.; Brunkhorst, F.M.; Rea, T.D.; Scherag, A.; Rubenfeld, G.; Kahn, J.M.; Shankar-Hari, M.; Singer, M.; et al. Assessment of Clinical Criteria for Sepsis: For the Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3). JAMA 2016, 315, 762–774. [Google Scholar] [CrossRef] [Scilit]
  39. Eickelberg, G.; Sanchez-Pinto, L.N.; Luo, Y. Predictive modeling of bacterial infections and antibiotic therapy needs in critically ill adults. J. Biomed. Inform. 2020, 109, 103540. [Google Scholar] [CrossRef] [Scilit]
  40. Desautels, T.; Calvert, J.; Hoffman, J.; Jay, M.; Kerem, Y.; Shieh, L.; Shimabukuro, D.; Chettipally, U.; Feldman, M.D.; Barton, C.; et al. Prediction of Sepsis in the Intensive Care Unit With Minimal Electronic Health Record Data: A Machine Learning Approach. JMIR Med. Inform. 2016, 4, e5909. [Google Scholar] [CrossRef] [Scilit]
  41. Mao, Q.; Jay, M.; Hoffman, J.L.; Calvert, J.; Barton, C.; Shimabukuro, D.; Shieh, L.; Chettipally, U.; Fletcher, G.; Kerem, Y.; et al. Multicentre validation of a sepsis prediction algorithm using only vital sign data in the emergency department, general ward and ICU. BMJ Open 2018, 8, e017833. [Google Scholar] [CrossRef] [Scilit]
  42. Luo, X.Q.; Yan, P.; Zhang, N.Y.; Luo, B.; Wang, M.; Deng, Y.H.; Wu, T.; Wu, X.; Liu, Q.; Wang, H.S.; et al. Machine learning for early discrimination between transient and persistent acute kidney injury in critically ill patients with sepsis. Sci. Rep. 2021, 11, 20269. [Google Scholar] [CrossRef] [Scilit]
  43. Ma, P.; Liu, J.; Shen, F.; Liao, X.; Xiu, M.; Zhao, H.; Zhao, M.; Xie, J.; Wang, P.; Huang, M.; et al. Individualized resuscitation strategy for septic shock formalized by finite mixture modeling and dynamic treatment regimen. Crit. Care 2021, 25, 243. [Google Scholar] [CrossRef] [Scilit]
  44. Bellomo, R.; Kellum, J.A.; Ronco, C. Acute kidney injury. Lancet 2012, 380, 756–766. [Google Scholar] [CrossRef] [Scilit]
  45. Zhang, Z.; Ho, K.M.; Hong, Y. Machine learning for the prediction of volume responsiveness in patients with oliguric acute kidney injury in critical care. Crit. Care 2019, 23, 112. [Google Scholar] [CrossRef] [Scilit]
  46. Zimmerman, L.P.; Reyfman, P.A.; Smith, A.D.R.; Zeng, Z.; Kho, A.; Sanchez-Pinto, L.N.; Luo, Y. Early prediction of acute kidney injury following ICU admission using a multivariate panel of physiological measurements. BMC Med. Inform. Decis. Mak. 2019, 19, 16. [Google Scholar] [CrossRef] [Scilit]
  47. Morid, M.; Sheng, O.; Del Fiol, G.; Facelli, J.; Bray, B.; Abdelrahman, S. Temporal Pattern Detection to Predict Adverse Events in Critical Care: Case Study with Acute Kidney Injury. Preprints 2019. [Google Scholar] [CrossRef] [Scilit]
  48. Alfieri, F.; Ancona, A.; Tripepi, G.; Randazzo, V.; Paviglianiti, A.; Pasero, E.; Vecchi, L.; Politi, C.; Cauda, V.; Fagugli, R.M. External validation of a deep-learning model to predict severe acute kidney injury based on urine output changes in critically ill patients. J. Nephrol. 2022, 35, 2047–2056. [Google Scholar] [CrossRef] [Scilit]
  49. Sun, M.; Baron, J.; Dighe, A.; Szolovits, P.; Wunderink, R.G.; Isakova, T.; Luo, Y. Early Prediction of Acute Kidney Injury in Critical Care Setting Using Clinical Notes and Structured Multivariate Physiological Measurements. In Studies in Health Technology and Informatics; IOS Press: Amsterdam, The Netherlands, 2019. [Google Scholar] [CrossRef] [Scilit]
  50. Wang, Y.; Wei, Y.; Yang, H.; Li, J.; Zhou, Y.; Wu, Q. Utilizing imbalanced electronic health records to predict acute kidney injury by ensemble learning and time series model. BMC Med. Inform. Decis. Mak. 2020, 20, 238. [Google Scholar] [CrossRef] [Scilit]
  51. Pang, H.; Kumar, S.; Ely, E.W.; Gezalian, M.M.; Lahiri, S. Acute kidney injury-associated delirium: A review of clinical and pathophysiological mechanisms. Crit. Care 2022, 26, 258. [Google Scholar] [CrossRef] [Scilit]
  52. Lai, C.F.; Wu, V.C.; Huang, T.M.; Yeh, Y.C.; Wang, K.C.; Han, Y.Y.; Lin, Y.F.; Jhuang, Y.J.; Chao, C.T.; Shiao, C.C.; et al. Kidney function decline after a non-dialysis-requiring acute kidney injury is associated with higher long-term mortality in critically ill survivors. Crit. Care 2012, 16, R123. [Google Scholar] [CrossRef] [Scilit]
  53. Zweck, E.; Kanwar, M.; Li, S.; Sinha, S.S.; Garan, A.R.; Hernandez-Montfort, J.; Zhang, Y.; Li, B.; Baca, P.; Dieng, F.; et al. Clinical Course of Patients in Cardiogenic Shock Stratified by Phenotype. JACC Heart Fail. 2023, 11, 1304–1315. [Google Scholar] [CrossRef] [Scilit]
  54. Soussi, S.; Tarvasmäki, T.; Kimmoun, A.; Ahmadiankalati, M.; Azibani, F.; Dos Santos, C.C.; Duarte, K.; Gayat, E.; Jentzer, J.C.; Harjola, V.P.; et al. Identifying biomarker-driven subphenotypes of cardiogenic shock: Analysis of prospective cohorts and randomized controlled trials. eClinicalMedicine 2025, 79, 103013. [Google Scholar] [CrossRef] [Scilit]
  55. McWilliams, C.J.; Lawson, D.J.; Santos-Rodriguez, R.; Gilchrist, I.D.; Champneys, A.; Gould, T.H.; Thomas, M.J.; Bourdeaux, C.P. Towards a decision support tool for intensive care discharge: Machine learning algorithm development using electronic healthcare data from MIMIC-III and Bristol, UK. BMJ Open 2019, 9, e025925. [Google Scholar] [CrossRef] [Scilit]
  56. Hammer, M.; Grabitz, S.D.; Teja, B.; Wongtangman, K.; Serrano, M.; Neves, S.; Siddiqui, S.; Xu, X.; Eikermann, M. A Tool to Predict Readmission to the Intensive Care Unit in Surgical Critical Care Patients—The RISC Score. J. Intensive Care Med. 2021, 36, 1296–1304. [Google Scholar] [CrossRef] [Scilit]
  57. Rojas, J.C.; Carey, K.A.; Edelson, D.P.; Venable, L.R.; Howell, M.D.; Churpek, M.M. Predicting Intensive Care Unit Readmission with Machine Learning Using Electronic Health Record Data. Ann. Am. Thorac. Soc. 2018, 15, 846–853. [Google Scholar] [CrossRef] [Scilit]
  58. Lin, Y.W.; Zhou, Y.; Faghri, F.; Shaw, M.J.; Campbell, R.H. Analysis and prediction of unplanned intensive care unit readmission using recurrent neural networks with long short-term memory. PLoS ONE 2019, 14, e0218942. [Google Scholar] [CrossRef] [Scilit]
  59. Zebin, T.; Chaussalet, T.J. Design and implementation of a deep recurrent model for prediction of readmission in urgent care using electronic health records. In Proceedings of the 2019 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB); IEEE: New York, NY, USA, 2019; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
  60. Khan, W.A.; Chung, S.H.; Awan, M.U.; Wen, X. Machine learning facilitated business intelligence (Part I) Neural networks learning algorithms and applications. Ind. Manag. Data Syst. 2020, 120, 164–195. [Google Scholar]
  61. Khan, W.A.; Chung, S.H.; Awan, M.U.; Wen, X. Machine learning facilitated business intelligence (Part II) Neural networks optimization techniques and applications. Ind. Manag. Data Syst. 2020, 120, 128–163. [Google Scholar]
  62. Stephens, A.F.; Šeman, M.; Diehl, A.; Pilcher, D.; Barbaro, R.P.; Brodie, D.; Pellegrino, V.; Kaye, D.M.; Gregory, S.D.; Hodgson, C.; et al. ECMO PAL: Using deep neural networks for survival prediction in venoarterial extracorporeal membrane oxygenation. Intensive Care Med. 2023, 49, 1090–1099. [Google Scholar] [CrossRef] [Scilit]
  63. Lundberg, S.M.; Lee, S.I. A unified approach to interpreting model predictions. Adv. Neural Inf. Process. Syst. 2017, 30. Available online: https://proceedings.neurips.cc/paper/2017/file/8a20a8621978632d76c43dfd28b67767-Paper.pdf (accessed on 21 September 2026).
  64. Hravnak, M.; DeVita, M.A.; Clontz, A.; Edwards, L.; Valenta, C.; Pinsky, M.R. Cardiorespiratory instability before and after implementing an integrated monitoring system. Crit. Care Med. 2011, 39, 65. [Google Scholar] [CrossRef] [Scilit]
  65. Tarassenko, L.; Hann, A.; Young, D. Integrated monitoring and analysis for early warning of patient deterioration. Br. J. Anaesth. 2006, 97, 64–68. [Google Scholar] [CrossRef] [Scilit]
  66. Schmidt, M.; Bailey, M.; Sheldrake, J.; Aubron, C.; Rycus, P.; Scheinkestel, C.; Cooper, D.; Brodie, D.; Pellegrino, V.; Combes, A.; et al. Predicting survival after ECMO for severe acute respiratory failure: The Respiratory ECMO Survival Prediction (RESP) Score. Am. J. Respir. Crit. Care Med. 2014, 189, 1374–1382. [Google Scholar] [CrossRef] [Scilit]
  67. Schmidt, M.; Burrell, A.; Roberts, L.; Bailey, M.; Sheldrake, J.; Rycus, P.T.; Hodgson, C.; Scheinkestel, C.; Cooper, D.J.; Thiagarajan, R.R.; et al. Predicting survival after ECMO for refractory cardiogenic shock: The survival after veno-arterial-ECMO (SAVE)-score. Eur. Heart J. 2015, 36, 2246–2256. [Google Scholar] [CrossRef] [Scilit]
  68. Amin, F.; Lombardi, J.; Alhussein, M.; Posada, J.D.; Suszko, A.; Koo, M.; Fan, E.; Ross, H.; Rao, V.; Alba, A.C.; et al. Predicting Survival After VA-ECMO for Refractory Cardiogenic Shock: Validating the SAVE Score. CJC Open 2021, 3, 71–81. [Google Scholar] [CrossRef] [Scilit]
  69. Hatib, F.; Jian, Z.; Buddi, S.; Lee, C.; Settels, J.; Sibert, K.; Rinehart, J.; Cannesson, M. Machine-learning Algorithm to Predict Hypotension Based on High-fidelity Arterial Pressure Waveform Analysis. Anesthesiology 2018, 129, 663–674. [Google Scholar] [CrossRef] [Scilit]
  70. Paradkar, N.; Roy Chowdhury, S. Coronary artery disease detection using photoplethysmography. In Proceedings of the 2017 39th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC); IEEE: New York, NY, USA, 2017; pp. 100–103. ISSN 1558-4615. [Google Scholar] [CrossRef] [Scilit]
  71. Guo, X.; Yin, Y.; Dong, C.; Yang, G.; Zhou, G. On the class imbalance problem. In Proceedings of the 2008 Fourth International Conference on Natural Computation; IEEE: New York, NY, USA, 2008; Volume 4, pp. 192–201. [Google Scholar]
  72. Tarawneh, A.S.; Hassanat, A.B.; Altarawneh, G.A.; Almuhaimeed, A. Stop oversampling for class imbalance learning: A review. IEEE Access 2022, 10, 47643–47660. [Google Scholar] [CrossRef] [Scilit]
  73. Alkhawaldeh, I.M.; Albalkhi, I.; Naswhan, A.J. Challenges and limitations of synthetic minority oversampling techniques in machine learning. World J. Methodol. 2023, 13, 373. [Google Scholar] [CrossRef] [Scilit]
  74. Demircioğlu, A. Applying oversampling before cross-validation will lead to high bias in radiomics. Sci. Rep. 2024, 14, 11563. [Google Scholar] [CrossRef] [Scilit]
  75. Rossetti, S.C.; Dykes, P.C.; Knaplund, C.; Cho, S.; Withall, J.; Lowenthal, G.; Albers, D.; Lee, R.Y.; Jia, H.; Bakken, S.; et al. Real-time surveillance system for patient deterioration: A pragmatic cluster-randomized controlled trial. Nat. Med. 2025, 31, 1895–1902. [Google Scholar] [CrossRef] [Scilit]
  76. Muñoz, J.; Fernández-Araujo, N.J.; Ruíz-Cacho, R.; Muñoz-Visedo, J. Artificial intelligence and computerized decision support in adult intensive care: A systematic review of randomized controlled trials. J. Crit. Care 2026, 94, 155600. [Google Scholar] [CrossRef] [Scilit]
  77. Shimabukuro, D.W.; Barton, C.W.; Feldman, M.D.; Mataraso, S.J.; Das, R. Effect of a machine learning-based severe sepsis prediction algorithm on patient survival and hospital length of stay: A randomised clinical trial. BMJ Open Respir. Res. 2017, 4, e000234. [Google Scholar] [CrossRef] [Scilit]
  78. Adams, R.; Henry, K.E.; Sridharan, A.; Soleimani, H.; Zhan, A.; Rawat, N.; Johnson, L.; Hager, D.N.; Cosgrove, S.E.; Markowski, A.; et al. Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nat. Med. 2022, 28, 1455–1460. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Sayood, S.S.; Sayood, K. Machine Learning Applications in the ICU: Opportunities and Pitfalls. Algorithms 2026, 19, 836. https://doi.org/10.3390/a19100836

AMA Style

Sayood SS, Sayood K. Machine Learning Applications in the ICU: Opportunities and Pitfalls. Algorithms. 2026; 19(10):836. https://doi.org/10.3390/a19100836

Chicago/Turabian Style

Sayood, Sinan S., and Khalid Sayood. 2026. "Machine Learning Applications in the ICU: Opportunities and Pitfalls" Algorithms 19, no. 10: 836. https://doi.org/10.3390/a19100836

APA Style

Sayood, S. S., & Sayood, K. (2026). Machine Learning Applications in the ICU: Opportunities and Pitfalls. Algorithms, 19(10), 836. https://doi.org/10.3390/a19100836

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop