All findings discussed in this section are model-generated outputs produced under specific parameterization assumptions. They should be interpreted as directional, simulation-derived evidence for hypothesis generation and protocol design guidance, and should not be interpreted as direct predictions of clinical outcomes in real ICU populations, where confounding factors, implementation variability, organizational context, and patient heterogeneity would influence the observed results in ways the present model does not capture.
4.1. Research Questions Addressed
This study evaluated three ICU nurse staffing SOPs across four demand scenarios using a Discrete Event Simulation model parameterized from 48,495 real ICU stays in MIMIC-IV. The results provide clear, statistically robust answers to the study’s three research questions.
RQ1: Which SOP configuration produces the best patient safety outcomes under normal operating conditions? Dynamic Escalation consistently produced the lowest adverse event rates across all conditions, achieving 4.260 adverse events per 100 patient days at baseline, a 16.8% reduction relative to the Fixed Ratio protocol (5.119) and an 18.4% reduction relative to the Acuity-Adjusted protocol (5.221). These differences were statistically significant (p < 0.0001), with moderate-to-large effect sizes (Dynamic Escalation vs. Fixed Ratio: r = −0.745; Dynamic Escalation vs. Acuity-Adjusted: r = −0.781).
However, this primary-analysis advantage reflects differences in total staffing capacity rather than in decision logic. A resource-constrained comparison (
Section 2.8,
Table 13) equalized maximum nurse hours across all three SOPs by reducing Dynamic Escalation’s base staffing to 8 day/6 night nurses, such that on-call activation restores capacity to the same 10 day/8 night level as Fixed Ratio and Acuity-Adjusted. Under this equal-resource design, Dynamic Escalation no longer outperforms the static protocols, it performs significantly worse than both Fixed Ratio (6.488 vs. 5.119 per 100 patient-days;
p < 0.0001, r = −0.897) and Acuity-Adjusted (6.488 vs. 5.221;
p < 0.0001, r = −0.855), which were themselves statistically indistinguishable from each other (
p = 0.270, r = 0.090).
This indicates that Dynamic Escalation’s entire primary-analysis advantage is attributable to its access to greater total staffing capacity, not to any independent benefit of workload-triggered escalation logic. If anything, the reactive design is a liability with equal resources: reducing baseline staffing and waiting for a workload threshold to be crossed before activating additional nurses leaves patients exposed to a period of elevated risk that static, adequately staffed protocols avoid entirely. The primary analysis (
Table 4) should therefore be interpreted as a comparison of complete staffing system configurations that differ in total available capacity, not as evidence that reactive, threshold-based escalation is a superior allocation strategy relative to static assignment when resources are held constant.
RQ2: How do SOP configurations perform under demand disruption? All three SOPs showed statistically significant differences in adverse event rates across every demand scenario (Kruskal–Wallis H = 104.1–122.7,
p < 0.0001;
Table 8), and each SOP’s own adverse event rate varied significantly across scenarios (Fixed Ratio H = 34.4; Acuity-Adjusted H = 26.9; Dynamic Escalation H = 30.1; all
p < 0.0001), indicating that none of the three protocols operates near a structural capacity ceiling under the modeled conditions (
Table 5). This ranking was stable across the joint sensitivity analysis when varying the escalation and overload thresholds simultaneously (
Table 11), as well as under activation-delay sensitivity testing (
Table 12), confirming that Dynamic Escalation’s advantage in the primary, unconstrained comparison is not an artifact of specific parameter or timing assumptions, although, as established in
Section 3.5, this advantage reflects staffing capacity rather than decision logic. Even under Combined Disruption, Dynamic Escalation’s adverse event rate (4.701 per 100 patient days) remained below Fixed Ratio’s own baseline rate (5.119), an 8.2% reduction.
RQ3: Does nurse workforce shortage or patient census surge pose a greater systemic risk? At equal (14-day) duration, this depends on the staffing protocol rather than holding as a general finding. Fixed Ratio and Acuity-Adjusted showed no reliable difference between the two disruption types (
Table 7), while Dynamic Escalation showed shortages caused modestly greater degradation than surges. This protocol-specific pattern is consistent with the mechanism by which each disruption acts: census surge increases patient volume while nurse capacity remains intact, whereas workforce shortage directly reduces the capacity that workload-responsive protocols depend on to absorb demand, a mechanism that only manifests where such responsive capacity exists. We therefore do not find support for a general claim that workforce shortage is more harmful than census surge across ICU staffing protocols, and restrict this conclusion to Dynamic Escalation specifically.
4.2. Practical Deployment Recommendations
The simulation results support the following directional, simulation-based observations for ICU staffing SOP design. All suggestions should be interpreted as directional guidance rather than prescriptive policy, as real-world implementation depends on organizational culture, staffing regulations, labor agreements, financial constraints, and workforce availability, all of which are factors not incorporated in the present simulation and which must be addressed through implementation science research before clinical translation. For example, mandatory nurse-to-patient ratio legislation in some jurisdictions may constrain the flexibility required for Dynamic Escalation’s threshold-based activation mechanism. Labor agreements governing on-call obligations and overtime compensation will directly affect the feasibility and cost of maintaining an available on-call nurse pool. These contextual factors mean that the optimal protocol configuration identified under simulated conditions may differ from the optimal configuration in any specific real-world institutional setting.
Simulation findings indicate that Dynamic Escalation’s apparent safety advantage over static protocols (16.8% relative to Fixed Ratio in the primary analysis) is attributable predominantly to its access to additional total nurse hours via on-call activation, rather than to a superior underlying allocation rule (
Section 3.5,
Table 13) [
18]. We therefore do not recommend Dynamic Escalation’s threshold-triggered decision logic on safety grounds independent of staffing capacity: its benefit is functionally equivalent to increasing total nurse staffing, and the reactive design does not outperform simply maintaining adequate staffing levels or an acuity-based allocation rule when total capacity is held constant. The on-call activation threshold of 85% workload capacity remains operationally straightforward to implement using existing NAS-based workload monitoring tools [
6] for institutions adopting this model for scheduling flexibility, independent of the decision-logic finding reported here.
Fixed Ratio protocols should be treated as a minimum floor, not a staffing target. The New York State 2023 ICU staffing rule requires a minimum 1:2 ratio, increased as appropriate for patient acuity [
4], yet the results demonstrate that a pure fixed-ratio implementation produces the highest adverse event rates among the three primary-analysis configurations tested, with a workload index of 0.764 at baseline (
Table 5), consistently elevated relative to the other protocols, though not near a structural capacity ceiling [
19].
It is worth noting that Law et al. [
20] found no detectable patient outcome improvement from acuity-tool-guided staffing ratios in a real-world hospital setting, which appears to contrast with this paper’s simulation findings. The authors themselves acknowledged multiple contributing factors, including that staffing levels may have been adequate prior to the mandate and that hospitals had significant leeway in acuity tool selection and deployment. This discrepancy underscores the gap between protocol design and real-world implementation, a limitation that simulation-based evaluation cannot address, and highlights the need for implementation science research alongside SOP design optimization.
Workforce retention should be treated as a patient safety intervention. The finding that nurse workforce shortage causes greater modeled performance degradation than census surge held only for Dynamic Escalation at equal (14-day) duration; Fixed Ratio and Acuity-Adjusted showed no reliable difference between the two disruption types (
Section 3.2,
Table 7). This suggests workforce retention may be a more consequential patient safety lever specifically where on-call or reserve-capacity staffing is used, rather than a general finding across all ICU staffing approaches [
21]. It should be noted, however, that this study evaluates shortage scenarios rather than directly evaluating retention interventions. The inference that workforce retention functions as a patient safety intervention therefore warrants direct empirical investigation rather than being treated as an established finding. Burnout, chronic fatigue, and occupational stress have been identified as primary drivers of nursing attrition [
22], and future work should examine whether interventions targeting these factors produce measurable improvements in safety-relevant staffing outcomes.
Bed overflow is a capacity problem, not a staffing problem. The finding that overflow counts were identical across all three SOPs within each scenario indicates that no staffing protocol, however well-designed, can resolve structural capacity shortfalls. ICUs experiencing high diversion rates require capacity expansion solutions alongside staffing optimization. This finding highlights a fundamental operational distinction: nurse staffing protocols and bed capacity management address different bottlenecks within the ICU system, and optimizing one cannot compensate for deficiencies in the other, a systems-level insight with direct implications for hospital capacity planning.
4.3. Limitations and Future Work
This study has several limitations that should be considered when interpreting the results. The most significant is the absence of external predictive validation. This study demonstrates two levels of model credibility: parameter verification: confirming that the implemented distributions accurately reproduce the MIMIC-IV source statistics (
Table 3); and internal consistency: demonstrated by stable output distributions across 100 replications (CV < 4%). What has not been demonstrated is external predictive validity, or the capacity of the model to reproduce patient flow behavior and safety outcomes in ICU settings other than BIDMC. Parameterization fidelity was assessed against the same MIMIC-IV data used to derive model inputs, which confirms implementation correctness but does not constitute validation in independent settings. Researchers applying these findings to other ICU contexts should treat the results as directional, simulation-derived evidence rather than transferable predictions.
Second, nursing task demand rates and adverse event probability thresholds were calibrated from the NAS-based literature rather than extracted directly from MIMIC-IV. As a result, absolute adverse event rates reported in this study are conditional on these assumed constants and should not be interpreted as estimates of actual clinical adverse event frequencies. Conclusions regarding the relative ranking of SOP configurations are more robust to this uncertainty, as confirmed by the ±25% sensitivity analysis in
Section 3.4, which demonstrates that the SOP ranking is invariant across all tested parameter variations.
Third, patient acuity in the model is assigned probabilistically at arrival based on the empirical distribution of the first-24 h maximum SOFA scores observed in the MIMIC-IV cohort. In real ICU operations, this value is not prospectively available at the moment of admission; it is only known retrospectively after 24 h of clinical observation. The Acuity-Adjusted and Dynamic Escalation protocols therefore operate in the model with idealized advance knowledge of eventual patient acuity, which may overstate the precision achievable by acuity-based staffing rules at the point of admission in real clinical settings, where initial triage relies on presenting severity indicators rather than a retrospective summary score.
Fourth, the model does not capture several important human and organizational factors including nurse experience heterogeneity, skill mix variation, cumulative fatigue accumulation, teamwork quality, communication patterns, and broader organizational culture. These factors are likely to both amplify the adverse effects of prolonged shortage scenarios and moderate the benefits of adaptive staffing protocols in ways the present model cannot reproduce. Their omission means the model represents an idealized version of ICU nursing operations and should not be expected to capture the full complexity of real clinical environments. More specifically, the adverse event generation mechanism reduces a multifactorial clinical phenomenon to a binary Bernoulli process conditioned solely on whether the unit-level workload index exceeds a fixed threshold. This abstraction, while necessary for model tractability and consistent with the shift-level measurement scale of the NAS instrument, cannot capture the within-shift temporal dynamics of error accumulation, the differential vulnerability of patients at different acuity levels to nurse-sensitive harm, or the protective effects of teamwork and communication quality. Additionally, the workload–risk relationship is modeled as a binary threshold, elevated adverse-event probability applies uniformly once workload index exceeds 0.88, with no risk gradient below that point, rather than as a continuous function of workload, creating an artificial discontinuity at the threshold boundary that may not reflect the true, likely more gradual, relationship between nursing workload and adverse-event risk.
Absolute adverse event rates should therefore be interpreted as relative indicators of workload-driven risk rather than as estimates of actual clinical event frequencies.
Similarly, the workload index captures direct patient care task demand only and does not account for the substantial indirect workload components of ICU nursing practice, including documentation burden, family communication, medication preparation, mentoring of junior staff, participation in multidisciplinary rounds, and response to unexpected clinical deterioration or simultaneous emergencies. Real ICU nursing workload is therefore substantially higher than the model represents, meaning the model likely underestimates the cognitive load experienced by nurses, particularly under shortage conditions, and may consequently understate the safety consequences of staffing shortfalls.
Fifth, the on-call nurse activation mechanism in SOP 3 assumes instantaneous response, which may overestimate real-world effectiveness where call-in latency introduces delays. A sensitivity analysis introducing activation delays of 0.5, 1.0, and 2.0 h (
Section 3.4) demonstrates that Dynamic Escalation maintains its advantage over Fixed Ratio even at a 2 h delay (4.497 vs. 5.119 adverse events per 100 patient days; −12.1%), though the magnitude of advantage narrows with increasing latency. In practice, the effectiveness of Dynamic Escalation will depend on the availability of a sufficiently large and geographically accessible on-call nurse pool, which varies across institutions.
Sixth, the simulation model was parameterized from a single academic medical center (BIDMC, Boston), and the MIMIC-IV-derived patient flow parameters, including arrival rates, LOS distributions, and acuity proportions, may not generalize to community hospitals, rural ICUs, or non-US healthcare systems with different patient mix and operational structures. Multi-site parameterization using data from diverse ICU types is recommended as a priority for future work.
Beyond geographic and institutional generalizability, the fixed 20-bed configuration represents a median adult MICU size and may not reflect the operational dynamics of smaller community ICUs (typically 8–12 beds), where fixed staffing ratios represent a larger proportion of total available capacity, or larger academic units (30–40 beds), where economies of scale may alter the relative advantage of adaptive protocols. Specialty ICUs, including cardiac, neonatal, neurological, and surgical units, have substantially different acuity profiles, task demand rates, and nurse-to-patient ratio norms that would require unit-specific parameterization before the present findings could be applied. The fixed shift pattern (day 07:00–18:00, night 18:00–07:00) similarly reflects BIDMC’s operational structure and may not generalize to ICUs using 12 h shifts or flexible scheduling arrangements.
Future work should extend this framework in several directions. Multi-site parameterization using data from diverse ICU types, including community hospitals, rural ICUs, and specialty units, would improve generalizability and address the single-site limitation of the present study. Incorporating nurse fatigue accumulation as a time-varying parameter would improve the realism of shortage scenario modeling, particularly for extended workforce shortage conditions. Economic analysis quantifying the cost per adverse event avoided under each SOP configuration, accounting for the incremental on-call staffing hours associated with Dynamic Escalation, would provide the cost-effectiveness framework necessary to support adoption decisions. A full probabilistic sensitivity analysis varying all model inputs simultaneously, including arrival distributions, LOS distributions, acuity proportions, and staffing levels, would provide a more comprehensive characterization of model uncertainty than the targeted parameter analysis conducted here. Finally, integration with real-time NAS-based workload monitoring systems would enable prospective, rather than retrospective, SOP evaluation as part of clinical operations.