3.1. Baseline Characteristics of the Study Population
A total of 194 patients were enrolled in the study, stratified by disease severity into mild (
n = 59, 30.42%), moderately severe (
n = 82, 42.3%), and severe (
n = 53, 24.2%) forms, based on the Revised Atlanta Classification. During the analysis, binary endpoints were applied; consequently, mild (MILD) cases were compared with the combined group of moderately severe/severe cases (MODSEV) (
Figure 2).
The baseline demographic, clinical and anamnestic, hemodynamic, nutritional, and laboratory characteristics of the study population are summarized in
Table 1. Additional comorbidity, endoscopy, medication, etiology, and toxic-habit characteristics are summarized in
Supplementary Table S1.
Comparing the demographic aspects of the two groups, there were no significant differences in age (54.68 ± 16.32 vs. 57.30 ± 16.64 years; p = 0.270) or in gender, with a similar proportion of males (62.7% vs. 65.2%). This is likely due to the small sample size.
The etiological distribution of acute pancreatitis was comparable between the two groups (p = 0.849). Biliary etiology was the leading cause in MILD (39.0%) and MODSEV (37.0%) patients, followed by alcohol-induced pancreatitis (25.4% vs. 24.4%), metabolic-associated pancreatitis (20.3% vs. 25.9%), and other causes (15.3% vs. 12.6%). The absence of significant between-group differences in etiology suggests that disease progression in this cohort is not primarily etiology-dependent.
Similarly, the distribution of toxic habits did not differ significantly between groups (p = 0.408). Combined alcohol and tobacco use was numerically more prevalent in the MODSEV group (31.9% vs. 20.3%), while isolated alcohol and smoking exposure were approximately the same. Although statistical significance was not achieved, the trend towards higher combined toxic exposure in the MODSEV group warrants further investigation, given the established synergistic role of alcohol and tobacco in pancreatic injury.
Hypertension was the most prevalent cardiovascular comorbidity and was significantly more frequent in the MODSEV group (68.1% vs. 52.5%; p = 0.038). A non-significant trend toward a greater burden of chronic ischemic cardiopathy was observed in MODSEV patients (16.3% vs. 6.8%; p = 0.073). Other cardiovascular conditions, including atrial fibrillation (7.4% vs. 3.4%; p = 0.285), acute heart failure (12.6% vs. 8.5%; p = 0.405), valvular disease (9.6% vs. 11.9%; p = 0.638), deep vein thrombosis (6.7% vs. 1.7%; p = 0.150), and prior acute myocardial infarction (3.7% vs. 1.7%; p = 0.457), were numerically more common in the MODSEV group but did not reach statistical significance.
Diabetes mellitus was present in 26.7% of MODSEV patients compared to 15.3% in the MILD group, with a trend toward significance (p = 0.083). Similarly, chronic kidney disease (CKD) was more prevalent in the MODSEV group (17.0% vs. 6.8%; p = 0.058), approaching statistical significance. These findings are consistent with the fact that metabolic comorbidities increase the risk of more severe disease.
Liver steatosis was highly prevalent in both groups, affecting 81.5% of MODSEV and 69.5% of MILD patients (p = 0.064). Fibrosis risk, assessed by the FIB-4 score, did not differ significantly between groups (p = 0.501): high-risk fibrosis was identified in 50.8% of MILD and 48.9% of MODSEV patients, indeterminate risk in 18.6% and 25.9%, and low risk in 30.5% and 25.2%, respectively. The high burden of advanced fibrosis across both groups highlights the importance of metabolic liver disease as a comorbidity in this population.
The prevalence of chronic respiratory disease (COPD and asthma) was similar between the two groups (14.8% vs. 13.6%; p = 0.819). However, pleural effusion was significantly more frequent in the MODSEV group (43.7% vs. 13.6%; p < 0.001), reflecting higher systemic inflammation and fluid redistribution. Systemic inflammatory response syndrome (SIRS) was recorded in a significantly higher proportion of MODSEV patients (64.9% vs. 43.1%; p = 0.005). ARDS was numerically more frequent in MODSEV patients (11.9% vs. 3.4%; p = 0.062), although without reaching statistical significance, probably due to limited sample size.
Upper gastrointestinal endoscopy was performed in 37.3% of MILD and 26.7% of MODSEV patients (p = 0.137) during the hospital stay. Gastritis was the most frequently identified lesion, numerically more common in the MILD group (27.1% vs. 15.6%; p = 0.059). Statistically significant differences were observed for lesions described at the level of the cardia (6.8% vs. 0.7%; p = 0.015), fundus (6.8% vs. 0%; p = 0.002), and duodenal second portion (6.8% vs. 0.7%; p = 0.015), all more frequent in the MILD group. Endoscopically identified tumors reached the threshold for statistical significance (5.1% vs. 0.7%; p = 0.050). Esophagitis (15.3% vs. 10.4%; p = 0.333), erosions, gastropathy, ulcers, and Mallory-Weiss lesions were comparable between groups. Proton pump inhibitor therapy was administered in 74.6% of MILD and 83.0% of MODSEV patients (p = 0.176), without an apparent impact on disease severity.
All measured inflammatory markers were significantly elevated in the MODSEV group at admission. White blood cell count was higher in MODSEV patients (15.11 ± 6.26 vs. 11.09 ± 4.11 × 103/µL; p < 0.001), as was C-reactive protein (120.57 ± 126.97 vs. 56.58 ± 72.39 mg/L; p < 0.001). The neutrophil-to-lymphocyte ratio (NLR) at admission was more than twice as high in the MODSEV group (14.91 ± 13.08 vs. 6.69 ± 3.99; p < 0.001), and this elevation persisted at 24 h (12.88 ± 11.73 vs. 5.28 ± 3.60; p < 0.001) and 48 h (10.75 ± 9.66 vs. 4.06 ± 2.20; p < 0.001), indicating more pronounced and persistent systemic inflammation. Platelet-to-lymphocyte ratio (PLR) was similarly elevated at all three time points in the MODSEV group (all p ≤ 0.001).
Hematocrit levels at admission and at 24 h did not differ significantly between the groups (p = 0.233 and p = 0.624, respectively). Platelet counts were also comparable (228.33 ± 89.86 vs. 222.77 ± 89.99 × 103/µL; p = 0.785). In contrast, lymphocyte counts at admission were significantly lower in the MODSEV group (1286 ± 997.63 vs. 1427.95 ± 611.92 × 103/µL; p = 0.008), consistent with the elevated NLR observed in this group.
Nutritional impairment was significantly more prevalent in the MODSEV group. Serum albumin at admission was lower in MODSEV patients (3.46 ± 0.55 vs. 3.88 ± 0.37 g/dL; p < 0.001), as was the total CONUT score (4.32 ± 2.90 vs. 2.73 ± 2.02; p = 0.001). Categorical analysis of the CONUT score revealed a significantly worse nutritional profile in MODSEV patients (p = 0.001): normal nutritional status was present in only 14.8% of MODSEV patients versus 30.5% of MILD patients, while moderate-to-severe malnutrition was identified in 40.0% of MODSEV patients compared to 13.6% in the MILD group. Total cholesterol levels did not differ significantly between groups (169.49 ± 68.51 vs. 176.32 ± 76.32 mg/dL; p = 0.748).
Blood urea nitrogen (BUN) at admission was significantly elevated in the MODSEV group (23.68 ± 18.05 vs. 16.10 ± 8.72 mg/dL; p = 0.001), indicating greater renal stress in more severe cases. Serum creatinine was paradoxically higher in the MILD group (1.88 ± 7.84 vs. 1.15 ± 0.73 mg/dL; p = 0.004), likely attributable to outlier values reflected in the high standard deviation of that group. Hepatic transaminases such as AST (177.64 ± 219.53 vs. 168.38 ± 239.16 U/L; p = 0.411) and ALT (153.68 ± 204.09 vs. 137.34 ± 197.76 U/L; p = 0.454) did not differ significantly between groups.
Serum amylase at admission was comparably elevated in both groups (MILD: 1048.12 ± 1501.69 U/L; MODSEV: 1019.69 ± 1135.67 U/L; p = 0.404), confirming uniform elevation as a diagnostic feature irrespective of disease severity. In contrast, glycemic parameters were significantly worse in the MODSEV group: both admission blood glucose (170.39 ± 118.33 vs. 131.71 ± 57.35 mg/dL; p = 0.003) and peak glucose during hospitalization (187.48 ± 140.26 vs. 134.55 ± 58.32 mg/dL; p < 0.001) were substantially higher, indicating more pronounced metabolic dysregulation in severe cases.
Vital signs at admission, including body temperature, heart rate, respiratory rate, and systolic and diastolic blood pressure, did not differ significantly between the groups (p > 0.05). Blood pressure category distribution was also similar (p = 0.740), with Grade II hypertension being the most common category in both groups (28.8% MILD vs. 31.1% MODSEV).
Disease severity was strongly associated with adverse clinical outcomes. Patients in the MODSEV group were hospitalized significantly longer than those in the mild group (9.61 ± 5.25 vs. 6.97 ± 1.99 days; p < 0.001), indicating increased healthcare costs and higher severity. In-hospital mortality occurred exclusively among MODSEV patients, with 14 deaths (11.57%), whereas no deaths were recorded in the mild group (p = 0.01).
Complications also differed markedly between groups. Uncomplicated cases were more frequent in the mild group (20/59, 33.9%) compared with the MODSEV group (26/135, 19.3%). Peripancreatic fluid collections were the most common complication in both groups but were substantially more frequent in mild patients numerically (33/59, 55.9% vs. 49/135, 36.3%). Necrotic collections showed the most pronounced disparity, occurring predominantly in the MODSEV group (51/135 cases, 37.77%) and being rare in mild disease (3/54, 5.6%), highlighting a strong association with the moderately severe/severe (MODSEV) outcome. Pseudocysts were observed only in the MODSEV group (6/135, 4.4%), while de novo diabetes occurred at a similarly low frequency in both groups (3/59, 5.1% in mild vs. 3/135, 2.2% in MODSEV). Overall, complication patterns demonstrated a clear shift toward more structurally severe pancreatic sequelae in the MODSEV group, particularly necrotic collections, reflecting greater disease severity and systemic involvement.
Detailed comorbidity, endoscopic, medication, etiology, and toxic-habit data for the study population—including cardiovascular, respiratory, metabolic, renal, hepatic, and neurological comorbidities, and upper gastrointestinal endoscopy findings—are provided in the
Supplementary Materials (Table S1). Of these, pleural effusion (43.7% vs. 13.6%,
p < 0.001) and SIRS (64.9% vs. 43.1%,
p = 0.005) are of direct relevance to ProbScore2 and are discussed further below; hypertension (68.1% vs. 52.5%,
p = 0.038) was also significantly more frequent in the MODSEV group.
These findings demonstrate that MODSEV patients present with a significantly more severe clinical, inflammatory, nutritional, and metabolic profile than MILD patients. The most robust discriminators between severity groups at admission were inflammatory markers (NLR, PLR, WBC, and CRP), nutritional indices (albumin, CONUT score), glycemic status, renal stress indices (BUN, serum creatinine), complication burden (SIRS, pleural effusion), lymphocyte count, hospitalization length, and the presence of hypertension in the medical history (
Figure S1).
Boxplots comparing the MILD and MODSEV groups across the significant admission discriminators identified in this analysis are provided in the
Supplementary Materials (Figure S1).
3.2. Comparison with Established Severity Scoring Systems
We evaluated six established severity scoring systems (Ranson, Glasgow-Imrie, Balthazar, CTSI, BISAP, and HAPS) for predicting moderately severe/severe acute pancreatitis. Diagnostic accuracy was compared by calculating the area under the receiver operating characteristic curve (AUC) for each system (
Figure 3).
The Glasgow-Imrie score demonstrated the highest discriminative performance among all evaluated conventional scoring systems, with an AUC of 0.832 (95% CI: 0.765–0.883, p < 0.001), indicating good diagnostic accuracy; at the optimal Youden Index threshold, its sensitivity was 74.6% and specificity was 74.4%. The CTSI score achieved the second highest AUC of 0.775 (95% CI: 0.707–0.842, p < 0.001), with an optimal sensitivity of 70.2% and specificity of 70.4%. The Balthazar score followed closely, with an AUC of 0.758 (95% CI: 0.685–0.829, p < 0.001) and a sensitivity/specificity of 68.9%/68.8% at the optimal cut-off; these two radiological scores are complementary and are commonly used together in clinical practice.
The BISAP score showed moderate discriminative accuracy (AUC = 0.725, 95% CI: 0.640–0.792, p < 0.001; sensitivity 65.5%, specificity 65.8%), followed by the Ranson score (AUC = 0.676, 95% CI: 0.584–0.752, p < 0.001; sensitivity 61.8%, specificity 62.3%).
Among the evaluated scoring systems, the HAPS demonstrated the lowest discriminative performance, with an AUC of 0.574 (95% CI: 0.497–0.672), which did not reach statistical significance (p = 0.064), suggesting limited utility for predicting disease severity in the present cohort. This may be explained by the absence of peritonitis in the study population, a key component of this scoring system.
3.3. Development of the New Prognostic Score (ProbScore2)
To develop the new prognostic score, 12 parameters with robust discriminatory value were selected from the previously identified candidate variables. The inclusion criteria were determined subjectively, with the primary consideration being the routine availability of these parameters at hospital admission. The selected variables were: C-reactive protein (CRP), white blood cell count (WBC), presence of systemic inflammatory response syndrome (SIRS), blood urea nitrogen (BUN), history of hypertension (HTN), admission blood glucose level, presence of pleural effusion, neutrophil-to-lymphocyte ratio (NLR), lymphocyte count, serum albumin, platelet-to-lymphocyte ratio (PLR), and serum creatinine. This pre-selection reflected practical clinical availability at admission (prioritizing routine bloodwork, vital signs, and a chest radiograph over specialized or delayed investigations) combined with variables previously reported as significant discriminators of AP severity in the univariable analysis described above (
Section 3.2) or in the existing literature, rather than an unrestricted, purely data-driven search across all measured parameters. This strategy was intended to reduce the risk of chance findings from testing a very large number of candidate variables in a moderate-sized single-center cohort, at the cost of an element of subjectivity in the initial candidate list; we address this trade-off further below with a penalized-regression sensitivity analysis.
Missing data among the 12 candidate variables were minimal: platelet-to-lymphocyte ratio was missing in 3 patients (1.5%), SIRS status in 2 patients (1.0%), and CRP in 1 patient (0.5%). All other candidate variables (white blood cell count, BUN, hypertension, admission glucose, pleural effusion, NLR, lymphocyte count, albumin, and creatinine) were complete for all 194 patients. Overall, 6 of 194 patients (3.1%) had at least one missing value among the 12 candidates and were excluded via listwise deletion from the initial backward stepwise selection model (n = 188). The final six-variable model was subsequently evaluated in all patients with complete data for the six retained predictors (n = 191).
Among the 12 candidate variables, independent predictors were identified using binary logistic regression with a backward stepwise, likelihood-ratio elimination procedure (entry criterion: p < 0.05; removal criterion: p > 0.10; classification cut-off = 0.5). The initial (full) model included all 12 variables, with disease severity (mild vs. moderately severe/severe) as the dependent variable.
After seven elimination steps, the following variables were sequentially removed based on their minimal contribution to the model: NLR (step 2), SIRS (step 3), lymphocyte count (step 4), creatinine (step 5), hypertension (step 6), and CRP (step 7). Model performance remained stable throughout the elimination process (Nagelkerke R2 decreased only from 0.514 to 0.496), indicating that removal of these variables did not substantially reduce the explanatory power of the model.
Following backward elimination, six variables remained in the final model: white blood cell count, BUN, pleural effusion, admission glucose, platelet-to-lymphocyte ratio at admission, and serum albumin. These variables were subsequently re-evaluated using a confirmatory logistic regression model with the ENTER method, yielding the final coefficients shown in
Table 2.
As a sensitivity analysis addressing the potential instability of stepwise variable selection in a moderate-sized cohort, we additionally fitted an L1-penalized (LASSO) logistic regression model across the same 12 candidate variables, with the penalty strength selected via 10-fold cross-validation using the one-standard-error rule. This penalized model retained six variables (white blood cell count, CRP, hypertension, pleural effusion, neutrophil-to-lymphocyte ratio, and albumin; in-sample AUC = 0.850); three of these (white blood cell count, pleural effusion, and albumin) were identical to those retained by backward stepwise selection, while the remaining three differed (LASSO retained CRP, hypertension, and NLR, whereas backward stepwise retained BUN, admission glucose, and PLR), most likely reflecting collinearity among inflammatory and metabolic markers measured at admission. Despite this partial divergence in the selected variables, both approaches achieved similar discriminative performance, and the bootstrap-based internal validation reported in
Section 2.5 already provides an optimism-corrected estimate of the final model’s performance that does not depend on which specific variable-selection method was used.
The final model demonstrated good calibration (Hosmer–Lemeshow test: χ
2 = 4.87, df = 8,
p = 0.771), with a Nagelkerke R
2 of 0.478 and an overall classification accuracy of 82.7%. Because the Hosmer–Lemeshow test alone is an insufficient basis for assessing calibration, a more comprehensive evaluation—including a calibration plot, calibration slope and intercept, Brier score, and decision curve analysis—is reported later in the Results (
Section 3.5).
To allow the model to be applied directly to individual patients, the full fitted logistic regression equation is provided here. The predicted logit (log-odds) of moderately severe/severe acute pancreatitis (MODSEV) is logit(p) = 3.083 + 0.1776 × (white blood cell count, in 1000 cells/µL) + 0.0469 × (BUN, mg/dL) + 0.947 × (pleural effusion: 1 = present, 0 = absent) + 0.0098 × (admission glucose, mg/dL) + 0.0037 × (PLR at admission) − 2.068 × (serum albumin, g/dL). The predicted probability is then obtained as p = 1/(1 + e − logit(p)). This full formula, together with the simplified 0–6 point bedside score described below, allows the model to be applied at the individual patient level using either electronic calculation (full formula) or a manual bedside count (simplified score).
Serum albumin (lower values indicating greater disease severity; OR = 0.12), glucose level at admission, and blood urea nitrogen emerged as the strongest statistically significant predictors (p < 0.05). Pleural effusion (p = 0.058) and platelet-to-lymphocyte ratio at admission (p = 0.070) did not reach the conventional p < 0.05 threshold for statistical significance in the final model and should not be interpreted as independent statistically significant predictors on their own. They were retained because they satisfied the pre-specified removal criterion of the backward stepwise procedure (p > 0.10), which was defined before model fitting, and because of their established clinical relevance to acute pancreatitis severity. To assess whether their retention was justified beyond this pre-specified statistical rule, we compared the full six-variable model with a reduced four-variable model excluding both pleural effusion and PLR: the two variables jointly contributed significantly to model fit (likelihood ratio test, χ2 = 8.98, df = 2, p = 0.011) and their inclusion improved both discrimination (AUC 0.873 vs. 0.852) and model fit (AIC 168.5 vs. 173.5) relative to the reduced model, indicating that although neither variable was independently significant, their joint contribution to the model was not attributable to chance.
The discriminative performance of the new six-variable score (ProbScore2) was evaluated using ROC analysis (
Table 3). ProbScore2 alone demonstrated excellent discriminative ability for predicting moderately severe/severe acute pancreatitis.
An area under the ROC curve of 0.873 indicates excellent discriminatory performance, according to the widely accepted classification in which an AUC > 0.80 represents excellent diagnostic accuracy.
To quantify the degree of overfitting expected from same-cohort development and evaluation, internal validation of ProbScore2 was performed using bootstrap resampling. Across 1000 bootstrap replicates, the mean optimism in the AUC was 0.020, yielding an optimism-corrected AUC of 0.856. This estimate closely agreed with the AUC obtained from repeated (100×) 10-fold cross-validation (mean AUC = 0.853, SD = 0.009; 2.5th–97.5th percentile range: 0.836–0.870). The small difference between the apparent (0.873) and internally validated (0.853–0.856) AUC values indicates only modest overfitting and supports the internal robustness of ProbScore2, although these internal estimates cannot substitute for external validation in an independent cohort.
For each continuous variable included in the final model (white blood cell count, BUN, admission glucose, PLR, and serum albumin), optimal, clinically applicable cut-off values were determined using ROC analysis by maximizing the Youden Index (sensitivity − [1 − specificity]) (
Table 4).
The discriminatory performance of the new score was directly compared with the established severity scoring systems described above (
Table 5,
Figure 4).
Within this cohort, ProbScore2 achieved a numerically higher AUC (0.873) than the Glasgow-Imrie score (0.832), CTSI (0.775), Balthazar score (0.758), BISAP (0.725), Ranson (0.676), and HAPS (0.574); these comparisons were made within the same retrospective, single-center sample used for model development and were not adjusted for multiple testing. This numerical difference was particularly pronounced when comparing ProbScore2 to the traditionally less discriminative clinical scoring systems, including HAPS, Ranson, and BISAP, within this cohort.
3.4. Performance of the Simplified Bedside Score (0–6 Points)
Based on the threshold values shown in
Table 4, each of the six variables can be assigned a binary score (0 or 1), allowing the construction of a simple, unweighted bedside clinical scoring system alongside the original ProbScore2 logistic regression formula. One point was assigned for each of the six dichotomized criteria (total score range 0–6) for all patients in the cohort (
n = 194), and diagnostic accuracy was evaluated using ROC analysis (
Table 6).
The results demonstrated a clear, stepwise association between the score and the probability of a moderately severe/severe (MODSEV) outcome: 15.4% of patients with a score of 0 were classified as moderately severe/severe (MODSEV), whereas this proportion increased to 100% among those with a score ≥5.
The simplified bedside score demonstrated excellent discriminatory ability, with an AUC of 0.852 (95% CI: 0.797–0.902, bootstrap estimate). The optimal threshold, based on the maximum Youden Index (0.567), was ≥3 points, corresponding to a sensitivity of 77.0% and a specificity of 79.7% (
Table 7,
Figure 5).
The 95% CI reported above reflects sampling uncertainty around the apparent AUC but does not account for the fact that the six cut-off values (
Table 4) were themselves derived by maximizing the Youden Index in the same cohort subsequently used to evaluate the simplified score. This circularity can inflate apparent performance. To address this explicitly, we repeated the full cut-off derivation and scoring procedure within a bootstrap resampling framework (1000 replicates): in each replicate, Youden-optimal cut-offs were re-derived from the bootstrap sample, the simplified score was recomputed accordingly, and its AUC was evaluated both in the bootstrap sample and in the original cohort. This yielded a mean optimism of 0.026 and an optimism-corrected AUC of 0.827 for the simplified bedside score, compared with the apparent AUC of 0.853—a larger correction than that observed for the full logistic regression model (apparent 0.873 vs. internally validated 0.853–0.856), consistent with the additional optimism introduced by data-driven cut-off selection. The individual cut-offs varied in their bootstrap stability: those for white blood cell count, admission glucose, PLR, and albumin were relatively stable (bootstrap median equal to or close to the apparent cut-off), whereas the BUN cut-off was less stable (apparent 15.0 mg/dL; bootstrap median 18.6 mg/dL; 95% range 15.0–23.0 mg/dL), indicating that this particular threshold should be interpreted with caution until confirmed in an independent cohort.
When evaluated in the same cohort of 194 patients, the simplified score achieved an AUC of 0.852, only marginally lower than that of the full logistic regression model (ProbScore2; AUC = 0.873), indicating that the unweighted additive score preserves most of the predictive performance of the original model.
Overall, fulfillment of at least three of the six criteria (≥3/6)—white blood cell count >12.72 ×10
3/µL, BUN >15.0 mg/dL, admission glucose >140.5 mg/dL, PLR >219, serum albumin <3.71 g/dL, and/or presence of pleural effusion—represents the optimal statistical threshold for identifying patients at risk of developing moderately severe/severe acute pancreatitis. At this cut-off, approximately 77% of patients with moderately severe/severe acute pancreatitis (MODSEV) are correctly identified, while approximately 80% of patients with mild disease are correctly excluded. Higher scores (≥4–5) provide greater specificity (96.6–100%) and may therefore be useful for very high-risk patients, although their reduced sensitivity makes them less suitable as primary decision thresholds (
Figure 6).
Figure 6.
Schematic overview of the simplified ProbScore2 bedside score as applied in routine clinical practice.
Figure 6.
Schematic overview of the simplified ProbScore2 bedside score as applied in routine clinical practice.