Next Article in Journal
PANDA-PLUS-Bench: A Clinical Benchmark for Evaluating the Robustness of AI Foundation Models in Prostate Cancer Diagnosis
Next Article in Special Issue
Explainable Artificial Intelligence for Predicting Gastrointestinal Adverse Effects of GLP-1 Receptor Agonists
Previous Article in Journal
Explainable AI-Driven Identification of Multimodal Biomarkers for Early Prediction of Cognitive Decline
Previous Article in Special Issue
MedROAD V2: An AI-Integrated Electronic Medical Record System with Advanced Clinical Decision Support
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Operationalizing Instability in Rule-Based Complete Blood Count Phenotyping Using Uncertainty-Aware Machine Learning

Medical Faculty Mannheim, Institute for Clinical Chemistry, University of Heidelberg, 68167 Mannheim, Germany
*
Author to whom correspondence should be addressed.
AI Med. 2026, 1(2), 13; https://doi.org/10.3390/aimed1020013
Submission received: 2 April 2026 / Revised: 15 May 2026 / Accepted: 20 May 2026 / Published: 22 May 2026
(This article belongs to the Special Issue Machine Learning Applications for Risk Stratification in Healthcare)

Abstract

Background: Complete blood count (CBC) phenotypes are routinely assigned using deterministic rule-based thresholds. While operationally efficient, such rules may lead to unstable phenotype assignments for results close to clinical cutoffs in the presence of analytical variability. Methods: We analyzed routine CBC data from a tertiary care hospital laboratory. Rule-based phenotypes for anemia subtype, white blood cell (WBC) status, and platelet (PLT) status were assigned using established laboratory thresholds. A patient-independent development and holdout split was applied. A multi-output gradient boosting model was trained to reproduce rule-based labels and provide probabilistic outputs. Phenotype stability was assessed by perturbing CBC parameters under realistic analytical noise. Instability was defined as any change in phenotype assignment across perturbations. Distances to decision boundaries were grouped into quantile-based bins. Model uncertainty was evaluated for the triage of unstable cases. Results: Phenotype instability was strongly concentrated near decision boundaries. Under medium analytical variability, samples closest to hemoglobin cutoffs exhibited the highest instability, with the highest instability in the bin closest to the cutoff, a sharp decrease in the adjacent bin, and lower instability across more distant bins. Model uncertainty was enriched among unstable cases, enabling prioritization of borderline samples while reviewing only a subset of all cases. Conclusions: Rule-based CBC phenotyping exhibits intrinsic instability near decision thresholds. Uncertainty-aware machine learning supports a practical framework to identify and prioritize borderline cases without replacing existing laboratory rules, supporting workload-controlled post-analytical decision support.

Graphical Abstract

1. Introduction

Complete blood count (CBC) testing is among the most frequently performed investigations in clinical laboratory medicine and plays a central role in diagnostic assessment, risk stratification, and clinical triage across many clinical settings. CBC-derived phenotypes such as anemia status and subtype, leukocyte abnormalities, and platelet abnormalities are routinely assigned using deterministic rule-based thresholds defined by reference intervals or clinical guidelines.
Rule-based phenotyping is transparent, interpretable and remains the foundation of routine laboratory reporting. However, hard decision boundaries implicitly assume that measured laboratory values are stable representations of the underlying biological state. In practice, laboratory measurements are affected by analytical variability, pre-analytical influences, and rounding effects [1,2,3]. As a result, values close to clinical decision thresholds may change with small fluctuations, which can lead to inconsistent phenotype assignments across repeated measurements.
Machine learning approaches have increasingly been applied to CBC data to support disease classification and clinical decision making. Several studies have focused on differentiating anemia etiologies using supervised learning methods. Tepakhan et al. applied random forest and gradient boosting models to distinguish iron-deficiency anemia from thalassemia using routine hematological parameters, with ground truth labels derived from clinical diagnosis and confirmatory testing rather than direct thresholding of CBC values [4]. Similarly, Wang et al. developed an interpretable ensemble learning model for thalassemia detection in pregnant women, where reference labels were based on established diagnostic criteria incorporating clinical context beyond CBC measurements alone [5].
Other studies have used CBC parameters for broader clinical prediction tasks. Lin et al. investigated early sepsis detection in emergency settings using CBC data combined with clinical variables, with outcome labels defined by retrospective clinical diagnosis [6]. Kamalzadeh et al. focused on predicting iron-deficiency anemia using reticulocyte maturation indices, again relying on diagnosis-based reference labels [7]. In these approaches, machine learning models are primarily evaluated on predictive performance against clinically defined outcomes.
In contrast to disease-oriented prediction, some work has addressed post-analytical laboratory workflows. Zhu et al. applied machine learning to improve the MDS-CBC score in order to optimize peripheral blood smear review. In this setting, labels were linked to expert laboratory review decisions rather than disease diagnosis, and the model was used to support workflow prioritization [8].
Related work has used machine learning to detect specimen misidentification through multi-analyte delta checks [9,10], and broader perspectives have argued that artificial intelligence in laboratory medicine is most valuable as decision support rather than as autonomous diagnosis [11,12]. A parallel literature in clinical machine learning emphasizes the importance of well-calibrated probabilities [13,14] and explicit uncertainty quantification [15,16,17,18] when models are used in safety-critical decision pathways.
Despite this growing body of literature, comparatively little attention has been paid to the robustness of rule-based CBC phenotyping itself. In routine laboratory workflows, phenotypes defined by fixed thresholds are typically treated as stable categorical outcomes. Instability near decision boundaries is rarely quantified, and binary classification outputs do not convey information about the reliability of individual assignments. Consequently, laboratories lack systematic tools to identify borderline cases that may warrant additional review or contextual interpretation [19,20].
In this study, we focus on the intrinsic stability of rule-based CBC phenotyping under realistic analytical variability. We do not propose new diagnostic criteria and we do not replace existing laboratory rules. Instead, we quantify how phenotype instability relates to proximity to clinical decision thresholds. We further investigate whether probabilistic outputs from machine learning models can be used to operationalize this instability by supporting workload-controlled prioritization of borderline cases within established laboratory workflows.

2. Materials and Methods

2.1. Dataset and Study Design

We conducted a retrospective observational study using routine complete blood count (CBC) measurements extracted from the hospital laboratory information system of Mannheim University Medical Center. Data were collected for the period from January 2023 through December 2025. In total, 803,617 CBC records were initially retrieved.
Only adult patients aged 18 years or older at the time of sampling were included. Age and sex were used exclusively for eligibility filtering and for rule-based phenotype assignment according to standard laboratory thresholds. No additional clinical, diagnostic, or biochemical variables beyond those contained in the CBC panel were used for the analyses.
After application of eligibility criteria and phenotype labeling, the final analytic dataset comprised 328,324 CBC records. To prevent information leakage arising from repeated measurements, all analyses were performed using a patient-level split into a development cohort and an independent holdout cohort. All CBC records from a given patient were assigned exclusively to a single cohort.
The development cohort consisted of 263,331 CBC records and was used for model training and internal analyses. The holdout cohort comprised 64,993 CBC records and was reserved exclusively for evaluation of phenotype stability, calibration, and uncertainty-guided triage. No rebalancing or downsampling was applied prior to cohort splitting, and the holdout cohort therefore reflects the real-world distribution of CBC phenotypes observed in routine laboratory practice.
This study was conducted using pseudonymized laboratory data and was approved by the Ethics Committee of the Medical Faculty Mannheim, Heidelberg University (approval no. 2023-893).

2.2. Rule-Based Phenotyping

Each CBC record was assigned three parallel phenotype outputs rather than a single mutually exclusive label. The phenotype dimensions were anemia subtype, white blood cell (WBC) status, and platelet (PLT) status. All labels were derived exclusively from routinely reported numerical CBC measurements and patient sex using predefined deterministic thresholds.
Anemia status was defined using sex-specific hemoglobin (HGB) cutoffs, with anemia assigned for HGB concentrations below 13.0 g/dL in males and below 12.0 g/dL in females. For records meeting anemia criteria, anemia subtype was further classified based on mean corpuscular volume (MCV) as microcytic (MCV < 80 fL), normocytic (MCV 80–100 fL), or macrocytic (MCV > 100 fL). Records not meeting anemia criteria were assigned to the non-anemic category.
WBC status was categorized as low, normal, or high using absolute leukocyte count thresholds of <4.0 × 109/L, 4.0–11.0 × 109/L, and >11.0 × 109/L, respectively. Platelet status was analogously classified as low (<150 × 109/L), normal (150–450 × 109/L), or high (>450 × 109/L).
No manual adjudication or expert revision of phenotype assignments was performed. The rule-based labels were applied uniformly across all records and served as fixed reference phenotypes for all subsequent analyses. Records with missing values in any required labeling parameter and samples from pediatric patients younger than 18 years were excluded prior to analysis.

2.3. Machine Learning Model

We employed a multi-output gradient boosting classifier based on the XGBoost framework to model CBC-derived phenotypes. We trained the model exclusively on the development cohort using the rule-based phenotype assignments defined in Section 2.2 as target outputs. The purpose of the model was not to replace deterministic laboratory rules or to perform disease diagnosis, but to generate probabilistic outputs associated with rule-based phenotype assignments for subsequent uncertainty and stability analyses.
We formulated the task as a multi-output classification problem with three parallel output dimensions corresponding to anemia subtype (4 classes), WBC status (3 classes), and PLT status (3 classes). The multi-output framework was implemented as three independent XGBoost classifiers wrapped in scikit-learn’s MultiOutputClassifier, with one classifier per phenotype head; class probabilities were obtained head-wise from each underlying classifier. For each output dimension, the model produced class probability estimates for all possible categories.
The input feature set consisted of a predefined subset of routinely reported numerical complete blood count (CBC) parameters available at the time of laboratory analysis. These included HGB, MCV, mean corpuscular hemoglobin concentration (MCHC), red cell distribution width (RDW), red blood cell count (RBC), WBC, and PLT. Patient age was included as an additional numerical feature. No clinical diagnoses, outcome variables, expert annotations, or non-CBC laboratory measurements were used as predictors.
Model hyperparameters were selected by nested cross-validation on the development cohort using StratifiedGroupKFold (5 outer × 3 inner folds; patient identifier used as the group to prevent leakage) with 12 iterations of RandomizedSearchCV per inner fold. The search space comprised n_estimators ∈ {200, 300, 450}, max_depth ∈ {4, 6, 8}, learning_rate ∈ {0.05, 0.08, 0.12}, subsample ∈ {0.8, 0.9, 1.0}, colsample_bytree ∈ {0.8, 0.9, 1.0}, and reg_lambda ∈ {0.5, 1.0, 2.0}. The configuration selected by majority vote across outer folds was n_estimators = 200, max_depth = 6, learning_rate = 0.05, subsample = 1.0, colsample_bytree = 0.9, reg_lambda = 1.0, with objective = “multi:softprob”, eval_metric = “mlogloss”, and tree_method = “hist”. The nested cross-validation procedure yielded an outer-fold mean exact-match accuracy of 0.993 ± 0.002 on the development cohort. After hyperparameter selection, a single final model was trained on the full development cohort using the selected parameter set. Patient-level holdout splitting used GroupShuffleSplit (test_size = 0.20). All random operations (splitting, perturbation noise generation, subsampling) used numpy.random with random_state = 42 unless otherwise specified. The analysis was performed in Python 3.11 with scikit-learn 1.5, XGBoost 2.0, NumPy 1.26, pandas 2.2, and matplotlib 3.8.
After training, we applied the final model to the independent holdout cohort to generate probabilistic predictions. We used these probabilities for calibration assessment, perturbation-based stability analysis, and uncertainty-guided triage. All results reported in this study are derived from this single model trained on the development cohort with the real-world data distribution.

2.4. Simulation of Analytical Variability and Stability Analysis

To assess the robustness of rule-based CBC phenotypes under analytically plausible variability, we performed a perturbation-based stability analysis on the independent holdout cohort. Analytical variability was simulated by applying additive Gaussian perturbations to selected CBC parameters in the independent holdout cohort.
Perturbations were applied to hemoglobin (HGB), mean corpuscular volume (MCV), white blood cell count (WBC), and platelet count (PLT), the four parameters directly involved in rule-based phenotype assignment. For each CBC record, additive Gaussian noise with zero mean was applied independently to each parameter. Three perturbation regimes (LOW, MID, HIGH) were defined to span a clinically meaningful range of analytical variability, informed by published analytical performance specifications for hematological parameters [2,19,21] and in broad alignment with typical analyzer-performance data reported in the recent literature [22]. The parameter-specific standard deviations were: for HGB, 0.10 g/dL (LOW), 0.20 g/dL (MID), and 0.30 g/dL (HIGH); for MCV, 0.50 fL (LOW), 1.00 fL (MID), and 1.50 fL (HIGH); for WBC, 0.15 × 109/L (LOW), 0.30 × 109/L (MID), and 0.45 × 109/L (HIGH); and for PLT, 5 × 109/L (LOW), 10 × 109/L (MID), and 15 × 109/L (HIGH). LOW approximates state-of-the-art analytical imprecision routinely achievable on modern hematology platforms; MID corresponds to typical operating conditions; HIGH represents a deliberate stress test under degraded analytical performance for the parameters governing rule-based CBC phenotyping.
For each record and perturbation level, 30 independent perturbation realizations were generated. After each perturbation, phenotype labels were recomputed using the same rule-based thresholds described in Section 2.2, without any modification of cutoff values or classification logic. No retraining or recalibration of the machine learning model was performed during the perturbation process.
A record was defined as unstable if at least one phenotype assignment (anemia subtype, WBC status, or PLT status) differed from the original unperturbed assignment in any perturbation realization. Records for which all phenotype assignments remained unchanged across all perturbations were classified as stable. Stability analyses were conducted exclusively on the holdout cohort, which was not used for model training or hyperparameter selection, to ensure unbiased estimation of instability rates.
For computational efficiency, the stability analysis was limited to a random subset of up to 50,000 holdout records when the full holdout cohort exceeded this size.

2.5. Distance to Decision Boundaries

To contextualize phenotype instability, we quantified the proximity of each CBC record to relevant clinical decision thresholds. We focused on anemia phenotyping, where deterministic cutoffs define discrete categories and where instability is expected to be most pronounced near threshold values.
For each record in the holdout cohort, we computed the absolute distance between the measured hemoglobin concentration and the nearest sex-specific anemia decision threshold. This distance served as a continuous measure of proximity to the rule-based cutoff used for phenotype assignment. Distance calculations were performed using the original, unperturbed CBC values.
To enable robust comparison across the full measurement range and to avoid sparsity effects at extreme distances, we stratified records into quantile-based distance bins. Each bin therefore contained approximately equal numbers of records. This approach was chosen to characterize monotonic relationships between proximity to decision boundaries and phenotype instability, rather than to define absolute clinical distance thresholds.
Instability rates were summarized within each distance bin and used to evaluate how phenotype robustness varied as a function of proximity to the decision boundary. These bin-wise summaries formed the basis for subsequent analyses linking instability patterns to model-derived uncertainty and triage performance.

2.6. Uncertainty-Guided Triage

We used model-derived probabilistic outputs to quantify uncertainty associated with rule-based phenotype assignments. For each phenotype dimension d ∈ {anemia, WBC, PLT}, per-record uncertainty was defined as U_d = 1 − max_c p_d(c), where p_d(c) denotes the model-predicted probability of class c within dimension d. Higher values therefore indicated lower model confidence in the rule-based assignment. To support a single workload-controlled review queue across phenotypes, an aggregate uncertainty score was computed as U_agg = max(U_anemia, U_WBC, U_PLT), equivalently U_agg = 1 − min(max_c p_anemia(c), max_c p_WBC(c), max_c p_PLT(c)). This conservative aggregation flags a record for review whenever at least one phenotype head produces a low-confidence prediction. U_agg was used as the ranking score for all triage analyses spanning the three phenotype dimensions; per-dimension triage was performed using the corresponding per-dimension U_d.
Uncertainty estimates were computed exclusively on the independent holdout cohort using the final model trained on the development cohort. We did not recalibrate or retrain the model using holdout data. Uncertainty values were used to rank CBC records from highest to lowest uncertainty.
To assess the operational utility of uncertainty estimates, we performed a triage analysis in which increasing fractions of records were selected for review based on descending uncertainty. For each review fraction, we calculated the proportion of unstable records captured within the selected subset.
Triage performance was summarized using cumulative capture curves, relating the fraction of reviewed records to the fraction of unstable cases identified. These analyses were performed separately for each phenotype dimension and in aggregate, enabling assessment of uncertainty-guided prioritization across different CBC-derived phenotypes.

3. Results

3.1. Cohort Characteristics and Phenotype Distribution

All results are reported for the independent holdout cohort described in Section 2.1. Rule-based phenotype assignments in the holdout cohort reflected the real-world distribution of CBC-derived categories observed in routine laboratory practice. Most records were assigned to non-anemic, normal WBC, and normal platelet categories, with lower frequencies observed for abnormal phenotypes and specific anemia subtypes. The distribution of phenotype combinations in the holdout cohort is summarized in Table 1. The trained model achieved an exact-match accuracy of 99.30% (anemia subtype 99.60%, macro-F1 0.991; WBC 99.80%, macro-F1 0.998; PLT 99.90%, macro-F1 0.997) on the holdout cohort, confirming that the rule-based labels are essentially recoverable from the input features.

3.2. Instability of Rule-Based Phenotypes Under Analytical Perturbation

Across the 50,000-record stability subsample of the holdout cohort, the proportion of records with at least one phenotype change across the 30 perturbation realizations was 19.91% under LOW analytical variability, 36.80% under MID, and 51.36% under HIGH. Even under the LOW regime (which approximates state-of-the-art analyzer imprecision), approximately one in five records exhibited a phenotype change in at least one of the three rule-based dimensions, underscoring that decision-threshold instability is a structural property of deterministic CBC phenotyping rather than a worst-case artefact. Under MID (desirable APS) and HIGH (minimum-acceptable APS) regimes, the corresponding fractions rose to roughly one in three and one in two records, respectively. Instability was not uniformly distributed across records but was concentrated in a subset of samples whose phenotype assignments changed across perturbed realizations.

3.3. Relationship Between Instability and Distance to Decision Boundaries

Instability rates were strongly associated with proximity to clinical decision thresholds. For anemia phenotyping, records with hemoglobin values closest to the sex-specific anemia cutoffs exhibited the highest instability rates. Quantitatively, under the MID perturbation regime, the instability rate in the bin closest to the cutoff (absolute distance < 0.9 g/dL, n = 10,326) was 59.6%, falling sharply to 27.4% in the second bin (0.9–1.8 g/dL, n = 9715) and remaining within 28.7–34.7% across the third, fourth, and fifth quantile bins (Figure 1). The same boundary-dominated pattern was observed under LOW perturbation (closest bin 33.9%, other bins 14.7–18.5%) and under HIGH perturbation (closest bin 79.0%, other bins 40.3–49.3%), confirming that proximity to the decision threshold dominates instability across the full range of analytical performance considered. Instability in the closest bin was approximately twofold higher than the average across the remaining four bins. This pattern demonstrated a smooth gradient rather than an abrupt transition, indicating a continuous relationship between proximity to decision boundaries and phenotype instability (Figure 1).

3.4. Model Calibration and Uncertainty Patterns

The trained gradient boosting model produced well-calibrated probabilistic outputs across all phenotype dimensions. Reliability diagrams showed close agreement between predicted confidence and empirical accuracy, particularly at higher confidence levels (Figure 2). Deviations from perfect calibration were primarily observed in intermediate confidence ranges, which contained fewer samples.
Higher uncertainty values were predominantly observed among records located near clinical decision thresholds, whereas records farther from thresholds generally exhibited high confidence and stable phenotype assignments. This alignment between model uncertainty and phenotype instability was consistent across anemia subtype, white blood cell status, and platelet status, supporting the use of uncertainty estimates for downstream triage and prioritization.

3.5. Uncertainty-Guided Triage of Unstable Phenotypes

Records were ranked in descending order of aggregate model uncertainty U_agg, and the fraction of unstable records captured was computed as a function of the fraction of records reviewed. Under the MID perturbation regime, reviewing the top 10% of records by uncertainty captured 12.5% of all unstable records, reviewing 20% captured 24.2%, and reviewing 30% captured 36.2% (Figure 3). The corresponding operating points for the LOW regime were 13.1%, 25.3%, and 37.3% (at 10%, 20%, and 30% review fractions, respectively), and for the HIGH regime were 11.8%, 23.2%, and 35.2%, indicating that uncertainty-based ranking provides a consistent enrichment over random selection across all three perturbation regimes.
The magnitude of this enrichment reflects the high accuracy of the trained classifier on the deterministic rule-based labels, which saturates the maximum class probability near 1.0 for most records and consequently limits the dynamic range available for confidence-based ranking. Despite this constraint, uncertainty-based triage consistently outperformed random selection across all three perturbation regimes, confirming that model-derived confidence provides actionable enrichment for borderline case identification in workload-controlled review workflows.
This pattern was consistent across LOW, MID and HIGH analytical perturbation levels (Figure 3), demonstrating that uncertainty-guided prioritization remains effective under varying degrees of analytical variability. Together, these results indicate that model-derived uncertainty can be operationalized to prioritize borderline CBC results for review while maintaining control over post-analytical workload.

3.6. Summary of Key Findings

In summary, perturbation-based analysis revealed that rule-based CBC phenotypes exhibit intrinsic instability near clinical decision thresholds. This instability follows a structured relationship with proximity to decision boundaries, with the highest instability in records closest to the cutoff and lower but non-zero instability across more distant bins. Together, distance-to-cutoff and model uncertainty provide complementary information that supports workload-controlled review of borderline CBC phenotypes without modifying the underlying decision rules.

4. Discussion

In this study, we systematically analyzed the stability of rule-based complete blood count (CBC) phenotypes under analytically plausible perturbations and demonstrate that instability is an intrinsic and structured property of threshold-based laboratory interpretation. We further show that model-derived uncertainty provides a principled and operationally feasible mechanism to identify and triage borderline cases without modifying existing laboratory decision rules.

4.1. Intrinsic Instability of Rule-Based CBC Phenotypes

Our results demonstrate that rule-based CBC phenotyping is not uniformly stable across the measurement space. Instead, instability is concentrated in a subset of records whose numerical values lie close to predefined clinical decision thresholds. This behavior was most pronounced for anemia phenotyping, where sex-specific hemoglobin cutoffs define discrete categorical outcomes, whereas white blood cell and platelet classifications exhibited comparatively lower instability rates. These differences likely reflect broader reference intervals and larger distances between physiological values and decision boundaries for leukocyte and platelet counts.
Importantly, no manual relabeling or subjective adjudication was applied at any stage. Instability emerged solely from analytically plausible perturbations applied to numerical measurements, indicating that such instability is not a modeling artifact but a structural consequence of deterministic thresholding. This observation is consistent with established principles of analytical and biological variation in laboratory medicine [1,2,23], and with the broader recognition that hard cutoffs interact poorly with measurement variability when the underlying distribution has appreciable mass near the cutoff [19,20].

4.2. Distance to Decision Thresholds as a Determinant of Instability

We observed a strong and structured relationship between phenotype instability and proximity to clinical decision boundaries. Records closest to the hemoglobin cutoff exhibited the highest instability rates, with instability decreasing progressively as distance from the threshold increased. This pattern was consistent across perturbation levels and followed a smooth gradient rather than an abrupt transition.
This finding provides a mechanistic explanation for why borderline laboratory results are disproportionately prone to interpretive variability, a phenomenon well recognized in routine laboratory practice but rarely quantified systematically [1,19,20]. Moreover, it indicates that instability is not random noise but a predictable function of measurement geometry relative to decision thresholds. The quantile-based distance stratification used in this study enabled this relationship to be demonstrated in a distribution-agnostic manner, independent of absolute measurement scales.

4.3. Calibration of Model-Derived Probabilities Across Phenotype Dimensions

The gradient boosting model trained on rule-based phenotype labels produced well-calibrated probabilistic outputs across anemia, WBC, and platelet phenotypes, as demonstrated by reliability curves in the independent holdout cohort (Figure 2). Predicted confidence values showed close agreement with empirical outcome frequencies, particularly in high-confidence regions corresponding to stable phenotype assignments [24].
Deviations from perfect calibration were primarily observed in intermediate confidence ranges, which coincided with regions of increased instability and proximity to clinical decision thresholds. This alignment supports the interpretation that model uncertainty captures meaningful ambiguity inherent in the data rather than model miscalibration.
We chose U = 1 − max_c p(c) for three practical reasons. Firstly, it is model-agnostic and post hoc, which is necessary because gradient-boosted trees do not natively admit Monte Carlo dropout or variational approximations. Secondly, in a well-calibrated multi-class classifier, it directly estimates the expected misclassification probability for the assigned label, which is the quantity of interest for triage [13,14]. Finally, it is a validated baseline for misclassification detection across multiple domains [15]. Alternative metrics such as Shannon entropy, deep ensembles [17], and Bayesian or MC-dropout approximations [18] could provide finer epistemic decomposition but at substantially higher training and inference cost.

4.4. Uncertainty as a Tool for Operational Triage Rather than Label Replacement

A central contribution of this work is the demonstration that calibrated model-derived uncertainty estimates can be used to prioritize unstable records for review in a workload-controlled manner. Cumulative capture analyses showed that ranking records by aggregate uncertainty consistently enriched the reviewed subset for unstable cases across all three perturbation regimes (Figure 3). The steep initial slope of the capture curves indicates that instability is highly concentrated among records with the highest uncertainty scores.
Crucially, this approach does not seek to replace or override established laboratory decision rules. Instead, it augments existing workflows by selectively flagging cases where deterministic interpretation is most vulnerable to analytical variability. This positioning aligns with recent work emphasizing the role of machine learning as decision support in laboratory medicine rather than as an autonomous diagnostic system [8].

4.5. Clinical and Laboratory Implications

The instability rates reported here are clinically meaningful, not just statistical. Under the desirable analytical performance regime (MID), roughly one in three CBC records would receive a different rule-based phenotype assignment under analytically plausible measurement variation. Such reclassifications concentrate on the situations where the assigned phenotype directly determines downstream action: borderline anemia near the sex-specific hemoglobin cutoffs, which triggers diagnostic workup and longitudinal follow-up, and borderline WBC or platelet status, which influences hematology referral, preoperative clearance, and dosing decisions for cytotoxic or anticoagulant therapy. Uncertainty-guided triage provides a workload-controlled mechanism to surface this subset for contextual review before the rule-based result drives further clinical action.
From a laboratory medicine perspective, these findings provide a quantitative framework for understanding and managing borderline results. Rather than treating all results equally, uncertainty-guided triage allows laboratories to focus expert review on cases where interpretive instability is most likely, thereby optimizing resource allocation. This is particularly relevant in high-throughput environments, where manual review capacity is limited and indiscriminate flagging may contribute to alert fatigue.
More broadly, this framework reframes uncertainty as a useful signal rather than an undesirable by-product of modeling. By explicitly quantifying ambiguity, laboratories can move toward more nuanced result interpretation while preserving established reference ranges and regulatory-compliant decision thresholds.

4.6. Limitations and Future Directions

Several limitations should be acknowledged. Firstly, this study focused on routine CBC parameters and predefined phenotype rules; extension to other laboratory domains and more complex interpretive algorithms warrants further investigation. Secondly, perturbations were designed to reflect analytical variability but did not explicitly model biological within-subject variation over time, which is known to contribute substantially to total measurement uncertainty [1,2]. Incorporating longitudinal data may further refine instability assessment. Because phenotype labels are deterministically derived from CBC parameters, the model is expected to recover rule structure; the added value here is calibrated uncertainty for borderline cases rather than improved diagnostic accuracy. Finally, while this study demonstrates operational feasibility, prospective evaluation of uncertainty-guided triage in real-world laboratory workflows will be necessary to assess its impact on efficiency and clinical outcomes.

4.7. Conclusions

In conclusion, rule-based CBC phenotypes exhibit predictable and structured instability near clinical decision thresholds. This instability can be quantified and operationally managed using model-derived uncertainty and proximity to decision thresholds without altering existing laboratory rules. Together, these signals provide a transparent, scalable, and clinically compatible approach to prioritizing borderline laboratory results and represent a practical step toward uncertainty-aware laboratory medicine.

Author Contributions

K.S.: Conceptualization, Methodology, Software, Formal analysis, Investigation, Data curation, Writing original draft, Visualization. C.G.: Investigation, Validation and editing. O.E.: Investigation, Validation and editing. A.F.: Supervision, Investigation, Validation. S.N. Visualization, Project administration. All authors have read and agreed to the published version of the manuscript.

Funding

No external funding was received. All work was performed within the routine activities of the Institute of Clinical Chemistry, University Medical Centre Mannheim.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and approved by the Ethics Committee of the Medical Faculty Mannheim, Heidelberg University (approval no. 2023-893).

Informed Consent Statement

Patient consent was waived because this study used retrospective, pseudonymized laboratory data from routine clinical care; no patient contact, intervention, or identifiable data were involved, and the waiver was granted as part of the ethics approval (Ethics Committee of the Medical Faculty Mannheim, Heidelberg University, approval no. 2023-893).

Data Availability Statement

The CBC data analysed in this study were obtained from the clinical laboratory information system of Mannheim University Medical Center and contain pseudonymized patient laboratory results. These data are not publicly available due to privacy restrictions and the conditions of the institutional ethics approval. De-identified summary data and analysis outputs supporting the reported results are available from the corresponding author upon reasonable request, subject to institutional approval.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Fraser, C.G. Biological Variation: From Principles to Practice; AACC Press: Washington, DC, USA, 2001. [Google Scholar]
  2. Coşkun, A.; Braga, F.; Carobene, A.; Tejedor Ganduxé, X.; Aarsand, A.K.; Fernández-Calle, P.; Díaz-Garzón, J.; Bartlett, W.A.; Jonker, N.; Aslan, B.; et al. Systematic review and meta-analysis of within-subject and between-subject biological variation estimates of 20 haematological parameters. Clin. Chem. Lab. Med. 2020, 58, 25–32. [Google Scholar] [CrossRef] [Scilit]
  3. Lippi, G.; Chance, J.J.; Church, S.; Dazzi, P.; Fontana, R.; Giavarina, D.; Grankvist, K.; Huisman, W.; Kouri, T.; Palicka, V.; et al. Preanalytical quality improvement: From dream to reality. Clin. Chem. Lab. Med. 2011, 49, 1113–1126. [Google Scholar] [CrossRef] [Scilit]
  4. Tepakhan, W.; Srisintorn, W.; Penglong, T.; Saelue, P. Machine learning approach for differentiating iron deficiency anemia and thalassemia using random forest and gradient boosting algorithms. Sci. Rep. 2025, 15, 16917. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Wang, Q.; Dai, X.; Xu, K.; Xu, M.; Ying, J. Development of an interpretable ensemble learning model for thalassemia detection in pregnant women using routine hematological parameters. Digit. Health 2025, 11, 20552076251396982. [Google Scholar] [CrossRef] [Scilit]
  6. Lin, T.-H.; Chung, H.-Y.; Jian, M.-J.; Chang, C.-K.; Lin, H.-H.; Yen, C.-T.; Tang, S.-H.; Pan, P.-C.; Perng, C.-L.; Chang, F.-Y.; et al. AI-driven innovations for early sepsis detection by combining predictive accuracy with blood count analysis in an emergency setting: Retrospective study. J. Med. Internet Res. 2025, 27, e56155. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Kamalzadeh, H.; Choobin, N.; Haghighat, A.; Ardestani, M.T.; Farhanpoor, M. AI-assisted haematology: Machine learning-based prediction of iron-deficiency anaemia from reticulocyte maturation indices. BMC Med. Inform. Decis. Mak. 2025, 26, 24. [Google Scholar] [CrossRef] [Scilit]
  8. Zhu, J.; Lemaire, P.; Mathis, S.; Ronez, E.; Clauser, S.; Jondeau, K.; Fenaux, P.; Adès, L.; Bardet, V. Machine learning-based improvement of MDS-CBC score brings platelets into the limelight to optimize smear review in the hematology laboratory. BMC Cancer 2022, 22, 972. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Rosenbaum, M.W.; Baron, J.M. Using machine learning-based multianalyte delta checks to detect wrong blood in tube errors. Am. J. Clin. Pathol. 2018, 150, 555–566. [Google Scholar] [CrossRef] [Scilit]
  10. Seok, H.S.; Choi, Y.; Yu, S.; Shin, K.-H.; Kim, S.; Shin, H. Machine learning-based delta check method for detecting misidentification errors in tumor marker tests. Clin. Chem. Lab. Med. 2024, 62, 1421–1432. [Google Scholar] [CrossRef] [Scilit]
  11. Cabitza, F.; Campagner, A.; Soares, F.; García de Guadiana-Romualdo, L.; Challa, F.; Sulejmani, A.; Seghezzi, M.; Carobene, A. The importance of being external. Methodological insights for the external validation of machine learning models in medicine. Comput. Methods Programs Biomed. 2021, 208, 106288. [Google Scholar] [CrossRef] [Scilit]
  12. Padoan, A.; Plebani, M. Artificial intelligence: Is it the right time for clinical laboratories? Clin. Chem. Lab. Med. 2022, 60, 1859–1861. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (ICML), Sydney, Australia, 6–11 August 2017; pp. 1321–1330. [Google Scholar]
  14. Niculescu-Mizil, A.; Caruana, R. Predicting good probabilities with supervised learning. In Proceedings of the 22nd International Conference on Machine Learning (ICML), Bonn, Germany, 7–11 August 2005; ACM: New York, NY, USA, 2005; pp. 625–632. [Google Scholar] [CrossRef] [Scilit]
  15. Hendrycks, D.; Gimpel, K. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France, 24–26 April 2017. [Google Scholar]
  16. Begoli, E.; Bhattacharya, T.; Kusnezov, D. The need for uncertainty quantification in machine-assisted medical decision making. Nat. Mach. Intell. 2019, 1, 20–23. [Google Scholar] [CrossRef] [Scilit]
  17. Lakshminarayanan, B.; Pritzel, A.; Blundell, C. Simple and scalable predictive uncertainty estimation using deep ensembles. In Proceedings of the Advances in Neural Information Processing Systems 30 (NeurIPS), Long Beach, CA, USA, 4–9 December 2017; pp. 6402–6413. [Google Scholar]
  18. Gal, Y.; Ghahramani, Z. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of the 33rd International Conference on Machine Learning (ICML), New York, NY, USA, 19–24 June 2016; pp. 1050–1059. [Google Scholar]
  19. Sandberg, S.; Fraser, C.G.; Horvath, A.R.; Jansen, R.; Jones, G.; Oosterhuis, W.; Petersen, P.H.; Schimmel, H.; Sikaris, K.; Panteghini, M. Defining analytical performance specifications: Consensus Statement from the 1st Strategic Conference of the European Federation of Clinical Chemistry and Laboratory Medicine. Clin. Chem. Lab. Med. 2015, 53, 833–835. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  20. Paal, M.; Habler, K.; Vogeser, M. Estimation of inter-laboratory reference change values from external quality assessment data. Biochem. Med. 2021, 31, 030902. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Aarsand, A.K.; Fernandez-Calle, P.; Webster, C.; Coşkun, A.; Gonzales-Lao, E.; Diaz-Garzon, J.; Jonker, N.; Bartlett, W.A.; Sandberg, S.; on behalf of the EFLM Working Group on Biological Variation. The EFLM Biological Variation Database. Available online: https://biologicalvariation.eu (accessed on 12 May 2026).
  22. Čičak, H.; Radišić Biljak, V.; Šimundić, A.-M. Verification of a 6-part differential haematology analyser Siemens Advia 2120i. Biochem. Med. 2022, 32, 020710. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Clinical and Laboratory Standards Institute (CLSI). Evaluation of Precision of Quantitative Measurement Procedures; EP05-A3; CLSI: Wayne, PA, USA, 2014. [Google Scholar]
  24. Steyerberg, E.W.; Vickers, A.J.; Cook, N.R.; Gerds, T.; Gonen, M.; Obuchowski, N.; Pencina, M.J.; Kattan, M.W. Assessing the performance of prediction models: A framework for traditional and novel measures. Epidemiology 2010, 21, 128–138. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Relationship between phenotype instability and distance to the anemia decision threshold. Instability rates for rule-based anemia phenotyping in the independent holdout cohort are shown as a function of absolute distance to the nearest sex-specific hemoglobin cutoff under the MID analytical perturbation regime (HGB σ = 0.20 g/dL; see §2.4). Records are grouped into quantile-based distance bins (quantile breakpoints 0, 0.2, 0.4, 0.6, 0.8, 1.0), so that each bin contains approximately equal numbers of records; the lowest bin therefore corresponds to records closest to the cutoff. Points indicate the fraction of unstable records within each bin; numbers adjacent to points denote the number of records per bin (binomial 95% confidence interval ≈ ±1 percentage point at n ≈ 10,000 per bin). Instability is highest among records closest to the decision threshold, drops sharply in the adjacent bin, and then gradually increases with further distance from the threshold.
Figure 1. Relationship between phenotype instability and distance to the anemia decision threshold. Instability rates for rule-based anemia phenotyping in the independent holdout cohort are shown as a function of absolute distance to the nearest sex-specific hemoglobin cutoff under the MID analytical perturbation regime (HGB σ = 0.20 g/dL; see §2.4). Records are grouped into quantile-based distance bins (quantile breakpoints 0, 0.2, 0.4, 0.6, 0.8, 1.0), so that each bin contains approximately equal numbers of records; the lowest bin therefore corresponds to records closest to the cutoff. Points indicate the fraction of unstable records within each bin; numbers adjacent to points denote the number of records per bin (binomial 95% confidence interval ≈ ±1 percentage point at n ≈ 10,000 per bin). Instability is highest among records closest to the decision threshold, drops sharply in the adjacent bin, and then gradually increases with further distance from the threshold.
Aimed 01 00013 g001
Figure 2. Calibration of model-derived uncertainty for CBC phenotypes in the independent holdout cohort. Reliability diagrams illustrating the relationship between mean predicted confidence and empirical accuracy for (A) anemia subtype, (B) white blood cell (WBC) status, and (C) platelet (PLT) status. For each phenotype, predictions were grouped into quantile-based confidence bins defined by maximum class probability (number of bins differs per phenotype due to data distribution). Points represent observed accuracy within each bin, and the dashed diagonal indicates perfect calibration. Numbers adjacent to points denote the number of CBC records per bin. Across phenotypes, higher confidence values were generally associated with higher empirical accuracy, indicating that model-derived probabilities provide a meaningful measure of uncertainty for rule-based phenotype assignments.
Figure 2. Calibration of model-derived uncertainty for CBC phenotypes in the independent holdout cohort. Reliability diagrams illustrating the relationship between mean predicted confidence and empirical accuracy for (A) anemia subtype, (B) white blood cell (WBC) status, and (C) platelet (PLT) status. For each phenotype, predictions were grouped into quantile-based confidence bins defined by maximum class probability (number of bins differs per phenotype due to data distribution). Points represent observed accuracy within each bin, and the dashed diagonal indicates perfect calibration. Numbers adjacent to points denote the number of CBC records per bin. Across phenotypes, higher confidence values were generally associated with higher empirical accuracy, indicating that model-derived probabilities provide a meaningful measure of uncertainty for rule-based phenotype assignments.
Aimed 01 00013 g002
Figure 3. Uncertainty-guided triage of unstable CBC phenotypes. Fraction of unstable records captured (y-axis) as a function of the fraction of CBC records reviewed (x-axis), with records ranked in descending order by the aggregate model uncertainty U_agg = 1 − min(max p_anemia, max p_WBC, max p_PLT). Curves are shown for the LOW, MID, and HIGH analytical perturbation regimes; the diagonal y = x reference line represents the capture expected under random selection. Across regimes, uncertainty-based prioritisation provides a consistent lift over random selection (1.2–1.4×).
Figure 3. Uncertainty-guided triage of unstable CBC phenotypes. Fraction of unstable records captured (y-axis) as a function of the fraction of CBC records reviewed (x-axis), with records ranked in descending order by the aggregate model uncertainty U_agg = 1 − min(max p_anemia, max p_WBC, max p_PLT). Curves are shown for the LOW, MID, and HIGH analytical perturbation regimes; the diagonal y = x reference line represents the capture expected under random selection. Across regimes, uncertainty-based prioritisation provides a consistent lift over random selection (1.2–1.4×).
Aimed 01 00013 g003
Table 1. Most frequent rule-based complete blood count (CBC) phenotype combinations in the independent holdout cohort (N = 64,993). Phenotypes were assigned deterministically for anemia subtype (none, microcytic, normocytic, macrocytic), white blood cell (WBC) status (low/normal/high), and platelet (PLT) status (low/normal/high) using predefined laboratory thresholds. Percentages are calculated relative to all holdout CBC records. Remaining combinations were pooled as “Other combinations”.
Table 1. Most frequent rule-based complete blood count (CBC) phenotype combinations in the independent holdout cohort (N = 64,993). Phenotypes were assigned deterministically for anemia subtype (none, microcytic, normocytic, macrocytic), white blood cell (WBC) status (low/normal/high), and platelet (PLT) status (low/normal/high) using predefined laboratory thresholds. Percentages are calculated relative to all holdout CBC records. Remaining combinations were pooled as “Other combinations”.
Anemia SubtypeWBC StatusPLT StatusRecords, nPercentage (%)
NoneNormalNormal18,01727.7
NormocyticNormalNormal14,64422.5
NormocyticHighNormal648910.0
NoneHighNormal48837.5
NormocyticNormalLow39406.1
NormocyticLowLow29004.5
NormocyticHighLow18752.9
NoneNormalLow15772.4
MicrocyticNormalNormal14672.3
NormocyticHighHigh13362.1
Other combinations (remaining 26) 786512.1
Total 64,993100.0
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Shater, K.; Gerhards, C.; Evliyaoglu, O.; Nittka, S.; Fischer, A. Operationalizing Instability in Rule-Based Complete Blood Count Phenotyping Using Uncertainty-Aware Machine Learning. AI Med. 2026, 1, 13. https://doi.org/10.3390/aimed1020013

AMA Style

Shater K, Gerhards C, Evliyaoglu O, Nittka S, Fischer A. Operationalizing Instability in Rule-Based Complete Blood Count Phenotyping Using Uncertainty-Aware Machine Learning. AI in Medicine. 2026; 1(2):13. https://doi.org/10.3390/aimed1020013

Chicago/Turabian Style

Shater, Karim, Catharina Gerhards, Osman Evliyaoglu, Stefanie Nittka, and Andreas Fischer. 2026. "Operationalizing Instability in Rule-Based Complete Blood Count Phenotyping Using Uncertainty-Aware Machine Learning" AI in Medicine 1, no. 2: 13. https://doi.org/10.3390/aimed1020013

APA Style

Shater, K., Gerhards, C., Evliyaoglu, O., Nittka, S., & Fischer, A. (2026). Operationalizing Instability in Rule-Based Complete Blood Count Phenotyping Using Uncertainty-Aware Machine Learning. AI in Medicine, 1(2), 13. https://doi.org/10.3390/aimed1020013

Article Metrics

Back to TopTop