Next Article in Journal
Noise-Dependent Robustness of XGBoost, LightGBM, and CatBoost
Previous Article in Journal
Machine Learning-Based Early Warning System for Transport-Layer Bottlenecks in Open-Source 5G Testbeds
Previous Article in Special Issue
Exploratory Study of Triple-Entry Data (X-STATIS) for the Evaluation of Energy Quality in Induction Motors Under Simulated Faults
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

EDBERT: Predicting Emergency Department Disposition Using a BERT-Based Architecture

1
La Trobe Business School, La Trobe University, Plenty Rd, Melbourne, VIC 3086, Australia
2
Department of Computer Science & Engineering, Rajshahi University of Engineering and Technology, Rajshahi 6204, Bangladesh
3
Centre for Data Analytics and Cognition, La Trobe University, Plenty Rd, Melbourne, VIC 3086, Australia
4
Australian Centre for Artificial Intelligence in Medical Innovation, La Trobe University, Melbourne, VIC 3086, Australia
5
School of Engineering and Technology, Central Queensland University, Rockhampton, QLD 4702, Australia
6
Emergency Department, St Vincent’s Hospital Melbourne, Melbourne, VIC 3065, Australia
7
The Royal Melbourne Hospital (RMH), Melbourne, VIC 3050, Australia
8
Department of Critical Care, Melbourne Medical School, The University of Melbourne, Melbourne, VIC 3010, Australia
9
School of Psychology, Faculty of Health, Deakin University, Burwood, VIC 3125, Australia
*
Author to whom correspondence should be addressed.
Algorithms 2026, 19(9), 729; https://doi.org/10.3390/a19090729 (registering DOI)
Submission received: 17 July 2026 / Revised: 23 August 2026 / Accepted: 25 August 2026 / Published: 30 August 2026
(This article belongs to the Special Issue Algorithms in Data Classification (4th Edition))

Abstract

Objective: To develop and evaluate EDBERT (Emergency Department Bidirectional Encoder Representations from Transformers), a BERT-based architecture built around a multi-input fusion network that enriches free-text triage notes with the additional context of presenting complaints and patient age to predict Emergency Department (ED) disposition, supporting early decision-making and reduced ED length of stay (LOS). Methods: A retrospective cohort of 570,143 ED presentations to the Royal Melbourne Hospital, Australia, was used. At the core of EDBERT, a fusion network integrates three complementary input sources—triage notes, presenting complaints, and patient age—through learnable weights, with the two free-text inputs encoded by customised eight-layer BERT encoder stacks pre-trained on triage text with Masked Language Modelling (MLM). The resulting architecture is also lightweight, containing 66 million parameters, substantially fewer than the 110 million of BERT-BASE. Eighty percent of the dataset was used for training, and twenty percent for testing. Performance was benchmarked against BERT-Base configurations of varying depth. Results: EDBERT achieved an accuracy of 83.26%, a macro-averaged F1-score of 81.99%, and an area under the receiver operating characteristic curve of 0.91, outperforming all single-input BERT-Base variants, including a depth-matched eight-layer model (82.08%), indicating that the gain derives from the fusion of complementary inputs combined with domain-adaptive pre-training rather than from model capacity. Conclusions: A fusion-based BERT architecture that supplements the triage narrative with presenting complaint and age context predicts ED disposition more accurately than larger single-input BERT models, while requiring substantially less memory and computation, making it practical for deployment in hospitals with limited computing resources.

1. Introduction

Australia faces an ever-growing demand for emergency care and hospital admissions [1]. Emergency Departments (EDs) provide 24 h assistance for people with severe or life-threatening injuries or illnesses, and they are experiencing increased patient demand that is surpassing population growth [2]. In 2023–2024, an estimated 9 million patients presented to EDs in Australia [2]—24,658 patients per day—with approximately 31% subsequently admitted to an inpatient ward [3].
Emergency Department length of stay (LOS) is the duration between a patient’s arrival at the ED and their departure from it [4,5]. Emergency Departments are often overcrowded [6], leading to prolonged ED LOS, which correlates with increased morbidity and mortality [7]. Jones et al. reported a time-associated linear increase in all-cause 30-day mortality for patients who remain in the ED for more than 5 h after arrival, with one additional death occurring for every 82 patients delayed for over 6 to 8 h [8]. The early identification of patients who require admission could improve ED flow, optimise resource allocation, and improve patient outcomes.
The triage process serves as the first point of contact for patients in the ED [9]. Before assessment, demographic information and the reason for presentation are collected and documented in the triage note by a triage nurse [10]. The presenting problem, also referred to as the chief complaint [11], provides a description of the primary concern that prompted the patient to attend the ED. The triage note is a brief clinical record that typically includes the reason for presentation, relevant symptoms, medical history, and physiological observations used to determine clinical urgency. The quality and completeness of triage documentation can influence the efficiency and effectiveness of early ED assessment and care [12]. Nonetheless, ED triage notes are commonly written in a highly compressed style and may contain abbreviations, symbols, incomplete phrases, and locally used shorthand. The following de-identified example from the study dataset illustrates these characteristics:
Presenting problem: Left axillary pain and rectal bleeding.
Triage note:
SP: this PM, intermittent L) axilla pain, sharp localised, also c/o blood in stool few times this week, small, bright. PHx nil. GCS 15, RR 16, SpO2 100%, BP 126/72, T 36.3, pain 2/10--declines analgesia, BOC 0, nil COVID sx.
This short narrative includes clinically relevant information such as the timing, location, and characteristics of the pain; the additional report of rectal bleeding; previous medical history; physiological observations; pain severity; and the patient’s decision to decline analgesia. Although such documentation is efficient for rapid clinical recording, its abbreviated and locally variable form creates challenges for consistent interpretation and automated natural language processing (NLP).
With the recent significant advancements in NLP, the adoption of large language models (LLMs) in the medical context has gained considerable attention [13]. These models have shown significant efficacy in downstream medical NLP tasks when trained on large datasets of clinical text [14]. Large language models such as Bidirectional Encoder Representations from Transformers (BERT) [15] and its pre-trained variants, including ClinicalBERT [16] and BioBERT [17], have shown significant potential for medical applications. In 2021, Tahayori et al. reported the first Australian application of BERT to free-text ED triage notes, predicting patient disposition with an accuracy of 83% and an area under the receiver operating characteristic curve (AUROC) of 0.88 [18]. Ferri et al. developed a deep ensemble multitask model incorporating BERT to predict emergency jurisdiction from multimodal emergency medical call incidents, reporting an accuracy of 0.80 [19]. Building on these findings, Chang et al. demonstrated that machine learning models integrating both structured and unstructured triage data could outperform emergency physicians in disposition prediction [20], and Chen et al. introduced a deep neural network integrating physicians’ clinical narratives with vital signs that outperformed traditional logistic regression and early warning scores in terms of F1-score [21]. Boughorbel et al. further explored multimodal outcome prediction in the ED using a perceiver-based language model that combined textual and structured clinical information, highlighting the potential of multi-input representations for ED outcome prediction [22]. More recently, Chen et al. used BlueBERT, a biomedical domain-specific language model [23], and achieved an AUROC of 0.9014, demonstrating that domain-adapted pre-trained models provide considerable gains over general-purpose LLMs [24]. Task-specific architectural adaptations have also demonstrated benefits in other deep learning domains, including computer vision, where multi-scale feature extraction has been used to improve cross-view gait recognition [25].
However, two important gaps remain in this body of work: First, most existing BERT-based disposition models rely on the free-text triage note as a single input, without systematically integrating the readily available structured context that accompanies it at triage, such as the presenting complaint and patient age; indeed, Tahayori et al. anticipated that adding age as a distinct feature would further improve predictive performance [18]. Second, developing, fine-tuning, and maintaining LLMs for healthcare applications is resource-intensive [26]: BERT-BASE contains approximately 110 million parameters [27], demanding substantial computational power that makes it less practical for hospitals with limited computing resources.
In this study, we developed EDBERT (Emergency Department Bidirectional Encoder Representations from Transformers), a lightweight BERT-based framework organised around a multi-input fusion network. Rather than introducing an entirely new Transformer architecture, EDBERT adapts the BERT framework to the ED setting by combining triage notes with complementary information from the presenting complaint and patient age. These three inputs are integrated through learnable weights to form a unified representation for disposition prediction. The Transformer encoder is deliberately designed to be lightweight, with fewer layers, attention heads, and hidden dimensions than conventional BERT configurations, thereby reducing memory requirements and computational cost while retaining predictive performance. The main contributions of this study are therefore as follows: (i) a fusion-based BERT framework that integrates free-text triage notes with presenting complaints and patient age for ED disposition prediction; (ii) a lightweight, domain-tailored encoder architecture designed to reduce memory and computational overhead; and (iii) a comparative evaluation of EDBERT against other BERT-based architectures using real-world ED data.

2. Materials and Methods

2.1. Data

We used a retrospective dataset of electronic patient triage records from the ED at the Royal Melbourne Hospital, Australia. Approximately 5% of presentations had other disposition categories, such as transfer to another hospital, leaving the ED before completion of care, or death in the ED, which were excluded from this study. Following these exclusions, the final analytical dataset comprised 570,143 presentations, of which 211,810 (37.2%) resulted in admission and 358,333 (62.8%) had a “Home” or “discharge” disposition. Each record included the triage note and presenting complaint as free-text fields, and patient age as an integer value.
Eighty percent of the final analytical dataset was used for model training, while the remaining twenty percent was reserved as the test set to evaluate final predictive performance on unseen observations.

2.2. The EDBERT Model

EDBERT is organised around a fusion network that integrates three complementary input sources (triage notes, presenting complaints, and patient age) into a single representation for disposition prediction. The two free-text inputs are encoded by customised, lightweight BERT encoder stacks (66 million parameters in total). Below, we describe the encoder stack and the fusion network that forms the core of the architecture.

2.2.1. Customised BERT Encoder Stack

The Transformer encoder stack in EDBERT (Figure 1) is a modified version of the BERT architecture [15]. While BERT-BASE consists of 12 transformer layers, 12 attention heads, a hidden size of 768, and an intermediate size of 3072, EDBERT uses a more compact architecture comprising eight transformer layers, eight multi-head self-attention heads, a hidden size of 512, and a feedforward intermediate size of 2048. The model vocabulary was tailored to include additional domain-specific tokens and expanded to 40,000 tokens, while the maximum positional embedding length of 512 was retained. Table 1 summarises these architectural modifications. The empirical analysis of encoder depth that informed the selection of the 8-layer architecture is provided in the Supplementary Materials (Table S1, Figures S1 and S2).

2.2.2. Fusion Network

The fusion network integrates three distinct input sources—triage notes, presenting complaints, and patient age—each providing complementary information for disposition prediction. Triage notes serve as the primary source because they contain rich, unstructured clinical information, including detailed symptom descriptions, severity indicators, contextual factors, and triage nurse observations. However, presenting complaints and age were also included because they provide concise and standardised information that may not always be stated explicitly or consistently within a triage note. Presenting complaints capture a patient’s primary reason for attending an ED, while age provides an important indicator of baseline clinical risk. The inclusion of presenting complaint and age was informed by both the prior literature and routine ED practice, as these variables provide complementary, routinely available information that may not be fully captured in the triage note. Previous studies have similarly shown that routinely collected triage features, including age, triage category, arrival mode, and vital signs, are among the strongest contributors to predictive performance [28].
Each input is first processed through a dedicated encoding pipeline. The triage notes and presenting complaints are tokenised and encoded separately using two identical customised encoder stacks (Section 2.2.1), pre-trained on domain-specific triage data, to produce contextual embeddings. During training, each embedding passes through a dropout layer to reduce overfitting. The age input is projected into the same embedding space as the BERT hidden layers via a fully connected linear transformation.
To combine the input sources, the model applies learnable scalar weights to each of the three embeddings, and the weighted embeddings are summed to form a single fused representation:
Fusion = w text · E text + w pc · E pc + w age · E age
This fused embedding is then passed through the final classification layer, which predicts the ED disposition (Figure 2).

2.3. Model Training

EDBERT was trained in two stages: In the first stage, the two customised encoder stacks were pre-trained using the Masked Language Modelling (MLM) technique [15]: 15% of the input tokens were randomly selected for masking, and the model was trained to reconstruct their original values from the surrounding context. This enables the model to learn deep, context-sensitive representations and capture the complex semantic and syntactic relationships within triage texts. In the second stage, the full model was trained end-to-end by calculating the loss between the predicted and actual disposition labels using the cross-entropy loss function.
The weights applied to the three input sources were determined through grid-search optimisation using only the training partition, with the held-out 20% test set reserved exclusively for final model evaluation. The resulting weights were w text = 0.65 for triage notes, w pc = 0.20 for presenting complaints, and w age = 0.15 for age. The test set was not used for fusion-weight optimisation, hyperparameter selection, or Masked Language Modelling (MLM) pre-training. MLM pre-training was likewise restricted to the training partition to minimise information leakage. The Adam optimiser [29] was used with a learning rate of 2 × 10 5 and a weight decay of 0.01. The batch size was set to 64 for training and 32 for evaluation, and a dropout rate of 0.5 was applied during training to mitigate overfitting. Table 2 summarises the hyperparameter configuration.

2.4. Evaluation

Predictive performance on the held-out test set was assessed using accuracy, macro-averaged precision, recall, F1-score, the receiver operating characteristic (ROC) curve, and the area under the ROC curve (AUC). EDBERT was benchmarked against BERT-Base configurations with 2 to 12 transformer layers (in increments of two). Two ablation analyses were conducted to justify the final design: (i) the effect of encoder depth on predictive performance, parameter count, and training time; and (ii) the contribution of MLM pre-training. The complete results of these analyses are reported in the Supplementary Materials. Finally, we compared EDBERT’s performance against several pre-trained BERT variants, including BERT-Base [15], MedBERT [30], DistilBERT [31], TinyBERT [32], and ALBERT [33], on the held-out set.

3. Results

The triage notes had a maximum length of 101 words and an average length of 391 characters, including spaces.
EDBERT achieved the highest performance of all evaluated models, with an accuracy of 83.26% and macro-averaged precision, recall, and F1-score of 82.14%, 81.85%, and 81.99%, respectively, outperforming the standard 12-layer BERT-Base model across all metrics (Table 3). Notably, EDBERT also exceeded the depth-matched eight-layer BERT-Base model trained on triage notes alone (83.26% vs. 82.08% accuracy; Supplementary Table S1), indicating that the improvement stems primarily from the fusion of complementary inputs and domain-adaptive pre-training rather than from model capacity. The ROC curve of EDBERT on the test set is shown in Figure 3, with an AUC of 0.91, indicating strong discriminative ability between the two disposition classes.
The supporting analyses that informed the final architecture are summarised below, with the full details provided in the Supplementary Materials. First, varying the encoder depth of BERT-Base from 2 to 12 layers produced only marginal accuracy differences (81.47% to 82.23%; Supplementary Table S1 and Figure S1), while parameter count and training time per epoch increased substantially with depth (Supplementary Figure S2). Deeper models did not yield performance gains, which may be attributed to the low lexical variability of the triage notes, potentially leading to overfitting in larger models. The eight-layer configuration offered the most favourable balance between predictive performance and computational efficiency, and it was therefore adopted for EDBERT. Second, MLM pre-training on triage notes improved EDBERT’s accuracy from 82.18% to 83.26% (Supplementary Table S2 and Figure S4), confirming the value of domain-adaptive pre-training.
Table 4 shows the performance comparison of pre-trained BERT variants, including BERT-Base, MedBERT, DistilBERT, TinyBERT, and ALBERT. All variants performed within a comparable band of 80.9% to 82.1% accuracy, with macro-averaged precision, recall, and F1-scores clustered between 81% and 82%. MedBERT, having already been pre-trained on the medical domain, achieved the highest accuracy (82.19%), marginally ahead of BERT-Base (81.7%). The compressed architectures—DistilBERT (81.6%), ALBERT (81.4%), and TinyBERT (80.9%)—performed comparably despite having considerably fewer parameters. Importantly, every BERT variant evaluated on triage notes alone remained below the 83.26% accuracy achieved by EDBERT.

4. Discussion

In this study, we designed and evaluated a domain-tailored BERT architecture whose central component is a fusion network that enriches the free-text triage note with the additional context of the presenting complaint and patient age. EDBERT predicts ED disposition with an accuracy of 83.26% and has an AUC of 0.91, while using 66 million parameters—approximately 40% fewer than BERT-BASE. These results compare favourably with previous work. Tahayori et al. reported an accuracy of 83% and an AUROC of 0.88 using a full-size BERT model on triage notes alone [18], and Chen et al. achieved an AUROC of 0.9014 using BlueBERT, a domain-specific model pre-trained on large biomedical corpora [24]. EDBERT achieved comparable or better discriminative performance with a substantially smaller architecture, without relying on external biomedical pre-training corpora.
The improved performance of EDBERT relative to larger, general-purpose BERT configurations can be explained by several factors. First, and most importantly, the fusion network supplements the triage narrative with the complementary context of the presenting complaint and patient age. A depth-matched eight-layer BERT-Base model trained on triage notes alone reached 82.08% accuracy, whereas EDBERT reached 83.26%, indicating that enriching the input representation, rather than adding model capacity, drives the gain. The ablation results suggest this enrichment acts in combination with domain-adaptive pre-training, as the fusion-based model without MLM pre-training reached 82.18% (Supplementary Table S2). The grid-searched fusion weights confirm the dominant contribution of the triage narrative ( w text = 0.65 ), consistent with prior evidence that free-text triage narratives carry the greatest predictive signal for disposition [10,20], while the presenting complaint ( w pc = 0.20 ) and age ( w age = 0.15 ) contribute additional discriminative context that the note alone does not capture. Second, domain-adaptive MLM pre-training on triage text allowed the model to learn the specialised terminology, abbreviations, and contextual patterns characteristic of ED documentation, which a general-purpose model has less opportunity to capture; removing MLM pre-training reduced accuracy by more than one percentage point. Third, triage notes are short and telegraphic, with a maximum length of 101 words in our dataset, and exhibit very low lexical variability (1.01% unique word ratio); the representational capacity of a 12-layer, 110-million-parameter model is therefore excessive and may promote overfitting, whereas EDBERT’s reduced encoder depth is matched to the linguistic complexity of the input.
The lightweight design of EDBERT has direct practical implications. Full-size LLMs are resource-intensive to fine-tune and maintain [26], which is a barrier for hospitals with limited computing infrastructure. By reducing memory usage and training time while preserving predictive performance, EDBERT is more feasible for local training and deployment within existing hospital systems. Early identification of patients likely to require admission could allow bed allocation to begin at the time of presentation, streamline care, reduce the cognitive load on ED clinicians, and ultimately contribute to reduced ED LOS [7,8].
An analysis of misclassified cases identified two recurring patterns, illustrated in Table 5.
False positives commonly involved older patients with substantial comorbidity and high-acuity symptom language who were ultimately discharged through planned or established care pathways; for example, 110 false-positive cases contained references to liver disease or cirrhosis. In contrast, false negatives often involved younger patients with short, apparently low-acuity presentations whose need for specialist admission was not evident at triage; 189 false-negative cases involved burns. These findings indicate that final disposition may depend on downstream specialist assessment, investigations, or care pathways not captured in the triage information. This supports future investigation of additional clinical context and uncertainty-aware prediction mechanisms.
More recent research has also emphasised the importance of uncertainty-aware prediction and effective collaboration between clinicians and predictive models. Abdulai et al. proposed a conformal prediction framework that enables a model to abstain when a case is ambiguous, thereby reducing the risk of overconfident misclassification and supporting safer clinical use [34]. Similarly, Barak-Corren et al. directly compared clinician and algorithmic predictions in a prospective observational study. They showed that physicians correctly predicted 57% of admissions with a specificity of 88%, whereas a random forest model using electronic health record (EHR) data from the first hour of the ED visit achieved a sensitivity of 69% and a specificity of 90%. Combining clinician judgement with algorithmic predictions produced the highest positive predictive value of 86% [35]. These findings indicate that predictive performance alone is insufficient for safe deployment in a high-risk clinical environment. Future work should incorporate uncertainty estimation or abstention mechanisms, allowing EDBERT to defer ambiguous cases for clinician review and evaluate the model within a human-in-the-loop framework to determine whether its predictions complement, rather than replace, clinician judgement. Prospective shadow-mode evaluation and subsequent clinician–model collaboration studies would help establish the safety, reliability, and practical value of EDBERT in real-world ED workflows.

Limitations

This study has several limitations. First, it used data from a single site. Because triage documentation style varies considerably between hospitals, the model may need to be re-trained on local data before implementation elsewhere, and a multi-centre study is required to evaluate the generalisability of the proposed architecture. Second, presentations with dispositions other than “Admit” and “Home” (approximately 5% of the cohort) were excluded, so the model’s behaviour on these less common outcomes remains untested.
Third, EDBERT currently produces point predictions without uncertainty estimates or an abstention mechanism. Recent work has emphasised the importance of uncertainty-aware prediction for safe clinical use: Abdulai et al. proposed a conformal prediction framework that enables a model to abstain when a case is ambiguous, reducing the risk of overconfident misclassification [34]. Incorporating such mechanisms would allow EDBERT to defer ambiguous cases for clinician review rather than producing potentially unreliable predictions.
Finally, EDBERT was evaluated retrospectively, without clinician interaction, so determining whether its predictions complement rather than replace clinician judgement remains untested. Barak-Corren et al. found that combining clinician judgement with algorithmic predictions produced the highest positive predictive value (86%) [35]. These findings suggest that predictive performance alone is insufficient for safe deployment in a high-risk clinical environment. Prospective shadow-mode evaluation and subsequent clinician–model collaboration studies within a human-in-the-loop framework would help establish the safety, reliability, and practical value of EDBERT in real-world ED workflows.

5. Conclusions

This study introduces EDBERT, a specialised Transformer-based architecture for ED disposition prediction whose central component is a fusion network that integrates the free-text triage note with the complementary context of the presenting complaint and patient age through learnable weights. Supplementing the triage narrative with this additional context, together with domain-adaptive MLM pre-training, enabled EDBERT to consistently outperform larger single-input BERT baselines. By further tailoring the encoder stacks to the specific characteristics of ED text, including reduced Transformer layers, attention heads, and hidden dimensions, EDBERT delivers this performance with fewer parameters and reduced training time. Overall, EDBERT offers a scalable, high-performing, and resource-efficient solution for predicting patient disposition in emergency settings, supporting timely and informed clinical decision-making. Future work will explore the generalisability of EDBERT across multiple hospital sites, the incorporation of additional patient features, uncertainty-aware prediction, and real-time deployment in ED triage workflows.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/a19090729/s1. Table S1: Performance of BERT-Base models with varying numbers of Transformer layers; Table S2: Performance of EDBERT with and without MLM pre-training; Figure S1: ROC curves of BERT-Base configurations; Figure S2: Parameter counts and training time per epoch; Figure S3: Confusion matrices of BERT-Base configurations; Figure S4: Performance of EDBERT without MLM pre-training.

Author Contributions

Conceptualisation, M.A.R., D.A. and H.A.; Methodology, M.A.H.; Formal Analysis, M.A.H.; Investigation, M.A.R., D.A. and H.A.; Data Curation, S.F. and H.A.; Writing—Original Draft Preparation, M.D.W.; Writing—Review and Editing, I.S., K.C., S.F., M.A.R., S.K., M.P., D.A. and H.A.; Visualisation, M.A.H.; Funding Acquisition, M.A.R. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the La Trobe Business School, La Trobe University, through the LBS 2024 SMALL GRANT, grant number 4.1165.09.

Institutional Review Board Statement

This study was conducted in accordance with the Declaration of Helsinki 1964 as revised in 2008, and the National Health and Medical Research Council National Statement on Ethical Conduct in Human Research (2023) Guidelines, and approved by the St Vincent’s Hospital Melbourne Human Research Ethics Committee (Project ID: 85929). Each site has received governance approval and site-specific assessment.

Informed Consent Statement

Due to the retrospective nature of this study, waiver of consent was approved via the HREC committee.

Data Availability Statement

De-identified data will be provided upon written request to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
EDEmergency Department
BERTBidirectional Encoder Representations from Transformers
EHRElectronic Health Record
EDBERTEmergency Department Bidirectional Encoder Representations from Transformers
LOSLength of Stay
NLPNatural Language Processing
LLMLarge Language Model
MLMMasked Language Modelling
AUROCArea Under the Receiver Operating Characteristic Curve
ROCReceiver Operating Characteristic
AUCArea Under the Curve

References

  1. Payne, K.; Risi, D.; O’Hare, A.; Binks, S.; Curtis, K. Factors that contribute to patient length of stay in the emergency department: A time in motion observational study. Australas. Emerg. Care 2023, 26, 321–325. [Google Scholar] [CrossRef] [Scilit]
  2. Australian Institute of Health and Welfare (AIHW). Emergency Department Care. 2023. Available online: https://www.aihw.gov.au/reports-data/myhospitals/sectors/emergency-department-care (accessed on 26 November 2024).
  3. Australian Institute of Health and Welfare. Emergency Department Care 2016–17: Australian Hospital Statistics. 2017. Available online: https://www.aihw.gov.au/reports/hospitals/ahs-2016-17-emergency-department-care/ (accessed on 26 November 2024).
  4. Farimani, R.M.; Karim, H.; Atashi, A.; Tohidinezhad, F.; Bahaadini, K.; Abu-Hanna, A.; Eslami, S. Models to predict length of stay in the emergency department: A systematic literature review and appraisal. BMC Emerg. Med. 2024, 24, 54. [Google Scholar] [CrossRef] [Scilit]
  5. Rahman, M.A.; Lim, D.Z.; Davoren, M.; Lok, I.; Rahman, S.; Hough, P.; Mosa, T.; Begum, S. Mapping the patient journey: Utilizing clinical informatics for a conceptual approach to identify aspects of emergency department access block. Netw. Model. Anal. Health Inform. Bioinform. 2024, 13, 54. [Google Scholar] [CrossRef] [Scilit]
  6. Kuo, K.M.; Lin, Y.L.; Chang, C.S.; Kuo, T.J. An ensemble model for predicting dispositions of emergency department patients. BMC Med. Inform. Decis. Mak. 2024, 24, 105. [Google Scholar] [CrossRef] [Scilit]
  7. Gurazada, S.G.; Gao, S.; Burstein, F.; Buntine, P. Predicting patient length of stay in Australian emergency departments using Data Mining. Sensors 2022, 22, 4968. [Google Scholar] [CrossRef] [Scilit]
  8. Jones, S.; Moulton, C.; Swift, S.; Molyneux, P.; Black, S.; Mason, N.; Oakley, R.; Mann, C. Association between delays to patient admission from the emergency department and all-cause 30-day mortality. Emerg. Med. J. 2022, 39, 168–173. [Google Scholar] [CrossRef] [Scilit]
  9. Janerka, C.; Leslie, G.D.; Gill, F.J. Patient experience of emergency department triage: An integrative review. Int. Emerg. Nurs. 2024, 74, 101456. [Google Scholar] [CrossRef] [Scilit]
  10. Sterling, N.W.; Patzer, R.E.; Di, M.; Schrager, J.D. Prediction of emergency department patient disposition based on natural language processing of triage notes. Int. J. Med. Inform. 2019, 129, 184–188. [Google Scholar] [CrossRef] [Scilit]
  11. Arvig, M.D.; Mogensen, C.B.; Skjøt-Arkil, H.; Johansen, I.S.; Rosenvinge, F.S.; Lassen, A.T. Chief complaints, underlying diagnoses, and mortality in adult, non-trauma emergency department visits: A population-based, multicenter cohort study. West. J. Emerg. Med. 2022, 23, 855. [Google Scholar] [CrossRef] [Scilit]
  12. Group, M.T. Emergency Triage; John Wiley & Sons: Hoboken, NJ, USA, 2008. [Google Scholar]
  13. Wang, D.; Zhang, S. Large language models in medical and healthcare fields: Applications, advances, and challenges. Artif. Intell. Rev. 2024, 57, 299. [Google Scholar] [CrossRef] [Scilit]
  14. Luo, X.; Deng, Z.; Yang, B.; Luo, M.Y. Pre-trained language models in medicine: A survey. Artif. Intell. Med. 2024, 154, 102904. [Google Scholar] [CrossRef] [Scilit]
  15. Devlin, J.; Chang, M.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv 2018, arXiv:1810.04805. [Google Scholar]
  16. Alsentzer, E.; Murphy, J.R.; Boag, W.; Weng, W.H.; Jin, D.; Naumann, T.; McDermott, M. Publicly available clinical BERT embeddings. arXiv 2019, arXiv:1904.03323. [Google Scholar]
  17. Lee, J.; Yoon, W.; Kim, S.; Kim, D.; Kim, S.; So, C.H.; Kang, J. BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 2020, 36, 1234–1240. [Google Scholar] [CrossRef] [Scilit]
  18. Tahayori, B.; Chini-Foroush, N.; Akhlaghi, H. Advanced natural language processing technique to predict patient disposition based on emergency triage notes. Emerg. Med. Australas. 2021, 33, 480–484. [Google Scholar] [CrossRef] [Scilit]
  19. Ferri, P.; Sáez, C.; Félix-De Castro, A.; Juan-Albarracín, J.; Blanes-Selva, V.; Sánchez-Cuesta, P.; García-Gómez, J.M. Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch. Artif. Intell. Med. 2021, 117, 102088. [Google Scholar] [CrossRef] [Scilit]
  20. Chang, Y.H.; Lin, Y.C.; Huang, F.W.; Chen, D.M.; Chung, Y.T.; Chen, W.K.; Wang, C.C. Using machine learning and natural language processing in triage for prediction of clinical disposition in the emergency department. BMC Emerg. Med. 2024, 24, 237. [Google Scholar] [CrossRef] [Scilit]
  21. Chen, C.H.; Hsieh, J.G.; Cheng, S.L.; Lin, Y.L.; Lin, P.H.; Jeng, J.H. Emergency department disposition prediction using a deep neural network with integrated clinical narratives and structured data. Int. J. Med. Inform. 2020, 139, 104146. [Google Scholar] [CrossRef] [Scilit]
  22. Boughorbel, S.; Jarray, F.; Homaid, A.A.; Niaz, R.; Alyafei, K. Multi-Modal Perceiver Language Model for Outcome Prediction in Emergency Department. arXiv 2023, arXiv:2304.01233. [Google Scholar]
  23. Peng, Y.; Yan, S.; Lu, Z. Transfer learning in biomedical natural language processing: An evaluation of BERT and ELMo on ten benchmarking datasets. In Proceedings of the 18th BioNLP Workshop and Shared Task, Florence, Italy, 1 August 2019; pp. 58–65. [Google Scholar]
  24. Chen, T.Y.; Huang, T.Y.; Chang, Y.C. Using a clinical narrative-aware pre-trained language model for predicting emergency department patient disposition and unscheduled return visits. J. Biomed. Inform. 2024, 155, 104657. [Google Scholar] [CrossRef] [Scilit]
  25. Zhang, B.; Li, Z.; Ma, Q.; Zhang, J.; Xiang, Z.; Jiang, D. Silhouette-Based Cross-View Motion Gait Recognition via a Multi-Scale Temporal Difference Unit. Electronics 2026, 15, 2512. [Google Scholar] [CrossRef] [Scilit]
  26. Sharaf, S.; Anoop, V. An analysis on large language models in healthcare: A case study of BioBERT. arXiv 2023, arXiv:2310.07282. [Google Scholar]
  27. Koroteev, M.V. BERT: A review of applications in natural language processing and understanding. arXiv 2021, arXiv:2103.11943. [Google Scholar]
  28. Williams, E.L.; Huynh, D.; Estai, M.; Sinha, T.; Summerscales, M.; Kanagasingam, Y. Predicting inpatient admissions from emergency department triage using machine learning: A systematic review. Mayo Clin. Proc. Digit. Health 2025, 3, 100197. [Google Scholar] [CrossRef] [Scilit]
  29. Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. arXiv 2014, arXiv:1412.6980. [Google Scholar]
  30. Rasmy, L.; Xiang, Y.; Xie, Z.; Tao, C.; Zhi, D. Med-BERT: Pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction. npj Digit. Med. 2021, 4, 86. [Google Scholar] [CrossRef] [Scilit]
  31. Sanh, V.; Debut, L.; Chaumond, J.; Wolf, T. DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv 2019, arXiv:1910.01108. [Google Scholar]
  32. Jiao, X.; Yin, Y.; Shang, L.; Jiang, X.; Chen, X.; Li, L.; Wang, F.; Liu, Q. Tinybert: Distilling bert for natural language understanding. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2020, Online, 16 November 2020; pp. 4163–4174. [Google Scholar]
  33. Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; Soricut, R. Albert: A lite bert for self-supervised learning of language representations. arXiv 2019, arXiv:1909.11942. [Google Scholar]
  34. Abdulai, A.S.B.; Storm, J.; Ehrlich, M. “I don’t know”: An uncertainty-aware machine learning model for predicting patient disposition at emergency department triage. Int. J. Med. Inform. 2025, 201, 105957. [Google Scholar] [CrossRef] [Scilit]
  35. Barak-Corren, Y.; Agarwal, I.; Michelson, K.A.; Lyons, T.W.; Neuman, M.I.; Lipsett, S.C.; Kimia, A.A.; Eisenberg, M.A.; Capraro, A.J.; Levy, J.A.; et al. Prediction of patient disposition: Comparison of computer and human approaches and a proposed synthesis. J. Am. Med. Inform. Assoc. 2021, 28, 1736–1745. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Eight-layer custom BERT encoder stack of the EDBERT model.
Figure 1. Eight-layer custom BERT encoder stack of the EDBERT model.
Algorithms 19 00729 g001
Figure 2. Overall architecture of the proposed EDBERT model, consisting of two 8-layer custom encoder stacks to extract contextual and semantic features from triage notes and presenting complaints, followed by a fusion network that integrates three distinct embeddings to generate a unified representation for final patient disposition prediction.
Figure 2. Overall architecture of the proposed EDBERT model, consisting of two 8-layer custom encoder stacks to extract contextual and semantic features from triage notes and presenting complaints, followed by a fusion network that integrates three distinct embeddings to generate a unified representation for final patient disposition prediction.
Algorithms 19 00729 g002
Figure 3. Performance of the proposed EDBERT model on the test set: (a) ROC curve (AUC = 0.91) and (b) confusion matrix.
Figure 3. Performance of the proposed EDBERT model on the test set: (a) ROC curve (AUC = 0.91) and (b) confusion matrix.
Algorithms 19 00729 g003
Table 1. Architectural differences between BERT-BASE and EDBERT’s customised BERT encoder stack.
Table 1. Architectural differences between BERT-BASE and EDBERT’s customised BERT encoder stack.
ComponentBERT-BASEEDBERT
Hidden Size (H)768512
Number of Transformer Layers (L)128
Number of Multi-Head Self-Attention Heads (A)128
Maximum Position Embeddings512512
Intermediate Size (I)30722048
Table 2. Hyperparameter values used during model training.
Table 2. Hyperparameter values used during model training.
HyperparameterValue
Learning Rate 2 × 10 5
Weight Decay0.01
Train Batch Size64
Test Batch Size32
Dropout Percentage0.5
Table 3. Performance comparison of EDBERT and the standard BERT-Base model, evaluated using accuracy and macro-averaged precision, recall, and F1-score percentages. Values in parentheses represent the corresponding 95% confidence intervals (CIs). Confidence intervals were obtained by parametric bootstrap over the class-conditional counts. BERT-Base was evaluated on N = 114,029 and EDBERT on N = 118,957 .
Table 3. Performance comparison of EDBERT and the standard BERT-Base model, evaluated using accuracy and macro-averaged precision, recall, and F1-score percentages. Values in parentheses represent the corresponding 95% confidence intervals (CIs). Confidence intervals were obtained by parametric bootstrap over the class-conditional counts. BERT-Base was evaluated on N = 114,029 and EDBERT on N = 118,957 .
ModelAccuracyPrecisionRecallF1-Score
BERT-Base (12 layers)82.1981.2280.1680.62
(81.97–82.41)(80.98–81.46)(79.91–80.41)(80.38–80.86)
EDBERT (8 layers)83.2682.1481.8581.99
(83.05–83.47)(81.92–82.36)(81.63–82.07)(81.77–82.21)
Table 4. Performance comparison of different BERT variants, evaluated using accuracy and macro-averaged precision, recall, and F1-score. Values in parentheses represent estimated 95% confidence intervals (CIs).
Table 4. Performance comparison of different BERT variants, evaluated using accuracy and macro-averaged precision, recall, and F1-score. Values in parentheses represent estimated 95% confidence intervals (CIs).
ModelAccuracyPrecisionRecallF1-Score
BERT-Base81.70(81.48–81.92)80.73(80.51–80.95)79.68(79.45–79.91)80.13(79.90–80.36)
MedBERT82.19(81.97–82.41)81.22(81.00–81.44)80.16(79.93–80.39)80.16(79.93–80.39)
DistilBERT81.60(81.38–81.82)79.75(79.52–79.98)79.58(79.35–79.81)79.13(78.90–79.36)
TinyBERT80.90(80.68–81.12)79.75(79.52–79.98)78.90(78.67–79.13)79.13(78.90–79.36)
ALBERT81.40(81.18–81.62)79.75(79.52–79.98)79.39(79.16–79.62)79.13(78.90–79.36)
Table 5. Representative EDBERT misclassification examples.
Table 5. Representative EDBERT misclassification examples.
AgePresenting ComplaintTriage NotePredicted Disposition (Probability Score)Actual Disposition
63GastrointestinalDue for ascitic tap on 7/10. Awoke this AM with worsening SOB and severe abdominal distension. Last tap on 20/09 drained 5L. PMHx: liver CA, CKD, HTN, T2DM, IHD.Admission (98.3%)Discharge
20Other MedicalRight hand burned on hot oil; partial-thickness/superficial dermal burn with blistering to fingers. First aid applied. PHx: nil.Discharge (98.1%)Admission
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Hossain, M.A.; Chavinda, K.; Woodbright, M.D.; Senadheera, I.; Kogul, S.; Freeman, S.; Putland, M.; Akhlaghi, H.; Alahakoon, D.; Rahman, M.A. EDBERT: Predicting Emergency Department Disposition Using a BERT-Based Architecture. Algorithms 2026, 19, 729. https://doi.org/10.3390/a19090729

AMA Style

Hossain MA, Chavinda K, Woodbright MD, Senadheera I, Kogul S, Freeman S, Putland M, Akhlaghi H, Alahakoon D, Rahman MA. EDBERT: Predicting Emergency Department Disposition Using a BERT-Based Architecture. Algorithms. 2026; 19(9):729. https://doi.org/10.3390/a19090729

Chicago/Turabian Style

Hossain, Md Ali, Krishan Chavinda, Mitchell D. Woodbright, Isuru Senadheera, Srikandabala Kogul, Sam Freeman, Mark Putland, Hamed Akhlaghi, Damminda Alahakoon, and Md Anisur Rahman. 2026. "EDBERT: Predicting Emergency Department Disposition Using a BERT-Based Architecture" Algorithms 19, no. 9: 729. https://doi.org/10.3390/a19090729

APA Style

Hossain, M. A., Chavinda, K., Woodbright, M. D., Senadheera, I., Kogul, S., Freeman, S., Putland, M., Akhlaghi, H., Alahakoon, D., & Rahman, M. A. (2026). EDBERT: Predicting Emergency Department Disposition Using a BERT-Based Architecture. Algorithms, 19(9), 729. https://doi.org/10.3390/a19090729

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop