MMAC-Net: A Multi-Modal Multi-Label Attention-Based Deep Learning Approach for Automated ICD-9 Coding of Rare Disease Admissions from Electronic Health Records
Abstract
1. Introduction
1.1. Problem Statement
1.2. Motivation
1.3. Rare Disease Classification
1.4. State-of-the-Art in Automated Coding
- Transformer Models: Aden et al. [12] achieved state-of-the-art results using pre-trained ClinicalBERT [13] combined with RNNs and LSTMs. However, their evaluation was primarily focused on the top-10 and top-50 most common codes, achieving high precision (0.87) for common diseases but leaving the performance on rare codes less explored.
1.5. Research Gap and Questions
- RQ1. Does adding prescription and microbiology modalities to a per-label attention text encoder improve ICD-9 code assignment for rare codes, relative to a matched text-only model evaluated under an identical protocol?
- RQ2. Is any such improvement concentrated in the long tail, that is, does it persist when results are stratified by training frequency, including few-shot codes?
- RQ3. Can this be achieved at a parameter and inference cost below that of transformer-based coders of comparable accuracy?
1.6. Organization of the Article
- Section 2 (Related Works) describes the recent advances in identification of rare diseases using various technology and methods, especially with machine learning and deep learning-based approaches, and closes with a structured comparison of prior systems.
- Section 3 (Methodology) details the proposed pipeline, including the dataset and rare-subset construction, data pre-processing, feature fusion of structured and unstructured MIMIC-III data, and the architecture of the attention-based CNN model used for multi-label classification.
- Section 4 (Evaluation) defines the evaluation framework utilized during the experiments to measure the efficiency of the model. This section discusses the loss function and evaluation metrics used in details.
- Section 5 (Experimental Design) describes the experimental setup, including hyperparameter tuning, the computational environment, and the specific metrics used to evaluate performance on imbalanced data (Precision@k, Macro-F1 vs. Micro-F1).
- Section 6 (Experimental Results and Comparisons) presents the quantitative results of the model, analyzing the discrepancy between Micro-AUC and Macro-F1 scores, and discussing the model’s behavior regarding over-fitting and validation loss, together with validity checks and a stratification of performance by label frequency.
- Section 7 (Results Analysis) discusses about the actual impact of the model components and multi-modal efficiency of the model, and presents attention-based explanations and two contrasting case studies.
- Section 8 (Conclusions and Future Work) summarizes the contributions, states the limitations and ethical considerations, and outlines future research directions, such as the integration of graph neural networks and ablation studies for enhanced explainability.
2. Related Works
2.1. Literature Review
2.2. Traditional Machine Learning Approaches
2.3. Deep Learning
2.4. Graph-Based Deep Learning
2.5. Addressing Data Scarcity: Few-Shot and Synthetic Learning
2.6. Critical Comparison of Prior Work
3. Methodology
3.1. Dataset
3.2. Dataset Modalities
3.3. ICD Coding
3.4. Modalities and Features
- Clinical Notes: Extracted from NOTEEVENTS.csv, specifically filtering for discharge summaries. These free-text narratives provide the primary unstructured data source for the model.
- Prescriptions: Structured data loaded from PRESCRIPTIONS.csv. Key features extracted include drug_type, drug name, prod_strength (product strength), dose_val_rx (dosage value), and route of administration.
- Microbiology Events: Extracted from MICROBIOLOGYEVENTS.csv, capturing infectious disease data. Features include org_itemid (organism), ad_itemid (antibiotic), dilution_value, and interpretation (susceptibility).
3.5. Pre-Processing
- Section-header handling. Section headers are detected using the header lexicon. Headers are retained by default.
- Normalization. All text is lower-cased.
- Noise reduction. Punctuation and numeric characters are removed.
- Stopword filtering. The NLTK English stopword list is applied.
- Lemmatization. The nltk.WordNetLemmatizer reduces words to their base form.
- Tokenization and truncation. A SpaceTokenizer is applied and documents are truncated or padded to 1500 tokens, truncating from the end of the document.
- Vocabulary. The vocabulary is constructed with a minimum token frequency of 3 on the training split only.
3.6. Feature Conversion
- Text Embeddings: Clinical notes are truncated to a maximum sequence length and converted into indices based on a vocabulary loaded from pre-trained Word2Vec embeddings.
- Categorical Features: Auxiliary categorical features (e.g., drug types, organism) are mapped to integer indices, handling unknown values by assigning them to a specific index.
- Numerical Features: Numerical attributes are normalized using Z-score standardization (subtracting the mean and dividing by the standard deviation) to ensure numerical stability during training. Two sample numerical attributes are the prescription dosage value and the microbiology dilution value, both of which are heavy-tailed and contain implausible extreme values arising from unit inconsistencies and free-text entry. Z-score standardization was preferred to min-max scaling for this reason: a min-max range is determined entirely by the two extreme order statistics, so a single erroneous dosage compresses the whole bulk of the distribution towards zero, whereas standardization is driven by the mean and standard deviation and degrades more gracefully.
3.7. Design Rationale for Multimodal Fusion
3.8. Model Architecture
- Embeddings: A learnable embedding layer initialized with pre-trained Word2Vec weights converts token indices into dense vectors. A dropout layer is applied immediately after embedding to prevent overfitting.
- Convolutional Encoder: A 1D convolutional layer iterates over the document to capture local n-gram patterns. The architecture supports variable kernel sizes and depths, using Xavier uniform initialization for weights.
- Per-Label Attention Mechanism: The model employs a specialized attention mechanism where a matrix U projects the convolutional outputs to the output space of size equal to the number of classes. A softmax function is applied to generate attention weights , which are used to compute a weighted sum of the document representation, creating a unique document vector for each label.
- Depthwise Convolution Layers: The transition from standard 1D convolutions to depthwise separable convolutions (DWConv) in the deeper blocks is a design choice intended for balancing high-capacity feature extraction with parameter efficiency. By decoupling the sequential spatial filtering from the channel-wise feature mixing, depthwise convolutions drastically reduce both the computational overhead and the risk of overfitting.
- Squeeze and Excitation: The squeeze-and-excitation (SE) blocks integrated into the convolutional pipeline serve as a dynamic channel-wise attention mechanism, fundamentally enhancing the representational power of the MMACNet architecture. Rather than treating all extracted feature maps equally, the SE block explicitly models the inter-dependencies between channels to perform adaptive feature re-calibration.
- Tabular Fusion Branch: A distinct branch processes structured data. Categorical features are passed through separate embedding layers, while numerical features undergo batch normalization. These representations are concatenated and processed by a multi-layer perceptron (MLP) before being fused with the text-based representations.
- Output Classifier: The final classification is performed by a linear layer that maps the fused document-label representations to logits, followed by a sigmoid activation depending on the loss configuration.
3.9. Model Training
- 1.
- Initialization: The model optimizer deployed was Adam, and loss function are initialized based on the provided configuration.
- 2.
- Training Loop: The model iterates through the training dataset in batches. For each batch, gradients are computed via backpropagation, and the optimizer updates model parameters.
- 3.
- Checkpointing: The checkpointing scheme monitors performance and saves the model state at regular intervals or when a best metric () value is achieved.
- 4.
- Regularization: A label-description regularization loss is added to the objective function, enforcing similarity between the learned attention vectors and the embeddings of the ICD code descriptions.
4. Evaluation
4.1. Loss Function
- N is the batch size.
- L is the total number of ICD classes.
- is the ground truth binary label for class l and patient i.
- is the raw output logit from the model.
- is the sigmoid activation function.
- is the regularization coefficient (config parameter lmbda).
- is the learned embedding vector for label l.
- is the fixed description embedding of label l, obtained by averaging the Word2Vec vectors of the words of its ICD-9 description.
4.2. Evaluation Metrics
- Precision@k (P@k): The proportion of relevant labels in the top-k predictions, that is, the number of correctly ranked labels within the top k divided by k. This is crucial for clinical decision support, where a physician typically reviews only the top few suggestions.
- Micro-F1 Score: The harmonic mean of precision and recall calculated globally by counting the total true positives, false negatives, and false positives. This metric biases towards common disease classes.
- Macro-F1 Score: The unweighted mean of the F1 scores calculated for each label individually. This metric is particularly important for this study as it treats rare diseases equally to common ones, highlighting the model’s performance on the long tail of the distribution.where is the F1 score for class c.
- Macro-AUC is the arithmetic mean of the area under the ROC curve (AUC) calculated for each class individually.
- Micro-AUC is the area under the receiver operating characteristic curve calculated using the global true positive rate () and false positive rate (), defined as
5. Experimental Design
5.1. Data Preparation and Preprocessing
- Clinical Narratives: Discharge summaries extracted from NOTEEVENTS.csv.gz.
- Diagnostic and Procedural Codes: ICD-9 codes derived from DIAGNOSES_ICD.csv.gz and PROCEDURES_ICD.csv.gz.
- Pharmacological Data: Medication records from PRESCRIPTIONS.csv.gz, including drug types, dosages, and administration routes.
- Microbiology: Infectious disease data from MICROBIOLOGYEVENTS.csv.gz, including organism identifiers and antibiotic susceptibility.
5.2. Hyperparameter Tuning
5.3. Training Protocol
5.4. Evaluation Framework
- Precision@k: Specifically and . This metric is prioritized as the stopping criterion, reflecting the clinical need for accurate top-ranked suggestions.
- F1 Scores: Both Macro-F1 and Micro-F1 are calculated to assess performance across rare (macro) and frequent (micro) classes.
- AUC Scores: Macro-AUC and Micro-AUC provide an aggregate measure of classification thresholds.
5.5. Computational Environment and Cost
6. Experimental Results and Comparison
6.1. Ablation Study
6.2. Comparison with State-of-the-Art
6.3. Comparative Analysis Between Rare Subset vs. The Whole Dataset
7. Results Analysis
7.1. Impact of Multi-Modal Data Fusion
- Baseline Performance (Notes Only): The model relying solely on unstructured discharge summaries yielded the lowest performance across all metrics, with a Macro-F1 of and a Precision@8 of . This confirms that while clinical narratives contain rich information, they are often insufficient on their own to capture the full clinical picture required for accurate coding, particularly for rare conditions. This insufficiency occurs because physicians frequently omit objective, structured data, such as specific laboratory thresholds, vital signs, or demographic baselines from their narrative summaries, as these are already accessible elsewhere in the electronic health record (EHR). Furthermore, the high variance and noise inherent to free-text notes exacerbate the challenge of identifying rare diseases, which already suffer from a lack of training examples.
- Synergy of Structured Data: The integration of structured data significantly enhanced the model’s predictive capability. The combination of Notes and Tabular resulted in a substantial increase in Macro-AUC (0.8106 vs. 0.7670) and nearly tripled the Macro-F1 score (0.0295 vs. 0.0109). This suggests that numerical features, such as the normalized prescription dosage and microbiology dilution values processed by the tabular branch, provide critical signals that help disambiguate complex code assignments.
- Optimal Configuration: The fully integrated model (Notes, Tabular and Categorical) achieved the highest performance across the majority of key metrics. It reached a peak Micro-F1 of and a Precision@8 of . This configuration leverages the late fusion strategy, where categorical embeddings (e.g., medications, microbiology events) and numerical features are concatenated with the text representation before the final classification layer.
7.2. Effective Late Fusion Mechanism
7.3. Behavior in High-Dimensional Classification
7.4. Explainability of Predictions
7.5. Case Studies
8. Conclusions and Future Work
- Multi-Modal Fusion Architecture: We successfully engineered a late fusion neural architecture that processes unstructured clinical text via a deep convolutional attention mechanism while simultaneously encoding structured categorical and numerical features through parallel perceptron branches.
- Validation of Structured Features: Through a rigorous ablation study, we demonstrated that the inclusion of structured data significantly enhances predictive performance. The fully integrated model (Notes, Tabular and Categorical) achieved a Macro-F1 of , nearly tripling the performance of the text-only baseline (), thereby validating the importance of multi-modal signals in disambiguating complex diagnoses.
- Competitive Performance: The proposed framework achieved a Micro-AUC of and a Precision@8 of on the full MIMIC-III dataset. These results are higher than the figures reported for baselines such as CAML and DCAN, though as explained in Section 3 and in the caption of Table 7, those figures are transcribed from their source publications rather than produced under our protocol, so the comparison is indicative and does not establish a ranking.
8.1. Limitations
8.2. Ethical Considerations and Human-in-the-Loop Use
8.3. Future Work
- Hyperparameter search with cross-validation. A grid or random search over depth, embedding dimension, kernel size and dropout under k-fold cross-validation, with sensitivity curves, to place the configuration of Table 4 on a systematic footing.
- Imbalance-aware objectives. Weighted binary cross-entropy, focal, asymmetric and distribution-balanced losses, compared on the rare subset.
- External validation. Evaluation on MIMIC-IV in zero-shot and fine-tuned settings, noting that MIMIC-IV originates from the same institution and therefore provides temporal rather than cross-site validation; genuinely multi-site data would be required for the latter.
- Patient-disjoint split. Retraining under a SUBJECT_ID-disjoint partition to quantify memorization across repeat admissions.
- Temporal ablation. Censoring prescriptions and microbiology events at successive cut-offs to quantify how much of the multi-modal gain derives from end-of-stay evidence.
- Prospective formulation. Defining a prediction time before diagnosis and excluding evidence recorded after it, including the discharge summary itself, to move from retrospective coding towards decision support during the admission.
8.4. Constraints Observed
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A. Pseudocode for the MMAC-Net Architecture
- Notation.
| Algorithm A1 MMAC-Net: multimodal attentional convolutional forward pass |
| Require: token ids ; categorical codes ; numerical vector |
| Ensure: per-label logits |
|
| Algorithm A2 MMAC-Net: optimization step on one mini-batch |
| Require: batch ; regularization weight |
|
References
- Office of the Commissioner. Rare Diseases at FDA, n.d. Available online: https://www.fda.gov/patients/rare-diseases-fda (accessed on 6 November 2025).
- Wang, C.M.; Whiting, A.H.; Rath, A.; Anido, R.; Ardigò, D.; Baynam, G.; Dawkins, H.; Hamosh, A.; Le Cam, Y.; Malherbe, H.; et al. Operational description of rare diseases: A reference to improve the recognition and visibility of rare diseases. Orphanet J. Rare Dis. 2024, 19, 334. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rezaei, M.; Näppi, J.J.; Bischl, B.; Yoshida, H. Bayesian uncertainty estimation for detection of long-tailed and unseen conditions in medical images. J. Med. Imaging 2023, 10, 054501. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wu, Z.; Guo, K.; Luo, E.; Wang, T.; Wang, S.; Yang, Y.; Zhu, X.; Ding, R. Medical long-tailed learning for imbalanced data: Bibliometric analysis. Comput. Methods Programs Biomed. 2024, 247, 108106. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- INSERM. Orphanet: An Online Rare Disease and Orphan Drug Database. 1999. Available online: http://www.orpha.net (accessed on 21 November 2025).
- Government of Canada, CIHR. Rare Disease Research Finds Answers for Families. 2025. Available online: https://cihr-irsc.gc.ca/e/54515.html (accessed on 6 November 2025).
- Bauskis, A.; Strange, C.; Molster, C.; Fisher, C. The diagnostic odyssey: Insights from parents of children living with an undiagnosed condition. Orphanet J. Rare Dis. 2022, 17, 233. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cho, J.; Joo, Y.S.; Yoon, J.G.; Lee, S.B.; Kim, S.Y.; Chae, J.H.; Kwon, Y.J. Characterizing Families of Pediatric Patients with Rare Diseases and Their Diagnostic Odysseys: A Comprehensive Survey Analysis from a Single Tertiary Center in Korea. Ann. Child Neurol. 2024, 32, 167–175. [Google Scholar] [CrossRef] [Scilit]
- Adachi, T.; El-Hattab, A.W.; Jain, R.; Nogales Crespo, K.A.; Quirland Lazo, C.I.; Scarpa, M.; Summar, M.; Wattanasirichaigoon, D. Enhancing equitable access to rare disease diagnosis and treatment around the world: A review of evidence, policies, and challenges. Int. J. Environ. Res. Public Health 2023, 20, 4732. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Johnson, A.; Pollard, T.; Shen, L.; Lehman, L.W.; Feng, M.; Ghassemi, M.; Moody, B.; Szolovits, P.; Celi, L.; Mark, R. MIMIC-III, a freely accessible critical care database. Sci. Data 2016, 3, 160035. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- World Health Organization. International Classification of Diseases (ICD). 2022. Available online: https://www.who.int/standards/classifications/classification-of-diseases (accessed on 6 November 2025).
- Aden, I.; Child, C.H.T.; Reyes-Aldasoro, C.C. International Classification of Diseases Prediction from MIMIIC-III Clinical Text Using Pre-Trained ClinicalBERT and NLP Deep Learning Models Achieving State of the Art. Big Data Cogn. Comput. 2024, 8, 47. [Google Scholar] [CrossRef] [Scilit]
- Huang, K.; Altosaar, J.; Ranganath, R. ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission. arXiv 2020, arXiv:1904.05342. [Google Scholar] [CrossRef] [Scilit]
- Merchant, A.M.; Shenoy, N.; Lanka, S.; Kamath, S. Ensemble neural models for ICD code prediction using unstructured and structured healthcare data. Heliyon 2024, 10, e36569. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Elvas, L.B.; Almeida, A.; Ferreira, J.C. Natural language processing in medical text processing: A scoping literature review. Int. J. Med. Inform. 2025, 204, 106049. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Douglas, J.C.; Gan, Y.; Hachey, B.; Kummerfeld, J.K. Less is More: Explainable and Efficient ICD Code Prediction with Clinical Entities. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Association for Computational Linguistics: Vienna, Austria, 2025; pp. 30835–30847. [Google Scholar] [CrossRef] [Scilit]
- Yang, X.; Bian, J.; Hogan, W.R.; Wu, Y. Clinical concept extraction using transformers. J. Am. Med. Inform. Assoc. 2020, 27, 1935–1942. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schubach, M.; Re, M.; Robinson, P.N.; Valentini, G. Imbalance-aware machine learning for predicting rare and common disease-associated non-coding variants. Sci. Rep. 2017, 7, 2959. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rigg, J.; Lodhi, H.; Nasuti, P. Using Machine Learning to Detect Patients with Undiagnosed Rare Diseases: An Application of Support Vector Machines to A Rare Oncology Disease. Value Health 2015, 18, A705. [Google Scholar] [CrossRef] [Scilit]
- Park, S.H.; Song, S.H.; Burton, F.; Arsan, C.; Jobst, B.; Feldman, M. Machine learning characterization of a rare neurologic disease via electronic health records: A proof-of-principle study on stiff person syndrome. BMC Neurol. 2024, 24, 272. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cohen, A.M.; Chamberlin, S.; Deloughery, T.; Nguyen, M.; Bedrick, S.; Meninger, S.; Ko, J.J.; Amin, J.J.; Wei, A.J.; Hersh, W. Detecting rare diseases in electronic health records using machine learning and knowledge engineering: Case study of acute hepatic porphyria. PLoS ONE 2020, 15, e0235574. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Crisafulli, S.; Fontana, A.; L’Abbate, L.; Vitturi, G.; Cozzolino, A.; Gianfrilli, D.; Martino, M.C.D.; Amico, B.; Combi, C.; Trifirò, G. Machine learning-based algorithms applied to drug prescriptions and other healthcare services in the Sicilian claims database to identify acromegaly as a model for the earlier diagnosis of rare diseases. Sci. Rep. 2024, 14, 6186. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Revel-Vilk, S.; Shalev, V.; Gill, A.; Paltiel, O.; Manor, O.; Tenenbaum, A.; Azani, L.; Chodick, G. Assessing the diagnostic utility of the Gaucher Earlier Diagnosis Consensus (GED-C) scoring system using real-world data. Orphanet J. Rare Dis. 2024, 19, 71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Tenenbaum, A.; Revel-Vilk, S.; Gazit, S.; Roimi, M.; Gill, A.; Gilboa, D.; Paltiel, O.; Manor, O.; Shalev, V.; Chodick, G. A machine learning model for early diagnosis of type 1 Gaucher disease using real-life data. J. Clin. Epidemiol. 2024, 175, 111517. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Garg, R.; Dong, S.; Shah, S.; Jonnalagadda, S.R. A Bootstrap Machine Learning Approach to Identify Rare Disease Patients from Electronic Health Records. arXiv 2016, arXiv:1609.01586. [Google Scholar] [CrossRef] [Scilit]
- Ehsani-Moghaddam, B.; Queenan, J.A.; MacKenzie, J.; Birtwhistle, R.V. Mucopolysaccharidosis type II detection by Naïve Bayes Classifier: An example of patient classification for a rare disease using electronic medical records from the Canadian Primary Care Sentinel Surveillance Network. PLoS ONE 2018, 13, e0209018. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kim, J.; Sharma, A.; Shanbhogue, S.; Weiss, J.; Ravikumar, P. AnEMIC: A Framework for Benchmarking ICD Coding Models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations; Che, W., Shutova, E., Eds.; Association for Computational Linguistics: Abu Dhabi, United Arab Emerites, 2022; pp. 109–120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rolando, M.; Raggio, V.; Naya, H.; Spangenberg, L.; Cagnina, L. A labeled medical records corpus for the timely detection of rare diseases using machine learning approaches. Sci. Rep. 2025, 15, 6932. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, S.; Varghese, P.; Stephenson, E.; Tu, K.; Gronsbell, J. Machine learning approaches for electronic health records phenotyping: A methodical review. J. Am. Med. Inform. Assoc. 2023, 30, 367–381. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jia, J.; Wang, R.; An, Z.; Guo, Y.; Ni, X.; Shi, T. RDAD: A machine learning system to support phenotype-based rare disease diagnosis. Front. Genet. 2018, 9, 587. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Faviez, C.; Vincent, M.; Garcelon, N.; Boyer, O.; Knebelmann, B.; Heidet, L.; Saunier, S.; Chen, X.; Burgun, A. Performance and clinical utility of a new supervised machine-learning pipeline in detecting rare ciliopathy patients based on deep phenotyping from electronic health records and semantic similarity. Orphanet J. Rare Dis. 2024, 19, 55. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shen, F.; Liu, S.; Wang, Y.; Wen, A.; Wang, L.; Liu, H. Utilization of Electronic Medical Records and Biomedical Literature to Support the Diagnosis of Rare Diseases Using Data Fusion and Collaborative Filtering Approaches. JMIR Med. Inform. 2018, 6, e11301. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Colbaugh, R.; Glass, K.; Rudolf, C.; Tremblay, M. Robust Ensemble Learning to Identify Rare Disease Patients from Electronic Health Records. In Proceedings of the 2018 40th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Honolulu, HI, USA, 17–21 July 2018; pp. 4085–4088. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wilson, A.; Chiorean, A.; Aguiar, M.; Sekulic, D.; Pavlick, P.; Shah, N.; King, L.S.; Génin, M.; Rollot, M.; Blanchon, M.; et al. Development of a rare disease algorithm to identify persons at risk of Gaucher disease using electronic health records in the United States. Orphanet J. Rare Dis. 2023, 18, 280. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- García-García, E.; González-Romero, G.M.; Martín-Pérez, E.M.; de Dios Zapata Cornejo, E.; Escobar-Aguilar, G.; Bonnet, M.F.C. Real-World Data and Machine Learning to Predict Cardiac Amyloidosis. Int. J. Environ. Res. Public Health 2021, 18, 908. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Banerjee, J.; Taroni, J.N.; Allaway, R.J.; Prasad, D.V.; Guinney, J.; Greene, C. Machine learning in rare disease. Nat. Methods 2023, 20, 803–814. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Visibelli, A.; Roncaglia, B.; Spiga, O.; Santucci, A. The Impact of Artificial Intelligence in the Odyssey of Rare Diseases. Biomedicines 2023, 11, 887. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Alsentzer, E.; Li, M.M.; Kobren, S.N.; Noori, A.; Kohane, I.S.; Zitnik, M. Few shot learning for phenotype-driven diagnosis of patients with rare genetic diseases. npj Digit. Med. 2025, 8, 380. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lee, J.; Liu, C.; Kim, J.; Chen, Z.; Sun, Y.; Rogers, J.R.; Chung, W.K.; Weng, C. Deep learning for rare disease: A scoping review. J. Biomed. Inform. 2022, 135, 104227. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Brasil, S.; Pascoal, C.; Francisco, R.; Ferreira, V.D.R.; Videira, P.A.; Valadão, G. Artificial Intelligence (AI) in Rare Diseases: Is the Future Brighter? Genes 2019, 10, 978. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abugabah, A.; Shukla, P.K.; Shukla, P.K.; Pandey, A. An intelligent healthcare system for rare disease diagnosis utilizing electronic health records based on a knowledge-guided multimodal transformer framework. BioData Min. 2025, 18, 70. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mullenbach, J.; Wiegreffe, S.; Duke, J.; Sun, J.; Eisenstein, J. Explainable Prediction of Medical Codes from Clinical Text. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers); Walker, M., Ji, H., Stent, A., Eds.; Association for Computational Linguistics: New Orleans, LA, USA, 2018; pp. 1101–1111. [Google Scholar] [CrossRef] [Scilit]
- Schilcher, J.; Nilsson, A.; Andlid, O.; Eklund, A. Fusion of electronic health records and radiographic images for a multimodal deep learning prediction model of atypical femur fractures. Comput. Biol. Med. 2024, 168, 107704. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, Z.; Shikany, A.; Ni, Y.; Zhang, G.; Weaver, K.N.; Chen, J. Using deep learning and electronic health records to detect Noonan syndrome in pediatric patients. Genet. Med. 2022, 24, 2329–2337. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, Z.; Shikany, A.; Husami, A.; Wang, X.; Mendonca, E.; Weaver, K.N.; Chen, J. Sequencing validates deep learning models for EHR-based detection of Noonan syndrome in pediatric patients. npj Genom. Med. 2025, 10, 56. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lee, W.; Park, S.; Joo, W.; Moon, I.C. Diagnosis Prediction via Medical Context Attention Networks Using Deep Generative Modeling. In Proceedings of the 2018 IEEE International Conference on Data Mining (ICDM), Singapore, 17–20 November 2018; pp. 1104–1109. [Google Scholar] [CrossRef] [Scilit]
- Khalaf, M.; Hussain, A.J.; Keight, R.; Al-Jumeily, D.; Keenan, R.; Chalmers, C.; Fergus, P.; Salih, W.; Abd, D.H.; Idowu, I.O. Recurrent Neural Network Architectures for Analysing Biomedical Data Sets. In Proceedings of the 2017 10th International Conference on Developments in eSystems Engineering (DeSE), Paris, France, 14–16 June 2017; pp. 232–237. [Google Scholar] [CrossRef] [Scilit]
- Khalaf, M.; Hussain, A.J.; Keight, R.; Al-Jumeily, D.; Fergus, P.; Keenan, R.; Tso, P. Machine learning approaches to the application of disease modifying therapy for sickle cell using classification models. Neurocomputing 2017, 228, 154–164. [Google Scholar] [CrossRef] [Scilit]
- Yu, K.; Wang, Y.; Cai, Y.; Xiao, C.; Zhao, E.; Glass, L.; Sun, J. Rare Disease Detection by Sequence Modeling with Generative Adversarial Networks. arXiv 2019, arXiv:1907.01022. [Google Scholar] [CrossRef] [Scilit]
- Yu, K.; Wang, Y.; Cai, Y. Modelling Patient Sequences for Rare Disease Detection with Semi-supervised Generative Adversarial Nets. In Proceedings of the Advanced Analytics and Learning on Temporal Data; Lemaire, V., Malinowski, S., Bagnall, A., Bondu, A., Guyet, T., Tavenard, R., Eds.; Springer: Cham, Switzerland, 2020; pp. 141–150. [Google Scholar]
- Zhou, X. ComplicaCode: Enhancing Disease Complication Detection in Electronic Health Records Through ICD Path Generation. In Proceedings of the International Conference on Artificial Neural Networks; Springer: Berlin/Heidelberg, Germany, 2024; pp. 29–43. [Google Scholar] [CrossRef] [Scilit]
- Cui, L.; Biswal, S.; Glass, L.M.; Lever, G.; Sun, J.; Xiao, C. CONAN: Complementary Pattern Augmentation for Rare Disease Detection. AAAI Conf. Artif. Intell. 2020, 34, 614–621. [Google Scholar] [CrossRef] [Scilit]
- Li, R.; Wen, A.; Gao, J.; Liu, H. MLGAN: A Meta-Learning based Generative Adversarial Network adapter for rare disease differentiation tasks. In Proceedings of the 14th ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics, New York, NY, USA, 3–6 September 2023; BCB ’23. [Google Scholar] [CrossRef] [Scilit]
- Li, W.; Wang, Y.; Cai, Y.; Arnold, C.; Zhao, E.; Yuan, Y. Semi-supervised Rare Disease Detection Using Generative Adversarial Network. arXiv 2018, arXiv:1812.00547. [Google Scholar] [CrossRef] [Scilit]
- Kaliappan, S.; Balaji, V.; Socrates, S.; Yamsani, N. Enhancing Precision Medicine through Artificial Neural Networks for Phenotyping and Risk Prediction of Rare Genetic Disorders. In Proceedings of the 2024 International Conference on Advancements in Smart, Secure and Intelligent Computing (ASSIC), Bhubaneswar, India, 27–29 January 2024; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Palavali, D.R.; Pothireddy, S. Privacy-Preserving Federated Learning for Multi-Institutional Diagnosis of Rare Diseases Using Heterogeneous EHR Data. In Proceedings of the 2025 International Conference on Communication, Computer, and Information Technology (IC3IT), Mandya, India, 24–25 October 2025; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
- Segura-Bedmar, I.; Camino-Perdones, D.; Guerrero-Aspizua, S. Exploring deep learning methods for recognizing rare diseases and their clinical manifestations from texts. BMC Bioinform. 2022, 23, 263. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dong, H.; Suárez-Paniagua, V.; Zhang, H.; Wang, M.; Casey, A.; Davidson, E.; Chen, J.; Alex, B.; Whiteley, W.; Wu, H. Ontology-driven and weakly supervised rare disease identification from clinical notes. BMC Med. Inform. Decis. Mak. 2023, 23, 86. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dong, H.; Suárez-Paniagua, V.; Zhang, H.; Wang, M.; Whitfield, E.; Wu, H. Rare Disease Identification from Clinical Notes with Ontologies and Weak Supervision. In Proceedings of the 2021 43rd Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Virtual, 1–5 November 2021; pp. 2294–2298. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, T.; Huang, D.; Lin, Y.; Wu, P.; Wu, Z.; Ma, G.; Lu, Y.; Dong, X.; Li, D.; Ge, J.; et al. A Specialized Large Language Model for Clinical Reasoning and Diagnosis in Rare Diseases. arXiv 2025, arXiv:2511.14638. [Google Scholar] [CrossRef] [Scilit]
- Zhao, W.; Wu, C.; Fan, Y.; Qiu, P.; Zhang, X.; Sun, Y.; Zhou, X.; Zhang, S.; Peng, Y.; Wang, Y.; et al. An agentic system for rare disease diagnosis with traceable reasoning. Nature 2026, 651, 775–784. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, X.; Mao, X.; Guo, Q.; Wang, L.; Zhang, S.; Chen, T. RareBench: Can LLMs Serve as Rare Diseases Specialists? In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, New York, NY, USA, 25–29 August 2024; KDD ’24, pp. 4850–4861. [Google Scholar] [CrossRef] [Scilit]
- Lee, J.; Yoon, W.; Kim, S.; Kim, D.; Kim, S.; So, C.H.; Kang, J. BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 2019, 36, 1234–1240. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kraljevic, Z.; Bean, D.; Shek, A.; Bendayan, R.; Hemingway, H.; Yeung, J.A.; Deng, A.; Balston, A.; Ross, J.; Idowu, E.; et al. Foresight—A generative pretrained transformer for modelling of patient timelines using electronic health records: A retrospective modelling study. Lancet Digit. Health 2024, 6, e281–e290. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, X.; Jin, Y.; Mao, X.; Wang, L.; Zhang, S.; Chen, T. Rareagents: Autonomous multi-disciplinary team for rare disease diagnosis and treatment. In Proceedings of the AAAI Conference on Artificial Intelligence, Singapore, 20–27 January 2026; Volume 40, pp. 101–109. [Google Scholar] [CrossRef] [Scilit]
- Yang, J.; Shu, L.; Duan, H.; Li, H. RDguru: A Conversational Intelligent Agent for Rare Diseases. IEEE J. Biomed. Health Inform. 2025, 29, 6366–6378. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lee, J.; Yao, L.; Jeong, H.H.; Liu, Z. LA-MARRVEL: A Knowledge-Grounded, Language-Aware LLM Framework for Clinically Robust Rare Disease Gene Prioritization. arXiv 2026, arXiv:2511.02263. [Google Scholar] [CrossRef] [Scilit]
- Oniani, D.; Hilsman, J.; Dong, H.; Gao, F.; Verma, S.; Wang, Y. Large Language Models Vote: Prompting for Rare Disease Identification. arXiv 2024, arXiv:2308.12890. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Wu, D.; Nguyen, Q.; Wang, K. Integrating Chain-of-Thought and Retrieval Augmented Generation Enhances Rare Disease Diagnosis From Clinical Notes. Med. Bull. 2026, 2, 167–183. [Google Scholar] [CrossRef] [Scilit]
- Wu, J.; Dong, H.; Li, Z.; Wang, H.; Li, R.; Patra, A.; Dai, C.; Ali, W.; Scordis, P.; Wu, H. A hybrid framework with large language models for rare disease phenotyping. BMC Med. Inform. Decis. Mak. 2024, 24, 289. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ao, G.; Chen, M.; Li, J.; Nie, H.; Zhang, L.; Chen, Z. Comparative analysis of large language models on rare disease identification. Orphanet J. Rare Dis. 2025, 20, 150. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mao, X.; Huang, Y.; Jin, Y.; Wang, L.; Chen, X.; Liu, H.; Yang, X.; Xu, H.; Luan, X.; Xiao, Y.; et al. A phenotype-based AI pipeline outperforms human experts in differentially diagnosing rare diseases using EHRs. npj Digit. Med. 2025, 8, 68. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Reese, J.T.; Chimirri, L.; Bridges, Y.; Danis, D.; Caufield, J.H.; Gargano, M.A.; Kroll, C.; Schmeder, A.; Liu, F.; Wissink, K.; et al. Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools. Eur. J. Hum. Genet. 2026, 34, 498–504. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- AlDin, Z.E.; Wu, J.; Fung, J.P.; King, J.; Watts, M.; ONeill, L.; Cross, A.R.; Sun, J. MIMIC-RD: Can LLMs differentially diagnose rare diseases in real-world clinical settings? arXiv 2025, arXiv:2601.11559. [Google Scholar] [CrossRef] [Scilit]
- Grothey, B.; Odenkirchen, J.; Brkic, A.; Schömig-Markiefka, B.; Quaas, A.; Büttner, R.; Tolkach, Y. Comprehensive testing of large language models for extraction of structured data in pathology. Commun. Med. 2025, 5, 96. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Du, X.; Zhou, Z.; Wang, Y.; Chuang, Y.W.; Yang, R.; Zhang, W.; Wang, X.; Zhang, R.; Hong, P.; Bates, D.W.; et al. Generative Large Language Models in Electronic Health Records for Patient Care Since 2023: A Systematic Review. medRxiv 2024, 2024.08.11.24311828. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ren, W.; Zhu, J.; Liu, Z.; Zhao, T.; Honavar, V. A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models. arXiv 2025, arXiv:2507.12774. [Google Scholar] [CrossRef] [Scilit]
- Nie, P.; Wu, H.; Cai, Z. Towards automatic icd coding via label graph generation. Mathematics 2024, 12, 2398. [Google Scholar] [CrossRef] [Scilit]
- Michalopoulos, G.; Malyska, M.; Sahar, N.; Wong, A.; Chen, H. ICDBigBird: A Contextual Embedding Model for ICD Code Classification. In Proceedings of the 21st Workshop on Biomedical Language Processing, Dublin, Ireland, 26 May 2022; pp. 330–336. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Zhao, Y.; Zheng, Y.; Wu, X. RareSyn: Health Record Synthesis for Rare Disease Diagnosis. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China, 4–9 November 2025; pp. 12311–12327. [Google Scholar] [CrossRef] [Scilit]
- Centers for Medicare & Medicaid Services. ICD-10-CM/PCS General Equivalence Mappings (GEMs). 2018. Available online: https://www.cms.gov/medicare/coding/icd10/downloads/2018-icd-10-cm-general-equivalence-mappings.zip (accessed on 28 December 2025).
- Li, F.; Yu, H. ICD coding from clinical text using multi-filter residual convolutional neural network. In Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA, 7–12 February 2020; Volume 34, pp. 8180–8187. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Canadian Institutes of Health Research; Natural Sciences and Engineering Research Council of Canada; Social Sciences and Humanities Research Council of Canada. Tri-Council Policy Statement: Ethical Conduct for Research Involving Humans; Secretariat on Responsible Conduct of Research: Ottawa, ON, Canada, 2022. [Google Scholar]






| Study | Dataset | Modalities | Architecture | Tail Reported | Limitation Addressed Here |
|---|---|---|---|---|---|
| CAML [42] | MIMIC-III full | Text only | CNN with per-label attention | Aggregate only | No structured clinical context |
| Aden et al. [12] | MIMIC-III top-10/50 | Text only | ClinicalBERT + LSTM | No | Rare codes excluded by design |
| Merchant et al. [14] | MIMIC-III full | Text + structured | BiLSTM ensemble | Aggregate only | No frequency stratification |
| ICDBigBird [79] | MIMIC-III | Text + ICD hierarchy | BigBird + GCN | Partial | No structured EHR events |
| LabGraph [78] | MIMIC-III | Text + label graph | Graph generation | Yes | Text-only inputs |
| Abugabah et al. [41] | Institutional + ORDO | Imaging + text + genomic | Swin + Med-BERT + GNN | Yes | Not ICD coding; corpus not public |
| Schilcher et al. [43] | Radiographs + EHR | Imaging + tabular | CNN late fusion | Single disease | Single-label, one condition |
| RareAgents [65] | MIMIC-IV-Ext-Rare | Text only | LLM agent ensemble | Yes | Diagnosis, not code assignment |
| MMAC-Net (this work) | MIMIC-III full + rare subset | Text + prescriptions + microbiology | DWConv + SE + per-label attention, late fusion | Yes | — |
| Metric | All | Rare |
|---|---|---|
| Admissions | 58,976 | 39,304 |
| Unique patients | 46,520 | 31,515 |
| Distinct ICD-9 codes | 8930 | 568 |
| Mean age | 53.5 | 60.5 |
| Median age | 60.5 | 64.5 |
| 90th percentile age | 81.9 | 82.7 |
| Unique ICD-9 codes per admission (avg) | 11.04 | 4.72 |
| Symbol | Operation | Shape (in → out) |
|---|---|---|
| Conv1D | Stem convolution over embedded tokens | |
| DWConv | Depthwise separable convolution | |
| BN | Batch normalization over channels | unchanged |
| SE | Channel gate, Equation (1) | |
| ⊕ | Residual addition, Equation (2) | unchanged |
| ⊗ | Channel-wise or attention product | unchanged |
| U, | Per-label attention projection and softmax | |
| M | Label-specific document representation | |
| Tabular MLP | Structured-feature encoder | |
| Concat | Fusion across the label axis | |
| FC 0, FC 1 | Output classifier |
| Parameter | Value |
|---|---|
| Embedding dimension | 100 |
| Sequence length (L) | 1500 tokens |
| Kernel size | 100 |
| Number of filter maps | 128 |
| Convolutional block depth | 6 |
| Objective function | BCE |
| Regularization coefficient | 0.25 |
| Embedding dropout | 0.6 |
| Fully connected dropout | 0.3 |
| Batch size | 512 |
| Activation function | ReLU |
| Batch normalization | True |
| Target classes—rare subset () | 568 |
| Target classes—full dataset () | 8930 |
| Tabular Fusion | |
| Fusion strategy | Late Fusion |
| Tabular hidden dimension | 50 |
| Model | Source | Parameter | Peak Memory | Peak GPU Memory | Train Time | Modality |
|---|---|---|---|---|---|---|
| MMAC-Net (full label space) | measured | 15.5 M | 102.91 GB | 57 GB | 27 H | Multi-modal |
| CAML [42] | reported | 6.2 M | — | — | — | Text |
| MultiResCNN [82] | reported | 11.9 M | — | — | — | Text |
| Data Modalities | Precision | F1 | AUC | |||
|---|---|---|---|---|---|---|
| p@8 | p@15 | Macro | Micro | Macro | Micro | |
| Notes + Tabular | 0.096 ± 0.001 | 0.073 ± 0.001 | 0.029 ± 0.001 | 0.381 ± 0.002 | 0.811 ± 0.001 | 0.970 ± 0.003 |
| Notes + Categorical | 0.095 ± 0.002 | 0.087 ± 0.003 | 0.020 ± 0.001 | 0.381 ± 0.004 | 0.808 ± 0.004 | 0.971 ± 0.002 |
| Notes + Tabular + Categorical | 0.159 ± 0.003 | 0.163 ± 0.002 | 0.084 ± 0.001 | 0.513 ± 0.003 | 0.878 ± 0.006 | 0.985 ± 0.009 |
| Notes (Baseline) | 0.092 ± 0.001 | 0.053 ± 0.001 | 0.011 ± 0.001 | 0.368 ± 0.001 | 0.767 ± 0.001 | 0.966 ± 0.001 |
| Model | # of Labels | Macro-AUC | Micro-AUC | Macro-F1 | Micro-F1 | P@8 |
|---|---|---|---|---|---|---|
| CNN [42] | 8922 | 0.835 | 0.974 | 0.034 | 0.420 | 0.619 |
| CAML [42] | 8922 | 0.893 | 0.985 | 0.056 | 0.506 | 0.704 |
| MultiResCNN [82] | 8930 | 0.910 ± 0.002 | 0.986 ± 0.001 | 0.085 ± 0.007 | 0.552 ± 0.005 | 0.734 ± 0.002 |
| DCAN [27] | 8930 | 0.848 ± 0.009 | 0.979 ± 0.001 | 0.066 ± 0.005 | 0.533 ± 0.006 | 0.721 ± 0.001 |
| TransICD [27] | 8930 | 0.886 ± 0.010 | 0.983 ± 0.002 | 0.058 ± 0.001 | 0.497 ± 0.001 | 0.666 ± 0.000 |
| Fusion [27] | 8930 | 0.910 ± 0.003 | 0.986 ± 0.000 | 0.081 ± 0.002 | 0.560 ± 0.003 | 0.744 ± 0.002 |
| Ensemble (BiLSTM + FCN) [14] | 8930 | 0.922 | 0.988 | 0.120 | 0.571 | 0.739 |
| MMAC-Net (proposed) | 8930 | 0.981 ± 0.002 | 0.997 ± 0.003 | 0.641 ± 0.003 | 0.724 ± 0.005 | 0.875 ± 0.003 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ashrafi, A.F.; Alhajj, R.; Rokne, J.G. MMAC-Net: A Multi-Modal Multi-Label Attention-Based Deep Learning Approach for Automated ICD-9 Coding of Rare Disease Admissions from Electronic Health Records. Appl. Sci. 2026, 16, 8962. https://doi.org/10.3390/app16188962
Ashrafi AF, Alhajj R, Rokne JG. MMAC-Net: A Multi-Modal Multi-Label Attention-Based Deep Learning Approach for Automated ICD-9 Coding of Rare Disease Admissions from Electronic Health Records. Applied Sciences. 2026; 16(18):8962. https://doi.org/10.3390/app16188962
Chicago/Turabian StyleAshrafi, Adnan Ferdous, Reda Alhajj, and Jon George Rokne. 2026. "MMAC-Net: A Multi-Modal Multi-Label Attention-Based Deep Learning Approach for Automated ICD-9 Coding of Rare Disease Admissions from Electronic Health Records" Applied Sciences 16, no. 18: 8962. https://doi.org/10.3390/app16188962
APA StyleAshrafi, A. F., Alhajj, R., & Rokne, J. G. (2026). MMAC-Net: A Multi-Modal Multi-Label Attention-Based Deep Learning Approach for Automated ICD-9 Coding of Rare Disease Admissions from Electronic Health Records. Applied Sciences, 16(18), 8962. https://doi.org/10.3390/app16188962

