Fine-Tuning LLaMA2 for Summarizing Discharge Notes: Evaluating the Role of Highlighted Information
Abstract
1. Introduction
2. Background
2.1. Text Summarization and Related Work
2.2. Dataset
2.3. Automatic Highlighting
3. Method
3.1. Data Preparation
3.2. Fine-Tuning Procedure
3.3. Generating Summary from Fine-Tuned Models
3.4. Evaluation Metrics
3.4.1. Manual Evaluation
3.4.2. Automatic Evaluation
3.4.3. LLM-Based Evaluation
4. Results
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Menachemi, N.; Collum, T.H. Benefits and drawbacks of electronic health record systems. Risk Manag. Healthc. Policy 2011, 4, 47–55. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Madzime, R.; Nyirenda, C. Enhanced Electronic Health Records Text Summarization Using Large Language Models. arXiv 2024, arXiv:2410.09628. [Google Scholar] [CrossRef] [Scilit]
- Zhou, S.; Liu, H.; Sen, P.; Perl, Y.; Dehkordi, M.K. CFC annotator: A cluster-focused combination algorithm for annotating electronic health records by referencing interface terminology. In Proceedings of the 18th International Joint Conference on Biomedical Engineering Systems and Technologies (BIOSTEC 2025), Porto, Portugal, 19–21 February 2025. [Google Scholar] [CrossRef] [Scilit]
- Bowman, S. Impact of electronic health record systems on information integrity: Quality and safety implications. Perspect. Health Inf. Manag. 2013, 10, 1c. [Google Scholar]
- O’Malley, A.S.; Grossman, J.M.; Cohen, G.R.; Kemper, N.M.; Pham, H.H. Are electronic medical records helpful for care coordination? Experiences of physician practices. J. Gen. Intern. Med. 2010, 25, 177–185. [Google Scholar] [CrossRef] [Scilit]
- Apathy, N.C.; Rotenstein, L.; Bates, D.W.; Holmgren, A.J. Documentation dynamics: Note composition, burden, and physician efficiency. Health Serv. Res. 2023, 58, 674–685. [Google Scholar] [CrossRef] [Scilit]
- Zhao, W.X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z. A survey of large language models. arXiv 2023, arXiv:2303.18223. [Google Scholar]
- Thirunavukarasu, A.J.; Ting, D.S.J.; Elangovan, K.; Gutierrez, L.; Tan, T.F.; Ting, D.S.W. Large language models in medicine. Nat. Med. 2023, 29, 1930–1940. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hadi, M.U.; Qureshi, R.; Shah, A.; Irfan, M.; Zafar, A.; Shaikh, M.B.; Akhtar, N.; Wu, J.; Mirjalili, S. A survey on large language models: Applications, challenges, limitations, and practical usage. Authorea Prepr. 2023. [Google Scholar] [CrossRef] [Scilit]
- Kwag, K.H.; González-Lorenzo, M.; Banzi, R.; Bonovas, S.; Moja, L. Providing doctors with high-quality information: An updated evaluation of web-based point-of-care information summaries. J. Med. Internet Res. 2016, 18, e15. [Google Scholar] [CrossRef] [Scilit]
- Shestov, A.; Levichev, R.; Mussabayev, R.; Maslov, E.; Zadorozhny, P.; Cheshkov, A.; Mussabayev, R.; Toleu, A.; Tolegen, G.; Krassovitskiy, A. Finetuning large language models for vulnerability detection. IEEE Access 2025, 13, 38889–38900. [Google Scholar] [CrossRef] [Scilit]
- Rallapalli, S.; Gallagher, S.; Mellinger, A.O.; Ratchford, J.; Sinha, A.; Brooks, T.; Nichols, W.R.; Winski, N.; Brown, B. Fine-Tuning LLMs for Report Summarization: Analysis on Supervised and Unsupervised Data. arXiv 2025, arXiv:2503.10676. [Google Scholar]
- Pivovarov, R.; Elhadad, N. Automated methods for the summarization of electronic health records. J. Am. Med. Inform. Assoc. 2015, 22, 938–947. [Google Scholar] [CrossRef] [Scilit]
- Sarzynski, E.; Hashmi, H.; Subramanian, J.; Fitzpatrick, L.; Polverento, M.; Simmons, M.; Brooks, K.; Given, C. Opportunities to improve clinical summaries for patients at hospital discharge. BMJ Qual. Saf. 2017, 26, 372–380. [Google Scholar] [CrossRef] [Scilit]
- Casey, J.A.; Schwartz, B.S.; Stewart, W.F.; Adler, N.E. Using electronic health records for population health research: A review of methods and applications. Annu. Rev. Public Health 2016, 37, 61–81. [Google Scholar] [CrossRef] [Scilit]
- Wu, X.-K.; Chen, M.; Li, W.; Wang, R.; Lu, L.; Liu, J.; Hwang, K.; Hao, Y.; Pan, Y.; Meng, Q. LLM Fine-Tuning: Concepts, Opportunities, and Challenges. Big Data Cogn. Comput. 2025, 9, 87. [Google Scholar] [CrossRef] [Scilit]
- Hu, M.; He, B.; Wang, Y.; Li, L.; Ma, C.; King, I. Mitigating large language model hallucination with faithful finetuning. arXiv 2024, arXiv:2406.11267. [Google Scholar] [CrossRef] [Scilit]
- Rumiantsau, M.; Vertsel, A.; Hrytsuk, I.; Ballah, I. Beyond Fine-Tuning: Effective Strategies for Mitigating Hallucinations in Large Language Models for Data Analytics. arXiv 2024, arXiv:2410.20024. [Google Scholar] [CrossRef] [Scilit]
- Liu, C.; Sun, K.; Zhou, Q.; Duan, Y.; Shu, J.; Kan, H.; Gu, Z.; Hu, J. CPMI-ChatGLM: Parameter-efficient fine-tuning ChatGLM with Chinese patent medicine instructions. Sci. Rep. 2024, 14, 6403. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hamzah, F.; Sulaiman, N. Optimizing Llama 7B for Medical Question Answering: A Study on Fine-Tuning Strategies and Performance on the MultiMedQA Dataset. Available online: https://osf.io/preprints/osf/g5aes_v1 (accessed on 1 July 2025).
- Li, I.; Pan, J.; Goldwasser, J.; Verma, N.; Wong, W.P.; Nuzumlalı, M.Y.; Rosand, B.; Li, Y.; Zhang, M.; Chang, D. Neural natural language processing for unstructured data in electronic health records: A review. Comput. Sci. Rev. 2022, 46, 100511. [Google Scholar] [CrossRef] [Scilit]
- Perković, G.; Drobnjak, A.; Botički, I. Hallucinations in llms: Understanding and addressing challenges. In Proceedings of the 2024 47th MIPRO ICT and Electronics Convention (MIPRO), Opatija, Croatia, 20–24 May 2024. [Google Scholar]
- Jha, S.; Jha, S.K.; Lincoln, P.; Bastian, N.D.; Velasquez, A.; Neema, S. Dehallucinating large language models using formal methods guided iterative prompting. In Proceedings of the 2023 IEEE International Conference on Assured Autonomy (ICAA), Laurel, MD, USA, 6–8 June 2023. [Google Scholar]
- Parthasarathy, V.B.; Zafar, A.; Khan, A.; Shahid, A. The ultimate guide to fine-tuning llms from basics to breakthroughs: An exhaustive review of technologies, research, best practices, applied research challenges and opportunities. arXiv 2024, arXiv:2408.13296. [Google Scholar] [CrossRef] [Scilit]
- Ahmad, P.N.; Liu, Y.; Khan, K.; Jiang, T.; Burhan, U. BIR: Biomedical Information Retrieval System for Cancer Treatment in Electronic Health Record Using Transformers. Sensors 2023, 23, 9355. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, Z.; Pfaff, E.; Guo, S.J.; Guo, Y.; Wu, Y.; Tao, C.; Stiglic, G.; Bian, J. Enriching real-world data with social determinants of health for health outcomes and health equity: Successes, challenges, and opportunities. Yearb. Med. Inform. 2023, 32, 253–263. [Google Scholar] [CrossRef] [Scilit]
- Majdik, Z.P.; Graham, S.S.; Shiva Edward, J.C.; Rodriguez, S.N.; Karnes, M.S.; Jensen, J.T.; Barbour, J.B.; Rousseau, J.F. Sample Size Considerations for Fine-Tuning Large Language Models for Named Entity Recognition Tasks: Methodological Study. JMIR AI 2024, 3, e52095. [Google Scholar] [CrossRef] [Scilit]
- Aali, A.; Van Veen, D.; Arefeen, Y.I.; Hom, J.; Bluethgen, C.; Reis, E.P.; Gatidis, S.; Clifford, N.; Daws, J.; Tehrani, A.S. A dataset and benchmark for hospital course summarization with adapted large language models. J. Am. Med. Inform. Assoc. 2024, 32, ocae312. [Google Scholar] [CrossRef] [Scilit]
- Aali, A.; Van Veen, D.; Arefeen, Y.I.; Hom, J.; Bluethgen, C.; Reis, E.P.; Gatidis, S.; Clifford, N.; Daws, J.; Tehrani, A.S. MIMIC-IV-Ext-BHC: Labeled Clinical Notes Dataset for Hospital Course Summarization. 2024. Available online: https://physionet.org/content/labelled-notes-hospital-course/1.1.0/ (accessed on 1 July 2025).
- Du, X.; Zhou, Z.; Wang, Y.; Chuang, Y.-W.; Yang, R.; Zhang, W.; Wang, X.; Zhang, R.; Hong, P.; Bates, D.W. Generative large language models in electronic health records for patient care since 2023: A systematic review. medRxiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Acharya, A.; Shrestha, S.; Chen, A.; Conte, J.; Avramovic, S.; Sikdar, S.; Anastasopoulos, A.; Das, S. Clinical risk prediction using language models: Benefits and considerations. J. Am. Med. Inform. Assoc. 2024, 31, ocae030. [Google Scholar] [CrossRef] [Scilit]
- Koohi Habibi Dehkordi, M.; Perl, Y.; Deek, F.P.; He, Z.; Keloth, V.K.; Liu, H.; Elhanan, G.; Einstein, A.J. Improving Large Language Models’ Summarization Accuracy by Adding Highlights to Discharge Notes: Comparative Evaluation. JMIR Med. Inform. 2025, 13, e66476. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Dehkordi, M.K.H.; Einstein, A.J.; Zhou, S.; Elhanan, G.; Perl, Y.; Keloth, V.K.; Geller, J.; Liu, H. Using annotation for computerized support for fast skimming of cardiology electronic health record notes. In Proceedings of the 2023 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Istanbul, Turkey, 5–8 December 2023; pp. 4043–4050. [Google Scholar]
- Dehkordi, M.K.H.; Kollapally, N.M.; Perl, Y.; Geller, J.; Deek, F.P.; Liu, H.; Keloth, V.K.; Elhanan, G.; Einstein, A.J. Skimming of Electronic Health Records Highlighted by an Interface Terminology Curated with Machine Learning Mining. In Proceedings of the 17th International Joint Conference on Biomedical Engineering Systems and Technologies (BIOSTEC 2024), Rome, Italy, 21–23 February 2024. [Google Scholar] [CrossRef] [Scilit]
- Dehkordi, M.K.; Kollapally, N.M.; Perl, Y.; Geller, J.; Deek, F.P.; Liu, H.; Keloth, V.K.; Elhanan, G.; Einstein, A.J. Curation of a Cardiology Interface Terminology for Highlighting Electronic Health Records using Machine Learning. In Proceedings of the 17th International Joint Conference on Biomedical Engineering Systems and Technologies (BIOSTEC 2024), Rome, Italy, 21–23 February 2024. [Google Scholar]
- Ahmed, I.; Islam, S.; Datta, P.P.; Kabir, I.; Chowdhury, N.U.R.; Haque, A. Qwen 2.5: A comprehensive review of the leading resource-efficient llm with potentioal to surpass all competitors. TechRxiv 2025. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jin, H.; Zhang, Y.; Meng, D.; Wang, J.; Tan, J. A comprehensive survey on process-oriented automatic text summarization with exploration of llm-based methods. arXiv 2024, arXiv:2403.02901. [Google Scholar]
- Zhong, M.; Liu, P.; Chen, Y.; Wang, D.; Qiu, X.; Huang, X. Extractive summarization as text matching. arXiv 2020, arXiv:2004.08795. [Google Scholar] [CrossRef] [Scilit]
- Gupta, S.; Gupta, S.K. Abstractive summarization: An overview of the state of the art. Expert Syst. Appl. 2019, 121, 49–65. [Google Scholar] [CrossRef] [Scilit]
- Van Veen, D.; Van Uden, C.; Blankemeier, L.; Delbrouck, J.-B.; Aali, A.; Bluethgen, C.; Pareek, A.; Polacin, M.; Reis, E.P.; Seehofnerová, A. Adapted large language models can outperform medical experts in clinical text summarization. Nat. Med. 2024, 30, 1134–1142. [Google Scholar] [CrossRef] [Scilit]
- Ma, C.; Wu, Z.; Wang, J.; Xu, S.; Wei, Y.; Liu, Z.; Zeng, F.; Jiang, X.; Guo, L.; Cai, X. An Iterative Optimizing Framework for Radiology Report Summarization with ChatGPT. IEEE Trans. Artif. Intell. 2024, 5, 4163–4175. [Google Scholar] [CrossRef] [Scilit]
- Hake, J.; Crowley, M.; Coy, A.; Shanks, D.; Eoff, A.; Kirmer-Voss, K.; Dhanda, G.; Parente, D.J. Quality, Accuracy, and Bias in ChatGPT-Based Summarization of Medical Abstracts. Ann. Fam. Med. 2024, 22, 113–120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M. Huggingface’s transformers: State-of-the-art natural language processing. arXiv 2019, arXiv:1910.03771. [Google Scholar]
- Luo, Z. Knowledge-guided aspect-based summarization. In Proceedings of the 2023 International conference on communications, computing and artificial intelligence (CCCAI), Shanghai, China, 23–25 June 2023. [Google Scholar]
- Wang, H.; Liu, J.; Duan, M.; Gong, P.; Wu, Z.; Wang, J.; Han, B. Cross-modal knowledge guided model for abstractive summarization. Complex Intell. Syst. 2024, 10, 577–594. [Google Scholar] [CrossRef] [Scilit]
- Ming, S.; Guo, Y.; Kilicoglu, H. Towards Knowledge-Guided Biomedical Lay Summarization using Large Language Models. In Proceedings of the Second Workshop on Patient-Oriented Language Processing (CL4Health), Albuquerque, NM, USA, 3–4 May 2025. [Google Scholar]
- Li, C.-Y.; Chun, S.A.; Geller, J. Perspective-Based Microblog Summarization. Information 2025, 16, 285. [Google Scholar] [CrossRef] [Scilit]
- Bodenreider, O. The unified medical language system (UMLS): Integrating biomedical terminology. Nucleic Acids Res. 2004, 32, D267–D270. [Google Scholar] [CrossRef] [Scilit]
- Donnelly, K. SNOMED-CT: The advanced terminology and coding system for eHealth. Stud. Health Technol. Inf. 2006, 121, 279. [Google Scholar]
- Johnson, A.E.; Bulgarelli, L.; Shen, L.; Gayles, A.; Shammout, A.; Horng, S.; Pollard, T.J.; Hao, S.; Moody, B.; Gow, B. MIMIC-IV, a freely accessible electronic health record dataset. Sci. Data 2023, 10, 1. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Johnson, A.; Bulgarelli, L.; Pollard, T.; Horng, S.; Celi, L.A.; Mark, R. Mimic-iv. PhysioNet. 2020. Available online: https://physionet.org/content/mimiciv/1.0/ (accessed on 23 August 2021).
- Grady, C. Institutional review boards: Purpose and challenges. Chest 2015, 148, 1148–1155. [Google Scholar] [CrossRef] [Scilit]
- Dehkordi, M.K.H.; Zhou, S.; Perl, Y.; Deek, F.P.; Einstein, A.J.; Elhanan, G.; He, Z.; Liu, H. Enhancing patient Comprehension: An effective sequential prompting approach to simplifying EHRs using LLMs. In Proceedings of the 2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Lisbon, Portugal, 3–6 December 2024. [Google Scholar] [CrossRef] [Scilit]
- Kollapally, N.M.; Dehkordi, M.K.H.; Perl, Y.; Geller, J.; Deek, F.P.; Liu, H.; Keloth, V.K.; Elhanan, G.; Einstein, A.J.; Zhou, S. Using clinical entity recognition for curating an interface terminology to aid fast skimming of EHRs. In Proceedings of the 2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Lisbon, Portugal, 3–6 December 2024. [Google Scholar]
- Dehkordi, M.K.H.; Lu, J.; Perl, Y.; Deek, F.P. Enhancing Patient Comprehension of Discharge Notes with a Retrieval-Augmented LLM Approach. In Proceedings of the 2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Wuhan, China, 15–18 December 2024. [Google Scholar]
- Dehkordi, M.K.H.; Perl, Y.; Deek, F.P. Optimizing Manual Review Using Machine Learning in Interface Terminology Curation for Automatic EHR Highlighting. In Proceedings of the 2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Wuhan, China, 15–18 December 2024. [Google Scholar]
- SNOMED International Homepage. Available online: http://www.snomed.org/ (accessed on 17 December 2025).
- Alsentzer, E.; Murphy, J.R.; Boag, W.; Weng, W.-H.; Jin, D.; Naumann, T.; McDermott, M. Publicly available clinical BERT embeddings. arXiv 2019, arXiv:1904.03323. [Google Scholar] [CrossRef] [Scilit]
- Liashchynskyi, P.; Liashchynskyi, P. Grid search, random search, genetic algorithm: A big comparison for NAS. arXiv 2019, arXiv:1912.06059. [Google Scholar] [CrossRef] [Scilit]
- Agarap, A.F. Deep learning using rectified linear units (relu). arXiv 2018, arXiv:1803.08375. [Google Scholar]
- Jais, I.K.M.; Ismail, A.R.; Nisa, S.Q. Adam optimization algorithm for wide and deep neural network. Knowl. Eng. Data Sci. 2019, 2, 41–46. [Google Scholar] [CrossRef] [Scilit]
- Zhou, S.; Dehkordi, M.K.H.; Perl, Y.; Deek, F.P.; Liu, H. Enhancing Electronic Health Records Annotation with a Cluster-Focused Combination Algorithm and Interface Terminologies. In Springer Book of HEALTHINF, Proceedings of the 18th International Joint Conference on Biomedical Engineering Systems and Technologies (BIOSTEC 2025), Porto, Portugal, 20–22 February 2025; Springer Nature: Berlin/Heidelberger, Germany, 2025. [Google Scholar]
- Han, Z.; Gao, C.; Liu, J.; Zhang, J.; Zhang, S.Q. Parameter-efficient fine-tuning for large models: A comprehensive survey. arXiv 2024, arXiv:2403.14608. [Google Scholar]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. Lora: Low-rank adaptation of large language models. ICLR 2022, 1, 3. [Google Scholar]
- Llugsi, R.; El Yacoubi, S.; Fontaine, A.; Lupera, P. Comparison between Adam, AdaMax and Adam W optimizers to implement a Weather Forecast based on Neural Networks for the Andean city of Quito. In Proceedings of the 2021 IEEE Fifth Ecuador Technical Chapters Meeting (ETCM), Cuenca, Ecuador, 12–15 October 2021. [Google Scholar]
- Zhang, R.; Han, J.; Liu, C.; Gao, P.; Zhou, A.; Hu, X.; Yan, S.; Lu, P.; Li, H.; Qiao, Y. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv 2023, arXiv:2303.16199. [Google Scholar]
- Gema, A.; Minervini, P.; Daines, L.; Hope, T.; Alex, B. Parameter-efficient fine-tuning of llama for the clinical domain. In Proceedings of the 6th Clinical Natural Language Processing Workshop, Mexico City, Mexico, 21 June 2024. [Google Scholar]
- Lermen, S.; Rogers-Smith, C.; Ladish, J. Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b. arXiv 2023, arXiv:2310.20624. [Google Scholar]
- Koraş, O.A.; Bahnan, R.; Kleesiek, J.; Dada, A. Towards Conditioning Clinical Text Generation for User Control. arXiv 2025, arXiv:2502.17571. [Google Scholar] [CrossRef] [Scilit]
- Su, Y.; Lan, T.; Wang, Y.; Yogatama, D.; Kong, L.; Collier, N. A contrastive framework for neural text generation. Adv. Neural Inf. Process. Syst. 2022, 35, 21548–21561. [Google Scholar]
- Peng, C.; Yang, X.; Chen, A.; Smith, K.E.; PourNejatian, N.; Costa, A.B.; Martin, C.; Flores, M.G.; Zhang, Y.; Magoc, T. A study of generative large language model for medical research and healthcare. npj Digit. Med. 2023, 6, 210. [Google Scholar] [CrossRef] [Scilit]
- Holtzman, A.; Buys, J.; Du, L.; Forbes, M.; Choi, Y. The curious case of neural text degeneration. arXiv 2019, arXiv:1904.09751. [Google Scholar]
- Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K.Q.; Artzi, Y. Bertscore: Evaluating text generation with bert. arXiv 2019, arXiv:1904.09675. [Google Scholar]
- Lin, C.-Y. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out; Association for Computational Linguistics: Stroudsburg, PA, USA, 2004. [Google Scholar]
- Papineni, K.; Roukos, S.; Ward, T.; Zhu, W.-J. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, Philadelphia, PA, USA, 6–12 July 2002. [Google Scholar]
- Laban, P.; Schnabel, T.; Bennett, P.N.; Hearst, M.A. SummaC: Re-visiting NLI-based models for inconsistency detection in summarization. Trans. Assoc. Comput. Linguist. 2022, 10, 163–177. [Google Scholar] [CrossRef] [Scilit]
- MacCartney, B. Natural Language Inference. Ph.D. Thesis, Stanford University, Stanford, CA, USA, 2009. [Google Scholar]
- Sun, Z.; Shen, Y.; Zhou, Q.; Zhang, H.; Chen, Z.; Cox, D.; Yang, Y.; Gan, C. Principle-driven self-alignment of language models from scratch with minimal human supervision. Adv. Neural Inf. Process. Syst. 2024, 36, 2511–2565. [Google Scholar]
- He, Z.; Bhasuran, B.; Jin, Q.; Tian, S.; Hanna, K.; Shavor, C.; Arguello, L.G.; Murray, P.; Lu, Z. Quality of Answers of Generative Large Language Models vs Peer Patients for Interpreting Lab Test Results for Lay Patients: Evaluation Study. arXiv 2024, arXiv:2402.01693. [Google Scholar] [CrossRef] [Scilit]
- Upton, G.J. Fisher’s exact test. J. R. Stat. Soc. Ser. A Stat. Soc. 1992, 155, 395–402. [Google Scholar] [CrossRef] [Scilit]
- Test, F.E. Fisher Exact Test. Available online: https://www.socscistatistics.com/tests/fisher/default2.aspx (accessed on 1 July 2025).
- Adams, G.; Zuckerg, J.; Elhadad, N. A meta-evaluation of faithfulness metrics for long-form hospital-course summarization. In Proceedings of the Machine Learning for Healthcare Conference, New York, NY, USA, 11–12 August 2023. [Google Scholar]
- Adams, G.; Alsentzer, E.; Ketenci, M.; Zucker, J.; Elhadad, N. What’s in a summary? Laying the groundwork for advances in hospital-course summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Virtual, 6–11 June 2021. [Google Scholar]
- Searle, T.; Ibrahim, Z.; Teo, J.; Dobson, R.J. Discharge summary hospital course summarisation of in patient electronic health record text with clinical concept guided deep pre-trained transformer models. J. Biomed. Inform. 2023, 141, 104358. [Google Scholar] [CrossRef] [Scilit]
- Adams, G. Generating Faithful and Complete Hospital-Course Summaries from the Electronic Health Record. Ph.D. Thesis, Columbia University, New York, NY, USA, 2024. [Google Scholar]


| Word Count | BERTScore | ROUGE-L | BLEU | Sumac_CONV | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| HCB | U_S | H_S | U_S | H_S | U_S | H_S | U_S | H_S | U_S | H_S |
| 393.19 | 300.38 | 331.67 | 59.61 | 63.75 | 21.82 | 23.43 | 8.41 | 10.4 | 40.2 | 67.7 |
| Coherence | Fluency | Conciseness | Correctness | Average | |||||
|---|---|---|---|---|---|---|---|---|---|
| U_S | H_S | U_S | H_S | U_S | H_S | U_S | H_S | U_S | H_S |
| 4.84 | 4.91 | 4.91 | 4.92 | 4.65 | 4.81 | 4.79 | 4.91 | 4.79 | 4.88 |
| Metric | Summary | Min | Max | STD |
|---|---|---|---|---|
| Completeness | U_S | 23.2 | 73.3 | 14.01 |
| H_S | 29.3 | 80.0 | 12.36 | |
| BERTScore | U_S | 57.8 | 69.2 | 2.86 |
| H_S | 59.3 | 69.3 | 2.75 | |
| ROUGE-L | U_S | 14.1 | 34.3 | 7.19 |
| H_S | 19.2 | 39.3 | 5.54 | |
| BLEU | U_S | 1.2 | 13.7 | 5.35 |
| H_S | 4.2 | 22.4 | 4.97 | |
| Sumac | U_S | 7.5 | 63.4 | 15.13 |
| H_S | 60.1 | 89.3 | 8.21 | |
| Coherence | U_S | 4 | 5 | 0.36 |
| H_S | 4 | 5 | 0.22 | |
| Fluency | U_S | 4 | 5 | 0.36 |
| H_S | 4 | 5 | 0.22 | |
| Conciseness | U_S | 4 | 5 | 0.43 |
| H_S | 4 | 5 | 0.40 | |
| Correctness | U_S | 4 | 5 | 0.40 |
| H_S | 4 | 5 | 0.30 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Koohi Habibi Dehkordi, M.; Perl, Y.; Deek, F.P.; Liu, H. Fine-Tuning LLaMA2 for Summarizing Discharge Notes: Evaluating the Role of Highlighted Information. Big Data Cogn. Comput. 2026, 10, 4. https://doi.org/10.3390/bdcc10010004
Koohi Habibi Dehkordi M, Perl Y, Deek FP, Liu H. Fine-Tuning LLaMA2 for Summarizing Discharge Notes: Evaluating the Role of Highlighted Information. Big Data and Cognitive Computing. 2026; 10(1):4. https://doi.org/10.3390/bdcc10010004
Chicago/Turabian StyleKoohi Habibi Dehkordi, Mahshad, Yehoshua Perl, Fadi P. Deek, and Hao Liu. 2026. "Fine-Tuning LLaMA2 for Summarizing Discharge Notes: Evaluating the Role of Highlighted Information" Big Data and Cognitive Computing 10, no. 1: 4. https://doi.org/10.3390/bdcc10010004
APA StyleKoohi Habibi Dehkordi, M., Perl, Y., Deek, F. P., & Liu, H. (2026). Fine-Tuning LLaMA2 for Summarizing Discharge Notes: Evaluating the Role of Highlighted Information. Big Data and Cognitive Computing, 10(1), 4. https://doi.org/10.3390/bdcc10010004

