Next Article in Journal
Quantitative Ultrasound SWE of Carotid Plaque in Symptomatic and Asymptomatic Patients: A Systematic Review and Meta-Analysis
Next Article in Special Issue
AI in Hand and Wrist Radiography: Multimodal Large Language Models for Distal Radius Fracture Detection and Characterization
Previous Article in Journal
Ultrasound Assessment of the Anterolateral Ligament of the Knee: A Narrative Review of Current Evidence, Interpretative Limitations, and Clinical Context
Previous Article in Special Issue
Performance of ChatGPT-4o, Gemini 2.0 Pro, and DeepSeek-V3 in Patient-Facing Information on Chest Wall Deformities: A Comparative Evaluation of Accuracy, RELIABILITY, and Reproducibility
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Leveraging Large Language Models for Automated Extraction of Abdominal Aortic Aneurysm Features from Radiology Reports

1
Department of Radiology, Montefiore Medical Center and Albert Einstein College of Medicine, Bronx, NY 10461, USA
2
Renaissance School of Medicine, Stony Brook University, Stony Brook, NY 11794, USA
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
Diagnostics 2026, 16(7), 1083; https://doi.org/10.3390/diagnostics16071083
Submission received: 11 February 2026 / Revised: 22 March 2026 / Accepted: 28 March 2026 / Published: 3 April 2026

Abstract

Background/Objectives. Abdominal computed tomography (CT) radiology reports contain critical information for abdominal aortic aneurysm (AAA) management, including aneurysm presence, size, rupture status, and prior repair. However, this information is often embedded within lengthy, heterogeneous reports, making manual extraction inefficient. We evaluated the performance of multiple large language models (LLMs) for automated extraction of AAA-related findings from radiology reports. Methods. We retrospectively analyzed 500 abdominal CT reports mentioning AAA from an urban academic health system (2020–2024). Ground truth labels were established by manual review. Four open-source LLMs (Qwen2.5-7B-Instruct, Llama3-Med42-8B, GPT-OSS-20B, and MedGemma-27B-text-it) were evaluated for extraction of aneurysm presence, size, morphology, rupture status, impending rupture, and prior aortic repair. Model outputs were compared with ground truth using exact-match accuracy, and inter-model agreement was assessed using Fleiss’ kappa. Reasoning traces were examined to characterize correct and incorrect model behavior. Results. Accuracy for identifying AAA presence ranged from 0.90 to 0.95 (κ = 0.851), and prior aortic repair from 0.90 to 0.97 (κ = 0.793). Accuracy for aneurysm size ranged from 0.67 to 0.88 (κ = 0.340), with low κ’s due to class imbalance or dimension misselection. Rupture and impending rupture were identified with accuracies exceeding 0.90 across models, though agreement was lower (κ = 0.485 and 0.589), reflecting low event prevalence. Larger models (GPT-OSS-20B, MedGemma-27B) generally outperformed smaller models. Reasoning analysis revealed strengths in measurement prioritization but recurrent errors, including dimension misselection, over-inference of prior repair, and conservative classification of rupture-related findings. Conclusions. LLMs can accurately extract clinically relevant AAA information from radiology reports with interpretable reasoning, with larger and medically trained models outperforming smaller or general-purpose models. Performance varies by task and model, underscoring the need for careful validation and human-in-the-loop deployment in clinical settings.
Keywords: abdominal aortic aneurysm; large language model; information extraction; radiology report; computed tomography abdominal aortic aneurysm; large language model; information extraction; radiology report; computed tomography

Share and Cite

MDPI and ACS Style

Mukherjee, P.; Lee, R.C.; Hadidchi, R.; Henry, S.; Coard, M.; Davis, M.; Rubinov, Y.; Nguyen-Luong, H.; Katz, L.; Duong, T.Q. Leveraging Large Language Models for Automated Extraction of Abdominal Aortic Aneurysm Features from Radiology Reports. Diagnostics 2026, 16, 1083. https://doi.org/10.3390/diagnostics16071083

AMA Style

Mukherjee P, Lee RC, Hadidchi R, Henry S, Coard M, Davis M, Rubinov Y, Nguyen-Luong H, Katz L, Duong TQ. Leveraging Large Language Models for Automated Extraction of Abdominal Aortic Aneurysm Features from Radiology Reports. Diagnostics. 2026; 16(7):1083. https://doi.org/10.3390/diagnostics16071083

Chicago/Turabian Style

Mukherjee, Praneel, Ryan C. Lee, Roham Hadidchi, Sonya Henry, Michael Coard, Matthew Davis, Yossef Rubinov, Ha Nguyen-Luong, Leah Katz, and Tim Q. Duong. 2026. "Leveraging Large Language Models for Automated Extraction of Abdominal Aortic Aneurysm Features from Radiology Reports" Diagnostics 16, no. 7: 1083. https://doi.org/10.3390/diagnostics16071083

APA Style

Mukherjee, P., Lee, R. C., Hadidchi, R., Henry, S., Coard, M., Davis, M., Rubinov, Y., Nguyen-Luong, H., Katz, L., & Duong, T. Q. (2026). Leveraging Large Language Models for Automated Extraction of Abdominal Aortic Aneurysm Features from Radiology Reports. Diagnostics, 16(7), 1083. https://doi.org/10.3390/diagnostics16071083

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop