Next Article in Journal
Changes in the Epidemiology of Thoracic and Cardiovascular Diseases in Korea During the COVID-19 Pandemic: A Nationwide Analysis
Previous Article in Journal
Scratch-Based Isolation of Primary Cells (SCIP): A Novel Method to Obtain a Large Number of Human Dental Pulp Cells Through One-Step Cultivation
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Comparative Analysis of M4CXR, an LLM-Based Chest X-Ray Report Generation Model, and ChatGPT in Radiological Interpretation

by
Ro Woon Lee
1,
Kyu Hong Lee
1,*,
Jae Sung Yun
2,
Myung Sub Kim
3 and
Hyun Seok Choi
4
1
Department of Radiology, Inha University College of Medicine, Incheon 22332, Republic of Korea
2
Department of Radiology, Ajou University School of Medicine, Suwon 16499, Republic of Korea
3
Department of Radiology, Kangbuk Samsung Hospital, Sungkyunkwan University School of Medicine, Seoul 03181, Republic of Korea
4
Deepnoid Inc., Seoul 08376, Republic of Korea
*
Author to whom correspondence should be addressed.
J. Clin. Med. 2024, 13(23), 7057; https://doi.org/10.3390/jcm13237057
Submission received: 29 October 2024 / Revised: 15 November 2024 / Accepted: 21 November 2024 / Published: 22 November 2024
(This article belongs to the Section Nuclear Medicine & Radiology)

Abstract

Background/Objectives: This study investigated the diagnostic capabilities of two AI-based tools, M4CXR (research-only version) and ChatGPT-4o, in chest X-ray interpretation. M4CXR is a specialized cloud-based system using advanced large language models (LLMs) for generating comprehensive radiology reports, while ChatGPT, built on the GPT-4 architecture, offers potential in settings with limited radiological expertise. Methods: This study evaluated 826 anonymized chest X-ray images from Inha University Hospital. Two experienced radiologists independently assessed the performance of M4CXR and ChatGPT across multiple diagnostic parameters. The evaluation focused on diagnostic accuracy, false findings, location accuracy, count accuracy, and the presence of hallucinations. Interobserver agreement was quantified using Cohen’s kappa coefficient. Results: M4CXR consistently demonstrated superior performance compared to ChatGPT across all evaluation metrics. For diagnostic accuracy, M4CXR achieved approximately 60–62% acceptability ratings compared to ChatGPT’s 42–45%. Both systems showed high interobserver agreement rates, with M4CXR generally displaying stronger consistency. Notably, M4CXR showed better performance in anatomical localization (76–77.5% accuracy) compared to ChatGPT (36–36.5%) and demonstrated fewer instances of hallucination. Conclusions: The findings highlight the complementary potential of these AI technologies in medical diagnostics. While M4CXR shows stronger performance in specialized radiological analysis, the integration of both systems could potentially optimize diagnostic workflows. This study emphasizes the role of AI in augmenting human expertise rather than replacing it, suggesting that a combined approach leveraging both AI capabilities and clinical judgment could enhance patient care outcomes.
Keywords: chest X-ray; AI; LLM chest X-ray; AI; LLM

Share and Cite

MDPI and ACS Style

Lee, R.W.; Lee, K.H.; Yun, J.S.; Kim, M.S.; Choi, H.S. Comparative Analysis of M4CXR, an LLM-Based Chest X-Ray Report Generation Model, and ChatGPT in Radiological Interpretation. J. Clin. Med. 2024, 13, 7057. https://doi.org/10.3390/jcm13237057

AMA Style

Lee RW, Lee KH, Yun JS, Kim MS, Choi HS. Comparative Analysis of M4CXR, an LLM-Based Chest X-Ray Report Generation Model, and ChatGPT in Radiological Interpretation. Journal of Clinical Medicine. 2024; 13(23):7057. https://doi.org/10.3390/jcm13237057

Chicago/Turabian Style

Lee, Ro Woon, Kyu Hong Lee, Jae Sung Yun, Myung Sub Kim, and Hyun Seok Choi. 2024. "Comparative Analysis of M4CXR, an LLM-Based Chest X-Ray Report Generation Model, and ChatGPT in Radiological Interpretation" Journal of Clinical Medicine 13, no. 23: 7057. https://doi.org/10.3390/jcm13237057

APA Style

Lee, R. W., Lee, K. H., Yun, J. S., Kim, M. S., & Choi, H. S. (2024). Comparative Analysis of M4CXR, an LLM-Based Chest X-Ray Report Generation Model, and ChatGPT in Radiological Interpretation. Journal of Clinical Medicine, 13(23), 7057. https://doi.org/10.3390/jcm13237057

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop