The Application of Large Language Models (LLMs) and Vision-Language Models (VLMs) in Healthcare

A Special Issue of Diagnostics (ISSN 2075-4418) belonging to the section "Machine Learning and Artificial Intelligence in Diagnostics".

Deadline for manuscript submissions: closed (31 May 2026) | Viewed by 6570

Editors


E-Mail Website
Guest Editor
Department of Software Engineering and Management, National Kaohsiung Normal University, Kaohsiung, Taiwan
Interests: artificial intelligence; data visualization; natural language processing; CDSS alert system
Special Issues, Collections and Topics in MDPI journals

E-Mail Website
Guest Editor
School of Health Care Administration, Taipei Medical University, Taipei 11031, Taiwan
Interests: Internet of Things; healthcare management; artificial intelligence; data visualization
Special Issues, Collections and Topics in MDPI journals

Special Issue Information

Dear Colleagues,

The rapid development of Large Language Models (LLMs) and Vision-Language Models (VLMs) is reshaping the landscape of healthcare research and practice. These models enable new levels of performance in natural language understanding, multimodal reasoning, and clinical data interpretation. Their applications range from automated clinical documentation and medical imaging analysis to personalized treatment recommendations and healthcare management, offering both opportunities and challenges for integration into real-world clinical environments.

This Special Issue invites submissions that explore theoretical advances, practical implementations, and case studies on the application of LLMs and VLMs in healthcare. Topics of interest include, but are not limited to, natural language processing for electronic health records, multimodal analysis for medical imaging and diagnostics, clinical decision support, healthcare informatics, and AI-driven patient engagement. We also welcome contributions addressing issues of interpretability, fairness, ethics, and governance in the deployment of these technologies.

By contributing to this Special Issue, authors will help illuminate the current progress and future directions of AI in healthcare, fostering a multidisciplinary dialogue that connects researchers, clinicians, and policymakers. This collaborative effort aims to accelerate innovation while ensuring safe, equitable, and impactful adoption of LLMs and VLMs in healthcare systems worldwide.

Dr. Shuo-Chen Chien
Prof. Dr. Wen-Shan Jian
Guest Editors

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Diagnostics is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2600 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • large language models (LLMs)
  • vision-language models (VLMs)
  • artificial intelligence in healthcare
  • natural language processing (NLP)
  • clinical decision support systems
  • medical imaging analysis
  • multimodal learning
  • healthcare data analytics
  • patient-centered care
  • ethical AI in medicine

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (4 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

23 pages, 2879 KB  
Article
Digital Registrar: A Schema-First Framework for Multi-Cancer Privacy-Preserving Pathology Abstraction via Local LLMs
by Nan-Haw Chow, Han Chang, Hung-Kai Chen, Chen-Yuan Lin, Ying-Lung Liu, Po-Yen Tseng, Li-Ju Shiu, Yen-Wei Chu, Pau-Choo Chung and Kai-Po Chang
Diagnostics 2026, 16(11), 1644; https://doi.org/10.3390/diagnostics16111644 - 27 May 2026
Viewed by 1209
Abstract
Background/Objectives: Free-text surgical pathology reports hinder automated cancer registry entry and secondary analytics. This study introduces a clinically governed schema layer for interoperability, testing whether a locally-deployable Large Language Model (LLM) pipeline can deliver robust registry-grade extraction across institutions. Methods: We developed a [...] Read more.
Background/Objectives: Free-text surgical pathology reports hinder automated cancer registry entry and secondary analytics. This study introduces a clinically governed schema layer for interoperability, testing whether a locally-deployable Large Language Model (LLM) pipeline can deliver robust registry-grade extraction across institutions. Methods: We developed a College of American Pathologists (CAP)-aligned clinical ontology encompassing 10 cancer types, 192 per-organ scalar fields, key biomarkers, and nested structures for lymph nodes and margins. Encoded via Declarative Self-improving Python (DSPy) signatures with grammar-constrained decoding using DSPy v3.2.1, this model-agnostic pipeline was benchmarked on 893 internal reports against a pathologist-adjudicated gold standard. External validation utilized 242 The Cancer Genome Atlas (TCGA) reports. Hardware feasibility was confirmed on a single 48-gigabyte (GB) Graphics Processing Unit (GPU), ensuring suitability for privacy-preserving on-premises deployment. Results: Using the gpt-oss-20b model, the framework achieved 92.0% macro-mean exact-match accuracy on internal data, demonstrating near-perfect run-to-run reliability. Critical prognostic indicators, including breast estrogen receptor/progesterone receptor (ER/PR) (98.7%) and margin positivity (>93%), maintained high fidelity. On the external TCGA cohort, accuracy was 77.5%, rising to 88.0% after excluding structurally silent fields absent in older narratives. Operationally, the model processed reports in 40–70 s, optimally balancing speed and accuracy. Conclusions: This schema-first abstraction layer successfully decouples clinical logic from specific Artificial Intelligence (AI) models. By reliably transforming narrative reports into machine-readable structures, it establishes a portable privacy-preserving foundation for automated cancer surveillance, institutional data reuse, and future multimodal clinical systems. Full article
Show Figures

Figure 1

14 pages, 1947 KB  
Article
Prompt Framing Modulates Safety in Shoulder and Elbow Red-Flag Vignettes: A Large Language Model Study
by Mehmet Yiğit Gökmen, Mehmet Maden and Onur Zengin
Diagnostics 2026, 16(10), 1439; https://doi.org/10.3390/diagnostics16101439 - 8 May 2026
Viewed by 493
Abstract
Background: Large language models (LLMs) are increasingly used for musculoskeletal health information, yet their safety in time-sensitive shoulder and elbow presentations with red-flag features remains insufficiently defined. We evaluated safety behavior using standardized vignettes, focusing on safety-critical under-triage and prompt-dependent performance differences. [...] Read more.
Background: Large language models (LLMs) are increasingly used for musculoskeletal health information, yet their safety in time-sensitive shoulder and elbow presentations with red-flag features remains insufficiently defined. We evaluated safety behavior using standardized vignettes, focusing on safety-critical under-triage and prompt-dependent performance differences. Methods: Eighty fictional vignettes (40 shoulder, 40 elbow) were created and classified a priori as red-flag (n = 24) or non-urgent (n = 56). Each vignette was queried in a single-turn format using three fixed prompt types (patient-, general physician-, and specialist-oriented), yielding 240 responses. Two blinded orthopedic surgeons rated outputs using a prespecified 0–8 rubric across four domains. Safety-critical under-triage was defined as failure to recommend timely urgent evaluation in red-flag presentations. Decision stability was assessed using 20 paired vignette sets differing by one predefined clinical variable. Results: The overall mean score was 6.42 ± 1.12 and was lower for red-flag than for non-urgent responses (5.28 ± 1.21 vs. 6.93 ± 0.81). Across the 72 prompt-specific responses generated for the 24 red-flag vignettes, urgency was correctly recognized in 53 responses (73.6%). Safety-critical under-triage occurred in 19 of 72 red-flag responses (26.4%) and was most frequent with patient-oriented prompts (10/24, 41.7%), followed by general physician-oriented prompts (6/24, 25.0%) and specialist-oriented prompts (3/24, 12.5%). Decision instability, defined as an inconsistent directional change after modification of a single risk-related variable, occurred in 6 of 20 paired vignette sets (30.0%). Conclusions: The evaluated LLM performed consistently well in non-urgent scenarios but showed prompt-dependent safety vulnerabilities in red-flag conditions, driven primarily by under-recognition of urgency. These findings support caution for unsupervised patient-facing use, highlight the need for explicit safeguards in high-risk presentations, and underscore the value of safety-focused evaluation frameworks in musculoskeletal care. Full article
Show Figures

Figure 1

17 pages, 933 KB  
Article
Comparative Evaluation of Five Multimodal Large Language Models for Medical Laboratory Image Recognition: Impact of Prompting Strategies on Diagnostic Accuracy
by Hui-Ru Yang, Kuei-Ying Lin, Ping-Chang Lin, Jih-Jin Tsai and Po-Chih Chen
Diagnostics 2026, 16(9), 1258; https://doi.org/10.3390/diagnostics16091258 - 22 Apr 2026
Viewed by 854
Abstract
Background: Multimodal large language models (MLLMs) show promise in medical imaging, but their performance is highly dependent on prompt engineering. This study systematically evaluates how different prompting strategies affect diagnostic accuracy in clinical laboratory image interpretation. Methods: We evaluated five MLLMs (ChatGPT-4o, Gemini [...] Read more.
Background: Multimodal large language models (MLLMs) show promise in medical imaging, but their performance is highly dependent on prompt engineering. This study systematically evaluates how different prompting strategies affect diagnostic accuracy in clinical laboratory image interpretation. Methods: We evaluated five MLLMs (ChatGPT-4o, Gemini 2.0 Flash, Claude 3.5 Sonnet, Grok-2, and Perplexity Pro (Claude 3.5 Sonnet)) using 177 proficiency testing images across three domains: blood smears (n = 78), urinalysis (n = 50), and parasitology (n = 49). Three prompting approaches were compared: (1) complex multi-choice prompts with 20 diagnostic options, (2) zero-shot open-ended prompts, and (3) two-step descriptive-reasoning prompts. Images were sourced from the Taiwan Society of Laboratory Medicine external quality assurance archives with expert consensus diagnoses. Results: Zero-shot prompting significantly outperformed complex multi-choice prompts across all models and domains (p < 0.001). With zero-shot prompts, Gemini achieved 78.5% overall accuracy (urinalysis: 92.0%; parasitology: 75.5%; blood smears: 64.1%), representing a 17% improvement over complex prompts. Two-step descriptive-reasoning prompts further improved blood smear accuracy by 8–12% for top-performing models, but showed minimal benefit in urinalysis and parasitology. The re-query mechanism (“please reconsider”) improved urinalysis accuracy by 7.6% but had a negligible effect on blood smears and parasitology. Conclusions: Prompting strategy critically determines MLLM diagnostic performance. Zero-shot approaches with minimal constraints consistently outperform complex multi-choice formats. The remarkable performance of general-purpose models in structured domains like urinalysis (>90% accuracy) demonstrates the considerable progress of multimodal AI. However, complex morphological tasks like blood smear interpretation require either specialized prompting techniques or domain-specific fine-tuning. These findings provide evidence-based guidance for optimizing AI integration in clinical laboratories. Full article
Show Figures

Figure 1

16 pages, 650 KB  
Article
Evaluating Medical Text Summaries Using Automatic Evaluation Metrics and LLM-as-a-Judge Approach: A Pilot Study
by Yuriy Vasilev, Irina Raznitsyna, Anastasia Pamova, Tikhon Burtsev, Tatiana Bobrovskaya, Pavel Kosov, Anton Vladzymyrskyy, Olga Omelyanskaya and Kirill Arzamasov
Diagnostics 2026, 16(1), 3; https://doi.org/10.3390/diagnostics16010003 - 19 Dec 2025
Cited by 4 | Viewed by 3106
Abstract
Background: Electronic health records (EHRs) remain a vital source of clinical information, yet processing these heterogeneous data is extremely labor-intensive. Summarization of these data using Large Language Models (LLMs) is considered a promising tool to support practicing physicians. Unbiased, automated quality control is [...] Read more.
Background: Electronic health records (EHRs) remain a vital source of clinical information, yet processing these heterogeneous data is extremely labor-intensive. Summarization of these data using Large Language Models (LLMs) is considered a promising tool to support practicing physicians. Unbiased, automated quality control is crucial for integrating the tools into routine practice, saving time and labor. This pilot study aimed to assess the potential and constraints of self-contained evaluation of summarization quality (without expert involvement) based on automatic evaluation metrics and LLM-as-a-judge. Methods: The summaries of text data from 30 EHRs were generated by six open-source low-parameter LLMs. The medical summaries were evaluated using standard automatic metrics (BLEU, ROUGE, METEOR, BERTScore) as well as the LLM-as-a-judge approach using the following criteria: relevance, completeness, redundancy, coherence and structure, grammar and terminology, and hallucinations. Expert evaluation was conducted using the same criteria. Results: The results showed that LLMs hold great promise for summarizing medical data. Nevertheless, neither the evaluation metrics nor LLM judges are reliable in detecting factual errors and semantic distortions (hallucinations). In terms of relevance, the Pearson correlation between the summary quality score and the expert opinions was 0.688. Conclusions: Completely automating the evaluation of medical summaries remains challenging. Further research should focus on dedicated methods for detecting hallucinations, along with investigating larger or specialized models trained on medical texts. Additionally, the potential integration of retrieval-augmented generation (RAG) within the LLM-as-a-judge architecture deserves attention. Nevertheless, even now, the combination of LLMs and the automatic evaluation metrics can underpin medical decision support systems by performing initial evaluations and highlighting potential shortcomings for expert review. Full article
Show Figures

Figure 1

Back to TopTop