Next Article in Journal
An Explainable Deep Learning Framework for Musculoskeletal Abnormality Detection
Previous Article in Journal
A Fast and Generalizable Deep Neural Network for the Detection of Atrial Fibrillation
Previous Article in Special Issue
Digital Twins for Hospital and Healthcare Operations: A Systematic Review of Resource Allocation, Infection Control, and Workflow Optimization
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Artificial Intelligence in Healthcare: From Predictive Models to Generative, Multimodal, and Agentic AI

1
Department of Audiovisual Techniques and Media Production, Vocational School of Technical Sciences, Fırat University, Elazığ 23119, Türkiye
2
Vocational School of Technical Sciences, Fırat University, Elazığ 23119, Türkiye
3
Department of Psychiatry, Elazığ Fethi Sekin City Hospital, Elazığ 23280, Türkiye
4
Department of Digital Forensics Engineering, Technology Faculty, Fırat University, Elazığ 23119, Türkiye
*
Authors to whom correspondence should be addressed.
Bioengineering 2026, 13(10), 1175; https://doi.org/10.3390/bioengineering13101175
Submission received: 7 September 2026 / Revised: 3 October 2026 / Accepted: 7 October 2026 / Published: 8 October 2026

Abstract

Artificial intelligence (AI) is increasingly used across the diagnostic pathway, from screening and disease detection to differential diagnosis, prognostic stratification, and treatment-response assessment. This narrative review examines predictive, generative, multi-modal, and agentic AI through a diagnostic-science lens and evaluates the evidence required for translation into clinical practice. Predictive models can identify complex patterns in images, physiological signals, laboratory data, and electronic health records, but high retrospective accuracy does not establish diagnostic utility. Clinical validity and utility depend on the intended use, target population, disease spectrum and prevalence, reference standard, operating threshold, calibration, and the consequences of false-positive and false-negative results. Generative AI may support problem representation, differential diagnosis, information synthesis, and documentation, yet fluent outputs can omit critical alternatives, amplify false premises, or convey unjustified certainty. Multimodal systems may better reflect clinical reasoning by integrating imaging, text, signals, pathology, and molecular data, but they must demonstrate incremental value over the best single-modality test and remain robust to missing data. Agentic systems can coordinate evidence retrieval, test selection, and sequential workflows, thereby increasing the need for bounded permissions, auditability, error recovery, and human escalation. Across all paradigms, a credible pathway to diagnostic implementation requires independent external and prospective validation, representative consecutive patients, subgroup analysis, workflow and human–AI evaluation, clinical impact studies, and post-deployment surveillance. The field should therefore be judged not by model capability alone, but by whether AI measurably improves diagnostic yield, timeliness, safety, equity, and patient-relevant outcomes in real care settings.

1. Introduction

Artificial intelligence (AI) is increasingly used across biomedical research and clinical care, including diagnosis, prognosis, risk stratification, medical imaging, physiological signal analysis, and decision support [1,2,3,4]. Much of this progress has been driven by task-specific machine-learning and deep-learning models developed for well-defined clinical problems. In some areas, particularly image-based diagnosis, these systems have achieved performance comparable to that of healthcare professionals [5]. However, strong benchmark results do not always translate into clinical benefit, and more complex models are not necessarily better than simpler, well-designed approaches [6]. The focus has therefore shifted from accuracy alone to whether AI systems remain reliable across patients, institutions, and clinical settings and whether their outputs can meaningfully support patient care [2,4].
From a diagnostic science perspective, AI should be evaluated as part of a test-to-decision pathway rather than as an isolated classifier. Relevant use cases include population or opportunistic screening, triage, disease detection, differential diagnosis among competing etiologies, diagnostic confirmation, and prognostic or treatment-response stratification after diagnosis. Each use case has distinct requirements for the intended population, index test and reference standard, disease spectrum and prevalence, operating threshold, uncertainty, and the clinical consequences of false-negative and false-positive results [5,7,8,9,10]. A credible translation pathway must therefore connect the model output to a defined user, clinical decision, downstream action, and measurable patient or workflow outcome.
The emergence of foundation models and generative AI has substantially broadened this landscape. Large language models (LLMs) can perform multiple language-intensive tasks within a shared pretrained architecture, including clinical summarization, question answering, documentation, information synthesis, and decision support [11,12,13,14]. Their flexibility contrasts with the narrower functional scope of many conventional predictive systems, but also introduces distinct limitations. Fluent outputs may contain unsupported or incorrect information, vary with prompts and context, and provide limited insight into the provenance or calibration of generated responses [11,13,15]. In parallel, multimodal AI is extending model inputs beyond a single data type by integrating combinations of text, medical images, electronic health records, physiological measurements, pathology, and molecular information [16]. More recently, agentic AI has introduced the prospect of systems that can combine reasoning with planning, memory, retrieval, tool use, and sequential action [17,18,19]. These developments are not simply successive generations of the same technology. Rather, they represent increasingly interconnected capabilities—prediction, generation, multimodal integration, and goal-directed action—that may coexist within the same clinical AI ecosystem.
The various forms of predictive, generative, multimodal and agentic AI are rapidly converging, so that it is increasingly less useful to consider each in isolation. The growing array of forms of AI in health care is in itself not a problem. Rather, the familiar problems of clinical translation persist: performance may decrease in application to new populations, in new institutions, with new data; seemingly strong results may be attenuated by poor calibration, by bias, by data leakage, by a lack of external validation, or by poor fit to clinical work flows [5,6]. The relevant question is therefore not only whether an AI system performs well, but whether it remains reliable, useful, and safe when introduced into actual care.
Although all clinical AI systems are susceptible to error, the consequences become particularly important when model outputs directly influence high-stakes recommendations or clinical actions. Evaluation should therefore consider not only model capability but also the intended clinical context and the potential consequences of failure. Randomized studies of machine-learning-based decision support systems have shown their potential to influence clinically relevant processes and thus must be evaluated within the respective care processes rather than on the basis of retrospective criteria [20]. Reporting frameworks such as CONSORT-AI and SPIRIT-AI provide standards for trials and trial protocols involving AI interventions [21,22], while FUTURE-AI extends this perspective across the AI lifecycle, emphasizing fairness, universality, traceability, usability, robustness, and explainability [23]. For multimodal and agentic systems, these considerations increasingly apply to the whole clinical system, including interactions between data, models, external tools, clinicians, and patients.
This state-of-the-science narrative synthesis integrates current evidence across predictive, generative, multimodal, and agentic AI and provides a clinically oriented framework for evaluating their translation into healthcare practice. We assess how they contribute to disease detection, differential diagnosis, prognostic stratification, and treatment-response assessment, while also considering documentation, workflow, and governance functions that determine whether diagnostic information can be used safely. The paradigms are treated as convergent capabilities rather than successive generations. Across them, we distinguish technical performance from diagnostic validity, diagnostic utility, and clinical impact, and organize the translational pathway around intended use, credible reference standards, independent external and prospective validation, human–AI interaction, workflow integration, patient-relevant outcomes, and post-deployment surveillance. The central question is not whether increasingly sophisticated systems can generate a prediction or recommendation, but whether their use improves diagnostic decisions and patient care reliably, safely, and equitably.
The scope extends beyond diagnostic model performance to developments that may alter how AI is incorporated into healthcare practice. Accordingly, treatment planning, digital twins, embodied and robotic systems, and autonomous scientific discovery are considered where they represent a meaningful extension from model-level inference toward patient-specific, interactive, or increasingly autonomous forms of clinical and biomedical decision support. Their inclusion is intended to examine the expanding translational boundary of healthcare AI rather than to provide a comprehensive review of each emerging field. This distinction also allows established clinical applications and more exploratory technologies to be discussed within the same framework without implying equivalent levels of evidence or clinical maturity.
This review is structured to move from conceptual evolution to clinical translation. We first establish the diagnostic and translational framework used to evaluate healthcare AI, then examine predictive, generative, multimodal, and agentic AI according to their distinct capabilities, evidence base, and failure modes. Subsequent sections consider applications across clinical practice, cross-cutting requirements for trustworthy implementation, and emerging directions that may extend current models of AI-enabled care.

Literature Search Strategy

A targeted literature search was conducted in Web of Science Core Collection, Scopus, and PubMed to identify relevant studies for each of the review themes. A set of appropriate search terms was defined for each section of the review. Where possible, primary studies providing clinically relevant data were given preference, particularly those reporting external validation, multi-center studies, prospective studies, randomized trials and implementation studies in clinical settings. Additionally, methodological papers and reviews from authoritative sources were included to address certain clinical areas where there is currently a lack of relevant clinical evidence and to provide context to the review of healthcare AI development and clinical application. Studies were selected on the basis of clinical relevance, methodological quality and relevance to the development and clinical application of healthcare AI. The literature was then narratively and integratively reviewed as opposed to a systematic review or meta-analysis. To strengthen the diagnostic focus, search terms and study selection also addressed screening and disease detection, differential diagnosis, diagnostic accuracy, reference standards, prognostic stratification, external and prospective validation, diagnostic workflow, and clinical implementation.
When multiple studies addressed similar questions, greater interpretive weight was given to evidence from independent external validation, multicenter studies, prospective evaluations, randomized or comparative studies, human–AI and workflow evaluations, and clinical implementation studies. Retrospective and proof-of-concept studies were included where they provided important evidence in emerging areas but were interpreted according to their corresponding level of clinical maturity. This evidence-prioritization approach was used to support a clinically oriented synthesis while maintaining the narrative and integrative nature of the review.

2. Evolution of Artificial Intelligence in Healthcare

Artificial intelligence in healthcare has evolved from relatively simple predictive models based on structured clinical data to more complex learning systems capable of processing large and heterogeneous healthcare datasets. Representation learning methods that enable direct extraction of relevant information from high-dimensional data such as medical images and signal data have been particularly successful. Furthermore, generative, multimodal, and agentic AI models have recently progressed to stages of content generation, multi-data fusion, and goal-directed actions, thereby significantly changing not only the scope of applications of AI in healthcare but also its clinical utilization and assessment. In diagnostic science, this evolution expands the claim under evaluation from single-endpoint classification to context-dependent differential diagnosis, multimodal evidence integration, and multi-step test-to-action pathways; accordingly, validation must remain anchored to the intended diagnostic use rather than to architectural novelty.

2.1. From Task-Specific Prediction to Representation Learning

The early expansion of artificial intelligence in medicine was dominated by task-specific models developed for a predefined input, endpoint, and clinical context. Conventional machine-learning pipelines typically depended on variables selected or engineered before model fitting, making performance closely coupled to the quality of data representation and the assumptions embedded in feature construction. This paradigm proved effective for structured clinical data and well-defined prediction problems, but became increasingly restrictive as biomedical datasets expanded in scale, dimensionality, and complexity [2,3,24].
Deep learning altered this relationship between data representation and prediction. Rather than relying primarily on manually specified features, multilayer neural networks enabled hierarchical representations to be learned directly from raw or minimally processed data. This shift was particularly consequential for information-rich modalities such as medical images and physiological signals, in which clinically relevant patterns are often distributed across high-dimensional inputs. Reviews of physiological-signal applications, for example, documented the increasing use of deep architectures across electrocardiographic, electroencephalographic, electromyographic, and electro-oculographic data and highlighted their capacity to exploit larger and more heterogeneous datasets. More broadly, the transition from conventional machine learning to deep learning has been characterized as a data-driven shift in which model development moved progressively from explicit feature design towards learned representations.
This transition did not eliminate the limitations of task-specific modelling. Most deep-learning systems remained optimized for a particular dataset and endpoint, and gains in representation learning did not inherently confer portability across institutions, populations, or tasks. As deployment experience accumulated, attention therefore moved from model architecture alone towards the data conditions under which models are developed and maintained. Zhang and colleagues framed this change as a shift from model-centric development towards a more data-centric perspective, emphasizing data availability, distribution shift, and the infrastructure required to sustain performance after deployment. Representation learning thus expanded the range of tractable medical data, but the prevailing paradigm remained largely one model for one clinical objective.

2.2. Transformers and Foundation Models as an İnflection Point

Transformers marked a more fundamental change because they enabled representations learned at scale to be reused across a wider range of downstream tasks. Initially developed for natural-language processing, transformer architectures were subsequently adapted to clinical text, electronic health records, medical imaging, physiological signals, and biomolecular sequences. A recent review of transformer-based healthcare applications documents this expansion across multiple forms of biomedical data and across tasks including diagnosis, reconstruction, report generation, and outcome prediction. Their importance for healthcare therefore extends beyond improvements in language modelling; transformers provided a scalable architecture for learning context-dependent representations from large and heterogeneous datasets.
Foundation models extended this principle further. Instead of training a model de novo for every clinical endpoint, large-scale pretraining creates a reusable representational substrate that can subsequently be adapted through fine-tuning, prompting, in-context learning, retrieval, or other task-specific mechanisms. This changes the unit of development from an isolated predictive model to a general pretrained model capable of supporting multiple downstream functions. In healthcare, this approach has been most visible in large language models, but the same principle increasingly applies to imaging, pathology, molecular data, and multimodal systems [11,15,25,26].
This distinction is crucial. In essence, the high parameter count of a foundation model does not automatically qualify every large neural network to be considered a foundation model. Rather, the key property of a foundation model is the transferability of the learned representations to other tasks, as well as (potentially) to other data domains. While the ability to learn to complete a large set of tasks in a single training run can potentially reduce the need to train many individual models to complete a single task, the sources of uncertainty in the model’s behavior are re-distributed. That is, in addition to the uncertainty inherent in any machine learning model, there is also uncertainty related to the pretraining data, the particular adaptation approach used for a given task, the model–task alignment, and how well the pretraining representations generalize to the specific clinical context in which they are applied. Thus, while the broad capability of a foundation or large language model to complete a wide variety of tasks is an important attribute, such broad capability does not automatically translate to clinical competence in all contexts [11,15,26,27]. Harrer, for example, argues that these systems are more appropriately developed as assistive technologies under human oversight than as substitutes for clinical decision makers.

2.3. From General-Purpose Models to İntegrated AI Systems

Once the performance of single tasks is improved by their corresponding AI models, the potential of various tasks is further increased by combining them. Generative models extended this capability by producing text, images, and structured outputs. Then, multimodal models can combine different data sources, while tool-enabled systems connect AI models with external knowledge, software and computing power. The various earlier AI approaches are not replaced but rather combined in clinical AI systems in order to increase their scope of application.
Multimodal AI systems aim to capture relationships between different sources of information that a clinician would normally interpret together. This could be, for example, medical images, narrative reports, structured health records, and molecular or physiological measurements over time. The type of information included would depend on the specific clinical problem being addressed [28]. Foundation-model approaches are increasingly being developed to align such information within shared or interoperable representational spaces, allowing a single system to support tasks that previously required separate modelling pipelines.
A further extension occurs when models are embedded within systems that maintain state, select tools, decompose tasks, and execute multi-step workflows. Biomedical AI-agent frameworks already describe systems in which language and generative models are combined with structured memory, scientific knowledge, computational tools, and experimental platforms. This architecture changes what constitutes an AI system: the clinically relevant object may no longer be one model producing one output, but an orchestrated sequence of models, data sources, tools, and human interventions. Recent healthcare literature similarly distinguishes single-turn conversational systems from modular, multistep agents capable of tool use and bounded workflow execution.
The more consequential change is the widening scope of what an AI system can coordinate. Healthcare AI has progressed from learning a mapping between inputs and outcomes towards building reusable representations and, increasingly, systems that can combine generation, heterogeneous evidence, external tools, and sequential computation. The corresponding unit of evaluation must therefore expand from individual model performance to the reliability of the integrated system, a distinction that becomes central in the sections that follow. This evolution from task-specific prediction towards increasingly integrated AI capabilities is summarized in Figure 1.
The conceptual trajectory of healthcare AI from task-specific predictive models to representation learning and foundation models, followed by the emergence of multimodal and generative capabilities and their integration into agentic systems. The progression represents an expansion of model scope and system-level capability rather than the replacement of earlier paradigms. The clinical and translational characteristics of these AI paradigms are summarized in Table 1.

3. Predictive AI in Diagnosis: From Performance to Clinical Utility

Predictive AI is one of the most established uses of artificial intelligence in medicine. It has been applied to disease risk estimation, prognosis, future clinical events, and treatment response, often using complex and heterogeneous clinical data. Its advantage lies in the ability to capture patterns that may be difficult to represent with conventional clinical scores. However, greater model complexity does not necessarily lead to better clinical performance. Machine-learning models have not consistently outperformed well-specified statistical approaches, particularly in studies with limited sample sizes, inadequate validation, or high risk of bias [6]. The key question is therefore no longer whether a model can predict accurately but whether its predictions are reliable across settings, well calibrated, clinically actionable, and capable of improving care.

3.1. Risk Prediction and Stratification

Risk prediction is a natural application of clinical AI because many medical decisions depend on estimating the likelihood of future events. Machine-learning models have been used to predict deterioration, complications, recurrence, hospitalisation, mortality, and other clinically important outcomes using electronic health records, laboratory data, imaging-derived features, physiological signals, and longitudinal information [29,30,31,32,33].
These models can capture nonlinear relationships and interactions across large numbers of predictors, which may be difficult to represent in conventional risk scores based on a limited set of predefined variables. However, greater complexity does not necessarily provide greater predictive value. Comparative studies have shown that the apparent advantage of machine learning over logistic regression often decreases when methodological quality, sample size, and risk of bias are taken into account [6]. The relevant comparison is therefore not between newer and older algorithms but between competing approaches evaluated under similar conditions.
High discrimination alone is insufficient for clinical use. A model may rank patients accurately while still producing poorly calibrated risk estimates or failing to identify clinically useful decision thresholds. A predictive model becomes clinically useful when its estimated probabilities are adequately calibrated to observed outcomes and its outputs meaningfully inform actionable clinical decisions. This requirement is particularly important when model outputs influence monitoring intensity, escalation of care, or allocation of limited healthcare resources.

3.2. Diagnostic Prediction: Disease Detection and Differential Diagnosis

Diagnostic AI has provided some of the clearest demonstrations of high-dimensional pattern recognition in medicine. Meta-analytic evidence shows that deep-learning systems can achieve high diagnostic performance across several image-intensive specialties, including radiology, ophthalmology, pathology, and endoscopy [7,34,35,36]. These findings established technical feasibility but also exposed an important limitation of the early literature: performance was frequently estimated in retrospective, curated datasets that differed substantially from the populations in which the systems would ultimately be used [7].
The distinction between benchmark accuracy and clinical diagnostic performance is consequential. Disease prevalence, case spectrum, acquisition protocols, reference standards, and referral pathways can all alter model behaviour. Comparisons between an algorithm and clinicians on a static test set therefore provide limited evidence about how the same system will perform when confronted with consecutive patients, uncertain presentations, workflow constraints, and changing data distributions.
Diagnostic studies should specify the target condition and relevant alternatives, care setting, intended population, index test, reference standard, and the model’s role as a replacement, triage, add-on, or second-reader test. A screening system may prioritize sensitivity and manageable referral burden, whereas a confirmatory aid may require high specificity and calibrated probabilities. A differential-diagnosis system must rank clinically plausible alternatives, identify discriminating evidence, and signal when available data are insufficient. False-negative and false-positive consequences should therefore be evaluated at clinically selected thresholds rather than inferred from AUROC alone [7,8,9,10].
External and prospective studies provide a more stringent test. Clinical validation of diabetic retinopathy screening in African populations demonstrated the importance of evaluating algorithms beyond the populations in which they were initially developed [37]. Similar principles have been examined in externally validated systems for pulmonary nodule malignancy [38], pancreatic cancer [39], coronary calcium assessment [40], and AI-assisted pathology [41,42]. Collectively, these studies mark a shift from asking whether an algorithm can recognise disease towards asking whether diagnostic performance survives changes in population and clinical setting.
The relevant endpoint is therefore no longer a marginal improvement in AUROC or sensitivity alone. For mature diagnostic AI, the evidentiary standard should increasingly include independent validation, representative case spectra, clinically relevant thresholds, and evaluation of how algorithmic information modifies human decisions.
For practical implementation, each output must be linked to a prespecified diagnostic action, such as repeat testing, specialist referral, biopsy, treatment initiation, or safe discharge, and to a failure or escalation pathway. Prospective diagnostic studies should recruit consecutive or otherwise representative patients, use blinded and clinically credible reference standards, retain indeterminate and missing cases in the analysis, and report diagnostic yield, time to diagnosis, additional testing, net benefit, downstream harms, and subgroup performance. This test-to-action chain distinguishes diagnostic validity from diagnostic utility.

3.3. Prognostic Stratification After Diagnosis and Treatment-Response Prediction

Prognostic AI addresses a different clinical question: not whether disease is present, but what is likely to happen after diagnosis. Predictive models have been developed to estimate survival, relapse, disease progression, complications, and functional recovery across a wide range of clinical conditions [30,31,32,33]. Oncology has been a particularly active area for AI-based prognostic modelling because of the availability of large, longitudinal, and increasingly multimodal clinical datasets. Prognosis refers to the expected course or probability of future clinical outcomes after a defined index point, such as diagnosis or treatment initiation. In clinical studies, the clinical outcome after a clinical decision may change with time, follow-up information may be incomplete and study endpoints may be defined differently by studies and institutions. Therefore, a good predictor of a negative clinical outcome does not necessarily provide insight into the reasons for the prediction or the potential effects of any clinical intervention to change the predicted outcome. A predictor of treatment response must distinguish between predicting poor outcome for a given patient and predicting that a given patient is more likely to benefit from one treatment vs. another. In other words, association of a clinical parameter with an outcome does not necessarily provide evidence for its therapeutic use.
While AI-based predictive models can be of great value to precision medicine by enabling decisions in the context of large amounts of data, their inclusion in treatment decisions requires more than good prognostic performance of a model.

3.4. External Validation and Generalisability

External validation is a critical step between model development and clinical use. Internal validation can estimate performance within the development setting, but it cannot show whether a model will perform similarly in different hospitals, populations, devices, or periods of clinical practice. External validation addresses this question by testing the transportability of model performance beyond the data on which it was developed [8].
This matters because clinical environments are rarely identical. Disease prevalence and referral patterns vary across regions, patient populations differ in baseline risk, and changes in clinical practice can alter treatment and coding. Differences in scanners, laboratory assays, and electronic health-record systems may further shift the data presented to the model.
Performance deterioration under these conditions is not an exceptional failure of AI but an expected consequence of dataset shift.
Prospective and multicentre evaluations provide increasingly important evidence of this transition. Studies have evaluated AI systems across independent populations for diabetic retinopathy [37], pulmonary nodules [38], left-ventricular systolic dysfunction [43], pancreatic cancer [39], coronary calcium scoring [40], developmental screening [44], diabetic-retinopathy programmes involving tens of thousands of patients [45], acute kidney injury [46], and breast-cancer prognostication [47]. Their collective importance lies less in the individual clinical domains than in demonstrating that model performance must be re-established when the data-generating environment changes.
Generalisability also cannot be judged by discrimination alone. AUROC describes ranking ability but does not establish whether a predicted probability corresponds to the observed event rate. Calibration is therefore essential whenever predictions support threshold-based decisions. Evaluation should consequently extend beyond accuracy, sensitivity, specificity, and AUROC to include calibration and, where appropriate, decision-analytic measures that quantify whether predictions improve decisions at clinically relevant thresholds [9].
External validation should thus be interpreted as a minimum requirement for translation rather than evidence of clinical effectiveness. A transportable model may still fail to improve care.

3.5. From Prediction to Diagnostic Utility and Patient Benefit

The most consequential gap in predictive AI lies between knowing that a model predicts accurately and knowing that using the model benefits patients. These are distinct evidentiary questions.
An externally validated model can remain clinically ineffective if its predictions arrive too late, duplicate information already available to clinicians, generate excessive alerts, encourage inappropriate intervention, or cannot be incorporated into existing workflows. Conversely, a model with only modest improvements in discrimination may be valuable if it reliably identifies a decision threshold at which management changes and patient benefit outweighs harm. The clinical evaluation is not a single validation step but rather a step in a chain of steps that need to be taken to determine the potential of AI in clinical settings (see Table 2). Technical evaluation establishes whether the model can reliably capture the intended signal, whereas external validation assesses whether performance is transportable to independent populations and settings. Workflow evaluation examines interactions among the AI system, clinicians, patients, and existing clinical processes, including potential safety and human-factor risks [10]. Clinical-impact studies subsequently determine whether AI-assisted care improves clinically meaningful processes or patient outcomes, ideally through comparative or randomized evaluations where appropriate [21,22]. This final transition remains comparatively underdeveloped. Prospective evaluations such as large-scale diabetic-retinopathy screening demonstrate that AI systems can be assessed under operational conditions [45], while deployment studies in pathology have begun to examine AI within clinician-facing diagnostic workflows [41,42]. Randomised evidence provides an even stronger test. In the HYPE trial, a machine-learning-derived early-warning system for intraoperative hypotension was evaluated against standard care, illustrating the methodological shift from assessing predictive accuracy to assessing the consequences of model-guided intervention [20].
The distinction is central to the interpretation of progress in predictive AI. A higher AUROC is not synonymous with better medicine. Clinical maturity requires evidence that predictions remain reliable outside the development dataset, provide information that changes an appropriate decision, and produce benefits that exceed the harms and costs introduced by the system. Evaluation must also continue after deployment because changes in populations, clinical practice, and data-generating systems can alter performance over time; contemporary consensus guidance therefore treats monitoring as part of the AI lifecycle rather than an endpoint after implementation [23].
Predictive AI has the most established translational evidence base, extending from retrospective and external validation to prospective and clinical-impact evaluation in selected applications. However, the level of evidence remains heterogeneous across clinical domains.

4. Generative AI in Diagnosis: From Language Models to Clinical Copilots

Generative artificial intelligence has altered the functional scope of clinical AI by moving beyond predefined outputs towards the synthesis, transformation, and contextualisation of medical information. Unlike conventional clinical prediction models designed to estimate a predefined class, probability, or risk endpoint, LLMs generate context-dependent sequences and can support a broader range of information-intensive clinical tasks. In clinical settings, these capabilities can support multistep tasks such as extracting relevant information, generating differential diagnoses, summarizing patient records, and presenting medical information in a form that is easier for patients to understand. They can be used to create clinical documents and respond to subsequent questions as more information becomes available. Early demonstrations established substantial medical knowledge within large pretrained models [48], but subsequent clinical studies have made clear that knowledge retrieval and clinically reliable reasoning are not equivalent. The relevant translational question is therefore no longer whether LLMs can generate plausible medical text but whether their outputs remain correct, context-sensitive, and useful when incorporated into human clinical reasoning and care workflows.

4.1. Clinical Reasoning Beyond Benchmark Performance

Medical examinations and vignette-based benchmarks initially provided convenient measures of LLM capability, but they test only a restricted subset of clinical reasoning. Real clinical decisions require sequential information gathering, interpretation of incomplete or conflicting evidence, adherence to guidelines, adjustment to changing context, and recognition of uncertainty. When these demands are incorporated into evaluation, performance becomes substantially less uniform.
Hager and colleagues evaluated state-of-the-art LLMs across 2400 real patient cases in a simulated clinical decision environment and found important deficiencies in diagnosis, guideline adherence, interpretation of laboratory information, and sensitivity to the amount and ordering of clinical information [49]. This distinction is important because fluent explanations can create an impression of reasoning even when the underlying inference is unstable.
At the same time, more recent systems demonstrate that generative models can contribute meaningful diagnostic information under appropriately constrained conditions. Prompt structures that require explicit diagnostic reasoning can expose intermediate clinical considerations [50], while newer models have shown improved performance in differential-diagnosis generation [51]. The resulting picture is therefore neither one of clinical equivalence nor of simple failure. LLM capability is highly task-dependent, and performance measured in isolated question-answering tasks cannot be assumed to generalise to longitudinal or high-consequence decisions.

4.2. Human-LLM Collaboration Rather than Autonomous Diagnosis

The clinically relevant unit of evaluation is increasingly the human–AI team rather than the model in isolation. This change is supported by emerging comparative and randomised evidence.
In a randomised clinical trial involving 50 physicians, access to an LLM did not significantly improve diagnostic-reasoning scores compared with conventional resources, despite the LLM performing strongly when evaluated alone [52]. This apparently paradoxical result is instructive. A capable model does not automatically produce a more capable clinical team. Effective augmentation depends on how recommendations are presented, when the system is queried, whether clinicians recognise incorrect suggestions, and how much weight is assigned to model outputs.
Generative AI should be evaluated not only by whether it produces the correct answer, but also by how it influences clinical reasoning. Important effects on information search, on framing the diagnostic problem, on dealing with uncertainty, on time to decide, on cognitive load, and on automation bias need to be considered. The clinically relevant question is therefore not whether an AI model can independently reach the correct diagnosis, but whether its use improves clinicians’ diagnostic reasoning and accuracy while reducing error. Direct model-versus-clinician comparisons alone provide limited insight into the clinical value of human–AI collaboration. Evaluation should instead determine which components of clinical reasoning can be effectively augmented by AI, which should remain under clinician control, and where independent verification is required.
For diagnostic use, LLMs should not be evaluated only on the correctness of a final answer. They should be tested on problem representation, breadth and prioritization of the differential diagnosis, identification of discriminating findings, selection of the next diagnostic tests, and calibration of uncertainty. A useful copilot should surface dangerous alternatives, ask for missing information, resist false premises, and abstain or escalate when evidence is insufficient. Studies should separately score omission of critical diagnoses, unsupported additions, harmful test recommendations, and clinician overreliance [49,50,51,52,53].

4.3. Generation as a Clinical İnformation İnterface

The near-term value of generative AI may be greatest in tasks centred on processing and communicating clinical information rather than making diagnoses. LLMs can summarise longitudinal records, draft clinical notes, translate technical language, retrieve information from institutional guidance, and adapt information for different users. These tasks make practical use of generative capabilities while preserving a clear role for human review.
Patient communication illustrates both the potential and the limitations of this approach. An evaluation of AI-assisted simplification of hospital discharge documentation found that generated summaries improved readability for patients; however, clinician review identified factual inaccuracies, omissions, and potentially safety-relevant errors [54]. Improved comprehensibility should therefore not be interpreted as equivalent to improved factual reliability.
Performance can also vary substantially across clinical tasks. In an evaluation using clinical vignettes, LLM performance was better for final diagnosis than for initial differential diagnosis or management [53]. Aggregate accuracy may therefore obscure weaknesses that become important when a model is used at different stages of clinical care.
More recent work has begun to evaluate generative AI in bounded clinical settings. PEACH, a perioperative chatbot grounded in 35 institutional protocols, was tested through silent deployment using real-world clinical interactions and achieved high guideline-concordant accuracy with few hallucinations, although its effect on patient outcomes remains unknown [55]. Such systems may offer a more realistic near-term model for clinical generative AI: constrained copilots operating within defined knowledge and workflow boundaries, with clinicians retaining oversight.

4.4. Hallucination, Factuality, and Context-Dependent Failure

Generative AI introduces a different type of failure from conventional predictive models. A predictive model may assign the wrong probability or class, whereas an LLM can produce incorrect information within a fluent and convincing explanation. This makes factual errors harder to recognize and potentially more consequential in clinical settings.
Hallucinations can result from missing knowledge, ambiguous prompts, unsupported inference, retrieval errors, conflicting context, or manipulated inputs. Standard benchmarks may not capture these vulnerabilities. Alber and colleagues showed that introducing a small amount of medical misinformation into training data increased harmful outputs without substantially affecting conventional benchmark performance [56]. Similarly, in physician-validated simulated cases containing false information, LLMs often elaborated on fabricated clinical details; mitigation prompting reduced but did not eliminate this behaviour [57].
In clinical evaluation of factual reliability of AI systems, it is important to test their ability to reject incorrect premises, to distinguish between evidence and assertion, and to express uncertainty. While retrieval augmentation, domain-specific grounding, constrained output spaces, and verification can all reduce errors, none of them on their own is sufficient. Factual reliability is a property of the overall information pathway consisting of the model, the retrieved evidence, the prompting, the user interaction with the system and the surrounding clinical system.

4.5. From General-Purpose Generation to Bounded Clinical Copilots

The evidence increasingly supports a distinction between general capability and clinical readiness. LLMs can demonstrate sophisticated medical knowledge and, in selected tasks, high diagnostic performance; however, the same models can fail under realistic information sequences, propagate false premises, or provide outputs that clinicians do not use effectively [49,52,56,57]. Clinical deployment therefore requires narrowing rather than simply expanding the model’s operating space.
A clinical copilot is defensible only when it is bounded to a clearly specified task, relies on explicitly defined information sources, interacts transparently with the EHR or institutional knowledge base, appropriately represents uncertainty, maintains an auditable record of its outputs and supporting evidence, and incorporates predefined points of clinician oversight and control. Recent studies have begun to evaluate bounded clinical LLMs under real-world or workflow-specific conditions, including perioperative care. For example, protocolized deployment of a simple, task-focused LLM in the form of a decision support tool yielded accurate and useful results [55]. In emergency care, language-model systems have also been developed around clinical records and workflow-specific tasks rather than unrestricted conversational use [58].
This shift has an important implication for the trajectory of generative AI in medicine. Progress should not be judged primarily by increases in parameter count, examination scores, or the apparent sophistication of generated responses. More meaningful advances will come from systems that reliably retrieve the right evidence, preserve clinically relevant context, expose uncertainty, support rather than distort human reasoning, and operate safely within clearly specified workflows. Clinical-system evaluation should therefore extend beyond model-level accuracy to include error propagation, abstention and escalation, clinician override, and the downstream consequences of incorrect outputs.
Generative models are also becoming increasingly capable of jointly processing text and other biomedical data. Early multimodal systems have shown the feasibility of combining clinical language with images for tasks such as dermatological assessment [59] and radiology-report generation [60]. SkinGPT-4, for example, was evaluated on real clinical dermatology cases with specialist input, whereas clinician-vision–language collaboration in radiology demonstrated both potential utility and clinically significant errors in generated reports. Because that transition changes the nature of both the available information and the possible failure modes, multimodal AI is considered separately in the next section rather than treated as an incremental extension of LLMs here.
Evidence for generative AI remains predominantly benchmark-based or retrospective, with comparatively limited prospective, workflow, and clinical-impact evaluation. Its clinical maturity therefore remains below that of established predictive applications.

5. Multimodal AI for Diagnostic Evidence Integration

Multimodal artificial intelligence extends clinical AI beyond the analysis of individual data types by integrating complementary information within a unified computational framework. This approach is particularly relevant to medicine, where diagnosis, prognosis, and treatment decisions routinely depend on the joint interpretation of heterogeneous evidence rather than any single measurement. However, multimodality should not be regarded as an objective in itself. Its clinical value depends on whether integration captures complementary information, remains robust when data are incomplete, and provides measurable benefit over well-performing unimodal approaches. This section examines these principles through multimodal patient representations, vision–language models, incremental clinical value, and the challenges associated with data fusion and missing modalities.

5.1. From İsolated Signals to İntegrated Patient Representations

In many clinical situations, decisions are not made based on the analysis of single data sources, such as imaging, lab results, EHRs, physiological signals, histology or molecular data from tumors. Instead, all of these sources contain partly redundant and partly complementary information on the patient. Instead of analyzing each of these sources for prediction in isolation, they can be modeled jointly by multimodal AI approaches. The value of a particular modality does not necessarily increase with the number of input modalities. Rather, its value is determined by whether it provides additional information that changes discrimination, calibration, or even clinical decisions made by the best unimodal model.
Primary studies from a number of diseases are illustrated in this section, including joint histology–genomic modelling across a number of cancer types to identify prognostic information [61], while integration of radiology, pathology, genomics, and clinical variables improved risk stratification in high-grade serous ovarian cancer [62]. Flexible multimodal frameworks have also combined routinely collected clinical information with neuroimaging for dementia assessment [63] and radiographs with clinical variables for osteoarthritis progression [64]. These studies move the field closer to the way clinicians actually reason, but they also expose a central methodological requirement: multimodal models should be judged against strong unimodal comparators, not against weak or historically convenient baselines.

5.2. Vision–Language Models as a Clinically İmportant Multimodal İnterface

Vision–language models (VLMs) extend multimodal learning by aligning images with natural language, enabling a model to interpret visual findings in clinical context and, in some settings, generate explanatory text. Their appeal is strongest in specialties in which images and narrative interpretation are inseparable. Transformer-based integration of chest radiographs and clinical parameters has shown that contextual non-imaging information can improve diagnostic performance over single-modality approaches [65]. In dermatology, SkinGPT-4 combined images with clinical concepts and physician notes and was evaluated on real clinical cases [59]. In pathology, in-context learning with a multimodal model has shown that image classification can be adapted to new cancer tasks with few examples and without conventional task-specific retraining [66].
However, multimodal fluency should not be equated with clinical reliability. In radiology report generation, clinician evaluation of Flamingo-CXR showed substantial potential for image-text assistance but also clinically significant errors in both AI-generated and human reports [60]. More recent benchmarking in emergency and critical care similarly indicates that VLM performance remains sensitive to task and model choice [67]. The most defensible near-term role is therefore not autonomous image interpretation, but context-aware assistance in which generated conclusions remain auditable against source images and clinical data.

5.3. Does Multimodality Provide İncremental Clinical Value?

The decisive question is whether multimodal integration adds clinically meaningful information. Several well-designed primary studies support this premise, but the magnitude of benefit is heterogeneous. In pulmonary embolism detection, fusion of CT pulmonary angiography with EHR data outperformed imaging-only and EHR-only models [68]. Multimodal chest-radiograph models similarly benefited from combining imaging with clinical parameters across diagnostic tasks [65]. In oncology, complementary information from imaging, pathology, and molecular data has improved prognostic modelling and risk stratification [61,62]. MultiSurv further demonstrated the feasibility of integrating clinical, imaging, and multiple omics modalities for pan-cancer survival prediction [69].
Yet a statistically superior multimodal model is not automatically a clinically superior model. Small absolute gains may not justify additional tests, data acquisition, computational complexity, or delayed decision-making. Apparent benefit can also reflect information leakage, unequal preprocessing, or a weak unimodal comparator. Incremental value should therefore be demonstrated using identical cohorts and endpoints, strong modality-specific baselines, uncertainty estimates, calibration, and where possible external or prospective validation. The relevant endpoint is not multimodal superiority per se, but whether the added information is sufficiently large, reproducible, and actionable to alter care.
For diagnostic use, incremental value should be expressed in terms that reflect the care pathway: change in sensitivity or specificity at the selected threshold, reclassification, diagnostic yield, avoided or added tests, time to diagnosis, and net benefit. The reference standard should be independent of the modalities supplied to the model whenever possible to reduce incorporation bias, and temporal alignment should prevent future information from leaking into the diagnostic prediction. If a required modality is missing or degraded, the system should provide a validated fallback, an uncertainty warning, or abstain rather than issue an unqualified conclusion [61,62,63,64,65,66,67,68,69,70].

5.4. Fusion and the Missing-Modality Problem

How modalities are combined matters because clinical data differ in structure, timing, reliability, and availability. Early fusion may obscure modality-specific information, whereas late fusion can miss interactions between data sources. More flexible approaches can address some of these limitations, but no fusion strategy eliminates a common problem in clinical practice: multimodal data are often incomplete.
Models trained only on complete cases may perform poorly when one or more modalities are unavailable at deployment. MultiSurv was designed to handle missing values and missing modalities [69] while more recent attention-based approaches use modality-specific representations and contrastive learning to preserve performance when inputs are incomplete [70]. Missing data should therefore be treated as part of the clinical setting rather than simply as a preprocessing problem. Evaluation should reflect realistic patterns of modality availability and consider that missingness itself may carry clinically relevant information. Evaluation should further determine whether integration failures or discordant inputs are detected and contained rather than propagated into downstream clinical decisions.

5.5. Clinical İmplications

Multimodal AI is most useful when it combines genuinely complementary information and remains reliable when some data are missing. The goal should not be to add as many modalities as possible, but to include those that provide clear incremental value. Each added data source should improve clinically relevant performance and the model should remain robust when that source is unavailable or degraded. This principle is particularly relevant for precision oncology, neurological disease, acute imaging, and longitudinal risk prediction, where clinically meaningful information is distributed across heterogeneous sources [61,62,63,65,68,69,70,71].
The evidence to date supports multimodal AI as an important direction for clinical decision support, but not as a universal replacement for unimodal systems. A well-calibrated single-modality model may remain preferable when additional data are costly, inconsistently available, or only marginally informative. Clinical maturity will therefore depend less on demonstrating that modalities can be fused than on showing that integration provides reproducible incremental value under the incomplete, heterogeneous, and shifting conditions of routine care. Representative primary studies illustrating the clinical value and remaining limitations of multimodal AI are summarized in Table 3.
Multimodal AI is supported mainly by retrospective studies, with external validation emerging in selected applications. Evidence of prospective performance and incremental clinical benefit over strong unimodal approaches remains limited.

6. Agentic AI in Diagnostic Workflows: From Assistance to Goal-Directed Clinical Action

In this Review, a system is considered agentic when it pursues an explicit goal through stateful and iterative interaction with its environment, supported by the following operational capabilities: goal-directed planning, maintenance of context or memory, selection of tools or actions, evaluation of intermediate observations, and adaptation of subsequent actions. The degree of autonomy is considered separately according to the extent of human control and oversight. Accordingly, the use of multiple agents, external tools, or a multi-step LLM pipeline alone is not sufficient for a system to be classified as agentic.
Agentic AI extends the trajectory of clinical artificial intelligence from information generation toward systems capable of coordinating multi-step tasks within defined goals and constraints. The relevance of this development to healthcare is not merely that it enables more autonomous behavior of computer programs. Of much greater relevance is the capability of such programs to link their reasoning to external knowledge, to application-specific tools and to the step-wise execution of workflows. Clinical evaluation of such systems therefore needs to be expanded from being centered on a single output to covering the entire decision–action path that such systems are able to produce. This section examines the progression from bounded clinical agents to multi-agent architectures and human–AI collaboration, together with the associated governance requirements.

6.1. From Generative Assistance to Agency

Agentic AI extends generative systems from producing individual responses to pursuing goals through planning, tool use, and sequential action. In healthcare, an agent may retrieve additional evidence, use specialised models or clinical calculators, revise its approach, and coordinate multiple steps within a workflow. Evaluation must therefore consider not only the quality of a single response, but also the sequence of decisions and actions taken to reach an outcome. This creates new opportunities for clinical workflow support, while also increasing the ways in which errors can arise and propagate.
Early primary studies suggest that bounded agents can perform clinically meaningful tasks when their objectives and tools are tightly specified. Autonomous oncology decision support has been developed and validated around structured clinical reasoning [72], while LLM-based agents have been tested for evidence-based medicine [73], radiotherapy planning [74], order-set optimisation [75], and disease-specific diagnostic or treatment planning [76]. These applications are more informative than generic conversational benchmarks because they require the system to operate within a defined clinical process rather than merely produce plausible text.

6.2. Planning, Tool Use, and Multi-Agent Reasoning

The principal advantage of an agentic architecture is decomposition. Complex clinical tasks can be divided into information extraction, evidence retrieval, differential generation, verification, and recommendation, with each step assigned to a specialised tool or agent. Multi-agent designs extend this principle by allowing several role-specific components to critique or refine one another. CARE-AD, for example, used a multi-agent LLM framework to analyse longitudinal clinical notes for Alzheimer’s disease prediction [77]. Similar architectures have been reported for neuro-ophthalmic diagnosis and personalised treatment planning [76] and for optimisation of clinical order sets [75].
However, additional agents do not inherently produce additional clinical value. Decomposition can improve traceability and error checking, but it can also introduce correlated reasoning, duplicated errors, communication failures, and greater computational cost. Comparisons of rule-based, single-agent, and multi-agent systems are therefore particularly important because they test whether orchestration itself contributes value rather than simply increasing system complexity [78]. The appropriate comparator for an agentic system is not an unassisted language model alone, but the strongest simpler workflow capable of performing the same task.
In diagnostic workflows, an agent may assemble longitudinal evidence, generate and reprioritize differential diagnoses, select or call diagnostic tools, and route referrals. Because an early patient-matching, retrieval, or anchoring error can propagate through every subsequent step, validation should evaluate each transition: data retrieval, patient identification, problem representation, differential generation, test selection, interpretation, and escalation. Explicit stop rules are required before invasive, costly, or otherwise high-consequence actions [72,73,74,75,76,77,78,79,80].

6.3. From Decision Support to Workflow Execution

Agentic systems become clinically distinctive when they move beyond recommendation towards workflow execution. This includes selecting and calling external tools, retrieving patient-specific information, generating intermediate artefacts, and determining the next action from previous results. A feasibility study of LLM agents for radiotherapy planning illustrates this transition: the model was embedded within a sequence of planning operations rather than evaluated only on textual knowledge [74]. Likewise, autonomous analysis of curated oncology data and agent-based clinical decision systems indicate a movement towards systems that coordinate multiple analytic steps before producing an output [72,79].
This transition should nevertheless remain bounded by the reversibility and consequence of the action. Drafting an order set for clinician approval and autonomously placing an order are not equivalent levels of agency. Clinical deployment therefore requires explicit action permissions, tool whitelisting, audit logs, provenance of retrieved evidence, and predefined escalation points. Human oversight is most meaningful when it is positioned before high-consequence or irreversible actions rather than appended as a nominal final review.

6.4. Safety and Evaluation of Agentic Systems

Agentic AI changes the object of safety evaluation. Simply assessing accuracy at the final output can be insufficient, since a seemingly correct output can hide an unsafe sequence of actions, and an early error in patient identity or retrieval can propagate to affect all subsequent actions. Instead, evaluation should assess the full completion of a task, the intermediate reasoning and all tool calls made during task completion, error recovery, adherence to specified constraints, patient identity, robustness to incorrect information, and the system’s ability to stop and/or escalate to a human as needed. Note that recent work has demonstrated silent failure of clinical natural language processing agents on patient identity, highlighting the insufficiency of answer-level evaluation for systems that have access to patients’ longitudinal records and/or to execute specific tools to support clinical tasks [80].
Safety must also be assessed at the system level. Multi-agent architectures introduce internal communication channels and new privacy and security boundaries; autonomous agents can amplify small upstream errors through repeated actions. For clinical use, the central question is therefore not whether an agent can complete a workflow under ideal conditions, but whether it fails detectably and recoverably when information is missing, contradictory, or incorrect. Prospective evaluation should report not only clinical performance but intervention frequency, override behaviour, failure severity, and the consequences of erroneous tool use.

6.5. Clinical Readiness

Agentic AI should be viewed as a relatively new and emerging layer of clinical automation rather than a fully developed autonomous clinician. The strongest near-term use cases for Agentic AI are bounded, auditable workflows—such as goal, tool, data source, action sets defined by humans—in which complex information processing is simplified and otherwise discrete AI applications are integrated by the agent. There is much less evidence to date to support the autonomous management of clinical objectives and their implementation by the system alone.
Clinical readiness therefore depends on controllability. Actions should be clearly bounded, the evidence supporting each action should be traceable, tool failures should be detectable, and predefined pathways for escalation to a human clinician should be available. As AI systems acquire greater capacity to perform actions within clinical workflows, their evaluation must extend beyond response quality to include the safety and reliability of the end-to-end clinical process. As the copilots of today’s AI systems transition to full agents that take actions, as opposed to simply generating information, validation of such systems will transition from assessing the quality of their responses to assessment of the safety of the end-to-end clinical workflow. These are issues of governance that go beyond assessing the performance of individual models, and are therefore discussed in the following section. Representative clinical studies evaluating different forms and applications of agentic AI are summarized in Table 4.
Agentic AI remains at an early translational stage, with evidence largely derived from proof-of-concept and task-level evaluations. Independent validation, prospective workflow studies, and evidence of clinical impact are currently limited.

7. Diagnostic and Clinical Applications Across Healthcare

Artificial intelligence is being introduced into the clinical setting in different ways depending on the medical discipline. For image-based disciplines such as dermatology, the translation of computer vision into clinical applications for diagnostic assistance is slowly being developed and evaluated. In contrast, data-rich disciplines are increasingly integrating heterogeneous sources, including patient records, physiological signals, laboratory findings, molecular profiles, and patient-generated health data, and are being used for screening, diagnostic, prognostic, therapeutic, procedural, alerting and management purposes of chronic diseases over long periods of time.
Measures of model performance alone do not establish the clinical importance of AI applications. Even highly accurate models may provide limited clinical benefit if they reproduce information already recognized by clinicians, deliver outputs too late to influence decision-making, or generate additional workload that outweighs their potential utility. Table S1 represents the various stages and forms of clinical translation that have been achieved in various clinical areas rather than representing the highest performing models for each clinical area.
Across specialties, AI occupies four diagnostically distinct roles: screening or triage, lesion or disease detection, differential diagnosis or etiologic classification, and prognostic or treatment-response stratification. These roles have different prevalence, reference-standard, operating-threshold, and downstream-action requirements and should not be compared using a single accuracy metric. Table S1 therefore maps each study to its AI task, clinical application, and stage of evidence rather than ranking algorithms by performance.
The clearest diagnostic translation pathways are seen when the AI output is connected to a defined next action: recall or referral in mammography and retinal screening [81,82,83], real-time lesion detection in colonoscopy [84,85], pathologic or intraoperative classification [41,86], etiologic dementia differentiation [87], and work-up of prostate cancer, aneurysm, fracture, melanoma, or pediatric respiratory disease [88,89,90,91,92,93]. These examples show why evaluation should progress from accuracy to prospective diagnostic yield, additional testing, time to diagnosis, decision change, and downstream outcomes.
Studies on clinical AI applications that are close to clinical practice show that clinical AI is not following a uniform translational trajectory. In some applications, the distance between algorithmic output and an actionable clinical decision is relatively short.
However, the vast majority of clinical AI applications require a substantially longer translational pathway to deliver clinical utility. For example, many prognostic models developed for use in oncology, nephrology, psychiatric and neurological care of chronic conditions may perform well at risk stratification but lack evidence of how the AI generated information results in changes to treatment or patient outcomes. Clinically mature AI applications are those that deliver information that is incremental, actionable, on time and most importantly of benefit within the specific clinical care pathway. The major data modalities underpinning AI applications across clinical specialties are summarized in Table 5. The table provides a conceptual mapping of commonly used data sources rather than a quantitative ranking of modality prevalence. Modalities increasingly coexist within multimodal systems, and their clinical relevance depends on the task, care setting, and availability of complementary information.

8. Trustworthy Clinical AI: Safety, Equity, Transparency, and Governance

The clinical value of AI-based decision support extends beyond predictive accuracy. Such a system must also support the clinician in decision making about diagnosis and/or treatment, in setting priorities and in monitoring patients. Safe implementation requires a sociotechnical system that remains reliable across relevant patient populations and clinical settings, provides appropriate transparency and traceability, and supports clear accountability.
For diagnostic AI, trustworthiness requires explicit management of false-negative and false-positive harms, indeterminate outputs, subgroup-specific performance, abstention and escalation criteria, and responsibility for the final interpretation. Prevalence drift requires particular attention because it changes positive and negative predictive values even when sensitivity and specificity appear stable [23,94,95,96,97,98,99,100,101].

8.1. Trustworthiness İs a System Property

Traditionally, AI has been developed in a retrospective fashion and then applied to clinical settings. As the technology begins to be used in more clinical scenarios, however, performance is not the only consideration for defining quality. A clinically trustworthy system must remain safe and useful for different populations, in different settings, over time. Importantly, all outputs must be traceable and all individuals involved in decision-making for critical clinical decision points must be clearly identified. The FUTURE-AI consensus framework outlines a set of characteristics that define trustworthy healthcare AI (fairness, universality, traceability, usability, robustness, and explainability) that must be embedded across the system lifecycle [23]. Importantly, trustworthiness is not an after-thought additional metric; it is a property of the data, model, interface, workflow, institution, and monitoring system.
While aggregate accuracy may hide failures in individual cases, a number of other types of failures are possible, such as systematic failure in certain subgroups of patients, using inappropriate stand-ins for important clinical variables, encouraging over-reliance, and failing after deployment. Governance must consequently focus on how performance is distributed and maintained, not simply on whether a model crossed a prespecified benchmark before release.

8.2. Bias, Fairness, and Health Equity

Clinical AI can reproduce inequities already embedded in healthcare data and can create new disparities through differential model performance. A widely used population-health algorithm was shown to underestimate the health needs of Black patients because healthcare cost was used as a proxy for illness; correcting the proxy substantially increased the proportion of Black patients identified for additional care [102]. In medical imaging, chest-radiograph classifiers have demonstrated systematically higher underdiagnosis rates in underserved groups, including Black and Hispanic patients and patients with lower socioeconomic status [95].
Fairness cannot therefore be reduced to demographic balance in the development dataset. Models can encode sensitive attributes even when those variables are not explicitly supplied. Deep learning systems can infer self-reported race from multiple medical imaging modalities with high accuracy, including under transformations that obscure obvious anatomical or acquisition-related explanations [96]. Removing protected variables is consequently not a sufficient mitigation strategy. Evaluation should examine clinically relevant error rates, calibration, and downstream consequences across prespecified and intersectional subgroups, while recognising that statistical fairness criteria may conflict and must be selected according to the clinical use case.

8.3. Explainability, Transparency, and Appropriate Reliance

Explainability is often seen as the key to winning users’ trust for AI systems. However, explaining a model’s behavior is not enough; the explanation must also be accurate. Most methods for explainable AI, including saliency maps, feature-attribution methods, and natural-language rationales, aim to make a system more interpretable than it actually is. The clinical objective of calibrated reliance therefore is to support users by describing the system’s intended use, limitations, degree of uncertainty, and evidence in order to enable them to decide whether to trust the system, to verify it, or even to reject it.
Human–AI interaction studies reveal the complexities of how human clinicians might be influenced by AI in clinical decision-making environments. While clinicians benefit from assistance provided by AI systems, incorrect information can mislead them. Moreover, clinicians can modify their clinical decisions in negative ways when they perceive an AI system to be authoritative [97]. Appropriate reliance is bidirectional. Clinical performance may be compromised not only by over-reliance on incorrect AI recommendations but also by algorithm aversion or under-reliance, particularly when clinicians perceive algorithmic advice as less trustworthy or as constraining professional judgment. Effective human–AI collaboration therefore requires calibrated reliance rather than either uncritical acceptance or systematic rejection of AI-supported recommendations [97,103]. Conversely, collaborative evaluation in dermatology has shown that human and AI performance depends on the interaction design and the clinician’s ability to integrate algorithmic information [98]. Transparency should therefore include provenance, versioning, data lineage, intended population, operating thresholds, and known failure modes rather than relying on post hoc visual explanations alone.

8.4. Privacy, Security, and Robustness

Clinical AI creates security risks at both the data and model levels. Training data may contain identifiable or inferable patient information, while deployed models can be exposed to adversarial inputs, data poisoning, prompt manipulation, or compromised retrieval sources. Medical machine-learning systems have been shown to be susceptible to adversarial perturbations that can alter clinically relevant predictions while remaining difficult for humans to detect [99]. As more systems are built to support patient care by using electronic patient records as well as external knowledge and tools, the distinction between a model’s security and clinical safety will decrease.
The robustness of a model needs to be tested against specific failures that may occur in clinical scenarios. In addition to testing against random numbers, a model must be able to handle acquisition variability, missing input variables, distribution shifts, corrupted input, malicious use, and also failures of other services that the model interacts with. The security measures should be aligned with the least privilege principle in order to limit access to data and functions. Authenticated actions must be monitored and logged. Measures to recover from failures must also be implemented. Privacy and cybersecurity are components of clinical risk management and thus are not purely IT issues.

8.5. Human Oversight, Accountability, and Lifecycle Governance

Human oversight is effective only when responsibilities and intervention points are explicitly defined. A nominal human-in-the-loop configuration is insufficient when clinicians cannot reliably detect errors, lack sufficient time to interrogate outputs, or face excessive volumes of automatically generated recommendations. Formal regulatory oversight complements these governance principles. FDA frameworks for AI-enabled medical devices increasingly emphasize lifecycle management, transparency, post-deployment monitoring, and controlled modification, thereby shaping both system accountability and human oversight in clinical use [104]. The intensity of human oversight should be proportionate to clinical risk administrative tasks might only need to be audited after the fact, whereas diagnostic, therapeutic or even executable decisions need stronger prospective control points and a clear escalation procedure.
In addition to being governed before the AI system is deployed, the AI system must also be governed after it has been deployed. The behavior of a model can change without it being modified, due to changes in the patient population and clinical practice, as well as in the devices, coding systems, prevalence of diseases and in the data from upstream systems. Thus, post-deployment monitoring of the AI system tracks discrimination, calibration, performance in subgroups, missing values, override rates for human-in-the-loop for different clinical scenarios and outcomes of clinical relevance. The monitoring has predefined rules for investigation, for recalibration, for suspension of the model or for retraining of the model. The monitoring is part of the lifecycle of trustworthy AI, as described in the FUTURE-AI framework, and not just a step towards regulatory approval [23].
As AI systems become multimodal, generative, and increasingly agentic, this lifecycle approach becomes increasingly important. Governance is therefore the mechanism by which technical capability is converted into accountable clinical practice. Although this Review primarily focuses on clinician-facing AI, patient-facing transparency, informed consent, shared decision-making, and community accountability remain important considerations for responsible clinical implementation. The core dimensions of trustworthy clinical AI and their corresponding governance requirements are summarized in Table 6.

9. Persistent Challenges to Clinical Translation

The barriers that continue to limit the translation of artificial intelligence into routine healthcare are increasingly system-level rather than purely algorithmic. As clinical AI expands from task-specific prediction toward generative, multimodal, and agentic systems, successful deployment depends on the integrity of the underlying data, the reproducibility of the development process, compatibility with clinical information infrastructure, and the resources required to operate increasingly complex models. These constraints are particularly important because improvements in model capability do not inherently resolve weaknesses in the ecosystem in which the model is developed and deployed.

9.1. Data Quality, Reproducibility, and Shortcut Learning

Clinical datasets are rarely created specifically for machine learning. Electronic health records, medical images, physiological signals, and other routinely collected data reflect clinical workflows, acquisition devices, institutional practices, coding conventions, and patterns of missingness. Consequently, increasing dataset size does not necessarily increase dataset quality. Large datasets can preserve systematic artefacts and confounding structures that models may exploit more efficiently as their capacity increases.
Diagnostic datasets introduce additional sources of bias, including spectrum bias, partial or differential verification, incorporation bias when the reference standard includes model inputs, label leakage from downstream care, and imperfect or time-dependent reference standards. These biases can make a system appear diagnostically accurate without showing that it recognizes disease in the intended population. Cohort assembly, reference-standard adjudication, patient-level partitioning, and handling of indeterminate cases must therefore be reported and tested explicitly [7,8,9,105,106,107].
A particularly important manifestation is shortcut learning, in which a model achieves apparently high performance by exploiting features correlated with the target rather than clinically meaningful characteristics of the underlying disease. This phenomenon has been demonstrated across multiple forms of medical data. Ly et al. evaluated 13 clinical datasets encompassing 207,487 patients and five modalities, including radiographs, CT, ECG, clinical text, and auscultation data. Conventional evaluation overestimated external performance by approximately 20% on average because models learned hidden data-acquisition biases [105]. In radiography, models developed to detect pneumothorax have similarly exploited the presence of chest drains a clinically inappropriate shortcut because the device may indicate that the condition has already been diagnosed and treated [106].
Shortcut behaviour is not limited to obvious imaging artefacts. Acquisition pathways, institutional identifiers, demographic proxies, documentation patterns, preprocessing artefacts, and other latent characteristics can become predictive signals. Importantly, shortcut learning can coexist with excellent conventional test-set performance. Brown et al. demonstrated that shortcut testing can identify such behaviour in clinical machine-learning applications, reinforcing the need to evaluate what information a model is using rather than considering predictive accuracy alone [107].
Reproducibility is closely linked to this problem. Model architecture and performance metrics are insufficient to reconstruct an AI pipeline if cohort selection, patient-level partitioning, preprocessing, annotation, missing-data handling, feature construction, or hyperparameter selection are incompletely reported. Data leakage can further inflate performance when information from the evaluation set enters model development directly or indirectly. These concerns become harder to audit for foundation and proprietary models because the composition and provenance of pretraining datasets may be only partially disclosed. Reproducibility should therefore be considered a property of the complete analytical pipeline rather than of the trained model alone.

9.2. Interoperability and İnfrastructure

Clinical AI must also operate within a fragmented digital environment. Patient information is distributed across electronic health records, laboratory information systems, imaging archives, physiological monitoring platforms, genomic databases, and increasingly patient-generated data. Differences in terminology, in coding schemes, in user interfaces, in concepts for time and in architectures for health care data pose a significant barrier to the transfer of successful software solutions into clinical work flows.
To be clinically relevant for multimodal and agentic systems, such systems typically consist of several data streams, retrieval mechanisms, external knowledge resources, computational tools and model components. Interoperability for such systems goes far beyond simple data exchange and requires consistent semantics for example for patient and encounter identification, data provenance, version management, authentication and stable communication.
Interoperability is therefore a prerequisite for the effective integration of AI into clinical workflows. Isolated optimization of data pipelines or individual system components is insufficient to ensure safe and effective AI deployment in real-world clinical environments. Systems validated in small-scale studies must also demonstrate reliable integration with existing clinical data infrastructures and operational systems before routine deployment.

9.3. Computational and Environmental Sustainability

Increasing model complexity also introduces substantial computational and operational resource requirements. Large pre-training, multi-modal processing, repeated fine-tuning, large numbers of inferences, and agent-centric workflows all require significant amounts of hardware, memory, cloud-based computing resources, energy and maintenance resources to support development and deployment. The economic and environmental costs associated with these requirements may also affect clinical scalability and translation.
Environmental impact must not be confused with model size. A recent study on the use of autonomous AI for diabetic-eye disease screening found that the emissions attributable to AI-based image processing were substantially lower than those associated with conventional in-person screening pathways that required patient travel to hospital facilities. Under the assumptions of that study, the AI-enabled care pathway was associated with substantially lower estimated emissions when the broader screening pathway, including patient travel, was considered [108].
However, as AI deployment expands across health systems, the sum of many small computational costs can end up having large consequences for environmental sustainability. Therefore, whilst minimizing computation is an important goal, it is equally important to also maximize clinical value per computational resource. Model compression, efficient architectures, invoking large models only when required, deployment to local edge devices where appropriate, and the use of hybrid systems where large models are reserved for specific clinical use cases are key to enabling sustainable clinical deployment of AI within health systems.
In summary, scaling up clinical AI in the near term will require more than just increasing the size of the models. Data, reproducibility, interoperability, and computational resources must all be engineered to support the growth of AI in clinical care. Translation of AI to the clinic will increasingly depend on engineering an AI ecosystem rather than improving individual algorithms.

10. Emerging Frontiers

Future healthcare AI systems are unlikely to be dominated by a single model class. Instead, they will learn over time, operate over distributed data, maintain a patient specific computational representation, interact with the physical world and aid scientific discovery. The maturity of these emerging healthcare capabilities is heterogeneous with some, like federated learning, already being demonstrated in multicentre clinical studies while others like autonomous discovery and highly personalized digital twins are very much emerging health care AI capabilities and should be viewed with that status rather than as simply another clinical system awaiting deployment.

10.1. Adaptive and Distributed AI

Most clinical models are trained on a single fixed data set and then deployed as static systems. Continual learning enables models to adapt to newly available information while preserving performance on previously learned tasks. In medical imaging, continual learning approaches have been developed to exploit dynamic model memory across sequential tasks. Frameworks for standardized continual learning have also started to be developed for segmentation models that learn to adapt to changing data distributions over time [109,110]. The clinical rationale for continual learning is that disease prevalence, imaging distributions, hardware platforms, and clinical workflows evolve over time; models that can adapt safely to such changes may therefore offer advantages over systems requiring complete retraining. However, there is still the problem of control: a model’s adaptations should not unintentionally remove previously learned behavior from memory or introduce unknown behavior; therefore, update policies, model versions, and re-evaluation must all be explicitly controlled as opposed to simply allowing online learning.
Distributed AI addresses a different constraint: clinically useful data are often dispersed across institutions that cannot pool patient-level information. Federated learning enables collaborative model training while data remain locally held. Sheller et al. demonstrated multi-institutional medical modelling without centralizing patient data, and Dayan et al. subsequently trained a federated model across international institutions to predict clinical outcomes in patients with COVID-19 [111,112]. Edge-oriented implementations extend this logic by moving inference or learning closer to the data source, potentially reducing latency and dependence on centralized infrastructure [113]. These approaches do not eliminate privacy, heterogeneity, or governance problems, but they shift the architecture of clinical AI from centralized data aggregation toward coordinated computation across sites and devices.

10.2. Digital Twins and Personalized AI

Digital twins seek to construct computational representations that evolve with an individual patient and can be used to simulate, forecast, or compare possible future states. This is more demanding than conventional risk prediction because the representation is intended to remain dynamically linked to patient-specific information. Current healthcare studies illustrate several components of this vision rather than a fully realized general-purpose patient twin. ClinicalGAN, for example, used patient digital twins to support monitoring in clinical trials, while multiscale modelling of immune surveillance in micrometastases demonstrated how mechanistic simulation can contribute to cancer patient digital twins [114,115]. Other work has explored patient-specific treatment computation through in silico trials and graph-based forecasting of longitudinal medical conditions.
The principal opportunity is personalization at the level of trajectories and interventions rather than only static risk categories. A mature digital twin could potentially combine longitudinal observations, mechanistic knowledge, and learned representations to test alternative management strategies before they are applied to the patient. However, the evidence base remains fragmented across disease-specific models, simulations, and proof-of-concept systems. Personalized AI should therefore not be equated with the existence of a comprehensive virtual patient. Its clinical credibility will depend on whether the twin remains synchronized with changing patient states and whether simulated counterfactuals are sufficiently accurate to inform real decisions.

10.3. Embodied AI and Robotics

Embodied AI extends computational intelligence into systems that sense and act in the physical clinical environment. Surgical and medical robotics provide the clearest current pathway toward this form of AI. Research has progressed from teleoperation and constrained automation toward autonomous subtasks such as suturing, tissue manipulation, navigation, and instrument handling. Pedram et al. developed and quantified an autonomous suturing framework using a cable-driven surgical robot, while Wang et al. demonstrated a task-autonomous medical robot capable of incision stapling and staple removal [116,117]. Importantly, early human translation has also occurred: a first-in-men study evaluated a robotic tool providing autonomous inner-ear access for cochlear implantation [118].
Error potential of embodied AI systems differs from that of information systems as errors could cause immediate physical harm. The performance of embodied AI systems depends on perception, control, human anatomy, robotic embodiment, instrumentation, and robustness to unexpected events. In the near term, clinical deployment is likely to favour bounded autonomy, i.e., autonomous subtasks that have been validated for specific applications and that are used within larger clinician-supervised procedures. Autonomous tasks have to function fail-safe, have to be recoverable and must also provide evidence of their potential to increase precision and to support workflows without causing new procedural hazards.

10.4. Autonomous Scientific Discovery

A more nascent frontier is the use of AI to participate directly in the scientific process. Self-driving laboratories combine machine learning with automated experimentation so that hypotheses or candidate solutions can be generated, tested, and iteratively refined with reduced manual intervention. Rapp et al. demonstrated an autonomous laboratory that navigated protein fitness landscapes by integrating machine-learning-guided selection with experimental measurement [119]. This shifts AI from analysing completed experiments toward selecting which experiment should occur next.
The frontier is now beginning to intersect more directly with biomedical discovery. A 2026 study reported an agentic framework for autonomous scientific discovery in cancer pathology, indicating that coordinated AI systems can move beyond isolated prediction toward multi-step scientific analysis [120]. To connect literature, computational studies, hypothesis generation, experimental design and the control of robots in closed or semi-closed loops by means of AI to generate reproducible biological knowledge, novel drugs or testable hypotheses of clinical relevance in medicine using research approaches more efficiently than currently practiced requires a body of evidence currently small in size and automation level.
Most research in AI today is moving across several fronts, and in our view the biggest transition happening currently is from fixed, centralized and for the most part task-specific AI systems, to more adaptive, distributed, patient-specific, physically interactive and even scientifically generative systems. While there is variable evidence of the maturity of these new systems, notable progress has been made recently in the areas of federated learning and bounded robotic autonomy, but evidence of the adequacy of current evidence for the bulk of more complex digital twins and fully autonomous scientific discovery remains largely at the level of proof-of-concept.

11. Conclusions

Artificial intelligence is expanding across the diagnostic continuum, from screening and triage to disease detection, differential diagnosis, prognostic stratification, and treatment-response assessment. Predictive models remain central to estimating classes and risks; generative models can synthesize histories and support differential reasoning; multimodal systems can integrate images, signals, text, pathology, and molecular data; and agentic architectures can coordinate evidence retrieval, tools, and sequential workflow steps. These paradigms are increasingly convergent, but technological convergence does not itself establish diagnostic value.
The diagnostic question must remain the organizing principle. Every system should specify its intended population, target condition and alternatives, role in the pathway, reference standard, operating threshold, and consequences of false-positive, false-negative, and indeterminate results. High retrospective accuracy, fluent generation, multimodal fusion, or autonomous task completion is insufficient if performance is poorly calibrated, fails under spectrum or prevalence shift, omits critical alternatives, or does not change an appropriate clinical decision. A credible implementation pathway therefore progresses from internal performance to independent external and prospective validation, human–AI and workflow evaluation, comparative clinical-impact studies, and post-deployment surveillance. Outcomes should include diagnostic yield, time to diagnosis, additional testing, downstream harm, equity, resource use, and patient-relevant benefit.
Trustworthiness is consequently a property of the complete diagnostic system rather than a checklist of model attributes. Fairness, transparency, privacy, security, accountability, and robustness must be engineered across the data, model, interface, users, institution, and monitoring process. In generative and agentic systems, provenance, uncertainty, bounded permissions, audit trails, abstention, error recovery, and escalation to a clinician are essential because errors can propagate across multiple steps.
Emerging approaches, including continual and federated learning, patient digital twins, embodied intelligence, and autonomous scientific discovery, may extend this ecosystem, but their diagnostic relevance must be demonstrated rather than inferred from novelty. The next generation of healthcare AI should be judged by whether rigorously evaluated human–AI systems improve disease detection, differential diagnosis, prognostic stratification, and care decisions safely, equitably, and reproducibly in real clinical settings. Clinical translation is achieved not when a model produces an impressive output, but when its use produces measurable and durable improvements in the diagnostic pathway, clinically relevant care processes, or patient-relevant outcomes.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/bioengineering13101175/s1, Table S1. Representative diagnostic and clinically translational applications of artificial intelligence across healthcare [121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140].

Author Contributions

Conceptualization, İ.B., B.T., G.T., S.D. and T.T.; methodology, İ.B., B.T., G.T., S.D. and T.T.; investigation, İ.B., B.T., G.T., S.D. and T.T.; writing—original draft preparation, İ.B., B.T. and G.T.; writing—review and editing, B.T., G.T., S.D. and T.T.; visualization, İ.B. and B.T.; supervision, B.T., S.D. and T.T. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable. This article is a review of previously published literature and did not involve new studies of humans or animals.

Informed Consent Statement

Not applicable. This review did not recruit participants or use identifiable individual—patient data.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare no conflict of interest.

Abbreviations

AIArtificial intelligence
MLMachine learning
DLDeep learning
LLMLarge language model
EHRElectronic health record
NLPNatural language processing
VLMVision–language model
RAGRetrieval-augmented generation
XAIExplainable artificial intelligence
CADComputer-aided diagnosis
CTComputed tomography
MRIMagnetic resonance imaging
ECGElectrocardiography
EEGElectroencephalography
IoTInternet of Things
FLFederated learning

References

  1. Yu, K.-H.; Beam, A.L.; Kohane, I.S. Artificial intelligence in healthcare. Nat. Biomed. Eng. 2018, 2, 719–731. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Topol, E.J. High-performance medicine: The convergence of human and artificial intelligence. Nat. Med. 2019, 25, 44–56. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Rajkomar, A.; Dean, J.; Kohane, I. Machine learning in medicine. N. Engl. J. Med. 2019, 380, 1347–1358. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Rajpurkar, P.; Chen, E.; Banerjee, O.; Topol, E.J. AI in health and medicine. Nat. Med. 2022, 28, 31–38. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Liu, X.; Faes, L.; Kale, A.U.; Wagner, S.K.; Fu, D.J.; Bruynseels, A.; Mahendiran, T.; Moraes, G.; Shamdas, M.; Kern, C.; et al. A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: A systematic review and meta-analysis. Lancet Digit. Health 2019, 1, e271–e297. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Evangelia, C.; Jie, M.; Collins, G.; Steyerberg, E.; Verbakel, J.; Van Calster, B. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J. Clin. Epidemiol. 2019, 110, 12–22. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Aggarwal, R.; Sounderajah, V.; Martin, G.; Ting, D.S.; Karthikesalingam, A.; King, D.; Ashrafian, H.; Darzi, A. Diagnostic accuracy of deep learning in medical imaging: A systematic review and meta-analysis. npj Digit. Med. 2021, 4, 65. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Cabitza, F.; Campagner, A.; Soares, F.; de Guadiana-Romualdo, L.G.; Challa, F.; Sulejmani, A.; Seghezzi, M.; Carobene, A. The importance of being external. methodological insights for the external validation of machine learning models in medicine. Comput. Methods Programs Biomed. 2021, 208, 106288. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Collins, G.S.; Moons, K.G.; Dhiman, P.; Riley, R.D.; Beam, A.L.; Van Calster, B.; Ghassemi, M.; Liu, X.; Reitsma, J.B.; Van Smeden, M. TRIPOD+ AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024, 385, e078378. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Vasey, B.; Nagendran, M.; Campbell, B.; Clifton, D.A.; Collins, G.S.; Denaxas, S.; Denniston, A.K.; Faes, L.; Geerts, B.; Ibrahim, M. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 2022, 28, 924–933. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Thirunavukarasu, A.J.; Ting, D.S.J.; Elangovan, K.; Gutierrez, L.; Tan, T.F.; Ting, D.S.W. Large language models in medicine. Nat. Med. 2023, 29, 1930–1940. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Shah, N.H.; Entwistle, D.; Pfeffer, M.A. Creation and adoption of large language models in medicine. JAMA 2023, 330, 866–869. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Clusmann, J.; Kolbinger, F.R.; Muti, H.S.; Carrero, Z.I.; Eckardt, J.-N.; Laleh, N.G.; Löffler, C.M.L.; Schwarzkopf, S.-C.; Unger, M.; Veldhuizen, G.P. The future landscape of large language models in medicine. Commun. Med. 2023, 3, 141. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Haug, C.J.; Drazen, J.M. Artificial intelligence and machine learning in clinical medicine, 2023. N. Engl. J. Med. 2023, 388, 1201–1208. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Wornow, M.; Xu, Y.; Thapa, R.; Patel, B.; Steinberg, E.; Fleming, S.; Pfeffer, M.A.; Fries, J.; Shah, N.H. The shaky foundations of large language models and foundation models for electronic health records. npj Digit. Med. 2023, 6, 135. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Acosta, J.N.; Falcone, G.J.; Rajpurkar, P.; Topol, E.J. Multimodal biomedical AI. Nat. Med. 2022, 28, 1773–1784. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Liu, F.; Niu, Y.; Zhang, Q.; Wang, K.; Dong, Z.; Wong, I.N.; Cheng, L.; Li, T.; Duan, L.; Li, K. A foundational architecture for AI agents in healthcare. Cell Rep. Med. 2025, 6, 102374. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Xu, G.; Li, X.; Chen, Y.; Duan, Y.; Wu, S.; Yu, H.; Chiu, C.-H.; Ni, J.; Tang, N.; Li, T.J.-J. A comprehensive survey of ai agents in healthcare. J. Biomed. Inform. 2026, 179, 105045. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Xu, J.; Ko, J.M.; Kvedar, J.C. AI agents in clinical practice: An evidence map. npj Digit. Med. 2026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  20. Wijnberge, M.; Geerts, B.F.; Hol, L.; Lemmers, N.; Mulder, M.P.; Berge, P.; Schenk, J.; Terwindt, L.E.; Hollmann, M.W.; Vlaar, A.P.; et al. Effect of a machine learning–derived early warning system for intraoperative hypotension vs standard care on depth and duration of intraoperative hypotension during elective noncardiac surgery: The HYPE randomized clinical trial. JAMA 2020, 323, 1052–1060. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Liu, X.; Rivera, S.C.; Moher, D.; Calvert, M.J.; Denniston, A.K.; Ashrafian, H.; Beam, A.L.; Chan, A.-W.; Collins, G.S.; Deeks, A.D.J. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. Lancet Digit. Health 2020, 2, e537–e548. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Rivera, S.C.; Liu, X.; Chan, A.-W.; Denniston, A.K.; Calvert, M.J.; Ashrafian, H.; Beam, A.L.; Collins, G.S.; Darzi, A.; Deeks, J.J. Guidelines for clinical trial protocols for interventions involving artificial intelligence: The SPIRIT-AI extension. Lancet Digit. Health 2020, 2, e549–e560. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Lekadir, K.; Frangi, A.F.; Porras, A.R.; Glocker, B.; Cintas, C.; Langlotz, C.P.; Weicken, E.; Asselbergs, F.W.; Prior, F.; Collins, G.S. FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ 2025, 388, e081554. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Faust, O.; Hagiwara, Y.; Hong, T.J.; Lih, O.S.; Acharya, U.R. Deep learning for healthcare applications based on physiological signals: A review. Comput. Methods Programs Biomed. 2018, 161, 1–13. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Chakraborty, C.; Bhattacharya, M.; Pal, S.; Lee, S.-S. From machine learning to deep learning: Advances of the recent data-driven paradigm shift in medicine and healthcare. Curr. Res. Biotechnol. 2024, 7, 100164. [Google Scholar] [CrossRef] [Scilit]
  26. Zhang, A.; Xing, L.; Zou, J.; Wu, J.C. Shifting machine learning for healthcare from development to deployment and from models to data. Nat. Biomed. Eng. 2022, 6, 1330–1345. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Nerella, S.; Bandyopadhyay, S.; Zhang, J.; Contreras, M.; Siegel, S.; Bumin, A.; Silva, B.; Sena, J.; Shickel, B.; Bihorac, A. Transformers and large language models in healthcare: A review. Artif. Intell. Med. 2024, 154, 102900. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Harrer, S. Attention is not all you need: The complicated case of ethically using large language models in healthcare and medicine. EBioMedicine 2023, 90, 104512. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Goldstein, B.A.; Navar, A.M.; Carter, R.E. Moving beyond regression techniques in cardiovascular risk prediction: Applying machine learning to address analytic challenges. Eur. Heart J. 2017, 38, 1805–1814. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Kourou, K.; Exarchos, T.P.; Exarchos, K.P.; Karamouzis, M.V.; Fotiadis, D.I. Machine learning applications in cancer prognosis and prediction. Comput. Struct. Biotechnol. J. 2014, 13, 8–17. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Tran, K.A.; Kondrashova, O.; Bradley, A.; Williams, E.D.; Pearson, J.V.; Waddell, N. Deep learning in cancer diagnosis, prognosis and treatment selection. Genome Med. 2021, 13, 152. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  32. Swanson, K.; Wu, E.; Zhang, A.; Alizadeh, A.A.; Zou, J. From patterns to patients: Advances in clinical machine learning for cancer diagnosis, prognosis, and treatment. Cell 2023, 186, 1772–1791. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  33. Senders, J.T.; Staples, P.C.; Karhade, A.V.; Zaki, M.M.; Gormley, W.B.; Broekman, M.L.; Smith, T.R.; Arnaout, O. Machine learning and neurosurgical outcome prediction: A systematic review. World Neurosurg. 2018, 109, 476–486.e471. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Nielsen, K.B.; Lautrup, M.L.; Andersen, J.K.; Savarimuthu, T.R.; Grauslund, J. Deep learning–based algorithms in screening of diabetic retinopathy: A systematic review of diagnostic performance. Ophthalmol. Retin. 2019, 3, 294–304. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Nazarian, S.; Glover, B.; Ashrafian, H.; Darzi, A.; Teare, J. Diagnostic accuracy of artificial intelligence and computer-aided diagnosis for the detection and characterization of colorectal polyps: Systematic review and meta-analysis. J. Med. Internet Res. 2021, 23, e27370. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Joseph, S.; Selvaraj, J.; Mani, I.; Kumaragurupari, T.; Shang, X.; Mudgil, P.; Ravilla, T.; He, M. Diagnostic accuracy of artificial intelligence-based automated diabetic retinopathy screening in real-world settings: A systematic review and meta-analysis. Am. J. Ophthalmol. 2024, 263, 214–230. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Bellemo, V.; Lim, Z.W.; Lim, G.; Nguyen, Q.D.; Xie, Y.; Yip, M.Y.; Hamzah, H.; Ho, J.; Lee, X.Q.; Hsu, W. Artificial intelligence using deep learning to screen for referable and vision-threatening diabetic retinopathy in Africa: A clinical validation study. Lancet Digit. Health 2019, 1, e35–e44. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Baldwin, D.R.; Gustafson, J.; Pickup, L.; Arteta, C.; Novotny, P.; Declerck, J.; Kadir, T.; Figueiras, C.; Sterba, A.; Exell, A. External validation of a convolutional neural network artificial intelligence tool to predict malignancy in pulmonary nodules. Thorax 2020, 75, 306–312. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Liu, K.-L.; Wu, T.; Chen, P.-T.; Tsai, Y.M.; Roth, H.; Wu, M.-S.; Liao, W.-C.; Wang, W. Deep learning to distinguish pancreatic cancer tissue from non-cancerous pancreatic tissue: A retrospective study with cross-racial external validation. Lancet Digit. Health 2020, 2, e303–e313. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  40. Eng, D.; Chute, C.; Khandwala, N.; Rajpurkar, P.; Long, J.; Shleifer, S.; Khalaf, M.H.; Sandhu, A.T.; Rodriguez, F.; Maron, D.J. Automated coronary calcium scoring using deep learning with multicenter external validation. npj Digit. Med. 2021, 4, 88. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Pantanowitz, L.; Quiroga-Garza, G.M.; Bien, L.; Heled, R.; Laifenfeld, D.; Linhart, C.; Sandbank, J.; Shach, A.A.; Shalev, V.; Vecsler, M. An artificial intelligence algorithm for prostate cancer diagnosis in whole slide images of core needle biopsies: A blinded clinical validation and deployment study. Lancet Digit. Health 2020, 2, e407–e416. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Raciti, P.; Sue, J.; Retamero, J.A.; Ceballos, R.; Godrich, R.; Kunz, J.D.; Casson, A.; Thiagarajan, D.; Ebrahimzadeh, Z.; Viret, J. Clinical validation of artificial intelligence–augmented pathology diagnosis demonstrates significant gains in diagnostic accuracy in prostate cancer detection. Arch. Pathol. Lab. Med. 2023, 147, 1178–1185. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  43. Attia, Z.I.; Kapa, S.; Yao, X.; Lopez-Jimenez, F.; Mohan, T.L.; Pellikka, P.A.; Carter, R.E.; Shah, N.D.; Friedman, P.A.; Noseworthy, P.A. Prospective validation of a deep learning electrocardiogram algorithm for the detection of left ventricular systolic dysfunction. J. Cardiovasc. Electrophysiol. 2019, 30, 668–674. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Tariq, Q.; Daniels, J.; Schwartz, J.N.; Washington, P.; Kalantarian, H.; Wall, D.P. Mobile detection of autism through machine learning on home video: A development and prospective validation study. PLoS Med. 2018, 15, e1002705. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. Heydon, P.; Egan, C.; Bolter, L.; Chambers, R.; Anderson, J.; Aldington, S.; Stratton, I.M.; Scanlon, P.H.; Webster, L.; Mann, S.; et al. Prospective evaluation of an artificial intelligence-enabled algorithm for automated diabetic retinopathy screening of 30 000 patients. Br. J. Ophthalmol. 2021, 105, 723–728. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Churpek, M.M.; Carey, K.A.; Edelson, D.P.; Singh, T.; Astor, B.C.; Gilbert, E.R.; Winslow, C.; Shah, N.; Afshar, M.; Koyner, J.L. Internal and external validation of a machine learning risk score for acute kidney injury. JAMA Netw. Open 2020, 3, e2012892. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Clift, A.K.; Dodwell, D.; Lord, S.; Petrou, S.; Brady, M.; Collins, G.S.; Hippisley-Cox, J. Development and internal-external validation of statistical and machine learning models for breast cancer prognostication: Cohort study. BMJ 2023, 381, e073800. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S.S.; Wei, J.; Chung, H.W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S. Large language models encode clinical knowledge. Nature 2023, 620, 172–180. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  49. Hager, P.; Jungmann, F.; Holland, R.; Bhagat, K.; Hubrecht, I.; Knauer, M.; Vielhauer, J.; Makowski, M.; Braren, R.; Kaissis, G. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat. Med. 2024, 30, 2613–2622. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  50. Savage, T.; Nayak, A.; Gallo, R.; Rangan, E.; Chen, J.H. Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine. npj Digit. Med. 2024, 7, 20. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  51. McDuff, D.; Schaekermann, M.; Tu, T.; Palepu, A.; Wang, A.; Garrison, J.; Singhal, K.; Sharma, Y.; Azizi, S.; Kulkarni, K. Towards accurate differential diagnosis with large language models. Nature 2025, 642, 451–457. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  52. Goh, E.; Gallo, R.; Hom, J.; Strong, E.; Weng, Y.; Kerman, H.; Cool, J.A.; Kanjee, Z.; Parsons, A.S.; Ahuja, N. Large language model influence on diagnostic reasoning: A randomized clinical trial. JAMA Netw. Open 2024, 7, e2440969. [Google Scholar] [PubMed]
  53. Rao, A.; Pang, M.; Kim, J.; Kamineni, M.; Lie, W.; Prasad, A.K.; Landman, A.; Dreyer, K.; Succi, M.D. Assessing the utility of ChatGPT throughout the entire clinical workflow: Development and usability study. J. Med. Internet Res. 2023, 25, e48659. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  54. Zaretsky, J.; Kim, J.M.; Baskharoun, S.; Zhao, Y.; Austrian, J.; Aphinyanaphongs, Y.; Gupta, R.; Blecker, S.B.; Feldman, J. Generative artificial intelligence to transform inpatient discharge summaries to patient-friendly language and format. JAMA Netw. Open 2024, 7, e240357. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Ke, Y.H.; Jin, L.; Elangovan, K.; Ong, B.W.X.; Oh, C.Y.; Sim, J.; Loh, K.W.T.; Soh, C.R.; Cheng, J.M.H.; Lee, A.K.Y. Real-world deployment and evaluation of PEri-operative AI CHatbot (PEACH): A large language model chatbot for peri-operative medicine. Anaesthesia 2026, 81, 62–71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  56. Alber, D.A.; Yang, Z.; Alyakin, A.; Yang, E.; Rai, S.; Valliani, A.A.; Zhang, J.; Rosenbaum, G.R.; Amend-Thomas, A.K.; Kurland, D.B. Medical large language models are vulnerable to data-poisoning attacks. Nat. Med. 2025, 31, 618–626. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  57. Omar, M.; Sorin, V.; Collins, J.D.; Reich, D.; Freeman, R.; Gavin, N.; Charney, A.; Stump, L.; Bragazzi, N.L.; Nadkarni, G.N. Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Commun. Med. 2025, 5, 330. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  58. Levra, A.G.; Gatti, M.; Mene, R.; Shiffer, D.; Costantino, G.; Solbiati, M.; Furlan, R.; Dipaola, F. A large language model-based clinical decision support system for syncope recognition in the emergency department: A framework for clinical workflow integration. Eur. J. Intern. Med. 2025, 131, 113–120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  59. Zhou, J.; He, X.; Sun, L.; Xu, J.; Chen, X.; Chu, Y.; Zhou, L.; Liao, X.; Zhang, B.; Afvari, S. Pre-trained multimodal large language model enhances dermatological diagnosis using SkinGPT-4. Nat. Commun. 2024, 15, 5649. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  60. Tanno, R.; Barrett, D.G.; Sellergren, A.; Ghaisas, S.; Dathathri, S.; See, A.; Welbl, J.; Lau, C.; Tu, T.; Azizi, S. Collaboration between clinicians and vision–language models in radiology report generation. Nat. Med. 2025, 31, 599–608. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  61. Chen, R.J.; Lu, M.Y.; Williamson, D.F.; Chen, T.Y.; Lipkova, J.; Noor, Z.; Shaban, M.; Shady, M.; Williams, M.; Joo, B. Pan-cancer integrative histology-genomic analysis via multimodal deep learning. Cancer Cell 2022, 40, 865–878.e866. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  62. Boehm, K.M.; Aherne, E.A.; Ellenson, L.; Nikolovski, I.; Alghamdi, M.; Vázquez-García, I.; Zamarin, D.; Long Roche, K.; Liu, Y.; Patel, D. Multimodal data integration using machine learning improves risk stratification of high-grade serous ovarian cancer. Nat. Cancer 2022, 3, 723–733. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  63. Qiu, S.; Miller, M.I.; Joshi, P.S.; Lee, J.C.; Xue, C.; Ni, Y.; Wang, Y.; De Anda-Duran, I.; Hwang, P.H.; Cramer, J.A. Multimodal deep learning for Alzheimer’s disease dementia assessment. Nat. Commun. 2022, 13, 3404. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  64. Tiulpin, A.; Klein, S.; Bierma-Zeinstra, S.M.; Thevenot, J.; Rahtu, E.; Meurs, J.v.; Oei, E.H.; Saarakkala, S. Multimodal machine learning-based knee osteoarthritis progression prediction from plain radiographs and clinical data. Sci. Rep. 2019, 9, 20038. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  65. Khader, F.; Müller-Franzes, G.; Wang, T.; Han, T.; Tayebi Arasteh, S.; Haarburger, C.; Stegmaier, J.; Bressem, K.; Kuhl, C.; Nebelung, S. Multimodal deep learning for integrating chest radiographs and clinical parameters: A case for transformers. Radiology 2023, 309, e230806. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  66. Ferber, D.; Wölflein, G.; Wiest, I.C.; Ligero, M.; Sainath, S.; Ghaffari Laleh, N.; El Nahhas, O.S.; Müller-Franzes, G.; Jäger, D.; Truhn, D. In-context learning enables multimodal large language models to classify cancer pathology images. Nat. Commun. 2024, 15, 10104. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  67. Kurz, C.F.; Merzhevich, T.; Eskofier, B.M.; Kather, J.N.; Gmeiner, B. Benchmarking vision-language models for diagnostics in emergency and critical care settings. npj Digit. Med. 2025, 8, 423. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  68. Huang, S.-C.; Pareek, A.; Zamanian, R.; Banerjee, I.; Lungren, M.P. Multimodal fusion with deep neural networks for leveraging CT imaging and electronic health record: A case-study in pulmonary embolism detection. Sci. Rep. 2020, 10, 22147. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  69. Vale-Silva, L.A.; Rohr, K. Long-term cancer survival prediction using multimodal deep learning. Sci. Rep. 2021, 11, 13505. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  70. Liu, J.; Capurro, D.; Nguyen, A.; Verspoor, K. Attention-based multimodal fusion with contrast for robust clinical prediction in the face of missing modalities. J. Biomed. Inform. 2023, 145, 104466. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  71. Koutsouleris, N.; Dwyer, D.B.; Degenhardt, F.; Maj, C.; Urquijo-Castro, M.F.; Sanfelici, R.; Popovic, D.; Oeztuerk, O.; Haas, S.S.; Weiske, J.; et al. Multimodal Machine Learning Workflows for Prediction of Psychosis in Patients With Clinical High-Risk Syndromes and Recent-Onset Depression. JAMA Psychiatry 2020, 78, 195–209. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  72. Ferber, D.; El Nahhas, O.S.; Wölflein, G.; Wiest, I.C.; Clusmann, J.; Leßmann, M.-E.; Foersch, S.; Lammert, J.; Tschochohei, M.; Jäger, D. Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology. Nat. Cancer 2025, 6, 1337–1349. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  73. Woo, J.J.; Yang, A.J.; Olsen, R.J.; Hasan, S.S.; Nawabi, D.H.; Nwachukwu, B.U.; Williams, R.J., III; Ramkumar, P.N. Custom large language models improve accuracy: Comparing retrieval augmented generation and artificial intelligence agents to noncustom models for evidence-based medicine. Arthroscopy 2025, 41, 565–573.e566. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  74. Wang, Q.; Wang, Z.; Li, M.; Ni, X.; Tan, R.; Zhang, W.; Wubulaishan, M.; Wang, W.; Yuan, Z.; Zhang, Z. A feasibility study of automating radiotherapy planning with large language model agents. Phys. Med. Biol. 2025, 70, 075007. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  75. Liu, S.; Huang, S.S.; McCoy, A.B.; Wright, A.P.; Horst, S.; Wright, A. Optimizing Order Sets With a Large Language Model–Powered Multiagent System. JAMA Netw. Open 2025, 8, e2533277. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  76. Wang, W. LLM-based multi-agent system for neuro-ophthalmic diagnosis and personalized treatment planning. Front. Neurosci. 2025, 19, 1688509. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  77. Li, R.; Wang, X.; Berlowitz, D.; Mez, J.; Lin, H.; Yu, H. CARE-AD: A multi-agent large language model framework for Alzheimer’s disease prediction using longitudinal clinical notes. npj Digit. Med. 2025, 8, 541. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  78. Yuan, T. Safety-aware AI for NSCLC trial pre-screening: A comparative proof-of-concept study of rule-based, single-agent, and multi-agent approaches. Int. J. Med. Inf. 2026, 220, 106628. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  79. Wang, J.; Swoboda, D.M.; Nazha, A. Autonomous Analysis of Curated Patient Data Using a Large Language Model–Based Multiagent Framework. JCO Clin. Cancer Inform. 2025, e2500176. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  80. Klang, E.; Glicksberg, B.S.; Gorenshtein, A.; Gavin, N.; Freeman, R.; Stump, L.; Charney, A.W.; Wei Ting, D.S.; Omar, M.; Nadkarni, G.N. Clinical agents fail silently on patient identity. Int. J. Med. Inform. 2026, 218, 106514. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  81. Lång, K.; Josefsson, V.; Larsson, A.-M.; Larsson, S.; Högberg, C.; Sartor, H.; Hofvind, S.; Andersson, I.; Rosso, A. Artificial intelligence-supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): A clinical safety analysis of a randomised, controlled, non-inferiority, single-blinded, screening accuracy study. Lancet Oncol. 2023, 24, 936–944. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  82. Ipp, E.; Liljenquist, D.; Bode, B.; Shah, V.N.; Silverstein, S.; Regillo, C.D.; Lim, J.I.; Sadda, S.; Domalpally, A.; Gray, G.; et al. Pivotal Evaluation of an Artificial Intelligence System for Autonomous Detection of Referrable and Vision-Threatening Diabetic Retinopathy. JAMA Netw. Open 2021, 4, e2134254. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  83. Ruamviboonsuk, P.; Tiwari, R.; Sayres, R.; Nganthavee, V.; Hemarat, K.; Kongprayoon, A.; Raman, R.; Levinstein, B.; Liu, Y.; Schaekermann, M.; et al. Real-time diabetic retinopathy screening by deep learning in a multisite national screening programme: A prospective interventional cohort study. Lancet Digit. Health 2022, 4, e235–e244. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  84. Wang, P.; Berzin, T.M.; Glissen Brown, J.R.; Bharadwaj, S.; Becq, A.; Xiao, X.; Liu, P.; Li, L.; Song, Y.; Zhang, D.; et al. Real-time automatic detection system increases colonoscopic polyp and adenoma detection rates: A prospective randomised controlled study. Gut 2019, 68, 1813. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  85. Repici, A.; Badalamenti, M.; Maselli, R.; Correale, L.; Radaelli, F.; Rondonotti, E.; Ferrara, E.; Spadaccini, M.; Alkandari, A.; Fugazza, A.; et al. Efficacy of Real-Time Computer-Aided Detection of Colorectal Neoplasia in a Randomized Trial. Gastroenterology 2020, 159, 512–520.e517. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  86. Hollon, T.C.; Pandian, B.; Adapa, A.R.; Urias, E.; Save, A.V.; Khalsa, S.S.S.; Eichberg, D.G.; D’Amico, R.S.; Farooq, Z.U.; Lewis, S.; et al. Near real-time intraoperative brain tumor diagnosis using stimulated Raman histology and deep neural networks. Nat. Med. 2020, 26, 52–58. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  87. Xue, C.; Kowshik, S.S.; Lteif, D.; Puducheri, S.; Jasodanand, V.H.; Zhou, O.T.; Walia, A.S.; Guney, O.B.; Zhang, J.D.; Poésy, S.; et al. AI-based differential diagnosis of dementia etiologies on multimodal data. Nat. Med. 2024, 30, 2977–2989. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  88. Saha, A.; Bosma, J.S.; Twilt, J.J.; van Ginneken, B.; Bjartell, A.; Padhani, A.R.; Bonekamp, D.; Villeirs, G.; Salomon, G.; Giannarini, G.; et al. Artificial intelligence and radiologists in prostate cancer detection on MRI (PI-CAI): An international, paired, non-inferiority, confirmatory study. Lancet Oncol. 2024, 25, 879–887. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  89. Shi, Z.; Miao, C.; Schoepf, U.J.; Savage, R.H.; Dargis, D.M.; Pan, C.; Chai, X.; Li, X.L.; Xia, S.; Zhang, X.; et al. A clinically applicable deep-learning model for detecting intracranial aneurysm in computed tomography angiography images. Nat. Commun. 2020, 11, 6090. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  90. Marchetti, M.A.; Cowen, E.A.; Kurtansky, N.R.; Weber, J.; Dauscher, M.; DeFazio, J.; Deng, L.; Dusza, S.W.; Haliasos, H.; Halpern, A.C.; et al. Prospective validation of dermoscopy-based open-source artificial intelligence for melanoma diagnosis (PROVE-AI study). npj Digit. Med. 2023, 6, 127. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  91. Menzies, S.W.; Sinz, C.; Menzies, M.; Lo, S.N.; Yolland, W.; Lingohr, J.; Razmara, M.; Tschandl, P.; Guitera, P.; Scolyer, R.A.; et al. Comparison of humans versus mobile phone-powered artificial intelligence for the diagnosis and management of pigmented skin cancer in secondary care: A multicentre, prospective, diagnostic, clinical trial. Lancet Digit. Health 2023, 5, e679–e691. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  92. Duron, L.; Ducarouge, A.; Gillibert, A.; Lainé, J.; Allouche, C.; Cherel, N.; Zhang, Z.; Nitche, N.; Lacave, E.; Pourchot, A.; et al. Assessment of an AI Aid in Detection of Adult Appendicular Skeletal Fractures by Emergency Physicians and Radiologists: A Multicenter Cross-sectional Diagnostic Study. Radiology 2021, 300, 120–129. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  93. Porter, P.; Abeyratne, U.; Swarnkar, V.; Tan, J.; Ng, T.-W.; Brisbane, J.M.; Speldewinde, D.; Choveaux, J.; Sharan, R.; Kosasih, K.; et al. A prospective multicentre study testing the diagnostic accuracy of an automated cough sound centred analytic system for the identification of common respiratory disorders in children. Respir. Res. 2019, 20, 81. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  94. Mao, Q.; Jay, M.; Hoffman, J.L.; Calvert, J.; Barton, C.; Shimabukuro, D.; Shieh, L.; Chettipally, U.; Fletcher, G.; Kerem, Y.; et al. Multicentre validation of a sepsis prediction algorithm using only vital sign data in the emergency department, general ward and ICU. BMJ Open 2018, 8, e017833. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  95. Seyyed-Kalantari, L.; Zhang, H.; McDermott, M.B.; Chen, I.Y.; Ghassemi, M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat. Med. 2021, 27, 2176–2182. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  96. Gichoya, J.W.; Banerjee, I.; Bhimireddy, A.R.; Burns, J.L.; Celi, L.A.; Chen, L.-C.; Correa, R.; Dullerud, N.; Ghassemi, M.; Huang, S.-C. AI recognition of patient race in medical imaging: A modelling study. Lancet Digit. Health 2022, 4, e406–e414. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  97. Gaube, S.; Suresh, H.; Raue, M.; Merritt, A.; Berkowitz, S.J.; Lermer, E.; Coughlin, J.F.; Guttag, J.V.; Colak, E.; Ghassemi, M. Do as AI say: Susceptibility in deployment of clinical decision-aids. npj Digit. Med. 2021, 4, 31. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  98. Tschandl, P.; Rinner, C.; Apalla, Z.; Argenziano, G.; Codella, N.; Halpern, A.; Janda, M.; Lallas, A.; Longo, C.; Malvehy, J. Human–computer collaboration for skin cancer recognition. Nat. Med. 2020, 26, 1229–1234. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  99. Finlayson, S.G.; Bowers, J.D.; Ito, J.; Zittrain, J.L.; Beam, A.L.; Kohane, I.S. Adversarial attacks on medical machine learning. Science 2019, 363, 1287–1289. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  100. Davis, S.E.; Lasko, T.A.; Chen, G.; Siew, E.D.; Matheny, M.E. Calibration drift in regression and machine learning models for acute kidney injury. J. Am. Med. Inform. Assoc. 2017, 24, 1052–1061. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  101. Finlayson, S.G.; Subbaswamy, A.; Singh, K.; Bowers, J.; Kupke, A.; Zittrain, J.; Kohane, I.S.; Saria, S. The clinician and dataset shift in artificial intelligence. N. Engl. J. Med. 2021, 385, 283–286. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  102. Obermeyer, Z.; Powers, B.; Vogeli, C.; Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 2019, 366, 447–453. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  103. Bankuoru Egala, S.; Liang, D. Algorithm aversion to mobile clinical decision support among clinicians: A choice-based conjoint analysis. Eur. J. Inf. Syst. 2024, 33, 1016–1032. [Google Scholar] [CrossRef] [Scilit]
  104. Food, U.; Administration, D. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions; FDA: Silver Spring, MD, USA, 2025. [Google Scholar]
  105. Ong Ly, C.; Unnikrishnan, B.; Tadic, T.; Patel, T.; Duhamel, J.; Kandel, S.; Moayedi, Y.; Brudno, M.; Hope, A.; Ross, H.; et al. Shortcut learning in medical AI hinders generalization: Method for estimating AI model generalization without external data. npj Digit. Med. 2024, 7, 124. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  106. Seah, J.; Tang, C.; Buchlak, Q.D.; Milne, M.R.; Holt, X.; Ahmad, H.; Lambert, J.; Esmaili, N.; Oakden-Rayner, L.; Brotchie, P. Do comprehensive deep learning algorithms suffer from hidden stratification? A retrospective study on pneumothorax detection in chest radiography. BMJ Open 2021, 11, e053024. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  107. Brown, A.; Tomasev, N.; Freyberg, J.; Liu, Y.; Karthikesalingam, A.; Schrouff, J. Detecting shortcut learning for fair medical AI using shortcut testing. Nat. Commun. 2023, 14, 4314. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  108. Wolf, R.M.; Abramoff, M.D.; Channa, R.; Tava, C.; Clarida, W.; Lehmann, H.P. Potential reduction in healthcare carbon footprint by autonomous artificial intelligence. npj Digit. Med. 2022, 5, 62. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  109. Perkonigg, M.; Hofmanninger, J.; Herold, C.J.; Brink, J.A.; Pianykh, O.; Prosch, H.; Langs, G. Dynamic memory to alleviate catastrophic forgetting in continual learning with medical imaging. Nat. Commun. 2021, 12, 5678. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  110. González, C.; Ranem, A.; Pinto dos Santos, D.; Othman, A.; Mukhopadhyay, A. Lifelong nnU-Net: A framework for standardized medical continual learning. Sci. Rep. 2023, 13, 9381. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  111. Sheller, M.J.; Edwards, B.; Reina, G.A.; Martin, J.; Pati, S.; Kotrotsou, A.; Milchenko, M.; Xu, W.; Marcus, D.; Colen, R.R. Federated learning in medicine: Facilitating multi-institutional collaborations without sharing patient data. Sci. Rep. 2020, 10, 12598. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  112. Dayan, I.; Roth, H.R.; Zhong, A.; Harouni, A.; Gentili, A.; Abidin, A.Z.; Liu, A.; Costa, A.B.; Wood, B.J.; Tsai, C.-S. Federated learning for predicting clinical outcomes in patients with COVID-19. Nat. Med. 2021, 27, 1735–1743. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  113. Akter, M.; Moustafa, N.; Lynar, T.; Razzak, I. Edge intelligence: Federated learning-based privacy protection framework for smart healthcare systems. IEEE J. Biomed. Health Inform. 2022, 26, 5805–5816. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  114. Chandra, S.; Prakash, P.; Samanta, S.; Chilukuri, S. ClinicalGAN: Powering patient monitoring in clinical trials with patient digital twins. Sci. Rep. 2024, 14, 12236. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  115. Rocha, H.L.; Aguilar, B.; Getz, M.; Shmulevich, I.; Macklin, P. A multiscale model of immune surveillance in micrometastases gives insights on cancer patient digital twins. npj Syst. Biol. Appl. 2024, 10, 144. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  116. Pedram, S.A.; Shin, C.; Ferguson, P.W.; Ma, J.; Dutson, E.P.; Rosen, J. Autonomous suturing framework and quantification using a cable-driven surgical robot. IEEE Trans. Robot. 2020, 37, 404–417. [Google Scholar] [CrossRef] [Scilit]
  117. Wang, J.; Yue, C.; Wang, G.; Gong, Y.; Li, H.; Yao, W.; Kuang, S.; Liu, W.; Wang, J.; Su, B. Task autonomous medical robot for both incision stapling and staples removal. IEEE Robot. Autom. Lett. 2022, 7, 3279–3285. [Google Scholar] [CrossRef] [Scilit]
  118. Topsakal, V.; Heuninck, E.; Matulic, M.; Tekin, A.M.; Mertens, G.; Van Rompaey, V.; Galeazzi, P.; Zoka-Assadi, M.; Van de Heyning, P. First study in men evaluating a surgical robotic tool providing autonomous inner ear access for cochlear implantation. Front. Neurol. 2022, 13, 804507. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  119. Rapp, J.T.; Bremer, B.J.; Romero, P.A. Self-driving laboratories to autonomously navigate the protein fitness landscape. Nat. Chem. Eng. 2024, 1, 97–107. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  120. Trost, F.; Zhang, B.; Aring, I.; Bauer, M.; Glamann, L.; Wessolly, M.; Johnson, K.; Göbel, H.; Lerbs, T.; Sangenne, T. An agentic framework for autonomous scientific discovery in cancer pathology. Nat. Med. 2026, 32, 2254–2266. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  121. Yao, X.; Rushlow, D.R.; Inselman, J.W.; McCoy, R.G.; Thacher, T.D.; Behnken, E.M.; Bernard, M.E.; Rosas, S.L.; Akfaly, A.; Misra, A.; et al. Artificial intelligence–enabled electrocardiograms for identification of patients with low ejection fraction: A pragmatic, randomized clinical trial. Nat. Med. 2021, 27, 815–819. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  122. Noseworthy, P.A.; Attia, Z.I.; Behnken, E.M.; Giblon, R.E.; Bews, K.A.; Liu, S.; Gosse, T.A.; Linn, Z.D.; Deng, Y.; Yin, J.; et al. Artificial intelligence-guided screening for atrial fibrillation using electrocardiogram during sinus rhythm: A prospective non-randomised interventional trial. Lancet 2022, 400, 1206–1212. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  123. Klein, E.A.; Richards, D.; Cohn, A.; Tummala, M.; Lapham, R.; Cosgrove, D.; Chung, G.; Clement, J.; Gao, J.; Hunkapiller, N.; et al. Clinical validation of a targeted methylation-based multi-cancer early detection test using an independent validation set. Ann. Oncol. 2021, 32, 1167–1177. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  124. Feng, L.; Liu, Z.; Li, C.; Li, Z.; Lou, X.; Shao, L.; Wang, Y.; Huang, Y.; Chen, H.; Pang, X.; et al. Development and validation of a radiopathomics model to predict pathological complete response to neoadjuvant chemoradiotherapy in locally advanced rectal cancer: A multicentre observational study. Lancet Digit. Health 2022, 4, e8–e17. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  125. Fulmer, R.; Joerin, A.; Gentile, B.; Lakerink, L.; Rauws, M. Using Psychological Artificial Intelligence (Tess) to Relieve Symptoms of Depression and Anxiety: Randomized Controlled Trial. JMIR Ment. Health 2018, 5, e64. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  126. Webb, C.A.; Trivedi, M.H.; Cohen, Z.D.; Dillon, D.G.; Fournier, J.C.; Goer, F.; Fava, M.; McGrath, P.J.; Weissman, M.; Parsey, R.; et al. Personalized prediction of antidepressant v. placebo response: Evidence from the EMBARC study. Psychol. Med. 2019, 49, 1118–1127. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  127. Jayakumar, P.; Moore, M.G.; Furlough, K.A.; Uhler, L.M.; Andrawis, J.P.; Koenig, K.M.; Aksan, N.; Rathouz, P.J.; Bozic, K.J. Comparison of an Artificial Intelligence–Enabled Patient Decision Aid vs Educational Material on Decision Quality, Shared Decision-Making, Patient Experience, and Functional Outcomes in Adults With Knee Osteoarthritis: A Randomized Clinical Trial. JAMA Netw. Open 2021, 4, e2037107. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  128. Marcuzzi, A.; Nordstoga, A.L.; Bach, K.; Aasdahl, L.; Nilsen, T.I.L.; Bardal, E.M.; Boldermo, N.Ø.; Falkener Bertheussen, G.; Marchand, G.H.; Gismervik, S.; et al. Effect of an Artificial Intelligence–Based Self-Management App on Musculoskeletal Health in Patients With Neck and/or Low Back Pain Referred to Specialist Care: A Randomized Clinical Trial. JAMA Netw. Open 2023, 6, e2320400. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  129. Auloge, P.; Cazzato, R.L.; Ramamurthy, N.; de Marini, P.; Rousseau, C.; Garnon, J.; Charles, Y.P.; Steib, J.-P.; Gangi, A. Augmented reality and artificial intelligence-based navigation during percutaneous vertebroplasty: A pilot randomised clinical trial. Eur. Spine J. 2020, 29, 1580–1589. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  130. Adams, R.; Henry, K.E.; Sridharan, A.; Soleimani, H.; Zhan, A.; Rawat, N.; Johnson, L.; Hager, D.N.; Cosgrove, S.E.; Markowski, A.; et al. Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nat. Med. 2022, 28, 1455–1460. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  131. Levin, S.; Toerper, M.; Hamrock, E.; Hinson, J.S.; Barnes, S.; Gardner, H.; Dugas, A.; Linton, B.; Kirsch, T.; Kelen, G. Machine-Learning-Based Electronic Triage More Accurately Differentiates Patients With Respect to Clinical Outcomes Compared With the Emergency Severity Index. Ann. Emerg. Med. 2018, 71, 565–574.e562. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  132. Brennan, M.; Puri, S.; Ozrazgat-Baslanti, T.; Feng, Z.; Ruppert, M.; Hashemighouchani, H.; Momcilovic, P.; Li, X.; Wang, D.Z.; Bihorac, A. Comparing clinical judgment with the MySurgeryRisk algorithm for preoperative risk assessment: A pilot usability study. Surgery 2019, 165, 1035–1045. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  133. Ardila, D.; Kiraly, A.P.; Bharadwaj, S.; Choi, B.; Reicher, J.J.; Peng, L.; Tse, D.; Etemadi, M.; Ye, W.; Corrado, G.; et al. End-to-end lung cancer screening with three-dimensional deep learning on low-dose chest computed tomography. Nat. Med. 2019, 25, 954–961. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  134. Sinha, P.; Delucchi, K.L.; McAuley, D.F.; O’Kane, C.M.; Matthay, M.A.; Calfee, C.S. Development and validation of parsimonious algorithms to classify acute respiratory distress syndrome phenotypes: A secondary analysis of randomised controlled trials. Lancet Respir. Med. 2020, 8, 247–257. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  135. Chen, T.; Li, X.; Li, Y.; Xia, E.; Qin, Y.; Liang, S.; Xu, F.; Liang, D.; Zeng, C.; Liu, Z. Prediction and Risk Stratification of Kidney Outcomes in IgA Nephropathy. Am. J. Kidney Dis. 2019, 74, 300–309. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  136. Kers, J.; Bülow, R.D.; Klinkhammer, B.M.; Breimer, G.E.; Fontana, F.; Abiola, A.A.; Hofstraat, R.; Corthals, G.L.; Peters-Sengers, H.; Djudjaj, S.; et al. Deep learning-based classification of kidney transplant pathology: A retrospective, multicentre, proof-of-concept study. Lancet Digit. Health 2022, 4, e18–e26. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  137. Illingworth, P.J.; Venetis, C.; Gardner, D.K.; Nelson, S.M.; Berntsen, J.; Larman, M.G.; Agresta, F.; Ahitan, S.; Ahlström, A.; Cattrall, F.; et al. Deep learning versus manual morphology-based embryo selection in IVF: A randomized, double-blind noninferiority trial. Nat. Med. 2024, 30, 3114–3120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  138. Artzi, N.S.; Shilo, S.; Hadar, E.; Rossman, H.; Barbash-Hazan, S.; Ben-Haroush, A.; Balicer, R.D.; Feldman, B.; Wiznitzer, A.; Segal, E. Prediction of gestational diabetes based on nationwide electronic health records. Nat. Med. 2020, 26, 71–76. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  139. Pavel, A.M.; Rennie, J.M.; de Vries, L.S.; Blennow, M.; Foran, A.; Shah, D.K.; Pressler, R.M.; Kapellou, O.; Dempsey, E.M.; Mathieson, S.R.; et al. A machine-learning algorithm for neonatal seizure recognition: A multicentre, randomised, controlled trial. Lancet Child. Adolesc. Health 2020, 4, 740–749. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  140. Giannini, H.M.; Ginestra, J.C.; Chivers, C.; Draugelis, M.; Hanish, A.; Schweickert, W.D.; Fuchs, B.D.; Meadows, L.; Lynch, M.; Donnelly, P.J.; et al. A Machine Learning Algorithm to Predict Severe Sepsis and Septic Shock: Development, Implementation, and Impact on Clinical Practice. Crit. Care Med. 2019, 47, 1485–1492. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Evolution of artificial intelligence capabilities in healthcare.
Figure 1. Evolution of artificial intelligence capabilities in healthcare.
Bioengineering 13 01175 g001
Table 1. Clinical and translational profile of major AI paradigms in healthcare.
Table 1. Clinical and translational profile of major AI paradigms in healthcare.
AI ParadigmKey Diagnostic UsesPrincipal BenefitKey HazardHighest-Risk FailureCurrent Evidence LevelEvaluation Priority
Predictive AIDetection, classification, risk and prognostic stratificationConsistent estimation of predefined clinical endpointsPoor calibration and limited transportabilityFalse reassurance or inappropriate intervention under distribution shiftMost mature; external and prospective evidence available in selected applicationsExternal/prospective validation, calibration, clinical impact
Generative AIDifferential diagnosis, synthesis of clinical information, documentation and decision supportContextual synthesis and communication of complex informationHallucination, omission, and automation biasPlausible but incorrect output influencing clinical reasoningPredominantly benchmark and retrospective; prospective/workflow evidence remains limitedFactual reliability, uncertainty, abstention/escalation, human–AI evaluation
Multimodal AIIntegration of imaging, text, EHR, signals, pathology, and molecular dataIntegration of complementary evidence across modalitiesModality imbalance, missing data, and integration failureUndetected propagation of missing or discordant informationMainly retrospective; external validation emergingIncremental value over unimodal models, missing-modality robustness, external/prospective validation
Agentic AIMulti-step reasoning, evidence retrieval, tool use, workflow coordinationExecution and coordination of complex clinical tasksError propagation and unsafe tool/action executionIncorrect autonomous action affecting downstream careEarly; largely proof-of-concept and task-level evidenceEnd-to-end task safety, tool failures, recovery, escalation, human override
Table 2. Evidence domains for the clinical translation of diagnostic and predictive artificial intelligence.
Table 2. Evidence domains for the clinical translation of diagnostic and predictive artificial intelligence.
Evaluation DomainCore QuestionTypical Study DesignKey OutcomesInterpretationSupporting Framework
Internal performanceDoes the model perform reliably within the development setting?Bootstrap, cross-validation, held-out test setDiscrimination, calibration, prediction errorEstablishes predictive performance; does not demonstrate transportabilityTRIPOD+AI [9]
External validationDoes performance persist in independent data?Temporal, geographic, cross-institutional, or cross-population validationDiscrimination, calibration, subgroup performanceTests transportability beyond model developmentTRIPOD+AI [9]; Cabitza et al. [8]
Prospective evaluationDoes performance persist when applied to consecutively encountered patients?Prospective cohort or multicentre evaluationProspective discrimination, calibration, failure rateTests robustness under contemporary clinical conditionsDECIDE-AI [10]
Workflow and human–AI evaluationDoes AI alter decisions or clinical processes in a useful and safe manner?Silent deployment, clinician–AI study, implementation studyDecision changes, time, workload, alert burden, human factorsEstablishes operational usefulness and interaction with usersDECIDE-AI [10]
Clinical-impact evaluationDoes AI-assisted care improve clinically meaningful processes or patient outcomes?Controlled, pragmatic, cluster-randomised, or randomised trialPatient outcomes, adverse events, resource use, clinically relevant process outcomesProvides the strongest evidence of clinical utilityCONSORT-AI [21]; SPIRIT-AI [22]
Post-deployment surveillanceAre performance and safety maintained after implementation?Real-world monitoring, registry, continuous surveillanceDataset/model drift, calibration, subgroup performance, safety eventsAssesses durability, equity, and emerging harmsFUTURE-AI [23]
Note. These domains are not necessarily sequential or mutually exclusive; external, prospective, workflow, and impact evaluations may overlap within the same study. The framework synthesises principles from TRIPOD+AI, DECIDE-AI, CONSORT-AI, SPIRIT-AI, FUTURE-AI, and methodological work on external validation [8,9,10,21,22,23].
Table 3. Representative primary studies illustrating the clinical value and limitations of multimodal AI.
Table 3. Representative primary studies illustrating the clinical value and limitations of multimodal AI.
Study Clinical TaskModalitiesKey ContributionTranslational Limitation
Boehm et al., 2022 [62]Ovarian cancer risk stratificationCT + histopathology + genomics + clinical dataComplementary multimodal features improved risk stratification.Single disease setting; broader validation required.
Qiu et al., 2022 [63]Dementia assessmentClinical data + neuropsychology + neuroimaging + functional measuresFlexible combinations supported multi-step dementia assessment.Availability and composition of modalities vary across settings.
Khader et al., 2023 [65]Multi-label chest diagnosisChest radiographs + clinical parametersTransformer-based integration outperformed single-modality models for several tasks.Retrospective evaluation; gain is task-dependent.
Zhou et al., 2024 [59]Dermatological diagnosisSkin images + clinical concepts/textDemonstrated interactive multimodal diagnostic capability on real cases.Specialty-specific evaluation; deployment safety remains unresolved.
Tanno et al., 2025 [60]Radiology report generationChest radiographs + languageDemonstrated clinician-VLM collaboration and expert evaluation of generated reports.Clinically significant errors persisted.
Huang et al., 2020 [68]Pulmonary embolism detectionCT pulmonary angiography + EHRLate fusion outperformed imaging-only and EHR-only models.Single use case; workflow benefit was not directly tested.
Vale-Silva & Rohr, 2021 [69]Pan-cancer survivalClinical + imaging + multi-omicsIntegrated heterogeneous modalities and accommodated missing modalities.Retrospective datasets; clinical utility requires prospective testing.
Liu et al., 2023 [70]Clinical prediction with incomplete dataStructured + unstructured clinical modalitiesAttention/contrastive fusion explicitly addressed missing modalities.Robustness depends on missingness patterns represented during development.
Abbreviations: CT = computed tomography; EHR = electronic health record; VLM = vision–language model.
Table 4. Representative primary studies of agentic AI in healthcare.
Table 4. Representative primary studies of agentic AI in healthcare.
Study Clinical TaskAgentic PropertiesEvaluation FocusAutonomy LevelKey Limitation
Autonomous oncology agent [72]Oncology decision-makingAutonomous clinical reasoning and decision supportDevelopment and validationBounded autonomousBounded task; broader prospective deployment needed.
Custom LLM agents [73]Evidence-based medicineRetrieval/tool-augmented evidence synthesisAccuracy versus non-custom modelsHuman-supervisedTask performance does not establish workflow safety.
Radiotherapy planning agents [74]Radiotherapy planningSequential tool use and planningFeasibility of automated planningHuman-supervisedHigh-consequence actions require stringent oversight.
LLM multi-agent order sets [75]Clinical order-set optimisationMulti-agent critique and optimisationQuality of generated order setsHuman-supervisedDownstream patient outcomes not established.
Neuro-ophthalmic MAS [76]Diagnosis and treatment planningRole-specialised multi-agent reasoningDiagnostic and planning performanceHuman-supervisedSpecialty-specific evaluation.
CARE-AD [77].Alzheimer’s disease predictionMulti-agent reasoning over longitudinal notesPredictive performanceHuman-supervisedRetrospective clinical-note setting.
Safety-aware multi-agent screening [78]NSCLC trial pre-screeningRule-based vs single- vs multi-agent orchestrationSafety and comparative performanceHuman-supervisedProof-of-concept evidence.
Clinical identity failures [80]Patient-specific clinical tasksAgent interaction with patient contextSilent identity-error failureBounded autonomousDemonstrates system-level rather than answer-level risk.
Table 5. Major data modalities used by artificial intelligence across clinical specialties.
Table 5. Major data modalities used by artificial intelligence across clinical specialties.
Clinical SpecialtyMajor Data Modalities Used in AIRepresentative Data
Radiology and medical imagingMedical imaging; clinical/EHR data; textCT, MRI, radiography, mammography, ultrasound, PET, radiology reports
PathologyDigital pathology; imaging; molecular/omics; clinical dataWhole-slide images, histology, cytology, genomic and molecular profiles
CardiologyECG; imaging; physiological monitoring; EHR; wearables12-lead ECG, Holter ECG, echocardiography, cardiac CT/MRI, PPG, blood pressure
NeurologyEEG/EMG; neuroimaging; EHR; wearables; omicsEEG, EMG, evoked potentials, MRI, CT, PET, gait and movement sensors
OncologyImaging; digital pathology; omics; EHR; laboratory dataCT/MRI/PET, histology, genomics, transcriptomics, biomarkers, treatment records
Psychiatry and mental healthClinical text; EHR; EEG; audio; wearablesClinical notes, questionnaires, EEG, speech, sleep, activity and smartphone-derived signals
GastroenterologyEndoscopic video/images; pathology; EHR; laboratory dataColonoscopy, gastroscopy, capsule endoscopy, histology, biochemical markers
OphthalmologyOcular imaging; clinical dataFundus photography, OCT, OCTA, slit-lamp and retinal imaging
DermatologyClinical/dermoscopic imaging; pathology; clinical dataSmartphone photographs, dermoscopy, histopathology
Physical medicine and rehabilitationWearables/sensors; video; EMG; clinical outcomesAccelerometry, inertial sensors, gait, motion capture, surface EMG, functional scores
OrthopaedicsImaging; sensor/biomechanical data; EHRRadiography, CT, MRI, gait, force and motion measurements
Emergency and critical careEHR; laboratory data; ECG; continuous physiological signalsVital signs, ECG, SpO2, blood pressure, laboratory trajectories, medication and intervention data
Surgery and perioperative medicineVideo; imaging; physiological monitoring; EHROperative video, endoscopy, intraoperative imaging, ECG, blood pressure, SpO2
PulmonologyImaging; physiological signals; audio; EHRChest CT/X-ray, spirometry, SpO2, respiratory waveforms, cough and breath sounds
NephrologyEHR; laboratory data; pathology; imaging; omicsCreatinine/eGFR trajectories, urine biomarkers, renal histology, ultrasound, molecular profiles
Obstetrics and gynecologyUltrasound/imaging; physiological signals; EHR; laboratory/omicsFetal ultrasound, cardiotocography, maternal records, laboratory and genomic data
Pediatrics and neonatologyEHR; physiological monitoring; EEG; imaging; audio/videoNeonatal EEG, vital signs, imaging, cough/cry audio, developmental video
Infectious diseasesEHR; laboratory/microbiology; physiological monitoring; text/omicsMicrobiology, antimicrobial susceptibility, vital signs, laboratory trends, genomic sequencing
Table 6. Core dimensions of trustworthy clinical AI and evidence required for governance.
Table 6. Core dimensions of trustworthy clinical AI and evidence required for governance.
DimensionPrincipal RiskEvidence RequiredGovernance ResponseRepresentative Refs
Fairness and equityUnequal error rates; proxy discriminationSubgroup/intersectional discrimination, calibration, error severityPredefined fairness analysis; mitigation; equity monitoring[95,96,102]
Transparency and explainabilityPlausible but unfaithful explanations; inappropriate relianceProvenance, uncertainty, explanation fidelity, human–AI interactionIntended-use documentation; auditability; user training[23,97,98]
Privacy and securityData leakage, adversarial manipulation, compromised inputsThreat modelling, privacy testing, adversarial/robustness evaluationLeast privilege; access control; logging; incident response[23,99]
Human oversightAutomation bias; unclear responsibilityOverride behaviour, intervention effectiveness, failure recognitionRisk-proportionate human review; escalation; accountability[23,97]
Lifecycle governanceDataset shift, model drift, changing workflowProspective performance, calibration, subgroup and outcome monitoringMonitoring thresholds; recalibration, suspension or retraining[23,100,101]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Baydili, İ.; Tasci, B.; Tasci, G.; Dogan, S.; Tuncer, T. Artificial Intelligence in Healthcare: From Predictive Models to Generative, Multimodal, and Agentic AI. Bioengineering 2026, 13, 1175. https://doi.org/10.3390/bioengineering13101175

AMA Style

Baydili İ, Tasci B, Tasci G, Dogan S, Tuncer T. Artificial Intelligence in Healthcare: From Predictive Models to Generative, Multimodal, and Agentic AI. Bioengineering. 2026; 13(10):1175. https://doi.org/10.3390/bioengineering13101175

Chicago/Turabian Style

Baydili, İsmail, Burak Tasci, Gülay Tasci, Sengul Dogan, and Turker Tuncer. 2026. "Artificial Intelligence in Healthcare: From Predictive Models to Generative, Multimodal, and Agentic AI" Bioengineering 13, no. 10: 1175. https://doi.org/10.3390/bioengineering13101175

APA Style

Baydili, İ., Tasci, B., Tasci, G., Dogan, S., & Tuncer, T. (2026). Artificial Intelligence in Healthcare: From Predictive Models to Generative, Multimodal, and Agentic AI. Bioengineering, 13(10), 1175. https://doi.org/10.3390/bioengineering13101175

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop