Next Article in Journal
The Dual Imperative in AI for OCD: Bridging Ethical Frameworks and Explainable Diagnostics
Previous Article in Journal
Exploratory Image-Level Classification of a Public Chest Radiograph Dataset Using a Lightweight SqueezeNet-Based Pipeline
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Operationalizing WHO Ethical Principles for Healthcare AI: A Lifecycle-Aligned Governance-by-Design Framework

by
Kaaviyashri Saraboji
1,
Keerthy Gopalakrishnan
2,
Divyanshi Sood
3,
Anmolpreet Kaur
4,
Suganti Shivaram
4,
Scott A. Helgeson
4,5,6,
Shivaram P. Arunachalam
4,5,6,* and
Dipankar Mitra
1,*
1
Department of Computer Science and Computer Engineering, University of Wisconsin-La Crosse, La Crosse, WI 54601, USA
2
Department of Internal Medicine, Wright Medical Center, Scranton, PA 18503, USA
3
Department of Internal Medicine, UChealth Parkview Medical Center, Pueblo, CO 81005, USA
4
Digital Engineering & Artificial Intelligence Laboratory (DEAL), Mayo Clinic, Jacksonville, FL 32224, USA
5
Division of Pulmonology, Department of Medicine, Mayo Clinic, Jacksonville, FL 32224, USA
6
Department of Critical Care Medicine, Mayo Clinic, Jacksonville, FL 32224, USA
*
Authors to whom correspondence should be addressed.
AI Med. 2026, 1(2), 16; https://doi.org/10.3390/aimed1020016
Submission received: 28 March 2026 / Revised: 29 May 2026 / Accepted: 4 June 2026 / Published: 10 June 2026

Abstract

Artificial intelligence (AI) is rapidly transforming healthcare through applications in clinical decision support, diagnostic imaging, population health management, and workflow optimization. Despite these advances, real-world deployment continues to expose critical challenges related to safety, bias, transparency, and integration into clinical workflows. Algorithmic bias can exacerbate health disparities, limited explainability may undermine clinician trust, and insufficient validation and post-deployment monitoring can compromise patient safety. Although the World Health Organization (WHO) has established six ethical principles for AI in health, including autonomy, well-being and safety, transparency, accountability, equity, and sustainability, translating these high-level principles into practical and enforceable governance mechanisms remains a persistent challenge. This narrative review synthesizes insights from bioethics, health policy, computer science, and clinical medicine to identify gaps in current AI governance approaches and proposes a lifecycle-aligned governance-by-design framework that operationalizes WHO ethical principles across key stages of the healthcare AI lifecycle, including data collection, model development, validation, deployment, and post-deployment monitoring. The framework integrates concrete governance mechanisms such as consent governance, fairness evaluation, external validation, explainability, clinician oversight, and continuous performance monitoring. Overall, this work advances a practical, lifecycle-integrated approach to AI governance and provides a structured foundation for developing safe, equitable, and trustworthy AI systems in healthcare.

1. Introduction

1.1. Artificial Intelligence in Healthcare

Artificial intelligence is becoming an integral part of modern healthcare, with applications spanning clinical decision support, diagnostic imaging, population health management, workflow optimization, and personalized treatment planning [1,2,3]. These systems have shown strong potential to improve efficiency, reduce clinician workload, and enhance both diagnostic and operational accuracy across a wide range of care settings [4,5]. In particular, machine learning models now achieve performance comparable to, and in some cases exceeding, human experts in tasks such as diabetic retinopathy screening, pathology image analysis, radiographic caries detection, and cardiovascular risk prediction [6,7,8,9]. More broadly, artificial intelligence is increasingly recognized as a system-level transformation influencing clinical workflows, healthcare delivery models, and decision-making processes across the healthcare ecosystem [10,11,12]. Such advances offer promising solutions to persistent challenges in healthcare, including workforce shortages, delays in diagnosis, and variability in care quality [3,12].
Despite these advancements, real-world deployment has consistently revealed important limitations [13,14]. Ethical, safety, equity, and accountability concerns continue to emerge even when systems perform well in controlled settings [15,16]. For example, algorithmic bias has been shown to disadvantage certain patient populations, particularly in widely used healthcare risk prediction tools [17]. Additional studies demonstrate that machine learning systems trained on healthcare data may encode and amplify existing disparities, resulting in reduced performance for underrepresented populations [17,18,19]. Limited explainability further complicates clinical adoption, as clinicians may find it difficult to interpret or trust model outputs in high-stakes decision-making scenarios [20,21]. In addition, inadequate external validation remains a major issue, with models often failing to generalize across institutions or diverse patient populations [13,22]. Real-world studies further show that model performance can degrade when deployed across different healthcare settings due to variations in patient populations, data quality, and clinical practices [23,24]. Misalignment with clinical workflows can also introduce unintended risks, especially when AI systems are deployed without sufficient integration into existing care processes [10,25]. Together, these challenges highlight that strong technical performance alone does not guarantee safe or equitable use in practice [14,15].
These recurring issues suggest that the primary challenge is not only algorithmic performance but governance across the artificial intelligence lifecycle, including how systems are designed, evaluated, deployed, and monitored over time [26,27]. As a result, a growing trust gap is emerging among clinicians, patients, and policymakers [28,29]. Clinicians have raised concerns about liability, erosion of professional judgment, automation bias, and the broader implications for clinical decision-making and resource allocation [30,31]. Patients, in turn, express concerns related to privacy, potential discrimination, and reduced human interaction in care [32,33]. At the same time, policymakers are faced with the challenge of balancing innovation with the need to ensure safety, accountability, and public trust [34,35]. In addition, successful adoption depends on effective integration into clinical workflows and adequate training of healthcare professionals to interpret and appropriately use AI-generated outputs [10,11].
A growing body of literature further indicates that current approaches to ethical AI in healthcare remain difficult to implement in practice [36,37]. Many existing frameworks outline important principles but provide limited guidance on how these principles can be translated into real-world clinical settings [28,29]. In parallel, increasing data complexity and reliance on large-scale datasets introduce additional challenges related to data governance and long-term system monitoring [33,38]. Differences in regulatory approaches and institutional capabilities further contribute to inconsistent implementations across healthcare systems [35,39]. Taken together, these limitations highlight the need for more structured governance approaches that extend beyond model development and embed ethical considerations throughout the full lifecycle of healthcare AI systems [26,40].

1.2. WHO Ethical Principles and the Operationalizing Challenge

In response to growing governance challenges, the World Health Organization (WHO) outlined six ethical principles for artificial intelligence in health: protecting human autonomy; promoting human well-being and safety; ensuring transparency and explainability; fostering responsibility and accountability; ensuring inclusiveness and equity; and promoting responsive and sustainable AI [40]. These principles provide a high-level ethical foundation intended to guide the responsible development and use of AI systems in healthcare [28,40]. Figure 1 summarizes these principles and their role in shaping ethical governance.
These principles are grounded in public health priorities, human rights frameworks, and long-standing bioethical values such as autonomy, beneficence, and justice [37,40]. Their broad and principle-based design allows them to remain applicable across diverse healthcare settings and adaptable to rapidly evolving AI technologies [29]. Similar ethical frameworks proposed by international organizations and research communities reinforce these foundational commitments, reflecting a growing global consensus on the core values underlying trustworthy AI [28,29,41]. As a result, WHO guidance has emerged as a widely recognized reference point for ethical AI in medicine [40].
Despite this broad acceptance, a persistent gap remains between ethical principles and their practical implementation [36,37]. Healthcare organizations and AI development teams often lack clear guidance on how to translate these high-level commitments into concrete technical and governance practices [26,34]. For example, while principles such as fairness and transparency are widely endorsed, they do not explicitly define how these concepts should be measured, evaluated, or enforced during different stages of the AI lifecycle [18,21]. This challenge has been widely recognized in the literature, where scholars argue that high-level ethical principles alone are insufficient without operational mechanisms, measurable standards, and enforceable governance structures [15,36].
This challenge is further compounded by the multidisciplinary nature of healthcare AI governance [42]. Effective implementation requires coordination across clinical practice, technical development, regulatory oversight, ethical review, and patient engagement [26,34]. In reality, these domains often operate in silos, limiting communication and shared accountability among stakeholders [25]. Fragmentation across governance frameworks and institutional practices further complicates implementation, as organizations must navigate overlapping and sometimes inconsistent ethical, technical, and regulatory expectations [29,35]. As a result, ethical principles may be acknowledged at a conceptual level but remain insufficiently integrated into day-to-day development workflows and decision-making processes [10,37]. Without structured, lifecycle-aligned governance mechanisms, there is a risk that ethical guidance remains disconnected from how healthcare AI systems are designed, deployed, and maintained in practice [26,40].

1.3. Scope, Objectives, and Contribution

This review focuses on how the World Health Organization’s ethical principles can be operationalized across the lifecycle of predictive healthcare AI systems [40]. It is intended for artificial intelligence developers, clinical AI architects, healthcare informatics leaders, and governance stakeholders involved in designing, validating, and deploying AI systems in clinical settings. Rather than providing a purely philosophical analysis of ethical principles, the goal of this work is to translate established normative guidance into practical, lifecycle-aligned governance mechanisms [36].
Specifically, this review aims to synthesize insights from ethical frameworks, regulatory guidance, and technical standards to identify areas of convergence as well as persistent gaps in current approaches [29,34]. It maps WHO’s six ethical principles to concrete governance mechanisms across five key lifecycle stages: data collection, model development, validation, deployment, and post-deployment monitoring [26,40]. In addition, it examines documented implementation failures to better understand where governance breakdowns occur in practice, drawing on prior analyses of real-world AI deployment challenges and system-level limitations [13,14]. Building on this analysis, the study proposes structured and actionable guidance for embedding governance-by-design within healthcare AI systems.
This work contributes by providing a comprehensive lifecycle mapping of WHO’s ethical principles to stage-specific governance requirements [29,40]. It integrates ethical, regulatory, and technical perspectives that are often addressed separately, demonstrating how approaches such as bias detection, explainability methods, validation protocols, and continuous monitoring can support the practical implementation of ethical commitments [15,20]. It further situates these mechanisms within broader governance and regulatory contexts, aligning with emerging expectations for lifecycle-based oversight and risk management in healthcare AI systems [35,43]. Through case-based analysis, it illustrates how failures at specific stages of the AI lifecycle can lead to downstream risks, and how governance-by-design strategies can help mitigate these issues [25,30]. By explicitly aligning ethical principles with lifecycle checkpoints, the proposed framework aims to bridge the gap between high-level guidance and real-world implementation, addressing widely recognized limitations of existing principle-based approaches [36,37].

1.4. Structure of This Review

This review is organized as follows. Section 2 outlines the methodological approach used in this study. Section 3 provides background on WHO’s ethical principles, the healthcare AI lifecycle, relevant governance frameworks, and key technical and regulatory considerations. Section 4 introduces the proposed lifecycle-aligned operationalization framework. Section 5 discusses areas of convergence across existing approaches, identifies persistent gaps, and examines practical implementation challenges. Section 6 highlights emerging directions and future considerations. Finally, Section 7 concludes with a synthesis of findings and recommendations for key stakeholders.

2. Materials and Methods

This review employs a multidisciplinary narrative synthesis integrating literature from bioethics, health policy, computer science, clinical medicine, and regulatory scholarship [14,37]. The primary objective is to examine how the World Health Organization’s ethical principles can be operationalized across the lifecycle of healthcare AI systems [40].
The sources included in this review span multiple domains [29,34]. These comprise ethical frameworks and guidance documents from international organizations, regulatory authorities, and professional societies, as well as technical standards and reporting guidelines for healthcare AI development and evaluation. In addition, peer-reviewed research addressing algorithmic bias, fairness metrics, explainability methods, lifecycle management, and validation practices was included [20,21,30]. Documented case studies describing both successful and failed AI implementations were also considered to provide practical insights into real-world challenges [25,44].
Relevant literature was identified through structured searches of PubMed, Web of Science, IEEE Xplore, Scopus, and arXiv. Search terms included combinations of keywords related to healthcare AI (e.g., artificial intelligence, machine learning), ethics and governance (e.g., ethics, accountability, transparency), bias and fairness (e.g., bias, equity, disparities), validation (e.g., evaluation, performance), and regulation (e.g., policy, GDPR, FDA, EU AI Act). Priority was given to publications from 2020 to 2025 to capture recent developments, while also incorporating foundational earlier studies where conceptually relevant. The review also considered key regulatory and data governance frameworks relevant to healthcare AI, including data protection and privacy standards that influence model development and deployment practices [32,38,45]. Illustrative case examples were selected based on documented governance failures, availability of independent analysis, and their relevance to lifecycle-specific ethical and operational challenges in real-world healthcare AI deployment. Additional details regarding the narrative review search strategy, principal search concepts, publication timeframe, and case selection rationale are summarized in Table 1.
This study adopts a lifecycle perspective, recognizing that ethical governance should be integrated from project inception through post-deployment monitoring rather than applied retrospectively [26,40]. The synthesis is narrative rather than systematic, with an emphasis on thematic integration and practical operationalization rather than exhaustive literature coverage. Case selection focused on well-documented examples of governance failures that illustrate identifiable breakdowns across different stages of the AI lifecycle and have been independently analyzed or publicly scrutinized [13,25].

3. Background

3.1. WHO’s Six Ethical Principles for AI in Health

The World Health Organization’s ethical framework for AI in healthcare consists of six interconnected principles that collectively define trustworthy AI [40]. These principles build on established bioethical foundations while addressing challenges unique to algorithmic decision-making, offering guidance that can be applied across diverse healthcare systems and cultural contexts [28,37]. Similar frameworks proposed by international organizations and research communities reinforce these commitments, reflecting a growing convergence around core ethical values for responsible AI [29,41].
Principle 1: Protecting Human Autonomy emphasizes the importance of maintaining meaningful human control over healthcare decisions [40]. This includes both patient autonomy, such as informed consent, privacy, and the right to make decisions about one’s care, and professional autonomy, which allows clinicians to exercise independent judgment [30,32]. In practice, this means AI systems should support rather than replace human decision-making [12]. Key requirements include clear disclosure of AI use, the ability for clinicians to override system outputs, and consent processes that account for AI-specific considerations [34]. Without careful design, automation can gradually shift decision-making authority toward algorithmic outputs, making it essential to actively preserve human agency [25,30].
Principle 2: Promoting Human Well-being and Safety requires that AI systems deliver clinical benefit while minimizing harm [40]. This extends beyond individual patient safety to include broader population health outcomes and the prevention of algorithmic risks such as diagnostic errors or inappropriate recommendations [16]. Operationalizing this principle involves rigorous validation across diverse populations, continuous performance monitoring, and structured processes for identifying and responding to adverse events [13,22]. Importantly, AI systems introduce new categories of risk, including model drift, distribution shift, and system-level failures, which require governance approaches beyond traditional clinical safety frameworks [14,23].
Principle 3: Ensuring Transparency and Explainability addresses the often opaque nature of complex machine learning models and their impact on trust and accountability [21]. Transparency operates at multiple levels, including clear documentation of training data, disclosure of model development processes, and the ability to provide meaningful explanations of predictions [20]. While highly complex models may offer improved performance, they can also reduce interpretability [46]. Emerging work in explainable AI highlights both the importance and limitations of post hoc explanation techniques, emphasizing the need for clinically meaningful and context-aware interpretability rather than purely technical explanations [47,48]. This principle therefore requires balancing predictive performance with interpretability to support clinical trust and effective oversight [49].
Principle 4: Fostering Responsibility and Accountability highlights the need for clearly defined responsibility across the AI lifecycle [26]. Healthcare AI systems are developed and deployed by multiple stakeholders, including developers, vendors, healthcare organizations, clinicians, and regulators [34]. This distributed structure can lead to fragmented accountability if not explicitly addressed [37]. To mitigate this, the principle calls for explicit assignment of responsibilities, comprehensive documentation, audit trails, and mechanisms for addressing harm when it occurs [35]. However, existing literature emphasizes that accountability frameworks remain underdeveloped in practice, particularly in complex multi-actor systems [36].
Principle 5: Ensuring Inclusiveness and Equity focuses on preventing AI systems from reinforcing or amplifying existing health disparities [40]. Equity considerations include fair representation in training data, equitable model performance across populations, and access to AI-enabled care [18]. Bias can emerge from historical inequities in healthcare data, measurement limitations, and optimization strategies that prioritize aggregate performance over fairness [17]. Empirical studies have demonstrated that such systems may disproportionately affect underserved populations if not carefully evaluated and mitigated [44]. Addressing these issues requires proactive bias assessment, implementation of fairness metrics, and intentional deployment strategies that include underrepresented groups [44,50].
Principle 6: Promoting Responsive and Sustainable AI emphasizes the need for governance systems that can adapt to ongoing technological change while supporting long-term sustainability [40]. Responsiveness involves updating governance practices as new evidence emerges and as AI systems evolve [26]. Sustainability includes environmental considerations such as computational resource use, as well as economic and social factors including cost-effectiveness, workforce impact, and long-term system maintenance [42]. This principle encourages a balance between innovation and caution, ensuring that short-term benefits do not lead to long-term unintended consequences [37].
The key characteristics and governance requirements associated with these principles are summarized in Table 2.
While these principles articulate high-level expectations, they do not specify how ethical obligations can be systematically embedded across the technical, organizational, and governance stages of the AI lifecycle. This limitation has been widely noted in the literature, where principle-based frameworks are often criticized for lacking operational clarity, measurable criteria, and enforceable implementation pathways [36,37].

3.2. Overview of the AI Lifecycle in Healthcare

Understanding ethical governance requires recognizing that AI systems evolve through distinct lifecycle stages, each presenting unique ethical challenges and requiring stage-specific governance mechanisms [26,37]. While various lifecycle models exist, this review adopts a five-stage framework encompassing data collection and curation, model development and training, validation and clinical evaluation, deployment and clinical integration, and monitoring and continuous learning [40,45]. This framework reflects the understanding that healthcare AI is not a static product but an evolving sociotechnical system requiring continuous oversight and governance [14,42].
Stage 1: Data Collection and Curation involves assembling, annotating, and preparing datasets for model training and evaluation [17]. This foundational stage determines the patterns the AI system will learn and significantly influences downstream performance and fairness [18]. Ethical considerations include obtaining appropriate consent for data use, protecting patient privacy, ensuring representative sampling across demographic groups, maintaining data quality through rigorous annotation protocols, and documenting dataset characteristics for transparency [32,33]. Data governance challenges are further amplified by the scale and heterogeneity of healthcare data, requiring structured approaches to provenance tracking and privacy preservation [38,45]. Data practices at this stage have lasting lifecycle effects, as biases introduced during data collection such as selection bias, measurement bias, and historical bias are often amplified during subsequent model training [50].
Stage 2: Model Development and Training encompass algorithm selection, feature engineering, model training, and initial performance optimization [51]. Developers make numerous technical choices, including feature inclusion, optimization objectives, and class imbalance handling, that implicitly encode values and priorities [15]. Ethical governance at this stage includes incorporating fairness constraints into training objectives, implementing bias detection and mitigation techniques, documenting development processes for reproducibility, integrating explainability methods, and involving clinical experts in system design [18,20]. Advances in explainable AI provide tools for interpreting model behavior, but their limitations highlight the need for careful evaluation of interpretability in clinical contexts [21,48]. This stage presents opportunities to proactively address fairness, for example through adversarial debiasing or constrained optimization, but also risks embedding opacity or unintended bias if governance mechanisms are absent [50].
Stage 3: Validation and Clinical Evaluation rigorously assess model performance, safety, and clinical utility prior to deployment [40]. This stage moves beyond internal technical metrics to evaluate real-world effectiveness and potential harm [52]. Governance mechanisms include internal validation using holdout datasets, external validation across independent sites and populations, prospective clinical trials comparing AI-assisted care with standard practice, disaggregated subgroup performance analysis, safety benchmarking, and structured reporting aligned with standards such as TRIPOD-AI and CONSORT-AI [53,54]. Validation serves as a critical ethical checkpoint: unsafe or biased systems should be identified and remediated before clinical use. However, insufficient external validation remains a recurring failure mode in healthcare AI, particularly when systems are deployed across heterogeneous clinical environments [22,23].
Stage 4: Deployment and Clinical Integration implements validated AI systems within real clinical workflows [11]. This stage involves technical integration with electronic health records, user interface design, clinician training, workflow redesign, and establishment of protocols for human-AI collaboration [10]. Ethical governance includes preserving clinician override capabilities, providing meaningful explanations to support clinical reasoning, ensuring equitable access across patient populations, conducting usability testing with intended end users, and implementing structured incident reporting systems [34]. Real-world deployment studies demonstrate that systems that perform well during validation may still fail in practice due to workflow misalignment, insufficient user training, or contextual variability not captured during development [23,25].
Stage 5: Monitoring and Continuous Learning tracks system performance post-deployment, detects degradation or model drift, investigates adverse events, and governs model updates or retirement [14]. This often-neglected stage is essential for maintaining safety and effectiveness over time [22]. Real-world data distributions shift due to changes in patient populations, evolving clinical practices, and emerging diseases, potentially degrading model performance [26]. Governance mechanisms include automated monitoring dashboards, statistical drift detection tools, periodic revalidation, adverse event surveillance, structured user feedback collection, and clearly defined update protocols [26]. Increasingly, healthcare AI oversight is compared to pharmacovigilance, recognizing the need for ongoing post-market surveillance and continuous risk management [15].
The overall structure of this five-stage lifecycle is illustrated in Figure 2, which summarizes the key stages of healthcare AI development and deployment.
These five stages are iterative and interconnected rather than strictly linear [26]. Monitoring outcomes may necessitate retraining, requiring renewed data collection and development. Validation findings may prompt redesign of dataset construction or feature engineering strategies. This iterative structure reinforces the necessity of governance-by-design: ethical safeguards must be embedded from project inception rather than applied retrospectively [37].
While this lifecycle framework clarifies where ethical risks arise, it does not specify how ethical principles should be operationalized at each stage [36]. Addressing this gap is the purpose of the governance-by-design framework developed in the subsequent section.

3.3. Supporting Frameworks and Standards

The WHO’s ethical principles exist within a broader ecosystem of complementary frameworks, regulatory guidance, and technical standards that collectively shape healthcare AI governance [29,40]. Examining areas of convergence and divergence across these approaches helps identify both shared priorities and persistent gaps in how ethical AI is implemented in practice [28]. At the same time, prior analyses have highlighted that the proliferation of frameworks does not necessarily translate into consistent implementation, underscoring the need for operational alignment across governance approaches [36,37].
The U.S. Food and Drug Administration’s Good Machine Learning Practice (GMLP) guidelines outline key principles for the development and regulation of machine learning-based medical devices [43]. These guidelines emphasize multidisciplinary collaboration, clearly defined intended use, robust software engineering practices, high-quality and representative datasets, and rigorous validation procedures. They also highlight the importance of transparency, real-world performance monitoring, and governance of model updates after deployment. Notably, GMLP reflects a shift toward lifecycle-oriented regulatory thinking, recognizing that AI systems may evolve over time through continuous learning and iterative updates [26,43]. This lifecycle perspective aligns with broader efforts to incorporate continuous risk management and post-market oversight into AI governance [35]. These lifecycle-oriented regulatory approaches increasingly intersect with healthcare MLOps and operational AI governance practices, particularly in areas such as reproducibility, version control, post-market monitoring, and management of continuously evolving clinical AI systems. Recent healthcare MLOps literature has emphasized the importance of reproducible development pipelines, continuous monitoring, governance of model updates, and operational oversight to support the safe deployment and maintenance of clinical AI systems throughout their lifecycle [55,56,57].
The European Union Artificial Intelligence Act introduces a risk-based regulatory framework that categorizes AI systems based on their potential impact. Clinical AI systems are typically classified as high-risk, requiring strict compliance with requirements related to data governance, technical documentation, transparency, human oversight, and performance standards. The Act is designed to integrate with existing healthcare regulations, including the Medical Device Regulation (MDR), In Vitro Diagnostic Regulation (IVDR), and GDPR, thereby embedding AI governance within established regulatory structures [34]. However, challenges remain in defining risk thresholds and ensuring consistent implementation across jurisdictions, particularly as AI systems become more adaptive and complex [39].
The OECD AI Principles provide a broader, cross-sector ethical framework centered on inclusive growth, human-centered values, transparency, robustness, and accountability. While not specific to healthcare, these principles have influenced national AI strategies and international policy alignment [29]. Their emphasis on trustworthiness and lifecycle accountability aligns conceptually with the needs of healthcare AI governance, although additional domain-specific adaptation is required for clinical applications [28].
In parallel, several AI-specific reporting standards have emerged to address transparency and reproducibility challenges in healthcare AI research [53,54]. TRIPOD-AI provides structured guidance for reporting prediction model development and validation, while SPIRIT-AI and CONSORT-AI extend clinical trial standards to AI-based interventions. DECIDE-AI focuses on early-stage evaluation of decision support systems, and STARD-AI addresses reporting requirements for diagnostic AI tools. Together, these standards aim to improve the quality, transparency, and reproducibility of AI research, supporting more reliable translation into clinical practice.
A comparative overview of major international AI governance frameworks, including their scope, key principles, and enforcement mechanisms, is presented in Table 3.
These frameworks show substantial alignment around core themes such as safety, transparency, accountability, and fairness, although they differ in scope, level of detail, and enforcement mechanisms [29,40]. Healthcare-specific frameworks, such as WHO guidance and FDA GMLP, provide more context-sensitive direction, while broader frameworks such as OECD and IEEE offer general ethical foundations that require adaptation for clinical use [35]. Regulatory instruments, including the FDA framework and the EU AI Act, provide enforceable requirements but may evolve more slowly than the technologies they govern. Moreover, increasing complexity in AI systems, particularly with adaptive and data-intensive models, further challenges traditional regulatory approaches and necessitates more flexible governance structures [14,42].
Key AI-specific reporting standards and their purpose and reporting requirements are summarized in Table 4.
These reporting standards address well-documented limitations in transparency and reproducibility within healthcare AI research [20]. However, adherence remains inconsistent, with many studies still lacking sufficient detail on data sources, development processes, or external validation [13]. While increasing expectations from journals and regulators may improve compliance over time, these standards alone do not ensure consistent ethical implementation [36].
Despite the growing ecosystem of ethical frameworks, regulatory approaches, and technical standards, none provides a unified, lifecycle-aligned method for embedding ethical requirements directly into routine AI development, deployment, and monitoring workflows [26,37]. This limitation highlights the need for governance approaches that translate high-level principles into practical, stage-specific actions, motivating the governance-by-design framework proposed in the subsequent section [36].

3.4. Challenges in Operationalizing Ethical Principles

Despite broad consensus on ethical principles, translating normative commitments into operational practice remains difficult [36,40]. A central issue is the level of abstraction. High-level principles such as fairness and transparency provide limited guidance on specific technical decisions, governance structures, or accountability mechanisms. As a result, developers and healthcare organizations committed to ethical AI often struggle to determine what concrete actions these principles require in practice [37]. For example, different definitions of fairness such as demographic parity, equalized odds, or predictive parity may lead to conflicting outcomes depending on the clinical context [18,50]. Similarly, the level and type of explainability required for clinical decision support remains unclear, particularly when balancing model complexity with interpretability [20,21]. Questions of accountability are equally challenging, especially when responsibility is distributed across developers, healthcare providers, vendors, and regulators [34]. These challenges highlight that ethical implementation is inherently context-dependent and requires trade-offs informed by clinical, technical, and societal considerations [15]. More broadly, prior work has emphasized that principle-based approaches alone are insufficient without operational tools, measurable criteria, and enforceable governance mechanisms [36,37].
Fragmented governance ecosystems further complicate operationalization [29]. Healthcare AI developers must navigate multiple overlapping frameworks, including WHO guidance, national regulatory authorities such as the FDA and EMA, regional legislation such as GDPR and the EU AI Act, professional society recommendations, and institutional policies [35,40]. While these frameworks share common goals, they differ in scope, terminology, and enforcement mechanisms, creating inconsistencies in implementation [34]. For instance, compliance with one regulatory framework does not necessarily ensure alignment with others, particularly in cross-border deployment scenarios. This fragmentation increases the burden on developers and healthcare institutions, often requiring duplication of validation, documentation, and reporting efforts [13]. It also creates gaps in accountability, where responsibilities may be unclear or distributed across multiple entities without clear coordination [37]. These challenges are further exacerbated by differences in national regulatory maturity and evolving legal interpretations of AI systems across jurisdictions [39].
Resource and capacity constraints present additional barriers, particularly in low-resource settings and smaller healthcare organizations [42]. Effective governance requires multidisciplinary expertise, including data scientists, clinicians, ethicists, legal specialists, and patient representatives [26]. It also depends on technical infrastructure to support data governance, validation, monitoring, auditing, and ongoing system maintenance [38]. In practice, many institutions lack the financial resources, technical capabilities, or organizational structures needed to sustain such efforts. This can lead to uneven adoption of governance practices and may limit the safe deployment of AI systems in under-resourced environments [14]. As a result, there is a risk that the benefits of healthcare AI become concentrated in well-resourced systems, potentially reinforcing existing disparities rather than reducing them [44]. Similar concerns have been raised in global health contexts, where uneven access to data, infrastructure, and expertise may further widen existing inequities in AI adoption and impact [33].
In addition, the dynamic nature of healthcare data introduces ongoing challenges for maintaining system performance and reliability [22]. Changes in patient populations, clinical practices, or disease patterns can lead to distribution shifts that degrade model performance over time [23]. Without continuous monitoring and recalibration, models that were initially validated may become less accurate or even unsafe in real-world settings [14]. This highlights the importance of post-deployment oversight mechanisms, which are often underdeveloped or inconsistently implemented across institutions [26]. Broader analyses of real-world AI deployment have similarly identified monitoring and lifecycle oversight as critical yet under-addressed components of trustworthy AI systems [13].
Tension between innovation and regulation creates further complexity [40]. AI technologies continue to evolve rapidly, with new model architectures, training approaches, and applications emerging across clinical domains [58]. Traditional regulatory frameworks, which are typically designed for static medical devices, are not always well suited to systems that learn and adapt over time [35]. Overly rigid regulation may slow innovation and delay the adoption of beneficial technologies, while insufficient oversight increases the risk of deploying inadequately validated or biased systems [34]. Addressing this tension requires governance approaches that are both flexible and robust, allowing adaptation to technological change while maintaining safety and ethical integrity [26]. Emerging regulatory models, including adaptive oversight and lifecycle-based risk management approaches, seek to address this balance by integrating continuous evaluation and iterative governance into regulatory processes [35].
Taken together, these challenges demonstrate that ethical AI governance cannot rely solely on high-level normative commitments. Effective operationalization requires lifecycle-specific, technically grounded, and context-sensitive governance mechanisms that can adapt over time while remaining aligned with core ethical principles [36,37].

3.5. Technical Foundations for Lifecycle-Aligned AI Governance

Operationalizing ethical principles requires technical methods and tools that translate normative requirements into concrete implementation practices [26,36]. Several technical domains provide foundational mechanisms for embedding ethical governance across the AI lifecycle [37]. These domains have been widely discussed across the literature as essential enablers of trustworthy and accountable AI systems in healthcare.
Data quality and provenance tracking are central to ensuring that AI systems learn from reliable and representative information [17,18]. Data quality includes accuracy, completeness, consistency, and timeliness relative to the intended clinical context. Poor data quality can directly compromise model performance and patient safety, particularly when biases in electronic health record data or labeling practices are not adequately addressed [44]. Provenance tracking documents data origins, collection processes, transformations, and lineage, enabling transparency, reproducibility, and auditability [33]. Technical mechanisms include metadata standards, data version control systems, structured documentation practices, and automated quality assessment tools capable of detecting anomalies prior to model training [38]. These practices are increasingly supported by emerging data governance frameworks and standards aimed at improving traceability and accountability across the AI lifecycle [45,59].
Algorithmic bias detection and mitigation techniques address fairness considerations during development [50]. Bias assessment relies on fairness metrics that quantify performance disparities across demographic subgroups [18]. Common metrics include demographic parity, equalized odds, and calibration within groups. These metrics are not mutually compatible; optimizing one may compromise another, requiring context-sensitive selection aligned with clinical objectives and stakeholder priorities [50]. Bias mitigation techniques operate at multiple stages. Pre-processing methods adjust training data distributions through resampling or reweighting [17]. In-processing approaches incorporate fairness constraints into the optimization objective. Post-processing methods adjust decision thresholds or recalibrate outputs after model training. Effective bias mitigation requires iterative evaluation and domain-specific judgment rather than reliance on a single metric or intervention, as emphasized in recent fairness-focused machine learning research [15].
Explainability and interpretability methods address the opacity of complex machine learning systems [21]. Intrinsically interpretable models such as decision trees or linear models provide transparent logic but may sacrifice predictive performance [46]. For complex models, including deep neural networks, post hoc explanation techniques aim to provide interpretable approximations of model behavior [48]. Model-agnostic approaches such as LIME and SHAP approximate local decision boundaries or quantify feature contributions [48]. Model-specific approaches, including attention mechanisms, identify influential input regions within neural architectures. The clinical utility of explanations depends not only on technical validity but also on their interpretability, contextual relevance, and usability in real clinical workflows [20]. At the same time, emerging literature highlights limitations of current explainability methods, including instability and potential misinterpretation, underscoring the need for careful evaluation of explanation quality in clinical contexts [21,47].
Privacy-preserving and secure learning techniques are increasingly important for protecting sensitive health data while enabling model development [38]. Approaches such as federated learning allow models to be trained across distributed datasets without centralizing patient data, thereby reducing privacy risks and supporting collaboration across institutions [45]. Secure multi-party computation and differential privacy methods further enhance data protection while maintaining analytical utility [32]. These approaches are particularly relevant in multi-institutional and cross-border healthcare settings where data sharing is constrained by regulatory requirements and data protection laws [33].
Robustness and security considerations are also essential components of ethical AI systems [60]. Machine learning models can be vulnerable to adversarial attacks or data perturbations that lead to incorrect predictions. Ensuring robustness requires stress testing models under varying conditions, monitoring for unexpected behavior, and implementing safeguards against malicious inputs. These considerations are critical for maintaining trust and safety in high-stakes clinical environments and are increasingly emphasized in safety-focused AI research [15].
These technical domains collectively support the operationalization of ethical principles [37]. Data quality and provenance tracking contribute to safety and equity. Bias detection and mitigation enable fairness assessment. Explainability techniques promote transparency and accountability. Privacy-preserving methods support autonomy and data protection. Robustness and security measures strengthen system reliability and safety.
However, technical tools alone are insufficient [36,37]. Effective governance requires institutional mechanisms that ensure these tools are appropriately selected, validated, and integrated into organizational workflows and decision-making processes [26]. Without such integration, even well-designed technical solutions may fail to achieve their intended ethical outcomes.

3.6. Why Governance-by-Design Is Required

Traditional AI governance approaches often treat ethics as a post-development consideration, applying oversight or review primarily at the deployment stage [36,37]. This reactive model has important limitations. Retrofitting ethical safeguards into completed systems is technically difficult because foundational design decisions, such as data sourcing, feature selection, model architecture, and optimization objectives, shape system behavior in ways that are not easily reversible. For example, bias mitigation applied after model development is often less effective than addressing representational imbalances during data collection or incorporating fairness constraints during training [50]. Similarly, inadequate data documentation or labeling practices introduced early in the pipeline can propagate downstream and remain difficult to detect or correct later [17]. More broadly, prior studies have shown that early-stage design choices have persistent effects across the AI lifecycle, reinforcing the need for proactive governance rather than retrospective correction [14].
Reactive governance may also create misaligned incentives [15]. When ethical evaluation is deferred to later stages, development processes may prioritize performance metrics such as accuracy or efficiency, while fairness, transparency, and safety considerations receive less attention. Late-stage interventions are typically more costly, less effective, and more difficult to implement. In addition, ethical risks may arise during development itself, including privacy violations, insufficient consent processes, or lack of transparency in data use and annotation practices [33]. These risks are often overlooked when governance is treated as an endpoint rather than an ongoing process. This limitation has been widely recognized in the literature, where reactive governance models are criticized for failing to address systemic risks that emerge throughout the AI lifecycle [36].
Governance-by-design offers an alternative paradigm centered on the proactive integration of ethical requirements across all stages of the AI lifecycle [26]. Drawing on established approaches such as privacy-by-design and safety-by-design, this model embeds ethical safeguards directly into system architecture and development workflows from project inception [32]. In practice, governance-by-design includes structured consent and data governance mechanisms during data collection; integration of fairness-aware learning and explainability methods during model development; rigorous, equity-focused validation strategies; preservation of human oversight and clinician-in-the-loop decision-making during deployment; and continuous monitoring to detect performance degradation, dataset shift, and disparate impact over time [40]. This lifecycle-aligned approach reflects broader shifts in AI governance toward continuous risk management and adaptive oversight rather than static evaluation [35].
Proactive governance provides several advantages [37]. It enables earlier identification of risks, when mitigation is technically feasible and less costly. It distributes accountability across lifecycle stages rather than concentrating responsibility at deployment [26]. It also promotes multidisciplinary collaboration among developers, clinicians, ethicists, and patient representatives, improving both system design and real-world usability [34]. Importantly, governance-by-design aligns with emerging regulatory expectations, including lifecycle documentation, risk management, and post-market surveillance requirements outlined in frameworks such as the EU AI Act and evolving medical device regulations [35]. It also helps address fragmentation across governance frameworks by providing a unifying structure for integrating ethical, technical, and regulatory requirements into a coherent development process [29].
Beyond compliance, governance-by-design supports the development of more robust and trustworthy systems. By embedding monitoring and feedback mechanisms, it enables continuous learning and adaptation in response to real-world conditions, addressing known challenges such as model drift and performance variability across populations [14]. In this sense, trustworthy AI becomes a system-level property that emerges from sustained governance practices rather than a feature added after development [28]. More broadly, this approach reframes ethical governance as enabling infrastructure that supports safe, scalable, and equitable AI deployment rather than as an external constraint on innovation [37].
The practical implications of governance-by-design compared to traditional reactive approaches are summarized in Table 5.

3.7. Global Data Protection Requirements for Healthcare AI

The European Union’s General Data Protection Regulation (GDPR), effective since 2018, establishes comprehensive data protection obligations with direct implications for healthcare AI [34]. GDPR governs the processing of personal data, including health data, which is classified as a special category requiring enhanced protection. Core principles include lawfulness, fairness, transparency, purpose limitation, data minimization, accuracy, storage limitation, integrity and confidentiality, and accountability. These principles form a foundational framework for data governance in AI-driven healthcare systems and have influenced global approaches to data protection [32,33].
For healthcare AI systems, lawful processing requires a recognized legal basis, most commonly explicit consent or legitimate interest in research contexts [34]. However, obtaining valid consent can be challenging in AI settings where data may be reused, models iteratively updated, or secondary analyses conducted beyond the original purpose of collection [32]. Research exemptions exist but require appropriate safeguards, ethical oversight, and proportionality in data use. Data subject rights include access, rectification, erasure, restriction of processing, portability, and objection. These rights introduce technical challenges when individual data contributions are embedded within trained models, making it difficult to isolate or remove specific data points after training [45]. These challenges highlight tensions between data protection requirements and the technical characteristics of machine learning systems, particularly in large-scale and continuously evolving models [38].
GDPR provisions addressing automated decision-making state that individuals have the right not to be subject solely to automated decisions that produce significant effects, unless specific conditions are met. In clinical contexts, this requirement reinforces the importance of meaningful human oversight in AI-assisted decision-making [28]. Additionally, data protection impact assessments (DPIAs) are mandated for high-risk processing activities, including large-scale health data use and automated decision systems. DPIAs provide structured mechanisms for identifying, assessing, and mitigating risks associated with AI deployment, supporting more accountable and transparent governance practices [26]. These requirements align with broader shifts toward lifecycle-based risk assessment and continuous oversight in AI governance [35].
Beyond Europe, global data protection frameworks increasingly reflect GDPR-inspired principles, although implementation varies across jurisdictions [29]. In the United States, the Health Insurance Portability and Accountability Act (HIPAA) governs health data privacy but applies only to specific covered entities and does not provide the same breadth of individual rights as GDPR [32]. Other jurisdictions, including Canada and several Asia-Pacific and Latin American countries, have adopted hybrid regulatory approaches that combine sector-specific protections with broader data governance principles. This diversity creates challenges for cross-border AI development and deployment, requiring organizations to navigate multiple regulatory regimes simultaneously and reconcile differing legal interpretations of data use and consent [39].
Emerging privacy-enhancing technologies provide technical mechanisms to support compliance with data protection requirements [38]. Federated learning enables collaborative model training across institutions without centralizing sensitive patient data, reducing exposure risks [45]. Differential privacy techniques introduce controlled noise into datasets or model outputs to limit the re-identification of individuals, while synthetic data generation offers alternative approaches for model development when access to real-world data is restricted [32]. These approaches are increasingly important for enabling innovation while maintaining compliance with evolving privacy standards and supporting data sharing in regulated environments [33].
Compliance with data protection requirements therefore necessitates integration across the entire AI lifecycle. This includes robust consent governance, purpose limitation safeguards, data minimization strategies, secure data storage and processing, and transparent communication regarding how data are used in AI systems [40]. These requirements reinforce the need for governance-by-design approaches, where privacy and data protection are embedded into system design and development processes rather than addressed through retrospective compliance efforts [36,37].

3.8. The EU AI Act: Risk-Based Regulatory Governance for Healthcare AI

The European Union Artificial Intelligence Act, adopted in 2024 with phased implementation extending through 2026 and 2027, establishes a comprehensive risk-based regulatory framework for AI systems. Under this framework, regulatory obligations scale according to the level of risk posed to safety and fundamental rights, creating a structured approach to governing diverse AI applications. This risk-based model reflects a broader shift toward proactive, lifecycle-oriented governance that emphasizes continuous risk management and accountability across the AI system lifecycle [35].
At the highest level, the Act prohibits AI systems that violate fundamental rights, such as social scoring by public authorities or systems that exploit vulnerable populations. Although most healthcare AI applications do not fall within this category, these prohibitions establish clear ethical boundaries that inform broader governance expectations.
Most clinical AI systems are classified as high-risk due to their potential impact on patient safety and clinical decision-making. This includes diagnostic support tools, triage systems, treatment recommendation models, and AI-enabled medical devices. High-risk classification imposes a comprehensive set of obligations, including lifecycle risk management systems, high-quality data governance practices emphasizing relevance and representativeness, detailed technical documentation, transparency requirements, human oversight mechanisms, accuracy and robustness standards, cybersecurity safeguards, conformity assessment procedures, and post-market monitoring requirements. These requirements closely align with lifecycle-based governance approaches and reinforce the need for continuous oversight throughout system development and deployment [14,26].
Conformity assessment procedures are required to verify compliance prior to market placement. For AI systems that qualify as medical devices, these requirements are integrated with existing regulatory pathways such as the Medical Device Regulation (MDR) and In Vitro Diagnostic Regulation (IVDR), creating a layered regulatory structure [34]. Following deployment, post-market monitoring obligations require continuous performance tracking, incident reporting, and corrective actions where necessary, reinforcing the importance of real-world evaluation and governance [35]. This reflects growing recognition that model performance may change over time due to dataset shift, evolving clinical practices, or differences in deployment environments [22,23].
The different risk categories defined by the EU AI Act, along with their relevance to healthcare AI systems and associated regulatory requirements, are summarized in Table 6.
The EU AI Act represents the most comprehensive AI-specific regulatory instrument to date and is likely to influence global governance practices due to its extraterritorial scope. Its requirements extend beyond European borders, affecting organizations that develop or deploy AI systems within the EU market. This global influence is already shaping regulatory discussions, policy alignment efforts, and emerging governance models in other regions [29].
Despite providing important regulatory clarity, implementation challenges remain [34]. These include ambiguity in risk classification boundaries, coordination with existing sector-specific regulations, and the need for specialized regulatory expertise to evaluate complex AI systems. In addition, fragmentation across regulatory frameworks and variation in institutional capacity may complicate consistent enforcement across jurisdictions, particularly as AI technologies continue to evolve rapidly [39]. Ensuring that regulatory systems remain adaptive while maintaining safety and accountability therefore remains an ongoing challenge [35].
Together, GDPR, sector-specific healthcare regulations, and the EU AI Act underscore the necessity of lifecycle-aligned governance structures capable of operationalizing ethical, legal, and technical obligations in an integrated and coherent manner [26,40]. These regulatory developments further reinforce the need for governance-by-design approaches that embed compliance and ethical safeguards throughout the AI lifecycle rather than treating them as external requirements [36,37].
The key cross-cutting governance enablers that support the integration of ethical, legal, and technical requirements across the AI lifecycle are illustrated in Figure 3.

4. Operationalizing WHO Principles Across the AI Lifecycle

This section presents the core contribution of this review: a governance-by-design framework that operationalizes WHO’s ethical principles across the healthcare AI lifecycle [26,40]. For each lifecycle stage, relevant ethical principles, concrete governance mechanisms, and illustrative failure modes are identified to demonstrate how ethical alignment can be embedded within technical and organizational practices rather than treated as a post hoc compliance exercise. Unlike existing ethical guidance, which often presents principles at a conceptual level, this framework explicitly maps ethical requirements to concrete governance mechanisms across each stage of the healthcare AI lifecycle [36]. This lifecycle-oriented perspective is further supported by prior studies demonstrating that early design and governance decisions have persistent effects across system performance, safety, and fairness outcomes throughout the AI lifecycle [14].
This lifecycle-aligned approach builds on prior work emphasizing that ethical principles alone are insufficient without implementation mechanisms embedded in real-world systems [36]. It also aligns with emerging perspectives that view ethical AI as a socio-technical system requiring coordination across technical design, clinical workflows, and institutional governance structures [37]. In healthcare contexts, where decisions directly impact patient outcomes, such integration is particularly critical for ensuring both safety and trustworthiness [37].
The proposed governance-by-design framework, which integrates WHO ethical principles with lifecycle-specific governance mechanisms, is illustrated in Figure 4.
This conceptual framework depicts the integration of WHO’s six ethical principles across five healthcare AI lifecycle stages. The diagram is structured as a matrix with lifecycle stages (Data Collection, Model Development, Validation, Deployment, Monitoring) arranged horizontally and WHO principles (Autonomy; Well-being and Safety; Transparency; Accountability; Equity and Inclusiveness; Sustainability) arranged vertically. Each cell identifies specific governance mechanisms corresponding to the intersection of principle and lifecycle stage. Darker shading indicates stages of primary ethical salience.
Bidirectional arrows connect lifecycle stages to illustrate iterative feedback loops, emphasizing that governance is cyclical rather than linear. This reflects real-world system behavior, where model performance, data distributions, and clinical use patterns evolve over time, requiring continuous reassessment and adaptation [26]. Real-world deployment studies further confirm that model performance and reliability can vary significantly across clinical environments, reinforcing the need for continuous lifecycle governance rather than static evaluation approaches. Surrounding the matrix are four external influence layers: regulatory frameworks (GDPR, EU AI Act, FDA GMLP), reporting standards (TRIPOD-AI, SPIRIT-AI, CONSORT-AI), technical methods (bias detection, explainability tools, drift monitoring), and stakeholder engagement (clinicians, patients, ethicists, regulators). The inclusion of these layers highlights that governance is not confined to model development but is shaped by broader institutional and regulatory ecosystems [22,23].
Importantly, recent advances in generative AI and large language models further reinforce the need for lifecycle-integrated governance [58]. These systems introduce additional challenges related to hallucination, lack of determinism, and evolving capabilities, making post hoc evaluation insufficient for ensuring safety and reliability. Similarly, foundation models trained on large-scale heterogeneous datasets may exhibit unpredictable behavior across clinical contexts, requiring robust validation and continuous monitoring mechanisms [61]. These developments underscore the importance of embedding governance controls throughout the lifecycle rather than relying on static validation approaches. Emerging studies on generative AI systems further highlight risks related to hallucination, instability, and context sensitivity, emphasizing the need for continuous monitoring, validation, and human oversight mechanisms [62].
In addition, continuous monitoring and feedback mechanisms are essential for identifying performance degradation, model drift, and emerging risks during real-world use [14].
A detailed mapping of WHO ethical principles to lifecycle-specific governance mechanisms is provided in Table 7.

4.1. Data Collection and Curation

Data collection forms the foundation of healthcare AI systems, shaping downstream performance, safety, and fairness outcomes [17,18]. Decisions made at this stage influence not only model accuracy but also the ethical integrity of the system. Early-stage data choices often have persistent effects across the AI lifecycle, reinforcing the importance of proactive governance at the point of data acquisition and curation [14]. Effective governance must therefore ensure that data practices respect patient autonomy, maintain high data quality, support transparency, establish clear accountability, promote equity through representative sampling, and adopt sustainable stewardship approaches.
Respect for autonomy begins with meaningful informed consent, particularly when data are used for purposes beyond routine clinical care [32]. Consent processes should clearly communicate secondary uses of data, potential commercial involvement, and future model development objectives. In practice, this is challenging because AI systems often evolve over time. Dynamic consent models offer a more flexible approach, allowing patients to update their preferences as new uses emerge [33]. At the same time, privacy-preserving techniques such as federated learning and secure multi-party computation provide technical mechanisms for reducing privacy risks by enabling model training without centralizing sensitive data [38,45]. These approaches are increasingly important in multi-institutional settings where data sharing is constrained by regulatory and governance requirements.
Failures in this area demonstrate the consequences of inadequate governance. In the widely discussed DeepMind–Royal Free Hospital collaboration, patient data were used to develop an acute kidney injury detection system without sufficient transparency or appropriate legal basis [32]. Regulatory review highlighted gaps in consent, oversight, and communication, reinforcing the need for structured data governance processes, including clear legal justification and formal data protection impact assessments. Similar real-world cases have shown that weaknesses in data governance can propagate downstream risks that are difficult to correct during later lifecycle stages [23].
Ensuring well-being and safety depends heavily on data quality and representativeness [18]. Data used for model development must be accurate, complete, and clinically meaningful. Annotation processes should reflect real clinical conditions, and systematic quality assurance procedures are needed to detect errors or inconsistencies before training [13]. Without these safeguards, flawed data can propagate through the pipeline and result in unreliable or unsafe predictions. In addition, data derived from electronic health records may reflect historical biases in clinical practice, further emphasizing the need for careful evaluation and curation [44,50]. Emerging data governance standards and quality frameworks aim to address these challenges by improving traceability and consistency across datasets [45].
Equity considerations require deliberate attention to representation [18]. Training data should reflect the diversity of patient populations in which models will be deployed. Without proactive inclusion strategies, AI systems risk underperforming for marginalized or underrepresented groups, reinforcing existing disparities in care [44,50]. This challenge is particularly pronounced in low- and middle-income settings, where limited data availability may restrict model generalizability and widen global inequities in access to AI-enabled healthcare [33]. Broader global health analyses further highlight the importance of inclusive data practices to prevent the amplification of structural inequities through AI systems.
Transparency at this stage is supported through structured dataset documentation [34]. Approaches such as datasheets for datasets provide detailed descriptions of data sources, collection methods, preprocessing steps, intended uses, and known limitations. Complementing this, provenance tracking systems record data lineage and transformations, enabling reproducibility, auditability, and retrospective evaluation of data-related decisions [38].
Accountability requires clearly defined roles and governance structures [34]. Data stewardship responsibilities should be explicitly assigned, with institutional review boards overseeing data acquisition and governance committees regulating access and secondary use. These mechanisms ensure that ethical and legal responsibilities are not diffuse but embedded within organizational processes.
Finally, sustainability emphasizes responsible data stewardship [42]. Rather than continuously collecting new data, organizations should prioritize the reuse of high-quality existing datasets where appropriate, while ensuring continued relevance and ethical compliance. Minimizing redundant data collection reduces resource burden and supports more efficient, scalable AI development [14]. Sustainable data practices contribute not only to technical performance but also to long-term trust and resilience within healthcare systems.

4.2. Model Development and Training

Model development transforms curated data into algorithmic decision-making systems, embedding choices related to feature selection, model architecture, and optimization objectives that inherently reflect clinical priorities and value judgments [14,15]. Decisions made at this stage are not purely technical; they shape how the system behaves in practice, influencing safety, fairness, and interpretability. Design decisions made during development often have persistent downstream effects, reinforcing the need for early integration of governance mechanisms. As a result, governance during model development must be understood as both a technical and normative process.
Respect for autonomy is supported through structured human-in-the-loop development processes that incorporate clinical expertise into model design [20]. Clinicians play a critical role in guiding feature selection, ensuring that input variables reflect clinically meaningful constructs rather than spurious correlations or proxy variables that may encode bias [18]. Designing systems with the expectation of clinician oversight, including the ability to override model recommendations, helps preserve human agency and reduces the risk of automation bias [25]. These practices also align with broader efforts to maintain human-centered AI systems in safety-critical environments [12].
Ensuring well-being and safety requires embedding safeguards directly into the training process [15]. This includes incorporating safety constraints into optimization objectives, performing robustness testing under varying input conditions, and evaluating model performance across diverse and clinically relevant scenarios [60]. In addition, uncertainty quantification methods such as calibrated probabilities or confidence intervals should accompany predictions to communicate the level of reliability and reduce overreliance on deterministic outputs. Without such measures, highly accurate models may still produce misleading or unsafe recommendations in edge cases or unfamiliar settings [22]. Robustness and stress testing are increasingly emphasized in safety-focused AI research, particularly for high-stakes domains such as healthcare [60].
Transparency at this stage is supported through structured model documentation. Model cards provide standardized descriptions of intended clinical use, training data characteristics, subgroup performance, ethical considerations, and known limitations. For more complex models, explainability techniques such as SHAP and LIME offer local approximations of model behavior and feature contributions, supporting interpretability [48]. However, the value of these explanations depends on their clinical relevance and usability. Explanations that are technically accurate but not meaningful to clinicians may fail to improve trust or decision-making [20]. Emerging research also highlights limitations of current explainability methods, including instability and potential misinterpretation, underscoring the need for careful evaluation in clinical contexts [21,47].
Accountability relies on rigorous software engineering and quality assurance practices [26]. Version control systems, reproducible training pipelines, code review processes, and automated testing frameworks are essential for ensuring reliability and traceability. In regulated healthcare environments, these practices must align with quality management systems used for medical device development [34]. Maintaining detailed records of model iterations, training datasets, and performance evaluations enables auditability and supports regulatory compliance.
Equity considerations require fairness-aware machine learning approaches integrated throughout development [50]. Bias detection and mitigation should not be treated as isolated steps but as ongoing processes [17]. Pre-processing methods address imbalances in training data through resampling or reweighting. In-processing approaches incorporate fairness constraints directly into model optimization. Post-processing techniques adjust outputs to reduce disparities across groups. Because fairness metrics are often mathematically incompatible, selecting an appropriate metric requires context-sensitive judgment informed by clinical goals and stakeholder priorities [50]. Governance processes should explicitly document these choices to ensure transparency and accountability. These approaches are supported by a growing body of fairness-focused machine learning research emphasizing the need for context-aware evaluation and iterative mitigation strategies.
Recent advances in machine learning, particularly the use of large language models and foundation models, introduce additional considerations at the development stage. These models are often trained on large, heterogeneous datasets and may exhibit emergent behaviors that are difficult to predict or fully control [63,64]. Their deployment in healthcare contexts requires careful evaluation of reliability, bias, and generalizability, as well as safeguards against issues such as hallucinated outputs or inconsistent reasoning [65,66]. These characteristics further reinforce the need for structured governance during model development, particularly for systems that evolve over time or operate across diverse clinical settings [63,64,67,68,69,70,71,72,73,74].
In addition, generative AI systems may exhibit non-deterministic behavior, producing variable outputs in response to similar clinical prompts or contextual inputs. Such variability complicates reproducibility, reliability assessment, and regulatory evaluation in high-stakes healthcare environments. Hallucinated or fabricated outputs may further introduce significant patient safety risks when inaccurate information is presented with high confidence [62]. These limitations reinforce the importance of maintaining meaningful human oversight, particularly for clinical decision support applications involving diagnosis, treatment planning, or patient communication. Continuous post-deployment monitoring, structured audit mechanisms, and periodic reevaluation are therefore essential to identify emerging risks, assess system reliability over time, and support safe integration of generative AI systems into clinical workflows [55,63,75].
Sustainability emphasizes resource-conscious development strategies. Techniques such as model compression, pruning, quantization, and knowledge distillation can reduce computational demands while maintaining performance [42]. Transfer learning approaches, which adapt pre-trained models to new clinical tasks, can further improve efficiency and reduce the need for large-scale data collection and training [14]. These strategies not only reduce environmental impact but also support more scalable and accessible AI development across diverse healthcare settings.
These practices also align with emerging healthcare MLOps approaches that emphasize reproducible model development, automated validation workflows, continuous integration, and traceable lifecycle management for clinical AI systems [55,56,57].

4.3. Validation and Clinical Evaluation

Validation determines whether AI systems are sufficiently robust, clinically meaningful, and safe for real-world deployment [13]. It serves as a critical ethical checkpoint where unsafe, biased, or poorly generalizable systems must be identified and addressed before integration into clinical practice [15]. Importantly, strong performance during model development does not guarantee effectiveness in real-world settings, where differences in patient populations, clinical workflows, and data distributions can significantly impact performance [22]. This gap between development and deployment has been widely documented and underscores the need for rigorous, multi-stage validation strategies [14].
Respect for autonomy requires that patients participating in AI-supported clinical studies are provided with clear and meaningful informed consent [32]. This includes transparent explanations of the algorithm’s role in clinical decision-making, potential benefits and risks, data usage, and the right to withdraw. In practice, communicating these aspects can be challenging due to the complexity of AI systems, making clarity and accessibility in consent processes especially important for maintaining patient trust [20]. These challenges are further amplified in AI-driven studies where system behavior may not be fully interpretable or predictable.
Ensuring well-being and safety requires evaluation that extends beyond traditional performance metrics such as accuracy or area under the curve [52]. Validation should prioritize clinically meaningful outcomes, including diagnostic impact, treatment decisions, and patient outcomes [13]. Internal validation provides an initial assessment using held-out datasets, but external validation across independent institutions, populations, and care settings is essential for assessing generalizability [22]. Prospective clinical studies, including randomized controlled trials, offer stronger evidence by comparing AI-assisted care with standard clinical practice [53]. Evidence from real-world deployments shows that models can experience significant performance degradation outside their original development environment, highlighting the risks of overfitting and dataset shift [23]. These findings reinforce the importance of continuous evaluation and post-deployment monitoring as part of a broader lifecycle governance strategy [14].
Transparency at this stage is supported through adherence to established reporting standards such as TRIPOD-AI for prediction models and CONSORT-AI for randomized controlled trials [21,22]. These frameworks promote consistent and comprehensive reporting of study design, data characteristics, and performance outcomes. In particular, disaggregated reporting across demographic and clinical subgroups is essential for identifying disparities that may be masked in aggregate metrics [30]. Without such transparency, systems that appear effective overall may still produce inequitable outcomes in practice. Increasingly, transparency is also linked to reproducibility and independent verification of results, which remain ongoing challenges in healthcare AI research [34].
Accountability requires a clear separation between model development and evaluation processes. Independent validation, including third-party assessment where feasible, reduces bias and strengthens confidence in reported performance [26]. Regulatory review processes establish minimum safety and performance thresholds, while institutional ethics committees provide additional oversight [34]. Detailed documentation of validation methods and outcomes ensures traceability and supports both regulatory compliance and clinical accountability.
Equity considerations demand subgroup analyses with sufficient statistical power to detect meaningful differences in performance [18]. Validation across diverse healthcare systems, geographic regions, and resource settings improves generalizability and reduces the risk of inequitable outcomes [14]. Equity impact assessments can further help determine whether a system is likely to mitigate, maintain, or exacerbate existing disparities in care delivery [50]. These assessments are particularly important when models are deployed in populations that differ from those represented in training data. Broader evidence from healthcare AI studies highlights that failure to evaluate subgroup performance can lead to unintended harm in vulnerable populations [44].
Sustainability in validation emphasizes the development of reusable and scalable evaluation infrastructure [45]. Standardized benchmark datasets, shared test sets, and federated validation frameworks enable multi-institutional evaluation without requiring centralized data sharing [38]. These approaches reduce duplication of effort and support continuous evaluation as models evolve over time. Establishing shared validation practices also improves comparability across studies and contributes to a more robust and reliable evidence base for healthcare AI systems [13]. Such infrastructure is increasingly recognized as essential for supporting long-term governance and lifecycle monitoring of AI systems in healthcare [26].

4.4. Deployment and Clinical Integration

Deployment translates validated AI systems into real-world clinical workflows, introducing significant sociotechnical complexity related to usability, workflow integration, and the preservation of ethical safeguards in practice [11,14]. At this stage, technical performance alone is insufficient. AI systems must function effectively within dynamic clinical environments, where time constraints, cognitive load, and organizational processes shape how decisions are made. Governance mechanisms must therefore ensure that AI supports, rather than disrupts, clinical reasoning and patient care. Real-world evidence increasingly shows that systems performing well in controlled settings may behave differently in clinical practice, reinforcing the importance of context-aware deployment strategies [23].
Respect for autonomy requires the preservation of meaningful clinician oversight [12]. AI systems should support, rather than replace, clinical judgment by enabling override capabilities and providing sufficient contextual information for independent evaluation [25]. Training and education are also essential to help clinicians understand system limitations and reduce the risk of automation bias, where users may over-rely on algorithmic outputs [20]. Patient autonomy similarly requires transparency regarding AI involvement in care decisions, along with clear communication about how recommendations are generated. Where feasible, patients should have the option to request human-led decision pathways, particularly in high-stakes or sensitive clinical contexts [28].
Ensuring well-being and safety depends heavily on how AI systems are integrated into clinical workflows [10]. Poor integration can increase cognitive burden, contribute to alert fatigue, or lead to misinterpretation of outputs [14]. Systems should be designed to align with existing clinical processes and decision pathways rather than introducing unnecessary complexity. In addition, fail-safe mechanisms are essential to ensure safe operation in the event of system errors or downtime [15]. Structured incident reporting and response protocols allow healthcare organizations to identify, investigate, and learn from AI-related adverse events, supporting continuous improvement and risk mitigation [26]. Real-world studies have shown that even high-performing models can introduce unintended safety risks when deployed without adequate workflow alignment or user training [23]. These findings highlight the importance of ongoing monitoring and adaptive governance after deployment [35].
Transparency at deployment requires communication that is tailored to different stakeholders. For clinicians, explanations should highlight clinically relevant features influencing predictions in familiar and interpretable terms [20]. For patients, communication should focus on clarity and accessibility, explaining the role of AI in care decisions without requiring technical expertise. At the system level, transparency includes comprehensive user training, accessible documentation, and clear notification when algorithms are updated, recalibrated, or modified over time [34]. Such practices are particularly important in adaptive systems, where model behavior may evolve after deployment and require continuous oversight [35].
Accountability becomes distributed across multiple stakeholders in deployment settings [37]. Developers are responsible for algorithm design, validation, and documentation. Vendors manage system integration, updates, and technical support. Healthcare organizations oversee deployment decisions, clinical governance, and risk management. Clinicians retain responsibility for interpreting outputs and making final decisions in patient care. Clear governance structures, including service-level agreements, liability frameworks, and institutional oversight committees, are essential to define roles and prevent ambiguity in responsibility allocation [26]. Without such clarity, accountability gaps may emerge, particularly in complex, multi-actor systems, as highlighted in broader governance analyses [36].
Equity requires deliberate and inclusive deployment strategies [18]. AI systems should not be limited to well-resourced institutions but should be adapted for use in diverse healthcare settings, including underserved and resource-constrained environments. Real-time fairness monitoring systems can track disaggregated performance across demographic groups and trigger alerts when disparities exceed acceptable thresholds [50]. In addition, engaging patients and communities in deployment decisions can help ensure that systems are culturally appropriate and responsive to local needs, improving both trust and effectiveness [33]. Without such efforts, deployment may inadvertently reinforce existing inequities in healthcare access and outcomes [50].
Sustainability at deployment involves ensuring that systems remain functional, maintainable, and adaptable over time [45]. This includes scalable infrastructure, interoperability with electronic health record systems and other health IT platforms, and clearly defined maintenance and update strategies [45]. Continuous monitoring for model drift, performance degradation, and changing clinical conditions is essential to maintain long-term reliability [22]. Sustainable deployment also requires organizational commitment, including training, resource allocation, and governance structures that support ongoing system evaluation and adaptation. These considerations ensure that AI systems remain safe, effective, and aligned with clinical needs over time.

4.5. Monitoring and Continuous Learning

Monitoring is essential for sustaining the safety, effectiveness, and ethical alignment of AI systems over time, particularly as clinical environments, patient populations, and care practices evolve [13,14]. Although often underemphasized, this stage is critical for maintaining trustworthiness beyond initial validation and deployment [27]. AI systems are not static; their performance can change as underlying data distributions shift, making continuous oversight a core component of responsible governance [22]. Increasingly, monitoring is viewed as analogous to post-market surveillance in clinical medicine, emphasizing the need for ongoing evaluation throughout the system lifecycle [76].
Respect for autonomy is reinforced through structured feedback mechanisms that allow clinicians to report concerns, identify unintended consequences, and document override patterns [25]. These inputs provide valuable real-world insight into system behavior and usability. Similarly, patients should have access to clear and accessible grievance procedures related to AI-influenced care. Incorporating these feedback channels into governance processes ensures that lived experience informs ongoing system evaluation and improvement. Such participatory approaches also align with broader efforts to promote human-centered and accountable AI systems [28].
Ensuring well-being and safety requires continuous performance monitoring in real-world settings [52]. Monitoring systems should track predictive accuracy, calibration, and clinically meaningful outcomes over time, rather than relying solely on static validation results. Drift detection mechanisms compare current data distributions and outcome relationships with those observed during training to identify performance degradation or context shift [22]. In addition, structured incident investigation processes should be established to analyze AI-related harms using formal root cause analysis methods. Predefined retraining, recalibration, or rollback protocols are necessary to guide appropriate responses when performance declines or risks are identified [60]. Without such mechanisms, degradation may go unnoticed and compromise patient safety. Real-world evidence has shown that performance degradation and unintended effects may only become apparent after deployment, reinforcing the importance of continuous monitoring [23].
Transparency at this stage is supported through accessible performance dashboards that present current metrics, trends over time, and subgroup-specific outcomes to clinicians, administrators, and governance bodies [48]. Public reporting mechanisms, such as algorithm registries or institutional reporting systems, can further enhance accountability and foster trust by making AI performance visible beyond internal stakeholders [46]. Transparency is particularly important for adaptive systems, where model behavior may change over time without explicit user awareness. These practices also contribute to reproducibility and enable independent oversight of AI system performance [47].
Accountability requires sustained oversight through both internal and external governance mechanisms [77]. This includes compliance with post-market surveillance requirements, periodic audits, and re-certification processes where applicable [43]. Within healthcare organizations, AI performance monitoring should be integrated into existing quality improvement and clinical governance frameworks, ensuring that responsibility for oversight is clearly defined and continuously maintained [43,78]. These mechanisms help address known challenges related to distributed accountability in complex AI ecosystems [36].
Equity considerations extend beyond initial validation and require longitudinal monitoring of subgroup performance [19,78]. Disparities may emerge over time due to shifts in patient populations, care delivery practices, or data availability [16]. Continuous evaluation enables early detection of such disparities. Adaptive fairness interventions, such as recalibration, dynamic threshold adjustment, targeted retraining, or modification of input features, can help address inequities when they arise [50]. Governance structures should define thresholds and response strategies in advance to support consistent and proactive intervention. These approaches are supported by growing evidence highlighting the need for ongoing fairness evaluation in real-world AI systems [19,50].
Sustainability emphasizes responsible model evolution over time. Continuous updating must balance the benefits of incorporating new data with the operational, computational, and environmental costs of retraining [42]. Incremental learning approaches, transfer learning, and scheduled model review cycles can support system adaptation while minimizing disruption to clinical workflows [79,80]. Establishing clear update policies and version control practices ensures that model changes are traceable, controlled, and aligned with clinical needs [45]. These practices are increasingly recognized as essential components of long-term lifecycle governance and adaptive AI systems [35].

4.6. Applied Case Study: Operationalizing Governance-by-Design in a Healthcare Risk Prediction System

To demonstrate the practical implementation of the proposed governance-by-design framework, this section presents an illustrative applied case study based on a healthcare risk prediction system inspired by the widely discussed Optum population health management algorithm [81]. The original Optum case has been extensively cited in the healthcare AI literature as an example of how algorithmic bias can emerge through the use of healthcare spending as a proxy for healthcare need, thereby contributing to systematic disparities in care allocation [81].
The present case study is not intended to reconstruct the proprietary Optum system directly. Rather, it uses a representative healthcare machine learning workflow to illustrate how lifecycle-aligned governance mechanisms could be operationalized in practice across data collection, model development, validation, deployment, and post-deployment monitoring. The case therefore serves as an applied governance example demonstrating how WHO ethical principles may be embedded within real-world healthcare AI development and deployment processes.

4.6.1. System Overview and ML Pipeline

The representative healthcare AI system considered in this case study is a supervised machine learning model developed to predict future patient risk and support enrolment decisions for care management programs. Similar population health management systems are commonly used to identify patients who may benefit from additional clinical monitoring, preventive interventions, or coordinated care services. The representative workflow described in this case study uses structured electronic health record data, laboratory results, medication history, prior hospital admissions, chronic disease indicators, and healthcare utilization features to generate patient-level risk scores intended to support resource allocation and early intervention planning.
The operational machine learning pipeline consists of five lifecycle stages aligned with the proposed governance-by-design framework: (1) data collection and curation, (2) model development and training, (3) validation and clinical evaluation, (4) deployment and clinical integration, and (5) monitoring and continuous learning. During the data collection stage, patient data are aggregated from multiple clinical and administrative sources, including electronic health records, laboratory systems, claims databases, and population health registries. The model development stage includes data preprocessing, feature engineering, model training, performance optimization, and explainability assessment. Validation involves both internal and external evaluation across independent patient populations and healthcare settings, including subgroup-specific performance analysis. During deployment, the system is integrated into clinical decision support workflows used by clinicians and care management teams. Post-deployment monitoring includes continuous performance surveillance, fairness monitoring, drift detection, and periodic model recalibration to support ongoing safety, reliability, and accountability throughout the AI lifecycle [55,56,57,75].

4.6.2. Governance-by-Design Across the ML Lifecycle

Application During Data Collection and Curation
Within the proposed governance-by-design framework, governance mechanisms are embedded directly into the data collection and preprocessing pipeline to address risks related to data quality, representativeness, privacy protection, and proxy bias. In the original Optum case, healthcare spending was used as a proxy for healthcare need. However, because historical healthcare expenditures may reflect structural disparities in access to care, the use of cost-based proxy variables contributed to systematic underestimation of illness burden among Black patients [81].
Under a lifecycle-aligned governance approach, this risk would be addressed through multidisciplinary review processes involving data scientists, clinicians, health equity specialists, institutional review boards, and compliance teams prior to model development. Dataset audits and feature selection reviews would evaluate whether selected variables may indirectly encode socioeconomic or racial disparities. Data provenance documentation, subgroup representation analysis, and fairness-oriented preprocessing assessments would also be integrated into the data engineering workflow to identify potential sources of historical bias before model training occurs [57,75].
Governance responsibilities at this stage would therefore extend beyond technical preprocessing tasks and include institutional accountability mechanisms. Data scientists would oversee preprocessing and feature engineering activities, clinicians would evaluate clinical relevance and contextual appropriateness, health equity specialists would assess representation and fairness risks, and institutional governance bodies would provide ethical and regulatory oversight. Embedding these controls directly into the data engineering pipeline supports proactive governance by identifying high-risk design decisions before downstream deployment and clinical integration.
Application During Model Development and Training
During model development and training, the proposed governance-by-design framework embeds governance mechanisms directly into model optimization, evaluation, and documentation workflows. Rather than optimizing exclusively for predictive accuracy, the development process incorporates fairness assessment, explainability analysis, robustness testing, and human oversight considerations throughout iterative model refinement. This approach reflects the understanding that healthcare AI systems must balance predictive performance with broader ethical and clinical safety objectives.
Within the representative workflow, governance processes would include subgroup-specific performance evaluation, fairness metric assessment, explainability testing, and reproducibility documentation prior to deployment approval. For example, model performance would be evaluated across demographic subgroups to identify potential disparities in calibration, sensitivity, or false negative rates that could contribute to unequal care allocation. Explainability methods, such as feature attribution analyses, would also be incorporated to support clinician interpretability and facilitate auditing processes. In addition, version-controlled training pipelines, model documentation, and model cards would support traceability, accountability, and regulatory review [55,56,57].
This stage also requires explicit management of trade-offs between fairness, transparency, safety, and predictive performance. For example, fairness constraints designed to reduce subgroup disparities may modestly reduce aggregate predictive accuracy, while more interpretable models may provide greater clinical transparency at the cost of reduced model complexity. Under the proposed framework, such trade-offs would not be managed solely by technical teams but through multidisciplinary governance structures involving machine learning engineers, clinicians, institutional leadership, and AI governance committees. These stakeholders would collaboratively establish acceptable performance thresholds and determine whether the system satisfies predefined ethical, clinical, and operational requirements prior to validation and deployment.
Application During Validation and Clinical Evaluation
The proposed governance-by-design framework emphasizes independent validation and subgroup-specific clinical evaluation prior to deployment. In the retrospective Optum case, disparities became apparent only after external analysis demonstrated that Black patients with equivalent risk scores exhibited substantially greater illness burden than White patients [81]. This highlights the importance of validation processes that extend beyond aggregate performance metrics and explicitly evaluate fairness, calibration, and clinical reliability across diverse patient populations.
Within the representative workflow, validation procedures would include internal and external validation across multiple healthcare settings, subgroup-specific performance analysis, calibration testing, prospective workflow simulation, and clinician usability assessment. Governance mechanisms at this stage would evaluate whether model performance remains clinically reliable across demographic groups and whether deployment may introduce unintended disparities in care allocation or access to intervention programs. Validation protocols would also assess whether clinicians can appropriately interpret and act upon model outputs within real-world decision-making environments.
Governance responsibilities during validation would involve collaboration among machine learning teams, clinical evaluators, institutional governance committees, and healthcare quality and safety offices. Importantly, deployment approval would depend not only on predictive performance but also on satisfaction of predefined governance thresholds related to fairness, explainability, safety, and clinical usability. Under this approach, substantial subgroup disparities or clinically unsafe performance outcomes would trigger remediation, recalibration, or redesign before deployment authorization. Embedding governance mechanisms into validation workflows therefore supports proactive identification of sociotechnical risks prior to large-scale clinical integration [57,75].
Application During Deployment and Clinical Integration
During deployment and clinical integration, the governance-by-design framework embeds accountability, human oversight, and operational safety mechanisms directly into clinical workflows. Within the representative healthcare risk prediction system, model outputs would be integrated into clinical decision support interfaces used by clinicians and care management teams responsible for patient outreach, risk stratification, and intervention planning. Governance at this stage therefore extends beyond technical deployment and includes organizational processes that shape how AI-generated recommendations are interpreted and applied in practice.
To reduce automation bias and overreliance on algorithmic recommendations, clinicians would retain decision-making authority and the ability to override model outputs when clinically appropriate. Contextual explanations and supporting clinical information would accompany risk predictions to improve interpretability and support informed decision-making. In addition, governance mechanisms would include deployment approval checkpoints, audit logging, incident reporting procedures, user training programs, and clearly defined escalation pathways for adverse events or unexpected system behavior. These operational safeguards help ensure that AI recommendations remain subject to clinical judgment and institutional accountability throughout deployment.
Governance responsibilities during deployment would involve coordination among healthcare organizations, clinical leadership, MLOps teams, compliance offices, and AI governance committees. MLOps teams would oversee deployment infrastructure, version control, and operational monitoring systems, while clinicians and institutional leadership would evaluate workflow integration, usability, and patient safety considerations [55,56]. Governance committees would additionally review whether deployment conditions remain aligned with ethical and regulatory requirements established during earlier lifecycle stages. Embedding governance mechanisms directly into deployment workflows therefore supports continuous oversight and helps ensure that operational implementation remains aligned with fairness, safety, transparency, and accountability objectives.
Application During Monitoring and Continuous Learning
Continuous monitoring and post-deployment oversight are essential because healthcare AI systems may experience performance degradation over time due to population shifts, evolving clinical practices, changes in healthcare utilization patterns, or modifications to underlying data infrastructure [55,75]. The retrospective Optum case illustrates how governance failures may remain undetected when monitoring processes do not adequately evaluate subgroup disparities or long-term system impacts [81]. Accordingly, the proposed governance-by-design framework treats monitoring as a continuous operational responsibility rather than a one-time validation activity performed prior to deployment.
Within the representative workflow, monitoring infrastructure would include real-time performance surveillance, subgroup fairness monitoring, drift detection systems, periodic recalibration procedures, and adverse event reporting mechanisms [55,56,75]. Governance processes would continuously evaluate whether model outputs remain clinically reliable and equitable across patient populations. For example, fairness monitoring systems could assess subgroup calibration, false negative rates, or intervention allocation disparities and automatically trigger escalation procedures when predefined thresholds are exceeded. Drift detection mechanisms would similarly evaluate whether changes in patient populations or healthcare practices are affecting model reliability over time.
Governance responsibilities during this stage would involve collaboration among MLOps teams, healthcare quality and safety offices, clinicians, institutional governance committees, and regulatory oversight bodies [55,56]. MLOps teams would oversee technical monitoring infrastructure and retraining workflows, while healthcare organizations and governance committees would evaluate broader clinical, ethical, and operational implications of system behavior [55,56]. Importantly, continuous monitoring would also support retrospective governance analysis by enabling institutions to identify previously unrecognized harms, reassess governance assumptions, and implement corrective interventions when necessary. Embedding monitoring and adaptive oversight into post-deployment workflows therefore supports long-term accountability and continuous alignment between AI system performance and ethical healthcare objectives.

4.6.3. Integration into Real-World AI Development and Deployment Workflows

A central objective of the proposed governance-by-design framework is to integrate governance mechanisms directly into operational AI development and deployment workflows rather than treating governance as a separate or purely retrospective review process [55,56,57]. In practice, this requires embedding ethical, technical, clinical, and regulatory oversight mechanisms into existing healthcare AI infrastructure, including data engineering pipelines, model development environments, validation procedures, deployment orchestration systems, and post-deployment monitoring platforms.
Within the representative workflow presented in this case study, governance mechanisms are integrated across multiple operational layers. Dataset documentation and subgroup representation assessments are incorporated into data engineering workflows prior to model training. Fairness testing, explainability analysis, and reproducibility documentation are embedded into model development and validation pipelines. Deployment approval checkpoints, audit logging systems, and clinician-in-the-loop review processes are integrated into clinical implementation workflows. Similarly, drift detection systems, subgroup fairness monitoring, and incident escalation procedures are incorporated into post-deployment monitoring infrastructure to support ongoing oversight and accountability [55,56].
This operational integration also clarifies stakeholder responsibilities across the AI lifecycle. Data scientists and MLOps teams are responsible for technical implementation and monitoring processes, clinicians and healthcare quality teams oversee clinical appropriateness and patient safety considerations, institutional governance committees evaluate ethical and regulatory compliance, and organizational leadership establishes acceptable thresholds for balancing fairness, transparency, safety, and predictive performance. Embedding governance responsibilities within existing organizational and technical workflows therefore supports continuous accountability while reducing the risk that ethical oversight becomes disconnected from real-world AI deployment practices. Table 8 summarizes the integration of governance-by-design mechanisms across the representative healthcare AI lifecycle, including governance controls, stakeholder responsibilities, and operational workflow integration points.
These lifecycle-oriented governance mechanisms also align closely with emerging healthcare MLOps and operational AI governance frameworks, which emphasize reproducibility, continuous validation, deployment monitoring, auditability, and post-deployment lifecycle management for clinical AI systems [55,56,57,75]. In healthcare settings, operational AI governance increasingly requires integration of governance controls into model registries, CI/CD pipelines, monitoring infrastructure, version control systems, and post-market surveillance workflows to support ongoing safety, reliability, and regulatory compliance throughout real-world deployment [55,56,75]. Recent work on clinical AI deployment and AI-based medical device governance further highlights the importance of reproducibility, continuous monitoring, and adaptive lifecycle management as essential components of trustworthy healthcare AI systems [57,75].
The operational workflow integration points summarized in Table 8 are informed by emerging healthcare MLOps, operational AI governance, and lifecycle management frameworks that emphasize reproducibility, deployment monitoring, auditability, and continuous oversight of clinical AI systems [55,56,57,75].

4.6.4. Lessons from Retrospective Governance Failure

The retrospective Optum case demonstrates that governance failures in healthcare AI systems rarely emerge from a single isolated technical error. Instead, harms often arise through the accumulation of weaknesses across multiple stages of the AI lifecycle, including inappropriate proxy selection, insufficient subgroup validation, fragmented accountability structures, inadequate fairness monitoring, and limited post-deployment oversight [81]. Although the original system achieved strong predictive performance according to conventional cost-based metrics, the use of healthcare expenditure as a proxy for healthcare need contributed to systematic disparities in care allocation that were not adequately identified during earlier stages of development and deployment.
This case study therefore highlights the importance of treating healthcare AI governance as a continuous sociotechnical process rather than a narrow technical compliance exercise. Governance mechanisms focused exclusively on model accuracy or regulatory approval may fail to detect broader ethical and clinical risks if fairness, transparency, accountability, and long-term monitoring are not integrated throughout the AI lifecycle [55,75,82]. The applied governance-by-design framework proposed in this paper addresses these limitations by embedding multidisciplinary oversight, fairness-oriented evaluation, stakeholder accountability, and continuous monitoring processes directly into operational AI workflows.
Importantly, the case also demonstrates that governance interventions are not limited to post hoc ethical review. Many of the disparities identified in the retrospective analysis could potentially have been mitigated earlier through upstream governance mechanisms such as proxy variable assessment, subgroup-specific validation, fairness-oriented performance evaluation, and continuous post-deployment monitoring. Embedding these safeguards into healthcare AI development pipelines therefore supports more proactive and adaptive governance approaches capable of responding to evolving sociotechnical risks throughout the lifecycle of clinical AI systems.

5. Discussion

5.1. Convergence Across Frameworks

This review identifies a substantial degree of convergence across international ethical, regulatory, and technical frameworks governing healthcare AI [83,84,85,86]. Foundational initiatives, including WHO ethical principles, FDA Good Machine Learning Practice (GMLP), the EU AI Act, OECD recommendations, and guidance from professional societies, consistently emphasize safety, transparency, accountability, and fairness as core requirements for trustworthy AI [87]. This alignment suggests the emergence of a shared global baseline for responsible AI in healthcare and reflects a growing consensus that ethical governance must be embedded throughout the AI lifecycle rather than applied at isolated stages [31].
Healthcare-specific frameworks, such as WHO guidance, FDA GMLP, and recommendations from the National Academy of Medicine, offer more operational detail compared to cross-sectoral principles developed by organizations such as the OECD or IEEE [43,88]. This increased specificity reflects the safety-critical nature of clinical decision-making, where even minor system failures can have significant consequences. At the same time, regulatory instruments such as the EU AI Act and FDA oversight pathways introduce enforceable requirements that formalize these expectations, although they may evolve more slowly than rapidly advancing AI technologies [39]. This creates a dynamic tension between regulatory stability and technological change, requiring governance models that can remain both robust and adaptable over time.
In contrast, ethical guidelines and professional standards tend to be more flexible and adaptive, allowing them to respond more quickly to emerging challenges, including those associated with generative AI and continuously learning systems [49,89]. However, their non-binding nature limits enforceability, creating a gap between normative expectations and practical implementation [90]. Reporting standards such as TRIPOD-AI, CONSORT-AI, and SPIRIT-AI help bridge this gap by improving methodological rigor, transparency, and reproducibility in evidence generation [53,54]. These standards play a critical role in translating high-level principles into measurable and reportable practices, although adherence remains inconsistent across the literature [52].
Despite this convergence, alignment at the level of principles does not automatically translate into consistency in implementation [19,50,78]. Differences in terminology, scope, and enforcement mechanisms across frameworks can create ambiguity for developers and healthcare organizations navigating multiple regulatory environments [76]. This challenge is particularly evident in cross-border deployment scenarios, where systems must simultaneously satisfy overlapping but not fully harmonized requirements. Fragmentation across governance regimes can also lead to duplication of effort, inconsistent validation practices, and uncertainty in accountability allocation [51].
Unlike many existing trustworthy AI frameworks that primarily emphasize high-level ethical principles or isolated governance requirements, the framework proposed in this review adopts a lifecycle-aligned operational structure that explicitly maps ethical principles to stage-specific governance mechanisms across data collection, model development, validation, deployment, and post-deployment monitoring [26,82,88]. Rather than treating governance as a static compliance exercise, the framework integrates technical safeguards, organizational oversight, and continuous monitoring processes throughout the AI lifecycle [55,56,75]. This approach extends beyond principle-based guidance by emphasizing operational governance checkpoints, multidisciplinary accountability, and continuous lifecycle integration within real-world healthcare AI workflows.
The lifecycle-aligned governance framework developed in this review builds on this convergence by integrating ethical principles, regulatory expectations, technical methods, and stakeholder engagement into a unified operational structure. By explicitly mapping principles to stage-specific governance mechanisms, the framework provides a practical pathway for translating shared normative commitments into consistent and actionable practices across the AI lifecycle [80,91]. Importantly, it shifts the focus from alignment in principle to alignment in implementation, enabling organizations to operationalize ethical requirements within real-world clinical and technical workflows. In this way, the framework contributes not only to conceptual clarity but also to practical governance, addressing a critical gap between existing guidance and implementation in healthcare AI systems.

5.2. Persistent Gaps in Ethical Operationalization

Despite broad convergence at the level of ethical principles and regulatory intent, substantial gaps remain in practical implementation [92,93]. These gaps reflect the inherent difficulty of translating high-level commitments into consistent, enforceable practices within complex and distributed healthcare AI ecosystems. Prior work has emphasized that principle-based frameworks alone are insufficient without operational mechanisms, measurable criteria, and lifecycle integration strategies [33].
First, accountability fragmentation persists due to the distributed nature of AI development and deployment [43,94]. Multiple stakeholders, including developers, vendors, healthcare organizations, clinicians, and regulators, each exercise partial control over system design, implementation, and use. When harms occur, responsibility may become diffuse, complicating both liability determination and mechanisms for redress. While existing frameworks emphasize accountability, they often lack clear operational models for allocating responsibility across the AI lifecycle [92]. In practice, this can result in governance gaps where no single actor assumes full ownership of downstream outcomes, particularly in systems involving third-party vendors or continuously updated models. These challenges are further intensified by cross-jurisdictional deployment and regulatory fragmentation, which can obscure responsibility boundaries and complicate enforcement [39,78].
Second, equity evaluation remains underdeveloped relative to the strength of normative commitments [19,78]. Although fairness is widely recognized as a core principle, there is limited consensus on how it should be measured, enforced, and monitored in clinical settings. Different fairness metrics may produce conflicting results, and guidance on selecting context-appropriate metrics remains limited [50]. Many validation studies continue to report aggregate performance metrics without sufficient subgroup analysis, obscuring disparities across demographic and clinical populations. In addition, formal equity impact assessments are rarely mandated prior to deployment, limiting the ability to proactively identify and mitigate potential harms. Empirical evidence from healthcare AI studies demonstrates that such gaps can lead to systematic underperformance in marginalized populations, reinforcing existing inequities in care [16,17]. Given the potential for AI systems to amplify structural disparities, this gap represents a critical ethical and clinical concern.
Third, post-market surveillance infrastructure remains comparatively underdeveloped [76]. Regulatory and institutional efforts are often concentrated on pre-deployment validation, while continuous monitoring receives less structured attention [30,95]. In real-world settings, however, AI systems are exposed to evolving data distributions, changing clinical practices, and new patient populations. These dynamics can lead to model drift, context shift, and emergent biases that degrade performance over time [24,52]. Without robust and standardized monitoring mechanisms, such degradation may remain undetected until it results in patient harm or system failure. Current approaches to post-deployment monitoring are often fragmented or ad hoc, lacking the consistency and rigor applied during earlier stages of development. Broader analyses of real-world AI deployment further confirm that monitoring remains one of the weakest links in the lifecycle of healthcare AI systems [95].
More broadly, these gaps highlight a structural disconnect between ethical intent and operational practice. Existing frameworks provide strong normative guidance but insufficient direction on how to embed these principles into day-to-day development, deployment, and monitoring processes. This disconnect reflects a broader challenge in AI governance, where alignment at the level of principles does not automatically translate into alignment in implementation.
Taken together, these challenges reinforce the need for lifecycle-integrated governance approaches. Addressing accountability fragmentation, strengthening equity evaluation, and establishing robust post-market surveillance systems require coordinated mechanisms that span all stages of the AI lifecycle [80,91]. Without such integration, ethical principles risk remaining aspirational rather than actionable in real-world healthcare settings.

5.3. Implementation Barriers

Even when governance frameworks are conceptually robust, translating them into practice remains challenging due to systemic and organizational barriers [66,71,96]. These barriers extend beyond technical limitations and reflect broader constraints related to resources, coordination, institutional capacity, and the pace of technological change. Prior work has highlighted that implementation gaps often arise not from a lack of guidance, but from the difficulty of integrating governance requirements into complex real-world healthcare environments [83,84,87].
Capacity limitations represent a major obstacle, particularly in low-resource settings and smaller healthcare institutions [72,73]. Effective governance requires multidisciplinary expertise spanning data science, clinical practice, ethics, and regulatory compliance, as well as technical infrastructure for monitoring, documentation, and continuous model evaluation. Sustaining these capabilities over time demands ongoing investment and institutional commitment. In practice, many organizations lack the necessary workforce, infrastructure, or funding to implement comprehensive governance frameworks. As a result, the ability to operationalize ethical AI is often concentrated in well-resourced institutions, raising concerns that AI innovation may disproportionately benefit already advantaged healthcare systems while widening global inequities [64,74]. This imbalance highlights the need for scalable and accessible governance models that can be adapted across diverse healthcare contexts [71].
Governance fragmentation further complicates implementation [39,88,92]. Developers and healthcare organizations must navigate overlapping and sometimes inconsistent requirements across international ethical guidelines, national regulations, regional data protection laws, professional standards, and institutional policies. In practice, these frameworks may differ in scope, terminology, and enforcement mechanisms, creating uncertainty about how to achieve compliance across jurisdictions. This challenge is particularly pronounced in cross-border deployment, where systems must satisfy multiple regulatory regimes simultaneously [65,67]. Without greater harmonization, fragmentation increases compliance burden, slows deployment, and introduces ambiguity in responsibility allocation across stakeholders. Broader analyses of AI governance have shown that such fragmentation can lead to duplication of validation efforts, inconsistent documentation practices, and gaps in accountability across the AI lifecycle [33,68].
A further challenge arises from the tension between innovation and regulation [69,94]. AI technologies evolve rapidly, with new architectures, training paradigms, and deployment models emerging at a pace that often outstrips regulatory adaptation. Regulatory systems originally designed for static medical devices may struggle to accommodate adaptive, continuously learning, or foundation model-based systems that evolve after deployment. This mismatch creates a difficult balance: overly restrictive regulation may hinder beneficial innovation and delay clinical adoption, while insufficient oversight increases the risk of deploying inadequately validated or unsafe systems. Recent advances in generative AI further intensify this tension, as these systems introduce new uncertainties related to reliability, explainability, and clinical appropriateness [63,90]. These developments underscore the need for governance approaches that are both adaptive and risk-sensitive, capable of evolving alongside technological change without compromising patient safety.
These challenges are supported by a growing body of literature highlighting persistent gaps in multiple dimensions of healthcare AI governance. Studies on technical robustness and system reliability demonstrate vulnerabilities related to instability, adversarial inputs, and performance degradation in real-world settings [60,70]. In parallel, research on fairness evaluation highlights ongoing challenges in defining, measuring, and operationalizing equity across diverse clinical populations [19,50,78]. Additional work has identified limitations in existing governance frameworks and post-deployment monitoring practices, emphasizing the lack of consistent, lifecycle-integrated oversight mechanisms [51,76].
Taken together, these findings highlight the need for adaptive governance models that can evolve alongside technological innovation while maintaining safety, accountability, and ethical integrity [80,91]. Such models must be flexible enough to accommodate emerging technologies yet structured enough to provide clear guidance for implementation across diverse healthcare contexts. In particular, there is a need for governance approaches that integrate technical tools, regulatory requirements, and organizational processes into cohesive systems capable of supporting lifecycle-aligned oversight in real-world clinical environments.

5.4. The Value of Governance-by-Design

The governance-by-design approach proposed in this review addresses persistent gaps in healthcare AI oversight by embedding ethical requirements throughout the AI lifecycle rather than confining evaluation to isolated checkpoints [66,71]. By integrating safeguards at each stage, potential risks can be identified and mitigated early, when intervention remains technically feasible and organizationally practical. This approach distributes responsibility across stakeholders while maintaining clear accountability pathways, reducing ambiguity in how ethical obligations are assigned and enforced. Importantly, this lifecycle-oriented model aligns with emerging perspectives that emphasize continuous risk management, adaptive validation, and system-level governance as essential components of trustworthy AI [96].
Evidence from real-world implementations highlights the consequences of fragmented governance [69]. Cases such as the DeepMind–Royal Free collaboration and proprietary clinical prediction systems that demonstrated degraded performance outside their development environments illustrate how failures at early lifecycle stages can propagate downstream harms [97]. These failures were not solely technical in nature; they reflected deficiencies in consent governance, insufficient external validation, and inadequate post-deployment monitoring. From a technical perspective, such failures are often linked to issues such as dataset bias, lack of distributional robustness, and absence of drift detection mechanisms, all of which can lead to performance degradation in real-world settings [97]. A lifecycle-integrated governance structure could have identified these vulnerabilities earlier and reduced the likelihood of harm.
Governance-by-design also reframes ethical oversight as enabling infrastructure rather than an external constraint [64,74]. By embedding transparency, fairness evaluation, and monitoring mechanisms directly into development workflows, ethical considerations become part of routine quality assurance rather than an afterthought. Technically, this includes integrating bias detection pipelines, explainability modules, uncertainty quantification, and validation protocols into model development and deployment processes. This alignment supports regulatory readiness, enhances institutional trust, and strengthens clinical legitimacy while ensuring that ethical safeguards are systematically enforced rather than retrofitted.
From a technical standpoint, governance-by-design enables the integration of key methodological safeguards across the AI lifecycle [46,47,48,49,89]. During data collection, representational auditing and dataset documentation (e.g., datasheets for datasets) improve transparency and mitigate bias. During model development, fairness-aware optimization, adversarial robustness testing, and uncertainty estimation techniques enhance model reliability and safety. Explainability methods such as SHAP and LIME support interpretability, although their limitations require careful clinical contextualization [21,47]. During validation, external validation protocols, subgroup analysis, and standardized reporting frameworks (TRIPOD-AI, CONSORT-AI) improve generalizability and transparency [9,53]. At deployment, workflow-aware system design, clinician-in-the-loop oversight, and fail-safe mechanisms reduce risks such as automation bias and misinterpretation. Finally, continuous monitoring mechanisms, including drift detection, performance dashboards, and adaptive retraining strategies, ensure long-term system reliability and safety.
However, governance-by-design requires sustained organizational commitment, resource allocation, and cultural change [67,68]. It demands not only technical tools but also formal governance structures, clearly defined responsibility chains, and meaningful stakeholder engagement processes. The framework presented here provides a structured roadmap, but its effectiveness ultimately depends on institutional uptake, integration into clinical workflows, and ongoing adaptation as technologies and healthcare environments evolve.
In real-world healthcare settings, many failures of AI systems have emerged not from algorithmic limitations alone, but from gaps in governance across the lifecycle. Issues such as biased training data, lack of external validation, poor workflow integration, and insufficient post-deployment monitoring have been repeatedly observed in clinical AI deployments. The governance-by-design framework directly addresses these challenges by embedding safeguards at each stage of system development and use. For example, representational auditing during data collection, fairness-aware optimization during model development, and subgroup-specific validation protocols help mitigate bias before deployment [33]. Similarly, structured external validation and continuous monitoring mechanisms reduce the risk of performance degradation across institutions and patient populations.
The framework also responds to practical challenges faced by clinicians and healthcare organizations. In clinical environments, AI systems must function within time-constrained workflows and support decision-making without increasing cognitive burden. Governance-by-design incorporates clinician-in-the-loop design, explainability mechanisms, and workflow-aligned deployment strategies to improve usability and reduce risks such as automation bias and alert fatigue [16,54,91]. In addition, continuous feedback loops and monitoring systems enable healthcare organizations to detect emerging issues, adapt models over time, and integrate AI oversight into existing quality improvement processes. These mechanisms ensure that governance is not a one-time requirement but an ongoing component of clinical practice.
From a broader systems perspective, the framework supports alignment with evolving regulatory and technological landscapes. As healthcare increasingly adopts advanced AI systems, including generative models and foundation models, new risks related to hallucination, instability, and context sensitivity have emerged [58,61,63,98]. These systems require additional safeguards, including output verification, human oversight, and continuous validation across diverse clinical contexts [59]. Governance-by-design provides a structured approach to managing these risks through lifecycle-integrated monitoring and adaptive governance mechanisms. By aligning ethical principles with real-world implementation processes, the framework offers a practical pathway for translating normative guidance into safe, scalable, and trustworthy AI deployment in clinical medicine.
In this sense, governance-by-design positions ethical AI not as a barrier to innovation, but as a foundational infrastructure for safe, equitable, and scalable healthcare AI. By embedding ethical considerations into the full lifecycle of system development and deployment, it enables innovation that is both responsible and sustainable. Figure 5 illustrates how governance-by-design addresses key real-world challenges in healthcare AI and enables improved clinical and system-level outcomes through lifecycle-integrated governance.
Despite the strengths of this multidisciplinary narrative synthesis, several limitations should be acknowledged. First, the review was narrative rather than systematic and therefore was not intended to provide exhaustive literature coverage. Although structured searches and representative case examples were incorporated, thematic synthesis and study selection involved interpretive judgment, which may introduce selection bias. In addition, the proposed governance-by-design framework remains conceptual and has not yet been prospectively evaluated across real-world institutional implementation settings. Future studies should focus on validating and operationalizing the framework through measurable governance indicators, implementation toolkits, and audit-oriented assessment approaches across diverse healthcare environments.

6. Future Directions

Advancing ethical governance in healthcare AI requires moving beyond high-level principles toward measurable, scalable, and globally coordinated implementation strategies. While significant progress has been made in defining ethical expectations, the next phase must focus on translating these commitments into operational, evaluable, and sustainable practices across real-world clinical environments. Increasingly, this shift is understood as a transition from principle-based governance to systems-level governance, where ethical requirements are embedded into technical, organizational, and regulatory infrastructures across the AI lifecycle [82,99,100].
First, greater international harmonization of regulatory frameworks is essential to reduce fragmentation and support safe cross-border deployment. Healthcare AI systems are increasingly developed and deployed across jurisdictions, yet developers must currently navigate heterogeneous regulatory requirements with differing terminology, documentation standards, and enforcement mechanisms. Coordinated efforts among organizations such as WHO, FDA, EMA, and other national authorities could align lifecycle expectations, risk classification approaches, and post-market surveillance requirements [99,100]. Importantly, harmonization should aim for interoperability rather than uniformity, preserving contextual flexibility while reducing redundancy and ambiguity. Such alignment is particularly important as AI systems become more adaptive and continuously updated, requiring consistent oversight across jurisdictions.
Second, the development of quantifiable ethics metrics is critical to move from qualitative commitments toward measurable performance indicators [79,80]. Although fairness, transparency, and accountability are widely endorsed, there remains limited consensus on how to evaluate these dimensions consistently across systems [8,101]. Emerging research highlights the need for standardized metrics assessing subgroup performance disparities, robustness across clinical contexts, explanation usefulness for end users, and responsiveness of monitoring systems [100]. In addition, advances in uncertainty quantification, calibration, and reliability assessment provide opportunities to operationalize safety in measurable terms. Establishing such metrics would enable benchmarking, comparative evaluation, and continuous improvement, while also supporting regulatory and institutional accountability.
Third, governance tooling must be expanded and made accessible to support practical implementation [91]. Advances in open-source ecosystems have produced tools for bias detection, explainability, and model monitoring; however, their adoption remains uneven. Broader dissemination of interoperable toolkits, automated auditing pipelines, and standardized documentation frameworks can lower technical barriers [79]. At the same time, these tools must be accompanied by clear interpretive guidance and institutional oversight, as technical solutions alone cannot ensure ethical alignment and may otherwise lead to superficial or checkbox-style compliance. Future work should also focus on integrating these tools directly into clinical systems and development pipelines to support real-time governance rather than retrospective auditing.
Fourth, capacity building is essential, particularly in low-resource and underrepresented healthcare settings [72,73]. Effective governance requires multidisciplinary expertise spanning clinical medicine, data science, ethics, and regulatory compliance, as well as infrastructure for validation, monitoring, and auditing. Without targeted investment, there is a risk that ethically robust AI systems will remain concentrated in well-resourced institutions, reinforcing existing global disparities in healthcare innovation [102]. Collaborative networks, shared validation platforms, federated evaluation infrastructures, and targeted funding initiatives can help democratize access to governance capabilities and promote more equitable participation in AI development and deployment.
Fifth, governance frameworks must evolve to address the emergence of generative AI and foundation models in healthcare [58,61,62,63,90,98]. These systems introduce new challenges, including hallucinated outputs, sensitivity to prompt variation, domain misalignment, and limited transparency in training and fine-tuning processes. In clinical settings, such risks may directly affect patient safety and decision-making. Extending lifecycle governance approaches to include output verification mechanisms, uncertainty-aware interfaces, context-sensitive deployment strategies, and robust human-in-the-loop safeguards will be essential to ensure responsible use of these technologies [103,104,105,106]. These requirements further reinforce the need for continuous validation and monitoring rather than static approval processes.
Finally, adaptive governance models are required to balance innovation with patient protection in rapidly evolving technological environments [59,107]. Traditional regulatory approaches, often designed for static medical devices, may be insufficient for continuously learning systems that evolve post-deployment. Emerging strategies such as regulatory sandboxes, conditional approvals with enhanced monitoring, iterative certification pathways, and real-time post-market evaluation offer promising approaches to maintaining safety while enabling innovation [108,109]. These models emphasize ongoing oversight rather than one-time approval and align closely with lifecycle-based governance approaches that recognize AI systems as dynamic rather than fixed technologies [107].
Collectively, these directions highlight the need for sustained investment in governance infrastructure capable of translating ethical principles into practical, measurable, and continuously evolving systems of oversight across the full healthcare AI lifecycle. Advancing these efforts will be essential to ensure that AI systems are not only innovative, but also safe, equitable, and trustworthy in real-world clinical practice.

7. Conclusions

This narrative review demonstrates that operationalizing the WHO’s ethical principles for healthcare AI requires systematic integration of ethical commitments into lifecycle-aligned governance mechanisms rather than reliance on abstract normative endorsement. While the six WHO principles, such as autonomy, well-being and safety, transparency, accountability, inclusiveness and equity, and sustainability, provide a robust ethical foundation, their effective realization depends on structured implementation across data collection, model development, validation, deployment, and monitoring stages [99,100]. The governance-by-design framework presented in this review embeds ethical safeguards directly into technical and organizational workflows from project inception through post-deployment surveillance, offering a practical pathway for translating ethical theory into operational practice [82]. By synthesizing international regulatory frameworks, technical standards, and documented implementation failures, this review highlights both convergence and persistent gaps in current governance approaches [99,110]. Although global consensus increasingly emphasizes safety, transparency, accountability, and fairness, challenges such as fragmented responsibility across stakeholders, insufficient equity evaluation, limited post-market surveillance infrastructure, and uneven governance capacity continue to hinder consistent implementation [100,101]. Addressing these challenges requires coordinated action across the healthcare AI ecosystem, including developers integrating governance into system design, healthcare institutions establishing robust oversight and monitoring processes, regulators advancing adaptive and harmonized frameworks, and policymakers investing in governance capacity and equitable access, while ensuring meaningful patient and community involvement [101,102]. As AI technologies continue to evolve, particularly with the emergence of foundation models and generative systems, governance strategies must also adapt to address new forms of uncertainty, opacity, and risk [103,104,105,108]. Advancing toward trustworthy healthcare AI will require the development of quantifiable ethical metrics, interoperable regulatory standards, and scalable governance tools that support continuous oversight [79,80]. Embedding WHO’s ethical principles across the AI lifecycle through governance-by-design ultimately reframes ethics not as a constraint on innovation but as essential infrastructure for safe, equitable, and sustainable AI-enabled healthcare, providing a structured roadmap for aligning technological advancement with patient-centered values and public trust [99,110].

Author Contributions

Conceptualization, K.S.; methodology, K.S.; investigation, K.S.; writing—original draft preparation, K.S.; writing—review and editing, K.G., D.S., A.K., S.S., S.A.H., S.P.A. and D.M.; supervision, D.M. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Acknowledgments

This work was supported by resources within the Digital Engineering and Artificial Intelligence Laboratory (DEAL), Department of Medicine, Mayo Clinic, Jacksonville, Florida, USA.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial intelligence
EMAEuropean Medicines Agency
EUEuropean Union
FDAU.S. Food and Drug Administration
GDPRGeneral Data Protection Regulation
GMLPGood machine learning practice
LLMLarge language model
MDRMedical device regulation
MLOpsMachine learning operations
OECDOrganization for Economic Co-operation and Development
WHOWorld Health Organization

References

  1. Yu, K.-H.; Beam, A.L.; Kohane, I.S. Artificial Intelligence in Healthcare. Nat. Biomed. Eng. 2018, 2, 719–731. [Google Scholar] [CrossRef]
  2. Jiang, F.; Jiang, Y.; Zhi, H.; Dong, Y.; Li, H.; Ma, S.; Wang, Y.; Dong, Q.; Shen, H.; Wang, Y. Artificial Intelligence in Healthcare: Past, Present and Future. Stroke Vasc. Neurol. 2017, 2, 230–243. [Google Scholar] [CrossRef] [PubMed]
  3. Davenport, T.; Kalakota, R. The Potential for Artificial Intelligence in Healthcare. Future Healthc. J. 2019, 6, 94–98. [Google Scholar] [CrossRef] [PubMed]
  4. Topol, E.J. High-Performance Medicine: The Convergence of Human and Artificial Intelligence. Nat. Med. 2019, 25, 44–56. [Google Scholar] [CrossRef]
  5. The Learning Health System Series; National Academy of Medicine. Artificial Intelligence in Health Care: The Hope, the Hype, the Promise, the Peril; Matheny, M., Israni, S.T., Ahmed, M., Whicher, D., Eds.; National Academies Press: Washington, DC, USA, 2019; p. 27111. ISBN 978-0-309-70513-4. [Google Scholar]
  6. Esteva, A.; Kuprel, B.; Novoa, R.A.; Ko, J.; Swetter, S.M.; Blau, H.M.; Thrun, S. Dermatologist-Level Classification of Skin Cancer with Deep Neural Networks. Nature 2017, 542, 115–118. [Google Scholar] [CrossRef]
  7. Albano, D.; Galiano, V.; Basile, M.; Di Luca, F.; Gitto, S.; Messina, C.; Cagetti, M.G.; Del Fabbro, M.; Tartaglia, G.M.; Sconfienza, L.M. Artificial Intelligence for Radiographic Imaging Detection of Caries Lesions: A Systematic Review. BMC Oral Health 2024, 24, 274. [Google Scholar] [CrossRef] [PubMed]
  8. Cai, Y.; Cai, Y.-Q.; Tang, L.-Y.; Wang, Y.-H.; Gong, M.; Jing, T.-C.; Li, H.-J.; Li-Ling, J.; Hu, W.; Yin, Z.; et al. Artificial Intelligence in the Risk Prediction Models of Cardiovascular Disease and Development of an Independent Validation Screening Tool: A Systematic Review. BMC Med. 2024, 22, 56. [Google Scholar] [CrossRef] [PubMed]
  9. Liu, X.; Faes, L.; Kale, A.U.; Wagner, S.K.; Fu, D.J.; Bruynseels, A.; Mahendiran, T.; Moraes, G.; Shamdas, M.; Kern, C.; et al. A Comparison of Deep Learning Performance against Health-Care Professionals in Detecting Diseases from Medical Imaging: A Systematic Review and Meta-Analysis. Lancet Digit. Health 2019, 1, e271–e297. [Google Scholar] [CrossRef] [PubMed]
  10. He, J.; Baxter, S.L.; Xu, J.; Xu, J.; Zhou, X.; Zhang, K. The Practical Implementation of Artificial Intelligence Technologies in Medicine. Nat. Med. 2019, 25, 30–36. [Google Scholar] [CrossRef]
  11. Shortliffe, E.H.; Sepúlveda, M.J. Clinical Decision Support in the Era of Artificial Intelligence. JAMA 2018, 320, 2199. [Google Scholar] [CrossRef]
  12. Topol, E.J. Deep Medicine: How Artificial Intelligence Can Make Healthcare Human Again, 1st ed.; Basic Books: New York, NY, USA, 2019; ISBN 978-1-5416-4463-2. [Google Scholar]
  13. Kelly, C.J.; Karthikesalingam, A.; Suleyman, M.; Corrado, G.; King, D. Key Challenges for Delivering Clinical Impact with Artificial Intelligence. BMC Med. 2019, 17, 195. [Google Scholar] [CrossRef] [PubMed]
  14. Ghassemi, M.; Naumann, T.; Schulam, P.; Beam, A.L.; Chen, I.Y.; Ranganath, R. A Review of Challenges and Opportunities in Machine Learning for Health. AMIA Summits Transl. Sci. Proc. 2020, 2020, 191–200. [Google Scholar] [PubMed]
  15. Wiens, J.; Saria, S.; Sendak, M.; Ghassemi, M.; Liu, V.X.; Doshi-Velez, F.; Jung, K.; Heller, K.; Kale, D.; Saeed, M.; et al. Do No Harm: A Roadmap for Responsible Machine Learning for Health Care. Nat. Med. 2019, 25, 1337–1340. [Google Scholar] [CrossRef]
  16. Challen, R.; Denny, J.; Pitt, M.; Gompels, L.; Edwards, T.; Tsaneva-Atanasova, K. Artificial Intelligence, Bias and Clinical Safety. BMJ Qual. Saf. 2019, 28, 231–237. [Google Scholar] [CrossRef]
  17. Gianfrancesco, M.A.; Tamang, S.; Yazdany, J.; Schmajuk, G. Potential Biases in Machine Learning Algorithms Using Electronic Health Record Data. JAMA Intern. Med. 2018, 178, 1544. [Google Scholar] [CrossRef]
  18. Rajkomar, A.; Hardt, M.; Howell, M.D.; Corrado, G.; Chin, M.H. Ensuring Fairness in Machine Learning to Advance Health Equity. Ann. Intern. Med. 2018, 169, 866–872. [Google Scholar] [CrossRef]
  19. Seyyed-Kalantari, L.; Zhang, H.; McDermott, M.B.A.; Chen, I.Y.; Ghassemi, M. Underdiagnosis Bias of Artificial Intelligence Algorithms Applied to Chest Radiographs in Under-Served Patient Populations. Nat. Med. 2021, 27, 2176–2182. [Google Scholar] [CrossRef]
  20. Tonekaboni, S.; Joshi, S.; McCradden, M.D.; Goldenberg, A. What Clinicians Want: Contextualizing Explainable Machine Learning for Clinical End Use. In Machine Learning for Healthcare Conference; PMLR: New York, NY, USA, 2019. [Google Scholar]
  21. Rudin, C. Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead. Nat. Mach. Intell. 2019, 1, 206–215. [Google Scholar] [CrossRef] [PubMed]
  22. Zech, J.R.; Badgeley, M.A.; Liu, M.; Costa, A.B.; Titano, J.J.; Oermann, E.K. Variable Generalization Performance of a Deep Learning Model to Detect Pneumonia in Chest Radiographs: A Cross-Sectional Study. PLoS Med. 2018, 15, e1002683. [Google Scholar] [CrossRef]
  23. Sendak, M.P.; Ratliff, W.; Sarro, D.; Alderton, E.; Futoma, J.; Gao, M.; Nichols, M.; Revoir, M.; Yashar, F.; Miller, C.; et al. Real-World Integration of a Sepsis Deep Learning Technology into Routine Clinical Care: Implementation Study. JMIR Med. Inf. 2020, 8, e15182. [Google Scholar] [CrossRef]
  24. Eertink, J.J.; Heymans, M.W.; Zwezerijnen, G.J.C.; Zijlstra, J.M.; De Vet, H.C.W.; Boellaard, R. External Validation: A Simulation Study to Compare Cross-Validation versus Holdout or External Testing to Assess the Performance of Clinical Prediction Models Using PET Data from DLBCL Patients. EJNMMI Res. 2022, 12, 58. [Google Scholar] [CrossRef]
  25. Cabitza, F.; Rasoini, R.; Gensini, G.F. Unintended Consequences of Machine Learning in Medicine. JAMA 2017, 318, 517. [Google Scholar] [CrossRef] [PubMed]
  26. Reddy, S.; Allan, S.; Coghlan, S.; Cooper, P. A Governance Model for the Application of AI in Health Care. J. Am. Med. Inform. Assoc. 2020, 27, 491–497. [Google Scholar] [CrossRef]
  27. Morley, J.; Elhalal, A.; Garcia, F.; Kinsey, L.; Mökander, J.; Floridi, L. Ethics as a Service: A Pragmatic Operationalisation of AI Ethics. Minds Mach. 2021, 31, 239–256. [Google Scholar] [CrossRef] [PubMed]
  28. Floridi, L.; Cowls, J.; Beltrametti, M.; Chatila, R.; Chazerand, P.; Dignum, V.; Luetge, C.; Madelin, R.; Pagallo, U.; Rossi, F.; et al. AI4People—An Ethical Framework for a Good AI Society: Opportunities, Risks, Principles, and Recommendations. Minds Mach. 2018, 28, 689–707. [Google Scholar] [CrossRef]
  29. Jobin, A.; Ienca, M.; Vayena, E. The Global Landscape of AI Ethics Guidelines. Nat. Mach. Intell. 2019, 1, 389–399. [Google Scholar] [CrossRef]
  30. London, A.J. Artificial Intelligence and Black-Box Medical Decisions: Accuracy versus Explainability. Hastings Cent. Rep. 2019, 49, 15–21. [Google Scholar] [CrossRef]
  31. Ahadian, P.; Xu, W.; Liu, D.; Guan, Q. Ethics of Trustworthy AI in Healthcare: Challenges, Principles, and Practical Pathways. Neurocomputing 2026, 661, 131942. [Google Scholar] [CrossRef]
  32. Price, W.N.; Cohen, I.G. Privacy in the Age of Medical Big Data. Nat. Med. 2019, 25, 37–43. [Google Scholar] [CrossRef]
  33. Vayena, E.; Blasimme, A. Biomedical Big Data: New Models of Control Over Access, Use and Governance. Bioethical Inq. 2017, 14, 501–513. [Google Scholar] [CrossRef]
  34. Gerke, S.; Minssen, T.; Cohen, G. Ethical and Legal Challenges of Artificial Intelligence-Driven Healthcare. In Artificial Intelligence in Healthcare; Elsevier: Amsterdam, The Netherlands, 2020; pp. 295–336. ISBN 978-0-12-818438-7. [Google Scholar]
  35. Zhou, K.; Gattinger, G. The Evolving Regulatory Paradigm of AI in MedTech: A Review of Perspectives and Where We Are Today. Ther. Innov. Regul. Sci. 2024, 58, 456–464. [Google Scholar] [CrossRef]
  36. Mittelstadt, B. Principles Alone Cannot Guarantee Ethical AI. Nat. Mach. Intell. 2019, 1, 501–507. [Google Scholar] [CrossRef]
  37. Morley, J.; Machado, C.C.V.; Burr, C.; Cowls, J.; Joshi, I.; Taddeo, M.; Floridi, L. The Ethics of AI in Health Care: A Mapping Review. Soc. Sci. Med. 2020, 260, 113172. [Google Scholar] [CrossRef]
  38. Kaissis, G.A.; Makowski, M.R.; Rückert, D.; Braren, R.F. Secure, Privacy-Preserving and Federated Machine Learning in Medical Imaging. Nat. Mach. Intell. 2020, 2, 305–311. [Google Scholar] [CrossRef]
  39. Han, Y.; Ceross, A.; Bergmann, J. Regulatory Frameworks for AI-Enabled Medical Device Software in China: Comparative Analysis and Review of Implications for Global Manufacturer. JMIR AI 2024, 3, e46871. [Google Scholar] [CrossRef]
  40. Guidance, WHO. Ethics and Governance of Artificial Intelligence for Health: WHO Guidance, 1st ed.; World Health Organization: Geneva, Switzerland, 2021; ISBN 978-92-4-002920-0. [Google Scholar]
  41. Mittelstadt, B.D.; Allo, P.; Taddeo, M.; Wachter, S.; Floridi, L. The Ethics of Algorithms: Mapping the Debate. Big Data Soc. 2016, 3, 2053951716679679. [Google Scholar] [CrossRef]
  42. Leslie, D. Understanding Artificial Intelligence Ethics and Safety: A Guide for the Responsible Design and Implementation of AI Systems in the Public Sector. arXiv 2019, arXiv:1906.05684. [Google Scholar] [CrossRef]
  43. Benjamens, S.; Dhunnoo, P.; Meskó, B. The State of Artificial Intelligence-Based FDA-Approved Medical Devices and Algorithms: An Online Database. NPJ Digit. Med. 2020, 3, 118. [Google Scholar] [CrossRef] [PubMed]
  44. Vokinger, K.N.; Feuerriegel, S.; Kesselheim, A.S. Mitigating Bias in Machine Learning for Medicine. Commun. Med. 2021, 1, 25. [Google Scholar] [CrossRef]
  45. Rieke, N.; Hancox, J.; Li, W.; Milletarì, F.; Roth, H.R.; Albarqouni, S.; Bakas, S.; Galtier, M.N.; Landman, B.A.; Maier-Hein, K.; et al. The Future of Digital Health with Federated Learning. NPJ Digit. Med. 2020, 3, 119. [Google Scholar] [CrossRef]
  46. Samek, W.; Wiegand, T.; Müller, K.-R. Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models. arXiv 2017, arXiv:1708.08296. [Google Scholar] [CrossRef]
  47. Guidotti, R.; Monreale, A.; Ruggieri, S.; Turini, F.; Pedreschi, D.; Giannotti, F. A Survey of Methods For Explaining Black Box Models. ACM Comput. Surv. (CSUR) 2018, 51, 1–42. [Google Scholar] [CrossRef]
  48. Holzinger, A. Explainable AI and Multi-Modal Causability in Medicine. i-com 2021, 19, 171–179. [Google Scholar] [CrossRef] [PubMed]
  49. Ahmad, M.A.; Eckert, C.; Teredesai, A. Interpretable Machine Learning in Healthcare. In Proceedings of the 2018 ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics 15 August 2018; ACM: Washington, DC, USA, 2018; pp. 559–560. [Google Scholar]
  50. Mehrabi, N.; Morstatter, F.; Saxena, N.; Lerman, K.; Galstyan, A. A Survey on Bias and Fairness in Machine Learning. ACM Comput. Surv. 2022, 54, 1–35. [Google Scholar] [CrossRef]
  51. Miotto, R.; Wang, F.; Wang, S.; Jiang, X.; Dudley, J.T. Deep Learning for Healthcare: Review, Opportunities and Challenges. Brief. Bioinform. 2018, 19, 1236–1246. [Google Scholar] [CrossRef]
  52. Nagendran, M.; Chen, Y.; Lovejoy, C.A.; Gordon, A.C.; Komorowski, M.; Harvey, H.; Topol, E.J.; Ioannidis, J.P.A.; Collins, G.S.; Maruthappu, M. Artificial Intelligence versus Clinicians: Systematic Review of Design, Reporting Standards, and Claims of Deep Learning Studies. BMJ 2020, 368, m689. [Google Scholar] [CrossRef]
  53. Liu, X.; Cruz Rivera, S.; Moher, D.; Calvert, M.J.; Denniston, A.K.; The Spirit-AI and Consort-AI Working Group; Spirit-AI and Consort-AI Steering Group; Chan, A.-W.; Darzi, A.; Holmes, C.; et al. Reporting Guidelines for Clinical Trial Reports for Interventions Involving Artificial Intelligence: The Consort-AI Extension. Nat. Med. 2020, 26, 1364–1374. [Google Scholar] [CrossRef]
  54. Cruz Rivera, S.; Liu, X.; Chan, A.-W.; Denniston, A.K.; Calvert, M.J.; The Spirit-AI and Consort-AI Working Group; Spirit-AI and Consort-AI Steering Group; Darzi, A.; Holmes, C.; Yau, C.; et al. Guidelines for Clinical Trial Protocols for Interventions Involving Artificial Intelligence: The Spirit-AI Extension. Nat. Med. 2020, 26, 1351–1363. [Google Scholar] [CrossRef] [PubMed]
  55. Rajagopal, A.; Ayanian, S.; Ryu, A.J.; Qian, R.; Legler, S.R.; Peeler, E.A.; Issa, M.; Coons, T.J.; Kawamoto, K. Machine Learning Operations in Health Care: A Scoping Review. Mayo Clin. Proc. Digit. Health 2024, 2, 421–437. [Google Scholar] [CrossRef] [PubMed]
  56. Li, Y.; Tian, J.; Xu, A.; Greiner, R.; Hayward, J.; Greenshaw, A.J.; Cao, B. Maturity Framework for Operationalizing Machine Learning Applications in Health Care: Scoping Review. J. Med. Internet Res. 2025, 27, e66559. [Google Scholar] [CrossRef]
  57. Lu, C.; Strout, J.; Gauriau, R.; Wright, B.; Marcruz, F.B.D.C.; Buch, V.; Andriole, K. An Overview and Case Study of the Clinical AI Model Development Life Cycle for Healthcare Systems. arXiv 2020, arXiv:2003.07678. [Google Scholar] [CrossRef]
  58. Bommasani, R.; Hudson, D.A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M.S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. On the Opportunities and Risks of Foundation Models. arXiv 2021, arXiv:2108.07258. [Google Scholar] [CrossRef]
  59. Xi, Z.; Chen, W.; Guo, X.; He, W.; Ding, Y.; Hong, B.; Zhang, M.; Wang, J.; Jin, S.; Zhou, E.; et al. The Rise and Potential of Large Language Model Based Agents: A Survey. Sci. China Inf. Sci. 2023, 68, 12110. [Google Scholar] [CrossRef]
  60. Finlayson, S.G.; Bowers, J.D.; Ito, J.; Zittrain, J.L.; Beam, A.L.; Kohane, I.S. Adversarial Attacks on Medical Machine Learning. Science 2019, 363, 1287–1289. [Google Scholar] [CrossRef]
  61. Moor, M.; Banerjee, O.; Abad, Z.S.H.; Krumholz, H.M.; Leskovec, J.; Topol, E.J.; Rajpurkar, P. Foundation Models for Generalist Medical Artificial Intelligence. Nature 2023, 616, 259–265. [Google Scholar] [CrossRef]
  62. Nori, H.; King, N.; McKinney, S.M.; Carignan, D.; Horvitz, E. Capabilities of GPT-4 on Medical Challenge Problems. arXiv 2023, arXiv:2303.13375. [Google Scholar] [CrossRef]
  63. Babic, B.; Glenn Cohen, I.; Stern, A.D.; Li, Y.; Ouellet, M. A General Framework for Governing Marketed AI/ML Medical Devices. NPJ Digit. Med. 2025, 8, 328. [Google Scholar] [CrossRef]
  64. Lee, P.; Bubeck, S.; Petro, J. Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. N. Engl. J. Med. 2023, 388, 1233–1239. [Google Scholar] [CrossRef]
  65. De Micco, F.; Di Palma, G.; Ferorelli, D.; De Benedictis, A.; Tomassini, L.; Tambone, V.; Cingolani, M.; Scendoni, R. Artificial Intelligence in Healthcare: Transforming Patient Safety with Intelligent Systems—A Systematic Review. Front. Med. 2025, 11, 1522554. [Google Scholar] [CrossRef]
  66. Shabani, M.; Borry, P. Rules for Processing Genetic Data for Research Purposes in View of the New EU General Data Protection Regulation. Eur. J. Hum. Genet. 2018, 26, 149–156. [Google Scholar] [CrossRef]
  67. Raji, I.D.; Smart, A.; White, R.N.; Mitchell, M.; Gebru, T.; Hutchinson, B.; Smith-Loud, J.; Theron, D.; Barnes, P. Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency 27 January 2020; ACM: Barcelona, Spain, 2020; pp. 33–44. [Google Scholar]
  68. Adamson, A.S.; Smith, A. Machine Learning and Health Care Disparities in Dermatology. JAMA Dermatol. 2018, 154, 1247. [Google Scholar] [CrossRef]
  69. Mohsen, F.; Ali, H.; El Hajj, N.; Shah, Z. Artificial Intelligence-Based Methods for Fusion of Electronic Health Records and Imaging Data. Sci. Rep. 2022, 12, 17981. [Google Scholar] [CrossRef]
  70. Khan, N.; Nauman, M.; Almadhor, A.S.; Akhtar, N.; Alghuried, A.; Alhudhaif, A. Guaranteeing Correctness in Black-Box Machine Learning: A Fusion of Explainable AI and Formal Methods for Healthcare Decision-Making. IEEE Access 2024, 12, 90299–90316. [Google Scholar] [CrossRef]
  71. Grande, D.; Luna Marti, X.; Feuerstein-Simon, R.; Merchant, R.M.; Asch, D.A.; Lewson, A.; Cannuscio, C.C. Health Policy and Privacy Challenges Associated with Digital Technology. JAMA Netw. Open 2020, 3, e208285. [Google Scholar] [CrossRef]
  72. Nebeker, C.; Harlow, J.; Espinoza Giacinto, R.; Orozco-Linares, R.; Bloss, C.S.; Weibel, N. Ethical and Regulatory Challenges of Research Using Pervasive Sensing and Other Emerging Technologies: IRB Perspectives. AJOB Empir. Bioeth. 2017, 8, 266–276. [Google Scholar] [CrossRef]
  73. Bouderhem, R. Shaping the Future of AI in Healthcare through Ethics and Governance. Humanit. Soc. Sci. Commun. 2024, 11, 416. [Google Scholar] [CrossRef]
  74. Sandmann, S.; Hegselmann, S.; Fujarski, M.; Bickmann, L.; Wild, B.; Eils, R.; Varghese, J. Benchmark Evaluation of DeepSeek Large Language Models in Clinical Decision-Making. Nat. Med. 2025, 31, 2546–2549. [Google Scholar] [CrossRef]
  75. Obermeyer, Z.; Powers, B.; Vogeli, C.; Mullainathan, S. Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations. Science 2019, 366, 447–453. [Google Scholar] [CrossRef]
  76. Collins, B.X.; Bélisle-Pipon, J.-C.; Evans, B.J.; Ferryman, K.; Jiang, X.; Nebeker, C.; Novak, L.; Roberts, K.; Were, M.; Yin, Z.; et al. Addressing Ethical Issues in Healthcare Artificial Intelligence Using a Lifecycle-Informed Process. JAMIA Open 2024, 7, ooae108. [Google Scholar] [CrossRef]
  77. Murphy, K.; Di Ruggiero, E.; Upshur, R.; Willison, D.J.; Malhotra, N.; Cai, J.C.; Malhotra, N.; Lui, V.; Gibson, J. Artificial Intelligence for Good Health: A Scoping Review of the Ethics Literature. BMC Med. Ethics 2021, 22, 14. [Google Scholar] [CrossRef]
  78. Gorelik, A.J.; Li, M.; Hahne, J.; Wang, J.; Ren, Y.; Yang, L.; Zhang, X.; Liu, X.; Wang, X.; Bogdan, R.; et al. Ethics of AI in Healthcare: A Scoping Review Demonstrating Applicability of a Foundational Framework. Front. Digit. Health 2025, 7, 1662642. [Google Scholar] [CrossRef]
  79. Farhud, D.D.; Zokaei, S. Ethical Issues of Artificial Intelligence in Medicine and Healthcare. Iran. J. Public Health 2021, 50, i–v. [Google Scholar] [CrossRef]
  80. Elendu, C.; Amaechi, D.C.; Elendu, T.C.; Jingwa, K.A.; Okoye, O.K.; John Okah, M.; Ladele, J.A.; Farah, A.H.; Alimi, H.A. Ethical Implications of AI and Robotics in Healthcare: A Review. Medicine 2023, 102, e36671. [Google Scholar] [CrossRef]
  81. Singh, M.P.; Keche, Y.N. Ethical Integration of Artificial Intelligence in Healthcare: Narrative Review of Global Challenges and Strategic Solutions. Cureus 2025, 17, e84804. [Google Scholar] [CrossRef]
  82. Hasan, A.; Prizant, N.; Kim, J.Y.; Rao, S.; Vidal, D.; Shaw, K.; Tobey, D.; Valladares, A.; Zilberstein, S.; Patel, M.; et al. Aligning AI Principles and Healthcare Delivery Organization Best Practices to Navigate the Shifting Regulatory Landscape. NPJ Digit. Med. 2025, 8, 278. [Google Scholar] [CrossRef]
  83. Doshi-Velez, F.; Kim, B. Towards A Rigorous Science of Interpretable Machine Learning. arXiv 2017, arXiv:1702.08608. [Google Scholar] [CrossRef]
  84. Ning, Y.; Teixayavong, S.; Shang, Y.; Savulescu, J.; Nagaraj, V.; Miao, D.; Mertens, M.; Ting, D.S.W.; Ong, J.C.L.; Liu, M.; et al. Generative Artificial Intelligence and Ethical Considerations in Health Care: A Scoping Review and Ethics Checklist. Lancet Digit. Health 2024, 6, e848–e856. [Google Scholar] [CrossRef]
  85. Bansal, G.; Wu, T.; Zhou, J.; Fok, R.; Nushi, B.; Kamar, E.; Ribeiro, M.T.; Weld, D. Does the Whole Exceed Its Parts? The Effect of AI Explanations on Complementary Team Performance. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems 6 May 2021; ACM: Yokohama, Japan, 2021; pp. 1–16. [Google Scholar]
  86. Zhang, J.; Zhang, Z. Ethics and Governance of Trustworthy Medical Artificial Intelligence. BMC Med. Inf. Decis. Mak. 2023, 23, 7. [Google Scholar] [CrossRef]
  87. Mooghali, M.; Stroud, A.M.; Yoo, D.W.; Barry, B.A.; Grimshaw, A.A.; Ross, J.S.; Zhu, X.; Miller, J.E. Trustworthy and Ethical AI-Enabled Cardiovascular Care: A Rapid Review. BMC Med. Inf. Decis. Mak. 2024, 24, 247. [Google Scholar] [CrossRef]
  88. Pham, T. Ethical and Legal Considerations in Healthcare AI: Innovation and Policy for Safe and Fair Use. R. Soc. Open Sci. 2025, 12, 241873. [Google Scholar] [CrossRef]
  89. Mun, S.K.; Wong, K.H.; Lo, S.-C.B.; Li, Y.; Bayarsaikhan, S. Artificial Intelligence for the Future Radiology Diagnostic Service. Front. Mol. Biosci. 2021, 7, 614258. [Google Scholar] [CrossRef] [PubMed]
  90. Ahmed, M.I.; Spooner, B.; Isherwood, J.; Lane, M.; Orrock, E.; Dennison, A. A Systematic Review of the Barriers to the Implementation of Artificial Intelligence in Healthcare. Cureus 2023, 15, e46454. [Google Scholar] [CrossRef]
  91. Mennella, C.; Maniscalco, U.; De Pietro, G.; Esposito, M. Ethical and Regulatory Challenges of AI Technologies in Healthcare: A Narrative Review. Heliyon 2024, 10, e26297. [Google Scholar] [CrossRef]
  92. Ratti, E.; Morrison, M.; Jakab, I. Ethical and Social Considerations of Applying Artificial Intelligence in Healthcare—A Two-Pronged Scoping Review. BMC Med. Ethics 2025, 26, 68. [Google Scholar] [CrossRef]
  93. Dankwa-Mullan, I. Health Equity and Ethical Considerations in Using Artificial Intelligence in Public Health and Medicine. Prev. Chronic Dis. 2024, 21, 240245. [Google Scholar] [CrossRef]
  94. Boudi, A.L.; Boudi, M.; Chan, C.; Boudi, F.B. Ethical Challenges of Artificial Intelligence in Medicine. Cureus 2024, 16, e74495. [Google Scholar] [CrossRef]
  95. Chinta, S.V.; Wang, Z.; Palikhe, A.; Zhang, X.; Kashif, A.; Smith, M.A.; Liu, J.; Zhang, W. AI-Driven Healthcare: A Review on Ensuring Fairness and Mitigating Bias. PLoS Digit. Health 2025, 4, e0000864. [Google Scholar] [CrossRef]
  96. Chen, I.Y.; Pierson, E.; Rose, S.; Joshi, S.; Ferryman, K.; Ghassemi, M. Ethical Machine Learning in Healthcare. Annu. Rev. Biomed. Data Sci. 2021, 4, 123–144. [Google Scholar] [CrossRef]
  97. Elgin, C.Y.; Elgin, C. Ethical Implications of AI-Driven Clinical Decision Support Systems on Healthcare Resource Allocation: A Qualitative Study of Healthcare Professionals’ Perspectives. BMC Med. Ethics 2024, 25, 148. [Google Scholar] [CrossRef] [PubMed]
  98. Thirunavukarasu, A.J.; Ting, D.S.J.; Elangovan, K.; Gutierrez, L.; Tan, T.F.; Ting, D.S.W. Large Language Models in Medicine. Nat. Med. 2023, 29, 1930–1940. [Google Scholar] [CrossRef] [PubMed]
  99. Wang, Y.; Federman, A.; Wurtz, H.; Manchester, M.A.; Morgado, L.; Scipion, C.E.A.; Adjini, M.; Williams, K.; Wills, B.; Helmly, V.; et al. The Evolving Literature on the Ethics of Artificial Intelligence for Healthcare: A PRISMA Scoping Review. Front. Digit. Health 2025, 7, 1701419. [Google Scholar] [CrossRef] [PubMed]
  100. Tang, L.; Li, J.; Fantus, S. Medical Artificial Intelligence Ethics: A Systematic Review of Empirical Studies. Digit. Health 2023, 9, 20552076231186064. [Google Scholar] [CrossRef]
  101. Saheb, T.; Saheb, T.; Carpenter, D.O. Mapping Research Strands of Ethics of Artificial Intelligence in Healthcare: A Bibliometric and Content Analysis. Comput. Biol. Med. 2021, 135, 104660. [Google Scholar] [CrossRef] [PubMed]
  102. SeyedAlinaghi, S.; Habibi, P.; Mehraeen, E. Ethical Considerations for AI Use in Healthcare Research. Healthc. Inform. Res. 2024, 30, 286–289. [Google Scholar] [CrossRef] [PubMed]
  103. Zavaleta-Monestel, E.; Anchía-Alfaro, A.; Rojas-Chinchilla, C.; Quesada-Loria, D.F.; Arguedas-Chacón, S. Ethical and Practical Dimensions of Artificial Intelligence (AI) in Healthcare: A Comprehensive Study of Professional Perceptions. Cureus 2025, 17, e78416. [Google Scholar] [CrossRef]
  104. Park, M.K.; Ashwood, N.; Capes, N. Ethics of Artificial Intelligence in Medicine. Cureus 2025, 17, e83567. [Google Scholar] [CrossRef]
  105. Ayers, J.W.; Poliak, A.; Dredze, M.; Leas, E.C.; Zhu, Z.; Kelley, J.B.; Faix, D.J.; Goodman, A.M.; Longhurst, C.A.; Hogarth, M.; et al. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Intern. Med. 2023, 183, 589. [Google Scholar] [CrossRef]
  106. Abdullahi, T.; Singh, R.; Eickhoff, C. Learning to Make Rare and Complex Diagnoses with Generative AI Assistance: Qualitative Study of Popular Large Language Models. JMIR Med. Educ. 2024, 10, e51391. [Google Scholar] [CrossRef]
  107. MacIntyre, M.R.; Cockerill, R.G.; Mirza, O.F.; Appel, J.M. Ethical Considerations for the Use of Artificial Intelligence in Medical Decision-Making Capacity Assessments. Psychiatry Res. 2023, 328, 115466. [Google Scholar] [CrossRef]
  108. Wang, L.; Ma, C.; Feng, X.; Zhang, Z.; Yang, H.; Zhang, J.; Chen, Z.; Tang, J.; Chen, X.; Lin, Y.; et al. A Survey on Large Language Model Based Autonomous Agents. Front. Comput. Sci. 2024, 18, 186345. [Google Scholar] [CrossRef]
  109. Jeblick, K.; Schachtner, B.; Dexl, J.; Mittermeier, A.; Stüber, A.T.; Topalis, J.; Weber, T.; Wesp, P.; Sabel, B.O.; Ricke, J.; et al. ChatGPT Makes Medicine Easy to Swallow: An Exploratory Case Study on Simplified Radiology Reports. Eur. Radiol. 2023, 34, 2817–2825. [Google Scholar] [CrossRef]
  110. Andrew, A. Potential Applications and Implications of Large Language Models in Primary Care. Fam. Med. Com. Health 2024, 12, e002602. [Google Scholar] [CrossRef] [PubMed]
Figure 1. WHO Ethical Principles for artificial intelligence in Health.
Figure 1. WHO Ethical Principles for artificial intelligence in Health.
Aimed 01 00016 g001
Figure 2. Five-stage lifecycle of healthcare AI systems.
Figure 2. Five-stage lifecycle of healthcare AI systems.
Aimed 01 00016 g002
Figure 3. Cross-Cutting Governance Enablers for Healthcare AI.
Figure 3. Cross-Cutting Governance Enablers for Healthcare AI.
Aimed 01 00016 g003
Figure 4. Governance-by-design framework for healthcare AI.
Figure 4. Governance-by-design framework for healthcare AI.
Aimed 01 00016 g004
Figure 5. Translating governance-by-design into real-world clinical impact.
Figure 5. Translating governance-by-design into real-world clinical impact.
Aimed 01 00016 g005
Table 1. Summary of Narrative Review Search Strategy and Case Selection Criteria.
Table 1. Summary of Narrative Review Search Strategy and Case Selection Criteria.
ComponentDescription
Review TypeMultidisciplinary narrative review
Databases SearchedPubMed, Web of Science, IEEE Xplore, Scopus, arXiv
Publication TimeframePrimarily 2020–2025 with inclusion of foundational earlier studies
Principal Search ConceptsHealthcare AI, machine learning, ethics, governance, fairness, explainability, validation, regulation, GDPR, FDA, EU AI Act
Included Literature TypesEthical frameworks, regulatory guidance documents, technical standards, reporting guidelines, peer-reviewed studies, and implementation case studies.
Case Selection CriteriaDocumented governance failures, independent analysis availability, lifecycle relevance, and real-world healthcare AI implementation challenges
Table 2. WHO ethical principles for AI in health and their core governance requirements.
Table 2. WHO ethical principles for AI in health and their core governance requirements.
PrincipleDefinitionCore RequirementsPrimary Lifecycle Focus
Human AutonomyPreserve meaningful human control in decision-makingInformed consent; privacy safeguards; clinician overrideData collection; deployment
Well-being and SafetyEnsure clinical benefit while minimizing harmValidation; adverse event reporting; continuous monitoringValidation; monitoring
Transparency and ExplainabilityEnsure system processes and outputs are understandableDataset documentation; explainability tools; reporting standardsAll stages
AccountabilityAssign clear responsibility across stakeholdersAudit trails; documentation; regulatory oversightAll stages
Equity and InclusivenessEnsure fair and unbiased outcomes across populationsRepresentative data; fairness metrics; bias mitigationData; validation; deployment
Responsiveness and SustainabilitySupport adaptive governance and long-term system viabilityEnvironmental assessment; adaptive governanceDevelopment; monitoring
Table 3. Comparison of International AI Governance Frameworks.
Table 3. Comparison of International AI Governance Frameworks.
FrameworkIssuing BodyYearScopeHealthcare SpecificityKey Principles/ElementsEnforcement MechanismLifecycle Coverage
WHO GuidanceWorld Health Organization2021Healthcare AI ethics and governanceHighAutonomy; well-being and safety; transparency; accountability; equity and inclusiveness; sustainabilityNormative guidanceComprehensive across lifecycle
OECD PrinciplesOECD2019General AI ethicsLowInclusive growth; human-centered values; transparency; robustness; accountabilityVoluntary adoptionHigh-level lifecycle principles
FDA GMLPU.S. FDA2021ML-based medical device regulationHighIntended use clarity; data quality; validation; monitoring; transparencyRegulatory authorityDevelopment validation, post-market
EU AI ActEuropean Union2024Risk-based AI regulationMediumRisk tiers; governance obligations; oversight; monitoringLegal EnforcementComprehensive high-risk AI
IEEE Ethically Aligned DesignIEEE2019Technical ethics guidanceMediumHuman rights; well-being; accountability; transparencyVoluntary standardsDesign and development emphasis
NAM AI CodeNational Academy of Medicine2023Healthcare AI principlesHighSafe and effective; equitable; trustworthy; human-centered; sustainableProfessional guidanceComprehensive lifecycle orientation
Table 4. AI Reporting Standards for Healthcare.
Table 4. AI Reporting Standards for Healthcare.
StandardFull NameYearPurposeKey Reporting ElementsTarget Audience
TRIPOD-AITransparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis—AI Extension2024Reporting prediction model development and validation27 items covering study design, data sources, model development, validation, and resultsResearchers; journal editors
SPIRIT-AIStandard Protocol Items: Recommendations for Interventional Trials—AI Extension2020Clinical trial protocol designAlgorithm description; version control; intended use; performance characteristicsTrial designers; funders
CONSORT-AIConsolidated Standards of Reporting Trials—AI Extension2020RCT reporting for AI interventionsIntervention details; implementation fidelity; algorithm trackingClinical researchers
DECIDE-AIDevelopmental and Exploratory Clinical Investigations of Decision Support Systems Driven by AI2021Early-stage AI transparencyIntended context; development methods; preliminary evaluationEarly-stage developers; regulators
STARD-AIStandards for Reporting Diagnostic Accuracy Studies—AI Extension2023Diagnostic AI evaluationDiagnostic evaluation methods; reference standards; study populationDiagnostic researchers
Table 5. Ethical governance outcomes with and without lifecycle operationalization.
Table 5. Ethical governance outcomes with and without lifecycle operationalization.
DimensionWithout Lifecycle OperationalizationWith Governance-by-Design
DataLimited consent transparency; potential privacy violationsStructured consent governance; privacy-preserving data practices
Bias and fairnessBias detected late or after deploymentBias detection and mitigation integrated during development
TransparencyLimited documentation and explainabilityModel cards, dataset documentation, and explainability tools integrated
ValidationReliance on internal performance metricsExternal validation and subgroup performance analysis
DeploymentPoor workflow integration; automation bias risksHuman oversight and clinician-in-the-loop design
MonitoringMinimal post-deployment monitoringContinuous performance surveillance and drift detection
AccountabilityDiffuse responsibility across stakeholdersDefined governance roles and audit mechanisms
Table 6. EU AI Act Risk Categories and Healthcare AI Classification.
Table 6. EU AI Act Risk Categories and Healthcare AI Classification.
Risk CategoryDefinitionHealthcare ExamplesKey RequirementsConformity AssessmentPenalties
Unacceptable RiskSystems violating fundamental rightsSocial scoring; exploitative targetingProhibitedNot applicableSignificant penalties
High RiskSystems affecting safety or fundamental rightsDiagnostic AI; treatment recommendation systems; AI-based medical devicesRisk management; data governance; documentation; transparency; human oversight; monitoringThird-party conformity assessment integrated with MDR/IVDRSubstantial penalties
Limited RiskSystems requiring transparencyHealth chatbots; symptom checkersTransparency disclosureSelf-assessmentLower penalties
Minimal RiskLow-impact systemsAdministrative toolsNo specific AI Act requirementsNoneNot applicable
Table 7. Operationalization Matrix: WHO Principles Across AI Lifecycle Stages.
Table 7. Operationalization Matrix: WHO Principles Across AI Lifecycle Stages.
Lifecycle StageAutonomyWell-Being & SafetyTransparencyAccountabilityEquity & InclusivenessSustainability
Data CollectionInformed consent; privacy safeguards; data subject rightsData quality protocols; representative samplingDataset documentation; provenance trackingData stewardship; ethics review; access governanceRepresentative sampling; addressing historical gapsEfficient data practices; minimal collection
Model DevelopmentHuman-in-the-loop design; override capabilitySafety constraints; robustness testing; uncertainty quantificationModel cards; explainability integrationVersion control; quality systemsFairness metrics; bias mitigationResource-efficient training; model compression
ValidationTrial informed consent; opt-in protocolsExternal validation; safety benchmarking; adverse event monitoringTRIPOD-AI reporting; disaggregated resultsIndependent validation; regulatory reviewSubgroup analysis; equity impact assessmentReusable test sets; federated validation
DeploymentClinician override; patient disclosureWorkflow integration; fail-safe mechanismsClinician/patient explanations; documentationClinical governance; liability frameworksEquitable access; disparity monitoringInteroperability; cost-effectiveness
MonitoringFeedback loops; grievance mechanismsDrift detection; performance trackingDashboards; public reportingPost-market surveillance; auditsLongitudinal equity monitoringAdaptive updating; sustainable infrastructure
Table 8. Operationalization of Governance-by-Design Across the Healthcare AI Lifecycle.
Table 8. Operationalization of Governance-by-Design Across the Healthcare AI Lifecycle.
AI Lifecycle StageGovernance MechanismsKey StakeholdersOperational Workflow Integration
Data Collection and CurationDataset audits, data provenance documentation, subgroup representation analysis, proxy bias review, privacy and compliance safeguardsData scientists, clinical domain experts, health equity specialists, institutional review boards, compliance teamsData engineering pipelines, preprocessing workflows, dataset approval checkpoints
Model Development and TrainingFairness testing, explainability assessment, reproducibility documentation, robustness testing, model cards, version controlML engineers, clinicians, AI governance committees, MLOps teams, institutional leadershipModel development environments, CI/CD pipelines, model registries
Validation and Clinical EvaluationExternal validation, subgroup-specific performance analysis, calibration testing, clinician usability assessment, deployment approval reviewClinical evaluators, healthcare quality teams, governance committees, patient safety officesPre-deployment validation workflows, clinical simulation and evaluation procedures
Deployment and Clinical IntegrationClinician-in-the-loop oversight, audit logging, deployment approval checkpoints, user training, incident escalation proceduresClinicians, care management teams, MLOps engineers, institutional leadership, compliance officesClinical decision support systems, deployment orchestration workflows, operational governance processes
Monitoring and Continuous LearningDrift detection, subgroup fairness monitoring, post-deployment surveillance, periodic recalibration, adverse event reportingMLOps teams, healthcare quality and safety offices, clinicians, governance committees, regulatorsMonitoring dashboards, observability infrastructure, retraining and incident response workflows
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Saraboji, K.; Gopalakrishnan, K.; Sood, D.; Kaur, A.; Shivaram, S.; Helgeson, S.A.; Arunachalam, S.P.; Mitra, D. Operationalizing WHO Ethical Principles for Healthcare AI: A Lifecycle-Aligned Governance-by-Design Framework. AI Med. 2026, 1, 16. https://doi.org/10.3390/aimed1020016

AMA Style

Saraboji K, Gopalakrishnan K, Sood D, Kaur A, Shivaram S, Helgeson SA, Arunachalam SP, Mitra D. Operationalizing WHO Ethical Principles for Healthcare AI: A Lifecycle-Aligned Governance-by-Design Framework. AI in Medicine. 2026; 1(2):16. https://doi.org/10.3390/aimed1020016

Chicago/Turabian Style

Saraboji, Kaaviyashri, Keerthy Gopalakrishnan, Divyanshi Sood, Anmolpreet Kaur, Suganti Shivaram, Scott A. Helgeson, Shivaram P. Arunachalam, and Dipankar Mitra. 2026. "Operationalizing WHO Ethical Principles for Healthcare AI: A Lifecycle-Aligned Governance-by-Design Framework" AI in Medicine 1, no. 2: 16. https://doi.org/10.3390/aimed1020016

APA Style

Saraboji, K., Gopalakrishnan, K., Sood, D., Kaur, A., Shivaram, S., Helgeson, S. A., Arunachalam, S. P., & Mitra, D. (2026). Operationalizing WHO Ethical Principles for Healthcare AI: A Lifecycle-Aligned Governance-by-Design Framework. AI in Medicine, 1(2), 16. https://doi.org/10.3390/aimed1020016

Article Metrics

Back to TopTop