5.1. Convergence Across Frameworks
This review identifies a substantial degree of convergence across international ethical, regulatory, and technical frameworks governing healthcare AI [
83,
84,
85,
86]. Foundational initiatives, including WHO ethical principles, FDA Good Machine Learning Practice (GMLP), the EU AI Act, OECD recommendations, and guidance from professional societies, consistently emphasize safety, transparency, accountability, and fairness as core requirements for trustworthy AI [
87]. This alignment suggests the emergence of a shared global baseline for responsible AI in healthcare and reflects a growing consensus that ethical governance must be embedded throughout the AI lifecycle rather than applied at isolated stages [
31].
Healthcare-specific frameworks, such as WHO guidance, FDA GMLP, and recommendations from the National Academy of Medicine, offer more operational detail compared to cross-sectoral principles developed by organizations such as the OECD or IEEE [
43,
88]. This increased specificity reflects the safety-critical nature of clinical decision-making, where even minor system failures can have significant consequences. At the same time, regulatory instruments such as the EU AI Act and FDA oversight pathways introduce enforceable requirements that formalize these expectations, although they may evolve more slowly than rapidly advancing AI technologies [
39]. This creates a dynamic tension between regulatory stability and technological change, requiring governance models that can remain both robust and adaptable over time.
In contrast, ethical guidelines and professional standards tend to be more flexible and adaptive, allowing them to respond more quickly to emerging challenges, including those associated with generative AI and continuously learning systems [
49,
89]. However, their non-binding nature limits enforceability, creating a gap between normative expectations and practical implementation [
90]. Reporting standards such as TRIPOD-AI, CONSORT-AI, and SPIRIT-AI help bridge this gap by improving methodological rigor, transparency, and reproducibility in evidence generation [
53,
54]. These standards play a critical role in translating high-level principles into measurable and reportable practices, although adherence remains inconsistent across the literature [
52].
Despite this convergence, alignment at the level of principles does not automatically translate into consistency in implementation [
19,
50,
78]. Differences in terminology, scope, and enforcement mechanisms across frameworks can create ambiguity for developers and healthcare organizations navigating multiple regulatory environments [
76]. This challenge is particularly evident in cross-border deployment scenarios, where systems must simultaneously satisfy overlapping but not fully harmonized requirements. Fragmentation across governance regimes can also lead to duplication of effort, inconsistent validation practices, and uncertainty in accountability allocation [
51].
Unlike many existing trustworthy AI frameworks that primarily emphasize high-level ethical principles or isolated governance requirements, the framework proposed in this review adopts a lifecycle-aligned operational structure that explicitly maps ethical principles to stage-specific governance mechanisms across data collection, model development, validation, deployment, and post-deployment monitoring [
26,
82,
88]. Rather than treating governance as a static compliance exercise, the framework integrates technical safeguards, organizational oversight, and continuous monitoring processes throughout the AI lifecycle [
55,
56,
75]. This approach extends beyond principle-based guidance by emphasizing operational governance checkpoints, multidisciplinary accountability, and continuous lifecycle integration within real-world healthcare AI workflows.
The lifecycle-aligned governance framework developed in this review builds on this convergence by integrating ethical principles, regulatory expectations, technical methods, and stakeholder engagement into a unified operational structure. By explicitly mapping principles to stage-specific governance mechanisms, the framework provides a practical pathway for translating shared normative commitments into consistent and actionable practices across the AI lifecycle [
80,
91]. Importantly, it shifts the focus from alignment in principle to alignment in implementation, enabling organizations to operationalize ethical requirements within real-world clinical and technical workflows. In this way, the framework contributes not only to conceptual clarity but also to practical governance, addressing a critical gap between existing guidance and implementation in healthcare AI systems.
5.2. Persistent Gaps in Ethical Operationalization
Despite broad convergence at the level of ethical principles and regulatory intent, substantial gaps remain in practical implementation [
92,
93]. These gaps reflect the inherent difficulty of translating high-level commitments into consistent, enforceable practices within complex and distributed healthcare AI ecosystems. Prior work has emphasized that principle-based frameworks alone are insufficient without operational mechanisms, measurable criteria, and lifecycle integration strategies [
33].
First, accountability fragmentation persists due to the distributed nature of AI development and deployment [
43,
94]. Multiple stakeholders, including developers, vendors, healthcare organizations, clinicians, and regulators, each exercise partial control over system design, implementation, and use. When harms occur, responsibility may become diffuse, complicating both liability determination and mechanisms for redress. While existing frameworks emphasize accountability, they often lack clear operational models for allocating responsibility across the AI lifecycle [
92]. In practice, this can result in governance gaps where no single actor assumes full ownership of downstream outcomes, particularly in systems involving third-party vendors or continuously updated models. These challenges are further intensified by cross-jurisdictional deployment and regulatory fragmentation, which can obscure responsibility boundaries and complicate enforcement [
39,
78].
Second, equity evaluation remains underdeveloped relative to the strength of normative commitments [
19,
78]. Although fairness is widely recognized as a core principle, there is limited consensus on how it should be measured, enforced, and monitored in clinical settings. Different fairness metrics may produce conflicting results, and guidance on selecting context-appropriate metrics remains limited [
50]. Many validation studies continue to report aggregate performance metrics without sufficient subgroup analysis, obscuring disparities across demographic and clinical populations. In addition, formal equity impact assessments are rarely mandated prior to deployment, limiting the ability to proactively identify and mitigate potential harms. Empirical evidence from healthcare AI studies demonstrates that such gaps can lead to systematic underperformance in marginalized populations, reinforcing existing inequities in care [
16,
17]. Given the potential for AI systems to amplify structural disparities, this gap represents a critical ethical and clinical concern.
Third, post-market surveillance infrastructure remains comparatively underdeveloped [
76]. Regulatory and institutional efforts are often concentrated on pre-deployment validation, while continuous monitoring receives less structured attention [
30,
95]. In real-world settings, however, AI systems are exposed to evolving data distributions, changing clinical practices, and new patient populations. These dynamics can lead to model drift, context shift, and emergent biases that degrade performance over time [
24,
52]. Without robust and standardized monitoring mechanisms, such degradation may remain undetected until it results in patient harm or system failure. Current approaches to post-deployment monitoring are often fragmented or ad hoc, lacking the consistency and rigor applied during earlier stages of development. Broader analyses of real-world AI deployment further confirm that monitoring remains one of the weakest links in the lifecycle of healthcare AI systems [
95].
More broadly, these gaps highlight a structural disconnect between ethical intent and operational practice. Existing frameworks provide strong normative guidance but insufficient direction on how to embed these principles into day-to-day development, deployment, and monitoring processes. This disconnect reflects a broader challenge in AI governance, where alignment at the level of principles does not automatically translate into alignment in implementation.
Taken together, these challenges reinforce the need for lifecycle-integrated governance approaches. Addressing accountability fragmentation, strengthening equity evaluation, and establishing robust post-market surveillance systems require coordinated mechanisms that span all stages of the AI lifecycle [
80,
91]. Without such integration, ethical principles risk remaining aspirational rather than actionable in real-world healthcare settings.
5.3. Implementation Barriers
Even when governance frameworks are conceptually robust, translating them into practice remains challenging due to systemic and organizational barriers [
66,
71,
96]. These barriers extend beyond technical limitations and reflect broader constraints related to resources, coordination, institutional capacity, and the pace of technological change. Prior work has highlighted that implementation gaps often arise not from a lack of guidance, but from the difficulty of integrating governance requirements into complex real-world healthcare environments [
83,
84,
87].
Capacity limitations represent a major obstacle, particularly in low-resource settings and smaller healthcare institutions [
72,
73]. Effective governance requires multidisciplinary expertise spanning data science, clinical practice, ethics, and regulatory compliance, as well as technical infrastructure for monitoring, documentation, and continuous model evaluation. Sustaining these capabilities over time demands ongoing investment and institutional commitment. In practice, many organizations lack the necessary workforce, infrastructure, or funding to implement comprehensive governance frameworks. As a result, the ability to operationalize ethical AI is often concentrated in well-resourced institutions, raising concerns that AI innovation may disproportionately benefit already advantaged healthcare systems while widening global inequities [
64,
74]. This imbalance highlights the need for scalable and accessible governance models that can be adapted across diverse healthcare contexts [
71].
Governance fragmentation further complicates implementation [
39,
88,
92]. Developers and healthcare organizations must navigate overlapping and sometimes inconsistent requirements across international ethical guidelines, national regulations, regional data protection laws, professional standards, and institutional policies. In practice, these frameworks may differ in scope, terminology, and enforcement mechanisms, creating uncertainty about how to achieve compliance across jurisdictions. This challenge is particularly pronounced in cross-border deployment, where systems must satisfy multiple regulatory regimes simultaneously [
65,
67]. Without greater harmonization, fragmentation increases compliance burden, slows deployment, and introduces ambiguity in responsibility allocation across stakeholders. Broader analyses of AI governance have shown that such fragmentation can lead to duplication of validation efforts, inconsistent documentation practices, and gaps in accountability across the AI lifecycle [
33,
68].
A further challenge arises from the tension between innovation and regulation [
69,
94]. AI technologies evolve rapidly, with new architectures, training paradigms, and deployment models emerging at a pace that often outstrips regulatory adaptation. Regulatory systems originally designed for static medical devices may struggle to accommodate adaptive, continuously learning, or foundation model-based systems that evolve after deployment. This mismatch creates a difficult balance: overly restrictive regulation may hinder beneficial innovation and delay clinical adoption, while insufficient oversight increases the risk of deploying inadequately validated or unsafe systems. Recent advances in generative AI further intensify this tension, as these systems introduce new uncertainties related to reliability, explainability, and clinical appropriateness [
63,
90]. These developments underscore the need for governance approaches that are both adaptive and risk-sensitive, capable of evolving alongside technological change without compromising patient safety.
These challenges are supported by a growing body of literature highlighting persistent gaps in multiple dimensions of healthcare AI governance. Studies on technical robustness and system reliability demonstrate vulnerabilities related to instability, adversarial inputs, and performance degradation in real-world settings [
60,
70]. In parallel, research on fairness evaluation highlights ongoing challenges in defining, measuring, and operationalizing equity across diverse clinical populations [
19,
50,
78]. Additional work has identified limitations in existing governance frameworks and post-deployment monitoring practices, emphasizing the lack of consistent, lifecycle-integrated oversight mechanisms [
51,
76].
Taken together, these findings highlight the need for adaptive governance models that can evolve alongside technological innovation while maintaining safety, accountability, and ethical integrity [
80,
91]. Such models must be flexible enough to accommodate emerging technologies yet structured enough to provide clear guidance for implementation across diverse healthcare contexts. In particular, there is a need for governance approaches that integrate technical tools, regulatory requirements, and organizational processes into cohesive systems capable of supporting lifecycle-aligned oversight in real-world clinical environments.
5.4. The Value of Governance-by-Design
The governance-by-design approach proposed in this review addresses persistent gaps in healthcare AI oversight by embedding ethical requirements throughout the AI lifecycle rather than confining evaluation to isolated checkpoints [
66,
71]. By integrating safeguards at each stage, potential risks can be identified and mitigated early, when intervention remains technically feasible and organizationally practical. This approach distributes responsibility across stakeholders while maintaining clear accountability pathways, reducing ambiguity in how ethical obligations are assigned and enforced. Importantly, this lifecycle-oriented model aligns with emerging perspectives that emphasize continuous risk management, adaptive validation, and system-level governance as essential components of trustworthy AI [
96].
Evidence from real-world implementations highlights the consequences of fragmented governance [
69]. Cases such as the DeepMind–Royal Free collaboration and proprietary clinical prediction systems that demonstrated degraded performance outside their development environments illustrate how failures at early lifecycle stages can propagate downstream harms [
97]. These failures were not solely technical in nature; they reflected deficiencies in consent governance, insufficient external validation, and inadequate post-deployment monitoring. From a technical perspective, such failures are often linked to issues such as dataset bias, lack of distributional robustness, and absence of drift detection mechanisms, all of which can lead to performance degradation in real-world settings [
97]. A lifecycle-integrated governance structure could have identified these vulnerabilities earlier and reduced the likelihood of harm.
Governance-by-design also reframes ethical oversight as enabling infrastructure rather than an external constraint [
64,
74]. By embedding transparency, fairness evaluation, and monitoring mechanisms directly into development workflows, ethical considerations become part of routine quality assurance rather than an afterthought. Technically, this includes integrating bias detection pipelines, explainability modules, uncertainty quantification, and validation protocols into model development and deployment processes. This alignment supports regulatory readiness, enhances institutional trust, and strengthens clinical legitimacy while ensuring that ethical safeguards are systematically enforced rather than retrofitted.
From a technical standpoint, governance-by-design enables the integration of key methodological safeguards across the AI lifecycle [
46,
47,
48,
49,
89]. During data collection, representational auditing and dataset documentation (e.g., datasheets for datasets) improve transparency and mitigate bias. During model development, fairness-aware optimization, adversarial robustness testing, and uncertainty estimation techniques enhance model reliability and safety. Explainability methods such as SHAP and LIME support interpretability, although their limitations require careful clinical contextualization [
21,
47]. During validation, external validation protocols, subgroup analysis, and standardized reporting frameworks (TRIPOD-AI, CONSORT-AI) improve generalizability and transparency [
9,
53]. At deployment, workflow-aware system design, clinician-in-the-loop oversight, and fail-safe mechanisms reduce risks such as automation bias and misinterpretation. Finally, continuous monitoring mechanisms, including drift detection, performance dashboards, and adaptive retraining strategies, ensure long-term system reliability and safety.
However, governance-by-design requires sustained organizational commitment, resource allocation, and cultural change [
67,
68]. It demands not only technical tools but also formal governance structures, clearly defined responsibility chains, and meaningful stakeholder engagement processes. The framework presented here provides a structured roadmap, but its effectiveness ultimately depends on institutional uptake, integration into clinical workflows, and ongoing adaptation as technologies and healthcare environments evolve.
In real-world healthcare settings, many failures of AI systems have emerged not from algorithmic limitations alone, but from gaps in governance across the lifecycle. Issues such as biased training data, lack of external validation, poor workflow integration, and insufficient post-deployment monitoring have been repeatedly observed in clinical AI deployments. The governance-by-design framework directly addresses these challenges by embedding safeguards at each stage of system development and use. For example, representational auditing during data collection, fairness-aware optimization during model development, and subgroup-specific validation protocols help mitigate bias before deployment [
33]. Similarly, structured external validation and continuous monitoring mechanisms reduce the risk of performance degradation across institutions and patient populations.
The framework also responds to practical challenges faced by clinicians and healthcare organizations. In clinical environments, AI systems must function within time-constrained workflows and support decision-making without increasing cognitive burden. Governance-by-design incorporates clinician-in-the-loop design, explainability mechanisms, and workflow-aligned deployment strategies to improve usability and reduce risks such as automation bias and alert fatigue [
16,
54,
91]. In addition, continuous feedback loops and monitoring systems enable healthcare organizations to detect emerging issues, adapt models over time, and integrate AI oversight into existing quality improvement processes. These mechanisms ensure that governance is not a one-time requirement but an ongoing component of clinical practice.
From a broader systems perspective, the framework supports alignment with evolving regulatory and technological landscapes. As healthcare increasingly adopts advanced AI systems, including generative models and foundation models, new risks related to hallucination, instability, and context sensitivity have emerged [
58,
61,
63,
98]. These systems require additional safeguards, including output verification, human oversight, and continuous validation across diverse clinical contexts [
59]. Governance-by-design provides a structured approach to managing these risks through lifecycle-integrated monitoring and adaptive governance mechanisms. By aligning ethical principles with real-world implementation processes, the framework offers a practical pathway for translating normative guidance into safe, scalable, and trustworthy AI deployment in clinical medicine.
In this sense, governance-by-design positions ethical AI not as a barrier to innovation, but as a foundational infrastructure for safe, equitable, and scalable healthcare AI. By embedding ethical considerations into the full lifecycle of system development and deployment, it enables innovation that is both responsible and sustainable.
Figure 5 illustrates how governance-by-design addresses key real-world challenges in healthcare AI and enables improved clinical and system-level outcomes through lifecycle-integrated governance.
Despite the strengths of this multidisciplinary narrative synthesis, several limitations should be acknowledged. First, the review was narrative rather than systematic and therefore was not intended to provide exhaustive literature coverage. Although structured searches and representative case examples were incorporated, thematic synthesis and study selection involved interpretive judgment, which may introduce selection bias. In addition, the proposed governance-by-design framework remains conceptual and has not yet been prospectively evaluated across real-world institutional implementation settings. Future studies should focus on validating and operationalizing the framework through measurable governance indicators, implementation toolkits, and audit-oriented assessment approaches across diverse healthcare environments.