Next Article in Journal
A Hybrid Anomaly Detection Framework for Reliable Physiological Signal Extraction in Multimodal Wearable Sleep Monitoring
Previous Article in Journal
Explainable Artificial Intelligence for Predicting Gastrointestinal Adverse Effects of GLP-1 Receptor Agonists
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Data Stewardship Barriers to Building Digital Twin Technology for Precision Medicine

1
Texas A&M Institute for Bioscience and Technology, Houston, TX 77030, USA
2
Texas A&M Health, Bryan, TX 77807, USA
3
Department of Translational Medical Sciences, Vashisht College of Medicine, Texas A&M University, Houston, TX 77030, USA
4
Section of Visual Computing & Computational Media, College of Performance, Visualization & Fine Arts, Texas A&M University, College Station, TX 77843, USA
5
Texas A&M Institute of Data Science, College Station, TX 77843, USA
6
School of Engineering Medicine, Texas A&M University, Houston, TX 77030, USA
7
Department of Statistics, College of Arts and Sciences, Texas A&M University, College Station, TX 77843, USA
8
Department of Medical Physiology Medicine, Vashisht College of Medicine, Texas A&M University, Bryan, TX 77807, USA
9
Department of Primary Care & Rural Medicine, Vashisht College of Medicine, Texas A&M University, Bryan, TX 77807, USA
*
Author to whom correspondence should be addressed.
AI Med. 2026, 1(3), 20; https://doi.org/10.3390/aimed1030020
Submission received: 25 March 2026 / Revised: 26 May 2026 / Accepted: 9 June 2026 / Published: 27 July 2026

Abstract

Digital twins (DTs) are dynamic, virtual representations of individual patients that could support predictive diagnostics and personalized therapeutic optimization. Their development depends on patient-level data from real-world data (RWD) sources and electronic medical record (EMR) data, but major barriers in data-stewardship, interoperability, provenance, and governance persist. This review examines critical bottlenecks within the current healthcare ecosystem, with particular attention to fragmented EMRs, limited longitudinal data continuity, and missing, incomplete, or inaccurate information. We address data availability, stewardship, and provenance rather than model construction itself. We also explore ethical imperatives of mitigating representational bias, where over-representation of European ancestry amplifies existing health inequalities. We evaluate blockchain technology as a decentralized trust anchor to ensure provenance, automate consent, and incentivize longitudinal data stewardship. In parallel, we acknowledge that clinically useful DTs also depend on substantial advances in model specification, calibration, and validation, especially for biologically complex diseases. By synthesizing computational, regulatory, and ethical challenges, we provide a roadmap for developing robust, equitable, and interoperable ecosystems for precision medicine.

1. Introduction

Precision medicine relies on the high-fidelity integration of clinical, molecular, environmental, lifestyle, and behavioral data to tailor interventions at the individual patient level. At the core of this evolution is digital twin (DT) technology, which uses real-time and longitudinal data to create computational representations of patients that can support simulation, prediction, and therapeutic optimization [1]. Unlike static predictive algorithms, DTs are intended to evolve as patient data change over time [2,3,4]. If successfully implemented, DTs could help clinicians compare treatment scenarios, support shared decision-making, and improve care planning at both the individual and population levels [5].
Evidence for DT applications is emerging across several domains, although the maturity of these systems varies substantially. In oncology, DT models have been tested on retrospective cases for patient-specific tumor dose planning [6]. Cardiac DTs model hemodynamic responses in heart failure management are in the concept stage of development [7]. In diabetes, DTs simulating glucose dynamics under various insulin regimens [8] have been evaluated in close loop systems. In Alzheimer’s disease clinical trials DTs have enabled synthetic controls in clinical trials [9] which has undergone regulatory qualification with the European Medicines Agency (EMA).
Clinical implementation of DT requires an unprecedented degree of data fluidity [10]. Medical data has emerged as a potentially transformative production factor, comparable to land and labor, yet it remains siloed (Figure 1) within proprietary institutional architectures [11]. Current data stewardship practices are characterized by fragmentation and institutional risk-aversion [11,12,13,14]. Regulations such as the Health Insurance Portability and Accountability Act (HIPAA) and the General Data Protection Regulation (GDPR), while essential for privacy, often lead to a “don’t share” posture due to inherent risk liability in sharing and the high cost and complexity of auditing data provenance across heterogeneous ecosystems [12,15].
As a result, these frameworks structurally prioritize privacy over portability. Sharing frictions stifle the collaborative environment required for DT training and model refreshing. Greater data liquidity and incentives for sharing and tools for multiagent stewardship are warranted [11]. Realizing this vision confronts profound technical, organizational, and ethical challenges that are fundamentally multi-agent in character. Healthcare data generation, storage, and utilization occur across a distributed ecosystem of hospitals, laboratories, pharmacies, payer organizations, device manufacturers, patients, researchers, and regulators, each operating under distinct incentives, governance frameworks, and technical standards.
In this manuscript, we use the term electronic health record (EHR) to describe the longitudinal record assembled across organizations, and electronic medical record (EMR) to describe records maintained within a single institution or health system. In principle, DT development is best supported by EHR-scale longitudinal data; in practice, many current efforts remain limited to EMR-bounded datasets because they are more tractable under existing governance structures.
Representational bias further threatens the integrity and fairness of clinical DTs. Existing clinicogenomic databases over-represent individuals of European ancestry, a limitation that risks training AI models that amplify rather than mitigate health inequities [13,16]. Addressing these challenges requires governance frameworks that can enable the interactions between patients, providers, and research institutions across time [17,18].
In this review, we argue that improving data stewardship and governance are prerequisites for scalable and equitable DT development but are not the only barriers to clinical translation [12,19]. We examine emerging technologies, such as federated learning and blockchain, as a means to track provenance and consent, thereby providing the “trust anchor” necessary for decentralized scientific collaboration [20,21,22]. We recognize that substantial technical challenges in model specification, calibration, validation, and drift management remain critical, even when data-sharing barriers are addressed. Our primary emphasis, however, is on the data-governance foundations required for DT systems to become clinically credible and socially trustworthy.

2. Real-World Data Silos: The Fragmentation of Clinical Truth

A fundamental tension exists between the rigor of prospective clinical trial data and the breadth of real-world observational data (Table 1). Clinical trials impose stringent inclusion and exclusion criteria, standardized treatment protocols, regular monitoring schedules, and adjudicated outcome assessments, yielding high-quality, internally valid datasets. However, these datasets typically involve small, homogeneous samples followed for limited periods of time, restricting their generalizability and external validity to broader populations and longer-term predictions [9]. Real-world data (RWD) offer greater diversity and longer follow-up, but lack appropriate controls and the data quality assurances afforded by trials [23]. For DT, this tradeoff is acute: models trained on trial data may fail to generalize to routine practice, while models trained on RWD likely inherit biases and errors embedded in observational records unless appropriate statistical techniques are implemented to account for prominent sources of bias and error in the analysis [24,25,26]. This tradeoff is especially pronounced in pharmacovigilance, where models trained on RWD may inherit systematic misclassification of drug exposure and outcomes, limiting their ability to support causal inference in medication safety.
The efficacy of a DT is fundamentally constrained by the quality and continuity of the RWD used during construction [27,28]. However, medical data remains largely fragmented, with critical information siloed across disparate regions and healthcare institutions [19,29,30]. This fragmentation serves as a primary barrier to unlocking the transformative potential of medical data in the digital health era [19,31,32]. For example, a pharmacogenomics test result obtained in a specialty cardiology practice may not be available to a primary care practitioner managing chronic disease and associated polypharmacy [10,33]. The appeal of RWD for training DT lies in its ecological validity: these data reflect the heterogeneity of actual patient populations, the complexity of real-world treatment patterns, and the influence of social determinants and adherence behaviors that are often excluded or attenuated in trial settings [34,35]. Building robust DTs therefore requires multi-modal data fusion architectures capable of integrating heterogeneous data streams, including structured EMR data, radiological and pathological imaging, genomic sequences, patient-generated health data, and social determinants, within a unified, temporally coherent patient representation.

2.1. EMR Limitations and the Longitudinal Gap

EMRs were originally designed for administrative and billing purposes rather than high-resolution physiological modeling. This design limitation has significant implications. EMRs are not structured to capture the full complexity of medication use, including adherence patterns, dose adjustments, treatment interruptions, or patient-reported side effects. As a result, medication records frequently represent prescribed intent rather than actual exposure. Diagnoses recorded in claims data may reflect reimbursement optimization rather than clinical precision, and medication records often lack information about adherence, over-the-counter use, or discontinuation reasons [36,37]. Consequently, EMR systems frequently fail to capture the continuous, longitudinal data required for dynamic DT simulation (Figure 2).
EMRs document discrete clinical events such as laboratory results and ICD-10 codes but often fail to record the critical intervals between these encounters. Laboratory results, vital signs, and symptom assessments are typically recorded opportunistically when patients present for care, rather than at regular intervals, introducing irregular sampling patterns and informative missingness that can bias time-series models [36,38].
The absence of standardized data formats across different manufacturers prevents clinical interoperability and complicates collaborative research [19]. A significant portion of the corpus of clinical truth resides in unstructured clinical notes, such as pathology reports, which require intensive manual abstraction to be useful for AI training [39]. Unstructured clinical notes, such as free-text pathology reports, require labor-intensive manual abstraction to extract structured information, leading to variability and inefficiencies. Retinal imaging AI models, for instance, struggle with inconsistent protocols, limiting generalizability [40]. Medication management, adverse drug reactions, and pharmacovigilance represent a domain of healthcare burden that remains computationally intractable [41].
A critical challenge for DT-enabled pharmacovigilance is the lack of reliable “ground truth” regarding medication behavior [41]. Unlike well-defined endpoints such as mortality, adverse drug events frequently require clinical adjudication and may present with nonspecific symptoms, making them difficult to capture in structured datasets. EMRs typically record medication orders or perhaps medications a patient tells a provider they have taken, not actual ingestion [42], creating a significant data gap in longitudinal medication adherence. Current EMR systems are considered a “poor source of truth” for these events, failing to capture medication adherence, over-the-counter supplement use, or patient-reported outcomes regarding side effects [41]. Without longitudinal adherence data, a DT may erroneously attribute a lack of therapeutic efficacy to biological resistance rather than simple non-compliance or the introduction of a new medication that perturbs prior stable pharmacotherapy [43]. For example, a complexity in real-world pharmacotherapy is phenoconversion, where environmental, clinical, or polypharmacy-related factors alter the functional activity of drug-metabolizing enzymes independent of genotype. DT models that rely solely on static genomic data without accounting for these dynamic influences risk generating clinically inaccurate predictions of drug response and toxicity. A clinically significant example is the inhibition of CYP2C19 by omeprazole and other proton pump inhibitors, which can convert a genotypically normal metabolizer phenotype to a poor metabolizer phenotype, critically reducing the antiplatelet efficacy of clopidogrel in patients who may be unknowingly managed with both agents.

2.2. The Impact of Multi-Agent Stewardship

The movement of patient-level health data across organizational boundaries is a formidable multi-agent challenge [12]. Healthcare institutions often adopt a “don’t share” posture, viewing data as a proprietary asset rather than a shared resource for precision medicine [19]. Institutional risk management, dominated by the complexities of HIPAA and GDPR, often leads to an aversion to data sharing due to the potential for or perception of liabilities [44]. Institutional data governance frequently supersedes patient agency, often restricting data use regardless of the patient’s individual wishes or consent [11,12]. Arguably, this state of affairs is an impediment to the highest social utility of health data while simultaneously denying the patient the agency informed consent is intended to provide. Unfortunately, dynamic and longitudinal patient consent that is truly informed about granular instances of data use is easier said than accomplished [11,45,46].
This “don’t share” posture is less a function of regulatory prohibition and more a reflection of misaligned incentives. Current systems do not adequately reward data curation, validation, or sharing, despite these activities being essential for developing high-quality, generalizable models. In the context of pharmacovigilance, this results in fragmented safety signals that fail to aggregate across institutions, limiting the early detection of medication-related harm.
To overcome these silos for DT implementation, more robust federated architectures are necessary, allowing for collaborative analysis of fragmented data without the need to move raw, sensitive patient records across boundaries [1,47]. Here, we examine the multi-agent challenge through the lens of RWD limitations, recognizing that the quality, completeness, and representativeness of training data fundamentally constrain the robustness and generalizability of AI-powered digital twins. We analyze how data fragmentation, annotation burdens, governance complexity, bias, and trust deficits shape the feasibility and equity of multi-agent digital twin deployment, and we identify technological and policy innovations—including federated learning, differential privacy, dynamic consent mechanisms, and blockchain-based provenance—that could enable more scalable and trustworthy systems.

3. Data Governance & Regulatory Compliance: FAIR or “Don’t Share”?

The transition from localized data silos to integrated DTs requires sophisticated navigation of the competitive dynamics that shape healthcare delivery. Beyond technical interoperability, organizational and economic factors limit data sharing across institutions. Health systems may view patient data as a competitive asset, fearing that sharing will benefit rival organizations or expose weaknesses in care quality. Hospitals and health systems may negotiate exclusive data access agreements with commercial partners, further restricting the data available to multi-agent DT networks [14,48,49].
Advocacy for a data sharing transaction is usually overcome by bureaucratic process and risk aversion. Data use agreements (DUAs) are often trapped in prolonged negotiations, unstructured workflows between many units in health organizations, with institutions adopting a protective stance to stave off potential disruption or legal exposure. Liability concerns also discourage sharing: institutions worry that errors in shared data could lead to legal claims, and the administrative burden of negotiating data use agreements, business associate contracts, and liability indemnification clauses create high transaction costs. In the absence of standardized contracting templates and dispute resolution mechanisms, institutions default to non-sharing as the safest stance [10,48,49,50]. The resulting administrative friction slows model development, limits the diversity of training data, and prevents agents from accessing the population-scale datasets needed to achieve robust performance [10,12,18]. Multi-agent architectures could reduce these barriers by enabling selective, auditable data sharing under explicit consent and governance rules, with agents responsible for enforcing institutional policies and tracking data lineage [51]. However, achieving consensus on these governance frameworks across heterogeneous stakeholders remains an open challenge [52].

3.1. Findability as a Barrier to DTs

Even though the FAIR (Findable, Accessible, Interoperable, Reusable) ethos [53] of data sharing is growing in the biomedical ecosystem, fulfilling these principles in practical terms can be extremely challenging in most settings. The intent of FAIR is a framework to make data machine-accessible and a kind of social currency [11]. A fundamental knowledge gap between data providers (health institutions) and data users (AI developers) is the need to ensure the quality or completeness of the data collected for clinical use, rather than AI model development. Attempts to verify data quality face the risk of violating current privacy boundaries. The “findability ” aspect is in direct opposition to privacy mandates in many ways, making cohort discovery and study feasibility highly challenging. This problem can be partially mitigated with thoughtful approaches to query metadata to assess completeness, interoperability, and statistical features [54]. However, curation of the metadata can be time-consuming and expensive.

3.2. Accessibility and the Regulatory-Compliance Landscape

While frameworks such as HIPAA in the United States and the GDPR in the European Union were designed to safeguard patient privacy, their implementation has frequently created significant legal and administrative friction that impedes data fluidity [44,55]. Institutions must manage non-overlapping and sometimes conflicting requirements between different jurisdictions [44]. This jurisdictional patchwork is particularly problematic for multi-national DT research networks: data flows that are HIPAA-compliant within the U.S. may simultaneously violate GDPR data residency requirements when crossing into European jurisdictions, effectively precluding the large-scale transatlantic collaboration required for diverse, adequately powered DT training cohorts. HIPAA primarily focuses on the protection of Identifiable Health Information (PHI) through the removal of specific identifiers, whereas GDPR adopts a broader “rights-based” approach, emphasizing the right to be forgotten (aka Right to Erasure under Article 17 of GDPR) [56] and explicit granular consent [45]. While not absolute, the Right to be Forgotten is a term describing rights for consumers akin to a digital reset button with companies that have obtained their data. Such rights include revocation of consent, request for deletion of data, and use in marketing. There are practical and legal limitations to the circumstances in which these rights can be exercised. Germane to AI models, it is impractical (though possible in certain instances) for a model to “unlearn” data it may have been trained on [57].
Institutional risk aversion is particularly pronounced for research and commercial uses of data, where the boundaries of permitted secondary uses under the Institutional Review Board approval and informed consent are unclear [12] and the potential for unethical re-identification or discriminatory profiling is high [47,58]. Institutions often cite regulatory complexities as a justification for limiting data reuse, effectively stifling collaborative research and innovation cycles. The immense cost and complexity of auditing the provenance and consent status across heterogeneous data elements from multiple sources discourage institutions from participating in cross-organizational data exchanges [12].

3.3. Interoperability

A key step toward interoperability has been the adoption of unified data standards across institutions. Implementing widely accepted formats such as HL7 FHIR (Fast Healthcare Interoperability Resources) for data exchange, WHO ICD-10 (International Classification of Diseases, 10th Revision) for diagnostic coding, and SNOMED-CT diagnostic coding, and SNOMED-CT (Systematized Nomenclature of Medicine—Clinical Terms) allows disparate EMR systems to exchange clinical summaries and sensor data seamlessly. These linkages in essence, enable a corpus of data known as the EHR [19,31]. Standardized message formats such as FHIR resources provide syntactic interoperability, but semantic alignment demands ontology mapping and reasoning capabilities. However, implementation variability and differences in local extensions limit FHIR’s effectiveness in practice. Organizations may support different versions of the standard, map local codes to national terminologies inconsistently, or omit optional fields that are critical for DT applications. Semantic ambiguity remains a challenge even when syntactic interoperability is achieved, as the meaning of clinical terms may vary across contexts [44].
To achieve deeper semantic interoperability among multi-agent DT components, developers are utilizing biomedical knowledge graphs and ontologies [3]. Knowledge graphs explicitly map relationships between heterogeneous data sources (like genes, diseases, and phenotypes), allowing DT agents to accurately reason over incomplete data and integrate varying data formats into a unified representation [59,60]. Effective coordination requires robust communication protocols and shared ontologies that enable agents to exchange information, align beliefs, and negotiate actions. In healthcare, terminological heterogeneity [61]—different institutions using different coding systems, vocabularies, and semantic models—complicates ontological alignment and increases the risk of misinterpretation [62]. Agents must be able to translate local concepts into shared representations, infer implicit relationships, and resolve ambiguities when multiple interpretations are plausible [63]. Negotiation protocols allow agents to resolve conflicts when goals or beliefs diverge, such as when one agent recommends a treatment that another considers contraindicated based on different data sources or risk assessments [64]. Multi-agent negotiation can draw on game-theoretic frameworks [65], auction mechanisms, or consensus algorithms, but in safety-critical healthcare applications, these protocols must also support explainability and human oversight [65].
Instead of attempting to pool siloed data into a central repository, which causes regulatory and interoperability friction, federated learning allows DT models to be trained locally across disparate institutions. Only the updated model weights (gradients) are shared and aggregated, effectively bypassing the need to move or standardize raw, sensitive patient records across organizational boundaries [12,66].

3.4. Reusability and DTs

Researchers must be able to trust data origin, provenance, quality, and context for data to be reusable. Blockchain technologies provide tamper-evident audit trails that guarantee data provenance. This enables the validation and replication of results, which is a critical requirement for reusing datasets to advance AI and machine learning models [47,54]. Initiatives like the European Union’s FAIR4Health aim to establish workflows that convert raw clinical data specifically into FAIR-compliant formats [19]. Traditional broad consent models are often inadequate for the evolving, secondary uses of data required by DTs. Additionally, bidirectional “prosent” mechanisms enable stakeholders to manage their own data. Prosent is a portmanteau or combination of the word proactive and the word consent [67]. For example, rare disease patients can actively pool their data and issue requests for researchers to reuse it for specific analyses [12,45].

4. Representational Bias & Health Equity: The Genome Data Gap

The development of equitable DTs is fundamentally compromised by the lack of demographic diversity in the datasets used for model training [68]. The current clinicogenomic landscape is characterized by a profound lack of diversity [13]. The lack of diverse medical datasets is mostly a reflection of where biomedical research infrastructure has historically been concentrated. Studies tend to enroll the populations most accessible to the institutions conducting them, and because much of the global evidence base originates in the U.S. and Europe, those populations are overrepresented.

4.1. Over-Representation of European Ancestry

Most existing genomic and clinical databases heavily over-represent populations of European ancestry. This bias reflects historical research practices, funding priorities, and access barriers that have concentrated genomic sequencing and phenotyping efforts in affluent, majority-White populations [69]. For example, the 5729 samples in the Cancer Genome Atlas were obtained from subjects identifying as White (77%), Black (12%) or Asian (3%), and Hispanic (3%) [70]. Because AI models depend on large, harmonized datasets for validation, this lack of diversity leads to “shortcut” learning where models may fail to generalize across non-European ancestral groups. If left unaddressed, representational bias within these systems amplifies existing health inequalities rather than mitigating them [12,13,71]. The problem is compounded by the inherent degree of genetic admixture within any one of the groups defined based on social context rather than genetic makeup.
For multi-agent DT, trust deficits manifest as asymmetric data availability, with agents serving affluent, educated, or dominant populations possessing richer data than those within underserved communities [28,72]. This asymmetry perpetuates health inequalities and undermines the scientific validity of models that purport to generalize across groups or populations [73]. As a result, AI models trained on these datasets perform poorly when applied to individuals from under-represented groups, missing causal variants, misclassifying pathogenic mutations, and generating inaccurate risk predictions [69,74,75]. The bias is compounded in genome-wide association studies (GWAS) and polygenic risk scores, where the lack of diverse training data means that models amplify inaccuracies by systematically underestimating risk or misguiding treatment decisions for underrepresented patients [76].
For multi-agent DT systems, these data quality issues create asymmetric information in which different environments—hospital systems, primary care networks, specialty clinics—possess partial, inconsistent, or outdated views of the same patient. Reconciling these views requires probabilistic data fusion techniques, but uncertainty propagates through the system, degrading the reliability of predictions and complicating decision-making [17]. Multi-agent DT systems that aggregate models trained on biased data propagate and potentially amplify these errors across the network. In multi-agent environments, bias can propagate through several mechanisms. Agents that rely on overlapping training datasets inherit common biases, and when these agents share model parameters or predictions, biases reinforce rather than average out. Agents optimizing on local objectives, such as institutional performance metrics or reimbursement targets, may inadvertently prioritize well-represented groups for whom data are richer and models more accurate, further marginalizing under-served populations [77,78]. For DTs, which aspire to personalized treatment, this representational bias may undermine clinical utility and equity [79].

4.2. Annotation and Phenotyping Burdens

Patients with rare or under-characterized conditions often experience a “diagnostic odyssey”, a prolonged, arduous, and often traumatizing journey that patients experience from the onset of their first symptoms to the moment they receive an accurate diagnosis. This is a problem of discrimination by omission: disease rarity, phenotypic heterogeneity, and limited clinician familiarity. Similar to the diagnostic odyssey being from a lack of information, the rare disease information gap can propagate through static clinical decision support. DT implementation could shorten this journey by supporting earlier pattern recognition and more individualized inference [80]. However, DTs can also inherit existing gaps in the underlying data and knowledge base. The approaches provided here do not directly address existing mechanistic knowledge gaps for rare disease challenges but might indirectly help by improving research participation and data availability. Training supervised machine learning models for DTs often requires richly annotated datasets precisely labeling clinical events, disease states, treatment responses, and outcomes. In healthcare, annotation depends on expert clinical judgment to interpret imaging studies, adjudicate diagnoses, classify phenotypes, or delineate anatomical structures, making the process labor-intensive, expensive, and slow. Phenopackaging—the systematic extraction and structuring of detailed clinical phenotypes from heterogeneous data sources—represents a critical bottleneck for DT development [12]. Constructing cohorts [10,80] with sufficient granularity to support precision models, such as distinguishing molecular subtypes of cancer or identifying treatment-resistant patient subgroups, requires extensive chart review, integration of genomic and imaging data, and longitudinal tracking of outcomes, all of which are difficult to scale [5,10,13,14,25,71]. Inter-rater variability [81,82], ambiguous findings, borderline measurements, and rare diseases exacerbate annotation and increase label noise [25]. Exemplified by the adverse drug reaction ground truth problem [83], annotation burdens limit the creation of the detailed and standardized cohorts necessary for agents to learn shared representations across institutions.

4.3. Socioeconomic and Environmental Confounders

Health equity in precision medicine requires moving beyond purely biological determinants to include non-medical determinants of health [12]. The realization of precision health through DTs is fundamentally dependent on the integration of high-resolution longitudinal datasets that extend beyond biological substrates [36] to include social determinants of health (SDOH) [84,85]. Despite the transformative potential of DTs to virtually simulate disease dynamics and treatment responses, these models frequently fail to include “invisible” variables, such as socioeconomic, lifestyle, and environmental factors that act as potent confounders but remain poorly represented in EMRs [86].
A central challenge in DT training is the “Zip Code Axiom,” wherein geospatial and environmental stressors, such as local air quality, noise pollution, and proximity to health-enhancing resources like healthy foods, are primary determinants of health outcomes [87]. While biological factors are critical, external influences often exert a greater longitudinal impact on a patient’s health trajectory. However, these variables are frequently siloed within disparate informatics infrastructures, creating significant data quality gaps, blind spots for models, and ultimately biases resulting in inequity.
Furthermore, the introduction of “Label Noise” during the annotation of these complex datasets remains a critical bottleneck [88]. Inconsistent protocols for integrating environmental data with clinical findings lead to systematic biases that compromise algorithmic performance. The laborious nature of “phenopackaging”, extracting detailed phenotypic annotation from unstructured notes, often relies on inconsistent manual processes, further amplifying noise [89]. Without standardized frameworks to control for these “invisible” variables, DT models risk misinterpreting correlations as causality, thereby entrenching clinical errors within high-stakes decision support systems.

4.4. Mistrust and the Barrier of Historical Extraction

A major impediment to expanding representation is the persistent mistrust of medical research institutions among historically underserved communities. The practical realization of multi-agent DT systems necessitates broad data movement across institutional boundaries, yet this movement is hindered by a legacy of extraction and systemic mistrust. Historically, socially and economically marginalized populations have been viewed as passive repositories of data rather than active stakeholders, leading to a pervasive reluctance to participate in biobanks and clinicogenomic registries [90].
Historical abuses of medical research subjects have created deep-seated mistrust of healthcare institutions and reluctance to share health data for research or commercial purposes [38]. This mistrust directly limits participation in biobanks, registries, and digital health initiatives that supply training data for DTs. Under-represented populations are often reluctant to contribute data, leading to biased datasets that systematically exclude the very groups who might benefit most from precision medicine. High-profile data breaches, opaque consent processes, and the commercialization of patient data without transparent benefit-sharing further erode trust [71]. Current static informed consent models are often unable to adapt to evolving research questions or commercial secondary uses, depriving participants of meaningful control over their digital replicas. In decentralized informatics ecosystems, institutional risk-management postures frequently default to “don’t share” due to the immense complexity and cost of auditing data provenance and participant consent across heterogeneous sources [91].
There is often no direct win for patients who contribute data, leading to a system where they bear the privacy risks while institutional stewards or commercial developers capture the economic utility. A significant incentive misalignment exists between the various agents in the data stewardship ecosystem. While commercial entities and research institutions seek population-scale data to train robust models, marginalized cohorts often perceive the benefits of participation as being too abstract or diffuse. This misalignment is compounded by the “many hands” problem of multi-agent systems, where it remains unclear who is accountable if data is misused or if an AI-driven recommendation results in clinical harm [92]. Likewise, the social and individual benefits of sharing data suffer from an appropriability conundrum [11,93]. Collectively, these factors provide major friction to developing the socially beneficial data commons necessary to train and validate population-scale models.

4.5. Pathways to Digital Equity

Addressing representational bias demands proactive data diversification strategies, including targeted recruitment of under-represented populations, partnerships with community health centers and safety-net providers, and investment in culturally tailored outreach and engagement. For example, the Multi-Ethnic Study of Atherosclerosis (MESA) intentionally enrolled a large, community-based cohort balanced across four major racial and ethnic groups, illustrating how proactive, multiethnic recruitment can mitigate representational bias in cardiovascular research [94]. To ensure that DTs serve as a force for health equity, the field must transition toward patient-centric data governance models. Open science through population-scale data commons is an ideal remedy to bias mitigation [11,13].
Federated learning architectures can facilitate these efforts by enabling agents at diverse sites, including under-resourced settings, to contribute to model training without centralizing sensitive data [95]. Transfer learning and domain adaptation techniques allow models trained on larger, albeit biased, datasets to be fine-tuned on smaller, more representative cohorts, improving performance for under-served groups while leveraging existing investments [96,97]. Multi-task learning, in which agents jointly optimize across multiple demographic strata or clinical endpoints, can encourage models to learn shared representations that generalize more equitably [98]. Synthetic data generation approaches, including Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs), offer a complementary strategy by creating statistically plausible patient records that augment under-represented training cohorts while preserving privacy, though rigorous validation is required to ensure that synthetic samples do not introduce distributional artifacts [98].
Governance agents responsible for fairness auditing must continuously monitor model performance across demographic dimensions, race, ethnicity, socioeconomic status, geography, and enforce accountability mechanisms when disparities exceed predefined thresholds [95]. Transparency in reporting performance stratified by group, and involving affected communities in governance and design, are essential for building trust and ensuring that DT systems advance rather than undermine health equity [12]. Implementing federated meta-learning can allow for the training of models on diverse, multi-hospital datasets while maintaining local residency, thereby enabling institutions in resource-limited areas to contribute to and benefit from high-fidelity precision medicine [99]. Successful adoption of DTs requires proactive engagement through idealized clinicogenomic registries that utilize innovative technology to build trust and provide patient-centric oversight [13]. AI systems in clinical practice must undergo periodic “bias stress-tests” and continuous subgroup-performance auditing across sex, race, and socioeconomic strata [100]. Incentives for the underlying data and model stewardship activities remain a thorny challenge but in Section 6.3 we discuss features of blockchain ledgers to address these challenges.

5. Technology & Computational Tools: Validation, Safety, and the Challenge of Model Drift

The technical architecture of a precision medicine DT requires raw computational power, rigorous validation frameworks and continuous monitoring to ensure clinical safety. As these models transition from research environments to routine practice, they must navigate the complexities of algorithmic drift, explainability, and multi-agent orchestration. Implementation is constrained by regulatory frameworks. Importantly, DT outputs must align with the realities of clinical decision-making. Clinicians routinely synthesize incomplete and conflicting data while incorporating patient preferences, cost considerations, and feasibility constraints that are not captured in structured datasets. Models that fail to account for this complexity risk low adoption regardless of technical performance. Here, frame the realities and challenges of DT implementation using the context regulatory risk doctrine for clinical development and intended uses for drugs.

5.1. Safety Frameworks: Phase 1

Regulatory doctrine provides that higher-risk AI applications in healthcare should follow an implementation roadmap informed by traditional clinical trial phases to promote safety and impact. The FDA’s Action Plan for Artificial Intelligence/Machine Learning-Based Software as a Medical Device (AI/ML-based SaMD), which introduced the concept of a Predetermined Change Control Plan (PCCP), provides an emerging regulatory scaffold for managing the iterative update cycles intrinsic to DT systems [101].
Ideal. Initial technical validation focuses on identifying potential harms and ensuring the algorithm performs within defined safety parameters. Deploying DT-driven decision support in clinical practice requires rigorous validation to ensure that models are accurate, reliable, and safe across the intended patient population and clinical contexts.
Reality. Traditional validation approaches rely on retrospective evaluation against held-out test sets and prospective performance monitoring in pilot deployments, but these methods are often insufficient for complex, adaptive AI systems. Multi-agent DTs introduce additional validation challenges because system behavior emerges from interactions among components that may be developed and updated independently. A modeling agent and a decision agent, each validated in isolation, may produce harmful recommendations when combined due to misaligned assumptions, incompatible probabilistic representations, or conflicting objectives. Comprehensive validation strategies must therefore include integration testing, in which multi-agent interactions are evaluated under realistic scenarios, and stress testing, in which edge cases, adversarial inputs, and failure modes are systematically explored. Simulation environments that replicate clinical workflows, patient trajectories, and system dynamics provide sandboxes for exercising agents before live deployment.

5.2. Efficacy and Effectiveness Frameworks: Phases 2 and 3

Phase 2 (Efficacy).
Ideal. Models are evaluated for their ability to achieve intended diagnostic or therapeutic outcomes under controlled conditions.
Phase 3 (Effectiveness).
Ideal. AI solutions must be compared to existing clinical standards of care to ensure they provide measurable improvement in “the wild”. Clinicians think the models are easy to use and add value to patient care. The AI solutions are trustworthy and included in clinical care.
Reality for Efficacy and Effectiveness. A growing need for verification, validation, and uncertainty quantification (VVUQ) methods are needed to assess the performance of these models [102]. In controlled settings like randomized controlled trials (RCTs), validation focuses on internal consistency and predictive accuracy using structured datasets. Strategies include k-fold cross-validation to assess model performance across subsets, minimizing overfitting, and tools like the Prediction Model Risk of Bias Assessment Tool (PROBAST) for bias evaluation [103]. These approaches allow isolation of variables, facilitating mechanistic verification through simulations against known outcomes. A point of conflict is the scalability of such methods, as some studies highlight computational demands limiting broad application.
An example is the LifeTIME project, where DTs predict multimorbidity risk using data from RCTs like the Children and Young People’s Health Partnership trial [104]. Models were trained on longitudinal cohorts, validated via k-fold cross-validation and PROBAST, demonstrating improved risk stratification in simulated personas. However, data linkage, data quality, and lack of individualization are cited as limitations to rigorous evaluation. In organ transplantation, DTs simulate graft viability using machine perfusion data [105]. Validation involves prospective integration into clinical workflows, comparing simulated outcomes with real-time immunosuppression adjustments, though limited by early-stage evidence.
In healthcare settings, validation shifts to prospective, adaptive approaches delivered amid heterogeneous data streams. Key strategies encompass uncertainty quantification to account for confounders like environmental factors, continuous monitoring via federated learning for privacy-preserving updates, and real-time benchmarking against EHRs. Consensus, in turn, supports integrating VVUQ to quantify prediction errors in dynamic environments and to enhance generalizability [102]. However, authors disagree on maturity: while mechanistic simulations show promise, prospective clinical validation remains sparse due to ethical and infrastructural barriers [106]. Even an accurate AI model will not be effective if clinicians do not use the model and/or do not include AI recommendations in clinical care. AI models must be easy to use and seamlessly fit into established clinical workflows.

5.3. Phase 4 (Monitoring): Model Drift and Diagnostic Safety

Ideal. Post-deployment surveillance is essential to systematically address regulatory compliance and ongoing safety.
Reality. A significant barrier to building reliable DTs is “calibration drift,” where model performance degrades over time due to changes in clinical practice, patient demographics, or data collection protocols [107]. Healthcare is characterized by non-stationary environments in which disease prevalence, treatment guidelines, patient behaviors, and data collection practices evolve over time. Image-based AI systems, such as those used for pulmonary nodule detection, often suffer from “black box” opacity, which complicates legal liability if the tool fails to identify abnormalities [108]. Models that perform well initially may degrade as distributions shift, a phenomenon known as model drift or concept drift [109]. For DTs, which aspire to maintain accurate patient representations over extended periods, detecting and correcting drift is essential. Monitoring agents must track performance metrics—calibration, discrimination, fairness—continuously across demographic strata and clinical contexts, triggering alerts when degradation exceeds thresholds. Updating models to address drift, however, introduces additional risks and metrics are not standardized to ascertain when sufficient RWD are available to inform the AI model. Retraining on recent data may ignore historical patterns, while deploying updated models without adequate validation may introduce new errors or biases [110]. Lifecycle management agents must therefore orchestrate versioning, rollback mechanisms, and phased deployment strategies, ensuring that updates are tested and documented before clinical use. Mitigation requires continuous subgroup-performance auditing and periodic “bias stress-tests” across different socioeconomic and other demographic strata. Quantitative evaluation tools like MedAgentAudit [111] are increasingly necessary to diagnose collaborative failure modes in multi-agent medical systems [111].

5.4. Explainability and Clinician Autonomy

For a DT to be clinically useful, its predictions must be explainable to ensure that clinician autonomy is preserved and to improve adoption viability [80]. Clinicians and patients require understandable explanations of DT recommendations that in turn promote trust and increase adoption. Black-box models, even when highly accurate, pose barriers to adoption in high-stakes settings where liability and informed consent demand transparency [18,112,113]. Multi-agent architectures can support explainability by explicitly modeling reasoning processes, maintaining provenance of data and decisions, and providing agents dedicated to generating natural-language justifications [64]. Causal models, which represent mechanistic relationships among variables, offer inherently interpretable predictions that align with clinical reasoning, though they are more challenging to learn from observational data than purely correlational approaches [114]. Models based on causality tend to require a higher standard of evidence, such as RCTs or other controlled scientific experiments. Observational data like epidemiological data tend to be more correlative, with confounders not isolated. Accountability mechanisms specify which agents—and ultimately which institutions or individuals—are responsible when DT recommendations lead to adverse outcomes [115]. Governance agents that record decision trails, consent states, and policy compliance provide audit trails that support post hoc review and liability determination [116].
Interactive visualization tools that allow users to explore DT simulations, adjust parameters, and compare alternative scenarios can foster shared decision-making and trust [117]. Research on cognitive load, decision fatigue, and human factors is essential to design interfaces that support effective human-agent collaboration without overwhelming users with complexity [118,119]. Techniques like Shapley Additive Explanations (SHAP) and hybrid encoder–decoder schemes are being deployed to mitigate hallucinations and inject domain knowledge into diagnostic dialogs [120,121].
The orchestration layer of multi-agent systems must prioritize “explainability by design,” ensuring that every recommendation is linked to a DOI, PMID, or specific clinical guideline ID (provenance aware) [122]. These underlying annotations are embedded in the models and clinical decision support interfaces, and they are readily callable should a user wish to dig deeper. These annotations are curated by experts and software developers, not end users. For example, many clinical decision support tools for pharmacogenomics-based prescribing incorporate knowledge bases such as ClinPGx to provide alerts when drug-variant interactions might be relevant to a prescribing decision [123]. These knowledge bases incorporate known medication variant relationships like “poor metabolizer phenotypes” coupled with thorough citation of scientific evidence in both medical guidelines and the peer-reviewed literature, which often includes RCTs. Ensuring that DT systems augment rather than replace human judgment requires human-in-the-loop interfaces that enable clinicians and patients to understand, challenge, and shape agent recommendations [27,124,125]. Explainable AI techniques including counterfactual reasoning, feature attribution, and concept activation vectors, must be integrated into multi-agent architectures so that agents can generate comprehensible justifications tailored to user expertise and context [85,122,126].
Achieving deeper semantic interoperability among multi-agent DT components will benefit from advances in provenance-aware biomedical knowledge graphs that explicitly represent entities, relationships, and ontologies spanning genes, proteins, diseases, drugs, and phenotypes. Knowledge graph embedding techniques enable agents to reason over incomplete or noisy data, infer missing relationships, and integrate heterogeneous data sources in a unified representation.

5.5. Multi-Agent Orchestration and Scaling

Modern blockchain and AI model architectures construct specialized agents for metadata management, data integration, and quality assurance [18]. Such systems are widely used for cryptocurrency transactions and other blockchain-based data transactions. A less sophisticated example of such an architecture in healthcare is the visICU virtual ICU platform [127]. The Philips eICU Program (formerly Visicu) employs a hierarchical, “hub-and-spoke” architecture designed for high-fidelity remote monitoring [128]. The core engine, eCareManager, functions as a vendor-agnostic integration layer and ingests high-frequency data from disparate hospital systems (Epic, Cerner, etc.) and bedside monitors using HL7 v2 and FHIR APIs. This “data fabric” architecture allows for event-driven processing, where clinical events (e.g., a verified lab result) trigger immediate automated risk assessments rather than waiting for batch uploads. Standardization is achieved through a Global Metadata Schema. Local data (e.g., “K+” vs. “potassium”) is mapped to standardized clinical concepts to ensure consistency across multi-center networks. The system utilizes unique identifiers to maintain data lineage and temporal accuracy, recording all physiological events as offsets (in minutes) from the time of ICU admission. Quality functions include APACHE scoring to compare actual vs. predicted outcomes, AI models are monitored for “drift” to ensure clinical reliability, and real-time auditing of “best practice” care bundles (e.g., Sepsis protocols). While not a DT per se, the system has transitioned toward a cloud-native, FHIR-first platform to support real-time, agentic AI-driven clinical decision support.
In a computational context, multi-agent DT systems comprise heterogeneous agents with distinct roles and capabilities [38]. In a medical context, sensing agents interface with medical devices, wearable sensors, and information systems to acquire real-time physiological and clinical data, performing preprocessing, quality checks, and anomaly detection before forwarding data to modeling agents [38,129]. Modeling agents maintain the mathematical or machine learning representations that constitute the DT, updating parameters as new data arrive and propagating uncertainty estimates to downstream decision agents [18,27]. Decision agents use the DT to generate recommendations of diagnostic tests, treatment plans, resource allocations, and interact with clinician and patient agents who retain ultimate authority in safety-critical contexts [18]. Governance agents monitor compliance with institutional policies, regulatory requirements, and consent preferences, mediating access requests, auditing data usage, and tokenizing policy violations [130]. Communication agents facilitate message passing, negotiation, and consensus protocols among distributed agents, while orchestration agents coordinate complex workflows such as virtual clinical trials or population health campaigns [38,131]. For example, systems like “Agent Hospital” illustrate the potential for evolvable medical agents that learn from simulated environments to improve diagnostic accuracy before real-world deployment [132]. Agent Hospital is a cutting-edge, fully virtual hospital simulation developed by Tsinghua University in China. It uses large language models (LLMs) to power autonomous AI agents—doctors, nurses, and patients—enabling large-scale, risk-free training and evolution of medical AI systems. LLMs are increasingly being explored as orchestration agents capable of natural language reasoning, translating clinical narratives into structured data, and coordinating inter-agent workflows, though their integration raises additional concerns around hallucination, auditability, and computational overhead in safety-critical settings.
Decentralized coordination is defined as agents acting based on local information and interactions with peers, without centralized command, offering scalability and resilience advantages. However, decentralization introduces risks of emergent behaviors that are difficult to predict or control [18,27,133]. Complex interactions among agents can produce system-level phenomena such as cascading errors, oscillatory instabilities, or unintended biases that are not evident when components are validated in isolation [133]. Agent coordination schemas are already being developed. For DT systems, emergent behaviors pose safety risks when agents’ local optimizations or conflicting objectives lead to harmful recommendations or resource misallocations. Detecting and mitigating these risks requires simulation-based validation, in which multi-agent interactions are exercised across diverse scenarios before clinical deployment, and runtime monitoring agents that detect anomalies and trigger corrective interventions [18,133]. Hierarchical coordination architectures, in which higher-level orchestration agents supervise lower-level operational agents, can provide structure and accountability while retaining distributed execution benefits [60,134,135]. However, hierarchies introduce latency and single points of failure, necessitating tradeoffs between control and resilience [85,126]. Put another way, more hierarchy allows control mechanisms at the cost of resilience due to failure cascades.
Scaling these models requires significant investment in GPU-accelerated computing and standardized protocols like FHIR to ensure interoperability across institutional AI agents [136]. Real-time DT applications, such as intensive care monitoring, surgical navigation, or closed-loop drug delivery, impose strict latency and reliability requirements that challenge distributed multi-agent architectures [133,137]. Centralizing computation in cloud data centers introduces network delays and dependence on internet connectivity, which may be inadequate in resource-limited settings or during emergencies [38,138,139]. Edge and fog computing paradigms, in which computational agents are deployed on local servers, gateways, or devices closer to data sources, reduce latency and enable offline operation [138]. In multi-agent DT systems, edge agents can perform time-critical data processing, inference, and control, while cloud agents handle model training, long-term storage, and population-level analytics [138]. However, distributing agents across edge, fog, and cloud layers complicates coordination, as agents must synchronize state (i.e., ensure working with the same instance of database), manage inconsistencies due to network partitions, and reconcile updates when connectivity is restored [126,138]. Consistency protocols—such as eventual consistency, conflict-free replicated data types, or consensus algorithms—help agents maintain coherent views despite distributed operation, but each introduces tradeoffs between performance, availability, and correctness [55,140].

5.6. Privacy and Computational Mitigation

Traditional de-identification, the removal of the 18 HIPAA identifiers, is increasingly viewed as insufficient, as machine learning models can potentially re-identify individuals through large-scale data aggregation or genomic linkage [1,32,54,141]. The health data field is shifting toward privacy-preserving computational frameworks. Technical approaches to privacy preservation offer pathways to enable collaborative DT development while mitigating regulatory and institutional concerns. This bodes well for alleviating some of the frictions of data sharing as most barriers arise due to privacy concerns.
Differential privacy mechanisms add noise to model updates or query results to probably limit the risk of inferring individual-level information from aggregate statistics [51,142]. By calibrating noise levels according to a privacy budget, these techniques provide mathematically rigorous guarantees that participating in model training does not significantly increase an individual’s re-identification risk [51]. However, privacy-preserving methods introduce tradeoffs between privacy guarantees and model accuracy. Strong differential privacy requires substantial noise, which can degrade prediction performance, especially when training data are limited or highly heterogeneous [51]. For multi-agent DTs, balancing these tradeoffs demands careful tuning of privacy parameters and ongoing monitoring of fairness and performance across demographic groups [143].

6. Blockchain as a Trust Anchor: Provenance, Consent, and Incentivization

The realization of high-fidelity precision medicine requires a foundational infrastructure capable of managing data security, interoperability, and patient privacy. We argue that blockchain technology serves as a pivotal decentralized trust anchor, providing a technical solution to the “don’t share” institutional posture by ensuring the integrity of patient-level data across heterogeneous ecosystems [11,12,13]. Blockchain and distributed ledger technologies provide immutable, transparent, and decentralized records of transactions, making them attractive for establishing trust and provenance in multi-agent DT systems where participants may be mutually distrusting [22,144,145]. By recording data access events, consent decisions, model training activities, and prediction requests as cryptographically signed transactions, blockchains create tamper-evident audit logs visible to authorized agents.
Blockchain-enabled tokenization offers a mechanism for incentivizing the labor-intensive work of data stewardship, curation, and validation that underpins high-quality DT training datasets [32,146,147]. Institutions, clinicians, or patients who contribute curated data, provide annotations, or validate model predictions could receive tokens representing fractional ownership or usage rights, which accrue value as digital twin products generate revenue. Decentralized validation nodes, analogous to cryptocurrency mining, could distribute the work of verifying data provenance, auditing model performance, and enforcing governance policies across a network of participants, with rewards allocated based on contributions [146,148,149]. This distributed approach reduces the concentration of power and gatekeeping roles traditionally held by centralized institutions, potentially lowering transaction costs and enabling broader participation. However, designing incentive structures that are robust to gaming, equitable across participants, and aligned with ethical norms is difficult [32,146,147]. Token economies risk financializing patient data in ways that may be exploitative or coercive, and ensuring that economically disadvantaged individuals are not disproportionately incentivized to relinquish privacy protections requires careful governance and oversight [54,150,151].

6.1. Tracking Data Provenance and Lineage

DTs are vulnerable to the “fruit from a poison tree” problem, where AI models trained on data with unknown provenance or undocumented consent are considered legally and ethically tainted [148,152]. Smart contracts—self-executing programs deployed on blockchains—can automate governance policies, enforcing rules for data sharing, access control, and benefit distribution without relying on centralized intermediaries [148,153]. These self-executing agreements computationally encode “if-then” conditions to manage access rights based on real-time patient preferences [154]. For example, a smart contract could permit a modeling agent to access patient data only if valid consent exists and the agent’s institution has paid required fees, with all transactions recorded transparently [153]. Distributed ledgers connect the provenance of data elements with cryptographic validation at every stage of the model lifecycle, ensuring that researchers can back-test predictive tools with confidence [109,152]. Blockchain creates a tamper-proof evidence trail of every data entry and transaction, mitigating the inherent potential for fraud, deception, and misappropriation in multiagent architectures [148]. The use of cryptographic hashes guarantees that documents and clinical records have not been altered since their original entry into the chain. Consensus protocols underlying blockchains, such as proof-of-stake or Byzantine fault tolerance algorithms [155], ensure that transaction records are agreed upon by network participants even in the presence of malicious or faulty agents, providing robustness against data tampering and single points of failure [153].

6.2. Automating Dynamic Consent

Current static informed consent models are inadequate for the long-term, evolving nature of DT research and implementation [11]. Dynamic consent frameworks aim to address these limitations by enabling individuals to grant, modify, or revoke permissions for specific uses of their data on an ongoing basis [22,154,156]. Digital platforms can present participants with granular options—consenting to aggregate research but not commercial use, or to genomic studies but not behavioral tracking—and update these preferences as circumstances change [154]. Every instance of data access and editing is recorded on-chain, creating a transparent audit trail that prevents post-facto consent falsification [148,154,156]. Implementing dynamic consent in multi-agent DT systems requires digital consent agents (smart contracts or consensus protocols in the case of blockchain systems) that interface with patients, present choices in understandable terms, propagate preference updates to all relevant agents, and enforce access controls in real-time [148,154,156]. Blockchain-based consent registries offer one approach, recording consent states as immutable transactions visible to authorized agents and enabling fine-grained, auditable access control [148,154,156]. These interactions with the primary data can generate digital metadata breadcrumbs recorded on the ledger under consensus protocols specified by smart contracts [148,154,156].
Blockchain technology facilitates cross-platform interoperability and manages decentralized identity and the use of smart contracts to enable granular data agency across research networks and data sharing networks [22,148,154,156]. In this paradigm, each organization, its data stewards, patients, research subjects, and researchers can digitally audit a consent ledger and enforce smart contracts representing the data governance concepts representing their priorities [22]. In this architecture, blockchain enables patients to provide or revoke consent for specific secondary data uses, effectively returning agency to the individual [22].
However, dynamic consent also introduces challenges [157]. Frequent permission updates impose cognitive and administrative burdens on participants, and overly granular choices may confuse or overwhelm individuals who lack technical or scientific expertise [157,158]. Balancing flexibility with usability demands user-centered design and ongoing support, and ensuring that consent platforms remain accessible to individuals with limited digital literacy or language proficiency is critical for equity [158].

6.3. Incentivizing Data Stewardship Through Digitized Validation

The work of data curation and validation is currently undervalued and siloed. The same features of blockchain ledgers that enable provenance awareness and appropriability make incentivization and data economics feasible [11]. To overcome the lack of motivation for curation of data, there is a push to transition from altruistic data donation to a “shared economic model.” Blockchain enables the tokenization of data elements that can be managed as digital assets, analogous to tokens or cryptocurrencies, facilitating infrastructure for a patient-centric data economy. Platform funding models can also rely on data users purchasing utility tokens to cover transaction fees and compensate data owners [159].
Digital health data marketplaces incentivize data sharing by compensating patients and institutions for the commercial use of their medical data with tokens (including nonfungible tokens) [160], rewards, or currency [31]. Institutions, clinicians, and patients who contribute to the labor-intensive work of data stewardship, curation, and validation (essential for high-quality digital twin training) can be rewarded. Rewards can include utility tokens or fractional ownership rights that accrue value and are administered digitally by smart contracts as the DT products generate revenue [11,51]. Such models have been deployed in consumer genomics, though none has yet achieved significant commercial traction or sustainability yet [11]. However, approaches in brand loyalty programs might provide a model upon which to build web3 communities of passionate participants [161].
Blockchain incentivizes platform providers and data stewards through dynamic valuation models that allocate returns based on market-assessed contributions. In fact, blockchain ledgers can address the appropriability conundrum and align incentives and rewards for data sharing and data stewardship [11,144]. A use case would be to provide a patient who has shared data with a clinicogenomic registry a dashboard of scientific publications that have used their data or specimens using the cryptographic hash to backtrack through the blockchain ledger and knowledge graphs [13,146]. That appropriability ledger can just as easily be leveraged to convey digital or financial rewards in much the same way validator notes do for crypto mining. By removing the need for intermediary trust, blockchain architectures allow diverse stakeholders—including clinics, labs, and innovative industry partners—to altruistically exchange data insights for the common good [13,146,147].

6.4. Implementation Challenges and Hybrid Models

Despite their promise, blockchain-based approaches face throughput limitations. Ethereum can process far fewer transactions per second than would be required for population-scale DT systems ingesting continuous clinical data streams [162,163]. Infrastructure costs, latency introduced by consensus protocols, and the computational burden of validating large files, such as high-density imaging or genomic sequence data (.bam, .dicom, .vcf), can impede blockchain network performance and these transactions [162]. To address these limitations, some have experimented with hybrid models that leverage off-chain storage with on-chain security and object-based tokenization to ensure rapid data access without compromising integrity [163]. These constraints help explain why blockchain remains limited in real-world clinical systems despite strong conceptual appeal. For many healthcare settings, hybrid models that combine off-chain storage, conventional institutional controls, and selectively deployed ledger functions may be more realistic than fully decentralized architectures.
Permissioned or consortium blockchains offer better performance by limiting validators to trusted institutions but sacrifice some of the decentralization and censorship-resistance properties that motivate blockchain adoption [162,164]. Integrating blockchain with existing healthcare information systems is technically and organizationally complex, requiring middleware agents, API development, and coordination among institutions with heterogeneous legacy infrastructures [162,165,166]. Legal and regulatory questions about the status of blockchain records, liability for smart contract failures, and cross-border data sovereignty remain unresolved [44,146]. Successful deployment requires supportive policies and sector-wide coordination to ensure that blockchain-based systems comply with evolving HIPAA and GDPR standards [44].

7. Future Directions & Conclusion: Toward an Ethical AI Ecosystem

The transition from localized, static medical modeling to dynamic, high-fidelity DTs represents a critical juncture in the evolution of precision medicine. However, the successful clinical implementation of these systems depends on moving beyond purely computational advancements to address the systemic data-stewardship barriers identified in this review (Figure 3).

7.1. Future Research Directions

The future of clinical DTs will be defined by several key technological and socio-technical priorities. To overcome institutional “don’t share” postures, the field must adopt decentralized architectures like federated learning, which allow for multi-institutional collaboration without the movement of raw, sensitive patient data. Concerted efforts are required to move beyond the current over-representation of European ancestry in clinicogenomic databases, incorporating non-biological determinants such as socioeconomic, lifestyle, and environmental confounders. Research must prioritize AI-driven tools for real-time data cleaning, metadata management, and automated provenance tracking to minimize the administrative burden on institutional stewards. Future frameworks must replace static informed consent with dynamic, blockchain-enabled models that empower patients to manage their data as a digital asset, facilitating longitudinal engagement essential for DT recalibration. Swarm learning, in which agents collaboratively train models through peer-to-peer communication and local consensus, offers greater decentralization than hierarchical federated learning but demands robust protocols for agent discovery, trust establishment, Byzantine fault tolerance, and provenance-aware data and model architectures. Research is needed to develop hybrid architectures that balance scalability, privacy, fairness, and resilience across diverse healthcare deployment scenarios.

7.2. Building National and Global Infrastructure

We argue that the current fragmented approach to data stewardship is unsustainable for population-scale precision medicine. A transition to population-scale or national-scale medical data infrastructures is necessary to fully unlock the potential of medical data. The widespread adoption of unified standards, such as FHIR and common data models, is essential to ensure that DTs remain interoperable across disparate healthcare systems. Establishing proactive, constraint-based governance frameworks, such as the Health AI Consumer Consortium (HAIC2) [52], can help harmonize the perspectives of patients, clinicians, and industry partners. The HAIC2 has progressed as a consumer-facing consortium [167] similar to the VISICU database consortium to align stakeholders and incents in the AI model ecosystem.

7.3. Conclusions

The workflows associated with training, inference, validation, and quality monitoring involve many frictions associated with data governance (Figure 4). The barriers to building clinical DT are primarily rooted in frictions in data stewardship, including fragmentation within RWD silos, defensive institutional governance frameworks, and persistent representational bias. At the same time, important technical challenges in model specification, calibration, validation, and biological fidelity also shape the feasibility of clinically useful DTs, even though those modeling issues are not the primary focus of this review.
The multi-agent challenge in implementing digital twins for healthcare arises from the intersection of technical complexity, institutional fragmentation, regulatory constraints, and social imperatives for trust and equity. RWD, while essential for training patient-specific models, are plagued by incompleteness, inconsistency, and bias, limiting the robustness and generalizability of AI-driven DTs. Multi-agent system architectures offer a pathway to coordinate distributed data sources, computational processes, and human stakeholders, but realizing this vision demands advances in privacy-preserving computation, semantic interoperability, fairness-aware learning, and decentralized governance.
Safety, validation, and accountability mechanisms are critical for clinical deployment, particularly in multi-agent contexts where emergent behaviors and distributed decision-making complicate traditional regulatory and liability frameworks. Building public trust—especially among underserved communities with historical reasons for mistrust—demands transparency, benefit-sharing, and meaningful patient control over data and its uses. Addressing representational bias and health equity requires deliberate data diversification, inclusive governance, and continuous monitoring to ensure that DT systems do not perpetuate or amplify existing disparities.
Federated learning, differential privacy, dynamic consent mechanisms, and blockchain-based provenance are promising building blocks for future DT ecosystems but each introduces tradeoffs and implementation constraints that require careful evaluation. By improving provenance tracking, strengthening patient agency, and enabling coordination across institutional boundaries, these approaches may support a more trustworthy and interoperable health-data ecosystem. Progress, however, will require more than technical innovation. It will also depend on policy reform, regulatory adaptation, and sustained investment in equitable data infrastructure. Multi-stakeholder collaboration among researchers, clinicians, patients, policymakers, and industry will be essential to translate DTs from promising concepts into clinically credible tools. A critical next step is the development of minimum data standards for DT–ready datasets, including structured medication exposure, longitudinal outcomes, and standardized phenotype definitions. Without such standards, interoperability alone will be insufficient to support reliable and equitable model development. Ultimately, the promise of the DT paradigm will depend on whether these systems can become personalized but also clinically valid, operationally feasible, and just.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/aimed1030020/s1.

Author Contributions

Conceptualization, P.J.S. and J.T.; methodology, P.J.S. and J.T.; validation, P.J.S.; investigation, P.J.S. and J.T.; resources, K.S.R. and P.J.S.; Writing—original draft preparation, P.J.S. and J.T.; writing—review and editing, P.J.S., J.T., S.L.R., Q.H., J.D.R., L.B., S.A.B., P.K.S. and K.S.R.; visualization, P.J.S., J.T. and J.D.R.; supervision, K.S.R.; project administration, P.J.S.; funding acquisition, K.S.R. and P.J.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded in part by CPRIT RP230204 and a gift from the Genentech Foundation.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No external or patient data were used in this work.

Acknowledgments

This manuscript was drafted with the assistance of an AI-based scientific writer to ensure adherence to MDPI structural guidelines and formal stylistic requirements. Claude Opus 4.6 and Chat GPT5.4mini were used for literature searching and drafting assistance. Gemini 3.1 Pro was used to validate reference and citation alignment. Copilot was used for editorial reviews. Notebook LLM (powered by GemeniPro3.1) was used to generate the graphical abstract and figures. The prompts used to generate Figure 1, Figure 2, Figure 3 and Figure 4 are detailed in the Supplementary Materials. All factual assertions and clinical synthesis were derived exclusively from the provided research library, and the authors maintain full responsibility for the accuracy and integrity of the final text.

Conflicts of Interest

The funders had no role in the design of the project; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results. PJS holds equity equivalent investments in Mysten Labs, Sus Health, N3XT and is a scientific advisory board member for BioPath Holdings Inc.

Abbreviations

The following abbreviations are used in this manuscript:
DT/DTsDigital Twin/Digital Twins
RWDReal-World Data
EMRElectronic Medical Record
AIArtificial Intelligence
EHRElectronic Health Record
HIPAAHealth Insurance Portability and Accountability Act
GDPRGeneral Data Protection Regulation
MADTMulti-Agent Digital Twin
FAIRFindable, Accessible, Interoperable, Reusable
ICD-10International Classification of Diseases, 10th Revision
SNOMED-CTSystematized Nomenclature of Medicine—Clinical Terms
FHIRFast Healthcare Interoperability Resources
GWASGenome-Wide Association Studies
SDOHSocial Determinants of Health
RCTRandomized Controlled Trial
SaMDSoftware as a Medical Device
FDAU.S. Food and Drug Administration
PCCPPredetermined Change Control Plan
VVUQVerification, Validation, and Uncertainty Quantification
PROBASTPrediction Model Risk of Bias Assessment Tool
SHAPShapley Additive Explanations
LLM/LLMsLarge Language Model/Large Language Models
APIApplication Programming Interface
HL7Health Level Seven
APACHEAcute Physiology and Chronic Health Evaluation
GPUGraphics Processing Unit
IoTInternet of Things
GANGenerative Adversarial Network
VAEVariational Autoencoder
NFTNon-Fungible Token
IRBInstitutional Review Board
DOIDigital Object Identifier
PHIProtected Health Information
CYPCytochrome P450
CYP2C19Cytochrome P450 2C19 enzyme
PGxPharmacogenomics/Pharmacogenetics
ICUIntensive Care Unit
eICUElectronic Intensive Care Unit
BFTByzantine Fault Tolerance
IPFSInterPlanetary File System
HAIC2 (HAIC2)Health AI Consumer Consortium
MLMachine Learning

References

  1. Nagaraj, D.; Khandelwal, P.; Steyaert, S.; Gevaert, O. Augmenting digital twins with federated learning in medicine. Lancet Digit. Health 2023, 5, e251–e253. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Sadée, C.; Testa, S.; Barba, T.; Hartmann, K.; Schuessler, M.; Thieme, A.; Church, G.M.; Okoye, I.; Hernandez-Boussard, T.; Hood, L.; et al. Medical digital twins: Enabling precision medicine and medical artificial intelligence. Lancet Digit. Health 2025, 7, 100864. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Kalyani, Y.; Collier, R. The Role of Multi-Agents in Digital Twin Implementation: Short Survey. ACM Comput. Surv. 2024, 57, 72. [Google Scholar] [CrossRef] [Scilit]
  4. Elgammal, Z.; Albrijawi, M.T.; Alhajj, R. Digital twins in healthcare: A review of AI-powered practical applications across health domains. J. Big Data 2025, 12, 234. [Google Scholar] [CrossRef] [Scilit]
  5. Silva, P.; Janjan, N.; Ramos, K.S.; Udeani, G.; Zhong, L.; Ory, M.G.; Smith, M.L. External control arms: COVID-19 reveals the merits of using real world evidence in real-time for clinical and public health investigations. Front. Med. 2023, 10, 1198088. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Master, R.; Rubin, N.; Sampson, J.; Yadav, K.K.; Pandita, S.; Sabbagh, A.; Krishnan, A.; Silva, P.J.; Ramos, K.S.; Gregoire, V.; et al. Advances in Artificial Intelligence for Glioblastoma Radiotherapy Planning and Treatment. Cancers 2025, 17, 3762. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Geddes, J.R.; Jensen, C.W.; Tanade, C.; Ghorbannia, A.; Fudim, M.; Patel, M.R.; Randles, A. Digital twins for noninvasively measuring predictive markers of right heart failure. npj Digit. Med. 2025, 8, 545. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Pantalone, K.M.; Xiao, H.; Bena, J.; Morrison, S.; Downie, S.; Boyd, A.M.; Shah, L.; Willis, B.; Beharry-Diaz, J.; Milinovich, A. Type 2 diabetes pharmacotherapy de-escalation through AI-enabled lifestyle modifications: A randomized clinical trial. NEJM Catal. Innov. Care Deliv. 2025, 6, CAT.25.0016. [Google Scholar]
  9. Fisher, C.K.; Smith, A.M.; Walsh, J.R. Machine learning for comprehensive forecasting of Alzheimer’s Disease progression. Sci. Rep. 2019, 9, 13622. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Silva, P.J.; Sweitzer, N.K. 6.07—Chimeric cohorts and consortia can power and scale precision medicine. In Comprehensive Precision Medicine, 1st ed.; Ramos, K.S., Ed.; Elsevier: Oxford, UK, 2024; pp. 264–282. [Google Scholar]
  11. Silva, P.; Silva, P.A.; Ramos, K.S. Genomic and Health Data as Fuel to Advance a Health Data Economy for Artificial Intelligence. BioMed Res. Int. Sci. Policy Innov. Driven Med. Res. 2025, 2025, 565955. [Google Scholar] [CrossRef] [Scilit]
  12. Silva, P.J.; Rahimzadeh, V.; Powell, R.; Husain, J.; Grossman, S.; Hansen, A.; Hinkel, J.; Rosengarten, R.; Ory, M.G.; Ramos, K.S. Health equity innovation in precision medicine: Data stewardship and agency to expand representation in clinicogenomics. Health Res. Policy Syst. 2024, 22, 170. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Silva, P.; Dahlke, D.V.; Smith, M.L.; Charles, W.; Gomez, J.; Ory, M.G.; Ramos, K.S. An Idealized Clinicogenomic Registry to Engage Underrepresented Populations Using Innovative Technology. J. Pers. Med. 2022, 12, 713. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Silva, P.J.; Ramos, K.S. Precision medicine at the academic-industry interface. In Precision Medicine for Investigators, Practitioners and Providers; Academic Press: Cambridge, MA, USA, 2020; pp. 545–560. [Google Scholar]
  15. Abdulkadir, U.; Waziri, V.O.; Alhassan, J.K.; Ismaila, I. Electronic Medical Records Management and Administration: Current Trends, Issues, Solutions, and Future Directions. SN Comput. Sci. 2024, 5, 460. [Google Scholar] [CrossRef] [Scilit]
  16. Kessler, M.D.; Yerges-Armstrong, L.; Taub, M.A.; Shetty, A.C.; Maloney, K.; Jeng, L.J.B.; Ruczinski, I.; Levin, A.M.; Williams, L.K.; Beaty, T.H. Challenges and disparities in the application of personalized genomic medicine to populations with African ancestry. Nat. Commun. 2016, 7, 12521. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Lee, H.M.; Jamali, R.; Lazarova-Molnar, S. A Conceptual Framework for Digital Twins of Multi-Agent Systems. Procedia Comput. Sci. 2025, 257, 321–328. [Google Scholar] [CrossRef] [Scilit]
  18. Croatti, A.; Gabellini, M.; Montagna, S.; Ricci, A. On the Integration of Agents and Digital Twins in Healthcare. J. Med. Syst. 2020, 44, 161. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Wang, J.; Lin, A.; Huang, Y.; Li, G.; Chen, T.; Sun, C.; Qian, W.; Ren, S.; Wong, H.Z.; Ding, Y. Medical Data as a Key Asset in the Digital Health Era: A Framework for Challenges and Strategies. iMetaMed 2025, 1, e70014. [Google Scholar] [CrossRef] [Scilit]
  20. Rubeis, G. Ethical implications of blockchain technology in biomedical research. Ethik Der Med. 2024, 36, 493–506. [Google Scholar] [CrossRef] [Scilit]
  21. Rahimzadeh, V.; Serpico, K.; Gelinas, L. Institutional review boards need new skills to review data sharing and management plans. Nat. Med. 2023, 29, 1307–1309. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Charles, W.M.; van der Waal, M.B.; Flach, J.; Bisschop, A.; van der Waal, R.X.; Es-Sbai, H.; McLeod, C.J. Blockchain-based dynamic consent and its applications for patient-centric research and health information sharing: Protocol for an integrative review. JMIR Res. Protoc. 2024, 13, e50339. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Concato, J.; Corrigan-Curay, J. Real-World Evidence—Where Are We Now? N. Engl. J. Med. 2022, 386, 1680–1682. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Thangaraj, P.M.; Shankar, S.V.; Huang, S.; Nadkarni, G.N.; Mortazavi, B.J.; Oikonomou, E.K.; Khera, R. A Novel Digital Twin Strategy to Examine the Implications of Randomized Clinical Trials for Real-World Populations. medRxiv 2024. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Ouedraogo, E.B.; Hawbani, A.; Wang, X.; Liu, Z.; Zhao, L.; Al-qaness, M.A.A.; Alsamhi, S.H. Digital Twin Data Management: A Comprehensive Review. IEEE Trans. Big Data 2025, 11, 2224–2243. [Google Scholar] [CrossRef] [Scilit]
  26. Hasani, N.; Farhadi, F.; Morris, M.A.; Nikpanah, M.; Rhamim, A.; Xu, Y.; Pariser, A.; Collins, M.T.; Summers, R.M.; Jones, E.; et al. Artificial Intelligence in Medical Imaging and its Impact on the Rare Disease Community: Threats, Challenges and Opportunities. PET Clin. 2022, 17, 13–29. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Khoshfekr Rudsari, H.; Tseng, B.; Zhu, H.; Song, L.; Gu, C.; Roy, A.; Irajizad, E.; Butner, J.; Long, J.; Do, K.-A. Digital twins in healthcare: A comprehensive review and future directions. Front. Digit. Health 2025, 7, 1633539. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Ringeval, M.; Etindele Sosso, F.A.; Cousineau, M.; Paré, G. Advancing Health Care With Digital Twins: Meta-Review of Applications and Implementation Challenges. J. Med. Internet Res. 2025, 27, e69544. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Chaffer, T.J.; Littlejohn, J.; Nadarasa, A.; Lamschtein, C. The Self-Sovereign Patient as a Cornerstone of Healthcare 4.0. Blockchain Healthc. Today 2025, 8, 414. [Google Scholar]
  30. Kumar, A.H.; Venkatram, C.P.; N, S.; Daniel, D.; Joe, I.R.P. Decentralized digital health ecosystems: A unified architecture for AI-enhanced medical record management. Front. Digit. Health 2025, 7, 1685628. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Khan, I.; Maher, M.; Khurshid, A. Creating a Health Data Marketplace for the Digital Health Era. Blockchain Healthc. Today 2024, 7, 338. [Google Scholar] [CrossRef] [Scilit]
  32. Damar, M.; Aydın, Ö.; Erenay, F.S. Blockchain Technology in Digital Health and Medical Technologies. Blockchain Healthc. Today 2025, 8, 409. [Google Scholar] [CrossRef] [Scilit]
  33. Silva, P.J.; Ramos, K.S. 7.03—Trends and implementation of preemptive pharmacogenomic testing. In Comprehensive Precision Medicine, 1st ed.; Ramos, K.S., Ed.; Elsevier: Oxford, UK, 2024; pp. 363–381. [Google Scholar]
  34. Makady, A.; de Boer, A.; Hillege, H.; Klungel, O.; Goettsch, W. What Is Real-World Data? A Review of Definitions Based on Literature and Stakeholder Interviews. Value Health 2017, 20, 858–865. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Makady, A.; Ham, R.t.; de Boer, A.; Hillege, H.; Klungel, O.; Goettsch, W. Policies for Use of Real-World Data in Health Technology Assessment (HTA): A Comparative Study of Six HTA Agencies. Value Health 2017, 20, 520–532. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. De Domenico, M.; Allegri, L.; Caldarelli, G.; d’Andrea, V.; Di Camillo, B.; Rocha, L.M.; Rozum, J.; Sbarbati, R.; Zambelli, F. Challenges and opportunities for digital twins in precision medicine from a complex systems perspective. npj Digit. Med. 2025, 8, 37. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. You, J.G.; Hernandez-Boussard, T.; Pfeffer, M.A.; Landman, A.; Mishuris, R.G. Clinical trials informed framework for real world clinical implementation and deployment of artificial intelligence applications. npj Digit. Med. 2025, 8, 107. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Nadeem, M.; Kostic, S.; Dornhöfer, M.; Weber, C.; Fathi, M. A comprehensive review of digital twin in healthcare in the scope of simulative health-monitoring. Digit. Health 2025, 11, 20552076241304078. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Mitchell, J.R.; Szepietowski, P.; Howard, R.; Reisman, P.; Jones, J.D.; Lewis, P.; Fridley, B.L.; Rollison, D.E. A Question-and-Answer System to Extract Data From Free-Text Oncological Pathology Reports (CancerBERT Network): Development Study. J. Med. Internet Res. 2022, 24, e27210. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  40. Chen, Q.; Leng, T.; Niu, S.; Trucco, E. Editorial: Generalizable and explainable artificial intelligence methods for retinal disease analysis: Challenges and future trends. Front. Med. 2024, 11, 1465369. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Silva, P.J.; Rogers, S.L.; Hassan-Toufique, Z.; Tao, J.; Bruce, S.A.; Shireman, P.K.; Ramos, K.S. Idealized Framework for Pharmacovigilance in an Ambulatory Primary Care and Chronic Disease Management Clinic. Future Pharmacol. 2025; submitted. [CrossRef] [Scilit]
  42. Tibble, H.; Sheikh, A.; Tsanas, A. Estimating medication adherence from Electronic Health Records: Comparing methods for mining and processing asthma treatment prescriptions. BMC Med. Res. Methodol. 2023, 23, 167. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  43. Rogers, S.; Silva, P.J.; Udeani, G.; Deleon, M.; Mutyala, S.; Panahi, L.; Abu-Baker, A.; Neal, G.; Ramos, K.S. Case Report: Life-Threatening Fluoxetine-Linked Postoperative Bleeding Informed by Pharmacogenetic Evaluation. Drugs RD 2024, 24, 117–121. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Barbaria, S.; Jemai, A.; Ceylan, H.İ.; Muntean, R.I.; Dergaa, I.; Boussi Rahmouni, H. Advancing Compliance with HIPAA and GDPR in Healthcare: A Blockchain-Based Strategy for Secure Data Exchange in Clinical Research Involving Private Health Information. Healthcare 2025, 13, 2594. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. Gross, M.S.; Hood, A.J.; Miller, R.C., Jr. Nonfungible Tokens as a Blockchain Solution to Ethical Challenges for the Secondary Use of Biospecimens: Viewpoint. JMIR Bioinform. Biotech. 2021, 2, e29905. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Barnes, C.; Aboy, M.R.; Minssen, T.; Allen, J.W.; Earp, B.D.; Savulescu, J.; Mann, S.P. Enabling Demonstrated Consent for Biobanking with Blockchain and Generative AI. Am. J. Bioeth. 2025, 25, 96–111. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Kostick-Quenet, K.M.; Compagnucci, M.C.; Aboy, M.; Minssen, T. Patient-centric federated learning: Automating meaningful consent to health data sharing with smart contracts. J. Law Biosci. 2025, 12, lsaf003. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Wan, W.; Silva, R.; Odenweller, D.J.; Leeuwon, S. 6.03—Proprietary strategies in precision medicine. In Comprehensive Precision Medicine, 1st ed.; Ramos, K.S., Ed.; Elsevier: Oxford, UK, 2024; pp. 197–220. [Google Scholar]
  49. Silva, P.J.; Ramos, K.S. Academic Medical Centers as Innovation Ecosystems: Evolution of Industry Partnership Models Beyond the Bayh–Dole Act. Acad. Med. 2018, 93, 1135–1141. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  50. Silva, P.J.; Schaibley, V.M.; Ramos, K.S. Academic medical centers as innovation ecosystems to address population –omics challenges in precision medicine. J. Transl. Med. 2018, 16, 28. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  51. Li, J.; Wang, D. Federated learning for digital twin applications: A privacy-preserving and low-latency approach. PeerJ Comput. Sci. 2025, 11, e2877. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  52. Rozenblit, L.; Price, A.; Solomonides, A.; Joseph, A.L.; Koski, E.; Srivastava, G.; Labkoff, S.; Bray, D.; Lopez-Gonzalez, M.; Singh, R. Toward responsible AI governance: Balancing multi-stakeholder perspectives on AI in healthcare. Int. J. Med. Inform. 2025, 203, 106015. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Wilkinson, M.D.; Dumontier, M.; Aalbersberg, I.J.; Appleton, G.; Axton, M.; Baak, A.; Blomberg, N.; Boiten, J.W.; Santos, L.B.D.S.; Bourne, P.E. The FAIR Guiding Principles for scientific data management and stewardship: Comment. Sci. Data 2016, 3, 160018. [Google Scholar] [PubMed]
  54. Charles, W.M.; Delgado, B.M. Health datasets as assets: Blockchain-based valuation and transaction methods. Blockchain Healthc. Today 2022, 5. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Katsoulakis, E.; Wang, Q.; Wu, H.; Shahriyari, L.; Fletcher, R.; Liu, J.; Achenie, L.; Liu, H.; Jackson, P.; Xiao, Y.; et al. Digital twins for health: A scoping review. npj Digit. Med. 2024, 7, 77. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  56. Nordberg, A. AI, Big Data, e-Health and the Right to be Forgotten. In Kunstig Intelligens og Big Data i Helsesektoren: Rettslige Perspektiver; Gyldendal Norsk Forlag A/S: Oslo, Norway, 2020. [Google Scholar]
  57. Zhang, H.; Nakamura, T.; Isohara, T.; Sakurai, K. A review on machine unlearning. SN Comput. Sci. 2023, 4, 337. [Google Scholar] [CrossRef] [Scilit]
  58. Krishappa, N.; Shivappa, G.G.; Zachariah, S.; Thanushree; Pattan, K.I.; Paria, A.; Hiremath, S.; Vaithiyanathan, R. A Blockchain-Based Framework with Zero-Knowledge Proof Incorporated for Safeguarded Sharing of Genomic Data Through Health Record Systems. Blockchain Healthc. Today 2025, 8, 419. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  59. Zhang, Y.; Sui, X.; Pan, F.; Yu, K.; Li, K.; Tian, S.; Erdengasileng, A.; Han, Q.; Wang, W.; Wang, J.; et al. A comprehensive large-scale biomedical knowledge graph for AI-powered data-driven biomedical research. Nat. Mach. Intell. 2025, 7, 602–614. [Google Scholar] [CrossRef] [Scilit]
  60. Cheng, H.; Wu, Y.; Khatwani, S.; Kruse, M.; Dligach, D.; Miller, T.A.; Afshar, M.; Gao, Y. Scaling Biomedical Knowledge Graph Retrieval for Interpretable Reasoning: Applications to Clinical Diagnosis Prediction. medRxiv 2026. medRxiv:2026.2001.2012.26343957. [Google Scholar] [CrossRef] [Scilit]
  61. Kreuzthaler, M.; Brochhausen, M.; Zayas, C.; Blobel, B.; Schulz, S. Linguistic and ontological challenges of multiple domains contributing to transformed health ecosystems. Front. Med. 2023, 10, 1073313. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  62. Adimulam, A.; Gupta, R.; Kumar, S. The Orchestration of Multi-Agent Systems: Architectures, Protocols, and Enterprise Adoption. arXiv 2026, arXiv:2601.13671. [Google Scholar]
  63. Carbonaro, A.; Marfoglia, A.; Nardini, F.; Mellone, S. CONNECTED: Leveraging digital twins and personal knowledge graphs in healthcare digitalization. Front. Digit. Health 2023, 5, 1322428. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  64. Wang, Y.; Guo, S.; Pan, Y.; Su, Z.; Chen, F.; Luan, T.H.; Li, P.; Kang, J.; Niyato, D. Internet of agents: Fundamentals, applications, and challenges. IEEE Trans. Cogn. Commun. Netw. 2025, 12, 4476–4501. [Google Scholar] [CrossRef] [Scilit]
  65. Wu, J.; Or, C.K. Position paper: Towards open complex human-AI agents collaboration systems for problem solving and knowledge management. arXiv 2025, arXiv:2505.00018. [Google Scholar]
  66. Milani, L.; Alver, M.; Laur, S.; Reisberg, S.; Haller, T.; Aasmets, O.; Abner, E.; Alavere, H.; Allik, A.; Annilo, T.; et al. From Biobanking to Personalized Medicine: The journey of the Estonian Biobank. medRxiv 2024. medRxiv:2024.2009.2022.24313964. [Google Scholar] [CrossRef] [Scilit]
  67. Porsdam Mann, S.; Savulescu, J.; Ravaud, P.; Benchoufi, M. Blockchain, consent and prosent for medical research. J. Med. Ethics 2021, 47, 244. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  68. Ghatti, S.; Yurish, L.A.; Shen, H.; Rheuban, K.; Enfield, K.B.; Facteau, N.R.; Engel, G.; Dowdell, K. Digital twins in healthcare: A survey of current methods. Arch. Clin. Biomed. Res. 2023, 7, 365–381. [Google Scholar] [CrossRef] [Scilit]
  69. Smith, L.A.; Cahill, J.A.; Lee, J.-H.; Graim, K. Equitable machine learning counteracts ancestral bias in precision medicine. Nat. Commun. 2025, 16, 2144. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  70. Spratt, D.E.; Chan, T.; Waldron, L.; Speers, C.; Feng, F.Y.; Ogunwobi, O.O.; Osborne, J.R. Racial/Ethnic Disparities in Genomic Sequencing. JAMA Oncol. 2016, 2, 1070–1074. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  71. Ory, M.G.; Adepoju, O.E.; Ramos, K.S.; Silva, P.S.; Vollmer Dahlke, D. Health equity innovation in precision medicine: Current challenges and future directions. Front. Public Health 2023, 11, 1119736. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  72. Meijer, C.; Uh, H.W.; El Bouhaddani, S. Digital Twins in Healthcare: Methodological Challenges and Opportunities. J. Pers. Med. 2023, 13, 1522. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  73. Paik, K.E.; Hicklen, R.; Kaggwa, F.; Puyat, C.V.; Nakayama, L.F.; Ong, B.A.; Shropshire, J.N.; Villanueva, C. Digital Determinants of Health: Health data poverty amplifies existing health disparities—A scoping review. PLoS Digit. Health 2023, 2, e0000313. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  74. Vaidya, A.; Chen, R.J.; Williamson, D.F.K.; Song, A.H.; Jaume, G.; Yang, Y.; Hartvigsen, T.; Dyer, E.C.; Lu, M.Y.; Lipkova, J.; et al. Demographic bias in misdiagnosis by computational pathology models. Nat. Med. 2024, 30, 1174–1190. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  75. Changalidis, A.; Barbitoff, Y.; Nasykhova, Y.; Glotov, A. A systematic review on the generative AI applications in human medical genetics. Front. Genet. 2026, 16, 1694070. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  76. Martin, A.R.; Kanai, M.; Kamatani, Y.; Okada, Y.; Neale, B.M.; Daly, M.J. Clinical use of current polygenic risk scores may exacerbate health disparities. Nat. Genet. 2019, 51, 584–591. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  77. Yang, J.; Soltan, A.A.S.; Eyre, D.W.; Yang, Y.; Clifton, D.A. An adversarial training framework for mitigating algorithmic biases in clinical machine learning. npj Digit. Med. 2023, 6, 55. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  78. Norori, N.; Hu, Q.; Aellen, F.M.; Faraci, F.D.; Tzovara, A. Addressing bias in big data and AI for health care: A call for open science. Patterns 2021, 2, 100347. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  79. Walton, N.A.; Nagarajan, R.; Wang, C.; Sincan, M.; Freimuth, R.R.; Everman, D.B.; Walton, D.C.; McGrath, S.P.; Lemas, D.J.; Benos, P.V.; et al. Enabling the clinical application of artificial intelligence in genomics: A perspective of the AMIA Genomics and Translational Bioinformatics Workgroup. J. Am. Med. Inform. Assoc. 2024, 31, 536–541. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  80. Janjan, N.; Silva, P.J.; Ramos, K.S.; Ory, M.G.; Smith, M.L. 6.02—Alternative evidence in drug development and regulatory science. In Comprehensive Precision Medicine, 1st ed.; Ramos, K.S., Ed.; Elsevier: Oxford, UK, 2024; pp. 180–196. [Google Scholar]
  81. Schilling, M.P.; Scherr, T.; Münke, F.R.; Neumann, O.; Schutera, M.; Mikut, R.; Reischl, M. Automated Annotator Variability Inspection for Biomedical Image Segmentation. IEEE Access 2022, 10, 2753–2765. [Google Scholar] [CrossRef] [Scilit]
  82. Tanno, R.; Saeedi, A.; Sankaranarayanan, S.; Alexander, D.C.; Silberman, N. Learning from noisy labels by regularized estimation of annotator confusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 11244–11253. [Google Scholar]
  83. Rogers, S.S. Artificial Intelligence in Pharmacogenomics: From Discovery to Clinical Decision Support. In Pharmacogenomics: A Primer for Clinical Practice, 2nd ed.; Lam, J., Ed.; McGraw-Hill: New York, NY, USA, 2026; in Press. [Google Scholar]
  84. Vallée, A. Envisioning the Future of Personalized Medicine: Role and Realities of Digital Twins. J. Med. Internet Res. 2024, 26, e50204. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  85. Vallée, A. Digital Twins for Personalized Medicine Require Epidemiological Data and Mathematical Modeling: Viewpoint. J. Med. Internet Res. 2025, 27, e72411. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  86. Weinberger, N.; Hery, D.; Mahr, D.; Adler, S.O.; Stadlbauer, J.; Ahrens, T.D. Beyond the gender data gap: Co-creating equitable digital patient twins. Front. Digit. Health 2025, 7, 1584415. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  87. Dwyer-Lindgren, L.; Bertozzi-Villa, A.; Stubbs, R.W.; Morozoff, C.; Mackenbach, J.P.; van Lenthe, F.J.; Mokdad, A.H.; Murray, C.J.L. Inequalities in Life Expectancy Among US Counties, 1980 to 2014: Temporal Trends and Key Drivers. JAMA Intern. Med. 2017, 177, 1003–1011. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  88. Zhang, S.; Chu, S.; Qiang, Y.; Zhao, J.; Wang, Y.; Wei, X. Combating Medical Label Noise through more precise partition-correction and progressive hard-enhanced learning. Comput. Methods Programs Biomed. 2025, 265, 108734. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  89. Oh, I.Y.; Schindler, S.E.; Ghoshal, N.; Lai, A.M.; Payne, P.R.O.; Gupta, A. Extraction of clinical phenotypes for Alzheimer’s disease dementia from clinical notes using natural language processing. JAMIA Open 2023, 6, ooad014. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  90. Kim, P.; Milliken, E.L. Minority Participation in Biobanks: An Essential Key to Progress. Methods Mol. Biol. 2019, 1897, 43–50. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  91. De-la-Rosa-Martinez, D.; Rivera-Buendía, F.; Cornejo-Juárez, P.; García-Pineda, B.; Nevárez-Luján, C.; Vilar-Compte, D. Risk factors and clinical outcomes for Clostridioides difficile infections in a case control study at a large cancer referral center in Mexico. Am. J. Infect. Control 2022, 50, 1220–1225. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  92. Geneviève, L.D.; Martani, A.; Shaw, D.; Elger, B.S.; Wangmo, T. Structural racism in precision medicine: Leaving no one behind. BMC Med. Ethics 2020, 21, 17. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  93. Press, W.H. What’s So Special About Science (And How Much Should We Spend on It?). Science 2013, 342, 817–822. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  94. Bild, D.E.; Bluemke, D.A.; Burke, G.L.; Detrano, R.; Diez Roux, A.V.; Folsom, A.R.; Greenland, P.; Jacob, D.R., Jr.; Kronmal, R.; Liu, K.; et al. Multi-Ethnic Study of Atherosclerosis: Objectives and design. Am. J. Epidemiol. 2002, 156, 871–881. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  95. Venkatesh, K.P.; Brito, G.; Kamel Boulos, M.N. Health digital twins in life science and health care innovation. Annu. Rev. Pharmacol. Toxicol. 2024, 64, 159–170. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  96. Hasanzadeh, F.; Josephson, C.B.; Waters, G.; Adedinsewo, D.; Azizi, Z.; White, J.A. Bias recognition and mitigation strategies in artificial intelligence healthcare applications. npj Digit. Med. 2025, 8, 154. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  97. Gu, T.; Pan, W.; Yu, J.; Ji, G.; Meng, X.; Wang, Y.; Li, M. Mitigating bias in AI mortality predictions for minority populations: A transfer learning approach. BMC Med. Inform. Decis. Mak. 2025, 25, 30. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  98. Li, C.; Ding, S.; Zou, N.; Hu, X.; Jiang, X.; Zhang, K. Multi-task learning with dynamic re-weighting to achieve fairness in healthcare predictive modeling. J. Biomed. Inform. 2023, 143, 104399. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  99. Jia, Z.; Zhou, T.; Yan, Z.; Hu, J.; Shi, Y. Personalized Meta-Federated Learning for IoT-Enabled Health Monitoring. IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024, 43, 3157–3170. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  100. Muhammad, J.; Ghergherehchi, M.; Ali, S.; Seung Song, H.; Rahim, N. Trustworthy AI for medical decisions: Adversarially robust and fair machine learning prediction for Parkinson’s disease. PLoS ONE 2026, 21, e0342062. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  101. Carvalho, E.; Mascarenhas, M.; Pinheiro, F.; Correia, R.; Balseiro, S.; Barbosa, G.; Guerra, A.; Oliveira, D.; Moura, R.; Martins dos Santos, A.; et al. Predetermined Change Control Plans: Guiding Principles for Advancing Safe, Effective, and High-Quality AI-ML Technologies. JMIR AI 2025, 4, e76854. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  102. Sel, K.; Hawkins-Daarud, A.; Chaudhuri, A.; Osman, D.; Bahai, A.; Paydarfar, D.; Willcox, K.; Chung, C.; Jafari, R. Survey and perspective on verification, validation, and uncertainty quantification of digital twins for precision medicine. npj Digit. Med. 2025, 8, 40. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  103. Wolff, R.F.; Moons, K.G.M.; Riley, R.D.; Whiting, P.F.; Westwood, M.; Collins, G.S.; Reitsma, J.B.; Kleijnen, J.; Mallett, S. PROBAST: A Tool to Assess the Risk of Bias and Applicability of Prediction Model Studies. Ann. Intern. Med. 2019, 170, 51–58. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  104. Milne-Ives, M.; Fraser, L.K.; Khan, A.; Walker, D.; van Velthoven, M.H.; May, J.; Wolfe, I.; Harding, T.; Meinert, E. Life course digital twins–intelligent monitoring for early and continuous intervention and prevention (LifeTIME): Proposal for a retrospective cohort study. JMIR Res. Protoc. 2022, 11, e35738. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  105. Olawade, D.B.; Odetayo, A.; Egbon, E.; Olasilola, O.R.; Makanjuola, B.D.; Daniel, R.I.A. Implementing digital twin technology in organ transplantation: Concepts, emerging evidence, and clinical translation pathways. Transpl. Rev. 2026, 40, 101004. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  106. Askin, S.; Burkhalter, D.; Calado, G.; El Dakrouni, S. Artificial Intelligence Applied to clinical trials: Opportunities and challenges. Health Technol. 2023, 13, 203–213. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  107. Davis, S.E.; Greevy, R.A., Jr.; Lasko, T.A.; Walsh, C.G.; Matheny, M.E. Detection of calibration drift in clinical prediction models to inform model updating. J. Biomed. Inform. 2020, 112, 103611. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  108. Maheshkar, J.A.; Vankayala, H.; Jakkula, V.K.; Raj, L.D.; Khedekar, P.; Laheri, R. Agentic Ai-Powered Autonomous Software Engineering Framework for Automated Code Generation and Debugging. Sci. Cult. 2026, 12, 2816–2822. [Google Scholar]
  109. Tiwari, A.; Mishra, S.; Kuo, T.-R. Current AI technologies in cancer diagnostics and treatment. Mol. Cancer 2025, 24, 159. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  110. Süalp, E.; Rezaei, M. Mitigating catastrophic forgetting in continual learning through model growth. arXiv 2025, arXiv:2509.01213. [Google Scholar]
  111. Gu, L.; Zhu, Y.; Sang, H.; Wang, Z.; Sui, D.; Tang, W.; Harrison, E.; Gao, J.; Yu, L.; Ma, L. MedAgentAudit: Diagnosing and Quantifying Collaborative Failure Modes in Medical Multi-Agent Systems. arXiv 2025, arXiv:2510.10185. [Google Scholar]
  112. Duffourc, M.N.; Gerke, S. The proposed EU Directives for AI liability leave worrying gaps likely to impact medical AI. npj Digit. Med. 2023, 6, 77. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  113. Zhang, C. The legal dilemma of medical artificial intelligence in China: Challenges to physicians’ duty to inform and a typology-based response. Front. Public Health 2025, 13, 1747635. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  114. Ferreira-da-Silva, R.; Cruz-Correia, R.; Ribeiro, I. Beyond black boxes: Using explainable causal artificial intelligence to separate signal from noise in pharmacovigilance. Int. J. Clin. Pharm. 2025, 48, 677–681. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  115. Edwards, C.; Murphy, A.; Singh, A.; Daniel, S.; Chamunyonga, C. The role of patient outcomes in shaping moral responsibility in AI-supported decision making. Radiography 2025, 31, 102948. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  116. Kim, J.; Lui, B.; Goldstein, P.A.; Rubin, J.E.; White, R.S.; Jotwani, R. From Data to Decisions: Harnessing Multi-Agent Systems for Safer, Smarter, and More Personalized Perioperative Care. J. Pers. Med. 2025, 15, 540. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  117. Cáceres-Gutiérrez, D.A.; Bonilla-Bonilla, D.M.; Liscano, Y.; Díaz Vallejo, J.A. From Architecture to Outcomes: Mapping the Landscape of Digital Twins for Personalized Diabetes Care-A Scoping Review. J. Pers. Med. 2025, 15, 504. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  118. Elkefi, S.; Asan, O. Digital Twins for Managing Health Care Systems: Rapid Literature Review. J. Med. Internet Res. 2022, 24, e37641. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  119. Krupas, M.; Kajati, E.; Liu, C.; Zolotova, I. Towards a human-centric digital twin for human–machine collaboration: A review on enabling technologies and methods. Sensors 2024, 24, 2232. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  120. Boudi, A.L.; Boudi, M.; Chan, C.; Boudi, F.B. Ethical Challenges of Artificial Intelligence in Medicine. Cureus 2024, 16, e74495. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  121. He, Q.; Xu, S.; Zhu, Z.; Wang, P.; Li, K.; Zheng, Q.; Li, Y. KRP-DS: A knowledge graph-based dialogue system with inference-aided prediction. Sensors 2023, 23, 6805. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  122. Mohsin, M.T.; Abdulrashid, I. Provenance-Aware Explainable Digital Twin for Personalized Health Management. medRxiv 2025. [Google Scholar] [CrossRef] [Scilit]
  123. Kwon, D. ActX and CompuGroup Partner to Bring Genomic Data to Electronic Health Records. Clin. OMICs 2017, 4, 31. [Google Scholar] [CrossRef] [Scilit]
  124. Riahi, V.; Diouf, I.; Khanna, S.; Boyle, J.; Hassanzadeh, H. Digital twins for clinical and operational decision-making: Scoping review. J. Med. Internet Res. 2025, 27, e55015. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  125. Alharthi, S. AI-powered in silico twins: Redefining precision medicine through simulation, personalization, and predictive healthcare. Saudi Pharm. J. 2025, 34, 1. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  126. Vallée, A. From prediction to intervention: Causal digital twins for personalized clinical decision support. J. Transl. Med. 2026, 24, 441. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  127. Udeh, C.; Udeh, B.; Rahman, N.; Canfield, C.; Campbell, J.; Hata, J.S. Telemedicine/Virtual ICU: Where Are We and Where Are We Going? Methodist. Debakey Cardiovasc. J. 2018, 14, 126–133. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  128. Pollard, T.J.; Johnson, A.E.; Raffa, J.D.; Celi, L.A.; Mark, R.G.; Badawi, O. The eICU Collaborative Research Database, a freely available multi-center database for critical care research. Sci. Data 2018, 5, 180178. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  129. Zhang, K.; Zhou, H.Y.; Baptista-Hon, D.T.; Gao, Y.; Liu, X.; Oermann, E.; Xu, S.; Jin, S.; Zhang, J.; Sun, Z.; et al. Concepts and applications of digital twins in healthcare and medicine. Patterns 2024, 5, 101028. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  130. Tasmurzayev, N.; Amangeldy, B.; Imanbek, B.; Baigarayeva, Z.; Imankulov, T.; Dikhanbayeva, G.; Amangeldi, I.; Sharipova, S. Digital Cardiovascular Twins, AI Agents, and Sensor Data: A Narrative Review from System Architecture to Proactive Heart Health. Sensors 2025, 25, 5272. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  131. Barat, S.; Parchure, R.; Darak, S.; Kulkarni, V.; Paranjape, A.; Gajrani, M.; Yadav, A.; Kulkarni, V. An Agent-Based Digital Twin for Exploring Localized Non-pharmaceutical Interventions to Control COVID-19 Pandemic. Trans. Indian. Natl. Acad. Eng. 2021, 6, 323–353. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  132. Li, J.; Lai, Y.; Li, W.; Ren, J.; Zhang, M.; Kang, X.; Wang, S.; Li, P.; Zhang, Y.-Q.; Ma, W. Agent hospital: A simulacrum of hospital with evolvable medical agents. arXiv 2024, arXiv:2405.02957. [Google Scholar]
  133. Kuruppu Appuhamilage, G.D.K.; Hussain, M.; Zaman, M.; Ali Khan, W. A health digital twin framework for discrete event simulation based optimised critical care workflows. npj Digit. Med. 2025, 8, 376. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  134. Yu, C.; He, Y.; Cheng, H.; Liu, Z.; Mu, D.; Shen, Z.; Jin, Z. From Passive to Proactive: A Multi-Agent System with Dynamic Task Orchestration for Intelligent Medical Pre-Consultation. arXiv 2025, arXiv:2511.01445. [Google Scholar]
  135. Pellegrino, G.; Gervasi, M.; Angelelli, M.; Corallo, A. A Conceptual Framework for Digital Twin in Healthcare: Evidence from a Systematic Meta-Review. Inf. Syst. Front. 2025, 27, 7–32. [Google Scholar] [CrossRef] [Scilit]
  136. Papachristou, K.; Katsakiori, P.F.; Papadimitroulas, P.; Strigari, L.; Kagadis, G.C. Digital Twins’ Advancements and Applications in Healthcare, Towards Precision Medicine. J. Pers. Med. 2024, 14, 1101. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  137. Halpern, G.A.; Nemet, M.; Gowda, D.M.; Kilickaya, O.; Lal, A. Advances and utility of digital twins in critical care and acute care medicine: A narrative review. J. Yeungnam Med. Sci. 2025, 42, 9. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  138. Mutlag, A.A.; Ghani, M.K.A.; Mohammed, M.A.; Lakhan, A.; Mohd, O.; Abdulkareem, K.H.; Garcia-Zapirain, B. Multi-Agent Systems in Fog-Cloud Computing for Critical Healthcare Task Management Model (CHTM) Used for ECG Monitoring. Sensors 2021, 21, 6923. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  139. Haider, E.; Mujahid, M.U.F.; Hidig, S.M. Harnessing hybrid edge-fog computing framework for emergency medical response management. Ann. Med. Surg. 2026, 88, 1098–1099. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  140. Mahmood, K.; Khan, S.; Abdelhaq, M.; Hassan, M.U.; Uddin, M.; Alsaqour, R.; Awan, K.A.; Alsoufi, M.A. Adaptive resource aware and privacy preserving federated edge learning framework for real time internet of medical things applications. Sci. Rep. 2025, 15, 36468. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  141. Gross, M.S.; Miller, R.C., Jr. Ethical implementation of the learning healthcare system with blockchain technology. Blockchain Healthc. Today Forthcom. 2019, 2. [Google Scholar] [CrossRef] [Scilit]
  142. Hemdan, E.E.; Sayed, A. Smart and Secure Healthcare with Digital Twins: A Deep Dive into Blockchain, Federated Learning, and Future Innovations. Algorithms 2025, 18, 401. [Google Scholar] [CrossRef] [Scilit]
  143. Ali, M.; Naeem, F.; Tariq, M.; Kaddoum, G. Federated Learning for Privacy Preservation in Smart Healthcare Systems: A Comprehensive Survey. IEEE J. Biomed. Health Inform. 2023, 27, 778–789. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  144. Charles, W. 1.09—Data continuity and linkage in the healthcare ecosystem. In Comprehensive Precision Medicine, 1st ed.; Ramos, K.S., Ed.; Elsevier: Oxford, UK, 2024; pp. 120–143. [Google Scholar]
  145. Charles, W. Data continuity and linkage in the healthcare ecosystem. J. Med. Internet Res. 2024, 26, e60258. [Google Scholar]
  146. Vasiliu-Feltes, I.; Mylrea, M.; Zhang, C.Y.; Wood, T.C.; Thornley, B. Impact of Blockchain-Digital Twin Technology on Precision Health, Pharmaceutical Industry, and Life Sciences: Conference Proceedings, Conv2X 2023. Blockchain Heal. Today 2023, 6. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  147. Esmaeilzadeh, P.; Mirzaei, T. Role of Incentives in the Use of Blockchain-Based Platforms for Sharing Sensitive Health Data: Experimental Study. J. Med. Internet Res. 2023, 25, e41805. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  148. Amofa, S.; Xia, Q.; Xia, H.; Obiri, I.A.; Adjei-Arthur, B.; Yang, J.; Gao, J. Blockchain-secure patient Digital Twin in healthcare using smart contracts. PLoS ONE 2024, 19, e0286120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  149. Suleiman, R.; Maradapu Vera Venkata Sai, A.; Yu, W.; Wang, C. Blockchain for Security in Digital Twins. Future Internet 2025, 17, 385. [Google Scholar] [CrossRef] [Scilit]
  150. Ni, E.; Tang, X.; Zhou, X.; Lee, D.; Elhussein, A.; Knight, E.; Gürsoy, G.; Gerstein, M. Recent advances and future prospects for blockchain in biomedicine. Cell Rep. Methods 2025, 5, 101114. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  151. Anik, F.I.; Sakib, N.; Shahriar, H.; Xie, Y.; Nahiyan, H.A.; Ahamed, S.I. Unraveling a blockchain-based framework towards patient empowerment: A scoping review envisioning future smart health technologies. Smart Health 2023, 29, 100401. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  152. Brogan, J.; Baskaran, I.; Ramachandran, N. Authenticating Health Activity Data Using Distributed Ledger Technologies. Comput. Struct. Biotechnol. J. 2018, 16, 257–266. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  153. Kuo, T.T.; Pham, A. Quorum-based model learning on a blockchain hierarchical clinical research network using smart contracts. Int. J. Med. Inform. 2023, 169, 104924. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  154. Bonotis, P.; Angelidis, P.; Tzimourta, K.D.; Bibi, S. Enabling Dynamic Consent Through AI and Blockchain: The CONSENT Platform. Stud. Health Technol. Inform. 2025, 332, 330–334. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  155. Xu, G.; Yao, T.; Zhang, K.; Meng, X.; Liu, X.; Xiao, K.; Chen, X. An Optimized Byzantine Fault Tolerance Algorithm for Medical Data Security. Electronics 2023, 12, 5045. [Google Scholar] [CrossRef] [Scilit]
  156. Albalwy, F.; Brass, A.; Davies, A. A blockchain-based dynamic consent architecture to support clinical genomic data sharing (ConsentChain): Proof-of-concept study. JMIR Med. Inform. 2021, 9, e27816. [Google Scholar] [PubMed]
  157. Lee, A.R.; Koo, D.; Kim, I.K.; Lee, E.; Yoo, S.; Lee, H.Y. Opportunities and challenges of a dynamic consent-based application: Personalized options for personal health data sharing and utilization. BMC Med. Ethics 2024, 25, 92. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  158. Prictor, M.; Teare, H.J.A.; Kaye, J. Equitable Participation in Biobanks: The Risks and Benefits of a “Dynamic Consent” Approach. Front. Public Health 2018, 6, 253. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  159. Niu, Y.A. Leveraging Blockchain Technology for Enhancing Genomic Data Management: A Multidisciplinary Framework for Privacy, Trust, Identity Protection, and Equity; Massachusetts Institute of Technology: Cambridge, MA, USA, 2025. [Google Scholar]
  160. Kaczynski, S.; Kominers, S.D. The Everything Token: How NFTs and Web3 Will Transform the Way We Buy, Sell, and Create; Penguin: New York, NY, USA, 2024. [Google Scholar]
  161. Dixon, C. Read Write Own: Building the Next Era of the Internet; Random House Publishing Group: New York, NY, USA, 2024. [Google Scholar]
  162. Yang, Y.; Liu, M.; Chen, H.; Chen, L. Application of blockchain-based digital twin technology in healthcare: A scoping review. Comput. Methods Programs Biomed. 2026, 276, 109231. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  163. Mallick, S.R.; Lenka, R.K.; Sobhanayak, S. Secure and scalable dual blockchain and IPFS driven IoT ecosystem for next gen healthcare systems. Sci. Rep. 2025, 15, 41064. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  164. Tawfik, A.M.; Al-Ahwal, A.; Eldien, A.S.T.; Zayed, H.H. ACHealthChain blockchain framework for access control and privacy preservation in healthcare. Sci. Rep. 2025, 15, 16696. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  165. Schmeelk, S.; Kanabar, M.; Peterson, K.; Pathak, J. Electronic health records and blockchain interoperability requirements: A scoping review. JAMIA Open 2022, 5, ooac068. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  166. Kasyapa, M.S.B.; Vanmathi, C. Blockchain integration in healthcare: A comprehensive investigation of use cases, performance issues, and mitigation strategies. Front. Digit. Health 2024, 6, 1359858. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  167. Lim, E.C.N.; Lim, C.E.D. Polycentric AI Governance: A Multi-Stakeholder Approach to Distributed Responsibility and Ethical Technology Management. Int. J. Adv. AI Appl. 2025, 1, 77–97. [Google Scholar]
Figure 1. Barriers to Clinical Implementation of Health Digital Twins. This figure emphasizes data stewardship, governance, interoperability, and provenance barriers.
Figure 1. Barriers to Clinical Implementation of Health Digital Twins. This figure emphasizes data stewardship, governance, interoperability, and provenance barriers.
Aimed 01 00020 g001
Figure 2. Striking the balance among data sources for health digital health twins. Clinical trial and research data provide internal rigor but may lack ecological and external validity, while RWD offers broader real-world representation but may suffer from missingness, bias, and limited grounding in adjudicated truth. This figure focuses on data-source and governance tradeoffs rather than the full workflow of model construction, validation, or updating.
Figure 2. Striking the balance among data sources for health digital health twins. Clinical trial and research data provide internal rigor but may lack ecological and external validity, while RWD offers broader real-world representation but may suffer from missingness, bias, and limited grounding in adjudicated truth. This figure focuses on data-source and governance tradeoffs rather than the full workflow of model construction, validation, or updating.
Aimed 01 00020 g002
Figure 3. Schematic of technology tactics to address data curation and governance impediments to DT implementation. Fragmentation, privacy preservation, representational bias, annotation burden, phenotypic truth, governance frictions, model and data provenance all represent impediments. Challenges and tactics to address them are provided.
Figure 3. Schematic of technology tactics to address data curation and governance impediments to DT implementation. Fragmentation, privacy preservation, representational bias, annotation burden, phenotypic truth, governance frictions, model and data provenance all represent impediments. Challenges and tactics to address them are provided.
Aimed 01 00020 g003
Figure 4. Outlines how digital twins move from data collection through training, validation, and deployment, while showing that governance, provenance, interoperability, and data-stewardship barriers can disrupt every stage of the process.
Figure 4. Outlines how digital twins move from data collection through training, validation, and deployment, while showing that governance, provenance, interoperability, and data-stewardship barriers can disrupt every stage of the process.
Aimed 01 00020 g004
Table 1. Comparison of Real-World Data and Electronic Medical Records as Data.
Table 1. Comparison of Real-World Data and Electronic Medical Records as Data.
Traditional Randomized TrialsTraditional Real-World Data (EMRs)
Cohort DiversityPoor (Strict inclusion/exclusion criteria)High (Reflects actual populations)
Longitudinal ContinuityLimited (Short follow-up durations)Fragmented (Lost between provider networks)
Ecological ValidityLow (Fails to reflect routine practice)High (Captures social determinants)
Ground Truth QualityHigh (Adjudicated outcomes)Low (Administrative coding, missing adherence)
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Silva, P.J.; Tao, J.; Rogers, S.L.; He, Q.; Robert, J.D.; Black, L.; Bruce, S.A.; Shireman, P.K.; Ramos, K.S. Data Stewardship Barriers to Building Digital Twin Technology for Precision Medicine. AI Med. 2026, 1, 20. https://doi.org/10.3390/aimed1030020

AMA Style

Silva PJ, Tao J, Rogers SL, He Q, Robert JD, Black L, Bruce SA, Shireman PK, Ramos KS. Data Stewardship Barriers to Building Digital Twin Technology for Precision Medicine. AI in Medicine. 2026; 1(3):20. https://doi.org/10.3390/aimed1030020

Chicago/Turabian Style

Silva, Patrick J., Jian Tao, Sara L. Rogers, Qiang He, Joshua D. Robert, Lance Black, Scott A. Bruce, Paula K. Shireman, and Kenneth S. Ramos. 2026. "Data Stewardship Barriers to Building Digital Twin Technology for Precision Medicine" AI in Medicine 1, no. 3: 20. https://doi.org/10.3390/aimed1030020

APA Style

Silva, P. J., Tao, J., Rogers, S. L., He, Q., Robert, J. D., Black, L., Bruce, S. A., Shireman, P. K., & Ramos, K. S. (2026). Data Stewardship Barriers to Building Digital Twin Technology for Precision Medicine. AI in Medicine, 1(3), 20. https://doi.org/10.3390/aimed1030020

Article Metrics

Back to TopTop