Abstract
Randomized controlled trials (RCTs) are often considered the gold standard for causal inference, but their implementation can be costly, time-consuming, and sometimes infeasible due to ethical or practical constraints. The so-called target trial emulation framework introduced the systematic use of observational data for treatment effect. This approach necessitates the detailed specification of a hypothetical trial protocol including eligibility criteria, treatment strategies, and outcome measures, which are then emulated by utilizing observational data. We expanded the target trial framework by integrating drug pathway embeddings and causal modeling, enabling prediction of treatment outcomes for unseen or held-out mechanisms of action based on the embedding relationships among existing therapies. We demonstrate that embedding-based models can reliably predict the direction of observed clinical outcomes across diverse therapeutic classes (e.g., small molecules, biologics), even when masking the observational data for the particular mechanism being estimated, though the precise magnitude of treatment effect remains hard to recover. This approach illustrates the potential to estimate the clinical efficacy of new drug mechanisms and to enhance the precision of future trial design and operations.
1. Introduction
1.1. Context
Randomized controlled trials (RCTs) are universally recognized as the gold standard for evaluating the efficacy and safety of investigational drugs in humans. However, RCTs remain expensive, time-consuming, and operationally complex to implement [1,2]. Given this pressure, clinical development leaders are seeking ways to improve efficiency and probability of success of their RCTs through the use of AI and external sources of data. One such source of external evidence is real-world data (RWD), encompassing health-related information collected from routine clinical practice [3]. Nonetheless, when used to assess the effectiveness of a drug versus a comparator, RWD are limited by the absence of randomization, which can introduce bias into naïve head-to-head comparisons since the two groups may represent different patient populations.
Overcoming this challenge has been the subject of a large body of literature [4]. This line of work has recently culminated in the target trial emulation framework [5]: a comprehensive playbook on the systematic use of observational data for treatment effect estimation. This approach necessitates the detailed specification of a hypothetical trial protocol including eligibility criteria, treatment strategies, and outcome measures, which is then emulated by utilizing observational data [6,7]. Examples of this work have been applied in screening colonoscopy in Medicare [8], and in patients with type 2 diabetes, where a propensity score-based target trial compared the impact of SGLT2 inhibitors versus other therapies [9].
1.2. Related Work
Within this framework, causal inference techniques are employed that often extend beyond the simple propensity score methods advocated in earlier approaches. For instance, G-computation is an approach relying on a model (often a machine learning one) trained to predict clinical outcomes for different observed treatments based on baseline patient characteristics [10]. This model is then employed to produce counterfactual predictions by systematically varying treatment assignments while maintaining constant patient characteristics, thus isolating the causal effect of therapeutic interventions and simulating randomization.
Its success notwithstanding, target trial emulation can only answer causal questions about treatments that are observed in the dataset. However, early clinical development often necessitates an ability to predict clinical outcomes on novel pharmaceuticals that are not yet available in RWD, or even where very little to no clinical testing has taken place. Such predictions are not intended to replace RCTs, but, if sufficiently accurate, they can effectively guide their design, ranging from high-level decisions such as go/no-go decisions, to granular decisions such as optimizing the choice of inclusion/exclusion criteria and comparator arm. As a result, this ability to not just emulate target trials between approved treatments but also to simulate future trials between an approved treatment and a novel mechanism of action (MoA) is one of the most valuable contributions AI can make in clinical drug development, and it is in line with a broader mandate for enhancing the role of AI and machine learning modelling in trial design and clinical development, under the more general rubric of model-informed drug development (MIDD). Regulatory bodies including the FDA recognize the potential of such methods to improve trial efficiency, increase probability of success (PoS), optimize drug dosing, better define potential comparators, subpopulations, and endpoints, and ultimately ensure that clinical trials become faster, more efficient, and more equitable.
To approach this problem of estimating the clinical efficacy of novel drugs that have not yet been tested in humans, the data science community has looked beyond RWD, into different sources of bioinformatic and molecular biology data. Already to date, bioinformatics databases have been instrumental in elucidating treatment mechanisms and identifying novel indications for existing drugs [11]. Often, bioinformatics datasets are organized into knowledge graphs, which are databases systematically capturing relationships between heterogeneous biological entities such as proteins, targets, pathways, diseases, and symptoms [12]. Molecular knowledge graphs have increasingly been leveraged in drug discovery in other use cases as well, such as target identification [13,14].
Recent advancements in machine learning have facilitated the creation of innovative molecular representations of drugs, such as knowledge graph embeddings [15], which convert drug or target names into numerical representations for easier mathematical manipulation (e.g., to assess similarity or overlap) [16,17].
1.3. Case Studies
1.3.1. Multiple Myeloma
Multiple myeloma is a hematological malignancy characterized by the clonal proliferation of plasma cells within the bone marrow. Despite notable therapeutic progress, multiple myeloma remains incurable, with most patients experiencing recurrent relapses requiring multiple lines of therapy [18]. The wealth of available and emerging MoAs, alongside extensive RWD generated from widespread clinical use and registry studies, positions MM as an exemplary setting for developing and validating innovative approaches for predictive modeling of treatment effects.
1.3.2. Asthma
Asthma is a prevalent chronic respiratory condition characterized by airway inflammation and hyperresponsiveness, affecting over a quarter of a billion people worldwide [19]. Despite the availability of effective therapies, including inhaled corticosteroids (ICS), long-acting beta-agonists (LABAs), and biologic agents, between one-third and one-half of patients continue to experience symptoms [19]. The prevalence of asthma and available treatments, in contrast to MM, provides a complementary setting for validating the approaches previously described.
1.4. Contribution of This Paper
This paper presents an integrated framework leveraging patient-level RWD and drug embeddings designed to predict the performance of new treatments prior to clinical testing (Figure 1). The proposed framework forecasts the effect of novel MoAs on a specific clinical endpoint that is observed in RWD, albeit only for approved treatments. The fundamental insight driving this approach is that numerical drug embeddings enable extrapolation of efficacy from approved to novel drugs, by leveraging overlaps between different drugs in terms of the molecular pathways involved in their mechanisms of action.
Figure 1.
Overall approach from treatments to predicted outcomes. The overall approach involves representing both existing and novel Mechanisms of Action (MoAs, i.e., drugs or classes of drugs) in an embedding space of fixed length. Several embedding methods are available; in this instance, each dimension of that embedding represents one human biological pathway. The ability to represent both approved and novel MoAs in the same numerical space then makes it possible to utilize causal predictive modelling, trained on a patient-level dataset enriched with embedding information, to extrapolate clinical efficacy from the former to the latter. EP: Endpoint; MoAs: Mechanisms of Action; Pw: Pathways.
As explained above, both key components to our approach, drug embeddings and causal inference, have been widely studied in the literature. However, to the best of our knowledge, this is the first approach that combines the two approaches to perform counterfactual predictions of the impact of a novel MoA on a clinical endpoint for a given patient, or on average for a given cohort, by enriching patient-level RWD with drug embeddings. This approach integrates the breadth of and generalizability of category embeddings, with the robustness and theoretical validity of causal inference methods. We discuss how the resulting predictions can inform several key trial design decisions, such as the choice of comparator, endpoint, target product profile, and go/no-go decisions at various development stages. The proposed significance of this work is summarized in Table 1.
Table 1.
Statement of Significance.
2. Materials and Methods
This section describes the methods used in this study. We first present the real-world datasets used for the multiple myeloma and asthma case studies, followed by the target trial protocols and cohort definitions applied to each dataset. We then describe the construction and validation of knowledge graph-derived treatment embeddings, the G-computation framework used for treatment effect estimation, and the evaluation strategy used to assess predictive performance and replicate historical clinical trial findings.
2.1. Real-World Datasets Used
2.1.1. COTA Dataset (Multiple Myeloma)
COTA is an electronic health records dataset that includes anonymized patient-level data from a combination of academic and community healthcare networks in the United States. The dataset used here spans from 2015 to May 2023 and contains detailed information on 10,776 patients with an active multiple myeloma diagnosis filtered from a broader network of over 2 million patients (Figure 2). The dataset encompasses a wide range of patient information, including demographics, medical history, diagnoses, medications, lab test results, pathology reports, and disease outcomes such as mortality and transplants.
Figure 2.
Steps to design final multiple myeloma cohort from broader COTA dataset. RRMM: relapsed and refractory multiple myeloma, LoT: Line of Treatment. ORR: Overall Response Rate.
2.1.2. Merative MarketScan Mortality Edition (Asthma)
Merative MarketScan Mortality Edition is a closed claims dataset that includes anonymized patient-level data from commercially insured and Medicare populations in the United States. The dataset includes information on patient demographics, medical history, diagnoses, prescriptions, and outcomes for over 47 million patients between 2016 and 2023.
2.2. Target Trial Protocol and Cohort Definition
Here, we describe how the cohorts were defined for multiple myeloma and asthma, respectively.
2.2.1. Target Trial and Cohort Definition for Multiple Myeloma
Patients were identified between 1988 and 2023 and were followed up to 319 months/until loss to follow-up, for an average of 29 months. To emulate the target trial eligibility criteria, patients had to have been exposed to three lines of therapy (LoTs) with a diagnosis of relapsed and refractory multiple myeloma (RRMM) at least once before starting the fourth LoT. Relapse refractory was defined as starting at least once a new LoT less than 60 days after the end of the previous LoT. Patients were eligible for cohort entry at each new LoT initiation after RRMM diagnosis, meaning a patient could enter the cohort multiple times as a new individual, to maximize the data to train the model.
LoTs were defined based on a COTA multiple myeloma LoT algorithm accounting for response to treatment, typical cycle duration, renewable prescriptions, and type of therapies, as per International Myeloma Working Group (IMWG) guidelines [20]. A LoT started when patient either (1) started a new regimen (excluding maintenance, conditioning, and consolidating regimens) more than 30 days after the start of a prior regimen or after its end, or (2) started a similar regimen more than 1 year after the end of a prior one or more than 100 days after stem cell transplant, or (3) started any salvage therapy. The LoT ended when there was discontinuation for progression or inadequate response or the start of a new LoT.
Patients were also required to be older than 18 years at diagnosis, to have at least two prior LoTs (i.e., starting a third LoT), to have been treated by one of the retained treatments (i.e., we retained only treatments with more than 50 patients to enable learnability for the model) (Figure 3), and to have an overall response rate (ORR) endpoint available.
Figure 3.
Breakdown of drugs within the COTA dataset after applying our target trial study design.
We selected ORR as the endpoint for this study based on its historical use as primary and secondary endpoint in external control arms (ECAs) and RCTs. ORR was defined as the percentage of patients with a response (at least partial) within one year of the start date of the LoT of interest. Response was defined in COTA according to IMWG criteria.
We extracted the full list of individual drugs available in the COTA dataset and labelled them in drug classes based on the classification available at RxNorm [21]. Drug categories were as follows: Proteasome inhibitors: bortezomib, carfilzomib, ixazomib; Thalidomide analog: thalidomide, lenalidomide, pomalidomide; CD38-directed cytolytic antibodies: daratumumab, isatuximab; Alkylating drugs: melphalan, cyclophosphamide, bendamustine; SLAMF7-directed immunostimulatory antibody: elotuzumab; anti-BCMA antibody: belantamab mafodotin, teclistamab; Anthracycline topoisomerase inhibitors: doxorubicin, daunorubicin; Vinca alkaloids: vincristine, vinblastine; Platinum-based drugs: cisplatin, carboplatin; Topoisomerase inhibitors: etoposide, irinotecan.
2.2.2. Target Trial Protocol and Cohort Definition for Asthma
A cohort of patients with moderate-to-severe asthma was designed to emulate target trial eligibility criteria (Figure 4). Adult patients with a documented diagnosis of asthma initiating GINA 4 or GINA 5 therapies were identified for inclusion in the cohort. Patients were required to have had moderate-to-severe asthma for at least twelve months, defined as receipt of medium to high doses of inhaled corticosteroids prior to the twelve-month baseline period and at least one exacerbation during the baseline period. Patients were also required to be enrolled continuously at baseline.
Figure 4.
Steps to design final asthma cohort from broader MarketScan dataset.
Treatment blocks were created for long-term maintenance treatments such as biologics and ICS, LABA, and LAMA (Figure 5). The date of the first prescription in a treatment block was set as the index date. Days’ supply and refill number from the prescription record were used to determine dosage and length of the treatment blocks, with allowable gaps of 30 days for all treatments except 60 days for benralizumab (biologics). Moderate and high doses of ICS were identified based on the GINA 2025 guidelines [19]. Treatment blocks shorter than ninety days were excluded to ensure adequate exposure. Similar to the multiple myeloma cohort, patients were eligible for cohort entry for each subsequent treatment block, meaning a patient could enter the cohort multiple times as a new individual, to maximize the data to train the model.
Figure 5.
Breakdown of drugs within the MarketScan dataset after applying our target trial study design.
We extracted the full list of individual drugs available in the MarketScan dataset and labelled them in drug classes based on the classification of the 2022 GINA guidelines [22]. Drug categories were as follows:
- Inhaled corticosteroids (ICSs);
- Long-acting beta-adrenoceptor agonists (LABAs);
- Long-acting muscarinic antagonists (LAMAs);
- Leukotriene receptor antagonists (LTRAs);
- Systemic corticosteroids (SCSs);
- Biologics;
- Theophylline.
2.3. Knowledge Graph Representations
2.3.1. Knowledge Graph Construction and Embedding Provenance
The treatment embeddings were derived from a heterogeneous biological knowledge graph constructed primarily using data sourced from the GeneCards suite [23]. The graph schema comprised three core node types: drugs, genes, and biological pathways (including SuperPaths derived from PathCards). These entities were connected by specific biological relationships, specifically drug–gene interactions (e.g., binds/targets) and gene–pathway associations (e.g., participates in).
To build the drug embeddings from the knowledge graph we used the following four-step method:
- Generation of high-dimensional embedding;
- Dimensionality reduction;
- Aggregating embeddings to treatment-combo level;
- Validation of the embeddings and of the new MoA.
2.3.2. Generation of High-Dimension Embedding
Embeddings were derived from the GeneCards knowledge graphs using the degree-weighted pathway count (DWPC) algorithm as the embedding method [24]. The underlying graph schema utilized three core node types—drugs, genes, and biological pathways—connected by drug–gene binding interactions and gene–pathway participation associations to define the structural metapaths. The output was a tabular matrix where each row corresponded to an MoA and each column to a biological pathway. The value in each cell represented the DWPC score linking the drug associated with the MoA to the given pathway. The total number of pathways ranged from a few hundred to several thousand.
The steps required to generate the DWPC score for a drug and pathway are outlined as follows:
- Calculate the number of paths (paths) between the drug and pathway that follow a given metapath;
- The degree (d) of the nodes (Nodesp) in these paths (paths) is obtained by counting the number of connections it has with other nodes.
The DWPC score is then calculated as follows:
w represents the dampening factor, which determines how much the degree affects the weight of a given node.
The DWPC score between a drug and its pathway is a measure of the strength of its relationship based on the number of paths between them. The paths are down-weighted proportionally to their connectivity in the graph. Descriptive statistics were computed for each drug present in the RWE dataset that had available nodes/edges in the knowledge graph and sufficient information to derive a reliable embedding.
2.3.3. Dimensionality Reduction
Given that the number of embedding dimensions is expected to exceed 1000, dimensionality reduction is essential in order to incorporate embeddings into the outcome model in a way that remains tractable for a cohort of this sample size, and to mitigate excessive sparsity of the embeddings, which may impede the model’s ability to learn meaningful patterns and adversely affect the predictive power of other important features, including patient covariates and prognostic factors. Therefore, a balance must be achieved between reducing dimensionality and retaining as much information as possible through reduced embeddings.
The initial approach involved filtering pathways based on descriptive statistics, namely, a lower threshold on standard deviation to ensure sufficient variability of each retained dimension, and excluding pathways with insufficient treatment representation as determined by percentile thresholds. Subsequently, dimensionality reduction techniques were applied. Various models were evaluated, including principal component analysis (PCA), singular value decomposition (SVD), and uniform manifold approximation and projection (UMAP), and selected using the steps outlined in the validation section, below. This process reduced the embedding dimension from an initial p = 4327 to p′ = 5.
The validation section provides a detailed discussion on the selection of the optimal approach and tuning of hyperparameters.
2.3.4. Embedding Aggregation to Treatment-Combination Level
The final step before integrating these embeddings into the RWD was to decide on the appropriate level of granularity of representing different treatments. Treatments were initially aggregated at the treatment class level. Subsequently, individual treatments were grouped together into combinations of treatments, allowing a one-to-one correspondence with combination treatments observed in the RWD. With regard to newly identified MoAs, their potential combinations with other drugs were also incorporated to enable the use of these embeddings in the counterfactual analysis step of the G-computation.
2.3.5. Model Selection and Validation of Drug Embeddings
The embedding generation process described above is fundamentally an unsupervised method, making it non-trivial to define a performance metric that can be used for model selection (e.g., selecting the dimensionality reduction algorithm) or hyperparameter tuning (e.g., the number of dimensions). One of the strengths of our overall approach is that the embeddings are in fact later employed in a supervised learning step, which can serve as an anchor for model selection and fine-tuning. Nonetheless, optimizing the entire system based on the predictive performance of the supervised model would result in a very high-dimensional optimization problem, and would detract from the modularity of the approach. We hence opted to introduce a standalone validation step for the embeddings, in two steps, as follows:
- Expert validation: Two medical doctors with clinical oncology experience recorded which pairs of treatments were expected to lie close to each other based on their known mechanisms of action and target profiles. Moreover, treatments from the same class were also recorded as expected neighbors. A 2-dimensional representation of the drugs created via UMAP was then used to assess whether these similarities were preserved by the embedding.
- Scoring metrics: We additionally considered the following scores:
- ○
- Adjusted mutual information (AMI) [25]. To quantify the alignment between clinically defined treatment classes and clusters obtained in the reduced-dimensional space. This metric evaluates the correspondence between two different ways of grouping the same set of elements—in this case, treatments.
- ○
- Metrics to evaluate how the distances are preserved post-dimension reduction
- ▪
- Spearman correlation between two sets of distances: similarity of distances in the full dimensional space and in reduced-dimensional space.
- ○
- Trustworthiness [26]. To evaluate how well k-nearest neighbors are preserved after transformation.
Another metric considered was mean squared error (MSE), which evaluates how well the embedding preserves the original information by measuring the average squared difference between the original input and its reconstruction. However, MSE was ultimately disregarded due to its extremely high correlation with Spearman correlation (98% and 99% for MM and asthma, respectively), rendering it redundant in the evaluation process.
Given that the basis for extrapolation of clinical outcomes to novel MoAs is their pathway overlap with one or more approved drugs, for predictions to be reliable, the new MoA should not constitute an outlier in the space of existing MoAs. To evaluate this, we computed the average pairwise distance between each treatment and all others, then compare how the new MoA performs on this metric. If its average pairwise distance is comparable to those of existing treatments, we can be more confident that the outcome model’s predictions for the new MoA are reliable.
2.4. Treatment Effect Estimation with G-Computation
2.4.1. Overview of Different Methods to Estimate the Treatment Effect in RWD
Estimating treatment effects in observational studies requires sophisticated methodological approaches to address confounding and selection bias. The selection of an appropriate method depends on multiple factors, including data characteristics, sample size constraints, and inferential objectives.
The primary approaches include inverse-propensity-to-treat weighting (IPTW) in which the propensity score is used to reweight the existing observed data to create a pseudo-population [27], effectively balancing confounding variables across the treatment groups while maintaining the original sample size. Propensity score matching takes a different approach by identifying pairs of patients with outcomes in the matched pairs [28].
Doubly robust estimation represents a hybrid approach that combines both propensity score models and outcome regression models [29]. The method involves fitting canonical link generalized linear models via IPTW maximum likelihood estimation followed by standardization for average treatment effect across treated and control groups. These approaches are becoming increasingly popular with implementations such as augmented inverse probability weighting (AIPW) or targeted maximum likelihood estimation (TMLE).
G-computation [30] demonstrates superior scalability to large datasets and complex confounding structures through its model-based prediction framework [31,32]. The method scales effectively because it models the outcome process directly, avoiding implementational complexity with propensity score estimation in settings with a large number of competing treatments, as is typical in real-world data. It is also the only approach that allows simulation of novel mechanisms of action (MoAs) via embeddings. As a result, we decided to leverage this framework.
2.4.2. Training of an Outcome Model Leveraging Embedding as Treatment Representation
G-computation employs a direct outcome modeling approach that estimates the conditional expectation of the response variable given treatment assignment and measured confounders [33]. It involves specifying generalized linear models [34] or more flexible machine learning approaches that capture the relationship between patient characteristics, treatment exposure, and outcomes [35,36]. The outcome model serves as the foundation for treatment effect estimation by explicitly modeling the data-generating process rather than treatment assignment mechanism. The approach requires careful specification of covariates, interaction terms to capture the underlying clinical relationships, with feature engineering and regularization terms ensuring robust fit to observed RWD [37].
Any regression model can be employed to accommodate diverse endpoint types, including continuous, binary, and time-to-event outcomes, the latter often involving censoring. Standard regression performance metrics such as R2, area under the curve (AUC), and concordance index (C-index), or other contextually relevant metrics, are used to evaluate model performance. For time-to-event endpoints, survival models were used instead with C-index being a more suitable performance metric. The modeling pipeline also incorporates several optimization steps, including imputation of missing data, feature selection, iteration over different model classes with hyperparameter tuning, and stability assessments of model performance. Leveraging survival models also allows us to account for censoring in the data—a common challenge with RWE data—as these models can natively handle censored observations and learn from this information.
2.4.3. Counterfactual Leveraging Through Embeddings and Novel MoAs
In our approach, during counterfactual inference, G-computation leverages learned embeddings of treatment mechanisms to predict outcomes, enabling extrapolation to novel MoAs. This method creates counterfactual predictions by systematically varying treatment assignments while holding patient characteristics constant, thereby isolating the causal effect of therapeutic intervention. It is worth noting that propensity-based methods are not as amenable to this approach, because they require a categorical representation of drugs, whereas G-computation simply requires the user to specify the counterfactuals regardless of whether drugs are represented as categories, or as numerical vectors (as in the case of embeddings). This represents a fundamental departure from traditional approaches which inherently assume treatments are independent and provide no mechanism for knowledge transfer between related therapeutic interventions. In contrast, the transition from categorical treatment variables to dense embedding representations fundamentally reframes treatment effect estimation as a machine learning problem with enhanced generalization capabilities.
2.4.4. Treatment Effect Estimation
The generation of counterfactuals enables the creation of synthetic patient groups with identical characteristics and distributions, along with their respective responses to each treatment. Classical treatment effect estimation methods can then be applied, such as difference in means, t-testing, or Cox proportional hazard models for time-to-event endpoints, yielding hazard ratio estimates. This also makes it possible to estimate the impact of different inclusion/exclusion criteria or choice of comparators on the average treatment effect, by appropriately filtering the cohort or selecting the counterfactuals which match the scientific question. Confidence intervals for the treatment effect can be natively extracted using methods such as t-testing or the Cox model (via the p-value of the treatment coefficient).
2.5. Feature Engineering and Feature Selection
All features used in the model are defined prior to the index date, i.e., before treatment initiation, to avoid data leakage and challenges related to time-varying confounders.
All features are either numerical or binary (Boolean); categorical features are encoded using either one-hot encoding or ordinal encoding, depending on the clinical nature of the feature.
Multiple windowed features were created capturing counts, means, and sums over the preceding 1, 3, 6, and 12 months, with the exact window definitions determined by clinical relevance.
A missing value imputation step was performed after the train/test split, with the specific approach also determined by clinical context. Missing values were imputed as 0 to represent the absence of an event (e.g., a diagnosis or procedure not recorded), or using MICE (multiple imputation by chained equations). More conventional approaches, such as mean and median imputation, were also tested for comparison.
The list of features (confounders) used in the model follows a two-step process. First, features are selected based on data availability and clinical expertise. They are then filtered using statistical tests and data quality checks, including a minimum variance threshold, a missing value threshold, and autocorrelation analysis to remove highly correlated features.
2.5.1. Target Causal Estimand
To formalize our target trial emulation, we defined an explicit causal estimand for both cohorts. Our target causal parameter is the average treatment effect (ATE), representing the observational analogue of the intention-to-treat effect. For the multiple myeloma cohort, the target population consisted of patients with RRMM initiating a fourth line of therapy. The treatment strategies compared were the initiation of the respective treatment arms. The outcome was the ORR, evaluated based on treatment initiation. Similarly, for the asthma cohort, the target population included adult patients with moderate-to-severe asthma initiating GINA 4 or 5 therapies, with the outcome evaluated based on the index date of the initial prescription block. The outcome was the time to first exacerbation.
2.5.2. Assumptions for Emulating Randomized Clinical Trials with RWE Data
Emulating a randomized trial from observational data relies on several key causal assumptions, which we address below for both the MM and asthma settings. To support the no-unmeasured-confounding assumption, we followed the Target Trial Emulation (TTE) framework, aligning inclusion/exclusion criteria and endpoint definitions as closely as possible with a reference trial for each indication. Confounders were selected based on clinical expertise and broad data availability, capturing relevant prognostic and treatment-selection factors prior to the index date. Positivity was assessed both globally and between pairs of treatments of interest, using standardized mean difference (SMD) to identify and address regions of poor overlap. Consistency was supported by clearly specifying treatment strategies and aligning them with the target trial protocol to minimize ambiguity in treatment exposure definitions. Correct model specification was addressed through a comprehensive evaluation protocol (next section) assessing model calibration, and robustness. E-values were also calculated to quantify the potential impact of unmeasured (hidden) confounders.
2.5.3. Evaluation Protocol
Machine learning models were trained in enriched datasets combining patient baseline characteristics with treatment embedding vectors. To assess the model’s ability to generalize to unseen treatment combinations, we applied two approaches.
We used a holdout strategy, where we excluded one treatment at a time from training and evaluated its predictive performance on the holdout set. This can give us a sense of the ability of the model to generalize to a treatment not observed during model training.
Key Steps:
- Data Preparation: Split into model data (excluding the holdout treatment combo) and a holdout dataset (patients receiving the excluded treatment combo). The holdout dataset is never used to train or assess the performance of the model.
- Model Training and Fine Tuning: Usual train pipeline using the model data (we do 80/20 train/test split multiple times to select the best model) and assess that the model is properly trained, identifying any over/under fitting.
- Prediction and Assessment: Compare model performance on the training and test datasets using task-appropriate metrics (e.g., AUC, R2, or C-index).
- Weighted Aggregation: The process is repeated holding out different treatment combinations, with final performance calculated as a weighted average based on population sizes.
A second important approach was to try replicating the results of known clinical trials by holding some treatments published in recent RCTs. This step is required as predictive performance considers only the ability of the model to recover observed outcomes (though in unseen patients), whereas causal inference additionally requires counterfactual accuracy. For this process, as mentioned above, when replicating a trial, all samples containing the treatment drug were completely removed from the RWE data to emulate a novel MoA. As a result, the training dataset was slightly different for each trial being emulated, and these samples were not used for any model training, feature selection, feature engineering.
2.5.4. Gathering of Prior Historical Trials for Validation
For historical trials to be comparable to our predictions, we needed to identify trials that had at least two treatment arms and where both high-level cohort definition criteria and endpoint matched the target trial and modelling approach. For multiple myeloma, this required trials where ORR was a characterized endpoint with a quantitative readout, where patients were at the 2nd to 4th LoT, and which compared drugs/drug combinations present in our dataset. We identified four trials (Table 2), with a broad range of drug MoAs, with antibodies (SLAMF7 targeted immunostimulatory, anti-CD38) and small molecules (BCL-2 inhibitor) vs. typical comparators (proteasome inhibitors and thalidomide analogs). For asthma, this involved selecting trials that focused on patients with moderate or severe asthma, evaluated biologics vs. standard of care (ICS combination therapies), and included the time to first exacerbation endpoint. Three trials were identified for inclusion (Table 3).
Table 2.
Trial emulation for multiple myeloma drug embeddings. Each historical comparator trial is reported with its clinical trial identifier (NCT) and characteristics are reported from clinicaltrials.gov. TA: thalidomide analog, PI: proteasome inhibitor, NCT: National Clinical Trial number, LOT: Line of Therapy, ORR: overall response rate, CI: confidence interval. ECOG: Eastern Cooperative Oncology Group, describes patient’s ability to self-care (0 fully active and independent, 1, restricted physically, 2, unable to carry work activities, 3, confined to bed, 4, completely disabled, 5 dead).
Table 3.
Trial emulation for asthma drug embeddings. Each historical comparator trial is reported with its clinical trial identifier (NCT) and characteristics are reported from clinicaltrial.gov. NCT: National Clinical Trial number, TTE HR: time to event hazard ratio, CI: confidence interval.
3. Results
3.1. KG Embedding and Embedding Validations
3.1.1. Multiple Myeloma
To validate knowledge graph embeddings, we first compared how well given drug classes clustered together in the full embedding space. The overall adjusted mutual information in the 1626 embedding dimensions (post filtering for embedding axes with no variation), for all drug classes, was 0.249. Because AMI is adjusted for chance (where a score of 0.0 indicates completely random clustering and 1.0 indicates perfect agreement), the observed positive values confirm that the high-dimensional embeddings captured non-random, pharmacologically meaningful structure. While absolute AMI scores are inherently bounded by the biological pleiotropy of drugs spanning multiple pathways, this metric served as a valuable relative diagnostic to compare and select the optimal dimensionality reduction technique [38,39]. A large proportion of the drugs in our datasets are legacy chemical compounds which have broad MoAs (e.g., thalidomide analogs, cyclosporine) that touch upon many different pathways; they may hence be less well captured by pathway-driven embeddings, which further explains the significant yet modest absolute AMI (Table S1). We also visually inspected the overall distribution of drug classes in a UMAP-generated 2D space representative of the full space. We reassuringly observed that several important multiple myeloma drugs co-clustered within their classes, such as anti-BCMA and CD-38 antibodies, and vinca alkaloids (Figure 6).
Figure 6.
Overall 2D UMAP representation of the drug-embedding space for each drug and their respective classes, from original full space 1626 dimensions for 44 drugs. Each combination of color–shape describes one drug class.
To further validate embeddings and their ability to be used in a modelling context, we then sought to reduce their dimensions to be used as features in our causal model and tested three different dimension reduction methods. For this dataset, we observed that PCA, compared with UMAP and truncated SVD, best captured the full dimension in the reduced 5D space (Table 4) across all informative metrics: AMI, Spearman correlation and trustworthiness. The high Spearman correlation and trustworthiness suggest that despite projecting more than 1600 dimensions onto 5 dimensions, the core signal was preserved. We therefore retained the top 5 PCA components (explaining 72% of the variance) and used them for downstream model-training tasks.
Table 4.
Multiple Myeloma Embedding Validation Results. AMI: adjusted mutual information, PCA 5D: principal component analysis focused on the first 5 components, SVD: singular value decomposition, UMAP: uniform manifold approximation and projection. See Methods.
The predictive performance of the outcome model, described in the following section, also serves as a final statistical validation of the embeddings.
3.1.2. Asthma
The overall adjusted mutual information in the 5142 embedding dimensions for all classes was 0.39, confirming that the embeddings were meaningful. The embeddings were also visually inspected within a 2D space, as shown in Figure 7. As seen in the figure, biologic therapies clustered together as did the other major therapy classes (LABA, LAMA, steroids). A full listing of individual drugs is presented in Table S2.
Figure 7.
Overall 2D UMAP representation of the asthma drug embedding space for each drug and their respective classes, from original full space 5142 dimensions for 1789. Each combination of color–shape describes one drug class.
Of the three methods evaluated for dimensionality reduction, truncated SVD 4D was the best performing when considering all three metrics. While the UMAP approach had a higher AMI, a lower Spearman correlation suggests reduced similarity with the original embeddings (Table 5).
Table 5.
Asthma Embedding Validation Results. AMI: adjusted mutual information, PCA 4D: principal component analysis focused on the first 4 components, SVD: singular value decomposition, UMAP: uniform manifold approximation and projection. See Methods.
Finally, the validity of the embeddings was assessed through the outcome model itself. As explained in additional detail below, the model performed well with a C-index of 0.67 (for the training set) and 0.65 (for the test set).
For both MM and asthma, a sensitivity analysis was run to measure the impact of choice of the embedding on downstream applications, in this case, impact on the average treatment effect. These results are presented in the Section 3: Validation with historical trials.
3.2. Downstream Outcome Causal Modelling
3.2.1. Multiple Myeloma
To test the ability of the embeddings and the real-world data to predict trial outcomes, we trained a causally aware model leveraging G-computation. In the first instance, the outcome model can be assessed as any other machine learning model, based on its ability to accurately predict its target (clinical outcomes) on unseen data. Here, we observed that the outcome model demonstrated robust predictive performance across multiple validation approaches, achieving an AUC of 0.805 on the training dataset and 0.719 on the test dataset, suggesting a minor effect of overfitting and robust overall performance (Table 6). Moreover, the treatment embedding dimensions were significant contributors to the overall performance, as evidenced by their high feature importance as measured by SHAP values (Figure 8).
Table 6.
Model performance and key metrics. AUC: area under the curve. See Methods.
Figure 8.
SHAP feature importance plot for the clinical model, showing the mean absolute SHAP values for all included features.
Systematic evaluations of different model architectures revealed the superior performance of treatment embedding approaches compared to traditional categorical encoding methods. The model incorporating all covariates with treatment embeddings achieved a test AUC of 0.719, matching the performance of the categorical encoded treatment approach on approved drugs while providing enhanced interpretability, reduced dimensionality and extended generalizability. We additionally ran a hold-out test where each drug class was masked during training, and the model was asked to predict outcomes on that masked class on the test set. The overall AUC for the hold-out prediction across all drug classes was 0.685, which suggests an ability to predict patient outcomes with no existing information about the drug itself (Table 7). Notably, the baseline model using only clinical covariates without treatment information achieved a lower test AUC, demonstrating that treatment-specific information contributed meaningfully to predictive accuracy.
Table 7.
Model performance and key metrics using a hold validation approach to mimic an unseen treatment mechanism. AUC: area under the curve. See Methods.
The treatment embedding methodology showed value in capturing complex treatment interactions, with four embedding features ranking among the top 15 most influential predictors based on SHAP value analysis (Figure 9). These embedding dimensions effectively encoded treatment combination patterns that would be difficult to capture through traditional categorical encoding approaches.
Figure 9.
SHAP beeswarm summary plot displaying the distribution of SHAP values for all model features.
3.2.2. Asthma
We took a similar approach with asthma as with multiple myeloma, with a causally aware G-computation model. The outcome model achieved a C-index of 0.67 on the training dataset and 0.65 on the test dataset, suggesting good concordance (Table 8).
Table 8.
Asthma model performance and key metrics. C-index, see Methods.
Similar to the multiple myeloma use case, the model that included all covariates with treatment embeddings had a test C-index of 0.65, matching the performance of the categorical encoded treatment approach. The average C-index across three hold-out datasets was 0.64, suggesting an ability to predict outcomes without information on the drug itself (Table 9).
Table 9.
Asthma model performance and key metrics using a hold validation approach to mimic a new MoA. C-index, see Methods.
3.3. Validation of the Treatment Effect Estimation vs. Historical Trials
3.3.1. Multiple Myeloma
Focusing on available public trial data that could be compared to our results in the COTA dataset, we identified four trials that focused on relapsed refractory multiple myeloma for patients undergoing two to four lines of treatments with a set of drugs represented in our dataset and which had a two-arm design with a readout on the ORR (see Section 2).
We compared trials in terms of average treatment effect, reporting the difference in ORR for the first arm of the trial vs. the second arm. A positive value indicated that the ORR was higher for the first arm while a negative value indicated a lower ORR vs. the second arm.
Overall, we found that predicted outcomes systematically matched the overall direction of the outcome, i.e., which arm performed better overall at the endpoint (Figure 10). The predictive accuracy for SLAMF-7 was the highest, while it was within confidence intervals for both anti-CD38 trials. The predicted outcome for BCL-2 remained directionally correct but out of bounds of the actual outcome vs. the published results (Figure 10). This might be attributed to BCL-2 being underrepresented in terms of patient journeys to learn from and potentially misrepresented in our embeddings (small molecule and a single drug, vs. multiple within the class for most other drugs). This highlights a key boundary condition; accurate quantitative extrapolation requires the novel mechanism to share at least partial biological pathway overlap with established drugs in the training data. For isolated or truly first-in-class mechanisms, such as BCL-2 in our cohort, the model may lack the necessary embedding anchors to confidently predict effect magnitude.
Figure 10.
Trial prediction vs. historical outcomes.
Overall, comparison with historical trials suggests that embeddings predict overall directionality of outcomes, although they may still lack predictive power to enable prediction at a quantitative level.
For all predictions, we report the confidence interval of the outcome model for this particular outcome. For published RCT results, we report the confidence interval when reported by the trial. The drug class used for arm A is always reported first in the legend, followed by “vs” the second arm.
E-values were also calculated. For MM (binary endpoint), E-values were also calculated from counterfactual predictions expressed as probabilities (Table 10). Because the control-arm outcome averages approximately 0.5, the resulting risk ratio—and therefore, the E-value—is mathematically capped (RR ≤ 2, E-value ≤ ~3.4) regardless of the true strength of confounding. We therefore interpret our E-values relative to this context-specific ceiling rather than against conventional thresholds developed for rarer outcomes, where a comparatively modest E-value does not necessarily indicate weaker robustness to unmeasured confounding.
Table 10.
E-values for treatment comparisons in multiple myeloma.
As explained in Section 2, we also ran a set of analyses to validate that we could properly emulate a randomized controlled trial (RCT) using observational data. This included assessing positivity violations between treatment arms, even though, in these experiments, the intervention treatment was completely removed from the RWE dataset to simulate a novel MoA. Despite this, we still wanted to test whether the patient cohorts were comparable between arms, as well as against patients treated with drugs mechanistically similar to those in the intervention arm.
The SMD scores of the top confounders/prognostic factors of these models for MM are shown in Table S3.
Some SMD values are high (value greater than 1), reflecting genuine differences in patient characteristics between these treatment groups in the RWE data. Other comparisons show considerably lower SMDs, indicating better covariate overlap between those specific populations. Because the intervention drug was fully removed prior to training the outcome model, positivity was assessed only across the remaining treatment options, and in that context, we did not observe strong violations overall.
Another key consideration in this simulation is the choice of embedding method and its impact on ATE estimation (Figure S1). We conducted a sensitivity analysis across different dimensionality reduction methods (PCA, UMAP, and truncated SVD) and varying numbers of dimensions (from 4 to 8) and found that the sign of the ATE remained consistent, even though the magnitude varied. This is consistent with the objective of this approach for novel MoAs, to predict directionally whether a drug is likely to be effective.
3.3.2. Asthma
After identifying three trials for inclusion in the asthma analysis, we predicted the outcome of each trial (Figure 11).
Figure 11.
Asthma—trial prediction vs. historical outcomes. The left bars display the published hazard ratios for time to first exacerbation from each trial, while the right bars present the corresponding model-predicted values.
We found that the predicted outcomes systematically matched the direction of the outcome, similar to the multiple myeloma use case above. The accuracy was highest for benralizumab trial and lowest for omalizumab. We predicted omalizumab to be more effective in reducing time to first exacerbation than it was in the clinical trials. This difference in results may be due to residual confounding, as the observed baseline exacerbation rate in omalizumab patients was lowest across the biologics in the cohort, suggesting these patients were more adequately controlled in the real world. Similar to multiple myeloma, these results suggest that embeddings can be leveraged to predict the direction of trial outcomes. However, the results also suggest the importance of fit-for-purpose datasets with adequate measures on disease severity to address unmeasured confounding. While the multiple myeloma use case was conducted in an oncology-specific dataset, the asthma dataset lacked specific severity measures related to lung function that would allow prediction at the quantitative level.
We also ran the same sensitivity analysis and positivity violation checks for the asthma cohort. Results were consistent. No strong positivity violations were observed, and the choice of embedding did not affect the direction of the ATE.
4. Discussion
4.1. Interpretation of Results
We demonstrated the feasibility of leveraging RWD and pathway embeddings derived from knowledge graphs to predict clinical trial outcomes for investigational molecules with no prior clinical data. By integrating pathway activation signatures with causal modeling via G-computation, our framework achieved promising predictive performance and additionally demonstrated the ability to reproduce the direction of treatment effect estimates observed in historical trials, mitigating observational bias present in RWD even in this challenging context where the objective is to extrapolate beyond drugs whose clinical outcomes are observable.
Our findings suggest that embedding biological knowledge into RWD can provide an actionable layer of evidence in early drug development in two ways: (1) predicting the direction of trial outcomes and (2) informing trial design. The pathway-based representations allowed us to capture drug–disease interactions beyond direct clinical data, enabling predictions for molecules that are pre-market. Prediction of trial outcomes before first-in-human facilitates go/no-go decisions in early development pipelines, as well as informing the target product profile of an asset by way of identifying subpopulations that are most likely to benefit from a certain treatment, via investigating feature importance in the outcome model. The method is also applicable in settings where there is some limited clinical data on the novel drug (e.g., from a Phase II or even Phase I read-out), which can be seamlessly integrated into the approach given that the causal inference framework mitigates any selection bias, and the drug embeddings enable both direct learning from observed outcomes with the novel drug and indirect learning from extrapolation from other drugs, without any modifications to the approach needed. This can increase the quality of evidence available for pivotal go/no-go decisions, whereas an improved understanding of the comparative efficacy between a novel drug and different approved comparators can provide a sponsor with confidence to electively add a head-to-head comparison to a comparator in the trial, thus improving the quality and clarity of evidence generated for patients. Statistics on feature importance can also serve as the basis for inclusion/exclusion criteria optimization. Finally, the ability to run this pipeline for multiple different endpoints can help choose an appropriate endpoint in cases where there are multiple options with regulatory and clinical validation, or guide the choice of secondary or hierarchical endpoints when the choice of primary is fixed.
These questions fall squarely under the remit of MIDD, offering an integrated framework for disease, drug, and trial modelling that is both empirical and biologically motivated, as it can combine patient-level RWD and molecular information stored in knowledge graphs (as well as limited amounts of patient-level RCT data where available). This grounding in biological pathways is a key strength, addressing a common challenge in ‘explainable AI’ for regulatory submissions by providing a mechanistic, interpretable basis for its predictions. This may, in turn, enhance regulatory trust compared to a pure ‘black box’ prediction based on RWD alone. It therefore also addresses some of the other key challenges present in MIDD, such as lack of patient-level granularity and often weak connection to clinical outcomes. This work extends previous efforts in predictive modeling of trial outcomes, which often rely primarily on trial metadata, chemical structure, or limited clinical features, without explicitly representing the relationship between the clinical outcome and the patient characteristics, which lies at the core of the ability of this model to inform clinical trials in the ways mentioned above.
4.2. Limitations
Despite encouraging results, several limitations should be acknowledged. First, while our model consistently predicted the direction of outcomes, overall prediction of the endpoint difference was often out of range of observed data. This highlights the inherent complexity of trial-level effect estimation and suggests that while embeddings capture important signals, additional layers of information (e.g., pharmacokinetics, safety profiles, or trial execution variables) may be necessary to fully recapitulate trial outcomes. Second, our validation relied on a holdout strategy against historical trials. Although this provides a pragmatic benchmark, prospective validation against ongoing or future trials will be essential to establish generalizability and real-world utility. Relatedly, our masking experiments evaluate mechanisms that, while withheld from training, remain embedded near existing biological pathway neighborhoods; performance on truly first-in-class mechanisms with sparse or no representation in the underlying knowledge graph has not been directly tested and should be considered exploratory until validated against emerging clinical data. Third, while UMAP visualization and expert feedback were useful for embedding validation, more systematic evaluation against curated biological pathway knowledge, or preclinical results, would strengthen confidence in interpretability. Additionally, the two-dimensional UMAP representation of multidimensional embeddings introduces interpretative limitations for clinical teams, as complex relationships may be oversimplified in this format. Furthermore, embedding proximity reflects mechanistic resemblance but does not formally guarantee exchangeability in the causal inference sense; mechanistically similar treatments may still differ in patient selection or confounding structure, and this assumption should be evaluated further as additional data become available. Because the embedding features used in the predictive model are reduced latent components derived from the original pathway-level DWPC matrix, their SHAP importance reflects composite biological signaling rather than the contribution of individual pathways directly. Confidence in these features therefore relies on the combination of upstream biological validation of the embedding structure and downstream empirical evaluation in held-out treatments and historical trial emulations, while acknowledging the limitations discussed above regarding effect-size precision and extrapolation to truly novel mechanisms.
Fourth, more sensitivity analyses could be undertaken to identify unmeasured confounders with approaches such as, e.g., E-values, negative controls, and bias formulas, with more extensive checks on positivity violations and model misspecifications.
Additionally, there were limitations associated with the RWD used. Firstly, for multiple myeloma, the analysis was constrained by a restricted patient sample size due to incomplete ORR data availability, limiting the robustness of subgroup inference. Second, the LoT identification relied on a proxy algorithm rather than direct clinical documentation, and a thorough review of COTA methodology documentation is recommended for a full understanding of this approach. Third, our analytical definition of ORR may differ from those used in clinical trials, as standardized definitions are not consistently documented across sources. For the asthma use case, while detailed prescription-level data were available to appropriately calculate treatment blocks, measures of disease severity, specifically related to lung function, were lacking. To mitigate this, we utilized established claims-based proxy variables for asthma severity, such as baseline exacerbation history, systemic corticosteroid use, and maintenance therapy intensity, consistent with standard pharmacoepidemiologic practices for claims data. Nevertheless, because our G-computation framework relies on the assumption of conditional exchangeability, some residual unmeasured confounding inevitably remains across both use cases.
Finally, while our hold-out experiments (AUC 0.685 for multiple myeloma, C-index 0.65 for asthma) demonstrated general predictive stability for novel mechanisms, we observed that quantitative accuracy degraded for mechanisms that were highly distant in the embedding space. Future research must prioritize formal out-of-distribution (OOD) evaluation metrics, specifically, calibration, discrimination, and coverage probabilities conditioned on embedding distance, and move to a cross-fitting framework for the ATE estimation, to provide robust uncertainty quantification and foster clinical trust before such models can inform prospective trial design.
Future research should aim to expand the range of RWD sources incorporated into the modeling pipeline, including longitudinal patient-level data, multi-omics datasets, and global trial registries. Enhancing causal modeling frameworks to better address confounding and effect heterogeneity across populations may also improve predictive fidelity. Furthermore, while our methodology prioritizes clinical interpretability and utilizes external GeneCards embeddings to inject biological priors into classical G-computation, recent advancements in deep learning offer promising alternatives. Integrating our approach with state-of-the-art causal representation learning architectures (such as CEVAE or Dragonnet) or end-to-end graph neural networks could allow future models to simultaneously learn optimal treatment representations and causal outcome mappings. Comparing our baseline framework against these advanced graph-based treatment modeling approaches, particularly as larger RWD cohorts become available to support such data-intensive architectures, is a critical next step to further optimize predictive fidelity. Finally, embedding interpretability and uncertainty quantification should be prioritized to foster clinical trust.
5. Conclusions
In this work, we employed three complementary methodologies to validate our drug-pathway embeddings: expert validation through UMAP visualizations, statistical validation using metrics such as Spearman correlation, and machine learning validation through CatBoost hold-out prediction performance. Across all validation strategies, the results consistently demonstrated that the embeddings effectively captured the pharmacological meaning underlying drug–pathway interactions.
Building on these validated embeddings, we trained a G-computation framework that integrated real-world data with the learned representations, enabling the model to infer mechanisms of action that were not explicitly present in the training set.
Finally, validation of the end-to-end model confirmed its ability to generalize to unseen scenarios. When evaluated against historical randomized clinical trials excluded from both embedding and real-world datasets, the model consistently predicted the correct directionality of treatment effects. Specifically, it accurately identified which drug in each trial would achieve a higher or lower ORR. These findings collectively suggest that our approach provides an interpretable framework for predicting pharmacological mechanisms and treatment outcomes in the absence of direct experimental data that can be used to guide clinical trial design.
Supplementary Materials
The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/mca31050179/s1, Table S1: List of individual drugs for multiple myeloma and corresponding drugs in RxNorm; Table S2: List of individual drugs for asthma and corresponding drugs in RxNorm; Table S3: Standardized Mean Differences (SMDs) key baseline confounders across treatment comparisons; Figure S1: Sensitivity analysis of embedding model choice across multiple myeloma treatment comparisons.
Author Contributions
Conceptualization, Writing—original draft, Writing—review and editing: M.B. and C.H.; Conceptualization, Formal Analysis, Writing—original draft, Writing—review and editing: M.K. and P.-Y.M.; Investigation, Writing—review and editing: F.F.; Conceptualization, Formal Analysis, Writing—review and editing: J.B.; Formal Analysis, Writing—review and editing: I.S. and L.D.; Supervision, Writing—review and editing: T.D.; Supervision, Conceptualization, Methodology, Writing—original draft, Writing—review and editing: A.P.; Writing—review and editing: F.D.; Conceptualization, Methodology, Writing—review and editing: B.R.; Writing—review and editing: L.H.; Writing—review and editing, Project Administration: E.H.; Supervision, Conceptualization, Methodology, Writing—review and editing: R.H.V. and C.A. All authors have read and agreed to the published version of the manuscript.
Funding
This study was funded by Sanofi. Publication management support was provided by Mir-Masoud Pourrahmat of Evidinno Outcomes Research Inc. (Vancouver, BC, Canada), with funding from Sanofi. The sponsor participated in the conception and design of the analysis, as well as in the decision to submit the article for publication. Additionally, the sponsor had the opportunity to review the manuscript for medical and scientific accuracy and intellectual property considerations. The authors are responsible for all content and editorial decisions, and they received no honoraria related to the development of this article.
Data Availability Statement
The data supporting the findings of this study were obtained from COTA, Inc. (Boston, MA, USA) and Merative™ MarketScan® (Ann Arbor, MI, USA) under restricted data use agreements and cannot be publicly shared due to licensing and patient privacy restrictions. The knowledge graph embeddings were derived using GeneCards, which is publicly accessible for academic use but was utilized under a commercial license for this study; therefore, the resulting DWPC embeddings and trained model weights are subject to commercial licensing restrictions and are not available for distribution.
Conflicts of Interest
Margot Blanchon, Carrie Heller, Maksim Kriukov, Jonathan Broadbent, Francesca Frau, Flavio Dormont, Brandon Rufino, Lichen Hao, and Ramon Hernandez Vecino are employees and/or shareholders of Sanofi. Pierre-Yves Mousset, Ilaria Sartori, Lise Diagne, Thomas Devenyns, Alex Peluffo, Edouard Hatton, and Chris Anagnostopoulos report employment with QuantumBlack—AI by McKinsey.
Abbreviations
The following abbreviations are used in this manuscript:
| AIPW | Augmented inverse probability weighting |
| AI | Artificial intelligence |
| AMI | Adjusted mutual information |
| AUC | Area under the curve |
| BCMA | B-cell maturation antigen |
| BCL-2 | B-cell lymphoma 2 |
| C-index | Concordance index |
| CD38 | Cluster of differentiation 38 |
| CI | Confidence interval |
| COTA | Clinical Outcomes Tracking and Analysis dataset |
| DWPC | Degree-weighted pathway count |
| ECA | External control arm |
| ECOG | Eastern Cooperative Oncology Group |
| EP | Endpoint |
| FDA | Food and Drug Administration |
| GINA | Global Initiative for Asthma |
| ICS | Inhaled corticosteroids |
| IMWG | International Myeloma Working Group |
| IPTW | Inverse propensity-to-treat weighting |
| KG | Knowledge graph |
| LABA | Long-acting beta-agonist |
| LAMA | Long-acting muscarinic antagonists |
| LoT | Line of treatment |
| LTRA | Leukotriene receptor antagonist |
| MIDD | Model-informed drug development |
| MM | Multiple myeloma |
| MoA | Mechanism of action |
| NCT | National Clinical Trial identifier |
| ORR | Overall response rate |
| PCA | Principal component analysis |
| PoS | Probability of success |
| R2 | Coefficient of determination |
| RCT | Randomized controlled trial |
| RRMM | Relapsed/refractory multiple myeloma |
| RWD | Real-world data |
| SCS | Systemic corticosteroids |
| SHAP | Shapley additive explanations |
| SLAMF7 | Signaling lymphocytic activation molecule family member 7 |
| SVD | Singular value decomposition |
| TMLE | Targeted maximum likelihood estimation |
| TTE HR | Time-to-event hazard ratio |
| UMAP | Uniform manifold approximation and projection |
References
- Mulcahy, A.; Rennane, S.; Schwam, D.; Dickerson, R.; Baker, L.; Shetty, K. Use of clinical trial characteristics to estimate costs of new drug development. JAMA Netw. Open 2025, 8, e2453275. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fernald, K.D.S.; Förster, P.C.; Claassen, E.; van de Burgwal, L.H.M. The pharmaceutical productivity gap—Incremental decline in r&d efficiency despite transient improvements. Drug Discov. Today 2024, 29, 104160. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wilson, B.E.; Booth, C.M. Real-world data: Bridging the gap between clinical trials and practice. eClinicalMedicine 2024, 78, 102915. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rubin, D.B. Estimating causal effects of treatments in randomized and nonrandomized studies. J. Educ. Psychol. 1974, 66, 688–701. [Google Scholar] [CrossRef] [Scilit]
- Hernán, M.A.; Wang, W.; Leaf, D.E. Target trial emulation: A framework for causal inference from observational data. In Jama Guide to Statistics and Methods; Livingston, E.H., Lewis, R.J., Eds.; McGraw-Hill Education: New York, NY, USA, 2019. [Google Scholar]
- Dickerman, B.A.; García-Albéniz, X.; Logan, R.W.; Denaxas, S.; Hernán, M.A. Avoidable flaws in observational analyses: An application to statins and cancer. Nat. Med. 2019, 25, 1601–1606. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hernán, M.A.; Robins, J.M. Using big data to emulate a target trial when a randomized trial is not available. Am. J. Epidemiol. 2016, 183, 758–764. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- García-Albéniz, X.; Hsu, J.; Hernán, M.A. The value of explicitly emulating a target trial when using real world evidence: An application to colorectal cancer screening. Eur. J. Epidemiol. 2017, 32, 495–500. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xie, Y.; Bowe, B.; Gibson, A.K.; McGill, J.B.; Maddukuri, G.; Yan, Y.; Al-Aly, Z. Comparative effectiveness of sglt2 inhibitors, glp-1 receptor agonists, dpp-4 inhibitors, and sulfonylureas on risk of kidney outcomes: Emulation of a target trial using health care databases. Diabetes Care 2020, 43, 2859–2869. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shin, J.I.; Grams, M.E. Trial emulation methods. Am. J. Kidney Dis. Off. J. Natl. Kidney Found. 2024, 83, 264–267. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nakagawa, C.; Yokoyama, S.; Hosomi, K.; Takada, M. Repurposing haloperidol for the treatment of rheumatoid arthritis: An integrative approach using data mining techniques. Ther. Adv. Musculoskelet. Dis. 2021, 13, 1759720x211047057. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nicholson, D.N.; Greene, C.S. Constructing knowledge graphs and their biomedical applications. Comput. Struct. Biotechnol. J. 2020, 18, 1414–1428. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- James, T.; Hennig, H. Knowledge graphs and their applications in drug discovery. Methods Mol. Biol. 2024, 2716, 203–221. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nabirotchkin, S.; Peluffo, A.E.; Rinaudo, P.; Yu, J.; Hajj, R.; Cohen, D. Next-generation drug repurposing using human genetics and network biology. Curr. Opin. Pharmacol. 2020, 51, 78–92. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- An, X.; Chen, X.; Yi, D.; Li, H.; Guan, Y. Representation of molecules for drug response prediction. Brief. Bioinform. 2021, 23, bbad393. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Celebi, R.; Uyar, H.; Yasar, E.; Gumus, O.; Dikenelli, O.; Dumontier, M. Evaluation of knowledge graph embedding approaches for drug-drug interaction prediction in realistic settings. BMC Bioinform. 2019, 20, 726. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wen, J.; Zhang, X.; Rush, E.; Panickan, V.A.; Li, X.; Cai, T.; Zhou, D.; Ho, Y.L.; Costa, L.; Begoli, E.; et al. Multimodal representation learning for predicting molecule-disease relations. Bioinform 2023, 39, btad085. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Girvan, A.; Yu, J.; Kamalakar, R.; Mearns, E.S.; Nixon, M.; Cornell, R.F. Treatment patterns and outcomes in patients with relapsed/refractory multiple myeloma receiving ≥3 lines of therapy: A real-world evaluation in the united states. Blood 2022, 140, 5288–5290. [Google Scholar] [CrossRef] [Scilit]
- Global Initiative for Asthma. Global Strategy for Asthma Management and Prevention. 2025. Available online: www.ginasthma.org (accessed on 15 November 2025).
- Moreau, P.; Kumar, S.K.; San Miguel, J.; Davies, F.; Zamagni, E.; Bahlis, N.; Ludwig, H.; Mikhael, J.; Terpos, E.; Schjesvold, F.; et al. Treatment of relapsed and refractory multiple myeloma: Recommendations from the international myeloma working group. Lancet Oncol. 2021, 22, e105–e118. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nelson, S.J.; Zeng, K.; Kilbourne, J.; Powell, T.; Moore, R. Normalized names for clinical drugs: Rxnorm at 6 years. J. Am. Med. Inform. Assoc. 2011, 18, 441–448. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Global Initiative for Asthma. Global Strategy for Asthma Management and Prevention, 2022. Available online: https://ginasthma.org/wp-content/uploads/2022/05/GINA-2022-Main-Report-Tracked-20220503-WMS.pdf (accessed on 15 November 2025).
- Stelzer, G.; Rosen, N.; Plaschkes, I.; Zimmerman, S.; Twik, M.; Fishilevich, S.; Stein, T.I.; Nudel, R.; Lieder, I.; Mazor, Y.; et al. The genecards suite: From gene data mining to disease genome sequence analyses. Curr. Protoc. Bioinform. 2016, 54, 1.30.1–1.30.33. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Himmelstein, D.S.; Baranzini, S.E. Heterogeneous network edge prediction: A data integration approach to prioritize disease-associated genes. PLoS Comput. Biol. 2015, 11, e1004259. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Adjusted Mutual Information Between Two Clusterings. Available online: https://ogrisel.github.io/scikit-learn.org/sklearn-tutorial/modules/generated/sklearn.metrics.adjusted_mutual_info_score.html# (accessed on 1 October 2025).
- Trustworthiness. Available online: https://scikit-learn.org/stable/modules/generated/sklearn.manifold.trustworthiness.html (accessed on 1 October 2025).
- Austin, P.C. An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivar. Behav. Res. 2011, 46, 399–424. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rosenbaum, P.R.; Rubin, D.B. The central role of the propensity score in observational studies for causal effects. Biometrika 1983, 70, 41–55. [Google Scholar] [CrossRef] [Scilit]
- Funk, M.J.; Westreich, D.; Wiesen, C.; Stürmer, T.; Brookhart, M.A.; Davidian, M. Doubly robust estimation of causal effects. Am. J. Epidemiol. 2011, 173, 761–767. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Robins, J.M.; Hernán, M.A.; Brumback, B. Marginal structural models and causal inference in epidemiology. Epidemiology 2000, 11, 550–560. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Le Borgne, F.; Chatton, A.; Léger, M.; Lenain, R.; Foucher, Y. G-computation and machine learning for estimating the causal effects of binary exposure statuses on binary outcomes. Sci. Rep. 2021, 11, 1435. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Seaman, S.R.; Keogh, R.H.; Dukes, O.; Vansteelandt, S. Using generalized linear models to implement g-estimation for survival data with time-varying confounding. Stat. Med. 2021, 40, 3779–3790. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Loh, W.W.; Ren, D. A tutorial on causal inference in longitudinal data with time-varying confounding using g-estimation. Adv. Methods Pract. Psychol. Sci. 2023, 6, 25152459231174029. [Google Scholar] [CrossRef] [Scilit]
- Naimi, A.I.; Cole, S.R.; Kennedy, E.H. An introduction to g methods. Int. J. Epidemiol. 2017, 46, 756–762. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Brand, J.E.; Zhou, X.; Xie, Y. Recent developments in causal inference and machine learning. Annu. Rev. Sociol. 2023, 49, 81–110. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hahn, P.; Murray, J.; Carvalho, C. Bayesian regression tree models for causal inference: Regularization, confounding, and heterogeneous effects. Bayesian Anal. 2020, 15, 965–1056. [Google Scholar] [CrossRef] [Scilit]
- Xu, Y.; Kim, J.S.; Hummers, L.K.; Shah, A.A.; Zeger, S.L. Causal inference using multivariate generalized linear mixed-effects models. Biometrics 2024, 80, ujae100. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kim, V.; Adaloglou, N.; Osterland, M.; Morelli, F.M.; Halawa, M.; König, T.; Gnutt, D.; Marin Zapata, P.A. Self-supervision advances morphological profiling by unlocking powerful image representations. Sci. Rep. 2025, 15, 4876. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Connell, W.; Garcia, K.; Goodarzi, H.; Keiser, M.J. Learning chemical sensitivity reveals mechanisms of cellular response. Commun. Biol. 2024, 7, 1149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.










