1. Introduction
Vaccines are indispensable for preventing disease, averting deaths, and improving global health. The World Health Organization (WHO) estimates that vaccination has averted an average of 4 million deaths annually in recent years [
1]. Beyond reducing disease burden, rapid vaccine development and deployment played a critical role in restoring socioeconomic activity during the COVID-19 pandemic [
2].
Yet many infectious diseases still lack effective vaccine protection, and the threat posed by emerging pathogens persists. Vaccine research and development therefore remain essential to safeguarding population health [
3]. In public health emergencies, rapid vaccine development is especially critical, clinical trials often account for a substantial proportion of the pre-licensure timeline, with estimates exceeding 50% in some analyses [
4,
5]. Improving the understanding of clinical trial design features may contribute to characterizing current development patterns. Notably, many setbacks in vaccine development may be related not only to biological or technical hurdles, but also to the absence of robust reference metrics and quantitative frameworks. This gap may contribute to trial designs that are effectively ad hoc, leading to high costs, prolonged timelines, and unsustainable burdens for sponsors [
6,
7].
Data-driven reference metrics and quantitative frameworks may be useful for characterizing patterns in vaccine trial design and progression outcomes. Although the number of vaccine clinical trials has generally increased year-on-year, progression rates have not improved markedly [
8], highlighting substantial heterogeneity in clinical development pathways across diseases, phases, and development settings. In broader clinical development research, data-driven analytical approaches have been used to examine historical trial characteristics, progression patterns, and development variability across therapeutic areas [
9,
10,
11]. However, such approaches remain less commonly applied in vaccine trial research, which has focused primarily on immunogenicity prediction, antigen selection, and safety signal detection [
12]. Comparatively less attention has been given to large-scale characterization of trial feature structures and empirical progression patterns across heterogeneous vaccine development settings.
Existing studies are also frequently constrained by limited scope (e.g., single vaccine platforms or narrow therapeutic areas), relatively small datasets, or analytical approaches focused mainly on isolated associations rather than multidimensional feature structure [
5,
8,
13,
14]. As a result, current evidence remains fragmented and provides limited insight into how observable trial characteristics cluster and vary across historical vaccine development programs.
To address these gaps, we assembled a near-comprehensive dataset of vaccine clinical trials conducted over the past decade. This study provides three descriptive analytical components: (1) a structured characterization of trial design features and observed progression outcomes; (2) identification of recurrent design configurations within disease- and phase-specific subsets, using random forests as an exploratory pattern-recognition tool (not as a predictive engine); and (3) exploratory scenario-based comparisons of projected time and cost distributions under different historical configuration distributions. Conventional regression analyses were used to examine statistical associations between observable trial characteristics and progression outcomes, whereas machine learning methods were applied to summarize recurring high-dimensional configuration patterns within the observed dataset. The resulting analyses were retrospective and exploratory in nature and were intended to support structured characterization of historical trial design patterns rather than prospective optimization or prescriptive trial design guidance. Collectively, this work contributes to a data-enabled framework for examining variability and recurrent configuration structures in vaccine clinical trial development.
2. Methods
2.1. Study Design
We assembled a comprehensive registry-based dataset of vaccine clinical trials conducted over the past decade. Trial progression outcomes were first descriptively summarized across major trial design features, followed by univariable and multivariable analyses to examine statistical associations between trial characteristics and progression outcomes. Multivariable models were specified based on clinical and methodological relevance together with data availability considerations.
Machine learning analyses were conducted within disease- and phase-specific subsets to identify important trial design features associated with observed progression outcomes. A random forest model was used to estimate variable importance and to summarize key feature patterns related to higher progression probabilities.
All analyses were retrospective and exploratory. Robust-oriented configurations were then identified exploratorily based on consistency of important features and relatively stable predicted progression probabilities across disease and phase subsets, supported by sensitivity analyses using different minimum sample size thresholds.
Exploratory scenario-based comparisons were performed to examine differences in projected development time and cost between robust-oriented configurations and historical configuration patterns.
2.2. Data Sources and Extraction
As shown in
Figure 1, we identified vaccine clinical trials registered on ClinicalTrials.gov between 1 January 2012 and 31 December 2022 through a structured registry search using the term “vaccine” [
15]. Registry-derived data were supplemented with publicly available regulatory and development-status sources, including FDA records and related databases [
16,
17,
18,
19]. Variables were extracted from predefined registry fields, and records with missing or ambiguous key variables required for analysis were excluded according to prespecified criteria. After exclusion of incomplete records and non-human studies, 1618 clinical trials were included in the final dataset. Detailed inclusion and exclusion procedures are provided in the
Supplementary file (pp. 3–4). Because the dataset was based on registry-reported information, reporting delays and incomplete updates may remain.
Operational definitions of trial progression outcomes were prespecified and interpreted as proxy measures of observable development progression rather than definitive indicators of clinical or regulatory success. For phase I–II trials, progression was defined as documented advancement to a subsequent trial phase. Trials without evidence of progression were classified as non-progressed based on available registry records, although absence of follow-up may reflect delayed reporting or incomplete updates. For phase III trials, progression was defined as trial completion with subsequent regulatory authorization where identifiable from public records. Follow-up was extended through 1 June 2025 to reduce potential right-censoring.
To improve interpretability and reduce overcategorization, feature variables were harmonized using predefined grouping rules based on registry fields and methodological considerations [
20,
21]. Feature definitions and grouping methods were prespecified based on trial registry fields and relevant methodological considerations. Continuous variables (e.g., sample size) were categorized into quantile-based groups to ensure balanced sample sizes and improve interpretability, while categorical variables (e.g., trial phase, vaccine purpose, disease type, age group, study design, and funding source) were harmonized using standard classifications from ClinicalTrials.gov and established frameworks (e.g., ICD-10). Categories with small sample sizes were combined where appropriate to improve stability and interpretability. Detailed definitions and grouping procedures are provided in the
Supplementary file (pp. 5–12). Potential collinearity among variables was considered conceptually, and results were interpreted with caution where structural overlap may exist. We summarized trial characteristics and progression outcomes across 13 feature domains, including year, continent, and sample size.
2.3. Identification of Factors Associated with Trial Progress
Univariable logistic regression analyses were initially performed to examine associations between individual trial design characteristics and progression outcomes across major feature domains, including sample size, vaccine purpose, disease type, age group, study center structure, funding source, randomization, control, and blinding [
22]. Crude odds ratios (ORs) with 95% confidence intervals (CIs) were estimated for each variable.
Candidate variables for multivariable modeling were determined based on a combination of observed associations, methodological relevance, and data availability considerations. These variables were subsequently included in multivariable logistic regression models to estimate adjusted associations within the observed dataset. Adjusted ORs with 95% CIs were reported. For categorical variables, overall
p values were calculated using likelihood ratio tests (LRTs) to assess the overall association between each variable and progression outcomes. Reference categories were defined using common or representative groups to facilitate interpretation. Results are presented as odds ratios (ORs) with 95% confidence intervals (CIs), where ORs greater than 1 indicate positive associations with progression outcomes [
23].
All regression analyses were exploratory and descriptive in nature and were intended to characterize statistical associations within the historical dataset rather than to establish causation or independent predictive effects.
2.4. Machine Learning-Based Analysis of Trial Design Features and Robustness Assessment
Machine learning methods were applied to characterize recurrent patterns in trial feature configurations and to evaluate the stability of progression behavior across disease- and phase-specific subsets within the historical dataset. The framework aimed to support retrospective pattern characterization under the observed data distribution, not to generate prescriptive recommendations or prospective predictions [
24,
25].
A Random Forest classifier was chosen as the primary model due to its flexibility with mixed-type variables and its ability to model complex non-parametric relationships without explicit functional assumptions [
26]. The model consists of an ensemble of 1000 decision trees grown using bootstrap sampling, with final predictions averaged across all trees to improve the stability of cross-validated outputs [
27].
To mitigate overfitting and generate cross-validated estimates, we applied stratified 5-fold cross-validation [
28]. In each fold, the model was trained on a subset of the data and evaluated on the held-out fold, producing out-of-fold (OOF) estimates for all observations. These OOF estimates were then used for configuration-level aggregation and robustness-oriented characterization analyses.
Variable importance was quantified using the mean decrease in Gini impurity from the fitted Random Forest model. To reduce dimensionality and avoid excessive fragmentation of configuration groups, only the five most influential variables were retained for subsequent aggregation and robustness analyses. Trials were grouped by identical combinations of these predefined strategy variables. For each configuration group, we calculated the observed progression proportion and Wilson-type 95% confidence intervals to quantify uncertainty [
29]. Additionally, empirical Bayes shrinkage with a Beta(1,1) prior was applied to regularize observed proportions in sparse groups.
Configuration groups were ordered using a robustness-oriented framework primarily based on the lower bound of the Wilson type 95% confidence interval, together with empirical progression proportions and agreement between empirical and cross-validated model-derived estimates.
In the primary analysis, only groups containing at least 10 observations were summarized at the configuration level to reduce instability from sparse feature combinations. This restriction was applied only during post hoc aggregation and did not affect model fitting or OOF estimation.
Configuration groups that simultaneously showed higher empirical progression proportions, higher lower-confidence-bound estimates, and closer agreement between empirical and model-derived estimates were interpreted as robustness-oriented configuration patterns.
Sensitivity analyses were performed using alternative minimum sample-size thresholds of n = 5 and n = 15, in addition to the primary threshold of n = 10 (
Supplementary file pp. 14–15). Consistency of high-ranking configuration groups across thresholds was examined to evaluate the robustness of observed patterns under varying sparsity conditions.
2.5. Simulation-Based Characterization of Configuration-Level Modeled Distributions
Exploratory simulation analyses were conducted to examine how different historical configuration distributions were associated with projected trial counts, cumulative development time, and cumulative development cost under standardized hypothetical assumptions.
Two configuration-distribution scenarios were evaluated. The first was based on historically observed configuration frequencies within each disease–phase subgroup. The second used robustness-oriented configuration distributions identified through the machine learning aggregation framework. The comparison was intended to explore whether concentration toward configurations showing greater empirical consistency would yield different modeled resource use distributions under simplified assumptions.
Progression probabilities were modeled using Beta distributions parameterized from observed progression proportions and corresponding sample sizes. Trial progression was then simulated iteratively until progression to the subsequent development stage occurred, with the number of required trials generated using a geometric process framework.
Phase-specific development time and cost parameters were derived from previously published estimates of vaccine and oncology clinical development studies [
7,
30,
31,
32]. To improve comparability across studies conducted in different periods, monetary values were inflation-adjusted to 2025 US dollars using Consumer Price Index data from the US Bureau of Labor Statistics [
33]. Time and cost parameters were sampled from log-normal distributions to incorporate uncertainty in the simulation framework.
Simulated phase I–III results were subsequently aggregated to estimate projected cumulative trial counts, development time, and development cost across repeated iterations. The results are presented as summary distributions, including means, medians, and empirical uncertainty intervals.
The simulation framework generated standardized scenario-based projections of trial counts, development time, and development cost across historical and robustness-oriented configuration distributions.
All analyses were conducted in R (version 4.3.2).
3. Results
3.1. Overall Progression Rates of Vaccine Clinical Trials
Among 1618 vaccine clinical trials included in the analysis, 579 demonstrated documented progression to the subsequent development stage, corresponding to an overall observed progression rate of 35.80%, broadly consistent with previous reports [
8]. From 2012 to 2021, the annual number of registered vaccine trials generally increased, followed by a decline in 2022. Observed progression rates decreased between 2012 and 2017 and subsequently increased during 2018–2021, reaching the highest level in 2021
Supplementary file (p. 2).
Regionally, the Americas contributed the largest number of trials. Oceania showed the highest observed progression rate despite having the fewest trials, whereas Africa showed the lowest observed progression rate. These differences should be interpreted cautiously, as regional variation in trial composition, disease focus, and development context may contribute to the observed patterns. Detailed results across trial characteristics are presented in
Table 1.
3.2. Factors Associated with Vaccine Trial Progression
In univariable logistic regression analyses, larger sample size, preventive vaccine purpose, COVID-19 indication, enrollment across all age groups, industry sponsorship, multi-arm trial structures, and the presence of blinding were associated with higher observed odds of progression within the study dataset. Randomization scheme was not significantly associated with progression outcomes.
In multivariable logistic regression analyses, larger sample size, preventive vaccine purpose, COVID-19 indication, and enrollment across all age groups remained associated with higher observed odds of progression. In contrast, associations for study center structure, intervention model, and blinding were attenuated after adjustment and were no longer statistically significant. Funding source remained associated with progression outcomes overall, although estimates for network/collaborative alliance-sponsored trials were imprecise because of the small sample size in this subgroup.
These findings should be interpreted as descriptive associations within the historical registry dataset rather than evidence of causal or independently predictive relationships. Full regression results are presented in
Table 1.
3.3. Machine Learning-Based Analysis of Trial Design Features and Model-Based Comparisons
In the machine learning models, the variables with the top 5 highest importance included phase, disease type, age group, funding source, and purpose of vaccine (
Supplementary file p. 13). After applying the minimum sample size threshold (n ≥ 10) to the original dataset, a total of 42 unique feature combinations were retained for strategy-level analysis across all disease phase subsets. Examples of the highest-ranked trial design configurations (by disease and phase) identified within the modeling framework are shown in
Table 2.
As shown in
Figure 2, configurations identified by the model as robustness-oriented were associated with relatively higher model-derived progression probabilities compared with the overall distribution of historical configurations. The pooled mean predicted progression probability was 48.93% for robust-oriented configurations versus 39.44% in the historical distribution; however, the two distributions showed substantial overlap.
Under standardized assumptions for per-trial duration and cost, scenario-based analyses suggested differences in projected time and cost distributions between the two groups. The weighted mean projected duration was 106.87 months for robust-oriented configurations compared with 128.25 months for historical configurations. The corresponding weighted mean projected cost was USD 100.67 million versus USD 108.33 million, respectively.
Monte Carlo simulation scatter plots comparing model-based and historical strategy distributions are presented in the
Supplementary file (p. 16). Across the simulated scenarios, the model-based configurations showed a modest tendency toward higher progression probabilities together with lower projected cumulative time and cost, although substantial overlap between the two distributions remained.
4. Discussion
Using 1618 global vaccine clinical trials, this study provides a structured characterization of historical associations between observable trial design features and progression outcomes within registry-based data. Larger sample size, preventive vaccine purpose, enrollment across all age groups, and COVID-19 indication remained associated with higher observed progression odds after multivariable adjustment, whereas several operational characteristics, including multicenter structure and blinding, showed attenuated associations after adjustment. Some observed associations likely reflect broader scientific, operational, and regulatory contexts surrounding vaccine development rather than isolated effects of individual design characteristics. For example, differences in blinding strategy, study structure, and funding source may partially capture variation in trial phase, development urgency, endpoint selection, or sponsor resources. Accordingly, these findings should be interpreted as descriptive patterns embedded within historical development environments rather than evidence of causal relationships or prescriptive trial design advantages. Temporal trends indicated that vaccine trial activity increased steadily between 2012 and 2021, coinciding with expansion of global vaccine development efforts during the COVID-19 period [
31,
34,
35]. However, COVID-19-related trials did not uniformly exhibit the highest progression proportions across all disease categories [
15]. This further suggests that progression outcomes are shaped by complex interactions among scientific, epidemiological, regulatory, and operational factors that cannot be fully disentangled within registry-based observational data.
The machine learning framework summarized recurrent high-dimensional feature combinations observed across historical vaccine trials and highlighted disease–phase-specific configuration patterns associated with comparatively stable progression behavior within the dataset. By focusing on the five most influential variables identified by the model, the analysis highlighted configuration groups showing higher empirical progression proportions together with greater consistency between empirical and model-derived estimates. Robust-oriented configurations within this framework were associated with higher model-estimated progression probabilities and modestly lower simulated cumulative time and cost under standardized assumptions compared with historical strategy distributions. However, substantial overlap remained between the simulated distributions, indicating that the observed differences were incremental and scenario-dependent. Collectively, descriptive configuration assessment may provide a useful exploratory framework for characterizing historical progression patterns across heterogeneous vaccine development settings.
This study has limitations. First, reliance on ClinicalTrials.gov and supplementary public databases may introduce selection bias, reporting delays, and incomplete outcome capture [
21]. Trials lacking documented progression were operationally classified as non-progression, which may underestimate true progression rates when follow-up information was delayed or unavailable [
8,
36]. Although follow-up through 1 June 2025 was used to reduce right-censoring, some outcomes likely remained unobserved. Second, as an observational registry-based study, the analysis is inherently vulnerable to confounding and omitted variable bias. Important determinants of vaccine development outcomes—including immunogenicity, prior preclinical evidence, sponsor expertise, manufacturing feasibility, regulatory interactions, and epidemic dynamics—were not comprehensively captured in the available data. Consequently, observed associations may largely reflect broader development environments rather than direct relationships between individual trial features and progression outcomes. Third, although Random Forest modeling and internal cross-validation were used to summarize configuration patterns, the ranking framework depended on predefined aggregation procedures, weighting choices, and sparsity thresholds. Alternative specifications could yield different configuration rankings and simulation outputs. Future work should focus on external validation across independent datasets, alternative model specification strategies, and more formal causal or decision-analytic frameworks to assess whether any observed configuration patterns remain stable under broader development conditions.
5. Conclusions
This study describes progression patterns in vaccine clinical trials using a large historical registry dataset. Sample size, vaccine purpose, target population, and funding source showed variable associations with trial progression across disease areas and development phases.
A machine learning approach was used to summarize common combinations of trial design features observed in historical data and to identify groups of configurations that showed relatively consistent progression outcomes within the study dataset. Scenario analyses showed that these robust-oriented configurations were associated with small differences in projected development time and cost compared with historical trial design patterns.
Overall, this study provides a descriptive overview of how vaccine trial design features relate to progression outcomes in historical data. Further validation using external datasets and prospective evaluation is needed before any broader application.
Author Contributions
Conceptualization, S.C. and D.Z. (Dachuang Zhou); Methodology, D.Z. (Dachuang Zhou); Software, S.C.; Validation, S.C., D.Z. (Di Zhang) and Y.X.; Formal Analysis, D.Z. (Di Zhang); Investigation, W.T.; Resources, S.C.; Data Curation, S.C.; Writing—Original Draft Preparation, D.Z. (Dachuang Zhou); Writing—Review& Editing, S.C.; Visualization, Y.X.; Supervision, W.T.; Project Administration, W.T.; Funding Acquisition, W.T. All authors have read and agreed to the published version of the manuscript.
Funding
This work was supported by the National Natural Science Foundation of China, International Cooperative Research Project under Grant 2023YFVA1002.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
All data used in this study are publicly available and can be freely accessed and used by other researchers to replicate or extend the findings of this study. All data sources and references are explicitly cited in the manuscript or in the
Supplementary Materials.
Conflicts of Interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
References
- World Health Organization. World Immunization Week 2024. Available online: https://www.who.int/campaigns/world-immunization-week/2024 (accessed on 1 March 2026).
- World Bank Global Economy to Expand by 4 Percent in 2021; Vaccine Deployment and Investment Key to Sustaining the Recovery. Available online: https://www.worldbank.org/en/news/press-release/2021/01/05/global-economy-to-expand-by-4-percent-in-2021-vaccine-deployment-and-investment-key-to-sustaining-the-recovery (accessed on 14 October 2025).
- Van Kerkhove, M.D.; Ryan, M.J.; Ghebreyesus, T.A. Preparing for “Disease X”. Science 2021, 374, 377. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Artaud, C.; Kara, L.; Launay, O. Vaccine Development: From Preclinical Studies to Phase 1/2 Clinical Trials. In Malaria Control and Elimination; Ariey, F., Gay, F., Ménard, R., Eds.; Springer: New York, NY, USA, 2019; Volume 2013, pp. 165–176. [Google Scholar]
- MacPherson, A.; Hutchinson, N.; Schneider, O.; Oliviero, E.; Feldhake, E.; Ouimet, C.; Sheng, J.; Awan, F.; Wang, C.; Papenburg, J.; et al. Probability of Success and Timelines for the Development of Vaccines for Emerging and Reemerged Viral Infectious Diseases. Ann. Intern. Med. 2021, 174, 326–334. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Glasziou, P.; Altman, D.G.; Bossuyt, P.; Boutron, I.; Clarke, M.; Julious, S.; Michie, S.; Moher, D.; Wager, E. Reducing Waste from Incomplete or Unusable Reports of Biomedical Research. Lancet 2014, 383, 267–276. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sertkaya, A.; Beleche, T.; Jessup, A. New Estimates of the Cost of Preventive Vaccine Development and Potential Implications from the COVID-19 Pandemic: Brief; Office of the Assistant Secretary for Planning and Evaluation (ASPE): Washington, DC, USA, 2024. [Google Scholar]
- Wong, C.H.; Siah, K.W.; Lo, A.W. Estimation of Clinical Trial Success Rates and Related Parameters. Biostatistics 2019, 20, 273–286. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Robinson, C.H.; Parekh, R.S.; Cuthbertson, B.; Fan, E.; Ouyang, Y.; Heath, A. Using Bayesian Pre-Trial Simulations to Optimize the Design of Adaptive Clinical Trials in Childhood Nephrotic Syndrome. Contemp. Clin. Trials 2025, 153, 107918. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Siah, K.W.; Kelley, N.W.; Ballerstedt, S.; Holzhauer, B.; Lyu, T.; Mettler, D.; Sun, S.; Wandel, S.; Zhong, Y.; Zhou, B.; et al. Predicting Drug Approvals: The Novartis Data Science and Artificial Intelligence Challenge. Patterns 2021, 2, 100312. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yu, H.; Hutson, A.; Ma, X. Machine Learning Assisted Adjustment Boosts Efficiency of Exact Inference in Randomized Controlled Trials. Sci. Rep. 2025, 15, 24454. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- European Medicines Agency. Guideline on Clinical Evaluation of Vaccines: Revision 1; European Medicines Agency: Amsterdam, The Netherlands, 2023. [Google Scholar]
- Iyer, A.; Narayanaswami, S. A Novel Model Using ML Techniques for Clinical Trial Design and Expedited Patient Onboarding Process. Clin. Outcomes Res. 2025, 17, 1–18. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Doane, M.R. Statistical NLP for Optimization of Clinical Trial Success Prediction in Pharmaceutical R&D. Doctoral Dissertation, The George Washington University, Washington, DC, USA, 2025. [Google Scholar]
- U.S. National Library of Medicine. National Institutes of Health ClinicalTrials.Gov. Available online: https://clinicaltrials.gov/ (accessed on 4 September 2025).
- World Health Organization. Vaccines and Immunization. Available online: https://www.who.int/health-topics/vaccines-and-immunization#tab=tab_1 (accessed on 14 October 2025).
- Vaccines, Blood & Biologics|FDA. Available online: https://www.fda.gov/vaccines-blood-biologics (accessed on 1 August 2025).
- Medicines|European Medicines Agency. Available online: https://www.ema.europa.eu/en/medicines (accessed on 1 August 2025).
- U.S. Food and Drug Administration Emergency Use Authorization. Available online: https://www.fda.gov/emergency-preparedness-and-response/mcm-legal-regulatory-and-policy-framework/emergency-use-authorization (accessed on 1 August 2025).
- National Center for Health Statistics; Centers for Disease Control and Prevention. ICD-10, International Statistical Classification of Diseases and Related Health Problems. Tabular List, 2022; Centers for Disease Control and Prevention (CDC); National Center for Health Statistics: Hyattsville, MD, USA, 2022. [Google Scholar]
- Zarin, D.A.; Tse, T.; Williams, R.J.; Califf, R.M.; Ide, N.C. The ClinicalTrials. Gov Results Database—Update and Key Issues. N. Engl. J. Med. 2011, 364, 852–860. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wasserstein, R.L.; Lazar, N.A. The ASA Statement on p -Values: Context, Process, and Purpose. Am. Stat. 2016, 70, 129–133. [Google Scholar] [CrossRef] [Scilit]
- Bland, J.M. Statistics Notes: The Odds Ratio. BMJ 2000, 320, 1468. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Feijoo, F.; Palopoli, M.; Bernstein, J.; Siddiqui, S.; Albright, T.E. Key Indicators of Phase Transition for Clinical Trials through Machine Learning. Drug Discov. Today 2020, 25, 414–421. [Google Scholar] [CrossRef] [Scilit]
- Aliper, A.; Kudrin, R.; Polykovskiy, D.; Kamya, P.; Tutubalina, E.; Chen, S.; Ren, F.; Zhavoronkov, A. Prediction of Clinical Trials Outcomes Based on Target Choice and Clinical Trial Design with Multi-Modal Artificial Intelligence. Clin. Pharmacol. Ther. 2023, 114, 972–980. [Google Scholar] [CrossRef] [Scilit]
- Liu, S.; Du, H.; Feng, M. Robust Predictive Models in Clinical Data—Random Forest and Support Vector Machines. In Leveraging Data Science for Global Health; Celi, L.A., Majumder, M.S., Ordóñez, P., Osorio, J.S., Paik, K.E., Somai, M., Eds.; Springer International Publishing: Cham, Switzerland, 2020; pp. 219–228. [Google Scholar]
- Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
- Kerman, J. Neutral Noninformative and Informative Conjugate Beta and Gamma Prior Distributions. Electron. J. Stat. 2011, 5, 1450–1470. [Google Scholar] [CrossRef] [Scilit]
- Agresti, A.; Coull, B.A. Approximate Is Better than “Exact” for Interval Estimation of Binomial Proportions. Am. Stat. 1998, 52, 119–126. [Google Scholar] [CrossRef] [Scilit]
- McGuire, R. Impact of Clinical Development on Oncology Drug Prices. Available online: https://pharmaphorum.com/views-and-analysis/impact-of-clinical-development-on-oncology-drug-prices (accessed on 14 October 2025).
- Lurie, N.; Saville, M.; Hatchett, R.; Halton, J. Developing Covid-19 Vaccines at Pandemic Speed. N. Engl. J. Med. 2020, 382, 1969–1973. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- MedPath Rising Costs and Success Rates: The Complex Economics of Oncology Drug Development. Available online: https://trial.medpath.com/news/e4b7b37aabae4852/rising-costs-and-success-rates-the-complex-economics-of-oncology-drug-development (accessed on 14 October 2025).
- U.S. Bureau of Labor Statistics Consumer Price Index Databases. Available online: https://www.bls.gov/cpi/data.htm (accessed on 14 October 2025).
- Le, T.T.; Cramer, J.P.; Chen, R.; Mayhew, S. Evolution of the COVID-19 Vaccine Development Landscape. Nat. Rev. Drug Discov. 2020, 19, 667–668. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, M.; Wang, H.; Tian, L.; Pang, Z.; Yang, Q.; Huang, T.; Fan, J.; Song, L.; Tong, Y.; Fan, H. COVID-19 Vaccine Development: Milestones, Lessons and Prospects. Signal Transduct. Target. Ther. 2022, 7, 146. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Huang, J.J.; Reznik, A.S.; Sonis, S.; Franzmann, E.J.; Villa, A. Clinical Trial Termination or Withdrawal in Head and Neck Squamous Cell Carcinoma. JAMA Otolaryngol.-Head Neck Surg. 2026, 152, 242–248. [Google Scholar] [CrossRef] [Scilit] [PubMed]
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |