Editor’s Choice Articles

Editor’s Choice articles are based on recommendations by the scientific editors of MDPI journals from around the world. Editors select a small number of articles recently published in the journal that they believe will be particularly interesting to readers, or important in the respective research area. The aim is to provide a snapshot of some of the most exciting work published in the various research areas of the journal.

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
19 pages, 1922 KB  
Article
Validated Transfer Learning Peters–Belson Methods for Survival Analysis: Ensemble Machine Learning Approaches with Overfitting Controls for Health Disparity Decomposition
by Menglu Liang and Yan Li
Stats 2025, 8(4), 114; https://doi.org/10.3390/stats8040114 - 10 Dec 2025
Viewed by 1053
Abstract
Background: Health disparities research increasingly relies on complex survey data to understand survival differences between population subgroups. While Peters–Belson decomposition provides a principled framework for distinguishing disparities explained by measured covariates from unexplained residual differences, traditional approaches face challenges with complex data patterns [...] Read more.
Background: Health disparities research increasingly relies on complex survey data to understand survival differences between population subgroups. While Peters–Belson decomposition provides a principled framework for distinguishing disparities explained by measured covariates from unexplained residual differences, traditional approaches face challenges with complex data patterns and model validation for counterfactual estimation. Objective: To develop validated Peters–Belson decomposition methods for survival analysis that integrate ensemble machine learning with transfer learning while ensuring logical validity of counterfactual estimates through comprehensive model validation. Methods: We extend the traditional Peters–Belson framework through ensemble machine learning that combines Cox proportional hazards models, cross-validated random survival forests, and regularized gradient boosting approaches. Our framework incorporates a transfer learning component via principal component analysis (PCA) to discover shared latent factors between majority and minority groups. We note that this “transfer learning” differs from the standard machine learning definition (pre-trained models or domain adaptation); here, we use the term in its statistical sense to describe the transfer of covariate structure information from the pooled population to identify group-level latent factors. We develop a comprehensive validation framework that ensures Peters–Belson logical bounds compliance, preventing mathematical violations in counterfactual estimates. The approach is evaluated through simulation studies across five realistic health disparity scenarios using stratified complex survey designs. Results: Simulation studies demonstrate that validated ensemble methods achieve superior performance compared to individual models (proportion explained: 0.352 vs. 0.310 for individual Cox, 0.325 for individual random forests), with validation framework reducing logical violations from 34.7% to 2.1% of cases. Transfer learning provides additional 16.1% average improvement in explanation of unexplained disparity when significant unmeasured confounding exists, with 90.1% overall validation success rate. The validation framework ensures explanation proportions remain within realistic bounds while maintaining computational efficiency with 31% overhead for validation procedures. Conclusions: Validated ensemble machine learning provides substantial advantages for Peters–Belson decomposition when combined with proper model validation. Transfer learning offers conditional benefits for capturing unmeasured group-level factors while preventing mathematical violations common in standard approaches. The framework demonstrates that realistic health disparity patterns show 25–35% of differences explained by measured factors, providing actionable targets for reducing health inequities. Full article
Show Figures

Figure 1

31 pages, 3426 KB  
Article
Maximum Likelihood and Calibrating Prior Prediction Reliability Bias Reference Charts
by Stephen Jewson
Stats 2025, 8(4), 109; https://doi.org/10.3390/stats8040109 - 6 Nov 2025
Viewed by 2277
Abstract
There are many studies in the scientific literature that present predictions from parametric statistical models based on maximum likelihood estimates of the unknown parameters. However, generating predictions from maximum likelihood parameter estimates ignores the uncertainty around the parameter estimates. As a result, predictive [...] Read more.
There are many studies in the scientific literature that present predictions from parametric statistical models based on maximum likelihood estimates of the unknown parameters. However, generating predictions from maximum likelihood parameter estimates ignores the uncertainty around the parameter estimates. As a result, predictive probability distributions based on maximum likelihood are typically too narrow, and simulation testing has shown that tail probabilities are underestimated compared to the relative frequencies of out-of-sample events. We refer to this underestimation as a reliability bias. Previous authors have shown that objective Bayesian methods can eliminate or reduce this bias if the prior is chosen appropriately. Such methods have been given the name calibrating prior prediction. We investigate maximum likelihood reliability bias in more detail. We then present reference charts that quantify the reliability bias for 18 commonly used statistical models, for both maximum likelihood prediction and calibrating prior prediction. The charts give results for a large number of combinations of sample size and nominal probability and contain orders of magnitude more information about the reliability biases in predictions from these methods than has previously been published. These charts serve two purposes. First, they can be used to evaluate the extent to which maximum likelihood predictions given in the scientific literature are affected by reliability bias. If the reliability bias is large, the predictions may need to be revised. Second, the charts can be used in the design of future studies to assess whether it is appropriate to use maximum likelihood prediction, whether it would be more appropriate to reduce the reliability bias by using calibrating prior prediction, or whether neither maximum likelihood prediction nor calibrating prior prediction gives an adequately low reliability bias. Full article
Show Figures

Figure 1

20 pages, 386 KB  
Article
A High Dimensional Omnibus Regression Test
by Ahlam M. Abid, Paul A. Quaye and David J. Olive
Stats 2025, 8(4), 107; https://doi.org/10.3390/stats8040107 - 5 Nov 2025
Cited by 3 | Viewed by 1261
Abstract
Consider regression models where the response variable Y only depends on the p×1 vector of predictors x=(x1,,xp)T through the sufficient predictor SP=α+xTβ. [...] Read more.
Consider regression models where the response variable Y only depends on the p×1 vector of predictors x=(x1,,xp)T through the sufficient predictor SP=α+xTβ. Let the covariance vector Cov(x,Y)=ΣxY. Assume the cases (xiT,Yi)T are independent and identically distributed random vectors for i=1,,n. Then for many such regression models, β=0 if and only if ΣxY=0 where 0 is the p×1 vector of zeroes. The test of H0:ΣxY=0 versus H1:ΣxY0 is equivalent to the high dimensional one sample test H0:μ=0 versus HA:μ0 applied to w1,,wn where wi=(xiμx)(YiμY) and the expected values E(x)=μx and E(Y)=μY. Since μx and μY are unknown, the test of H0:β=0 versus H1:β0 is implemented by applying the one sample test to vi=(xix¯)(YiY¯) for i=1,,n. This test has milder regularity conditions than its few competitors. For the multiple linear regression one component partial least squares and marginal maximum likelihood estimators, the test can be adapted to test H0:(βi1,,βik)T=0 versus H1:(βi1,,βik)T0 where 1kp. Full article
(This article belongs to the Section Regression Models)
9 pages, 590 KB  
Article
Predictions of War Duration
by Glenn McRae
Stats 2025, 8(4), 92; https://doi.org/10.3390/stats8040092 - 9 Oct 2025
Viewed by 3415
Abstract
The durations of wars fought between 1480 and 1941 A.D. were found to be well represented by random numbers chosen from a single-event Poisson distribution with a half-life of (1.25 ± 0.1) years. This result complements the work of L.F. Richardson who found [...] Read more.
The durations of wars fought between 1480 and 1941 A.D. were found to be well represented by random numbers chosen from a single-event Poisson distribution with a half-life of (1.25 ± 0.1) years. This result complements the work of L.F. Richardson who found that the frequency of outbreaks of wars can be described as a Poisson process. This result suggests that a quick return on investment requires a distillation of the many stressors of the day, each one of which has a small probability of being included in a convincing well-orchestrated simple call-to-arms. The half-life is a measure of how this call wanes with time. Full article
Show Figures

Figure 1

19 pages, 1013 KB  
Article
A Simulation-Based Comparative Analysis of Two-Parameter Robust Ridge M-Estimators for Linear Regression Models
by Bushra Haider, Syed Muhammad Asim, Danish Wasim and B. M. Golam Kibria
Stats 2025, 8(4), 84; https://doi.org/10.3390/stats8040084 - 24 Sep 2025
Cited by 4 | Viewed by 2559
Abstract
Traditional regression estimators like Ordinary Least Squares (OLS) and classical ridge regression often fail under multicollinearity and outlier contamination respectively. Although recently developed two-parameter ridge regression (TPRR) estimators improve efficiency by introducing dual shrinkage parameters, they remain sensitive to extreme observations. This study [...] Read more.
Traditional regression estimators like Ordinary Least Squares (OLS) and classical ridge regression often fail under multicollinearity and outlier contamination respectively. Although recently developed two-parameter ridge regression (TPRR) estimators improve efficiency by introducing dual shrinkage parameters, they remain sensitive to extreme observations. This study develops a new class of Two-Parameter Robust Ridge M-Estimators (TPRRM) that integrate dual shrinkage with robust M-estimation to simultaneously address multicollinearity and outliers. A Monte Carlo simulation study, conducted under varying sample sizes, predictor dimensions, correlation levels, and contamination structures, compares the proposed estimators with OLS, ridge, and the most recent TPRR estimators. The results demonstrate that TPRRM consistently achieves the lowest Mean Squared Error (MSE), particularly in heavy-tailed and outlier-prone scenarios. Application to the Tobacco and Gasoline Consumption datasets further validates the superiority of the proposed methods in real-world conditions. The findings confirm that the proposed TPRRM fills a critical methodological gap by offering estimators that are not only efficient under multicollinearity, but also robust against departures from normality. Full article
Show Figures

Figure 1

27 pages, 1173 KB  
Article
The Unit-Modified Weibull Distribution: Theory, Estimation, and Real-World Applications
by Ammar M. Sarhan, Thamer Manshi and M. E. Sobh
Stats 2025, 8(3), 81; https://doi.org/10.3390/stats8030081 - 12 Sep 2025
Cited by 4 | Viewed by 2193
Abstract
This paper introduces the Unit-Modified Weibull (UMW) distribution, a novel probability model defined on the unit interval (0, 1). We derive its key statistical properties and estimate its parameters using the maximum likelihood method. The performance of the [...] Read more.
This paper introduces the Unit-Modified Weibull (UMW) distribution, a novel probability model defined on the unit interval (0, 1). We derive its key statistical properties and estimate its parameters using the maximum likelihood method. The performance of the estimators is assessed via a simulation study based on mean squared error, coverage probability, and average confidence interval length. To evaluate the practical utility of the model, we analyze three real-world data sets. Both parametric and nonparametric goodness-of-fit techniques are employed to compare the UMW distribution with several well-established competing models. In addition, nonparametric diagnostic tools such as total time on test transform plots and violin plots are used to explore the data’s behavior and assess the adequacy of the proposed model. Results indicate that the UMW distribution offers a competitive and flexible alternative for modeling bounded data. Full article
Show Figures

Figure 1

22 pages, 1533 KB  
Article
A Markov Chain Monte Carlo Procedure for Efficient Bayesian Inference on the Phase-Type Aging Model
by Cong Nie, Xiaoming Liu, Serge Provost and Jiandong Ren
Stats 2025, 8(3), 77; https://doi.org/10.3390/stats8030077 - 27 Aug 2025
Cited by 2 | Viewed by 2630
Abstract
The phase-type aging model (PTAM) belongs to a class of Coxian-type Markovian models that can provide a quantitative description of well-known aging characteristics that are part of a genetically determined, progressive, and irreversible process. Due to its unique parameter structure, estimation via the [...] Read more.
The phase-type aging model (PTAM) belongs to a class of Coxian-type Markovian models that can provide a quantitative description of well-known aging characteristics that are part of a genetically determined, progressive, and irreversible process. Due to its unique parameter structure, estimation via the MLE method presents a considerable estimability issue, whereby profile likelihood functions are flat and analytically intractable. In this study, a Markov chain Monte Carlo (MCMC)-based Bayesian methodology is proposed and applied to the PTAM, with a view to improving parameter estimability. The proposed method provides two methodological extensions based on an existing MCMC inference method. First, we propose a two-level MCMC sampling scheme that makes the method applicable to situations where the posterior distributions do not assume simple forms after data augmentation. Secondly, an existing data augmentation technique for Bayesian inference on continuous phase-type distributions is further developed in order to incorporate left-truncated data. While numerical results indicate that the proposed methodology improves parameter estimability via sound prior distributions, this approach may also be utilized as a stand-alone statistical model-fitting technique. Full article
Show Figures

Figure 1

16 pages, 2240 KB  
Article
Pattern Classification for Mixed Feature-Type Symbolic Data Using Supervised Hierarchical Conceptual Clustering
by Manabu Ichino and Hiroyuki Yaguchi
Stats 2025, 8(3), 76; https://doi.org/10.3390/stats8030076 - 25 Aug 2025
Viewed by 1049
Abstract
This paper describes a region-oriented method of pattern classification based on the Cartesian system model (CSM), a mathematical model that allows manipulating mixed feature-type symbolic data. We use the supervised hierarchical conceptual clustering to generate class regions for respective pattern class based on [...] Read more.
This paper describes a region-oriented method of pattern classification based on the Cartesian system model (CSM), a mathematical model that allows manipulating mixed feature-type symbolic data. We use the supervised hierarchical conceptual clustering to generate class regions for respective pattern class based on the evaluation of the generality of the regions and the separability of the regions against other classes in each clustering step. We can easily find the robustly informative features to describe each pattern class against other pattern classes. Some examples show the effectiveness of the proposed method. Full article
Show Figures

Figure 1

19 pages, 750 KB  
Article
Evaluating Estimator Performance Under Multicollinearity: A Trade-Off Between MSE and Accuracy in Logistic, Lasso, Elastic Net, and Ridge Regression with Varying Penalty Parameters
by H. M. Nayem, Sinha Aziz and B. M. Golam Kibria
Stats 2025, 8(2), 45; https://doi.org/10.3390/stats8020045 - 31 May 2025
Cited by 13 | Viewed by 4119
Abstract
Multicollinearity in logistic regression models can result in inflated variances and yield unreliable estimates of parameters. Ridge regression, a regularized estimation technique, is frequently employed to address this issue. This study conducts a comparative evaluation of the performance of 23 established ridge regression [...] Read more.
Multicollinearity in logistic regression models can result in inflated variances and yield unreliable estimates of parameters. Ridge regression, a regularized estimation technique, is frequently employed to address this issue. This study conducts a comparative evaluation of the performance of 23 established ridge regression estimators alongside Logistic Regression, Elastic-Net, Lasso, and Generalized Ridge Regression (GRR), considering various levels of multicollinearity within the context of logistic regression settings. Simulated datasets with high correlations (0.80, 0.90, 0.95, and 0.99) and real-world data (municipal and cancer remission) were analyzed. Both results show that ridge estimators, such as kAL1, kAL2, kKL1, and kKL2, exhibit strong performance in terms of Mean Squared Error (MSE) and accuracy, particularly in smaller samples, while GRR demonstrates superior performance in large samples. Real-world data further confirm that GRR achieves the lowest MSE in highly collinear municipal data, while ridge estimators and GRR help prevent overfitting in small-sample cancer remission data. The results underscore the efficacy of ridge estimators and GRR in handling multicollinearity, offering reliable alternatives to traditional regression techniques, especially for datasets with high correlations and varying sample sizes. Full article
Show Figures

Figure 1

12 pages, 882 KB  
Article
mbX: An R Package for Streamlined Microbiome Analysis
by Utsav Lamichhane and Jeferson Lourenco
Stats 2025, 8(2), 44; https://doi.org/10.3390/stats8020044 - 29 May 2025
Cited by 3 | Viewed by 3576
Abstract
Here, we introduce the mbX package: an R-based tool designed to streamline 16S rRNA gene microbiome data analysis following taxonomic classification. It automates key post-sequencing steps, including taxonomic data cleaning and visualization, addressing the need for reproducible and user-friendly microbiome workflows. mbX’s core [...] Read more.
Here, we introduce the mbX package: an R-based tool designed to streamline 16S rRNA gene microbiome data analysis following taxonomic classification. It automates key post-sequencing steps, including taxonomic data cleaning and visualization, addressing the need for reproducible and user-friendly microbiome workflows. mbX’s core functions, ezclean and ezviz, take raw taxonomic output (such as those from QIIME 2) and sample metadata to produce a cleaned relative abundance dataset and high-quality stacked bar plots with minimal manual intervention. We validated mbX on 14 real microbiome datasets, demonstrating significant improvements in efficiency and consistency of post-processing of DNA sequence data. The results show that mbX ensures uniform taxonomic formatting, eliminates common manual errors, and quickly generates publication-ready figures, greatly facilitating downstream analysis. For a dataset with 20 samples, both functions of mbX ran in less than 1 s and used less than 1 GB of memory. For a dataset with more than 1170 samples, the functions ran within 125 s and used less than 4.5 GB of memory. By integrating seamlessly with existing pipelines and emphasizing automation, mbX fills a critical gap between sequence classification and statistical analysis. An upcoming version will have an added function which will further extend mbX to automated statistical comparisons, aiming for an end-to-end microbiome analysis solution by integrating mbX with currently available pipelines. This article presents the design of mbX, its workflow and features, and a comparative discussion positioning mbX relative to other microbiome bioinformatics tools. The contributions of mbX highlight its significance in accelerating microbiome research through reproducible and streamlined data analysis. Full article
(This article belongs to the Section Statistical Software)
Show Figures

Figure 1

13 pages, 497 KB  
Article
Modeling Uncertainty in Ordinal Regression: The Uncertainty Rating Scale Model
by Gerhard Tutz
Stats 2025, 8(2), 42; https://doi.org/10.3390/stats8020042 - 23 May 2025
Cited by 1 | Viewed by 1823
Abstract
In questionnaires, respondents sometimes feel uncertain about which category to choose and may respond randomly. Including uncertainty in the modeling of response behavior aims to obtain more accurate estimates of the impact of explanatory variables on actual preferences and to avoid bias. Additionally, [...] Read more.
In questionnaires, respondents sometimes feel uncertain about which category to choose and may respond randomly. Including uncertainty in the modeling of response behavior aims to obtain more accurate estimates of the impact of explanatory variables on actual preferences and to avoid bias. Additionally, variables that have an impact on uncertainty can be identified. A model is proposed that explicitly considers this uncertainty but also allows stronger certainty, depending on covariates. The developed uncertainty rating scale model is an extended version of the adjacent category model. It differs from finite mixture models, an approach that has gained popularity in recent years for modeling uncertainty. The properties of the model are investigated and compared to finite mixture models and other ordinal response models using illustrative datasets. Full article
Show Figures

Figure 1

17 pages, 646 KB  
Article
A Smoothed Three-Part Redescending M-Estimator
by Alistair J. Martin and Brenton R. Clarke
Stats 2025, 8(2), 33; https://doi.org/10.3390/stats8020033 - 30 Apr 2025
Viewed by 1593
Abstract
A smoothed M-estimator is derived from Hampel’s three-part redescending estimator for location and scale. The estimator is shown to be weakly continuous and Fréchet differentiable in the neighbourhood of the normal distribution. Asymptotic assessment is conducted at asymmetric contaminating distributions, where smoothing is [...] Read more.
A smoothed M-estimator is derived from Hampel’s three-part redescending estimator for location and scale. The estimator is shown to be weakly continuous and Fréchet differentiable in the neighbourhood of the normal distribution. Asymptotic assessment is conducted at asymmetric contaminating distributions, where smoothing is shown to improve variance and change-of-variance sensitivity. Other robust metrics compared are largely unchanged, and therefore, the smoothed functions represent an improvement for asymmetric contamination near the rejection point with little downside. Full article
(This article belongs to the Section Statistical Methods)
Show Figures

Figure 1

21 pages, 326 KB  
Article
Quantum-Inspired Latent Variable Modeling in Multivariate Analysis
by Theodoros Kyriazos and Mary Poga
Stats 2025, 8(1), 20; https://doi.org/10.3390/stats8010020 - 28 Feb 2025
Cited by 2 | Viewed by 2976
Abstract
Latent variables play a crucial role in psychometric research, yet traditional models often struggle to address context-dependent effects, ambivalent states, and non-commutative measurement processes. This study proposes a quantum-inspired framework for latent variable modeling that employs Hilbert space representations, allowing questionnaire items to [...] Read more.
Latent variables play a crucial role in psychometric research, yet traditional models often struggle to address context-dependent effects, ambivalent states, and non-commutative measurement processes. This study proposes a quantum-inspired framework for latent variable modeling that employs Hilbert space representations, allowing questionnaire items to be treated as pure or mixed quantum states. By integrating concepts such as superposition, interference, and non-commutative probabilities, the framework captures cognitive and behavioral phenomena that extend beyond the capabilities of classical methods. To illustrate its potential, we introduce quantum-specific metrics—fidelity, overlap, and von Neumann entropy—as complements to correlation-based measures. We also outline a machine-learning pipeline using complex and real-valued neural networks to handle amplitude and phase information. Results highlight the capacity of quantum-inspired models to reveal order effects, ambivalent responses, and multimodal distributions that remain elusive in standard psychometric approaches. This framework broadens the multivariate analysis theoretical and methodological toolkit, offering a dynamic and context-sensitive perspective on latent constructs while inviting further empirical validation in diverse research settings. Full article
(This article belongs to the Section Multivariate Analysis)
20 pages, 3127 KB  
Article
A New Weighted Lindley Model with Applications to Extreme Historical Insurance Claims
by Morad Alizadeh, Mahmoud Afshari, Gauss M. Cordeiro, Ziaurrahman Ramaki, Javier E. Contreras-Reyes, Fatemeh Dirnik and Haitham M. Yousof
Stats 2025, 8(1), 8; https://doi.org/10.3390/stats8010008 - 15 Jan 2025
Cited by 23 | Viewed by 2633
Abstract
In this paper, we propose a weighted Lindley (NWLi) model for the analysis of extreme historical insurance claims. It extends the classical Lindley distribution by incorporating a weight parameter, enabling more flexibility in modeling insurance claim severity. We provide a comprehensive theoretical overview [...] Read more.
In this paper, we propose a weighted Lindley (NWLi) model for the analysis of extreme historical insurance claims. It extends the classical Lindley distribution by incorporating a weight parameter, enabling more flexibility in modeling insurance claim severity. We provide a comprehensive theoretical overview of the new model and explore two practical applications. First, we investigate the mean-of-order P (MOOP(P)) approach for quantifying the expected claim severity based on the NWLi model. Second, we implement a peaks over a random threshold (PORT) analysis using the value-at-risk metric to assess extreme claim occurrences under the new model. Further, we provide a simulation study to evaluate the accuracy of the estimators under various methods. The proposed model and its applications provide a versatile tool for actuaries and risk analysts to analyze and predict extreme insurance claim severity, offering insights into risk management and decision-making within the insurance industry. Full article
(This article belongs to the Section Reliability Engineering)
Show Figures

Figure 1

18 pages, 1035 KB  
Article
Bidirectional f-Divergence-Based Deep Generative Method for Imputing Missing Values in Time-Series Data
by Wen-Shan Liu, Tong Si, Aldas Kriauciunas, Marcus Snell and Haijun Gong
Stats 2025, 8(1), 7; https://doi.org/10.3390/stats8010007 - 14 Jan 2025
Cited by 9 | Viewed by 2893
Abstract
Imputing missing values in high-dimensional time-series data remains a significant challenge in statistics and machine learning. Although various methods have been proposed in recent years, many struggle with limitations and reduced accuracy, particularly when the missing rate is high. In this work, we [...] Read more.
Imputing missing values in high-dimensional time-series data remains a significant challenge in statistics and machine learning. Although various methods have been proposed in recent years, many struggle with limitations and reduced accuracy, particularly when the missing rate is high. In this work, we present a novel f-divergence-based bidirectional generative adversarial imputation network, tf-BiGAIN, designed to address these challenges in time-series data imputation. Unlike traditional imputation methods, tf-BiGAIN employs a generative model to synthesize missing values without relying on distributional assumptions. The imputation process is achieved by training two neural networks, implemented using bidirectional modified gated recurrent units, with f-divergence serving as the objective function to guide optimization. Compared to existing deep learning-based methods, tf-BiGAIN introduces two key innovations. First, the use of f-divergence provides a flexible and adaptable framework for optimizing the model across diverse imputation tasks, enhancing its versatility. Second, the use of bidirectional gated recurrent units allows the model to leverage both forward and backward temporal information. This bidirectional approach enables the model to effectively capture dependencies from both past and future observations, enhancing its imputation accuracy and robustness. We applied tf-BiGAIN to analyze two real-world time-series datasets, demonstrating its superior performance in imputing missing values and outperforming existing methods in terms of accuracy and robustness. Full article
Show Figures

Figure 1

17 pages, 1524 KB  
Article
Exact Inference for Random Effects Meta-Analyses for Small, Sparse Data
by Jessica Gronsbell, Zachary R. McCaw, Timothy Regis and Lu Tian
Stats 2025, 8(1), 5; https://doi.org/10.3390/stats8010005 - 7 Jan 2025
Cited by 3 | Viewed by 2198
Abstract
Meta-analysis aggregates information across related studies to provide more reliable statistical inference and has been a vital tool for assessing the safety and efficacy of many high-profile pharmaceutical products. A key challenge in conducting a meta-analysis is that the number of related studies [...] Read more.
Meta-analysis aggregates information across related studies to provide more reliable statistical inference and has been a vital tool for assessing the safety and efficacy of many high-profile pharmaceutical products. A key challenge in conducting a meta-analysis is that the number of related studies is typically small. Applying classical methods that are asymptotic in the number of studies can compromise the validity of inference, particularly when heterogeneity across studies is present. Moreover, serious adverse events are often rare and can result in one or more studies with no events in at least one study arm. Practitioners remove studies in which no events have occurred in one or both arms or apply arbitrary continuity corrections (e.g., adding one event to arms with zero events) to stabilize or define effect estimates in such settings, which can further invalidate subsequent inference. To address these significant practical issues, we introduce an exact inference method for random effects meta-analysis of a treatment effect in the two-sample setting with rare events, which we coin “XRRmeta”. In contrast to existing methods, XRRmeta provides valid inference for meta-analysis in the presence of between-study heterogeneity and when the event rates, number of studies, and/or the within-study sample sizes are small. Extensive numerical studies indicate that XRRmeta does not yield overly conservative inference. We apply our proposed method to two real-data examples using our open-source R package. Full article
Show Figures

Figure 1

13 pages, 694 KB  
Article
Multiple Imputation of Composite Covariates in Survival Studies
by Lily Clements, Alan C. Kimber and Stefanie Biedermann
Stats 2022, 5(2), 358-370; https://doi.org/10.3390/stats5020020 - 29 Mar 2022
Cited by 2 | Viewed by 3965
Abstract
Missing covariate values are a common problem in survival studies, and the method of choice when handling such incomplete data is often multiple imputation. However, it is not obvious how this can be used most effectively when an incomplete covariate is a function [...] Read more.
Missing covariate values are a common problem in survival studies, and the method of choice when handling such incomplete data is often multiple imputation. However, it is not obvious how this can be used most effectively when an incomplete covariate is a function of other covariates. For example, body mass index (BMI) is the ratio of weight and height-squared. In this situation, the following question arises: Should a composite covariate such as BMI be imputed directly, or is it advantageous to impute its constituents, weight and height, first and to construct BMI afterwards? We address this question through a carefully designed simulation study that compares various approaches to multiple imputation of composite covariates in a survival context. We discuss advantages and limitations of these approaches for various types of missingness and imputation models. Our results are a first step towards providing much needed guidance to practitioners for analysing their incomplete survival data effectively. Full article
(This article belongs to the Special Issue Survival Analysis: Models and Applications)
Show Figures

Figure 1

19 pages, 878 KB  
Article
A Bayesian Approach for Imputation of Censored Survival Data
by Shirin Moghaddam, John Newell and John Hinde
Stats 2022, 5(1), 89-107; https://doi.org/10.3390/stats5010006 - 26 Jan 2022
Cited by 9 | Viewed by 7337
Abstract
A common feature of much survival data is censoring due to incompletely observed lifetimes. Survival analysis methods and models have been designed to take account of this and provide appropriate relevant summaries, such as the Kaplan–Meier plot and the commonly quoted median survival [...] Read more.
A common feature of much survival data is censoring due to incompletely observed lifetimes. Survival analysis methods and models have been designed to take account of this and provide appropriate relevant summaries, such as the Kaplan–Meier plot and the commonly quoted median survival time of the group under consideration. However, a single summary is not really a relevant quantity for communication to an individual patient, as it conveys no notion of variability and uncertainty, and the Kaplan–Meier plot can be difficult for the patient to understand and also is often mis-interpreted, even by some physicians. This paper considers an alternative approach of treating the censored data as a form of missing, incomplete data and proposes an imputation scheme to construct a completed dataset. This allows the use of standard descriptive statistics and graphical displays to convey both typical outcomes and the associated variability. We propose a Bayesian approach to impute any censored observations, making use of other information in the dataset, and provide a completed dataset. This can then be used for standard displays, summaries, and even, in theory, analysis and model fitting. We particularly focus on the data visualisation advantages of the completed data, allowing displays such as density plots, boxplots, etc, to complement the usual Kaplan–Meier display of the original dataset. We study the performance of this approach through a simulation study and consider its application to two clinical examples. Full article
(This article belongs to the Special Issue Survival Analysis: Models and Applications)
Show Figures

Figure 1

25 pages, 467 KB  
Article
Resampling Plans and the Estimation of Prediction Error
by Bradley Efron
Stats 2021, 4(4), 1091-1115; https://doi.org/10.3390/stats4040063 - 20 Dec 2021
Cited by 7 | Viewed by 6437
Abstract
This article was prepared for the Special Issue on Resampling methods for statistical inference of the 2020s. Modern algorithms such as random forests and deep learning are automatic machines for producing prediction rules from training data. Resampling plans have been the key [...] Read more.
This article was prepared for the Special Issue on Resampling methods for statistical inference of the 2020s. Modern algorithms such as random forests and deep learning are automatic machines for producing prediction rules from training data. Resampling plans have been the key technology for evaluating a rule’s prediction accuracy. After a careful description of the measurement of prediction error the article discusses the advantages and disadvantages of the principal methods: cross-validation, the nonparametric bootstrap, covariance penalties (Mallows’ Cp and the Akaike Information Criterion), and conformal inference. The emphasis is on a broad overview of a large subject, featuring examples, simulations, and a minimum of technical detail. Full article
(This article belongs to the Special Issue Re-sampling Methods for Statistical Inference of the 2020s)
Show Figures

Figure 1

17 pages, 429 KB  
Article
Survival Augmented Patient Preference Incorporated Reinforcement Learning to Evaluate Tailoring Variables for Personalized Healthcare
by Yingchao Zhong, Chang Wang and Lu Wang
Stats 2021, 4(4), 776-792; https://doi.org/10.3390/stats4040046 - 27 Sep 2021
Cited by 5 | Viewed by 4278
Abstract
In this paper, we consider personalized treatment decision strategies in the management of chronic diseases, such as chronic kidney disease, which typically consists of sequential and adaptive treatment decision making. We investigate a two-stage treatment setting with a survival outcome that could be [...] Read more.
In this paper, we consider personalized treatment decision strategies in the management of chronic diseases, such as chronic kidney disease, which typically consists of sequential and adaptive treatment decision making. We investigate a two-stage treatment setting with a survival outcome that could be right censored. This can be formulated through a dynamic treatment regime (DTR) framework, where the goal is to tailor treatment to each individual based on their own medical history in order to maximize a desirable health outcome. We develop a new method, Survival Augmented Patient Preference incorporated reinforcement Q-Learning (SAPP-Q-Learning) to decide between quality of life and survival restricted at maximal follow-up. Our method incorporates the latent patient preference into a weighted utility function that balances between quality of life and survival time, in a Q-learning model framework. We further propose a corresponding m-out-of-n Bootstrap procedure to accurately make statistical inferences and construct confidence intervals on the effects of tailoring variables, whose values can guide personalized treatment strategies. Full article
(This article belongs to the Special Issue Survival Analysis: Models and Applications)
Show Figures

Figure 1

Back to TopTop