Previous Issue
Volume 9, June
 
 

Stats, Volume 9, Issue 4 (August 2026) – 16 articles

  • Issues are regarded as officially published after their release is announced to the table of contents alert mailing list.
  • You may sign up for e-mail alerts to receive table of contents of newly released issues.
  • PDF is the official format for papers published in both, html and pdf forms. To view the papers in pdf format, click on the "PDF Full-text" link, and use the free Adobe Reader to open them.
Order results
Result details
Section
Select all
Export citation of selected articles as:
48 pages, 1391 KB  
Article
Modeling Various Data Structures via the New Type II Exponentiated Half Logistic-Odd Log-Logistic-G Power Series Class of Distributions
by Thatayaone Moakofi, Broderick Oluyede, Neo Dingalo and Bakang Tlhaloganyang
Stats 2026, 9(4), 82; https://doi.org/10.3390/stats9040082 - 6 Aug 2026
Viewed by 69
Abstract
In this paper, we introduce the type II exponentiated half logistic-odd log-logistic-G power series class of distributions for modeling symmetric, skewed and heavy-tailed data with diverse hazard rate shapes. The proposed class of distributions is obtained by compounding the generalized family of distributions [...] Read more.
In this paper, we introduce the type II exponentiated half logistic-odd log-logistic-G power series class of distributions for modeling symmetric, skewed and heavy-tailed data with diverse hazard rate shapes. The proposed class of distributions is obtained by compounding the generalized family of distributions involving the type II exponentiated half logistic-G and odd log-logistic-G families with a discrete power series distribution. Various statistical properties of the proposed class of distributions, including moments, survival and hazard rate functions, order statistics, probability weighted moments, and Rényi entropy are derived. The model parameters are estimated using different estimation methods, and their performance is evaluated through Monte Carlo simulation studies. Finally, the flexibility and applicability of the proposed class of distributions are illustrated using real data sets. The results demonstrate that the proposed model provides a better fit than several existing competing models. Full article
Show Figures

Figure 1

29 pages, 4576 KB  
Article
Modeling Healthcare Data with Logistic Quantile and Uniform-Based Mixture Polynomial Distributions
by Mohan D. Pant, Aditya Chakraborty and Jovanna A. Tracz
Stats 2026, 9(4), 81; https://doi.org/10.3390/stats9040081 - 4 Aug 2026
Viewed by 127
Abstract
Continuous healthcare data often deviate from normality, which can substantially increase the risk of making invalid inferences, given that many inferential statistical procedures rely on normality assumption. To obviate this issue, we propose a new family of non-normal distributions based on a linear [...] Read more.
Continuous healthcare data often deviate from normality, which can substantially increase the risk of making invalid inferences, given that many inferential statistical procedures rely on normality assumption. To obviate this issue, we propose a new family of non-normal distributions based on a linear combination of the quantile functions of standard logistic and uniform (0, 1) distributions. This new family of non-normal distributions is studied within three different methods: L-moments, conventional moments, and percentiles. Its performance is compared among the three methods in the context of parameter estimation and data modeling. The results of Monte Carlo simulation and bootstrapping techniques indicate that the L-moment-based estimates of parameters of L-skewness and L-kurtosis are substantially less biased than their percentile-based estimates of left–right tail-weight ratio (a measure of skewness) and tail-weight factor (a measure of kurtosis), which in turn are superior to their moment-based counterparts of skewness and kurtosis, especially for small sample sizes and higher-order moments. On the other hand, the data modeling results indicate that the percentile-based fits of the proposed distributions provide slightly better approximations to real-world healthcare data than their L-moment-based counterparts, whereas both percentile- and L-moment-based methods are superior to their conventional moment-based counterparts. Full article
(This article belongs to the Topic Statistics and Data Science)
Show Figures

Figure 1

23 pages, 1071 KB  
Article
Censored-Data Inference for Combined Consecutive-Type Systems with Imperfect Cold Standby Coverage
by Ioannis S. Triantafyllou
Stats 2026, 9(4), 80; https://doi.org/10.3390/stats9040080 - 29 Jul 2026
Viewed by 144
Abstract
In the present work, we study combined m-consecutive-k-out-of-n and consecutive kc-out-of-n reliability systems under imperfect cold standby redundancy. The proposed framework extends the ordinary perfect standby assumption by allowing the activation of the spare system to [...] Read more.
In the present work, we study combined m-consecutive-k-out-of-n and consecutive kc-out-of-n reliability systems under imperfect cold standby redundancy. The proposed framework extends the ordinary perfect standby assumption by allowing the activation of the spare system to be successful with a given coverage probability. The system-level redundancy policy is considered, while the classical perfect cold standby model is obtained as a special case. Exact reliability representations for the resulting structures are discussed through signature-based arguments. In particular, expressions for the reliability function, the Mean Time to Failure and the Mean Residual Lifetime are provided. Special emphasis is also placed on statistical inference under right-censored lifetime data. A numerical study is carried out to illustrate the effect of imperfect coverage, censoring, and design parameters on the performance of the proposed reliability schemes. Full article
Show Figures

Figure 1

31 pages, 1993 KB  
Article
Flexible Bivariate Generalized Shifted Inverse Trinomial Distributions for Over- and Under-Dispersed Count Data
by Shin-Zhu Sim, Seng-Huat Ong, Hong-Seng Sim, Yong-Kheng Goh and Hari Mohan Srivastava
Stats 2026, 9(4), 79; https://doi.org/10.3390/stats9040079 - 24 Jul 2026
Viewed by 204
Abstract
Modeling bivariate count data with complex dispersion and dependence structures remains a significant challenge in statistical data analysis. This article introduces two new bivariate count distributions derived from the generalized shifted inverse trinomial distribution. The proposed models, denoted by BGIT-I and BGIT-II, are [...] Read more.
Modeling bivariate count data with complex dispersion and dependence structures remains a significant challenge in statistical data analysis. This article introduces two new bivariate count distributions derived from the generalized shifted inverse trinomial distribution. The proposed models, denoted by BGIT-I and BGIT-II, are constructed using convolution and trivariate reduction methods. They provide flexible joint frameworks for modeling correlated count data while accommodating different marginal dispersion patterns. BGIT-I allows negative, near-zero, and positive dependence, whereas BGIT-II induces non-negative dependence through a common component. The proposed models have simple, tractable probability generating functions, which facilitate the derivation of probabilistic properties and motivate a probability-generating-function-based estimation approach alongside maximum-likelihood estimation. The finite-sample performance of the estimators is further examined through a Monte Carlo simulation study. The practical utility of the proposed models is illustrated using two real bivariate count data sets involving shunter accidents and patient counts in critical care and emergency room settings. The results show that the proposed BGIT models provide competitive alternatives for modeling bivariate count data with different dispersion and dependence characteristics. Full article
Show Figures

Figure 1

13 pages, 3069 KB  
Article
An Interpretable Decomposition-Based Framework for U.S. Influenza-like Illness Surveillance
by Changzhi Ma and Zheng Xu
Stats 2026, 9(4), 78; https://doi.org/10.3390/stats9040078 - 24 Jul 2026
Viewed by 245
Abstract
Public-health respiratory-disease surveillance requires distinguishing expected seasonal illness activity from atypical changes that may warrant further investigation. Weekly influenza-like illness (ILI) data support situational monitoring, retrospective anomaly screening, and short-term forecasting, but these tasks are often conducted using separate statistical baselines. We empirically [...] Read more.
Public-health respiratory-disease surveillance requires distinguishing expected seasonal illness activity from atypical changes that may warrant further investigation. Weekly influenza-like illness (ILI) data support situational monitoring, retrospective anomaly screening, and short-term forecasting, but these tasks are often conducted using separate statistical baselines. We empirically evaluated whether seasonal–trend decomposition using LOESS (STL) could provide a common interpretable baseline using weekly national U.S. ILI data from 1997 to 2025. Expected activity was represented by the estimated trend plus recurring seasonal component, while residuals represented baseline-adjusted deviations. Residual exceedance and persistence summaries were examined retrospectively, and forecast performance was evaluated through rolling-origin comparisons with SARIMA, seasonal naive, and ETS benchmarks at 4-, 8-, and 12-week horizons. The decomposition highlighted major departures from the recurring annual pattern, including the 2009 H1N1 pandemic and COVID-19-era disruptions. At the 4-week horizon, STL had the lowest RMSE (0.925), slightly below SARIMA (0.929), seasonal naive (1.270), and ETS (1.710). At the 8- and 12-week horizons, seasonal naive performed best, with RMSEs of 1.270 and 1.280, compared with 1.280 and 1.450 for STL. The two-week persistence rule produced no pre-peak warnings at any tested threshold, indicating that retrospective residual screening did not by itself provide effective early warning. The empirical study demonstrates that STL provides a transparent common baseline for linking monitoring, retrospective anomaly screening, and short-term forecasting while distinguishing expected seasonal progression from baseline-adjusted deviations. Full article
(This article belongs to the Topic Statistics and Data Science)
Show Figures

Figure 1

2 pages, 194 KB  
Correction
Correction: Cerqueti, R.; Lupi, C. Some New Tests of Conformity with Benford’s Law. Stats 2021, 4, 745–761
by Roy Cerqueti and Claudio Lupi
Stats 2026, 9(4), 77; https://doi.org/10.3390/stats9040077 - 22 Jul 2026
Viewed by 157
Abstract
There were several errors in the original publication [...] Full article
24 pages, 931 KB  
Article
BSTZINB: A Bayesian Framework for Negative-Binomial Modeling of Spatio-Temporal Zero-Inflated Count Data in Epidemiology
by Suman Majumder, Yoonbae Jun, Sounak Chakraborty, Chae Young Lim and Tanujit Dey
Stats 2026, 9(4), 76; https://doi.org/10.3390/stats9040076 - 20 Jul 2026
Viewed by 244
Abstract
Modern Bayesian hierarchical methodologies allow us to leverage spatio-temporal dependencies between observations, enhancing both health effect estimation and map visualization in efficient and flexible ways. However, the necessary levels of statistical software are often unavailable or difficult to access. We have recently examined [...] Read more.
Modern Bayesian hierarchical methodologies allow us to leverage spatio-temporal dependencies between observations, enhancing both health effect estimation and map visualization in efficient and flexible ways. However, the necessary levels of statistical software are often unavailable or difficult to access. We have recently examined Bayesian spatio-temporal models to estimate the association between COVID-19 death counts and various social and environmental risk factors, including ambient air pollution exposure. Typically, it is very common that in an infection disease mapping problem with count data, we have excessive zeros, and it is usually for over-dispersed count outcome variables. Furthermore, the theory suggests that the excess zeros are generated by a separate process from the count values and that the excess zeros need to be modeled independently. Our proposed models are specially designed to handle the zero-inflation and over-dispersion in count data through Zero-Inflated Negative Binomial regression with random effects that vary across time and space within a Markov Chain Monte Carlo framework. Drawing on our knowledge and experience, we aim to provide a simple, unified, and publicly available software that can be applied in various disease mapping studies under the contemporary Bayesian framework. Full article
Show Figures

Figure 1

20 pages, 617 KB  
Article
E-CVWMD and E-CVWMD-Pairwise: Novel Joint Performance Metrics for Mixed-Type Multivariate Hydroclimatic Models
by David Arango-Londoño, Delia Ortega-Lenis, Mauricio A. Mazo-Lopera and Paula Moraga
Stats 2026, 9(4), 75; https://doi.org/10.3390/stats9040075 - 16 Jul 2026
Viewed by 217
Abstract
Evaluating joint predictive performance for multivariate hydroclimatic models requires metrics that simultaneously assess marginal accuracy and cross-variable dependence recovery. Existing metricsthe Energy Score, Variogram Score, and their derivativesdo not adapt to the structural complexity of the residual correlation matrix, treating a single correlated [...] Read more.
Evaluating joint predictive performance for multivariate hydroclimatic models requires metrics that simultaneously assess marginal accuracy and cross-variable dependence recovery. Existing metricsthe Energy Score, Variogram Score, and their derivativesdo not adapt to the structural complexity of the residual correlation matrix, treating a single correlated pair identically to a fully dense dependence structure. We propose two novel metric families: Metric E (E-CVWMD: Enhanced Coefficient-of-Variation Weighted Marginal-Dependence) and Metric E2 (E-CVWMD-Pairwise), which are designed for mixed-type multivariate responses combining continuous and binary outcomes within a cross-validation framework. We position Metrics E and E2 as diagnostic ranking tools for comparing competing models rather than as strictly proper scoring rules, and we provide a strictly proper Log-Loss variant (E-LL/E2-LL) for applications that require the full properness guarantee. Metric E assigns variable-level weights proportional to the coefficient of variation (CV) of each outcome on the training partition and adaptively calibrates the marginal-dependence trade-off parameter α via a global distance-correlation test. Metric E2 refines this by replacing the global test with a pairwise Spearman screening index π^, the proportion of variable pairs with significant residual correlationwhich maps linearly to α(π^)=1π^/2[0.5,1]. Applied to the validation of a Generalized Multivariate Functional Additive Mixed Model (GMFAMM) on 62 Valle del Cauca meteorological stations (Ntest 31,663), the naive significance-based index saturates (π^=1.0) at this large sample sizeevery pair, including correlations as small as |ρ^s| 0.01, is flagged “significant”which is precisely the sample-size sensitivity we address. Under the effect-size screening (|ρ^s| 0.05), three negligibly correlated pairs are excluded, yielding π^=0.70 and αE2=0.65, a better-calibrated weight than Metric E’s αE0.797 under the same data. A large-scale simulation study with 37,440 model evaluations confirms that Metric E inverts the correct ranking at correlation levels ρ0.40 (CDR = 0%), while E2 maintains correct discrimination in 14 of 15 simulation conditions (M1 vs. M3). We also delimit the metrics’ scope: E2 degrades under near-saturated uniform dependencea regime in which the strictly proper Energy Score remains preferableand the pairwise index is sensitive to sample size, for which we provide an effect-size-based variant. An R package (mvmetrics v0.2.0) implementing both metrics, the Log-Loss variant, alternative weighting schemes, and the effect-size screening is publicly available. Full article
Show Figures

Figure 1

16 pages, 2355 KB  
Article
A Simulation-Based Evaluation of the DR-GEE Approach Based on Flexible Cluster-Size Weighting Under Hybrid Informative Cluster Size Structures
by Betül Dağoğlu Hark and Zeliha Nazan Alparslan
Stats 2026, 9(4), 74; https://doi.org/10.3390/stats9040074 - 12 Jul 2026
Viewed by 217
Abstract
This study evaluates a DR-GEE approach based on Doubly Robust Generalized Estimating Equations for marginal inference under a hybrid informative cluster size structure. Hybrid informative cluster size refers to situations where cluster size can be related to both the marginal response variable and [...] Read more.
This study evaluates a DR-GEE approach based on Doubly Robust Generalized Estimating Equations for marginal inference under a hybrid informative cluster size structure. Hybrid informative cluster size refers to situations where cluster size can be related to both the marginal response variable and the distribution of covariates associated with the response. The proposed approach aims to achieve a cluster-balanced marginal estimator designed to mitigate the excessive influence of cluster sizes. To this end, the method combines a flexible cluster-size weighting component, defined by the α adjustment parameter, with an augmentation term derived from the study’s outcome model. Thus, both the direct cluster size–outcome relationship and the imbalance in the distribution of covariates are taken into account. A comprehensive Monte Carlo simulation evaluated performance under varying informativeness levels (γ = 0.1, 0.5, 1.0) and average cluster sizes (λ = 3, 5, 8). DR-GEE was compared with standard GEE, CWGEE, WCR, and DWGEE across different tuning parameters (α = 0.25, 0.50, 1.0). The results show that DR-GEE with α = 1 generally achieved the most favorable performance under hybrid ICS conditions. For example, when γ = 1 and λ = 5, GEE exhibited substantial bias (0.538) and high RMSE (0.566), whereas DR-GEE (α = 1) markedly reduced bias (0.076) and RMSE (0.224). Unlike CWGEE and DWGEE, which address only one dimension of informativeness, DR-GEE balances both cluster size–outcome dependence and covariate information. Full article
(This article belongs to the Section Biostatistics)
Show Figures

Figure 1

20 pages, 692 KB  
Article
Bayesian Integration of Probability and Non-Probability Samples
by Qi Chen and Balgobin Nandram
Stats 2026, 9(4), 73; https://doi.org/10.3390/stats9040073 - 6 Jul 2026
Viewed by 340
Abstract
Non-probability samples (NPS) are increasingly used in practice because they are relatively inexpensive and often contain the study outcome variable. However, NPS may suffer from selection bias and usually do not provide survey weights, making finite-population inference difficult. Probability samples (PS), on the [...] Read more.
Non-probability samples (NPS) are increasingly used in practice because they are relatively inexpensive and often contain the study outcome variable. However, NPS may suffer from selection bias and usually do not provide survey weights, making finite-population inference difficult. Probability samples (PS), on the other hand, provide survey weights under known sampling designs but may not contain the study outcome variable. This paper develops a Bayesian approach for integrating probability and non-probability samples under a superpopulation model. The proposed method jointly models latent survey weights and missing probability-sample outcomes within a hierarchical Bayesian framework. The method is compared with a naive Bayesian bootstrap estimator and a weighted regression estimator. A BMI data application and simulation studies show that the proposed method reduces bias relative to the naive estimator while accounting for uncertainty in latent weights and missing outcomes. Full article
Show Figures

Figure 1

30 pages, 3622 KB  
Article
Central Limit Theorem of the Recursive Estimate of Density Function Under Randomly Censored Data
by Meraou Mohammed Amine and Rabhi Abbes
Stats 2026, 9(4), 72; https://doi.org/10.3390/stats9040072 - 3 Jul 2026
Viewed by 376
Abstract
Kernel density estimation for right-censored data has been extensively studied in the non-recursive setting, whereas recursive approaches adapted to censoring remain largely unexplored despite their considerable computational advantages in sequential data environments. In this paper, we introduce a recursive kernel density estimator for [...] Read more.
Kernel density estimation for right-censored data has been extensively studied in the non-recursive setting, whereas recursive approaches adapted to censoring remain largely unexplored despite their considerable computational advantages in sequential data environments. In this paper, we introduce a recursive kernel density estimator for independent right-censored observations through a Kaplan-Meier weighting scheme. The proposed estimator can be updated incrementally as new observations become available, avoiding repeated re-computation of the entire estimator and substantially reducing memory and computational requirements. Under mild regularity conditions, we establish the asymptotic normality of the estimator and derive its asymptotic variance, which explicitly reflects the effect of the recursive weighting mechanism and the censoring process. We also construct asymptotic confidence intervals for the underlying density using a plug-in variance estimator. An extensive Monte Carlo study, including Gaussian, exponential, heavy-tailed, multimodal, contaminated, and severely censored scenarios, demonstrates that the proposed estimator achieves estimation accuracy comparable to that of the classical censored Parzen-Rosenblatt estimator while offering substantial computational gains. In particular, the recursive procedure remains stable under high censoring levels and exhibits excellent scalability for large and sequentially collected datasets. The proposed methodology provides an efficient and theoretically justified alternative for nonparametric density estimation under right censoring and is particularly suited to applications involving streaming data, such as survival analysis, reliability engineering, medical monitoring, and online forecasting. Full article
(This article belongs to the Special Issue Nonparametric Inference: Methods and Applications)
Show Figures

Figure 1

17 pages, 2910 KB  
Article
Hybrid Regime-Switching Models for Cryptocurrency Prices: An Asset-Dependent Performance Analysis Using Markov Chains and Random Forests
by Steve Karam, Joseph El Maalouf and Nadine Dirani
Stats 2026, 9(4), 71; https://doi.org/10.3390/stats9040071 - 30 Jun 2026
Viewed by 521
Abstract
This study develops a leakage-free hybrid Markov–Random Forest framework for cryptocurrency price forecasting and evaluates it on Bitcoin and Ethereum. Daily OHLCV features are lagged by one trading day to prevent look-ahead bias, while regime labels are assigned from observed price changes using [...] Read more.
This study develops a leakage-free hybrid Markov–Random Forest framework for cryptocurrency price forecasting and evaluates it on Bitcoin and Ethereum. Daily OHLCV features are lagged by one trading day to prevent look-ahead bias, while regime labels are assigned from observed price changes using a two-state Markov chain with increasing and decreasing states. Regime-specific Random Forest models are then tuned independently via time-series cross-validation, allowing the predictive structure to adapt to regime-specific market conditions. The empirical results exhibit clear asset dependence. For Ethereum, the hybrid model outperforms the standalone Random Forest on magnitude-based metrics, attaining lower MAE and RMSE while also delivering a modest improvement in directional accuracy. Regime-specific tuning further identifies distinct optimal hyperparameter configurations across the increasing and decreasing states, suggesting that Ethereum’s upward and downward dynamics are structurally heterogeneous and can be better captured through regime-aware learning. By contrast, for Bitcoin, the standalone Random Forest delivers superior magnitude forecasting performance, while the regime-specific models differ only in tree depth and share the remaining tuning parameters, indicating that regime conditioning adds limited incremental value in a more persistent market. Statistical tests reinforce these findings. For Ethereum, Diebold–Mariano tests show that the hybrid significantly outperforms the standalone Random Forest under squared loss, while the absolute-loss comparison is only marginal. Across both assets, directional accuracy remains close to random chance, confirming the limited predictability of next-day price direction from lagged OHLCV features. Overall, the hybrid framework is most valuable when regime-specific dynamics are sufficiently distinct, offering improved forecasting performance and greater interpretability than a single global model. Full article
(This article belongs to the Topic Statistics and Data Science)
Show Figures

Figure 1

12 pages, 1463 KB  
Article
Modeling Exposure Mixtures and Spatiotemporal Dependence in Count Data Using Bayesian Kernel Machine Regression
by Ning Sun, Zoran Bursac and Boubakari Ibrahimou
Stats 2026, 9(4), 70; https://doi.org/10.3390/stats9040070 - 26 Jun 2026
Viewed by 312
Abstract
We propose a Bayesian kernel machine regression (BKMR) framework for count outcomes with dynamic spatiotemporal dependence. The proposed model, termed Negative Binomial BKMR with spatiotemporal effects (NB-BKMR), integrates (i) a negative binomial likelihood to accommodate overdispersion, (ii) a kernel-based exposure–response surface for complex [...] Read more.
We propose a Bayesian kernel machine regression (BKMR) framework for count outcomes with dynamic spatiotemporal dependence. The proposed model, termed Negative Binomial BKMR with spatiotemporal effects (NB-BKMR), integrates (i) a negative binomial likelihood to accommodate overdispersion, (ii) a kernel-based exposure–response surface for complex mixtures, (iii) hierarchical group-wise variable selection and (iv) a dynamic spatiotemporal random effect structure based on a Leroux conditional autoregressive (CAR) prior evolving over time. Posterior inference is conducted in a fully Bayesian framework using Polya-Gamma data augmentation. Through simulation studies, under varying nonlinear exposure–response functions, correlation structures, and spatiotemporal dependence patterns, we show that NB-BKMR yields well-calibrated uncertainty quantification and robust identification of dominant mixture drivers, even when exposures are highly correlated. An application to the U.S. state-level traffic fatality counts (1982–1988) illustrates how the model uncovers nonlinear effects and interactions among socioeconomic and behavioral predictors while improving predictive performance relative to generalized additive models with spatiotemporal smooths. This work extends existing BKMR methodology by unifying mixture modeling, count outcomes, and dynamic spatial dependence in a single coherent framework, with particular relevance for areal public health surveillance data. Full article
Show Figures

Figure 1

17 pages, 280 KB  
Article
Statistics of Non-Conserved Observables in Lindblad Master Equations
by Giovanni Modanese
Stats 2026, 9(4), 69; https://doi.org/10.3390/stats9040069 - 25 Jun 2026
Viewed by 268
Abstract
We study the dynamics of observables that are conserved under the Hamiltonian evolution of a closed quantum system, but cease to be conserved when the system is coupled to a Markovian environment and described by a Lindblad master equation. Starting from the adjoint [...] Read more.
We study the dynamics of observables that are conserved under the Hamiltonian evolution of a closed quantum system, but cease to be conserved when the system is coupled to a Markovian environment and described by a Lindblad master equation. Starting from the adjoint Lindblad equation, we derive elementary expressions for the time derivatives of the expectation value and second moment of an observable O, with particular emphasis on the case [H,O]=0 but L(O)0. These formulae provide a direct assessment of how collapse operators break Hamiltonian conservation laws and generate fluctuations of formerly conserved quantities. The discussion is illustrated by analytic examples: one-qubit amplitude damping, a two-qubit excitation-number model, a momentum-diffusion model in which the mean is conserved while the variance grows, and the Jaynes–Cummings model. The latter also shows the complementary case of a reservoir coupled through a conserved quantity, where dephasing can occur without changing the statistics of that quantity. We finally comment on the relation between Lindblad source terms and idealized wave-function reduction models in which local conservation may hold only statistically. Full article
Show Figures

Figure 1

15 pages, 298 KB  
Article
Pertinent Prediction Intervals in Linear Regression
by Dimitris N. Politis
Stats 2026, 9(4), 68; https://doi.org/10.3390/stats9040068 - 25 Jun 2026
Viewed by 334
Abstract
In linear regression, a point predictor Y^f of a future response Yf associated with a regressor value of interest x̲f can easily be constructed. Since Y^f will always incur a prediction error, it is desirable to [...] Read more.
In linear regression, a point predictor Y^f of a future response Yf associated with a regressor value of interest x̲f can easily be constructed. Since Y^f will always incur a prediction error, it is desirable to accompany the point predictor by a prediction interval, say C(x̲f), that will contain the target Yf with a pre-specified high probability, e.g., 90%. An estimated prediction interval, say C^(x̲f), is called pertinent if its construction incorporates the variability of all estimators that are employed in the prediction problem. So far, pertinent prediction intervals have only been constructed via some form of bootstrap. However, resampling can be quite computationally expensive since the estimation/prediction problem has to be re-calculated on a large number of pseudo-scatterplots, each having the same sample size as the original one. The paper at hand proposes a short-cut that directly employs the asymptotic normal distribution of relevant estimators—as opposed to a bootstrap histogram—in order to capture their variability. The resulting prediction interval achieves pertinence without full-scale resampling, thus offering computational savings of orders of magnitude. Full article
Show Figures

Figure 1

9 pages, 279 KB  
Article
Changes in Variance and the Detection of Trends
by Markus Neuhäuser
Stats 2026, 9(4), 67; https://doi.org/10.3390/stats9040067 - 24 Jun 2026
Viewed by 412
Abstract
Background: Tests for a trend in location are appropriate when there is an ordered alternative such as, for example, when it is assumed that the effect does not decrease with increasing doses of a drug or fertilizer. Classical trend tests for normally distributed [...] Read more.
Background: Tests for a trend in location are appropriate when there is an ordered alternative such as, for example, when it is assumed that the effect does not decrease with increasing doses of a drug or fertilizer. Classical trend tests for normally distributed data as well as the nonparametric Jonckheere trend test can have inflated type I error rates when variances differ between groups. Here, different approaches suggested to handle heterogeneous variances are investigated in combination with the Williams trend test. Methods: A simulation study was performed to compare the Jonckheere trend test with competing tests. The different tests were investigated for normal and non-normal data and also applied to a data set on sizes of walnuts opened by birds in various stages of a winter. Results: With one exception, all investigated trend tests can have an inflated type I error rate when variances differ. Only a nonparametric multiple contrast test based on relative effects showed an acceptable type I error rate in all scenarios considered in the simulation. Conclusions: The Williams trend test in combination with the nonparametric multiple contrast test based on relative effects can be suggested for routine use. With this procedure, an increase in variance cannot cause a significant result in the test for trend. Full article
(This article belongs to the Section Biostatistics)
Show Figures

Figure 1

Previous Issue
Back to TopTop