Next Article in Journal
A Copula-Based Framework for Multivariate Count Time Series with Mixed Marginal Distributions
Previous Article in Journal
BLIDE: Bayesian Learning of Infectious Disease Emerging in COVID-19 Studies
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Predictive Power Analysis of Multiple Test Procedures Under Arbitrary Dependence

by
George Karabatsos
Departments of Mathematics, Statistics, and Computer Sciences, and Educational Statistics, University of Illinois Chicago, 1040 W. Harrison St. (M/C 147), Chicago, IL 60607, USA
Stats 2026, 9(3), 56; https://doi.org/10.3390/stats9030056
Submission received: 1 April 2026 / Revised: 19 May 2026 / Accepted: 21 May 2026 / Published: 29 May 2026
(This article belongs to the Section Statistical Theory and Methods)

Abstract

Many statistical problems can be addressed by applying a multiple testing procedure (MTP) that controls either the Family-Wise Error Rate (FWER) or False Discovery Rate (FDR) under unknown arbitrarily interdependent p-values, without explicitly modeling these inter-correlations. They include the FWER-controlling Bonferroni MTP and Holm MTP; the FDR-controlling Benjamini-Yekutieli MTP; and the DP-MTP, based on a Dirichlet process (DP) prior distribution supporting the entire space of MTPs that control either the FWER or the FDR. For such an MTP, this study introduces a new and congenial method for Bayesian predictive power analysis, for power calculation and sample size determination for any given planned future (e.g., replication or interim) study. This novel MTP predictive power analysis method is based on a joint prior distribution defining a scale matrix mixture of asymmetric multivariate normal mean-variance mixture distributions, factorized as a general prior distribution for effect sizes (e.g., obtained from expert judgment or results of prior studies), and a uniform prior distribution for correlation matrices representing arbitrary dependencies between p-values of test statistics of given multiple hypothesis tests under their alternative hypotheses. The new MTP power analysis method also results in p-value weights which can be used to minimize the relative impacts of and assess for significance-chasing biases (e.g., publication bias, p-hacking, etc.) in multiple testing, without needing to assume that p-values (effect sizes) are independent. Previous MTP power analysis methods are conditional in that they assume fixed effect sizes and a fixed correlation matrix for test statistics (p-values) which may be difficult to specify; while, incongenially, not fully accounting for their uncertainty, especially for MTPs that control either the FWER or FDR for arbitrarily correlated p-values. The new simulation-based MTP predictive power analysis method is illustrated through the analysis of p-values obtained by a famous study of lead exposure and re-analyzed by the previous MTP literature, using R package bnpMTP (v1.0).

1. Introduction

Multiple hypothesis testing problems continue to arise in many fields of the natural (life, physical), social (e.g., psychology, education), and the formal sciences, especially the bio- and statistical sciences where various hypothesis testing and multiple testing procedures (MTPs) have been developed over decades.
A hypothesis testing procedure is based on a test statistic, typically defined as a ratio of an effect size to its standard error (for example, refs. [1,2]). Examples of effect size include a raw or standardized mean difference; correlation or regression coefficient; and (log-)odds ratio; etc.; each of which can be converted from one metric to another [3]. Under the given null hypothesis ( H 0 ) tested (against the corresponding alternative hypothesis H 1 , respectively), such a test statistic follows a central-zero-mean (non-central, resp.) t-distribution or normal distribution (e.g., asymptotically for large samples) under the null hypothesis tested, perhaps after transforming another test statistic that follows a central (non-central, resp.) F, χ ν 2 , or other standard distribution under H 0 ( H 1 , resp.) [4]. A p-value is a standardized deterministic transformation of the given test statistic onto the [ 0 , 1 ] interval ([5], p. 17), such that the smaller the p, the more decisively the given null hypothesis H 0 is rejected ([6], p. 31). The p-value ( p ( x ) ), obtained from the test statistic computed from the given sample data set x x n of size n, is declared as “significant”, i.e., rejecting the given null H 0 , if  p α , based on a small specified fixed Type I error probability ( α ) of incorrectly rejecting H 0 over imaginary datasets x n * randomly sampled from a population data-generating process where H 0 is true. Meanwhile, p-values provide a common “bottom line” language for the communication of statistical results that are still very often communicated by scientific journals, despite concerns [7]. The p-value is basic to statistical inference in that it can be used to compute the upper bound of the Bayes factor (for example, refs. [8,9]) and hence can provide approximate Bayesian computation.
The conditional statistical power ( 1 β ¯ ) of a null hypothesis testing procedure is the probability that it will detect an effect size ( θ ) to reject the given null H 0 , given specified θ , n, α , and the Type II error probability ( β ¯ ) of incorrectly not rejecting H 0 over imaginary datasets x n * randomly sampled from a population data-generating process where H 1 is true. The conditional power of virtually any classical χ ν 2 -, F, z-, or t-test procedure can be calculated by a closed-form formula (for example, ref. [1]), which can be easily rearranged to solve for any of the four quantities (power, n, θ , and  α ) given the other three. Power analysis is used to calculate the power ( 1 β ¯ ) of a test procedure, or the minimum data sample size needed to achieve some minimum desired power (e.g., 0.80 ) for a planned future (e.g., replication or interim) study, given anticipated or estimated θ and α ; and less often, to evaluate the sensitivity of studies, or to make decisions about criteria of statistical significance [10]. However, conditional power analysis assumes a specified fixed effect size ( θ ) without accounting for its uncertainty. This issue can be addressed theoretically by predictive power analysis ([11,12,13], Section 6.5), which averages the conditional power over a specified prior distribution representing available knowledge for the effect size (for example, refs. [14,15,16,17]).
Widely available software packages now easily enable common statistical analyses of data to routinely output results of typically correlated p-values from multiple hypothesis tests (often, on many variables), either as the main data analysis goal, and/or as automatic by-products of other main statistical modeling objectives. Also, an MTP for marginal p-values can reduce the results of different (e.g., t, F, χ 2 , z, Wilcoxon, and/or log-rank, etc.) test statistics to a common interpretable p-value scale, while conveniently, not requiring the statistician to assume nor to specify and estimate an explicit model for the potentially complex joint distributions of the test statistics, having typically unknown correlations. All things considered, it is no wonder why MTPs for marginal p-values are popular and relevant in applied statistics ([18], and references therein), which are easily usable and computable, while not requiring direct access to the original data, so that such an MTP can be readily used for meta-analysis.
Given that there are very many MTPs, and that p-values in applied statistics are typically dependent (correlated), we henceforth focus on the subset of MTPs for marginal p-values, where each MTP provides either Family-Wise Error Rate (FWER) or False Discovery Rate (FDR) control under unknown arbitrary dependencies between p-values, even when the given observed p-values arise from a highly heterogeneous mix of different hypothesis testing procedures, and without requiring direct access to the original data. They include the un/weighted: Bonferroni [19] MTP (B-MTP) and Holm [20] MTP (H-MTP), each of which control the FWER; the Benjamini & Yekutieli [21] MTP (BY-MTP), which controls the FDR; and the Dirichlet process MTP (DP-MTP) [22], defined by a DP prior distribution [23] which supports the entire space of MTPs controlling either the FWER or the FDR under arbitrary dependencies between p-values. A weighted version of the B-MTP, H-MTP, BY-MTP, or DP-MTP is defined by importance weights assigned to respective p-values and null hypotheses [18,22,24,25]. All these MTPs are reviewed in Section 2.
MTP power analysis methods have a shorter history and focus on conditional instead of predictive power. For any MTP for testing a given set of m 1 hypothesis with FWER or FDR control, power can be defined by any one of the following four ways, with respect to the m-variate distribution of their test statistics under the alternative hypotheses (null hypothesis, resp.) for all m tests (resp.) with m × m correlation matrix. The marginal (individual) power of each test of hypotheses H 0 , j (for j = 1 , , m ) is the marginal probability that the corresponding p-value leads to a rejection under level α by the MTP, with respect to the marginal distribution of the corresponding jth test statistic under the alternative hypotheses [26,27]. The average power is the un/weighted average over these m marginal probabilities (resp.) under the alternative hypotheses (weights sum to m over all m hypotheses tested) [24,28,29,30]; which relates to the criterion of expected number of rejected hypotheses [31]. The disjunctive ( d ( = r ) = 1 /minimal/any-pair) power is the probability that at least one false null hypotheses is rejected at level α by the given MTP [28,32,33,34,35]. Finally, the conjunctive (complete/all-pairs) power is the probability that all false null hypotheses are rejected at level α by the MTP [26,28,34].
For modern applied statistics which routinely outputs multiple p-values, MTP power analysis is more relevant than power analysis of a single test procedure. A power analysis can be used to calculate the given MTP’s power, or the data sample size required to achieve some minimum desired power for the MTP for a planned future (e.g., replication or interim) study, and for the other mentioned uses of power analysis for a single test. For some FWER-controlling MTPs applied to up to three hypotheses, explicit formulas are available to compute (conditional) conjunctive power and disjunctive power [32]. Simulation methods can be used to calculate the power(s) for any FWER- or FDR-controlling MTP (for example, ref. [26]).
For any MTP that controls either the FWER or FDR for m given hypothesis tests, optimal weights for respective p-values (hypothesis tests) can be determined through maximization of a defined power function(s) of interest (e.g., disjunctive, conjunctive, or marginal power, etc.) based on the specified marginal power for each test ([36,37], and references therein). In principle, an MTP that accounts for arbitrary dependence between p-values, while incorporating optimized weights for the respective p-values, provides an MTP analysis that relatively down-weights the impacts of any present significance-chasing biases (due to publication bias and/or p-hacking, etc.) while accounting for any heterogeneous effects, and without needing to assume independence among p-values (effect sizes) as done in traditional meta-analysis [38]. This is because the significance-chasing bias level of any null hypothesis test’s binary-valued result (1 = significant, p α ; or 0 = non-significant, p > α ) can be measured by the binary result’s distance from the test’s marginal power probability, given the estimated or anticipated effect size of the test, as done by the χ 2 test of excess significance [39], while other tests and models of publication bias can have limited power and be difficult to use [39,40,41]. As an aside, while the test of excess significance is based on a χ 2 type distance, in principle, any other distance measure from the general class of power divergence statistics [42] can be used, such as the Hellinger distance. While the χ 2 test of excess significance aims to detect the presence of significance-chasing biases instead of correcting for them, one key insight of this established bias testing procedure is that the marginal power of each of the m hypothesis tests of a given multiple hypothesis testing application provide standards by which to address significance-chasing biases. To elaborate, a significant result of a null hypothesis test with low marginal power is more likely to be a result of significance-chasing biases, compared to a significant result of a null hypothesis test with high marginal power which is more likely to be a result of a true effect or signal and less likely to be a result of bias. It is less likely because such a high powered test already makes it easier for the data analyst to detect a true effect or signal and thus lowers the need for the analyst to engage in significance-chasing bias behavior. Therefore, assigning weights to p-values as increasing functions of their respective marginal powers would produce results of multiple hypothesis testing that would place greater relative weight to p-values that are less susceptible to significance-chasing biases.
For a given weighted multiple testing problem, optimal p-value weights can be found using any one of the available optimization algorithms that maximize any of the mentioned criteria of MTP power. See the Appendix A for a more detailed review (based on [37]). However, while each of the available optimization methods are based on a clear objective, they are not fully satisfactory, as they are limited to Bonferroni, graphical or chain MTPs; while numerically optimizing disjunctive or conjunctive power can be computationally costly when the number of hypothesis tests (m) is large, it can yield non-unique p-value weights under equal (specified) marginal powers and correlated test statistics; it can be difficult to determine or interpret (e.g., numerically optimizing for weighted or average power is easier, but involves introducing another set of weights for power that may be difficult to determine or interpret); and it employs only conditional power, in that each method assumes that both the effect sizes and the correlation matrix of the m-variate distribution of the test statistics are fixed and known.
More ideally, a predictive MTP power analysis can be used, which would assign a prior distribution on both of these multivariate quantities in order to account for their uncertainty. Indeed, for specialized multiple testing problems employing m hypothesis tests, the  m × m correlation matrix of the test statistics can be directly derived or calculated, either from uncorrelated test statistics (p-values) or orthogonal designs or contrasts; permutation- or bootstrap-based estimates of test statistics; or for pairwise mean comparisons done via a Dunnett or Tukey–Kramer type MTP (see the Appendix A for a more detailed review). However, more broadly speaking, in the modern data analysis era where MTP inferences need to be made from multiple p-values which often result from applications of highly heterogeneous combinations of different hypothesis testing procedures (e.g., not only for group mean comparisons), it can be very challenging to derive explicit equations or a sharp prior distribution about the correlations among corresponding test statistics, especially when the number (m) of hypothesis tests is large. Moreover, as mentioned, this paper focuses on MTPs for marginal p-values, which each accounts for arbitrary dependencies between p-values while not needing to model or estimate these correlations directly from data.
Further, while each of the available MTP power analysis methods provides only conditional power analysis, i.e., based on fixed effect sizes and a fixed specified correlation matrix, such a method is uncongenial with any MTP (e.g., B-MTP, H-MTP, or DP-MTP) that controls either the FWER or FDR under arbitrary dependence (correlations) between the p-values (test statistics). This is because such an MTP accounts for all possible correlation matrices of the test statistics, not only one. While it is tempting to view equation- and optimization-based methods for computing MTP power as more expediently convenient and precise compared to simulation methods, such an optimization method typically needs to assume a fixed correlation matrix in order to be computationally tractable. A simulation-based MTP power analysis can flexibly model more complex random correlation structures beyond a fixed correlation matrix.
Based on all the above considerations and the current state of the related MTP power analysis literature, we propose the first MTP predictive power analysis method, a straightforward simulation-based method of predictive power analysis for any MTP that provides FWER or FDR control under arbitrary-dependent p-values. This new MTP power analysis method, for any m given hypothesis tests, is based on a joint prior distribution for the parameters of a newly developed m-variate (upper) non-central t distribution, defined by an m × 1 location vector, m × m scale matrix (mainly determined by an underlying correlation matrix), and a vector of m degrees of freedom parameters, for the distribution of the m 1 test statistics under their alternative hypotheses. The first part of this joint prior is defined by a prior distribution for the location vector representing the respective ratios of m effect sizes to their respective standard errors, which can represent either: prior subjective clinical judgment information; be centered on the estimated effect sizes from a previous study; or more generally by a power prior, defined as the posterior distribution of the effect sizes from prior historical data [43] for, e.g., a future replication (or interim) study. The second part of the joint prior is defined by a uniform prior distribution for the underlying m × m correlation matrix for the m test statistics under their respective alternative hypotheses, which can fully account for uncertainty and typical lack of prior information about this matrix, and ensure a congenial power analysis of the given MTP that controls FWER or FDR under arbitrarily dependent p-values (test statistics). Therefore, the new MTP power analysis method is based on a scale matrix mixture of asymmetric multivariate normal mean-variance mixtures of the test statistics under the alternative hypotheses of all the given multiple hypothesis tests (resp.). Indeed, this is the first paper that provides a method of MTP power analysis for arbitrary-correlated p-values, which can rapidly compute power results using an efficient algorithm [44] for sampling the uniform distribution of correlation matrices.
In Bayesian statistical inference, the prior distribution must reflect the current state of knowledge and prior information about model parameters available to the data analyst. When there is a lack of prior information about model parameters, then Bayesian theory implies that a coherent choice of prior distribution is an “objective” non-informative prior distribution for the correlation matrix. Meanwhile, most priors used in statistical applications are objective instead of subjective, in part, because subjective elicitation is too difficult to be done even in a limited way for more than a few unknown parameters in the given statistical problem, so that such unknown parameters must be handled via objective Bayesian methods ([45,46], p. 216). Correspondingly, for the current context of predictive MTP power analysis, the “objective” uniform prior distribution on the correlation matrices often honestly reflects the statistician’s available prior knowledge in many MTP statistical situations. This is because of the fact that in such common situations, the MTP statistician only has access to the p-values of the m hypothesis tests, without any access to the raw data used to compute these p-values, as in traditional meta-analysis. This thus leaves the statistician without the ability to estimate this dependence structure and to elicit a more informative prior distribution, and renders the ”objective” uniform prior over all valid correlation matrices as a reasonable choice prior in such a situation because it would honestly represent the available prior information. Indeed, such common MTP data analysis scenarios are in large part what makes MTPs that control either the FWER or FDR under arbitrarily dependent p-values quite relevant and useful ([21], for example), including the B-MTP, H-MTP, BY-MTP, and DP-MTP. While it is arguable that the uniform prior may not meaningfully capture the more likely dependence structures among test statistics, such as equi-correlated or AR(1) correlation structures, more informative priors can be difficult to determine in more common MTP scenarios. The uniform prior is not “wrong” because it assigns non-zero probability to all arbitrary valid correlations structures for test statistics supported by the B-MTP, H-MTP, BY-MTP, and DP-MTP.
This new MTP predictive power analysis method: (1) determines estimates of the marginal powers of the m given hypothesis tests (resp.), which then can either be used for: (2) sample size determination for each of these tests given desired minimum marginal powers; (3)–(4) the other two mentioned purposes of power analysis of a single testing procedure; (5) constructing normalized weights of the corresponding m p-values (resp.) while optimizing the power of the MTP, and down-weighting the impact of significance-chasing biases in the given MTP analysis; and/or (6) comparing with observed p-values to provide an assessment of such biases for each individual p-value using Hellinger distance, in a similar spirit of the original seminal χ 2 test of excess significance (but here, not using a measure of bias over all m hypothesis tests or ”studies” based on χ 2 type distance). It is well-known in the MTP field that a weighted MTP, which gives relatively higher weights to the subset of tested null hypotheses that are likely to be false, tends to have higher statistical power compared to unweighted multiple testing procedures, for various correlation structures according to extensive simulation studies (e.g., ref. [47]). Further, the MTP predictive power analysis method achieves all these six objectives without needing to assume independence of the m p-values, if the power-analyzed MTP controls for the FWER or FDR under arbitrarily dependent p-values.
Next, Section 2 reviews MTPs that each strongly control FWER and/or FDR under arbitrarily dependent p-values, in order to contextualize the description of the new predictive MTP power analysis method for such an MTP in Section 3. Then, Section 4 illustrates the new predictive MTP power analysis method to evaluate the power of the DP-MTP, and of the B-MTP and H-MTP for comparison purposes, through the MTP analysis of p-values arising from 41 hypothesis tests reported by a famous study on the neuropsychologic effects of unidentified childhood exposure to lead. Section 5 proves a theorem on the convergence and convergence speed of the algorithm. Section 6 concludes by summarizing the main ideas of the paper.
Central to the new predictive MTP power analysis method is Algorithm 1, presented in Section 3, which can be used to simulate draws of samples from the prior predictive distribution of the joint prior mentioned above, in order to yield estimates of the various measures of powers of the m test statistics, respectively, and the associated statistics mentioned above. This includes the corresponding m marginal powers, being well-defined concepts in the MTP literature [26,27] that can be used for estimating the Hellinger distance measures of significant-chasing bias, and for constructing p-value weights that can be used in a subsequent weighted MTP analysis. Meanwhile, the univariate marginal distributions of a prior predictive distribution are unique given a specifically defined prior and sampling distribution (likelihood), implying that the corresponding marginal powers and associated functions of them are uniquely defined, including the associated Hellinger distance bias measures and p-value weights, all while accounting for the uncertainty for the parameters underlying MTP power analysis by assigning them a joint prior distribution which averages conditional power over this prior. As mentioned earlier, this is unlike some of the previous approaches to MTP power analysis, all based on conditional power, which can yield non-unique p-value weights. While Algorithm 1 by default is based on the uniform prior on all valid correlation matrices for the m given test statistics, the algorithm can be easily modified to implement more informative prior, such as priors supporting equi-correlated or AR(1) correlation structures.

2. Review of MTPs Valid Under Arbitrary Dependence

Here, a formal multiple hypothesis testing framework [48] is reviewed, useful for MTP inference [22]. Consider the probability space, ( X , X , P ) , with probability function P a member of a set P of distributions representing a parametric or non-parametric model. A null hypothesis, H 0 , is a subset (submodel), H 0 P , of distributions on ( X , X ) , where P H 0 denotes that P satisfies H 0 , and  X is the sample space of any observable dataset sampled from the distribution P, with  X the corresponding sigma algebra of events. Any application of multiple testing, performed on a sample dataset x P P , aims to determine whether P satisfies distinct null hypotheses, belonging to a certain set (family) H of candidate null hypotheses, with this set being typically countable, or even uncountable.
For any given sampled dataset, x P P , a multiple testing procedure (MTP) is a decision function R : x X R ( x ) H that returns the subset of rejected null hypotheses; R c ( x ) = H R ( x ) is the subset of non-rejected hypotheses; and the MTP R ( x ) commits a Type I error if H 0 R ( x ) H 0 ( P ) ; and commits a Type II error if H 0 R ( x ) H 1 ( P ) , where H 1 ( P ) is the subset of the alternative hypotheses that are true under the given true data-generating distribution P. In a typical multiple testing application, H is a finite or countable set of m null hypotheses, H = { H 0 , 1 , , H 0 , m } , tested respectively against given alternative hypotheses { H 1 , j } j = 1 m , with m = | H | Z + the number of candidate null hypotheses, m 0 ( P ) = | H 0 ( P ) | m true null hypotheses, m 1 ( P ) = | H 1 ( P ) | = m m 0 m (truly) false null hypotheses, and  π 0 = m 0 / m , under any P P . For simplicity and with no loss of generality, it is henceforth assumed that H is countable, while the subsequently described hypothesis testing-related ideas can easily be extended to continuous hypothesis testing after adopting some more complex notation.
A typical MTP is a function R ( p ) of a family of p-values, p = ( p j , H 0 , j H ) , where each p j -value measures how probable are the observed data x, given that the null hypothesis H 0 , j is true ([5], p. 19, Definition 2.1). Assume that for each null hypothesis H 0 , j H there exists a p-value function, p j : X [ 0 , 1 ] , having a marginally super-uniform probability distribution ( P ) when H 0 , j is true, i.e.,
P X P [ p j ( X ) t ] t , for P P , H 0 , j H 0 ( P ) , and t [ 0 , 1 ] ,
with strict equality ( P X P [ p H ( X ) t ] = t for all t [ 0 , 1 ] ) when the p j -value is calibrated, i.e., uniform U [ 0 , 1 ] distributed under the null hypothesis H 0 , j , which can be achieved (or improved on) using any of the available p-value calibration methods, e.g., for a test of a discrete model or a composite null hypothesis; or for model checking ([5,49,50,51,52], and references therein).
Thus, the assumption (1) for each p-value does not require a calibrated p-value for MTPs considered in the paper, since the U [ 0 , 1 ] distribution is a special case of the super-uniform distribution, even though a calibrated p-value would allow for a universal interpretation of the p-value, e.g., based on Fisher’s scale of evidence for p-values ([6], p. 31). The p-values analyzed in the case study illustration in Section 4 are calibrated asymptotically. Indeed, it is possible to define p-values only via the super-uniformity property (1) ([53], Definition 8.3.26), which does not make specific requirements on the distribution of the underlying test statistic under the null hypothesis. Though, the predictive MTP power analysis method introduced in Section 3 requires the specification of an explicit functional relationship between each p-value and its test statistic, between the joint distribution of test statistics and corresponding p-values for any multiple testing problem applying m hypothesis tests. Then, a p-value function can be written as p ( x ) p { t ( x ) } from a test statistic t t ( x ) of any random dataset x.
When using any MTP R to test a set of hypotheses, H , a traditional criterion for Type I error control is the FWER, the probability ( P ) of making at least one false discovery [54]:
FWER P ( R ) = P X P [ Reject any true hypothesis , H 0 , j H 0 ( P ) H ] ;
while an alternative criterion is the FDR [55], the expected ( E ) proportion of false rejections of null hypotheses out of the total number of rejected hypotheses:
FDR P ( R ) = E X P [ FDP P ( R ( X ) ) ] = E X P | R ( X ) H 0 ( P ) | max { | R ( X ) | , 1 } .
The practice of multiple hypothesis testing aims to maximize the expected number of rejections while controlling the FWER or FDR at a preset level α ( 0 , 1 ) , typically α = 0.05 or 0.01 , etc. An MTP strongly controls the FWER (FDR, resp.) if FWER P ( R ) α (if FDR P ( R ) π 0 α α , resp.) for any α ( 0 , 1 ) and all P P (all π 0 [ 0 , 1 ] , resp.), while FDR FWER , with  FDR = FWER if all null hypotheses are true, and thus FWER α implies FDR α , i.e., FDR control is more liberal [18].
Many MTPs can each be characterized as a step-up MTP, R S U Δ α , which for the order statistics p ( 1 ) p ( m ) of m p-values of the given multiple testing problem, specifies a non-decreasing sequence of thresholds 0 Δ α ( H 0 , ( 1 ) ) Δ α ( H 0 , ( m ) ) 1 for the respective m ordered tested null hypotheses, and then rejects the null hypotheses having the R α ( x ) smallest p-values, with
R α ( x ) = max r { 0 , 1 , , m } { r : p ( r ) ( x ) Δ α ( H 0 , ( r ) ) } .
where p ( 0 ) 0 , and given a specified threshold function:
Δ α , ν ( H 0 . ( r ) ) α π ( H 0 , ( r ) ) β ν ( r ) ,
with weight function π : H [ 0 , 1 ] and β ν : R + R + being a shape (reshaping) function. MTPs that strongly control FWER under arbitrary dependencies between p-values, include: the (unweighted) Bonferroni [19] MTP (B-MTP), defined by thresholds Δ α , ν ( H 0 , ( r ) ) α β ν ( r ) / m = α / m with shapes β ν ( r ) 1 and weights π ( H 0 , ( r ) ) w ( r ) = 1 / m ; the Holm [20] step-down MTP (H-MTP), defined by thresholds Δ α , ν ( H 0 , ( r ) ) α m r + 1 . The weighted B-MTP is defined by thresholds Δ α , ν ( H 0 , ( r ) ) = α π ( H 0 , ( r ) ) β ν ( r ) = α w ( r ) using weights π ( H j ) w j [ 0 , 1 ] assigned to respective null hypotheses H 0 , j H such that i = 1 m w j = 1 , where π ( H j ) is an arbitrary probability distribution on H defining the relative importances of the m null hypotheses tested [18]. The weighted H-MTP is defined by thresholds Δ α , ν ( H 0 , ( r ) ) w ( r ) k = r m w ( k ) α with weights w j 0 , j = 1 , , m .
For any probability measure ν on ( 0 , ) , the step-up procedure (4) and (5) based on the shape function
β ν ( r ) 0 r x d ν ( x )
strongly controls FDR α π 0 α under arbitrary dependencies between p-values, where π : H [ 0 , 1 ] is a probability mass function with respect to a counting measure Λ on H and π 0 = H 0 , j H 0 Λ ( { H 0 , j } ) π ( H 0 , j ) . Using different weights π ( H 0 , j ) (weights Λ ( { H 0 , j } ) , resp.) over countable H gives rise to weighted p-values (weighted FDR; Benjamini & Hochberg [24], resp.). For any finite set of m null hypotheses H and p-values, the (BY; Benjamini & Yekutieli [21]) distribution-free step-up MTP (i.e., BY-MTP) is defined by probability measure ν ( { k } ) = ( k j = 1 m 1 j ) 1 with support in { 1 , , m } and threshold function Δ α , ν ( H 0 , ( r ) , r ) = α β ν ( r ) / m , based on linear shape function β ν ( r ) = k = 1 r k ν ( { k } ) = r / ( j = 1 m 1 j ) and p-value weights π ( H 0 , j ) w j = 1 / m for j = 1 , , m ([48], Lemma 3.2, p. 976). The weighted BY-MTP is instead defined by an arbitrary probability distribution π ( H 0 , j ) on H .
Each of the above MTPs provides conservatively control the FWER or FDR, especially when the number of hypothesis tests (m) is large, while other choices of ν can sometimes improve the power of the corresponding MTP ([48], Section 4.2). With this in mind, the DP-MTP method was developed [22]: this MTP is defined by a Bayesian nonparametric, Dirichlet process (DP) [23] prior distribution that supports the entire space of random probability measures (r.p.m.s), ν , which in turn, for an observed set of p values { p j } j = 1 m , induces a DP prior distribution of random MTP thresholds Δ α , ν ( H 0 , ( r ) ) α π ( H 0 , ( r ) ) β ν ( r ) via (4)–(6), where each random threshold Δ α , ν (and DP random ν ), and thus the DP-MTP, strongly controls either the FWER or FDR under arbitrary dependence among p-values. Thus, the DP-MTP procedure naturally accounts for the uncertainty in the selection of MTPs and their respective decisions regarding which numbers of the smallest p-values are significant discoveries, from any set of hypotheses tested. Further, DP-MTP also measures each p-value’s probability of significance relative to the DP prior predictive distribution of this space of all MTPs.
Specifically, for any probability space ( X , A , G ) , an r.p.m. ν follows a Dirichlet process (DP) prior with baseline probability measure ν 0 and mass parameter M, denoted ν DP ( M ν 0 ) , if 
( ν ( B 1 ) , , ν ( B m ) ) Dirichlet m ( M ν 0 ( B 1 ) , , M ν 0 ( B m ) )
for any (pairwise-disjoint) partition B 1 , , B m of the sample space X , with expectation E [ ν ( · ) ] = ν 0 ( · ) , variance V [ ν ( · ) ] = ν 0 ( · ) [ 1 ν 0 ( · ) ] M + 1 , and almost-sure support of the space of discrete r.p.m.s ν [23]. The DP-MTP is based on the DP prior (7) centered on baseline measure ν 0 ( r 1 , r ] ( r j = 1 m 1 j ) 1 (for r = 0 , 1 , , m ) chosen to match in prior expectation (given M) the probability measure ν defining the BY MTP; and based on an exponential hyper-prior distribution of the DP mass parameter M in (7), given by
M Exponential ( μ ) ,
and a default hyper-prior is chosen by μ 1 . Karabatsos [22] provides further discussions of the choices and characterizations of the DP prior for DP-MTP.
With respect to the joint hierarchical DP prior distribution (7) and (8) with BY-MTP baseline ν 0 , the DP-MTP, for each p-value from a given set of m ordered p-values, { p ( r ) } r = 1 m , counts the proportion of times the p-value is significant, represented by the following vector (denoted [ · ] ) of prior predictive probabilities:
[ Pr { 1 ( r R α , ν ( x ) ; p ( r ) ) = 1 } : r = 0 , 1 , , m ] = [ 1 ( r R α , ν ( x ) ; p ( r ) ) : r = 0 , 1 , , m ] D m ( d ν 1 , , d ν m M ν 0 ) ,
where 1 ( · ) is the (1 or 0 valued) indicator function, D m ( ν 1 , , ν m M ν 0 ) (with ν r B j = ν ( r 1 , r ] for r = 1 , , m ) denotes the cumulative distribution function (CDF) of the Dirichlet distribution (7), and
R α , ν ( x ) = max r { 0 , 1 , , m } { r : p ( r ) ( x ) Δ α , ν ( H 0 , ( r ) ) } = max r { 0 , 1 , , m } r : p ( r ) ( x ) α π ( H 0 , ( r ) ) β ν ( r ) = max r { 0 , 1 , , m } r : p ( r ) ( x ) α π ( H 0 , ( r ) ) j = 1 r j ν ( j 1 , j ] ,
with p ( 0 ) 0 , while inducing a prior predictive distribution of the number R α , ν in (10) of the smallest p-values from { p ( r ) } r = 1 m that represent significant discoveries. Also, (10) is defined by weights π ( H 0 , ( r ) ) and corresponding test levels α ( r ) = α π ( H 0 , ( r ) ) given the overall prespecified level α (e.g., α = 0.05 or 0.01 , etc.) for the respective m ordered p-values { p ( r ) } r = 1 m satisfying r = 1 m π ( H 0 , ( r ) ) 1 , such that for r = 1 , , m , equal weights π ( H 0 , ( r ) ) = 1 / m defines the unweighted DP-MTP [22], while a weighted DP-MTP is defined by an arbitrary probability distribution on H (as in the weighted B-MTP).
The joint hierarchical DP prior distribution (7) and (8) of the DP-MTP induces a prior predictive distribution for the shape parameter β ν and corresponding threshold parameter Δ α , ν and number of discoveries R α from (4), thereby treating these functions as random instead of fixed as done by the standard MTPs, while accounting for uncertainty in the selection of MTPs and their respective cut-off points and decisions regarding which of the smallest p-values are significant discoveries from a given set H of null hypotheses tested. The DP-MTP method thus emphasizes prior predictive hypothesis testing [56] and multiple bias modeling [57] while extending Vibrations of Effects analysis [58] to provide multiple hypothesis testing with uncertainty quantification. See Karabatsos [22] for further discussion.
The DP-MTP method can be run by R package bnpMTP [59] using the code line bnpMTP(·), which, given m input p-values (and perhaps optional inputs of α , p-value weights, N, and  μ ), uses standard methods to generate N Monte Carlo samples of r.p.m.s { ν g } g = 1 m from the joint prior distribution (7) and (8) (with BY-MTP ν 0 ) and of corresponding N samples from the prior predictive distribution of (9) and (10). In principle, DP-MTP can be extended to online multiple hypothesis testing ([60], and references therein) of a potentially infinite number of hypothesis tests, by relaxing the constraints of the weights π ( H 0 , ( r ) ) of the respective m ordered p values { p ( r ) } r = 1 m ( t ) obtained at any given time point t to satisfy r = 1 m π ( H 0 , ( r ) ) < 1 (with total level r = 1 m α π ( H 0 , ( r ) ) < α ) before all tests are performed, and to satisfy r = 1 m π ( H 0 , ( r ) ) 1 (with intended total level r = 1 m α π ( H 0 , ( r ) ) = α ) after all tests are done (the same method can be used to define an online version of any other MTP mentioned in this section). This defines the only online MTP that controls FWER and/or FDR under arbitrarily dependent p-values with uncertainty quantification in multiple testing. Further, it is straightforward to extend the DP-MTP method to handle tests of continuous hypotheses, either by modeling ν as a discrete r.p.m. (e.g., assigned a DP prior), or by assigning ν a more general Bayesian nonparametric prior distribution which supports the space of continuous random probability measures, such as a DP mixture of continuous densities [61].

3. Predictive MTP Power Analysis Under Arbitrary Dependence

Consider any (off/online, un/weighted, and discrete or continuous) multiple hypothesis testing scenario employing m 1 hypothesis tests using a preset total Type I error rate control α , and corresponding random vector of m 1 test statistics T = ( T 1 , , T m ) with possible realizations T = t = ( t 1 , , t m ) R m . Each test statistic T j T typically has (or can be specified to have) the general form T j θ j / σ θ j , with  θ an effect size parameter of interest (e.g., Section 1). Any realized value t j t n j = θ ^ j / σ θ ^ j of T j is determined by some estimate θ ^ j of θ with (estimated) standard error σ ^ θ ^ j based on a dataset of size n j .
Further, assume that each of the m test procedures tests a null hypothesis H 0 , j H against an alternative hypothesis H 1 , j using a test statistic T j T which, under H 0 , j , follows a standard central location( μ )-scale( σ ) Student t-distribution, T j T ( μ 0 , σ 1 , ν j ) with degrees of freedom ν j > 0 an increasing function of n j ; and which, under H 1 , j , follows a non-central t-distribution, T j NT ( μ j , σ 1 , ν j ) , with non-zero non-centrality parameter μ j R / { 0 } ; perhaps after transforming the original test statistic to have a normal-like distribution under H 0 , j and under H 1 j . The non-central N T ( μ j , 1 , ν ) distribution (and asymptotic normal distribution N ( μ j , 1 ) distribution, resp.) is thus the basis for power analysis for a t-test (z-test, resp.) of a mean difference, correlation or regression coefficient, (log-)odds ratio; etc., as mentioned in Section 1 and by statistics textbooks. Further examples using z-tests and t-tests are illustrated in Section 4.
Recall that if μ 0 , then the N T ( μ , 1 , ν ) distribution is unimodal and asymmetric; and if μ = 0 , this distribution is symmetric and coincides with the T ( 0 , 1 , ν ) distribution. For each test procedure j = 1 , , m , we have that n j implies ν j d and asymptotic convergences to normal distributions, T ( 0 , 1 , ν ) d N ( 0 , 1 ) under H 0 , j , and  N T ( μ j , 1 , ν ) d N ( μ j , 1 ) under H 1 , j , while the test procedure may alternatively employ these asymptotic normal distributions; the two-tailed p-value is given by p j = 2 F ν j ( | t j | ) with asymptotic p-value p j = 2 Φ ( | t j | ) ; and alternatively, a one-sided lower-tailed (upper-tailed, resp.) p-value is given by p j = F ν j ( t ) with asymptotic p-value p j = Φ ( t j ) (or by p j = 1 F ν j ( t j ) with asymptotic p-value p j = 1 Φ ( t j ) , resp.), where F ν is the cumulative distribution function (CDF) of the T ( 0 , 1 , ν ) distribution and Φ is the normal N ( 0 , 1 ) distribution CDF.
Consider m ( 1 ) variate generalizations of the normal, t, and non-central t distributions. Let N m ( μ , Σ ) be the m-variate normal distribution of a random vector Y = ( Y 1 , , Y m ) , defined by parameters of mean vector μ = ( μ 1 , , μ m ) and m × m symmetric positive-definite (s.p.d.) covariance matrix, Σ = ( σ j , k ) m × m , with diagonal elements ( σ 1 , 1 2 σ 1 2 , , σ m , m 2 σ m 2 ) the respective variances of ( T 1 , , T m ) , and each off-diagonal element σ j , k is the covariance C ( T j , T k ) for each distinct j , k { 1 , , m } . It can be convenient to model the (co)variance matrix Σ through either of the three following separation strategies: Σ diag ( σ ) P diag ( σ ) , Σ σ 2 P , or  Σ P ; where diag ( σ ) is the m × m diagonal matrix of standard deviations σ = ( σ 1 , 1 , , σ m , m ) , σ 2 is a common variance, and  P = ( ρ j , k ) m × m is the m × m correlation matrix (i.e., a s.p.d. matrix with ρ j , j = 1 and ρ j , k [ 1 , 1 ] ) for all distinct j , k { 1 , , m } ). A sample random vector Y N m ( μ , Σ ) can be generated by Y = μ + A Z , with  Z = ( Z 1 , , Z m ) and { Z j } j = 1 m iid N ( 0 , 1 ) , using the Cholesky decomposition Σ = A A of the s.p.d. matrix Σ , where A = ( A j , k ) m × m is the lower-triangular matrix with A j j > 0 for j = 1 , , m .
Let T m ( μ , Σ , ν ) be the m-variate Student’s t-distribution of random T = ( T 1 , , T m ) , a symmetric distribution defined by parameters of location (mean) vector μ = ( μ 1 , , μ m ) , scale matrix Σ = ( σ j , k ) m × m , and degrees of freedoms ν = ( ν 1 , , ν m ) ; with corresponding pairwise covariances C ( T j , T k ) = σ j , k ν j ν k 2 Γ { ( ν j 1 ) / 2 } Γ ( ν j / 2 ) Γ { ( ν k 1 ) / 2 } Γ ( ν k / 2 ) (if ν j > 1 , ν k > 1 ) and correlations P ( T j , T k ) = ρ j , k 2 Γ { ( ν j 1 ) / 2 } ν j 2 Γ ( ν j / 2 ) Γ { ( ν k 1 ) / 2 } ν k 2 Γ ( ν k / 2 ) (if ν j > 2 , ν k > 2 ; and ρ j , k σ j , k σ j . j σ k , k ) for all distinct j , k { 1 , , m } , where Γ ( · ) is the gamma function. If  Σ is a diagonal matrix, then the univariate marginal distributions of T m ( μ , Σ , ν ) are independent symmetric Student t-distributions, T ( μ j , σ j , j , ν j ) for j = 1 , , m ([62], Corollary 3). A sample random vector T T m ( μ , Σ , ν ) can be generated by
T = μ + A Z ( V 1 / ν 1 , , V m / ν m ) ,
with Z = ( Z 1 , , Z m ) , Z j N ( 0 , 1 ) and V j χ ν j 2 for j = 1 , , m ([62], Definition 1), where ⊘ denotes the Hadamard (element-wise) division of vectors. A special case of T m ( μ , Σ , ν ) is the T m ( μ , Σ , ν ) distribution ([63,64], Equation (1.1)), defined by stochastic representation T = μ + A Z / V / ν , Z = ( Z 1 , , Z m ) , { Z j } j = 1 m iid N ( 0 , 1 ) , V χ ν 2 , ν > 0 , and non-independent, m univariate marginal T ( μ j , σ j , ν ) distributions, j = 1 , , m .
Now, introduce the m-variate upper non-central t-distribution, T NT m ( μ , Σ , ν ) , defined by parameters of non-centrality vector μ R m , scale matrix Σ , and degrees of freedoms, ν = ( ν 1 , , ν m ) . This m-variate distribution is a member of the class of normal mean-variance mixtures ([65], Section 3.2.2), and defined by the stochastic representation
T = ( μ + A Z ) ( V 1 / ν 1 , , V m / ν m ) ,
with Z = ( Z 1 , , Z m ) , Z j N ( 0 , 1 ) and V j χ ν j 2 for j = 1 , , m . The  NT m ( μ , Σ , ν ) distribution is asymmetric when μ 0 , and coincides with the symmetric T m ( μ , Σ , ν ) distribution when μ = 0 . A special case of NT m ( μ , Σ , ν ) is the m-variate upper non-central t-distribution, NT m ( μ , Σ , ν ) , defined by a scalar degrees of freedom ν ν and univariate marginal non-central t-distributions NT ( μ j , ν ) for j = 1 , , m ([64,66,67], p. 81; Section 3.1; Equation (5.1); resp.), and by a complicated m-variate probability density function (pdf) even if Σ σ 2 P . Therefore, most applications of this distribution T NT m ( μ , Σ , ν ) employ its stochastic representation, given by T = ( V ν ) 1 / 2 ( μ + A Z ) , where Z = ( Z 1 , , Z m ) with { Z j } j = 1 m iid N ( 0 , 1 ) and V χ ν 2 ([66,68], Section 3.1; p.133, Equation (8); resp.), e.g., this distribution can be sampled using the mvtnorm R package code line rmvt(..., type = “Kshirsagar”) [69]. A straightforward modification of this code to divide a m-variate normal random ( μ + A Z ) by a random ( V 1 / ν 1 , , V m / ν m ) instead of V / ν , enables sampling from the NT m ( μ , Σ , ν ) distribution with vector ν via its stochastic representation (12).
Let θ ^ = ( θ ^ 1 , , θ ^ m ) be the vector of anticipated or estimated effect sizes for a planned future study that will conduct tests of a specific set of m null hypotheses. In particular, θ ^ may represent the data analyst’s anticipated effect sizes, or may be the estimates of effect sizes obtained from a prior study (used for a future replication or interim study) of the same m hypothesis tests. Also, let σ θ ^ = ( σ θ ^ 1 , , σ θ ^ m ) be the standard errors of these m effect sizes θ ^ , which are respectively decreasing functions of the (anticipated or actual) sample sizes ( n 1 , , n m ) , and are the square roots of the diagonal elements of the (e.g., inverse Fisher information) variance–covariance matrix Σ θ ^ of θ ^ . In more common applications of power analysis, σ θ ^ is known instead of the full matrix Σ θ ^ , though, covariances can sometimes be derived for certain effect size measures [70]. Specifying such covariances may not be crucial as this paper emphasizes power analysis of MTPs controlling FWER or FDR under arbitrarily dependent p-values (test statistics and effect sizes).
A conditional MTP power analysis m null hypotheses tests is based on test statistics distributed as T T m ( 0 , P , ν ) under all null hypotheses { H 0 , j } j = 1 m jointly, and distributed as T NT m ( μ , P , ν ) under all alternative hypotheses { H 1 , j } j = 1 m jointly, given a specified fixed location vector μ = | θ ^ σ θ ^ | ; m × m correlation matrix P of the underlying m-variate normal N m ( μ , P ) distribution; and degrees of parameters ν = ( ν 1 , , ν m ) (where ν j based on sample size n j for any test based on a normal distribution approximation of its test statistic T j under the null H 0 , j ). For m = 1 hypothesis test, NT m ( μ , P , ν ) becomes the univariate non-central NT ( μ , 1 , ν ) distribution ( N ( μ , 1 ) distribution, resp.) used by classical conditional power analysis of a t-test (z-test if ν , resp.) [1]. However, conditional MTP power analysis relies on a prespecified fixed parameters ( μ , P , ν ) which can be difficult to specify when the number of tests m is sufficiently large, and does not account for their uncertainty; and cannot provide a congenial power analysis of an MTP that controls the FWER or FDR under arbitrary dependence between p-values, because such an MTP accounts for all possible m × m s.p.d. correlation matrices of test statistics underlying these m p-values, instead of only one fixed correlation matrix.
These issues can be addressed by an predictive power analysis of such an MTP, based on a joint prior distribution for the parameters ( μ , P ) driving the m-variate distribution of test statistics, T NT m ( μ , P , ν ) , and corresponding p values under all alternative hypotheses { H 1 , j } j = 1 m . Such a prior accounts for uncertainty in these parameters, and leads to the calculation of marginal powers of m hypothesis tests, thereby addressing all six aims of MTP power analysis (mentioned at the end of Section 1) without needing to assume independent p-values. The joint prior distribution (probability density function) for ( μ , P ) is specified by
π ( μ , P ) = N m ( μ θ ^ σ θ ^ , I m ) LKJ ( P η 1 ) ,
based on an m-variate normal distribution for the mean parameters μ , with prior mean vector defined by anticipated or estimated effect sizes and their respective standard errors, and with (co)variances given by the m × m identity matrix ( I m ); and a uniform prior distribution for the m × m (s.p.d.) correlation matrix parameters P (i.e., the LKJ distribution [71] with parameter η 1 ), which ensures a congenial (predictive) power analysis for such an MTP that controls the FWER or FDR under arbitrary dependent p-values, because this uniform prior supports all possible m × m correlation matrices of the test statistics underlying these p-values. In contrast, a conditional MTP power analysis assigns a point-mass prior at a chosen point ( μ ^ = θ ^ σ θ ^ , P ) with a fixed correlation matrix P .
The default, m-variate normal prior N m ( μ θ ^ σ θ ^ , I m ) in (13) is equivalent to the prior for μ θ σ ^ θ based on a prior θ N m ( θ ^ , diag ( σ θ ^ 2 ) ) on the m effect sizes θ , previously considered for m = 1 effect size [15]. The  N m ( μ θ ^ σ θ ^ , I m ) prior is convenient because it allows for prior mean specification without needing explicit knowledge of standard errors possibly unavailable from a given published study, as illustrated in the case study in Section 4.
That said, publication bias and the winner’s curse often lead to overestimated original effect estimates (for example, refs. [72,73,74]), implying that for a power analysis of a replication study, the  N m ( μ θ ^ σ θ ^ , I m ) prior may be over-optimistic and lead to underpowered replication studies. A simple correction for this over-optimism is to multiply the effect sizes θ ^ in this prior by a vector factor d = ( d 1 , , d m ) that takes on values between the vector of zeros 0 m and a vector of ones 1 m , with shrinkage factor s = 1 m d that can be chosen based on previous replication studies in the same field (as in [15] for the m = 1 case). Further, a prior for μ can be based on a (non-diagonal) s.p.d. covariance matrix Σ θ ^ for θ . More generally, a power prior distribution [43] can be specified for ( μ ) (or even for ( μ , P , ν ) ), defined as a posterior distribution constructed from previous historical data (likelihood) and a prior density for the parameters [16].
Simulation Algorithm 1 calculates the predictive powers for either the B-MTP, H-MTP, BY-MTP, and/or DP-MTP, for m given hypothesis tests, applicable to either an offline or online multiple testing scenario. This algorithm efficiently samples from the specified prior distribution for μ , and rapidly samples from the uniform prior distribution on m × m  correlation matrices using the Pourahmadi & Wang [44] algorithm, run by the randcorr package (v1.0) [75] of the R software (v4.5.1) [76].
Algorithm 1 employs a default “objective” uniform prior that support all valid correlation matrices, for reasons explained in Section 1. However, when necessary, the algorithm can be easily modified in obvious ways to multiple testing situations where the prior on the correlation matrix is informative, such as a prior supporting correlation matrices that have either an equi-correlated structure, or alternatively, an AR(1) structure (e.g., for time-series or repeated measures designs).
Algorithm 1 Simulation algorithm to compute the predictive marginal powers for each MTP.
Inputs: The m 1 test procedures used to test a given set of null hypotheses { H 0 , j } j = 1 m , and their:
               Test statistics T = ( T 1 , , T m ) ; Type(s) of test(s) (lower-, upper, or two-tailed test);
               Ratios, ( θ ^ σ θ ^ ) = ( θ ^ 1 / σ θ ^ 1 , , θ ^ m / σ θ ^ m ) , of anticipated or estimated
               effect sizes θ ^ = ( θ ^ 1 , , θ ^ m ) divided by their corresponding standard errors;
               Degrees of freedoms, ν = ( ν 1 , , ν m ) ; Total Type I error, α ; indicator function 1 ( · ) ;
               Choice of MTP to use: B-MTP, H-MTP, BY-MTP, and/or DP-MTP;
               Number of samples of DP r.p.m.s for DP-MTP if used: N (e.g., N 1000 );
               Number of sampling iterations for power analysis, S (e.g., S 5000 ).
for s = 1 , , S do:
        (a) Draw μ ( s ) N m ( θ ^ σ θ ^ , I m )   (or draw from another prior for μ mentioned in paper).
        (b) Draw P ( s ) LKJ ( η 1 )        (using Pourahmadi & Wang [44] algorithm).
        (c) Draw t ( s ) NT m ( μ ( s ) , P ( s ) , ν )     (using stochastic representation Equation (12)).
        (d) From each t j ( s ) t m ( s ) = ( t 1 ( s ) , , t m ( s ) ) , for  j = 1 , , m , calculate (as intended) the
               lower-, upper-, and/or 2-tailed p-values p 1 ( s ) p ( t 1 ( s ) ) , , p j ( s ) p ( t j ( s ) ) , p m ( s ) p ( t m ( s ) ) .
        (e) Apply each used B-MTP, H-MTP, BY-MTP, and/or DP-MTP, on  p ( s ) , at level α , and
               for sorted p-values p ( 1 ) ( s ) p ( r ) ( s ) p ( m ) ( s ) , calculate for r = 1 , , m ,
               d α , M , ( r ) = 1 { p ( r ) Δ α , ν , M ( H 0 , ( r ) ) } , for MTP M { B , H , BY } ,
               d α , DP , ( r ) ( s ) 1 N g = 1 N d α , DP , ( r ) , g ( s ) , where d α , DP , ( r ) , g ( s ) 1 { p ( r ) ( s ) Δ α , ν g ( s ) , M ( H 0 , ( r ) ) } ,
               to yield: ( d α , B , ( r ) ( s ) ) r = 1 m , ( d α , H , ( r ) ( s ) ) r = 1 m , ( d α , BY , ( r ) ( s ) ) r = 1 m , and/or ( d α , DP , ( r ) ( s ) ) r = 1 m , as intended.
end for
Output: For each MTP used, M { B , H , BY , DP } , the calculated estimates of:
               Predictive marginal powers:   pmp ¯ α , M 1 S s = 1 S ( d α , M , ( r ) ( s ) ) r = 1 m = ( d ¯ α , M , ( r ) ) r = 1 m ;
               Predictive average power:      pap ¯ α , M 1 S 1 m s = 1 S r = 1 m d α , M , ( r ) ( s ) ;
               Predictive disjunctive power: pdp ¯ α , M 1 S s = 1 S 1 ( r = 1 m d α , M , ( r ) ( s ) 1 ) ;
               Predictive conjunctive power: pcp ¯ α , M 1 S s = 1 S 1 ( r = 1 m d α , M , ( r ) ( s ) = m ) .
               p-value weights:                                π ^ M ( H 0 , ( r ) ) exp [ log ( d ¯ α , M , ( r ) ) a ^ ] k = 1 m exp [ log ( d ¯ α , M , ( k ) ) a ^ ] , for  r = 1 , , m ,
                                                                                                            where a ^ max k = 1 , , m [ log ( d ¯ α , M , ( k ) ) ] .
               Significance Chasing Bias (Hellinger Distance):
               SigChase ( r ) 1 2 [ { d α , M , ( r ) 1 / 2 ( x ) d ¯ α , M , ( r ) 1 / 2 } 2 + { ( 1 d α , M , j ( x ) ) 1 / 2 ( 1 d ¯ α , M , ( r ) ) 1 / 2 } 2 ] ;
               for sorted p-values p ( 1 ) p ( r ) p ( m ) and nulls { H 0 , ( r ) } r = 1 m tested on data x,
               with d α , M , ( r ) ( x ) = 1 { p ( r ) Δ α , ν , M ( H 0 , ( r ) ) } , for  r = 1 , , m , for MTP M { B , H , BY } ,
               and/or d α , DP , ( r ) 1 N g = 1 N d α , DP , ( r ) , g = 1 N g = 1 N 1 { p ( r ) Δ α , ν g , M ( H 0 , ( r ) ) } , as intended.
               Speed of Monte Carlo convergence of  h ¯ S , M = 1 S s = 1 s h ( T s ) :
                                                                            Assessed by sample variance v g , S = 1 S 2 s = 1 S { h ( t m ( s ) ) h ¯ S } 2 ,
                                                                            for each g pmp ¯ α , M and each g { pap ¯ α , M , pdp ¯ α , M , pcp ¯ α , M } .
Also, a typical hypothesis test is most often defined by a test statistic that has a known distribution under the null hypothesis, most often a normal or Student t distribution (at least approximately). This is why Section 3 and Section 4 and the MTP power analysis Algorithm 1 emphasize the use of these distributions by default. In other situations, the hypothesis test statistic may have a non-normal, discrete, and/or unknown distribution under the null hypothesis (the latter which can be estimated through permutation- or bootstrap-based simulation), for one or more hypothesis test(s) among the m total tests performed. In these other hypothesis testing situations, it is straightforward to use the non-normal and/or estimated distribution of the hypothesis test statistic under the null hypothesis, instead of the normal or Student t- null hypothesis distribution(s) within Algorithm 1 in order to conduct an MTP power analysis that would be more accurate than a power analysis assuming a null or Student t null hypothesis distribution. Recall from Section 2 that the super-uniformity condition does not make specific requirements on the distribution of the test statistic under the null hypothesis.

4. Case Study Illustration

The applicability of the new MTP predictive power analysis method is showcased through the analysis of two-tailed p-values from forty-one hypothesis tests, presented in the third column of Table 1.
These results were reported by a study [77] of the neuropsychologic effects of childhood lead exposure measured from the shed baby teeth that teachers collected from first- and second-grade Massachusetts schoolchildren during 1975–1978. This landmark study applied an innovative non-invasive way to collect data on accumulated lead content in children (e.g., lead from ingested wall paint), and yielded statistical results that motivated much further research on lead poisoning and related areas, and inspired actions by the Center for Disease Control to lower the blood lead standard for children, and by the U.S. Congress and Environmental Protection Agency (EPA) to eliminate lead from gasoline, paint, plumbing, and other uses. Still, this study was criticized for methodological flaws, including that it reported p-values that did not jointly control for multiple comparisons for all reported p-values (for further discussion, see [78], pp. 279–283). Therefore, these p-values (subsets) have been reanalyzed several times by the MTP literature ([21], Table 1).
The tooth measurements identified a group of 100 children with low lead exposure (<6 ppm; lowest 10th percentile) and a group of 58 children with high lead exposure (>24 ppm; highest 10th percentile). Both groups were compared on each of 11 items of a teachers’ behavior ratings scale, using respective ( χ 2 ) z-tests of equal group proportions of negative behavior responses. These groups were also compared using analysis of covariance (ANCOVA) t-tests, comparing group mean: total sum behavioral rating score; sub-scale and full-scale (or sum) scores on the Wechsler Intelligence Scales for Children measuring verbal IQ and performance IQ; and scores from Seashore and Token tests of verbal processing; and reaction times varied by four time intervals. Each ANCOVA controlled for five covariates: mother’s age at subject’s birth; mother’s educational level; father’s socioeconomic status; number of pregnancies; and parental IQ.
In total, m = 41 hypothesis tests were performed, which Table 1 summarizes, including the number of tests performed for each kind of testing procedure and variable. The fourth column of Table 1 shows the test statistic (t) derived from each two-tailed p-value, obtained by t j = Φ 1 ( 1 p j / 2 ) for each of the z-tests (i.e., tests j = 1 , , 11 in Table 1), and by t j = T 1 ( 1 p j / 2 0 , 1 , ν j = 151 ) for each of the ANCOVA t-tests (tests j = 12 , , 41 = m ), where Φ denotes the standard normal N ( 0 , 1 ) CDF, and  T is the standard t CDF on ν j = n j ( k + 2 ) = 158 ( 5 + 2 ) = 151 degrees of freedom, based on k = 5 covariates and 2 treatment groups. Each test statistic t is an absolute ratio of an effect size measured by a raw group mean difference, divided by the standard error of this difference, where as appropriate, the effect size is either a between-group difference in proportions for a z-test, or a difference in dependent outcome means for an ANCOVA t-test.
The predictive MTP power analysis Algorithm 1 using α = 0.05 was applied for S = 5000 sampling iterations, to evaluate the powers of the DP-MTP (for N 1000 iterations), the Bonferroni MTP, Holm MTP, and the Benjamini–Yekutieli MTP (BY-MTP), based on 41-variate normal N 41 ( θ ^ σ θ ^ , I 41 ) prior for μ , with mean vector θ ^ σ θ ^ specified by the fourth column of Table 1, the vector of test statistics. Further, while a uniform LKJ prior is assigned to the correlation matrix of the 41 test statistics, the degrees of freedom for these test statistics were set by ν j for the z-tests of proportions indexed by j = 1 , , 11 for each of the 11 behavior items, and set by ν j 151 for each of the ANCOVA t-tests ( j = 12 , , 41 ). (For each ANCOVA t-test, Needleman et al. [77] did not provide r 2 of the combined 5 adjusting covariates nor the residual variance of the unadjusted dependent responses for both treatments, which could permit a more exact ANCOVA power analysis [79].) Overall, for each of the four MTPs, this MTP predictive power analysis is for a future replication study applying the same 41 null hypothesis tests and total sample sizes n j = 158 used for each test, based on an ignorable shrinkage factor vector, d 1 m (a vector of m ones) implying no over-optimism in the vector of m effect sizes θ ^ . The code used to run this MTP power analysis is available at: https://github.com/GeorgeKarabatsos/Predictive-Power-Analysis-of-Multiple-Test-Procedures-Under-Arbitrary-Dependence- provide the software code used to run this MTP power analysis, including R code from the bnpMTP package [59] and from the other cited packages.
The following results of the predictive MTP power analysis provided by Algorithm 1 are as follows, for the DP-MTP, B-MTP, H-MTP, and BH-MTP. The DP-MTP can be considered as the baseline MTP method, because as mentioned earlier, the DP prior supports the B-MTP, H-MTP, and BH-MTP, and all other MTPs which control the FWER or FDR under arbitrarily correlated p-values. For the 41 hypothesis tests, the DP-MTP, B-MTP, H-MTP, and BH-MTP obtained average power 0.25, 0.21, 0.22, and 0.26, respectively; all with disjunctive power 1.00 and conjunctive power 0.
Table 1 shows the marginal power for each of the 41 hypothesis tests, estimated by Algorithm 1. These are respectively the marginal powers of the 41 test statistics, with respect to the prior predictive distribution of the test statistics under their alternative hypotheses, under the prior centered on the observed effect sizes, and the uniform prior for the correlation matrices. The table shows that the marginal power tends to increase with the absolute t-statistic, a function of the absolute effect size, as intuitively expected. In other words, these are the marginal powers of the 41 hypothesis tests after accounting for arbitrarily correlated p-values (test statistics), powers which are different across these tests. The marginal powers are not very large for the tests, which reflects the relatively small sample size of the study and confirms the known fact that there is a decrease in power for individual tests in multiple testing scenarios [27]. Further, the DP-MTP marginal powers (Table 1) are similar to those of the three other MTPs (Figure 1, left side).
For DP-MTP, Table 1 also shows the p-value weights, and probability of significance discovery (PrSig and PrSig.w) for each unweighted and weighted p-value among the 41 total tests. These results are compared with the significance discovery indicators for each of the unweighted and weighted versions of B-MTP, H-MTP, and BY-MTP, obtained from the p.adjust() and p.adjust.w() codes of R package someMTP (v1.4.1.1) [80]. DP-MTP indicates that any test results in a discovery by PrSig > 0 (or by PrSig.w > 0 for weighted testing), while providing uncertainty quantification in multiple testing. The table shows that zero values of PrSig (and PrSig.w) correspond to non-discoveries found by the other MTPs.
For DP-MTP, Table 1 shows the significance-chasing bias index for each of the 41 p-values (tests), where non-bias is indicated by a small index value. The 41 bias indices of DP-MTP tended to be lower than the corresponding indices of B-MTP, H-MTP, and BY-MTP (Figure 1, right side), perhaps because of the DP prior supports all MTPs controlling the FWER or FDR instead of one significance cutoff for p-values.
We now consider a prior sensitivity analysis for the DP-MTP, with respect to four different levels of the shrinkage factor (Section 3). The shrinkage (and corresponding sensitivity analysis) can address the fact that, for a future replication setting, the previously observed test statistics defining the prior mean can possibly be affected by overestimation and the winner’s curse, and so the shrinkage can help rule out any systematic optimism in the reported power estimates.
Specifically, we consider shrinkage factors s = 0 (already considered above), 1 / 4 , 1 / 2 , and  3 / 4 , such that each shrinkage factor s is defined as the same for all 41 hypothesis tests, and they are respectively referred to as zero, low, medium and large shrinkage factors. For these shrinkage factors, the DP-MTP respectively obtained: average powers of 0.25, 0.13, 0.05, and 0.02; disjunctive powers of 1.00, 0.98, 0.69, 0.29; and conjunctive powers of zero.
Figure 2 compares the marginal powers, p-value weights, and significance-chasing Hellinger distance measures across these four shrinkage levels. Overall, as expected, the measures of power decreases with increasing shrinkage level. Interestingly, the estimated p-value weights are rather insensitive to varying shrinkage. The significance-chasing measures mostly increased with decreasing shrinking level, but not always, and not necessarily in a strictly ordered fashion. As reasonably expected, these measures can vary with respect to the specified prior distribution on the effect sizes.

5. Theorem

Algorithm 1 addresses MTP predictive power analysis problem by evaluating the expectation
E f { h ( T ) } = T h ( t m ) f ( t m ) d t 1 , , d t m ,
where f is the marginal prior predictive probability density function of test statistics t m = ( t 1 , , t m ) :
f ( t m ) NT m ( t m μ , P , ν ) LKJ ( P η 1 ) d μ 1 d μ m d ρ 1 , 2 d ρ m 1 , m ,
and where h : R m R + is a bounded function h ( · ) of t m corresponding to m ordered p-values p ̲ ( t m ) ( p ( r ) p ( t ( r ) ) ) r = 1 m , and defined either for a vector of predictive marginal powers: pmp α , M = ( h { p ( t ( r ) ) } d α , M , ( r ) ) r = 1 m ( 1 { p ( r ) Δ α , ν , M ( H 0 , ( r ) ) } ) r = 1 m ; for predictive average power: h { p ̲ ( t m ) } pap = 1 m r = 1 m d α , M , ( r ) ( x ) ; for disjunctive power: h { p ̲ ( t m ) } pdp = 1 ( r = 1 m d α , M , ( r ) 1 ) ; or for conjunctive power: h { p ̲ ( t m ) } pcp = 1 ( r = 1 m d α , M , ( r ) = m ) .
For any choice of these functions h, the power analysis algorithm estimates E f { h ( T ) } by h ¯ S = 1 S s = 1 S h ( t m ( s ) ) , and assesses the speed of convergence of h ¯ S by the sample variance:
v h , S = 1 S 2 s = 1 S { h ( t m ( s ) ) h ¯ S } 2 .
Theorem 1. 
The estimator h ¯ S almost surely (a.s.) converges h ¯ S a . s . E f { h ( T ) } by the Strong Law of Large Numbers (SLLN), and  v S is an estimate of the variance
V { h ¯ S } = 1 S T [ h ( t m ) E f { h ( T ) } ] 2 f ( t m ) d t 1 , , d t m .
Proof. 
For the iid h ( t m ( s ) ) obtained from t m ( s ) i i d f , h < implies E f { h ( T ) } < with finite second moment. Then, a.s. convergence of h ¯ S under SLLN holds ([81], Section 22), and h 2 has a finite expectation under f, which ensures that v S estimates (17). □
Further, given that each of the above choices of function h is binary (0 or 1) valued, its variance attains maximum 1 / 4 when E f { h ( T ) } = 1 / 2 , which implies the upper bound 1 / 4 S for its Monte Carlo variance (16). For example, the case study in Section 4 was based on running S 5000 sampling iterations of the algorithm, yielding an upper bound of 1 / 4 S = 1 / 20,000 = 0.00005 for the Monte Carlo variance of each of the presented estimates of power, implying fast convergence of the Monte Carlo power analysis algorithm.

6. Conclusions

This study proposed and illustrated a practical automatic method for congenially evaluating the power of any MTP which controls FWER or FDR under unknown and unanalyzed arbitrary dependencies (inter-correlations) between p-values. The predictive power analysis method is defined by a general prior distribution for the effect sizes and a joint uniform prior for the correlation matrix for the test statistics to fully account for uncertainty in these parameters.
The new method not only can be used to marginal powers of tests (resp.) or corresponding sample size determinations (given desired marginal powers) for a future planned (e.g., replication or interim) study, but also, the calculated marginal powers can either be used to weight p-values to increase MTP power while minimizing the relative impacts of any significance-chasing biases, or compared with raw p-values to evaluate for the presence of such biases.
As mentioned, Algorithm 1 outputs consistent estimates of the marginal predictive powers of the given m hypothesis tests, which can then be used as a basis for assigning weights to p-values, respectively, in a subsequent application of a weighted MTP based on these p-value weights. Recall that a weighted MTP, which assigns relatively higher weights to the subset of tested null hypotheses that are likely to be false, tends to have higher statistical power compared to unweighted multiple testing procedure. An interesting open question is how accurately such an MTP can detect the subset of truly false null hypotheses in any given multiple hypothesis testing problem, for a range of plausible correlation structures for the test statistics (p-values). This question deserves to be addressed extensively in future research.

Funding

This research is supported in part by National Institute for Health grant 1R01AA028483-01.

Data Availability Statement

All the research data and software code that produced all the empirical results presented in this article are provided in: https://github.com/GeorgeKarabatsos/Predictive-Power-Analysis-of-Multiple-Test-Procedures-Under-Arbitrary-Dependence-.

Acknowledgments

The author gives special thanks to the journal editors and anonymous reviewers for editorial suggestions which have helped improve the presentation of this paper, first made publicly available as an arXiv arXiv:2603.07312 on 7 March 2026. Section 2 of this paper was presented at the 14th International Conference on Bayesian Nonparametrics (BNP14) at UCLA on 26 June 2025.

Conflicts of Interest

The author declares no conflicts of interest.

Appendix A. A Review of Existing Algorithms for Optimal p-Value Weighting, and for Calculating the Correlation Matrix of Test Statistics

For Bonferroni MTPs, optimal p-value weights can be found using any one of various optimization algorithms that maximize either the: weighted power [30]; expected number of rejections [82]; average power based on explicit formulae [36,83]; and disjunctive or conjunctive power, while maximizing the asymptotic m-variate normal distribution of z-scores of one-sided p-values for hypothesis tests, given their respective specified levels of marginal (conditional) powers [37]. For graphical or chain MTPs, optimal p-value weights can be found using grid search or simulation algorithms [84,85,86] or deep learning algorithms maximizing weighted power [87].
For specialized multiple testing problems employing m hypothesis tests, the  m × m correlation matrix of the test statistics can be derived either: for uncorrelated test statistics (p-values), perhaps from an orthogonal design or contrasts; or when their covariance matrix or joint distribution can be estimated, perhaps using permutation or bootstrap sampling methods via the Westfall & Young [78] MTP or related MTPs; while such a covariance matrix estimate can be used to whiten-transform [88] the original test statistics into m independent asymptotic normal N ( 0 , 1 ) test statistics; or for pairwise mean comparisons done via a Dunnett or a Tukey–Kramer type MTP. The Dunnett [89] MTP, which compares the outcome mean between each of m given treatment groups to the outcome mean of the control group, the correlation matrix can be calculated as a simple square-root function of group sample sizes, such that, for example, a balanced design produces a m × m matrix with all distinct pairwise correlations equaling 0.50 . Then, the resulting m pairwise mean comparison test statistics under the null hypotheses follows a m-variate t-distribution with zero-location(mean) vector and with scale matrix determined by these correlations and a common degrees of freedom ([89], pp. 1101–1103). For the Tukey [90]–Kramer [91] MTP, which compares means between the distinct pairs of all treatment groups, explicit formulas for the correlation matrix of m test statistics can be derived for certain, e.g., balanced incomplete block, ANCOVA, incomplete Block Lattice, or partially balanced designs; [92]. However, the results of this MTP can conflict with the result of the F-test of equality of means between all three or more treatment groups [93].

References

  1. Cohen, J. Statistical Power Analysis for the Behavioral Sciences (Eds. 1–2); Academic Press: New York, NY, USA, 1969; Lawrence Earlbaum Associates: Hillsdale, NJ, USA, 1988. [Google Scholar]
  2. Hedges, L. Effect sizes for experimental research. Br. J. Math. Stat. Psychol. 2026, 79, 31–45. [Google Scholar] [CrossRef] [Scilit]
  3. Borenstein, M. Effect sizes for continuous data. In The Handbook of Research Synthesis and Meta-Analysis, 2nd ed.; Cooper, H., Hedges, L., Valentine, J., Eds.; Russell Sage Foundation: New York, NY, USA, 2009; pp. 221–235. [Google Scholar]
  4. Hoyle, M. Transformations: An introduction and a bibliography. Int. Stat. Rev. 1973, 41, 203–223. [Google Scholar] [CrossRef] [Scilit]
  5. Dickhaus, T. Simultaneous Statistical Inference; Springer: Berlin/Heidelberg, Germany, 2014. [Google Scholar]
  6. Efron, B. Large Scale Inference: Empirical Bayes Methods for Estimation, Testing, and Prediction; Cambridge University Press: Cambridge, UK, 2010. [Google Scholar]
  7. Wasserstein, R.; Lazar, N. The ASA’s statement on p–values: Context, process, and purpose. Am. Stat. 2016, 70, 129–133. [Google Scholar] [CrossRef] [Scilit]
  8. Benjamin, D.; Berger, J. Three recommendations for improving the use of p-values. Am. Stat. 2019, 73, 186–191. [Google Scholar] [CrossRef] [Scilit]
  9. Held, L.; Ott, M. On p-values and Bayes factors p-values and Bayes factors. Annu. Rev. Stat. Appl. 2018, 5, 393–419. [Google Scholar] [CrossRef] [Scilit]
  10. Murphy, K. Power analysis. In International Encyclopedia of Statistical Science; Lovric, M., Ed.; Springer: Berlin/Heidelberg, Germany, 2025; pp. 1924–1926. [Google Scholar]
  11. Spiegelhalter, D.; Abrams, K.; Myles, J. Bayesian Approaches to Clinical Trials and Health-Care Evaluation; John Wiley and Sons: Hoboken, NJ, USA, 2004. [Google Scholar]
  12. Spiegelhalter, D.; Freedman, L. A predictive approach to selecting the size of a clinical trial, based on subjective clinical opinion. Stat. Med. 1986, 5, 1–13. [Google Scholar] [CrossRef] [Scilit]
  13. Spiegelhalter, D.; Freedman, L.; Blackburn, P. Monitoring clinical trials: Conditional or predictive power? Control. Clin. Trials 1986, 7, 8–17. [Google Scholar] [CrossRef] [Scilit]
  14. Demartino, R.; Egidi, L.; Held, L.; Pawel, S. Mixture Priors for Replication Studies. Stat. Sci. 2026; to appear.
  15. Micheloud, C.; Held, L. Power calculations for replication studies. Stat. Sci. 2022, 37, 369–379. [Google Scholar] [CrossRef] [Scilit]
  16. Pawel, S.; Aust, F.; Held, L.; Wagenmakers, E. Power priors for replication studies. Test 2024, 33, 127–154. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Pawel, S.; Consonni, G.; Held, L. Bayesian approaches to designing replication studies. Psychol. Methods 2023, 31, 22–39. [Google Scholar] [CrossRef] [Scilit]
  18. Tamhane, A.; Gou, J. Multiple test procedures based on p-values. In Handbook of Multiple Comparisons; CRC Press: Boca Raton, FL, USA, 2022; pp. 11–34. [Google Scholar]
  19. Bonferroni, C. Teoria statistica delle classi e calcolo delle probabilità. Pubbl. R Ist. Super. Sci. Econ. Commer. Firenze 1936, 8, 3–62. [Google Scholar]
  20. Holm, S. A simple sequentially rejective multiple test procedure. Scand. J. Stat. 1979, 6, 65–70. [Google Scholar]
  21. Benjamini, Y.; Yekutieli, D. The control of the false discovery rate in multiple testing under dependency. Ann. Stat. 2001, 29, 1165–1188. [Google Scholar] [CrossRef] [Scilit]
  22. Karabatsos, G. Bayesian nonparametric sensitivity analysis of multiple test procedures under dependence. Biom. J. 2025, 67, e70101. [Google Scholar] [CrossRef] [Scilit]
  23. Ferguson, T. A Bayesian analysis of some nonparametric problems. Ann. Stat. 1973, 1, 209–230. [Google Scholar] [CrossRef] [Scilit]
  24. Benjamini, Y.; Hochberg, Y. Multiple hypotheses testing with weights. Scand. J. Stat. 1997, 24, 407–418. [Google Scholar] [CrossRef] [Scilit]
  25. Kang, G.; Ye, K.; Liu, N.; Allison, D.; Gao, G. Weighted multiple hypothesis testing procedures. Stat. Appl. Genet. Mol. Biol. 2009, 8, 23. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Porter, K. Statistical power in evaluations that investigate effects on multiple outcomes: A guide for researchers. J. Res. Educ. Eff. 2018, 11, 267–295. [Google Scholar] [CrossRef] [Scilit]
  27. Senn, S.; Bretz, F. Power and sample size when multiple endpoints are considered. Pharm. Stat. 2007, 6, 161–170. [Google Scholar] [CrossRef] [Scilit]
  28. Bretz, F.; Hothorn, T.; Westfall, P. Multiple Comparisons Using R; Chapman and Hall/CRC: Boca Raton, FL, USA, 2011. [Google Scholar]
  29. Dudoit, S.; Shaffer, J.; Boldrick, J. Multiple hypothesis testing in microarray experiments. Stat. Sci. 2003, 18, 71–103. [Google Scholar] [CrossRef] [Scilit]
  30. Westfall, P.; Krishen, A. Optimally weighted, fixed sequence and gatekeeper multiple testing procedures. J. Stat. Plan. Inference 2001, 99, 25–40. [Google Scholar] [CrossRef] [Scilit]
  31. Spjøtvoll, E. On the optimality of some multiple comparison procedures. Ann. Math. Stat. 1972, 43, 398–411. [Google Scholar] [CrossRef] [Scilit]
  32. Chen, J.; Luo, J.; Liu, K.; Mehrotra, D. On power and sample size computation for multiple testing procedures. Comput. Stat. Data Anal. 2011, 55, 110–122. [Google Scholar] [CrossRef] [Scilit]
  33. Maurer, W.; Mellein, B. On new multiple tests based on independent p-values and the assessment of their power. In Multiple Hypotheses Testing; Springer: Berlin/Heidelberg, Germany, 1988; pp. 48–66. [Google Scholar]
  34. Ramsey, P. Power differences between pairwise multiple comparisons. J. Am. Stat. Assoc. 1978, 73, 479–485. [Google Scholar] [CrossRef]
  35. Westfall, P.; Tobias, R.; Wolfinger, R. Multiple Comparisons and Multiple Tests Using SAS; SAS Institute: Cary, NC, USA, 2011. [Google Scholar]
  36. Roeder, K.; Wasserman, L. Genome-wide significance levels and weighted hypothesis testing. Stat. Sci. 2009, 24, 398–413. [Google Scholar] [CrossRef] [Scilit]
  37. Xi, D.; Chen, Y. Optimal weighted Bonferroni tests and their graphical extensions. Stat. Med. 2024, 43, 475–500. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Cooper, H.; Hedges, L.; Valentine, J. The Handbook of Research Synthesis and Meta-Analysis, 3rd ed.; Russell Sage Foundation: New York, NY, USA, 2019. [Google Scholar]
  39. Ioannidis, J.; Trikalinos, T. An exploratory test for an excess of significant findings (with discussion). Clin. Trials 2007, 4, 245–257. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  40. Elliott, G.; Kudrin, N.; Wüthrich, K. The power of tests for detecting p-hacking. Rev. Econ. Stat. 2026, 1–44. [Google Scholar] [CrossRef] [Scilit]
  41. Marks-Anglin, A.; Chen, Y. A historical review of publication bias. Res. Synth. Methods 2020, 11, 725–742. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Cressie, N.; Read, T. Multinomial goodness-of-fit tests. J. R. Stat. Soc. Ser. B 1984, 46, 440–464. [Google Scholar] [CrossRef] [Scilit]
  43. Ibrahim, J.; Chen, M.; Gwon, Y.; Chen, F. The power prior: Theory and applications. Stat. Med. 2015, 34, 3724–3749. [Google Scholar] [CrossRef] [Scilit]
  44. Pourahmadi, M.; Wang, X. Distribution of random correlation matrices: Hyperspherical parameterization of the Cholesky factor. Stat. Probab. Lett. 2015, 106, 5–12. [Google Scholar] [CrossRef] [Scilit]
  45. Berger, J.; Bernardo, J.; Sun, D. Objective Bayesian Inference; World Scientific: Singapore, 2024. [Google Scholar]
  46. Berger, J.; Wolpert, R. A conversation with James O. Berger. Stat. Sci. 2004, 19, 205–218. [Google Scholar] [CrossRef] [Scilit]
  47. Xie, C. Weighted multiple testing correction for correlated tests. Stat. Med. 2012, 31, 341–352. [Google Scholar] [CrossRef] [Scilit]
  48. Blanchard, G.; Roquain, E. Two simple sufficient conditions for FDR control. Electron. J. Stat. 2008, 2, 963–992. [Google Scholar] [CrossRef] [Scilit]
  49. Dickhaus, T. Randomized p-values for multiple testing of composite null hypotheses. J. Stat. Plan. Inference 2013, 143, 1968–1979. [Google Scholar] [CrossRef] [Scilit]
  50. Dickhaus, T.; Straßburger, K.; Schunk, D.; Morcillo-Suarez, C.; Illig, T.; Navarro, A. How to analyze many contingency tables simultaneously in genetic association studies. Stat. Appl. Genet. Mol. Biol. 2012, 11, 12. [Google Scholar] [CrossRef] [Scilit]
  51. Gosselin, F. A new calibrated Bayesian internal goodness-of-fit method: Sampled posterior p-values as simple and general p-values that allow double use of the data. PLoS ONE 2011, 6, e14770. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  52. Moran, G.; Blei, D.; Ranganath, R. Holdout predictive checks for Bayesian model criticism. J. R. Stat. Soc. Ser. B 2024, 86, 194–214. [Google Scholar] [CrossRef]
  53. Casella, G.; Berger, R. Statistical Inference, 2nd ed.; Duxbury: Pacific Grove, CA, USA, 2002. [Google Scholar]
  54. Hochberg, Y.; Tamhane, A. Multiple Comparison Procedures; John Wiley and Sons: New York, NY, USA, 1987. [Google Scholar]
  55. Benjamini, Y.; Hochberg, Y. Controlling the false discovery rate: A practical and powerful approach to multiple testing. J. R. Stat. Soc. Ser. B 1995, 57, 289–300. [Google Scholar] [CrossRef] [Scilit]
  56. Box, G. Sampling and Bayes’ inference in scientific modelling and robustness (with discussion). J. R. Stat. Soc. Ser. A 1980, 143, 383–430. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  57. Greenland, S. Multiple-bias modelling for analysis of observational data. J. R. Stat. Soc. Ser. A 2005, 168, 267–306. [Google Scholar] [CrossRef] [Scilit]
  58. Vinatier, C.; Hoffmann, S.; Patel, C.; DeVito, N.; Cristea, I.; Tierney, B.; Ioannidis, J.; Naudet, F. What is the vibration of effects? BMJ Evid.-Based Med. 2025, 30, 61–65. [Google Scholar] [CrossRef] [Scilit]
  59. Karabatsos, G. bnpMTP: Bayesian Nonparametric Sensitivity Analysis of Multiple Testing Procedures for p-Values, R package version 1.0.0; R Foundation for Statistical Computing: Vienna, Austria, 2025. [Google Scholar]
  60. Robertson, D.; Wason, J.; Ramdas, A. Online multiple hypothesis testing. Stat. Sci. 2023, 38, 557–575. [Google Scholar] [CrossRef] [Scilit]
  61. Lo, A. On a class of Bayesian nonparametric estimates. Ann. Stat. 1984, 12, 351–357. [Google Scholar] [CrossRef] [Scilit]
  62. Ogasawara, H. The multivariate t-distribution with multiple degrees of freedom. Commun. Stat.-Theory Methods 2024, 53, 144–169. [Google Scholar] [CrossRef] [Scilit]
  63. Cornish, E. The multivariate t-distribution associated with a set of normal sample deviates. Aust. J. Phys. 1954, 7, 531–542. [Google Scholar] [CrossRef] [Scilit]
  64. Kotz, S.; Nadarajah, S. Multivariate t-Distributions and Their Applications; Cambridge University Press: Cambridge, UK, 2004. [Google Scholar]
  65. McNeil, A.; Frey, R.; Embrechts, P. Quantitative Risk Management: Concepts, Techniques, Tools; Princeton University Press: Princeton, NJ, USA, 2005. [Google Scholar]
  66. Juritz, J. Aspects of Non-Central Multivariate t Distributions. Ph.D. Thesis, University of Cape Town, Cape Town, South Africa, 1973. [Google Scholar]
  67. Kshirsagar, A. Some extensions of the multivariate t-distribution and the multivariate generalization of the distribution of the regression coefficient. Math. Proc. Camb. Philos. Soc. 1961, 57, 80–85. [Google Scholar] [CrossRef] [Scilit]
  68. Hofert, M. On sampling from the multivariate t distribution. R J. 2013, 5, 129–136. [Google Scholar] [CrossRef] [Scilit]
  69. Genz, A.; Bretz, F. Computation of Multivariate Normal and t Probabilities; Lecture Notes in Statistics; Springer: Berlin/Heidelberg, Germany, 2009. [Google Scholar]
  70. Gleser, L.; Olkin, I. Stochastically dependent effect sizes. In The Handbook of Research Synthesis and Meta-Analysis, 2nd ed.; Cooper, H., Hedges, L., Valentine, J., Eds.; Russell Sage Foundation: New York, NY, USA, 2009; pp. 357–376. [Google Scholar]
  71. Lewandowski, D.; Kurowicka, D.; Joe, H. Generating random correlation matrices based on vines and extended onion method. J. Multivar. Anal. 2009, 100, 1989–2001. [Google Scholar] [CrossRef] [Scilit]
  72. Anderson, S.; Maxwell, S. Addressing the ’Replication Crisis’: Using original studies to design replication studies with appropriate statistical power. Multivar. Behav. Res. 2017, 52, 305–324. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  73. Button, K.; Ioannidis, J.; Mokrysz, C.; Nosek, B.; Flint, J.; Robinson, E.; Munafò, M. Power failure: Why small sample size undermines the reliability of neuroscience. Nat. Rev. Neurosci. 2013, 14, 365–376. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  74. Ioannidis, J. Why most discovered true associations are inflated. Epidemiology 2008, 19, 640–648. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  75. Makalic, E.; Schmidt, D. An efficient algorithm for sampling from sink(x) for generating random correlation matrices. arXiv 2018, arXiv:1809.05212. [Google Scholar] [CrossRef] [Scilit]
  76. R Core Team. R: A Language and Environment for Statistical Computing; R Foundation for Statistical Computing: Vienna, Austria, 2025. [Google Scholar]
  77. Needleman, H.; Gunnoe, C.; Leviton, A.; Reed, R.; Peresie, H.; Maher, C.; Barrett, P. Deficits in psychologic and classroom performance of children with elevated dentine lead levels. N. Engl. J. Med. 1979, 300, 689–695. [Google Scholar] [CrossRef] [Scilit]
  78. Westfall, P.; Young, S. Resampling-Based Multiple Testing: Examples and Methods for p-Value Adjustment; John Wiley and Sons: New York, NY, USA, 1993. [Google Scholar]
  79. Shieh, G. Power analysis and sample size planning in ANCOVA designs. Psychometrika 2020, 85, 101–120. [Google Scholar] [CrossRef] [Scilit]
  80. Finos, L. someMTP: Some Multiple Testing Procedures, R package version 1.4.1.1; R Foundation for Statistical Computing: Vienna, Austria, 2021. [Google Scholar]
  81. Billingsley, P. Probability and Measure, 3rd ed.; John Wiley and Sons: New York, NY, USA, 1995. [Google Scholar]
  82. Zhang, Z.; Wang, C.; Troendle, J. Optimizing the order of hypotheses in serial testing of multiple endpoints in clinical trials. Stat. Med. 2015, 34, 1467–1482. [Google Scholar] [CrossRef] [Scilit]
  83. Rubin, D.; Dudoit, S.; van der Laan, M. A method to increase the power of multiple testing procedures through sample splitting. Stat. Appl. Genet. Mol. Biol. 2006, 5, 1–18. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  84. Bretz, F.; Maurer, W.; Hommel, G. Test and power considerations for multiple endpoint analyses using sequentially rejective graphical procedures. Stat. Med. 2011, 30, 1489–1501. [Google Scholar] [CrossRef] [Scilit]
  85. Dmitrienko, A.; Paux, G.; Brechenmacher, T. Power calculations in clinical trials with complex clinical objectives. J. Jpn. Soc. Comput. Stat. 2015, 28, 15–50. [Google Scholar] [CrossRef] [Scilit]
  86. Wiens, B.; Dmitrienko, A.; Marchenko, O. Selection of hypothesis weights and ordering when testing multiple hypotheses in clinical trials. J. Biopharm. Stat. 2013, 23, 1403–1419. [Google Scholar] [CrossRef] [Scilit]
  87. Zhan, T.; Hartford, A.; Kang, J.; Offen, W. Optimizing graphical procedures for multiplicity control in a confirmatory clinical trial via deep learning. Stat. Biopharm. Res. 2022, 14, 92–102. [Google Scholar] [CrossRef] [Scilit]
  88. Kessy, A.; Lewin, A.; Strimmer, K. Optimal whitening and decorrelation. Am. Stat. 2018, 72, 309–314. [Google Scholar] [CrossRef] [Scilit]
  89. Dunnett, C. A multiple comparison procedure for comparing several treatments with a control. J. Am. Stat. Assoc. 1955, 50, 1096–1121. [Google Scholar] [CrossRef]
  90. Tukey, J. Comparing individual means in the analysis of variance. Biometrics 1949, 5, 99–114. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  91. Kramer, C. Extension of multiple range tests to group means with unequal numbers of replications. Biometrics 1956, 12, 307–310. [Google Scholar] [CrossRef] [Scilit]
  92. Kramer, C. Extension of multiple range tests to group correlated adjusted means. Biometrics 1957, 13, 13–18. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  93. Gurvich, V.; Naumova, M. Logical contradictions in the one-way ANOVA and Tukey–Kramer multiple comparisons tests with more than two groups of observations. Symmetry 2021, 13, 1387. [Google Scholar] [CrossRef] [Scilit]
Figure 1. For the 41 tests, comparing marginal powers and significance-chasing biases between the 4 MTPs.
Figure 1. For the 41 tests, comparing marginal powers and significance-chasing biases between the 4 MTPs.
Stats 09 00056 g001
Figure 2. For the 41 tests analyzed by the DP-MTP, comparing marginal powers, p-value weights, and significance-chasing biases across four levels of shrinkage, namely: zero-shrinkage (0), low shrinkage of s = 1 / 4 (L), medium shrinkage of s = 1 / 2 (M), and high shrinkage of s = 3 / 4 (H).
Figure 2. For the 41 tests analyzed by the DP-MTP, comparing marginal powers, p-value weights, and significance-chasing biases across four levels of shrinkage, namely: zero-shrinkage (0), low shrinkage of s = 1 / 4 (L), medium shrinkage of s = 1 / 2 (M), and high shrinkage of s = 3 / 4 (H).
Stats 09 00056 g002
Table 1. Two-tailed p-values from (Needleman et al. [77], Tables 3, 7 and 8); test (t) statistics; and DP-MTP results of marginal predictive power (MargPwr); p-value weights; probability of discovery for each p-value, PrSig (each weighted p-value PrSig.w); and significance-chasing bias index in Hellinger distance. Superscripts indicate significant discoveries according to either the un/weighted B-MTP, H-MTP, or BY-MTP.
Table 1. Two-tailed p-values from (Needleman et al. [77], Tables 3, 7 and 8); test (t) statistics; and DP-MTP results of marginal predictive power (MargPwr); p-value weights; probability of discovery for each p-value, PrSig (each weighted p-value PrSig.w); and significance-chasing bias index in Hellinger distance. Superscripts indicate significant discoveries according to either the un/weighted B-MTP, H-MTP, or BY-MTP.
TestScalep-Valuet Stat.MargPwrp-WeightsPrSigPrSig.wsigChase
1Behavior 10.0032.970.460.050.280.53 BY0.13
2Behavior 20.051.960.230.020.000.000.35
3Behavior 30.051.960.230.020.000.000.35
4Behavior 40.141.480.150.010.000.000.28
5Behavior 50.081.750.180.020.000.000.31
6Behavior 60.012.580.370.040.020.08 BY0.35
7Behavior 70.042.050.250.030.000.000.37
8Behavior 80.012.580.380.040.020.08 BY0.36
9Behavior 90.051.960.230.020.000.000.35
10Behavior 100.0032.970.470.050.280.53 BY0.14
11Behavior 110.0032.970.470.050.280.53 BY0.14
12Sum Behavior0.022.350.300.030.000.000.40
13Verbal IQ 10.042.070.250.030.000.000.37
14Verbal IQ 20.051.980.230.020.000.000.35
15Verbal IQ 30.022.350.310.030.000.000.41
16Verbal IQ 40.490.690.060.010.000.000.18
17Verbal IQ 50.081.760.180.020.000.000.30
18Verbal IQ 60.360.920.080.010.000.000.20
19Performance IQ 10.032.190.280.030.000.000.39
20Performance IQ 20.380.880.070.010.000.000.19
21Performance IQ 30.151.450.140.010.000.000.27
22Performance IQ 40.540.610.050.010.000.000.17
23Performance IQ 50.900.130.030.000.000.000.13
24Performance IQ 60.370.900.080.010.000.000.20
25Full Verbal IQ0.032.190.270.030.000.000.38
26Full Perf. IQ0.032.190.280.030.000.000.39
27Full VerbalPerf.IQ0.081.760.200.020.000.000.32
28Seashore 10.0023.150.500.050.380.70 B,BY0.09
29Seashore 20.032.190.260.030.000.000.37
30Seashore 30.071.820.190.020.000.000.32
31Total Seashore0.0023.150.510.050.380.70 B,BY0.09
32Token 10.370.900.070.010.000.000.19
33Token 20.900.130.040.000.000.000.14
34Token 30.420.810.070.010.000.000.19
35Token 40.051.980.210.020.000.000.34
36Total Token0.091.710.180.020.000.000.30
37Sentence0.042.070.240.020.000.000.36
38Reaction Time 10.321.000.080.010.000.000.20
39Reaction Time 20.0013.360.560.060.57 B,H0.76 B,H,BY0.01
40Reaction Time 30.0013.360.550.050.57 B,H0.76 B,H,BY0.02
41Reaction Time 40.012.610.370.040.020.08 BY0.35
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Karabatsos, G. Predictive Power Analysis of Multiple Test Procedures Under Arbitrary Dependence. Stats 2026, 9, 56. https://doi.org/10.3390/stats9030056

AMA Style

Karabatsos G. Predictive Power Analysis of Multiple Test Procedures Under Arbitrary Dependence. Stats. 2026; 9(3):56. https://doi.org/10.3390/stats9030056

Chicago/Turabian Style

Karabatsos, George. 2026. "Predictive Power Analysis of Multiple Test Procedures Under Arbitrary Dependence" Stats 9, no. 3: 56. https://doi.org/10.3390/stats9030056

APA Style

Karabatsos, G. (2026). Predictive Power Analysis of Multiple Test Procedures Under Arbitrary Dependence. Stats, 9(3), 56. https://doi.org/10.3390/stats9030056

Article Metrics

Back to TopTop