Abstract
We propose a Bayesian sequential procedure to test hypotheses concerning the relative risk between two specific treatments based on the binary data obtained from the two-arm clinical trial. Our development is based on an optimal sequential test cast within the Bayesian framework. This approach enables us to provide, in a straightforward manner based on the Stopping Rule Principle (SRP), an assessment of the various error probabilities via posterior probabilities and conditional error probabilities. An attractive feature of our approach is the relative simplicity of the calculations involved without having to resort to cumbersome iterative methods of ‘spending the alpha’ and the like. The proposed methods are illustrated using a sequential safety study of adverse events following H1N1 influenza vaccination and are compared with existing sequential approaches. The findings demonstrate that the proposed Bayesian framework provides an efficient and flexible approach for the early detection of adverse events under several different prior distributions of the parameters involved.
MSC:
62L10; 62L12; 62C10
1. Introduction
The post-market safety surveillance of vaccines and drugs increasingly relies on sequential monitoring of adverse events, in which data accumulate over time as patients are enrolled, treated, and observed for a pre-specified side effect. Regulatory and public health programs such as the Vaccine Safety Datalink and the FDA’s Sentinel Initiative routinely use a two-arm trail comparison, typically contrasting an exposed (treated) group with an unexposed (control) group to detect an elevated rate of an adverse event as early as possible while controlling the chance of a false alarm. A leading example that motivates this paper is the sequential safety surveillance of the H1N1 influenza vaccine analyzed by ref. [1] and later analyzed by refs. [2,3]. In such studies, the main focus is the relative risk (RR) of an adverse event in the exposed group versus the control group. Because enrollment and follow-up occur sequentially, any valid inferential procedure must account for it in the decision to stop (or continue) the study. For such studies, ref. [2] has developed a linear programming framework and proposed an elaborate alpha-spending approach that minimizes the expected time to signal or the expected sample size of the relative risk (RR) under the sequential hypothesis testing procedure. Extension to the bounded-width confidence interval was developed in ref. [3]. However, this approach requires substantial and cumbersome calculations in order to determine the sequence of stopping boundaries, and may not achieve the desired overall Type I error probability.
In this paper, we utilize the Bayesian framework to construct a sequential testing procedure of the hypotheses concerning the RR. As will be seen below, this Bayesian approach provides substantially simpler calculations of the probabilities of the various errors associated with these tests of hypotheses. By the Stopping Rule Principle (SRP) (see ref. [4]), posterior distributions only depend on the data actually observed, but not on the stopping rule itself. This property is shown to be particularly attractive for the sequential monitoring of adverse events in which the stopping rule is incidental to the evidentiary content of the data collected up to that point. Ref. [5] exploited this property to construct a modified Bayesian sequential test that augmented the standard Bayesian sequential test with a ‘no-decision’ region and admitted both a Bayesian and a conditional-frequentist interpretation of the resulting error probabilities. More recently, refs. [6,7] introduced and established the existence of a Uniformly Most Powerful Bayesian Test (UMPBT), which provides an explicit bridge between the Bayesian rejection threshold and a corresponding frequentist rejection region and the significance level. The main contribution of this paper is the construction of a complete Bayesian sequential testing framework for the relative risk in a two-arm adverse event surveillance study with binary data. This approach simplifies the computational workflow and enriches the outcome with directly interpretable measures of evidence instead of using cumbersome alpha-spending calculations.
In Section 2, we consider the testing problem of general hypotheses and related inference problems to the RR, , entirely within the Bayesian framework. Instead of determining a value of under to calculate the sequential test statistics, the Bayesian sequential test uses the entire alternative support to calculate the corresponding Bayes factor. The resulting sequential Bayesian test does not necessitate any alpha-spending calculations (of the type provided by ref. [2]). Additionally, we present a modified Bayesian test (as was advocated in ref. [5]; see their Section 3 in particular) that benefits from an easily calculated conditional error probabilities and has both a classical (conditional) frequentist as well as Bayesian interpretation. In Section 3, we apply our Bayesian sequential test and the modified Bayesian sequential test to the data discussed in refs. [2,3] under different alternatives, as well as different priors. Specifically, based on the modified Bayesian test, we calculate the conditional Type I or Type II error probability for each data point with the corresponding decision we make. Furthermore, we connect our results to UMPBT and some simulations that can be compared from the frequentist perspective. In Section 4, we compare our results with the results presented in refs. [2,3] and provide some closing remarks.
2. The Bayesian Framework for Binomial Probabilities
2.1. Statistical Framework
Consider a clinical experiment involving a finite population of patients. The patients arrive sequentially and are classified according to two binary attributes: whether they were assigned to receive one of the two available treatments A or B and whether they exhibited a particular side effect (Yes or No) of interest. Having classified in this manner the first n patients, , the results are summarized in the following table, Table 1 below.
Table 1.
Table of n items with two treatment groups.
For a give n, the distribution of these four counts is the multinomial distribution, so that
where is the vector of the corresponding probabilities, along with . Note that the counts in Table 1 follow, marginally for a given n, the Binomial distribution, . In fact, as our sampling of units for classification is conducted sequentially, one at a time, we find it useful to denote by , with , for each sampling stage , in which case,
In the design of the two-arm clinical trial, in which the patients (who are arriving sequentially) are assigned at random to receive one of two available treatments, with some known probability, and , respectively, resulting with of the patients receiving treatment A and of them receiving treatment B. Accordingly, for a given n,
Similarly, given n, , once assigned to one of the two treatment groups, the patients are subsequently observed during a fixed-length monitoring window for an ‘event’ of interest, such as the post-treatment expression of an adverse side effect, which we generically denote for ‘present’ and for a ‘absent’. We denote by and by the conditional probability for expressing (presenting) the monitored side effect given that the patient received Treatment A or Treatment B, respectively. Accordingly, we have
as well as
so that , as desired by design.
For such a study, one is interested in making an inference concerning the unknown probabilities and , and, more importantly, concerning the relative risk (RR) of the two treatments to express the monitored side effect, namely, . Note that, under this two-arm clinical trial, the (total) probability for presenting the monitored side effect, which we denote by , is therefore
Hence, we immediately obtain that, given an observed side effect, the conditional probability it may be attributed to Treatment A as
where we have put to indicate the known matching odds of patients assigned to the treatment B rather than to treatment A. The matching ratio could be thought as the ratio of the number of patients assigned to treatment B relative to the number of patients assigned to treatment A.
Having classified the first n patients according to the above table, we denote by the total number of patients presenting the monitored side effect (i.e., ), and we denote by the number of patients in Treatment group A that presented the monitored side effect. Clearly, given n, . However, since it is also given that n, , we give n and , . Refs. [2,3,8] refer to m as the ‘time’ index for this two-arm binomial experiment. Accordingly, given n, the statistic can be used for inference on the unknown parameter , whereas, given n and , the count X (out of m) can be used for inference on the unknown parameter and hence of the relative risk parameter .
For the ‘simple’ case, ref. [9] considered the sequential test
for some specified , , which continues the sampling process as long as the null hypothesis, , is not rejected, and stops the sampling process once there is sufficient ‘evidence’ to reject it. More specially, for given desired probabilities of the Type I error and the Type II error at a given , they determined the corresponding sample size and a critical (boundary) test value needed to achieve an ‘optimal’ fixed-sample UMP test of the hypotheses in (2). Since the observations leading to the data in Table 1 become available, sequentially, one at a time or in batches, the collection of the data ceases once the observed total number of patients exhibiting the monitored side effect, , exceeds the predefined threshold , leading to the stopping time
which has the Negative Binomial distribution, , on the integers . Accordingly, the proposed -optimal sequential test of (2) may be written as
Note that, in either case, the final number of observations at ‘termination’ is , which is the total number of observations collected and used to arrive at a conclusion for the test for (2). Note, however, that, unlike the Sequential Probability Ratio Test (SPRT), the sequential collection of the data under the Stopping Rule (3) is continued as long as is not rejected. Accordingly, upon ’termination’ (either by reaching the maximal number of observations or by rejecting beforehand), .
Refs. [2,10] have considered the above two-arm binomial experiment and the relative risk (RR), , of the two treatments to present the monitored side effect. Thus, they considered the testing problem of the hypotheses concerning the parameter in the form
which are equivalent to the hypotheses against .
For a specified matching allocation ratio between the control group and the treatment group, , we denote . We note that the hypotheses in (5) can be stated equivalently in terms of as
For instance, if the matching allocation ratio is 1:1 (balanced), then and therefore , leading to the test of
Hence, the optimal sequential test of ref. [9] for hypotheses in the form of (2) can be applied in this current situation also.
2.2. The Standard Bayesian Test
Suppose that for a given m and unknown , a binomial random variable, with a given by
We will see below that the structure of the hypotheses in (2) can generally be formulated as
for some and . We assume that, given that the hypothesis is true, has a proper prior , over , , so that
Accordingly, the test of the hypotheses in (7) can equivalently be formulated as the simple hypotheses concerning the marginal distribution of , namely,
where, for a given m and , is the marginal distribution of under , given by
respectively. In the Bayesian framework, one typically assumes that the prior probability that is true is for some and then proceeds to obtain the posterior probability of being true, given the data , namely , where, in the case of the hypotheses (8),
where we have substituted and . The ratio is the so-called Bayes factor of to (see, for example, ref. [5]). Note that when , , which indicates that the prior probability of the null hypothesis is equal to the prior probability of the alternative hypothesis. The posterior probability of being true, given the data , simplifies to
Hence, in this basic fixed-sample setup, the Bayesian would simply reject in favor of if , or, equivalently, if . That is, for a given m and , the Bayesian (fixed-sample) test is
However, the Bayesian test in (11) might lead to certain ambiguities, especially whenever the value of the Bayes factor is close to the boundary of 1. Based on ref. [11], ref. [12] suggests varying levels of ‘rejection of ’, all according to the amount of evidence in the data measured by the index , as presented in Table 2.
Table 2.
The values of Bayes factor as a summary of the evidence provided by the data.
Accordingly, one would substantially reject if and would strongly reject if and decisively reject if .
As is apparent from (9), the calculations of the Bayes factor require specification of the prior distribution for . For the binomial model in (7), a natural choice of prior distributions for would be the conjugate class of distributions over so that for some parameters , . Hence, the prior of is given by
where and where, for any ,
is the incomplete beta function. Note that . With this (conjugate) Beta prior distribution for over , the posterior distribution of , given the data , is the distribution, so that, given m and ,
Furthermore, in this case, the marginal probability distribution of is the Beta-Binomial distribution, given by
As an immediate application, consider the binomial experiment, as summarized in Table 1 above, in which the sequential collection of the data was terminated after n observations according to the stopping rule in (3), so that, for the given n and , , where, as in (6), with the known value . We remind the reader that, according to the Stopping Rule Principle (SRP), posterior probabilities, as evidentiary data measures, remain unaffected by the optional Stopping Rule for the data collection; see ref. [4] (p. 74) or ref. [13] (p. 502). Hence, in this case also, with the stopping time M in (3), the expression (9) for the posterior probability of holds true, also given the event for some and . That is
In similarity to the hypotheses in (6), we consider two additional cases of hypotheses concerning the RR parameter . In all cases, we will assume equal prior probabilities of the null and the alternative hypotheses so that, in all cases, we take (so that ) in (9).
- Case 1: We are interested in testing of against , which, with , can readily be seen as equivalent to the testing ofIn this case, for , the marginal distributions of under and areandrespectively. Hence, the Bayes factor of to in (13) isAccordingly, upon termination of the observation process at for some n, with a given total number of observed side effects , of which of them are ‘attributed’ to Treatment A, we may calculate the value of the Bayes factor , in (14), for any choice of and . Therefore, we can obtain the corresponding posterior probability of being true (given the data ), as given in (19) below.
- Case 2: In a similar fashion, the testing of against , which, with , can readily be seen as equivalent to the testing ofIn this case also, with equal prior probabilities for and (i.e., ), we have for ,andHence, the Bayes factor of to in (15) isHere also, upon termination of the observation process at for some n, with a given total number of observed side effects , of which of them are ‘attributed’ to Treatment A, we may calculate the value of the Bayes factor , in (16), for any choice of a, b and and therefore obtain the corresponding posterior probability of being true (given the data ), as given in (19) below.
- Case 3: Here, we test against , which is equivalent to testingwith . In this case also, with equal prior probabilities for and , (i.e., ), we have, for ,andHence, the Bayes factor of to in (17) isClearly, the calculation of the posterior probability is straightforward, as in the previous cases. In fact, in all Cases 1–3, we reject if and report the respective posterior probabilities of and being true asAccordingly, in all of these three cases, the Bayesian test can be presented asfor . Clearly, the Bayesian tests in (21) can incorporate Jeffrey’s criteria, as in Table 2, to determine the strength of the data in evidence to support the rejection of the respective .
Note that, to fully apply the test of the hypotheses in (8), one would need to specify the values of the prior distribution parameters a and b and matching allocation ratio between the two treatment groups, , in the calculations of , .
2.3. The Modified Bayesian Test
In the spirit of Table 2, we modify the standard Bayesian test to include a ‘no-decision region’. This would often be a useful approach in a situation when there is an ambiguous result that appears.
As advocated and suggested in ref. [5], we modify the Bayesian test in (8) to include a no-decision region. The modified Bayesian test takes the form
where are two constants, defined and calculated as described in (23) below. For more details about the ‘no-decision region’ and a further explanation and justification in the sequential setting, see the details in Section 3 of ref. [5].
Let and be the cdf of under and , respectively. Their inverse and exist in the range of . For any , we let
Accordingly, the decision constants r and u are calculated as
As stated in ref. [5], the posterior probabilities of and could also be interpreted as the conditional error frequentist Type I and Type II probabilities. is also a conditional frequentist test arising from the use of the conditioning statistic
over the domain (the complement of is the no-decision region). Therefore, the conditional error probabilities upon accepting or rejecting , given , denoted as and , are given by
which are in complete agreement with the Bayesian posterior probabilities of , given the data provided in (10) (see also (19) and (20)).
3. Example
Ref. [2] reported the results of a sequential study on the side effects of the H1N1 influenza vaccine during the 2009–2010 influenza season (see ref. [1]), which was conducted with an equal allocation between the exposed group, A, and the not-exposed group, B, so that the study design matching constant was and, hence, . The 24 successive data points on from the study cited are presented in the table below, along with the values of . Applying their iterative ‘sequential’ procedure to the hypotheses of Case 2, ref. [2] determined, for , that the testing procedure should have been terminated at the 19th-group data point, with and , which resulted with an estimated RR of . For Case 3 and a related confidence interval, ref. [3] determined, for , that the testing procedure should have been terminated at the 18th-group data point, with and , which resulted with an estimated RR of .
For comparison, we apply the above Bayesian Beta-Binomial models we propose to these 24 data points under several prior distributions; however, all had a prior mean of . Bayes factors are known to depend on the prior specification under the alternative hypothesis, and the sensitivity of Bayes factors to prior assumptions has been widely discussed in the Bayesian literature. For this reason, we consider both non-informative and informative prior specifications in our analysis. The purpose of these examples is not to advocate a universally objective prior, but rather to illustrate the behavior of the proposed sequential testing procedure under commonly used prior structures.
3.1. The Standard Bayesian Test
- (I)
- Case of non-informative (uniform) prior distribution
We assume a uniform prior distribution for , so that . We calculated the Bayes factor and the corresponding posterior probability , for the hypotheses tests on the RR, in the three cases we considered above. We always assume that the prior probability that is true, that is, . The numerical results are provided in Table 3 below.
Table 3.
The values of Bayes factor and posterior probability of for each data point among the three different cases with the uniform prior.
As can be seen, for the hypotheses of Case 1, our Bayesian model above has led to the potentially earlier ‘termination’ at the 18th-group data point with a total of ‘events’, the observed side effect of which was noted in the treatment group with an estimated relative risk of . For the hypotheses of Case 2, our Bayesian model led to the potentially earlier ‘termination’ at the 17th-group data point with a total of ‘events’, the observed side effect of which were noted in the treatment group with an estimated relative risk of . The resulting Bayes factors are and , which are less than based on the standard shown in Table 2. This leads to a rejection of , and the posterior probabilities of the null hypothesis, and the alternative hypotheses, given the data, are
and
which are clearly smaller than the assumed prior probability of , being true (marked in a red line in Figure 1).
Figure 1.
The posterior probability of for the 24 data points of Table 3: (a) Case 1; (b) Case 2; (c) Case 3.
On the other hand, for the hypotheses of Case 3, one would ‘terminate’ the observation process much earlier, already at the 14th-group data point with a total of ‘events’, the observed side effects of which were noted in the treatment group, leading to the estimated relative risk of . The resulting Bayes factor is , which is less than based on the standard shown in Table 2, and the posterior probabilities of the null hypothesis and the alternative hypotheses given the data are
which are clearly smaller than the assumed prior probability of being true (marked in a red line in Figure 1).
Remark 1.
We point out that, in Case 3, we choose the early ‘termination’ at the 14th-group data point instead of 8th-group data point. At the 8th-group data point with a total of ‘events’, the observed side effects of which were noted in the treatment group, leading to the estimated relative risk of . The resulting Bayes factor is , and the posterior probabilities of the null hypothesis and the alternative hypothesis given the data are
Although the value of is less than , following the standard shown in Table 2, the evidence against at the 8th-group data point is not worth more than a bare mention. Therefore, we would substantially reject and stop at the 14th-group data point.
Remark 2.
In Case 3, we observe that the values of Bayes factor at the 4th and 9th-group data point are exactly equal to 1. Due to the data structure, we have the matching allocation ratio , and, at the 4th-group and 9th-group data points, . Moreover, since the mean of the chosen prior distributions is , it is not surprising that, at these two specific data points, the values of the Bayes factor are 1.
We further applied Jeffrey’s prior, discussed in ref. [14], as an ‘objective’ and non-informative prior to above three cases since a uniform prior on the parameter does not generally remain uniform under nonlinear transformations of the parameter . Specifically, for the binomial model, Jeffrey’s prior is . Since the results under Jeffrey’s prior are similar to the results we obtained under the uniform prior, we would not give a detailed interpretation of the results under Jeffrey’s prior (shown in Appendix A).
- (II)
- Case of informative prior distribution
Here, we assume an informative prior distribution for that is induced by determining the values of by the following set of two equations:
for some given and and where by (1). In the current example, we have and assume and . Hence, we ‘solve’ for our prior parameters to obtain . Note that the informative prior information may be obtained from expert suggestion or by borrowing from previous similar studies. The informative prior presented in this case is only an illustration. Applying the procedure we proposed above, we calculated the Bayes factor and the corresponding posterior probability , , for the hypotheses testing on in the three cases we considered above. The numerical results are provided in Table 4 below.
Table 4.
The values of Bayes factor and posterior probability of for each data point among different cases with the informative prior.
As can be seen, for the hypotheses of Case 1, based on the standard shown in Table 2, our Bayesian model above led to the potentially earlier ‘termination’ at the 17th-group data point with a total of ‘events’, the observed side effect of which was noted in the treatment group with an estimated relative risk of . For the hypotheses of Case 2, our Bayesian model led to the potentially earlier ‘termination’ at the 16th-group data point with a total of ‘events’, the observed side effect of which was noted in the treatment group with an estimated relative risk of . The resulting Bayes factors are and . These lead to a rejection of , , and the posterior probabilities of the null hypothesis and the alternative hypothesis given the data are
and
which are clearly smaller than the assumed prior probability of , being true (marked in a red line in Figure 2).
Figure 2.
The posterior probability of for the 24 data points of Table 4: (a) Case 1; (b) Case 2; (c) Case 3.
Furthermore, for the hypotheses of Case 3, based on the standard shown in Table 2, the 15th-group data point with a total of ‘events’, the observed side effect of which was noted in the treatment group, lead to an estimated relative risk of . The resulting Bayes factor is , and the posterior probabilities of the null hypothesis and the alternative hypothesis given the data are
which is clearly smaller than the assumed prior probability of being true (marked in red line in Figure 2).
In conclusion, with the uniform prior, Case 1 would be rejected and ‘terminated’ at the 18th-group data point, Case 2 would be rejected and ‘terminated’ at the 17th-group data point, and Case 3 would be rejected and ‘terminated’ at the 14th-group data point. However, with the informative prior we chose, Case 1 would be rejected and ‘terminated’ at the 17th-group data point, Case 2 would be rejected and ‘terminated’ at the 16th-group data point, and Case 3 would be rejected and ‘terminated’ at the 15th-group data point.
3.2. The Modified Bayesian Test
Applying the modified Bayesian test in (22) to Case 1, 2, and 3, along with the same prior structures we discussed above, we analyze the results of the above example again. Under both types of priors, we obtained the same conclusions that were obtained utilizing the standard Bayesian test for the hypotheses , (see Section 3.1). The results corresponding to the non-informative (uniform) prior distributions are shown in Figure 3, Table 5 and Table 6. The results corresponding to the informative prior distributions are shown in Figure 4, Table 7 and Table 8. However, under both prior structures, utilizing the modified Bayesian test, we see that there are several cases with some data points falling into the ‘no-decision’ region. In all cases, we provide the calculated conditional Type I and Type II error probabilities, and , corresponding to the ‘Reject’ or ‘Accept’ decision (see Table 6 and Table 8).
Figure 3.
The acceptance and rejection boundaries of for the 24 data points of Table 5: (a) Case 1; (b) Case 2; (c) Case 3.
Table 5.
The values of Bayes factor and the boundaries based on it for each data point among different cases with the uniform prior.
Table 6.
The decision and corresponding conditional and for each data point among different cases with the uniform prior (Decision: ‘R’ indicates Reject, ‘A’ indicates Accept, ‘ND’ indicates No Decision, ‘NA’ indicates No Answer).
Figure 4.
The acceptance and rejection boundaries of for the 24 data points of Table 7: (a) Case 1; (b) Case 2; (c) Case 3.
Table 7.
The values of Bayes factor and the decision based on it for each data point among different cases with the informative prior.
Table 8.
The decision and corresponding conditional and for each data point among different cases with the informative prior (Decision: ‘R’ indicates Reject, ‘A’ indicates Accept, ‘ND’ indicates No Decision, ‘NA’ indicates No Answer).
One may interpret the results of Case 2 under the informative prior case, for instance, as a guide for explaining the results shown in these 4 tables. According to the results in Table 7, at the 10th-group data point, since the Bayes factor is greater than , one would accept the null hypothesis that the relative risk is equal to 1 and would report the corresponding conditional Type II error probability , as shown in Table 8. On the other hand, at the 13th-group data point, since is less than , along with the standard mentioned in Table 2, we would reject the null hypothesis at the 13th-group data point and report the conditional Type I error probability . Additionally, there are some data points within the ‘no-decision’ region, since the Bayes factor of the corresponding data point is within the values of and .
- (I)
- Results under non-informative (uniform) prior distribution
- (II)
- Results under informative prior distribution
3.3. HPD Region
Once the ‘termination’ point of the study is known (with a given prior distribution), it is possible to obtain the Highest Posterior Density (HPD) region for the RR, , as now one has its posterior available, given the data. We illustrate this by calculating a HPD region for each ‘terminated’ data point we discuss above.
- Case 1 ()
At the 18th-group data point, which is the ‘terminated’ data point for Case 1 with the uniform prior, we obtain the posterior distribution of , where . The calculated HPD region of is , shown in Figure 5, which indicates that the corresponding HPD region of , the relative risk, is . Note that the conditional Type I error probability is for Case 1 under the modified Bayesian test. Although the HPD region excludes , the conditional Type I error probabilities indicate that there are around in favor of the null hypothesis for Case 1.
Figure 5.
The HPD region of where at the 18th-group data point with the posterior of in Case 1.
At the 17th-group data point, which is the ‘terminated’ data point for Case 1 with the informative prior, we obtain the posterior distribution of , where . The calculated HPD region of is , which indicates that the corresponding HPD region of is . Note that the conditional Type I error probability is for Case 1 under the modified Bayesian test.
- Case 2 ()
At the 17th-group data point, which is the ‘terminated’ data point for Case 2 with the uniform prior, we obtain the posterior distribution of , where . The calculated HPD region of is , which indicates the corresponding HPD region of , the relative risk, is . Note that the conditional Type I error probability is for Case 2 under the modified Bayesian test.
At the 16th-group data point, which is the ‘terminated’ data point for Case 2 with the informative prior, we obtain the posterior distribution of , where . The calculated HPD region of is , which indicates that the corresponding HPD region of is . Note that the conditional Type I error probability is for Case 2 under the modified Bayesian test.
- Case 3 ()
At the 14th-group data point, which is the ‘terminated’ data point for Case 3 with the uniform prior, we obtain the posterior distribution of , where . The calculated HPD region of is , which indicates that the corresponding HPD region of is . The conditional Type I error probability is under the modified Bayesian test.
At the 15th-group data point, which is the ‘terminated’ data point for Case 3 with the informative prior, we obtain the posterior distribution of , where . The calculated HPD region of is , which indicates that the corresponding HPD region of is . The conditional Type I error probability is under the modified Bayesian test.
Remark 3.
We notice that some HPD regions may include , the value of γ under the null hypothesis. Since the Bayes factor shows that the overall evidence from the data favors the alternative compared to the null, the HPD region reflects that the null is not entirely implausible. Even if the null value is within the HPD region, the posterior may still assign higher credibility to other values (those outside the null), which explains why the Bayes factor rejects the null. Hence, the null hypothesis may still be within a credible interval, but its likelihood is overshadowed by the alternative. It also indicates that the null hypothesis is rejected substantially.
3.4. Connection to the UMPBT
Considering Case 2, now treated as the UMPBT, we construct the corresponding ‘group’ sequential Bayesian test with a certain ‘termination’ point. However, we note that, in the current context, the threshold of the Bayes factor to reject the null hypothesis takes the form for some . By Lemma 1 in ref. [6], the UMPBT() is constructed by finding that and
which can be solved numerically. For instance, by applying the suggestion of the Jeffrey’s evidence level shown in Table 2, we may reject the null hypothesis if , so that in this case.
- (I)
- Case of non-informative (uniform) prior distribution
Since the ‘termination’ point of Case 2 with the uniform prior is the 17th-group data point, we have and , and . We use and obtain the corresponding . At this value of and , whenever 119 or more of the patients enrolled in this fixed sample size study. Accordingly, for the fixed-sample-size test (with ), once the number of the observed side effects was , we would reject the null hypothesis, incurring a (classical) Type I error probability . The rejection region for this significance test also corresponds to the region for which the Bayes factor corresponding to the UMPBT() for all values of . With equal prior probabilities of the null and the alternative hypotheses we assumed previously (), the implied significant level of for this UMPBT approximately corresponds to the Bayesian test, resulting in
This serves as illustration of the fact that, unlike the posterior probabilities, the p-value or attained significant level is not a true measure of the strength of the evidence in favor/against the null hypothesis; see ref. [15] for a discussion of this point.
However, note that, if we choose Jeffrey’s evidence level of , the calculated and , which would not match the results obtained from the UMPBT procedure in this case. This may be attributed to the group sequential nature of the data and the uniform prior that we considered here.
- (II)
- Case of informative prior distribution
Since the ‘termination’ point of Case 2 with the informative prior is the 16th-group data point, we have and , and . Applying Jeffrey’s evidence level of with the corresponding (equivalent) fixed sample size of this UMPBT , we obtain the corresponding . With these values of and , whenever 110 or more of the patients enrolled in this fixed sample size study. Accordingly, for the fixed-sample-size test, once the number of the observed side effects was , we would reject the null hypothesis, incurring a (classical) Type I error . The rejection region for this significance test also corresponds to the region for which the Bayes factor corresponding to the UMPBT() for all values of . With equal prior probabilities of the null and the alternative hypotheses we assumed previously (), the implied significant level of for this UMPBT, which approximately corresponds to the Bayesian test, resulted in
3.5. Simulation Results
In this subsection, we provide the results of some simulations of Case 2 of and (see (15)), aimed at illustrating further the properties of the sequential Bayesian testing procedure. For the simulations, we utilize the method introduced in ref. [9] to simulate the values of the relevant stopping time for the corresponding -UMP test of versus , with an optimal sample size , as was determined by and for various alternative choices of . Relation (1) for was then utilized to express these hypotheses in terms of the relative risk . The simulated values of the data, along with , were then used to calculate the values of the Bayes factor in (16). Based on the standard shown in Table 2, would be rejected if the calculated . Throughout the 10,000 simulation runs, we assumed that the prior probability of is . The results are presented in Table 9, which provide the calculated power and probability of the Type I error of the test for different sample sizes (different ) and different priors choices for , as were used in the standard Bayesian test. To that end, we simulated the data under the alternative choices of (equivalently of ) of and as well. As can be seen, as the sample size increases, the power increases as expected, regardless of the priors. Moreover, the Type I error probability decreases with the sample size under a uniform prior, but increases when the sample size increases under informative prior, which indicates that the Bayes factors may exhibit sensitivity to prior specification, giving support to default type choices for a prior distribution of . One should also note that, when the value of the Bayes factor we calculated is much less than the critical value , the simulated value of Type I error probability is much less than the nominal value of that we predefined for the UMP test.
Table 9.
Power and Type I error probability of the standard Bayesian test for in and in of Case 2 for different sample sizes under different priors. ‘avg’ indicates the average of the corresponding simulated values.
The results from applying the Modified Bayesian Test in (22) for Case 2 with these choices of values of and with are presented in Table 10. The values of r and u in (22) were obtained according to (23). As we can see, when , we have an chance to correctly reject , with the average conditional Type I error probability being under the uniform prior. Compared with the result shown in Table 9, when , we have a chance to correctly reject , with the average posterior probability being under the uniform prior. Moreover, we note that these simulations resulted in a chance to reject the null, a chance to accept the null, and a chance to conclude with no decision when under the chosen informative prior. We attribute this particular outcome to the particular choice of the parameters of the prior distribution and the likely inadequacy of the sample size in this discrete case. To allow a comparison with an ‘objective-type’ of prior choice, we augmented the simulation study of the Uniform and the Informative type choices of priors, with the Jefferey’s prior for the Binomial case having (see Table 10). Clearly, further simulations are needed to explore the full spectrum of the potential sensitivity to the prior choices. However, for the sake of space, we leave that to future studies. To facilitate reproducibility, we provide the sample simulation code as Supplementary Material.
Table 10.
Decision and the conditional error probability of Bayesian modified hypotheses test for in and in of Case 2 in sample size under different priors. ‘avg’ indicates the average of the corresponding simulated values. ‘NA’ indicates No Answer.
4. Summary and Discussion
In this paper, we propose the standard Bayesian test and the modified Bayesian test of hypotheses concerning the relative risk, , of a two-arm clinical trial. In this context, the relative risk can be represented as —a parameter representing an event probability. Since the output of each observation collected is binary data, the sequential experimental process can be viewed as a sequential binomial process with a probability of . Note that, under the Bayesian framework, based on the SRP, each data point remains unaffected by the optimal stopping rule used in the process (unlike the ‘classical’ sequential test). Within the Bayesian framework, we consider the sequential testing of hypotheses concerning the parameter . By utilizing the conjugate beta-binomial model, we are able to obtain the corresponding decision ‘rule’ for the test at each observed data point as determined from the calculated value of the Bayes factor.
To deal with the values of the Bayes factors that are too close to the decision boundaries, Jefferey’s criteria for the evidentiary level for a ‘rejection of ’ are taken into consideration. Additionally, we consider a modification of the standard Bayesian test ref. [5], which includes a ‘no-decision’ region, and calculate the corresponding conditional Type I and Type II error probabilities and . To illustrate our approach and the methods discussed, we analyze the data presented in ref. [2] under the three different hypothesis testing cases and several different priors. It is worth pointing out that the informative prior we propose is designed to use the knowledge of the bounded-width probability. For each case with different priors, we analyze the calculated Bayes factor corresponding to each data point and make the decision within two Bayesian tests.
Specifically, we consider, in Case 2, the very same hypothesis testing problem as shown in ref. [2]. The result presented in ref. [2] indicates that the testing procedure should have been terminated at the 19th-group data point. However, the results obtained from the standard Bayesian test and the modified Bayesian test lead to a potentially earlier ‘termination’ at the 17th-group data point under the uniform prior and at the 16th-group data point under an informative prior. Furthermore, utilizing the modified Bayesian test with a uniform prior, at the 17th-group data point, the Type I error conditional probability is (given the data). Similarly, with the modified Bayesian test and under an informative prior, at the 16th-group data point, the conditional Type I error probability is (given the data). From Section 3.4, we find our ‘termination’ based on the Bayesian test with a uniform prior corresponding to a p-value of and with an informative prior corresponding to a p-value of from the frequentist point of view.
Case 3 has the very same hypothesis testing problem as shown in ref. [3]. The result presented in ref. [3] indicates that the testing procedure should have been terminated at the 18th-group data point. However, the results obtained from both the standard Bayesian test and the modified Bayesian test lead to the potentially earlier ‘termination’ at the 14th-group data point with a uniform prior and at the 15th-group data point with an informative prior. Furthermore, with the modified Bayesian test under a uniform prior, at the 14th-group data point, the conditional Type I error probability (given the data). Similarly, with the modified Bayesian test under an informative prior, at the 15th-group data point, the conditional Type I error probability is (given the data).
These cases illustrate the benefits of Bayesian testing in this context. Unlike the classical sequential test, which typically requires heavy calculations to determine ‘stopping boundaries’ as well as ‘decision boundaries’ (for a-priory specified Type I and Type II error probabilities), the Bayesian test is easy to use and implement, and it could obtain the test results in a relatively simple manner. Compared with the results obtained in refs. [2,3], we conclude that our results are consistent with their results. Although our decisions for Case 2 and Case 3 arrive slightly earlier than theirs, the values of the data points are similar.
Clearly, our Bayesian approach, which allows for a posterior-based analysis (as well as interpretations) of the available data as was collected, is simpler and, one could argue, more intuitive. It is attractive to those who prescribe to implementing the Bayesian paradigm, especially in clinical studies, where expert opinion and prior information is amply available. However, to others, the Bayesian framework and the testing procedure we advocate for here could potentially be seen as particularly limiting. The potential dependency on the choice of the prior distribution of the relevant parameters, the inclusion of the no-decision region as part of the terminal testing procedure, and the interpretation of the posterior measure of evidence could all seen as limitations to the practitioner. The reader may find the Discussion and the Rejoinder parts of ref. [5] as valuable resources for these considerations, as well as ref. [16] as a resource for a general discussion concerning the applicability of the Modified Bayesian test and the no-decision region in sequential clinical trail settings.
In addition, we are able to illustrate, using the very same data, the connection of our Bayesian test and the UMPBT. With this connection, the results we obtained from the Bayesian test can be approximately interpreted by the rejection region and the p-value from the frequentist perspective of the UMP. Moreover, by applying the modified Bayesian test, we could even obtain conditional error probabilities as well, which give us a measure of the data evidence in making decisions (of ‘Reject’ or ‘Accept’ ). Indeed, for such binary data hypothesis testing, utilizing Bayesian framework could provide a straightforward and easy path to analyze the data.
Supplementary Materials
The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/math14142486/s1, File S1: Simulation R code demo.
Author Contributions
Methodology, J.W. and B.B.; Software, J.W. and B.B.; Investigation, J.W.; Writing—review & editing, J.W. and B.B.; Supervision, B.B. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Data Availability Statement
The original data presented in the study are openly available in ref. [2].
Conflicts of Interest
The authors declare no conflicts of interest.
Appendix A. Case of Jeffrey’s (Non-Informative) Prior
In this section, we present the results obtained from the standard Bayesian test and the modified Bayesian test under the Jeffrey’s (non-informative) prior of the data from ref. [2].
Appendix A.1. The Standard Bayesian Test
Introduced in ref. [14], the Jeffrey’s prior, as considered an ‘objective’ or non-informative prior distribution, is for binomial distributions in our case. We calculated the Bayes factor and the corresponding posterior probability , for the hypotheses tests on the RR, , in the three cases we considered. The numerical results are provided in Table A1 below.
As can be seen, for the hypotheses of Case 1 and Cse 2, our Bayesian model led to the potentially earlier ‘termination’ at the 18th-group data point with a total of ‘events’, the observed side effects of which were noted in the treatment group with an estimated relative risk of . The resulting Bayes factors are in Case 1 and in Case 2, which is less than based on the standard shown in Table 2. This leads to a rejection of , , and the posterior probabilities of the null hypothesis and the alternative hypothesis given the data are
which are clearly smaller than the assumed prior probability of , being true (marked in a red line in Figure A1).
On the other hand, for the hypotheses of Case 3, the observation process would ‘terminate’ much earlier at the 14th-group data point with a total of ‘events’, the observed side effects of which were noted in the treatment group, leading to the estimated relative risk of . The resulting Bayes factor is , which is less than based on the standard shown in Table 2, and the posterior probabilities of the null hypothesis and the alternative hypothesis given the data are
which are clearly smaller than the assumed prior probability of being true (marked in a red line in Figure A1).
Figure A1.
The posterior probability of for the 24 data points of Table A1: (a) Case 1; (b) Case 2; (c) Case 3.
Table A1.
The values of Bayes factor and posterior probability of for each data point among the three different cases with the Jeffrey’s prior.
In conclusion, under Jeffrey’s prior, both Case 1 and Case 2 would reject and be ‘terminated’ at the 18th-group data point, and Case 3 would reject and ‘terminated’ at the 14th-group data point.
Appendix A.2. The Modified Bayesian Test
In the modified Bayesian Sequential test, we obtain the same conclusions that were obtained utilizing the standard Bayesian test for the hypotheses , (see Appendix A.1). The results corresponding to the Jeffrey’s prior distribution are shown in Figure A2, Table A2 and Table A3. The calculated conditional Type I and Type II error probabilities, and , corresponding to the ‘Reject’ or ‘Accept’ decision (see Table A3).
Figure A2.
The acceptance and rejection boundaries of for the 24 data points of Table A2: (a) Case 1; (b) Case 2; (c) Case 3.
Table A2.
The values of the Bayes factor and the boundaries based on it for each data point among different cases with Jeffrey’s prior.
Table A3.
The decision and corresponding conditional and for each data point among different cases with Jeffrey’s prior (Decision: ‘R’ indicates Reject, ‘A’ indicates Accept, ‘ND’ indicates No Decision, ‘NA’ indicates No Answer).
References
- Lee, G.M.; Greene, S.K.; Weintraub, E.S.; Baggs, J.; Kulldorff, M.; Fireman, B.H.; Baxter, R.; Jacobsen, S.J.; Irving, S.; Daley, M.F.; et al. H1N1 and seasonal influenza vaccine safety in the vaccine safety datalink project. Am. J. Prev. Med. 2011, 41, 121–128. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Silva, I.R.; Kulldorff, M.; Katherine Yih, W. Optimal alpha spending for sequential analysis with binomial data. J. R. Stat. Soc. Ser. B Stat. Methodol. 2020, 82, 1141–1164. [Google Scholar] [CrossRef] [Scilit]
- Silva, I.R.; Zhuang, Y. Bounded-width confidence interval following optimal sequential analysis of adverse events with binary data. Stat. Methods Med. Res. 2022, 31, 2323–2337. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Berger, J.O.; Wolpert, R.L. The Likelihood Principle; IMS: Waite Hill, OH, USA, 1988. [Google Scholar]
- Berger, J.O.; Boukai, B.; Wang, Y. Unified frequentist and Bayesian testing of a precise hypothesis. Stat. Sci. 1997, 12, 133–160. [Google Scholar] [CrossRef] [Scilit]
- Johnson, V.E. Uniformly most powerful Bayesian tests. Ann. Stat. 2013, 41, 1716. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nikooienejad, A.; Johnson, V.E. On the existence of uniformly most powerful bayesian tests with application to non-central chi-squared tests. Bayesian Anal. 2021, 16, 93. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Silva, I.R.; Maro, J.; Kulldorff, M. Exact sequential test for clinical trials and post-market drug and vaccine safety surveillance with Poisson and binary data. Stat. Med. 2021, 40, 4890–4913. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, J.; Boukai, B. Early detection of treatment’s acute side effect: A sequential approach: J. Wang, B. Boukai. Stat. Pap. 2026, 67, 22. [Google Scholar] [CrossRef] [Scilit]
- Kulldorff, M.; Davis, R.L.; Kolczak, M.; Lewis, E.; Lieu, T.; Platt, R. A maximized sequential probability ratio test for drug and vaccine safety surveillance. Seq. Anal. 2011, 30, 58–78. [Google Scholar] [CrossRef] [Scilit]
- Jeffreys, H. Theory of Probability, 3rd ed.; Clarendon Press: Oxford, UK, 1961. [Google Scholar]
- Kass, R.E.; Raftery, A.E. Bayes factors. J. Am. Stat. Assoc. 1995, 90, 773–795. [Google Scholar] [CrossRef]
- Berger, J.O. Statistical Decision Theory and Bayesian Analysis; Springer Science & Business Media: New York, NY, USA, 2013. [Google Scholar]
- Jeffreys, H. An invariant form for the prior probability in estimation problems. Proc. R. Soc. Lond. Ser. A Math. Phys. Sci. 1946, 186, 453–461. [Google Scholar] [CrossRef] [Scilit]
- Sellke, T.; Bayarri, M.J.; Berger, J.O. Calibration of ρ values for testing precise null hypotheses. Am. Stat. 2001, 55, 62–71. [Google Scholar] [CrossRef] [Scilit]
- Berger, J.O.; Boukai, B.; Wang, Y. Simultaneous Bayesian-frequentist sequential testing of nested hypotheses. Biometrika 1999, 86, 79–92. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.






