1. Introduction
Even in the era of big data, small data sets are common. There are ethical, financial and practical reasons and constrains which lead to limited sample sizes. For instance, studies with endangered species often have small sample sizes. Moreover, sample sizes for subgroup analyses might be small.
For parametric statistical methods one usually has to assume that the underlying data are—at least approximatively—normally distributed. However, this requirement is often not justified in practical applications [
1,
2], and cannot be verified based on small samples [
3].
Therefore, nonparametric statistical methods are often recommended for small sample sizes because they require no normality assumption. Obviously, asymptotic methods should be avoided in the case of small samples. Therefore, exact permutation tests were suggested. These tests are flexible nonparametric alternatives, are powerful and have become a standard method of statistical inference [
4]. The
p-value is computed as the proportion of permutations with a test statistic at least as extreme as the value of the test statistic for the actually observed data. This approach can also be applied in non-standard situations for which standard methods are not available [
5].
Permutation tests require the exchangeability of observations under the null hypothesis. This requirement is fulfilled for independent and identically distributed random variables, but also if any permutation of the random variables has the same joint distribution function. For more details on permutation tests and examples we refer to [
4,
5,
6,
7]. Bonnini et al. [
4] pointed out that the improvement in computers’ power and speed made permutation tests readily available. Various software systems offer permutation tests, in R there are packages such as coin [
8] where tests can also be carried out in case of ties. Permutation tests can also be performed approximately based on a simple random sample out of all possible permutations. Further approximations are based on the analytical moments of the exact permutation distribution [
9] and on the tail of the distribution which can be approximated by a generalized Pareto distribution [
10].
Instead to draw permutations (without replacement), bootstrap sampling (with replacement) is a further alternative. When both methods are available, the permutation tests are usually more powerful [
4]. An exception are studies with very small sample size. In this case bootstrap tests might be less conservative because there are more possible bootstrap samples than permutations [
11].
In the Behrens-Fisher problem the aim is to test for a difference in location although variances might differ between groups [
12]. Then, different variances might occur under the null hypothesis, so that exchangeability is no longer fulfilled. Bootstrap tests are possible in the Behrens-Fisher problem because bootstrap samples can be drawn separately for the different samples [
7]. However, permutation tests based on studentized test statistics might also be possible, even in case of small samples [
4,
12,
13].
The aim of this paper is to reassess the validity and practical usefulness of so-called Chebby Checker procedures proposed by Beasley et al. [
14,
15]. It shall be shown that the modified Chebyshev inequality underlying the Chebby Checker 3 method (CC3) does not hold in general. The other Chebby Checker methods CC1 and CC2 are compared with alternative procedures. The following section describes the Chebby checker methods. Moreover, an example and a simulation study are presented in later sections, before some further methods are briefly reviewed in
Section 5.
2. Chebby Checker Tests
Beasley et al. [
14,
15] introduced a different approach for nonparametric testing which they called Chebby checker method. Three different methods were introduced based on three variants of the Chebyshev inequality [
15]. To be precise, the following three inequalities were used: the original Chebyshev inequality
a modified Chebyshev inequality proposed by DasGuptas [
16]
and a new modified inequality. Beasley et al. [
14] combined Chebyshev’s original inequality with a modification introduced by Saw et al. [
17] and claimed
Beasley et al. [
14] considered the case that the classical two-sample
t statistic is used. Under the null hypothesis of equality the
t statistic has a central
t distribution with
N − 2 degrees of freedom (df), where
N is the total sample size. Consequently, we have
µτ = 0 and
. Then, using the original Chebyshev inequality, one obtains
as the
p-value of the Chebby Checker 1 method (CC1), where
is the observed value of the
t statistic [
15]. The Chebby Checker 2 method (CC2) is based on Das Gupta’s modification [
16], so that one obtains
as the
p-value [
15].
The modified inequality introduced by Beasley et al. [
14] yields
as the
p-value. However, as the Chebby Checker 3 (CC3) method, Beasley et al. [
14,
15] recommended using the maximum of this
p-value and the
p-value of the classical
t test based on the
t distribution.
Let’s consider the case
N = 49 and
T = 1. Then,
and, consequently, the modified Chebyshev inequality from Equation (1) gives
However, for a random variable τ that has a t distribution with N − 2 = 47 degrees of freedom.
Thus, the modified Chebyshev inequality (1) does not hold in general. This is even easier to see for a standard normally distributed random variable τ. In that case, we have and for our example with N = 49 and T = 1 we have .
Beasley et al. [
14,
15] applied their modified Chebyshev inequality (1) for small sample sizes
N. In such a case the inequality can hold. However, with
T = 1 we need a value of
N as small as 2 (
) in order to get
. For
T = 2, we have
for a standard normally distributed random variable
τ. In this case we need
N ≤ 4 (
) in order to get
. To be precise, when
N = 4 and
T = 2, the modified Chebyshev inequality gives
.
As shown by Beasley et al. [
14] (Table 2), the
p-value based on the modified Chebyshev inequality can be larger or smaller than the usually used
p-value based on the
t distribution. It depends on the values of
N and
tobs which of the two
p-values is larger. The possibility that the
p-value based on the modified Chebyshev inequality can be smaller than the one based on the
t distribution also demonstrates that the modified Chebyshev inequality (1) cannot hold in general.
Since it was recommended to use the maximum of the two possible
p-values based on the modified Chebyshev inequality and based on the
t distribution [
14,
15], the resulting test called Chebby Checker 3 has no inflated type I error rate when the assumptions of the classical
t test hold. However, in that case the
t test can safely be applied without any modification.
Whether or not the assumptions of the
t test hold, we do not recommend the Chebby Checker 3 approach introduced by Beasley et al. [
14] because the underlying modified Chebyshev inequality (1) does not hold in general. However, the two remaining tests, Chebby Checker 1 and Chebby Checker 2, might be applied and will be illustrated using example data in the next section.
3. Example
Zar [
18] (p. 131) presented blood-clotting times of adult rabbits treated with two different drugs called G and B, and sample sizes 7 and 6. The values in the drug G group were: 9.9, 9.0, 11.1, 9.6, 8.7, 10.4, and 9.5 min. In the drug B group the times 8.8, 8.4, 7.9, 8.7, 9.1, and 9.6 min were observed. The classical
t test as presented by Zar [
18], that is the two-sample
t test assuming equal variances, gives
tobs = 2.4765 and a
p-value of 0.0308 based on the
t distribution with 11 degrees of freedom.
Using the value tobs = 2.4765 and N = 13 one can obtain the p-values of the Chebby Checker tests. The Chebby checker method CC1 gives p = 0.1993 and the Chebby checker method CC2 gives p = 0.0664. Hence, the significance at α = 0.05 observed using the t test cannot be confirmed with the Chebby Checker tests CC1 and CC2.
If we carry out the Fisher-Pitman permutation test, i.e., the permutation test with the
t statistic, we do not need to assume that the underlying data are normal; based on
permutations the
p-value of this test is 0.0297. A permutation test with the Wilcoxon rank sum also gives a significance at α = 0.05: the
p-value is 0.0478. A bootstrap test with the
t statistic based on 100,000 bootstrap samples yields
p = 0.0315. Thus, the example reflects the low power of the Chebby Checker tests CC1 and CC2 noted by [
14] and shown in the following section.
4. Simulation Study
A Monte Carlo simulation study was performed using R (version 4.5.2); 100,000 simulation runs were generated for each configuration. The Chebby Checker tests CC1 and CC2 are compared with the classical
t test (i.e., the two-sample
t test assuming equal variances) and the Fisher-Pitman permutation test for the sample sizes of the example, i.e.,
6 and
7, where
denotes the sample size in group
i. The exact Wilcoxon rank-sum test (i.e., a permutation test with the rank sum) as well as a permutation test based on the difference of 20% trimmed means are also included in the comparison. The trimmed mean is a robust location estimator. In additional simulations, balanced sample sizes between 5 and 10 per group and the larger sample sizes
are examined. All tests are performed two-sided with the nominal significance α = 0.05, and five different continuous distributions are investigated (see
Table 1).
Table 1 displays the simulated actual type I error rates, thus the simulation evaluates whether the tests maintain the nominal significance level (α = 0.05) when no effect is present. The size of the Fisher-Pitman permutation test and the trimmed-mean permutation test are very close to the nominal significance level, whereas the
t test is somewhat conservative for skewed distributions. The Wilcoxon rank-sum test is also conservative. However, the Chebby Checker methods CC1 and CC2 are very conservative as already noted by Beasley et al. [
14,
15]. This marked conservatism also holds for larger sample sizes: With
and standard normally distributed data, the actual type I error rate for CC1 is <0.001 and the one for CC2 is 0.001, whereas the
t test has a size of 0.050 in this case.
Results for balanced sample sizes between 5 and 10 per group are displayed in
Supplementary Table S1. The results are similar with the exception that the exact Wilcoxon rank-sum test is hardly conservative for
The reason is that the conservatism of the rank sum depends on the pattern of the steps of its distribution function, and for
the distribution function is very close to 0.975 between two steps [
19].
The statistical power is evaluated under the alternative hypothesis, where a true location shift of 2 is introduced between groups, see
Table 2. The Chebby Checker tests CC1 and CC2 have a much lower power than the
t test and the Fisher-Pitman permutation test. In particular, CC1 has a low power which is not surprising because this test is extremely conservative. However, although CC2 is less conservative and more powerful than CC1, it cannot be recommended because competitive tests are much more powerful. The trimmed-mean permutation test is more powerful than the Fisher-Pitman permutation test for the skewed distribution, but less powerful for the investigated symmetric distributions. Results for balanced sample sizes between 5 and 10 are similar (see
Supplementary Table S2). For the scenarios with
, the exact Wilcoxon rank-sum test does not suffer much from conservatism and has a competitive power.
5. Further Methods
As mentioned above, the Chebby Checker tests CC1 and CC2 cannot be recommended. Therefore, Beasley et al. [
14,
15] proposed the Chebby Checker test CC3. However, this test cannot be recommended because the underlying modified Chebyshev inequality does not hold in general. Beasley et al. [
14] particularly proposed their method for small sample sizes in combination with small nominal significance levels lower than 5%. In these cases, it may be not possible to obtain a significance when performing permutations tests. For instance, when assuming a significance level of α = 0.005, we have 1/α = 200, so that at least 200 permutations are needed for a minimal obtainable
p-value of 0.005, or at least 400 permutations when applying a two-tailed test with a symmetrically distributed test statistic. In a two-group situation with balanced sample sizes, one would need at least five observations per group to achieve more than 200 permutations, or six observations per group to achieve more than 400 permutations. However, sample sizes less than five might be rare and are sometimes generally not suggested [
20].
Permutation tests might be conservative, especially for small sample sizes. However, the Fisher-Pitman permutation test and the 20% trimmed-mean permutation test are hardly conservative even for small sample sizes such as 5 per group. In contrast, the Wilcoxon rank-sum test can be conservative. This conservatism can lead to a loss of power. The mid-
p-value [
21] was proposed to reduce the conservatism, it is calculated by subtracting half of the probability of the observed value of the test statistic from the
p-value. However, this approach cannot guarantee the nominal significance level. Moreover, it would not help to reduce the conservatism when applying CC1 or CC2 using the continuously distributed
t statistic.
The approach to apply a permutation test in the Behrens-Fisher problem can also result in a test that might not control the significance level. For this case, Francis and Manly [
22] suggested a bootstrap calibration for a test that can have an excessive size. The bootstrap calibration changes the significance criterion, if necessary, so that actual size and nominal significance level coincide. A bootstrap calibration can also be used to determine whether a test is reliable [
23].
There can be situations where permutation or bootstrap tests are not available. An example might be a nonparametric rank-based procedure for factorial designs [
13,
24]. How can one assess whether such an approach is reliable when a test is applied to small or sparse data so that it is unclear how large the actual type I error rate might be? This question was addressed with a simulation study in a recently published example [
25]. For a nonparametric rank-based procedure for a two-factor model, a simulation study showed that the size could be as large as 11% for a nominal significance level of α = 5%, when sample sizes are as small as observed. However, an observed
p-value was 0.0059. Since this
p-value is strikingly smaller than 0.05, it might be significant although the test is liberal. In a simulation study, it was counted how often a
p-value is smaller than or equal to the observed
p-value 0.0059, under the null hypothesis, and for sample sizes as small as observed. In the scenarios considered in the simulation, the estimated proportion of
p-values smaller than or equal to the observed
p-value (0.0059) was between 0.039 to 0.042. Hence, it could be confirmed that there is a significance at the 5% level, when the observed
p-value was 0.0059 [
25].
By the way, the data set analyzed in [
25] is another example where small or sparse data can occur: When bones or other remains from prehistoric times such as the Paleolithic are examined, the number of available finds and therefore the sample sizes are usually small.
6. Discussion and Conclusions
For nonparametric testing, Beasley et al. [
14,
15] proposed a modified Chebyshev inequality. However, as shown above, this modified Chebyshev inequality does not hold in general and, therefore, it should not be used to construct a statistical test. The Chebby checker methods CC1 and CC2 cannot be recommend due to their conservatism and low power. Alternative methods in case of small sample sizes are permutation and bootstrap tests. Due to fast computers, user-friendly software and efficient algorithms, these more computer-intensive methods are readily available nowadays, not only for small samples. Consequently, these methods are not limited to small samples and can be recommended more general, in particular for clinical trials and other studies where experimental units are randomized to different groups [
26,
27].
Hence, exact tests can be performed more widely. When the exact
p-value can easily be obtained, as e.g., when analysing a 2 × 2 table with Pearson’s χ
2 test, there is no reason to approximate it [
28]. Permutation tests can also be applied for complex designs [
29,
30]. One option is to permute residuals. For example, when a response variable might depend on both, a continuous covariate and a categorical factor, an analysis of covariance can be applied. However, one could also perform a linear regression (without considering the categorical variable) and compare the resulting residuals between the different categories with a permutation test. Similar procedures are also possible for more complex models [
31,
32] and have a relatively large power [
32]. These methods can be carried out using an R package called RRPP (residual randomization in permutation procedures) [
33].
Bootstrap tests are a further approach with some additional applications due to the sampling with replacement [
7]. However, sometimes neither a permutation nor a bootstrap test is possible. In that case, a simulation study can help to assess the reliability of the applied method.
Other issues related to small data sets are power and the positive predictive value, that is the proportion of true positive results among all positive results. When sample sizes are small both, the power and the positive predictive value, might be small. With regard to these issues and ways how the power might be increased without changing the sample size, it is referred to Neuhäuser and Ruxton [
7].