Next Article in Journal
Score Cloud Analysis for Rule-Aware Ranking Robustness Under Discrete Judgment Uncertainty
Previous Article in Journal
A Simulation-Based Modified Singular Spectrum Analysis Framework for Signal Extraction and the Exploration of Structured Nonlinear Temporal Behaviour
Previous Article in Special Issue
Central Limit Theorem of the Recursive Estimate of Density Function Under Randomly Censored Data
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Modified Chebyshev Inequality and Its Appropriateness for Nonparametric Testing

by
Markus Neuhäuser
Department of Mathematics, Informatics, and Technology, RheinAhrCampus, Koblenz University of Applied Sciences, Joseph-Rovan-Allee 2, 53424 Remagen, Germany
Stats 2026, 9(5), 88; https://doi.org/10.3390/stats9050088
Submission received: 21 July 2026 / Revised: 20 August 2026 / Accepted: 22 August 2026 / Published: 24 August 2026
(This article belongs to the Special Issue Nonparametric Inference: Methods and Applications)

Abstract

Sample sizes might be small in a variety of applications. Many classical and often-applied statistical methods rely on asymptotic approximations and, therefore, should not be used in the case of small sample sizes. Approaches that can be applied to small data sets often involve computer-intensive methods such as permutation tests and bootstrapping. As an alternative, a modified Chebyshev inequality was proposed by Beasley et al. (Applied Statistics 2004; 53, 95–108) and suggested for nonparametric testing. Here, it is shown that this modified Chebyshev inequality does not hold in general and, therefore, it should not be used to construct a statistical test. Available nonparametric tests that can be recommended for small sample sizes are discussed.

1. Introduction

Even in the era of big data, small data sets are common. There are ethical, financial and practical reasons and constrains which lead to limited sample sizes. For instance, studies with endangered species often have small sample sizes. Moreover, sample sizes for subgroup analyses might be small.
For parametric statistical methods one usually has to assume that the underlying data are—at least approximatively—normally distributed. However, this requirement is often not justified in practical applications [1,2], and cannot be verified based on small samples [3].
Therefore, nonparametric statistical methods are often recommended for small sample sizes because they require no normality assumption. Obviously, asymptotic methods should be avoided in the case of small samples. Therefore, exact permutation tests were suggested. These tests are flexible nonparametric alternatives, are powerful and have become a standard method of statistical inference [4]. The p-value is computed as the proportion of permutations with a test statistic at least as extreme as the value of the test statistic for the actually observed data. This approach can also be applied in non-standard situations for which standard methods are not available [5].
Permutation tests require the exchangeability of observations under the null hypothesis. This requirement is fulfilled for independent and identically distributed random variables, but also if any permutation of the random variables has the same joint distribution function. For more details on permutation tests and examples we refer to [4,5,6,7]. Bonnini et al. [4] pointed out that the improvement in computers’ power and speed made permutation tests readily available. Various software systems offer permutation tests, in R there are packages such as coin [8] where tests can also be carried out in case of ties. Permutation tests can also be performed approximately based on a simple random sample out of all possible permutations. Further approximations are based on the analytical moments of the exact permutation distribution [9] and on the tail of the distribution which can be approximated by a generalized Pareto distribution [10].
Instead to draw permutations (without replacement), bootstrap sampling (with replacement) is a further alternative. When both methods are available, the permutation tests are usually more powerful [4]. An exception are studies with very small sample size. In this case bootstrap tests might be less conservative because there are more possible bootstrap samples than permutations [11].
In the Behrens-Fisher problem the aim is to test for a difference in location although variances might differ between groups [12]. Then, different variances might occur under the null hypothesis, so that exchangeability is no longer fulfilled. Bootstrap tests are possible in the Behrens-Fisher problem because bootstrap samples can be drawn separately for the different samples [7]. However, permutation tests based on studentized test statistics might also be possible, even in case of small samples [4,12,13].
The aim of this paper is to reassess the validity and practical usefulness of so-called Chebby Checker procedures proposed by Beasley et al. [14,15]. It shall be shown that the modified Chebyshev inequality underlying the Chebby Checker 3 method (CC3) does not hold in general. The other Chebby Checker methods CC1 and CC2 are compared with alternative procedures. The following section describes the Chebby checker methods. Moreover, an example and a simulation study are presented in later sections, before some further methods are briefly reviewed in Section 5.

2. Chebby Checker Tests

Beasley et al. [14,15] introduced a different approach for nonparametric testing which they called Chebby checker method. Three different methods were introduced based on three variants of the Chebyshev inequality [15]. To be precise, the following three inequalities were used: the original Chebyshev inequality
P τ μ τ σ τ T 1 T 2 ,
a modified Chebyshev inequality proposed by DasGuptas [16]
P τ μ τ σ τ T 1 3 T 2 ,
and a new modified inequality. Beasley et al. [14] combined Chebyshev’s original inequality with a modification introduced by Saw et al. [17] and claimed
P τ μ τ σ τ T 1 N + 1 T 2
Beasley et al. [14] considered the case that the classical two-sample t statistic is used. Under the null hypothesis of equality the t statistic has a central t distribution with N − 2 degrees of freedom (df), where N is the total sample size. Consequently, we have µτ = 0 and σ τ = N 2 N 4 . Then, using the original Chebyshev inequality, one obtains σ τ 2 t o b s 2 = ( N 2 ) ( N 4 ) t o b s 2 as the p-value of the Chebby Checker 1 method (CC1), where t o b s is the observed value of the t statistic [15]. The Chebby Checker 2 method (CC2) is based on Das Gupta’s modification [16], so that one obtains σ τ 2 3 t o b s 2 = ( N 2 ) 3 ( N 4 ) t o b s 2 as the p-value [15].
The modified inequality introduced by Beasley et al. [14] yields ( N 2 ) / ( N 4 ) ( N + 1 ) t o b s 2 as the p-value. However, as the Chebby Checker 3 (CC3) method, Beasley et al. [14,15] recommended using the maximum of this p-value and the p-value of the classical t test based on the t distribution.
Let’s consider the case N = 49 and T = 1. Then, 1 N + 1 T 2 = 0.02 and, consequently, the modified Chebyshev inequality from Equation (1) gives
P τ μ τ σ τ 1 0.02 .
However, P τ μ τ σ τ 1 = P τ σ τ = P τ 47 45 = P τ 1.022 = 0.3120 > 0.02 for a random variable τ that has a t distribution with N − 2 = 47 degrees of freedom.
Thus, the modified Chebyshev inequality (1) does not hold in general. This is even easier to see for a standard normally distributed random variable τ. In that case, we have τ μ τ σ τ = τ and for our example with N = 49 and T = 1 we have P τ 1 = 0.3173 > 0.02 .
Beasley et al. [14,15] applied their modified Chebyshev inequality (1) for small sample sizes N. In such a case the inequality can hold. However, with T = 1 we need a value of N as small as 2 ( N N ) in order to get 1 N + 1 T 2 0.3173 . For T = 2, we have P τ 2 = 0.0455 for a standard normally distributed random variable τ. In this case we need N ≤ 4 ( N N ) in order to get 1 N + 1 T 2 0.0455 . To be precise, when N = 4 and T = 2, the modified Chebyshev inequality gives P τ 2 1 4 + 1 2 2 = 0.05 .
As shown by Beasley et al. [14] (Table 2), the p-value based on the modified Chebyshev inequality can be larger or smaller than the usually used p-value based on the t distribution. It depends on the values of N and tobs which of the two p-values is larger. The possibility that the p-value based on the modified Chebyshev inequality can be smaller than the one based on the t distribution also demonstrates that the modified Chebyshev inequality (1) cannot hold in general.
Since it was recommended to use the maximum of the two possible p-values based on the modified Chebyshev inequality and based on the t distribution [14,15], the resulting test called Chebby Checker 3 has no inflated type I error rate when the assumptions of the classical t test hold. However, in that case the t test can safely be applied without any modification.
Whether or not the assumptions of the t test hold, we do not recommend the Chebby Checker 3 approach introduced by Beasley et al. [14] because the underlying modified Chebyshev inequality (1) does not hold in general. However, the two remaining tests, Chebby Checker 1 and Chebby Checker 2, might be applied and will be illustrated using example data in the next section.

3. Example

Zar [18] (p. 131) presented blood-clotting times of adult rabbits treated with two different drugs called G and B, and sample sizes 7 and 6. The values in the drug G group were: 9.9, 9.0, 11.1, 9.6, 8.7, 10.4, and 9.5 min. In the drug B group the times 8.8, 8.4, 7.9, 8.7, 9.1, and 9.6 min were observed. The classical t test as presented by Zar [18], that is the two-sample t test assuming equal variances, gives tobs = 2.4765 and a p-value of 0.0308 based on the t distribution with 11 degrees of freedom.
Using the value tobs = 2.4765 and N = 13 one can obtain the p-values of the Chebby Checker tests. The Chebby checker method CC1 gives p = 0.1993 and the Chebby checker method CC2 gives p = 0.0664. Hence, the significance at α = 0.05 observed using the t test cannot be confirmed with the Chebby Checker tests CC1 and CC2.
If we carry out the Fisher-Pitman permutation test, i.e., the permutation test with the t statistic, we do not need to assume that the underlying data are normal; based on 13 7 = 1716 permutations the p-value of this test is 0.0297. A permutation test with the Wilcoxon rank sum also gives a significance at α = 0.05: the p-value is 0.0478. A bootstrap test with the t statistic based on 100,000 bootstrap samples yields p = 0.0315. Thus, the example reflects the low power of the Chebby Checker tests CC1 and CC2 noted by [14] and shown in the following section.

4. Simulation Study

A Monte Carlo simulation study was performed using R (version 4.5.2); 100,000 simulation runs were generated for each configuration. The Chebby Checker tests CC1 and CC2 are compared with the classical t test (i.e., the two-sample t test assuming equal variances) and the Fisher-Pitman permutation test for the sample sizes of the example, i.e., n 1 = 6 and n 2 = 7, where n i denotes the sample size in group i. The exact Wilcoxon rank-sum test (i.e., a permutation test with the rank sum) as well as a permutation test based on the difference of 20% trimmed means are also included in the comparison. The trimmed mean is a robust location estimator. In additional simulations, balanced sample sizes between 5 and 10 per group and the larger sample sizes n 1 =   n 2 = 50 are examined. All tests are performed two-sided with the nominal significance α = 0.05, and five different continuous distributions are investigated (see Table 1).
Table 1 displays the simulated actual type I error rates, thus the simulation evaluates whether the tests maintain the nominal significance level (α = 0.05) when no effect is present. The size of the Fisher-Pitman permutation test and the trimmed-mean permutation test are very close to the nominal significance level, whereas the t test is somewhat conservative for skewed distributions. The Wilcoxon rank-sum test is also conservative. However, the Chebby Checker methods CC1 and CC2 are very conservative as already noted by Beasley et al. [14,15]. This marked conservatism also holds for larger sample sizes: With n 1 =   n 2 = 50 and standard normally distributed data, the actual type I error rate for CC1 is <0.001 and the one for CC2 is 0.001, whereas the t test has a size of 0.050 in this case.
Results for balanced sample sizes between 5 and 10 per group are displayed in Supplementary Table S1. The results are similar with the exception that the exact Wilcoxon rank-sum test is hardly conservative for n 1 =   n 2 = 8   and   α = 0.05 . The reason is that the conservatism of the rank sum depends on the pattern of the steps of its distribution function, and for n 1 =   n 2 = 8 the distribution function is very close to 0.975 between two steps [19].
The statistical power is evaluated under the alternative hypothesis, where a true location shift of 2 is introduced between groups, see Table 2. The Chebby Checker tests CC1 and CC2 have a much lower power than the t test and the Fisher-Pitman permutation test. In particular, CC1 has a low power which is not surprising because this test is extremely conservative. However, although CC2 is less conservative and more powerful than CC1, it cannot be recommended because competitive tests are much more powerful. The trimmed-mean permutation test is more powerful than the Fisher-Pitman permutation test for the skewed distribution, but less powerful for the investigated symmetric distributions. Results for balanced sample sizes between 5 and 10 are similar (see Supplementary Table S2). For the scenarios with n 1 =   n 2 8 , the exact Wilcoxon rank-sum test does not suffer much from conservatism and has a competitive power.

5. Further Methods

As mentioned above, the Chebby Checker tests CC1 and CC2 cannot be recommended. Therefore, Beasley et al. [14,15] proposed the Chebby Checker test CC3. However, this test cannot be recommended because the underlying modified Chebyshev inequality does not hold in general. Beasley et al. [14] particularly proposed their method for small sample sizes in combination with small nominal significance levels lower than 5%. In these cases, it may be not possible to obtain a significance when performing permutations tests. For instance, when assuming a significance level of α = 0.005, we have 1/α = 200, so that at least 200 permutations are needed for a minimal obtainable p-value of 0.005, or at least 400 permutations when applying a two-tailed test with a symmetrically distributed test statistic. In a two-group situation with balanced sample sizes, one would need at least five observations per group to achieve more than 200 permutations, or six observations per group to achieve more than 400 permutations. However, sample sizes less than five might be rare and are sometimes generally not suggested [20].
Permutation tests might be conservative, especially for small sample sizes. However, the Fisher-Pitman permutation test and the 20% trimmed-mean permutation test are hardly conservative even for small sample sizes such as 5 per group. In contrast, the Wilcoxon rank-sum test can be conservative. This conservatism can lead to a loss of power. The mid-p-value [21] was proposed to reduce the conservatism, it is calculated by subtracting half of the probability of the observed value of the test statistic from the p-value. However, this approach cannot guarantee the nominal significance level. Moreover, it would not help to reduce the conservatism when applying CC1 or CC2 using the continuously distributed t statistic.
The approach to apply a permutation test in the Behrens-Fisher problem can also result in a test that might not control the significance level. For this case, Francis and Manly [22] suggested a bootstrap calibration for a test that can have an excessive size. The bootstrap calibration changes the significance criterion, if necessary, so that actual size and nominal significance level coincide. A bootstrap calibration can also be used to determine whether a test is reliable [23].
There can be situations where permutation or bootstrap tests are not available. An example might be a nonparametric rank-based procedure for factorial designs [13,24]. How can one assess whether such an approach is reliable when a test is applied to small or sparse data so that it is unclear how large the actual type I error rate might be? This question was addressed with a simulation study in a recently published example [25]. For a nonparametric rank-based procedure for a two-factor model, a simulation study showed that the size could be as large as 11% for a nominal significance level of α = 5%, when sample sizes are as small as observed. However, an observed p-value was 0.0059. Since this p-value is strikingly smaller than 0.05, it might be significant although the test is liberal. In a simulation study, it was counted how often a p-value is smaller than or equal to the observed p-value 0.0059, under the null hypothesis, and for sample sizes as small as observed. In the scenarios considered in the simulation, the estimated proportion of p-values smaller than or equal to the observed p-value (0.0059) was between 0.039 to 0.042. Hence, it could be confirmed that there is a significance at the 5% level, when the observed p-value was 0.0059 [25].
By the way, the data set analyzed in [25] is another example where small or sparse data can occur: When bones or other remains from prehistoric times such as the Paleolithic are examined, the number of available finds and therefore the sample sizes are usually small.

6. Discussion and Conclusions

For nonparametric testing, Beasley et al. [14,15] proposed a modified Chebyshev inequality. However, as shown above, this modified Chebyshev inequality does not hold in general and, therefore, it should not be used to construct a statistical test. The Chebby checker methods CC1 and CC2 cannot be recommend due to their conservatism and low power. Alternative methods in case of small sample sizes are permutation and bootstrap tests. Due to fast computers, user-friendly software and efficient algorithms, these more computer-intensive methods are readily available nowadays, not only for small samples. Consequently, these methods are not limited to small samples and can be recommended more general, in particular for clinical trials and other studies where experimental units are randomized to different groups [26,27].
Hence, exact tests can be performed more widely. When the exact p-value can easily be obtained, as e.g., when analysing a 2 × 2 table with Pearson’s χ2 test, there is no reason to approximate it [28]. Permutation tests can also be applied for complex designs [29,30]. One option is to permute residuals. For example, when a response variable might depend on both, a continuous covariate and a categorical factor, an analysis of covariance can be applied. However, one could also perform a linear regression (without considering the categorical variable) and compare the resulting residuals between the different categories with a permutation test. Similar procedures are also possible for more complex models [31,32] and have a relatively large power [32]. These methods can be carried out using an R package called RRPP (residual randomization in permutation procedures) [33].
Bootstrap tests are a further approach with some additional applications due to the sampling with replacement [7]. However, sometimes neither a permutation nor a bootstrap test is possible. In that case, a simulation study can help to assess the reliability of the applied method.
Other issues related to small data sets are power and the positive predictive value, that is the proportion of true positive results among all positive results. When sample sizes are small both, the power and the positive predictive value, might be small. With regard to these issues and ways how the power might be increased without changing the sample size, it is referred to Neuhäuser and Ruxton [7].

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/stats9050088/s1, Table S1: Simulated type I error rates, i.e., the proportions of false rejections when there is no difference between groups, of different tests for different balanced sample sizes (two-sided α = 0.05, based on 10,000 simulation runs); Table S2: Simulated power, i.e., the proportions of correct rejections when there is a difference between groups, of different tests for different balanced sample sizes, group 2 has the same distribution as group 1 except for a location shift of 2 (two-sided α = 0.05, based on 10,000 simulation runs); R code.

Funding

This work was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation)—project number 569634273.

Data Availability Statement

R code for the analysis of the example data and the simulation study is available as Supplementary Material. The R code also includes the raw data of the example.

Acknowledgments

The authors thank the anonymous reviewers for their valuable comments and suggestions.

Conflicts of Interest

The author declares no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
CC1Chebby Checker method 1 (according to [15])
CC2Chebby Checker method 2 (according to [15])
CC3Chebby Checker method 3 (according to [15])
RRPPresidual randomization in permutation procedures

References

  1. Büning, H. Robust analysis of variance. J. Appl. Stat. 1997, 24, 319–332. [Google Scholar] [CrossRef] [Scilit]
  2. Nanna, M.J.; Sawilowsky, S.S. Analysis of Likert scale data in disability and medical rehabilitation research. Psychol. Methods 1998, 3, 55–67. [Google Scholar] [CrossRef] [Scilit]
  3. Bellara, C.A.; Julien, M.; Hanley, J.A. Normal approximations to the distributions of the Wilcoxon statistics: Accurate to what N? Graphical insights. J. Stat. Educ. 2010, 18, 1. [Google Scholar] [CrossRef] [Scilit]
  4. Bonnini, S.; Assegie, G.M.; Trzcinska, K. Review about the permutation approach in hypothesis testing. Mathematics 2024, 12, 2617. [Google Scholar] [CrossRef] [Scilit]
  5. Manly, B.F.J. Randomization, Bootstrap and Monte Carlo Methods in Biology, 3rd ed.; Chapman & Hall/CRC: London, UK, 2007. [Google Scholar]
  6. Bonnini, S.; Corain, L.; Marozzi, M.; Salmaso, L. Nonparametric Hypothesis Testing: Rank and Permutation Methods with Applications in R; Wiley: Chichester, UK, 2014. [Google Scholar]
  7. Neuhäuser, M.; Ruxton, G.D. The Statistical Analysis of Small Data Sets; Oxford University Press: Oxford, UK, 2024. [Google Scholar]
  8. Hothorn, T.; Hornik, K.; van de Weil, M.A.; Zeileis, A. Implementing a class of permutation tests: The coin package. J. Stat. Softw. 2008, 28, 1–23. [Google Scholar] [CrossRef] [Scilit]
  9. Zhou, C.; Wang, H.J.; Wang, Y.M. Efficient moment-based permutation tests. Adv. Neural Inf. Process. Syst. 2009, 22, 2277–2285. [Google Scholar] [PubMed]
  10. Knijnenburg, T.A.; Wessels, L.F.A.; Reinders, M.J.T.; Shmulevich, I. Fewer permutations, more accurate P-values. Bioinformatics 2009, 25, i161–i168. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Neuhäuser, M.; Jöckel, K.H. A bootstrap test for the analysis of microarray experiments with a very small number of replications. Appl. Bioinform. 2006, 5, 173–179. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Neubert, K.; Brunner, E. A studentized permutation test for the nonparametric Behrens-Fisher problem. Comput. Stat. Data Anal. 2007, 51, 5192–5204. [Google Scholar] [CrossRef] [Scilit]
  13. Brunner, E.; Bathke, A.C.; Konietschke, F. Rank and Pseudo-Rank Procedures for Independent Observations in Factorial Designs; Springer: Cham, Switzerland, 2018. [Google Scholar]
  14. Beasley, T.M.; Page, G.P.; Brand, J.P.L.; Gadbury, G.L.; Mountz, J.D.; Allison, D.B. Chebyshev’s inequality for nonparametric testing with small N and α in microarray research. Appl. Stat. 2004, 53, 95–108. [Google Scholar] [CrossRef] [Scilit]
  15. Beasley, T.M.; Brand, J.P.L.; Long, J.D. The use of nonparametric procedures in the statistical analysis of microarray data. In DNA Microarrays and Related Genomics Techniques: Design, Analysis, and Interpretation of Experiments; Allison, D.B., Page, G.P., Beasley, T.M., Edwards, J.W., Eds.; Taylor & Francis/CRC: Boca Raton, FL, USA, 2006; pp. 245–265. [Google Scholar]
  16. DasGuptas, A. Best constants in Chebyschev inequalities with various applications. Metrika 2000, 51, 185–200. [Google Scholar] [CrossRef] [Scilit]
  17. Saw, J.G.; Yang, M.C.K.; Mo, T.C. Chebyshev inequality with estimated mean and variance. Am. Stat. 1984, 38, 130–132. [Google Scholar] [CrossRef] [Scilit]
  18. Zar, J.H. Biostatistical Analysis; Pearson Education: Upper Saddle River, NJ, USA, 2010. [Google Scholar]
  19. Neuhäuser, M. Nonparametric Statistical Tests: A Computational Approach; CRC Press: Boca Raton, FL, USA, 2012. [Google Scholar]
  20. Curtis, M.J.; Alexander, S.; Cirino, G.; Docherty, J.R.; George, C.H.; Giembycz, M.A.; Hoyer, D.; Insel, P.A.; Izzo, A.A.; Ji, Y.; et al. Experimental design and analysis and their reporting II: Updated and simplified guidance for authors and peer reviewers. Br. J. Pharmacol. 2018, 175, 987–993. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Lancaster, H.O. Significance tests in discrete distributions. J. Am. Stat. Assoc. 1961, 56, 223–234. [Google Scholar] [CrossRef]
  22. Francis, R.I.C.C.; Manly, B.F.J. Bootstrap calibration to improve the reliability of tests to compare sample means and variances. Environmetrics 2001, 12, 713–729. [Google Scholar] [CrossRef] [Scilit]
  23. Manly, B.F.J.; Francis, R.I.C.C. Testing for mean and variance differences with samples from distributions that may be non-normal with unequal variances. J. Stat. Comput. Simul. 2002, 72, 633–646. [Google Scholar] [CrossRef] [Scilit]
  24. Brunner, E.; Konietschke, F.; Pauly, M.; Puri, M.L. Rank-based procedures in factorial designs: Hypotheses about non-parametric treatment effects. J. R. Stat. Soc. Ser. B Stat. Methodol. 2017, 79, 1463–1485. [Google Scholar] [CrossRef] [Scilit]
  25. Neuhäuser, M. Violence and warfare in the European Mesolithic and Paleolithic: A re-analysis. Commun. Stat.-Case Stud. Data Anal. Appl. 2026. Epub ahead of printing. [Google Scholar] [CrossRef] [Scilit]
  26. Berger, V.W. Pros and cons of permutation tests in clinical trials. Stat. Med. 2000, 19, 1319–1328. [Google Scholar] [CrossRef]
  27. Berger, V.W.; Lunneborg, C.; Ernst, M.D.; Levine, J.G. Parametric analyses in randomized clinical trials. J. Mod. Appl. Stat. Methods 2002, 1, 74–82. [Google Scholar] [CrossRef] [Scilit]
  28. Neuhäuser, M.; Ruxton, G.D. The choice between Pearson’s χ2 test and Fisher’s exact test for 2x2 tables. Pharm. Stat. 2025, 24, e70012. [Google Scholar] [PubMed]
  29. Pesarin, F. Multivariate Permutation Tests; Wiley: New York, NY, USA, 2001. [Google Scholar]
  30. Pesarin, F.; Salmaso, L. Permutation Tests for Complex Data: Theory, Applications and Software; Wiley: New York, NY, USA, 2010. [Google Scholar]
  31. Oja, H. On permutation tests in multiple regression and analysis of covariance problems. Aust. J. Stat. 1987, 29, 91–100. [Google Scholar] [CrossRef] [Scilit]
  32. Anderson, M.J.; ter Braak, C.J.F. Permutation tests for multi-factorial analysis of variance. J. Stat. Comput. Simul. 2003, 73, 85–113. [Google Scholar] [CrossRef] [Scilit]
  33. Collyer, M.L.; Adams, D.C. RRPP: An R package for fitting linear models to high-dimensional data using residual randomization. Methods Ecol. Evol. 2018, 9, 1772–1779. [Google Scholar] [CrossRef] [Scilit]
Table 1. Simulated type I error rates, i.e., the proportions of false rejections when there is no difference between groups, of different tests for sample sizes n 1   =   6 and n 2   =   7 (two-sided α = 0.05).
Table 1. Simulated type I error rates, i.e., the proportions of false rejections when there is no difference between groups, of different tests for sample sizes n 1   =   6 and n 2   =   7 (two-sided α = 0.05).
Distributiont TestFisher-Pitman
Permutation Test
CC1CC2Trimmed-Mean
Permutation Test
Exact Rank-
Sum Test *
Standard normal0.0470.0500.0010.0170.0500.035
Lognormal with µ = 0, σ = 10.0220.0500.00010.0080.0490.035
Exponential with rate 10.0320.0500.00030.0110.0500.035
Chi-square with df = 10.0220.0490.00020.0070.0490.035
Uniform on [0, 4]0.0500.0500.0010.0190.0500.035
* values are exactly determined for the rank test [19].
Table 2. Simulated power, i.e., the proportions of correct rejections when there is a difference between groups, of different tests for sample sizes n 1   =   6 and n 2   =   7 , group 2 has the same distribution as group 1 except for a location shift of 2 (two-sided α = 0.05).
Table 2. Simulated power, i.e., the proportions of correct rejections when there is a difference between groups, of different tests for sample sizes n 1   =   6 and n 2   =   7 , group 2 has the same distribution as group 1 except for a location shift of 2 (two-sided α = 0.05).
Distributiont TestFisher-Pitman
Permutation Test
CC1CC2Trimmed-Mean
Permutation Test
Exact Rank-
Sum Test
Standard normal0.8940.9030.1990.7570.8910.844
Lognormal with µ = 0, σ = 10.5720.6320.1260.4420.7130.617
Exponential with rate 10.8630.8900.3480.7740.9290.836
Chi-square with df = 10.7020.7410.2100.5820.8160.706
Uniform on [0, 4]0.8000.8010.0890.5850.7440.687
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Neuhäuser, M. A Modified Chebyshev Inequality and Its Appropriateness for Nonparametric Testing. Stats 2026, 9, 88. https://doi.org/10.3390/stats9050088

AMA Style

Neuhäuser M. A Modified Chebyshev Inequality and Its Appropriateness for Nonparametric Testing. Stats. 2026; 9(5):88. https://doi.org/10.3390/stats9050088

Chicago/Turabian Style

Neuhäuser, Markus. 2026. "A Modified Chebyshev Inequality and Its Appropriateness for Nonparametric Testing" Stats 9, no. 5: 88. https://doi.org/10.3390/stats9050088

APA Style

Neuhäuser, M. (2026). A Modified Chebyshev Inequality and Its Appropriateness for Nonparametric Testing. Stats, 9(5), 88. https://doi.org/10.3390/stats9050088

Article Metrics

Back to TopTop