Next Article in Journal
A Practical Framework for Incorporating Complex Survey Design in Bayesian Kernel Machine Regression
Previous Article in Journal
Assessing the Accuracy of Bootstrap-Based Standard Errors in Regression Models with Unobserved Heterogeneity
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Coverage and Precision of Net Promoter Score Confidence Intervals Across Sampling Distributions

1
Clinical and Translational Research Institute, Northeast Ohio Medical University, 4209 St. Rt 44, Rootstown, OH 44272, USA
2
Executive Education, Northeast Ohio Medical University, 4209 St Rt 44, Rootstown, OH 44272, USA
3
Department of Data Science, University of Mississippi Medical Center, 2500 N State St., Jackson, MS 39216, USA
*
Author to whom correspondence should be addressed.
Stats 2026, 9(2), 45; https://doi.org/10.3390/stats9020045
Submission received: 24 December 2025 / Revised: 14 April 2026 / Accepted: 16 April 2026 / Published: 21 April 2026

Abstract

The Net Promoter Score (NPS) is a widely used metric for customer loyalty in business. However, the current theoretical gaps in the literature suggest practical refinements for real-world applications. In this simulation study, we use an unbiased estimator of the variance for the sample NPS to examine coverage and width for three different confidence interval methods: Wald, bootstrap t, and adjusted Wald with weights corresponding to four underlying population distribution shapes: extreme (E), left-skewed (LS), triangular (T), and uniform (U). As the sample size increased, all methods approached the nominal 95% coverage rate with an exception for the extreme population; the adjusted Wald method with triangular and uniform weights is particularly robust among the representative population shapes examined. All adjusted Wald methods performed comparably in width, especially at a larger n. The confidence interval width depended on the population shape. Overall, the Wald and bootstrap t methods should be avoided at small sample sizes and are not recommended. Our methods raise awareness of the sampling distribution of the NPS statistic, provide a theoretical basis for an unbiased estimator of the variance, and assess reliable confidence interval construction. These results provide an informed application of NPS and lay the foundation for future methodological development.

1. Introduction

The Net Promoter Score (NPS) is a simple but powerful tool used to gauge customer loyalty in business contexts [1]. It is based on one question: “How likely are you to recommend our company/product/service to a friend or colleague?” The responses are given on an 11-point ordinal scale {0, 1, …, 10}, and customers are grouped into three classes: Promoters (those giving 9 or 10), Passives (those giving 7 or 8), or Detractors (those giving 0 up to and including 6). The counts for each of the three classes are taken from a sample survey, and percentages are computed. The NPS is defined to be:
NPS = % Promoters − % Detractors = (proportion of Promoters − proportion of Detractors) × 100.
Without a loss of generality, just the proportions can be used. In this case, we see that the NPS [ 1.00 , 1.00 ] . In terms of interpretation, an NPS of 0 means that Promoters and Detractors are proportionally balanced, while positive scores are “good”, demonstrating proportionally more Promoters than Detractors. While there are no widely accepted absolute benchmark values, and these vary among industries, many would consider an NPS over 0.50 to be excellent [2].
The NPS has seen rapid adoption as a key performance indicator of customer loyalty and satisfaction. In practice, organizations track the NPS over time and compare scores across customer segments, making a statistical inference about the NPS (e.g., confidence intervals and hypothesis tests) an important concern. Despite its popularity in business, marketing, healthcare, etc., uncertainty quantification is rarely reported in practice, and guidance on which confidence intervals are reliable for NPS’s categorical nature is limited. The NPS metric has faced scrutiny in the general and peer-reviewed literature. For example, researchers have questioned whether the NPS is truly superior to other customer metrics, with some finding that the NPS typically has no stronger positive association with business outcomes (like customer retention or revenue growth) than other loyalty or satisfaction metrics. Other recent research termed the NPS “ambiguous and unreliable if not accompanied by other tools,” since identical NPS values may mask differing customer sentiment compositions [3]. Rocks [4] gave a brief summary of several of the critiques of the NPS. An extensive and comprehensive historical literature review of the NPS can be found at Measuring U [5].
Because the NPS is so widely used for making strategic decisions across many organizations spanning various sectors, one might question, “how much can we trust this metric?” A way to think about assessing the NPS is to cast its interpretation in the context of estimation under a probability sampling design. For example, suppose that a large healthcare system is considering deploying a workforce development course to its constituents at great expense in terms of time and money. The 10,000 employees represent the population, while 20 individuals attending the inaugural course represent the sample. Hypothetically, if the entire population took the course, the true (and unknown) NPS would be the so-called parameter, using the parlance of a statistician. The NPS computed from the 20 individuals would be the statistic, our point estimate of the parameter. Much like a sampling poll, interest lies in inference on the parameter. How comfortable should leadership be in the system-wide success of the course with the NPS from an inaugural cohort of 20 employees? A single point estimate of NPS with no measure of uncertainty will not be very helpful to the leadership of the healthcare system and could even lead to risky conclusions and decisions. In practice, however, NPS is almost always reported as a single point estimate, with formal measures of uncertainty, such as confidence intervals, rarely if ever given. Providing the leadership with a range of plausible values for the NPS parameter, i.e., a confidence interval, would be useful in making better informed strategic planning, policy making, etc. A confidence interval shifts the NPS from a solitary, potentially misleading number into a more nuanced decision-support tool, enabling organizations to allocate resources, assess initiatives, and gauge risk with greater clarity and conviction.

1.1. Background and Theory

Because it is a difference of two proportions, the NPS itself can be thought of as a statistic derived from a sample from a three-category multinomial distribution (Detractor, Passive, and Promoter counts from a multinomial distribution with three categories). More formally, we have three possible mutually exclusive outcomes, with outcome probabilities p , p 0 , and p + , where these correspond to “success” probabilities for the Detractor, Passive, and Promoter outcomes, respectively. In addition, we assume we have n { 1 , 2 , } independent multinoulli trials, the outcome probabilities are all non-negative, and that p + p 0 + p + = 1 . If the random variables X , X 0 , and X + indicate the number of times the Detractor, Passive, and Promoter outcomes are observed over the n trials, respectively, then the vector ( X , X 0 , X + ) follows a multinomial distribution with parameters n and the vector ( p , p 0 , p + ) . While the trials are independent, their outcomes, X , X 0 , and X + , are dependent because they must sum to n. The probability mass function for the multinomial distribution in this case is:
n x ! x 0 ! x + ! p x p 0 x 0 p + x + for x , x 0 , x + 0 .
Marginally, the count for any single category in a k-category multinomial distribution with a fixed n is binomially distributed:
X i Binomial ( n , p i )
for i = −, 0, or + with probability mass function:
n x i ! ( n x i ) ! p i x i ( 1 p i ) n x i for x i 0 .
Now, suppose we survey a random sample of n respondents. Let X + Binomial ( n , p + ) and X Binomial ( n , p ) . Then, the sample NPS is:
NPS ^ = X + X n .
We now show important results for this statistic using basic probability theory [6].
Theorem 1.
NPS ^ is an unbiased estimator for NPS.
Proof. 
Let p ^ + = X + / n be the sample proportion of Promoters and p ^ = X / n be the sample proportion of Detractors. Therefore, NPS ^ = p ^ + p ^ .
Hence, the mean, or expected value, of NPS ^ is:
E [ NPS ^ ] = E [ p ^ + p ^ ] = E [ p ^ + ] E [ p ^ ] = p + p .
Theorem 2.
The variance of  NPS ^  for a sample of size n is:
Var [ NPS ^ ] = p + + p ( p + p ) 2 n .
Proof. 
It is known that Var [ p ^ i ] = p i ( 1 p i ) n for i = − or +, and that Cov [ p ^ + , p ^ ] = p + p n . Hence, the variance of NPS ^ is:
Var [ NPS ^ ] = Var [ p ^ + p ^ ] = Var [ p ^ + ] + Var [ p ^ ] 2 Cov [ p ^ + , p ^ ] = 1 n [ p + ( 1 p + ) + p ( 1 p ) + 2 p + p ] = 1 n [ p + + p ( p + p ) 2 ] .
We now consider how to derive an estimator for the variance of NPS ^ .
Theorem 3.
The estimator for the variance of NPS ^ is:
Var ^ 1 [ NPS ^ ] = p ^ + + p ^ ( p ^ + p ^ ) 2 n .
The estimator is biased for a small sample size, n, but the bias goes to zero as n becomes large.
Proof. 
Even though E [ p ^ + ] = p + and E [ p ^ ] = p , note that E [ ( p ^ + p ^ ) 2 ] ( p + p ) 2 . In fact, E [ ( p ^ + p ^ ) 2 ] = Var [ p ^ + p ^ ] + ( p + p ) 2 . Therefore, the estimator underestimates the true variance and is not correctable in practice. The bias involves the true variance, which is what we are trying to estimate in the first place, making correction circular without additional modeling assumptions. □
While it can be shown that the estimator, as expressed in Theorem 3, is asymptotically unbiased for the variance of NPS ^ , there is an easier way to prove this. Nonetheless, showing this for the estimator as expressed in Theorem 3 may be of some independent interest. We, therefore, include this result in Appendix A.
By simple algebraic rearrangement, we examine an alternative expression of the estimator for the variance of NPS ^ .
Theorem 4.
An alternative expression for the estimator for the variance of NPS ^ is:
Var ^ 2 [ NPS ^ ] = p ^ + ( 1 p ^ + ) + p ^ ( 1 p ^ ) + 2 p ^ + p ^ n .
The estimator is biased for a small sample size, n, but the bias goes to zero as n becomes large.
Proof. 
From Theorem 2, this estimator directly mimics the form:
Var [ p ^ + p ^ ] = [ p + ( 1 p + ) + p ( 1 p ) + 2 p + p ] n .
Taking the expectation of Var ^ 2 [ NPS ^ ] , we compute:
E [ p ^ + ( 1 p ^ + ) ] = E [ p ^ + ] + E [ p ^ + 2 ] = p + + Var [ p ^ + ] + ( E [ p ^ + ] ) 2 = p + p + ( 1 p + ) n + p + 2 = p + ( 1 p + ) 1 1 n
and the result for E [ p ^ ( 1 p ^ ) ] is similar. We also see:
E [ 2 p ^ + p ^ ] = 2 ( E [ p ^ + ] E [ p ^ ] + Cov [ p ^ + , p ^ ] ) = 2 p + p p + p n = 2 p + p 1 1 n .
Accordingly, this gives us the following expression:
E [ Var ^ 2 [ NPS ^ ] ] = Var [ NPS ^ ] 1 1 n ,
which shows that Var ^ 2 [ NPS ^ ] is biased low for Var [ NPS ^ ] . Clearly, the finite-sample correction 1 1 n 1 as n , thus affirming that Var ^ 2 [ NPS ^ ] is asymptotically unbiased. □
Theorem 5.
The unbiased estimator for the variance of NPS ^ is:
Var ^ 3 [ NPS ^ ] = p ^ + ( 1 p ^ + ) + p ^ ( 1 p ^ ) + 2 p ^ + p ^ n 1 .
Remark 1.
The estimated standard deviation of NPS ^ is simply SD ^ 3 [ NPS ^ ] = Var ^ 3 [ NPS ^ ] .
Proof. 
This estimator is merely a bias-corrected version of the alternative expression of the estimator given in Theorem 4, using a known multiplicative factor. The derivation immediately follows by multiplying both sides of Equation (4) by n / ( n 1 ) . □

1.2. Confidence Intervals for NPS

Confidence intervals for the difference between independent proportions have been well studied (e.g., see [7]). To derive a confidence interval for NPS, which is the difference between two dependent proportions, there are three approaches we can take.
First, we can compute Wald confidence intervals under the assumption of approximate normality for NPS ^ as:
NPS ^ ± z 1 α / 2 * · Var ^ 3 [ NPS ^ ] ,
where z 1 α / 2 * is the 1 α / 2 quantile of the standard normal distribution, and Var ^ 3 [ NPS ^ ] is the unbiased estimator for the variance of NPS ^ from Theorem 5. Wald confidence intervals for NPS are simple and reasonable in moderate-to-large sample sizes but can be too narrow in small samples, especially when p ^ + or p ^ is near the boundary 0 or 1. This is due to the reliance on the normal approximation and poor variance estimation at boundaries. In fact, it is possible that Wald confidence intervals can extend beyond the boundaries −1 or 1, which is not valid for NPS.
An important work that addressed large-sample confidence intervals for pairwise comparisons of multinomial proportions was given by Goodman [8]. Goodman used the (biased) variance given in Theorem 2, along with the square root of the 1 α / M quantile of the χ 2 distribution with 1 degree of freedom for α [ 0 , 1 ] . This is a Bonferroni correction on the χ 2 critical point to adjust for the M = k ( k 1 ) / 2 possible pairwise differences available across all k categories. In our case, since NPS was pre-specified, and we are not interested in the other two pairwise differences, we use M = 1 , and Goodman’s method reduces to a Wald confidence interval (note the square root of the 1 α quantile of the χ 2 distribution with 1 degree of freedom is equivalent to the 1 α / 2 quantile of the standard normal distribution).
Second, we can compute bootstrapped confidence intervals when we have small sample sizes and the assumption of normality for NPS ^ is tenuous. There are several different types of nonparametric bootstrap confidence intervals that we can consider, including the bootstrap percentile interval and the t interval with bootstrap standard error, for example. Efron and Tibshirani [9] described “confidence intervals based on bootstrap tables”, also commonly referred to as studentized bootstrap intervals or bootstrap t confidence intervals. While bootstrapping does not overcome the weakness of small samples as a basis for inference, Hesterberg [10] showed that bootstrap t confidence intervals can perform relatively well compared to the other confidence interval types, with respect to coverage, width, and robustness when n is small. Bootstrap t confidence intervals could work well for NPS because there is no reliance on the normal approximation, and the procedure will account for the skewness of the sampling distribution and heteroskedasticity with respect to NPS ^ . Also, bootstrap t confidence intervals fully account for the dependence between p ^ + and p ^ because we are resampling across all observations. Bootstrapping assumes the empirical distribution approximates the population well, which can be suspect if n is very small.
Rocks [4] proposed and systematically evaluated several interval estimation techniques for NPS via simulation and empirical data. One finding was that the Wald confidence interval performed poorly. By contrast, so-called “adjusted” Wald (AW) confidence intervals (simply adding a pseudo-count to each category before generating a Wald confidence interval) provided substantial improvements in accuracy. A well-performing method that was identified was a specific adjusted Wald variant denoted as AW(3, T) (we will describe this notation in greater detail in Section 2). This method involves adding 3/4 to the counts of both Promoters and Detractors and 3/2 to the count of Passives before the construction of a Wald confidence interval. AW(3, T) achieved the lowest mean absolute error in coverage among several methods and “good performance across the n values and confidence levels examined, especially for data likely to be observed in practice.”
When we incorporate pseudo-counts, this leads to an adjusted sample size and adjusted proportions, which means we are working with a different estimator, that is, a smoothed version of NPS ^ , denoted as NPS ^ s . The adjusted Wald method uses NPS ^ s , not the original NPS ^ , as the center of the confidence interval. While the smoothed estimator can be biased for the true NPS, this is intentionally done to improve the behavior of the confidence interval. Note that Var ^ 3 [ NPS ^ ] is only unbiased for the variance of the original NPS estimator based on the original multinomial sample of size n. If we apply the AW(3, T) confidence interval with the estimator in Theorem 5 using the adjusted sample size and adjusted proportions, then it is important to recognize that Var ^ 3 [ NPS ^ s ] will be unbiased only for the variance of NPS ^ s . Also, it can be shown that the variance of NPS ^ s (using Theorem 2) will generally be smaller than the variance of NPS ^ , with the difference depending on the alignment of the prior multinomial distribution implied by the added pseudo-counts and the true outcome probabilities. More will be said about this in Section 2. The key takeaway is that the adjusted Wald method sacrifices the accuracy of the estimator (bias) for the sake of a relatively larger gain in precision (smaller variance), a classic bias–variance tradeoff.
When considering our confidence interval approaches, one of the first questions we must answer is, “When will approximate normality hold for the sampling distribution of NPS ^ ?” The sampling distribution of NPS ^ = p ^ + p ^ will begin to look approximately normal as the sample size increases because it is the difference between two (correlated) binomial sample proportions. This relies on an extension of the classic binomial normal approximation rule, which is a specific application of the Central Limit Theorem, and it allows us to state that each expected count for the Promoter and Detractor multinomial cells should be at least 5, with 10 being better [11].
So, a rough rule of thumb in practice to determine whether NPS ^ is approximately normal would be:
min ( n p ^ + , n p ^ ) 5 ( preferably 10 for a better approximation ) .
For example, if p ^ + = 0.80 and p ^ = 0.10 , then at n = 50 , the sampling distribution for NPS ^ would be borderline normal since n p ^ + = X + = 40 and n p ^ = X = 5 . This works because both p ^ + and p ^ are binomial proportions (based on marginals of a multinomial) and the difference of two approximately normal random variables is also approximately normal. It can be shown that the two binomial sample proportions’ being correlated only affects the variance of the limiting distribution, not its normality.

1.3. Value Added from Our Approach

After reviewing the scope and limitations of the current literature on the NPS, we identify several unresolved inferential issues and provide a focused, statistically grounded critique that motivates the methodological contributions developed in this paper. While some existing work, most notably Rocks [4], has proposed several confidence interval methods and sparked interest in inference for NPS, important questions remain underexplored.
Understanding the sampling distribution of NPS ^ is crucial for assessing inference quality. This includes exploring its actual shape or behavior conditional on various multinomial distributions. The sampling distribution is discrete and bounded, with symmetry or skewness depending on the underlying outcome probabilities. Simulating from multiple multinomial distributions and directly visualizing the sampling distributions of NPS ^ yields crucial insight (e.g., visualizing sampling variability).
As a corollary, we provide diagnostic guidance for when the sampling distribution of NPS ^ is approximately normal. The Wald method relies heavily on the approximate normality of the sampling distribution. In small samples, the sampling distribution of NPS ^ could be far from normal, so ignoring this jeopardizes interval coverage and hypothesis test validity.
Bootstrap methods offer a flexible approach that is standard in modern practice. Specifically, bootstrap t confidence intervals are easy to implement and offer a nonparametric alternative for highly discrete or skewed settings, especially for small sample sizes.
The estimator of Var [ NPS ^ ] commonly used in the literature is biased, as we rigorously demonstrate in this paper. This could negatively impact confidence interval performance assessments, especially for small sample sizes. We derive and employ a finite-sample unbiased variance estimator that is not standard in the existing NPS inference, using it for our assessment of each confidence interval method under consideration.
In confidence interval assessments, precision matters as much as coverage. Accordingly, we examine both coverage and width. This is a key point because two intervals may both attain nominal 95% coverage, say, but have different precisions. Focusing on coverage and width provides a fuller picture of interval performance.
Studies often focus on the marginal (averaged) performance of estimation over the parameter space. However, when employing an inference in practice, we primarily focus on performance conditional on a single observed or assumed multinomial distribution that is emblematic of real-world data settings. Averaging over many multinomial distributions can mask poor performance in specific but important cases (e.g., very high NPS).
The performance of the adjusted Wald method is sensitive to the choice of pseudo-counts, which implicitly encode prior beliefs about the underlying multinomial distribution. If the parameterization for the actual distribution of counts is far from the implied prior used to determine the adjusted counts, then the estimate of NPS could be substantially biased, and the difference in variances between smoothed and unsmoothed estimators of the NPS could be substantially impacted. Therefore, the resulting confidence interval may achieve poor coverage or inflated width. This risk is important to characterize and understand.
Taken together, the points above represent a set of distinct methodological and practical contributions to inference for the NPS:
  • We characterize and visualize the finite-sample sampling distribution of NPS ^ under realistic multinomial response profiles, demonstrating when normal approximations are unreliable and why this directly affects confidence interval validity.
  • We show that the variance estimator commonly used for NPS inference is biased in finite samples and introduce a finite-sample unbiased variance estimator, using it to reassess confidence interval performance across methods.
  • Rather than averaging performance over the parameter space, we evaluate confidence interval methods conditional on realistic data-generating scenarios, jointly considering coverage and width to provide actionable guidance for robust interval selection in practice.
These contributions motivate the simulation and analytical results that follow, which are designed to quantify the coverage, precision, and robustness of competing confidence interval procedures under realistic finite-sample conditions.
Section 2 describes the data-generating framework, variance estimators, and confidence interval methods under study. Section 3 presents simulation results assessing finite-sample coverage and width across realistic scenarios, and Section 4 discusses these findings in the context of applied settings.

2. Materials and Methods

This section formally defines the confidence interval procedures and populations under consideration evaluated in the simulation study.
To compute bootstrap t confidence intervals, we resample the studentized statistic:
t * = NPS ^ * NPS ^ SD ^ 3 * ,
which is the defining feature of the bootstrap t method, where:
  • NPS ^ = the estimate from the original sample
  • NPS ^ * = the bootstrap estimate from the bootstrap sample (resampled with replacement)
  • SD ^ 3 * = estimated standard deviation of NPS ^ * from the bootstrap sample using Theorem 5.
To construct the confidence interval, for each bootstrap sample, b = 1 , 2 , , B , we compute t * . For each simulated dataset, we used B = 1000 bootstrap samples. Random number generation was initialized using a fixed seed to ensure reproducibility. We next obtain t 1 α / 2 * and t α / 2 * , which are quantiles of the B values of t * . The bootstrap t interval is:
[ NPS ^ t 1 α / 2 * · SD ^ 3 [ NPS ^ ] , NPS ^ t α / 2 * · SD ^ 3 [ NPS ^ ] ] ,
where SD ^ 3 [ NPS ^ ] is the estimated standard deviation from the original sample using Theorem 5. Note that, for all confidence interval methods considered in this paper, no truncation to the natural bounds of the NPS ([−1, 1]) was applied. This choice was made to ensure a fair comparison across methods and to avoid introducing method-specific distortions in coverage or the interval width. In our simulations, we observed no evidence of degenerate bootstrap samples causing the collapse of the bootstrap t interval. In particular, bootstrap t interval widths were never near zero. Accordingly, no special handling or filtering of near-zero variance bootstrap samples was required.
Following Rocks [4], we use the notation AW(w, shape) to denote an adjusted Wald confidence interval, where w is the total weight added across the observed counts, and shape describes the assumed distribution of outcome probabilities corresponding to a population that is either extreme (E), triangular (T), or uniform (U). We also introduce a population that is left-skewed (LS), a shape likely to be observed in the field. In our research, w is set equal to 3. Details are given below in Table 1. Each population configuration serves as a fixed data-generating mechanism from which repeated samples are drawn to induce a sampling distribution for NPS ^ , allowing us to study finite-sample variability and departures from normality under realistic response profiles.
To compute adjusted Wald confidence intervals, we first consider the pseudo-counts added to the original vector of counts ( X , X 0 , X + ) , respectively. For example, for AW(3, LS) we would have ( X + 1 / 4 , X 0 + 1 / 4 , X + + 5 / 2 ) . Notice all pseudo-counts sum to w = 3 for each population, so that the “sample size” used to compute the adjusted Wald confidence interval is equal to n + 3 . By dividing the pseudo-counts by 3, we obtain the implied prior distribution of outcome probabilities. Notice that the three populations E, T, and U have vectors ( p , p 0 , p + ) , such that the NPS = 0. However, the left-skewed population with vector ( p , p 0 , p + ) = ( 1 / 12 , 1 / 12 , 5 / 6 ) has a NPS = 5 / 6 1 / 12 = 0.75 . The variance (of NPS ^ ) listed in Table 1 is for a single multinoulli trial using the formula given in Theorem 2.
We use simulation ( 100 , 000 draws) to obtain twelve sampling distributions of NPS ^ associated with different populations using different sample sizes. With R = 100 , 000 replications per scenario, the maximum Monte Carlo standard error for coverage estimates is 0.25 / R 0.0016 , making simulation variability negligible for the comparisons of interest. For each of the four populations in Table 1, we use sample sizes of n = 20 , 50 , 100 . In this research, the confidence coefficient C will be set to 0.95. For each sample, we compute a Wald confidence interval using Equation (5), and a bootstrap t confidence interval using Equation (7). We also compute four adjusted Wald confidence intervals; specifically, AW(3, E), AW(3, LS), AW(3, T), and AW(3, U). As described previously, this simply means using pseudo-counts and, subsequently, the adjusted sample size and adjusted proportions, to obtain NPS ^ s and Var ^ 3 [ NPS ^ s ] (using Theorem 5), which are then input into Equation (5). Critically, for each of the four populations, three of the adjusted Wald confidences will be misspecified to the extent that the implied prior distribution of outcome probabilities will not align with the true distribution of outcome probabilities for the population from which the sample were drawn. We then compute the average width and proportion coverage of the confidence intervals across the 100,000 samples from the four populations for each sample size and for all six confidence interval types. Note that coverage will be computed with respect to the respective true NPS for each population.
Lastly, we generate three visuals. First, we generate a faceted plot of the twelve sampling distributions of NPS ^ . Second, we generate a line plot of coverage rates that shows how often the six different 95% confidence intervals contain the true NPS across different sample sizes and populations. This evaluates empirical coverage. Third, we generate a line plot of average interval widths that displays the average width of the 95% confidence intervals for each method, again by sample size and population. This reflects the precision and efficiency of the intervals.
All analyses were performed using R Statistical Software (v4.4.2; R Core Team 2024) [12]. The code to reproduce our analysis is openly available at (https://github.com/philturk/Net_Promoter_Score_Confidence_Intervals, accessed on 15 April 2026).

3. Results

Using the methods and range of realistic data-generating scenarios defined in Section 2, we now evaluate the sampling distribution and finite-sample confidence interval coverage and width associated with the estimation of NPS.

3.1. Sampling Distributions of NPS ^

The sampling distributions of NPS ^ are displayed in Figure 1. These sampling distributions represent the distribution of NPS ^ obtained from repeated samples drawn from a fixed population configuration in Table 1 and illustrate when approximate normality is plausible and when it fails. The red dashed vertical lines represent the true NPS for each of the populations. Notice that the mean of each sampling distribution approximately equals the true NPS (unbiasedness). For the left-skewed population, the sampling distributions corresponding to n = 20 and, to a lesser extent, n = 50 show non-normality. These are conditions under which the rule of thumb given in Equation (6) would not hold. Not surprisingly, the extreme population shows the greatest variance. Notice that the variability could be substantial for smaller sample sizes. This could have a dramatic effect on decision making for leaders who use NPS.

3.2. Confidence Interval Coverage

Looking at Figure 2, we see that the simulation results depend on the sample size and population. Correctly matching the adjusted Wald variant to the corresponding population type provides no guarantee that one will attain the specified confidence interval coverage. For example, AW(3, LS) yields confidence intervals that are too narrow and performs no better, or sometimes even worse, than the three other adjusted Wald variants. In general, as n increases, all six methods yield confidence intervals with coverage closer to the advertised 95%. Note that the extreme population is an exception since it is a boundary case on the multinomial simplex ( p 0 = 0 ), which induces a highly discrete sampling distribution for NPS ^ (grid spacing 2 / n ) and maximal variance. Because NPS can only take values on this coarse grid, confidence interval coverage cannot adjust smoothly and instead exhibits oscillatory behavior as n changes, particularly near the boundary of the NPS scale. These features together slow the apparent finite-sample convergence of empirical coverage toward the nominal level. We see that, in the absence of prior knowledge and among the population shapes examined, AW(3, T) and AW(3, U) are good choices for a confidence interval method, followed closely by AW(3, E). The Wald and AW(3, LS) methods cannot be recommended as sensible choices generalizable over all n and population types.
To complement the graphical summaries in Figure 2, Table 2 reports mean empirical coverage across all population shapes, sample sizes, and confidence interval procedures.

3.3. Confidence Interval Width

The simulation results for the average width of the 95% confidence intervals in Figure 3 also depend on the sample size and population. The extreme population gave the broadest confidence intervals, while the left-skewed population gave the narrowest. Looking over population types at n = 20 , and also comparing back to the coverage in Figure 2, reveals that the Wald and bootstrap t methods should be avoided. The adjusted Wald methods are fairly comparable in width for each population types, especially as n gets bigger.
Corresponding numerical summaries of mean confidence interval width are reported in Table 3, allowing a direct comparison of precision across methods and sample sizes.

3.4. Applied Illustration

To reinforce how small-sample inference can differ across interval constructions, we consider a simple worked example motivated by the LS population in Table 1. Suppose a survey of size n = 20 yields counts:
( X , X 0 , X + ) = ( 1 , 2 , 17 ) ,
which is consistent with an LS setting. Clearly, the normal approximation is questionable. The sample proportions are p ^ + = 17 / 20 = 0.85 and p ^ = 1 / 20 = 0.05 , giving:
NPS ^ = p ^ + p ^ = 17 1 20 = 0.80 .
We first compute the Wald confidence interval. Using the finite-sample unbiased variance estimator from Theorem 5, we obtain:
Var ^ 3 [ NPS ^ ] = 0.85 ( 0.15 ) + 0.05 ( 0.95 ) + 2 ( 0.85 ) ( 0.05 ) 19 = 0.2600 19 0.01368 ,
so SD ^ 3 [ NPS ^ ] 0.01368 = 0.117 . A nominal 95% Wald interval (Equation (5)) is therefore:
0.80 ± 1.96 ( 0.117 ) = [ 0.571 , 1.029 ] .
In this illustration, the interval extends beyond the natural parameter space for NPS, which is bounded above by 1.
Next, consider the adjusted Wald procedure AW(3, T), which adds pseudo-counts ( 3 / 4 , 3 / 2 , 3 / 4 ) to ( X , X 0 , X + ) . This yields adjusted counts:
( X * , X 0 * , X + * ) = ( 1.75 , 3.5 , 17.75 ) , n * = n + 3 = 23 ,
and adjusted proportions p ^ + * = 17.75 / 23 and p ^ * = 1.75 / 23 . The smoothed NPS estimator is:
NPS ^ s = p ^ + * p ^ * = 17.75 1.75 23 = 16 23 0.696 .
Applying the same unbiased variance estimator (Theorem 5) with the adjusted sample size gives:
Var ^ 3 [ NPS ^ s ] 0.01654 , SD ^ 3 [ NPS ^ s ] 0.129 ,
and thus a nominal 95% adjusted Wald interval:
0.696 ± 1.96 ( 0.129 ) = [ 0.444 , 0.948 ] .
Compared to the Wald interval, AW(3, T) is slightly wider in this example, but it remains within [ 1 , 1 ] (n.b., as with any normal-approximation interval, adjusted Wald intervals can also extend beyond the parameter bounds in small samples). This simple illustration aligns with the simulation findings for the LS population at small n; specifically, while Wald intervals can be narrower in this setting, adjusted Wald procedures (e.g., AW(3, T)) provide more stable inference under discreteness and skewness. Because AW(3, U) behaved similarly to AW(3, T) in our simulations, we omit it here for brevity.

4. Discussion

There are limitations we acknowledge in this research. In practice, the assumption that we have obtained a true simple random sample might be plainly suspect. Quite often, we are dealing with a representative sample or, perhaps, a voluntary response sample. Nevertheless, our methods give practitioners a benchmark that could be useful, provided that any interpretation is tempered with caution.
To complement the coverage and width results, we briefly examine the bias of the NPS estimators at n = 20 . As expected, the original estimator NPS ^ and the bootstrap-based approach exhibit a negligible degree of bias across all population configurations. In contrast, the smoothed estimator underlying adjusted Wald intervals can exhibit non-negligible bias when the implied pseudo-count distribution is misaligned with the true population, with a magnitude approaching 0.10 in our simulations. When the pseudo-counts are well aligned with the underlying population, the bias is effectively zero. These results highlight the bias–variance tradeoff inherent to adjusted Wald methods; improved coverage and stability may come at the cost of bias under model misspecification.
As mentioned in Section 1, an NPS value of zero represents a natural interpretive benchmark separating Promoter and Detractor subpopulations. In practice, testing whether NPS differs from zero aligns with how the metric would be used in organizational benchmarking and decision-making. For hypothesis testing (e.g., checking if a NPS is significantly different from zero or if two NPS scores differ), we note that similar principles apply. Because NPS ^ is essentially a sample mean (of coded values 1 , 0 , and 1), a one-sample t-test can test if NPS equals zero, or some other value under the null hypothesis, by using the estimated standard deviation and n 1 degrees of freedom. In practice, this would be a useful tactic for assessing whether customer loyalty to a product, say, is comparable to that of some predefined standard. Likewise, comparing two independent NPS scores (from, say, two companies or two time periods) can be done with a two-sample t statistic comparing the difference in NPS ^ to zero. For example, a two-sample t-test of NPS scores of a product or service over a period of time is a way to assess whether a statistically significant improvement in customer loyalty has been achieved. These procedures assume approximate normality of the NPS estimator, an assumption that can be checked using our rule of thumb. The main caution is to use the correct variance formula for NPS and not to treat the two proportions ( p ^ + and p ^ ) as if they were independent.
A line of inquiry we have explored in our research has been to raise awareness about the sampling distribution of the NPS statistic and to more rigorously address its inferential properties. Because NPS ^ is a linear combination of multinomial counts, its exact sampling distribution is discrete (with a granularity of 1 / n in NPS units), being derived from the multinomial probability mass function. Our results confirm that the uncertainty in NPS ^ can be considerable, leading to confidence intervals for NPS that can be very wide for small sample sizes. Overall, while NPS is valued for simplicity, we caution against over-reliance on a single-number metric without proper statistical context and validation when making an inference about a customer base.
Sampling variability in NPS ^ is something that should not be ignored. Much like high-quality sampling polls, the recommendation we make is to always pair an estimate of NPS with a margin of error or confidence interval. However, in practice, the authors note that this is rarely done. The consequences of ignoring sampling variability can be practically profound and lead to “chasing noise”. Let us reconsider our example with the healthcare system that is considering deploying a new workforce development course to its constituents. Consider a scenario under which, had all the constituents taken the course (the population), the true unknown NPS would have been 0.75. Now, suppose a class consists of 20 students. On average, we would expect NPS ^ 0.75 , and yet, it is entirely possible by chance to see an estimate of 0.25 (Figure 1). The leadership at the healthcare system might wrongly conclude that the course is ineffective when, in fact, the healthcare system at large would enthusiastically embrace the course. As a second scenario, it is possible that a population could have a true unknown NPS of 0, and yet, a class of 20 students might yield an estimate of 0.50 due to chance (Figure 1). The leadership at the healthcare enterprise might wrongly conclude that the course is excellent when, on average, all of their constituents actually view the course with equanimity.
Many times in practice, there might be neither intuition nor prior knowledge as to the type of population under consideration. Based on our simulation results for the coverage and width of confidence intervals for NPS that account for the multinomial nature of the data, the adjusted Wald methods with triangular or uniform pseudo-counts are good choices across a range of practically relevant population shapes, followed by extreme pseudo-counts.
The mean squared error (MSE) of an estimator is defined to be the sum of the variance of the estimator and the squared bias of the estimator. This provides a useful way to compare estimators with respect to precision and accuracy to see which one is “best”, at least from an MSE perspective. We did not compute the MSEs for the alternative and unbiased estimators for Var [ NPS ^ ] . As expected, the smoothed estimator for NPS underlying adjusted Wald intervals exhibits finite-sample bias when the implied pseudo-count distribution is misaligned with the true population, particularly at very small sample sizes. However, the inferential consequences of this bias are already captured through the empirical coverage results reported here. A systematic evaluation of estimator bias and MSE represents a natural direction for future work.
Our methods are developed assuming an infinite, or “large”, population. If samples are drawn from a finite population, such that the sampling fraction becomes excessively high (e.g., ≥0.10, say), then the resulting confidence intervals for NPS using the methods given in our research will be conservative. Future research could explore attenuating the variance estimator using a finite-population correction factor [13].
In practice, when we have missing data or nonresponse (e.g., 10 out of 20 students in a class respond to a survey), our treatment of the issue depends on the reason for the missingness and our analysis goals. Practical approaches for resolving the issue depend on the setting and can become quite complex. This would represent another fruitful area of future research. A complete case analysis, the most common approach, assumes that responses are missing completely at random and proceeds with estimation using just the respondents. However, there is a risk of bias in estimation if the nonrespondents differ systematically from the respondents, which occurs when missingness is not completely at random. When in doubt, seek the consultation of a trained statistician.

Author Contributions

P.T.: conceptualization, formal analysis, methodology, project administration, software, supervision, visualization, and writing (original draft, review and editing). J.C.: writing (original draft, review and editing). E.M.: formal analysis, software, visualization, and writing (original draft, review and editing). All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

No empirical data were created or analyzed in this study. All results are based on simulated data that were generated as described in Section 2. Data sharing is not applicable to this article.

Acknowledgments

The authors gratefully acknowledge the helpful feedback received from Jennifer Reneker and Kiersten Winter (Northeast Ohio Medical University) in the preparation of this manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

This appendix provides supporting details showing that Var ^ 1 [ NPS ^ ] is asymptotically unbiased; it is not required for implementation.
In statistics, an estimator is considered consistent if, as the sample size increases, the estimator’s value converges to the true value of the parameter being estimated in probability. More formally, we want to show that Var ^ 1 [ NPS ^ ] p Var [ NPS ^ ] as n . To show that this is true, we rely on the fact that p ^ + p p + and p ^ p p through the Law of Large Numbers. Hence, any continuous function g, say, of p ^ + and p ^ , will also converge in probability to the continuous function g evaluated at the parameters p + and p .
Because the numerator of the variance estimator is a continuous function of p ^ + and p ^ , we have:
p ^ + + p ^ ( p ^ + p ^ ) 2 p p + + p ( p + p ) 2
and multiplying each side by 1 / n :
Var ^ 1 [ NPS ^ ] p Var [ NPS ^ ] .
(By the same logic, Var ^ 2 [ NPS ^ ] is a consistent estimator of Var [ NPS ^ ] .)
In the case of the estimator for the variance of NPS ^ from Theorem 3, we can conclude asymptotic unbiasedness because the two key properties of the Bounded Convergence Theorem are met. First, we showed that the estimator is consistent (convergence in probability). Second, the estimator is bounded. To see this, note that the estimator for the variance of NPS ^ is:
Var ^ 1 [ NPS ^ ] = p ^ + + p ^ ( p ^ + p ^ ) 2 n ,
which contains all numerator terms between 0 and 1. Therefore, it is easy to see that the estimator is bounded between 0 and 1 n , and so:
sup n Var ^ 1 [ NPS ^ ] 1 n 0 as n ,
Because the estimator is bounded, it is also trivially uniformly integrable. To verify this, we check:
ϵ > 0 , K > 0 such that sup n E [ Var ^ 1 [ NPS ^ ] · 1 { Var ^ 1 [ NPS ^ ] > K } ] < ϵ
where 1 { · } denotes an indicator function. Since sup n Var ^ 1 [ NPS ^ ] 1 n , then, for any K > 1 n , we have:
1 { Var ^ 1 [ NPS ^ ] > K } = 0 almost surely E [ Var ^ 1 [ NPS ^ ] · 1 { Var ^ 1 [ NPS ^ ] > K } ] = 0 .
Therefore:
sup n E [ Var ^ 1 [ NPS ^ ] · 1 { Var ^ 1 [ NPS ^ ] > K } ] = 0 < ϵ for all ϵ > 0
Said in another way, Var ^ 1 [ NPS ^ ] is uniformly integrable because its tail expectation is always zero for any K > 1 n .
In conclusion, the boundedness of Var ^ 1 [ NPS ^ ] implies uniform integrability. In turn, convergence in probability and uniform integrability implies convergence in the mean; that is, Var ^ 1 [ NPS ^ ] is asymptotically unbiased for Var [ NPS ^ ] .

References

  1. Reichheld, F.F. The one number you need to grow. Harv. Bus. Rev. 2003, 81, 46–55. [Google Scholar] [PubMed]
  2. Qualtrics. What Is a Good Net Promoter Score? 2021. Available online: https://www.qualtrics.com/experience-management/customer/good-net-promoter-score/ (accessed on 4 July 2025).
  3. Cazzaro, M.; Chiodini, P.M. Statistical validation of critical aspects of the Net Promoter Score. TQM J. 2023, 35, 191–209. [Google Scholar] [CrossRef]
  4. Rocks, B. Interval estimation for the “net promoter score”. Am. Stat. 2016, 70, 365–372. [Google Scholar] [CrossRef]
  5. Measuring U. Has the Net Promoter Score Been Discredited in the Academic Literature? 2019. Available online: https://measuringu.com/nps-discredited/#:~:text=Across%2022%2...emic%20papers%2C%20we,best%20predictor%2C%20it%20was%20close (accessed on 4 July 2025).
  6. Casella, G.; Berger, R.L. Statistical Inference, 2nd ed.; Duxbury: Pacific Grove, CA, USA, 2002. [Google Scholar]
  7. Newcombe, R.G. Interval estimation for the difference between independent proportions: Comparison of eleven methods. Stat. Med. 1998, 17, 873–890. [Google Scholar] [CrossRef]
  8. Goodman, L.A. On Simultaneous Confidence Intervals for Multinomial Proportions. Technometrics 1965, 7, 247–254. [Google Scholar] [CrossRef]
  9. Efron, B.; Tibshirani, R. An Introduction to the Bootstrap; Chapman and Hall/CRC: New York, NY, USA, 1993. [Google Scholar]
  10. Hesterberg, T.C. What teachers should know about the bootstrap: Resampling in the undergraduate statistics curriculum. Am. Stat. 2015, 69, 371–386. [Google Scholar] [CrossRef] [PubMed]
  11. Ott, L.; Longnecker, M. An Introduction to Statistical Methods and Data Analysis; Cengage Learning: Boston, MA, USA, 2016. [Google Scholar]
  12. R Core Team. R: A Language and Environment for Statistical Computing; R Foundation for Statistical Computing: Vienna, Austria, 2024. [Google Scholar]
  13. Thompson, S.K. Sampling, 3rd ed.; Wiley Series in Probability and Statistics Sampling; Wiley: Hoboken, NJ, USA, 2012. [Google Scholar]
Figure 1. Sampling distributions of NPS ^ .
Figure 1. Sampling distributions of NPS ^ .
Stats 09 00045 g001
Figure 2. Coverage of 95% confidence intervals for NPS.
Figure 2. Coverage of 95% confidence intervals for NPS.
Stats 09 00045 g002
Figure 3. Average width of 95% confidence intervals for NPS.
Figure 3. Average width of 95% confidence intervals for NPS.
Stats 09 00045 g003
Table 1. Populations Under Consideration.
Table 1. Populations Under Consideration.
PopulationPseudo-Counts ( p , p 0 , p + ) Variance
Extreme(3/2, 0, 3/2)(1/2, 0, 1/2)1.00
Left-Skewed(1/4, 1/4, 5/2)(1/12, 1/12, 5/6)0.35
Triangular(3/4, 3/2, 3/4)(1/4, 1/2, 1/4)0.50
Uniform(1, 1, 1)(1/3, 1/3, 1/3)0.67
Table 2. Mean empirical coverage (nominal 95%) across population shapes (E, LS, T, U) and sample sizes n { 20 , 50 , 100 } .
Table 2. Mean empirical coverage (nominal 95%) across population shapes (E, LS, T, U) and sample sizes n { 20 , 50 , 100 } .
MethodELSTU
20 50 100 20 50 100 20 50 100 20 50 100
Wald0.9580.9330.9420.8660.9240.9370.9280.9450.9460.9370.9410.947
Bootstrap t0.9600.9520.9510.8300.9630.9660.9550.9530.9510.9700.9570.953
AW(3,E)0.9580.9330.9420.9750.9650.9590.9660.9550.9540.9620.9530.953
AW(3,LS)0.9370.9310.9370.9180.9240.9400.9100.9320.9440.9200.9340.944
AW(3,T)0.9580.9330.9420.9740.9540.9530.9570.9540.9510.9570.9490.950
AW(3,U)0.9580.9330.9420.9750.9610.9560.9570.9540.9510.9570.9490.950
Table 3. Mean confidence interval width across population shapes (E, LS, T, U) and sample sizes n { 20 , 50 , 100 } .
Table 3. Mean confidence interval width across population shapes (E, LS, T, U) and sample sizes n { 20 , 50 , 100 } .
MethodELSTU
20 50 100 20 50 100 20 50 100 20 50 100
Wald0.8760.5540.3920.4960.3250.2320.6150.3910.2770.7130.4520.320
Bootstrap t0.9140.5710.3950.5700.3780.2470.6590.4000.2790.7690.4630.323
AW(3,E)0.8190.5390.3860.5790.3470.2400.6150.3910.2770.6900.4460.318
AW(3,LS)0.8100.5370.3860.4730.3160.2280.6030.3880.2760.6790.4430.317
AW(3,T)0.7910.5310.3830.5370.3350.2350.5770.3800.2730.6560.4360.314
AW(3,U)0.8010.5330.3840.5510.3390.2370.5900.3840.2740.6670.4390.315
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Turk, P.; Cinderich, J.; McNeill, E. Coverage and Precision of Net Promoter Score Confidence Intervals Across Sampling Distributions. Stats 2026, 9, 45. https://doi.org/10.3390/stats9020045

AMA Style

Turk P, Cinderich J, McNeill E. Coverage and Precision of Net Promoter Score Confidence Intervals Across Sampling Distributions. Stats. 2026; 9(2):45. https://doi.org/10.3390/stats9020045

Chicago/Turabian Style

Turk, Philip, Jordan Cinderich, and Emma McNeill. 2026. "Coverage and Precision of Net Promoter Score Confidence Intervals Across Sampling Distributions" Stats 9, no. 2: 45. https://doi.org/10.3390/stats9020045

APA Style

Turk, P., Cinderich, J., & McNeill, E. (2026). Coverage and Precision of Net Promoter Score Confidence Intervals Across Sampling Distributions. Stats, 9(2), 45. https://doi.org/10.3390/stats9020045

Article Metrics

Back to TopTop