1. Introduction
The Net Promoter Score (NPS) is a simple but powerful tool used to gauge customer loyalty in business contexts [
1]. It is based on one question: “How likely are you to recommend our company/product/service to a friend or colleague?” The responses are given on an 11-point ordinal scale {0, 1, …, 10}, and customers are grouped into three classes: Promoters (those giving 9 or 10), Passives (those giving 7 or 8), or Detractors (those giving 0 up to and including 6). The counts for each of the three classes are taken from a sample survey, and percentages are computed. The NPS is defined to be:
Without a loss of generality, just the proportions can be used. In this case, we see that the NPS
. In terms of interpretation, an NPS of 0 means that Promoters and Detractors are proportionally balanced, while positive scores are “good”, demonstrating proportionally more Promoters than Detractors. While there are no widely accepted absolute benchmark values, and these vary among industries, many would consider an NPS over 0.50 to be excellent [
2].
The NPS has seen rapid adoption as a key performance indicator of customer loyalty and satisfaction. In practice, organizations track the NPS over time and compare scores across customer segments, making a statistical inference about the NPS (e.g., confidence intervals and hypothesis tests) an important concern. Despite its popularity in business, marketing, healthcare, etc., uncertainty quantification is rarely reported in practice, and guidance on which confidence intervals are reliable for NPS’s categorical nature is limited. The NPS metric has faced scrutiny in the general and peer-reviewed literature. For example, researchers have questioned whether the NPS is truly superior to other customer metrics, with some finding that the NPS typically has no stronger positive association with business outcomes (like customer retention or revenue growth) than other loyalty or satisfaction metrics. Other recent research termed the NPS “ambiguous and unreliable if not accompanied by other tools,” since identical NPS values may mask differing customer sentiment compositions [
3]. Rocks [
4] gave a brief summary of several of the critiques of the NPS. An extensive and comprehensive historical literature review of the NPS can be found at Measuring U [
5].
Because the NPS is so widely used for making strategic decisions across many organizations spanning various sectors, one might question, “how much can we trust this metric?” A way to think about assessing the NPS is to cast its interpretation in the context of estimation under a probability sampling design. For example, suppose that a large healthcare system is considering deploying a workforce development course to its constituents at great expense in terms of time and money. The 10,000 employees represent the population, while 20 individuals attending the inaugural course represent the sample. Hypothetically, if the entire population took the course, the true (and unknown) NPS would be the so-called parameter, using the parlance of a statistician. The NPS computed from the 20 individuals would be the statistic, our point estimate of the parameter. Much like a sampling poll, interest lies in inference on the parameter. How comfortable should leadership be in the system-wide success of the course with the NPS from an inaugural cohort of 20 employees? A single point estimate of NPS with no measure of uncertainty will not be very helpful to the leadership of the healthcare system and could even lead to risky conclusions and decisions. In practice, however, NPS is almost always reported as a single point estimate, with formal measures of uncertainty, such as confidence intervals, rarely if ever given. Providing the leadership with a range of plausible values for the NPS parameter, i.e., a confidence interval, would be useful in making better informed strategic planning, policy making, etc. A confidence interval shifts the NPS from a solitary, potentially misleading number into a more nuanced decision-support tool, enabling organizations to allocate resources, assess initiatives, and gauge risk with greater clarity and conviction.
1.1. Background and Theory
Because it is a difference of two proportions, the NPS itself can be thought of as a statistic derived from a sample from a three-category multinomial distribution (Detractor, Passive, and Promoter counts from a multinomial distribution with three categories). More formally, we have three possible mutually exclusive outcomes, with outcome probabilities
, and
, where these correspond to “success” probabilities for the Detractor, Passive, and Promoter outcomes, respectively. In addition, we assume we have
independent multinoulli trials, the outcome probabilities are all non-negative, and that
. If the random variables
, and
indicate the number of times the Detractor, Passive, and Promoter outcomes are observed over the
n trials, respectively, then the vector
follows a multinomial distribution with parameters
n and the vector
. While the trials are independent, their outcomes,
, and
, are dependent because they must sum to
n. The probability mass function for the multinomial distribution in this case is:
Marginally, the count for any single category in a
k-category multinomial distribution with a fixed
n is binomially distributed:
for
i = −, 0, or + with probability mass function:
Now, suppose we survey a random sample of
n respondents. Let
and
. Then, the sample NPS is:
We now show important results for this statistic using basic probability theory [
6].
Theorem 1. is an unbiased estimator for NPS.
Proof. Let be the sample proportion of Promoters and be the sample proportion of Detractors. Therefore, .
Hence, the mean, or expected value, of
is:
□
Theorem 2. The variance of for a sample of size n is: Proof. It is known that
for
i = − or +, and that
. Hence, the variance of
is:
□
We now consider how to derive an estimator for the variance of .
Theorem 3. The estimator for the variance of is:The estimator is biased for a small sample size, n, but the bias goes to zero as n becomes large. Proof. Even though and , note that . In fact, . Therefore, the estimator underestimates the true variance and is not correctable in practice. The bias involves the true variance, which is what we are trying to estimate in the first place, making correction circular without additional modeling assumptions. □
While it can be shown that the estimator, as expressed in Theorem 3, is asymptotically unbiased for the variance of
, there is an easier way to prove this. Nonetheless, showing this for the estimator as expressed in Theorem 3 may be of some independent interest. We, therefore, include this result in
Appendix A.
By simple algebraic rearrangement, we examine an alternative expression of the estimator for the variance of .
Theorem 4. An alternative expression for the estimator for the variance of is:The estimator is biased for a small sample size, n, but the bias goes to zero as n becomes large. Proof. From Theorem 2, this estimator directly mimics the form:
Taking the expectation of
, we compute:
and the result for
is similar. We also see:
Accordingly, this gives us the following expression:
which shows that
is biased low for
. Clearly, the finite-sample correction
as
, thus affirming that
is asymptotically unbiased. □
Theorem 5. The unbiased estimator for the variance of is: Remark 1. The estimated standard deviation of is simply .
Proof. This estimator is merely a bias-corrected version of the alternative expression of the estimator given in Theorem 4, using a known multiplicative factor. The derivation immediately follows by multiplying both sides of Equation (
4) by
. □
1.2. Confidence Intervals for NPS
Confidence intervals for the difference between independent proportions have been well studied (e.g., see [
7]). To derive a confidence interval for NPS, which is the difference between two
dependent proportions, there are three approaches we can take.
First, we can compute Wald confidence intervals under the assumption of approximate normality for
as:
where
is the
quantile of the standard normal distribution, and
is the unbiased estimator for the variance of
from Theorem 5. Wald confidence intervals for NPS are simple and reasonable in moderate-to-large sample sizes but can be too narrow in small samples, especially when
or
is near the boundary 0 or 1. This is due to the reliance on the normal approximation and poor variance estimation at boundaries. In fact, it is possible that Wald confidence intervals can extend beyond the boundaries −1 or 1, which is not valid for NPS.
An important work that addressed large-sample confidence intervals for pairwise comparisons of multinomial proportions was given by Goodman [
8]. Goodman used the (biased) variance given in Theorem 2, along with the square root of the
quantile of the
distribution with 1 degree of freedom for
. This is a Bonferroni correction on the
critical point to adjust for the
possible pairwise differences available across all
k categories. In our case, since NPS was pre-specified, and we are not interested in the other two pairwise differences, we use
, and Goodman’s method reduces to a Wald confidence interval (note the square root of the
quantile of the
distribution with 1 degree of freedom is equivalent to the
quantile of the standard normal distribution).
Second, we can compute bootstrapped confidence intervals when we have small sample sizes and the assumption of normality for
is tenuous. There are several different types of nonparametric bootstrap confidence intervals that we can consider, including the bootstrap percentile interval and the
t interval with bootstrap standard error, for example. Efron and Tibshirani [
9] described “confidence intervals based on bootstrap tables”, also commonly referred to as studentized bootstrap intervals or bootstrap
t confidence intervals. While bootstrapping does not overcome the weakness of small samples as a basis for inference, Hesterberg [
10] showed that bootstrap
t confidence intervals can perform relatively well compared to the other confidence interval types, with respect to coverage, width, and robustness when
n is small. Bootstrap
t confidence intervals could work well for NPS because there is no reliance on the normal approximation, and the procedure will account for the skewness of the sampling distribution and heteroskedasticity with respect to
. Also, bootstrap
t confidence intervals fully account for the dependence between
and
because we are resampling across all observations. Bootstrapping assumes the empirical distribution approximates the population well, which can be suspect if
n is very small.
Rocks [
4] proposed and systematically evaluated several interval estimation techniques for NPS via simulation and empirical data. One finding was that the Wald confidence interval performed poorly. By contrast, so-called “adjusted” Wald (AW) confidence intervals (simply adding a pseudo-count to each category before generating a Wald confidence interval) provided substantial improvements in accuracy. A well-performing method that was identified was a specific adjusted Wald variant denoted as AW(3, T) (we will describe this notation in greater detail in
Section 2). This method involves adding 3/4 to the counts of both Promoters and Detractors and 3/2 to the count of Passives before the construction of a Wald confidence interval. AW(3, T) achieved the lowest mean absolute error in coverage among several methods and “good performance across the
n values and confidence levels examined, especially for data likely to be observed in practice.”
When we incorporate pseudo-counts, this leads to an adjusted sample size and adjusted proportions, which means we are working with a different estimator, that is, a smoothed version of
, denoted as
. The adjusted Wald method uses
, not the original
, as the center of the confidence interval. While the smoothed estimator can be biased for the true NPS, this is intentionally done to improve the behavior of the confidence interval. Note that
is only unbiased for the variance of the original NPS estimator based on the original multinomial sample of size
n. If we apply the AW(3, T) confidence interval with the estimator in Theorem 5 using the adjusted sample size and adjusted proportions, then it is important to recognize that
will be unbiased only for the variance of
. Also, it can be shown that the variance of
(using Theorem 2) will generally be smaller than the variance of
, with the difference depending on the alignment of the prior multinomial distribution implied by the added pseudo-counts and the true outcome probabilities. More will be said about this in
Section 2. The key takeaway is that the adjusted Wald method sacrifices the accuracy of the estimator (bias) for the sake of a relatively larger gain in precision (smaller variance), a classic bias–variance tradeoff.
When considering our confidence interval approaches, one of the first questions we must answer is, “When will approximate normality hold for the sampling distribution of
?” The sampling distribution of
will begin to look approximately normal as the sample size increases because it is the difference between two (correlated) binomial sample proportions. This relies on an extension of the classic binomial normal approximation rule, which is a specific application of the Central Limit Theorem, and it allows us to state that each expected count for the Promoter and Detractor multinomial cells should be at least 5, with 10 being better [
11].
So, a rough rule of thumb in practice to determine whether
is approximately normal would be:
For example, if
and
, then at
, the sampling distribution for
would be borderline normal since
and
. This works because both
and
are binomial proportions (based on marginals of a multinomial) and the difference of two approximately normal random variables is also approximately normal. It can be shown that the two binomial sample proportions’ being correlated only affects the variance of the limiting distribution, not its normality.
1.3. Value Added from Our Approach
After reviewing the scope and limitations of the current literature on the NPS, we identify several unresolved inferential issues and provide a focused, statistically grounded critique that motivates the methodological contributions developed in this paper. While some existing work, most notably Rocks [
4], has proposed several confidence interval methods and sparked interest in inference for NPS, important questions remain underexplored.
Understanding the sampling distribution of is crucial for assessing inference quality. This includes exploring its actual shape or behavior conditional on various multinomial distributions. The sampling distribution is discrete and bounded, with symmetry or skewness depending on the underlying outcome probabilities. Simulating from multiple multinomial distributions and directly visualizing the sampling distributions of yields crucial insight (e.g., visualizing sampling variability).
As a corollary, we provide diagnostic guidance for when the sampling distribution of is approximately normal. The Wald method relies heavily on the approximate normality of the sampling distribution. In small samples, the sampling distribution of could be far from normal, so ignoring this jeopardizes interval coverage and hypothesis test validity.
Bootstrap methods offer a flexible approach that is standard in modern practice. Specifically, bootstrap t confidence intervals are easy to implement and offer a nonparametric alternative for highly discrete or skewed settings, especially for small sample sizes.
The estimator of commonly used in the literature is biased, as we rigorously demonstrate in this paper. This could negatively impact confidence interval performance assessments, especially for small sample sizes. We derive and employ a finite-sample unbiased variance estimator that is not standard in the existing NPS inference, using it for our assessment of each confidence interval method under consideration.
In confidence interval assessments, precision matters as much as coverage. Accordingly, we examine both coverage and width. This is a key point because two intervals may both attain nominal 95% coverage, say, but have different precisions. Focusing on coverage and width provides a fuller picture of interval performance.
Studies often focus on the marginal (averaged) performance of estimation over the parameter space. However, when employing an inference in practice, we primarily focus on performance conditional on a single observed or assumed multinomial distribution that is emblematic of real-world data settings. Averaging over many multinomial distributions can mask poor performance in specific but important cases (e.g., very high NPS).
The performance of the adjusted Wald method is sensitive to the choice of pseudo-counts, which implicitly encode prior beliefs about the underlying multinomial distribution. If the parameterization for the actual distribution of counts is far from the implied prior used to determine the adjusted counts, then the estimate of NPS could be substantially biased, and the difference in variances between smoothed and unsmoothed estimators of the NPS could be substantially impacted. Therefore, the resulting confidence interval may achieve poor coverage or inflated width. This risk is important to characterize and understand.
Taken together, the points above represent a set of distinct methodological and practical contributions to inference for the NPS:
We characterize and visualize the finite-sample sampling distribution of under realistic multinomial response profiles, demonstrating when normal approximations are unreliable and why this directly affects confidence interval validity.
We show that the variance estimator commonly used for NPS inference is biased in finite samples and introduce a finite-sample unbiased variance estimator, using it to reassess confidence interval performance across methods.
Rather than averaging performance over the parameter space, we evaluate confidence interval methods conditional on realistic data-generating scenarios, jointly considering coverage and width to provide actionable guidance for robust interval selection in practice.
These contributions motivate the simulation and analytical results that follow, which are designed to quantify the coverage, precision, and robustness of competing confidence interval procedures under realistic finite-sample conditions.
Section 2 describes the data-generating framework, variance estimators, and confidence interval methods under study.
Section 3 presents simulation results assessing finite-sample coverage and width across realistic scenarios, and
Section 4 discusses these findings in the context of applied settings.
2. Materials and Methods
This section formally defines the confidence interval procedures and populations under consideration evaluated in the simulation study.
To compute bootstrap
t confidence intervals, we resample the studentized statistic:
which is the defining feature of the bootstrap
t method, where:
= the estimate from the original sample
= the bootstrap estimate from the bootstrap sample (resampled with replacement)
= estimated standard deviation of from the bootstrap sample using Theorem 5.
To construct the confidence interval, for each bootstrap sample,
, we compute
. For each simulated dataset, we used
bootstrap samples. Random number generation was initialized using a fixed seed to ensure reproducibility. We next obtain
and
, which are quantiles of the
B values of
. The bootstrap
t interval is:
where
is the estimated standard deviation from the original sample using Theorem 5. Note that, for all confidence interval methods considered in this paper, no truncation to the natural bounds of the NPS ([−1, 1]) was applied. This choice was made to ensure a fair comparison across methods and to avoid introducing method-specific distortions in coverage or the interval width. In our simulations, we observed no evidence of degenerate bootstrap samples causing the collapse of the bootstrap
t interval. In particular, bootstrap
t interval widths were never near zero. Accordingly, no special handling or filtering of near-zero variance bootstrap samples was required.
Following Rocks [
4], we use the notation AW(
w, shape) to denote an adjusted Wald confidence interval, where
w is the total weight added across the observed counts, and shape describes the assumed distribution of outcome probabilities corresponding to a population that is either extreme (E), triangular (T), or uniform (U). We also introduce a population that is left-skewed (LS), a shape likely to be observed in the field. In our research,
w is set equal to 3. Details are given below in
Table 1. Each population configuration serves as a fixed data-generating mechanism from which repeated samples are drawn to induce a sampling distribution for
, allowing us to study finite-sample variability and departures from normality under realistic response profiles.
To compute adjusted Wald confidence intervals, we first consider the pseudo-counts added to the original vector of counts
, respectively. For example, for AW(3, LS) we would have
. Notice all pseudo-counts sum to
for each population, so that the “sample size” used to compute the adjusted Wald confidence interval is equal to
. By dividing the pseudo-counts by 3, we obtain the implied prior distribution of outcome probabilities. Notice that the three populations E, T, and U have vectors
, such that the NPS = 0. However, the left-skewed population with vector
has a NPS =
. The variance (of
) listed in
Table 1 is for a single multinoulli trial using the formula given in Theorem 2.
We use simulation (
draws) to obtain twelve sampling distributions of
associated with different populations using different sample sizes. With
replications per scenario, the maximum Monte Carlo standard error for coverage estimates is
, making simulation variability negligible for the comparisons of interest. For each of the four populations in
Table 1, we use sample sizes of
. In this research, the confidence coefficient
C will be set to 0.95. For each sample, we compute a Wald confidence interval using Equation (
5), and a bootstrap
t confidence interval using Equation (
7). We also compute four adjusted Wald confidence intervals; specifically, AW(3, E), AW(3, LS), AW(3, T), and AW(3, U). As described previously, this simply means using pseudo-counts and, subsequently, the adjusted sample size and adjusted proportions, to obtain
and
(using Theorem 5), which are then input into Equation (
5). Critically, for each of the four populations, three of the adjusted Wald confidences will be misspecified to the extent that the implied prior distribution of outcome probabilities will not align with the true distribution of outcome probabilities for the population from which the sample were drawn. We then compute the average width and proportion coverage of the confidence intervals across the 100,000 samples from the four populations for each sample size and for all six confidence interval types. Note that coverage will be computed with respect to the respective true NPS for each population.
Lastly, we generate three visuals. First, we generate a faceted plot of the twelve sampling distributions of . Second, we generate a line plot of coverage rates that shows how often the six different 95% confidence intervals contain the true NPS across different sample sizes and populations. This evaluates empirical coverage. Third, we generate a line plot of average interval widths that displays the average width of the 95% confidence intervals for each method, again by sample size and population. This reflects the precision and efficiency of the intervals.
4. Discussion
There are limitations we acknowledge in this research. In practice, the assumption that we have obtained a true simple random sample might be plainly suspect. Quite often, we are dealing with a representative sample or, perhaps, a voluntary response sample. Nevertheless, our methods give practitioners a benchmark that could be useful, provided that any interpretation is tempered with caution.
To complement the coverage and width results, we briefly examine the bias of the NPS estimators at . As expected, the original estimator and the bootstrap-based approach exhibit a negligible degree of bias across all population configurations. In contrast, the smoothed estimator underlying adjusted Wald intervals can exhibit non-negligible bias when the implied pseudo-count distribution is misaligned with the true population, with a magnitude approaching 0.10 in our simulations. When the pseudo-counts are well aligned with the underlying population, the bias is effectively zero. These results highlight the bias–variance tradeoff inherent to adjusted Wald methods; improved coverage and stability may come at the cost of bias under model misspecification.
As mentioned in
Section 1, an NPS value of zero represents a natural interpretive benchmark separating Promoter and Detractor subpopulations. In practice, testing whether NPS differs from zero aligns with how the metric would be used in organizational benchmarking and decision-making. For hypothesis testing (e.g., checking if a NPS is significantly different from zero or if two NPS scores differ), we note that similar principles apply. Because
is essentially a sample mean (of coded values
, and 1), a one-sample
t-test can test if NPS equals zero, or some other value under the null hypothesis, by using the estimated standard deviation and
degrees of freedom. In practice, this would be a useful tactic for assessing whether customer loyalty to a product, say, is comparable to that of some predefined standard. Likewise, comparing two independent NPS scores (from, say, two companies or two time periods) can be done with a two-sample
t statistic comparing the difference in
to zero. For example, a two-sample
t-test of NPS scores of a product or service over a period of time is a way to assess whether a statistically significant improvement in customer loyalty has been achieved. These procedures assume approximate normality of the NPS estimator, an assumption that can be checked using our rule of thumb. The main caution is to use the correct variance formula for NPS and not to treat the two proportions (
and
) as if they were independent.
A line of inquiry we have explored in our research has been to raise awareness about the sampling distribution of the NPS statistic and to more rigorously address its inferential properties. Because is a linear combination of multinomial counts, its exact sampling distribution is discrete (with a granularity of in NPS units), being derived from the multinomial probability mass function. Our results confirm that the uncertainty in can be considerable, leading to confidence intervals for NPS that can be very wide for small sample sizes. Overall, while NPS is valued for simplicity, we caution against over-reliance on a single-number metric without proper statistical context and validation when making an inference about a customer base.
Sampling variability in
is something that should not be ignored. Much like high-quality sampling polls, the recommendation we make is to always pair an estimate of NPS with a margin of error or confidence interval. However, in practice, the authors note that this is rarely done. The consequences of ignoring sampling variability can be practically profound and lead to “chasing noise”. Let us reconsider our example with the healthcare system that is considering deploying a new workforce development course to its constituents. Consider a scenario under which, had all the constituents taken the course (the population), the true unknown NPS would have been 0.75. Now, suppose a class consists of 20 students. On average, we would expect
, and yet, it is entirely possible by chance to see an estimate of 0.25 (
Figure 1). The leadership at the healthcare system might wrongly conclude that the course is ineffective when, in fact, the healthcare system at large would enthusiastically embrace the course. As a second scenario, it is possible that a population could have a true unknown NPS of 0, and yet, a class of 20 students might yield an estimate of 0.50 due to chance (
Figure 1). The leadership at the healthcare enterprise might wrongly conclude that the course is excellent when, on average, all of their constituents actually view the course with equanimity.
Many times in practice, there might be neither intuition nor prior knowledge as to the type of population under consideration. Based on our simulation results for the coverage and width of confidence intervals for NPS that account for the multinomial nature of the data, the adjusted Wald methods with triangular or uniform pseudo-counts are good choices across a range of practically relevant population shapes, followed by extreme pseudo-counts.
The mean squared error (MSE) of an estimator is defined to be the sum of the variance of the estimator and the squared bias of the estimator. This provides a useful way to compare estimators with respect to precision and accuracy to see which one is “best”, at least from an MSE perspective. We did not compute the MSEs for the alternative and unbiased estimators for . As expected, the smoothed estimator for NPS underlying adjusted Wald intervals exhibits finite-sample bias when the implied pseudo-count distribution is misaligned with the true population, particularly at very small sample sizes. However, the inferential consequences of this bias are already captured through the empirical coverage results reported here. A systematic evaluation of estimator bias and MSE represents a natural direction for future work.
Our methods are developed assuming an infinite, or “large”, population. If samples are drawn from a finite population, such that the sampling fraction becomes excessively high (e.g., ≥0.10, say), then the resulting confidence intervals for NPS using the methods given in our research will be conservative. Future research could explore attenuating the variance estimator using a finite-population correction factor [
13].
In practice, when we have missing data or nonresponse (e.g., 10 out of 20 students in a class respond to a survey), our treatment of the issue depends on the reason for the missingness and our analysis goals. Practical approaches for resolving the issue depend on the setting and can become quite complex. This would represent another fruitful area of future research. A complete case analysis, the most common approach, assumes that responses are missing completely at random and proceeds with estimation using just the respondents. However, there is a risk of bias in estimation if the nonrespondents differ systematically from the respondents, which occurs when missingness is not completely at random. When in doubt, seek the consultation of a trained statistician.