Next Article in Journal
Umbral Methods, Function Factorisation and Mittag–Leffler Fourier-Type Integral Transform
Next Article in Special Issue
Local-Time Sensitivity and Burst Instability for Threshold Functionals of One-Dimensional Diffusions
Previous Article in Journal
Existence of Measurable Versions of Stochastic Processes
Previous Article in Special Issue
Quantile Reparameterized Regression Model and Machine Learning with Long-Term Survivors
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

On the Structural Distortion Induced by the Inverse Box–Cox Transformation

Department of Civil and Georesources Engineering, Faculty of Engineering, Universidade do Porto, 4200-465 Porto, Portugal
Axioms 2026, 15(7), 519; https://doi.org/10.3390/axioms15070519
Submission received: 29 May 2026 / Revised: 2 July 2026 / Accepted: 6 July 2026 / Published: 10 July 2026
(This article belongs to the Special Issue Probability Theory and Stochastic Processes: Theory and Applications)

Abstract

The Box–Cox transformation is widely used to improve normality, stabilize variance, and enable Gaussian-based modelling in a transformed scale. After model fitting, conditional summaries are often mapped back to the original scale by applying the inverse transformation. This paper shows that this transform–fit–inverse procedure has a structural limitation: nonlinear inverse transformations do not, in general, preserve conditional expectations. Equivalently, conditional expectation and nonlinear inversion do not commute. Within the Box–Cox Gaussian framework, the admissible domain of the inverse transformation leads naturally to a truncated normal formulation in the transformed scale. Under this formulation, we derive a second-order decomposition showing that the original-scale conditional mean differs from the inverse-transformed truncated conditional mean by a curvature-driven correction term depending on the truncated conditional variance. The usual untruncated Gaussian expression is recovered as a local approximation when the inadmissible probability is negligible. A numerical sensitivity analysis, focused on 0 λ Y 1 , illustrates how the distortion depends on the transformation parameter, correlation, and conditional dispersion. A real-data illustration using medical insurance charges further shows that the discrepancy can be visible in an applied regression setting and is not removed by changing the transformation of the explanatory variable. The results distinguish this structural invariance problem from classical retransformation bias and show that inverse-transformed fitted curves should be interpreted as transformation-induced structural curves, not automatically as conditional mean functions on the original scale.

1. Introduction

Transformation-based modelling is a standard strategy in statistics. A nonlinear transformation is often applied to the response variable, or to several variables, in order to improve normality, stabilize variance, or simplify the dependence structure. After fitting a model in the transformed scale, fitted values, regression curves, or conditional summaries are commonly mapped back to the original scale by applying the inverse transformation. This transform–fit–inverse procedure is natural and widely used, but it raises a fundamental question: which statistical structures are preserved when a model fitted in the transformed scale is returned to the original scale?
This question is particularly important for conditional expectations. In regression analysis, the conditional mean is often the primary object of interpretation and prediction. However, conditional expectation is a linear operator, whereas the inverse of a nonlinear transformation is itself nonlinear. As a consequence, the inverse transformation of a conditional mean in the transformed scale does not generally coincide with the conditional mean in the original scale. In operator terms, nonlinear transformations do not commute with conditional expectation.
The Box–Cox transformation, introduced by [1], provides a classical and analytically tractable setting in which this issue can be studied. It is widely used for variance stabilization, improvement of normality, and simplification of functional relationships in statistical modelling [2,3,4]. It is also closely related to the Power Normal distribution family introduced by [5], which provides a probabilistic framework for transformation-based Gaussian modelling [6,7,8]. The transformation continues to be used in modern applications, including data-driven modelling and machine-learning contexts [9,10].
The phenomenon studied here is related to, but distinct from, the classical retransformation bias problem. Retransformation bias is primarily concerned with corrected prediction or mean estimation on the original scale after fitting a model in a transformed scale [11,12,13,14,15]. In contrast, the present paper addresses a structural invariance question: whether the conditional mean operator itself is preserved by a transform–fit–inverse modelling pipeline.
The contribution of this paper is to make explicit this structural non-commutativity principle in the Box–Cox framework. If h is a nonlinear inverse transformation and V is a transformed response, then in general
E [ h ( V ) U ] h { E [ V U ] } .
Equality can hold under restrictive conditions, such as affine transformations or degenerate conditional distributions, but it is not a generic property of nonlinear transformation-based models. Thus, the inverse image of a conditional mean in the transformed scale is generally a structural curve induced by the transformation, not the conditional mean function in the original scale.
Within the Box–Cox setting, the admissible-domain issue must also be taken into account. The inverse Box–Cox map is defined only on
D λ = { u R : 1 + λ u > 0 } .
Thus, an unrestricted Gaussian latent variable is not globally compatible with the inverse Box–Cox map whenever λ 0 . For this reason, the exact probabilistic formulation considered in this paper is based on the Gaussian density restricted to the admissible domain, that is, on a truncated normal formulation in the transformed scale. In that formulation, the discrepancy between the original-scale conditional mean and the inverse-transformed conditional mean is governed, to second order, by the curvature of the inverse transformation and by the conditional variance in the transformed domain.
Transformation methods have a long history in statistical modelling and have been used not only to improve approximate normality, but also to reduce skewness, stabilize variance and simplify regression [1,16]. The Box–Cox transformation is one of the most widely used members of this tradition, but it is part of a broader class of transformation-based methods, including later extensions such as the Yeo–Johnson transformation for variables that may take nonpositive values [17]. Multivariate versions of the Box–Cox transformation have also been proposed in order to address joint normality and dependence structure [18].
Related transformation ideas also appear in Gaussian copula modelling and in modern statistical learning procedures, where nonlinear marginal transformations are often used to separate marginal behaviour from dependence modelling. The present paper focuses on a different, but related, issue: the effect of nonlinear inverse transformations on conditional means after fitting a model in the transformed scale.
The transform–fit–invert pipeline considered in this paper can be summarized schematically as
( X , Y ) ( U , V ) = { g λ X ( X ) , g λ Y ( Y ) } E [ V U = u ] g λ Y 1 { E [ V U = u ] } .
The structural discrepancy arises in the final step because
g λ Y 1 { E [ V U = u ] } E { g λ Y 1 ( V ) U = u }
unless the inverse transformation is affine or the conditional dispersion is degenerate.
The paper is organized as follows. Section 2 formulates the general non-commutativity principle for conditional expectations under nonlinear transformations. Section 3 introduces the Box–Cox transformation, its inverse, and the corresponding Gaussian framework with admissible-domain truncation. Section 4 derives the second-order distortion formula. Section 5 presents a numerical sensitivity analysis. Section 6 provides a real-data illustration and a sensitivity analysis with respect to the transformation of the explanatory variable. Section 7 discusses practical implications and concludes.

2. Conditional Expectation, Projection, and Non-Commutativity

Let ( Ω , F , P ) be a probability space and let V L 2 ( Ω , F , P ) . For a sub- σ -algebra G F , the conditional expectation E [ V G ] is the orthogonal projection of V onto the closed subspace L 2 ( G ) of G -measurable square-integrable random variables. Thus, if  P G denotes this projection operator, then
P G V = E [ V G ] .
In particular, conditional expectation is a linear operator.
In contrast, if  h : R R is nonlinear, the map V h ( V ) is nonlinear. Hence there is, in general, no reason to expect
P G { h ( V ) } = h ( P G V ) ,
or equivalently
E [ h ( V ) G ] = h { E [ V G ] } .
The failure of this identity is the non-commutativity mechanism studied in this paper.
Theorem 1. 
Non-commutativity of nonlinear transformations and conditional expectation. Let h : I R be a continuous function on an interval I R . Suppose that, for every square-integrable random variable V taking values in I, and for every sub-σ-algebra G F ,
E [ h ( V ) G ] = h { E [ V G ] } .
Then h is affine on I. Equivalently, no genuinely nonlinear transformation can preserve conditional expectations for all admissible conditional distributions.
Proof. 
It is enough to consider the trivial σ -algebra G = { , Ω } . Then the assumed identity becomes
E [ h ( V ) ] = h { E [ V ] } .
Let a , b I and p ( 0 , 1 ) . If V is such that
P ( V = a ) = p , P ( V = b ) = 1 p ,
then
E [ V ] = p a + ( 1 p ) b , E [ h ( V ) ] = p h ( a ) + ( 1 p ) h ( b ) .
The assumed identity therefore implies
p h ( a ) + ( 1 p ) h ( b ) = h { p a + ( 1 p ) b } .
Thus h satisfies Jensen’s equality on I. Since h is continuous, it follows that h is affine on I.    □
The continuity assumption is used only to exclude pathological solutions of Jensen’s functional equation. The same conclusion holds under weaker regularity assumptions, such as measurability or local boundedness on an interval. Under such assumptions, Jensen’s equality implies that h is affine. In the present statistical setting this regularity requirement is natural, since the transformations used in transformation-based modelling, including the Box–Cox transformation, are continuous and smooth on their admissible domains.
The affine case is the exceptional case. If  h ( v ) = a + b v , then
E [ h ( V ) G ] = a + b E [ V G ] = h { E [ V G ] } .
Thus affine transformations preserve conditional expectations. The Taylor expansions derived below do not establish the non-commutativity itself; rather, they quantify its leading contribution in smooth transformation models.

3. The Box–Cox Transformation and Gaussian Framework

For λ 0 , the Box–Cox transformation is defined by
g λ ( x ) = x λ 1 λ , x > 0 ,
with inverse
h λ ( u ) = ( 1 + λ u ) 1 / λ , 1 + λ u > 0 .
For λ = 0 , the transformation is understood in the limiting sense,
g 0 ( x ) = log x , h 0 ( u ) = exp ( u ) .
The admissible domain of the inverse transformation is
D λ = { u R : 1 + λ u > 0 } .
Thus D λ = ( 1 / λ , ) if λ > 0 , D λ = ( , 1 / λ ) if λ < 0 , and  D 0 = R .
The first two derivatives of the inverse transformation are
h λ ( u ) = ( 1 + λ u ) 1 / λ 1 , h λ ( u ) = ( 1 λ ) ( 1 + λ u ) 1 / λ 2 .
Consequently, the inverse Box–Cox transformation is affine only when λ = 1 , for which
g 1 ( x ) = x 1 , h 1 ( u ) = 1 + u , h 1 ( u ) = 0 .
For every λ 1 , the inverse map has nonzero curvature on its admissible domain. Since the leading contribution to the discrepancy between
E [ h λ ( V ) U ]
and
h λ { E [ V U ] }
depends on the curvature of h λ , the parameter λ controls not only the marginal transformation of the data but also the induced conditional-mean distortion.
We now specialize to a transformation-based Gaussian framework. Let X > 0 and Y > 0 be random variables on the original scale, and define
U = g λ X ( X ) , V = g λ Y ( Y ) ,
where the transformation parameters may be different. A natural Gaussian reference model in the transformed scale is
( U , V ) N 2 ( μ , Σ ) ,
with
μ = μ U μ V , Σ = σ U 2 ρ σ U σ V ρ σ U σ V σ V 2 ,
where σ U > 0 , σ V > 0 , and  1 < ρ < 1 .
Because the inverse Box–Cox maps are defined only on their admissible domains, the exact probabilistic formulation is the Gaussian density restricted to
A = D λ X × D λ Y .
Thus the transformed vector is interpreted as
( U , V ) ( U , V ) A .
The unrestricted Gaussian distribution is used only as a reference density, or as an approximation when the probability of the inadmissible region is negligible. For a fixed admissible value of u, the Gaussian reference conditional distribution is
V U = u N ( m ( u ) , τ 2 ) ,
where
m ( u ) = μ V + ρ σ V σ U ( u μ U ) , τ 2 = σ V 2 ( 1 ρ 2 ) .
Here τ 2 denotes the conditional variance in the transformed scale, distinct from the marginal variance σ V 2 . In the exact admissible-domain formulation, this conditional normal distribution is truncated to D λ Y . We denote the corresponding truncated conditional mean and variance by
m T ( u ) = E [ V U = u , V D λ Y ]
and
τ T 2 ( u ) = Var ( V U = u , V D λ Y ) .
When λ Y = 0 , no truncation is required and m T ( u ) = m ( u ) , τ T 2 ( u ) = τ 2 .
The structural discrepancy in the exact admissible-domain formulation is
Δ T ( u ) = E [ h λ Y ( V ) U = u , V D λ Y ] h λ Y { m T ( u ) } .
Since X = h λ X ( U ) , conditioning on X = x is equivalent to conditioning on U = g λ X ( x ) . Thus,
E [ Y X = x ] = E [ h λ Y ( V ) U = g λ X ( x ) , V D λ Y ] ,
whereas the inverse-transformed truncated conditional mean is
h λ Y m T { g λ X ( x ) } .

4. Second-Order Distortion Formula

We now derive an explicit local representation of the structural distortion. In the exact admissible-domain formulation, we write
V = m T ( u ) + ε T ,
where
E [ ε T U = u , V D λ Y ] = 0 , E [ ε T 2 U = u , V D λ Y ] = τ T 2 ( u ) .
Since Y = h λ Y ( V ) , the original-scale conditional mean is
E [ Y U = u , V D λ Y ] = E [ h λ Y ( V ) U = u , V D λ Y ] ,
whereas the inverse-transformed truncated conditional mean is h λ Y { m T ( u ) } . The following result should therefore be understood as a local small-dispersion expansion. More precisely, it describes the leading contribution to the discrepancy as the truncated conditional distribution of V U = u concentrates around its truncated conditional mean m T ( u ) . In this sense, the remainder term is interpreted in the asymptotic regime τ T 2 ( u ) 0 . The second-order formula should be interpreted as a local small-dispersion approximation. Its accuracy depends on the concentration of the truncated conditional distribution of V U = u around m T ( u ) and on the smoothness of the inverse transformation over the region where this conditional distribution places most of its mass.
If h λ Y is three times continuously differentiable on an interval I u D λ Y containing the relevant support of the truncated conditional distribution, Taylor’s theorem with remainder gives
h λ Y ( V ) = h λ Y { m T ( u ) } + h λ Y { m T ( u ) } ( V m T ( u ) ) + 1 2 h λ Y { m T ( u ) } ( V m T ( u ) ) 2 + R 3 ( u ) ,
with
| R 3 ( u ) | M 3 ( u ) 6 | V m T ( u ) | 3 , M 3 ( u ) = sup z I u h λ Y ( 3 ) ( z ) .
Taking conditional expectations gives
E [ R 3 ( u ) U = u , V D λ Y ] M 3 ( u ) 6 E | V m T ( u ) | 3 U = u , V D λ Y .
For the Box–Cox inverse transformation,
h λ ( 3 ) ( z ) = ( 1 λ ) ( 1 2 λ ) ( 1 + λ z ) 1 / λ 3 ,
with the limiting exponential case obtained when λ = 0 . Therefore, the approximation is expected to be most reliable when the conditional dispersion in the transformed scale is small and when m T ( u ) is not close to the boundary of the admissible domain D λ Y .
Proposition 1. 
Second-order structural distortion. Assume that m T ( u ) D λ Y and that the truncated conditional distribution of V U = u is concentrated in a neighbourhood where h λ Y is twice differentiable. Then, in the local small-dispersion regime τ T 2 ( u ) 0 ,
Δ T ( u ) = 1 2 τ T 2 ( u ) h λ Y { m T ( u ) } + o τ T 2 ( u ) ,
where
Δ T ( u ) = E [ h λ Y ( V ) U = u , V D λ Y ] h λ Y { m T ( u ) } .
Proof. 
Using
V = m T ( u ) + ε T ,
we expand h λ Y around m T ( u ) :
h λ Y { m T ( u ) + ε T } = h λ Y { m T ( u ) } + ε T h λ Y { m T ( u ) } + ε T 2 2 h λ Y { m T ( u ) } + o ( ε T 2 ) .
Taking conditional expectations under U = u and V D λ Y , and using the fact that E [ ε T U = u , V D λ Y ] = 0 , gives
E [ h λ Y ( V ) U = u , V D λ Y ] = h λ Y { m T ( u ) } + 1 2 h λ Y { m T ( u ) } τ T 2 ( u ) + o τ T 2 ( u ) .
   □
Using
h λ Y ( z ) = ( 1 λ Y ) ( 1 + λ Y z ) 1 / λ Y 2 ,
the leading term becomes
Δ T ( u ) 1 2 τ T 2 ( u ) ( 1 λ Y ) { 1 + λ Y m T ( u ) } 1 / λ Y 2 .
Thus the distortion is governed by the transformed-scale conditional variance, the departure from the affine case, and the location of m T ( u ) within the admissible domain. Expressed as a function of the original covariate x, with  u = g λ X ( x ) ,
Δ T , X ( x ) 1 2 τ T 2 { g λ X ( x ) } ( 1 λ Y ) 1 + λ Y m T { g λ X ( x ) } 1 / λ Y 2 .
This shows that the discrepancy between the true conditional mean in the original scale and the inverse-transformed transformed-scale conditional mean is not a constant bias, but a structural function of the covariate.
When the probability assigned to the inadmissible region is negligible, the truncated conditional moments are close to their untruncated Gaussian counterparts:
m T ( u ) m ( u ) , τ T 2 ( u ) τ 2 = σ V 2 ( 1 ρ 2 ) .
In this local approximation,
Δ ( u ) 1 2 σ V 2 ( 1 ρ 2 ) ( 1 λ Y ) { 1 + λ Y m ( u ) } 1 / λ Y 2 .
The logarithmic case λ Y = 0 provides a useful benchmark. Then h 0 ( z ) = exp ( z ) and h 0 ( z ) = exp ( z ) . Since D 0 = R , no truncation is required. If  V U = u is Gaussian, the exact expression is
E [ exp ( V ) U = u ] = exp m ( u ) + 1 2 σ V 2 ( 1 ρ 2 ) .
Hence
Δ ( u ) = exp { m ( u ) } exp 1 2 σ V 2 ( 1 ρ 2 ) 1 ,
showing explicitly that the inverse-transformed conditional mean exp { m ( u ) } is not the conditional mean in the original scale.
Figure 1 illustrates the same mechanism numerically. The true conditional mean on the original scale is compared with the inverse-transformed linear predictor obtained from the transformed scale. Although the two curves are nearly indistinguishable over most of the domain, the zoom on the right tail reveals a small but systematic discrepancy. This confirms visually that the original-scale conditional mean is not, in general, simply the inverse image of the linear conditional mean in the transformed scale.
Quantitative summary measures for representative parameter configurations are reported in Table 1. The discrepancy was also evaluated quantitatively over the plotted grid, using the maximum absolute difference, the root-mean-square difference and the maximum relative difference between the two curves. These measures confirm that the discrepancy is numerically small over most of the domain, but systematic and more pronounced in the right-tail region.
The table confirms the analytical predictions and highlights the most extreme configurations. The largest relative discrepancies occur when the conditional dispersion in the transformed scale is larger and the inverse transformation is more strongly nonlinear. For example, increasing σ V from 0.5 to 2.0 , with  λ Y = 0.5 and ρ = 0.8 , substantially increases all distortion measures. Conversely, larger values of ρ reduce the distortion by reducing the conditional dispersion of V U . The affine case λ Y = 1 provides the benchmark in which the distortion is essentially zero. The scale-free measure R max is particularly useful for comparing parameter configurations, since it expresses the maximum discrepancy relative to the corresponding original-scale conditional mean.

5. Numerical Sensitivity Analysis

The numerical analysis focuses on the range
0 λ Y 1 .
This range was chosen in order to include the logarithmic case λ Y = 0 , genuinely nonlinear positive transformations, and the affine benchmark λ Y = 1 , for which the structural distortion vanishes. The parameters varied in the numerical experiments are
λ Y { 0 , 0.25 , 0.5 , 0.75 , 1 } , ρ { 0.2 , 0.5 , 0.8 , 0.95 } ,
and
σ V { 0.5 , 1 , 1.5 , 2 } .
Negative values of λ Y are also relevant in applied Box–Cox modelling, but they lead to an upper-bounded admissible domain
D λ Y = ( , 1 / λ Y ) ,
and may involve stronger interactions between truncation, boundary behaviour and curvature of the inverse transformation. A systematic study of negative transformation parameters is therefore left for future work. Unless otherwise stated, we set λ X = 0.3 .
The quantities M ( u ) , m T ( u ) , and  τ T 2 ( u ) are evaluated by Monte Carlo integration from the conditional truncated normal distribution.
For each parameter configuration and each value of u on the numerical grid, the Monte Carlo integration was performed using M = 2 × 10 5 simulated draws from the corresponding conditional normal distribution, retaining only draws in the admissible domain D λ Y . The same simulation size was used across all parameter configurations. A fixed random-number seed was used to ensure reproducibility. The computations were carried out in MATLAB R2023a using double-precision arithmetic.
The MATLAB script used for the numerical sensitivity analysis and simulated examples is provided as Supplementary Material as Code S1. The second-order approximation to the distortion is given by
Δ 2 ( u ) = 1 2 τ T 2 ( u ) h λ Y { m T ( u ) } .
Figure 2 shows the effect of λ Y , with  ρ = 0.8 and σ V = 1 . As expected, the distortion decreases as λ Y approaches the affine case λ Y = 1 .
Figure 3 shows the effect of the correlation parameter. Larger values of ρ reduce the conditional dispersion in the transformed scale and therefore attenuate the curvature-induced distortion.
Figure 4 shows the effect of σ V . Increasing σ V increases the conditional dispersion of V U , thereby amplifying the distortion induced by the nonlinear inverse transformation.
Here U denotes the grid of admissible u-values used in the numerical evaluation. To summarize the numerical results, Table 1 reports
D max = max u U | Δ ( u ) | , D rms = 1 | U | u U Δ ( u ) 2 1 / 2 ,
and
R max = max u U Δ ( u ) M ( u ) .

6. A Real-Data Illustration

Let Y denote the variable charges. The MATLAB script used for the real-data insurance illustration is provided as Supplementary Material Code S2. Since Y is strictly positive and strongly right-skewed, it provides a natural empirical setting for a Box–Cox transformation. We fitted a Gaussian linear regression model in the transformed response scale. In the main graphical display, BMI was transformed using λ X = 0.5 , and the model included age, transformed BMI, number of children, smoking status, sex, region, and the interactions age × smoker and transformed BMI × smoker as covariates. The response transformation parameter was selected by profile likelihood. The maximum was attained at
λ ^ Y = 0.125 ,
indicating a clearly nonlinear transformation of the response variable. For graphical purposes, the fitted curves are displayed as functions of BMI, separately for smokers and non-smokers, with the remaining covariates fixed at reference values. Thus, the figure should be interpreted as a two-dimensional section of the fitted multivariable regression surface, rather than as a simple univariate regression of charges on BMI. For each value of BMI in a grid, and separately for smokers and non-smokers, we computed two fitted curves on the original scale. The first is the usual transform–fit–invert curve,
m ^ TFI ( x ) = g λ ^ Y 1 E ^ g λ ^ Y ( Y ) X = x .
The second is the second-order structural approximation to the conditional mean on the original scale,
m ^ SA ( x ) = g λ ^ Y 1 { μ ^ ( x ) } + 1 2 g λ ^ Y 1 { μ ^ ( x ) } σ ^ 2 ,
where μ ^ ( x ) is the fitted conditional mean in the transformed scale and σ ^ 2 is the residual variance in that scale. For the Box–Cox inverse transformation,
g λ 1 ( z ) = ( 1 + λ z ) 1 / λ ,
the second derivative is
g λ 1 ( z ) = ( 1 λ ) ( 1 + λ z ) 1 / λ 2 .
Figure 5 compares the usual transform–fit–invert curves with the corresponding structural second-order approximations. The two curves are close, but they are not identical. The discrepancy is systematic and has the same origin as in the theoretical development: the nonlinear inverse Box–Cox transformation does not commute with conditional expectation.
The relative discrepancy
100 m ^ SA ( x ) m ^ TFI ( x ) m ^ TFI ( x )
is shown in Figure 6. Over the displayed BMI range, the structural approximation differs from the usual transform–fit–invert predictor by approximately 4 % to 7 % . For non-smokers, the discrepancy is nearly constant, around 6.5 % , whereas for smokers it ranges from slightly above 5 % at lower BMI values to approximately 4 % at higher BMI values.
To examine whether the transformation of the explanatory variable affects the magnitude of the discrepancy, we performed a sensitivity analysis with respect to λ X . The response transformation was kept fixed at the profile-likelihood estimate λ ^ Y = 0.125 , while the BMI transformation parameter was varied over the grid
λ X { 0 , 0.125 , 0.25 , 0.5 , 0.75 , 1 } .
For each value of λ X , the regression model was refitted in the transformed scale, and the relative discrepancy between the transform–fit–invert predictor and the second-order structural approximation was computed over the displayed BMI range and for both smoking groups. The results are reported in Table 2.
In this empirical example, the effect of λ X on the magnitude of the discrepancy is limited. The mean relative discrepancy varies only from approximately 5.81 % to 5.84 % , while the maximum discrepancy remains close to 7 % for all values of λ X . Thus, although the transformation of the explanatory variable can slightly modify the numerical magnitude of the discrepancy, it does not remove its structural origin. The discrepancy remains a consequence of applying a nonlinear inverse transformation to a conditional summary computed in the transformed response scale. This empirical illustration supports the main message of the paper. Even in a simple real-data regression setting, the curve obtained by fitting a linear Gaussian model in the transformed scale and then applying the inverse Box–Cox transformation should not automatically be interpreted as the conditional mean curve on the original scale. The observed discrepancy is not a numerical artefact of the theoretical examples nor a consequence of a particular choice of λ X in this illustration; it is a structural consequence of applying a nonlinear inverse transformation to a conditional summary computed in another scale.

7. Practical Implications and Conclusions

The main practical implication is that an inverse-transformed fitted value should not automatically be interpreted as a conditional mean on the original scale. In the admissible-domain formulation, the inverse-transformed fitted curve is
h λ Y { m T ( u ) } ,
whereas the conditional mean of the original response is
E [ h λ Y ( V ) U = u , V D λ Y ] .
These quantities coincide only under special conditions, such as an affine inverse transformation or a degenerate conditional distribution. In general, they answer different statistical questions: the former is the inverse image of a transformed-scale conditional mean, whereas the latter is the mean predictor under squared-error loss on the original scale.
The second-order expansion provides a diagnostic for this discrepancy:
C T ( u ) = 1 2 h λ Y { m T ( u ) } τ T 2 ( u ) .
Thus transformation-based regression should distinguish between the transformed-scale conditional mean, the inverse-transformed fitted curve, and the original-scale conditional mean. Conflating these objects may lead to incorrect interpretation of regression curves, fitted values, conditional effects, and uncertainty summaries on the original scale.
The analysis also clarifies the relation with classical retransformation bias. Retransformation bias is usually concerned with correcting predictions after returning from a transformed scale to the original scale. The perspective developed here is structural: nonlinear inverse transformations do not, in general, preserve the conditional mean operator. This discrepancy is not a finite-sample effect, an estimation error, or a numerical artefact.
Although the paper focuses on conditional means, the same invariance question is relevant for other regression summaries. The answer depends on the statistical functional being considered. Conditional quantiles behave differently from means: under strictly monotone transformations, quantiles are equivariant, so that the inverse transformation of a transformed-scale conditional quantile corresponds to the conditional quantile on the original scale, provided that the conditioning structure is treated consistently. In contrast, conditional variances, prediction intervals and other uncertainty summaries are not generally preserved in a simple way by nonlinear inverse transformations, because such transformations change scale, curvature and dispersion asymmetrically. Thus, the present analysis should be viewed as the conditional-mean case of a broader invariance problem: which statistical summaries are preserved, and which are structurally altered, by transform–fit–inverse modelling procedures. A systematic treatment of quantiles, variances, prediction intervals and other regression summaries lies beyond the scope of the present paper, but represents a natural direction for future work.
The results do not argue against nonlinear transformations. Transformations such as Box–Cox may remain useful for improving normality, stabilizing variance, or simplifying dependence structures. However, once fitted relationships are mapped back to the original scale, the target of interpretation changes unless the non-commutativity with conditional expectation is explicitly accounted for.
From a practical point of view, the structural approximation is most relevant when the target of inference is the conditional mean of the response on the original scale. This includes situations in which inverse-transformed fitted values are used as mean predictions, regression curves are interpreted on the original scale, or covariate effects are discussed after back-transformation. In such cases, the transform–fit–invert curve and the original-scale conditional mean represent different statistical objects.
The magnitude of the discrepancy is expected to be larger when the inverse transformation is strongly nonlinear over the relevant range and when the conditional dispersion in the transformed scale is not negligible. In contrast, if the inverse transformation is nearly affine over the region of interest, or if the transformed-scale conditional variance is very small, the discrepancy may be numerically minor. Practitioners should therefore distinguish between using a transformation as a modelling device in the transformed scale and interpreting inverse-transformed fitted values as mean predictions on the original scale.
Future work may proceed in several directions. One direction is to develop higher-order correction formulas and assess their finite-sample behaviour. Another is to study analogous structural distortions for other transformation families, including transformations designed for zero or negative data. More broadly, the non-commutativity framework may be extended beyond Box–Cox models to other nonlinear transformations, filtering operations, and dependence structures.
Overall, the paper shows that inverse Box–Cox distortion is not merely a technical bias-correction issue. It reflects a structural limitation of nonlinear transformation-based modelling: transformations may simplify the data in one scale, but they do not generally preserve conditional mean structure when mapped back to the original scale.

Supplementary Materials

The following supporting information is available online: https://www.mdpi.com/article/10.3390/axioms15070519/s1, Code S1: MATLAB script for the numerical sensitivity analysis and simulated examples, including all required local functions; Code S2: MATLAB script for the real-data insurance illustration, including all required local functions; Text S1: README file describing the supplementary materials.

Funding

This research received no external funding.

Data Availability Statement

The medical insurance charges dataset used in this study is publicly available online from the Kaggle repository, Medical Cost Personal Datasets, at https://www.kaggle.com/datasets/mirichoi0218/insurance (accessed on 5 July 2026). The numerical illustrations were based on simulated data generated using MATLAB R2023a scripts. The MATLAB scripts used to generate the simulated data, numerical sensitivity analysis, real-data illustration, and figures are provided as Supplementary Materials accompanying this manuscript.

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. Box, G.E.; Cox, D.R. An Analysis of Transformations. J. R. Stat. Soc. Ser. B 1964, 26, 211–243. [Google Scholar] [CrossRef] [Scilit]
  2. Draper, N.R.; Smith, H. Applied Regression Analysis, 1st ed.; Wiley: Hoboken, NJ, USA, 1966. [Google Scholar]
  3. Carroll, R.J.; Ruppert, D. Transformation and Weighting in Regression; Chapman and Hall/CRC: Boca Raton, FL, USA, 1988. [Google Scholar]
  4. Sakia, R.M. The Box–Cox Transformation Technique: A Review. Statistician 1992, 41, 169–178. [Google Scholar] [CrossRef] [Scilit]
  5. Goto, M.; Inoue, T. Some Properties of the Power-Normal Distribution. J. Biom. 1980, 1, 28–54. [Google Scholar] [CrossRef] [Scilit]
  6. Freeman, J.; Modarres, R. Inverse Box-Cox: The Power-Normal Distribution. Stat. Probab. Lett. 2006, 76, 764–772. [Google Scholar] [CrossRef] [Scilit]
  7. Goto, M.; Inoue, T.; Tsuchya, Y. On the Estimation of Parameters in the Power-Normal Distribution. Bull. Inf. Cybern. 1984, 21, 41–53. [Google Scholar] [CrossRef] [Scilit]
  8. Goto, M.; Hamasaki, T. The Bivariate Power-Normal Distribution. Bull. Inform. Cybernet. 2002, 34, 29–49. [Google Scholar] [CrossRef] [Scilit]
  9. Blum, L.; Elgendi, M.; Menon, C. Impact of Box-Cox Transformation on Machine-Learning Algorithms. Front. Artif. Intell. 2022, 5, 877569. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Alshamrani, A. Machine Learning Techniques Improving the Box–Cox Transformation in Breast Cancer Prediction. Electronics 2025, 14, 3173. [Google Scholar] [CrossRef] [Scilit]
  11. Duan, N. Smearing Estimate: A Nonparametric Retransformation Method. J. Am. Stat. Assoc. 1983, 78, 605–610. [Google Scholar] [CrossRef]
  12. Miller, D.M. Reducing Transformation Bias in Curve Fitting. Am. Stat. 1984, 38, 124–126. [Google Scholar] [CrossRef] [Scilit]
  13. Taylor, J.M.G. The Retransformed Mean after a Fitted Power Transformation. J. Am. Stat. Assoc. 1986, 81, 114–118. [Google Scholar] [CrossRef]
  14. Sakia, R.M. Retransformation Bias: A Look at the Box–Cox Transformation to Linear Models. Commun. Stat. Simul. Comput. 1990, 19, 189–203. [Google Scholar]
  15. Manning, W.G.; Mullahy, J. Estimating Log Models: To Transform or Not to Transform? J. Health Econ. 2001, 20, 461–494. [Google Scholar] [CrossRef] [Scilit]
  16. Hoyle, M.H. Transformations: An Introduction and a Bibliography. Int. Stat. Rev. 1973, 41, 203–223. [Google Scholar] [CrossRef] [Scilit]
  17. Yeo, I.-K.; Johnson, R.A. A New Family of Power Transformations to Improve Normality or Symmetry. Biometrika 2000, 87, 954–959. [Google Scholar] [CrossRef] [Scilit]
  18. Velilla, S. A Note on the Multivariate Box–Cox Transformation to Normality. Stat. Probab. Lett. 1993, 17, 259–263. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Comparison between the true conditional mean on the original scale and the inverse-transformed linear predictor obtained from the transformed scale. The curves are nearly indistinguishable over most of the domain, while the right-tail zoom reveals a small but systematic discrepancy.
Figure 1. Comparison between the true conditional mean on the original scale and the inverse-transformed linear predictor obtained from the transformed scale. The curves are nearly indistinguishable over most of the domain, while the right-tail zoom reveals a small but systematic discrepancy.
Axioms 15 00519 g001
Figure 2. Structural distortion Δ ( u ) for different values of λ Y [ 0 , 1 ] , with ρ = 0.8 and σ V = 1 .
Figure 2. Structural distortion Δ ( u ) for different values of λ Y [ 0 , 1 ] , with ρ = 0.8 and σ V = 1 .
Axioms 15 00519 g002
Figure 3. Structural distortion Δ ( u ) for different values of ρ , with λ Y = 0.5 and σ V = 1 .
Figure 3. Structural distortion Δ ( u ) for different values of ρ , with λ Y = 0.5 and σ V = 1 .
Axioms 15 00519 g003
Figure 4. Structural distortion Δ ( u ) for different values of σ V , with λ Y = 0.5 and ρ = 0.8 .
Figure 4. Structural distortion Δ ( u ) for different values of σ V , with λ Y = 0.5 and ρ = 0.8 .
Axioms 15 00519 g004
Figure 5. Medical insurance charges data. Comparison between the usual transform–fit–invert predictor and the second-order structural approximation to the conditional mean on the original scale, as a function of BMI and smoking status, for λ X = 0.5 . The displayed curves are two-dimensional sections of a multivariable fitted model. The example is intended only as an empirical illustration of the structural discrepancy, not as an optimal predictive model for medical charges.
Figure 5. Medical insurance charges data. Comparison between the usual transform–fit–invert predictor and the second-order structural approximation to the conditional mean on the original scale, as a function of BMI and smoking status, for λ X = 0.5 . The displayed curves are two-dimensional sections of a multivariable fitted model. The example is intended only as an empirical illustration of the structural discrepancy, not as an optimal predictive model for medical charges.
Axioms 15 00519 g005
Figure 6. Relative difference between the second-order structural approximation and the usual transform–fit–invert predictor in the medical insurance charges data.
Figure 6. Relative difference between the second-order structural approximation and the usual transform–fit–invert predictor in the medical insurance charges data.
Axioms 15 00519 g006
Table 1. Summary measures of the structural distortion for selected parameter configurations. Here R max denotes the maximum relative distortion over the numerical grid.
Table 1. Summary measures of the structural distortion for selected parameter configurations. Here R max denotes the maximum relative distortion over the numerical grid.
λ Y ρ σ V D max D rms R max (Relative)
0.250.501.00 4.47495 × 10 1 3.27621 × 10 1 2.86238 × 10 1
0.250.801.00 2.65228 × 10 1 1.72324 × 10 1 1.92296 × 10 1
0.250.951.00 7.95341 × 10 2 4.90362 × 10 2 6.68236 × 10 2
0.500.800.50 2.26690 × 10 2 2.24999 × 10 2 3.75044 × 10 2
0.500.801.00 9.07455 × 10 2 8.81265 × 10 2 1.99848 × 10 1
0.500.802.00 3.62352 × 10 1 3.04694 × 10 1 3.58205 × 10 1
0.750.801.00 4.43805 × 10 2 3.76786 × 10 2 9.15514 × 10 2
1.000.801.00 8.88178 × 10 16 2.38561 × 10 16 6.04498 × 10 16
Table 2. Sensitivity of the relative discrepancy to different values of λ X , with λ Y = 0.125 .
Table 2. Sensitivity of the relative discrepancy to different values of λ X , with λ Y = 0.125 .
λ X MinimumMeanMaximumStandard Deviation
04.19095.80756.94341.1302
0.1254.18495.80876.94271.1301
0.2504.17995.81076.94301.1301
0.5004.17325.81756.94711.1304
0.7504.17065.82806.95541.1313
1.0004.17245.84196.96801.1328
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Gonçalves, R. On the Structural Distortion Induced by the Inverse Box–Cox Transformation. Axioms 2026, 15, 519. https://doi.org/10.3390/axioms15070519

AMA Style

Gonçalves R. On the Structural Distortion Induced by the Inverse Box–Cox Transformation. Axioms. 2026; 15(7):519. https://doi.org/10.3390/axioms15070519

Chicago/Turabian Style

Gonçalves, Rui. 2026. "On the Structural Distortion Induced by the Inverse Box–Cox Transformation" Axioms 15, no. 7: 519. https://doi.org/10.3390/axioms15070519

APA Style

Gonçalves, R. (2026). On the Structural Distortion Induced by the Inverse Box–Cox Transformation. Axioms, 15(7), 519. https://doi.org/10.3390/axioms15070519

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop