1. Introduction
Transformation-based modelling is a standard strategy in statistics. A nonlinear transformation is often applied to the response variable, or to several variables, in order to improve normality, stabilize variance, or simplify the dependence structure. After fitting a model in the transformed scale, fitted values, regression curves, or conditional summaries are commonly mapped back to the original scale by applying the inverse transformation. This transform–fit–inverse procedure is natural and widely used, but it raises a fundamental question: which statistical structures are preserved when a model fitted in the transformed scale is returned to the original scale?
This question is particularly important for conditional expectations. In regression analysis, the conditional mean is often the primary object of interpretation and prediction. However, conditional expectation is a linear operator, whereas the inverse of a nonlinear transformation is itself nonlinear. As a consequence, the inverse transformation of a conditional mean in the transformed scale does not generally coincide with the conditional mean in the original scale. In operator terms, nonlinear transformations do not commute with conditional expectation.
The Box–Cox transformation, introduced by [
1], provides a classical and analytically tractable setting in which this issue can be studied. It is widely used for variance stabilization, improvement of normality, and simplification of functional relationships in statistical modelling [
2,
3,
4]. It is also closely related to the Power Normal distribution family introduced by [
5], which provides a probabilistic framework for transformation-based Gaussian modelling [
6,
7,
8]. The transformation continues to be used in modern applications, including data-driven modelling and machine-learning contexts [
9,
10].
The phenomenon studied here is related to, but distinct from, the classical retransformation bias problem. Retransformation bias is primarily concerned with corrected prediction or mean estimation on the original scale after fitting a model in a transformed scale [
11,
12,
13,
14,
15]. In contrast, the present paper addresses a structural invariance question: whether the conditional mean operator itself is preserved by a transform–fit–inverse modelling pipeline.
The contribution of this paper is to make explicit this structural non-commutativity principle in the Box–Cox framework. If
h is a nonlinear inverse transformation and
V is a transformed response, then in general
Equality can hold under restrictive conditions, such as affine transformations or degenerate conditional distributions, but it is not a generic property of nonlinear transformation-based models. Thus, the inverse image of a conditional mean in the transformed scale is generally a structural curve induced by the transformation, not the conditional mean function in the original scale.
Within the Box–Cox setting, the admissible-domain issue must also be taken into account. The inverse Box–Cox map is defined only on
Thus, an unrestricted Gaussian latent variable is not globally compatible with the inverse Box–Cox map whenever
. For this reason, the exact probabilistic formulation considered in this paper is based on the Gaussian density restricted to the admissible domain, that is, on a truncated normal formulation in the transformed scale. In that formulation, the discrepancy between the original-scale conditional mean and the inverse-transformed conditional mean is governed, to second order, by the curvature of the inverse transformation and by the conditional variance in the transformed domain.
Transformation methods have a long history in statistical modelling and have been used not only to improve approximate normality, but also to reduce skewness, stabilize variance and simplify regression [
1,
16]. The Box–Cox transformation is one of the most widely used members of this tradition, but it is part of a broader class of transformation-based methods, including later extensions such as the Yeo–Johnson transformation for variables that may take nonpositive values [
17]. Multivariate versions of the Box–Cox transformation have also been proposed in order to address joint normality and dependence structure [
18].
Related transformation ideas also appear in Gaussian copula modelling and in modern statistical learning procedures, where nonlinear marginal transformations are often used to separate marginal behaviour from dependence modelling. The present paper focuses on a different, but related, issue: the effect of nonlinear inverse transformations on conditional means after fitting a model in the transformed scale.
The transform–fit–invert pipeline considered in this paper can be summarized schematically as
The structural discrepancy arises in the final step because
unless the inverse transformation is affine or the conditional dispersion is degenerate.
The paper is organized as follows.
Section 2 formulates the general non-commutativity principle for conditional expectations under nonlinear transformations.
Section 3 introduces the Box–Cox transformation, its inverse, and the corresponding Gaussian framework with admissible-domain truncation.
Section 4 derives the second-order distortion formula.
Section 5 presents a numerical sensitivity analysis.
Section 6 provides a real-data illustration and a sensitivity analysis with respect to the transformation of the explanatory variable.
Section 7 discusses practical implications and concludes.
2. Conditional Expectation, Projection, and Non-Commutativity
Let
be a probability space and let
. For a sub-
-algebra
, the conditional expectation
is the orthogonal projection of
V onto the closed subspace
of
-measurable square-integrable random variables. Thus, if
denotes this projection operator, then
In particular, conditional expectation is a linear operator.
In contrast, if
is nonlinear, the map
is nonlinear. Hence there is, in general, no reason to expect
or equivalently
The failure of this identity is the non-commutativity mechanism studied in this paper.
Theorem 1. Non-commutativity of nonlinear transformations and conditional expectation. Let be a continuous function on an interval . Suppose that, for every square-integrable random variable V taking values in I, and for every sub-σ-algebra ,Then h is affine on I. Equivalently, no genuinely nonlinear transformation can preserve conditional expectations for all admissible conditional distributions. Proof. It is enough to consider the trivial
-algebra
. Then the assumed identity becomes
Let
and
. If
V is such that
then
The assumed identity therefore implies
Thus
h satisfies Jensen’s equality on
I. Since
h is continuous, it follows that
h is affine on
I. □
The continuity assumption is used only to exclude pathological solutions of Jensen’s functional equation. The same conclusion holds under weaker regularity assumptions, such as measurability or local boundedness on an interval. Under such assumptions, Jensen’s equality implies that h is affine. In the present statistical setting this regularity requirement is natural, since the transformations used in transformation-based modelling, including the Box–Cox transformation, are continuous and smooth on their admissible domains.
The affine case is the exceptional case. If
, then
Thus affine transformations preserve conditional expectations. The Taylor expansions derived below do not establish the non-commutativity itself; rather, they quantify its leading contribution in smooth transformation models.
3. The Box–Cox Transformation and Gaussian Framework
For
, the Box–Cox transformation is defined by
with inverse
For
, the transformation is understood in the limiting sense,
The admissible domain of the inverse transformation is
Thus
if
,
if
, and
.
The first two derivatives of the inverse transformation are
Consequently, the inverse Box–Cox transformation is affine only when
, for which
For every
, the inverse map has nonzero curvature on its admissible domain. Since the leading contribution to the discrepancy between
and
depends on the curvature of
, the parameter
controls not only the marginal transformation of the data but also the induced conditional-mean distortion.
We now specialize to a transformation-based Gaussian framework. Let
and
be random variables on the original scale, and define
where the transformation parameters may be different. A natural Gaussian reference model in the transformed scale is
with
where
,
, and
.
Because the inverse Box–Cox maps are defined only on their admissible domains, the exact probabilistic formulation is the Gaussian density restricted to
Thus the transformed vector is interpreted as
The unrestricted Gaussian distribution is used only as a reference density, or as an approximation when the probability of the inadmissible region is negligible. For a fixed admissible value of
u, the Gaussian reference conditional distribution is
where
Here
denotes the conditional variance in the transformed scale, distinct from the marginal variance
. In the exact admissible-domain formulation, this conditional normal distribution is truncated to
. We denote the corresponding truncated conditional mean and variance by
and
When
, no truncation is required and
,
.
The structural discrepancy in the exact admissible-domain formulation is
Since
, conditioning on
is equivalent to conditioning on
. Thus,
whereas the inverse-transformed truncated conditional mean is
4. Second-Order Distortion Formula
We now derive an explicit local representation of the structural distortion. In the exact admissible-domain formulation, we write
where
Since
, the original-scale conditional mean is
whereas the inverse-transformed truncated conditional mean is
. The following result should therefore be understood as a local small-dispersion expansion. More precisely, it describes the leading contribution to the discrepancy as the truncated conditional distribution of
concentrates around its truncated conditional mean
. In this sense, the remainder term is interpreted in the asymptotic regime
. The second-order formula should be interpreted as a local small-dispersion approximation. Its accuracy depends on the concentration of the truncated conditional distribution of
around
and on the smoothness of the inverse transformation over the region where this conditional distribution places most of its mass.
If
is three times continuously differentiable on an interval
containing the relevant support of the truncated conditional distribution, Taylor’s theorem with remainder gives
with
Taking conditional expectations gives
For the Box–Cox inverse transformation,
with the limiting exponential case obtained when
. Therefore, the approximation is expected to be most reliable when the conditional dispersion in the transformed scale is small and when
is not close to the boundary of the admissible domain
.
Proposition 1. Second-order structural distortion. Assume that and that the truncated conditional distribution of is concentrated in a neighbourhood where is twice differentiable. Then, in the local small-dispersion regime ,where Proof. Using
we expand
around
:
Taking conditional expectations under
and
, and using the fact that
, gives
□
Using
the leading term becomes
Thus the distortion is governed by the transformed-scale conditional variance, the departure from the affine case, and the location of
within the admissible domain. Expressed as a function of the original covariate
x, with
,
This shows that the discrepancy between the true conditional mean in the original scale and the inverse-transformed transformed-scale conditional mean is not a constant bias, but a structural function of the covariate.
When the probability assigned to the inadmissible region is negligible, the truncated conditional moments are close to their untruncated Gaussian counterparts:
In this local approximation,
The logarithmic case
provides a useful benchmark. Then
and
. Since
, no truncation is required. If
is Gaussian, the exact expression is
Hence
showing explicitly that the inverse-transformed conditional mean
is not the conditional mean in the original scale.
Figure 1 illustrates the same mechanism numerically. The true conditional mean on the original scale is compared with the inverse-transformed linear predictor obtained from the transformed scale. Although the two curves are nearly indistinguishable over most of the domain, the zoom on the right tail reveals a small but systematic discrepancy. This confirms visually that the original-scale conditional mean is not, in general, simply the inverse image of the linear conditional mean in the transformed scale.
Quantitative summary measures for representative parameter configurations are reported in
Table 1. The discrepancy was also evaluated quantitatively over the plotted grid, using the maximum absolute difference, the root-mean-square difference and the maximum relative difference between the two curves. These measures confirm that the discrepancy is numerically small over most of the domain, but systematic and more pronounced in the right-tail region.
The table confirms the analytical predictions and highlights the most extreme configurations. The largest relative discrepancies occur when the conditional dispersion in the transformed scale is larger and the inverse transformation is more strongly nonlinear. For example, increasing from to , with and , substantially increases all distortion measures. Conversely, larger values of reduce the distortion by reducing the conditional dispersion of . The affine case provides the benchmark in which the distortion is essentially zero. The scale-free measure is particularly useful for comparing parameter configurations, since it expresses the maximum discrepancy relative to the corresponding original-scale conditional mean.
5. Numerical Sensitivity Analysis
The numerical analysis focuses on the range
This range was chosen in order to include the logarithmic case
, genuinely nonlinear positive transformations, and the affine benchmark
, for which the structural distortion vanishes. The parameters varied in the numerical experiments are
and
Negative values of
are also relevant in applied Box–Cox modelling, but they lead to an upper-bounded admissible domain
and may involve stronger interactions between truncation, boundary behaviour and curvature of the inverse transformation. A systematic study of negative transformation parameters is therefore left for future work. Unless otherwise stated, we set
.
The quantities , , and are evaluated by Monte Carlo integration from the conditional truncated normal distribution.
For each parameter configuration and each value of u on the numerical grid, the Monte Carlo integration was performed using simulated draws from the corresponding conditional normal distribution, retaining only draws in the admissible domain . The same simulation size was used across all parameter configurations. A fixed random-number seed was used to ensure reproducibility. The computations were carried out in MATLAB R2023a using double-precision arithmetic.
The MATLAB script used for the numerical sensitivity analysis and simulated examples is provided as
Supplementary Material as Code S1. The second-order approximation to the distortion is given by
Figure 2 shows the effect of
, with
and
. As expected, the distortion decreases as
approaches the affine case
.
Figure 3 shows the effect of the correlation parameter. Larger values of
reduce the conditional dispersion in the transformed scale and therefore attenuate the curvature-induced distortion.
Figure 4 shows the effect of
. Increasing
increases the conditional dispersion of
, thereby amplifying the distortion induced by the nonlinear inverse transformation.
Here
denotes the grid of admissible
u-values used in the numerical evaluation. To summarize the numerical results,
Table 1 reports
and
6. A Real-Data Illustration
Let
Y denote the variable
charges. The MATLAB script used for the real-data insurance illustration is provided as
Supplementary Material Code S2. Since
Y is strictly positive and strongly right-skewed, it provides a natural empirical setting for a Box–Cox transformation. We fitted a Gaussian linear regression model in the transformed response scale. In the main graphical display, BMI was transformed using
, and the model included age, transformed BMI, number of children, smoking status, sex, region, and the interactions age × smoker and transformed BMI × smoker as covariates. The response transformation parameter was selected by profile likelihood. The maximum was attained at
indicating a clearly nonlinear transformation of the response variable. For graphical purposes, the fitted curves are displayed as functions of BMI, separately for smokers and non-smokers, with the remaining covariates fixed at reference values. Thus, the figure should be interpreted as a two-dimensional section of the fitted multivariable regression surface, rather than as a simple univariate regression of
charges on BMI. For each value of BMI in a grid, and separately for smokers and non-smokers, we computed two fitted curves on the original scale. The first is the usual transform–fit–invert curve,
The second is the second-order structural approximation to the conditional mean on the original scale,
where
is the fitted conditional mean in the transformed scale and
is the residual variance in that scale. For the Box–Cox inverse transformation,
the second derivative is
Figure 5 compares the usual transform–fit–invert curves with the corresponding structural second-order approximations. The two curves are close, but they are not identical. The discrepancy is systematic and has the same origin as in the theoretical development: the nonlinear inverse Box–Cox transformation does not commute with conditional expectation.
The relative discrepancy
is shown in
Figure 6. Over the displayed BMI range, the structural approximation differs from the usual transform–fit–invert predictor by approximately
to
. For non-smokers, the discrepancy is nearly constant, around
, whereas for smokers it ranges from slightly above
at lower BMI values to approximately
at higher BMI values.
To examine whether the transformation of the explanatory variable affects the magnitude of the discrepancy, we performed a sensitivity analysis with respect to
. The response transformation was kept fixed at the profile-likelihood estimate
, while the BMI transformation parameter was varied over the grid
For each value of
, the regression model was refitted in the transformed scale, and the relative discrepancy between the transform–fit–invert predictor and the second-order structural approximation was computed over the displayed BMI range and for both smoking groups. The results are reported in
Table 2.
In this empirical example, the effect of on the magnitude of the discrepancy is limited. The mean relative discrepancy varies only from approximately to , while the maximum discrepancy remains close to for all values of . Thus, although the transformation of the explanatory variable can slightly modify the numerical magnitude of the discrepancy, it does not remove its structural origin. The discrepancy remains a consequence of applying a nonlinear inverse transformation to a conditional summary computed in the transformed response scale. This empirical illustration supports the main message of the paper. Even in a simple real-data regression setting, the curve obtained by fitting a linear Gaussian model in the transformed scale and then applying the inverse Box–Cox transformation should not automatically be interpreted as the conditional mean curve on the original scale. The observed discrepancy is not a numerical artefact of the theoretical examples nor a consequence of a particular choice of in this illustration; it is a structural consequence of applying a nonlinear inverse transformation to a conditional summary computed in another scale.
7. Practical Implications and Conclusions
The main practical implication is that an inverse-transformed fitted value should not automatically be interpreted as a conditional mean on the original scale. In the admissible-domain formulation, the inverse-transformed fitted curve is
whereas the conditional mean of the original response is
These quantities coincide only under special conditions, such as an affine inverse transformation or a degenerate conditional distribution. In general, they answer different statistical questions: the former is the inverse image of a transformed-scale conditional mean, whereas the latter is the mean predictor under squared-error loss on the original scale.
The second-order expansion provides a diagnostic for this discrepancy:
Thus transformation-based regression should distinguish between the transformed-scale conditional mean, the inverse-transformed fitted curve, and the original-scale conditional mean. Conflating these objects may lead to incorrect interpretation of regression curves, fitted values, conditional effects, and uncertainty summaries on the original scale.
The analysis also clarifies the relation with classical retransformation bias. Retransformation bias is usually concerned with correcting predictions after returning from a transformed scale to the original scale. The perspective developed here is structural: nonlinear inverse transformations do not, in general, preserve the conditional mean operator. This discrepancy is not a finite-sample effect, an estimation error, or a numerical artefact.
Although the paper focuses on conditional means, the same invariance question is relevant for other regression summaries. The answer depends on the statistical functional being considered. Conditional quantiles behave differently from means: under strictly monotone transformations, quantiles are equivariant, so that the inverse transformation of a transformed-scale conditional quantile corresponds to the conditional quantile on the original scale, provided that the conditioning structure is treated consistently. In contrast, conditional variances, prediction intervals and other uncertainty summaries are not generally preserved in a simple way by nonlinear inverse transformations, because such transformations change scale, curvature and dispersion asymmetrically. Thus, the present analysis should be viewed as the conditional-mean case of a broader invariance problem: which statistical summaries are preserved, and which are structurally altered, by transform–fit–inverse modelling procedures. A systematic treatment of quantiles, variances, prediction intervals and other regression summaries lies beyond the scope of the present paper, but represents a natural direction for future work.
The results do not argue against nonlinear transformations. Transformations such as Box–Cox may remain useful for improving normality, stabilizing variance, or simplifying dependence structures. However, once fitted relationships are mapped back to the original scale, the target of interpretation changes unless the non-commutativity with conditional expectation is explicitly accounted for.
From a practical point of view, the structural approximation is most relevant when the target of inference is the conditional mean of the response on the original scale. This includes situations in which inverse-transformed fitted values are used as mean predictions, regression curves are interpreted on the original scale, or covariate effects are discussed after back-transformation. In such cases, the transform–fit–invert curve and the original-scale conditional mean represent different statistical objects.
The magnitude of the discrepancy is expected to be larger when the inverse transformation is strongly nonlinear over the relevant range and when the conditional dispersion in the transformed scale is not negligible. In contrast, if the inverse transformation is nearly affine over the region of interest, or if the transformed-scale conditional variance is very small, the discrepancy may be numerically minor. Practitioners should therefore distinguish between using a transformation as a modelling device in the transformed scale and interpreting inverse-transformed fitted values as mean predictions on the original scale.
Future work may proceed in several directions. One direction is to develop higher-order correction formulas and assess their finite-sample behaviour. Another is to study analogous structural distortions for other transformation families, including transformations designed for zero or negative data. More broadly, the non-commutativity framework may be extended beyond Box–Cox models to other nonlinear transformations, filtering operations, and dependence structures.
Overall, the paper shows that inverse Box–Cox distortion is not merely a technical bias-correction issue. It reflects a structural limitation of nonlinear transformation-based modelling: transformations may simplify the data in one scale, but they do not generally preserve conditional mean structure when mapped back to the original scale.