1. Introduction and Summary
Traditionally, time series use models such as moving average and autoregressive models. The danger of such parametric models is that if the wrong model is chosen, the results will be wrong as well. It is of crucial importance for the asymptotic variance to be correct. Otherwise, the model will not even give asymptotic accuracy. Here, I take a nonparametric approach. The degree of accuracy available is limited to the cases where the results of [
1] can be applied, that is, when the total order of the cross-cumulants is less than nine. For example, this allows the computation of the asymptotic variance and the Central Limit Theorem (CLT) for the sample autocovariance as well as its distribution and quantiles to magnitude
, written as
, where
n is the sample size. This is also the first time a nonparametric CLT has been given for the sample autocorrelation. For higher-order accuracy, or when the number of terms is large, software may be needed to apply or extend [
1], as spelt out in
Appendix B.
CLTs for the sample mean from a stationary series,
, were given in Chapter 4 of [
2], Chapter 18 of [
3], and some of my own papers.
Let
be a stationary series.
the autocovariance. Let
be the mean of
. Then, under regularity conditions,
the normal distribution. We call
and
of (
2) the
approximate variance and the
asymptotic variance of
. It was not until [
4] that this result for
was extended to Edgeworth–Cornish–Fisher (ECF) expansions, that is, the expansions in powers of
for its distribution, density and quantiles. Ref. [
5] did give these for smooth functions of the sample cross-moments of a
linear process, and ref. [
4] extended these results to a general stationary process. For any
I give the
Ith-order ECF expansions for
, that is, these expansions to
, in terms of the cross-cumulants of
up to order
.
For any smooth function of sample cross-moments, this paper gives three types of ECF expansions: (1) the
mean-type expansions, i.e., those about the normal using the exact variance, when this is available (
of (
1) in the case of the sample mean); (2) the
standard-type expansions, i.e., those about the normal using the approximate variance (
of (
1) in the case of
); and (3) the
asymptotic-type expansions, i.e., those about the normal using the asymptotic variance (
of (
1) in the case of
). For
, these were given only recently, in [
4]: for
, they use
for
.
Now suppose that
is a
standard estimator based on a sample of size
n of an unknown parameter
in a statistical model. That is,
as
and for
, its
rth-order cross-cumulants are
and can be expanded in powers of
, with explicit forms for the coefficients needed in these expansions. (The classic example is when
is a smooth function of a sample mean.) Then ECF expansions are available for its distribution that improve upon the first-order expansions, that is, the CLT. In order to be self-contained,
Section 2 summarises these ECF expansions to
.
Per [
6], smooth functions of a standard estimator, such as the sample cross-cumulants and functions of them, are also standard estimators. Thus, these also have the ECF expansions described in
Section 2.
Section 3 gives the cumulant coefficients for a function of an unbiased standard estimator that are needed for the third-order ECF expansions.
In
Section 4, I show that the sample non-central cross-moments of a stationary process are unbiased standard estimators. (In fact, their cumulant expansions have exactly two terms.)
The most important functions of cross-cumulants, are the sample autocovariance,
, and the sample autocorrelation,
. These have been useful tools in time series since [
7]: see [
8].
Corollary 1 gives three choices for the cumulant coefficients needed for these expansions: their mean-type values, their approximate values and their asymptotic values. These are given in terms of cross-cumulants that I call the K functions.
Applications of stationary time series fall into two classes. The first class, dealt with in
Section 5, is when the mean,
, is zero. This includes many branches of radio, radar, sonar, speech, image analysis, communications, control and seismology. This class is dealt with much more easily than the second class, stationary series with an unknown mean. For example, the autocovariance is a cross-cumulant if the mean is zero, but otherwise it is a function of two cross-cumulants.
Section 6,
Section 7,
Section 8,
Section 9 and
Section 10 deal with the second class, that is, when
is unknown.
Section 6 gives the ECF expansions for the distribution to any order
I of
.
Section 7 gives the ECF expansions for the distribution to order two of the sample autocovariance (
).
Section 8 gives the ECF expansions to order one (the CLT) for the sample autocorrelation (
). An exact form is given for
) but is not available for
. Corollaries 3 and 4 spell out for the first time how the asymptotic variances of the sample autocovariance and the sample autocorrelation depend on the mean and cross-cumulants up to order four.
Section 9 gives multivariate CLTs for
and
. These extensions of (
1) and (
2) are all new. However, the examples go beyond CLTs: they provide the distribution, density and quantiles of the estimates up to order
I,
, for various
I.
Other important examples are the sample versions of the autocovariances of degree
,
and their autocorrelations. These have not been studied for
, perhaps because for odd values of
r they are zero for a Gaussian process. CLTs for
are given in Example 4 for
and in
Section 10 for
.
Section 11 shows how to extend the results of
Section 4 to a multivariate stationary process.
Section 12 discusses our results and identifies some future directions worth pursuing.
As noted, our nonparametric approach gives an alternative to traditional parametric time series models, such as a moving average (MA), autoregressive (AR), or ARMA process. Another difficulty with such parametric models is estimating their parameters. A vast collection of software has been developed for economists and others to use for this purpose. Potentially, these parametric time series models and the software for them could be replaced by the nonparametric method given here, where the results of [
1] can be applied or extended to the desired degree of accuracy.
This paper does not deal with observations that have a signal beyond the mean (for example, or a sinusoidal signal.) Furthermore, I do not provide software to help with the large number of terms sometimes needed.
Turning to the literature on ECF expansions, refs. [
9,
10,
11] showed that Cornish–Fisher expansions gave substantial improvements to the CLT, even for lattice estimators such as the binomial and Poisson estimators. Ref. [
12] presented an extension to nonstationary processes. Ref. [
13] obtained an Edgeworth correction by bootstrapping in autoregressions. Ref. [
14] gave an Edgeworth expansion for U-statistics with dependent observations. Ref. [
15] used an Edgeworth expansion to study empirical likelihood methods for dependent processes. For a review of a paper by Taniguchi and Kakizawa on Edgeworth and saddlepoint expansions for time series, see [
16]. Ref. [
17] gave expansions for a maximum likelihood estimator for a stationary Gaussian process.
See [
7,
18,
19,
20,
21,
22,
23,
24] for consistency and some theory and applications. The variance formula of [
7] can produce invalid confidence intervals: see [
25,
26,
27]. This is not surprising because it neglects the mean and higher-order cumulants.
Ref. [
28] gave an Edgeworth expansion for
under long-range dependence. Ref. [
29] gave an Edgeworth expansion for the maximum likelihood estimator of a stationary long-memory Gaussian time series.
Note 1. Refs. [5,30] considered the general stationary linear processwhere are constants, while are independent and identically distributed (i.i.d.) random variables from a distribution on R with finite cumulants . Its cross-cumulants were shown to beand that is finite for processes such as ARMA where exponentially as See [5] for a simulation study. For an autoregressive process , see p. 158 of [5]. Ref. [
31] considered the model
where
are constants, while
are i.i.d. with finite cumulants
where
. In this case, (
5) holds with
2. ECF Expansions for Standard Estimators
This section summarises the Edgeworth–Cornish–Fisher expansions that can be found in [
32,
33].
Univariate estimators. Suppose that
is
a standard estimator of an unknown
with respect to
n, typically the sample size. That is,
as
, and its cumulants can be expanded as
where the
cumulant coefficients may depend on
n but are bounded as
,
is bounded away from 0, and
means that
is bounded in
n. Here and below, ≈ indicates an asymptotic expansion that need not converge. For non-lattice
, the distribution and quantiles of
have expansions in powers of
of the form
where
is a unit normal random variable with density
;
are polynomials in
x and
; and
I call these
the Ith-order ECF expansions. They are given in [
32] to order five, that is, to
, starting as follows:
where, for
is the
kth Hermite polynomial,
Thus, (
8)–(
15) give the
Ith-order ECF expansions for the distribution, density and quantiles of
of (
7), in terms of
for
; in terms of
for
; and in terms of
for
Note 2. Refs. [34,35] gave (11) for for a special case. Their results were extended to standard estimators in [32]. In their examples, (11) gave quantiles to a large number of decimal places. However, for small enough n, divergence can start after the first or second term. In practice, one terminates (8)–(11) if they start to diverge. Mean-type univariate estimators. We say that
is
a mean-type estimator if
and, for
its
rth cumulant has magnitude
, say,
where
is bounded as
. (The classic example is when
is the mean of a random sample from a distribution on
R with mean
w and
rth cumulant
.) For this special type of standard estimator, the above results simplify. If
is replaced with
and other
are replaced with zero, then
and
in (
13)–(
15).
Multivariate estimators. Set
the density and distribution of the multivariate normal
.
I now reserve for any sequence from . To avoid double subscripts, I introduce the bar notation.
The
multivariate Hermite polynomial,
, is
while
is the
element of
Their integrated form is
Now suppose that
is a
standard estimator of
with respect to
n. That is,
as
, and for
,
the
rth order cumulants of
can be expanded as
and the
cumulant coefficients may depend on
n but are bounded as
. Therefore,
Then,
V may depend on
n, but I assume that
is bounded away from zero. By [
6], the distribution and density of
can be expanded as
where
and for
Statements (
20) and (
21) use
the tensor summation convention of implicitly summing
over their range
. The coefficients
are called
the Edgeworth coefficients. These are polynomials in the cumulant coefficients
of (
16). See [
33] for their definition. There, I show that the Edgeworth coefficients needed for Edgeworth expansions to order three, that is, to
, are
where
is the operator that symmetrizes
. Thus,
For
and more details, see [
33]. These give the Edgeworth expansions for the distribution of
to order four, that is, to
, for non-lattice
.
ECF expansions for parametric and nonparametric standard estimators were first given in [
32].
In
Section 4, I show that for the sample cross-moments of a stationary time series, only the first two terms in (
6) and (
16) are non-zero for
Mean-type multivariate estimators. We say that
is
a mean-type estimator if
and, for
its
rth-order cumulants have magnitude
, say,
where
is bounded as
. (The classic example is when
is the mean of a random sample from a distribution on
with mean
w and
rth-order cumulants
.) For this special type of standard estimator, the above results simplify. If
is replaced
and other
are replaced with zero, then (
18)–(
21) hold with
That is, the classic Edgeworth expansions for a sample mean can be applied with this re-interpretation of
.
4. The Cumulants of the Sample Cross-Moments
Let
be any real stationary process with finite mean, non-central cross-moments, central cross-moments, and cross-cumulants,
where
denotes a string of
r zeros. For the relationships between them, see Section 3.30 of [
36]. Multivariate relations can be written down from their univariate versions. For example,
Given a sequence of integers
and
, set
These are not changed by permuting subscripts. Furthermore, at least one
is zero. For
and
of (
41), the
ath autocovariance and autocorrelation are
Transforming from
to
,
Now suppose that one observes only
. For
of (
42) and
, define the
sample non-central cross-moment
This is an unbiased estimator of
of (
40). For example, if
, then
and
have unbiased estimators
Let us write (
44) as
For example, for
, and for
. Given
and sequences of integers
, set
Thus,
. We shall see that (
16) holds with
for
So
Example 1. If and then , of (45), of (46), I now rewrite (
49) in the form (
16).
Definition 1. Given let be a symmetric function of integers such thatI call K a stationary function of degree r.
An example is
of (
41).
The next theorem will give us two choices of
needed for (
6)–(
11): the exact values
of (
55), (
64) below and the asymptotic values
of (
56), (
58) below.
Theorem 1. Let K be a stationary function of degree . Take and integers . SetThen, for ,as , when finite, andNow suppose that Then Proof. Transform from
to
for
. Thus,
, and (
52) holds with the number of permissible
equal to
For example,
(From here, (
54)) and (
64) follow. □
(Thus,
can be negative.) (
52) reduces the
r summations to
r − 1 summations. Note that I use an overline to indicate a limit as
. For the mean-type expansions,
of (
54) is used.
Since
,
of (
59) can be written as
We now come to the crux of this paper. Corollaries 1 and 2 give the cumulant coefficients needed for the leading terms of the Edgeworth expansions for
of (
48). Of fundamental importance are the approximate and asymptotic variances, both for CLTs and for Edgeworth expansions. We apply Theorem 1 to show that
of (
48) is both a standard estimator and a mean-type estimator.
Corollary 1. Take of (48), and of (51)–(53). Then , and for K of (50),Thus, (54)–(62) hold when is replaced with . For example,Thus, the Edgeworth expansions (18) and (19) hold for with , where of (67), and its other elements are defined similarly. For example,Furthermore, is a mean-type estimator. Thus, the Edgeworth expansions (18) and (19) hold for with , where of (68),and the other elements are defined similarly. Proof. . By (
49), (
52), and (
55), for
of (
66),
LHS (
49) =
□
Note that (
66) is an exact result. Alternatively, under mild conditions, one can use the mean-type expansions about
.
When
, there are another two options for
, namely, the exact value
of (
64) or its limit,
of (
58). When
,
of (
63) is simpler than
, and so
is simpler than
, and the expansions of
Section 2 may be simpler if
n is replaced with
as in Example 2. This is where the role of
and
in (
51) is made clear.
We now show that for the univariate version of Corollary 1, one has the option of ECF expansions in powers of
, rather than in powers of
as in
Section 2.
Corollary 2. Take of (63). Given a sequence of integers π, and , of (47), setfor K of (50). Thus, is a standard estimator, and the ECF expansions for of (7), that is, (8)–(11), hold when are replaced with ,For example, is given by (65) with Furthermore, is a mean-type estimator. Thus, the ECF expansions (8)–(11) hold for of (71) with of (69), and replaced by the exact variance of , Alternatively, under mild conditions, one can use the asymptotic-type expansions about .
Let us compare the standard Edgeworth expansion for
of (
71) with the mean-type Edgeworth expansion for
where
is replaced with
.
Under mild conditions, this also holds when
is replaced with
while
is replaced with
.
A similar result holds for the versions of in Corollary 1.
of (
50) can be written in terms of the cross-cumulants of
using p. 254–265 of [
1] if
, where
the length of the sequence
. See their p. 58 and
Appendix B below for some examples. This covers
, which has
, and
, which has
L = 6. but not
, which has
L = 9. Thus, at present, the ECF expansions for
can be obtained to
, but the ECF expansions for
can be obtained only to
. These are both important extensions to their CLTs.
7. Inference for the Autocovariance with Unknown
Consider the standard-type ECF expansions to
for the sample autocovariance,
, where
. Take
in (
22),
The non-zero derivatives are
For
and
r,
of
Section 7, and
of Example 3. Therefore,
By (
34), for
,
This needs
. By (
66) with
This can be inferred by identifying
with
in (
A15).This gives the
needed for
of (
84), the approximate variance of
.
Alternatively, one can use (
84) with
of
Section 6,
of Example 3, and
of (
85), or one can use their asymptotic values under mild conditions
The approximate and asymptotic variance of the sample autocovariance. As noted, the exact, approximate, and asymptotic variances are the most important measures of an estimate. I now spell out for the first time how the approximate and asymptotic variances of the sample autocovariance of a stationary process depends on the cross-cumulants up to order four.
Corollary 3. Set , , and . Then the approximate variance of isThus, the asymptotic variance of is If
, then as
of (
1), the asymptotic variance of
! Therefore, if
, then for large
,
For a Gaussian process, only second-order covariances are non-zero, so that
.
Figure 1,
Figure 2,
Figure 3 and
Figure 4 plot
for
against
for a Gaussian process with
, for lags
(the bottom curves),
and
, (the top curves). The four figures are for
and
and for
and
.
Note how increases rapidly with and correlation r. Doubling n has minor effect.
I now give
and
needed for the second-order standard-type ECF expansions, that is, the expansions to
. By (
34), (
35) and (
83),
for
of (
81), and
of (
60). This still needs
and
. By (
66) with
By (
66) with
This has 25 terms, or eight if
is a Gaussian process. This completes
of (
89). With
of (
88), this gives
of (
8)–(
13), and the standard-type ECF expansions for
to
. For these expansions to
,
of (
14) and (
15) need
and
. By (
36),
By (
37),
is the sum of
Therefore,
is needed.
of (
82).
of (
64) with
K of (
77). Set
By (
66) with
Therefore,
if
is a Gaussian process, but in general, there are nine terms. By (
66) with
By (
66) with
Software is needed for the many terms in (
94) and (
95). This completes
, giving
of (
8)–(
11). These now give the standard-type ECF expansions for the distribution, density and quantiles of
to
. The mean-type and asymptotic-type ECF expansions follow from Note 3.
For the CLT for
, see
Section 9.
8. Inference for the Autocorrelation with Unknown
Consider the standard-type ECF expansions to
for the
ath sample autocorrelation,
. Take
,
Accordingly,
are given by (
38) and (
39) with
and
are
of (
80) and
of (
81).
and
are
and
of (
72).
and
are simply
and
with
.
is
of (
86), and
is
with
.
is
of (
90).
is
with
.
is
of (
91).
is
with
.
is
of Example 4.6 of [
5]. (This corrects the derivatives of
t given in Example 4.5 of [
5]).
The approximate and asymptotic variance of the sample autocorrelation. I now show for the first time how the asymptotic variance of the sample autocorrelation of a stationary process depends on its cross-cumulants up to order four.
Corollary 4. Set and . Take and of (A1) and (87). The approximate variance of the sample autocorrelation with lag a, is , whereThus the asymptotic variance of is , whereFor example, if is a Gaussian process with mean μ and non-centrality parameter then Thus, is linear in the non-centrality parameter .
Proof. Substitute
into
of (
38).
is given by identifying
with
in (
A16). For a Gaussian process, only second-order
are non-zero, and so (
97) and (
98) follow from
□
If
, (
98) simplifies to
If
, where
, then
of (
99) and the asymptotic variance of (
98) are given by
To see that
, note that
Contrast (
98) with the CLT for
given in Theorems 7.2.1 and 7.2.2 of [
31], under certain conditions,
To see that these are wrong, note that they do not include the third-order
from covariance
or the fourth-order
from variance
. Similarly,
appears in more than one of the six terms in (
38) but not in [
31].
Figure 5 plots
of (
98) against lag
a for
equal to
.
This shows that increases much more rapidly with the noncentrality parameter than with the lag a.
I now move on to
needed for the second-order standard-type ECF expansions for
.
of (
39) also needs
Set
. By (
66) with
with 17 terms, or six if
is Gaussian. This gives
. Next I give
. By (
66) with
also needs
. By (
66) with
However, its large number of terms requires software. These now give the distribution and quantiles of
to
. To extend this to
, write out
similarly.
The mean-type and asymptotic-type ECF expansions follow from Note 3.
9. Multivariate CLTs for Sample Autocovariances and Autocorrelations
We now come to the most important example: multivariate CLTs for the estimates of autocovariances and autocorrelations. This example concerns the asymptotic normality of the sample autocovariances, , and the sample autocorrelations, .
Fix and take . For set
I first approximate
as
of (
26). By (
68),
For
, where
Thus, for
, the non-zero first derivatives are
and by (
26), for
of (
87) and for
Thus, by (
25), for
,
One can replace
A with its asymptotic limit
, where
Theorem 8.3.3 of [
37] gave a comparable result that also allowed for the fourth-order cumulant of
when
; see (3.3) of [
38].
Now apply (
25) to
. Under these conditions, the non-zero first derivatives are
and by (
26),
Note how
A and
B depend on the second, third, and fourth-order cross-cumulants, as well as on
.
See the figures following Corollaries 3 and 4 above for special cases of
and
. One can replace
B with its asymptotic limit
, where
This asymptotic result corrects the celebrated Bartlett’s formula, (13) of [
7]. Ref. [
37] and (1.2) of ref. [
38] made corrections when
. Ref. [
38] (p. 53) recognised that Bartlett’s formula did not allow for ’higher-order cumulants’ and gave a correction for when
.
Alternatively, one can use the mean-type forms of A and B.
, is used as a diagnostic for model development in time series; see [
31].
12. Discussion
The CLT (
2) for
, the mean of a sample from a stationary process, is well known. However, not until very recently, in [
4], were its ECF expansions given. CLTs for the sample autocorrelation,
, and for
were given in [
31]. However these did not allow for the first-, third- or fourth-order cross-cumulants. I corrected these shortcomings in
Section 8 and
Section 9.
Section 8 also gives the ECF expansions of
to
, and explains how to improve this to
, as well as how to extend it to the distribution of
.
Empirical distribution functions, kernel density estimation, and resampling techniques such as bootstrapping are all actively used in the analysis of stationary processes. However, while these are nonparametric techniques, unlike the content of this paper, they do not provide the nonparametric theory for the distribution of the general standard estimator based on a sample from a stationary process, given here. Consequently, econometrics, the foremost user of time series, has been dominated by software constructing and applying estimators for parametric ARMA processes and the like. However, as noted, parametric models suffer from a severe drawback: if the model is wrong, the results will generally not even give first-order (CLT) accuracy. This is where a nonparametric method is superior. Furthermore, the Kalman filter and prediction in ARMA time series with their Wold and Kolmogorov representations assume that
; see 8.6 and 10.9 of [
39]. We are less interested in this case, although it is important because it is used frequently in radio, radar, controls and communications work; see [
40] and 1.2 of [
41]. In this case, the terms with
in Corollaries 3 and 4 and, more generally, in
Appendix B are zero.
Software is not needed for third-order ECF expansions for the sample mean,
, the sample autocovariance,
, or the CLT for the sample autocorrelations. However, software is needed for the second-order ECF expansions for the sample autocorrelations because of the large numbers of terms for the
K functions. Software is also needed in Example 4 and
Section 10 when computing the asymptotic variance of the estimates of third-order cross-moments and cross-cumulants.
However in many areas of signal detection, including radio, radar, sonar, speech, image analysis, communications, control and seismology, the time series can often be assumed to be Gaussian; see, for example, [
39,
40,
42]. In this important case, many of the terms in
Appendix B are zero; see, for example, Corollaries 3 and 4.
Future directions.
1. The CLTs for
and
in
Section 7 and
Section 8 and the CLTs of
and
in (
104) and (
105) have many applications, as Bartlett’s formula generally ignores the effect of non-zero mean and higher cross-cumulants.
2. I included only one example in
Section 11 for multivariate stationary series, as it deserves its own paper.
3.
Section 9 gives CLTs for
and
. These can be used for model development in time series, as in [
31]. Extending
Section 9 to Edgeworth expansions of order two or three for
, that is, to
or
, will require software.
4. Extensions to signal plus noise problems, to partial autocorrelation, and to autocorrelation in regression are needed.See [
27,
43] and Chapter VII of [
44].
5. These results can easily be extended to both weighted estimators and estimators based on series with missing data, as performed in [
4].
6. Ref. [
45] gave
and a CLT for weighted sums of a spatial process. It should be feasible to extend this to give the other leading
. A CLT for the spatial autocorrelation would also be useful. See [
46,
47,
48].
7. To extend these results to confidence intervals for a general function of cross-cumulants, say,
, would be an important advancement. The first step is to extend them to Edgeworth expansions for the studentized form
is a consistent estimator of
of (
59), and one can take
for some
. Ref. [
49] gave this to
for the case
, so that
, using
of
Section 7.
8. Similarly, one could estimate
of (
81) with
for
of
Section 10. Other
can be estimated similarly.
9. How sensitive are the results to the estimation of the many higher-order cross-cumulants required? A numerical study to consider this for some examples would be useful.
10. Software to implement the results of
Section 3 for any smooth function of
w would be useful.
11. Software to extend the results of [
1] and to spell out its succinct notation as shown in the
Appendix A below would be useful. For example, results from [
1] allowed me to obtain
for
but not
for
.
12. One could replace my empirical estimator of the cross-cumulant with a less biased estimator. While a jack-knife or bootstrap could be used, analytic results are more easily obtained by adjusting it with an estimator of its bias. However, its advantage is limited, as this may only remove the first term in
of (
30).
13. There is also the prospect of tilted extensions. These extend the applicability of univariate Edgeworth expansions to the whole line; see [
16].