1. Introduction
Forecast evaluation relies on a variety of statistical metrics, each capturing different aspects of predictive quality. While Mean Squared Prediction Error (MSPE) is the most commonly used criterion, other measures, such as Mean Absolute Error (MAE), correlations between forecasts and the target variable, and Mean Directional Accuracy (MDA), are also widely employed (MDA is also known as hit-rate, success rate or sign agreement ratio, and it is considered in many articles studying predictability, including [
1,
2,
3,
4,
5,
6], among many others). MDA, in particular, measures a forecast’s ability to correctly anticipate the direction of change in the target variable, making it especially relevant for investors interested in trading strategies based on directional accuracy rather than MSPE. (See, for instance, Ref. [
7]).
Despite the usefulness of these metrics, there are clear and fundamental differences among them. Most importantly, when comparing two competing forecasts, say, forecasts A and B, it is entirely possible for one metric to favor forecast A while another favors forecast B. A well-known example is the so-called MSPE Paradox: under certain inefficiency conditions, the forecast with the lowest Mean Squared Prediction Error may also exhibit the lowest correlation with the target variable (see [
8]). This paradox challenges the notion that lower MSPE necessarily signals a more informative or useful forecast, particularly when correlation is interpreted as a proxy for signal content.
This paper uncovers another set of paradoxical results, involving the relationship between the correlation of a forecast with the target variable and Mean Directional Accuracy (MDA). We show that, under certain conditions, a forecast that is more strongly correlated with the target may deliver worse sign predictions than a less correlated alternative. In other words, improvements in correlation may coincide with a deterioration in directional performance.
Within a Gaussian framework, we derive analytical expressions that clarify this phenomenon. In particular, we show that sign predictability depends not only on the correlation between the forecast and the target variable, but also on their standardized means, a result that partly echoes the findings of [
9]. These additional dimensions of forecast behavior may distort the relationship between correlation and directional accuracy. As a consequence, higher correlation does not guarantee superior MDA, and in some regions of the parameter space, the relationship can even reverse. An empirical application to exchange rate forecasting illustrates that these paradoxes are not merely theoretical curiosities, but empirically relevant phenomena that can affect forecast evaluation and model selection in practice.
It is important to emphasize that [
9] primarily focuses on what they describe as “the subtle and little-understood connection between sign dynamics and volatility dynamics”, as discussed on page 1276. Our focus is different. We study the relationship between sign predictability, as measured by MDA, and the correlation between the forecast and the target variable, a connection that, to our knowledge, has not been explicitly examined in the existing literature.
The rest of this paper is organized as follows:
Section 2 presents the analytical framework.
Section 3 derives a decomposition of MDA in terms of the correlation between the forecast and the target variable. In that section, we also provide several analytical examples in which the MDA Paradox arises.
Section 4 derives formal conditions under which the MDA Paradox cannot occur, including the case of Mincer–Zarnowitz [
10] efficient forecasts.
Section 5 presents an empirical illustration, and
Section 6 concludes the study.
2. Analytical Framework
Let
denote the time series of interest, which we assume to be integrated of order one, i.e., I(1). Our primary focus is on the variation in this time series over a forecast horizon h, that is, the change between t and t + h. For notational simplicity, we assume in the theoretical sections of this paper that h = 1 (in the empirical illustration, we consider longer horizons as well). Accordingly, we define the first difference as
Under the I(1) assumption on
, the series
is stationary. At time t, we consider two competing forecasts for
and
. These forecasts are constructed using information available at time t, and they are treated as primitives. That is, we do not model the process by which these forecasts are generated, nor do we account for parameter estimation uncertainty. In this respect, our framework is similar in spirit to the setup in [
11], which is further emphasized in [
12]. (Forecast evaluation has been extensively reviewed by [
13,
14]. The literature on nested model comparisons has evolved through the contributions of [
15,
16,
17,
18,
19], with more recent insights provided by [
20]. At the population level, Ref. [
21] discusses the implications of parameter uncertainty, while Ref. [
22] provides a cornerstone framework for evaluating predictive ability in finite-sample and under conditional settings.) For clarity of exposition, we drop the time subscript t in what follows. We will also assume that the joint vector (Y,X,Z) is weakly stationary, ergodic and jointly normally distributed. While the Gaussian assumption may appear restrictive, it provides analytical clarity and tractability. Moreover, it may offer a reasonable approximation to the distribution of actual financial returns at medium to long horizons. While short-term returns are typically characterized by fat-tailed distributions, longer-horizon log returns can be viewed as averages of shorter-term returns. By the Central Limit Theorem, this aggregation process justifies the use of the Gaussian assumption as a plausible approximation for both long-run returns and their forecasts.
We also show results for efficient forecasts in the Mincer–Zarnowitz [
10] sense, that is, unbiased forecasts whose errors are uncorrelated with the information available at the time they are made. This orthogonality condition ensures that no obvious predictable structure remains in the forecast errors and is particularly relevant to our analysis. Mincer–Zarnowitz efficiency plays a central role in preventing the MSPE Paradox from arising. Here, we also examine whether Mincer–Zarnowitz efficiency is sufficient to prevent paradoxical discrepancies between MDA and correlation-based measures of forecast performance. For doing so, it is convenient to recall the exact definition of Mincer–Zarnowitz efficiency and to see how it looks in our framework.
Remark: Mincer–Zarnowitz [10] efficiency. A forecast
is said to be Mincer–Zarnowitz efficient if it has no bias and it is
auto-efficient. A forecast
is unbiased when
. A forecast
is
auto-efficient whenever
; otherwise, it is said to be
auto-inefficient.
In our context, in which we have two forecasts X and Z for the same target variable Y, they will be unbiased if their expected values
and
coincide with the expected value of Y, which we will denote
. So, unbiasedness means
X will be
auto-efficient as long as
where
denotes the correlation between forecast X and Y, and
represent the standard deviations of X and Y respectively. Applying the same reasoning to forecast Z, we have that Z will be
auto-efficient as long as
where
denotes the correlation between forecast Z and Y, and
represent the standard deviations of Z and Y respectively. In the next section, we present our first proposition and a set of examples illustrating the MDA Paradox.
4. Conditions Under Which the MDA Paradox Is Impossible
In this section we focus on establishing two results that provide different sets of conditions under which the MDA Paradox is impossible. We also show a very interesting proposition that allows for a simple computation of MDA differentials when forecasts are unbiased and the target variable has zero mean. Before presenting all these results, we show and prove the following Lemma 2, which is very simple, but we use it later in Proposition 2.
Lemma 2. Let denote the cumulative distribution function of the standard bivariate normal distribution with correlation coefficient , evaluated at the fixed point Then
Proof of Lemma 2. Let (X,Y) be a random vector that follows a standard bivariate normal distribution with correlation coefficient . It follows that
Therefore
and
which is the desired result. □
The following proposition establishes that, under Mincer–Zarnowitz efficiency, paradoxical results with respect to MDA and correlations cannot arise. This result is in line with the MSPE Paradox, which is likewise impossible when forecasts are efficient (see [
6,
8]).
Proposition 2. In the context of Proposition 1, if both forecasts X and Z are Mincer–Zarnowitz [10] efficient, that is to say Then the MDA Paradox is impossible.
Proof of Proposition 2. We begin by expressing the MDA associated with forecast X
By symmetry, the MDA associated with forecast Z can be written as
Let us continue analyzing the function
keeping in mind that the case of
follows by symmetry. Given
and
is an increasing function of
. To see this, we will compute its derivative
The derivative of the univariate Gaussian distribution in the right-hand side is straightforward:
For the derivative of the bivariate Gaussian distribution, we have
where
;
, and
.
From Lemmas 1 and 2 we obtain
Therefore, we have
or simply
Rearranging terms we have
but
Therefore
which means that
Thus,
is an increasing function in
Consequently, whenever
we will have
, which leads to
Hence, the MDA-Correlation Paradox cannot arise when both forecasts are Mincer–Zarnowitz [
10] efficient, which completes the proof. □
Proposition 2 indicates that the MDA Paradox cannot occur when comparing Mincer–Zarnowitz [
10] efficient forecasts. This is an important result. However, since violations of efficiency are fairly common in the forecasting literature, it is important to explore some conditions under which the MDA Paradox may still be prevented in the presence of inefficient forecasts. (Violations of forecast efficiency have been documented across a wide range of variables in numerous studies. See [
29,
30,
31,
32,
33] among others.) This is precisely addressed in the following Lemma 3.
Lemma 3. In the context of Proposition 1, and assuming if the standardized means of both forecasts are the samethen paradoxical results between MDA and correlations are impossible. Proof of Lemma 3: Using (5) in (1) we get
so
If
, Lemma 1 implies that the right-hand side in the previous expression is non-negative, which in turn means that
and no paradoxical results between correlations and MDA can emerge as stated. □
Lemma 3 and Proposition 2 establish clear sufficient conditions under which the MDA Paradox is impossible. It remains to investigate whether a simpler condition, such as unbiasedness, is also sufficient to prevent this paradox. Although the simple example in
Section 3.2.5 shows that this is not the case, it turns out that unbiased forecasts, in the context of a mean-zero target variable, fall within the framework of Lemma 3. (In the numerical illustration in example
Section 3.2.5, we already showed that the MDA Paradox emerges when both forecasts are unbiased. In this particular example we assumed
.) Interestingly, in this setting, a closed-form expression for the difference in MDA is also available, as we shall see next.
Proposition 3. In the context of Proposition 1, if and both forecasts X and Z are unbiased for the mean-zero target variable Y, that is to say
Thenand the MDA Paradox is impossible. Proof of Proposition 3. If X and Y are jointly normal random variables with mean zero, unit variances and correlation coefficient then we have
Using polar coordinates
, we can write this probability as
Let
, then using the following change in variables
we get
Let .
Moreover,
. Thus, the probability
equals
In this last step we have used that arctan is an odd function. Let us express
for
Given that
+
=1, we get
which means
Therefore
which means
. Yet, we started with
which in turn means,
. This leads to the equality
. So we finally have
which is an increasing function of
as given by the derivative
Recalling that
and using the symmetry of the bivariate normal distribution with zero means and equal variances, we get
which finally leads to
which is the desired result.
Figure 2 represents
as a function of
. □
5. Empirical Illustration
In this section, we illustrate the MDA Paradox with an application in the context of the exchange rate forecasting literature. Our target variable is the Chilean Peso (CLP) vis-à-vis the US Dollar. We evaluate forecasts from the Survey of Professional Forecasters (SPF) conducted by the Central Bank of Chile. Specifically, we use monthly observations of the Chilean Peso and SPF forecasts for the period April 2012–April 2024. For exchange rate data, we extract the daily closing price of the CLP from Bloomberg and convert them to monthly frequencies by sampling from the last day of each month. We also sample the closing price from the day before the survey is released. The latter time series is simply denoted as CLP+. We will explain in brief why we consider two different time series for the Chilean Peso. Data from the SPF are directly obtained from the Central Bank of Chile. During the sample period, the survey release date varied between the 9th and 13th of each month.
We will use the following notation: denotes the Chilean Peso at month t. Being more specific, represents the amount of Chilean Pesos required to buy one U.S. Dollar at the closing price of the last day of month t. So, and just to give an example, if month t corresponds to January, corresponds to the closing price of the CLP in January 31st. If follows in this example, that would represent the closing price of the CLP in February 28. We have already mentioned that the survey is released to the public sometime in the middle of each month. We will use the notation to generically denote the closing price of the CLP the day before the survey is released. So, following with the previous example, if the survey is released in February the 10th, then would represent the closing price of the CLP in February the 9th. So, the following inequality describes the timeline in the flow of information: .
We have spent a few lines describing this timeline because we want to assess the predictive ability of the survey, which is released around the middle of each month. So, following with the previous example, if the survey is released in February 10th the difference
is not entirely unknown. This variation represents how the CLP changes from the end of January to the end of February. Yet, at the moment the survey is released we already know what happened in the first nine days of February. More generally
The only unknown variable in the right-hand side of the previous expression is
, so that would be the focus of our interest. The survey asks for forecasts at three different horizons: 2, 11 and 23 months ahead. The Central Bank of Chile provides the median values across all respondents. We label SPF2, SPF11 and SPF23 the time-series containing the median of these forecasts 2, 11 and 23 months ahead, respectively. An important observation is in order: we consider each time-series SPF2, SPF11 and SPF23 as independent forecasts for the Chilean exchange rate and we explore their ability to correctly predict the Chilean Peso at different forecasting horizons
h, varying from
h = 1, 2, 3, 6, 9, 11, 12, 18 and 24 months ahead. So, despite SPF2 being the median of the respondents regarding the question for the Chilean exchange rate 2 months ahead, we explore also if this answer is useful to predict at the shortest horizon of one month, as well as the longer horizons of 3, 6, 9, 11, 12, 18 and 24 months ahead. This is in line with previous work by [
28] who showed that SPF2, SPF11 and SPF23 could be useful forecasts of the Chilean Peso at several horizons.
Figure 3 depicts the Chilean peso (our target variable) during the sample period of interest and the time series of our three forecasts SPF2, SPF11 and SPF23. This figure reveals an upward trend in all four series, which is a clear evidence of a non-stationary behavior. To avoid any potential spurious results that are so frequent in the analysis of trending data or, more generally, non-stationary data, we focus on forecasting Chilean Peso log-returns, rather than the Chilean Peso itself.
Figure 3 shows that all three forecasts, SPF2, SPF11 and SPF23, closely tracked the Chilean Peso during the early part of our sample period, from April 2012 to May 2018. After that, however, the three forecasts display a systematic downward bias, consistently underestimating the value of the US Dollar. Furthermore, in the latter half of the sample, SPF2 tends to lie above SPF11, which in turn tends to lie above SPF23. This ordering suggests that the median forecaster expected the Chilean Peso to strengthen as the forecasting horizon increased, an outcome that did not materialize during this period. Our predictive analysis focuses on the second part of the sample period, from June 2018 to April 2024. We do so partly because Ref. [
28] analyzes the same data set only up to May 2018, and partly because of the previously documented downward bias in the survey. As will become evident below, this biased behavior appears to be a key element underlying the paradoxical results that follow.
We use lower-case letters to denote the natural logarithm of a variable, so that:
The
h-period log-return of the Chilean Peso is defined as
We focus on forecasts of this variable at horizons of
h = 1, 2, 3, 6, 9, 11, 12, 18, 24 months ahead. Notice that this expression represents the logarithmic return of the CLP between the middle of month
t + 1 and the last day of month
t + h. When
h = 1, the return spans only a few days; when
h = 3, it covers roughly two and a half months. For longer horizons, the extension is straightforward. We consider the following forecasts obtained from the survey:
where
and
denotes the forecast of the nominal exchange rate reported by the
at month t + 1. The index
represents the corresponding identifiers of the survey. Recall that we treat SPF2, SPF11, and SPF23 as distinct forecasts for
. Accordingly, in what follows, we evaluate
and
as different forecasts on their own merits.
When choosing a benchmark to compare our forecasts, it is natural to consider the Driftless Random Walk, which predicts zero returns at every forecast horizon. This zero forecast has been a landmark in the exchange rate literature since the seminal article by [
25]. More recently, Ref. [
27] emphasizes that, in exchange rate forecasting, the zero forecast is the toughest benchmark to outperform. Ref. [
34] provides a more recent example in which the zero forecast is used as a benchmark when forecasting exchange rates returns.
Despite its popularity, one of the drawbacks of the zero forecast is that it has no variance, so its correlation with the target variable is not defined. Along the same lines, the zero forecast also exhibits a poor MDA, which is exactly zero. Therefore, this particular forecast requires a slight modification to be used more sensibly.
To address this, let us consider the following benchmark Forecast
, where
is an independent Gaussian white noise process. We denote its variance simply as
. It follows that the correlation between this benchmark forecast and
is exactly zero, while its MDA is 0.5, just as in example
Section 3.2.6Thus, our simple generalization of the traditional zero forecast recovers the standard pure-luck benchmark for MDA, which—to the best of our knowledge—is among the most widely used benchmarks in MDA evaluations. See for instance, Refs. [
33,
35,
36]. It also serves as a benchmark for correlations, because it now has some variance. Of course, the benchmark correlation is exactly zero. Notice that the Gaussian assumption for our benchmark is not essential, as we only need independence, zero mean, some non-negligible variance and a symmetric pdf. around zero. Nevertheless, this assumption is consistent with the analytical framework that we have used in this paper. (In a nutshell, our simple generalization of the traditional zero forecast is designed to achieve three objectives. First, it is closely related to the traditional zero forecast. Second, it delivers an MDA of 0.5, thereby recovering the standard pure-luck benchmark commonly used in MDA evaluations. Third, it provides a meaningful benchmark for correlation-based evaluation. Indeed, it is impossible to compute a correlation between the zero forecast and the target variable, as the former has zero variance. Likewise, adopting 0.5 as a benchmark for MDA alone leaves open the question of what correlation such an implicit benchmark would have with the target variable. Our generalized zero forecast resolves these issues in a unified framework.)
Table 1 next reports the MDA of our survey-based forecasts (Panel 1) and their correlation with Chilean Peso returns (Panel 2). In Panel 1, we present the MDA for SPF2, SFF11 and SPF23 at different horizons. Figures with stars indicate that the null hypothesis of the survey’s MDA being superior to the pure luck benchmark (0.5) is rejected at usual significance levels. In other words, figures with stars indicate that the 0.5 benchmark outperforms the corresponding survey with statistical significance. Panel 2 reports the correlation between each survey and Chilean Peso returns. Figures with stars indicate correlations that are statistically significant and positive.
In the first panel of
Table 1, only two entries show MDA rising slightly above 50%, while the overall average MDA is a disappointing 44.5%, illustrating the survey’s poor performance on this metric. Moreover, eight entries in this panel display statistically significant results in favor of the 0.5 benchmark. In stark contrast, we see only positive figures in the second panel of the table, with an average correlation between the survey and the target variable of 29%. The maximum correlation is as high as 60%, achieved with SPF23 when forecasting 9 months ahead. Moreover, 19 out of 27 entries in this second panel report statistically significant and positive correlations.
The MDA Paradox emerges clearly when comparing the two panels of
Table 1: in the first panel, the survey is almost always outperformed by our naïve benchmark in terms of MDA, whereas in the second panel, the survey consistently outperforms the same benchmark in terms of correlations. The MDA Paradox is most evident in SPF2 forecasts at 9, 11, and 12 months ahead: despite MDA figures being significantly lower than the naïve 0.5 benchmark, their correlations with Chilean Peso returns remain positive and statistically significant. (As a robustness check, we split the predictive sample into two balanced subperiods and re-estimated both MDA and correlation measures separately for each window. While this exercise reveals some heterogeneity in the levels of MDA and correlations across subsamples, the qualitative message remains unchanged. Evidence of the MDA Paradox—both in its weak and, in some cases, strong form—continues to emerge in both subperiods, reinforcing the robustness and empirical relevance of our main findings. A detailed version of this robustness analysis is available upon request.)
Results in
Table 1 can be better understood in light of our example in
Section 3.2.6, where the MDA Paradox emerges in an environment in which we compare two forecasts X and Z for our target variable Y, such that Z and Y are totally independent, Z is mean zero with a symmetric distribution, X is positively correlated with the target variable (
), and
. As a matter of fact,
Table 2 next shows averages of both our target variable and survey-based forecasts as a proxy for
and
.
Results in
Table 2 show that while forecasts are centered in negative territory, our target variable has a positive mean at every forecast horizon. Nonetheless, this last fact alone is not enough to obtain the MDA Paradox. In
Section 3.2.6, we additionally used the sufficient condition
to guarantee the Paradox. Yet, it is not hard to derive from Expression (1) the following equality that is necessary and sufficient for the MDA Paradox to occur in our specific case
Expression (6) indicates that we will have the MDA Paradox whenever
As a final exercise, we can evaluate the following expression
using sample estimates of the parameters
,
,
and
from the data of our forecasts and target variable, to check if our Gaussian approximation matches the empirical MDA presented in
Table 1. On average, our Gaussian approach is extremely accurate. The median MDA reported from
Table 1 is 46.5%, while the median MDA obtained from Expression (7) across all forecasts and horizons is 47.1%. For particular entries of
Table 1 the Gaussian approach is extremely close (46.4% vs. 46.5% for SPF 23 when h = 3), but for other entries this Gaussian approximation is not very accurate (48.5% vs. 35.2% for SPF2 when h = 6). All in all, our Gaussian approximation seems useful to explain the MDA Paradox in this example on the aggregate level.