Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
On Testing Equal Conditional Predictive Ability Under Measurement Error
\baselineskip18pt
\setcounter{totalnumber}{50}
\setcounter{topnumber}{50}
\setcounter{bottomnumber}{50}
\abovedisplayskip1.5ex plus1ex minus1ex
\belowdisplayskip1.5ex plus1ex minus1ex
\abovedisplayshortskip1.5ex plus1ex minus1ex
\belowdisplayshortskip1.5ex plus1ex minus1ex
abstractLoss functions are widely used to compare several competing forecasts. However, forecast comparisons are often based on mismeasured proxy variables for the true target. We introduce the concept of exact robustness to measurement error for loss functions and fully characterize this class of loss functions as the Bregman class.
For such exactly robust loss functions, forecast loss differences are on average unaffected by the use of proxy variables and, thus, inference on conditional predictive ability can be carried out as usual. Moreover, we show that more precise proxies give predictive ability tests higher power in discriminating between competing forecasts. Simulations illustrate the different behavior of exactly robust and non-robust loss functions. An empirical application to US GDP growth rates demonstrates that it is easier to discriminate between forecasts issued at different horizons if a better proxy for GDP growth is used. \\
Keywords: Equal Predictive Ability, Forecasting, Hypothesis Testing, Measurement Error \\
JEL classification: C12 (Hypothesis Testing), C52 (Model Evaluation, Validation, and Selection), C53 (Forecasting and Prediction Methods)
bibunit\section{Motivation}
\doublespacing
Due to the central role of forecasts in economic policy, business, climate research and beyond, forecast comparisons have a long tradition. Such comparisons rely on a (statistically or economically motivated) loss function that measures the loss as a function of the issued forecast and the realization of the target variable.
Since the seminal contribution of DM95, tests of equal predictive ability (EPA) have played a central role in comparing competing forecasts. GW06 extend EPA tests by introducing tests of equal conditional predictive ability (ECPA), where the conditioning is, e.g., on current economic conditions. The null hypothesis of ECPA tests is that the conditional mean of the forecast losses are identical.
E(C)PA tests are studied extensively under estimation error in the forecasts (see, e.g., Wes96, ClarkMcCracken2001, Pat20). However, the effect of measurement error in the observed target variable has not received as much attention. Some exceptions---to be discussed below---are the works of Pat11, Laurent2013 and LP18.
Nonetheless, measurement error is present in many economic and financial time series. Examples in economics include the gross domestic product (GDP) Aea16, inflation rates CM09,FS16, and job earnings AS13.
In finance, the conditional variance can only be approximated by the squared return or high-frequency measures such as realized volatility Aea13.
Examples beyond economics and finance include, among others, meteorological applications such as the measurement of precipitation or wind speeds Fer17.
As a consequence, many forecast comparisons are carried out with approximated, mismeasured target variables, also called proxies.
In such a case, the forecast losses---as measured by some loss function---may be systematically different from those obtained using the actual, but latent, target variable. Hence, differences in predictive ability may be clouded by the use of such proxies. In this paper, we derive conditions under which ECPA tests can be validly carried out if only some (conditionally unbiased) proxy for the target variable is available. If several alternative proxies are available, we further derive conditions which proxy entails the most powerful tests.
To do so, we define a loss function to be exactly robust to measurement error if the (conditional) expectation of the forecast loss differences is unchanged when using the proxy instead of the true target variable. Since most of the literature on forecast evaluation is concerned with univariate quantities Gne11,Pat11, it is worth stressing that the target variable and the forecasts may be multivariate here. Some work on characterizing strictly consistent loss functions for the specific multivariate mean functional can be found in BGW05, Laurent2013 and FK15.
Our first main contribution is to characterize the loss functions that are exactly robust to measurement error. We show that the class of exactly robust loss functions coincides with the Bregman loss functions BGW05, and the class to which Pat11 refers as “robust” loss functions in the univariate case.
This implies that only conditional mean forecasts can be compared robustly in ECPA tests.
While this is of course rather restrictive, for many economic variables that are measured with error, the conditional mean is precisely the object of interest; e.g., the conditional mean of GDP growth or inflation, or the conditional mean of squared asset returns (which commonly coincides with the conditional return variance).
The importance of the conditional mean is further underscored by the prevalence in economics of (V)ARMA-type models, which are designed to dynamically model the conditional mean.
The studies most closely related to our first contribution are those of Pat11 (in the univariate case) and Laurent2013 (in the multivariate case), both of which are devoted to the specific task of comparing conditional variance forecasts for financial returns using high-frequency (HF) proxies.
Pat11 (later on extended by Laurent2013) characterizes loss functions that give consistent relative rankings when only a conditionally unbiased proxy for the conditional variance is available. Our exact robustness is a stronger requirement than Pat11's Pat11 ordering robustness. While under exact robustness two forecasts have the same conditional predictive ability for the true target and the proxy, ordering robustness merely implies that the ordering of the two forecasts is preserved when a proxy is used instead of the true target; however, the magnitude in predictive ability may be changed. For instance, the average forecast loss differences may be large for the true target, yet very small when the proxy is used. Thus, while the ranking (in population) is preserved, it may be much harder to discriminate between the two forecasts in finite samples when only the proxy is available. In contrast, under exact robustness, the magnitude of the expected differences is identical for the true target and the proxy. Thus, our concept of exact robustness almost immediately implies that ECPA tests can be carried out as usual under measurement error, in particular allowing us to do a local power analysis. While our exact robustness is a stronger requirement than Pat11's Pat11 ordering robustness, we show in Theorem (ref) that---surprisingly---the respective classes of loss functions coincide.
By doing so, we also refine the results of Pat11 along several dimensions; see Remarks (ref) and (ref) for details.
Another important aspect of our first main contribution is that---unlike Pat11 and Laurent2013---we do not restrict attention to the mean functional from the outset. Thus, by narrowing down the class of functionals that can be evaluated robustly to the mean functional, we are able to show that (e.g.) the ranking of median forecasts is affected by noisy proxies.
In Appendix (ref), we strengthen this result by showing that median (and more generally, quantile) forecasts cannot be evaluated robustly, even when the conditional unbiasedness assumption is replaced by any other “resemblance condition” on the proxy. In other words, a robust evaluation of quantile forecasts requires the proxy and the true target to coincide, that is, it requires the absence of measurement error. However, the absolute error loss---pertaining to median forecasts---has regularly been used in comparing forecasts for mismeasured variables, such as inflation Han05,Mea21, GDP growth RW09,BK14 and integrated variances HL05. In each case, our results suggest that these comparisons should be interpreted with extreme caution, due to the non-robustness of the median.
Our second main contribution is to study the local power of ECPA tests using proxy variables and exactly robust loss functions. We demonstrate that power increases for more accurate proxies. The (infeasible) upper bound for the test power is obtained when evaluating forecasts with the most accurate proxy. In our case, this “proxy” is the---generally even ex post---latent target functional, i.e., the conditional mean of the target variable.
This supports the intuition that it is easier to discriminate between competing forecasts if the target is approximated more precisely.
Our simulations show that the asymptotic local power of the proxy-based ECPA test provides a good approximation in finite samples. We further demonstrate the dangers of using non-robust loss functions for comparing predictive accuracy with proxy variables. Specifically, size distortions may arise and, for certain alternatives, a loss of power occurs. These drawbacks, instead of getting less serious in larger samples, get more pronounced as the sample size increases.
We apply our proxy-based ECPA test to GDP growth rates in the US. GDP is a measure of the aggregate real output of the economy and, as such, is perhaps the most important macroeconomic indicator. However, GDP (and, hence, also GDP growth) cannot be measured exactly for various reasons. For instance, tax returns are incorporated into the national accounts only over time, leading to frequent revisions of GDP estimates (and thus different vintages, i.e., series of GDP releases). Also, the US Bureau of Economic Analysis relies on economic census data collected only once every five years for computing GDP. Hence, GDP estimates are inherently based on some extrapolation, leading to error. Thus, for comparing forecasts, we can only use approximations of true GDP growth.
Several proxies for true GDP growth (denoted $\Delta\text{GDP}$) are available LSF08. The arguably most popular proxy is the expenditure-side approximation $\Delta\text{GDP}_E$, followed by the income-side proxy $\Delta\text{GDP}_I$.
Our third proxy, $\Delta\text{GDP}_+$ from Aea16 combines both of these information sources and can be regarded as a more precise proxy of latent GDP growth.
Informed by our theory, we anticipate that using $\Delta\text{GDP}_+$ in ECPA tests leads to better discrimination between different forecasts. This is indeed what we find when comparing mean predictions from the Survey of Professional Forecasters (SPF) issued at different horizons. Naturally, we expect that $\Delta\text{GDP}$ forecasts for some time $t$ that were issued one quarter ago to be superior to those issued two or even four quarters ago. We confirm this and find that the evidence in favor of shorter horizon SPF forecasts is more convincing, the more precise the proxy. We obtain similar results for Greenbook forecasts, and also for more recent vintages of $\Delta\text{GDP}_E$, where later vintages typically provide more accurate proxies of true GDP growth.
The remainder of the paper proceeds as follows.
In Section (ref), we define \textit{exact/ordering robustness to measurement error} for loss functions, characterize these loss functions and show that exact and ordering robustness are equivalent.
Then, we derive the local power of ECPA tests based on proxy variables and robust loss functions.
Section (ref) numerically illustrates the local power results in Monte Carlo experiments. Section (ref) applies our test to US GDP growth forecasts. Finally, Section (ref) concludes.
Appendix (ref) establishes a non-robustness result for quantile forecasts under any “resemblance condition” on the proxy, and the Appendices (ref) and (ref) contain all proofs and some further technical derivations.
\section{Main Results}
\subsection{Characterizing Robust Loss Functions}
Let $(\Omega, \mathcal{A}, \operatorname{P})$ be a probability space.
If not stated otherwise, all (in-)equalities involving conditional expectations are tacitly assumed to hold $\operatorname{P}$-almost surely (a.s.) in the following.
Let $\mathcal{P}$ denote some class of distribution functions on $\mathbb{R}^{k}$ ($k\in\mathbb{N}$), which is specified later on.
For some time point $t\in\mathbb{N}$, we denote the random variable of interest (e.g., GDP growth) by $Y_t:\Omega\rightarrow\mathsf{O}\subset\mathbb{R}^{k}$, the time-$(t-1)$ information set by $\mathcal{F}_{t-1}$, and the conditional distribution of $Y_{t}$ given $\mathcal{F}_{t-1}$ by $F_t(\omega,\cdot)=\operatorname{P}\{Y_t\leq\cdot\mid\mathcal{F}_{t-1}\}(\omega)$, where we assume that $F_t(\omega,\cdot)\in\mathcal{P}$ for all $\omega\in\Omega$.
We denote by $F_t(\cdot)$ the random variable defined by the mapping $\omega \mapsto F_t(\omega,\cdot)$.
\begin{rem}
The theory of this article can be extended to $\tau$-step ahead forecasts for $\tau \ge 2$ in a straightforward fashion by considering the information set $\mathcal{F}_{t-\tau}$ instead of $\mathcal{F}_{t-1}$. Since we leave $\mathcal{F}_{t-1}$ unspecified in the following, this can be seen by letting $\mathcal{F}_{t-1}$ contain only information available at time $(t-\tau)$.
We deliberately do not make this explicit in the notation in order to keep the exposition as simple as possible.
\end{rem}
The aim in a forecasting situation is to predict some target functional $T:\mathcal{P}\rightarrow\mathsf{A}\subset\mathbb{R}^{k}$ of $F_t$. We denote the target functional by $x_t^\ast:\Omega\rightarrow\mathsf{A}$, $\omega\mapsto T(F_t(\omega,\cdot))$. For the leading case $k=1$, this may be the conditional mean of GDP growth, $x_t^\ast=\operatorname{E}_{t-1}[Y_t]$, where we write $\operatorname{E}_{t-1}[\,\cdot\,]=\operatorname{E}[\ \cdot\mid\mathcal{F}_{t-1}]$ for short. We assume that there exists a \textit{strictly $\mathcal{P}$-consistent} (or simply \textit{strictly consistent}) loss function $L:\mathsf{O}\times\mathsf{A}\rightarrow\mathbb{R}$ for the functional $T$, that is
\begin{equation}
\operatorname{E}_{Y \sim F}\big[L(Y,T(F))\big]\leq\operatorname{E}_{Y \sim F}\big[L(Y,x)\big],
\end{equation}
for all $F\in\mathcal{P}$ and all $x\in\mathsf{A}$, and equality in (ref) implies $x=T(F)$; see Gne11.
Here, the notation $\operatorname{E}_{Y \sim F}[\,\cdot\,]$ denotes the expectation with respect to $Y$ with distribution $F$.
A functional for which a strictly consistent loss function exists is termed \textit{elicitable}.
Assuming $T$ to be elicitable is not restrictive in our context, because when no strictly consistent scoring function exists, forecasts cannot be compared validly Gne11.
Denote by $x_{it}:\Omega\rightarrow\mathsf{A}$ ($i=1,2$) the two competing, $\mathcal{F}_{t-1}$-measurable forecasts of $x_t^\ast$. The forecast loss difference
\[
d(Y_t, x_{1t}, x_{2t}) = L(Y_t, x_{1t}) - L(Y_t, x_{2t})
\]
measures the relative performance of $x_{1t}$ and $x_{2t}$. Since the loss function is negatively oriented, a negative (positive) loss difference favors $x_{1t}$ ($x_{2t}$).
As pointed out in the Motivation, the true $Y_t$ may often not be available for comparing forecasts due to measurement error. Instead, one has to rely on a proxy $\widehat{Y}_t:\Omega\rightarrow\mathsf{O}$ for $Y_t$ and use $d(\widehat{Y}_t, x_{1t}, x_{2t})$. Of course, the proxy has to bear some resemblance to the target. To ensure this, we make the assumption that $\widehat{Y}_t$ is \textit{conditionally unbiased} for $Y_t$, i.e., that $\operatorname{E}_{t-1}[\widehat{Y}_t]\overset{\text{a.s.}}{=}\operatorname{E}_{t-1}[Y_t]$.
This also implies that (mean) differences between $Y_t$ and $\widehat{Y}_t$ cannot be predicted from information available at time $(t-1)$. We consider this to be a natural assumption for a proxy variable, since otherwise one could predict the average measurement error $\widehat{Y}_t-Y_t$ based on $\mathcal{F}_{t-1}$. We refer to the empirical application for a modeling framework of GDP growth rates, where the conditional unbiasedness assumption is satisfied for $Y_t=\Delta\text{GDP}_{t}$ and $\widehat{Y}_t\in\big\{\Delta\text{GDP}_{E,t}, \Delta\text{GDP}_{I,t}\big\}$; see in particular (ref).
We denote the conditional distribution function of $\widehat{Y}_t\mid\mathcal{F}_{t-1}$ by $\widehat{F}_{t}(\omega,\cdot)$ for all $\omega \in \Omega$, and the corresponding random variable by $\widehat{F}_{t}(\cdot)$. The following definition ensures that the expected forecast loss differences are the same for $Y_t$ and a conditionally unbiased proxy $\widehat{Y}_t$.
\begin{defn}
$L(\cdot,\cdot)$ is \textit{\underline{exactly} robust to measurement error} (or simply: \textit{exactly robust}) with respect to $\mathcal{P}$, if
\[
\operatorname{E}_{t-1}[d(Y_t, x_{1t}, x_{2t})]=\operatorname{E}_{t-1}[d(\widehat{Y}_t, x_{1t}, x_{2t})]\qquad\text{a.s.}
\]
for all $\mathcal{F}_{t-1}$-measurable forecasts $x_{1t}$ and $x_{2t}$, and all $Y_t$ and all conditionally unbiased proxies $\widehat{Y}_t$ with $F_{t}(\omega,\cdot)\in\mathcal{P}$ and $\widehat{F}_{t}(\omega,\cdot)\in\mathcal{P}$ for all $\omega\in\Omega$.
\end{defn}
Exact robustness to measurement error suggests that we can learn as much about the relative merits of $x_{1t}$ and $x_{2t}$ by observing $d(\widehat{Y}_t, x_{1t}, x_{2t})$ instead of $d(Y_t, x_{1t}, x_{2t})$. However, this is only correct in expectation, or equivalently, asymptotically under a suitable law of large numbers. In finite samples, this is unfortunately not true, since $d(\widehat{Y}_t, x_{1t}, x_{2t})$ may have a larger (or smaller) variance than $d(Y_t, x_{1t}, x_{2t})$.
The more restricted the class $\mathcal{P}$ in Definition (ref), the richer the class of exactly robust loss functions. An extreme case arises if $\mathcal{P}=\{F:\mathbb{R}^{k
}\rightarrow[0,1] \mid F(x)=I_{\{x\geq o\}}\text{ for some }o\in\mathsf{O}\}$ is the class of degenerate distributions. (The inequality $x\geq o$ is to be understood component-wisely if the quantities involved are vector-valued.) Then, by (a.s.) constancy of $Y_t$ and $\widehat{Y}_t$, we necessarily have that $Y_t \overset{\text{a.s.}}{=} \widehat{Y}_t$ by conditional unbiasedness. The latter implies that \textit{any} loss function is exactly robust with respect to this rather restricted class $\mathcal{P}$.
\begin{rem}
Consider univariate log-returns $r_t$ on some speculative asset and assume that $\operatorname{E}[r_t\mid\mathcal{F}_{t-1}]=0$, as is common for financial data. In this setting, forecasts of the (latent) conditional variance $x_t^\ast=\operatorname{E}[r_t^2\mid\mathcal{F}_{t-1}] = \operatorname{Var}(r_t\mid\mathcal{F}_{t-1})$ are essential for risk management purposes Aea13. Thus, in our setting, the target variable is $Y_t=r_t^2$ and $T$ is the mean functional. To compare two volatility forecasts $x_{1t}$ and $x_{2t}$, the natural choice is then $Y_t=r_t^2$. However, in the---different, but related---context of evaluating volatility forecasts of GARCH models, AB98 advocate the use of less noisy high-frequency proxies $\widehat{Y}_t$ of the conditional variance (such as realized volatility or the range) satisfying $\operatorname{E}_{t-1}[\widehat{Y}_t]=\operatorname{E}_{t-1}[Y_t]=x_t^\ast$. This example shows that (other than our notation for $Y_t$ and $\widehat{Y}_t$ suggests) $\widehat{Y}_t$ does not necessarily have to be regarded as coming “close” to $Y_t$. Instead, we may sometimes interpret $Y_t$ \textit{and} $\widehat{Y}_t$ as providing two different conditionally unbiased estimates of the target functional $x_t^\ast$.
\end{rem}
\begin{rem}
In the framework of Remark (ref), Pat11 investigates conditions under which loss functions produce consistent \textit{relative rankings} in the sense that
\begin{equation*}
\operatorname{E}[L(Y_t, x_{1t})]\lesseqgtr\operatorname{E}[L(Y_t, x_{2t})] \quad\Longleftrightarrow\quad \operatorname{E}[L(\widehat{Y}_t, x_{1t})]\lesseqgtr\operatorname{E}[L(\widehat{Y}_t, x_{2t})]
\end{equation*}
for any (volatility) proxy satisfying $\operatorname{E}_{t-1}[\widehat{Y}_t]=\operatorname{E}_{t-1}[Y_t](=x_t^\ast)$.
This property is conceptually different from \textit{exact} robustness to measurement error in that, first, only the ranking of forecasts is concerned and, second, unconditional means are considered. Laurent2013 generalize Pat11's Pat11 results by considering multivariate $r_t$, where $Y_t=\operatorname{vech}(r_{t}r_t^\prime)$ is the target variable and the target functional $x_t^\ast=\operatorname{E}[Y_t\mid\mathcal{F}_{t-1}]$ is the ($\operatorname{vech}$-transformed) conditional variance-covariance matrix of the returns. Here, $\operatorname{vech}(\cdot)$ stacks the lower triangular part of a matrix into a vector.
\end{rem}
Despite the conceptual differences it will be insightful to transfer the idea of loss functions that produce consistent relative rankings to our framework. To do so, we generalize the definition of Pat11 as follows:
\begin{defn}
$L(\cdot,\cdot)$ is \textit{\underline{ordering} robust to measurement error} (or simply: \textit{ordering robust}) with respect to $\mathcal{P}$, if
\begin{equation*}
\operatorname{E}_{t-1}[d(Y_t, x_{1t}, x_{2t})]\lesseqgtr0 \quad \text{a.s.}
\quad\Longleftrightarrow\quad
\operatorname{E}_{t-1}[d(\widehat{Y}_t, x_{1t}, x_{2t})]\lesseqgtr0 \quad \text{a.s.}
\end{equation*}
for all $\mathcal{F}_{t-1}$-measurable forecasts $x_{1t}$ and $x_{2t}$, and all $Y_t$ and all conditionally unbiased proxies $\widehat{Y}_t$ with $F_{t}(\omega,\cdot)\in\mathcal{P}$ and $\widehat{F}_{t}(\omega,\cdot)\in\mathcal{P}$ for all $\omega\in\Omega$.
\end{defn}
We argue that exact robustness is a more useful concept in practice than the weaker ordering robustness. To see why, assume that $\operatorname{E}_{t-1}[d(Y_t, x_{1t}, x_{2t})]=\nu_t$ for some large $\nu_t>0$, such that $x_{2t}$ is a much better forecast. Under exact robustness, we then have $\operatorname{E}_{t-1}[d(\widehat{Y}_t, x_{1t}, x_{2t})]=\nu_t$, such that $x_{2t}$ is again clearly superior when judged using the proxy $\widehat{Y}_t$. However, under ordering robustness, we may merely have $\operatorname{E}_{t-1}[d(\widehat{Y}_t, x_{1t}, x_{2t})]=\widetilde{\nu}_t$ for some small $\widetilde{\nu}_t>0$.
Hence, in finite samples, it may be much harder to identify $x_{2t}$ as the better forecast when only the $\widehat{Y}_t$ are available. In other words, two forecasts may be easier to separate when making an exactly robust comparison instead of an ordering robust comparison. Thus, when the forecasts and the proxy are taken as given in applied work, we advocate the use of exactly robust loss functions. However, this recommendation is based solely on conceptual considerations, because---foreshadowing one of the main results of Theorem (ref)---the classes of exactly robust and ordering robust loss functions coincide.
Before we can state Theorem (ref), recall the concept of a \textit{subgradient} from convex analysis. To do so, we let $\langle\cdot,\cdot\rangle$ denote the standard scalar product in $\mathbb{R}^{k}$. Then, the subgradient at value $x \in \mathsf{A}$ of some convex function $\phi:\mathsf{A}\rightarrow\mathbb{R}$, denoted $\,\mathrm{d}\phi(x)$, is any vector $s\in\mathsf{A}$ satisfying $\phi(y)\geq\phi(x)+\langle s, y-x\rangle$ for all $y\in\mathsf{A}$ HL01. Recall from HL01 that such a vector always exists.
\begin{thm}
Let $\mathcal{P}$ be a convex set of distribution functions. Assume that $T:\mathcal{P}\rightarrow\mathsf{A}$ is surjective, $\mathsf{A}$ is convex, and $L(\cdot,\cdot)$ is strictly $\mathcal{P}$-consistent for $T$. Then, the following are equivalent:
\begin{enumerate}
• $L(\cdot,\cdot)$ is of the form
\begin{equation}
L(Y, x)=\phi(x) + \langle \,\mathrm{d}\phi(x), Y-x \rangle + a(Y),
\end{equation}
where $\phi:\mathsf{A}\rightarrow\mathbb{R}$ is strictly convex with subgradient $\,\mathrm{d}\phi(\cdot)$, and $a:\mathsf{O}\rightarrow\mathbb{R}$ is integrable with respect to all $F\in\mathcal{P}$;
• $L(\cdot,\cdot)$ is exactly robust with respect to $\mathcal{P}$;
• $L(\cdot,\cdot)$ is ordering robust with respect to $\mathcal{P}$;
• For all $Y_t$ with $F_t(\omega,\cdot)\in\mathcal{P}$ for all $\omega\in\Omega$, it holds that $x_t^\ast=T(F_{t}(\cdot))\overset{\text{a.s.}}{=}\operatorname{E}_{t-1}[Y_t]$.
\end{enumerate}
\end{thm}
Theorem (ref) offers three main insights. First, it characterizes the exactly robust loss functions, while making \textit{no} assumption on the functional to be evaluated in advance. Exact robustness is essential to theoretically study ECPA tests under the alternative, where there is a difference in predictive ability, i.e., $\operatorname{E}_{t-1}[d(Y_t,x_{1t}, x_{2t})]\neq0$. In the absence of the equivalence of (b) and (c) established in Theorem (ref), mere ordering robustness would not be sufficient to do so, because the magnitude of the deviation from zero in $\operatorname{E}_{t-1}[d(\widehat{Y}_t,x_{1t}, x_{2t})]\neq0$ may be changed. Thus, a second contribution of Theorem (ref) is the equivalence of exact and ordering robustness, and we merely speak of robust loss functions in the following when no confusion can arise. The third insight is that if, in practice, there is measurement error in the target variable, then \textit{only} conditional mean forecasts can be evaluated robustly. For other functionals, such as the median, the forecast ranking is affected by the use of proxies. In particular, many commonly used loss functions, such as absolute error (AE) loss $L(Y,x)=|Y-x|$, are not robust. Nonetheless, the AE loss has been used in comparing forecasts for mismeasured variables, such as inflation Han05,Mea21, GDP growth RW09,BK14 and integrated variances HL05. Thus, Theorem (ref) casts doubt on the rankings obtained by these comparisons.
Theorem (ref) crucially depends on the conditional unbiasedness condition, $\operatorname{E}_{t-1} [ \widehat Y_t ] \overset{\text{a.s.}}{=} \operatorname{E}_{t-1} [ Y_t ]$.
This raises the question if forecasts for other target functionals than the mean can be compared robustly under alternative “resemblance conditions” on $Y_t$ and $\widehat Y_t$.
In Appendix (ref), we show that conditional quantile forecasts cannot be evaluated (exactly) robustly, unless one imposes the degenerate resemblance condition that the distributions of $Y_t$ and $\widehat Y_t$ coincide.
This reinforces our interpretation of the mean being the only target functional that allows for robust evaluation.
Thus, we have the surprising result that while the mean is more robust than the median in forecast evaluation, the opposite is well-known to hold in classical estimation theory.
\begin{rem}
The equivalence of (a) and (c) in Theorem (ref) may be viewed as a generalization of Proposition 1 in Pat11 for $k=1$ and of Proposition 2 in Laurent2013 for $k\in\mathbb{N}$. Since Laurent2013 use very similar regularity conditions as Pat11, we only highlight the main improvements on the latter.
First, while Patton's characterization builds on the \textit{unconditional} expectation, we consider \textit{conditional} expectations, which is crucial for tests of equal \textit{conditional} predictive ability GW06 and tests of superior \textit{conditional} predictive ability LLQ21+.
Second, similar to the property of strict consistency for loss functions Gne11, robustness should be considered with respect to a specified class $\mathcal{P}$ of distributions. While we merely require $\mathcal{P}$ to be convex, Pat11 restricts attention to absolutely continuous distributions.
Third, Pat11 only considers to continuously differentiable losses, which ignores important classes of loss functions, such as the generalized piecewise linear (GPL) losses. In contrast,
our Theorem (ref) dispenses with \textit{any} regularity conditions on the class of possible loss functions. Thus, our class of loss functions in (ref) is broader than his \textit{and} we do not rule out non-differentiable losses from the outset. Fourth, by considering $x_t^\ast=\operatorname{E}[r_t^2\mid\mathcal{F}_{t-1}]$ (cf. Remark (ref)), Pat11 specifies $T$ to be the mean functional as an \textit{assumption}, such that implicitly only Bregman loss functions are considered at the outset. In contrast, we do not specify $T$ in advance, allowing us to show the non-robustness of \textit{any} functional apart from the mean. We view this final improvement as the most important one due to its practical implication: when there is measurement error, one can only evaluate mean forecasts robustly. The forecast ranking of any other other functional (e.g., the median) will be affected.
\end{rem}
\begin{rem}
The classification of loss functions in Theorem (ref) can be employed (by invoking Theorem 2.5 in DFZ2020) for the objective functions in M-\textit{estimation} of semiparametric models, when only a proxy $\widehat Y_t$ of the response variable, $Y_t$, is observable.
Theorem (ref) then implies that for conditionally unbiased proxies, the M-estimator is consistent for, and only for, conditional \textit{mean} models; in contrast to, e.g., conditional \textit{quantile} models.
\end{rem}
\subsection{Testing Equal Predictive Accuracy}
Section (ref) shows that for loss functions of the form (ref), the expected loss differences are unchanged when using a conditionally unbiased proxy $\widehat{Y}_t$. Here, we consider the implications of Theorem (ref) for statistical tests of ECPA, and show that the robustness property leads to valid tests (in the sense that size is kept) whose power increases for more accurate proxies.
To that end, we outline the framework of \textit{conditional} predictive ability testing pioneered by GW06, which extends the classical predictive ability tests of DM95. Recall that interest in ECPA tests centers on the null hypothesis
\[
H_0\ :\quad \operatorname{E}_{t-1}[d(Y_t, x_{1t}, x_{2t})]\overset{\text{a.s.}}{=}0\qquad\text{for all }t=1,2,\ldots.
\]
To test this \textit{conditional} moment condition based on a finite sample of length $n$, GW06 propose to test the implication of $H_0$ that $\operatorname{E}[h_{t-1}d(Y_t, x_{1t}, x_{2t})]=0$ for a $\mathcal{F}_{t-1}$-measurable test function $h_{t-1}$, taking values in $\mathbb{R}^{q}$. They do so using the Wald-type test statistic $T_n=n\overline{Z}_n^{\prime}\widetilde{\Omega}_n^{-1}\overline{Z}_n$, where $Z_t=h_{t-1}d(Y_t, x_{1t}, x_{2t})$, $\overline{Z}_n=1/n\sum_{t=1}^{n}Z_t$, and $\widetilde{\Omega}_n$ is an invertible and consistent estimator of $\operatorname{Var}(\sqrt{n} \, \overline{Z}_n)$. Under weak regularity conditions, GW06 show that $T_n$ is asymptotically $\chi^2_{q}$-distributed under $H_0$.
However, when $Y_t$ is not observed, $T_n$ cannot be computed and we have to rely on the feasible test statistic
\[
\widehat{T}_n=n\overline{\widehat{Z}}_n^{\prime}\widehat{\Omega}_n^{-1}\overline{\widehat{Z}}_n,
\]
where $\widehat{Z}_{t}=h_{t-1}d(\widehat{Y}_t, x_{1t}, x_{2t})$, $\overline{\widehat{Z}}_n=1/n\sum_{t=1}^{n}\widehat{Z}_t$, and $\widehat{\Omega}_n$ is an invertible and consistent estimator of $\Omega_n=\operatorname{Var}(\sqrt{n}\overline{\widehat{Z}}_n)$. Arguing as before, $\widehat{T}_n$ follows a $\chi^2_{q}$-distribution asymptotically under the proxy hypothesis
\[
\widehat{H}_0\ :\ \operatorname{E}_{t-1}[d(\widehat{Y}_t, x_{1t}, x_{2t})]\overset{\text{a.s.}}{=}0\qquad\text{for all }t=1,2,\ldots.
\]
Hence, in the presence of measurement error, $\widehat{T}_n$ can validly test $H_0$ if $\operatorname{E}_{t-1}[d(Y_t, x_{1t}, x_{2t})]\overset{\text{a.s.}}{=}\operatorname{E}_{t-1}[d(\widehat{Y}_t, x_{1t}, x_{2t})]$, i.e., if $H_0$ and $\widehat{H}_0$ are equivalent, which is obviously implied by our exact robustness property.
Under exact robustness to measurement error, $H_0$ implies that $\operatorname{E}[d(Y_t,x_{1t}, x_{2t})] = 0 = \operatorname{E}[d(\widehat{Y}_t,x_{1t}, x_{2t})]$ by the law of iterated expectations. Thus, when $L(\cdot,\cdot)$ is of the form given in Theorem (ref), EPA tests of the hypothesis $\operatorname{E}[d(Y_t,x_{1t}, x_{2t})] = 0$ can also be validly carried out using a conditionally unbiased proxy $\widehat{Y}_t$ of $Y_t$.
\begin{rem}
For HF proxies in finance, LP18 establish that equality of the expected loss differences is not strictly necessary for the validity of equal predictive ability tests under measurement error.
Instead, it suffices that the expected loss differences converge at rate $o(\sqrt{n})$, which they call the “convergence-of-hypotheses” condition.
However, verification of the latter often requires non-trivial primitive conditions. E.g., when comparing volatility forecasts with high-frequency proxies, the sample size of intraday returns for computing volatility proxies must diverge faster than the number of out-of-sample volatility forecasts. In macroeconomic applications, where the sampling frequency cannot be arbitrarily increased, the “convergence-of-hypotheses” condition is not applicable.
One conclusion of our Theorem (ref) is that the convergence-of-hypotheses condition is not required (because it holds trivially) for mean forecasts when conditionally unbiased proxies are used.
\end{rem}
Next, we derive the limit of $\widehat{T}_n$ under the local alternative
\begin{align}
H_{a,\operatorname{loc}}\ :\ \operatorname{E}[h_{t-1}d(Y_t, x_{1t}, x_{2t})]=\frac{\delta}{\sqrt{n}}\quad\text{for all }t=1,2,\ldots,
\end{align}
where $\delta\in\mathbb{R}^{q}$. The magnitude of the local alternative, $\delta/\sqrt{n}$, converges to the null hypothetical value of zero as $n\to\infty$, thus making it harder for our test to reject $H_0$ for increasing $n$. Since $\operatorname{E}[Z_t]=\delta/\sqrt{n}$ under $H_{a,\operatorname{loc}}$, $Z_t$ depends on $n$. We reflect this in our notation by writing $Z_t=Z_{n,t}=(Z_{n,t}^{(1)},\ldots,Z_{n,t}^{(q)})^\prime$. Note that for $\delta=0$, $H_{a,\operatorname{loc}}$ is strictly speaking not an alternative as it reduces to the null. However, the method of proof for deriving the asymptotic limit of $\widehat{T}_n$ is the same for all $\delta\in\mathbb{R}^{q}$. Hence, we leave $\delta$ unrestricted in $H_{a,\operatorname{loc}}$. Values of $\delta$ with larger norm correspond to local alternatives with larger magnitudes. Note that under fixed alternatives, ECPA tests are consistent, i.e., power converges to one \textit{no matter} which conditionally unbiased proxy is used. Thus, only by considering \textit{local} alternatives, we are able to derive analytical results on the power of tests that use different proxies.
If the underlying loss function is robust, then it follows under $H_{a,\operatorname{loc}}$ by the law of iterated expectations (LIE) that
\begin{equation}
\frac{\delta}{\sqrt{n}}=\operatorname{E}[h_{t-1}d(Y_t, x_{1t}, x_{2t})]
= \operatorname{E}[h_{t-1}\operatorname{E}\{d(Y_t, x_{1t}, x_{2t})\mid\mathcal{F}_{t-1}\}]
= \operatorname{E}[h_{t-1}d(\widehat{Y}_t, x_{1t}, x_{2t})]
\end{equation}
for any conditionally unbiased proxy $\widehat{Y}_t$. The property in (ref), implied by exact robustness, is essential for our local power results, because it allows to test $\operatorname{E}[h_{t-1}d(Y_t, x_{1t}, x_{2t})]=\delta/\sqrt{n}$ via the $\widehat{Z}_{n,t}=h_{t-1}d(\widehat{Y}_t, x_{1t}, x_{2t})$. Thus, one key benefit of exact robustness is that it allows to compare “how difficult” it is to assess the relative merits of two forecasts, when only an imperfect proxy $\widehat{Y}_{t}$ is available. Note that even though for different proxies $\widehat{Y}_t$ the deviations under (ref) are equal, better proxies may still give less variable $\widehat{Z}_{n,t}$, such that departures from $H_0$ of the form in (ref) may be detected more easily. To investigate this, we derive the local power of the test based on $\widehat{T}_n$.
To do so, we make the following assumptions, which are similar to those of GW06. As only some proxy variable $\widehat{Y}_t$ is available, we specify these assumptions in terms of $\widehat{Z}_{n,t}=h_{t-1}d(\widehat{Y}_t, x_{1t}, x_{2t})$. Nonetheless, the specific (conditionally unbiased) choices $\widehat{Y}_t=Y_t$ and $\widehat{Y}_t=x_t^\ast$ are also allowed. Let $\widehat{W}_t=(\widehat{Y}_t^\prime,X_t^\prime)^\prime$, where the $\mathcal{F}_{t-1}$-measurable $X_t$ is $\mathbb{R}^{s}$-valued and contains all predictors that the forecasts $x_{1t}$ and $x_{2t}$ are based on. For instance, $X_t$ may contain lagged $\widehat{Y}_{t}$'s.
\begin{enumerate}
• $\big\{(\widehat{W}_t^\prime,h_t^\prime)^\prime\big\}$ is $\alpha$-mixing of size $-2r/(r-2)$ or $\phi$-mixing of size $-r/(r-1)$.
• $\operatorname{E}|\widehat{Z}_{n,t}^{(i)}|^{2r}\leq\Delta_{Z}<\infty$ for $i=1,\ldots,q$ and $r$ from D1.
• $\Omega_n=\operatorname{Var}\big({n}^{-1/2}\sum_{t=1}^{n}\widehat{Z}_{n,t}\big)\longrightarrow\Omega$, as $n\to\infty$, where $\Omega$ is positive definite.
• The forecasts $x_{1t}=f_{n,t}^{(1)}(X_t,\ldots,X_{t-m_1+1})$ and $x_{2t}=f_{n,t}^{(2)}(X_t,\ldots,X_{t-m_2+1})$ are measurable functions of a finite number of predictors, where $m=\max(m_1,m_2)\leq\overline{m}<\infty$.
\end{enumerate}
Our conditions D1--D4 are standard in the literature, and are essentially those of GW06. Assumption D4, that forecasts are based on a finite number of lags of the predictor variables, is a technical convenience used to expedite the proof of Theorem (ref). Nonetheless, this assumption accommodates many different estimation schemes, where $m_1$ and $m_2$ may be deterministic or data-driven (and, hence, stochastic). We refer to GW06 for more detail. Note that in D4 we suppress the dependence of the forecasts on $n$ for notational brevity.
\begin{thm}
Let Assumptions \textup{D1}--\textup{D4} hold. Then, if $\widehat{\Omega}_n-\Omega_n=o_{\operatorname{P}}(1)$, it holds for any exactly robust loss function under $H_{a,\operatorname{loc}}$ that
\[
\widehat{T}_n\overset{d}{\longrightarrow}\chi_{q}^2(\delta^\prime\Omega^{-1}\delta),\qquad\text{as }n\to\infty,
\]
where $\chi_{q}^2(c)$ denotes a $\chi_q^2$-distribution with non-centrality parameter $c\in\mathbb{R}$.
\end{thm}
In the special case when $Y_t=\widehat{Y}_t$, Theorem (ref) is a refinement of the consistency under fixed alternatives established by GW06. For $\delta=0$, Theorem (ref) shows that $\widehat{T}_n$ has a standard $\chi^2$-limit under $H_0$. Moreover, it implies that the asymptotic local power (ALP), i.e., the asymptotic probability of rejecting $H_0$ under $H_{a,\operatorname{loc}}$, is
\begin{equation}
\operatorname{P}\big\{\chi_q^{2}(\delta^\prime\Omega^{-1}\delta)>\chi_{q,1-\tau}^{2}(0)\big\},
\end{equation}
where $\chi_{q,1-\tau}^{2}(0)$ is the $(1-\tau)$-quantile of the $\chi_{q}^{2}(0)$-distribution and $\tau \in (0,1)$ is the significance level of the test.
This shows that the ALP depends on the magnitude of the local alternative (via $\delta$) \textit{and} on the proxy via the limit $\Omega$ of $\Omega_n=\operatorname{Var}\big({n}^{-1/2} \sum_{t=1}^{n}\widehat{Z}_{n,t}\big)$.
Thus, more precise proxies that give less variable $\widehat{Z}_{n,t}$ lead to tests with higher local power. We shed more light on this in the next subsection.
To make the test based on $\widehat{T}_n$ operational, we need a consistent estimator $\widehat{\Omega}_n$ of $\Omega_n$.
To that end, we consider a heteroskedasticity and autocorrelation consistent (HAC) estimator Newey/West:87a
\begin{equation}
\widehat{\Omega}_n=\frac{1}{n}\sum_{t=1}^{n}\widehat{Z}_{n,t}\widehat{Z}_{n,t}^\prime + \frac{1}{n}\sum_{h=1}^{m_n}w_{n,h}\sum_{t=h+1}^{n}\Big(\widehat{Z}_{n,t}\widehat{Z}_{n,t-h}^\prime + \widehat{Z}_{n,t-h}\widehat{Z}_{n,t}^\prime\Big),
\end{equation}
where $m_n$ is a sequence of integers, and $w_{n,h}$ is a scalar triangular array of weights. We restrict $m_n$ and $w_{n,h}$ as follows:
\begin{itemize}
• The sequence of integers $m_n$ satisfies $m_n\rightarrow\infty$ and $m_n=o(n^{1/4})$, as $n\to\infty$.
• It holds that $|w_{n,h}|\leq\Delta_{w}<\infty$ for all $n\in\mathbb{N}$ and $h\in\{1,\ldots,m_n\}$, and $w_{n,h}\rightarrow1$, as $n\to\infty$, for all $h=1,\ldots,m_n$.
\end{itemize}
\begin{prop}
Let Assumptions \textup{D1}--\textup{D6} hold. Then, it holds under $H_{a,\operatorname{loc}}$ that $\widehat{\Omega}_n-\Omega_n=o_{\operatorname{P}}(1)$, as $n\to\infty$.
\end{prop}
The estimator $\widehat{\Omega}_n$ is the omnibus choice. It works for EPA tests (where $H_0$ reduces to $\operatorname{E}[d(Y_t,x_{1t}, x_{2t})] = 0$ by the LIE) but also for ECPA tests. For the latter tests, simpler estimators may be used, because---under $H_0$---the sequences $\{h_{t-1}d(\widehat{Y}_t,x_{1t}, x_{2t}),\mathcal{F}_t\}$ are martingale differences and thus, uncorrelated. This implies that $\Omega_n$ simplifies to $\Omega_n=1/n \sum_{t=1}^{n}\operatorname{E}[\widehat{Z}_{n,t}\widehat{Z}_{n,t}^\prime]$, rendering $\widehat{\Omega}_n= 1/n \sum_{t=1}^{n}\widehat{Z}_{n,t}\widehat{Z}_{n,t}^\prime$ the estimator of choice. However, even for EPA tests of one-step-ahead forecasts considered here, DM95 recommend to use $m_n=0$ in (ref) (with an empty sum defined to be zero). Thus, using $\widehat{\Omega}_n=1/n\sum_{t=1}^{n}\widehat{Z}_{n,t}\widehat{Z}_{n,t}^\prime$ for both EPA and ECPA tests seems reasonable, and we opt for this choice in the numerical experiments in Section (ref).
\subsection{Finding Optimal Proxies}
Under the local alternative $H_{a,\operatorname{loc}}$, Theorem (ref) shows that the test's asymptotic local power is maximized by choosing a proxy $\widehat Y_t$ that minimizes $\Omega$.
We discuss such choices in the following by utilizing the linearity of the Bregman loss functions in $Y_t$. Specifically, we have
\begin{prop}
Suppose that the assumptions of Theorem (ref) hold and that any of (a)--(d) are in force. Then, it holds that
\begin{align}
\Omega_n
&= \frac{1}{n} \sum_{t=1}^n \operatorname{Var} \Big( h_{t-1} d(x_t^\ast, x_{1t}, x_{2t}) \Big)
+ \frac{1}{n} \sum_{t=1}^n \operatorname{E} \Big[ h_{t-1} b_{t-1}^\prime \operatorname{Var}_{t-1} \big( \widehat Y_t \big) b_{t-1} h_{t-1}^\prime\Big] \notag\\
&\qquad+ \frac{2}{n} \sum_{s < t} \operatorname{Cov} \Big( h_{s-1} d(\widehat Y_s, x_{1s}, x_{2s}) ,\; h_{t-1} d(x_t^\ast, x_{1t}, x_{2t}) \Big),
\end{align}
where $b_{t-1}=\,\mathrm{d}\phi(x_{1t})-\,\mathrm{d}\phi(x_{2t})$. Furthermore, if the covariance terms vanish asymptotically,
\begin{align}
\Omega_n
= \frac{1}{n} \sum_{t=1}^n \operatorname{Var} \Big( h_{t-1} d(x_t^\ast, x_{1t}, x_{2t}) \Big)
+ \frac{1}{n} \sum_{t=1}^n \operatorname{E} \Big[ h_{t-1} b_{t-1}^\prime \operatorname{Var}_{t-1} \big( \widehat Y_t \big)b_{t-1} h_{t-1}^\prime \Big]+o(1),
\end{align}
with an asymptotic lower bound of $\Omega^\ast=\lim_{n\to\infty}1/n \sum_{t=1}^n \operatorname{Var} \left( h_{t-1} d(x_t^\ast, x_{1t}, x_{2t}) \right)$ (in the sense that $\Omega-\Omega^{\ast}$ is positive semi-definite), which is attained if and only if $\widehat Y_t = x_t^\ast$ a.s.
\end{prop}
The covariance terms in (ref) vanish (and, hence, (ref) holds) if the $\big\{ h_{t-1} d(\widehat{Y}_t, x_{1t}, x_{2t}) \big\}$ are serially uncorrelated, which holds under $H_0$. But (ref) may also hold under $H_{a,\operatorname{loc}}$ as the derivation of $\Omega$ in Proposition (ref) of Appendix (ref) shows. If (ref) holds, the generally infeasible choice of $\widehat Y_t = x_t^\ast$ gives an upper bound for the local test power as then, $\operatorname{Var}_{t-1}\big( \widehat Y_t\big) \overset{\text{a.s.}}{=} 0$, such that the second (positive semi-definite) term on the right-hand side of (ref) vanishes. This is very intuitive: It is easiest to assess the relative merits of two forecasts $x_{1t}$ and $x_{2t}$ if they can be compared against the target $x_t^{\ast}$ itself. This supports the intuition that it is easier to distinguish between two forecasts if the target functional is approximated more precisely. We stress once again that this result relies on the newly introduced exact robustness concept. While comparing forecasts with $x_t^{\ast}$ itself is best from a theoretical point of view, we are not aware of a practical situation where $x_t^\ast$ is observable, even ex post at time $t$.
Since the test power decreases with increasing (in terms of the Loewner order) average variances, $\operatorname{Var}_{t-1} \big( \widehat Y_t \big)$, the proxy $ \widehat Y_t$ with smallest possible \textit{conditional} variance should be employed in practice.
However, these \textit{conditional} variances are generally unknown in practice.
Then, one often has to resort to domain-specific knowledge for choosing the best proxy.
We now discuss this exemplarily in the univariate case ($k=1$) for the volatility forecasting example of Remark (ref), and for our macroeconomic application in Section (ref).
In the macroeconomic application, $Y_t$ denotes true GDP growth.
As discussed in the Motivation, true GDP growth cannot be observed, even ex post.
However, several different proxies $\widehat Y_t$ are available.
Suppose, as is plausible, that there is some additive measurement error, such that $\widehat Y_t = Y_t + \widehat{\varepsilon}_t$; see, e.g., FRW05 for such a modeling approach. Here, the estimation error $\widehat{\varepsilon}_t$ is independent of $\mathcal{F}_{t-1}$, i.e., $\widehat{\varepsilon}_t$ cannot be predicted from information available at time $t-1$. Then, $\operatorname{Var}_{t-1} (\widehat Y_t) = \operatorname{Var}_{t-1}(Y_t) + \operatorname{Var}(\widehat{\varepsilon}_t)$.
Hence, choosing the most accurately estimated proxy $\widehat Y_t$ (i.e., one with smallest possible $\operatorname{Var}(\widehat{\varepsilon}_t)$) gives ECPA tests higher (local) power.
Observing a proxy $\widehat Y_t$ closer to $x_t^\ast$ than $Y_t$---in the sense that $\operatorname{Var}_{t-1}(\widehat Y_t)$ is closer to $\operatorname{Var}_{t-1}(x_t^{\ast})(=0)$ than $\operatorname{Var}_{t-1}(Y_t)$---seems delusive in this application.
While in the previous example, true GDP growth as the target variable is the best proxy, one can sometimes “get closer” to the ideal $x_t^{\ast}$.
To illustrate this, consider the classical situation of forecasting the conditional variance $x_t^\ast = \operatorname{Var}_{t-1}(r_t)$ of the univariate log-return $r_t$ on a risky asset with $\operatorname{E}[r_t\mid\mathcal{F}_{t-1}]=0$. In this case, $x_t^\ast=\operatorname{E}[r_t^2\mid\mathcal{F}_{t-1}]$, such that $Y_t = r_t^2$ is the natural target against which to compare forecasts. However, for a standard diffusion process, as in ABDL03, one can show that $x_t^{\ast}=\operatorname{E}_{t-1}[IV_t]$ with $IV_t$ the latent integrated variance.
Let $\widehat Y_t = RV_t$ be the realized variance, which estimates $IV_t$ from HF returns.
Then, under ABDL03 and given that $RV_t$ is an unbiased estimator of $IV_t$, it holds that $\operatorname{E}_{t-1} [Y_t] = \operatorname{E}_{t-1} [\widehat Y_t] = x_t^\ast$.
In this setting, it is already well documented that $\widehat Y_t=RV_t$ exhibits much smaller conditional variance than the squared return $Y_t=r_t^2$. Hence, $RV_t$ should be favored for testing ECPA of volatility forecasts as shown by our Theorem (ref) and Proposition (ref).
Indeed, this is common practice in the literature when comparing volatility forecasts BPT01,HL05,KJH05,Laurent2013. However, our Theorem (ref) and Proposition (ref) are the first theoretical results supporting this practice.
(We mention that AB98 provide early arguments for the use of HF proxies in the \textit{absolute} evaluation of volatility forecasts via Mincer--Zarnowitz regressions, whereas our focus is on forecast \textit{comparison}.)
Since applications in finance comparing RV to the squared return are already available in the literature, we focus on a macroeconomic application in Section (ref).
Summing up our results of Section (ref), we find that for conditional mean forecasts, ECPA tests based on $\widehat{T}_n$ are asymptotically valid under $H_0$ as long as the proxy $\widehat{Y}_t$ is conditionally unbiased. However, Theorem (ref) and Proposition (ref) show that ECPA tests lose local power when $\operatorname{Var}_{t-1}(\widehat Y_t)>\operatorname{Var}_{t-1}(Y_t)$, (as is the case in the macroeconomic example) or gain local power when $\operatorname{Var}_{t-1}(\widehat Y_t)<\operatorname{Var}_{t-1}(Y_t)$ (as in the finance application).
In either case, we urge applied researchers to compare forecasts based on the most accurate proxy available. The next section illustrates the power loss from using imprecise proxies in simulations.
\section{Numerical Experiments}
\subsection{Simulations for Robust Loss Functions}
We let $k=1$ and consider the following data-generating process (DGP) and conditional mean forecasts:
\begin{align}
Y_t &= \mu(1-\phi)+\phi Y_{t-1}+\varepsilon_t,&&\varepsilon_t\overset{\text{i.i.d.}}{\sim}N(0,\sigma_{\varepsilon}^2),\\
\widehat{Y}_t &= Y_{t}+\widehat{\varepsilon}_{t},&&\widehat{\varepsilon}_{t}\overset{\text{i.i.d.}}{\sim}N(0,\sigma_{\widehat{\varepsilon}}^2),\\
x_{1t} &= \mu(1-\phi)+\phi Y_{t-1}+\varepsilon_{1,t-1},&&\varepsilon_{1t}\overset{\text{i.i.d.}}{\sim}N(0,\sigma_{1}^2),\\
x_{2t} &= \phi Y_{t-1},&&
\end{align}
were $\{\varepsilon_t\}$, $\{\widehat{\varepsilon}_t\}$ and $\{\varepsilon_{1t}\}$ are mutually independent of each other, $|\phi|<1$, $\mu\in\mathbb{R}$, and $\sigma_{1}^2=\mu^2(1-\phi)^2+\xi/\sqrt{n}$ for $\xi\in\mathbb{R}$. Note that the forecasts are measurable with respect to $\mathcal{F}_{t-1}=\sigma(\widehat{Y}_{t-1},\widehat{Y}_{t-2},\ldots;\varepsilon_{1,t-1},\varepsilon_{1,t-2},\ldots;\widehat{\varepsilon}_{t-1},\widehat{\varepsilon}_{t-2},\ldots)$. Of course, it is a convenience (used to render some subsequent computations more tractable) to assume that the latent target $Y_{t-1}$ is known to the forecasters issuing $x_{1t}$ and $x_{2t}$. The main point of these simulations is to investigate the power loss in the comparison from using the proxy $\widehat{Y}_t$. Note that with our choice of $\mathcal{F}_{t-1}$, we have $\operatorname{E}_{t-1}[\widehat{Y}_{t}]=\operatorname{E}_{t-1}[Y_t]$.
We first consider the squared error (SE) as a robust loss function, i.e., $L(Y,x)=(Y-x)^2$, which arises for $\phi(x)=x^2$ and $a(Y)=-Y^2$ in (ref). As the SE loss elicits the mean, we have $x_t^\ast=\operatorname{E}_{t-1}[Y_t]=\mu(1-\phi)+\phi Y_{t-1}$. Hence, $x_{1t}$ equals the optimal forecast confounded by some additive noise with variance $\sigma_1^2$. This implies that the larger $\sigma_1^2$, the worse the forecast $x_{1t}$. On the other hand, $x_{2t}$ is a biased forecast, yet with no additive noise. Note that when $\sigma_1^2<\mu^2(1-\phi)^2$ ($\sigma_1^2>\mu^2(1-\phi)^2$), $x_{1t}$ ($x_{2t}$) is the preferred forecast; see (ref).
Following GW06, we choose the $\mathcal{F}_{t-1}$-predictable test function $h_{t-1}=(1, \widehat{Y}_{t-1})^\prime$. To assess the magnitude of the local alternative, Appendix (ref) shows that
\begin{equation}
\operatorname{E}[h_{t-1}d(Y_t, x_{1t}, x_{2t})] = \begin{pmatrix}\operatorname{E}[d(Y_t, x_{1t}, x_{2t})]\\ \operatorname{E}[\widehat{Y}_{t-1}d(Y_t, x_{1t}, x_{2t})]\end{pmatrix} =\begin{pmatrix}\sigma_1^2 - \mu^2(1-\phi)^2\\ \operatorname{E}[\widehat{Y}_{t-1}]\operatorname{E}[d(Y_t, x_{1t}, x_{2t})]\end{pmatrix} = \begin{pmatrix}\xi/\sqrt{n}\\ \mu\xi/\sqrt{n}\end{pmatrix}.
\end{equation}
Thus, we simulate under the local alternative
\[
H_{a,\operatorname{loc}}\ :\ \operatorname{E}[h_{t-1}d(Y_t,x_{1t}, x_{2t})]=\delta/\sqrt{n},
\]
where $\delta=(\xi, \mu\xi)^\prime$. For $\xi=0$, our results correspond to size. To compute the theoretical ALP in (ref), we calculate $\Omega$ from D3 in Proposition (ref) of Appendix (ref).
We run our simulations on a grid of $\xi$ values in the interval $[-4,4]$, and compare the empirical rejection frequencies based on $10,000$ simulation replications with the theoretical ALP at a significance level of $5\%$.
We do so for $\mu=1$, $\phi=0.2$, and $\sigma_{\varepsilon}^2=1$. To investigate the impact of noise in the proxy, we consider several values of $\sigma_{\widehat{\varepsilon}}^2$. Specifically, we choose $\sigma_{\widehat{\varepsilon}}^2$ such that $\zeta=\operatorname{Var}(Y_t)/\sigma_{\widehat{\varepsilon}}^2\in\{1/5,\, 1/2,\, 1,\, 2,\, 5,\, \infty\}$ with $\zeta=\infty$ corresponding to the case of no noise, i.e., $\sigma_{\widehat{\varepsilon}}^2=0$.
The ratio $\zeta$ can be interpreted as the \textit{signal to noise ratio} (SNR) of the noisy target; for $\zeta=\infty$, we accurately observe the target, while for $\zeta=1$, the signal $Y_t$ is of the same magnitude as the noise $\widehat{\varepsilon}_t$.
\begin{figure}
\caption{Rejection frequencies for ECPA tests using SE loss and different $\xi$.}
\caption*{ \textit{Notes:} Empirical rejection frequencies of $\widehat{T}_n$-based ECPA test for sample sizes $n\in\{50,\, 100,\, 500\}$. Results based on DGP and forecasts from (ref)--(ref) with $\mu=1$, $\phi=0.2$, $\sigma_{\varepsilon}^2=1$, $\sigma_1^2=\mu^2(1-\phi)^2+\xi/\sqrt{n}$, and $\sigma_{\widehat{\varepsilon}}^2$ chosen to yield SNR of $\zeta\in\{1/5,\, 1/2,\, 1,\, 2,\, 5,\, \infty\}$. Forecasts are evaluated using SE loss.}
\end{figure}
The six panels in Figure (ref) correspond to the six different values of $\zeta$. In each panel, the empirical rejection frequencies are displayed as a function of $\xi$ for sample sizes $n\in\{50,\ 100,\ 500\}$. The theoretical ALP is displayed as the black reference line. Note that by (ref) the ALP curve is symmetric around $\delta=0$ and, hence, also around $\xi=0$. We draw two conclusions from Figure (ref).
\begin{enumerate}
• As predicted by Theorem (ref), the empirical rejection frequencies converge to the ALP curve as $n\to\infty$. In particular, no matter how noisy the proxy, size is approximately accurate. As could be expected from a local power result, the rejection frequencies do not increase monotonically with the sample size as is typically the case for fixed alternatives. Here, while for larger $n$ rejections for $\xi>0$ are more frequent in Figure (ref), for smaller $n$ rejections for $\xi<0$ occur more often.
• As also expected from Theorem (ref), more precise proxies (i.e., with higher $\zeta$) lead to easier discrimination between forecasts in the sense of higher power; see in particular the upper two panels. On the other hand, when $\sigma_{\widehat{\varepsilon}}^2\to\infty$, the signal in $Y_t$ is eventually buried by the noise in $\widehat{\varepsilon}_t$, giving only close to trivial power in the lower two panels.
\end{enumerate}
To summarize, our simulations show that for noisy target variables, ECPA tests are \textit{valid} when using \textit{robust} loss functions in the sense of unaffected rejection rates under the null hypothesis. However, even for robust loss functions, noisy targets negatively influence the test power, which suggests that applied researchers should utilize the most accurate proxy in forecast comparisons.
\subsection{Simulations for Non-Robust Loss Functions}
Here, we illustrate the behavior of the ECPA test based on $\widehat{T}_n$ for a non-robust loss function. We again consider the DGP and the two forecasts from (ref)--(ref). However, we now use the AE as our loss function $L(\cdot, \cdot)$, which is well-known to elicit the median Gne11. Since the mean and the median coincide for the symmetric conditional distribution of our employed DGP, the optimal forecast is again $x_t^\ast=\mu(1-\phi)+\phi Y_{t-1}$. We deliberately choose a simulation setting where the mean coincides with the median and, hence, both the SE and AE loss elicit the mean. This allows us to specifically focus on \textit{robustness} of the losses, detached from their elicited functional.
For the AE loss differences we obtain from straightforward calculations using properties of the folded normal distribution that
\begin{align}
\operatorname{E}[d(Y_t, x_{1t}, x_{2t})] &= (\sigma_{\varepsilon}^2+\sigma_{1}^2) \sqrt{\frac{2}{\pi}} - \sigma_{\varepsilon}^2\sqrt{\frac{2}{\pi}}\exp\left\{-\frac{\mu^2(1-\phi)^2}{2\sigma_{\varepsilon}^2}\right\}\notag\\
&-\mu(1-\phi)\left[1-2\Phi\left(-\frac{\mu(1-\phi)}{\sigma_{\varepsilon}}\right)\right],
\end{align}
where $\Phi(\cdot)$ denotes the standard normal distribution function. Except for $\sigma_1^2$ and $\sigma_{\widehat{\varepsilon}}^2$, we choose the parameters as in the previous subsection, i.e., $\mu=1$, $\phi=0.2$, and $\sigma_{\varepsilon}^2=1$. We vary $\sigma_{\widehat{\varepsilon}}^2$ to yield SNR parameters $\zeta\in\{2,\infty\}$.
For $\sigma_1^2=0.70...$, we obtain $\operatorname{E}[d(Y_t, x_{1t}, x_{2t})]=0$ by (ref). Thus, for $\sigma_1^2=0.70...+\xi/\sqrt{n}$ we get that
\begin{align*}
\operatorname{E}[h_{t-1}d(Y_t, x_{1t}, x_{2t})]&=\begin{pmatrix}\operatorname{E}[d(Y_t, x_{1t}, x_{2t})]\\
\operatorname{E}[\widehat{Y}_{t-1}d(Y_t, x_{1t}, x_{2t})]\end{pmatrix}= \begin{pmatrix}\sqrt{2/\pi}\xi/\sqrt{n}\\
\sqrt{2/\pi}\mu\xi/\sqrt{n}\end{pmatrix}=:\delta/\sqrt{n}.
\end{align*}
This again amounts to simulating under $H_{a,\operatorname{loc}}$, with $\xi=0$ corresponding to the null. We vary $\xi$ in the interval $[-4, 4]$.
\begin{figure}[t!]
\caption{Rejection frequencies for ECPA tests using AE loss and different $\xi$.}
\caption*{ \textit{Notes:} Empirical rejection frequencies of $\widehat{T}_n$-based ECPA test for sample sizes $n\in\{5000,\, 10000,\, 20000\}$. Results based on DGP and forecasts from (ref)--(ref) with $\mu=1$, $\phi=0.2$, $\sigma_{\varepsilon}^2=1$, $\sigma_1^2=0.70...+\xi/\sqrt{n}$, and $\sigma_{\widehat{\varepsilon}}^2$ chosen to yield SNR of $\zeta\in\{2,\infty\}$. Forecasts are evaluated using AE loss.}
\end{figure}
Due to the non-robustness of the AE, we obtain different values for the expectations when $\widehat{Y}_t$ is used. E.g., for $\sigma_1^2=0.70...$ we obtain for $\zeta=2$ that
\begin{equation}
\operatorname{E}[d(\widehat{Y}_t, x_{1t}, x_{2t})]=0.0050... \neq 0=\operatorname{E}[d(Y_t, x_{1t}, x_{2t})].
\end{equation}
Since the difference is rather small, we carry out the simulations for large sample sizes $n\in\{5000,\ 10000,\ 20000\}$ to highlight the implications of (ref) under the null.
If the AE loss were robust, a test based on $\widehat{T}_n$ would have ALP given in (ref). Thus, the empirical local power curve would closely resemble the U-shape of the ALP curves in Figure (ref), with a minimum in $\xi=0$ at the nominal level. As reference lines, the left panel of Figure (ref) shows the test results based on $T_n$, which uses the true value $Y_t$ instead of the proxy $\widehat{Y}_t$, i.e., $\zeta=\infty$. The results based on $T_n$ have the characteristic U-shape, as could be expected from Theorem (ref) applied for $Y_t=\widehat{Y}_t$. However, as the AE loss is non-robust, qualitative differences emerge when using $\widehat{T}_n$, as shown in the right panel of Figure (ref) ($\zeta=2$). First, the tests are oversized under the null ($\xi=0$), which only gets worse for increasing $n$. Thus, even in this setting where the AE is a judicious choice for evaluating median forecasts, employing the noisy observations $\widehat{Y}_t$ impairs the empirical test size, which contrasts with the SE loss results in Figure (ref). Second, and perhaps more seriously, the non-robust loss function leads to decreased power for negative $\xi$, i.e., for $\sigma_1^2<0.7$. We again have the undesirable result that this power loss \textit{increases} in $n$.
\section{Comparing Forecasts for US GDP Growth}
As a measure of total real activity, GDP is arguably the most important macroeconomic indicator. However, as pointed out in the Motivation, GDP---and, by extension, GDP growth---can only be measured with error.
In this application, we focus on continuously-compounded growth rates for US GDP, denoted by $Y_t=\Delta\text{GDP}_t$.
Specifically, we compare SPF and Greenbook forecasts issued at different horizons, yet for the same quarter $t$.
We evaluate the forecasts using six different GDP growth proxies and we investigate, if---as indicated by Theorem (ref) and Proposition (ref)---it is indeed easier to discriminate between two forecasts when a more precise proxy is used.
To describe the growth proxies, recall that there are three ways to compute GDP: the production, income, and expenditures approach; see LSF08.
All these methods may be regarded as providing estimates of the true latent value of GDP.
We use proxies based on the income or expenditures approach, or a combination of the two.
Our first four $\Delta\text{GDP}$ proxies are based on expenditure-side estimates, obtained from \url{https://tinyurl.com/43fk9ms2}.
For this approach, the Federal Reserve Bank of Philadelphia provides data for all available \textit{vintages}, i.e., in each quarter, they report updated $\Delta\text{GDP}$ estimates for all previous quarters.
As information on past GDP accumulates over time, it is reasonable to assume that more recent vintages estimate true growth $\Delta\text{GDP}$ more accurately.
Here, we use the first, second, third, and most recent vintage as our $\widehat{Y}_t$'s, and denote them by $\Delta\text{GDP}_{E1}$, $\Delta\text{GDP}_{E2}$, $\Delta\text{GDP}_{E3}$ and $\Delta\text{GDP}_{E}$, respectively.
The most recent vintage refers to the latest available data as of September 30, 2020.
Our fifth growth proxy $\Delta\text{GDP}_I$ is based on the most recent vintage of the income approach, and is obtained from \url{https://tinyurl.com/4tjvcb8w}.
While proxies based on the income method feature less prominently in economics, Nal10 nonetheless finds $\Delta\text{GDP}_I$ to better reflect the growth in real economic activity during the business cycle.
As Aea16 note in their conclusion that early vintages of $\Delta\text{GDP}_I$ often provide less accurate information than their counterparts from the expenditure side (mainly due to the delayed availability of accurate tax returns), we only use the most recent vintage $\Delta\text{GDP}_I$.
As our sixth proxy, we use the “GDP Plus‘‘ approach of Aea16, which combines estimates from the income and expenditure side, and which is available at \url{https://tinyurl.com/4tjvcb8w}.
The authors argue that $\Delta\text{GDP}_I$ and $\Delta\text{GDP}_E$ provide complementary information on GDP growth.
The recent popularity of $\Delta\text{GDP}_+$ is documented by the Federal Reserve Bank of Philadelphia reporting it alongside the more classical measures $\Delta\text{GDP}_E$ and $\Delta\text{GDP}_I$.
The methodology behind $\Delta\text{GDP}_+$ has also spurred the development of new GDP proxies Almuzara2021,Jacobs2020.
In more detail, Aea16 propose a dynamic factor model, where $\Delta\text{GDP}_E$ and $\Delta\text{GDP}_I$ both load on the single (latent) factor $\Delta\text{GDP}$:
\begin{equation}
\begin{pmatrix}
\Delta\text{GDP}_{E,t}\\ \Delta\text{GDP}_{I,t}
\end{pmatrix}=\begin{pmatrix}
1\\ 1
\end{pmatrix}\Delta\text{GDP}_{t}+\begin{pmatrix}
\varepsilon_{E,t}\\ \varepsilon_{I,t}
\end{pmatrix},
\end{equation}
where $(\varepsilon_{E,t}, \varepsilon_{I,t})^\prime$ are assumed to be i.i.d. with mean zero, implying in particular conditional unbiasedness of the growth proxies, i.e., $\operatorname{E}_{t-1}[\Delta\text{GDP}_{E,t}]=\operatorname{E}_{t-1}[\Delta\text{GDP}_{I,t}]=\operatorname{E}_{t-1}[\Delta\text{GDP}_{t}]$.
The authors extract GDP growth estimates, denoted $\Delta\text{GDP}_+$, by using the Kalman smoother.
Aea16 argue that the Kalman filter extractions $\Delta\text{GDP}_+$ can be regarded as more precise approximations of true growth than either $\Delta\text{GDP}_E$ or $\Delta\text{GDP}_I$.
As for $\Delta\text{GDP}_I$, we only use the most recent vintage of $\Delta\text{GDP}_+$.
\begin{figure}[tb]
\begin{center}
\caption{{GDP Measurements}}
\begin{subfigure}{\linewidth}
\caption{GDP Vintages based on Expenditure Approach}
\end{subfigure}
\begin{subfigure}{\linewidth}
\caption{GDP Measurement Approaches}
\end{subfigure}
\end{center}
\end{figure}
We consider the six GDP proxies from 1985Q1 until 2019Q3 to avoid structural breaks in our sample due to the Great Moderation and the Corona crisis.
One may wonder whether the differences in our proxies are large enough during this period to suspect substantially altered results of ECPA tests. To shed light on this, Figure (ref) displays the six GDP proxies over time. The upper panel displays the four vintages of the expenditure approach. The lower panel shows $\Delta\text{GDP}_{I}$, $\Delta\text{GDP}_{E}$ and $\Delta\text{GDP}_+$. In both panels, we see a clear joint behavior while the exact measurements differ---sometimes even substantially. For instance, in 2008Q1 in panel (a) the first three vintages suggest a growing economy, whereas the most recent vintage indicates an almost 2.5% decline. Hence, the results of ECPA tests may plausibly depend on which proxy is used in the comparison.
Valid ECPA tests using Theorem (ref) rely on \textit{conditionally unbiased} proxies satisfying
\begin{equation}
\operatorname{E}_{t-1}[\Delta\text{GDP}_t - \Delta\text{GDP}_{M,t}]\overset{\text{a.s.}}{=}0\qquad\text{for all}\quad M\in\mathcal{M}:=\{E1,\, E2,\, E3,\, E,\, I,\, +\}.
\end{equation}
As true $\Delta\text{GDP}$ is latent, this cannot be tested directly.
Instead, we test the implication of (ref) that the proxy differences have zero conditional mean
\begin{equation}
\operatorname{E}_{t-1}[\Delta\text{GDP}_{M,t}-\Delta\text{GDP}_{N,t}]\overset{\text{a.s.}}{=}0\qquad\text{for all}\quad M\neq N,\ M,N\in\mathcal{M}.
\end{equation}
Of course, it may be the case that all proxies $\Delta\text{GDP}_{M}$ ($M\in\mathcal{M}$), while satisfying (ref), are biased in the same direction, thus invalidating (ref). However, as our proxies are based on inherently different approaches to GDP measurement, a common bias seems unlikely. Thus, we view a passed test of (ref) as a strong indication for (ref) to hold as well.
We test (ref) by using multiple sets of instruments in a standard conditional moment test.
As the sequence $\{\Delta\text{GDP}_{M,t}-\Delta\text{GDP}_{N,t}\}$ is a martingale difference under the null hypothesis in (ref), we follow GW06 and base our test on the sample variance estimator (instead of on a HAC estimator).
We use the following five instrument choices: (1) a constant; (2) a constant and the lagged first GDP proxy; (3) a constant and the lagged second GDP proxy; (4) a constant plus the difference of the lagged GDP proxies; and (5) a constant, the first lagged GDP proxy, and the difference of the lagged GDP proxies.
\begin{table}[tb]
\caption{$p$-Values of Tests for Conditional Unbiasedness of GDP Proxies}
\begin{tabularx}{0.76\linewidth}{XX @ l rrrrr}
\toprule
& & & \multicolumn{5}{c}{Instruments} \\
\cmidrule(lr){3-8}
$GDP_M$ & $GDP_N$ & & Inst. 1 & Inst. 2 & Inst. 3 & Inst. 4 & Inst. 5 \\
\midrule
$\Delta\text{GDP}_{E1}$ & $\Delta\text{GDP}_{E2}$ & & 0.363 & 0.516 & 0.527 & 0.825 & 0.706\\
$\Delta\text{GDP}_{E1}$ & $\Delta\text{GDP}_{E3}$ & & 0.227 & 0.379 & 0.269 & 1.000 & 0.590\\
$\Delta\text{GDP}_{E1}$ & $\Delta\text{GDP}_{E}$ & & 0.106 & 0.089 & 0.276 & 0.551 & 0.635\\
$\Delta\text{GDP}_{E1}$ & $\Delta\text{GDP}_{I}$ & & 0.171 & 0.061 & 0.106 & 0.978 & 0.524\\
$\Delta\text{GDP}_{E1}$ & $\Delta\text{GDP}_{+}$ & & 0.097 & 0.045 & 0.031 & 0.982 & 0.350\\
$\Delta\text{GDP}_{E2}$ & $\Delta\text{GDP}_{E3}$ & & 0.649 & 0.892 & 0.878 & 0.987 & 0.962\\
$\Delta\text{GDP}_{E2}$ & $\Delta\text{GDP}_{E}$ & & 0.244 & 0.289 & 0.532 & 0.627 & 0.797\\
$\Delta\text{GDP}_{E2}$ & $\Delta\text{GDP}_{I}$ & & 0.306 & 0.205 & 0.226 & 0.998 & 0.622\\
$\Delta\text{GDP}_{E2}$ & $\Delta\text{GDP}_{+}$ & & 0.237 & 0.208 & 0.105 & 0.937 & 0.486\\
$\Delta\text{GDP}_{E3}$ & $\Delta\text{GDP}_{E}$ & & 0.308 & 0.421 & 0.613 & 0.825 & 0.918\\
$\Delta\text{GDP}_{E3}$ & $\Delta\text{GDP}_{I}$ & & 0.377 & 0.350 & 0.321 & 0.993 & 0.778\\
$\Delta\text{GDP}_{E3}$ & $\Delta\text{GDP}_{+}$ & & 0.321 & 0.389 & 0.199 & 0.843 & 0.698\\
$\Delta\text{GDP}_{E}$ & $\Delta\text{GDP}_{I}$ & & 0.862 & 0.407 & 0.974 & 0.246 & 0.986\\
$\Delta\text{GDP}_{E}$ & $\Delta\text{GDP}_{+}$ & & 0.935 & 0.574 & 0.962 & 0.282 & 0.988\\
$\Delta\text{GDP}_{I}$ & $\Delta\text{GDP}_{+}$ & & 0.736 & 0.612 & 0.335 & 0.646 & 0.808\\
\bottomrule
\addlinespace
\multicolumn{8}{p{.73\linewidth}}{ \textit{Notes:} This table presents $p$-values for the tests of conditional unbiasedness of all 15 combinations of the employed GDP proxies. The five columns refer to the instrument choices given in the main text.
}
\end{tabularx}
\end{table}
Table (ref) reports $p$-values of these conditional moment restriction tests for all 15 distinct pairwise combinations of the six GDP proxies.
We find that at the $5\%$ ($10\%$) significance level, the null hypothesis can only be rejected in 2 (5) out of the 75 cases, which is below the nominal level for both choices.
We now describe the forecasts that we compare using our six proxies $\widehat{Y}_t$. The first set of growth forecasts is taken from the SPF. Since 1968, the SPF publishes quarterly forecasts---prepared by private sector economists---of macroeconomic variables in the US. The SPF is widely used for forecast comparisons in the academic literature Car03,Cam07,EMW09. Specifically, we consider the (cross-respondent) mean GDP growth predictions of the SPF, available at \url{https://tinyurl.com/y9p8oylx}.
Second, we employ the so-called “Greenbook” forecasts of GDP growth, available at \url{https://tinyurl.com/y7b6pm2f}.
By using substantial resources, the staff of the Board of Governors of the Federal Reserve prepares these forecasts for each meeting of the Federal Open Market Committee RomerRomer2000.
We use the forecast closest to the middle of each of the respective quarters.
Notice that the Greenbook forecasts are only available to the public after a five-year lag.
Our goal is to show that ECPA tests can more easily discriminate between two forecasts $x_{1t}$ and $x_{2t}$ if a more precise proxy is used. To this end, we require one forecast to be clearly superior to the other. For this, we employ forecasts for the same quarter $t$ (from \textit{either} the SPF \textit{or} Greenbook), however with varying horizons, i.e., forecasts for $\Delta\text{GDP}_t$ issued at different time points $t-\tau$. HE14 formally show that forecasts based on larger information sets are superior to those based on smaller information sets.
The forecast with the shorter horizon ($x_{1t}$) naturally nests the information set of the longer horizon forecast ($x_{2t}$) that was issued a longer time ago, implying superiority of the former.
As we consider multi-step ahead forecasts, we employ a HAC covariance estimator Newey/West:87a, Andrews:91.
\begin{table}[t!]
\caption{Loss Differences and $p$-Values of ECPA Tests for Multiple GDP Measurements}
\begin{tabularx}{1\linewidth}{X @ lrrr @ lrrr @ lrrr}
\toprule
& \multicolumn{4}{c}{1Q vs.\ 2Q Ahead} & \multicolumn{4}{c}{1Q vs.\ 4Q Ahead} & \multicolumn{4}{c}{2Q vs.\ 4Q Ahead} \\
\cmidrule(lr){2-5} \cmidrule(lr){6-9} \cmidrule(lr){10-13}
&& Loss & \multicolumn{2}{c}{$p$-value} && Loss & \multicolumn{2}{c}{$p$-value} && Loss & \multicolumn{2}{c}{$p$-value} \\
\cmidrule(lr){4-5} \cmidrule(lr){8-9} \cmidrule(lr){12-13}
&& Diff. & Inst. $1$ & Inst. $2$ && Diff. & Inst. $1$ & Inst. $2$ && Diff. & Inst. $1$ & Inst. $2$ \\
\midrule
\\
& & \multicolumn{11}{l}{Panel A: Greenbook Forecasts and GDP Vintages} \\
\cmidrule(lr){2-13}
$\Delta\text{GDP}_{E1}$ & & -0.565 & 0.115 & 0.265 & & -0.874 & 0.075 & 0.208 & & -0.278 & 0.194 & 0.374\\
$\Delta\text{GDP}_{E2}$ & & -0.611 & 0.081 & 0.209 & & -0.835 & 0.106 & 0.301 & & -0.179 & 0.404 & 0.462\\
$\Delta\text{GDP}_{E3}$ & & -0.570 & 0.096 & 0.255 & & -0.823 & 0.107 & 0.310 & & -0.218 & 0.329 & 0.406\\
$\Delta\text{GDP}_{E}$ & & -0.554 & 0.091 & 0.206 & & -0.911 & 0.070 & 0.151 & & -0.412 & 0.087 & 0.196\\
\midrule
\\
& & \multicolumn{11}{l}{Panel B: SPF Forecasts and GDP Vintages} \\
\cmidrule(lr){2-13}
$\Delta\text{GDP}_{E1}$ & & -0.551 & 0.127 & 0.344 & & -1.146 & 0.111 & 0.158 & & -0.595 & 0.116 & 0.221\\
$\Delta\text{GDP}_{E2}$ & & -0.508 & 0.157 & 0.406 & & -1.234 & 0.101 & 0.135 & & -0.725 & 0.075 & 0.078\\
$\Delta\text{GDP}_{E3}$ & & -0.464 & 0.197 & 0.432 & & -1.156 & 0.126 & 0.267 & & -0.692 & 0.098 & 0.145\\
$\Delta\text{GDP}_{E}$ & & -0.509 & 0.118 & 0.251 & & -1.286 & 0.094 & 0.129 & & -0.777 & 0.091 & 0.151\\
\midrule
\\
& & \multicolumn{11}{l}{Panel C: Greenbook Forecasts and GDP Measurements} \\
\cmidrule(lr){2-13}
$\Delta\text{GDP}_I$ & & -0.665 & 0.168 & 0.414 & & -0.770 & 0.178 & 0.441 & & -0.178 & 0.427 & 0.302\\
$\Delta\text{GDP}_E$ & & -0.554 & 0.091 & 0.206 & & -0.911 & 0.070 & 0.151 & & -0.412 & 0.087 & 0.196\\
$\Delta\text{GDP}_{+}$ & & -0.555 & 0.152 & 0.286 & & -0.747 & 0.160 & 0.391 & & -0.250 & 0.189 & 0.202\\
\midrule
\\
& & \multicolumn{11}{l}{Panel D: SPF Forecasts and GDP Measurements} \\
\cmidrule(lr){2-13}
$\Delta\text{GDP}_I$ & & -0.557 & 0.117 & 0.245 & & -1.304 & 0.092 & 0.096 & & -0.747 & 0.093 & 0.078\\
$\Delta\text{GDP}_E$ & & -0.509 & 0.118 & 0.251 & & -1.286 & 0.094 & 0.129 & & -0.777 & 0.091 & 0.151\\
$\Delta\text{GDP}_{+}$ & & -0.531 & 0.093 & 0.182 & & -1.258 & 0.086 & 0.063 & & -0.727 & 0.086 & 0.129\\
\bottomrule
\addlinespace
\multicolumn{13}{p{.98\linewidth}}{ \textit{Notes:} This table reports the loss differences (column“Loss Diff.”) together with the $p$-values of the tests of ECPA based on the two instrument choices given in the text for the different comparisons given in Panels A--D.
The first column panel reports results for the comparison of one- against two-quarter ahead forecasts, whereas the second and third column panels compare one- against four-, and two- against four-quarter ahead forecasts.
}
\end{tabularx}
\end{table}
Table (ref) shows $p$-values of ECPA tests (based on $\widehat{T}_n$) for the three combinations of one-, two-, and four-quarter ahead GDP growth forecasts. It does so individually for the SPF and the Greenbook forecasts.
As test functions $h_{t-1}$, we use a constant only (Inst. 1), and a constant jointly with the loss difference of the forecast, lagged by the horizon of the shorter of the two forecast horizons (Inst. 2).
Panels A and B of Table (ref) report results for the four different vintages, while Panels C and D are for $\Delta\text{GDP}_{E}$, $\Delta\text{GDP}_{I}$ and $\Delta\text{GDP}_{+}$.
We find that all loss differences in Table (ref) are negative, implying---as expected---better predictive ability of the shorter horizon forecasts.
For the upper two panels based on the different GDP vintages, we find substantially lower $p$-values for more recent vintages, which we explain by the theoretically higher test power.
E.g., the most recent vintage exhibits a smaller $p$-value than the first vintage in all 12 instances; and it attains overall the smallest $p$-value in nine of the 12 cases.
The results are qualitatively similar---but less pronounced---for the lower two panels of Table (ref), which analyze $\Delta\text{GDP}_{E}$, $\Delta\text{GDP}_{I}$ and $\Delta\text{GDP}_{+}$.
Here, the $\Delta\text{GDP}_+$ approach, which is claimed to be superior in the literature, exhibits the lowest $p$-value only in five instances, but it seems to be generally lower, especially compared to the tests based on $\Delta\text{GDP}_{I}$.
\section{Conclusion}
In this paper, we answer the following question: what can we learn from forecast comparisons if the target variable cannot be observed precisely, and only mismeasured proxy variables are available?
We show that the classical forecast evaluation tools of loss functions and inference thereon can be used without modification if (a) a conditionally unbiased proxy is available and (b) the target functional is the conditional \textit{mean}.
In contrast, for other target functionals such as e.g., the conditional median, the use of approximated proxy variables generally distorts the loss differences and hence, inference in tests for ECPA.
This leads to the perhaps surprising conclusion that when evaluating forecasts in the presence of measurement error, the mean is “more robust” than the median, whereas the converse is well-known in classical estimation theory.
Hence, using the mean as target functional is particularly attractive in forecasting settings that are prone to measurement error, such as in our empirical application on GDP growth.
Further applications are widespread and include, among others, macroeconomic variables as inflation rates, volatility forecasts in finance, meteorological quantities as precipitation or wind speeds, and case or death counts in infectious disease forecasting.
We further demonstrate that even though standard inference in classical ECPA tests for mean forecasts is valid under measurement error, the test's (local) power decreases with the magnitude of the error.
This gives theoretical content to the empirical observation that “[a]lthough consistency of the ordering is ensured by an appropriate choice of the loss function independently of the quality of the proxy, a high precision proxy allows to efficiently discriminate between models” Laurent2013.
We also confirm this in Monte Carlo experiments and provide an empirical illustration on GDP growth.
We emphasize that this increase in power afforded by more precise proxies is particularly important in economics, where sample sizes are often limited---due to low-frequency data collection, structural breaks, etc.
\singlespacing
\putbib[thebib]
\setcounter{page}{1}
\doublespacing