Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
149,975 characters · 12 sections · 52 citation commands
A Novel Approach to Predictive Accuracy Testing in Nested Environments -1ex
{\bf Running Head:} PREDICTIVE ACCURACY IN NESTED ENVIRONMENTS
This paper is concerned with comparing the forecasting performance of two nested models through tests that rely on out of sample mean squared error (MSE) loss differentials. Our proposed approach bypasses the widely documented complications caused by the degenerate asymptotic variances of these differentials that occur in nested environments while also leading to nuisance parameter free standard normal asymptotics. Our approach remains valid under both stationary and persistent predictors thus also greatly expanding its practical relevance in economics and finance.
Since the early work of dm1995 and w1996 a vast body of theoretical research has been concerned with developing new methods for comparing the out of sample predictive ability of competing models. Such tests typically compare the out of sample forecast errors generated from two models under a variety of loss functions and forecasting schemes (e.g. recursive, rolling or fixed updating of model parameter estimates) with the aim of testing the null hypothesis of equal predictive accuracy. Most of the test statistics introduced in this literature are based on estimated out of sample MSE loss differentials associated with the two competing forecast error series and have been shown to be asymptotically normally distributed provided that the models being compared are non-nested and a set of standard regularity conditions hold.
The fundamental difficulties that arise as one moves from a non-nested to a nested environment have also generated a vast and growing literature aiming to operationalise and adapt the above approach to nested models. In a nested modelling context a key complication comes from the fact that under the null of interest the population errors of the two models are identical thus leading to sample MSE loss differentials that are identically zero in the limit with null asymptotic variances. These in turn result in test statistics that are not well defined asymptotically and in the failure of normal approximations for popular test statistics such as the Diebold-Mariano statistic (henceforth referred to as DM).
Alternative normalisations applied to the MSE loss differentials in nested contexts have subsequently been shown to lead to test statistics with well defined but no longer Gaussian limiting null distributions expressed as functionals of stochastic integrals in Brownian Motions (cm2001, cm2005, m2007, ht2015). With the exception of restrictive frameworks that rule out heteroskedasticity or allow only a single additional predictor in the nesting model these distributions typically depend on a variety of model specific parameters that cannot be eliminated via standard HAC type corrections, requiring simulation based approaches for their implementation (see w2006, and cm2013). The asymptotics of these test statistics are further influenced by how the in-sample observations are allowed to grow relative to the out of sample observations and the particular choice of the forecasting scheme used to generate forecasts.
Rather than relying on these non-standard and non-Gaussian distributions this same literature has also proposed to bypass the difficulties underlying nested model comparisons by continuing to use normal approximations for adjusted versions of DM type statistics. In cw2007 for instance the authors introduced an adjustment to the spread of the out of sample MSEs of the two competing models and argued that although asymptotic normality cannot be established per se such an approach results in reasonably accurate inferences with acceptable size distortions. The adjustment essentially corrects for the fact that under the null hypothesis of equal predictive accuracy the MSE of the larger model is contaminated with estimation noise. This adjusted DM type statistic proposed in cw2007 has become the norm in economic applications involving out of sample forecast comparisons with recent examples found in mp2009, imp2016, ew2021 amongst numerous others.
In this paper we introduce an alternative formulation of the out of sample MSE loss differential between two models that is not subject to the variance degeneracy problem of existing procedures. This subsequently allows us to construct novel test statistics for testing the null hypothesis of equal out of sample population MSEs which are shown to have simple nuisance parameter free normal distributions. The main idea underlying our proposed approach is based on the observation that MSE comparisons across two competing models need not be performed within the same out of sample span of available forecast error observations. These can be performed over partially overlapping segments instead, leading to test statistics that accumulate MSE spreads over all possible such segments. This new setting can trivially accommodate desirable features such as conditional heteroskedasticity and persistent predictors and is also shown to lead to both consistent and locally powerful test statistics. As we discuss further below, our approach can also be adapted to broader contexts where the nestedness of models is an important consideration for inferences such as model selection testing.
Besides conventional forecasting objectives, nested models are commonly encountered environments when it comes to testing economic hypotheses and validating theories. Notable examples include forecast accuracy comparisons against random walk models in the exchange rate literature spurred by the early work of mr1983 and more recently reconsidered in r2005, mp2009 amongst others, equity premium predictability issues as recently investigated in fnb2013, aw2017 and numerous others. Our key aim here is to propose a way of addressing and resolving a long standing issue that has generated a vast agenda on the formal comparison of such models via their out-of-sample predictive accuracy. The important auxiliary debate on the advantages or disadvantages of using out-of-sample versus in-sample approaches is not part of our focus. It is also important to emphasise that our interest here is on testing population level predictive ability when forecasts are generated recursively as opposed to finite sample based predictive ability as considered for instance in gw2006. This latter approach is able to avoid the complications induced by the nestedness of models being compared by proceeding via a rolling fixed window based forecasting scheme so that the issue of competing models becoming identical in the limit can be bypassed.
Throughout this paper we also followed the common practice of referring to statistics based on MSE differentials obtained from competing estimated models as Diebold-Mariano type statistics. We must acknowledge however that the specific testing approach initially developed by these authors was not concerned with model specific considerations or specification testing motives as its underlying theory was developed for given sequences of forecast errors assumed to satisfy certain regularity conditions (see d2015). Nevertheless, the forecasting literature of the past decade has generally amalgamated the notion of forecast evaluation with the evaluation of models on the basis of their forecasting abilities.
The paper is organised as follows. Section 2 introduces the nested forecasting environment and establishes the limiting null distributions of two novel test statistics. Section 3 concentrates on their asymptotic power properties, establishing their consistency and ability to detect local departures from the null. Section 4 introduces a simple adjustment to the same statistics shown to further enhance their power properties without affecting their null distributions. Section 5 provides a comprehensive finite sample evaluation of our methods based on two DGPs calibrated to commonly encountered applications. Section 6 illustrates the use of our proposed methods via an application to exchange rate models. Section 7 overviews our key results and discusses extensions. Proofs are given in the Appendix. Further simulation results are provided in an online supplement.
We consider the following predictive regressions
where the ${\bm x}_{it}$'s are the $(p_{i}\times 1)$ vectors of predictors, ${\bm \delta}_{1}$ and ${\bm \beta}_{i}$ the $(p_{1}\times 1)$ and $(p_{i}\times 1)$ parameter vectors and $v_{t}$ and $u_{t}$ the random disturbance terms. We let ${\bm x}_{t}=({\bm x}_{1t}',{\bm x}_{2t}')'$ and ${\bm \beta}=({\bm \beta}_{1}',{\bm \beta}_{2}')'$ and set $p=p_{1}+p_{2}$. Here model ((ref)) is nested within the larger model in ((ref)) and under ${\bm \beta}_{2}=0$ we have ${\bm \delta}_{1}\equiv {\bm \beta}_{1}$ and $v_{t+1}\equiv u_{t+1}$. The formulation of the above two nested models is standard and parallels closely the most commonly encountered setting considered in the predictive accuracy testing literature as for instance in ht2015.
One step ahead forecasts of $y_{t+1}$ from ((ref)) and ((ref)) are generated recursively as $\hat{y}_{1,t+1|t}={\bm x}_{1t}'\hat{\bm \delta}_{1t}$ and $\hat{y}_{2,t+1|t}={\bm x}_{t}'\hat{\bm \beta}_{t}$ for $t=k_{0},\ldots,T-1$ where $\hat{\bm \delta}_{1t}=(\sum_{j=1}^{t}{\bm x}_{1j-1}{\bm x}_{1j-1}')^{-1}\sum_{j=1}^{t}{\bm x}_{1j-1}'y_{j}$, $\hat{\bm \beta}_{t}=(\sum_{j=1}^{t}{\bm x}_{j-1}{\bm x}_{j-1}')^{-1}\sum_{j=1}^{t}{\bm x}_{j-1}'y_{j}$ and the resulting pseudo out of sample forecast errors are then obtained as $\hat{e}_{1,t+1}=y_{t+1}-{\bm x}_{1t}'\hat{\bm \delta}_{1t}$ and $\hat{e}_{2,t+1}=y_{t+1}-{\bm x}_{t}'\hat{\bm \beta}_{t}$. Here $k_{0}$ is the sample location used to initiate the first recursive forecasts that lead to the first out of sample forecast errors $\hat{e}_{1,k_{0}+1}$ and $\hat{e}_{2,k_{0}+1}$ and subsequently resulting in $(T-k_{0})$ out of sample forecast error observations. Throughout this paper we take $k_{0}$ to be a given fraction of the sample size, setting $k_{0}=[T \pi_{0}]$ for some $\pi_{0} \in (0,1)$.
Following the early work of dm1995, w1996 and others, a common approach for comparing the predictive accuracy of the two models under MSE loss involves testing the null hypothesis
using a test statistic based on suitably normalised versions of the sample average MSE loss differentials
Within a non-nested setting and a strictly stationary and ergodic environment dm1995 and w1996 established a standard normal limit theory for this class of test statistics (e.g. $\sqrt{T-k_{0}} \ \overline{D}_{T}/\hat{\sigma}_{{\overline{D}}_{T}}$ with $\hat{\sigma}^{2}_{{\overline{D}}_{T}}$ denoting some suitable long run variance estimator) leading to their systematic use in applied work and a voluminous literature on their refinements. Within a nested context where $u_{t+1}\equiv v_{t+1}$ however it is straightforward to observe that $\sqrt{T-k_{0}} \ \overline{D}_{T}\stackrel{p}\rightarrow 0$ and $\hat{\sigma}_{{\overline{D}}_{T}} \stackrel{p}\rightarrow 0$ invalidating the limiting standard normal approximation and the use of these test statistics for inference purposes. This degeneracy problem is not solely confined to the Diebold-Mariano type statistics but universally affects all existing methods that compare forecast errors (or models) in nested settings with recursively generated forecasts.
These observations have led to a vast body of research on out of sample predictive accuracy testing in nested models due to their importance in empirical applications in areas such as asset pricing and the modelling of expected returns in particular. For their validity, inferences in nested contexts such as ((ref)) and ((ref)) must rely on the observation that under $H_{0}$ it is $(T-k_{0})\overline{D}_{T}$ rather than $\sqrt{T-k_{0}} \ \overline{D}_{T}$ that turns out to have a non-degenerate limit which could be used for developing suitable inferences (see cm2001, cm2005, m2007, ht2015). This however is also problematic due to the non-standard and non-pivotal nature of the resulting asymptotic distributions. These take the form of functionals of stochastic integrals in Brownian Motions and with the exception of some special cases contain nuisance parameters that are difficult to remove via standard HAC type normalisations. Even under special instances such as conditional homoskedasticity these distributions continue to depend on the number of extra predictors included in the nesting models and the fraction of the sample used to build the first recursive forecasts. Equally importantly these results have been obtained under stationarity and ergodicity assumptions ruling out the important and frequently encountered case of predictors having roots near unity in their autoregressive representations.
Instead of evaluating the two sequences of squared forecast errors $\{\hat{e}_{1,t+1}^{2}\}$ and $\{\hat{e}_{2,t+1}^{2}\}$ over the entire {\it and} same interval $[k_{0}+1,T]$ as it is done in the formulation of all commonly used test statistics based on $\overline{D}_{T}$ we here propose to compare the two out of sample MSEs over partially overlapping segments of the $[k_{0}+1,T]$ interval instead. For this purpose we introduce the following generalised MSE spread
where $\ell_{1}$ and $\ell_{2}$ control the range over which the two squared forecast error sequences are evaluated. Note that setting $\ell_{1}=\ell_{2}=T-k_{0}$ in ((ref)) reduces it to $\overline{D}_{T}$ which can be viewed as a special case of $\widetilde{D}_{T}(\ell_{1},\ell_{2})$. In line with the analysis based on (4) we take $\ell_{1}=[(T-k_{0})\lambda_{1}]$ and $\ell_{2}=[(T-k_{0})\lambda_{2}]$ with $\lambda_{1}$ and $\lambda_{2}$ referring to the fraction of the $(T-k_{0})$ squared forecast errors associated with models ((ref)) and ((ref)) respectively.
From a theoretical standpoint, proceeding with the use of $\widetilde{D}_{T}(\ell_{1},\ell_{2})$ instead of $\overline{D}_{T}$ has no bearing on the null hypothesis being tested in the sense that when $u_{t+1}\equiv v_{t+1}$ the population counterpart of $\widetilde{D}_{T}(\ell_{1},\ell_{2})$ also equals zero. A key feature of ((ref)) that distinguishes it from $\overline{D}_{T}$ however is that the variance of its suitably normalised version will no longer be degenerate provided that $\ell_{1}\neq \ell_{2}$ (equivalently $\lambda_{1}\neq \lambda_{2}$). This normalised version of ((ref)) which forms the basis of our proposed test statistics is given by $Z_{T}(\ell_{1},\ell_{2})=\sqrt{T-k_{0}} \ \widetilde{D}_{T}(\ell_{1},\ell_{2})$,
Note that ((ref)) is simply the normalised difference in the means of the two sample MSEs evaluated over the two relevant segments of the effective sample size. \\
REMARK 1: A key point to observe here is that the variance of ((ref)) is well-defined and no longer collapses to zero in the limit provided that $\lambda_{1}$ and $\lambda_{2}$ are bounded away from zero and bounded away from each other. To illustrate and motivate this point heuristically let us replace both $\hat{e}_{1t+1}^{2}$ and $\hat{e}_{2t+1}^{2}$ in ((ref)) with $(u_{t+1}^{2}-\sigma^{2}_{u})$ for $\sigma^{2}_{u}\equiv E[u_{t}^{2}]$. Taking the $u_{t}'s$ to be IID(0,1) with $E[u_{t+1}^{4}]<\infty$ it follows that
suggesting that a test statistic based on $\widetilde{D}_{T}(\ell_{1},\ell_{2})$ will not have a degenerate distribution as it was the case with the use of $\overline{D}_{T}$ in nested contexts. We may also wish to point out that having $\lambda_{1}$ and $\lambda_{2}$ bounded away from zero is merely a technical requirement in the asymptotics that follow as in practice these parameters will naturally be set at or near their maximum boundary of one.
The quantity in ((ref)) forms the building block of our proposed test statistics for testing the null hypothesis in ((ref)) against one sided right tail based alternatives as it is the norm in this literature. We consider two types of test statistics that operationalise ((ref)). Our choice is guided by the simplicity of the ensuing asymptotics and their intuitive interpretation while recognising that alternative constructions/normalisations of $\widetilde{D}_{T}(\ell_{1},\ell_{2})$ may also be considered.
The first test statistic that we consider is denoted $Z_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and is based on implementing inferences for given magnitudes $\ell_{1}^{0}=[(T-k_{0})\lambda_{1}^{0}]$ and $\ell_{2}^{0}=[(T-k_{0})\lambda_{2}^{0}]$. We write
with $\hat{\sigma}^{2}$ denoting a consistent estimator of $V[u_{t+1}^{2}]$.
Our second test statistic is based on averaging ((ref)) across the $\ell_{j}$'s. The averaging can be implemented over $\ell_{1}\in [1,T-k_{0}]$ for a given $\ell_{2}^{0}$ (e.g., $\ell_{2}^{0}=T-k_{0}$) so that the MSE of the smaller model accumulates progressively as $\ell_{1}$ increases. More generally, this averaging can be performed over any desired and feasible range of $\ell_{1}$. To allow such level of generality we introduce the fractional parameter $\tau_{0}$ and write
where the choice of $\tau_{0}$ determines the user-chosen range of $\ell_{1}$ over which the average of $Z_{T}(\ell_{1}, \ell_{2}^{0})$ is taken (given $\ell_{2}^{0}=[(T-k_{0})\lambda_{2}^{0}]$). Rather than imposing a fixed and given $\lambda_{1}^{0}$ as in ((ref)) this average based statistic essentially considers a range of such magnitudes and subsequently aggregates outcomes via averaging. Given the role played by $\ell_{1}$ and $\ell_{2}$ in our inferences we can expect that choosing the averaging range in a way that excludes low magnitudes of $\ell_{1}$ (so that the estimated MSEs associated with model 1 remain sufficiently accurate) will result in more reliable inferences. The issue of how best to select these user-inputs is postponed until Section 3 where we provide precise guidelines informed by a theoretical local power analysis. One motivation behind this average based statistic when compared with ((ref)) is that one can remain partly more agnostic about the specific magnitude to use for one of the two required user-inputs in $Z_{T}([(T-k_{0})\lambda_{1}^{0}],[(T-k_{0})\lambda_{2}^{0}])$ while setting the other one (e.g., $\lambda_{2}^{0}$) at or near its maximum boundary of one. Although our context is different, this is also reminiscent of the various approaches used in the structural break testing literature when one does not wish to take a stance on the location of a potential change-point. More importantly, and borrowing from the same literature, we may also conjecture that the averaging process may result in tests with more favorable size-power trade-offs.
At this stage it is also important to point out that there are numerous alternative possibilities for designing test statistics in the spirit of ((ref)) and ((ref)) (e.g., double averaging across $\ell_{1}$ and $\ell_{2}$, alternative functional forms etc.). An interesting avenue for future research could be the design of a class of test statistics based on $Z_{T}(\ell_{1},\ell_{2})$ and having desirable optimality properties as it has been attempted in the structural break literature.
Although both ((ref)) and ((ref)) allow for a broad range of theoretically feasible magnitudes for $(\lambda_{1}^{0},\lambda_{2}^{0})$ in $Z^{0}_{T}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $(\tau_{0},\lambda_{2}^{0})$ in $\overline{Z}_{T}(\tau_{0};\lambda_{2}^{0})$ one naturally expects that choosing $(\lambda_{1}^{0},\lambda_{2}^{0})$ and $(\tau_{0},\lambda_{2}^{0})$ to lie in the vicinity of unity would capture the greatest amount of information from the two competing models. As we show further below such a choice does indeed lead to remarkably powerful tests with excellent size control. Given a sequence of forecast errors available to the investigator the practical implementation of either ((ref)) or ((ref)) is also as straightforward as calculating standard DM type test statistics.
To establish the limiting properties of our test statistics under the null hypothesis in ((ref)) we introduce a set of high level assumptions ensuring a flexible environment that encompasses the vast majority of settings considered in the literature while also allowing for a richer temporal structure. As we wish to highlight the generality and usefulness of our methods based on the use of ((ref))-((ref)) we abstain from primitive conditions that may unnecessarily suggest a restrictive scope for their use. More importantly our use of high level assumptions is motivated by the fact that our proposed methods can be immediately seen to be robust to a very rich dynamic structure of predictors including highly persistent processes, strictly stationary and ergodic processes, long memory processes etc.\\
ASSUMPTION A: \\ {\em
(i) $\displaystyle \sup_{\lambda \in (0,1]} \left| \frac{\sum\limits_{t=k_{0}}^{k_{0}-1+[(T-k_{0})\lambda]}\hat{e}_{j,t+1}^{2}}{\sqrt{T-k_{0}}}- \frac{\sum\limits_{t=k_{0}}^{k_{0}-1+[(T-k_{0})\lambda ]}u_{t+1}^{2}}{\sqrt{T-k_{0}}} \right| \stackrel{H_{0}}=o_{p}(1)$ for $j=1,2$.\\
(ii) The sequence of demeaned squared errors $\eta_{t}=u_{t+1}^{2}-\sigma^{2}_{u}$ has autocovariances $\gamma_{j}^{\eta}$ that satisfy $\sum_{j=0}^{\infty}|\gamma_{j}^{\eta}|<\infty$ and fulfills a functional central limit theorem, that is $T^{-\frac{1}{2}}\sum_{t=1}^{[Ts]} (u_{t+1}^{2}-\sigma^{2}_{u})\stackrel{\cal{D}}\rightarrow \sigma W_{\eta}(s)$ on $D_{\mathbb{R}}([0,1])$ the space of cadlag functions on $[0,1]$ with $W_{\eta}(.)$ denoting a standard Brownian Motion and $\sigma^{2}=\gamma_{0}^{\eta}+2 \sum_{j=1}^{\infty}\gamma_{j}^{\eta}$>0. \\
(iii) A consistent estimator $\hat{\sigma}^{2}$ of $\sigma^{2}$ exists, that is $\hat{\sigma}^{2} \stackrel{p}\rightarrow \sigma^{2} \in (0,\infty)$.\\ }
We note that condition A(i) holds for both $j=1$ and $j=2$ highlighting the fact that we operate within a nested environment with ((ref)) being the true model. A1(i) is trivially satisfied under a very broad range of settings used to obtain the large sample properties of DM type statistics in nested models. An important feature to also highlight here is the fact that A(i) does not restrict the persistence properties of the predictors which could be highly persistent in the sense of following local to unit root processes for instance. This greatly expands and enriches the environment in which predictive accuracy inferences have commonly been introduced. The robustness of A(i) to the persistence properties of the predictors is an important and useful feature of the squared forecast errors as opposed to their level for which a result such as A(i) would not hold. For more primitive conditions illustrating specialised environments under which A(i) holds see dp2008a and ht2015 for strictly stationary and ergodic/mixing settings and brn2019 for environments where A(i) is shown to hold under both stationary and unit-root or near unit-root regressors.
Assumption A(ii) requires the centered squared errors driving ((ref))-((ref)) to satisfy a functional central limit theorem with $\sigma^{2}$ referring to their long run variance. The absolute summability of the autocovariances of $\eta_{t}$ ensures that $\sigma^{2}$ the limit of $V[\sum_{t=k_{0}}^{T-1}\eta_{t+1}/\sqrt{T-k_{0}}]$ exists. Examples of processes which satisfy Assumption A(ii) include a broad range of conditionally heteroskedastic ARCH/GARCH processes under suitable existence of moments restrictions. For a detailed set of primitive assumptions ensuring that the stated FCLT holds see gkl2000, gkl2001, bhh2008, l2009 and references therein.
Assumption A(iii) requires that the long run variance of $\eta_{t}$ be estimated consistently. Such an estimator could be trivially constructed using least squares residuals from either the null or alternative models. Under conditional homoskedasticity an obvious candidate would be $\hat{\sigma}^{2}_{hom}=\sum_{t=k_{0}}^{T-1}\hat{\eta}_{t+1}^{2}/(T-k_{0})$ while under dependent errors (e.g. if the $u_{t}'s$ follow a GARCH type process) a Newey-West type formulation as in dp2008b would be suitable and ensure that A(iii) holds. \\
REMARK 2: As pointed out in Remark 1, our asymptotic theory for $Z_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ in ((ref)) imposes $\lambda_{1}^{0}$ and $\lambda_{2}^{0}$ to be bounded away from zero and to be bounded away from each other, say $0<\underline{\lambda}\leq \lambda_{i}^{0}\leq 1$ for $i=1,2$ and $|\lambda_{1}^{0}-\lambda_{2}^{0}| \geq \epsilon$ for some positive fraction $\epsilon$. In what follows we refer to such a set from which these two user-inputs can be selected as $\Lambda^{0}$. The implementation of the average based statistic $\overline{Z}_{T}(\tau_{0};\lambda_{2}^{0})$ in ((ref)) requires setting $\lambda_{2}^{0}$ as above and averaging $Z_{T}(\ell_{1},[(T-k_{0})\lambda_{2}^{0}])$ across $\ell_{1}=[(T-k_{0})\tau_{0}]+1,\ldots,(T-k_{0})$ for some $\tau_{0}$ bounded away from zero and one. We refer to this set as $\overline{\Lambda}^{0}$.
The following two propositions summarise the large sample behavior of our two test statistics under the null hypothesis stated in ((ref)).
{\bf PROPOSITION 1}. {\em Under Assumptions A(i)-(iii), the null hypothesis in ((ref)), and for given $(\lambda_{1}^{0},\lambda_{2}^{0})\in \Lambda^{0}$, we have as $T \rightarrow \infty$
where
}
{\bf PROPOSITION 2}. {\em Under Assumptions A(i)-(iii), the null hypothesis in ((ref)), and for given $(\tau_{0},\lambda_{2}^{0}) \in \overline{\Lambda}^{0}$, we have as $T \rightarrow \infty$
where {
} }
The variance components of the distributional outcomes in ((ref)) and ((ref)) are of course known to the investigator so that both $Z_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{Z}_{T}(\tau_{0};\lambda_{2}^{0})$ can be trivially standardised as
and
to proceed with standard normal inferences for testing $H_{0}$.
The results in ((ref))-((ref)) highlight the simplicity and practicality of our proposed inferences while at the same time offering a solution to an important problem that has not been satisfactorily resolved in this literature. The test statistics in ((ref)) and ((ref)) allow us to generalise the widely used DM style forecast accuracy testing approach to a broad class of empirically relevant models including specifications with highly persistent predictors with or without conditional heteroskedasticity.
We naturally expect the quality of inferences (e.g. power, size vs power trade-offs) to be influenced by the specific choices of ($\lambda_{1}^{0},\lambda_{2}^{0}$) in ((ref)) and ($\tau_{0},\lambda_{2}^{0}$) in ((ref)). Although the above null asymptotics hold under a very broad range of parameterisations for those user-inputs a formal analysis of their local asymptotic power allows us to provide precise and tight guidelines ensuring excellent power properties with good size control.
We here deviate from Assumption A(i) in order to evaluate the large sample behavior of ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T}(\tau_{0};\lambda_{2}^{0})$ when the DGP is given by ((ref)). Assumption A(i) continues to hold for $j=2$ but no longer for $j=1$ since model ((ref)) is misspecified due to the omitted ${\bm x}_{2,t}$ predictors. As $\hat{e}_{1,t+1}^{2}$ will now be contaminated by those omitted predictors we expect the stochastic properties of the latter (e.g. the variance of the ${\bm x}_{2,t}$'s and their correlation with the ${\bm x}_{1,t}$'s) to influence the power properties of both test statistics. Unlike their null distributions we thus also expect the test statistics to diverge at different rates depending on whether the predictors are stationary or highly persistent. Our analysis of the consistency and local power properties of ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T}(\tau_{0};\lambda_{2}^{0})$ is guided by these two distinct scenarios which we consider separately.
We initially concentrate on the case where the predictors driving both ((ref)) and ((ref)) are stationary and ergodic. Specifically, we operate under the following set of high level assumptions that mirror closely the most common environments considered in the predictive accuracy testing literature. \\
ASSUMPTION B1: \\
{\em (i) $\displaystyle \sup_{\lambda\in [0,1]}\left\lVert \dfrac{\sum_{t=1}^{[T\lambda]} {\bm x}_{t}{\bm x}_{t}'}{T}-\lambda \ {\bm Q}\right\rVert=o_{p}(1)$ with ${\bm Q}$ a $p\times p$ nonrandom positive definite matrix,
(ii) $\displaystyle \dfrac{\sum_{t=1}^{[T\lambda]} {\bm x}_{t}u_{t+1}}{\sqrt{T}} \stackrel{{\cal D}}\rightarrow {\bm \Omega}^{1/2} \ {\bm W}(\lambda)$ with ${\bm W}(.)$ denoting a p-dimensional standard Brownian Motion and ${\bm \Omega}=E[{\bm x}_{t}{\bm x}_{t}'u_{t+1}^{2}]>0$,
(iii) Assumptions A(ii)-(iii) hold.
(iv) The user-inputs in ${\cal S}_{T}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{\cal S}_{T}(\tau_{0};\lambda_{2}^{0})$ are such that $(\lambda_{1}^{0},\lambda_{2}^{0})\in \Lambda^{0}$ and $(\tau_{0},\lambda_{2}^{0})\in \overline{\Lambda}^{0}$ respectively. \\
}
Assumptions B1(i)-(iii) mirror closely the environment of ht2015 and can be viewed as more primitive conditions ensuring that Assumption A(i) holds. B1(i) requires that the predictors satisfy a uniform law of large numbers and rules out trending or local to unit root predictors while B1(ii) ensures that $\{{\bm x}_{t}u_{t+1}\}$ satisfies a multivariate functional central limit theorem. Our main result regarding the asymptotic power properties of the two tests within such a stationary environment is now summarised in Proposition 3 below. \\
{\bf PROPOSITION 3}. {\it (i) Suppose model ((ref)) holds with ${\bm \beta}_{2}\neq 0$ and fixed, then under assumption B1 and as $T \rightarrow \infty$ we have ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0}) \stackrel{p}\rightarrow \infty$ and $\overline{{\cal S}}_{T}(\tau_{0};\lambda_{2}^{0}) \stackrel{p}\rightarrow \infty$. (ii) Suppose model ((ref)) holds with ${\bm \beta}_{2}={\bm \gamma}/T^{1/4}$ for ${\bm \gamma\neq 0}$. Under assumption B1, $\lim_{||{\bm \gamma}||\rightarrow \infty}\lim_{T \rightarrow \infty}{\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})=\infty$ and $\lim_{||{\bm \gamma}||\rightarrow \infty}\lim_{T \rightarrow \infty}\overline{{\cal S}}_{T}(\tau_{0};\lambda_{2}^{0})=\infty$ in probability.} \\
The above results highlight the consistency of both test statistics as well as their ability to detect local departures from the null hypothesis under stationary settings. It is here also important to point out that the local to the null parameterisation of ${\bm \beta}_{2}$ based on $T^{1/4}$ rather than the usual $T^{1/2}$ rate commonly encountered in stationary settings is not in any way due to our specific test statistics or assumptions. The same scenario would also occur in a conventional regression based testing environment and is due to the fact that we are dealing with inferences about the behavior of {\it squared} errors rather than their level.
To gain further insights into the specific role played by key factors influencing power it is useful to also present the explicit asymptotic local power functions of the two tests for a given size $\alpha \in (0,1)$. These will in turn be used to provide explicit guidance on selecting suitable parameterisations of our two test statistics (i.e. $(\lambda_{1}^{0},\lambda_{2}^{0})$ in ((ref)) and $(\tau_{0};\lambda_{2}^{0})$ in ((ref))). In what follows it is useful to also recall that $\pi_{0}$ refers to the given fraction of the sample size used to initiate the recursive computation of forecasts. \\
{\bf COROLLARY 1}. {\it Suppose model ((ref)) holds with ${\bm \beta}_{2}={\bm \gamma}/T^{1/4}$ for ${\bm \gamma\neq 0}$. Under assumption B1 and letting $q_{\alpha}$ denote the upper $\alpha$-quantile of the standard normal distribution with CDF $\Phi(.)$, the asymptotic local power functions of the tests based on ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T}(\tau_{0};\lambda_{2}^{0})$ are given by $1-\Phi(q_{\alpha}-\psi^{0})$ and $1-\Phi(q_{\alpha}-\overline{\psi})$ respectively, where
with $v^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\bar{v}(\tau_{0};\lambda_{2}^{0})$ as in ((ref)) and ((ref))-((ref)) and the ${\bm Q}_{ij}$'s referring to the components of the population moment matrix ${\bm Q}$ in assumption B1(i).
}
We note that power is monotonic in the sense that both $\psi^{0}$ and $\overline{\psi}$ are non-decreasing as $||{\bm \gamma}||$ gets large. For a given significance level, the larger the two non-centrality parameters are the greater the associated probabilities of rejecting the null hypothesis.
The expressions in ((ref))-((ref)) are particularly useful for highlighting the factors that influence power by shifting the center of the null asymptotic standard normal distributions away from zero. Viewing the asymptotic local power functions $\Phi(\psi^{0}-q_{\alpha})$ and $\Phi(\overline{\psi}-q_{\alpha})$ in Corollary 1 as providing approximations to the correct decision frequencies of the two test statistics under a sufficiently large $T$ and {\emph specific alternatives}, we note that for a given size $\alpha$ both test statistics are expected to exhibit a stronger ability to detect departures from the null when the variances of the omitted predictors are large and their correlation with the included predictors small. This feature is particularly important since it hints at the fact that the presence of nearly integrated predictors may help enhance power, a scenario we formally consider further below.
To highlight these points with greater clarity it is useful to focus on the simplified case of two centered predictors ${\bm x}_{t}=(x_{1,t},x_{2.t})$ so that ((ref))-((ref)) simplify as
with $\rho_{12}=Corr[x_{1,t},x_{2,t}]$. All other things being equal, power is expected to deteriorate under a noisy omitted predictor that has low variance (low $E[x_{2,t}^{2}]$) and/or that is highly correlated with the included predictor (e.g. $|\rho_{12}| \approx 1$). Interestingly, this also suggests that an ideal setting in terms of power implications is one where omitted predictors are highly persistent while included predictors are stationary so that $\rho_{12}\approx 0$ with $E[x_{2,t}^{2}]$ large. Another important factor affecting power is the variance of the $u_{t}^{2}$'s which impacts the magnitudes of $\psi^{0}$ and $\overline{\psi}$ via $\sigma \equiv \sqrt{V[u_{t+1}^{2}]}$. Within an NID errors setting for instance we have $V[u_{t+1}^{2}]=E[u_{t+1}^{4}]-\sigma^{4}_{u}=2 \sigma^{4}_{u}$ so that {\it all other things being equal}, an environment with high kurtosis will have a detrimental impact on the power properties of both test statistics.\\
{\it Power enhancing choices for $(\lambda_{1}^{0},\lambda_{2}^{0})$ and $(\tau_{0};\lambda_{2}^{0})$} \\
The non-centrality parameters in ((ref))-((ref)) are also useful for assessing the impact of $(\lambda_{1}^{0}, \lambda_{2}^{0})$ and $(\tau_{0},{\lambda}_{2}^{0})$ on both the absolute and relative local powers of the two tests and for providing useful guidance on suitable choices for those user-inputs. From Corollary 1, since the mapping $m \mapsto P[Z>q_{\alpha}-m]$ is increasing in $m$ on $[0,\infty)$, a test of size $\alpha$ based on $S_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ will be preferable, in terms of its local power, to a test of the same size based on $S_{T}^{0}({\lambda_{1}^{0}}',{\lambda_{2}^{0}}')$ whenever $\psi^{0}(\lambda_{1}^{0},\lambda_{2}^{0})>\psi^{0}({\lambda_{1}^{0}}',{\lambda_{2}^{0}}')$, holding all other parameters entering $\psi^{0}$ constant. Given $\psi^{0}$ in ((ref)) with $v^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ defined as in ((ref)) it follows that those two parameters should be set near their boundary of 1 and in close vicinity of one another (e.g. ${\cal S}_{T}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ for $\lambda_{2}^{0}\approx 0.9$ as a possibility). \\
Regarding the average based statistic $\overline{{\cal S}}_{T}(\tau_{0};\lambda_{2}^{0})$ we note from ((ref))-((ref)) that $\overline{\psi}$ in ((ref)) viewed as a function of $\lambda_{2}^{0}$ and $\tau_{0}$ (holding all other parameters constant) reaches its unique maximum for
supporting the use of $\overline{{\cal S}}_{T}(\tau_{0};\lambda_{2}^{0}=0.5\tau_{0}+0.5)$ in its practical implementation. If $\tau_{0}=0.5$ for instance, which corresponds to a test statistic that averages across the largest half of the $\ell_{1}$ magnitudes, this approximate asymptotic power based metric points to an implementation based on $\overline{{\cal S}}_{T}(\tau_{0}=0.5;\lambda_{2}^{0}=0.75)$. Since $\overline{\psi}$ is also a monotonically increasing function of $\tau_{0}$ however, it also follows that the same average based statistic should be operationalised with a choice of $\tau_{0}$ that is in the vicinity of 1 (e.g. $\overline{{\cal S}}_{T}(\tau_{0}=0.8;\lambda_{2}^{0}=0.5(0.8)+0.5)$ or $\overline{{\cal S}}_{T}(\tau_{0}=0.9;\lambda_{2}^{0}=0.5(0.9)+0.5)$ as possibilities). A practical side to this power enhancing choice of $\lambda_{2}^{0}$ is that the implementation of $\overline{\cal S}_{T}(\tau_{0};\lambda_{2}^{0})$ essentially requires only a single user-input. \\
Given these preferred parameterisations of the two test statistics it is also useful to evaluate whether either of the two statistics is expected to dominate the other in the sense of $\psi^{0}$ being greater or smaller than $\overline{\psi}$ over particular regions of the pairs $(\lambda_{1}^{0},\lambda_{2}^{0})$ and $(\tau_{0},\lambda_{2}^{0}=0.5\tau_{0}+0.5)$, holding all other parameters constant. Given the standard normal asymptotics of both test statistics a useful metric for comparing their local powers is Pitman's Asymptotic Relative Efficiency (ARE) which here takes particularly simple forms, following directly from Corollary 1. To avoid confusion between the $\lambda_{2}^{0}$ parameter used in ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\lambda_{2}^{0}$ used in $\overline{\cal S}_{T}(\tau_{0},\lambda_{2}^{0})$ we write the two statistics as ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{\cal S}_{T}(\tau_{0};\overline{\lambda}_{2}^{0})$ with $\overline{\lambda}_{2}^{0}=0.5 \tau_{0}+0.5$ as in ((ref)). Their ${\text{ARE}}$ is now given by
and more specifically
From ((ref)) we can observe a clear trade-off between $\tau_{0}$ and the magnitudes of $\lambda_{1}^{0}$ and $\lambda_{2}^{0}$ used in ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$. If we focus on $\lambda_{1}^{0}=1$ it follows from ((ref)) that ${\text{ARE}}\geq 1$ for
which is a monotonically increasing function of $\tau_{0}$ and highlights the fact that the average based statistic will dominate ${\cal S}_{T}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ in terms of its local power (i.e. ${\text{ ARE}} <1$) unless impractically large magnitudes of $\lambda_{2}^{0}$ are used in its implementation. If the average based statistic is implemented with $\tau_{0}=0.8$ for instance, its power properties will dominate ${\cal S}_{T}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ unless $\lambda_{2}^{0}>0.9798$. If it is implemented with $\tau_{0}=0.9$ the average based statistic will again dominate ${\cal S}_{T}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ unless $\lambda_{2}^{0}>0.9908$. These values suggest that the average based statistic with $\tau_{0}$ set in the vicinity of unity (e.g. $\overline{\cal S}(\tau_{0}=0.8;\overline{\lambda}_{2}^{0}=0.9)$) will dominate ${\cal S}^{0}_{T}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ in terms of its local power unless impractically large magnitudes of $\lambda_{2}^{0}$ are used in ${\cal S}^{0}_{T}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$.
We now consider an environment where the $p$ predictors ${\bm x}_{t}$ entering ((ref))-((ref)) are modelled as local to unit root processes specified as
where ${\bm C}=diag(c_{1},\ldots,c_{p})$ for $c_{i}>0$, $i=1,\ldots,p$ and $\epsilon_{t}$ some stationary and ergodic random disturbance process. The new set of assumptions under which we establish our results are now summarised in Assumption B2 below where ${\bm J}_{C}(s)=({\bm J}_{1C}(s),{\bm J}_{2C}(s))'$ denotes a p-dimensional Ornstein-Uhlenbeck process whose two components ${\bm J}_{1C}(s)$ and ${\bm J}_{2C}(s)$ are associated with the dynamics of ${\bm x}_{1,t}$ and ${\bm x}_{2,t}$ respectively. \\
ASSUMPTION B2: \\
{\em (i) $\left(\dfrac{{\bm x}_{[Ts]}}{\sqrt{T}},\dfrac{\sum_{t=1}^{[Ts]} u_{t}}{\sqrt{T}}, \dfrac{\sum_{t=1}^{[Ts]} (u_{t}^{2}-\sigma^{2}_{u})}{\sqrt{T}}\right) \stackrel{\cal{D}}\rightarrow ({\bm J}_{C}(s),\sigma_{u} W_{u}(s), \sigma W(s))$, $s \in [0,1]$.
(ii) Assumption A(iii) holds.
(iii) The user-inputs in ${\cal S}_{T}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{\cal S}_{T}(\tau_{0};\lambda_{2}^{0})$ are such that $(\lambda_{1}^{0},\lambda_{2}^{0})\in \Lambda^{0}$ and $(\tau_{0},\lambda_{2}^{0})\in \overline{\Lambda}^{0}$ respectively.
}
The asymptotic power properties of ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T}(\tau_{0};\lambda_{2})$ are now summarised in Proposition 4 below. \\
{\bf PROPOSITION 4}. {\it (i) Suppose model ((ref)) holds with ${\bm \beta}_{2}\neq 0$ and fixed, then under Assumption B2 and as $T \rightarrow \infty$ we have $ {\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0}) \stackrel{p}\rightarrow \infty$ and $\overline{{\cal S}}_{T}(\tau_{0};\lambda_{2}) \stackrel{p}\rightarrow \infty$. (ii) Suppose model ((ref)) holds with ${\bm \beta}_{2}={\bm \gamma}/T^{3/4}$ for ${\bm \gamma\neq 0}$. Under Assumption B2, $\lim_{||{\bm \gamma}||\rightarrow \infty}\lim_{T \rightarrow \infty}{\cal S}^{0}_{T}(\lambda_{1}^{0},\lambda_{2}^{0})=\infty$ and $\lim_{||{\bm \gamma}||\rightarrow \infty}\lim_{T \rightarrow \infty} \overline{{\cal S}}_{T}(\tau_{0};\lambda_{2})=\infty$ in probability.} \\
A key message that is conveyed by Proposition 4 when contrasted with Proposition 3 is the important impact of persistence on the power properties the test statistics. The presence of persistent predictors leads to a faster divergence rate for both statistics as reflected in the faster convergence rate towards zero of ${\bm \beta}_{2}$ that can be accommodated. With highly persistent predictors, both test statistics diverge at the same $T^{3/2}$ rate compared with a rate of $T^{1/2}$ when predictors were stationary.
A more explicit formulation of the departure from the null distribution in this local to unit-root context can also be highlighted through the following formulations of the limiting distributions of the two test statistics under the local alternative of interest. \\
{\bf COROLLARY 2}. {\it Suppose model ((ref)) holds with ${\bm \beta}_{2}={\bm \gamma}/T^{3/4}$ for ${\bm \gamma\neq 0}$. Under assumption B2 and as $T \rightarrow \infty$ we have
with
where ${\bm J}_{C}^{*}(s)={\bm J}_{2C}(s)-{\bm M}(s) {\bm J}_{1C}(s)$ and ${\bm M}(s)=(\int_{0}^{s}{\bm J}_{1C}{\bm J}_{1C}')^{-1}(\int_{0}^{s}{\bm J}_{1C}{\bm J}_{2C}')$. } \\
It is here interesting to compare ((ref))-((ref)) with the non-centrality parameters ((ref))-((ref)) obtained in the stationary context. The two pairs are essentially analogous in the sense that the constant population moments of predictors (i.e. the ${\bm Q}_{i,j}$'s) are now replaced by stochastic integrals in Ornstein-Uhlenbeck processes (i.e. ${\bm J}_{i,C}$). The homogeneity throughout the sample of the limit moment matrix in Assumption B1(i) is of course no longer valid in the context of the stochastic integrals in ((ref))-((ref)). The above results also imply that the role played by the pairs $(\lambda_{1}^{0}, \lambda_{2}^{0})$ and $(\tau_{0},\lambda_{2}^{0})$ in this local to unit-root context will mirror our earlier analysis based on a stationary setting, supporting the same practical implementation of both test statistics in terms of their parameterisations i.e. the power enhancing choices for $(\lambda_{1}^{0}, \lambda_{2}^{0})$ and $(\tau_{0},\lambda_{2}^{0})$ discussed above continue to hold in the current context.
Here we explore a particular adjustment that can be applied to our two test statistics ${\cal S}_{T}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{\cal S}_{T}(\tau_{0};\lambda_{2}^{0})$ with the purpose of boosting their asymptotic local power properties without affecting their limiting null distributions. The theoretical principle underlying our proposed approach mirrors the idea in fly2015 where the authors proposed to augment Wald type statistics with a component that vanishes asymptotically under the null while diverging under alternatives of interest. More formally we seek to augment our proposed two test statistics as
for some suitably chosen $h_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{h}_{T}(\tau_{0};\lambda_{2}^{0})$ terms which are such that these adjusted versions of our two test statistics maintain the same limiting null distributions as in Proposition 1 while at the same time displaying more favorable power properties.
In what follows we show that a particular transformation of the forecast errors $\hat{e}_{2,t+1}$ associated with the larger forecasting model can be used to design such augmentation terms in a way that fulfils the requirement that $h_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0}) $ and $\overline{h}_{T}(\tau_{0};\lambda_{2}^{0})$ vanish asymptotically under the null while diverging at a desirable rate under the alternative. The augmentation we propose to consider is motivated by the well known Clark and West adjustment to DM type statistics introduced in cw2007. The original motivation behind Clark and West's adjustment relied on the intuition that under the null hypothesis estimation noise contaminates the ${\hat{e}_{2,t+1}}$'s due to the estimation of parameters that are zero in the population. This in turn translates into an inflated $\text{MSE}_{2}$ resulting in test statistics that are severely undersized. Clark and West proposed to correct for such distortions by suitably adjusting the magnitudes of the forecast errors estimated from the larger model.
To lay down the context and with no loss of generality it is useful to operate within a simplified version of ((ref))-((ref)), setting ${\bm \delta}_{1}=0$ and ${\bm \beta}_{1}=0$ so that $\hat{e}_{1,t+1}=u_{t+1}$ and $\hat{e}_{2,t+1}=u_{t+1}-{\bm x}_{2,t}'(\hat{\bm \beta}_{2}-\bm \beta_{2})$. We can now write $Z_{T}(\ell_{1},\ell_{2})$ in ((ref)) as
Although it is implicit in our Assumption A(i) that the last two terms in the right hand side of ((ref)) vanish asymptotically under the null hypothesis, in finite samples the rightmost quadratic form is likely to pull down the spread in MSEs causing their null distribution to be mis-centered. Noting that $(\hat{\bm \beta}_{2,t}-\bm \beta_{2})' {\bm x}_{2,t} {\bm x}_{2,t}'(\hat{\bm \beta}_{2,t}-\bm \beta_{2}) \equiv (\hat{e}_{1,t+1}-\hat{e}_{2,t+1})^{2}$, cw2007's proposal was to reformulate the sample MSE spreads between Models 1 and 2 with an adjusted version of $\hat{e}_{2,t+1}^{2}$, say $\widetilde{e}_{2,t+1}^{2}$, given by
It turns out that implementing the adjustment in ((ref)) within our two test statistics (i.e., using $\widetilde{e}_{2,t+1}^{2}$ instead of $\hat{e}_{2,t+1}^{2}$ in ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{\cal S}_{T}(\tau_{0};\lambda_{2}^{0})$) allows us to reformulate them as in ((ref))-((ref)) with $h_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{h}_{T}(\tau_{0};\lambda_{0}^{2})$ fulfilling the desirable requirements in fly2015 in the sense that the adjustments do not alter the asymptotic null distributions of the test statistics while at the same time leading to an increase in the associated non-centrality parameters under the local alternatives of interest.
The expression in ((ref)) is also useful for highlighting what distinguishes our framework that operates under $\ell_{1}\neq \ell_{2}$ with a standard approach that sets $\ell_{1}=\ell_{2}=T-k_{0}$ as for instance in all Diebold-Mariano type statistics. Under $\ell_{1}=\ell_{2}$ we note that the first two terms in the right hand side of ((ref)) cancel out so that the asymptotic behavior of the expression is determined by the two rightmost quadratic forms whose {\it non-normalised} versions have been shown to be $O_{p}(1)$ with non-standard limits (see cm2001, cm2005). Allowing $\ell_{1}\neq \ell_{2}$ essentially forces the asymptotics of the MSE spreads to be driven solely by the first two components in the right hand side of ((ref)).
Letting ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T,adj}(\tau_{0};\lambda_{2}^{0})$ denote the adjusted versions of our two test statistics it immediately follows from ((ref)) and standard algebra that
and
The expressions in ((ref)) and ((ref)) highlight the fact that the adjustment to the MSE of the larger model results in test statistics that are augmented versions of ${\cal S}_{T}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{\cal S}_{T}(\tau_{0};\lambda_{2}^{0})$. As we establish formally below the presence of the additional terms $h_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{h}_{T}(\tau_{0};\lambda_{2}^{0})$ leaves the limiting null distributions unchanged as both quantities vanish asymptotically. Under the alternative both $h_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{h}_{T}(\tau_{0};\lambda_{2}^{0})$ diverge to infinity at the same rate as ${\cal S}_{T}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{\cal S}_{T}(\tau_{0};\lambda_{2}^{0})$ implying that ${\cal S}_{T,adj}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{\cal S}_{T,adj}(\tau_{0};\lambda_{2}^{0})$ will also share the consistency and local power characteristics of their unadjusted counterparts in the sense of diverging to infinity as $||{\bm \gamma}||\rightarrow \infty$. More importantly however, the presence of $h_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{h}_{T}(\tau_{0};\lambda_{2}^{0})$ does result in different (strictly larger) non-centrality parameters that make these adjusted statistics have more favorable power properties. These features are formalised in Proposition 5 and Corollary 3 below. \\
{\bf PROPOSITION 5}. {\em The results in Propositions 1-4 continue to hold when ${\cal S}^{0}_{T}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T}(\tau_{0};\lambda_{2}^{0})$ are replaced with ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T,adj}(\tau_{0};\lambda_{2}^{0})$ respectively. } \\
{\bf COROLLARY 3}. {\em (i) Under the assumptions of Corollary 1 (stationary predictors), the asymptotic local power functions of the tests based on ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T,adj}(\tau_{0};\lambda_{2}^{0})$ are given by $1-\Phi(q_{\alpha}-2 \psi^{0})$ and $1-\Phi(q_{\alpha}-2 \overline{\psi})$ with $\psi^{0}$ and $\overline{\psi}$ as in ((ref)) and ((ref)). (ii) Under the assumptions of Corollary 2 (persistent predictors) we have ${\cal \bm S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})\stackrel{\cal D}\rightarrow \mathcal{N}(0,1)+\xi_{adj}^{0}$ and $\overline{\cal \bm S}_{T,adj}(\tau_{0};\lambda_{2}^{0})\stackrel{\cal D}\rightarrow \mathcal{N}(0,1)+ \overline{\xi}_{adj}$ where
with $\xi^{0}$ and $\overline{\xi}$ as in ((ref)) and ((ref)). } \\
Proposition 5 essentially implies that all our results regarding the null limiting distributions of ${\cal S}_{T}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{\cal S}_{T}(\tau_{0};\lambda_{2}^{0})$ and their general power properties (consistency and detectability of local departure from the null) continue to hold for their adjusted counterparts while Corollary 3 documents important differences in their specific non-centrality terms.
Indeed, the results in Corollary 3 are particularly interesting and useful for the practical assessment of the power properties of the adjusted versus unadjusted statistics. We have an environment whereby the limiting distributions of the two types of test statistics are the same under the null hypothesis while their non-centrality parameters differ under the alternative, pointing to a more favorable behavior for the adjusted statistics when it comes to detecting local departures from the null.
Letting $\psi^{0}_{adj}$ and $\overline{\psi}_{adj}$ denote the non-centrality parameters associated with the adjusted statistics, Corollary 3(i) establishes that in a stationary context we have $\psi^{0}_{adj}=2 \psi^{0}$ and $\overline{\psi}_{adj}=2\overline{\psi}$ so that $\psi^{0}_{adj}/\psi_{0}=2$ and $\overline{\psi}_{adj}/\overline{\psi}=2$. In the case of persistent predictors the comparison between $\xi^{0}$ and $\xi^{0}_{adj}$ and between $\overline{\xi}$ and $\overline{\xi}_{adj}$ also indicates that the adjusted quantities will stochastically dominate their non-adjusted counterparts in the sense that $P[\xi^{0}_{adj}>q]\geq P[\xi^{0}>q]$ and $P[\overline{\xi}_{adj}>q]\geq P[\overline{\xi}>q]$ for some given critical value $q$ and this is again expected to translate into more favorable power outcomes for the adjusted statistics under persistent predictors as well.
In this section we investigate the size and power properties of ${\cal S}_{T}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T}(\tau_{0};\lambda_{2}^{0})$ together with their adjusted versions across two DGPs calibrated to commonly encountered applications and sample sizes in macroeconomics and finance. The experiments are designed to emphasise the role of the pairs ($\lambda_{1}^{0}, \lambda_{2}^{0}$) and $(\tau_{0};\lambda_{2}^{0})$ on inferences with the choice of their magnitudes guided by the analysis surrounding our results in Corollaries 1 and 2.
More specifically, the implementation of ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ is restricted to $\lambda_{1}^{0}=1$ across $\lambda_{2}^{0} \in \{0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95\}$ and similarly for ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$, thus providing a very broad coverage across a range of user-inputs. The average based statistic $\overline{{\cal S}}_{T}(\tau_{0};{\lambda}_{2}^{0})$ and its adjusted version $\overline{{\cal S}}_{T,adj}(\tau_{0};{\lambda}_{2}^{0})$ are implemented for $\tau_{0}\in \{0.5, 0.8\}$ across ${\lambda}_{2}^{0} \in \{0.50, 0.60, 0.70, 0.75, 0.80, 0.85, 0.90, 0.95, 1.00\}$.
All our size and power simulations below set $\pi_{0}=0.25$ (i.e., $k_{0}=[T \ 0.25]$) to initiate the recursively generated forecasts.
A specification that mimics a frequently encountered setting in the asset pricing literature is one where the null model is the martingale difference sequence, $y_{t+1}=u_{t+1}$, and the larger model the single predictor based predictive regression, $y_{t+1}=\beta x_{t}+u_{t+1}$ with $x_{t}=\phi_{1} x_{t-1}+v_{t}$. Letting $\Sigma=\{\{\sigma^{2}_{u},\rho_{uv}\sigma_{u}\sigma_{v}\},\{\rho_{uv}\sigma_{u}\sigma_{v},\sigma^{2}_{v}\}\}$ denote the covariance of $(u_{t},v_{t})$, in line with commonly encountered magnitudes from the equity premium predictability literature we set $\sigma^{2}_{u}=3$, $\sigma^{2}_{v}=0.01$, $\rho_{uv}=-0.8$ and experiment with $\phi_{1} \in \{0.75,0.95, 0.98\}$. The conditionally homoskedastic setting takes $(u_{t},v_{t})\sim NID(0,\Sigma)$ while conditional heteroskedasticity is modelled via an ARCH(1) specification, writing $u_{t}=\epsilon_{t}\sqrt{h_{t}}$ with $h_{t}=\alpha_{0}+\alpha_{1}u_{t-1}^{2}$ and $\epsilon_{t} \sim NID(0,1)$. This latter choice naturally influences the magnitudes of $\sigma^{2}_{u}$ and $\rho_{uv}$ chosen above and we parameterise $\{\alpha_{0},\alpha_{1}\}$ in a way that maintains the same magnitude for $\sigma^{2}_{u}$ as in the conditionally homoskedastic case i.e. $\alpha_{0}/(1-\alpha_{1})=\sigma^{2}_{u}$. For this purpose we set $(\alpha_{0},\alpha_{1})=(1.8, 0.4)$ throughout.
Size experiments set $\beta=0$ while for the power properties of the tests we fix the sample size at $T=500$ and evaluate correct decision frequencies as $\beta$ moves away from the null with $\beta \in \{0,-1.5,-1.75,-2.0,-2.25, -2.5,-3,-3.5\}$. For $\beta=\gamma/T^{1/4}$ this is equivalent to $|\gamma|$ increasing with $\gamma \in \{0,-7.1,-8.3,-9.5,-10.6,-11.8,-14.2,-16.6\}$. Lastly, all of the above experiments are conducted using two alternative estimators for $\sigma$. The first one denoted $\hat{\sigma}^{2}_{hom}$ is suitable under conditional homoskedasticity while the second one denoted $\hat{\sigma}^{2}_{nw}$ is its robustified version \`{a} la Newey-West. Both estimators are based on the residuals from the model estimated under the alternative.
As our Monte-Carlo simulations encompass a very broad range of scenarios and test statistic parameterisations we provide an extensive selection of outcomes in a supplementary online appendix accompanying this paper. Our focus below is on a selection of key size/power results under conditional homoskedasticity and test statistic parameterisations that mainly rely on our recommendations based on our theoretical local power analysis above. \\
{\it Empirical Size} \\
Table (ref) presents size estimates for ${\cal S}_{T}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ and ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ across a broad range of parameterisations in a conditionally homoskedastic setting. For both test statistics we note good to excellent matches of the nominal size of 10% across almost all choices of $\lambda_{2}^{0}$ for $T\geq 500$. The adjusted statistic ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ in particular has empirical sizes that almost perfectly match 10% for virtually all magnitudes of $\lambda_{2}^{0}$. Under $\phi_{1}=0.75$ for instance, ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0}=1,{\lambda_{2}^{0}=0.8})$ has resulted in an empirical size of 10.4% for $T=1000$ and 11.0% for $T=500$. The corresponding figures for $\phi_{1}=0.95$ were 10.6% and 11.2% respectively, thus also highlighting the robustness of the test statistics to the degree of persistence of the predictors as expected from our results in Propositions 1-2. Similar outcomes also characterise ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0}=1,{\lambda_{2}^{0}=0.9})$ suggesting that these test statistics maintain good size control for $T\geq 500$ even when $\lambda_{2}^{0}$ is as large as 0.90 or 0.95 and $\lambda_{1}^{0}$ is set equal to one.
Regarding the size properties of the unadjusted ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ statistic we note a mild undersizeness for magnitudes of $\lambda_{2}^{0}$ that are in the vicinity of 1, with its empirical sizes clustered around 7-8%. Overall the outcomes in Table (ref) have highlighted remarkably stable size properties for both the unadjusted and adjusted statistics across the different magnitudes of $\lambda_{2}^{0}$ including when it is set at 0.90 or 0.95. This is particularly reassuring given our earlier theoretical power analysis which pointed at desirable parameterisations that satisfy $\lambda_{1}^{0}\approx \lambda_{2}^{0}$ with both $\lambda_{1}^{0}$ and $\lambda_{2}^{0}$ set in the vicinity of 1 in the practical implementation of ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$.
Before proceeding further it is also useful to briefly rationalise the size behavior of these two statistics when $\lambda_{2}^{0}$ is chosen to lie almost at its boundary as when we set $\lambda_{2}^{0}=0.95$. In such instances we noted the mild undersizeness of ${\cal S}_{T}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ and mild oversizeness of ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ when operating with small to moderately sized samples. A magnitude of $\lambda_{2}^{0}$ that is close to 1 essentially translates into more `MSE content' from the larger model and hence a greater exposure to estimation noise when the null model holds true. This results in the unadjusted ${\cal S}_{T}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ test statistic's distribution being pushed leftward with fewer than expected rejections of the null. On the other hand the adjusted statistic which aims to correct for estimation error via ((ref)) sees its correction factor's contribution increase as $\lambda_{2}^{0}\rightarrow 1$, a correction factor that is overly inflated in small samples. At this stage it is also useful to point out that the effective sample size is given by $T-k_{0}$ so that under $\pi_{0}=0.25$ and $T=250$ we have only about 188 data points when implementing the tests. A highly persistent predictor combined with such a small sample size can be seen to result in some degree of oversizeness for $S_{T,adj}(\lambda_{1}^{0},\lambda_{2}^{0})$ based inferences when $|\lambda_{1}^{0}-\lambda_{2}^{0}|$ is particularly small (e.g., for $(\lambda_{1}^{0},\lambda_{2}^{0})=(1,0.95)$). Nevertheless, these finite sample distortions quickly fade away as we increase the sample size to $T=500$.
For comparison purposes the last column of Table (ref) also includes the corresponding size estimates for the DM and CW statistics. These conform with the consensus view that the DM statistic is severely undersized under such nested settings while the CW statistics' empirical sizes are clustered around 5% for a nominal size of 10%, in line with the simulation results in Clark and West (2007).
We next consider the finite sample size properties of the average based statistics $\overline{{\cal S}}_{T}(\tau_{0};\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T,adj}(\tau_{0};\lambda_{2}^{0})$. Recall that the averaging is performed across a portion of the null model's MSE as captured by $\tau_{0}$ and for a given fraction of the second model's MSE $\lambda_{2}^{0}$. Here we present outcomes obtained under $\tau_{0}=0.8$ which only sum across the larger magnitudes of $\lambda_{1}$. Such a choice is theoretically justified by our earlier power analysis with further scenarios presented in the supplementary appendix. Results are presented in Table (ref) from which we note that the adjusted statistic $\overline{{\cal S}}_{T,adj}({\tau_{0}=0.8};\lambda_{2}^{0})$ displays good to excellent size control (e.g. empirical size estimates near 10% under $\lambda_{2}^{0}=1$) under moderate to large sample size choices.
An exception to this is when $\lambda_{2}^{0}\approx 0.5 \ \tau_{0}+0.5 (\approx 0.9 \ here)$ under which it shows a tendency to overreject the null hypothesis in smaller samples. This is in complete agreement with our earlier theoretical power analysis where we showed that holding all else constant the power of the test statistic must peak under $\lambda_{2}^{0}=0.5 \tau_{0}+0.5$. Thus the empirical sizes peaking for $\lambda_{2}^{0}$ in the vicinity of $0.5(0.8)+0.5=0.90$ highlight the size vs power trade-off that will characterise this average based test statistic.
Regarding the unadjusted statistic $\overline{{\cal S}}_{T}({\tau_{0}=0.8};\lambda_{2}^{0})$ we can note a tendency to underreject (e.g. empirical sizes in the vicinity of 7% under $\tau_{0}=0.8$) and this undersizeness deteriorating as $\lambda_{2}^{0} \rightarrow 1$ and $\phi_{1}$ gets closer to 1. This behavior conforms with the intuition that estimation noise caused by the estimation of parameters that are zero in the population pushes the test statistic too much to the left, a feature that was the key motivation behind Clark and West's adjustment to the DM statistic. Note for instance that these distortions are substantially dampened when the test statistic is implemented with smaller magnitudes of $\lambda_{2}^{0}$ for which it shows good to excellent size control.\\
{\it Empirical Power} \\
Table (ref) presents empirical power estimates for ${\cal S}_{T}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ and ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ across $\lambda_{2}^{0}\in \{0.80, 0.85, 0.90, 0.95\}$. The sample size is fixed at $T=500$ and power is evaluated as the DGP moves further away from the null hypothesis. The choices of $\lambda_{2}^{0}$ are dictated by our theoretical results in Corrollaries 1-2 which pointed to magnitudes satisfying $\lambda_{1}^{0}\approx \lambda_{2}^{0}\approx 1$. For both test statistics we note the tendency of their empirical power to converge to 1 as $|\gamma|$ is allowed to increase. We can also clearly observe the particularly favorable impact that the degree of persistence of predictors has on power. As expected from our findings in Propositions 3-4 and their corollaries, power improves as $\lambda_{2}^{0}\rightarrow 1$ and as $\phi_{1}\rightarrow 1$.
Given the good overall size control displayed by ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ Table (ref) clearly highlights the benefits of basing inferences on this adjusted version of our first test statistic calibrated to $\lambda_{1}^{0}=1$ and $\lambda_{2}^{0}\approx 0.9$, also noting that its performance improves considerably as $\phi_{1}\rightarrow 1$. For $\beta_{2}=-2$ for instance it displays power in the vicinity of 70%-80% under $\phi_{1}=0.75$ and 100% under $\phi_{1}=0.95$ or $\phi_{1}=0.98$. It is here important to relate our simulation outcomes with the theoretical results of Corollary 3 where we established an ARE of 2 for the adjusted statistic ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ relative to its unadjusted counterpart. This theoretical power enhancement is clearly supported by the empirical power estimates of Table (ref). Compare for instance the empirical power of 51.2% for the unadjusted statistic under $\phi_{1}=0.75$ and $\beta=-2.25$ with 81.2% for its adjusted version, a power gain of 30 percentage points.
At this stage it is also important to recall that our theoretical analysis based on the asymptotic relative efficiency of ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ versus $\overline{{\cal S}}_{T,adj}({\tau_{0}};\lambda_{2}^{0})$ clearly pointed to the potentially superior power performance of the average based statistic, for larger magnitudes of $\tau_{0}$ in particular. This is clearly corroborated by the comparison between the power outcomes in Table (ref) and Table (ref) with the latter presenting power outcomes for the $\overline{\cal S}_{T}(\tau_{0}=0.8;\lambda_{2}^{0})$ and $\overline{\cal S}_{T,adj}(\tau_{0}=0.8;\lambda_{2}^{0})$ statistics. Focusing on the “optimal” choice of $\lambda_{2}^{0}=0.5 \tau_{0}+0.5=0.90$ when $\tau_{0}=0.8$ we note from Table (ref) that $\overline{{\cal S}}_{T,adj}({\bm \tau_{0}=0.8};\lambda_{2}^{0}=0.9)$ clearly dominates all configurations of ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ in terms of its power properties, typically resulting in relative power gains in excess of 10 percentage points.
Before proceeding further it is also useful to comment on the power behavior of the DM and CW statistics in comparison to $\overline{{\cal S}}_{T}({\bm \tau_{0}=0.8};\lambda_{2}^{0}=0.90)$. Despite being severely undersized and theoretically unsuitable in the present nested context we note that the DM statistic does show a reasonable ability to detect departures from the null. However we can also observe that it is uniformly dominated by $\overline{{\cal S}}_{T,adj}({\bm \tau_{0}=0.8};\lambda_{2}^{0}=0.90)$ which under $\phi_{1}=0.75$ for instance exceeds its power by about ten percentage points. Comparing the power performance of $\overline{{\cal S}}_{T,adj}({\bm \tau_{0}=0.8};\lambda_{2}^{0}=0.90)$ with that of the CW statistic we note that these two test statistics display very similar power outcomes across most scenarios. Although the CW statistic does not have a well defined limiting distribution due to the nestedness of the competing models it appears to display reasonably good power properties across the DGPs we have considered, despite being far-off the standard normal distribution under the null (as implied by its size properties).
The second DGP allows for multiple predictors and is calibrated to mimic US inflation based predictive regressions as considered in sw2010 and ghm2014. We use the same setting as in ghm2014 and consider a DGP given by $y_{t+1}=\mu+\rho y_{t}+\beta_{1}x_{1,t}+\beta_{2}x_{2,t}+\beta_{3}x_{3,t}+u_{t+1}$ with $\mu=1$ and $\rho=0.25$. The predictors ${\bm x}_{t}=(x_{1,t},x_{2,t},x_{3,t})'$ follow the VAR(1) process $\bm x_{t}= \Phi \ \bm x_{t-1}+\bm v_{t}$ with ${\Phi}=\{\{0.6,0.1,0\},\{0.6,0.25,0\},\{0,0,0.9\}\}$ thus encompassing both persistent and much noisier processes while also being interdependent. The conditionally homoskedastic scenario takes $(u_{t},v_{1,t},v_{2,t},v_{3,t})'\sim NID(0,I_{4})$ while conditional heteroskedasticity is captured via an ARCH(1) process for $u_{t}$ as in DGP1 with $\alpha_{0}=0.6$ and $\alpha_{1}=0.4$ so that its unconditional variance matches unity as in the homoskedastic scenario.
For our size experiments we set $\beta_{1}=\beta_{2}=\beta_{3}=0$ and our power analysis focuses on alternatives to $\beta_{1}=\beta_{2}=\beta_{3}=0$ by fixing $(\beta_{1},\beta_{2},\beta_{3})=(0.15,0.15,-0.15)$ and evaluating rejection rates of the null hypothesis for $T=250, 500, 1000$. \\
{\it Empirical Size} \\
Tables (ref) and (ref) present empirical size estimates corresponding to the null DGP under $\beta_{1}=\beta_{2}=\beta_{3}=0$ for ${\cal S}_{T}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T}(\tau_{0}=0.80;\lambda_{2}^{0})$ respectively together with their adjusted versions. As DGP2 contains a larger number of predictors than DGP1 we expect the impact of estimation error on the MSE of the second/larger model to be more pronounced under the null. This is indeed corroborated by the size estimates in Table (ref) where we note the undersizeness of the unadjusted ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ statistic which is biased downward and thus results in too few rejections of the null (e.g. 6.4% under $T=1000$ and $\lambda_{2}^{0}=0.9$ versus a nominal size of 10%). Furthermore, its undersizeness tends to deteriorate for larger magnitudes of $\lambda_{2}^{0}$ as this translates into an increased influence of the larger model's MSE.
The adjusted version of the test statistic ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ on the other hand is quite effective in adjusting for estimation noise (e.g. the earlier empirical size of 6.4% is now pushed up to 12%), while also showing a tendency to “over-adjust” in small to moderately size samples, in particular for larger magnitudes of $\lambda_{2}^{0}$. Table (ref) presents the corresponding size estimates for the average based statistic $\overline{{\cal S}}_{T,adj}({\bm \tau_{0}=0.8};\lambda_{2}^{0})$. We note that this latter statistic maintains good to excellent size control for moderate sample sizes across all magnitudes of $\lambda_{2}^{0}$ but requires larger samples when $\lambda_{2}^{0}$ is set near one. \\
{\it Empirical Power} \\
For this DGP our power experiments focus on documenting the rejection frequencies of the null hypothesis for a fixed alternative as the sample size is allowed to increase. Results are presented in Tables (ref)-(ref) for ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T}({\bm \tau_{0}=0.8};\lambda_{2}^{0})$ and their adjusted versions. Under either $T=500$ or $T=1000$ all test statistics have powers at or near 100%.
Tables (ref)-(ref) also clearly corroborate our theoretical power analysis that highlighted peaking powers under $\lambda_{2}^{0}=0.5 \tau_{0}+0.5$. Focusing on $\overline{{\cal S}}_{T}({\bm \tau_{0}=0.8};\lambda_{2}^{0})$ we can clearly observe the empirical powers to be largest under $\lambda_{2}^{0}=0.9$ for all sample sizes (e.g. 98.2% versus 95.7% when $\lambda_{2}^{0}=1$ and for $T=250$). \\
The main findings from our simulation experiments can be summarised as follows. (i) The adjusted versions of the two test statistics ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T,adj}({\bm \tau_{0}};\lambda_{2}^{0})$ displayed good to excellent size control across most parameterisations of their respective inputs ($\lambda_{1}^{0}, \lambda_{2}^{0})$ and ($\tau_{0} ,\lambda_{2}^{0})$. (ii) Both test statistics are consistent and have non-trivial local asymptotic power while their finite sample power properties are strongly influenced by the respective magnitudes of those same inputs. The guidelines provided by our theoretical power analysis do however lead to highly favorable power outcomes with implementations such as ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0}\approx 0.9)$ and $\overline{{\cal S}}_{T,adj}({\bm \tau_{0} \approx 0.8};\lambda_{2}^{0}\approx 0.9)$ standing out in terms of their size/power trade-offs, especially for moderately sized samples such as $T \geq 500$. (iii) The proposed methods are valid irrespective of the degree of persistence of the predictors as also corroborated by our finite sample simulations. (iv) Our Monte-Carlo analysis did show that Clark and West's CW statistic which although not grounded on formal standard normal asymptotics also performed particularly well in terms of its power, despite relatively important size distortions. Strictly speaking the CW statistic has been introduced for handling nested models estimated via a rolling as opposed to a recursive approach since from a theoretical standpoint it continues to suffer from the variance degeneracy problem characterising DM type constructions.
The above simulation based outcomes combined with our earlier local power analysis point to precise guidelines for the choice of tuning parameters required in the implementation of our test statistics. For the ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ statistic we argued that $\lambda_{1}^{0}$ and $\lambda_{2}^{0}$ should be set near their boundary of one and in close vicinity of one another. Our simulations based on ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ for $\lambda_{2}^{0}$ set in the 0.80-0.95 range have indeed resulted in good to excellent size-power tradeoffs and good to excellent size control. The robustness of the empirical size outcomes to a much broader range of the $\lambda_{2}^{0}$ magnitudes is also noteworthy as illustrated by the outcomes in Tables 1 and 5. \\
Regarding the $\overline{\cal S}_{T,adj}(\tau_{0};\lambda_{2}^{0})$ statistic, our theoretical local power analysis led us to argue for $\tau_{0}$ to be set in the vicinity of unity and $\lambda_{2}^{0}$ as $\lambda_{2}^{0}=0.5 \ \tau_{0}+0.5$. Simulations based on $\tau_{0}=0.8$ did indeed result in good to excellent size and power properties. As the above choice for $\lambda_{2}^{0}$ is a maximiser of local power it is perhaps natural to expect some size distortions in smaller samples when $\lambda_{2}^{0}$ is set in this way and in particular when this is combined with the presence of highly persistent predictors. This is indeed confirmed by our size experiments implemented under $(\tau_{0};\lambda_{2}^{0})=(0.8,0.9)$ and a local-to-unity type predictor. Nevertheless, these specific finite sample distortions can also be seen to progressively vanish as the sample size is allowed to grow.
We illustrate the implementation of our proposed methods by revisiting a widely considered puzzle in the international economics literature, namely the random walk like behavior of exchange rates. Our goal is to use our new test statistics in order to evaluate whether past exchange rate levels have any predictive power for subsequent exchange rate changes. Letting $s_{t}$ denote the log of the spot exchange rate, we compare the out of sample predictive accuracy of the larger model $\Delta s_{t+1} = \alpha + \beta \ s_{t}+u_{t+1}$ (model 2) with the random walk with drift specification $\Delta s_{t+1} = \alpha+u_{t+1}$ (model 1).
We consider six major currencies (EUR, YEN, GBP, CHF, AUD and CAD) and implement our tests on daily spot rates spanning the period between 1999/01/04 and 2021/07/16, sourced from the Saint-Louis Fred database. An important advantage of the methods developed in this paper is their robustness to the persistence properties of predictors which is particularly relevant when considering exchange rate series. Indeed, for all three daily series we have considered the first order autocorrelation coefficient from an AR(1) fit is 0.99.
Predictive accuracy testing outcomes (p-values) based on ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{{\cal S}}_{T,adj}(\tau_{0};\lambda_{2}^{0})$ are presented in Table (ref) where we used $\pi_{0}=0.5$ to initiate the expanding window estimation (i.e., starting from the middle of the sample). For robustness considerations inferences based on ${\cal S}_{T,adj}(\lambda_{1}^{0}=1,\lambda_{2}^{0})$ are implemented across $\lambda_{2}^{0}=\{0.80,0.85,0.90,0.95\}$ while for $\overline{\cal S}_{T,adj}(\tau_{0};\lambda_{2}^{0}=1)$ we consider $\tau_{0}=0.80$ and $\tau_{0}=0.90$. Looking at the top and middle panels of Table (ref) we note that results unanimously corroborate the fact that the level of exchange rates does not have any meaningful forecasting power for future currency returns over the period considered and across all major currencies. This result based on the use of daily data also corroborates the recent findings in ew2021 based on monthly frequencies. The bottom panel of Table (ref) displays the p-values associated with the standard DM and the CW statistics. It is here interesting to point out that inferences based on the standard CW statistic lead to a rejection of the random walk specification for the EURUSD and CADUSD series when implementing the test at a 5% level or above which is in sharp contrast with the large p-values obtained using our two test statistics.
The main motivation of this paper was to provide a way of bypassing the variance degeneracy problem that arises in the context of out-of-sample nested model comparisons. We did so by developing two new test statistics shown to have nuisance parameter free standard normal asymptotics and good power properties including in the close vicinity of the null hypothesis. Our proposed inferences can trivially accommodate conditional heteroskedasticity and are also shown to be robust to the presence of highly persistent predictors. Although our power analysis has ruled out the case of deterministic trends via Assumption B1 for instance, these can also be easily accommodated within our framework without any changes to the implementation of the tests provided that the trends are formulated in a scaled form as $(t/T)$ and its powers. Nested comparisons in purely deterministic environments would be particularly relevant in areas such as temperature modelling (e.g., wuzhao2017, grg2020) or the recent literature on modelling pandemic dynamics (e.g., jiangshao2020, lilinton2021).
Although our proposed test statistics require two user inputs each, both have been shown to display good to excellent size control across a very broad range of such parameterisations in a multitude of empirically relevant settings. Although these user-inputs do have considerable influence on the finite sample power properties of both test statistics their choices can be accurately guided by examining the power functions associated with each test statistic, as demonstrated by our simulations and their theoretical backing.
It is important to recognise that our proposed test statistics do involve the discarding of some information albeit very limited and with its amount under the control of the user. As a result some power loss is of course unavoidable but the absence of any alternative approach that uses more information while achieving the same purpose in an environment that can accommodate both stationary and persistent predictors as well as conditional heteroskedasticity makes such power losses only notional. Our simulation results have indeed shown that very little needs to be discarded for our methods to work well and to provide reliable inferences. In this sense they are not subject to the disadvantages of sample splitting based techniques used for instance in the goodness of fit literature (e.g. half-sample methods).
The principles underlying our proposed inferences based on ((ref)) should also be portable beyond out of sample forecasting considerations to areas involving model selection testing \`{a} la v1989 where nestedness versus non-nestedness or the overlapping nature of models being compared influences test procedures due to variance degeneracy problems (see also s2015). In sw2017 for instance, the authors developed a model selection test for choosing between two parametric likelihoods based on sample splitting principles which although different from our approach based on MSE comparisons on overlapping intervals was driven by similar concerns. Adapting the analysis of this paper to such model selection testing contexts is a promising avenue currently being explored.
\setcounter{table}{0}
{ {\bf PROOF OF PROPOSITION 1}. We consider the asymptotic behavior of $Z_{T}(\ell_{1},\ell_{2})$ in ((ref)). Rescaling the time axis we write $Z_{T}(\lambda_{1},\lambda_{2})\equiv Z_{T}([(T-k_{0})\lambda_{1}], [(T-k_{0})\lambda_{2}])$ and focus on $Z_{T}(\lambda_{1},\lambda_{2})$. Using $\hat{e}_{j,t+1}^{2}=u_{t+1}^{2}+(\hat{e}_{j,t+1}^{2}-u_{t+1}^{2})$ ($j=1,2$) in ((ref)) yields
From Assumption A(i) we have $\sup_{\lambda_{1}}|{\cal N}_{2T}(\lambda_{1})|=o_{p}(1)$ and $\sup_{\lambda_{2}}|{\cal N}_{3T}(\lambda_{2})|=o_{p}(1)$. Combining with
and
gives
It is now convenient to reformulate ((ref)) as
and note that the second component in the right hand side of ((ref)) is $O(1/\sqrt{T-k_{0}})$. We now recall that our setting operates under fixed and given magnitudes of $(\lambda_{1},\lambda_{2})$, say $(\lambda_{1}^{0},\lambda_{2}^{0})$ chosen such that $(\lambda_{1}^{0},\lambda_{2}^{0}) \in \Lambda^{0}$. We have
It follows from Assumptions A(ii)-A(iii), the continuous mapping theorem and Slutsky's theorem that
The right hand side of ((ref)) is a centered Gaussian random variable with variance $v^{0}(\lambda_{1}^{0},\lambda_{2}^{0})=|\lambda_{1}^{0}-\lambda_{2}^{0}|/\lambda_{1}^{0}\lambda_{2}^{0}$ as stated in ((ref)). Specifically, the statement in ((ref)) is equivalent to $Z_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0}) \stackrel{\cal D}{\rightarrow} N(0,|\lambda_{1}^{0}-\lambda_{2}^{0}|/\lambda_{1}^{0}\lambda_{2}^{0})$. This also establishes that ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})\equiv Z_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})/\sqrt{v^{0}(\lambda_{1}^{0},\lambda_{2}^{0})}\stackrel{\cal D}{\rightarrow} N(0,1)$. \qed \\
{\bf PROOF OF PROPOSITION 2}. We view $Z_{T}(\lambda_{1},\lambda_{2})$ in ((ref)) as a functional of $\lambda_{1}$ whose range is determined by the choice of $\tau_{0}$, and for a given $\lambda_{2}=\lambda_{2}^{0}$, satisfying $(\tau_{0},\lambda_{2}^{0})\in \overline{\Lambda}^{0}$. Assumption A(ii) combined with standard continuous mapping arguments applied to ((ref)) yields
The asymptotic behavior of $\overline{Z}_{T}(\tau_{0};\lambda_{2}^{0})$ in ((ref)) now follows by appealing to Assumptions A(i)-(iii), the continuity of the average operation and ((ref)). Specifically,
Note that
so that ((ref)) is well defined almost surely and by construction centered Gaussian. It now suffices to obtain its variance. We have
where we appealed to Fubini's Theorem for interchanging expectations with integration in the second row of ((ref)). Standard integral calculus now leads to ((ref))-((ref)). \qed \\
The following lemma collects some key results used in the proofs of Proposition 3 and Corollary 1 on the power properties of the proposed tests under stationarity. \\
{\bf LEMMA A1}. Suppose model ((ref)) holds with ${\bm \beta}_{2}={\bm \gamma}/T^{1/4}$. Under Assumption B1 and as $T \rightarrow \infty$ we have
{\bf PROOF OF LEMMA A1}. (i) As we operate under model ((ref)) with ${\bm \beta}_{2}={\bm \gamma}/T^{1/4}$ we have $\hat{e}_{1,t+1}-u_{t+1}={\bm x'}_{2,t}{\bm \beta}_{2}-{\bm x'}_{1,t} (\hat{\bm \delta}_{1,t}-{\bm \beta}_{1})$ so that the following identity holds
where
with $t=k_{0},\ldots,k_{0}-1+[(T-k_{0})\lambda]$ in all of the above summations and below, unless otherwise indicated. \\
For $A_{1T}(\lambda)$, we write
and as $|\sqrt{(T-[T\pi_{0}])/T}-\sqrt{1-\pi_{0}}|=o(1)$ we have
so that Assumption B1(i) directly implies
Before focusing on the remainder quantities we consider the limiting behavior of $(\hat{\bm \delta}_{1,t}-{\bm \beta}_{1})$. Setting $t=[Ts]$ we write
For the second term in the right hand side of ((ref)) we have
due to Assumptions B1(i)-(ii). For the first term in the right hand side of ((ref)) we can write
{
}
so that Assumptions B1(i)-(ii) also ensure that
Combining ((ref)) and ((ref)) and using the triangle inequality in ((ref)) yields
We now focus on $A_{2T}(\lambda)$. Using suitable normalisations and appealing to ((ref)), we can express $A_{2T}(\lambda)$ as
Assumption B1(i) combined with the result in ((ref)) give
For $A_{3T}(\lambda)$ we write
so that using ((ref)), Assumption B1(i) and the triangle inequality yields
Next, as an immediate consequence of Assumption B1(ii) we have
Finally, using ((ref)) together with Assumption B1(ii) yields
Combining ((ref)), ((ref)) and ((ref))-((ref)) with successive uses of the triangle inequality yields the stated result in Lemma A1(i). The statement in ((ref)) follows an identical line of argument as above and details are therefore omitted from the exposition here. \\
{\bf PROOF OF PROPOSITION 3}. (i) We initially consider the case of a fixed and non-zero ${\bm \beta}_{2}$ and establish that ${\cal S}^{0}_{T}(\lambda_{1}^{0},\lambda_{2}^{0})\stackrel{p}\rightarrow \infty$. Using ((ref)) and appealing to Lemma A1(ii) we have
We can now note from ((ref))-((ref)) that the first term in the right hand side of ((ref)) is $O_{p}(T^{-1/2})$ so that
It is now straightforward to adapt the result in Lemma A1(i) to a fixed ${\bm \beta}_{2}$ setting and infer that
yielding (for fixed ${\bm \beta}_{2}$)
where we also made use of Assumption B1(iii) ensuring that $\hat{\sigma}\stackrel{p}\rightarrow \sigma \in (0, \infty)$. It now follows that
with $v^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ given by ((ref)), thus leading to ${\cal S}_{T}^{0}(\lambda_{1}^{0},\lambda_{2}^{0}) \stackrel{p}\rightarrow \infty$ as stated. Proceeding similarly for $\overline{Z}_{T}(\tau_{0};\lambda_{2}^{0})$ we have
with $\overline{v}(\tau_{0};\lambda_{2}^{0})$ as in ((ref))-((ref)), thus also establishing that $\overline{\cal S}_{T}(\tau_{0};\lambda_{2}^{0})\stackrel{p}\rightarrow \infty$. \\
(ii) We next focus on the local asymptotic behavior of the two test statistics with ${\bm \beta}_{2}$ parameterised as ${\bm \beta}_{2}={\bm \gamma}/T^{1/4}$. Using ((ref)) in conjunction with Lemma A1(i)-(ii) we have
It now follows directly from ((ref)) in Lemma A1, Assumption B1(iii) and Slutsky's theorem that
as required. The result for $\overline{Z}_{T}(\tau_{0};\lambda_{2}^{0})$ follows identical arguments and is therefore omitted. \qed \\
{\bf PROOF OF COROLLARY 1}. Follows directly from ((ref))-((ref)) in the proof of Proposition 3. \qed \\
{\bf LEMMA A2}.
{\bf PROOF OF LEMMA A2}. For ((ref)) we have
due to Assumption B2(i). For ((ref)) we have
which also follows from Assumption B2(i). For ((ref)) we write
From ((ref)) it also follows that
and the statement in ((ref)) follows directly using ((ref)) in ((ref)) and appealing to the continuous mapping theorem. \\
{\bf LEMMA A3}.
{\bf PROOF OF LEMMA A3}. We consider ((ref)) first. We operate under ${\bm \beta}_{2}={\bm \gamma}/T^{3/4}$. Recalling that $\hat{e}_{1,t+1}-u_{t+1}={\bm x}_{2,t}'{\bm \beta}_{2}-{\bm x}_{1,t}'(\hat{\bm \delta}_{1,t}-{\bm \beta}_{1})$ and using $\lim_{T\rightarrow \infty} ((T-k_{0})/T)^{j} \rightarrow (1-\pi_{0})^{j}$, we write
It next follows from ((ref))-((ref)) that the last two terms in the right hand side of ((ref)) are $O_{p}(T^{-1/4})$ so that we also have
Using ((ref)) and ((ref)) from Lemma A2 together with the continuous mapping theorem, ((ref)) leads to the required result in ((ref)). The result in ((ref)) is established following similar arguments and details are omitted. \qed \\
{\bf PROOF OF PROPOSITION 4 and COROLLARY 2}. We focus on Part (ii) of the Proposition as the test consistency property stated in Part (i) follows as its direct consequence. From ((ref)) and Lemma A2, under the local alternative ${\bm \beta}_{2}={\bm \gamma}/T^{3/4}$ we have
using ((ref)), Slutsky and the continuous mapping theorems (note that the standard normality of the first component in the right hand side of ((ref)) has been established in Proposition 1). It now follows directly from ((ref)) that $\lim_{||\gamma||\rightarrow \infty}\lim_{T\rightarrow \infty} {\cal S}_{T}(\lambda_{1}^{0},\lambda_{2}^{0})$ as required. The result for $\overline{\cal S}_{T}(\tau_{0};\lambda_{2}^{0})$ follows identical lines and its details are omitted. \qed \\
{\bf PROOF OF PROPOSITION 5 and COROLLARY 3}. We have $\hat{e}_{1,t+1}^{2}-\tilde{e}_{2,t+1}^{2}=(\hat{e}_{1,t+1}^{2}- \hat{e}_{2,t+1}^{2})+(\hat{e}_{1,t+1}-\hat{e}_{2,t+1})^{2}$ which leads to the formulations of ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})$ and $\overline{\cal S}_{T,adj}(\tau_{0};\lambda_{2}^{0})$ in ((ref)) and ((ref)) respectively. Under the null hypothesis and for both test statistics the result follows by verifying that
Noting that
the statement in ((ref)) follows directly from Assumption A(i) since we operate under the null hypothesis noting also that $(\hat{e}_{1,t+1}\hat{e}_{2,t+1}-u_{t+1}^{2})=(\hat{e}_{1,t+1}-\hat{e}_{2,t+1})\hat{e}_{2,t+1}+(\hat{e}_{2,t+1}^{2}-u_{t+1}^{2})$ from which we infer the $o_{p}(1)$'ness of the third component in the right hand side of ((ref)). It now follows that Propositions 1 and 2 continue to hold for the two adjusted statistics. \\
For the behavior of the adjusted statistics under the alternative we initially consider the case of stationary predictors and operate under ${\bm \beta}_{2}={\bm \gamma}/T^{1/4}$ as in the setting of Corollary 1. Using ((ref)) with Lemma A1 we can write
and
Using ((ref))-((ref)) it follows that ${\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})\stackrel{\cal D}\rightarrow N(2\psi^{0},1)$ and similarly for $\overline{\cal S}_{T,adj}^{0}(\lambda_{1}^{0},\lambda_{2}^{0})\stackrel{\cal D}\rightarrow N(2\overline{\psi},1)$ which establishes the fact that Proposition 3 continues to hold for the two adjusted statistics in addition to part (i) of Corollary 3. The result for the case of persistent predictors follows identical lines, making use of Lemma A3 and Proposition 3 which in turn establishes part (ii) of Corollary 3 and that Proposition 4 also holds for the two adjusted statistics. \qed }