EconBase
← Back to paper

Semiparametric Testing with Highly Persistent Predictors

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

100,261 characters · 17 sections · 80 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametric Testing with Highly Persistent Predictors

\abstract{ We address the issue of semiparametric efficiency in the bivariate regression problem with a highly persistent predictor, where the joint distribution of the innovations is regarded an infinite-dimensional nuisance parameter. Using a structural representation of the limit experiment and exploiting invariance relationships therein, we construct invariant point-optimal tests for the regression coefficient of interest. This approach naturally leads to a family of feasible tests based on the component-wise ranks of the innovations that can gain considerable power relative to existing tests under non-Gaussian innovation distributions, while behaving equivalently under Gaussianity. When an i.i.d.\ assumption on the innovations is appropriate for the data at hand, our tests exploit the efficiency gains possible. Moreover, we show by simulation that our test remains well behaved under some forms of conditional heteroskedasticity.

JEL classification: C12, C14

Keywords: predictive regression, limit experiment, LABF, maximal invariant, rank statistics.

Introduction

Over the past two decades, inference for the bivariate regression model with a highly persistent predictor has been well studied under the assumption of bivariate Gaussian innovations. Several procedures have been proposed in the econometric literature, see CavanaghElliotStock1995, CampbellYogo2005, JanssonMoreira2006, EMW2015, and MoreiraMourao2017. These inference procedures are all constructed based on the assumption of Gaussian innovations and, while their validity has been established under weaker assumptions, the asymptotic power of all these procedures cannot go beyond the Gaussian power envelope.

In the present paper we show that, when the application supports an additional assumption of serially independent innovations, sizable power gains are possible beyond the Gaussian power envelope. We establish this result by studying in detail the invariance structures that are present in the limiting experiment associated with the predictive regression model. This leads to a semiparametric power envelop which, under non-Gaussian innovation distributions, lies above the Gaussian power envelope. In that case, even without knowing the innovation distribution, our method dominates existing QMLE-based methods.

Our results precisely quantify the statistical efficiency gains from non-Gaussian innovation distributions when innovations are serially independent in predictive regression models. Under such, arguably restrictive assumption, we construct semiparametrically optimal (in a sense to be made precise later) tests. Whether in concrete applications the assumption of serially independence is warranted, is an empirical question. When it is, it can, as our results show, be exploited leading to sizable power gains (of, as Section (ref) shows, up to 30% under Student-$t_3$ innovation distributions). Symmetrically, to make an informed choice, we study the behavior of our test when the innovations are not i.i.d.\ but exhibit conditional heteroskedasticity as often found in (financial) applications. Section (ref) shows that, for the deviations studied, our test still has desirable size and power properties.

We note that our conceptual ideas reach further. We could, for instance, allow for serial dependence along the lines of ZvdAW2016 where an AR-type model on the error is imposed. Conditional heterogeneity could formally be addressed along the lines of LingMcaleer2003 where a GARCH-type structure on the error is imposed; or following Boswijk2005 where the (potentially nonstationary) volatility is estimated nonparametrically. These relaxations would technically be non-trivial and are left for future research. Note that, in view of the robustness-efficiency trade-off Muller2011, an i.i.d.\ assumption on the innovations ultimately driving the error term is not avoidable. Our test gives the empirical researchers an additional option: an improved power when innovations are i.i.d.\ and non-Gaussian.

The study of (optimal) semiparametric inference in the predictive regression model is complicated by the nonstandard asymptotic behavior induced by the local-to-unity asymptotics on the persistence parameter. More precisely, the associated likelihood ratios are of the Locally Asymptotically Brownian Functional (LABF) form in Jeganathan1995 and henceforth outside the conventional Locally Asymptotically Normality (LAN) world. As a consequence, the usual semiparametric approach based on projecting the score of the parameter of interest on the tangent space of nuisance scores is not straightforward. In particular, the model does not feature an adaptiveness property, which complicates its analysis. Jansson2008 deals with the unit root testing problem, which also admits the LABF form, by guessing and then proving a least favorable direction of parametric submodels. An alternative approach has been proposed for the unit root testing problem in ZvdAW2016 and generalized to other common types of limiting experiments in Zhou2020. In the present paper we apply these techniques to the predictive regression model.

The key idea is to exploit invariance structures in a so-called “structural” representation of the limit experiment. This approach sets us apart from most of the statistical and econometric literature where invariance arguments are used in the sequence of experiments. Instead, we obtain procedures which are invariant in the limit experiment, thereby making the analysis tractable and applicable to many models. Furthermore, the unique bivariate nature of the predictive regression model leads to a nonstandard multivariate structure in the associated limit experiment (see Theorem (ref)). Therefore, we present the approach in detail in the present paper.

Our contribution is twofold. First, we derive the semiparametric power envelope for (asymptotically) invariant tests in case the predictor's persistence level is assumed to be known, based on the structural LABF limit experiments. More precisely, Girsanov's theorem, combined with the limiting likelihood ratios for LABF experiments, leads to a description of the limit experiment by stochastic differential equations (SDEs). The observations in the limit experiment correspond to the limits of partial-sum processes of the innovations and score functions in the predictive regression model. In this structural representation of the limit experiment, we find that the nuisance parameters induced by the density function of the innovations only appear in the drifts of the driving Brownian motions. This leads to an invariance restriction by taking the Brownian bridges (which are invariant with respect to these drifts) of these processes, and allows us to remove the nonparametric nuisance parameter (the density $f$ of the innovations). We show that this also generates the maximal invariant. In this way, we avoid the problem of explicitly finding the least-favorable submodel. The likelihood of the maximal invariant immediately, by the Neyman-Pearson lemma, leads to the semiparametric power envelope.

Second, we propose a family of semiparametric feasible tests that has desirable properties. These tests are constructed using (asymptotically) sufficient statistics that are based on the increments of innovations, their component-wise ranks, and a pair of chosen marginal reference densities for both innovations including a reference correlation parameter. The ranks appear naturally as rank-based partial-sum score processes which weakly converge to the Brownian bridge that is invariant w.r.t.\ the density perturbation parameters. To further eliminate the remaining nuisance parameter, namely the predictor's persistence level, we employ the Approximate Least Favorable Distribution (ALFD) approach proposed by EMW2015. We also follow their suggestion to switch to standard asymptotic approximations when the persistence parameter is far from unity. This helps to control the size of our tests uniformly under both non-stationarity and stationarity, see Appendix (ref). The tests thus obtained are semiparametric in the sense that they have correct asymptotic sizes (under all innovation densities allowed) regardless of the choices of the marginal reference densities or the reference correlation.

Next to their uniform (relative to our model) validity, our test are more powerful than existing tests when the true innovation density is non-Gaussian. In particular, we compare our test to EMW2015 (henceforth denoted as EMW), which is based on Gaussian likelihood ratios JanssonMoreira2006. Our asymptotic analysis using invariance arguments shows that, under non-Gaussian innovations, the EMW test actually is measurable with respect to an invariant in the limit that is not maximally invariant. As a result, under non-Gaussianity, we can construct tests that outperform the Gaussian power envelope and, thus, outperform the EMW test; see Remark (ref). The power improvement depends on the choices of the marginal reference densities: when they are “closer” to the true marginal densities, we gain more power (and, again, while always having the desired size). Additionally, if one fixes the marginal reference densities to be Gaussian, our test is generally still more powerful than the EMW test under non-Gaussian innovation density; while under Gaussian innovation density, our test performs equivalently to the EMW test. This property is often referred to as the Chernoff-Savage result (see ChernoffSavage1958). In the present LABF setting we have not been able to formally prove this Chernoff-Savage result, but our simulations indicate that this property nevertheless may hold.

Our rank-based test can be regarded a generalized version of quasi-likelihood ratio tests which take the reference density to be Gaussian. The extra freedom to choose the reference density also comes with the cost of actually choosing it. However, we note that, in line with traditional quasi-likelihood methods, one can always choose the Gaussian reference density. Based on the classical Chernoff-Savage result, we conjecture that our rank-based procedure will then always outperform the quasi-likelihood procedure. This is confirmed by simulations and intuition, but, as discussed below, given the non-standard limiting experiment structure, we have not been able to prove this formally. Alternative, one could study a plug-in estimator where the reference density is nonparametrically estimated. We do not study this formally in the present paper; however, see Section (ref) for some simulation results. Similarly, one may envision an approach where one pre-tests the residuals for, e.g., high kurtosis and chooses a references density based on that pre-test result.

The paper is organized as follows. In Section (ref), we introduce the model and testing problem under consideration. In Section (ref), we develop the asymptotic power envelope for test that are (asymptotically) invariant with respect to the innovation density $f$, assuming the predictor's persistence parameter $\gamma$ is known. This development is based on the theory of limit experiments (see, e.g., LeCam1986 and vdVaart2000) and a structural version for models of LABF likelihood ratios (see ZvdAW2016). In particular, this section explains where our power gains come from, see Remark (ref). In Section (ref), we employ the ALFD approach proposed by EMW2015, among several available choices in the literature, to eliminate the nuisance parameter $\gamma$. In Section (ref), we report large- and small-sample performances of our tests under both i.i.d.\ and conditional heteroskedastic errors. Section (ref) concludes. All proofs are gathered in the appendix.

Model

Let $y_t$ denote a random variable, observable at time $t$, that we wish to predict at time $t-1$ using an observable explanatory variable $x_{t-1}$. We consider the predictive regression model

align[align omitted — 158 chars of source]

with $x_0=0$.\footnote{Note that this assumption on the initial value $x_0$ could possibly be relaxed to the weaker assumption $T^{-1/2}x_0=o_\mathrm{P}(1)$ under $\beta=0$ and $\gamma=1$. One can possibly proceed along the lines of MullerElliott2003; see also a remark on this point in Section 4 of JanssonMoreira2006. We keep the assumption $x_0=0$ for simplicity.} The parameter space is given by $\mu\in\mathbb{R}$, $\alpha\in\mathbb{R}$, $\beta\in\mathbb{R}$, and $\gamma\in(-1,1]$. We have observations available for $t=1,\dots,T$.

Equation ((ref)) features, along the lines of CavanaghElliotStock1995 and JanssonMoreira2006, an intercept $\alpha$. However, as $\mu$ is a nuisance parameter in our model, the intercept $\alpha$ can be subsumed in $\mu$ without affecting inference on $\beta$. Indeed, our test statistics will only depend on the increments of $x_t$, denoted by $\Delta x_t$, and their associated ranks and, thus, they are invariant with respect to $\alpha$. We therefore omit $\alpha$ in the rest of this paper.

To eliminate the nuisance intercept parameter $\mu$ in ((ref)), one can directly impose an invariance restriction in the sequence of predictive regression experiments. For instance, the JanssonMoreira2006 test is based on the maximal invariant statistic $(y_2- y_1, y_3 - y_1, \dots, y_T - y_1)^\prime$. In the present paper, our statistic is only based on $y_t$'s through their ranks and, thus, also enjoys finite-sample invariance w.r.t.\ $\mu$. To simplify notation, we set $\mu=0$ throughout the paper and nowhere assume $\mathrm{E}_f(\varepsilon^y_t)=0$. We will need to impose $\mathrm{E}_f(\varepsilon^x_t)=0$: allowing for deterministic trends in $x_t$ would lead to an entirely different asymptotic analysis.

Summarizing, as outlined in the introduction, we assume that the innovations $\varepsilon_t=(\varepsilon^y_t,\varepsilon^x_t)^\prime$ are independent and identically distributed (i.i.d.) with (bivariate) density $f$ satisfying the following condition.

assumption\begin{enumerate} • $\mathrm{E}_f(\varepsilon^x_t)=0$ and $\operatorname{Var}_f(\varepsilon_t)=\begin{pmatrix}\sigma_{y}^2 & \rho\sigma_{y}\sigma_{x} \\ \rho\sigma_{y}\sigma_{x} & \sigma_{x}^2\end{pmatrix}$ is a finite positive-definite matrix. • The density $f$ is absolutely continuous with a.e.\ derivative $\dot{f}=\begin{pmatrix}\dot{f}_y \\ \dot{f}_x\end{pmatrix}$. • The (standardized) Fisher information for location, \begin{align*} J_f=\begin{pmatrix}J_{f_{yy}} & J_{f_{yx}} \\ J_{f_{yx}} & J_{f_{xx}}\end{pmatrix}= \mathrm{E}_f\left(\ell_f \ell_f^\prime\right), \end{align*} where $\ell_f$ is the (standardized) location score function \begin{align*} \ell_f =\begin{pmatrix}\sigma_y\ell_{f_y} \\ \sigma_x\ell_{f_x}\end{pmatrix} =\begin{pmatrix}-\sigma_y\dot{f}_y/f \\ -\sigma_x\dot{f}_x/f\end{pmatrix}, \end{align*} is finite.\footnote{Being a Fisher information for location, $J_f$ is automatically nonsingular and positive definite, see MayerWolf1990.} • $f>0$. $\hfill \Box$ \end{enumerate}

Let $\mathfrak{F}$ denote the set of densities satisfying Assumption (ref).

The Fisher information $J_f$ and scores $\ell_f$ for location are standardized in the sense that they are actually those related to $\varepsilon^y_t/\sigma_y$ and $\varepsilon^x_t/\sigma_x$. As a result, $\ell_f$ and $J_f$ do not depend on $\sigma_y$ or $\sigma_x$. Note, however, that they both still depend on the correlation between the innovations $\varepsilon^y_t$ and $\varepsilon^x_t$, i.e., they still depend on $\rho$.

We are interested in (optimal) tests for the (composite) null hypothesis

align[align omitted — 75 chars of source]

versus the one-sided alternative

align[align omitted — 75 chars of source]

As the literature focuses on test derived using an assumed Gaussian innovation density, we will throughout this paper consider Gaussian densities as a special case. This will allow us to make explicit where the power improvements come from in the case of non-Gaussian, serially independent, innovations $(\varepsilon^y,\varepsilon^x)$.

remark[Gaussian $f$] In case $f$ is zero-mean bivariate Gaussian with correlation matrix $\mathbf{R} = \begin{pmatrix} 1 & \rho \\ \rho & 1\end{pmatrix}$, Assumption (ref) is satisfied with $\ell_f(\varepsilon^y,\varepsilon^x) = \mathbf{R}^{-1}\begin{pmatrix} \varepsilon^y/\sigma_y \\ \varepsilon^x/\sigma_x \end{pmatrix}$ and $J_f=\mathbf{R}^{-1}$.

Local perturbations

Following the by now standard approach in the literature, we study the limit experiment in the sense of H\a'{a}jek-Le Cam by considering local alternatives for all model parameters, that is, for both the parameter of interest $\beta$ and the nuisance parameters ($\gamma$ and $f$). For $\beta$ and $\gamma$ the appropriate rates of convergence are well known, see, e.g., ElliottStock1994, CampbellYogo2005, or JanssonMoreira2006. More precisely, we consider a $T^{-1}$-localization rate for $\beta$ and $\gamma$, i.e.,

align[align omitted — 140 chars of source]

with $b\in\mathbb{R}$ and $c\in(-\infty,0]$.\footnote{We use here the common approach in the literature to restrict the nuisance parameter $c$ to $(-\infty,0]$. We conjecture that all results remain valid, with the obvious modifications, in case one would choose the larger parameter space $c\in\mathbb{R}$; see, e.g., MoreiraMourao2017.} Observe that the local perturbation for $b$ features a scaling by $\sigma_y/\sigma_x$. This ensures that the limit experiment will not depend on $\sigma_y$ and $\sigma_x$ (although it still depends on $\rho$).

The nuisance parameter $f$ is infinite dimensional, so it is somewhat more involved to describe its relevant local perturbations. Introduce the separable Hilbert space

equation[equation omitted — 228 chars of source]

where $ \mathrm{L}_2^f(\mathbb{R}^2,\mathcal{B})$ denotes, the space of Borel-measurable functions $h:\,\mathbb{R}^2\to\mathbb{R}$ satisfying $\mathrm{E}_f h^2(\varepsilon)=\int_{\mathbb{R}^2}h^2(\varepsilon) f(\varepsilon)\mathrm{d}\varepsilon<\infty$. The model assumption $\mathrm{E}_f(\varepsilon^x_t)=0$ induces the restriction that local perturbations for $f$ are orthogonal to the first component of $\varepsilon$: $\mathrm{E}_f \exh(\varepsilon)=0$.

The separability of the Hilbert space $\mathrm{L}_2^{0,f}$ ensures the existence of a countable orthonormal basis $h_k$, $k\in\mathbb{N}$, such that each $h_k$ is bounded and two times continuously differentiable with bounded derivatives; see, e.g., Rudin1987. Therefore, any function $h\in \mathrm{L}_2^{0,f}$ can be written as $h= \sum_{k=1}^\infty \eta_k h_k $, for some $\eta = (\eta_k)_{k\in\mathbb{N}} \in \ell_2=\{ (z_k)_{k\in\mathbb{N}}\, | \, \sum_{k=1}^\infty z_k^ 2<\infty\}$. Besides the space $\ell_2$, we also need the space $c_{00}$ which is defined as the subset of sequences with finite support, i.e.,

align[align omitted — 161 chars of source]

Observe that $c_{00}$ is a dense subspace of $\ell_2$. It is introduced only in the asymptotic analysis to avoid convergence of infinite-dimensional processes and possibly induced mathematical complications, see Section (ref). However, the restriction $\eta\in c_{00}$ will not affect our conclusions. Indeed, considering $\eta\in c_{00}$ restricts our analysis to a subset of all semiparametric models which potentially makes the obtained upper bound higher. However, as we are able to show that this higher upper bound is (point-wisely) attainable by feasible tests for arbitrary innovation density in sequence, see Remark (ref), it constitutes the semiparametric power envelope and the test is semiparametrically optimal.

We model local perturbations to the innovation density $f$ as

align[align omitted — 161 chars of source]

where $\eta\in c_{00}$. We thus use a standard localization rate $T^{-1/2}$ for the bivariate density $f$. Indeed, Proposition (ref) below shows that all the above rates are appropriate in the sense that they lead to contiguous alternatives for the induced probability measures as $T$ tends to infinity.

In order to show that the above localization of the innovation density is valid, we need to establish that $f^{(T)}_{\eta}\in\mathfrak{F}$. This is the content of the next proposition.

propositionLet $f\in\mathfrak{F}$ and $\eta\in c_{00}$, then there exists a finite integer $\widetilde{T}$ such that for all $T\geq\widetilde{T}$ we have $f^{(T)}_{\eta}\in\mathfrak{F}$.

The proof uses exactly the same arguments as in the proof of Proposition 3.1 in ZvdAW2016, but with support $\mathbb{R}^2$ instead of $\mathbb{R}$. It is therefore omitted.

In terms of the local parameters $b$, $c$, and $\eta$, the hypothesis of interest becomes

align[align omitted — 95 chars of source]

versus the one-sided alternative

align[align omitted — 67 chars of source]

Partial-sum processes

In order to derive the limiting experiment for the predictive regression model, we need to introduce some partial-sum processes and study their asymptotic behavior. We denote by $\mathrm{P}^{(T)}_{b,c,\eta;f}$ the law of $(y_1,x_1)^\prime,\dots,(y_T,x_T)^\prime$ under the model ((ref))--((ref)), where the parameters $\beta$ and $\gamma$ are given by ((ref)) and the innovation density is given by ((ref)). Formally, we define the sequence of experiments of interest as

equation[equation omitted — 193 chars of source]

where $\Omega^{(T)}:=\mathbb{R}^{2\times T}$ and $\mathcal{F}^{(T)}:=\mathcal{B}(\mathbb{R}^{2\times T})$. We denote the expectation taken under the measure $\mathrm{P}^{(T)}_{0,0,0;f}$ by $\mathrm{E}^{(T)}$.

Let us already mention that we will also introduce a collection of probability measures $\mathbb{P}_{b,c,\eta}$, defined on a probability space $(\Omega,\mathcal{F})$, representing the limit experiment $\mathcal{E}\left(f\right)$ in Section (ref) below; see ((ref)). We will denote he expectation taken under the measure $\mathbb{P}_{0,0,0}$ by $\mathbb{E}$. That is, $\mathrm{P}^{(T)}$ and $\mathrm{E}^{(T)}$ refer to finite-sample distributions in the sequence of experiments, while $\mathbb{P}$ and $\mathbb{E}$ refer to distributions in the limit experiment.

As a final ingredient for our analysis, we introduce some partial-sum processes that we use throughout to link the sequence of experiments $\mathcal{E}^{(T)}\left(f\right)$ to the limit experiment $\mathcal{E}\left(f\right)$. In particular, define, with $\Delta x_t := x_t - x_{t-1}$, the partial-sum processes\footnote{One may consider partial sum processes that start at $t=2$ in order to make them exactly invariant to translations in $x_t$. This would, clearly, have no effect on our asymptotic results.}

align[align omitted — 546 chars of source]

Here we standardize the first three partial-sum processes by the standard deviations $\sigma_y$ and $\sigma_x$ in order to make their limits scale invariant. Under $\mathrm{P}^{(T)}_{0,0,0;f}$, by the Functional Central Limit Theorem (see also Lemma (ref)), we have

align[align omitted — 305 chars of source]

where the Brownian motions $W_{\varepsilon}$, $W_{\ell_{f_y}}$, $W_{\ell_{f_x}}$ and $W_{h}$ are defined on the common probability space $\left(\Omega,\mathcal{F},\mathbb{P}_{0,0,0}\right)$. We have to be precise about the notion of weak convergence adopted in ((ref)) as $W_{h}$ is infinite dimensional. In line with stochastic process theory, we mean that all finite-dimensional subprocesses of $W^{(T)}_{h}$ weakly converges in the space $D^{M+3}[0,1]$ with the uniform topology, where $M$ is the dimension of the finite-dimensional subprocess considered. This is precisely because we take the local parameter $\eta$ to be in $c_{00}$. For the sake of convenient notation, we write the seemingly infinite-dimensional convergence ((ref)). As argued above, we are ultimately able to attain the semiparametric power envelope induced under the restriction $\eta\in c_{00}$ so that we can claim semiparametric optimality.

Next, define the column vectors $J_{f_{y}h}=(J_{f_{y}h_k})_{k\in\mathbb{N}}$ and $J_{f_{x}h}=(J_{f_{x}h_k})_{k\in\mathbb{N}}$, where $J_{f_{y}h_k}:=\mathrm{E}_f\left[\sigma_y\ell_{f_y}(\varepsilon_t)h_k(\varepsilon_t)\right]$ and $J_{f_{x}h_k}:=\mathrm{E}_f\left[\sigma_x\ell_{f_x}(\varepsilon_t)h_k(\varepsilon_t)\right]$. As we have the equalities $\mathrm{E}_f\left[\varepsilon^x_t\ell_{f_y}(\varepsilon_t)\right]=-\sigma_y\int_{\mathbb{R}^2}\varepsilon^x\frac{\dot{f}_y(\varepsilon)}{f(\varepsilon)}f(\varepsilon)\mathrm{d}\varepsilon=-\sigma_y\int_{\mathbb{R}^2}\varepsilon^x\dot{f}_y(\varepsilon)\mathrm{d}\varepsilon=0$ and $\mathrm{E}_f\left[\varepsilon^x_t\ell_{f_x}(\varepsilon_t)\right]=-\sigma_x\int_{\mathbb{R}^2}\varepsilon^x\frac{\dot{f}_x(\varepsilon)}{f(\varepsilon)}f(\varepsilon)\mathrm{d}\varepsilon=-\sigma_x\int_{\mathbb{R}^2}\varepsilon^x\dot{f}_x(\varepsilon)\mathrm{d}\varepsilon=\sigma_x\int_{\mathbb{R}^2}f(\varepsilon)\mathrm{d}\varepsilon=\sigma_x$, the behavior of the Brownian motions $W_{\varepsilon}$, $W_{\ell_{f_y}}$, $W_{\ell_{f_x}}$ and $W_{h}$ is described by the covariance matrix

align[align omitted — 419 chars of source]

where $I_{\infty}$ denotes the $\infty$-dimensional identity matrix. The scaling by $\sigma_x$ and $\sigma_y$ introduced in ((ref))--((ref)) is indeed such that the covariance matrix ((ref)) does not depend on $\sigma_x$ or $\sigma_y$. Again, it still depends on $\rho$ through the various $J$ matrices.

Recall that the functions $h_k$ form an orthonormal basis for all zero-mean finite-variance functions that are orthogonal to $\varepsilon^x_t$. In view of the covariance matrix ((ref)), we may thus write, for $s\in[0,1]$,

align[align omitted — 193 chars of source]

Consequently, we also have

align[align omitted — 403 chars of source]

We again consider the special case of a Gaussian density $f$.

remark[Gaussian $f$] In the situation of Gaussian $f$ as discussed in Remark (ref), we may write the decomposition ((ref)) as $W_{\ell_{f_x}}=W_{\varepsilon}-\frac{\rho}{\sqrt{1-\rho^2}}W_\perp$ where $W_\perp$ is the standard Brownian motion generated by the increments $\left(\varepsilon^y/\sigma_y-\rho\varepsilon^x/\sigma_x\right)/\sqrt{1-\rho^2}$. Indeed, $W_{\varepsilon}$ and $W_\perp$ are independent (calculate the correlation of the increments that generate both processes). Thus, we also find $J_{f_{x}h}^\primeW_{h}(s)=-\rhoW_\perp$ and the decomposition ((ref)) becomes $J_{f_{xx}}=1+\frac{\rho^2}{1-\rho^2}=\frac{1}{1-\rho^2}=J_{f_{yy}}$. Moreover, we have $W_{\ell_{f_y}}=\frac{1}{\sqrt{1-\rho^2}}W_\perp$ and $J_{f_{yx}}=-\frac{\rho}{1-\rho^2}$.

Eliminating the nuisance parameter \lowercase{$f$} by invariance

We first focus on eliminating the nuisance parameter $f$ from the testing problem outlined in Section (ref). We will see that this can be handled using invariance arguments in the limit experiment, which we derive in Section (ref). In Section (ref), we consider the nuisance parameter $\gamma$.

We take the following steps in this section:

enumerate• Provide a structural representation of the limit experiment (Section (ref)). • Characterize maximally invariant test statistics in this limit experiment (Section (ref)). • Provide a structural representation of the invariant limit experiment (Section (ref)). • Provide a feasible version of the asymptotically invariant test statistics to be applied in the sequence of predictive regression experiments (Section (ref)).

These steps also show that, to eliminate the nuisance parameter $f$, instead of studying invariance restrictions in the sequence of finite-sample experiments, we only impose them in the limit experiment. Unlike for the location parameter $\mu$ (of $\varepsilon^y_t$), this limiting invariance property of the parameter $f$ does not follow directly from exact finite-sample invariance properties. Notably, the existing tests in the literature share this feature, as they also (implicitly) impose the invariance restriction in the limit, though not in the sequence; see Remark (ref). As far as we know, all existing tests belong to the class of asymptotically invariant (w.r.t.\ $f$) tests, while our test is semiparametrically optimal in the model we study. Section (ref) shows that this approach leads to considerable power gains in case the innovations are non-Gaussian, while no power is lost under Gaussianity.

A Structural Representation of the Limit Experiment

We consider the limit experiment corresponding to the predictive regression model ((ref))--((ref)) using the local perturbations ((ref)) and ((ref)), i.e., the limit of the experiments $\mathcal{E}^{(T)}\left(f\right)$ indexed by $T$, by studying the asymptotic behavior of the induced likelihood ratios. We expand the likelihood ratio around $(\beta,\gamma,\eta)=(0,1,0)$ and derive its limit in the following proposition, which can be interpreted as a generalization of Lemma 4 in JanssonMoreira2006 by including non-Gaussian distributions and perturbations thereof.\footnote{As preparation for the results in Section (ref), we allow in this proposition for local perturbations with respect to $\gamma$ even though, in the present section, $\gamma$ is assumed to be known.}

propositionFix $f\in\mathfrak{F}$. Consider the local parameters $b\in\mathbb{R}$, $c\in\mathbb{R}$, and $\eta\in c_{00}$. Then, \begin{itemize} • Under $\mathrm{P}^{(T)}_{0,0,0;f}$, the log-likelihood ratio of the predictive regression experiment satisfies, as $T\to\infty$, \begin{align} \log\frac{\mathrm{d}\mathrm{P}^{(T)}_{b,c,\eta;f}}{\mathrm{d}\mathrm{P}^{(T)}_{0,0,0;f}} = \Delta^{(T)}(b,c,\eta) - \frac12\mathcal{Q}^{(T)}(b,c,\eta) + o_\mathrm{P}(1), \end{align} where \begin{align*} \Delta^{(T)}(b,c,\eta) =& \frac{b}{T}\sum_{t=1}^T \frac{x_{t-1}}{\sigma_x} \sigma_y\ell_{f_y}(y_t, \Delta x_t) + \frac{c}{T}\sum_{t=1}^T x_{t-1} \ell_{f_x}(y_t, \Delta x_t) \\ & + \frac{1}{\sqrt{T}}\sum_{t=1}^T \sum_{k}\eta_kh_{k}(y_t, \Delta x_t), \\ \mathcal{Q}^{(T)}(b,c,\eta) =& \left(b^2J_{f_{yy}}+c^2J_{f_{xx}}+2bc J_{f_{yx}}\right)\frac{1}{T^2}\sum_{t=1}^T \frac{x_{t-1}^2}{\sigma_x^2} \\ & + \left(2bJ_{f_{y}h}^\prime\eta+2cJ_{f_{x}h}^\prime\eta\right)\frac{1}{T^{3/2}}\sum_{t=1}^{T}\frac{x_{t-1}}{\sigma_x} + \eta^\prime\eta. \end{align*} • Still under $\mathrm{P}^{(T)}_{0,0,0;f}$, as $T\to\infty$, we have \begin{align} \log\frac{\mathrm{d}\mathrm{P}^{(T)}_{b,c,\eta;f}}{\mathrm{d}\mathrm{P}^{(T)}_{0,0,0;f}} \Rightarrow \mathcal{L}(b,c,\eta) = \Delta(b,c,\eta) - \frac12\mathcal{Q} (b,c,\eta), \end{align} where \begin{align*} \Delta(b,c,\eta) =& b\int_0^1 W_{\varepsilon}(s)\mathrm{d}W_{\ell_{f_y}}(s) + c\int_0^1 W_{\varepsilon}(s)\mathrm{d}W_{\ell_{f_x}}(s) + \eta^\primeW_{h}(1)\\ =& \int_0^1 W_{\varepsilon}(s)\left(bJ_{f_{y}h}+cJ_{f_{x}h}\right)^\prime\mathrm{d}W_{h}(s) + c\int_0^1 W_{\varepsilon}(s)\mathrm{d}W_{\varepsilon}(s) + \eta^\primeW_{h}(1), \\ \mathcal{Q}(b,c,\eta) =& \left(b^2J_{f_{yy}}+c^2J_{f_{xx}}+2bcJ_{f_{yx}}\right)\int_0^1 W_{\varepsilon}(s)^2\mathrm{d}s \\ & + \eta^\prime\eta + \left(2bJ_{f_{y}h}^\prime\eta+2cJ_{f_{x}h}^\prime\eta\right)\int_0^1W_{\varepsilon}(s)\mathrm{d}s\\ =& \int_0^1\left|\left(bJ_{f_{y}h}+cJ_{f_{x}h}\right)W_{\varepsilon}(s)+\eta\right|^2\mathrm{d}s + c^2\int_0^1W_{\varepsilon}(s)^2\mathrm{d}s. \end{align*} • For every $b,c\in\mathbb{R}$ and $\eta\in c_{00}$, under $\mathbb{P}_{0,0,0}$, $\mathbb{E}[\exp\left(\mathcal{L}(b,c,\eta)\right)]=1$. \end{itemize}

A proof of Proposition (ref) is provided in Appendix (ref), but let us give a brief sketch here. Part (i) is immediate from an informal Taylor expansion of the log-likelihood ratios and, formally, follows from HvdAW2015, which provides generally applicable sufficient conditions for the quadratic expansion of likelihood ratios with densities that are differentiable in quadratic mean (DQM). This DQM condition is implied, for location models, by the absolutely continuity of the innovation density function and finiteness of the associated Fisher information, i.e., precisely the content of Assumption (ref). A detailed discussion can be found in LeCam1986 or LeCamYang2000. Part (ii) follows from the continuous mapping theorem applied to the weak convergence in ((ref)). Both forms of the central sequence $\Delta$ and quadratic term $\mathcal{Q}$ follow from ((ref)) and ((ref)). Part (iii) follows from standard stochastic calculations concerning Dol\a'{e}ans-Dade exponentials. To see this, note that $W_{h}$ and $W_{\varepsilon}$ are independent in view of ((ref)) and, thus, have vanishing quadratic covariation.

Part (iii) of Proposition (ref) ensures that we can introduce a collection of probability measures $\mathbb{P}_{b,c,\eta}$ on the measurable space $\left(\Omega,\mathcal{F}\right)$ (on which the Brownian motions $W_{\varepsilon}$, $W_{\ell_{f_y}}$, $W_{\ell_{f_x}}$ and $W_{h}$ are defined) by the Radon-Nikodym derivative

align[align omitted — 135 chars of source]

where $\mathcal{L}(b,c,\eta)$ is defined in ((ref)). Then, in the sense of H\a'{a}jek-Le Cam (see, for instance, vdVaart2000, Chapter 9), the sequence of predictive regression experiments, indexed by sample size $T$, weakly converges to the limit experiment described by the measures $\mathbb{P}_{b,c,\eta}$. We formally define this limit experiment by

align[align omitted — 167 chars of source]

where $\Omega:=C[0,1]\times C[0,1]\times C[0,1]\times C^{\mathbb{N}}[0,1]$ and $\mathcal{F}:=\mathcal{B_C}\otimes\mathcal{B_C}\otimes\mathcal{B_C}\otimes(\otimes_{k=1}^\infty\mathcal{B_C})$.

The following statement is an immediate consequence of Proposition (ref).

corollaryLet $f\in\mathfrak{F}$, then the sequence of experiments $\mathcal{E}^{(T)}\left(f\right)$ converges to the limit experiment $\mathcal{E}\left(f\right)$ as $T\to\infty$.

Although the log-likelihood ratios $\mathcal{L}(b,c,\eta)$ formally describe the limiting experiment, it is more insightful to provide, what we call, a structural representation. This structural representation provides a fixed-horizon continuous-time model for which the likelihoods are exactly equal to $\exp\left(\mathcal{L}(b,c,\eta)\right)$. From a statistical point of view, the induced experiments are thus equal. The result follows from an immediate application of Girsanov's theorem to the Radon-Nikodym derivates ((ref)). Its proof is therefore omitted.

theoremFix $f\in\mathfrak{F}$. Let, under $\mathbb{P}_{0,0,0}$, $Z_{\varepsilon}$, and $Z_{h}$ be zero-drift Brownian motions with covariance according to the first and last row and column of ((ref)). The limit experiment $\mathcal{E}\left(f\right)$ can be described as: observe $\left\{\left(W_{\varepsilon}(s),W_{h}(s)\right):s\in[0,1]\right\}$ generated by \begin{align} \mathrm{d}W_{\varepsilon}(s) &= c W_{\varepsilon}(s)\mathrm{d} s + \mathrm{d} Z_{\varepsilon}(s),\\ \mathrm{d}W_{h}(s) &= (bJ_{f_{y}h}+cJ_{f_{x}h})W_{\varepsilon}(s)\mathrm{d} s + \eta\mathrm{d} s + \mathrm{d} Z_{h}(s). \end{align}

A few remarks can be made in relation to Theorem (ref). First, note that for $b=c=0$ and $\eta=0$, we obtain $W_{\varepsilon}=Z_{\varepsilon}$ and $W_{h}=Z_{h}$. Secondly, the theorem essentially states that while $\left(W_{\varepsilon},W_{h}^\prime\right)^\prime$ is a zero-drift Brownian motion under $\mathbb{P}_{0,0,0}$, it becomes an Ornstein-Uhlenbeck process under $\mathbb{P}_{b,c,\eta}$, where the log-likelihood ratio $\log\left(\mathrm{d}\mathbb{P}_{b,c,\eta}/\mathrm{d}\mathbb{P}_{0,0,0}\right)$ equals $\mathcal{L}(b,c,\eta)$. Observe in particular that local perturbations of the innovation density $f$, as described by $\eta$, only affect the drift in ((ref)). We will consider inference procedures that are invariant with respect to $\eta$ in the limit experiment. In terms of the (sequence of) predictive regression model(s) this consequently translates into invariance with respect to (local perturbations in) the innovation density $f$.

In view of ((ref))--((ref)), we may also write

align[align omitted — 365 chars of source]

where $Z_{\ell_{f_x}}$ and $Z_{\ell_{f_y}}$ are zero-drift Brownian motions under $\mathbb{P}_{0,0,0}$. However, these equations do not contain any additional information, precisely given ((ref)) and ((ref)). Nevertheless, they will turn out useful when describing the likelihood ratio of the maximal invariant $\mathcal{M}$ to be introduced below in ((ref)).

Maximal Invariant

In the limit experiment $\mathcal{E}\left(f\right)$, the parameter $b\in\mathbb{R}$ is the parameter of interest, while $c\in\mathbb{R}$ and $\eta\in c_{00}$ are nuisance parameters. Observe that the nuisance parameter $\eta$ appears only in the drift of the SDEs in Theorem (ref). This suggests an invariance restriction in line with the approach in ZvdAW2016 for unit root testing.

To be specific, we first introduce, for $\eta\in c_{00}$, the transformations $\mathfrak{g}_{\eta}:C^\mathbb{N}[0,1]\to C^\mathbb{N}[0,1]$ by

align[align omitted — 56 chars of source]

for $W\in C^{\mathbb{N}}[0,1]$ and all $s\in[0,1]$. The transformation $\mathfrak{g}_{\eta}$ adds a drift $s\mapsto-\etas$ to $W$. Thus, Theorem (ref) implies that the law of $\left(W_{\varepsilon},(\mathfrak{g}_{\eta}(W_{h}))^\prime\right)^\prime$ under $\mathbb{P}_{b,c,0}$ is the same as the law of $\left(W_{\varepsilon},W_{h}^\prime\right)^\prime$ under $\mathbb{P}_{b,c,\eta}$.\footnote{By ((ref)) and ((ref)), the same holds for $W_{\ell_{f_x}}$ and $W_{\ell_{f_y}}$.} Denote by $\mathfrak{G}_{\eta}$ the group of transformations $\mathfrak{g}_{\eta}$ for $\eta\in c_{00}$. We can now characterize the maximal invariant with respect to $\mathfrak{G}_{\eta}$ in the limit experiment $\mathcal{E}\left(f\right)$.

For any process $W$, we define the associated bridge process by

align[align omitted — 67 chars of source]

for all $s\in[0,1]$. Then, one readily verifies

align*[align* omitted — 176 chars of source]

As a result, the bridges $B^{W_{h}}$ are invariant under the transformations $\mathfrak{g}_{\eta}$.

Define the mapping $M$ by $M(W_{\varepsilon},W_{h}):=(W_{\varepsilon},B^{W_{h}})$. It then follows that statistics that are measurable with respect to the $\sigma$-field

align[align omitted — 147 chars of source]

are invariant with respect to $\mathfrak{g}_{\eta}$ for all $\eta\in c_{00}$. Moreover, in the following theorem, we show $\mathcal{M}$ to be maximally invariant. Its proof is, again, provided in Appendix (ref).

theoremIn the limit experiment $\mathcal{E}\left(f\right)$, for $\eta\in c_{00}$, the $\sigma$-field $\mathcal{M}$ in ((ref)) is maximally invariant with respect to $\mathfrak{G}_{\eta}$.

A Structural Representation of the Invariant Limit Experiment

Theorem (ref) implies that any inference invariant with respect to $\mathfrak{G}_{\eta}$ must be measurable with respect to $\mathcal{M}$; see, e.g., LehmannRomano2005. Therefore, by the Neyman-Pearson lemma, inference based on the likelihood ratio with respect to $\mathcal{M}$ yields the power envelope for invariant tests in the limit experiment $\mathcal{E}\left(f\right)$. The following result provides this likelihood ratio.

theoremFix $f\in\mathfrak{F}$. Then the likelihood ratios in the limit experiment $\mathcal{E}\left(f\right)$ restricted to the maximal invariant $\mathcal{M}$ are given by \begin{align} \exp{\mathcal{L}_\mathcal{M}(b,c)} := \frac{\mathrm{d}\mathbb{P}_{b,c}^{\mathcal{M}}}{\mathrm{d}\mathbb{P}_{0,0}^{\mathcal{M}}} = \mathbb{E}\left[\frac{\mathrm{d}\mathbb{P}_{b,c,\eta}}{\mathrm{d}\mathbb{P}_{0,0,0}}|\mathcal{M}\right] = \exp\left(\Delta_{\mathcal{M}}(b,c)-\frac{1}{2}\mathcal{Q}_{\mathcal{M}}(b,c)\right), \end{align} where \begin{align} \Delta_{\mathcal{M}}(b,c) &= \int_0^1W_{\varepsilon}(s)\left(bJ_{f_{y}h}+cJ_{f_{x}h}\right)^\prime\mathrm{d}B^{W_{h}}(s) +c\int_0^1W_{\varepsilon}(s)\mathrm{d}W_{\varepsilon}(s)\\ &= \nonumber b\int_0^1W_{\varepsilon}(s)\mathrm{d}B_{\ell_{f_y}}(s) + c\left(\int_0^1W_{\varepsilon}(s)\mathrm{d}B_{\ell_{f_x}}(s) +W_{\varepsilon}(1)\overline{W_{\varepsilon}}\right),\\ \mathcal{Q}_{\mathcal{M}}(b,c) &= \left(bJ_{f_{y}h}+cJ_{f_{x}h}\right)^2\int_0^1\left(W_{\varepsilon}(s)-\overline{W_{\varepsilon}}\right)^2\mathrm{d}s + c^2\int_0^1W_{\varepsilon}(s)^2\mathrm{d}s\\ &= \nonumber \left(b^2J_{f_{yy}}+c^2(J_{f_{xx}}-1)+2bcJ_{f_{yx}}\right) \left(\overline{W_{\varepsilon}^2}-(\overline{W_{\varepsilon}})^2\right) + c^2\left(\overline{W_{\varepsilon}}\right)^2, \end{align} with $\overline{W_{\varepsilon}^2} = \int_0^1W_{\varepsilon}(s)^2\mathrm{d}s$ and $\overline{W_{\varepsilon}} = \int_0^1W_{\varepsilon}(s)\mathrm{d}s$.

The proof is provided in Appendix (ref). The first ways to write $\Delta_{\mathcal{M}}(b,c)$ and $\mathcal{Q}_{\mathcal{M}}(b,c)$ make explicit that the likelihood factorizes in a conditional likelihood given $W_{\varepsilon}$ and the marginal likelihood of $W_{\varepsilon}$. Both second ways to write $\Delta_{\mathcal{M}}(b,c)$ and $\mathcal{Q}_{\mathcal{M}}(b,c)$ follow from ((ref))--((ref)) and ((ref))--((ref)). Those are the versions that we use below to construct our feasible test statistics. Theorem (ref) also immediately yields the semiparametric power envelope, still for fixed $c$, that we do not present in detail for brevity.

The restriction to invariant tests removes the nuisance parameter $\eta$ from the testing problem. Indeed, the likelihood ratio ((ref)) no longer depends on $\eta$. Therefore, we can formally define the limit experiment restricted to the maximal invariance $\mathcal{M}$ as

equation[equation omitted — 178 chars of source]

Again, the likelihood ratios $\mathrm{d}\mathbb{P}_{b,c}^{\mathcal{M}}/\mathrm{d}\mathbb{P}_{0,0}^{\mathcal{M}}$ can also be interpreted as Girsanov transformations. We state this as a corollary as the result follows immediately from calculating the bridges corresponding to $W_{\ell_{f_y}}$ and $W_{\ell_{f_x}}$ in Theorem (ref).

corollaryFix $f\in\mathfrak{F}$. Let, under $\mathbb{P}_{0,0}^{\mathcal{M}}$, $Z_{\varepsilon}$ and $Z_{h}$ be zero-drift Brownian motions with covariance according to the first and last row and column of ((ref)). The limit experiment $\mathcal{E}_\mathcal{M}\left(f\right)$ can be described as follows: we observe, with $B^{W_{h}}(s)=W_{h}(s)-sW_{h}(1)$, $\left\{\left(W_{\varepsilon}(s),B^{W_{h}}(s)\right):s\in[0,1]\right\}$ with $\left(W_{\varepsilon},W_{h}\right)$ generated by \begin{align} \mathrm{d}W_{\varepsilon}(s) &= c W_{\varepsilon}(s)\mathrm{d} s + \mathrm{d} Z_{\varepsilon}(s),\\ \mathrm{d}W_{h}(s) &= (bJ_{f_{y}h}+cJ_{f_{x}h})W_{\varepsilon}(s)\mathrm{d} s + \mathrm{d} Z_{h}(s). \end{align}

The difference between Corollary (ref) and Theorem (ref) is twofold. First, besides the process $W_{\varepsilon}$, the observation in the invariant limit experiment in Corollary (ref) is only the Brownian bridge $B^{W_{h}}$ and not the complete Brownian motion $W_{h}$. Second, as a consequence of this, the nuisance parameter $\eta$ disappeared from ((ref)).

Corollary (ref) does not provide, as far as we know, a further invariance structure that can be used to eliminate the nuisance parameter $c$. As a result, we rely, in Section (ref), on the so-called Approximate Least Favorable Distribution method to deal with this last nuisance parameter.

We conclude this section by again considering the special case of a Gaussian innovation density $f$. This also shows where exactly our power gains, under serially independent innovations, come from relative to the Gaussian procedures in, for instance, JanssonMoreira2006.

remark[Attainability of the Semiparametric Power Envelope] One may expect the semiparametric power envelope to be formally attainable by a likelihood-ratio test constructed using a nonparametric estimate of the score function $\ell_f$. Intuitively, the argument is as follows. Rewrite $\int_0^1W_{\varepsilon}(s)\mathrm{d}B_{\ell_{f_y}}(s) = \int_0^1\left(W_{\varepsilon}(s) - \overline{W_{\varepsilon}}\right)\mathrm{d}W_{\ell_{f_y}}(s)$. Hence, even though there is a bias $a$ (at rate $\sqrt{T}$) in the estimated score function, this bias will be canceled out automatically since $\int_0^1\left(W_{\varepsilon}(s) - \overline{W_{\varepsilon}}\right)\mathrm{d}\left(as+W_{\ell_{f_y}}(s)\right) = \int_0^1\left(W_{\varepsilon}(s) - \overline{W_{\varepsilon}}\right)\mathrm{d}W_{\ell_{f_y}}(s)$. The same argument applies to the term $\int_0^1W_{\varepsilon}(s)\mathrm{d}B_{\ell_{f_x}}(s)$. Compare the discussion in Jansson2008 for the unit root testing problem and Zhou2020 for general LAN, LAMN, and LABF experiments.
remark[Gaussian $f$] In the situation of Gaussian $f$, Remark (ref) and Remark (ref) imply that $B_{\ell_{f_y}}$ and $B_{\ell_{f_x}}$ are linear combinations of $B_{\varepsilon}$ and $B_\perp$ (the Brownian bridges generated by $W_{\varepsilon}$ and $W_\perp$, respectively). As a result, the optimal invariant procedures are measurable with respect to $W_{\varepsilon}$ and $B_\perp$. Using the same conditional expectation calculation, the associated log-likelihood ratio of the Gaussian $\sigma$-field, $\mathcal{M}_{\rm Gaussian}=\sigma\left(W_{\varepsilon},B_\perp\right)$, leads to the Gaussian log-likelihood ratio in JanssonMoreira2006. As $B_\perp$ is spanned by $B^{W_{h}}$, the $\sigma$-field $\mathcal{M}_{\rm Gaussian}$ is also invariant w.r.t $\eta$ (or $f$), but it is not maximally invariant. As a consequence, under non-Gaussianity, this leads to an efficiency loss in statistical inference. Note that all existing tests in the literature are (essentially) based on the Gaussian likelihood of the generally non-maximally invariant $\mathcal{M}_{\rm Gaussian}$, e.g., JanssonMoreira2006 and EMW2015. Therefore, these tests belong to the class of asymptotically invariant tests. This invariance imposed in the limiting experiment is associated to invariance w.r.t.\ the innovation density $f$ in the sequence as $\eta$ represents local perturbations precisely of $f$. Indeed, we have the convergence $W^{(T)}_{\varepsilon}(s) \Rightarrow W_{\varepsilon}(s)$ and the one associated to $W_\perp$ for all $f \in \mathfrak{F}$, hence, $\eta$ will not enter the associated equation ((ref)) in the limiting experiment. See Muller2011 for a more comprehensive analysis of this convergence.

Rank-based asymptotically invariant statistics

The elimination of the nuisance parameter $\eta$ is performed in the limit experiment $\mathcal{E}\left(f\right)$ and leads to $\mathcal{E}_\mathcal{M}\left(f\right)$. We now show how this elimination can be mimicked in the actual predictive regression model of interest, i.e., in $\mathcal{E}^{(T)}\left(f\right)$. It is reasonable to expect that exploiting the asymptotic invariance structures also works “well” for the sequence of experiments. The claim will be substantiated by the simulation results in Section (ref).

In line with the vast literature on rank-based inference, the appearance of the Brownian Bridges $B_{\ell_{f_x}}$ and $B_{\ell_{f_y}}$ in Corollary (ref), naturally suggest to use statistics that are based on ranks of the innovations $\varepsilon^y_t$ and $\varepsilon^x_t$ in the predictive regression model. Indeed, we will follow that route. However, in the present situation we deal with bivariate innovations $(\varepsilon^y_t,\varepsilon^x_t)$ which complicates the analysis considerably relative to models with univariate innovations that are mostly studied in the literature.

As the true innovation density $f$ is unknown, we actually base our test statistic on an assumed (so-called reference) density $g$ that also satisfies Assumption (ref). Let $g_y$ and $g_x$ denote the marginal densities for the first, respectively, second component of $g$. The bivariate nature of the innovations $(\varepsilon^y_t,\varepsilon^x_t)$ implies that we cannot deal with a completely general reference bivariate density $g$. Thus, we choose marginal reference densities $g_y$ and $g_x$, and a reference correlation parameter $\rho_g$. For the marginal reference densities, we impose the standard condition in the rank-based inference literature, see, e.g., Theorem 13.5 in vdVaart2000.

assumptionThe marginal reference densities $g_i$, $i=\{y,x\}$, are strictly positive, absolutely continuous with derivative $\dot{g}_i$ and $J_{g_i}:=\int\left(\dot{g}_i/g_i\right)^2g_i<\infty$. Moreover, we have \begin{equation} \lim_{T\to\infty}\frac{1}{T}\sum_{t=1}^T \left(-\frac{\dot{g}_i}{g_i}\left(G_i^{-1}\left(\frac{t}{T+1}\right)\right)\right)^2 = J_{g_i}, \end{equation} where $G_i^{-1}$ is the inverse cumulative distribution function associated to $g_i$.

Moreover, given an additionally chosen reference correlation $\rho_g\in(-1,1)$, we define the associated bivariate reference score function

align[align omitted — 180 chars of source]

where

align*[align* omitted — 325 chars of source]

The linearity of the reference score functions $\ell_{g_y}$ and $\ell_{g_x}$ is key to the analysis that follows. It implies that, when using component-wise ranks of the innovations $(\varepsilon^y,\varepsilon^x)$, the resulting rank-based processes converge to a bivariate Brownian bridge. Despite its seemingly restrictive nature, the linearity allows use to fully exploit the invariance structures embedded in the predictive regression model of interest, leading to sizable power gains (see Section (ref)).

Now, let $R_{y,t}$ denote the rank of $y_t$ (among $y_1,\ldots,y_T$), while $R_{x,t}$ denotes the rank of $\Delta x_t=x_t-x_{t-1}$ (among $\Delta x_1,\ldots,\Delta x_T$). Note that the pairs $\left(R_{y,t},R_{x,t}\right)$ equal the (component-wise) ranks of $\left(\varepsilon^y_t,\varepsilon^x_t\right)$ under $\beta=0$ and $\gamma = 0$. We define the bivariate partial sum process of the rank-based scores by

align[align omitted — 312 chars of source]

for $s\in[0,1]$. The following result establishes the limiting behavior of $B_{\ell_{g}}^{(T)}$ under $\mathrm{P}^{(T)}_{0,0,\eta;f}$. Its proof is again provided in Appendix (ref).

propositionSuppose $\varepsilon_t=(\varepsilon^y_t,\varepsilon^x_t)^\prime$ are i.i.d.\ innovations with density $f\in\mathfrak{F}$. Let $g_y$ and $g_x$ be reference densities that satisfy Assumption (ref) and fix the reference correlation $\rho_g$. Then, under $\mathrm{P}^{(T)}_{0,0,\eta;f}$, we have \begin{align} B_{\ell_{g}}^{(T)}\RightarrowB_{\ell_{g}}, \end{align} where $B_{\ell_{g}}$ is a bivariate Brownian bridge, i.e., $B_{\ell_{g}}(s)=W_{\ell_{g}}(s)-sW_{\ell_{g}}(1)$, with $W_{\ell_{g}}$ a zero-drift Brownian motion. The covariance of $W_{\ell_{g}}$ with $W_{\varepsilon}$ and $W_{\ell_{f}}:=(W_{\ell_{f_y}},W_{\ell_{f_x}})^\prime$ is given by \begin{equation} \operatorname{Var}\begin{pmatrix} W_{\varepsilon}(1)W_{\ell_{f}}(1)W_{\ell_{g}}(1) \end{pmatrix} = \begin{pmatrix} 1&e_1^\prime&\bm{\sigma}_{\varepsilon g}^\prime\\e_1&J_f&J_{fg}\\\bm{\sigma}_{\varepsilon g}&J_{gf}&J_g \end{pmatrix}, \end{equation} where \begin{align*} e_1 &= (0,1)^\prime, \\ \bm{\sigma}_{\varepsilon g} &= (\sigma_{\varepsilon g_y},\sigma_{\varepsilon g_x})^\prime = \mathrm{E}_f\left[ \varepsilon^x_t \ell_g\left(G_y^{-1}(F_y(\varepsilon^y_t)),G_x^{-1}(F_x(\varepsilon^x_t))\right) \right],\\ J_{fg} &= J_{gf}^\prime = \mathrm{E}_f\left[ \ell_f(\varepsilon^y_t,\varepsilon^x_t)\ell_g\left(G_y^{-1}(F_y(\varepsilon^y_t)),G_x^{-1}(F_x(\varepsilon^x_t))\right)^\prime \right],\\ J_g &= \mathrm{E}_f\left[ \ell_g\left(G_y^{-1}(F_y(\varepsilon^y_t)),G_x^{-1}(F_x(\varepsilon^x_t))\right)\ell_g\left(G_y^{-1}(F_y(\varepsilon^y_t)),G_x^{-1}(F_x(\varepsilon^x_t))\right)^\prime \right]. \end{align*}

The above result is classical for univariate rank statistics. In the present paper, we use component-wise bivariate ranks. One complication is that the matrix $J_g$ depends on $f$ through its copula. This implies that, like $J_g$, it will have to be estimated in applications; compare also to the discussion of Theorem 3.1 in Zhou2020.

We use the rank-based processes $B_{\ell_{g}}^{(T)}$ to replace $B_{\ell_{f}}$ in the likelihood ratio in Theorem (ref); see Section (ref) for details. In line with Remark (ref)}, one could contemplate to use reference densities $\hat{f}$ based on a non-parametric estimate of the true innovation density, but we leave a formal analysis for future work. As we will see in Section (ref), even for incorrectly chosen reference densities (that is, for $g\neq f$), our procedure features power gains over existing Gaussian based procedures. These gains come from the assumption that the error term $\varepsilon_t$ is driven by some i.i.d.\ innovations, which may possibly be maintained in empirical work. It is important to note that choosing a reference density $g\neq f$ does not affect the validity of our test. The test will be of the appropriate level irrespective of the reference densities $g_y$ and $g_x$ chosen (provided they satisfy Assumption (ref)). But, likelihood ratio tests based on Theorem (ref) still feature the nuisance parameter $c$. We deal with this in the next section.

For completeness, we also provide the equivalent to Corollary (ref) when using the reference density $g$.

corollaryFix $f\in\mathfrak{F}$. Let $b\in\mathbb{R}$, $c\in(-\infty,0]$, and $\eta\in c_{00}$. Then, under $\mathbb{P}_{b,c,\eta}$, the behavior of $W_{\varepsilon}$ and $B_{\ell_{g}}$ follows \begin{align} \mathrm{d} W_{\varepsilon}(s) &= cW_{\varepsilon}(s)\mathrm{d} s + \mathrm{d} Z_{\varepsilon}(s),\\ \mathrm{d}W_{\ell_{g}}(s) &= J_{gf}\begin{pmatrix} b \\ c \end{pmatrix} W_{\varepsilon}(s)\mathrm{d} s + \mathrm{d}Z_{\ell_{g}}(s), \end{align} where, under $\mathbb{P}_{0,0,0}$, $Z_{\ell_{g}}$ is a bivariate Brownian motion with variance $J_g$ and covariance with $W_{\varepsilon}$ equal to $\bm{\sigma}_{\varepsilon g}$.

Eliminating the nuisance parameter $\gamma$ by ALFD

In the previous section, we have developed the semiparametric power envelope for tests on $b$ that are invariant with respect to $\eta$, under the assumption that $c$ is known. We now address the question of testing the regression coefficient $\beta$ in case $\gamma$ is treated as a nuisance parameter as well.

As argued in the the discussion following Corollary (ref), we conjecture that the nuisance parameter $c$ cannot be dealt with using invariance arguments. Various alternative methods to deal with nuisance parameters in testing problems have been used in the literature. In relation to the predictive regression model at hand, we mention the Bonferroni method (CavanaghElliotStock1995 and CampbellYogo2005); tests based on a conditional unbiasedness condition (JanssonMoreira2006); and tests based on a numerically calculated Approximate Least Favorable Distribution (ALFD) as more recently proposed in EMW2015. All these techniques apply to the Gaussian likelihood ratio statistic in Remark (ref).

These approaches have different advantages and disadvantages. CampbellYogo2005 proposes a modified Bonferroni method to eliminate the nuisance parameter $c$, leading to a simple yet more powerful test than the CavanaghElliotStock1995 test. However, as pointed out by Phillips2014, inference based on Bonferroni bounds can be severely undersized when the predictor is “far away” from being a unit root process ($\gamma<<1$). In such a case, confidence intervals obtained by inverting the test may end up having essentially zero coverage probability. JanssonMoreira2006 develops an approach conditional on specific auxiliary statistics---the terms (only) associated with $c$ in the Gaussian likelihood ratio---and derives an optimal test in the class of conditionally unbiased tests. Nevertheless, such a conditional unbiasedness constraint narrows the considered class and rules out some more powerful tests. Consequently, as shown by the simulation results of JanssonMoreira2006, the associated test has relatively low power compared to the CampbellYogo2005 test under most alternatives.

Recently, EMW2015 proposes a numerical algorithm to determine an ALFD of the nuisance parameter $c$ to optimize weighted average power over some compact interval (of $c$). Note that with respect to our parameter of interest $b$, we consider point-optimal test and do not use weighted powers over a discretized space to avoid the induced computational complexities. On one hand, the ALFD yields an upper bound of the weighted average power for all valid tests. On the other hand, integrating out the likelihood statistic w.r.t.\ the ALFD leads to a “nearly optimal” test whose power is close to the upper bound. Moreover, by switching to standard asymptotic approximations in case $\gamma$ appears to be far from unity, the associated test can achieve better size and power performances uniformly for all $c\in (-\infty, 0]$ (e.g., across the parameter space $\gamma\in(-1,1]$). Therefore, we employ this ALFD approach in the present paper, together with the switching mechanism (see Appendix (ref)), to our rank-based likelihood statistics in ((ref)).\footnote{We expect that other approaches based on likelihood ratios, e.g., the approaches of CampbellYogo2005 and JanssonMoreira2006, will apply here as well. This is because (i) the semiparametric likelihood ratio $\mathcal{L}_\mathcal{M}(b,c)$ in Theorem (ref) is as the general version of the Gaussian likelihood ratio in JanssonMoreira2006, thus when the true density is Gaussian, the former reduces to the latter; (ii) its rank-based proxy in ((ref)) has the same structure (exponential family); and (iii) the asymptotic behaviors of the associated rank-based processes are known and consistently estimable.} This leads to tests that are of correct size for all relevant $c$ and have good power performance. We confirm these properties by simulations in Section (ref).

The Approximately Least Favorable Distribution (ALFD) Approach

In Section (ref) we used invariance arguments to reduce the predictive regression testing problem towards log-likelihood ratio of the form ((ref)) where $b$ is the parameter of interest to be tested and $c$ is a nuisance parameter. We briefly outline, in the present section, how the Approximate Least Favorable Distribution approach in EMW2015 works in our setting.

Rewrite the log-likelihood ratio of the maximal invariant $\mathcal{M}$ in Theorem (ref) as

align*[align* omitted — 138 chars of source]

where

align[align omitted — 374 chars of source]

One can thus consider the four-dimensional sufficient statistic $S:=\left(S_1,S_2,S_3,S_4\right)$. For notational simplicity, in the present section, we denote by $F_{b,c}(S)$ the distribution of $S$ under $\mathbb{P}_{b,c}$. The hypothesis of interest is

align[align omitted — 143 chars of source]

Note that, thus, both the null and the alternative hypothesis are composite. We first discuss elimination of the nuisance parameter $c$ under the alternative and, subsequently, its elimination under the null.

To eliminate the nuisance parameter $c$ under the alternative, a standard approach is to consider a so-called weighted average power (see, e.g., AndrewsPloberger1994)

align[align omitted — 148 chars of source]

where $\varphi$ is some test function for the problem above and $\Lambda_1$ is a probability weighting measure for $c\in(-\infty,0]$. The weighting measure $\Lambda_1$ can be chosen by the researcher and reflects the weights that she assigns to various values of $c$ under the alternative. Due to Fubini's Theorem, we have

align[align omitted — 103 chars of source]

which leads to the simple alternative hypothesis $\mathrm{H}_{1;\Lambda_1}$, under which the distribution of $S$ is given by the mixture $F_{b;\Lambda_1}(S)=\int F_{b,c}(S)\mathrm{d}\Lambda_1(c)$. In this way, the testing problem is reduced to testing $\mathrm{H}_0$ against $\mathrm{H}_{1;\Lambda_1}$.

Subsequently, in order to eliminate the nuisance parameter $c$ under the null we proceed as follows. Again we impose a probability weighting measure $\Lambda_0$ for $c$ and introduce the simple null hypothesis, denoted $\mathrm{H}_{0;\Lambda_0}$, under which the distribution of $S$ is given by $F_{b;\Lambda_0}(S)=\int F_{b,c}(S)\mathrm{d}\Lambda_0(c)$. Now we define the test $\varphi_{\bar{b};\Lambda}$ by

align[align omitted — 307 chars of source]

where the critical value $\kappa$ is chosen to obtain the desired size. By the Neyman-Pearson Lemma, $\varphi_{\bar{b},\Lambda_0}$ is point optimal at $b=\bar{b}$, for the problem of testing the null $\mathrm{H}_{0;\Lambda_0}$ against the alternative $\mathrm{H}_{1;\Lambda_1}$.

The problem of choosing $\Lambda_0$ is, unfortunately, more complicated than that of choosing $\Lambda_1$. The reason is that we want to control the rejection probability of the test, not only under $\mathrm{H}_{0;\Lambda_0}$, but for all values of $c\in(-\infty,0]$. In general there is no reason to expect that a level-$\alpha$ test under $\mathrm{H}_{0;\Lambda_0}$ is of correct size for the entire null hypothesis $\mathrm{H}_0$. However, for some specific choices of $\Lambda_0$ this statement is true, and such a distribution is called a least-favorable distribution; see, e.g., LehmannRomano2005, Theorem 3.8.1. Formally, a distribution $\Lambda_0^*$ is called least favorable if the most powerful level-$\alpha$ test ((ref)) for testing $\mathrm{H}_{0;\Lambda_0^*}$ against $\mathrm{H}_{1;\Lambda_1}$ is of the desired size for the (entire) null hypothesis $\mathrm{H}_0$. Moreover, once more by Theorem 3.8.1 in LehmannRomano2005, the test $\varphi_{\bar{b},\Lambda_0^*}$ is also point optimal (at $b=\bar{b}$) for this problem. A least-favorable distribution $\Lambda_0^*$ exists in most of the usual statistical problems. conditions that ensure this and associated references can be found in Section 3.8 of LehmannRomano2005.

As, in most cases, the least-favorable distribution $\Lambda_0^*$ is not easily obtained, EMW2015 propose a numerical method to find, what they call, an “Approximate Least Favorable Distribution” (ALFD). The ALFD is defined as follows.

definitionAn $\epsilon$-ALFD is a probability distribution $\Lambda_0^{*\epsilon}$ over $(-\infty,0]$ satisfying \begin{itemize} • the Neyman-Pearson test ((ref)) with $\Lambda=\Lambda_0^{*\epsilon}$ and critical value $\kappa=\kappa^*$, i.e., $\varphi_{\bar{b},\Lambda_0^{*\epsilon}}$, is of size $\alpha$ under $\mathrm{H}_{0;\Lambda_0^{*\epsilon}}$ and has power $\bar{\pi}$ against $\mathrm{H}_{1;\Lambda_1}$; • there exists $\kappa^{*\epsilon}$ such that the test ((ref)) with $\Lambda=\Lambda_0^{*\epsilon}$ and $\kappa=\kappa^{*\epsilon}$, $\varphi_{\bar{b},\Lambda_0^{*\epsilon}}^\epsilon$, is of level $\alpha$ under $\mathrm{H}_0$, and has power of at least $\bar{\pi}-\epsilon$ against $\mathrm{H}_{1;\Lambda_1}$. \end{itemize}

The test $\varphi_{\bar{b},\Lambda_0^{*\epsilon}}^{\epsilon}$ (in particular, the ALFD $\Lambda_0^{*\epsilon}$ and the critical value $\kappa^{*\epsilon}$) is exactly what we are looking for, once we have set the weights $\Lambda_1$ of interest for the alternative hypothesis. Besides the size control under $\mathrm{H}_0$, the definition above also ensures that the test $\varphi_{\bar{b},\Lambda_0^{*\epsilon}}^{\epsilon}$ enjoys a near-optimality property with a relatively small power loss (less than $\epsilon$).

Note that even for a given (small) value of $\epsilon$, the ALFD $\Lambda_0^{*\epsilon}$ is not necessarily “close” to the least favorable distribution $\Lambda_0^*$. Actually, (possibly infinitely) many pairs of $(\Lambda_0^{*\epsilon},\kappa^{*\epsilon})$ may satisfy Definition (ref). The details about how to implement the numerical algorithm to determine a pair of $(\Lambda_0^{*\epsilon},\kappa^{*\epsilon})$ (henceforth the test $\varphi_{\bar{b},\Lambda_0^{*\epsilon}}$) for a small $\epsilon$ can be found in Section 3 and Appendix A of EMW2015. As the nuisance parameter space $c\in(-\infty,0]$ is unbounded, we also need to “switch” back to standard test statistics (i.e., in the stationary case) for large values of $|c|$. We provide in Appendix (ref) the details about our test for the standard part of the limit experiment $\mathcal{E}\left(f\right)$.

Putting it all together

Putting everything together, our test for the predictive regression model is based on applying the ALFD approach to the rank-based counterpart (using Proposition (ref)) of the asymptotically point-optimal invariant derived in Theorem (ref).

We thus replace, in the sufficient statistic $S=\left(S_1,S_2,S_3,S_4\right)$ in ((ref)), $W_{\varepsilon}$, $B_{\ell_{f_y}}$, and $B_{\ell_{f_x}}$ by $W^{(T)}_{\varepsilon}$, $B_{\ell_{g_y}}^{(T)}$, and $B_{\ell_{g_x}}^{(T)}$, leading to the feasible rank-based statistic

equation[equation omitted — 137 chars of source]

where

align*[align* omitted — 450 chars of source]

To make the log-likelihood ratio $\mathcal{L}_{\mathcal{M}}$ in ((ref)) fully feasible, we also have to deal with $J_f$. From KaganLandsman1999 we know that $J_f$ is diagonalized by the Cholesky root of the correlation matrix $\mathbf{R}_g$. Therefore, we replace $J_f$ by

align[align omitted — 220 chars of source]

where $J_{g_y}$ and $J_{g_x}$ are the Fisher information of the chosen marginal reference densities defined in Assumption (ref), and $\mathbf{R}_g$ is the correlation matrix based on the chosen reference correlation $\rho_g$, i.e., $\mathbf{R}_g := \left(

smallmatrix1 & \rho_g \\ \rho_g & 1

\right)$. We recommend to use a consistent estimate of $\rho$ as $\rho_g$ regarding the power of the test, although any choice of $\rho_g$ would lead to correct sizes. This leads to our feasible rank-based log-likelihood statistic

align[align omitted — 208 chars of source]

of which the limit is given by the proposition below.

propositionSuppose $\varepsilon_t=(\varepsilon^y_t,\varepsilon^x_t)^\prime$ are i.i.d.\ innovations with density $f\in\mathfrak{F}$. Let $g_y$ and $g_x$ be reference densities that satisfy Assumption (ref) and fix the reference correlation $\rho_g$. Then, for $b\in\mathbb{R}$ and $c\in(-\infty,0]$, under $\mathrm{P}^{(T)}_{0,0,\eta;f}$, we have \begin{align*} \mathcal{L}_g^{(T)}(b,c) \Rightarrow \mathcal{L}_g(b,c), \end{align*} where \begin{align} \mathcal{L}_g(b,c) := \lpinteS_{g,1}+\lppersS_{g,2}-\frac{1}{2}\left((b,c)J_p(b,c)^\prime-c^2 \right)S_{g,3}-\frac{1}{2}c^2S_{g,4} \end{align} with \begin{align*} S_{g,1} &= \int_0^1W_{\varepsilon}(s)\mathrm{d}B_{\ell_{g_y}}(s), S_{g,2} = \int_0^1W_{\varepsilon}(s)\mathrm{d}B_{\ell_{g_x}}(s) + W_{\varepsilon}(1)\overline{W_{\varepsilon}}, \\ S_{g,3} &= \overline{W_{\varepsilon}^2} - \left(\overline{W_{\varepsilon}}\right)^2 {\rm and } S_{g,4} = \overline{W_{\varepsilon}^2}. \end{align*}

We omit the proof of Proposition (ref) since it directly follows from the weak convergences in ((ref)) and ((ref)), the continuous mapping theorem, and the rank-based stochastic integral convergence argument in the proof of Lemma 4.1 of ZvdAW2016.

Although not explicit in the above, observe that the statistic $\mathcal{L}_g^{(T)}(b,c)$ in ((ref)) still depends on $\sigma_x$ through $W^{(T)}_{\varepsilon}$ defined in ((ref)). We will simply replace $\sigma_x$ by its sample counterpart below. As long as this estimator is consistent, the continuous mapping theorem shows that this replacement has no asymptotic consequences. The statistic does not depend on $\sigma_y$, but it does depends on the reference correlation $\rho_g$.

Now, applying the ALFD algorithm to $\mathcal{L}_g^{(T)}(b,c)$, we obtain a distribution $\Lambda^{*\epsilon}_{0,g}$ and critical value $\kappa_{g,n}$ such that the test

align[align omitted — 433 chars of source]

is of size $\alpha$. Here $\bar{b}$ serves as a fixed alternative point for the quasi-likelihood statistic; see ERS1996.

In order to get the appropriate critical values of the test, note that we need consistent estimates, under the null, of $J_g$ and $J_{fg}$. We need these in order to ensures the feasibility of the numerically determined pair $(\Lambda_0^{*\epsilon},\kappa^{*\epsilon})$. In applications $J_g$ and $J_{fg}$ can easily be estimated, however, in the Monte Carlo study below we estimate $J_g$ and $J_{fg}$ based on the known.\footnote{A simple consistent estimator for $J_g$ would be the sample covariance of the rank-based scores $\ell_g$ defined in ((ref)) and a direct rank-based estimator for $J_{fg}$ can be found in CassartHallinPain2010.} This is necessary as we cannot afford to determine a pair $(\Lambda_0^{*\epsilon},\kappa^{*\epsilon})$ for each repetition in the simulation. That would be too intensive computationally.

A Monte Carlo Study

In this section, we explore by Monte Carlo the size and power properties of our test ((ref)), combined with the switching approach detailed in Appendix (ref), (labeled WZ) relative to the Gaussian quasi-likelihood counterpart in EMW2015 (labeled EMW). From the theoretical results, both tests should enjoy good size properties but the WZ test should exhibit larger power in case the true innovation distribution is not Gaussian. Under Gaussian innovation distribution, both tests should have similar power.

Section (ref) provides simulations under the predictive regression model studied formally in this paper. Section (ref) provides results of our test under conditional heteroskedasticity. Finally, Section (ref) provides results when the reference density used in the test is estimated.

Simulations under maintained i.i.d.\ assumption

We simulate the model ((ref))--((ref)) with $\mu=2$, $\sigma_y = 3$, $\sigma_x=3$, and $\rho=-0.5$. All results reported in this section are based on 10,000 replications.

For the ALFD approach, we choose a discrete weighting distribution $\Lambda_1$ in ((ref)) where each of the 57 points

align*[align* omitted — 50 chars of source]

of the support have equal weight. The same 57 points are also as the support of $\Lambda_{0}^{*\epsilon}$. For the test statistic in ((ref)), we choose a fixed alternative $\bar{b}=B(1.645)$ where the power is about $50\%$. For the reference correlation $\rho_g$, we use the simple sample correlation of $\hat\epsilon^y_t$ and $\hat\epsilon^x_t$ under the null, where $\hat\epsilon^y_t = y_t - \sum_{t=1}^{T}y_t$ and $\hat\epsilon^x_t$ is the residual of the regression of $x_t$ on $x_{t-1}$.

We present the power curves in two ways. The first presentation follows EMW2015: We let the local nuisance parameter $c$ (which governs the persistence of the predictor) take 21 values $c\in\{0,-10,-20,\dots,-200\}$. And to have roughly similar power for each value of $c$, we transform the parameter $b$ by

align[align omitted — 106 chars of source]

Alternatives for $\beta$ are now characterized by different values of $\delta$. The null hypothesis $\mathrm{H}_0$ corresponds to $\delta=0$, and we let the parameter of interest $b$ take three alternatives: $\delta\in\{1,2,3\}$. Secondly, we present power curves where we fix the nuisance parameter $c = -25$ and plot the rejection rates for $\delta \in [0, 6]$. The significance level $\alpha$ is chosen to be $5\%$ in all cases.

figure[figure omitted — 403 chars of source]
figure[figure omitted — 398 chars of source]
figure[figure omitted — 362 chars of source]
figure[figure omitted — 360 chars of source]

In Figure (ref), we reports the large-sample ($T=2,000$) size and power properties of our rank-based WZ test and the EMW test, for different combinations of the true density $f$ and the marginal reference densities $g_y$ and $g_x$. The upper-left subplot reports the case where $f$ is a multivariate $t_3$ density, while $g_y$ and $g_x$ are both univariate $t_3$ densities. Both the EMW test and the WZ test are of correct size for all chosen values of $c$. Under the alternative hypothesis (i.e., for $\delta\in\{1,2,3\}$), the WZ test is more powerful than the EMW test. Taking the alternative $\delta=2$ as example, for most values of $c$, the power of the EMW test is about $65\%$ while the WZ test attains about $90\%$ power. In the upper-right subplot, we keep $f$ unchanged and let $g_y$ and $g_x$ both be Gaussian. Both tests provide correct size and, again, the WZ test is more powerful than the EMW test. However, compared to the upper-left subplot, we observe that the WZ test suffers a small power loss when choosing reference densities that are further away from the true ones. When $f$ is Gaussian, the WZ test with Gaussian marginal reference densities shares almost the same size and power performances as the EMW test, as shown by the bottom-left subplot. The bottom-right subplot presents the case when $f$ is Gaussian, while the marginal reference densities $g_y$ and $g_x$ are univariate $t_3$. In this case, the WZ test is less powerful than the EMW test. In practice, we may want to avoid this power loss by pre-testing the residuals under the null hypothesis. We study this in Section (ref).

figure[figure omitted — 394 chars of source]
figure[figure omitted — 390 chars of source]

Actually, one can always use Gaussian reference densities as a conservative choice, which is based on a (numerical) ChernoffSavage1958 result --- keeping the marginal reference densities $g_y$ and $g_x$ Gaussian, the WZ test is always more powerful than the EMW test when $f$ is non-Gaussian, and it works as well as the EMW test when $f$ is Gaussian. A formal proof of this result in LABF-type experiments is still an open question, but we show that this property holds in some more simulations. In Figure (ref), we fix $g_y$ and $g_x$ to be Gaussian, and choose four different multivariate innovation distributions: (i) Gaussian copula with Laplace marginal distributions (top-left, labeled Multi-Laplace); (ii) Multivariate Pearson distribution with skewness 3 and kurtosis 36 (top-right, labeled Multi-Pearson); (iii) Gaussian copula with $t_3$ distribution for the first dimension and Gaussian distribution for the second dimension (bottom-left, labeled Multi-combo1); and (iv) $t_3$ copula with Gaussian for the first dimension and $t_3$ for the second dimension (bottom-right, labeled Multi-combo2). These simulations support the Chernoff-Savage result and also show that the further away the true distribution is from Gaussian, the more power can be gained by the WZ test. Moreover, case (iv) in the bottom-right subplot shows that actually the power we gain by the WZ test is from the innovation of the first dimension, $\varepsilon^y_t$. When the distribution of $\varepsilon^y_t$ is Gaussian, we do as well as the EMW test. We conjecture that inference for $\beta$ in the predictive regression model ((ref))-((ref)) is adaptive with respect to the marginal density of $\varepsilon^x_t$, when $\gamma$ is eliminated by the ALFD approach in EMW2015.

In Figure (ref) and Figure (ref), we present the powers of the WZ test and the EMW test (for fixed $c = -15$ and for $\delta \in [0, 6]$) under the same settings as in Figure (ref) and Figure (ref), respectively. The results show the power gain of the WZ test over the EMW test uniformly for all alternative values.

We also provide some small-sample ($T=200$) results for both tests in Figure (ref) and Figure (ref) (the small-sample counterparts of Figure (ref) and Figure (ref), respectively). The conclusions are similar: both tests are of good size (all around $4.5\%$) using the same combinations of $\Lambda_{0}^{*\epsilon}$ and $\kappa_g$. The WZ test still gains considerable power in the case of non-Gaussian densities, though the gain is slightly smaller than in the large-sample case. This once more shows the additional information present, when supported by the application at hand, of an i.i.d.-ness assumption on the innovations. Appendix (ref) provides additional simulation results in Figure (ref) and Figure (ref) for Figure (ref) and Figure (ref) using $c = -15$ and $\delta \in [0, 6]$, respectively.

Finally, we repeat the simulations of Figure (ref) and Figure (ref), but for $\rho = -0.9$, in Figure (ref) and Figure (ref) respectively in Appendix (ref). These simulations confirm our previous conclusions about the WZ test: correct sizes, power gain under non-Gaussian $f$, the Chernoff-Savage result, and decent small-sample performances.

Simulations under conditional heteroskedasticity

In many (financial) applications the maintained assumption of i.i.d.\ innovations will not be satisfied. We therefore study, by simulation, the behavior of the tests when the innovations exhibit conditional heteroskedasticity. The tests are identical to those in the previous sections, thus not adapted to deal with possible heteroskedasticity.

Keeping everything else unchanged, we replace the i.i.d.\ innovations $(\varepsilon^y_t, \varepsilon^x_t)^\prime$, by a univariate GARCH(1,1) model (i) for $\varepsilon^y_t$ only; or (ii) for both $\varepsilon^y_t$ and $\varepsilon^x_t$ in the data generating process. Formally, we choose

align*[align* omitted — 131 chars of source]

where, for case (ii), $\varepsilon_{1,t}$ and $\varepsilon_{2,t}$ are independently generated by the GARCH(1,1) model

align*[align* omitted — 120 chars of source]

for $j = 1,2$, where $\nu_{j,t}$'s are i.i.d.\ innovations. For case (i), we let $\varepsilon_{2,t}$ be i.i.d.\ and independent of $\varepsilon_{1,t}$. The joint density of $\nu_{1,t}$ and $\nu_{2,t}$ is denoted by $f$. The GARCH parameters are chosen based on common empirical findings.

figure[figure omitted — 400 chars of source]
figure[figure omitted — 404 chars of source]

In Figure (ref), we present case (i) where only the innovations of the response variable, $\varepsilon^y$, exhibit conditional heteroskedasticity, while the predictor innovations $\varepsilon^x$ are still i.i.d. We show results for three density combinations as mentioned in the title of each subplot and three different values for the correlation of innovations ($\rho = -0.1$, $-0.5$, and $-0.9$). In all nine cases, we find that both the EMW and WZ tests still have decent sizes, i.e., heteroskedasticity appearing only in $\varepsilon^y$ will not affect their size performances much. In terms of power, the WZ test outperforms the EMW test under the $t_3$ distribution, and both tests have similar powers under Gaussianity. In addition, when $\varepsilon^y$ is exhibits more heteroskedasticity (i.e., when $\rho$ is close to 0), the WZ test gains more power as heteroskedasticity pushes the unconditional innovation distribution further away from Gaussianity.

Figure (ref) presents the results for case (ii) where both innovations, $\varepsilon^y$ and $\varepsilon^x$, are heteroskedastic and correlated as modeled above. When $\rho$ is close to zero, the size distortion becomes smaller, while for larger (absolute) values of $\rho$, both tests become more oversized, especially under heavy-tailed innovation distribution. The WZ test suffers less size distortion than the EMW test under $t_3$ distributions and, using $t_3$ reference marginal densities, the size distortion bcomes even smaller (see the bottom panel).

The small-sample counterparts of Figure (ref) and Figure (ref) with $T = 200$ are provided in Appendix (ref). We draw conclusions similar to the i.i.d.\ case in Section (ref). Additionally, we find that, when both $\varepsilon^y$ and $\varepsilon^x$ are heteroskedastic and their correlation is close to $-1$, both the EMW and WZ tests are less over-sized in the small-sample case.

These conclusions above also apply to other GARCH settings with different value chosen for parameters. These simulation results are available upon request.

Simulations under estimated reference density

In this section, we provide simulation results for the WZ test based on nonparametrically estimated reference densities, i.e., $g_y = \hat{f}_y$ and $g_x = \hat{f}_x$, under the i.i.d.\ setting as in Section (ref).

figure[figure omitted — 501 chars of source]
figure[figure omitted — 497 chars of source]

In Figure (ref), we compare the WZ test with $g_y = \hat{f}_y$ and $g_x = \hat{f}_x$ (dotted lines) with the EMW test (dashed lines) and with as the WZ test using correctly specified reference marginal densities (solid lines). When both the true and the reference densities are Gaussian (right plot), we see that all three tests perform similarly with decent size and power properties. When the true innovation distribution is Student-$t_3$, all three tests control the sizes well, while in terms of power, both WZ tests outperform the Gaussian-based EMW test. The WZ test with estimated reference densities suffers a small efficiency loss due to the nonparametric estimation.

Figure (ref) provides the small-sample results under the same setting but with sample size $T = 200$. In general, the smaller sample leads to lower size and power for the WZ test with estimated reference densities relative to the large-sample case. But again, when $f$ is heavy-tailed, it can be more powerful than the EMW test.

Conclusion

In this paper, we show that there is significant statistical information, when supported by the application at hand, in a maintained assumption of serially independent innovations in a predictive regression model. We exploit this information by deriving the (maximal) invariance structures in the associated limit experiment.

Specifically, we first derive the maximal invariant in the (structural) limit experiment where the predictor's persistence parameter is assumed to be known. This leads to the semiparametric power envelope for test that are invariant with respect to the innovation density. The associated likelihood ratio thus gives the semiparametric counterparts of the Gaussian sufficient statistics of JanssonMoreira2006. Under non-Gaussianity, larger powers are possible than under Gaussianity; a well-known result in many classical statistical models. To eliminate the predictor's persistence nuisance parameter, we employ the ALFD approach recently proposed in EMW2015.

Our analysis naturally leads to statistics based on the bivariate component-wise ranks of the innovations in the model. Our statistics involve a choice of reference densities that is, subject to some mild regularity conditions, largely arbitrary. Irrespective of the choice of reference densities, our test are of correct asymptotic size. Under non-Gaussianity, even with incorrectly specified reference densities, our test have better power properties than existing tests in the literature that are derived under the assumption of Gaussian innovation densities. These alternative tests do not need serially independent innovations and, as a result, we precisely quantify the power improvements possible when such an assumption is supported by the data. Monte Carlo simulations corroborate our asymptotic results and illustrate that the rank-based tests also work well in smaller samples.