EconBase
← Back to paper

Semiparametrically Point-Optimal Hybrid Rank Tests for Unit Roots

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

116,508 characters · 24 sections · 67 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametrically Point-Optimal Hybrid Rank Tests for Unit Roots

frontmatter\runtitle{Semiparametrically optimal unit root tests} \begin{aug} , \and \runauthor{Zhou, van den Akker, and Werker.} \end{aug} \begin{abstract} We propose a new class of unit root tests that exploits invariance properties in the Locally Asymptotically Brownian Functional limit experiment associated to the unit root model. The invariance structures naturally suggest tests that are based on the ranks of the increments of the observations, their average, and an assumed reference density for the innovations. The tests are semiparametric in the sense that they are valid, i.e., have the correct (asymptotic) size, irrespective of the true innovation density. For a correctly specified reference density, our test is point-optimal and nearly efficient. For arbitrary reference densities, we establish a Chernoff-Savage type result, i.e., our test performs as well as commonly used tests under Gaussian innovations but has improved power under other, e.g., fat-tailed or skewed, innovation distributions. To avoid nonparametric estimation, we propose a simplified version of our test that exhibits the same asymptotic properties, except for the Chernoff-Savage result that we are only able to demonstrate by means of simulations. \end{abstract} \begin{keyword}[class=MSC] \kwd[Primary ]{62G10}\kwd{62G20} \kwd[; secondary ]{62P20}\kwd{62M10} \end{keyword} \begin{keyword} \kwd{unit root test} \kwd{semiparametric power envelope} \kwd{limit experiment} \kwd{LABF} \kwd{maximal invariant} \kwd{rank statistic} \end{keyword}

Introduction

The monographs of Patterson2011 ({Patterson2011}, Patterson2012) and Choi2015 provide an overview of the literature on unit roots tests. This literature traces back to White58 and includes seminal papers as DickeyFuller79 (DickeyFuller79, DickeyFuller81), Phillips1987, PhillipsPerron88, and ERS1996. The present paper fits into the stream of literature that focuses on “optimal” testing for unit roots. Important earlier contributions here are DufourKing91, SaikkonenLuukkonen1993, and ERS1996. The latter paper derives the asymptotic power envelope for unit root testing in models with Gaussian innovations. RothenbergStock97 and Jansson2008 consider the non-Gaussian case.

The present paper considers testing for unit roots in a semiparametric setting. Following earlier literature, we focus on a simple AR(1) model driven by possibly serially correlated errors. The innovations driving these serially correlated errors are i.i.d., whose distribution is considered a nuisance parameter. Apart from some smoothness and the existence of relevant moments, no assumptions are imposed on this distribution. From earlier work it is known that the unit root model leads to Locally Asymptotically Brownian Functional (LABF) limit experiments Jeganathan1995. As a consequence, no uniformly most powerful test exists (even in case the innovation distribution would be known) -- see also ERS1996. In the semiparametric case the limit experiment becomes more difficult due to the infinite-dimensional nuisance parameter. Jansson2008 derives the semiparametric power envelope by mimicking ideas that hold for Locally Asymptotically Normal (LAN) models. However, the proposed test needs a nonparametric score function estimator which complicates its implementation. The point-optimal tests proposed in the present paper only require nonparametric estimation of a real-valued cross-information factor and we also provide a simplified version that avoids any nonparametric estimation.

The main contribution of this manuscript is twofold. First, we derive the semiparametric power envelopes of unit root tests with serially correlated errors for two cases: symmetric or possibly non-symmetric innovation distributions (Section (ref)). Our method of derivation is novel and exploits the invariance structures embedded in the semiparametric unit root model. To be precise, we use a “structural” description of the LABF limit experiment (Section (ref)), obtained from Girsanov's theorem. This limit experiment corresponds to observing an infinitely-dimensional Ornstein-Uhlenbeck process (on the time interval $[0,1]$). The unknown innovation density in the semiparametric unit root model takes the form of an unknown drift parameter in this limit experiment. Within the limit experiment, Section (ref) derives the maximal invariant, i.e., a reduction of the data which is invariant with respect to the nuisance parameters (that is, the unknown drift in the limiting Ornstein-Uhlenbeck experiment). It turns out that this maximal invariant takes a rather simple form: all processes associated to density perturbations have to be replaced by their associated bridges (i.e., consider $W(s) - sW(1)$ for the process $W(s)$ with $s\in[0, 1]$). The power envelopes for invariant tests in the limit experiment then readily follow from the Neyman-Pearson lemma. An application of the Asymptotic Representation Theorem (see, e.g., Theorem 15.1 in vdVaart00) subsequently yields the local asymptotic power envelope (Theorem (ref)). In case the innovation density is known to be symmetric, the semiparametric power envelope coincides with the parametric power envelope. This implies the existence of an adaptive testing procedure (see also Jansson2008). Moreover, we note that our analysis of invariance structures in the LABF experiment is also of independent interest and could, for example, be exploited in the analysis of optimal inference for cointegration or predictive regression models. Also, the analysis gives an alternative interpretation of the test proposed in ERS1996 --- the ERS test --- as this test is also based on an invariant, though not the maximal one (see Remark (ref)).

As a second contribution, we provide two new classes of easy-to-implement unit root tests that are semiparametrically optimal in the sense that their asymptotic power curves are tangent to the associated semiparametric power envelopes (Section (ref)). The form of the maximal invariant developed before suggests how to construct such tests based on the ranks/signed-ranks (depending on whether the innovation density is known to be symmetric or not) of the increments of the observations, the average of these increments, and an assumed reference density $g$. These tests are semiparametric in the sense that the reference density need not equal the true innovation density, while they are still valid (i.e., provide the correct asymptotic size). The reference density is not restricted to be Gaussian, which it generally needs to be in more classical QMLE results. When the reference density is correctly specified (i.e., $g$ happens to be equal to the true density $f$), the asymptotic power curve of our test is tangent to the semiparametric power envelope, and this in turn gives the optimality property. A feasible version of the oracle test using $g=f$ is obtained by using a nonparametrically estimated density $\hat{f}$, of which the corresponding simulation results are provided in Section (ref).

In relation to the classical literature on efficient rank-based testing (for instance, HajekSidak1967, HallinPuri1988, and HvdAW1) our approach can be interpreted as follows. In the aforementioned papers, the invariance arguments (that is, using the ranks of the innovations) are applied in the sequence of models at hand. We, on the other hand, only apply the invariance arguments in the limit experiment. In this way, we can extend these ideas to non-LAN experiments. For the LAN case, both approaches would effectively lead to the same results. Our tests, despite the absence of a LAN structure, satisfy a ChernoffSavage1958 type result (Corollary (ref)): for any reference density our test outperforms, at any true density, its classical counterpart which, in this case, is the ERS test. We provide, in Section (ref), even simpler alternative classes of tests that require no nonparametric estimations at all. These (simplified) classes of tests coincide with their corresponding originals for correctly specified reference density and, hence, share the same optimality properties. In case of misspecified reference density, the alternative classes still seem to enjoy the Chernoff-Savage type property, though only for a Gaussian reference density. This is in line with with the traditional Chernoff-Savage results for Locally Asymptotically Normal models.

The remainder of this paper is organized as follows. Section (ref) introduces the model assumptions and some notation. Next, Section (ref) contains the analysis of the limit experiment. In particular we study invariance properties in the limit experiment leading to our new derivation of the semiparametric power envelopes. The classes of hybrid rank tests we propose are introduced in Section (ref). Section (ref) provides the results of a Monte Carlo study and Section (ref) contains a discussion of possible extensions of our results. All proofs are organized in (ref).

The model

Consider observations $Y_1,\dots,Y_T$ generated from the classical component specification, for $t\in\mathbb{Z}_{+}$,

align[align omitted — 143 chars of source]

where $v_0=v_{-1}=\cdots=v_{1-p}=0$, the innovations $\{\varepsilon_t\}$ form an i.i.d.\ sequence defined for $t\in\mathbb{Z}$ with density $f$, and $\Gamma(L)$ is the AR(p) lag polynomial. Moreover, it is assumed that $Y_0 = \mu$. We impose the following assumptions on this innovation density.

assumption\begin{enumerate} • The density $f$ is absolutely continuous with a.e. derivative $f^\prime$, i.e., for all $a<b$ we have \begin{align*} f(b)-f(a)=\int_a^b f^\prime(e) \mathrm{d} e. \end{align*} • $\mathrm{E}_{f} [\varepsilon_t]=\int e f(e) \mathrm{d} e=0$ and $\sigma_f^2=\operatorname{Var}_f[\varepsilon_t]<\infty$. • The standardized Fisher-information for location, \begin{align*} J_{f}=\sigma_f^2\int\phi_f^2(e) f(e) \mathrm{d} e, \end{align*} where $\phi_f(e)= -(f^\prime /f)(e)$ is the location score, is finite. • The density $f$ is positive, i.e., $f>0$. \end{enumerate}

The imposed smoothness on $f$ is mild and standard (see, e.g., LeCam1986, vdVaart00). The finite variance assumption (b) is important to our asymptotic results as it is essential to the weak convergence, to a Brownian motion, of the partial-sum process generated by the innovations.\footnote{Let us already mention that, although not allowed for in our theoretical results, we will also assess the finite-sample performances of the proposed tests (Section (ref)) for innovation distributions with infinite variance. For tests specifically developed for such cases we refer to Hasan2001, AhnFotopoulosHe03, and CallegariCappuccioLubian03.} The zero intercept assumption in (b) excludes a deterministic trend in the model. Such a trend leads to an entirely different asymptotic analysis, see HvdAW1. The Fisher information $J_{f}$ in (c) has been standardized by premultiplying with the variance $\sigma_f^2$, so that it becomes scale invariant (i.e., invariant with respect to $\sigma_f$). In other words, $J_{f}$ only depends on the shape of the density $f$ and not on its variance $\sigma_f^2$. The positivity of the density $f$ in (d) is mainly made for notational convenience.

The assumption on the initial condition, $v_0=v_{-1}=\cdots=v_{1-p}=0$, is less innocent then it may appear. Indeed, it is known, see MullerElliott2003 and ElliottMuller2006, that, even asymptotically, the initial condition can contain non-negligible statistical information. Nevertheless, it is still stronger than necessary for the sake of simplicity, and it can be relaxed to the level of generality in ERS1996.

Let $\mathfrak{F}$ denote the set of densities satisfying Assumption (ref). We also investigate in the present paper the special case of symmetric densities. For that purpose, we denote by $\mathfrak{F}_\mathbb{S}$ the set of densities which satisfy Assumption (ref) and, at the same time, are symmetric about zero. Of course, it follows $\mathfrak{F}_\mathbb{S}\subset\mathfrak{F}$.

With respect to the autocorrelation structure $\Gamma(L)$, we impose the following assumption.

assumptionThe lag polynomial $\Gamma(z) := 1-\Gamma_1 z - \cdots - \Gamma_p z^p$ is of finite order $p\in\mathbb{Z}_{+}$ and satisfies $\min_{|z|\in\mathbb{C}:|z|\leq 1}|\Gamma(z)|>0$.

Let $\mathfrak{G}\subset\mathbb{R}^p$ denote the set of $\left(\Gamma_1,\ldots,\Gamma_p\right)'$ such that the induced lag polynomial satisfies Assumption (ref). For mathematical convenience we restrict the lag polynomial to be of finite order $p$. One may expect many of the results in the present paper to extend to the case $p=\infty$ (see, e.g., Jeganathan1997).

The main goal of this paper is to develop tests, with optimality features, for the semiparametric unit root hypothesis

align*[align* omitted — 195 chars of source]

i.e., apart from Assumption (ref)-(ref), no further structure is imposed on $f$, the intercept $\mu$, and the autocorrelation structure $\Gamma(L)$.

In the following section, we derive the (asymptotic) power envelope of tests that are (locally and asymptotically) invariant with respect to the nuisance parameters $\mu$, $\Gamma$, and $f$. We consider both the non-symmetric ($f\in\mathfrak{F}$) and the symmetric case ($f\in\mathfrak{F}_\mathbb{S}$). Section (ref) is subsequently devoted to tests, depending on a reference density $g$ that can be freely chosen, that are point optimal with respect to this power envelope and proves a Chernoff-Savage type result.

The power envelope for invariant tests

This section first introduces some notations and preliminaries (Section (ref)). Afterwards, we will derive the limit experiment (in the Le Cam sense) corresponding to the component unit root model ((ref))-((ref)) and provide a “structural” representation of this limit experiment (Section (ref)). In Section (ref) we discuss, exploiting this structural representation, a natural invariance restriction, to be imposed on tests for the unit root hypothesis with respect to the infinite-dimensional nuisance parameter associated to the innovation density. We derive the maximal invariant and obtain from this the power envelope for invariant tests in the limit experiment. At last, in Section (ref), we exploit the Asymptotic Representation Theorem to translate these results to obtain (asymptotically) optimal invariant test in the sequence of unit root models. Again we consider both the case of unrestricted densities $f\in\mathfrak{F}$ and that of symmetric densities $f\in\mathfrak{F}_\mathbb{S}$.

Preliminaries

We first introduce local reparameterizations for the parameter of interest $\rho$ and the nuisance autocorrelation structure $\Gamma(L)$. Then we discuss a convenient parametrization of perturbations to the innovation density $f$ which we use to deal with the semiparametric nature of the testing problem. These perturbations follow the standard approach of local alternatives in (semiparametric) models commonly used in experiments that are Locally Asymptotically Normal (LAN). We will see that, with respect to all parameters but $\rho$, the model is actually LAN; compare also Remark (ref) below. Moreover, we introduce some partial sum processes that we need in the sequel, as well as their Brownian limits.

Local reparameterizations of $\rho$ and $\Gamma(L)$

It is well-known, and goes back to Phillips1987, ChanWei88 and PhillipsPerron88, that the contiguity rate for the unit root testing problem, i.e., the fastest convergence rate at which it is possible to distinguish (with non-trivial power) the unit root $\rho=1$ from a stationary alternative $\rho<1$, is given by $T^{-1}$. Therefore, in order to compare performances of tests with this proper rate of convergence, we reparametrize the autoregression parameter $\rho$ into its local-to-unity form, i.e.,

equation[equation omitted — 77 chars of source]

The appropriate local reparameterization for the lag polynomial $\Gamma(L)$ is of a traditional form with rate $\sqrt{T}$, i.e.,

align[align omitted — 101 chars of source]

where the local perturbation $\gamma$ is defined by $\gamma(L) := - \gamma_1 L - \cdots - \gamma_p L^p$ with local parameter $\gamma := (\gamma_1,\dots,\gamma_p)'\in\mathbb{R}^p$. As $\mathfrak{G}$ is open, $\Gamma^{(T)}_\gamma \in \mathfrak{G}$ for $T$ large enough.

Perturbations to the innovation density

To describe the local perturbations to the density $f$, we need the separable Hilbert space

align*[align* omitted — 225 chars of source]

where $\mathrm{L}_2^f(\mathbb{R},\mathcal{B})$ denotes the space of Lebesgue-measurable functions $b:\,\mathbb{R}\to\mathbb{R}$ satisfying $\int b^2(e) f(e)\mathrm{d} e<\infty$. Because of the separability, there exists a countable orthonormal basis $b_k$, $k\in\mathbb{N}$, of $\mathrm{L}_2^{0,f}$ (see, e.g., Rudin1987). This basis can be chosen such that $b_k \in C_{2,b}(\mathbb{R}) $, for all $k$, i.e., each $b_k$ is bounded and two times continuously differentiable with bounded derivatives. Moreover, $\mathrm{E}_{f}b_k(\varepsilon)=0$ and $\operatorname{Var}_f{b_k(\varepsilon)}=1$. Hence each function $b\in \mathrm{L}_2^{0,f}$ can be written as $b= \sum_{k=1}^\infty \eta_k b_k $, for some $\eta := (\eta_k)_{k\in\mathbb{N}} \in \ell_2=\{ (x_k)_{k\in\mathbb{N}}\, | \, \sum_{k=1}^\infty x_k^ 2<\infty\}$. Besides the sequence space $\ell_2$ we also need the sequence space $c_{00}$ which is defined as the set of sequences with finite support, i.e., \[ c_{00}=\left\{ (x_k)_{k\in\mathbb{N}}\in\mathbb{R}^\mathbb{N} \, \left|\, \sum_{k=1}^\infty 1\{x_k\neq 0\}<\infty\right. \right\}. \] Of course, $c_{00}$ is a dense subspace of $\ell_2$. Given the orthonormal basis $b_k$ and $\eta\in c_{00}$, we introduce the following perturbation to the density $f$:

align[align omitted — 146 chars of source]

The rate $T^{-1/2}$ is already indicative of the standard LAN behavior of the nuisance parameter $f$ as will formally follow from Proposition (ref) below.

For the symmetric case, we can assume that the perturbations $b_k$, $k\in\mathbb{N}$, are also symmetric about zero.

The following proposition shows that the perturbations, both for the non-symmetric and for the symmetric case, are valid in the sense that they satisfy the conditions on the innovation density that we imposed throughout on the model (Assumption (ref)). The proof is organized in (ref).

propositionLet $f\in\mathfrak{F}$ and suppose $\eta\in c_{00}$. Then there exists $T^\prime\in\mathbb{N}$ such that for all $T\geq T^\prime$ we have $f^{(T)}_\eta\in\mathfrak{F}$. If we further restrict $f\in\mathfrak{F}_\mathbb{S}$ and $b_k$ is chosen symmetric about zero for $k\in\mathbb{N}$, then there exists $T^{\prime\prime}\in\mathbb{N}$ such that for all $T\geq T^{\prime\prime}$ we have $f^{(T)}_\eta\in\mathfrak{F}_\mathbb{S}$.
remarkIn semiparametric statistics one typically parametrizes perturbations to a density by a so-called “non-parametric” score function $h\in\mathrm{L}_2^{0,f}$, i.e., the perturbation takes the form $f(e) k(T^{-1/2} h(e)))\approx f(e)(1+T^{-1/2} h(e))$ for a suitable function $k$; see, for example, BKRW for details. By using the basis $b_k$, $k\in\mathbb{N}$, we instead tackle all such perturbations simultaneously via the infinite-dimensional nuisance parameter $\eta$. Of course, one would need to use $\ell_2$ as parameter space to “generate” all score functions $h$. We instead restrict to $c_{00}$ which ensures ((ref)) to be a density (for large $T$). For our purposes this restriction will be without cost. Intuitively, this is since $c_{00}$ is a dense subspace of $\ell_2$ (so if a property is “sufficiently continuous” one only needs to establish it on $c_{00}$ because it extends to the closure).

Partial sum processes

To describe the limit experiment in Section (ref), we introduce some partial sum processes and their limits. These results are fairly classical but, for completeness, precise statements are organized in Lemma A.1 in the supplementary material.

Define, for $s\in [0,1]$, the partial sum processes

align*[align* omitted — 547 chars of source]

Note that we pick the starting point of the sums at $t=p+2$, so that these partial sum processes are (maximally) invariant with respect to the intercept $\mu$ (otherwise, e.g., for $t=p+1$, the term $\Gamma(L)\Delta Y_t$ contains $\Delta Y_1 = Y_1 - \mu$).

Using Assumption (ref) we find, under the null hypothesis, joint weak convergence of observation processes $W^{(T)}_{\varepsilon}$, $W^{(T)}_{\phi_f}$, $W^{(T)}_{\Gamma}$, and $W^{(T)}_{b_k}$ to Brownian motions that we denote by $W_{\varepsilon}$, $W_{\phi_f}$, $W_{\Gamma}$ and $W_{b_k}$, respectively.\footnote{All weak convergences in this paper are in product spaces of $D[0,1]$ with the uniform topology.} These limiting Brownian motions are defined on a probability space $(\Omega,\mathcal{F},\mathbb{P}_{0,0,0})$. Let us already mention that we will introduce a collection of probability measures $\mathbb{P}_{h,\gamma,\eta}$, on $(\Omega, \mathcal{F})$, representing the limit experiment, in Section (ref). We use the notational convention that probability measures related to the limit experiment (i.e., to the “W-processes”) are denoted by $\mathbb{P}$, while probability measures related to the finite-sample unit root model, i.e., observing $Y_1, \dots ,Y_T$, will be denoted by $\mathrm{P}^{(T)}_{}$.

We remark that integrals like $\int_0^1W_{\varepsilon}^{(T)}(s-) \rdW_{\phi_f}^{(T)}(s)$ can be shown to converge weakly, under the null hypothesis, to the associated stochastic integral with the limiting Brownian motions, i.e., to $\int_0^1 W_{\varepsilon}(s)\rdW_{\phi_f}(s)$. Weak convergence of integrals like $\int_0^1(W_{\varepsilon}^{(T)}(s-))^2 \mathrm{d} s$ follows from an application of the continuous mapping theorem. Again, details can be found in the proof of Proposition (ref) in the (ref).

Behavior of $W_{\varepsilon}$, $W_{\phi_f}$, $W_{\Gamma}$ and $W_{b_k}$ under $\mathbb{P}_{0,0,0}$ when $f\in\mathfrak{F}$

As $\varepsilon$ and $b_k(\varepsilon)$ are orthogonal for each $k$, it holds that $W_{\varepsilon}$ and $W_{b_k}$, $k\in\mathbb{N}$, are all mutually independent. Moreover, we have

align*[align* omitted — 124 chars of source]

As $\phi_f(\varepsilon)$ is the score of the location model, it is well known (see, for example, BKRW) that we have (under Assumption (ref)) $\mathrm{E}_{f}[\phi_f(\varepsilon)]=0$ and $\mathrm{E}_{f}[\varepsilon\phi_f(\varepsilon)] =1$. Consequently, for $f\in \mathfrak{F}$, again because $\varepsilon$ and $b_k(\varepsilon)$ are orthogonal for each $k$, we can decompose

align[align omitted — 143 chars of source]

with coefficients $J_{f,k}:=\sigma_f\mathrm{E}_{f}[b_k(\varepsilon)\phi_f(\varepsilon)]$. This establishes

align[align omitted — 106 chars of source]

Moreover, we have, for $k\in\mathbb{N}$,

align[align omitted — 171 chars of source]

and

align[align omitted — 123 chars of source]

As for $W_{\Gamma}$: since, under the null, $\left(\Delta Y_{t-1},\dots,\Delta Y_{t-p}\right)'$ is independent of the innovation $\varepsilon_t$ and $\mathrm{E}_{0,0,0}\left[\Delta Y_{t-i}\right]=0$, $i=1,\dots,p$, it follows

align[align omitted — 164 chars of source]

and

align[align omitted — 110 chars of source]

We define the covariance matrix of $W_{\Gamma}(1)$ as

align[align omitted — 119 chars of source]

with

align*[align* omitted — 155 chars of source]

Behavior of $W_{\varepsilon}$, $W_{\phi_f}$, $W_{\Gamma}$ and $W_{b_k}$ under $\mathbb{P}_{0,0,0}$ when $f\in\mathfrak{F}_\mathbb{S}$

In this case, the density function $f(\varepsilon)$ is an even function and so are the perturbation functions $b_k(\varepsilon)$, $k\in\mathbb{N}$. The score $\phi_f(\varepsilon)$ is now an odd function. Therefore, $\phi_f(\varepsilon)$ cannot be decomposed by $\varepsilon$ and $b_k(\varepsilon)$ anymore as in ((ref)) and ((ref)). Instead of ((ref)) we now have

align*[align* omitted — 83 chars of source]

for all $f\in\mathfrak{F}_\mathbb{S}$ and $k\in\mathbb{N}$. All the other results mentioned above still hold.

A structural representation of the limit experiment

The results in the previous section are needed to study the asymptotic behavior of log-likelihood ratios. These in turn determine the limit experiment, which we use to study asymptotically optimal procedures invariant with respect to the nuisance parameters. Thus, fix $\mu\in\mathbb{R}$, $\Gamma\in\mathfrak{G}$, and $f\in\mathfrak{F}$. Let, for $h\in\mathbb{R}$, $\gamma\in\mathbb{R}^p$, and $\eta\in c_{00}$, $\mathrm{P}^{(T)}_{h,\gamma,\eta;\mu,\Gamma,f}$ denote the law of $Y_1,\dots,Y_T$ under ((ref))-((ref)) with parameter $\rho$ given by ((ref)), $\Gamma^{(T)}_\gamma(L)$ given by ((ref)), and innovation density ((ref)). The following proposition shows that the semiparametric unit root model is of the Locally Asymptotically Brownian Functional (LABF) type introduced in Jeganathan1995.

propositionLet $\mu\in\mathbb{R}$, $\Gamma\in\mathfrak{G}$, $f\in \mathfrak{F}$, $h\in\mathbb{R}$, $\gamma\in\mathbb{R}^p$, and $\eta\in c_{00}$. Let $\Delta$ denote differencing, i.e., $\Delta Y_t=Y_t-Y_{t-1}$. \begin{enumerate} • Then we have, under $\mathrm{P}^{(T)}_{0,0,0;\mu,\Gamma,f}$, \begin{align*} \log \frac{\mathrm{d} \mathrm{P}^{(T)}_{h,\gamma,\eta;\mu,\Gamma,f}}{\mathrm{d} \mathrm{P}^{(T)}_{0,0,0;\mu,\Gamma,f}} =& \sum_{t=1}^{T} \log \frac{ f_\eta^{(T)}\left(\Gamma^{(T)}_\gamma(L)\left(\Delta Y_t -\frac{h}{T} (Y_{t-1}-\mu)\right)\right)}{f\left(\Gamma(L)(\Delta Y_t)\right)} \\ =& h\Delta^{(T)}_{f} + \gamma'\Delta^{(T)}_\Gamma + \sum_{k=1}^\infty \eta_k \Delta^{(T)}_{b_{k}} - \frac{1}{2} \mathcal{I}^{(T)}_(h,\gamma,\eta) + o_P(1),\nonumber \end{align*} where the central-sequence $\Delta^{(T)} = \big(\Delta^{(T)}_{f},\Delta^{(T)}_\Gamma,\Delta^{(T)}_{b_{}}\big)$, with $\Delta^{(T)}_{b_{}} = \big(\Delta^{(T)}_{b_{k}}\big)_{k\in\mathbb{N}}$, is given by \begin{align*} \Delta^{(T)}_{f} & = \frac{1}{T}\sum_{t=p+2}^T \Gamma(L)(Y_{t-1}-Y_1) \phi_f(\Gamma(L)\Delta Y_t), \\ \Delta^{(T)}_\Gamma & = \frac{1}{\sqrt{T}}\sum_{t=p+2}^{T}\left(\Delta Y_{t-1},\dots,\Delta Y_{t-p}\right)' \phi_f(\Gamma(L)\Delta Y_t), \\ \Delta^{(T)}_{b_{k}} & = \frac{1}{\sqrt{T}}\sum_{t=p+2}^T b_k(\Gamma(L)\Delta Y_t), \quad k\in\mathbb{N}, \end{align*} and \begin{align*} \mathcal{I}^{(T)}_(h,\gamma,\eta) = & h^2 J_{f}\frac{1}{T^2}\sum_{t=p+2}^T \frac{\left(\Gamma(L)(Y_{t-1}-Y_1)\right)^2}{\sigma_f^2} + \gamma'\Sigma_\Gamma\gamma + \|\eta\|_2^2 \\ & + 2h\frac{1}{T^{3/2}}\sum_{t=p+2}^T \frac{\Gamma(L)(Y_{t-1}-Y_1)}{\sigma_f} \sum_{k=1}^\infty \eta_k J_{f,k}. \end{align*} • Moreover, with $\Delta_{f}=\int_0^1W_{\varepsilon}(s)\rdW_{\phi_f}(s)$, $\Delta_\Gamma = W_{\Gamma}(1)$, and $\Delta_{b_{k}}=W_{b_k}(1)$, $k\in\mathbb{N}$, we have, still under $\mathrm{P}^{(T)}_{0,0,0;\mu,\Gamma,f}$ and as $T\to\infty$, \begin{align} \frac{\mathrm{d} \mathrm{P}^{(T)}_{h,\gamma,\eta;\mu,\Gamma,f}}{\mathrm{d} \mathrm{P}^{(T)}_{0,0,0;\mu,\Gamma,f}} \Rightarrow \exp\left(h \Delta_{f} + \gamma'\Delta_\Gamma + \sum_{k=1}^\infty \eta_k \Delta_{b_{k}} -\frac{1}{2}\mathcal{I}(h,\gamma,\eta) \right), \end{align} where \begin{align*} \mathcal{I}(h,\gamma,\eta) =& h^2 J_{f}\int_0^1 (W_{\varepsilon}(s))^2 \mathrm{d} s + \gamma'\Sigma_\Gamma\gamma + \|\eta\|_2^2 \\ & + 2 h \int_0^1 W_{\varepsilon}(s) \mathrm{d} s \sum_{k=1}^\infty \eta_k J_{f,k}. \end{align*} • For all $h\in\mathbb{R}$, $\gamma\in\mathbb{R}^p$ and $\eta \in c_{00}$ the right-hand side of ((ref)) has unit expectation under $\mathbb{P}_{0,0,0}$. \end{enumerate}

Of course, Proposition (ref) still holds for $f\in\mathfrak{F}_\mathbb{S}$; in that case we have $J_{f,k}=0$. The proof of (i) follows by an application of Proposition 1 in HvdAW2 which provides generally applicable sufficient conditions for the quadratic expansion of log likelihood ratios. Of course, Part (ii) is not surprising and follows using the weak convergence of the partial sum processes to Brownian motions (and integrals involving the partial sum processes to stochastic integrals) discussed above. Finally, Part (iii) follows by verifying the Novikov condition. For the sake of completeness, detailed proofs are organized in (ref).

Part (iii) of the proposition implies that we can introduce, for $h\in\mathbb{R}$, $\gamma\in\mathbb{R}^p$, and $\eta\in c_{00}$, new probability measures $\mathbb{P}_{h,\gamma,\eta}$ on the measurable space $(\Omega,\mathcal{F})$ (on which the processes $W_{\varepsilon}$, $W_{\phi_f}$, $W_{\Gamma}$ and $W_{b_k}$ were defined) by their Radon-Nikodym derivatives with respect to $\mathbb{P}_{0,0,0}$:

align*[align* omitted — 227 chars of source]

Proposition (ref) then implies that the sequence of (local) unit root experiments (each $T\in\mathbb{N}$ yields an experiment) weakly converges (in the Le Cam sense) to the experiment described by the probability measures $\mathbb{P}_{h,\gamma,\eta}$. Formally, we define the sequence of experiments of interest by

align*[align* omitted — 211 chars of source]

for $T\in\mathbb{N}$, and the limit experiment by, with $\mathcal{B}_C$ the Borel $\sigma$-field on $C[0,1]$,

align*[align* omitted — 338 chars of source]

Note that the latter experiment indeed depends on $f$ as the measure $\mathbb{P}_{h,\gamma,\eta}$ depends on $f$.

corollaryLet $\mu\in\mathbb{R}$, $\Gamma\in\mathfrak{G}$, and $f\in\mathfrak{F}$. Then the sequence of experiments $\mathcal{E}^{(T)}(\mu,\Gamma,f)$, $T\in\mathbb{N}$, converges to the experiment $\mathcal{E}(f)$ as $T\to\infty$.

The Asymptotic Representation Theorem (see, for example, Chapter 9 in vdVaart00) implies that for any statistic $A_T$ which converges in distribution to the law $L_{h,\gamma,\eta}$, under $\mathrm{P}^{(T)}_{h,\gamma,\eta;\mu,\Gamma,f}$, there exists a (possibly randomized) statistic $A$, defined on $\mathcal{E}(f)$, such that the law of $A$ under $\mathbb{P}_{h,\gamma,\eta}$ is given by $L_{h,\gamma,\eta}$. This allows us to study (asymptotically) optimal inference: the “best” procedure in the limit experiment also yields a bound for the sequence of experiments. If one is able to construct a statistic (for the sequence) that attains this bound, it follows that the bound is sharp and the statistic is called (asymptotically) optimal. This is precisely what we do: Section (ref) establishes the bound and in Section (ref) we introduce statistics attaining it.

To obtain more insight in the limit experiment $\mathcal{E}(f)$ the following proposition, which follows by a direct application of Girsanov's theorem, provides a “structural” description of the limit experiment.

propositionLet $h\in\mathbb{R}$, $\gamma\in\mathbb{R}^p$, $\eta\in c_{00}$, and $f \in \mathfrak{F}$. The processes $Z_{\varepsilon}$, $Z_{\phi_f}$, $Z_{\Gamma}$ and $Z_{b_k}$, $k\in\mathbb{N}$, defined by the starting values $Z_{\varepsilon}(0)=Z_{\phi_f}(0)=Z_{b_k}(0)=0$, $Z_{\Gamma}(0) = \boldsymbol{0}_{p}$, and the stochastic differential equations, for $s\in[0,1]$, \begin{align*} \mathrm{d} Z_{\varepsilon}(s) &= \mathrm{d} W_{\varepsilon}(s) - h W_{\varepsilon}(s)\mathrm{d} s, \\ \mathrm{d} Z_{\phi_f}(s) &= \mathrm{d} W_{\phi_f}(s) - h J_{f}W_{\varepsilon}(s)\mathrm{d} s - \textstyle\sum_{k}\eta_kJ_{f,k}\mathrm{d} s, \\ \mathrm{d} Z_{\Gamma}(s) &= \mathrm{d} W_{\Gamma}(s) - \gamma \mathrm{d} s, \\ \mathrm{d} Z_{b_k}(s) &= \mathrm{d} W_{b_k}(s) - h J_{f,k} W_{\varepsilon}(s)\mathrm{d} s - \eta_k\mathrm{d} s, \quad k\in\mathbb{N}, \end{align*} are zero-drift Brownian motions under $\mathbb{P}_{h,\gamma,\eta}$. Their joint law is that of $(W_{\varepsilon},W_{\phi_f},W_{\Gamma},(W_{b_k})_{k\in\mathbb{N}})$ under $\mathbb{P}_{0,0,0}$.

For the case $f\in\mathfrak{F}_\mathbb{S}$, Proposition (ref) still applies with $J_{f,k}=0$. Moreover, for this case, we denote by $\mathcal{E}^{(T)}_{\mathbb{S}}(f)$ the associated sequence of experiments and by $\mathcal{E}_\mathbb{S}(f)$ the associated limit experiment.

remarkPart (i) and (ii) Proposition (ref) show that the parameter $\mu$ vanishes in the limit. More explicitly, in the proof of this proposition, we replace $\mu$ in the likelihood ratio term by $Y_1$ and then show that the difference term is $o_P(1)$. On the other hand, one could also “localize” the parameter $\mu$ as $\mu=\mu^{(T)}_d=\mu_0+d$ (with rate $T^0=1$) as in Jansson2008. As shown in that paper, the term associated to the parameter $d$ does not change with $T$ and is independent of the other terms of the likelihood ratio. By the additively separable structure, we can treat the parameter $\mu$ “as if” it is known. In either way, inference for $\rho$ would be invariant with respect to $\mu$ in the limit. Analogously, in the finite-sample experiment $\mathcal{E}^{(T)}(f)$, $\mu$ is eliminated (automatically) by using the increments $\Delta Y_t$, $t=2,\dots,T$, which are (maximally) invariant with respect to $\mu$ (see Section (ref)).

The limit experiment: invariance and power envelope

In this section, we consider the limit experiments $\mathcal{E}(f)$ and $\mathcal{E}_\mathbb{S}(f)$. In these experiments, we observe the processes $W_{\varepsilon}$, $W_{\phi_f}$, $W_{\Gamma}$, and $W_{b_k}$, $k\in\mathbb{N}$, continuously on the time interval $[0,1]$ from the model $(\mathbb{P}_{h,\gamma,\eta}\, |\, h\in\mathbb{R},\,\gamma\in\mathbb{R}^p,\,\eta\in c_{00})$, and we are interested in the power envelopes for testing the hypothesis

equation[equation omitted — 179 chars of source]

To eliminate the nuisance parameters $\gamma$ and $\eta$, we first propose a statistic that is sufficient for the parameter of interest $h$ and does not depend on $\gamma$. Afterwards, using Proposition (ref), we discuss a natural invariance structure with respect to the infinite-dimensional nuisance parameter $\eta$. We derive the maximal invariant and apply the Neyman-Pearson lemma to obtain the power envelopes of invariant tests within the experiments $\mathcal{E}(f)$ and $\mathcal{E}_\mathbb{S}(f)$, respectively in Section (ref) and Section (ref).

We begin with the elimination of $\gamma$. The statistic $\left(W_{\varepsilon},W_{\phi_f},(W_{b_k})_{k\in\mathbb{N}}\right)$ serves as a sufficient statistic for the parameter $h$. This is because, according to the structural version of $\mathcal{E}(f)$ (or $\mathcal{E}_\mathbb{S}(f)$) in Proposition (ref), the distribution of $W_{\Gamma}$ is only affected by $\gamma$ and so is the distribution of the process $\left(W_{\varepsilon},W_{\phi_f},(W_{b_k})_{k\in\mathbb{N}}\right)$ only affected by $h$ and $\eta$. It then follows that the distribution of the statistic $\left(W_{\varepsilon},W_{\phi_f},W_{\Gamma},(W_{b_k})_{k\in\mathbb{N}}\right)$ conditional on the statistic $\left(W_{\varepsilon},W_{\phi_f},(W_{b_k})_{k\in\mathbb{N}}\right)$ is only a function of $W_{\Gamma}$ and $\gamma$ and, in turn, does not depend on $h$. Next, observe that the sufficient statistic $\left(W_{\varepsilon},W_{\phi_f},(W_{b_k})_{k\in\mathbb{N}}\right)$ is independent of the process $W_{\Gamma}$ and thus has a distribution that does not depend on $\gamma$. This allows us to restrict attention to the sufficient statistic $\left(W_{\varepsilon},W_{\phi_f},(W_{b_k})_{k\in\mathbb{N}}\right)$ to conduct inference for $h$, and the nuisance parameter $\gamma$ disappears from the likelihoods.

Elimination of $\eta$ in $\mathcal{E}(f)$ and the associated power envelope

The elimination of the nuisance parameter $\eta$ is more involved and different for the limit experiments $\mathcal{E}(f)$ and $\mathcal{E}_\mathbb{S}(f)$. We start with $\mathcal{E}(f)$.

In Proposition (ref), if one applies the decompositions ((ref)) and ((ref)) to the first equation (of $W_{\varepsilon}$) and the fourth equation (of $W_{b_k}$), one retrieves the second equation (of $W_{\phi_f}$). This essentially allows us to omit the process $W_{\phi_f}$ and restrict the observations to the processes $W_{\varepsilon}$ and $W_{b_k}$, $k\in\mathbb{N}$.

Now we formalize the invariance structure with respect to $\eta$. Introduce, for $\eta\in c_{00}$, the transformation $g_\eta=(g_{\eta_k})_{k\in\mathbb{N}}: C^\mathbb{N}[0,1]\to C^\mathbb{N}[0,1]$ defined by, for $W\in C[0,1]$,

equation[equation omitted — 115 chars of source]

i.e., $g_{\eta_k}$ adds a drift $s\mapsto -\eta_k s$ to its argument process. Proposition (ref) implies that the law of $(W_{\varepsilon},(g_{\eta_k}(W_{b_k}))_{k\in\mathbb{N}})$ under $\mathbb{P}_{h,\gamma,0}$ is the same as the law of $(W_{\varepsilon},(W_{b_k})_{k\in\mathbb{N}})$ under $\mathbb{P}_{h,\gamma,\eta}$. Hence our testing problem ((ref)) is invariant with respect to the transformations $g_\eta$. Therefore, following the invariance principle, it is natural to restrict attention to test statistics that are invariant with respect to these transformations as well, i.e., test statistics $t$ that satisfy

equation[equation omitted — 190 chars of source]

Given a process $W$ let us define the associated bridge process by $B^W(s)=W(s)- s W(1)$. Now note that we have, for all $s\in [0,1]$ and $k\in\mathbb{N}$,

align*[align* omitted — 158 chars of source]

i.e., taking the bridge of a process ensures invariance with respect to adding drifts to that process. Define the mapping $M$ by

align*[align* omitted — 108 chars of source]

where $B_{b_k} := B^{W_{b_k}}$. It follows that statistics that are measurable with respect to the $\sigma$-field,

align[align omitted — 176 chars of source]

are invariant (with respect to $g_\eta$, $\eta\in c_{00}$). It is, however, not (immediately) clear that we did not throw away too much information. Formally, we need $\mathcal{M}$ to be maximally invariant which means that each invariant statistic is $\mathcal{M}$-measurable. The following theorem, which once more exploits the structural description of the limit experiment, shows that this indeed is the case.

theoremThe $\sigma$-field $\mathcal{M}$ in ((ref)) is maximally invariant for the group of transformations $g_\eta$, $\eta\in c_{00}$, in the experiment $\mathcal{E}(f)$.

The above theorem implies that invariant inference must be based on $\mathcal{M}$. An application of the Neyman-Pearson lemma, using $\mathcal{M}$ as observation, yields the power envelope for the class of invariant tests. To be precise, consider the likelihood ratios restricted to $\mathcal{M}$, which are given by

align*[align* omitted — 233 chars of source]

where the conditional expectation indeed does not depend on $\eta$ precisely because of the invariance. The conditional expectation also does not depend on $\gamma$ as a result of the arguments stated at the beginning of this section.

To calculate this conditional expectation, we first introduce $B_{\phi_f}=B^{W_{\phi_f}}$, i.e., the bridge process associated to $W_{\phi_f}$. Following the decomposition in ((ref)), we can decompose $\Delta_{f}=\int_0^1W_{\varepsilon}(s)\rdW_{\phi_f}(s)=I+II$ with

align*[align* omitted — 221 chars of source]

Note that part $I$ is $\mathcal{M}$-measurable. Under $\mathbb{P}_{0,0,0}$ the random variables $W_{b_k}(1)$, $k\in\mathbb{N}$, are independent of $W_{\varepsilon}$ and $B_{b_k}$, $k\in\mathbb{N}$. Indeed, the independence of $W_{\varepsilon}$ holds by construction and the independence of $B_{b_k}$ is a well-known, and easy to verify, property of Brownian bridges. We thus obtain, since $\mathcal{I}{}(h,\gamma,\eta)$ is $\mathcal{M}$-measurable as well,

align*[align* omitted — 802 chars of source]

with

align[align omitted — 507 chars of source]

where the last equality follows from ((ref)). Note that the distribution of this likelihood ratio indeed does not depend on the nuisance parameters $\gamma$ or $\eta$.

We can now formalize the notion of point-optimal invariant tests in the limit experiment. To that end, let us denote the $(1-\alpha)$-quantile of $\mathrm{d}\mathbb{P}_{h}^{\mathcal{M}}/\mathrm{d}\mathbb{P}_{0}^{\mathcal{M}}$ under $\mathbb{P}_{0,\gamma,\eta}$, which does not depend on either $\gamma$ or $\eta$, by $c(h,J_f;\alpha)$. Define the size-$\alpha$ test $\phi^{\mathfrak{F}*}_{f,\alpha}(\bar{h}) = \mathbbm{1}\left\{\mathrm{d}\mathbb{P}_{\bar{h}}^{\mathcal{M}} / \mathrm{d} \mathbb{P}_{0}^{\mathcal{M}}\geq c(\bar{h},J_f;\alpha)\right\}$, for a fixed value of $\bar{h}<0$. Note that this is an oracle test in $\mathcal{E}(f)$ that depends on the true value of $f$. A feasible test is provided in Section (ref). The power function of this oracle test is given by

align*[align* omitted — 365 chars of source]

An application of the Neyman-Pearson lemma yields the following.

corollaryLet $f\in\mathfrak{F}$ and $\alpha\in (0,1)$. Let $\phi$ be a (possibly randomized) test that is $\mathcal{M}$-measurable and is of size $\alpha$, i.e., $\mathbb{E}_{0} \phi \leq \alpha$. Let $\pi$ denote the power function of this test, i.e., $\pi(h)=\mathbb{E}_{h} \phi$. Then we have \begin{align*} \pi(\bar{h})\leq \pi^{\mathfrak{F}*}_{f,\alpha}(\bar{h};\bar{h}), \bar{h}<0. \end{align*}

The (oracle) test $\phi^{\mathfrak{F}*}_{f,\alpha}(\bar{h})$ thus is point optimal, i.e., its power function is tangent to the (semiparametric) power envelope $h\mapsto \pi^{\mathfrak{F}*}_{f,\alpha}(h; h)$ at $h=\bar{h}$.\footnote{Here and later in this section, the early usage of the concept “power envelope” is due to the fact that it is shown to be the upper bound later in this section and point-wisely attainable by tests in sequence in Section (ref).}

remarkThe notion of invariance in the limit experiment leads to another interpretation of the ERS1996 (ERS) test statistic. Note that $\sigma$-field $\mathcal{M}_{\varepsilon}=\sigma\left(W_{\varepsilon}(s);\, s\in [0,1]\right)$ is also invariant, though not maximally so. The likelihood ratio conditional on observing only $\mathcal{M}_{\varepsilon}$ is given by \begin{align*} \frac{\mathrm{d} \mathbb{P}_{h}^{\mathcal{M}_{\varepsilon}}}{\mathrm{d} \mathbb{P}_{0}^{\mathcal{M}_{\varepsilon}}} =& \mathbb{E}_0 \left[ \frac{\mathrm{d} \mathbb{P}_{h}^\mathcal{M} } {\mathrm{d} \mathbb{P}_{0}^\mathcal{M}} \mid \mathcal{M}_{\varepsilon} \right]\\ =& \exp\left( h\int_0^1W_{\varepsilon}(s) \mathrm{d} B_{\varepsilon}(s)+ h W_{\varepsilon}(1) \int_0^1 W_{\varepsilon}(s) \mathrm{d} s -\frac{1}{2} h^2 \mathcal{I}^*\right)\\ & \times\mathbb{E}_0 \left[ \exp\left(h\int_0^1W_{\varepsilon}(s)\mathrm{d} B_b(s)\right) \mid \mathcal{M}_{\varepsilon} \right]\\ =& \exp\left( h\int_0^1W_{\varepsilon}(s) \mathrm{d} W_{\varepsilon}(s) - \frac{1}{2} h^2 \mathcal{I}^*\right)\\ & \times \exp\left(\frac12h^2\left[\int_0^1W_{\varepsilon}^2(s)\mathrm{d} s - \left(\int_0^1W_{\varepsilon}(s)\mathrm{d} s\right)^2\right]\left(J_f-1\right)\right)\\ =& \exp\left( h\int_0^1 W_{\varepsilon}(s) \mathrm{d} W_{\varepsilon}(s) -\frac{1}{2}h^2\int_0^1W_{\varepsilon}^2(s)\mathrm{d} s\right), \end{align*} where $W_{b}(s) = \sum_{k=1}^\infty J_{f,k}W_{b_k}(s) $ and $B_{b}(s) = B^{W_{b}}(s)$ for notational simplicity. As a result, the ERS test statistic equals the likelihood ratio statistic using the (non-maximal) invariant $\mathcal{M}_{\varepsilon}$. This explains the improved power of our tests within the model we consider. Moreover, for Gaussian $f$, we have $\mathcal{M}_{\varepsilon}=\mathcal{M}$ and obtain point-optimality of the ERS test.\footnote{Similarly, one could try to derive the statistic resulting from using $\mathcal{M}_{B}=\sigma\left(B_{b_k}(s);\, s\in [0,1]\right)$ as an invariant. However, that does not seem to lead to an insightful result.}
remarkThe semiparametric power envelope derived above for this case of $f\in\mathfrak{F}$, of course, coincides with the one in Jansson2008 based on the invariance constraint. This can be seen by rewriting $\Delta^*_f=\int_0^1 W_{\varepsilon}(s) \mathrm{d} W_{\phi_f}(s) - (W_{\phi_f}(1) - W_{\varepsilon}(1)) \int_0^1 W_{\varepsilon}(s) \mathrm{d} s$. We feel our approach is attractive since, by describing the perturbations on $f$ with an orthonormal basis and a infinite-dimensional parameter instead of one single parameter, there is no need to find the least favorable direction.\footnote{The traditional semiparametric approach, developed for LAN-type experiments, accomplishes this by projecting the score function of the parameter of interest onto the tangent space of nuisance score functions. However, this approach seems not easily generalizable to LABF-type experiments.} Moreover, we feel that the use of the invariance principle, rather than the similarity constraint, more naturally suggests (partly) rank-based tests.

Elimination of $\eta$ in $\mathcal{E}_\mathbb{S}(f)$ and the associated power envelope

Since in this case we have $J_{f,k}=0$, the structural representation of experiment $\mathcal{E}_\mathbb{S}(f)$ in Proposition (ref) becomes

align*[align* omitted — 299 chars of source]

Note that here the process $W_{\Gamma}$ is again removed by sufficiency in order to eliminate the nuisance parameter $\gamma$ (see the discussion at the beginning of Section (ref)). Following the same argument for the case of $f\in\mathfrak{F}$, statistics that are measurable with respect to the $\sigma$-field

align[align omitted — 222 chars of source]

are invariant with respect to the transformations $g_\eta$, $\eta\in c_{00}$. Moreover, in the following theorem, we show that this $\sigma$-field is maximally invariant.

theoremThe $\sigma$-field $\mathcal{M}^{\mathbb{S}}$ in ((ref)) is maximally invariant for the group of transformations $g_\eta$, $\eta\in c_{00}$, in the experiment $\mathcal{E}_\mathbb{S}(f)$.

The likelihood ratio restricted to $\mathcal{M}^{\mathbb{S}}$ is given by

align*[align* omitted — 736 chars of source]

where

align[align omitted — 111 chars of source]

As before, we then formalize the point-optimal invariant tests in $\mathcal{E}_\mathbb{S}(f)$. Denote the $(1-\alpha)$-quantile of $\mathrm{d}\mathbb{P}_{h}^{\mathcal{M}^{\mathbb{S}}}/\mathrm{d}\mathbb{P}_{0}^{\mathcal{M}^{\mathbb{S}}}$ under $\mathbb{P}_{0,\gamma,\eta}$ by $c_{\mathbb{S}}(h,J_f;\alpha)$. Define the size-$\alpha$ test $\phi^{\mathfrak{F}_\mathbb{S}*}_{f,\alpha}(\bar{h}) = \mathbbm{1}\left\{\mathrm{d}\mathbb{P}_{\bar{h}}^{\mathcal{M}^{\mathbb{S}}} / \mathrm{d} \mathbb{P}_{0}^{\mathcal{M}^{\mathbb{S}}}\geq c_{\mathbb{S}}(\bar{h},J_f;\alpha)\right\}$, for a fixed value of $\bar{h}<0$. The associated power function is

align*[align* omitted — 231 chars of source]

Again, by the Neyman-Pearson lemma, the following corollary holds.

corollaryLet $f\in\mathfrak{F}_\mathbb{S}$ and $\alpha\in (0,1)$. Let $\phi$ be a (possibly randomized) test that is $\mathcal{M}^{\mathbb{S}}$-measurable and is of size $\alpha$, i.e., $\mathbb{E}_{0} \phi \leq \alpha$. Let $\pi$ denote the power function of this test, i.e., $\pi(h)=\mathbb{E}_{h} \phi$. Then we have \begin{align*} \pi(\bar{h}) \leq \pi^{\mathfrak{F}_\mathbb{S}*}_{f,\alpha}(\bar{h};\bar{h}). \end{align*}

The likelihood ratio of the maximal invariant $\mathcal{M}^{\mathbb{S}}$, $\mathrm{d} \mathbb{P}_{h}^{\mathcal{M}^{\mathbb{S}}} / \mathrm{d} \mathbb{P}_{0}^{\mathcal{M}^{\mathbb{S}}}$, equals the likelihood ratio of the full observation $(W_{\varepsilon},W_{\phi_f},W_{\Gamma},(W_{b_k})_{k\in\mathbb{N}})^\prime$ with known $\eta=0$ (i.e., known $f$) and $\gamma=0$, $\mathrm{d} \mathbb{P}_{h,0,0} / \mathrm{d} \mathbb{P}_{0,0,0}$. This shows that the “semiparametric” power envelope $\pi^{\mathfrak{F}_\mathbb{S}*}_{f,\alpha}$ actually coincides with the parametric power envelope, which is defined based on the likelihood ratio $\mathrm{d} \mathbb{P}_{h,0,0} / \mathrm{d} \mathbb{P}_{0,0,0}$. This verifies again the adaptivity result in Jansson2008 under the same conditions for this unit root testing problem.\footnote{A discussion about the notion of “adaptiveness” in this nonstandard testing problem can be found in Section 5 of Jansson2008.} We provide, in Sections (ref) and (ref), a class of adaptive unit root tests based on signed-rank statistics for this setting.

remarkThe semiparametric power envelopes $\pi^{\mathfrak{F}*}_{f,\alpha}$ and $\pi^{\mathfrak{F}_\mathbb{S}*}_{f,\alpha}$ are scale invariant, i.e., invariant with respect to the value of $\sigma_f>0$. This is easily seen from the fact that $W_{\varepsilon}$, $W_{\phi_f}$ and $J_f$ are all scale invariant.
remarkThe problem of eliminating the nuisance parameter $\eta$ in the symmetric density case ($f\in\mathfrak{F}_\mathbb{S}$) is actually the same as that of eliminating the autocorrelation parameter $\gamma$ in the beginning of this section: the nuisance parameter appears only in a process which is independent of all the other processes. To be specific, $\eta$ only affects the distribution of $W_{b_k}$, which is independent of $W_{\phi_f}$ as well as $W_{\varepsilon}$. Therefore, the distribution of $\left(W_{\varepsilon},W_{\phi_f},(W_{b_k})_{k\in\mathbb{N}}\right)$ conditional on the statistic $\left(W_{\varepsilon},W_{\phi_f}\right)$ does not depend on the parameter of interest $h$. Hence, the statistic $\left(W_{\varepsilon},W_{\phi_f}\right)$ serves as a sufficient statistic for $h$ which is also invariant with respect to $\eta$.

The asymptotic power envelope for asymptotically invariant tests

Now we translate the results for the limit LABF experiment to the unit root model of interest. To mimick the invariance in the limit experiment we introduce the following definition.

definitionLet $\mu\in\mathbb{R}$, $\Gamma\in\mathfrak{G}$, $f\in\mathfrak{F}$. A sequence of test statistics $\psi^{(T)}$ is said to be asymptotically invariant if the distribution of $\psi^{(T)}$ weakly converges, under $\mathrm{P}^{(T)}_{h,\gamma,\eta;\mu,\Gamma,f}$ for all $h\leq 0$, $\gamma\in\mathbb{R}^p$, and $\eta\in c_{00}$, to the distribution of an invariant test in the limit experiment $\mathcal{E}(f)$/$\mathcal{E}_\mathbb{S}(f)$, under $\mathbb{P}_{h,\gamma,\eta}$.

The Asymptotic Representation Theorem (see, e.g., vdVaart00 Chapter 9) now yields the following main result on the asymptotic power envelope.

theoremLet $\mu\in\mathbb{R}$, $\Gamma\in\mathfrak{G}$, $f\in\mathfrak{F}$, and $\alpha\in (0,1)$. Let $\phi_T(Y_1,\dots,Y_T)$, $T\in\mathbb{N}$, be an asymptotically invariant test of size $\alpha$, i.e., $\limsup_{T\to\infty} \mathrm{E}_{0,\gamma,\eta} \phi_T\leq \alpha$ for all $\gamma\in\mathbb{R}^p$ and $\eta\in c_{00}$. Let $\pi_T$ denote the power function of $\phi_T$, i.e., $\pi_T(h,\gamma,\eta)=\mathrm{E}_{h,\gamma,\eta;\mu,\Gamma,f} \phi_T$. Then, we have \begin{align*} \limsup_{T\to\infty} \pi_T(h,\gamma,\eta)\leq \pi^{\mathfrak{F}*}_{f,\alpha}(h ; h),\quad h<0, \gamma\in\mathbb{R}^p, {\rm and } \eta\in c_{00}. \end{align*} For the case of $f\in\mathfrak{F}_\mathbb{S}$, we have \begin{align*} \limsup_{T\to\infty} \pi_T(h,\gamma,\eta)\leq \pi^{\mathfrak{F}_\mathbb{S}*}_{f,\alpha}(h ; h),\quad h<0, \gamma\in\mathbb{R}^p, {\rm and } \eta\in c_{00}. \end{align*}

These power envelopes for invariant tests in the limit experiments $\mathcal{E}(f)$ and $\mathcal{E}_\mathbb{S}(f)$ thus provide upper bounds to the asymptotic powers of invariant tests for the unit root hypotheses in $\mathcal{E}^{(T)}(f)$ and $\mathcal{E}^{(T)}_{\mathbb{S}}(f)$, respectively. The next section introduces two classes of tests (based on rank and signed-rank statistics) that (in a point-wise sense) attain these power envelopes and, thereby, demonstrates that these bounds are sharp. We additionally provide a Chernoff-Savage type result for these classes of tests.

A class of semiparametrically optimal hybrid rank tests

The appearance of the bridge process $B_{\phi_f}$ in the “efficient central sequence” $\Delta^*_f$ naturally suggests the (partial) use of ranks in the construction of test statistics. Indeed, we can construct an empirical analogue of $B_{\phi_f}$ by considering a partial-sum process which only depends on the observations via the ranks $R_t$ of $\Gamma(L)\Delta Y_t$ amongst $\Gamma(L)\Delta Y_{p+2},\dots, \Gamma(L)\Delta Y_T$. We allow for the use of a reference density $g$ that may or may not be equal to the true underlying innovation density $f$. Our findings compare to Quasi-ML methods: if the true innovation density happens to be the same as the selected reference density the inference procedure is point-optimal. At the same time, the procedure remains valid, i.e., has proper asymptotic size, even in case the true innovation density does not coincide with the reference density. Note that these results also hold in case the reference density is non-Gaussian, while Quasi-ML results are generally restricted to Gaussian reference densities.

We need the following mild assumption on the reference density.

assumptionThe density $g\in\mathfrak{F}$, with finite variance $\sigma_g^2$, satisfies \begin{equation*} \lim_{T\rightarrow\infty}\frac{1}{T}\sum_{i=1}^T \sigma_g^2\phi_g^2\left(G^{-1}\left(\frac{i}{T+1}\right)\right) = J_g, \end{equation*} with location score function $\phi_g(\varepsilon):=-(g^\prime/g)(\varepsilon)$, where $J_g$ is the standardized Fisher information for location of $g$.\footnote{Similarly to the standardized Fisher information $J_f$ of $f$, the Fisher information $J_g$ of $g$ is standardized by the variance $\sigma_g^2$. As a result, it is scale invariant.}

Now we can formulate the following direct extension of Lemma A.1 in HvdAW1 and its signed-rank counterpart. The proof for the weak convergence of the stochastic integrals in ((ref)) and ((ref)) is provided in (ref).

lemmaLet $\mu\in\mathbb{R}$, $\Gamma\in\mathfrak{G}$, and $g$ satisfy Assumption (ref). \begin{itemize} • For the case $f\in\mathfrak{F}$, consider the partial sum process \begin{equation} B^{(T)}_{\phi_g}(s) = \frac{1}{\sqrt{T}}\sum_{t=p+2}^{\lfloor s T \rfloor} \sigma_g \left(\phi_g\left( G^{-1}\left(\frac{R_t}{T-p}\right)\right) - \bar{\phi}_g^{(T)}\right) \end{equation} for $s\in[0,1]$, where $\bar{\phi}_g^{(T)} := T^{-1}\sum_{i=1}^{T-p-1}\phi_g(G^{-1}(i/(T-p)))$, and $R_t$ denotes the rank of $\Gamma(L)\Delta Y_t$, $t=p+2,\dots,T$. Then, under $P^{(T)}_{0,0,0;\mu,\Gamma,f}$ and as $T\to\infty$, we have \begin{align} \left[ W^{(T)}_{\varepsilon}, W^{(T)}_{\phi_f}, B^{(T)}_{\phi_g} \right]^\prime \Rightarrow \left[ W_{\varepsilon}, W_{\phi_f}, B_{\phi_g} \right]^\prime, \end{align} and \begin{align} \int_0^1 W_{\varepsilon}^{(T)}(s-)\rdB_{\phi_g}^{(T)}(s) \Rightarrow \int_0^1 W_{\varepsilon}(s)\rdB_{\phi_g}(s). \end{align} Here, $B_{\phi_g}$ is the associated Brownian bridge of $W_{\phi_g}$, which itself is a Brownian motion defined on the same probability space $(\Omega,\mathcal{F},\mathbb{P}_{0,0,0})$ as $W_{\varepsilon}$ and $W_{\phi_f}$, with covariance matrix \begin{align} \operatorname{Cov}_{0,0,0} \begin{bmatrix} W_{\varepsilon}(1) \\ W_{\phi_f}(1) \\ W_{\phi_g}(1) \end{bmatrix} = \begin{pmatrix} 1 & 1 & \sigma_{\varepsilon\phi_g} \\ & J_f & J_{fg} \\ & & J_g \end{pmatrix}, \end{align} where \begin{align*} \sigma_{\varepsilon\phi_g} =& \sigma_f^{-1}\sigma_g\int_0^1 F^{-1}(u) \phi_g(G^{-1}(u))\mathrm{d} u, \\ J_{fg} =& \sigma_f\sigma_g\int_0^1 \phi_f(F^{-1}(u)) \phi_g(G^{-1}(u))\mathrm{d} u. \end{align*} • For the case $f\in\mathfrak{F}_\mathbb{S}$, consider the partial sum process \begin{align} W^{(T)}_{\phi_g}(s) =& \frac{1}{\sqrt{T}}\sum_{t=p+2}^{\lfloor sT\rfloor}s_t \sigma_g \left( \phi_{g}\left(G^{-1}\left(\frac{1}{2} + \frac{R_t^{+}}{2(T-p)}\right)\right) \right) \end{align} for $s\in[0,1]$, where $s_t$ and $R_t^{+}$ denote the sign of $\Gamma(L)\Delta Y_t$ and the rank of its absolute value, respectively, for $t=p+2,\dots,T$. Then, under $P^{(T)}_{0,0,0;\mu,\Gamma,f}$ and as $T\to\infty$, we have \begin{align} \left[W^{(T)}_{\varepsilon}, W^{(T)}_{\phi_f}, W^{(T)}_{\phi_g} \right]^\prime \Rightarrow \left[ W_{\varepsilon}, W_{\phi_f}, W_{\phi_g} \right]^\prime, \end{align} and \begin{align} \int_0^1 W_{\varepsilon}^{(T)}(s-)\rdB_{\phi_g}^{(T)}(s) \Rightarrow \int_0^1 W_{\varepsilon}(s)\rdB_{\phi_g}(s), \end{align} where the law of $W_{\phi_g}$ is given in ((ref)). \end{itemize}

In practice, the autoregressive structure $\Gamma(L)$ will not be known. In that case, one must rely on ranks (and signs) based on innovations calculated using an estimated autoregressive structure, i.e., using residuals. This is usually referred to as inference based on aligned ranks. For this purpose, we first introduce the following assumption on the estimation of $\Gamma(L)$ (see HallinPuri1994).

assumption\begin{itemize} • There exists, under the null hypothesis, a $\sqrt{T}$-consistent estimator $(\widehat\Gamma_{1},\dots,\widehat\Gamma_{p})^\prime$ of $(\Gamma_{1},\dots,\Gamma_{p})^\prime$. That is, for all $\mu\in\mathbb{R}$, $\Gamma\in\mathfrak{G}$, $f\in\mathfrak{F}$ and all $\epsilon>0$, there exists $b$ and $T_b$ such that \begin{align}P^{(T)}_{0,0,0;\mu,\Gamma,f}\left\{\|\sqrt{T}(\widehat\Gamma - \Gamma)\|>b\right\} < \epsilon, \forall t\geq T_b. \end{align} • The estimator $(\widehat\Gamma_{1},\dots,\widehat\Gamma_{p})^\prime$ is discretized on grids of mesh width $T^{-1/2}$. That is, the number of possible values of $(\widehat\Gamma_{1},\dots,\widehat\Gamma_{p})^\prime$ in balls of the form $\{\Upsilon\in\mathbb{R}^p ~ \| ~ T^{-1/2}(\Upsilon-\Upsilon_0)\leq c\}$ remains bounded, as $T\to\infty$, for all $\Upsilon_0\in\mathbb{R}^p$ and all $c>0$. \end{itemize}

Part (i) of Assumption (ref) is mild as many known estimators of $(\Gamma_{1},\dots,\Gamma_{p})^\prime$ exist for the AR($p$) model that can be applied to the increments $\Delta Y_t$, see, e.g., BrockwellDavis2002. Part (ii) is standard in the semiparametric literature and can easily be met by transforming the estimator in Part (i), see, e.g., Bickel1982 and Kreiss1987. Such an estimator is often called locally asymptotically discrete. The use of aligned ranks does not invalidate the conclusion of Lemma (ref). This is the content of the following result, of which the proof is provided in (ref).

lemmaUnder Assumption (ref), Lemma (ref) remains valid in case aligned signs and ranks are used, i.e., signs and ranks of innovations calculated using the estimated autoregressive coefficients $(\widehat\Gamma_{1},\dots,\widehat\Gamma_{p})$.

We denote by $\widehat{B}^{(T)}_{\phi_g}$ and $\widehat{W}^{(T)}_{\phi_g}$ the aligned-rank-based counterparts of the rank-based processes $B^{(T)}_{\phi_g}$ and $W^{(T)}_{\phi_g}$, respectively.

Hybrid rank tests based on a reference density

In this section, we propose our unit root tests, for both the non-symmetric case ($f\in\mathfrak{F}$) and the symmetric case ($f\in\mathfrak{F}_\mathbb{S}$). Section (ref) provides the basis for optimal invariant tests in the limit experiment $\mathcal{E}(f)$ and $\mathcal{E}_\mathbb{S}(f)$. Lemmas (ref) and (ref) can subsequently be used to approximate, in the sequences of unit root experiments $\mathcal{E}^{(T)}(f)$ and $\mathcal{E}^{(T)}_{\mathbb{S}}(f)$, the observable processes in these limit experiments. More specifically, these two lemmas indicate that constructing the rank-based score partial sum processes $B^{(T)}_{\phi_g}$ and $W^{(T)}_{\phi_g}$ as in ((ref)) and ((ref)), together with $W^{(T)}_{\varepsilon}$, intuitively corresponds to observing the $\sigma$-fields $\mathcal{M}_{g} := \sigma(W_{\varepsilon}, B_{\phi_g})$ and $\mathcal{M}_{g}^\mathbb{S} := \sigma(W_{\varepsilon}, W_{\phi_g})$ in the limit experiments $\mathcal{E}(f)$ and $\mathcal{E}_\mathbb{S}(f)$, respectively. This then leads to our asymptotically invariant tests based on $\mathcal{M}_{g}$ and $\mathcal{M}_{g}^\mathbb{S}$.

The following proposition establishes the likelihood ratio restricted to the information $\mathcal{M}_{g}$ and $\mathcal{M}_{g}^\mathbb{S}$. The Neyman-Pearson lemma implies that tests based on these likelihood ratios are point optimal amongst the class of invariant tests in the limit experiments.

propositionDefine $W_{\perp}$ implicitly via the decomposition \begin{equation} W_{\phi_g} = \covegW_{\varepsilon} + \sqrt{J_g - \sigma_{\varepsilon\phi_g}^2}W_{\perp}, \end{equation} which is a standard Brownian motion under $\mathbb{P}_{0,0,0}$ and denote the associated bridge by $B_{\perp}$.\footnote{The implicit requirement $J_g \geq \sigma_{\varepsilon\phi_g}^2$ in the decomposition ((ref)) is directly guaranteed the Cauchy-Schwarz inequality, $\sigma_g^2\int_0^1\left|\phi_g(G^{-1}(u))\right|^2\mathrm{d} u \cdot \sigma_f^{-2}\int_0^1 \left|F^{-1}(u)\right|^2 \mathrm{d} u \geq \big|\sigma_f^{-1}\sigma_g\int_0^1 F^{-1}(u) \phi_g(G^{-1}(u))\mathrm{d} u\big|^2$, and the fact that $\sigma_f^{-2}\int_0^1 [F^{-1}(u)]^2 \mathrm{d} u = 1$.} \begin{itemize} • The likelihood ratio $\mathrm{d}\mathbb{P}_h^\mathcal{M}/\mathrm{d}\mathbb{P}_0^\mathcal{M}$ restricted to the outcome space $\mathcal{M}_{g}$ is given by \begin{align} \frac{\mathrm{d}\mathbb{P}^{\mathcal{M}_{g}}_h}{\mathrm{d}\mathbb{P}^{\mathcal{M}_{g}}_0} = \mathbb{E}_0\left[\frac{\mathrm{d} \mathbb{P}_{h}^\mathcal{M} } {\mathrm{d} \mathbb{P}_{0}^\mathcal{M}}\mid\mathcal{M}_{g}\right]=\exp\left(h\Delta_g -\frac{1}{2}h^2\mathcal{I}_g\right), \end{align} where \begin{align*} \Delta_g =& \Delta_\varepsilon + \lambda\Delta_\perp, \\ \mathcal{I}_g =& \int_0^1W_{\varepsilon}^2(s)\mathrm{d} s + \lambda^2\left(\frac{J_g}{\sigma_{\varepsilon\phi_g}^2}-1\right)\left[\int_0^1W_{\varepsilon}(s)^2\mathrm{d} s - \left( \int_0^1 W_{\varepsilon}(s) \mathrm{d} s\right)^2\right], \end{align*} with $\Delta_\varepsilon = \int_0^1W_{\varepsilon}(s)\rdW_{\varepsilon}(s)$, $\Delta_\perp = \sqrt{J_g/\sigma_{\varepsilon\phi_g}^2-1}\int_0^1W_{\varepsilon}(s)\rdB_{\perp}(s)$, and \begin{align*} \lambda=(J_{fg}\sigma_{\varepsilon\phi_g}-\sigma_{\varepsilon\phi_g}^2)/(J_g-\sigma_{\varepsilon\phi_g}^2). \end{align*} • The likelihood ratio $\mathrm{d}\mathbb{P}^{\mathcal{M}_{g}^\mathbb{S}}_h/\mathrm{d}\mathbb{P}^{\mathcal{M}_{g}^\mathbb{S}}_0$ restricted to the outcome space $\mathcal{M}_{g}^\mathbb{S}$ is given by \begin{align} \frac{\mathrm{d}\mathbb{P}^{\mathcal{M}_{g}^\mathbb{S}}_h}{\mathrm{d}\mathbb{P}^{\mathcal{M}_{g}^\mathbb{S}}_0} =& \mathbb{E}_0\left[\frac{\mathrm{d} \mathbb{P}_{h}^{\mathcal{M}^{\mathbb{S}}} } {\mathrm{d} \mathbb{P}_{0}^{\mathcal{M}^{\mathbb{S}}}}\mid\mathcal{M}_{g}\right]=\exp\left(h\Delta_g^\mathbb{S} -\frac{1}{2}h^2\mathcal{I}_g^\mathbb{S}\right), \end{align} where \begin{align*} \Delta_g^\mathbb{S} =& \Delta_\varepsilon + \lambda\Delta_\perp^\mathbb{S}, \\ \mathcal{I}_g^\mathbb{S} =& \left(1 + \lambda^2\frac{J_g}{\sigma_{\varepsilon\phi_g}^2}-\lambda^2\right)\left[\int_0^1W_{\varepsilon}(s)^2\mathrm{d} s\right], \end{align*} with $\Delta_\perp^\mathbb{S} = \sqrt{J_g/\sigma_{\varepsilon\phi_g}^2-1}\int_0^1W_{\varepsilon}(s)\rdW_{\perp}(s)$. \end{itemize}
remarkTo construct the test statistic in the limit experiment $\mathcal{E}(f)$, instead of simply replacing $B_{\phi_f}$ by $B_{\phi_g}$, we rely on the likelihood ratio of $\mathcal{M}_{g}$. This, by the Neyman-Pearson Lemma, is the optimal way regardless of complexity. Clearly, it has $\mathcal{M}_{g}\subseteq\mathcal{M}$ so that $\mathcal{M}_{g}$ is invariant for the group of transformations $g_\eta$.\footnote{This is due to the decomposition $B_{\phi_g}=\covegB_{\varepsilon}+\sum_{k=1}^{\infty}J_{g,k}B_{b_k}$.} When $g=f$, $\mathcal{M}_{g}=\mathcal{M}$ so that it is maximally invariant, which means that we capture all available information about $h$, and this in turn gives the asymptotic optimality for the finite-sample counterpart below. The same argument also holds for the limit experiment $\mathcal{E}_\mathbb{S}(f)$.
remarkThe result of Proposition (ref) can also be achieved by first applying Girsanov's Theorem to the experiment associated to observing $W_{\varepsilon}$ and $W_{\phi_g}$ defined via \begin{align*} \mathrm{d} W_{\varepsilon}(s) &= h W_{\varepsilon}(s)\mathrm{d} s + \mathrm{d} Z_\varepsilon(s), \\ \mathrm{d} W_{\phi_g}(s) &= h J_{fg}W_{\varepsilon}(s)\mathrm{d} s + \mathrm{d} Z_{\phi_g}(s), \end{align*} to get the likelihood ratio of $\mathcal{M}_{g}^\mathbb{S}$. This experiment is obtained by combining the limit experiment in Proposition (ref) and the covariance matrix in ((ref)). In other words, this provides the limit experiment associated to $\mathcal{M}_{g}^\mathbb{S}$. Subsequently, one can take the expectation of the likelihood ratio of $\mathcal{M}_{g}^\mathbb{S}$ obtained above conditional on $\mathcal{M}_{g}$ to get the likelihood ratio of $\mathcal{M}_{g}$. The associated limit experiment is given by \begin{align*} \mathrm{d} W_{\varepsilon}(s) &= h W_{\varepsilon}(s)\mathrm{d} s + \mathrm{d} Z_\varepsilon(s), \\ \mathrm{d} B_{\phi_g}(s) &= h J_{fg}\left[W_{\varepsilon}(s) - \overline{W_{\varepsilon}}\right]\mathrm{d} s + \mathrm{d} \left[Z_{\phi_g}(s) - s Z_{\phi_g}(1)\right], \end{align*} where $\overline{W_{\varepsilon}} = \int_0^1W_{\varepsilon}(r)\mathrm{d} r$.

Observe that $W_{\perp}$ is a standard Brownian motion under $\mathbb{P}_{0,0,0}$ and independent of $W_{\varepsilon}$. When $g=f$, we have $J_{fg}=J_f=J_g$ and $\sigma_{\varepsilon\phi_g}=1$, so that $\lambda=1$ and $B_{\phi_g}=B_{\phi_f}$. As a result, we have $\Delta_g=\Delta^*_f$ and $\mathcal{I}_g=\mathcal{I}^*_f$, and $\Delta_g^\mathbb{S}=\Delta_f$ and $\mathcal{I}_g^\mathbb{S}=\mathcal{I}_f$.

The central idea to construct a hybrid rank test is to use a (quasi)-likelihood ratio test based on $L_{\mathcal{M}_{g}}(h,\lambda) := h\Delta_g -\frac{1}{2}h^2\mathcal{I}_g$ from ((ref)) for the case of $f\in\mathfrak{F}$, and $L_{\mathcal{M}_{g}^\mathbb{S}}(h,\lambda) := h\Delta_g^\mathbb{S} -\frac{1}{2}h^2\mathcal{I}_g^\mathbb{S}$ from ((ref)) for the case of $f\in\mathfrak{F}_\mathbb{S}$. In both cases, we then replace $W_{\varepsilon}$ and $B_{\phi_g}$ by their finite-sample counterparts in Lemma (ref) (or those in Lemma (ref) in case $\Gamma(L)$ would be known, for example, when $p=0$). The remaining unknown finite-sample parameters $\sigma_f^2$ and $\lambda$ are replaced by estimates that need to satisfy the following condition.

assumptionThere exist consistent, under the null hypothesis, estimators $\hat\sigma_f^2>0~a.s.$, $\hat\sigma_{\varepsilon\phi_g}$ and $\hat{J}_{fg}$ of $\sigma_f^2$, $\sigma_{\varepsilon\phi_g}$ and $J_{fg}$, respectively. More precisely, for all $\mu\in\mathbb{R}$, $\Gamma\in\mathfrak{G}$, $f\in\mathfrak{F}$, we have $\hat\sigma_f^2 \stackrel{p}{\to} \sigma_f^2$, $\hat\sigma_{\varepsilon\phi_g} \stackrel{p}{\to} \sigma_{\varepsilon\phi_g}$, and $\hat{J}_{fg} \stackrel{p}{\to} J_{fg}$, under $\mathrm{P}^{(T)}_{0,0,0;\mu,\Gamma,f}$ as $T\to\infty$.

Such estimators are easily constructed, although $\hat{J}_{fg}$ is somewhat more involved. Estimating the real-valued cross-information $J_{fg}$ requires nonparametric techniques, but is considerably simpler than a full nonparametric estimation of $\phi_f$. Estimating $J_{fg}$ can be done along similar lines as estimating the Fisher information $J_f$, see, e.g., Bickel1982, BKRW, Schick1986, and Klaassen1987. A direct rank-based estimator of $J_{fg}$ has been proposed in CassartHallinPain2010. It is also worth noting that the consistency automatically also holds under local alternatives due to Le Cam's third lemma.

Based on a chosen reference density $g$ satisfying Assumption (ref) and estimators $\hat\sigma_f$, $\hat\sigma_{\varepsilon\phi_g}$ and $\hat{J}_{fg}$ satisfying Assumption (ref), we introduce the following partial sum processes:

align*[align* omitted — 653 chars of source]

where $\widehat{B}^{(T)}_{\phi_g}(s)$ and $\widehat{W}^{(T)}_{\phi_g}(s)$ are defined by Lemma (ref). Note that in the case of known $\Gamma(L)$, e.g., the i.i.d.\ case, one can simply use $\widehat{\Gamma}(L) = \Gamma(L)$ (where $\widehat{B}^{(T)}_{\phi_g}(s)=B^{(T)}_{\phi_g}(s)$ and $\widehat{W}^{(T)}_{\phi_g}(s)=W^{(T)}_{\phi_g}(s)$). Now, given a fixed alternative $\bar{h}<0$, we define

align*[align* omitted — 138 chars of source]

with

align*[align* omitted — 444 chars of source]

where $\widehat{\Delta}_\varepsilon = \int_0^1\widehat{W}^{(T)}_\varepsilon(s-)\mathrm{d}\widehat{W}^{(T)}_\varepsilon(s)$, $\widehat{\Delta}_\perp = \sqrt{J_g/\hat\sigma_{\varepsilon\phi_g}^2-1}\int_0^1\widehat{W}^{(T)}_\varepsilon(s-)\mathrm{d}\widehat{B}^{(T)}_{\perp}(s)$ and $\hat\lambda=(\hat{J}_{fg}\hat\sigma_{\varepsilon\phi_g}-\hat\sigma_{\varepsilon\phi_g}^2)/(J_g-\hat\sigma_{\varepsilon\phi_g}^2)$. Define also for the case of $f\in\mathfrak{F}_\mathbb{S}$,

align*[align* omitted — 171 chars of source]

with

align*[align* omitted — 342 chars of source]

where $\widehat{\Delta}_\perp^\mathbb{S} = \sqrt{J_g/\hat\sigma_{\varepsilon\phi_g}^2-1}\int_0^1\widehat{W}^{(T)}_\varepsilon(s-)\mathrm{d}\widehat{W}^{(T)}_\perp(s)$.

By Slutsky's theorem, we have the convergences $\big(\widehat{W}^{(T)}_\varepsilon, \widehat{B}^{(T)}_{\perp} \big)^\prime \Rightarrow \big( W_{\varepsilon}, B_{\perp} \big)^\prime$ and $\big(\widehat{W}^{(T)}_\varepsilon, \widehat{W}^{(T)}_\perp \big)^\prime \Rightarrow \big( W_{\varepsilon}, W_{\perp} \big)^\prime$, and the convergences of stochastic integrals $\int_0^1\widehat{W}^{(T)}_\varepsilon(s-)\mathrm{d}\widehat{B}^{(T)}_{\perp}(s) \Rightarrow \int_0^1W_{\varepsilon}(s)\rdB_{\perp}(s)$ and $\int_0^1\widehat{W}^{(T)}_\varepsilon(s-)\mathrm{d}\widehat{W}^{(T)}_\perp(s) \Rightarrow \int_0^1W_{\varepsilon}(s)\rdW_{\perp}(s)$ (see the proof of Lemma (ref)). We thus also obtain $\widehat{L}_{\mathcal{M}_{g}}(\bar{h},\hat\lambda)\wtoL_{\mathcal{M}_{g}}(\bar{h},\lambda)$ and $\widehat{L}_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\hat\lambda)\wtoL_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\lambda)$ under $P^{(T)}_{0,0,0;\mu,\Gamma,f}$. Define the critical values $c_{\mathcal{M}_{g}}(\bar{h},\sigma_{\varepsilon\phi_g},\lambda,J_g;\alpha)$ and $c_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\sigma_{\varepsilon\phi_g},\lambda,J_g;\alpha)$ by the $(1-\alpha)$-quantiles of $L_{\mathcal{M}_{g}}(\bar{h},\lambda)$ and $L_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\lambda)$, respectively. This leads to the (feasible) tests

align*[align* omitted — 218 chars of source]

and

align*[align* omitted — 251 chars of source]

These tests are not only based on the ranks of $\Delta Y_t$ but also their average, therefore, we name them Hybrid Rank Tests (HRTs).

We can now state our main theoretical result.

theorem\begin{itemize} • Under the Assumptions 1 - 5, for each chosen $\alpha\in (0,1)$ and $\bar{h}\in(-\infty,0)$, we have \begin{itemize} • The HRT $\phi_{\mathcal{M}_{g}}(\bar{h},\alpha)$ is asymptotically of size $\alpha$. • The HRT $\phi_{\mathcal{M}_{g}}(\bar{h},\alpha)$ is asymptotically invariant. • The HRT $\phi_{\mathcal{M}_{g}}(\bar{h},\alpha)$ is point-optimal, at $h=\bar{h}$, if $g=f$. \end{itemize} • Under the Assumptions 1 - 5 and $g\in\mathfrak{F}_\mathbb{S}$, for each chosen $\alpha\in (0,1)$ and $\bar{h}\in(-\infty,0)$, we have \begin{itemize} • The HRT $\phi_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\alpha)$ is asymptotically of size $\alpha$. • The HRT $\phi_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\alpha)$ is asymptotically invariant. • The HRT $\phi_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\alpha)$ is point-optimal, at $h=\bar{h}$, if $g=f$. \end{itemize} \end{itemize}

Theorem (ref) shows the HRTs are valid irrespective of the choice of the reference density and point-optimal for a correctly specified reference density. Moreover, in the corollary below, we state that the HRTs enjoy a Chernoff-Savage type result.

corollaryFix $\alpha\in(0,1)$ and $\bar{h}<0$. The HRT $\phi_{\mathcal{M}_{g}}(\bar{h},\alpha)$ is, for any reference density $g$ satisfying Assumption (ref), more powerful, at $h=\bar{h}$ and for $\mu\in\mathbb{R}$, $\Gamma\in\mathfrak{G}$ and $f\in\mathfrak{F}$, than the ERS test except when $f$ is Gaussian where they have equal powers. The same argument holds for the HRT $\phi_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\alpha)$ with $g\in\mathfrak{F}_\mathbb{S}$ for any $f\in\mathfrak{F}_\mathbb{S}$.

Corollary (ref) is a particularly useful result for applied work. The HRT dominates its classical canonical Gaussian counterpart, i.e., the ERS test in the present model, for any reference density $g$. Traditionally, this claim can only be made for Gaussian reference densities, but the framework here even allows for a stronger result. Our formulation of the testing problem using invariance arguments is convenient in this respect: the larger the invariant $\sigma$-field that is used, the more powerful the test.

The situation can be compared to Quasi Maximum Likelihood methods. However, again, in classical situations these methods are restricted to Gaussian reference densities. In the present setup, any reference density $g$ (subject to the regularity conditions imposed) can be used. The resulting test will always be valid, but more powerful in case the reference density chosen is closer to the true underlying density $f$.

remarkIt is worth noting that the invariance constraint is only imposed in the limit and, therefore the maximal invariant needs only to be derived in the limit experiment. In other words, in the finite-sample unit root testing experiment $\mathcal{E}^{(T)}(f)$ (or $\mathcal{E}^{(T)}_{\mathbb{S}}(f)$ for $f\in\mathfrak{F}_\mathbb{S}$), we actually use statistics that are only asymptotically invariant (i.e., their limiting equivalents are measurable with respect to the maximally invariant sigma-field $\mathcal{M}$ (or $\mathcal{M}^{\mathbb{S}}$ for $f\in\mathfrak{F}_\mathbb{S}$), while, for finite $T$, they are not necessarily (maximally) invariant with respect to some transformation on the density $f$. In fact, a maximal invariant may very well not even exist in the finite-sample experiments. Specifically, $W^{(T)}_{\varepsilon}$ approximates $W_{\varepsilon}$ in the limit, whose distribution does not change with the density $f$, while the distribution of $W_{\varepsilon}^{(T)}$ does depend on the density $f$. Instead, the statistic $B^{(T)}_{\phi_g}$ (or $W^{(T)}_{\phi_g}$ for $f\in\mathfrak{F}_\mathbb{S}$) is distribution-free, that is, its distribution is not affected by any transformation on the density $f$. In Section (ref) below, we will show that the asymptotic approximations work well even in smaller samples.
remarkThe additional power of the HRT compared to the ERS test is not free due to the stronger weak convergence assumption employed. Consequently, the class of models for which the HRTs are valid forms a sub-class of the class where the ERS tests are valid. In this sub-class, the HRT dominates the ERS test, but outside they may even loose validity. In the opposite direction, the MullerWatson2008 low-frequency unit root test can be applied in a even larger class of models than the ERS tests. Again, within the class of models where the ERS test is valid, it has lower power. A more general and detailed discussion in this direction can be found in Muller2011. Our test will still be relevant in many applications, notably those where policy implications are derived under an i.i.d.\ assumption on the innovations. Also, our approach can most likely be extended to situations where the innovations are described by some explicit dynamic location-scale model. We come back to this point in Section (ref).

Approximate hybrid rank tests

A somewhat inconvenient aspect of the hybrid rank tests is that we need to estimate “the real-valued parameter” $J_{fg}$. As mentioned before, this is (much) less complicated than estimating the score function $\phi_f$ (as needed in Jansson2008), but might still be considered cumbersome, despite all references mentioned below Assumption (ref). Moreover, the critical value $c_{\mathcal{M}_{g}}(\bar{h},\hat\sigma_{\varepsilon\phi_g},\hat\lambda,J_g;\alpha)$ depends on the estimates $\hat\sigma_{\varepsilon\phi_g}$ and $\hat\lambda$ (henceforth $\hat{J}_{fg}$). Of course, this introduces no difficulty to implementing the test for a single dataset (though one would need to simulate a critical value), however, when it comes to a Monte Carlo study to access the performances of the HRTs, the computational effort will be significant. Therefore, we introduce a simplified version of the hybrid rank test. This simplified test is obtain by invoking $\lambda=1$, which holds in case $g=f$.

To be precise, define

align[align omitted — 186 chars of source]

where

align*[align* omitted — 522 chars of source]

and

equation[equation omitted — 201 chars of source]

where

align*[align* omitted — 323 chars of source]

Also define $L_g(\bar{h}):=L_{\mathcal{M}_{g}}(\bar{h},1)$ and $L_g^{\mathbb{S}}(\bar{h}):=L_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},1)$, then we have $\widehat{L}_g(\bar{h})\Rightarrow L_g(\bar{h})$ and $\widehat{L}_g^\mathbb{S}(\bar{h})\Rightarrow L_g^\mathbb{S}(\bar{h})$ under $\mathrm{P}^{(T)}_{0,0,0;\mu,\Gamma,f}$. Denoting the $(1-\alpha)$-quantiles of $L_g(\bar{h})$ and $L_g^\mathbb{S}(\bar{h})$ by $c_{g}(\bar{h},\sigma_{\varepsilon\phi_g},J_g;\alpha)$ and $c_{g}^\mathbb{S}(\bar{h},\sigma_{\varepsilon\phi_g},J_g;\alpha)$, respectively. These lead to the feasible tests

align*[align* omitted — 138 chars of source]

and

align*[align* omitted — 171 chars of source]

Since $\phi_g(\bar{h},\alpha)$ and $\phi_g^\mathbb{S}(\bar{h},\alpha)$ are approximate versions of the Hybrid Rank Tests $\phi_{\mathcal{M}_{g}}(\bar{h},\alpha)$ and $\phi_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\alpha)$, we refer to them as Approximate Hybrid Rank Tests (AHRTs).

theoremUnder the same conditions as Theorem (ref), the asymptotic properties of the Hybrid Rank Tests --- validity, invariance, and point-optimality when $g=f$ --- also hold for the Approximate Hybrid Rank Tests.

The proof of Theorem (ref) follows along the same lines as that of Theorem (ref) but using the weak convergences $\widehat{L}_g(\bar{h}) \Rightarrow L_g(\bar{h})$ and $\widehat{L}_g^\mathbb{S}(\bar{h}) \Rightarrow L_g^\mathbb{S}(\bar{h})$. The simulation results in Section (ref) show that these asymptotic properties carry over to finite samples.

remarkAlthough we are not able to provide a rigorous mathematical proof, the Monte-Carlo study indicates that the Chernoff-Savage property is also preserved for the AHRTs, at least in case the reference density $g$ is chosen to be Gaussian. Such a result would be more in line with applications of the Chernoff-Savage result in classical LAN situations.

From a computational point of view, the AHRTs have the advantage that nonparametric estimation of $J_{fg}$ is no longer needed. This significantly reduces the computational effort in the Monte Carlo study. Indeed, even though the critical value $c_{g}(\bar{h},\sigma_{\varepsilon\phi_g},J_g;\alpha)$ and $c_{g}^\mathbb{S}(\bar{h},\sigma_{\varepsilon\phi_g},J_g;\alpha)$ are still data dependent, it is, for given $\alpha$, $\bar{h}$, and reference density $g$, a function of only one argument --- the parameter $\sigma_{\varepsilon\phi_g}$. Observe, by Cauchy-Schwarz, that $\sigma_{\varepsilon\phi_g}$ is bounded by $\sqrt{J_g}$. For the chosen three reference densities, the critical value functions are listed in Table (ref). These are obtained by fitting a fourth-order polynomial to the exact critical values. In the Monte Carlo study (Section (ref)), we use these approximating critical value functions for computational speed.

table[table omitted — 1,257 chars of source]
remark[Nonparametrically estimated reference density] The Hybrid Rank Test and the Approximate Hybrid Rank Test are optimal when the reference density $g$ coincides with the actual innovation density $f$. It is therefore reasonable to consider these test using a nonparametric estimate of $f$, say $\hat{f}$, as reference density. Commonly such estimators are based on the order statistics of the residuals $\hat\varepsilon_t$. Under a suitable consistency condition, the HRT based on $\hat{f}$ asymptotically is conjectured to behave as the HRT based on the true innovation density $f$. Thus, such test achieves the optimality properties of Theorem (ref) and Theorem (ref) globally. Notably, even if there exists relatively large bias in the estimation of $f$, the usage of rank statistics ensures zero expectation of the feasible score function $\phi_{\hat{f}}[\widehat{F}^{-1}(R_t/(T+1))]$, which furthermore ensures the validity of the HRTs and the AHRTs.

Monte Carlo study

This section reports the results of a Monte Carlo study to corroborate our asymptotic results, and to analyze the small-sample performance of the Approximate Hybrid Rank Tests. As mentioned earlier, we use the Approximate Hybrid Rank Tests in this simulation to avoid having to simulate the critical value for each individual replication. For the fixed alternative, we choose $\bar{h}=-7\sigma_{\varepsilon\phi_g}$ for two reasons. First, when $g=f$, we have $\sigma_{\varepsilon\phi_g}=1$ and hence $\bar{h}=-7$, which is in line with ERS1996. Second, as $\hat\sigma_{\varepsilon\phi_g}$ appears in the denominator in the AHRT statistic ((ref)), the statistic becomes better behaved for values of $\hat\sigma_{\varepsilon\phi_g}$ close to zero. The corresponding critical value functions for various reference densities are provided in Table (ref). The estimators for $\sigma_f^2$ and $\sigma_{\varepsilon\phi_g}$ we use are

align*[align* omitted — 348 chars of source]

Moreover, to simplify the notation, we denote the Approximate Hybrid Rank Test with reference density $g$ by ${\rm AHRT}$-$g$ and, in particular, by ${\rm AHRT}$-$\phi$ for Gaussian reference density and by ${\rm AHRT}$-$\hat{f}$ for a (nonparametrically) estimated reference density. To be specific, for the nonparametric estimation of $f$, we employ a kernel-based method $\hat{f}(x) = (T\mathfrak{h})^{-1}\sum_{t=1}^{T}K\left((x-\hat{\varepsilon}_t)/\mathfrak{h}\right)$, where the kernel $K$ is chosen to be Gaussian, $\mathfrak{h}$ is the bandwidth chosen by the rule $\left(4/(3T)\right)^{\frac{1}{5}}\hat\sigma_f$, and $\hat{\varepsilon}_t$ are the residuals from the regression ((ref)) below. Throughout we use significance level $\alpha = 5\%$ and all results are based on 20,000 Monte-Carlo replications.

We compare the performances of the propsed AHRT tests with two alternatives. First we consider the Dickey-Fuller test (denoted by DF-$\rho$) from DickeyFuller79. This test is based on the statistic $T(\hat{\rho}-1)$ where $\hat{\rho}$ is the least-squares estimator in the regression

align[align omitted — 116 chars of source]

The critical values for this test are -13.52 for $T=100$ and -14.05 for $T=2,500$. The second competitor is the ERS test with $\bar{h}=-7$. This test is based on the statistic $[S(\bar{\alpha})-\bar{\alpha}S(1)]/\hat{\omega}^2$ with $\bar{\alpha}=1+T^{-1}\bar{h}$ and $S(a) = (Y_a-Z_a\hat\beta)'(Y_a-Z_a\hat\beta)$, with $Y_a$ and $Z_a$ defined as

align*[align* omitted — 87 chars of source]

where $\hat\beta$ is estimated by regressing $Y_{\bar{\alpha}}$ on $Z_{\bar{\alpha}}$. The long-run variance estimator (for ERS test), $\hat{\omega}^2$, is chosen to be $\hat{\omega}^2_{AR(p)} = \hat\sigma_e^2 \big/ \left(1-\sum_{i=1}^p \widehat{\Gamma}_i \right)^2$ with $\hat\sigma_e^2 = \sum_{t=1}^T\hat{\varepsilon}_t^2/T$, where the residuals $\hat{\varepsilon}_t$ and coefficient estimates $\widehat{\Gamma}_i$ are from the regression in ((ref)). The critical values for this test are 3.11 for $T=100$ and 3.26 for $T=2,500$. We do not consider the Dickey-Fuller $t$-test as it is dominated by the DF-$\rho$ test in the current model. Similarly, the DF-GLS test proposed in ERS1996 is also omitted as it behaves asymptotically the same as the ERS test, but can be oversized in smaller samples.

Simulation results with i.i.d.\ innocations

In this section, we start the Monte Carlo study with i.i.d.\ innovations, i.e., the data is generated by the model in ((ref))-((ref)) with $\Gamma(L) = 1$. Therefore, for the newly-proposed AHRT test and its competitors introduced above, we have $\widehat{\Gamma}(L) = 1$ (or, equivalently, $\widehat{\Gamma}_1 = \cdots = \widehat{\Gamma}_p = 0$).

Large-sample performance

We first use the large-sample performances to illustrate the asymptotic properties. In particular, the chosen sample size $T$ is $2,500$.

figure[figure omitted — 360 chars of source]

Figure (ref) shows the power curves for 9 combinations of 3 innovation densities $f$ and 3 reference densities $g$ (for AHRT-$g$): $f$ and $g$ are chosen to be Laplace, Student $t_3$, or Gaussian.\footnote{The semiparametric power envelopes are based on $40,000$ Monte Carlo replications where the W-processes are approximated by a simple Euler approximation using $2,500$ grid points.} In line with our theoretical results, we find that the AHRT-$g$ test outperforms the two competitors in most cases. More specifically, when $g=f$ (the graphs on the diagonal), the AHRT-$f$ has power very close to the semiparametric power envelope and it is tangent to it at the point $-h=7$. This verifies the point-optimal result of the AHRT-$f$ test in Theorem (ref). The AHRT-$\hat{f}$ test has a similar behavior as the AHRT-$f$ but with a slightly lower power due to the efficiency loss in estimating the density $f$. This small amount of power loss is also different from case to case, e.g., for $f = t_3$, this power loss is almost indistinguishable; and the amount decreases to zero as $T$ goes to infinity. Moreover, when the reference density $g$ is Gaussian (the three right-most graphs), the AHRT-$\phi$ outperforms its competitors for non-Gaussian $f$; while for Gaussian $f$, the AHRT-$\phi$ test and the ERS test have indistinguishable power (and they outperform the Dickey-Fuller-$\rho$ test). This corroborates the Chernoff-Savage property of the AHRT-$\phi$ test mentioned in Remark (ref).

figure[figure omitted — 356 chars of source]

In order to investigate the Chernoff-Savage result for the AHRT-$\phi$ test even further, we consider in Figure (ref) the AHRT-$\phi$ test for i.i.d errors generated with nine different true innovation densities $f$. These include innovation densities $f$ that are extremely heavy-tailed, skewed, or both. The first row of graphs shows three extremely heavy-tailed distributions: Student $t_2$, Student $t_1$, and a stable distribution with parameter values $0.5$ for stability, $0$ for skewness, $1$ for scale, and $0$ location. As these densities do not all satisfy our maintained assumptions, these graphs exclude power envelopes and the AHRT-$\hat{f}$ power functions. The top three graphs in Figure (ref) show that the AHRT-$\phi$ is much more powerful than its competitors and that its power increases with the heaviness of the tail. The second and third row show the effect of skewness in $f$. Specifically, the AHRT-$\phi$'s power is higher when $f$ is skewed-normal (with skewness 0.8145) than that when $f$ is normal (in Figure (ref)). This indicates that the AHRT-$\phi$ can acquire power from skewness. The same conclusion can be drawn from the comparison of the AHRT-$\phi$ power function for $t_4$ and that of a skewed $t_4$ with skewness $\approx 2.7$. To further remove the effects of the other moments, in the third row, we also employ the Pearson distributions with identical mean, variance and kurtosis, but different skewness --- ${\rm skewness}=1$ for Pearson-I, ${\rm skewness}=3$ for Pearson-II, and ${\rm skewness}=6$ for Pearson-III. Comparing the corresponding power functions, it validates again that the larger the skewness of the true distribution $f$ is, the more powerful the ${\rm AHRT}^\phi$ becomes.

A final remark on the size of the AHRT tests. In all cases where the true density $f$ satisfies our maintained assumption, i.e., $f\in\mathfrak{F}$ (that is all cases in Figure (ref) and the skewnormal, $t_4$, Pearson-I, Pearson-II, and Pearson-III in Figure (ref)), the simulated sizes are between $4.9\%$ and $5.1\%$. This verifies the validity of the AHRTs claimed in Theorem (ref). In the other cases, i.e., $f\not\in\mathfrak{F}$, the AHRT is somewhat conservative. More precisely, the simulated sizes of the AHRT-$\phi$ are $4.8\%$, $4.1\%$, $3.7\%$, and $4.7\%$ for the Student $t_2$, $t_1$, stable, and skew-$t_4$ distribution, respectively. This result seems consistent over all simulations.

Small-sample performance

We also report the performance of the AHRTs and the two competitors introduced above for smaller samples. Figures (ref) and (ref) are the small-sample versions, with $T = 100$, of Figures (ref) and (ref), respectively. We observe that, even with a slight downward shift of the power functions for all three tests considered, the findings of the large-sample case remain valid in this small-sample case. For larger values of $h$, the DF-$\rho$ test sometimes dominates the other two tests. This is due to the fact that the DF-$\rho$ in this Monte-Carlo setting appears to have a superior convergence speed (towards its asymptotic power as sample size $T$ increases) to those of the AHRTs and the ERS test. This phenomenon appears only in (ultra)-small-sample cases, and disappears when $T$ gets larger (e.g., $T = 200$). In cases with enough samples (say, $T \geq 100$) and when $f$ is significantly away from the Gaussian density, irrespective of the choice of $g$, the AHRTs performs favorably.

figure[figure omitted — 369 chars of source]
figure[figure omitted — 294 chars of source]

Concerning the small-sample size, we find it to range from about $4.0\%$ to $4.5\%$ for the cases where $f\in\mathfrak{F}$. Again, when $f$ does not satisfy our maintained assumptions ($f\not\in\mathfrak{F}$) the AHRT turns out to be conservative. More precisely, we find a size of $3.7\%$, $3.1\%$, $2.4\%$, and $4.1\%$ for the $t_2$, $t_1$, stable, and skew-$t_4$ distribution, respectively. This makes the improved power even more remarkable.

It may also be useful to illustrate the convergence of the power function of the AHRT-${f}$ to the semiparametric power envelope as sample size $T$ increases. This is the purpose of Figure (ref). For three cases: Gaussian, Laplace, and Student $t_3$, we find that the convergence indeed occurs already at relatively small samples, which is not always the case for alternative unit root tests.

figure[figure omitted — 212 chars of source]

Performance of signed-rank-based AHRT test under restriction $f \in \mathfrak{F}_\mathbb{S}$

In the case with the additional symmetric density restriction, $f \in \mathfrak{F}_\mathbb{S}$, we repeat the above Monte Carlo study above but for the signed-rank-based AHRT $\phi_g^\mathbb{S}(\bar{h},\alpha)$ introduced in Section (ref). Not surprisingly, we find very similar results as above, except for a slightly higher power for all the AHRT tests (essentially due to the symmetry restriction imposed, in which case we have the adaptivity result). Since it's a bit unfair to compare with the competitor which actually does not benefit from this constraint, and also to conserve space, we put the associated simulation results --- two figures that can be treated as the signed-rank-based AHRT's counterparts of Figure (ref) and Figure (ref) --- in the Appendix C in (ref). In short, the large-sample result shows that the power function of the signed AHRT-$f$ is tangent to the (parametric) power envelope, which corroborates the point-optimal property in Theorem (ref) and in turn the adaptive result in Section (ref); the Chernoff-Savage result still holds for the signed-rank version of the AHRT-$\phi$ test. The second figure therein shows that it works well in the small sample case. Furthermore, one can also check the convergence of the signed-rank-based AHRT-${f}$ power function to the parametric power envelope as sample size $T$ increases in Figure (ref).

Simulation results with ARMA innocations

figure[figure omitted — 344 chars of source]
figure[figure omitted — 340 chars of source]

In this section we provide simulation results for the cases where the innovations follow an ARMA model, i.e.,

align*[align* omitted — 146 chars of source]

where the assumption on $\varepsilon_t$ stay unchanged. This corresponds to a lag-polynomial $\Gamma(L)$ with an infinite order, which we employ here in order to show by simulations the potential that the finite order $p$ restriction could be relaxed. As for the estimation of $\widehat{\Gamma}(L)$ in ((ref)) used by the AHRT test and its competitors, we fix the AR regression order $p = 8$. Since there exists efficiency loss in the estimation of error autocorrelation parameters $\Gamma_1,\dots,\Gamma_p$ (for finite-sample cases), here we increase the sample sizes for both the large-sample case (to $T = 5,000$) and the small-sample case (to $T = 250$). Under the same organization of true and reference densities as for the i.i.d.\ case, we provide the simulation results for large-sample case and small-sample case in Figure (ref) and Figure (ref), respectively. Note that, in Figure (ref), we remove the powers of the ADF-$\rho$ test since it is severely oversized.

From these results, we draw similar conclusions to those in the i.i.d.\ case, except for a slight power loss for all these chosen unit root tests. These results validate that, on one hand, the serial correlation in the errors can be well handled, as for the cases of other unit root tests, with an additional auto-regression for the increments of the observed process. On the other hand, they also indicate that even though in the limit the unknown autocorrelation structure has no effects on the inference for $\rho$, its estimation consumes some efficiency in finite sample cases. Yet, Figure (ref) shows that the AHRT suffers less from this kind of efficiency loss or, in other words, the AHRT has a faster convergence speed of its power to the associated envelope as $T\to\infty$.

Conclusion

This paper has provided a structural representation of the limit experiment of the standard unit root model in a univariate but semiparametric setting. Using invariance arguments, we have derived the semiparametric power envelope. These invariance structures also lead, using the Neyman-Pearson lemma, to point-optimal semiparametric tests. The analysis naturally leads to the use of rank-based statistics.

Our tests are asymptotically valid, invariant, and (with a correctly chosen reference density) point-optimal. Moreover, we establish a Chernoff-Savage type property of our test: irrespective of the reference density chosen, our test outperforms its classical competitor which in this case is the ERS test. Finally, we introduced a simplified version of our test and show, in a Monte-Carlo study, that our theoretical results carry over to finite samples.

As potential future work we mention the use of similar ideas to construct hybrid rank-based tests in more general time-series models with, for instance, a deterministic time trend term, or stochastic volatility. Also, the structural representation of the limit experiment and its invariance properties could be applied to other non-stationary time-series models, for instance, cointegration and predictive regression models.

supplement\sname{Supplement A} \stitle{Supplement to “Semiparametrically optimal hybrid rank tests for unit roots”} \sdescription{This supplemental file contains technical proofs for propositions and theorems in the main context.}