Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
116,508 characters · 24 sections · 67 citation commands
Semiparametrically Point-Optimal Hybrid Rank Tests for Unit Roots
The monographs of Patterson2011 ({Patterson2011}, Patterson2012) and Choi2015 provide an overview of the literature on unit roots tests. This literature traces back to White58 and includes seminal papers as DickeyFuller79 (DickeyFuller79, DickeyFuller81), Phillips1987, PhillipsPerron88, and ERS1996. The present paper fits into the stream of literature that focuses on “optimal” testing for unit roots. Important earlier contributions here are DufourKing91, SaikkonenLuukkonen1993, and ERS1996. The latter paper derives the asymptotic power envelope for unit root testing in models with Gaussian innovations. RothenbergStock97 and Jansson2008 consider the non-Gaussian case.
The present paper considers testing for unit roots in a semiparametric setting. Following earlier literature, we focus on a simple AR(1) model driven by possibly serially correlated errors. The innovations driving these serially correlated errors are i.i.d., whose distribution is considered a nuisance parameter. Apart from some smoothness and the existence of relevant moments, no assumptions are imposed on this distribution. From earlier work it is known that the unit root model leads to Locally Asymptotically Brownian Functional (LABF) limit experiments Jeganathan1995. As a consequence, no uniformly most powerful test exists (even in case the innovation distribution would be known) -- see also ERS1996. In the semiparametric case the limit experiment becomes more difficult due to the infinite-dimensional nuisance parameter. Jansson2008 derives the semiparametric power envelope by mimicking ideas that hold for Locally Asymptotically Normal (LAN) models. However, the proposed test needs a nonparametric score function estimator which complicates its implementation. The point-optimal tests proposed in the present paper only require nonparametric estimation of a real-valued cross-information factor and we also provide a simplified version that avoids any nonparametric estimation.
The main contribution of this manuscript is twofold. First, we derive the semiparametric power envelopes of unit root tests with serially correlated errors for two cases: symmetric or possibly non-symmetric innovation distributions (Section (ref)). Our method of derivation is novel and exploits the invariance structures embedded in the semiparametric unit root model. To be precise, we use a “structural” description of the LABF limit experiment (Section (ref)), obtained from Girsanov's theorem. This limit experiment corresponds to observing an infinitely-dimensional Ornstein-Uhlenbeck process (on the time interval $[0,1]$). The unknown innovation density in the semiparametric unit root model takes the form of an unknown drift parameter in this limit experiment. Within the limit experiment, Section (ref) derives the maximal invariant, i.e., a reduction of the data which is invariant with respect to the nuisance parameters (that is, the unknown drift in the limiting Ornstein-Uhlenbeck experiment). It turns out that this maximal invariant takes a rather simple form: all processes associated to density perturbations have to be replaced by their associated bridges (i.e., consider $W(s) - sW(1)$ for the process $W(s)$ with $s\in[0, 1]$). The power envelopes for invariant tests in the limit experiment then readily follow from the Neyman-Pearson lemma. An application of the Asymptotic Representation Theorem (see, e.g., Theorem 15.1 in vdVaart00) subsequently yields the local asymptotic power envelope (Theorem (ref)). In case the innovation density is known to be symmetric, the semiparametric power envelope coincides with the parametric power envelope. This implies the existence of an adaptive testing procedure (see also Jansson2008). Moreover, we note that our analysis of invariance structures in the LABF experiment is also of independent interest and could, for example, be exploited in the analysis of optimal inference for cointegration or predictive regression models. Also, the analysis gives an alternative interpretation of the test proposed in ERS1996 --- the ERS test --- as this test is also based on an invariant, though not the maximal one (see Remark (ref)).
As a second contribution, we provide two new classes of easy-to-implement unit root tests that are semiparametrically optimal in the sense that their asymptotic power curves are tangent to the associated semiparametric power envelopes (Section (ref)). The form of the maximal invariant developed before suggests how to construct such tests based on the ranks/signed-ranks (depending on whether the innovation density is known to be symmetric or not) of the increments of the observations, the average of these increments, and an assumed reference density $g$. These tests are semiparametric in the sense that the reference density need not equal the true innovation density, while they are still valid (i.e., provide the correct asymptotic size). The reference density is not restricted to be Gaussian, which it generally needs to be in more classical QMLE results. When the reference density is correctly specified (i.e., $g$ happens to be equal to the true density $f$), the asymptotic power curve of our test is tangent to the semiparametric power envelope, and this in turn gives the optimality property. A feasible version of the oracle test using $g=f$ is obtained by using a nonparametrically estimated density $\hat{f}$, of which the corresponding simulation results are provided in Section (ref).
In relation to the classical literature on efficient rank-based testing (for instance, HajekSidak1967, HallinPuri1988, and HvdAW1) our approach can be interpreted as follows. In the aforementioned papers, the invariance arguments (that is, using the ranks of the innovations) are applied in the sequence of models at hand. We, on the other hand, only apply the invariance arguments in the limit experiment. In this way, we can extend these ideas to non-LAN experiments. For the LAN case, both approaches would effectively lead to the same results. Our tests, despite the absence of a LAN structure, satisfy a ChernoffSavage1958 type result (Corollary (ref)): for any reference density our test outperforms, at any true density, its classical counterpart which, in this case, is the ERS test. We provide, in Section (ref), even simpler alternative classes of tests that require no nonparametric estimations at all. These (simplified) classes of tests coincide with their corresponding originals for correctly specified reference density and, hence, share the same optimality properties. In case of misspecified reference density, the alternative classes still seem to enjoy the Chernoff-Savage type property, though only for a Gaussian reference density. This is in line with with the traditional Chernoff-Savage results for Locally Asymptotically Normal models.
The remainder of this paper is organized as follows. Section (ref) introduces the model assumptions and some notation. Next, Section (ref) contains the analysis of the limit experiment. In particular we study invariance properties in the limit experiment leading to our new derivation of the semiparametric power envelopes. The classes of hybrid rank tests we propose are introduced in Section (ref). Section (ref) provides the results of a Monte Carlo study and Section (ref) contains a discussion of possible extensions of our results. All proofs are organized in (ref).
Consider observations $Y_1,\dots,Y_T$ generated from the classical component specification, for $t\in\mathbb{Z}_{+}$,
where $v_0=v_{-1}=\cdots=v_{1-p}=0$, the innovations $\{\varepsilon_t\}$ form an i.i.d.\ sequence defined for $t\in\mathbb{Z}$ with density $f$, and $\Gamma(L)$ is the AR(p) lag polynomial. Moreover, it is assumed that $Y_0 = \mu$. We impose the following assumptions on this innovation density.
The imposed smoothness on $f$ is mild and standard (see, e.g., LeCam1986, vdVaart00). The finite variance assumption (b) is important to our asymptotic results as it is essential to the weak convergence, to a Brownian motion, of the partial-sum process generated by the innovations.\footnote{Let us already mention that, although not allowed for in our theoretical results, we will also assess the finite-sample performances of the proposed tests (Section (ref)) for innovation distributions with infinite variance. For tests specifically developed for such cases we refer to Hasan2001, AhnFotopoulosHe03, and CallegariCappuccioLubian03.} The zero intercept assumption in (b) excludes a deterministic trend in the model. Such a trend leads to an entirely different asymptotic analysis, see HvdAW1. The Fisher information $J_{f}$ in (c) has been standardized by premultiplying with the variance $\sigma_f^2$, so that it becomes scale invariant (i.e., invariant with respect to $\sigma_f$). In other words, $J_{f}$ only depends on the shape of the density $f$ and not on its variance $\sigma_f^2$. The positivity of the density $f$ in (d) is mainly made for notational convenience.
The assumption on the initial condition, $v_0=v_{-1}=\cdots=v_{1-p}=0$, is less innocent then it may appear. Indeed, it is known, see MullerElliott2003 and ElliottMuller2006, that, even asymptotically, the initial condition can contain non-negligible statistical information. Nevertheless, it is still stronger than necessary for the sake of simplicity, and it can be relaxed to the level of generality in ERS1996.
Let $\mathfrak{F}$ denote the set of densities satisfying Assumption (ref). We also investigate in the present paper the special case of symmetric densities. For that purpose, we denote by $\mathfrak{F}_\mathbb{S}$ the set of densities which satisfy Assumption (ref) and, at the same time, are symmetric about zero. Of course, it follows $\mathfrak{F}_\mathbb{S}\subset\mathfrak{F}$.
With respect to the autocorrelation structure $\Gamma(L)$, we impose the following assumption.
Let $\mathfrak{G}\subset\mathbb{R}^p$ denote the set of $\left(\Gamma_1,\ldots,\Gamma_p\right)'$ such that the induced lag polynomial satisfies Assumption (ref). For mathematical convenience we restrict the lag polynomial to be of finite order $p$. One may expect many of the results in the present paper to extend to the case $p=\infty$ (see, e.g., Jeganathan1997).
The main goal of this paper is to develop tests, with optimality features, for the semiparametric unit root hypothesis
i.e., apart from Assumption (ref)-(ref), no further structure is imposed on $f$, the intercept $\mu$, and the autocorrelation structure $\Gamma(L)$.
In the following section, we derive the (asymptotic) power envelope of tests that are (locally and asymptotically) invariant with respect to the nuisance parameters $\mu$, $\Gamma$, and $f$. We consider both the non-symmetric ($f\in\mathfrak{F}$) and the symmetric case ($f\in\mathfrak{F}_\mathbb{S}$). Section (ref) is subsequently devoted to tests, depending on a reference density $g$ that can be freely chosen, that are point optimal with respect to this power envelope and proves a Chernoff-Savage type result.
This section first introduces some notations and preliminaries (Section (ref)). Afterwards, we will derive the limit experiment (in the Le Cam sense) corresponding to the component unit root model ((ref))-((ref)) and provide a “structural” representation of this limit experiment (Section (ref)). In Section (ref) we discuss, exploiting this structural representation, a natural invariance restriction, to be imposed on tests for the unit root hypothesis with respect to the infinite-dimensional nuisance parameter associated to the innovation density. We derive the maximal invariant and obtain from this the power envelope for invariant tests in the limit experiment. At last, in Section (ref), we exploit the Asymptotic Representation Theorem to translate these results to obtain (asymptotically) optimal invariant test in the sequence of unit root models. Again we consider both the case of unrestricted densities $f\in\mathfrak{F}$ and that of symmetric densities $f\in\mathfrak{F}_\mathbb{S}$.
We first introduce local reparameterizations for the parameter of interest $\rho$ and the nuisance autocorrelation structure $\Gamma(L)$. Then we discuss a convenient parametrization of perturbations to the innovation density $f$ which we use to deal with the semiparametric nature of the testing problem. These perturbations follow the standard approach of local alternatives in (semiparametric) models commonly used in experiments that are Locally Asymptotically Normal (LAN). We will see that, with respect to all parameters but $\rho$, the model is actually LAN; compare also Remark (ref) below. Moreover, we introduce some partial sum processes that we need in the sequel, as well as their Brownian limits.
It is well-known, and goes back to Phillips1987, ChanWei88 and PhillipsPerron88, that the contiguity rate for the unit root testing problem, i.e., the fastest convergence rate at which it is possible to distinguish (with non-trivial power) the unit root $\rho=1$ from a stationary alternative $\rho<1$, is given by $T^{-1}$. Therefore, in order to compare performances of tests with this proper rate of convergence, we reparametrize the autoregression parameter $\rho$ into its local-to-unity form, i.e.,
The appropriate local reparameterization for the lag polynomial $\Gamma(L)$ is of a traditional form with rate $\sqrt{T}$, i.e.,
where the local perturbation $\gamma$ is defined by $\gamma(L) := - \gamma_1 L - \cdots - \gamma_p L^p$ with local parameter $\gamma := (\gamma_1,\dots,\gamma_p)'\in\mathbb{R}^p$. As $\mathfrak{G}$ is open, $\Gamma^{(T)}_\gamma \in \mathfrak{G}$ for $T$ large enough.
To describe the local perturbations to the density $f$, we need the separable Hilbert space
where $\mathrm{L}_2^f(\mathbb{R},\mathcal{B})$ denotes the space of Lebesgue-measurable functions $b:\,\mathbb{R}\to\mathbb{R}$ satisfying $\int b^2(e) f(e)\mathrm{d} e<\infty$. Because of the separability, there exists a countable orthonormal basis $b_k$, $k\in\mathbb{N}$, of $\mathrm{L}_2^{0,f}$ (see, e.g., Rudin1987). This basis can be chosen such that $b_k \in C_{2,b}(\mathbb{R}) $, for all $k$, i.e., each $b_k$ is bounded and two times continuously differentiable with bounded derivatives. Moreover, $\mathrm{E}_{f}b_k(\varepsilon)=0$ and $\operatorname{Var}_f{b_k(\varepsilon)}=1$. Hence each function $b\in \mathrm{L}_2^{0,f}$ can be written as $b= \sum_{k=1}^\infty \eta_k b_k $, for some $\eta := (\eta_k)_{k\in\mathbb{N}} \in \ell_2=\{ (x_k)_{k\in\mathbb{N}}\, | \, \sum_{k=1}^\infty x_k^ 2<\infty\}$. Besides the sequence space $\ell_2$ we also need the sequence space $c_{00}$ which is defined as the set of sequences with finite support, i.e., \[ c_{00}=\left\{ (x_k)_{k\in\mathbb{N}}\in\mathbb{R}^\mathbb{N} \, \left|\, \sum_{k=1}^\infty 1\{x_k\neq 0\}<\infty\right. \right\}. \] Of course, $c_{00}$ is a dense subspace of $\ell_2$. Given the orthonormal basis $b_k$ and $\eta\in c_{00}$, we introduce the following perturbation to the density $f$:
The rate $T^{-1/2}$ is already indicative of the standard LAN behavior of the nuisance parameter $f$ as will formally follow from Proposition (ref) below.
For the symmetric case, we can assume that the perturbations $b_k$, $k\in\mathbb{N}$, are also symmetric about zero.
The following proposition shows that the perturbations, both for the non-symmetric and for the symmetric case, are valid in the sense that they satisfy the conditions on the innovation density that we imposed throughout on the model (Assumption (ref)). The proof is organized in (ref).
To describe the limit experiment in Section (ref), we introduce some partial sum processes and their limits. These results are fairly classical but, for completeness, precise statements are organized in Lemma A.1 in the supplementary material.
Define, for $s\in [0,1]$, the partial sum processes
Note that we pick the starting point of the sums at $t=p+2$, so that these partial sum processes are (maximally) invariant with respect to the intercept $\mu$ (otherwise, e.g., for $t=p+1$, the term $\Gamma(L)\Delta Y_t$ contains $\Delta Y_1 = Y_1 - \mu$).
Using Assumption (ref) we find, under the null hypothesis, joint weak convergence of observation processes $W^{(T)}_{\varepsilon}$, $W^{(T)}_{\phi_f}$, $W^{(T)}_{\Gamma}$, and $W^{(T)}_{b_k}$ to Brownian motions that we denote by $W_{\varepsilon}$, $W_{\phi_f}$, $W_{\Gamma}$ and $W_{b_k}$, respectively.\footnote{All weak convergences in this paper are in product spaces of $D[0,1]$ with the uniform topology.} These limiting Brownian motions are defined on a probability space $(\Omega,\mathcal{F},\mathbb{P}_{0,0,0})$. Let us already mention that we will introduce a collection of probability measures $\mathbb{P}_{h,\gamma,\eta}$, on $(\Omega, \mathcal{F})$, representing the limit experiment, in Section (ref). We use the notational convention that probability measures related to the limit experiment (i.e., to the “W-processes”) are denoted by $\mathbb{P}$, while probability measures related to the finite-sample unit root model, i.e., observing $Y_1, \dots ,Y_T$, will be denoted by $\mathrm{P}^{(T)}_{}$.
We remark that integrals like $\int_0^1W_{\varepsilon}^{(T)}(s-) \rdW_{\phi_f}^{(T)}(s)$ can be shown to converge weakly, under the null hypothesis, to the associated stochastic integral with the limiting Brownian motions, i.e., to $\int_0^1 W_{\varepsilon}(s)\rdW_{\phi_f}(s)$. Weak convergence of integrals like $\int_0^1(W_{\varepsilon}^{(T)}(s-))^2 \mathrm{d} s$ follows from an application of the continuous mapping theorem. Again, details can be found in the proof of Proposition (ref) in the (ref).
As $\varepsilon$ and $b_k(\varepsilon)$ are orthogonal for each $k$, it holds that $W_{\varepsilon}$ and $W_{b_k}$, $k\in\mathbb{N}$, are all mutually independent. Moreover, we have
As $\phi_f(\varepsilon)$ is the score of the location model, it is well known (see, for example, BKRW) that we have (under Assumption (ref)) $\mathrm{E}_{f}[\phi_f(\varepsilon)]=0$ and $\mathrm{E}_{f}[\varepsilon\phi_f(\varepsilon)] =1$. Consequently, for $f\in \mathfrak{F}$, again because $\varepsilon$ and $b_k(\varepsilon)$ are orthogonal for each $k$, we can decompose
with coefficients $J_{f,k}:=\sigma_f\mathrm{E}_{f}[b_k(\varepsilon)\phi_f(\varepsilon)]$. This establishes
Moreover, we have, for $k\in\mathbb{N}$,
and
As for $W_{\Gamma}$: since, under the null, $\left(\Delta Y_{t-1},\dots,\Delta Y_{t-p}\right)'$ is independent of the innovation $\varepsilon_t$ and $\mathrm{E}_{0,0,0}\left[\Delta Y_{t-i}\right]=0$, $i=1,\dots,p$, it follows
and
We define the covariance matrix of $W_{\Gamma}(1)$ as
with
In this case, the density function $f(\varepsilon)$ is an even function and so are the perturbation functions $b_k(\varepsilon)$, $k\in\mathbb{N}$. The score $\phi_f(\varepsilon)$ is now an odd function. Therefore, $\phi_f(\varepsilon)$ cannot be decomposed by $\varepsilon$ and $b_k(\varepsilon)$ anymore as in ((ref)) and ((ref)). Instead of ((ref)) we now have
for all $f\in\mathfrak{F}_\mathbb{S}$ and $k\in\mathbb{N}$. All the other results mentioned above still hold.
The results in the previous section are needed to study the asymptotic behavior of log-likelihood ratios. These in turn determine the limit experiment, which we use to study asymptotically optimal procedures invariant with respect to the nuisance parameters. Thus, fix $\mu\in\mathbb{R}$, $\Gamma\in\mathfrak{G}$, and $f\in\mathfrak{F}$. Let, for $h\in\mathbb{R}$, $\gamma\in\mathbb{R}^p$, and $\eta\in c_{00}$, $\mathrm{P}^{(T)}_{h,\gamma,\eta;\mu,\Gamma,f}$ denote the law of $Y_1,\dots,Y_T$ under ((ref))-((ref)) with parameter $\rho$ given by ((ref)), $\Gamma^{(T)}_\gamma(L)$ given by ((ref)), and innovation density ((ref)). The following proposition shows that the semiparametric unit root model is of the Locally Asymptotically Brownian Functional (LABF) type introduced in Jeganathan1995.
Of course, Proposition (ref) still holds for $f\in\mathfrak{F}_\mathbb{S}$; in that case we have $J_{f,k}=0$. The proof of (i) follows by an application of Proposition 1 in HvdAW2 which provides generally applicable sufficient conditions for the quadratic expansion of log likelihood ratios. Of course, Part (ii) is not surprising and follows using the weak convergence of the partial sum processes to Brownian motions (and integrals involving the partial sum processes to stochastic integrals) discussed above. Finally, Part (iii) follows by verifying the Novikov condition. For the sake of completeness, detailed proofs are organized in (ref).
Part (iii) of the proposition implies that we can introduce, for $h\in\mathbb{R}$, $\gamma\in\mathbb{R}^p$, and $\eta\in c_{00}$, new probability measures $\mathbb{P}_{h,\gamma,\eta}$ on the measurable space $(\Omega,\mathcal{F})$ (on which the processes $W_{\varepsilon}$, $W_{\phi_f}$, $W_{\Gamma}$ and $W_{b_k}$ were defined) by their Radon-Nikodym derivatives with respect to $\mathbb{P}_{0,0,0}$:
Proposition (ref) then implies that the sequence of (local) unit root experiments (each $T\in\mathbb{N}$ yields an experiment) weakly converges (in the Le Cam sense) to the experiment described by the probability measures $\mathbb{P}_{h,\gamma,\eta}$. Formally, we define the sequence of experiments of interest by
for $T\in\mathbb{N}$, and the limit experiment by, with $\mathcal{B}_C$ the Borel $\sigma$-field on $C[0,1]$,
Note that the latter experiment indeed depends on $f$ as the measure $\mathbb{P}_{h,\gamma,\eta}$ depends on $f$.
The Asymptotic Representation Theorem (see, for example, Chapter 9 in vdVaart00) implies that for any statistic $A_T$ which converges in distribution to the law $L_{h,\gamma,\eta}$, under $\mathrm{P}^{(T)}_{h,\gamma,\eta;\mu,\Gamma,f}$, there exists a (possibly randomized) statistic $A$, defined on $\mathcal{E}(f)$, such that the law of $A$ under $\mathbb{P}_{h,\gamma,\eta}$ is given by $L_{h,\gamma,\eta}$. This allows us to study (asymptotically) optimal inference: the “best” procedure in the limit experiment also yields a bound for the sequence of experiments. If one is able to construct a statistic (for the sequence) that attains this bound, it follows that the bound is sharp and the statistic is called (asymptotically) optimal. This is precisely what we do: Section (ref) establishes the bound and in Section (ref) we introduce statistics attaining it.
To obtain more insight in the limit experiment $\mathcal{E}(f)$ the following proposition, which follows by a direct application of Girsanov's theorem, provides a “structural” description of the limit experiment.
For the case $f\in\mathfrak{F}_\mathbb{S}$, Proposition (ref) still applies with $J_{f,k}=0$. Moreover, for this case, we denote by $\mathcal{E}^{(T)}_{\mathbb{S}}(f)$ the associated sequence of experiments and by $\mathcal{E}_\mathbb{S}(f)$ the associated limit experiment.
In this section, we consider the limit experiments $\mathcal{E}(f)$ and $\mathcal{E}_\mathbb{S}(f)$. In these experiments, we observe the processes $W_{\varepsilon}$, $W_{\phi_f}$, $W_{\Gamma}$, and $W_{b_k}$, $k\in\mathbb{N}$, continuously on the time interval $[0,1]$ from the model $(\mathbb{P}_{h,\gamma,\eta}\, |\, h\in\mathbb{R},\,\gamma\in\mathbb{R}^p,\,\eta\in c_{00})$, and we are interested in the power envelopes for testing the hypothesis
To eliminate the nuisance parameters $\gamma$ and $\eta$, we first propose a statistic that is sufficient for the parameter of interest $h$ and does not depend on $\gamma$. Afterwards, using Proposition (ref), we discuss a natural invariance structure with respect to the infinite-dimensional nuisance parameter $\eta$. We derive the maximal invariant and apply the Neyman-Pearson lemma to obtain the power envelopes of invariant tests within the experiments $\mathcal{E}(f)$ and $\mathcal{E}_\mathbb{S}(f)$, respectively in Section (ref) and Section (ref).
We begin with the elimination of $\gamma$. The statistic $\left(W_{\varepsilon},W_{\phi_f},(W_{b_k})_{k\in\mathbb{N}}\right)$ serves as a sufficient statistic for the parameter $h$. This is because, according to the structural version of $\mathcal{E}(f)$ (or $\mathcal{E}_\mathbb{S}(f)$) in Proposition (ref), the distribution of $W_{\Gamma}$ is only affected by $\gamma$ and so is the distribution of the process $\left(W_{\varepsilon},W_{\phi_f},(W_{b_k})_{k\in\mathbb{N}}\right)$ only affected by $h$ and $\eta$. It then follows that the distribution of the statistic $\left(W_{\varepsilon},W_{\phi_f},W_{\Gamma},(W_{b_k})_{k\in\mathbb{N}}\right)$ conditional on the statistic $\left(W_{\varepsilon},W_{\phi_f},(W_{b_k})_{k\in\mathbb{N}}\right)$ is only a function of $W_{\Gamma}$ and $\gamma$ and, in turn, does not depend on $h$. Next, observe that the sufficient statistic $\left(W_{\varepsilon},W_{\phi_f},(W_{b_k})_{k\in\mathbb{N}}\right)$ is independent of the process $W_{\Gamma}$ and thus has a distribution that does not depend on $\gamma$. This allows us to restrict attention to the sufficient statistic $\left(W_{\varepsilon},W_{\phi_f},(W_{b_k})_{k\in\mathbb{N}}\right)$ to conduct inference for $h$, and the nuisance parameter $\gamma$ disappears from the likelihoods.
The elimination of the nuisance parameter $\eta$ is more involved and different for the limit experiments $\mathcal{E}(f)$ and $\mathcal{E}_\mathbb{S}(f)$. We start with $\mathcal{E}(f)$.
In Proposition (ref), if one applies the decompositions ((ref)) and ((ref)) to the first equation (of $W_{\varepsilon}$) and the fourth equation (of $W_{b_k}$), one retrieves the second equation (of $W_{\phi_f}$). This essentially allows us to omit the process $W_{\phi_f}$ and restrict the observations to the processes $W_{\varepsilon}$ and $W_{b_k}$, $k\in\mathbb{N}$.
Now we formalize the invariance structure with respect to $\eta$. Introduce, for $\eta\in c_{00}$, the transformation $g_\eta=(g_{\eta_k})_{k\in\mathbb{N}}: C^\mathbb{N}[0,1]\to C^\mathbb{N}[0,1]$ defined by, for $W\in C[0,1]$,
i.e., $g_{\eta_k}$ adds a drift $s\mapsto -\eta_k s$ to its argument process. Proposition (ref) implies that the law of $(W_{\varepsilon},(g_{\eta_k}(W_{b_k}))_{k\in\mathbb{N}})$ under $\mathbb{P}_{h,\gamma,0}$ is the same as the law of $(W_{\varepsilon},(W_{b_k})_{k\in\mathbb{N}})$ under $\mathbb{P}_{h,\gamma,\eta}$. Hence our testing problem ((ref)) is invariant with respect to the transformations $g_\eta$. Therefore, following the invariance principle, it is natural to restrict attention to test statistics that are invariant with respect to these transformations as well, i.e., test statistics $t$ that satisfy
Given a process $W$ let us define the associated bridge process by $B^W(s)=W(s)- s W(1)$. Now note that we have, for all $s\in [0,1]$ and $k\in\mathbb{N}$,
i.e., taking the bridge of a process ensures invariance with respect to adding drifts to that process. Define the mapping $M$ by
where $B_{b_k} := B^{W_{b_k}}$. It follows that statistics that are measurable with respect to the $\sigma$-field,
are invariant (with respect to $g_\eta$, $\eta\in c_{00}$). It is, however, not (immediately) clear that we did not throw away too much information. Formally, we need $\mathcal{M}$ to be maximally invariant which means that each invariant statistic is $\mathcal{M}$-measurable. The following theorem, which once more exploits the structural description of the limit experiment, shows that this indeed is the case.
The above theorem implies that invariant inference must be based on $\mathcal{M}$. An application of the Neyman-Pearson lemma, using $\mathcal{M}$ as observation, yields the power envelope for the class of invariant tests. To be precise, consider the likelihood ratios restricted to $\mathcal{M}$, which are given by
where the conditional expectation indeed does not depend on $\eta$ precisely because of the invariance. The conditional expectation also does not depend on $\gamma$ as a result of the arguments stated at the beginning of this section.
To calculate this conditional expectation, we first introduce $B_{\phi_f}=B^{W_{\phi_f}}$, i.e., the bridge process associated to $W_{\phi_f}$. Following the decomposition in ((ref)), we can decompose $\Delta_{f}=\int_0^1W_{\varepsilon}(s)\rdW_{\phi_f}(s)=I+II$ with
Note that part $I$ is $\mathcal{M}$-measurable. Under $\mathbb{P}_{0,0,0}$ the random variables $W_{b_k}(1)$, $k\in\mathbb{N}$, are independent of $W_{\varepsilon}$ and $B_{b_k}$, $k\in\mathbb{N}$. Indeed, the independence of $W_{\varepsilon}$ holds by construction and the independence of $B_{b_k}$ is a well-known, and easy to verify, property of Brownian bridges. We thus obtain, since $\mathcal{I}{}(h,\gamma,\eta)$ is $\mathcal{M}$-measurable as well,
with
where the last equality follows from ((ref)). Note that the distribution of this likelihood ratio indeed does not depend on the nuisance parameters $\gamma$ or $\eta$.
We can now formalize the notion of point-optimal invariant tests in the limit experiment. To that end, let us denote the $(1-\alpha)$-quantile of $\mathrm{d}\mathbb{P}_{h}^{\mathcal{M}}/\mathrm{d}\mathbb{P}_{0}^{\mathcal{M}}$ under $\mathbb{P}_{0,\gamma,\eta}$, which does not depend on either $\gamma$ or $\eta$, by $c(h,J_f;\alpha)$. Define the size-$\alpha$ test $\phi^{\mathfrak{F}*}_{f,\alpha}(\bar{h}) = \mathbbm{1}\left\{\mathrm{d}\mathbb{P}_{\bar{h}}^{\mathcal{M}} / \mathrm{d} \mathbb{P}_{0}^{\mathcal{M}}\geq c(\bar{h},J_f;\alpha)\right\}$, for a fixed value of $\bar{h}<0$. Note that this is an oracle test in $\mathcal{E}(f)$ that depends on the true value of $f$. A feasible test is provided in Section (ref). The power function of this oracle test is given by
An application of the Neyman-Pearson lemma yields the following.
The (oracle) test $\phi^{\mathfrak{F}*}_{f,\alpha}(\bar{h})$ thus is point optimal, i.e., its power function is tangent to the (semiparametric) power envelope $h\mapsto \pi^{\mathfrak{F}*}_{f,\alpha}(h; h)$ at $h=\bar{h}$.\footnote{Here and later in this section, the early usage of the concept “power envelope” is due to the fact that it is shown to be the upper bound later in this section and point-wisely attainable by tests in sequence in Section (ref).}
Since in this case we have $J_{f,k}=0$, the structural representation of experiment $\mathcal{E}_\mathbb{S}(f)$ in Proposition (ref) becomes
Note that here the process $W_{\Gamma}$ is again removed by sufficiency in order to eliminate the nuisance parameter $\gamma$ (see the discussion at the beginning of Section (ref)). Following the same argument for the case of $f\in\mathfrak{F}$, statistics that are measurable with respect to the $\sigma$-field
are invariant with respect to the transformations $g_\eta$, $\eta\in c_{00}$. Moreover, in the following theorem, we show that this $\sigma$-field is maximally invariant.
The likelihood ratio restricted to $\mathcal{M}^{\mathbb{S}}$ is given by
where
As before, we then formalize the point-optimal invariant tests in $\mathcal{E}_\mathbb{S}(f)$. Denote the $(1-\alpha)$-quantile of $\mathrm{d}\mathbb{P}_{h}^{\mathcal{M}^{\mathbb{S}}}/\mathrm{d}\mathbb{P}_{0}^{\mathcal{M}^{\mathbb{S}}}$ under $\mathbb{P}_{0,\gamma,\eta}$ by $c_{\mathbb{S}}(h,J_f;\alpha)$. Define the size-$\alpha$ test $\phi^{\mathfrak{F}_\mathbb{S}*}_{f,\alpha}(\bar{h}) = \mathbbm{1}\left\{\mathrm{d}\mathbb{P}_{\bar{h}}^{\mathcal{M}^{\mathbb{S}}} / \mathrm{d} \mathbb{P}_{0}^{\mathcal{M}^{\mathbb{S}}}\geq c_{\mathbb{S}}(\bar{h},J_f;\alpha)\right\}$, for a fixed value of $\bar{h}<0$. The associated power function is
Again, by the Neyman-Pearson lemma, the following corollary holds.
The likelihood ratio of the maximal invariant $\mathcal{M}^{\mathbb{S}}$, $\mathrm{d} \mathbb{P}_{h}^{\mathcal{M}^{\mathbb{S}}} / \mathrm{d} \mathbb{P}_{0}^{\mathcal{M}^{\mathbb{S}}}$, equals the likelihood ratio of the full observation $(W_{\varepsilon},W_{\phi_f},W_{\Gamma},(W_{b_k})_{k\in\mathbb{N}})^\prime$ with known $\eta=0$ (i.e., known $f$) and $\gamma=0$, $\mathrm{d} \mathbb{P}_{h,0,0} / \mathrm{d} \mathbb{P}_{0,0,0}$. This shows that the “semiparametric” power envelope $\pi^{\mathfrak{F}_\mathbb{S}*}_{f,\alpha}$ actually coincides with the parametric power envelope, which is defined based on the likelihood ratio $\mathrm{d} \mathbb{P}_{h,0,0} / \mathrm{d} \mathbb{P}_{0,0,0}$. This verifies again the adaptivity result in Jansson2008 under the same conditions for this unit root testing problem.\footnote{A discussion about the notion of “adaptiveness” in this nonstandard testing problem can be found in Section 5 of Jansson2008.} We provide, in Sections (ref) and (ref), a class of adaptive unit root tests based on signed-rank statistics for this setting.
Now we translate the results for the limit LABF experiment to the unit root model of interest. To mimick the invariance in the limit experiment we introduce the following definition.
The Asymptotic Representation Theorem (see, e.g., vdVaart00 Chapter 9) now yields the following main result on the asymptotic power envelope.
These power envelopes for invariant tests in the limit experiments $\mathcal{E}(f)$ and $\mathcal{E}_\mathbb{S}(f)$ thus provide upper bounds to the asymptotic powers of invariant tests for the unit root hypotheses in $\mathcal{E}^{(T)}(f)$ and $\mathcal{E}^{(T)}_{\mathbb{S}}(f)$, respectively. The next section introduces two classes of tests (based on rank and signed-rank statistics) that (in a point-wise sense) attain these power envelopes and, thereby, demonstrates that these bounds are sharp. We additionally provide a Chernoff-Savage type result for these classes of tests.
The appearance of the bridge process $B_{\phi_f}$ in the “efficient central sequence” $\Delta^*_f$ naturally suggests the (partial) use of ranks in the construction of test statistics. Indeed, we can construct an empirical analogue of $B_{\phi_f}$ by considering a partial-sum process which only depends on the observations via the ranks $R_t$ of $\Gamma(L)\Delta Y_t$ amongst $\Gamma(L)\Delta Y_{p+2},\dots, \Gamma(L)\Delta Y_T$. We allow for the use of a reference density $g$ that may or may not be equal to the true underlying innovation density $f$. Our findings compare to Quasi-ML methods: if the true innovation density happens to be the same as the selected reference density the inference procedure is point-optimal. At the same time, the procedure remains valid, i.e., has proper asymptotic size, even in case the true innovation density does not coincide with the reference density. Note that these results also hold in case the reference density is non-Gaussian, while Quasi-ML results are generally restricted to Gaussian reference densities.
We need the following mild assumption on the reference density.
Now we can formulate the following direct extension of Lemma A.1 in HvdAW1 and its signed-rank counterpart. The proof for the weak convergence of the stochastic integrals in ((ref)) and ((ref)) is provided in (ref).
In practice, the autoregressive structure $\Gamma(L)$ will not be known. In that case, one must rely on ranks (and signs) based on innovations calculated using an estimated autoregressive structure, i.e., using residuals. This is usually referred to as inference based on aligned ranks. For this purpose, we first introduce the following assumption on the estimation of $\Gamma(L)$ (see HallinPuri1994).
Part (i) of Assumption (ref) is mild as many known estimators of $(\Gamma_{1},\dots,\Gamma_{p})^\prime$ exist for the AR($p$) model that can be applied to the increments $\Delta Y_t$, see, e.g., BrockwellDavis2002. Part (ii) is standard in the semiparametric literature and can easily be met by transforming the estimator in Part (i), see, e.g., Bickel1982 and Kreiss1987. Such an estimator is often called locally asymptotically discrete. The use of aligned ranks does not invalidate the conclusion of Lemma (ref). This is the content of the following result, of which the proof is provided in (ref).
We denote by $\widehat{B}^{(T)}_{\phi_g}$ and $\widehat{W}^{(T)}_{\phi_g}$ the aligned-rank-based counterparts of the rank-based processes $B^{(T)}_{\phi_g}$ and $W^{(T)}_{\phi_g}$, respectively.
In this section, we propose our unit root tests, for both the non-symmetric case ($f\in\mathfrak{F}$) and the symmetric case ($f\in\mathfrak{F}_\mathbb{S}$). Section (ref) provides the basis for optimal invariant tests in the limit experiment $\mathcal{E}(f)$ and $\mathcal{E}_\mathbb{S}(f)$. Lemmas (ref) and (ref) can subsequently be used to approximate, in the sequences of unit root experiments $\mathcal{E}^{(T)}(f)$ and $\mathcal{E}^{(T)}_{\mathbb{S}}(f)$, the observable processes in these limit experiments. More specifically, these two lemmas indicate that constructing the rank-based score partial sum processes $B^{(T)}_{\phi_g}$ and $W^{(T)}_{\phi_g}$ as in ((ref)) and ((ref)), together with $W^{(T)}_{\varepsilon}$, intuitively corresponds to observing the $\sigma$-fields $\mathcal{M}_{g} := \sigma(W_{\varepsilon}, B_{\phi_g})$ and $\mathcal{M}_{g}^\mathbb{S} := \sigma(W_{\varepsilon}, W_{\phi_g})$ in the limit experiments $\mathcal{E}(f)$ and $\mathcal{E}_\mathbb{S}(f)$, respectively. This then leads to our asymptotically invariant tests based on $\mathcal{M}_{g}$ and $\mathcal{M}_{g}^\mathbb{S}$.
The following proposition establishes the likelihood ratio restricted to the information $\mathcal{M}_{g}$ and $\mathcal{M}_{g}^\mathbb{S}$. The Neyman-Pearson lemma implies that tests based on these likelihood ratios are point optimal amongst the class of invariant tests in the limit experiments.
Observe that $W_{\perp}$ is a standard Brownian motion under $\mathbb{P}_{0,0,0}$ and independent of $W_{\varepsilon}$. When $g=f$, we have $J_{fg}=J_f=J_g$ and $\sigma_{\varepsilon\phi_g}=1$, so that $\lambda=1$ and $B_{\phi_g}=B_{\phi_f}$. As a result, we have $\Delta_g=\Delta^*_f$ and $\mathcal{I}_g=\mathcal{I}^*_f$, and $\Delta_g^\mathbb{S}=\Delta_f$ and $\mathcal{I}_g^\mathbb{S}=\mathcal{I}_f$.
The central idea to construct a hybrid rank test is to use a (quasi)-likelihood ratio test based on $L_{\mathcal{M}_{g}}(h,\lambda) := h\Delta_g -\frac{1}{2}h^2\mathcal{I}_g$ from ((ref)) for the case of $f\in\mathfrak{F}$, and $L_{\mathcal{M}_{g}^\mathbb{S}}(h,\lambda) := h\Delta_g^\mathbb{S} -\frac{1}{2}h^2\mathcal{I}_g^\mathbb{S}$ from ((ref)) for the case of $f\in\mathfrak{F}_\mathbb{S}$. In both cases, we then replace $W_{\varepsilon}$ and $B_{\phi_g}$ by their finite-sample counterparts in Lemma (ref) (or those in Lemma (ref) in case $\Gamma(L)$ would be known, for example, when $p=0$). The remaining unknown finite-sample parameters $\sigma_f^2$ and $\lambda$ are replaced by estimates that need to satisfy the following condition.
Such estimators are easily constructed, although $\hat{J}_{fg}$ is somewhat more involved. Estimating the real-valued cross-information $J_{fg}$ requires nonparametric techniques, but is considerably simpler than a full nonparametric estimation of $\phi_f$. Estimating $J_{fg}$ can be done along similar lines as estimating the Fisher information $J_f$, see, e.g., Bickel1982, BKRW, Schick1986, and Klaassen1987. A direct rank-based estimator of $J_{fg}$ has been proposed in CassartHallinPain2010. It is also worth noting that the consistency automatically also holds under local alternatives due to Le Cam's third lemma.
Based on a chosen reference density $g$ satisfying Assumption (ref) and estimators $\hat\sigma_f$, $\hat\sigma_{\varepsilon\phi_g}$ and $\hat{J}_{fg}$ satisfying Assumption (ref), we introduce the following partial sum processes:
where $\widehat{B}^{(T)}_{\phi_g}(s)$ and $\widehat{W}^{(T)}_{\phi_g}(s)$ are defined by Lemma (ref). Note that in the case of known $\Gamma(L)$, e.g., the i.i.d.\ case, one can simply use $\widehat{\Gamma}(L) = \Gamma(L)$ (where $\widehat{B}^{(T)}_{\phi_g}(s)=B^{(T)}_{\phi_g}(s)$ and $\widehat{W}^{(T)}_{\phi_g}(s)=W^{(T)}_{\phi_g}(s)$). Now, given a fixed alternative $\bar{h}<0$, we define
with
where $\widehat{\Delta}_\varepsilon = \int_0^1\widehat{W}^{(T)}_\varepsilon(s-)\mathrm{d}\widehat{W}^{(T)}_\varepsilon(s)$, $\widehat{\Delta}_\perp = \sqrt{J_g/\hat\sigma_{\varepsilon\phi_g}^2-1}\int_0^1\widehat{W}^{(T)}_\varepsilon(s-)\mathrm{d}\widehat{B}^{(T)}_{\perp}(s)$ and $\hat\lambda=(\hat{J}_{fg}\hat\sigma_{\varepsilon\phi_g}-\hat\sigma_{\varepsilon\phi_g}^2)/(J_g-\hat\sigma_{\varepsilon\phi_g}^2)$. Define also for the case of $f\in\mathfrak{F}_\mathbb{S}$,
with
where $\widehat{\Delta}_\perp^\mathbb{S} = \sqrt{J_g/\hat\sigma_{\varepsilon\phi_g}^2-1}\int_0^1\widehat{W}^{(T)}_\varepsilon(s-)\mathrm{d}\widehat{W}^{(T)}_\perp(s)$.
By Slutsky's theorem, we have the convergences $\big(\widehat{W}^{(T)}_\varepsilon, \widehat{B}^{(T)}_{\perp} \big)^\prime \Rightarrow \big( W_{\varepsilon}, B_{\perp} \big)^\prime$ and $\big(\widehat{W}^{(T)}_\varepsilon, \widehat{W}^{(T)}_\perp \big)^\prime \Rightarrow \big( W_{\varepsilon}, W_{\perp} \big)^\prime$, and the convergences of stochastic integrals $\int_0^1\widehat{W}^{(T)}_\varepsilon(s-)\mathrm{d}\widehat{B}^{(T)}_{\perp}(s) \Rightarrow \int_0^1W_{\varepsilon}(s)\rdB_{\perp}(s)$ and $\int_0^1\widehat{W}^{(T)}_\varepsilon(s-)\mathrm{d}\widehat{W}^{(T)}_\perp(s) \Rightarrow \int_0^1W_{\varepsilon}(s)\rdW_{\perp}(s)$ (see the proof of Lemma (ref)). We thus also obtain $\widehat{L}_{\mathcal{M}_{g}}(\bar{h},\hat\lambda)\wtoL_{\mathcal{M}_{g}}(\bar{h},\lambda)$ and $\widehat{L}_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\hat\lambda)\wtoL_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\lambda)$ under $P^{(T)}_{0,0,0;\mu,\Gamma,f}$. Define the critical values $c_{\mathcal{M}_{g}}(\bar{h},\sigma_{\varepsilon\phi_g},\lambda,J_g;\alpha)$ and $c_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\sigma_{\varepsilon\phi_g},\lambda,J_g;\alpha)$ by the $(1-\alpha)$-quantiles of $L_{\mathcal{M}_{g}}(\bar{h},\lambda)$ and $L_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\lambda)$, respectively. This leads to the (feasible) tests
and
These tests are not only based on the ranks of $\Delta Y_t$ but also their average, therefore, we name them Hybrid Rank Tests (HRTs).
We can now state our main theoretical result.
Theorem (ref) shows the HRTs are valid irrespective of the choice of the reference density and point-optimal for a correctly specified reference density. Moreover, in the corollary below, we state that the HRTs enjoy a Chernoff-Savage type result.
Corollary (ref) is a particularly useful result for applied work. The HRT dominates its classical canonical Gaussian counterpart, i.e., the ERS test in the present model, for any reference density $g$. Traditionally, this claim can only be made for Gaussian reference densities, but the framework here even allows for a stronger result. Our formulation of the testing problem using invariance arguments is convenient in this respect: the larger the invariant $\sigma$-field that is used, the more powerful the test.
The situation can be compared to Quasi Maximum Likelihood methods. However, again, in classical situations these methods are restricted to Gaussian reference densities. In the present setup, any reference density $g$ (subject to the regularity conditions imposed) can be used. The resulting test will always be valid, but more powerful in case the reference density chosen is closer to the true underlying density $f$.
A somewhat inconvenient aspect of the hybrid rank tests is that we need to estimate “the real-valued parameter” $J_{fg}$. As mentioned before, this is (much) less complicated than estimating the score function $\phi_f$ (as needed in Jansson2008), but might still be considered cumbersome, despite all references mentioned below Assumption (ref). Moreover, the critical value $c_{\mathcal{M}_{g}}(\bar{h},\hat\sigma_{\varepsilon\phi_g},\hat\lambda,J_g;\alpha)$ depends on the estimates $\hat\sigma_{\varepsilon\phi_g}$ and $\hat\lambda$ (henceforth $\hat{J}_{fg}$). Of course, this introduces no difficulty to implementing the test for a single dataset (though one would need to simulate a critical value), however, when it comes to a Monte Carlo study to access the performances of the HRTs, the computational effort will be significant. Therefore, we introduce a simplified version of the hybrid rank test. This simplified test is obtain by invoking $\lambda=1$, which holds in case $g=f$.
To be precise, define
where
and
where
Also define $L_g(\bar{h}):=L_{\mathcal{M}_{g}}(\bar{h},1)$ and $L_g^{\mathbb{S}}(\bar{h}):=L_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},1)$, then we have $\widehat{L}_g(\bar{h})\Rightarrow L_g(\bar{h})$ and $\widehat{L}_g^\mathbb{S}(\bar{h})\Rightarrow L_g^\mathbb{S}(\bar{h})$ under $\mathrm{P}^{(T)}_{0,0,0;\mu,\Gamma,f}$. Denoting the $(1-\alpha)$-quantiles of $L_g(\bar{h})$ and $L_g^\mathbb{S}(\bar{h})$ by $c_{g}(\bar{h},\sigma_{\varepsilon\phi_g},J_g;\alpha)$ and $c_{g}^\mathbb{S}(\bar{h},\sigma_{\varepsilon\phi_g},J_g;\alpha)$, respectively. These lead to the feasible tests
and
Since $\phi_g(\bar{h},\alpha)$ and $\phi_g^\mathbb{S}(\bar{h},\alpha)$ are approximate versions of the Hybrid Rank Tests $\phi_{\mathcal{M}_{g}}(\bar{h},\alpha)$ and $\phi_{\mathcal{M}_{g}^\mathbb{S}}(\bar{h},\alpha)$, we refer to them as Approximate Hybrid Rank Tests (AHRTs).
The proof of Theorem (ref) follows along the same lines as that of Theorem (ref) but using the weak convergences $\widehat{L}_g(\bar{h}) \Rightarrow L_g(\bar{h})$ and $\widehat{L}_g^\mathbb{S}(\bar{h}) \Rightarrow L_g^\mathbb{S}(\bar{h})$. The simulation results in Section (ref) show that these asymptotic properties carry over to finite samples.
From a computational point of view, the AHRTs have the advantage that nonparametric estimation of $J_{fg}$ is no longer needed. This significantly reduces the computational effort in the Monte Carlo study. Indeed, even though the critical value $c_{g}(\bar{h},\sigma_{\varepsilon\phi_g},J_g;\alpha)$ and $c_{g}^\mathbb{S}(\bar{h},\sigma_{\varepsilon\phi_g},J_g;\alpha)$ are still data dependent, it is, for given $\alpha$, $\bar{h}$, and reference density $g$, a function of only one argument --- the parameter $\sigma_{\varepsilon\phi_g}$. Observe, by Cauchy-Schwarz, that $\sigma_{\varepsilon\phi_g}$ is bounded by $\sqrt{J_g}$. For the chosen three reference densities, the critical value functions are listed in Table (ref). These are obtained by fitting a fourth-order polynomial to the exact critical values. In the Monte Carlo study (Section (ref)), we use these approximating critical value functions for computational speed.
This section reports the results of a Monte Carlo study to corroborate our asymptotic results, and to analyze the small-sample performance of the Approximate Hybrid Rank Tests. As mentioned earlier, we use the Approximate Hybrid Rank Tests in this simulation to avoid having to simulate the critical value for each individual replication. For the fixed alternative, we choose $\bar{h}=-7\sigma_{\varepsilon\phi_g}$ for two reasons. First, when $g=f$, we have $\sigma_{\varepsilon\phi_g}=1$ and hence $\bar{h}=-7$, which is in line with ERS1996. Second, as $\hat\sigma_{\varepsilon\phi_g}$ appears in the denominator in the AHRT statistic ((ref)), the statistic becomes better behaved for values of $\hat\sigma_{\varepsilon\phi_g}$ close to zero. The corresponding critical value functions for various reference densities are provided in Table (ref). The estimators for $\sigma_f^2$ and $\sigma_{\varepsilon\phi_g}$ we use are
Moreover, to simplify the notation, we denote the Approximate Hybrid Rank Test with reference density $g$ by ${\rm AHRT}$-$g$ and, in particular, by ${\rm AHRT}$-$\phi$ for Gaussian reference density and by ${\rm AHRT}$-$\hat{f}$ for a (nonparametrically) estimated reference density. To be specific, for the nonparametric estimation of $f$, we employ a kernel-based method $\hat{f}(x) = (T\mathfrak{h})^{-1}\sum_{t=1}^{T}K\left((x-\hat{\varepsilon}_t)/\mathfrak{h}\right)$, where the kernel $K$ is chosen to be Gaussian, $\mathfrak{h}$ is the bandwidth chosen by the rule $\left(4/(3T)\right)^{\frac{1}{5}}\hat\sigma_f$, and $\hat{\varepsilon}_t$ are the residuals from the regression ((ref)) below. Throughout we use significance level $\alpha = 5\%$ and all results are based on 20,000 Monte-Carlo replications.
We compare the performances of the propsed AHRT tests with two alternatives. First we consider the Dickey-Fuller test (denoted by DF-$\rho$) from DickeyFuller79. This test is based on the statistic $T(\hat{\rho}-1)$ where $\hat{\rho}$ is the least-squares estimator in the regression
The critical values for this test are -13.52 for $T=100$ and -14.05 for $T=2,500$. The second competitor is the ERS test with $\bar{h}=-7$. This test is based on the statistic $[S(\bar{\alpha})-\bar{\alpha}S(1)]/\hat{\omega}^2$ with $\bar{\alpha}=1+T^{-1}\bar{h}$ and $S(a) = (Y_a-Z_a\hat\beta)'(Y_a-Z_a\hat\beta)$, with $Y_a$ and $Z_a$ defined as
where $\hat\beta$ is estimated by regressing $Y_{\bar{\alpha}}$ on $Z_{\bar{\alpha}}$. The long-run variance estimator (for ERS test), $\hat{\omega}^2$, is chosen to be $\hat{\omega}^2_{AR(p)} = \hat\sigma_e^2 \big/ \left(1-\sum_{i=1}^p \widehat{\Gamma}_i \right)^2$ with $\hat\sigma_e^2 = \sum_{t=1}^T\hat{\varepsilon}_t^2/T$, where the residuals $\hat{\varepsilon}_t$ and coefficient estimates $\widehat{\Gamma}_i$ are from the regression in ((ref)). The critical values for this test are 3.11 for $T=100$ and 3.26 for $T=2,500$. We do not consider the Dickey-Fuller $t$-test as it is dominated by the DF-$\rho$ test in the current model. Similarly, the DF-GLS test proposed in ERS1996 is also omitted as it behaves asymptotically the same as the ERS test, but can be oversized in smaller samples.
In this section, we start the Monte Carlo study with i.i.d.\ innovations, i.e., the data is generated by the model in ((ref))-((ref)) with $\Gamma(L) = 1$. Therefore, for the newly-proposed AHRT test and its competitors introduced above, we have $\widehat{\Gamma}(L) = 1$ (or, equivalently, $\widehat{\Gamma}_1 = \cdots = \widehat{\Gamma}_p = 0$).
We first use the large-sample performances to illustrate the asymptotic properties. In particular, the chosen sample size $T$ is $2,500$.
Figure (ref) shows the power curves for 9 combinations of 3 innovation densities $f$ and 3 reference densities $g$ (for AHRT-$g$): $f$ and $g$ are chosen to be Laplace, Student $t_3$, or Gaussian.\footnote{The semiparametric power envelopes are based on $40,000$ Monte Carlo replications where the W-processes are approximated by a simple Euler approximation using $2,500$ grid points.} In line with our theoretical results, we find that the AHRT-$g$ test outperforms the two competitors in most cases. More specifically, when $g=f$ (the graphs on the diagonal), the AHRT-$f$ has power very close to the semiparametric power envelope and it is tangent to it at the point $-h=7$. This verifies the point-optimal result of the AHRT-$f$ test in Theorem (ref). The AHRT-$\hat{f}$ test has a similar behavior as the AHRT-$f$ but with a slightly lower power due to the efficiency loss in estimating the density $f$. This small amount of power loss is also different from case to case, e.g., for $f = t_3$, this power loss is almost indistinguishable; and the amount decreases to zero as $T$ goes to infinity. Moreover, when the reference density $g$ is Gaussian (the three right-most graphs), the AHRT-$\phi$ outperforms its competitors for non-Gaussian $f$; while for Gaussian $f$, the AHRT-$\phi$ test and the ERS test have indistinguishable power (and they outperform the Dickey-Fuller-$\rho$ test). This corroborates the Chernoff-Savage property of the AHRT-$\phi$ test mentioned in Remark (ref).
In order to investigate the Chernoff-Savage result for the AHRT-$\phi$ test even further, we consider in Figure (ref) the AHRT-$\phi$ test for i.i.d errors generated with nine different true innovation densities $f$. These include innovation densities $f$ that are extremely heavy-tailed, skewed, or both. The first row of graphs shows three extremely heavy-tailed distributions: Student $t_2$, Student $t_1$, and a stable distribution with parameter values $0.5$ for stability, $0$ for skewness, $1$ for scale, and $0$ location. As these densities do not all satisfy our maintained assumptions, these graphs exclude power envelopes and the AHRT-$\hat{f}$ power functions. The top three graphs in Figure (ref) show that the AHRT-$\phi$ is much more powerful than its competitors and that its power increases with the heaviness of the tail. The second and third row show the effect of skewness in $f$. Specifically, the AHRT-$\phi$'s power is higher when $f$ is skewed-normal (with skewness 0.8145) than that when $f$ is normal (in Figure (ref)). This indicates that the AHRT-$\phi$ can acquire power from skewness. The same conclusion can be drawn from the comparison of the AHRT-$\phi$ power function for $t_4$ and that of a skewed $t_4$ with skewness $\approx 2.7$. To further remove the effects of the other moments, in the third row, we also employ the Pearson distributions with identical mean, variance and kurtosis, but different skewness --- ${\rm skewness}=1$ for Pearson-I, ${\rm skewness}=3$ for Pearson-II, and ${\rm skewness}=6$ for Pearson-III. Comparing the corresponding power functions, it validates again that the larger the skewness of the true distribution $f$ is, the more powerful the ${\rm AHRT}^\phi$ becomes.
A final remark on the size of the AHRT tests. In all cases where the true density $f$ satisfies our maintained assumption, i.e., $f\in\mathfrak{F}$ (that is all cases in Figure (ref) and the skewnormal, $t_4$, Pearson-I, Pearson-II, and Pearson-III in Figure (ref)), the simulated sizes are between $4.9\%$ and $5.1\%$. This verifies the validity of the AHRTs claimed in Theorem (ref). In the other cases, i.e., $f\not\in\mathfrak{F}$, the AHRT is somewhat conservative. More precisely, the simulated sizes of the AHRT-$\phi$ are $4.8\%$, $4.1\%$, $3.7\%$, and $4.7\%$ for the Student $t_2$, $t_1$, stable, and skew-$t_4$ distribution, respectively. This result seems consistent over all simulations.
We also report the performance of the AHRTs and the two competitors introduced above for smaller samples. Figures (ref) and (ref) are the small-sample versions, with $T = 100$, of Figures (ref) and (ref), respectively. We observe that, even with a slight downward shift of the power functions for all three tests considered, the findings of the large-sample case remain valid in this small-sample case. For larger values of $h$, the DF-$\rho$ test sometimes dominates the other two tests. This is due to the fact that the DF-$\rho$ in this Monte-Carlo setting appears to have a superior convergence speed (towards its asymptotic power as sample size $T$ increases) to those of the AHRTs and the ERS test. This phenomenon appears only in (ultra)-small-sample cases, and disappears when $T$ gets larger (e.g., $T = 200$). In cases with enough samples (say, $T \geq 100$) and when $f$ is significantly away from the Gaussian density, irrespective of the choice of $g$, the AHRTs performs favorably.
Concerning the small-sample size, we find it to range from about $4.0\%$ to $4.5\%$ for the cases where $f\in\mathfrak{F}$. Again, when $f$ does not satisfy our maintained assumptions ($f\not\in\mathfrak{F}$) the AHRT turns out to be conservative. More precisely, we find a size of $3.7\%$, $3.1\%$, $2.4\%$, and $4.1\%$ for the $t_2$, $t_1$, stable, and skew-$t_4$ distribution, respectively. This makes the improved power even more remarkable.
It may also be useful to illustrate the convergence of the power function of the AHRT-${f}$ to the semiparametric power envelope as sample size $T$ increases. This is the purpose of Figure (ref). For three cases: Gaussian, Laplace, and Student $t_3$, we find that the convergence indeed occurs already at relatively small samples, which is not always the case for alternative unit root tests.
In the case with the additional symmetric density restriction, $f \in \mathfrak{F}_\mathbb{S}$, we repeat the above Monte Carlo study above but for the signed-rank-based AHRT $\phi_g^\mathbb{S}(\bar{h},\alpha)$ introduced in Section (ref). Not surprisingly, we find very similar results as above, except for a slightly higher power for all the AHRT tests (essentially due to the symmetry restriction imposed, in which case we have the adaptivity result). Since it's a bit unfair to compare with the competitor which actually does not benefit from this constraint, and also to conserve space, we put the associated simulation results --- two figures that can be treated as the signed-rank-based AHRT's counterparts of Figure (ref) and Figure (ref) --- in the Appendix C in (ref). In short, the large-sample result shows that the power function of the signed AHRT-$f$ is tangent to the (parametric) power envelope, which corroborates the point-optimal property in Theorem (ref) and in turn the adaptive result in Section (ref); the Chernoff-Savage result still holds for the signed-rank version of the AHRT-$\phi$ test. The second figure therein shows that it works well in the small sample case. Furthermore, one can also check the convergence of the signed-rank-based AHRT-${f}$ power function to the parametric power envelope as sample size $T$ increases in Figure (ref).
In this section we provide simulation results for the cases where the innovations follow an ARMA model, i.e.,
where the assumption on $\varepsilon_t$ stay unchanged. This corresponds to a lag-polynomial $\Gamma(L)$ with an infinite order, which we employ here in order to show by simulations the potential that the finite order $p$ restriction could be relaxed. As for the estimation of $\widehat{\Gamma}(L)$ in ((ref)) used by the AHRT test and its competitors, we fix the AR regression order $p = 8$. Since there exists efficiency loss in the estimation of error autocorrelation parameters $\Gamma_1,\dots,\Gamma_p$ (for finite-sample cases), here we increase the sample sizes for both the large-sample case (to $T = 5,000$) and the small-sample case (to $T = 250$). Under the same organization of true and reference densities as for the i.i.d.\ case, we provide the simulation results for large-sample case and small-sample case in Figure (ref) and Figure (ref), respectively. Note that, in Figure (ref), we remove the powers of the ADF-$\rho$ test since it is severely oversized.
From these results, we draw similar conclusions to those in the i.i.d.\ case, except for a slight power loss for all these chosen unit root tests. These results validate that, on one hand, the serial correlation in the errors can be well handled, as for the cases of other unit root tests, with an additional auto-regression for the increments of the observed process. On the other hand, they also indicate that even though in the limit the unknown autocorrelation structure has no effects on the inference for $\rho$, its estimation consumes some efficiency in finite sample cases. Yet, Figure (ref) shows that the AHRT suffers less from this kind of efficiency loss or, in other words, the AHRT has a faster convergence speed of its power to the associated envelope as $T\to\infty$.
This paper has provided a structural representation of the limit experiment of the standard unit root model in a univariate but semiparametric setting. Using invariance arguments, we have derived the semiparametric power envelope. These invariance structures also lead, using the Neyman-Pearson lemma, to point-optimal semiparametric tests. The analysis naturally leads to the use of rank-based statistics.
Our tests are asymptotically valid, invariant, and (with a correctly chosen reference density) point-optimal. Moreover, we establish a Chernoff-Savage type property of our test: irrespective of the reference density chosen, our test outperforms its classical competitor which in this case is the ERS test. Finally, we introduced a simplified version of our test and show, in a Monte-Carlo study, that our theoretical results carry over to finite samples.
As potential future work we mention the use of similar ideas to construct hybrid rank-based tests in more general time-series models with, for instance, a deterministic time trend term, or stochastic volatility. Also, the structural representation of the limit experiment and its invariance properties could be applied to other non-stationary time-series models, for instance, cointegration and predictive regression models.