EconBase
← Back to paper

Semiparametrically Optimal Cointegration Test

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

97,436 characters · 14 sections · 83 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametrically Optimal Cointegration Test

\abstract{This paper aims to address the issue of semiparametric efficiency for cointegration rank testing in finite-order vector autoregressive models, where the innovation distribution is considered an infinite-dimensional nuisance parameter. Our asymptotic analysis relies on Le Cam's theory of limit experiment, which in this context takes the form of Locally Asymptotically Brownian Functional (LABF). By leveraging the structural version of LABF, an Ornstein-Uhlenbeck experiment, we develop the asymptotic power envelopes of asymptotically invariant tests for both cases with and without a time trend. We propose feasible tests based on a nonparametrically estimated density and demonstrate that their power can achieve the semiparametric power envelopes, making them semiparametrically optimal. We validate the theoretical results through large-sample simulations and illustrate satisfactory size control and excellent power performance of our tests under small samples. In both cases with and without time trend, we show that a remarkable amount of additional power can be obtained from non-Gaussian distributions.}

JEL classification: C12, C14

Keywords: cointegration, semiparametric efficiency, limit experiment, LABF.

Introduction

Cointegration has been a central topic in time series econometrics ever since its concept was introduced by granger1981some and engle1987co. Cointegration refers to the phenomenon where multiple nonstationary time series have stationary linear combinations, known as cointegrating relationships. Determining the number of these relationships, or the cointegration rank, is of utmost importance. This inferential problem is often addressed in a finite-order vector autoregressive (VAR) process. Early attempts include, among others, residual-based tests by engle1987co and phillips1990asymptotic, likelihood ratio tests by johansen1988statistical,johansen1991estimation and johansen1990maximum, and principal component tests by stock1988testing. Since then, extensive literature has been devoted to constructing cointegration rank tests with good power properties. Our paper shares the same goal, particularly investigating to what extent we can exploit testing power from the innovation distributions that deviate from Gaussianity.

The main objective of this paper is to develop semiparametrically optimal cointegration rank tests for a finite-order VAR model written in the ECM form. We assume that the innovations are independently and identically distributed, and we treat the innovation distribution as an infinite-dimensional nuisance parameter. Our analysis relying on Le Cam's asymptotic theory, where the concept of limit experiment --- as the limit of the sequence of experiments of interest (here, the cointegration experiments) --- plays the central role; see, e.g., le2012asymptotic and vdVaart00. The most recent work using this approach is hallin2016semiparametric (hereafter referred to as HvdAW), which provides a complete factorization for the cointegration parameter $\boldsymbol{\Pi}$ in model given by ((ref))--((ref)) below and characterizes all possible limit experiments. HvdAW shows that the associated limit experiments are of different types, including Locally Asymptotically Normality (LAN), Locally Asymptotically Mixed Normality (LAMN), and Locally Asymptotically Brownian Functional (LABF), depending on the parameter directions; see Jeganathan1995 for their definitions. HvdAW focuses on the LAN experiment brought by the time trend term, while some earlier works, such as phillips1991optimal and hodgson1998adaptiveET,hodgson1998adaptiveJOE, concentrate on on the LAMN direction. However, the LABF direction remains unexplored, and our aim is to fill this gap in the literature.

Our contribution to the literature is threefold. First, following zhou2019semiparametrically's structural representation technique for the LABF-type experiments, we develop the semiparametric power envelopes of asymptotically invariant tests in both cases with and without a time trend. Specifically, in this structural versions as an Ornstein-Uhlenbeck (OU) experiment, the nuisance density perturbation parameter ($\boldsymbol{\eta}$) appears as a constant drift, which can be eliminated by taking the associate `bridge' process. More importantly, we show that the $\sigma$-field consisting of OU processes (for elements unaffected by $\boldsymbol{\eta}$) and OU bridges (for elements affected by $\boldsymbol{\eta}$) is maximally invariant, which provides optimality in the limit. According to the Neyman-Pearson lemma, the corresponding likelihood ratio test is optimal among all invariant tests since every invariant statistic is a function of the maximal invariant. The Asymptotic Representation Theorem then translates the limiting optimality to the sequence of cointegration experiments (see, e.g., vdVaart00).

Our second contribution is to propose feasible tests whose powers can attain the semiparametric power envelopes asymptotically (as shown in Theorem (ref) and Corollary (ref)). This confirms that the derived semiparametric power envelopes are indeed `envelopes', rather than just upper bounds, and that our tests are semiparametrically optimal. To construct these tests, we follow the traditional semiparametric inference literature, where the unknown nonparametric part is replaced by its kernel estimates (as see in works such as Bickel1982, schick1986asymptotically, and Klaassen1987). To improve finite-sample performance, especially when the sample size is small, we adopt schick1987note's technique and use all samples for our semiparametric statistic without sample splitting. Our Monte Carlo study confirms the validity and optimality of our tests using large-sample simulations and exhibits satisfactory size control and excellent power performance in small-sample scenarios.

We regard the treatment of the linear time trend specification as our third contribution. Following the unit root and cointegration literature, we augment the stochastic term with a deterministic time trend term as in ((ref)). This additive structure, employed by many later works (e.g., elliott1996efficient and Jansson2008 for unit root testing, and saikkonen2000trend, lutkepohl2000testing, and boswijk2015improved for cointegration), has several advantages, including the ability to emphasize that the trend is at most linear. This trend specification is one of the main differences between our paper and HvdAW.\footnote{Another distinction of our paper from HvdAW is in terms of distribution assumption: HvdAW assumes the innovation distribution to be elliptical, while we allow it to be essentially unrestricted (up to the regularity conditions outlined in Assumption (ref) for our asymptotic analysis). The reason for this difference is that HvdAW's proposed distribution-free rank-based tests utilize the concept of Mahalanobis distance, which in this context is for multidimensional generalization of rank-based inference under the more restricted elliptical distribution assumption.} HvdAW employs the traditional trend-in-VAR representation as in ((ref)), making the time trend the main power source for their test (essentially due to the super-consistency rate $T^{-3/2}$ introduced by the time trend). In contrast, our paper regards the time trend (local) parameter $\boldsymbol{\delta}$ as a nuisance parameter to be eliminated (see Remark (ref) for more detailed discussions). We rely on the “profile likelihood” approach to eliminate this parameter, taking advantage of the simple quadratic structure of the likelihood with respect to $\boldsymbol{\tau}$. Additionally, we develop a limiting statistic $\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}^*)$ for the semiparametric case, analogue to the $\mathit{\Lambda}_{p,C}^{GLS}(\bar{C};\bar{C}^*)$ statistic proposed by boswijk2015improved for the Gaussian case. The statistic $\mathit{\Lambda}_{p,C}^{GLS}(\bar{C};\bar{C}^*)$ embeds and thus can help compare many existing Gaussian cointegration tests. Likewise, our semiparametric version $\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}^*)$ enables us to develop semiparametric versions of existing Gaussian tests that handle the time trend specification.

Finally, we discuss extensions that incorporate serially correlated errors and allow for more general reduced rank hypotheses, significantly expanding the applicability of our developed tests. Our discussion is based on existing limit experiment results, mostly provided by HvdAW, particularly their full limit experiment characterization in Proposition A.2 of their online supplementary appendix. Building on these results, we show that the inference for the parameter of interest, $\boldsymbol{\Pi}$, is adaptive to both the parameters governing the serial correlation in the errors and those governing the existing cointegrating relationships. Therefore, we `just' need to consistently estimate these parameters and replace them with their estimates.

The remainder of the paper is structured as follows. Section (ref) presents the model setup and assumptions. Section (ref) develops the limit experiment, sequentially eliminates the density perturbation and time trend parameters, and derives the corresponding semiparametric power envelopes for the cases with and without a time trend. Then, based on a nonparametrically estimated density, Section (ref) proposes feasible semiparametrically optimal tests, whose finite-sample performances are accessed by a Monte Carlo study in Section (ref). Section (ref) provides discussions on necessarily extensions to expand the empirical applicability of our tests. The proofs of our theoretical results can be found in the supplementary Appendices.

Model

We consider observations $\mathbf{y}_1,\dots,\mathbf{y}_{T} \in \mathbb{R}^{p}$ generated by the vector auto-regression (VAR) of order one in error correction form

align[align omitted — 219 chars of source]

where $\Delta$ denotes the first-order difference operator (i.e., $\Delta\mathbf{x}_t = \mathbf{x}_t - \mathbf{x}_{t-1}$), $\boldsymbol{\mu}$ and $\boldsymbol{\tau}$ (both in $\mathbb{R}^p$) are unknown parameters that govern the constant and linear time trend terms, respectively, $\boldsymbol{\Pi}\in\mathbb{R}^{p\timesp}$ is an unknown parameter of interest, and $\{\boldsymbol{\varepsilon}_t\}$ is a $p$-dimensional i.i.d.\ sequence of innovations with density $f$.

Throughout this paper, we assume $\mathbf{x}_0 = \boldsymbol{0}$ as the initial value condition. However, this assumption is less innocent than it may seem, as noted by elliott1996efficient (ERS), who observed that even asymptotically, the initial observations can carry information. Further investigations into this issue can be found in muller2003tests and elliott2006minimizing. That being said, since our paper employs the same local-to-unity asymptotics as ERS, we can relax this condition to the same extent.

In this paper, we use the notation $\mathfrak{F}_p$ to represent the family of densities that satisfy the following assumptions, which we impose on $f$.

assumption[] \begin{enumerate} • $f$ is absolutely continuous with a.e.\ gradient $\,\nabla f(\boldsymbol{\varepsilon}_1) = \big(\partialf(\boldsymbol{\varepsilon}_1)/\partial\varepsilon_{1,1},\dots,\partialf(\boldsymbol{\varepsilon}_1)/\partial\varepsilon_{1,p}\big)^\prime$. • $\mathrm{E}_{f}[\boldsymbol{\varepsilon}_t] = \boldsymbol{0}$ and the covariance $\boldsymbol{\Sigma} := \operatorname{Var}_f[\boldsymbol{\varepsilon}_1]$ is positive definite and finite. • The Fisher information $\boldsymbol{J}_f := \mathrm{E}_f[\boldsymbol{\ell}_f(\boldsymbol{\varepsilon}_1)\boldsymbol{\ell}_f(\boldsymbol{\varepsilon}_1)^{\prime}]$, where $\boldsymbol{\ell}_f(\boldsymbol{\varepsilon}_1) := -\nablaf(\boldsymbol{\varepsilon}_1)/f(\boldsymbol{\varepsilon}_1)$ denotes the location score of $f$, is finite. • $f$ is positive. \end{enumerate}

The absolute continuity assumption (a) on $f$ is a mild smoothness condition commonly imposed in the semiparametric literature. This condition is imposed for two reasons: Firstly, it allows us to proceed with the limit experiment approach (le2012asymptotic) since it implies the differentiability in quadratic mean (DQM) result, which is the exactly right condition needed for our log-likelihood ratio expansion (vdVaart00).\footnote{A deep discussion and appreciation of DQM can be found in pollard1997another.} Secondly, the absolute continuity assumption enables us to perform nonparametric estimation of the score function $\boldsymbol{\ell}_f$, which will be used to construct a feasible semiparametrically optimal test. The finite-variance condition (b) guarantees that the Fisher information matrix $\boldsymbol{J}_f$ is nonsingular (mayer1990cramer). Together with the finite-Fisher-information condition (c), they ensure the the weak convergence of the partial-sum processes of $\boldsymbol{\varepsilon}_t$ and $\boldsymbol{\ell}_f(\boldsymbol{\varepsilon}_t)$ to Brownian motions. The positive density condition (d) is merely for notational convenience, e.g., when defining the score function $\boldsymbol{\ell}_f$.

As the primary focus of this paper is to address the issue of semiparametric efficiency in the context of cointegration, we begin by considering a simple case of testing the null hypothesis of no cointegration against the alternative of the existence of at least one cointegrating relationship, stated as:

align[align omitted — 116 chars of source]

In Section (ref), we will briefly discuss the extension of our results to more general reduced rank hypotheses on $\boldsymbol{\Pi}$, drawing upon existing literature.

remarkThe choice of time trend specification can have significant implications for the asymptotic results of cointegration tests. In this paper, we adopt the specification from the branch of unit root and cointegration literature which adds a level constant and a linear time trend to the stochastic component, as given in ((ref)). See, for example, elliott1996efficient for unit root testing, and lutkepohl2000testing and boswijk2015improved for cointegration testing. It differs from the specification used in hallin2016semiparametric, which follows the traditional “trend-in-VAR” specification given by \begin{align} \Delta\mathbf{y}_t = \mathbf{v} + \mathbf{v}_1t + \boldsymbol{\Pi}\mathbf{y}_{t-1} + \boldsymbol{\varepsilon}_t. \end{align} In particular, the authors focus on the case where $\mathbf{v}_1 = \boldsymbol{0}$. Although model ((ref)) is equivalent to model ((ref))--((ref)) under the parameter constraints $\mathbf{v} = -\boldsymbol{\Pi}\boldsymbol{\mu} + (\mathbf{I}_{p} + \boldsymbol{\Pi})\boldsymbol{\tau}$ and $\mathbf{v}1 = -\boldsymbol{\Pi}\boldsymbol{\tau}$, it may generate quadratic time trends without these constraints. That being said, model ((ref))--((ref)) has the advantage of emphasizing that the time trend considered in $\mathbf{y}_t$ is at most linear (see lutkepohl2000testing for further discussion). These different time trend specifications lead to distinct asymptotic results. Under model ((ref)), the cointegration test of hallin2016semiparametric will have asymptotic power that depends on $\mathbf{v}$. Specifically, the test is more powerful when $\mathbf{v}$ is larger but has low power when the time trend is close to zero. In contrast, our test's asymptotic power does not depend on $\boldsymbol{\mu}$ or $\boldsymbol{\tau}$, as they are eliminated using the invariance principle.

Semiparametric Power Envelopes

Preliminaries

Our asymptotic analysis relies on the limit experiment approach. The limit experiment of a sequence of experiments (in this case, cointegration experiments) is defined by the convergence of likelihood ratios under specific local perturbations. To ensure that the likelihood-ratio convergence is neither explosive nor degenerate, we need to localize $\boldsymbol{\mu}$, $\boldsymbol{\tau}$, $\boldsymbol{\Pi}$ and $f$ with appropriate rates. Such local perturbations are known as the contiguous alternative (see vdVaart00). In what follows, we introduce these local reparameterizations separately.

We follow the unit root and cointegration literature in adopting the local-to-unity asymptotics for the key parameter of interest, $\boldsymbol{\Pi}$. The associated local reparameterization is given by

align[align omitted — 111 chars of source]

where $\mathbf{C}\in\mathbb{R}^{p\timesp}$ is referred to the local parameter. The “super-consistency” contiguity rate $T^{-1}$ here is common in models for nonstationary time series (see Phillips1987 and ChanWei88), including unit root testing, cointegration, and predictive regression with persistent predictors.\footnote{For related predictive regression literature, see, e.g., elliott1994inference, jansson2006optimal, and werker2022semiparametric.} This nonstandard rate $T^{-1}$ leads to LABF-type experiments, as we show in Proposition (ref) below.

We localize the time trend parameter as follows:

align[align omitted — 162 chars of source]

where $\boldsymbol{\tau}_0$ represents the true value. This local reparameterization is the natural multivariate counterpart of the unit root testing problem, as discussed in Jansson2008. Following that paper, without loss of generality, we assume that $\boldsymbol{\tau}_0$ is zero.

We do not localize the constant term parameter $\boldsymbol{\mu}$ in our approach since we show in the proof of Proposition (ref) that it vanishes asymptotically. This result is in line with the findings of Jansson2008, as the univariate counterpart, and show that information about $\boldsymbol{\mu}$ only comes from the first few observations rather than the data flow.

Finally, we adopt the approach by zhou2019semiparametrically and introduce explicit nonparametric local perturbations to the innovation density $f$ as follows:

align[align omitted — 219 chars of source]

Here, $\boldsymbol{b} := (b_1,b_2,\dots)^\prime$ is a vector of functions that govern the perturbations in different directions, and $\boldsymbol{\eta} := (\eta_1,\eta_2,\dots)^\prime$ is the local parameter that determines the severity of these perturbations. We choose $\boldsymbol{b}$ to be a countable orthonormal basis of the separable Hilbert space

align*[align* omitted — 215 chars of source]

$\boldsymbol{\varepsilon}\in\mathbb{R}^{p}$, where $ \L_2^f(\mathbb{R}^p,\mathcal{B})$ denotes the space of Borel-measurable functions $b:\mathbb{R}^p\to\mathbb{R}$ that are square-integrable. Accordingly, we have $\operatorname{Var}_f[b_k(\boldsymbol{\varepsilon})] = 1$. The separability of the Hilbert space ensures the existence of such a countable orthonormal basis. We further assume that $b_k \in C_{2,b}(\mathbb{R})$ for all $k$, meaning that $b_k$'s are bounded and twice continuously differentiable with bounded derivatives.

We restrict the local perturbation parameter $\boldsymbol{\eta}$ to have only finitely many non-zero elements, i.e., $\boldsymbol{\eta} \in c_{00}$ where $c_{00} := \{(z_k)_{k\in\mathbb{N}}\in\mathbb{R}^{\mathbb{N}}\,|\,\sum_{k=1}^\infty \mathbbm{1}\{z_k \neq 0\} < \infty\}$, which is a dense subspace of the parameter space $\ell_2 = \{(z_k)_{k\in\mathbb{N}}\,|\,\sum_{k=1}^\infty z_k^ 2<\infty\}$. This restriction does not sacrifice generality but rather helps avoid dealing with the convergence of infinite-dimensional Brownian motions and possibly associated measurability complexities. Notably, when $\boldsymbol{\eta} = \boldsymbol{0}$, we obtain $f_{\boldsymbol{\eta}}^{(T)}=f$. We demonstrate in the following proposition that for any $\boldsymbol{\eta} \neq \boldsymbol{0}$, $f_{\boldsymbol{\eta}}^{(T)}$ satisfies Assumption (ref). The proof is detailed in Appendix (ref).

propositionFor any fixed $f\in\mathfrak{F}_p$ and $\boldsymbol{\eta}\in c_{00}$, there exists $T^\prime \in \mathbb{N}$ such that $f_{\boldsymbol{\eta}}^{(T)}\in\mathfrak{F}_p$ for all $T\geqT^\prime$.

However, we do not aim to demonstrate that the particular nonparametric form $f_{\boldsymbol{\eta}}^{(T)}$ with $\boldsymbol{\eta}\in c_{00}$ is capable of generating all possible local perturbations on $f$ in $\mathfrak{F}_p$. Nonetheless, this concrete perturbation specification does not affect our analysis of semiparametric optimality, particularly the derivation of the semiparametric power envelopes. Essentially, $f_{\boldsymbol{\eta}}^{(T)}$ can be regarded as a complex sub-model under which a local power upper bound can be derived. Once a feasible semiparametric test is constructed that can attain this upper bound (as we will demonstrate in Section (ref)), it becomes the semiparametric power envelope.

The limiting experiment

We define $\mathrm{P}^{(T)}_{\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta};\boldsymbol{\mu},f}$ as the law of $\mathbf{y}_1,\dots,\mathbf{y}_T$ generated by the error-correction model ((ref))--((ref)) with local reparameterizations in ((ref))--((ref)). Additionally, we introduce the probability measure of the associated limit experiment, denoted as $\mathbb{P}_{\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta}}$, which will be formally introduced later. In the following proposition, we demonstrate that the log-likelihood ratio process $\mathcal{L}^{(T)}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta}):= \log\big({\rm d}\mathrm{P}^{(T)}_{\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta};\boldsymbol{\mu},f}/{\rm d}\mathrm{P}^{(T)}_{\boldsymbol{0},\boldsymbol{0},\boldsymbol{0};\boldsymbol{\mu},f}\big)$ follows the Locally Asymptotically Brownian Functional form introduced by Jeganathan1995.

propositionConsider $f\in\mathfrak{F}_p$. Let $\mathbf{C}\in\mathbb{R}^{p\times p}$, $\boldsymbol{\mu}\in\mathbb{R}^{p}$, $\boldsymbol{\delta}\in\mathbb{R}^{p}$, and $\boldsymbol{\eta}\in c_{00}$. \begin{itemize} • Under $\mathrm{P}^{(T)}_{\boldsymbol{0},\boldsymbol{0},\boldsymbol{0};\boldsymbol{\mu},f}$, as $T\to\infty$, the log-likelihood ratio is decomposed as \begin{align} \mathcal{L}^{(T)}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta}) = \mathit{\Delta}^{(T)}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta})-\frac{1}{2}\mathcal{Q}^{(T)}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta})+o_{P}(1), \end{align} where \begin{align*} \mathit{\Delta}^{(T)}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta}) :=& \frac{1}{T}\sum_{t=2}^{T}(\mathbf{C}\mathbf{y}_{t-1,1})^{\prime}\boldsymbol{\ell}_{f}(\Delta\mathbf{y}_t) + \frac{1}{\sqrt{T}}\sum_{t=2}^{T}\left(\mathbf{d}_{\mathbf{C},t}\boldsymbol{\delta}\right)^\prime\boldsymbol{\ell}_f(\Delta \mathbf{y}_t) + \frac{1}{\sqrt{T}}\sum_{t=2}^{T}\boldsymbol{\eta}^{\prime}\boldsymbol{b}(\Delta \mathbf{y}_t), \\ \mathcal{Q}^{(T)}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta}) :=& \frac{1}{T^2}\sum_{t=2}^{T}(\mathbf{C}\mathbf{y}_{t-1,1})^{\prime}\boldsymbol{J}_f\mathbf{C}\mathbf{y}_{t-1,1} + \frac{1}{T}\sum_{t=2}^{T}(\mathbf{d}_{\mathbf{C},t}\boldsymbol{\delta})^{\prime}\boldsymbol{J}_f\mathbf{d}_{\mathbf{C},t}\boldsymbol{\delta} \\ & + \frac{2}{T^{3/2}}\sum_{t=2}^{T}(\mathbf{C}\mathbf{y}_{t-1,1})^{\prime}\boldsymbol{J}_f\mathbf{d}_{\mathbf{C},t}\boldsymbol{\delta} + \frac{2}{T^{3/2}}\sum_{t=2}^{T}\boldsymbol{\eta}^{\prime} \boldsymbol{J}_{\boldsymbol{b}f}\mathbf{C}\mathbf{y}_{t-1,1} \\ & + \frac{2}{T}\sum_{t=2}^{T}\boldsymbol{\eta}^{\prime} \boldsymbol{J}_{\boldsymbol{b}f}\mathbf{d}_{\mathbf{C},t}\boldsymbol{\delta} + \boldsymbol{\eta}^{\prime}\boldsymbol{\eta}, \end{align*} with $\mathbf{d}_{\mathbf{C},t} := \mathbf{I}_{p}-\frac{t-1}{T}\mathbf{C}$, $\mathbf{y}_{t-1,1} := \mathbf{y}_{t-1} - \mathbf{y}_{1}$, and $\boldsymbol{J}_{\boldsymbol{b}f} := \mathrm{E}_{f}[\boldsymbol{b}(\boldsymbol{\varepsilon}_1)\boldsymbol{\ell}_f(\boldsymbol{\varepsilon}_1)^\prime]$. • Let $\boldsymbol{W}_{\boldsymbol{\varepsilon}},\boldsymbol{W}_{\boldsymbol{\ell}_f},\boldsymbol{W}_{\boldsymbol{b}}$ be Brownian motions defined on the probability space $(\Omega,\mathcal{F},\mathbb{P}_{\boldsymbol{0},\boldsymbol{0},\boldsymbol{0}})$ with covariance \begin{align} \operatorname{Var}\begin{pmatrix}\boldsymbol{W}_{\boldsymbol{\varepsilon}}(1)\\\boldsymbol{W}_{\boldsymbol{\ell}_f}(1)\\\boldsymbol{W}_{\boldsymbol{b}}(1)\end{pmatrix}= \begin{pmatrix} \boldsymbol{\Sigma} & \mathbf{I}_{p} & \boldsymbol{0}_{p,\infty} \\ \mathbf{I}_{p} & \boldsymbol{J}_{f} & \boldsymbol{J}_{f\boldsymbol{b}} \\ \boldsymbol{0}_{\infty,p} & \boldsymbol{J}_{\boldsymbol{b}f} & \mathbf{I}_\infty \end{pmatrix}, \end{align} where $\mathbf{I}$ and $\boldsymbol{0}$ are the identity and zero matrices of dimensions shown in their superscripts, respectively. Then, under $\mathrm{P}^{(T)}_{\boldsymbol{0},\boldsymbol{0},\boldsymbol{0};\boldsymbol{\mu},f}$ and as $T\to\infty$, we have \begin{align} \mathcal{L}^{(T)}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta})\Rightarrow\mathcal{L}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta}) = \mathit{\Delta}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta})-\frac{1}{2}\mathcal{Q}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta}), \end{align} where \begin{align*} \mathit{\Delta}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta}) :=& \int_0^1(\mathbf{C}\boldsymbol{W}_{\boldsymbol{\varepsilon}}(u))^{\prime}{\rm d}\boldsymbol{W}_{\boldsymbol{\ell}_f}(u) + \int_0^1(\mathbf{d}_{\mathbf{C}}(u)\boldsymbol{\delta})^{\prime}{\rm d}\boldsymbol{W}_{\boldsymbol{\ell}_f}(u)+\boldsymbol{\eta}^{\prime}\boldsymbol{W}_{\boldsymbol{b}}(1), \\ \mathcal{Q}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta}) :=& \int_0^1(\mathbf{C}\boldsymbol{W}_{\boldsymbol{\varepsilon}}(u))^{\prime}\boldsymbol{J}_f C\boldsymbol{W}_{\boldsymbol{\varepsilon}}(u){\rm d}u + \int_0^1(\mathbf{d}_{\mathbf{C}}(u)\boldsymbol{\delta})^{\prime}\boldsymbol{J}_f\mathbf{d}_{\mathbf{C}}(u)\boldsymbol{\delta}{\rm d}u \\ & + 2\int_0^1(\mathbf{C}\boldsymbol{W}_{\boldsymbol{\varepsilon}}(u))^{\prime}\boldsymbol{J}_f\mathbf{d}_{\mathbf{C}}(u)\boldsymbol{\delta}{\rm d}u + 2\boldsymbol{\eta}^{\prime}\boldsymbol{J}_{\boldsymbol{b}f}\mathbf{C}\widebar\boldsymbol{W}_{\boldsymbol{\varepsilon}} + 2\boldsymbol{\eta}^{\prime}\boldsymbol{J}_{\boldsymbol{b}f} \begingroup \def\mathaccent#\mathbf{d}##2{ \kern0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax111{\mathbf{d}} \endgroup _{\mathbf{C}}\boldsymbol{\delta} + \boldsymbol{\eta}^\prime\boldsymbol{\eta}, \end{align*} with $\mathbf{d}_{\mathbf{C}}(u) := \mathbf{I}_{p} - u\mathbf{C}$, $ \begingroup \def\mathaccent#\mathbf{d}##2{ \kern0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax111{\mathbf{d}} \endgroup _{\mathbf{C}} := \int_0^1\mathbf{d}_{\mathbf{C}}(u){\rm d}u = \mathbf{I}_{p} - \mathbf{C}/2$, and $\widebar\boldsymbol{W}_{\boldsymbol{\varepsilon}} := \int_0^1\boldsymbol{W}_{\boldsymbol{\varepsilon}}(u){\rm d}u$. • Under $\mathbb{P}_{\boldsymbol{0},\boldsymbol{0},\boldsymbol{0}}$, $\forall \, \mathbf{C}\in\mathbb{R}^{p \times p}$, $\boldsymbol{\delta}\in\mathbb{R}^{p}$, and $\boldsymbol{\eta}\in c_{00}$, $\mathbb{E}\left[\exp(\mathcal{L}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta}))\right] = 1$. \end{itemize}

The proof of part (a) could have essentially followed from a Taylor expansion, assuming twice continuous differentiability on $f$. However, instead, we employed the framework of hallin2015quadratic, which is built upon the DQM condition and is implied by the absolute continuity assumption. This allows for a broader family of innovation density $f$, including distributions such the double exponential distribution. The proof for Part (b) is based on the functional central limit theorem, the continuous mapping theorem, and an application of hansen1992convergence. Part (c) follows from standard stochastic calculus of the Dol\'eans-Dade exponential, once the Novikov's condition is verified. The detailed proofs are organized in Appendix (ref) of the supplementary material.

Part (a) and Part (b) demonstrate that the limit experiment, specifically with respect to $\boldsymbol{\Pi}$, is LABF as defined by Jeganathan1995. In other words, the central sequence weakly converges to a stochastic integral where the integrand and integrator processes exhibit correlation. This is evident from the covariance matrix equation ((ref)), which shows that $\operatorname{Cov}(\boldsymbol{W}_{\boldsymbol{\varepsilon}}(1),\boldsymbol{W}_{\boldsymbol{\ell}_f}(1)) = \mathbf{I}_p$.

Part (c) allows us to introduce a new collection of probability measures, denoted by $\mathbb{P}_{\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta}}$, through its Radon-Nikodym derivative w.r.t.\ $\mathbb{P}_{\boldsymbol{0},\boldsymbol{0},\boldsymbol{0}}$, given by

align*[align* omitted — 226 chars of source]

The measurable space $(\Omega,\mathcal{F})$ is that of the Brownian motions $(\boldsymbol{W}_{\boldsymbol{\varepsilon}},\boldsymbol{W}_{\boldsymbol{\ell}_f},\boldsymbol{W}_{\boldsymbol{b}})$, which are defined as $\Omega := C^p[0,1]\timesC^p[0,1]\timesC^{\infty}[0,1]$ and $\mathcal{F} := (\otimes_p\mathcal{B}_C) \otimes (\otimes_p\mathcal{B}_C) \otimes (\otimes_{k=1}^{\infty}\mathcal{B}_C)$, where $\mathcal{B}_C$ denotes the Borel $\sigma$-field on $C[0,1]$.

Having introduced the necessary ingredients, we can now provide a formal definition of the limit experiment as

align*[align* omitted — 246 chars of source]

Then, in Le Cam's sense (see, e.g., vdVaart00), the sequence of cointegration experiments, denoted by $\mathcal{E}^{(T)}(f)$, converges to the limit experiment $\mathcal{E}(f)$ as the sample size $T$ tends to infinity. In the following proposition, we regard $\exp\mathcal{L}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta})$ as the Radon-Nikodym derivative and apply Girsanov's Theorem to obtain a structural representation of the limit experiment $\mathcal{E}(f)$.

propositionLet $\mathbf{C}\in\mathbb{R}^{p\timesp}$, $\boldsymbol{\delta}\in\mathbb{R}^{p}$, $\boldsymbol{\eta}\in c_{00}$, and fix $f\in\mathfrak{F}_{p}$. The limit experiment $\mathcal{E}(f)$ associated with the log-likelihood ratio $\mathcal{L}_f(\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta})$ can be described as follows. We observe processes $\boldsymbol{W}_{\boldsymbol{\varepsilon}}$, $\boldsymbol{W}_{\boldsymbol{\ell}_f}$ and $\boldsymbol{W}_{\boldsymbol{b}}$, which are generated according to the following stochastic differential equations (SDEs): \begin{align} {\rm d}\boldsymbol{W}_{\boldsymbol{\varepsilon}}(u) =& \mathbf{C}\boldsymbol{W}_{\boldsymbol{\varepsilon}}(u){\rm d}u + \mathbf{d}_{\mathbf{C}}(u)\boldsymbol{\delta}{\rm d}u + {\rm d}\boldsymbol{Z}_{\boldsymbol{\varepsilon}}(u), \\ {\rm d}\boldsymbol{W}_{\boldsymbol{\ell}_f}(u) =& \boldsymbol{J}_f \mathbf{C}\boldsymbol{W}_{\boldsymbol{\varepsilon}}(u){\rm d}u + \boldsymbol{J}_f\mathbf{d}_{\mathbf{C}}(u)\boldsymbol{\delta}{\rm d}u + \boldsymbol{J}_{f\boldsymbol{b}}\boldsymbol{\eta}{\rm d}u + {\rm d}\boldsymbol{Z}_{\boldsymbol{\ell}_f}(u), \\ {\rm d}\boldsymbol{W}_{\boldsymbol{b}}(u) =& \boldsymbol{J}_{\boldsymbol{b}f} \mathbf{C}\boldsymbol{W}_{\boldsymbol{\varepsilon}}(u){\rm d}u + \boldsymbol{J}_{\boldsymbol{b}f}\mathbf{d}_{\mathbf{C}}(u)\boldsymbol{\delta}{\rm d}u + \boldsymbol{\eta}{\rm d}u + {\rm d}\boldsymbol{Z}_{\boldsymbol{b}}(u), \end{align} where $\boldsymbol{Z}_{\boldsymbol{\varepsilon}}$, $\boldsymbol{Z}_{\boldsymbol{\ell}_f}$ and $\boldsymbol{Z}_{\boldsymbol{b}}$ are Brownian motions under $\mathbb{P}_{\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta}}$ with drift zero and covariance given by ((ref)).

Next, we will use this structural limit experiment to eliminate the nuisance parameters, $\boldsymbol{\eta}$ and $\boldsymbol{\delta}$, sequentially. By doing so, we will derive the semiparametric power envelopes for both cases, without and with a time trend.

\texorpdfstring{Eliminating $\boldsymbol{\eta}$} using Brownian bridge

Despite exhibiting LABF behavior with respect to the direction of $\mathbf{C}$, the limit experiment remains LAN with respect to $\boldsymbol{\eta}$. This is a critical feature, as it results in that $\boldsymbol{\eta}$ only appears as constant drifts in ((ref))--((ref)). To eliminate these drifts, we can simply “take the bridges” of the affected processes.

We formally define the transformation $\mathfrak{g}_{\boldsymbol{\eta}}$ as follows:

align[align omitted — 192 chars of source]

for a process $\boldsymbol{W}\in D^{\mathbb{N}}[0,1]$. We denote by $\mathfrak{G}_{\boldsymbol{\eta}}$ the group of $\mathfrak{g}_{\boldsymbol{\eta}}$ for $\boldsymbol{\eta}\in c_{00}$. Intuitively, the transformation $\mathfrak{g}_{\boldsymbol{\eta}}$ adds a constant drift $u \to -\boldsymbol{\eta}u$ to $\boldsymbol{W}$. To eliminate such a constant drift, we employ the bridge-taking operator defined as

align[align omitted — 109 chars of source]

For a fixed $\boldsymbol{\eta}\in c_{00}$, we have $\boldsymbol{B}^{[\mathfrak{g}_{\boldsymbol{\eta}}(\boldsymbol{W})]}(u) = [\mathfrak{g}_{\boldsymbol{\eta}}(\boldsymbol{W})](u) - u[\mathfrak{g}_{\boldsymbol{\eta}}(\boldsymbol{W})](1) = (\boldsymbol{W}(u) - \boldsymbol{\eta}u) - u(\boldsymbol{W}(1) - \boldsymbol{\eta}) = \boldsymbol{W}(u) - u\boldsymbol{W}(1) = \boldsymbol{B}^{\boldsymbol{W}}(u)$. This shows that the deduced bridge process is an invariant statistic with respect to $\mathfrak{G}_{\boldsymbol{\eta}}$.

Note that in the structural limit experiment described in ((ref))--((ref)), the parameter $\boldsymbol{\eta}$ appears in $\boldsymbol{W}_{\boldsymbol{\ell}_f}$ and $\boldsymbol{W}_{\boldsymbol{b}}$, but not in $\boldsymbol{W}_{\boldsymbol{\varepsilon}}$. Therefore, to eliminate the constant drifts caused by $\boldsymbol{\eta}$, we take the bridges of $\boldsymbol{W}_{\boldsymbol{\ell}_f}$ and $\boldsymbol{W}_{\boldsymbol{b}}$ while keeping $\boldsymbol{W}_{\boldsymbol{\varepsilon}}$ unchanged, resulting in an invariant $\sigma$-field given by

align[align omitted — 170 chars of source]

where $\boldsymbol{B}_{\boldsymbol{\ell}_f} := \boldsymbol{B}^{\boldsymbol{W}_{\boldsymbol{\ell}_f}}$ and $\boldsymbol{B}_{\boldsymbol{b}} := \boldsymbol{B}^{\boldsymbol{W}_{\boldsymbol{b}}}$. Furthermore, we show below that $\mathcal{M}$ is maximally invariant.

theoremIn the limit experiment $\mathcal{E}(f)$ modeled by ((ref))--((ref)), the $\sigma$-field $\mathcal{M}$ defined in ((ref)) is maximally invariant with respect to the group of transformations $\mathfrak{G}_{\boldsymbol{\eta}}$, where $\boldsymbol{\eta}\in c_{00}$.

The proof of Theorem (ref) is based on the definition of maximal invariant in Section 6.2 of LehmannRomano2005.\footnote{Note that $\sigma(\boldsymbol{W}_{\boldsymbol{\varepsilon}},\boldsymbol{W}_{\boldsymbol{\ell}_f},\boldsymbol{W}_{\boldsymbol{b}}) = \sigma(\boldsymbol{W}_{\boldsymbol{\varepsilon}},\boldsymbol{W}_{\boldsymbol{b}})$ due to the decomposition $\boldsymbol{W}_{\boldsymbol{\ell}_f} = \boldsymbol{\Sigma}^{-1}\boldsymbol{W}_{\boldsymbol{\varepsilon}} + \boldsymbol{J}_{f\boldsymbol{b}}\boldsymbol{W}_{\boldsymbol{b}}$ according to the covariance ((ref)). This fact simplifies the proof.} The detailed proof is provided in Appendix (ref) of the supplementary material. The maximal invariant result of $\mathcal{M}$ plays the vital role in our semiparametric optimality study, as every invariant statistic (w.r.t.\ $\boldsymbol{\eta}$) is $\mathcal{M}$-measurable (see LehmannRomano2005), and is every invariant test (w.r.t.\ $\boldsymbol{\eta}$). Therefore, by the Neyman-Pearson Lemma, the likelihood ratio test based on $\mathcal{M}$ is optimal among all invariant test (w.r.t.\ $\boldsymbol{\eta}$). The log-likelihood ratio of $\mathcal{M}$ is given by

align[align omitted — 236 chars of source]

where

align*[align* omitted — 1,849 chars of source]

The detailed derivation is provided in the supplementary Appendix (ref).

Up to this point, we have applied the invariance principle to eliminate the nuisance parameter $\boldsymbol{\eta}$, leaving us with another nuisance parameter $\boldsymbol{\delta}$ to address. In the following subsections, we will sequentially consider two cases: one in which there is no time trend ($\boldsymbol{\delta} = 0$), and another in which there is a time trend ($\boldsymbol{\delta}$ is unknown).

Semiparametric power envelope, no-time-trend case

We begin by considering the case of an intercept only, which corresponds to the error-correction model in ((ref))--((ref)) with $\boldsymbol{\tau} = \boldsymbol{0}$ (or equivalently, $\boldsymbol{\delta} = \boldsymbol{0}$). This specification is important in its own right and is perhaps the most commonly used model in applied works. The advantage of the no-time-trend specification is that the associated tests enjoy better power performance, albeit with a potential cost of model misspecification, when researchers do not find obvious evidence of a time trend in their datasets. For further discussions on this issue, we refer to lutkepohl2000testing and saikkonen2000trend.

We use $\mathcal{L}^{\boldsymbol{\mu}*}_f$ to denote the log-likelihood ratio associated with the maximal invariant under $\boldsymbol{\delta} = \boldsymbol{0}$, for which we have

align[align omitted — 271 chars of source]

where

align*[align* omitted — 724 chars of source]

We define the associated likelihood ratio test by $\phi^{\boldsymbol{\mu}*}_{f,\alpha}(\bar\mathbf{C}) := \mathbbm{1}\{\mathcal{L}^{\boldsymbol{\mu}*}_f(\bar\mathbf{C}) > \kappa^{\boldsymbol{\mu}}_\alpha(\bar\mathbf{C})\}$, where $\kappa^{\boldsymbol{\mu}}_\alpha(\bar\mathbf{C})$ is the $1-\alpha$ quantile of $\mathcal{L}^{\boldsymbol{\mu}*}_f(\bar\mathbf{C})$ under $\mathbb{P}_{\boldsymbol{0},\boldsymbol{0},\boldsymbol{\eta}}$. An upper power bound for $\boldsymbol{\eta}$-invariant tests can then be given as follows:

align*[align* omitted — 309 chars of source]

Call a test $\phi^{(T)}$ in $\mathcal{E}^{(T)}(f)$ asymptotically $\boldsymbol{\eta}$-invariant if it weakly converges, under $\mathrm{P}^{(T)}_{\mathbf{C},\boldsymbol{0},\boldsymbol{\eta};\boldsymbol{\mu},f}$, to a test $\phi$ in $\mathcal{E}(f)$ that is invariant w.r.t.\ $\boldsymbol{\eta}$. The combination of the Neyman-Pearson Lemma and the Asymptotic Representation Theorem (vdVaart00) yields the following theorem.

theoremAssuming that $\boldsymbol{\delta} = \boldsymbol{0}$ is known, let $f\in\mathfrak{F}_{p}$, $\boldsymbol{\mu}\in\mathbb{R}^{p}$, and $\alpha\in(0,1)$. If a test $\phi^{(T)}(\mathbf{y}_1,\dots,\mathbf{y}_T)$, where $T\in\mathbb{N}$, is an asymptotically $\boldsymbol{\eta}$-invariant test of size $\alpha$, i.e., $\limsup_{T\to\infty}\mathrm{E}_{\boldsymbol{0},\boldsymbol{0},\boldsymbol{\eta};\boldsymbol{\mu},f}[\phi^{(T)}]\leq\alpha$, then we have \begin{align*} \limsup_{T\to\infty} \mathrm{E}_{\mathbf{C},\boldsymbol{0},\boldsymbol{\eta};\boldsymbol{\mu},f}[\phi^{(T)}] \leq \pi^{\boldsymbol{\mu}*}_{f,\alpha}(\mathbf{C};\mathbf{C}), \forall \, \mathbf{C}\in\mathbb{R}^{p\timesp}, \, \boldsymbol{\eta}\in c_{00}. \end{align*}

The proof of the theorem involves two main steps. First, we use the Neyman-Pearson lemma to establish the upper bound for the power of $\boldsymbol{\eta}$-invariant tests in $\mathcal{E}(f)$ at point $\mathbf{C}$, which is $\pi^{\boldsymbol{\mu}*}_{f,\alpha}(\mathbf{C};\mathbf{C})$. This determines the maximum achievable power in the limit experiment. Second, the Asymptotic Representation Theorem states that any test in the sequence of experiments $\mathcal{E}^{(T)}(f)$ has a representation in the limit experiment $\mathcal{E}(f)$. Therefore, the best achievable power in $\mathcal{E}(f)$ is also the best possible power that can be achieved asymptotically in $\mathcal{E}^{(T)}(f)$. The proof is completed by applying this result to the class of asymptotically $\boldsymbol{\eta}$-invariant tests.

\texorpdfstring{Eliminating $\boldsymbol{\delta}$ using profile likelihood & semiparametric power envelope for time-trend case}

In this subsection, we consider case where the linear trend parameter $\boldsymbol{\delta}$ is unknown and treated as a nuisance parameter. To eliminate $\boldsymbol{\delta}$, observing that the log-likelihood ratio $\mathcal{L}^{\mathcal{M}}_f(\mathbf{C},\boldsymbol{\delta})$ is quadratic in $\boldsymbol{\delta}$, we use the profile likelihood method. Specifically, $\boldsymbol{\delta}$ is “profiled out” by

align[align omitted — 221 chars of source]

which results in an invariant statistic w.r.t $\boldsymbol{\delta}$.\footnote{More formally, the obtained statistic is invariant w.r.t.\ the group of transformations of the form $\mathbf{y}_t = \mathbf{y}_t + \mfcu$, $\forall\mathbf{c}\in\mathbb{R}^{p}$.} The optimality of this approach for quadratic-in-nuisance form is established in, e.g., LehmannRomano2005, where it is shown that the resulting profile likelihood-based test is the best among invariant tests. For further reference on this method in the unit root testing literature, see elliott1996efficient and Jansson2008.

We split the optimization of ((ref)) into two steps: (i) we derive the maximum likelihood estimate of $\boldsymbol{\delta}$ by solving the maximization under one alternative value of $\mathbf{C}$, and (ii) we plug this estimate into likelihood statistic $\mathcal{L}^{\mathcal{M}}_f$ under another alternative value. This splitting allows us to obtain the semiparametric analogue of the statistic $\mathit{\Lambda}_{p,C}^{GLS}(\bar{C};\bar{C}^*)$ introduced in boswijk2015improved (BJN). This statistic encompasses existing Gaussian cointegration test statistics obtained by choosing different values of $\bar{C}$ and $\bar{C}^*$. By interpreting the choices of $\bar{C}$ and $\bar{C}^*$ in BJN's $\mathit{\Lambda}_{p,C}^{GLS}(\bar{C};\bar{C}^*)$, we contribute to the literature by providing insights into how the current Gaussian cointegration tests for the time trend case are derived in the limiting perspective, and by proposing semiparametric efficient versions of these tests using our semiparametric analogue afterwards.

In detail, we select the reference alternative $\bar\mathbf{C}^*$ and derive the maximum likelihood estimate (MLE) of $\boldsymbol{\delta}$ as follows:

align[align omitted — 285 chars of source]

where

align*[align* omitted — 1,952 chars of source]

are the linear and quadratic term in $\boldsymbol{\delta}$, respectively. Once we have $\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}$, we use another reference alternative $\bar\mathbf{C}$ to construct the profile likelihood ratio as

align[align omitted — 681 chars of source]

with

align*[align* omitted — 1,018 chars of source]

where $\boldsymbol{W}_{\boldsymbol{\varepsilon}}^{\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}}(u) := \boldsymbol{W}_{\boldsymbol{\varepsilon}}(u) - u\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}$ is the de-drifted version (of $\boldsymbol{W}_{\boldsymbol{\varepsilon}}$) with drift estimate $\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}$, and $\widebar\boldsymbol{W}_{\boldsymbol{\varepsilon}}^{\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}} := \widebar\boldsymbol{W}_{\boldsymbol{\varepsilon}} - \frac{1}{2}\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}$ is its average.\footnote{Here the profile log-likelihood $\mathcal{L}^{\mathcal{M}}_f(\bar\mathbf{C},\hat{\boldsymbol{\delta}}_{\bar\mathbf{C}^*}) - \mathcal{L}^{\mathcal{M}}_f(\boldsymbol{0},\hat\boldsymbol{\delta}_{\boldsymbol{0}})$ is derived by a sum of two terms, $\mathcal{L}^{\mathcal{M}}_f(\boldsymbol{0},\hat{\boldsymbol{\delta}}_{\bar\mathbf{C}^*}) - \mathcal{L}^{\mathcal{M}}_f(\boldsymbol{0},\hat\boldsymbol{\delta}_{\boldsymbol{0}}) = -0.5\,\boldsymbol{W}^{\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}}_{\boldsymbol{\varepsilon}}(1)^\prime\boldsymbol{\Sigma}^{-1}\boldsymbol{W}^{\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}}_{\boldsymbol{\varepsilon}}(1)$ and $\mathcal{L}^{\mathcal{M}}_f(\bar\mathbf{C},\hat{\boldsymbol{\delta}}_{\bar\mathbf{C}^*}) - \mathcal{L}^{\mathcal{M}}_f(\boldsymbol{0},\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}) = \mathit{\Delta}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}^*) - \frac{1}{2}\mathcal{Q}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}^*)$, where $\hat\boldsymbol{\delta}_{\boldsymbol{0}}$ denotes the estimate for $\boldsymbol{\delta}$ under the chosen alternative $\bar\mathbf{C}^* = \boldsymbol{0}$. By ((ref)), we have $\hat\boldsymbol{\delta}_{\boldsymbol{0}} = \boldsymbol{W}_{\boldsymbol{\varepsilon}}(1)$. The second term is intuitive and easy-to-interpret --- it is nothing but the likelihood ratio of the sigma-field $\sigma(\boldsymbol{W}_{\boldsymbol{\varepsilon}}^{\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}}, \boldsymbol{B}_{\boldsymbol{\ell}_f})$.}

Our statistic $\mathcal{L}^{\boldsymbol{\tau}*}f(\bar\mathbf{C};\bar\mathbf{C}^*)$ serves as the semiparametric counterpart and reduces to BJN's $\mathit{\Lambda}{p,C}^{GLS}(\bar{C};\bar{C}^*)$ when $f$ is Gaussian. Specifically, under Gaussianity, the score function is given by $\boldsymbol{\ell}_f(\boldsymbol{\varepsilon}_t) = \Sigma^{-1}\boldsymbol{\varepsilon}_t$, and the limit experiment in Proposition (ref) simplifies to ((ref)). Using $\bar\mathbf{C}^*$ to estimate $\boldsymbol{\delta}$ and $\bar\mathbf{C}$ to construct the statistic, we arrive at BJN's $\mathit{\Lambda}_{p,C}^{GLS}(\bar{C};\bar{C}^*)$. It is worth noting again that $\mathit{\Lambda}{p,C}^{GLS}(\bar{C};\bar{C}^*)$ embodies asymptotic representations of existing Gaussian cointegration tests. In a similar vein, in Section (ref), we will base on our $\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}^*)$ to derive their semiparametric versions.

We conclude this section by presenting the semiparametric power envelope for the case of a linear time trend. We let $\bar\mathbf{C}^* = \bar\mathbf{C}$ and define the likelihood ratio test as $\phi^{\boldsymbol{\tau}*}_{f,\alpha}(\bar\mathbf{C}) := \mathbbm{1}\{\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}) > \kappa^{\boldsymbol{\tau}}_\alpha(\bar\mathbf{C})\}$, where $\kappa^{\boldsymbol{\tau}}_\alpha(\bar\mathbf{C})$ is the $1-\alpha$ quantile of $\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C})$. The power function is given by

align*[align* omitted — 316 chars of source]

We refer to a test $\phi_{T}$ in $\mathcal{E}^{(T)}(f)$ as asymptotically $(\boldsymbol{\delta},\boldsymbol{\eta})$-invariant if it weakly converges to a test in $\mathcal{E}(f)$ that is invariant w.r.t.\ both $\boldsymbol{\eta}$ and $\boldsymbol{\delta}$. Using reasoning analogous to Theorem (ref), we provide the semiparametric power envelope of asymptotically $(\boldsymbol{\delta},\boldsymbol{\eta})$-invariant tests for the time trend case in the following theorem.

theoremAssuming that $\boldsymbol{\delta}$ is unknown, let $f\in\mathfrak{F}_{p}$, $\boldsymbol{\mu}\in\mathbb{R}^{p}$, and $\alpha\in(0,1)$. If a test $\phi^{(T)}(\mathbf{y}_1,\dots,\mathbf{y}_T)$, where $T\in\mathbb{N}$, is an asymptotically $(\boldsymbol{\delta},\boldsymbol{\eta})$-invariant test of size $\alpha$, that is, $\limsup_{T\to\infty}\mathrm{E}_{\boldsymbol{0},\boldsymbol{\delta},\boldsymbol{\eta};\boldsymbol{\mu},f}[\phi^{T}]\leq\alpha$, then we have \begin{align*} \limsup_{T\to\infty}\mathrm{E}_{\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta};\boldsymbol{\mu},f}[\phi^{(T)}] \leq \pi^{\boldsymbol{\tau}*}_{f,\alpha}(\mathbf{C};\mathbf{C}), \forall \, \mathbf{C}\in\mathbb{R}^{p\timesp}, \, \boldsymbol{\delta}\in\mathbb{R}^p, \, \boldsymbol{\eta}\in c_{00}. \end{align*}

Semiparametric Inference

This section proposes semiparametrically optimal cointegration tests based on the asymptotic results above. We first employ $\mathcal{L}^{\boldsymbol{\mu}*}_f(\mathbf{C})$ (resp.\ $\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}^*)$) to derive (asymptotic) semiparametric counterparts of some existing Gaussian cointegration tests for the case without time trend (resp.\ the case with time trend) in Section (ref) (resp.\ Section (ref)). Then, in Section (ref), we construct feasible semiparametric cointegration tests primarily by replacing $\boldsymbol{B}_{\boldsymbol{\ell}_f}$ by its finite-sample counterpart, which is built upon a nonparametric estimate of the score function $\boldsymbol{\ell}_f$.

Semiparametric counterparts of existing tests (no-time-trend case)

When there is no time trend ($\boldsymbol{\delta} = \boldsymbol{0}$), we follow the line of likelihood ratio test by johansen1991estimation and saikkonen1997testing, and profile out the chosen alternative $\bar\mathbf{C}$. Specifically, we maximize the likelihood ratio statistic $\mathcal{L}^{\boldsymbol{\mu}*}_f(\bar\mathbf{C})$ in ((ref)) w.r.t.\ $\bar\mathbf{C}\in\mathbb{R}^{p\timesp}$, and obtain the semiparametric Johansen trace statistic

align[align omitted — 843 chars of source]

which does not require a specification for $\bar\mathbf{C}$ anymore. See jansson2012nearly for a comprehensive study on this operation for unit root testing. Borrowing their conclusion to our multivariate case, we argue that the resulting $\Lambda^{Johansen}_f$-based test attains $\pi^{\boldsymbol{\mu}*}_{f,\alpha}(\mathbf{C};\mathbf{C})$ pointwise at a random alternative point.

Under the Gaussianity of the true density, i.e., $f = \phi$, we have $\boldsymbol{B}_{\boldsymbol{\ell}_f} = \boldsymbol{\Sigma}^{-1}\boldsymbol{B}_{\boldsymbol{\varepsilon}}$ and $\boldsymbol{J}_f = \boldsymbol{\Sigma}^{-1}$, and the semiparametric statistic $\Lambda^{Johansen}_f$ simplifies to the original Gaussian Johansen trace statistic as follows:

align[align omitted — 572 chars of source]

Semiparametric counterparts of existing tests (time-trend case)

When dealing with a linear time trend (i.e., unknown $\boldsymbol{\delta}$), we consider the GLS trace test proposed by saikkonen2000trend (SL test) (see also lutkepohl2000testing). In this case, we choose $\bar\mathbf{C}^* = \boldsymbol{0}$ for the trend parameter estimate $\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}$, which leads to $\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*} = \hat\boldsymbol{\delta}_{\boldsymbol{0}} = \boldsymbol{W}_{\boldsymbol{\varepsilon}}(1)$. Consequently, $\boldsymbol{W}_{\boldsymbol{\varepsilon}}^{\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}}(u) = \boldsymbol{W}_{\boldsymbol{\varepsilon}}(u) - u\boldsymbol{W}_{\boldsymbol{\varepsilon}}(1) = \boldsymbol{B}_{\boldsymbol{\varepsilon}}(u)$ and $\boldsymbol{W}_{\boldsymbol{\varepsilon}}^{\hat\boldsymbol{\delta}_{\bar\mathbf{C}^*}}(1) = \boldsymbol{0}$. The likelihood ratio statistic is then given by:

align*[align* omitted — 401 chars of source]

By maximizing $\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\boldsymbol{0})$ w.r.t.\ $\bar\mathbf{C}$, we obtain the semiparametric SL trace statistic defined as

align[align omitted — 489 chars of source]

When the true density is Gaussian (i.e., $f = \phi$), we have $\boldsymbol{B}_{\boldsymbol{\ell}_f} = \boldsymbol{\Sigma}^{-1}\boldsymbol{B}_{\boldsymbol{\varepsilon}}$ and $\boldsymbol{J}_f = \boldsymbol{\Sigma}^{-1}$. In this case, the semiparametric SL trace statistic $\Lambda^{SL}_f$ reduces to the original Gaussian SL trace statistic, which is given by

align[align omitted — 579 chars of source]
remarkAlternatively, one can consider the statistic proposed by boswijk2015improved, which corresponds to the case where $\bar\mathbf{C}^* = \bar\mathbf{C}$.\footnote{After substantial rearrangement, one can achieve $\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}) = \mathcal{L}^{\boldsymbol{\mu}*}_f(\bar\mathbf{C}) + \frac{1}{2}\mathcal{L}^{d}_f(\bar\mathbf{C})-\frac{1}{2}\mathcal{L}^{d}_f(\boldsymbol{0})$, where $\mathcal{L}^{d}_f(\mathbf{C}) = \mathit{\Delta}^d_f(\mathbf{C})^\prime\mathcal{Q}^{dd}_f(\mathbf{C})^{-1}\mathit{\Delta}^{d}_f(\mathbf{C})$, with $\mathit{\Delta}^{d}_f(\mathbf{C}) := \int_0^1\mathbf{d}_{\mathbf{C}}(u)^{\prime}{\rm d}\big(\boldsymbol{B}_{\boldsymbol{\ell}_f}(u)+\boldsymbol{\Sigma}^{-1}\boldsymbol{W}_{\boldsymbol{\varepsilon}}(1)u - \boldsymbol{J}_f\mathbf{C}\widetilde{\boldsymbol{W}}_{\boldsymbol{\varepsilon}}(u)u - \boldsymbol{\Sigma}^{-1}\mathbf{C} \begingroup \def\mathaccent#\boldsymbol{W}##2{ \kern0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax111{\boldsymbol{W}} \endgroup _{\boldsymbol{\varepsilon}}u\big)$, $\mathcal{Q}^{dd}_f(\mathbf{C}) := \int_0^1\mathbf{d}_{\mathbf{C}}(u)^\prime\boldsymbol{J}_f\mathbf{d}_{\mathbf{C}}(u){\rm d}u + \begingroup \def\mathaccent#\mathbf{d}##2{ \kern0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax111{\mathbf{d}} \endgroup _{\mathbf{C}}^\prime\left(\boldsymbol{\Sigma}^{-1}-\boldsymbol{J}_f\right) \begingroup \def\mathaccent#\mathbf{d}##2{ \kern0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax111{\mathbf{d}} \endgroup _{\mathbf{C}}$, and $\widetilde\boldsymbol{W}_{\boldsymbol{\varepsilon}}(u) := \boldsymbol{W}_{\boldsymbol{\varepsilon}}(u) - \widebar\boldsymbol{W}_{\boldsymbol{\varepsilon}}$.} By taking the maximum of $\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C})$ w.r.t.\ $\bar\mathbf{C}$, one obtains the semiparametric BJN statistic \begin{align*} \Lambda^{BJN}_f := \max_{\bar\mathbf{C}\in\mathbb{R}^{p\timesp}}\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}). \end{align*} Similarly, when $f$ is Gaussian, the statistic $\Lambda^{BJN}_f$ reduces to its Gaussian counterpart introduced in their Section 2.2 of their paper.

Semiparametrically optimal cointegration tests

This subsection proposes feasible cointegration tests whose powers can achieve the power bounds developed earlier, establishing that these power bounds are sharp (globally in $f$) and our tests are semiparametrically optimal (pointwise in $\mathbf{C}$).

Our tests are constructed using a nonparametric estimate of the unknown score function $\boldsymbol{\ell}_f$. For this estimation, one could consider the sample splitting technique, which requires fairly minimal conditions on $f$ (see Bickel1982, schick1986asymptotically, and drost1997adaptive). However, to improve the finite-sample performance, especially when the sample size is moderately small, we follow the approach of schick1987note which uses the full sample without sample splitting. Unlike other works along this line that require the symmetric density assumption (e.g., kreiss1987adaptive for the ARMA model and ling2003adaptive for the ARMA-GARCH model), our test statistic construction is based on the framework of koul1997efficient, which does not need any additional condition on $f$.

Our estimate for $\boldsymbol{\ell}_f$ is constructed as follows:

align[align omitted — 173 chars of source]

where $\hatf$ is the usual kernel density estimator given by

align[align omitted — 168 chars of source]

Here $K(\boldsymbol{\varepsilon}) = k(\varepsilon_1) \cdots k(\varepsilon_p)$ is the kernel function, and $a_T$ and $b_T$ are positive sequences converging to zero. We impose the following mild assumptions.

assumption\begin{itemize} • The marginal kernel function $k(\cdot)$ is bounded, symmetric, continuously differentiable with $\int_{\mathbb{R}}r^2k(r)dr < \infty$ and, for some positive constant C, $|\dot{k}(r)|/k(r) < \infty$ for all $r\in\mathbb{R}$. • The sequences $\langle a_T \rangle$ and $\langle b_T \rangle$ satisfy $a_T \to 0$, $b_T \to 0$, and $T a_T^4 b_T^2 \to \infty$. \end{itemize}

Using the nonparametric estimate $\hat\boldsymbol{\ell}_f$, we construct estimators for $\boldsymbol{B}_{\boldsymbol{\ell}_f}$ and $\boldsymbol{J}_f$ as follows:

align[align omitted — 310 chars of source]

and

align[align omitted — 175 chars of source]

We also define

align[align omitted — 152 chars of source]

and assume that $\widehat\boldsymbol{\Sigma}$ is some consistent estimate of $\boldsymbol{\Sigma}$. With these tools in hand, we can define $\widehat\mathcal{L}^{\boldsymbol{\mu}*}_f(\mathbf{C})$ as the feasible version of the likelihood ratio statistic $\mathcal{L}^{\boldsymbol{\mu}*}_f(\mathbf{C})$ by replacing $\boldsymbol{W}_{\boldsymbol{\varepsilon}}$, $\boldsymbol{B}_{\boldsymbol{\ell}_f}$, $\boldsymbol{\Sigma}$, and $\boldsymbol{J}_f$ in the latter with their finite-sample counterparts $\boldsymbol{W}^{(T)}_{\boldsymbol{\varepsilon}}$, $\widehat\boldsymbol{B}_{\boldsymbol{\ell}_f}^{(T)}$, $\widehat\boldsymbol{\Sigma}$, and $\widehat{\boldsymbol{J}}_f$, respectively. In similar way, we define $\widehat\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}^*)$ as the feasible version of $\mathcal{L}^{\boldsymbol{\tau}}_f(\bar\mathbf{C};\bar\mathbf{C}^*)$. The following theorem summarizes the consistency results for these nonparametric estimators. The proof is presented in Appendix (ref).

theorem\begin{itemize} • Assuming that $f\in\mathfrak{F}_p$ and Assumption (ref) holds, let $\widehat\boldsymbol{B}_{\boldsymbol{\ell}_f}^{(T)}$ and $\widehat\boldsymbol{J}_f$ be defined as in ((ref))--((ref)). Then, as $T\to\infty$ under $\mathrm{P}^{(T)}_{\mathbf{C},\boldsymbol{0};f}$, we have \begin{align} \widehat\boldsymbol{B}_{\boldsymbol{\ell}_f}^{(T)} \Rightarrow \boldsymbol{B}_{\boldsymbol{\ell}_f} {\rm and } \widehat\boldsymbol{J}_f \stackrel{p}{\to} \boldsymbol{J}_f. \end{align} • Let $\widehat\boldsymbol{W}_{\boldsymbol{\varepsilon}}$ be defined by ((ref)) and let $\widehat\boldsymbol{\Sigma}$ be some consistent estimator for $\boldsymbol{\Sigma}$ such that $\widehat\boldsymbol{\Sigma}\stackrel{p}{\to}\boldsymbol{\Sigma}$. Then, under $\mathrm{P}^{(T)}_{\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta};\boldsymbol{\mu},f}$ and as $T\to\infty$, \begin{align} \widehat\mathcal{L}^{\boldsymbol{\mu}*}_f(\bar\mathbf{C}) \Rightarrow \mathcal{L}^{\boldsymbol{\mu}*}_f(\bar\mathbf{C}) {\rm and } \widehat\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}^*) \Rightarrow \mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}^*), \end{align} where $\mathcal{L}^{\boldsymbol{\mu}*}_f(\bar\mathbf{C})$ and $\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}^*)$ are defined respectively in ((ref)) and ((ref)), whose behaviors are characterized by Proposition (ref). \end{itemize}

In the proof, a key step is to address the bias, denoted as $\boldsymbol{\zeta}$, that arises when estimating $\boldsymbol{\ell}_f$ due to the lack of an $\sqrt{T}$-unbiased estimator for $\boldsymbol{\ell}_f$ when $f$ is not further restricted, such as being symmetric (see Klaassen1987). However, this bias is canceled out by $\widehat\boldsymbol{B}_{\boldsymbol{\ell}_f}^{(T)}$ as shown in the following calculation:

align*[align* omitted — 696 chars of source]

where $\check\boldsymbol{\ell}_f^{(T)}$ denotes the debiased version of $\hat\boldsymbol{\ell}_f^{(T)}$. Notably, as $\boldsymbol{B}_{\boldsymbol{\ell}_f}$ eliminates the density perturbation $\boldsymbol{\eta}$ in the limit, its finite-sample counterpart $\widehat\boldsymbol{B}_{\boldsymbol{\ell}_f}^{(T)}$ eliminates the nonparametric estimation bias in the sequence, effectively addressing the bias issue in the estimation of $\boldsymbol{\ell}_f$.

As a direct consequence of Theorem (ref), the following corollary demonstrates that the power upper bounds $\pi^{\boldsymbol{\mu}*}_{f,\alpha}(\mathbf{C};\mathbf{C})$ and $\pi^{\boldsymbol{\tau}*}_{f,\alpha}(\mathbf{C};\mathbf{C})$ are globally sharp in $f\in\mathfrak{F}_p$.

corollaryLet $f\in\mathfrak{F}_p$ and assume that Assumption (ref) holds. Define likelihood ratio tests as $\hat\phi^{\boldsymbol{\mu}*}_{\alpha}(\bar\mathbf{C}) := \mathbbm{1}\{\widehat\mathcal{L}^{\boldsymbol{\mu}*}_f(\bar\mathbf{C}) > \kappa^{\boldsymbol{\mu}}_\alpha(\bar\mathbf{C})\}$ and $\hat\phi^{\boldsymbol{\tau}*}_{\alpha}(\bar\mathbf{C}) := \mathbbm{1}\{\widehat\mathcal{L}^{\boldsymbol{\tau}*}_f(\bar\mathbf{C};\bar\mathbf{C}) > \kappa^{\boldsymbol{\tau}}_\alpha(\bar\mathbf{C})\}$. Then, under $\mathrm{P}^{(T)}_{\mathbf{C},\boldsymbol{\delta},\boldsymbol{\eta};\boldsymbol{\mu},f}$, we have \begin{align*} \lim_{T\to\infty} \mathrm{E}\big[\hat\phi^{\boldsymbol{\mu}*}_{\alpha}(\bar\mathbf{C})\big] = \pi^{\boldsymbol{\mu}*}_{f,\alpha}(\bar\mathbf{C};\mathbf{C}) {\rm and } \lim_{T\to\infty} \mathrm{E}\big[\hat\phi^{\boldsymbol{\tau}*}_{\alpha}(\bar\mathbf{C})\big] = \pi^{\boldsymbol{\tau}*}_{f,\alpha}(\bar\mathbf{C};\mathbf{C}), \end{align*} for all $\bar\mathbf{C},\mathbf{C}\in\mathbb{R}^{p\timesp}$.

To avoid the necessity of selecting a reference alternative $\bar\mathbf{C}$, we adopt the approach used in the aforementioned limiting Johansen test $\Lambda^{Johansen}_f$ and the limiting SL test $\Lambda_f^{SL}$. In other words, we construct the feasible versions of these tests by replacing $\boldsymbol{W}_{\boldsymbol{\varepsilon}}$, $\boldsymbol{B}_{\boldsymbol{\ell}_f}$, $\boldsymbol{\Sigma}$, and $\boldsymbol{J}_f$ with their finite-sample counterparts $\boldsymbol{W}^{(T)}_{\boldsymbol{\varepsilon}}$, $\widehat\boldsymbol{B}_{\boldsymbol{\ell}_f}^{(T)}$, $\widehat\boldsymbol{\Sigma}$, and $\widehat\boldsymbol{J}_f$, respectively. These feasible versions are based on the estimate $\hatf$ and are denoted as $\hat\Lambda^{Johansen}f$ and $\hat\Lambdaf^{SL}$, respectively.

remarkIn addition to the kernel estimation method used in this paper, an alternative approach for density estimation is the semi-nonparametric (SNP) method proposed by gallant1987semi. For the SNP method employed for the cointegration testing problem, see boswijk2002semi.

Simulations

In this section, we conduct a Monte Carlo study to validate the asymptotic results using large-sample simulations and evaluate the small-sample performance of our proposed semiparametrically optimal cointegration tests above. Specifically, we compare our tests, which are based on the statistics $\Lambda^{Johansen}_f$ for the case without time trend, and $\Lambda^{SL}_f$ for the case with time trend, to their original Gaussian versions, which are based on the statistics $\Lambda^{Johansen}_{\phi}$ and $\Lambda^{SL}_{\phi}$, respectively.

We consider a two-dimension VAR model in error correction form ((ref))--((ref)) with i.i.d.\ innovations $\boldsymbol{\varepsilon}_t$ of density $f\in\mathfrak{F}_2$. Following the asymptotic local power analysis, we set the parameter of interest to be $\boldsymbol{\Pi} = \mathbf{C}/T$ as in ((ref)). We specify the local parameter $\mathbf{C}$ using the form

align*[align* omitted — 63 chars of source]

where $c\in(-\infty,0]$. Thus, $c=0$ corresponds to the null hypothesis of $r=0$, while $c<0$ corresponds to the alternative hypothesis of $r=1$, recalling that $r$ denotes the rank of $\boldsymbol{\Pi}$. We consider $T=2500$ for the large-sample case and $T=250$ for the small-sample case. We base all results on $20,000$ independent replications, and set the significance level to be $5\%$ throughout this section.

We estimate the covariance matrix $\boldsymbol{\Sigma}$ using the following estimator:

align*[align* omitted — 201 chars of source]

where $\overline{\Delta\mathbf{y}} = \sum_{t=2}^{T}\Delta\mathbf{y}_t/(T-1)$. To construct the density estimator in ((ref)), which is used to obtain $\boldsymbol{\ell}_f$ in ((ref)) and subsequently $\widehat\boldsymbol{B}_{\boldsymbol{\ell}_f}^{(T)}$ and $\widehat\boldsymbol{J}_f$ as defined in ((ref))--((ref)), we choose standard Logistic kernels for $k(\cdot)$ and set $b_T = 0$. Note that the choice of $b_T = 0$ violates Assumption (ref)-(b), but we made this choice for simplicity and concreteness as the qualitative results appeared to be more sensitive to the choice of $a_T$ than to the choice of $b_T$. Following Silverman's rule of thumb (see silverman1986density), we set the bandwidth $a_T$ for each dimension $i = 1,\dots,p$ to be

align*[align* omitted — 84 chars of source]

where $\hat{\sigma}_i^2$ is the $i$-th diagonal element of $\widehat\boldsymbol{\Sigma}$.

figure[figure omitted — 385 chars of source]
figure[figure omitted — 334 chars of source]

Figure (ref) displays the large-sample performances ($T = 2500$) of the Johansen test (labeled Johansen-$\phi$) and the $\hat{f}$-based Johansen test (labeled Johansen-$\hat{f}$) for the case without a time trend (upper panel), and the SL test (labeled SL-$\phi$) and the $\hat{f}$-based SL test (labeled SL-$\hat{f}$) for the case with a time trend (bottom panel). We consider two different distributions: the Multivariate Student-$t_3$ distribution (left panel) and the Gaussian distribution (right panel). For the case without a time trend and under Gaussian $f$, both Johansen-type tests exhibit similar powers that are close to the Gaussian power envelope. Under the Student-$t_3$ distribution, the $\hat{f}$-based Johansen test outperforms the original Johansen test by a considerable amount of power gain. For instance, at $-c = 10$, the rejection rate of Johansen-$\phi$ is slightly below $60\%$, while that of Johansen-$\hat{f}$ is nearly $90\%$. This demonstrates that our semiparametric efficient Johansen test can effectively utilize the information contained in the innovation distribution, which is particularly advantageous when the true density deviates significantly from Gaussian, a limitation of the original version.

Furthermore, it is worth noting that a small discrepancy may be observed between the power of the Johansen-$\hat{f}$ test and the semiparametric power envelope. This slight reduction in power can be attributed to the finite-sample nature of the data, even with a relatively large sample size, and the nonparametric estimation of the density $f$. Nevertheless, this gap tends to decrease as the sample size $T$ increases towards infinity, and the power of the Johansen-$\hat{f}$ test converges to the semiparametric power envelope. These results are also applicable to the time trend case for the corresponding SL-$\phi$ and SL-$\hat{f}$ tests, which provides numerical evidence that our proposed $\hat{f}$-based tests are semiparametrically optimal.

Figure (ref) presents the small-sample performances with a sample size of $n=250$, which corresponds to the scenarios depicted in Figure (ref). These results confirm that all four tests under evaluation, namely Johansen-$\phi$ and Johansen-$\hat{f}$ for the no-time-trend case, and SL-$\phi$ and SL-$\hat{f}$ for the time-trend case, maintain satisfactory size control even with a relatively small sample size. Furthermore, we observe that the power properties observed in the large-sample case are also applicable in the small-sample case. Notably, the $\hat{f}$-based tests provide significant power improvement over their Gaussian counterparts when $f$ follows a Student-$t_3$ distribution, even with a small sample size. However, when $f$ is Gaussian, the $\hat{f}$-based tests may exhibit slightly lower power compared to their Gaussian counterparts, mainly due to the need for nonparametric density estimation.

figure[figure omitted — 342 chars of source]

We extend our investigation to the scenarios under asymmetric distributions, as illustrated in Figure (ref). Specifically, we examine the performance of the four tests (Johansen-$\phi$, Johansen-$\hat{f}$, SL-$\phi$, and SL-$\hat{f}$) under the Skewed-$t_4$ distribution in both large ($T=2500$, left panel) and small ($T=250$, right panel) samples for both the no-time-trend case (upper panel) and the time-trend case (bottom panel). Our results confirm the previously reported findings, including satisfactory size control and the remarkable power improvement of the $\hat{f}$-based tests over their Gaussian counterparts, are applicable to the asymmetric distribution case as well. Moreover, additional simulations unreported here indicate that the $\hat{f}$-based tests gain more and more power as the Skewed-$t_4$ distribution becomes more and more skewed (i.e., further away from Gaussian).

Discussions on extensions

The preceding sections have focused on a simple case for semiparametric efficiency, which assumed no serial correlation in the error term and considered only the null hypothesis of no cointegrating relationship (i.e., $H_0: \, r = 0$). To broaden the applicability of our tests, we briefly discuss whether our results remain valid when we relax the assumption of no serial correlation and consider a more general reduced rank hypothesis.\footnote{In the Gaussian case, boswijk2015improved have addressed the results of relaxing these two assumptions in their Sections 2.3 and 2.4.}

We consider a generalized model where the observations $\mathbf{y}_t = (y_{1,t},\dots,y_{p,t})^{\prime}$ are generated as follows:

align[align omitted — 223 chars of source]

where, in addition, the parameter $\boldsymbol{\Gamma} = \{\boldsymbol{\Gamma}_1,\dots,\boldsymbol{\Gamma}_{k-1}\}\in\mathbb{R}^{p\times(k-1)p}$ of (known and finite) order $p$ governs the lag terms. We are interested in testing the hypothesis

align*[align* omitted — 52 chars of source]

where $r_0\in\mathbb{N}_0$ and $r_0 < p$.

Our discussion is based on the results in HvdAW, specifically their `complete' version of the limit experiment for cointegration provided in their online supplementary appendix (Proposition A.2). Following HvdAW, we consider the following localized parameterizations:

itemize• For $\boldsymbol{\Gamma}$, we use \begin{align} \boldsymbol{\Gamma} = \boldsymbol{\Gamma}^{(T)}_{\mathbf{G}} = \boldsymbol{\Gamma}_0 + \frac{\mathbf{G}}{\sqrt{T}}, \end{align} where the local parameter $\mathbf{G} = \{\mathbf{G}_1,\dots,\mathbf{G}_{k-1}\} \in \mathbb{R}^{p\times(k-1)p}$ • For $\boldsymbol{\Pi}$, building upon the factorization $\boldsymbol{\Pi} = \boldsymbol{\alpha}_0\boldsymbol{\beta}_0$, where $\boldsymbol{\alpha}_0$ and $\boldsymbol{\beta}_0$ are full-rank $p\times r_0$ matrices, we consider \begin{align*} \boldsymbol{\Pi} = \boldsymbol{\Pi}^{(T)}_{\mathbf{A},\mathbf{B},\mathbf{C}} = \boldsymbol{\alpha}\boldsymbol{\beta}^\prime+\frac{\boldsymbol{\alpha}_\perp\mathbf{C}\boldsymbol{\alpha}_\perp^\prime}{T}, \end{align*} and \begin{align} \boldsymbol{\alpha} = \boldsymbol{\alpha}^{(T)}_{\mathbf{A}} = \boldsymbol{\alpha}_0 + \frac{\mathbf{A}}{\sqrt{T}} {\rm and } \boldsymbol{\beta} = \boldsymbol{\beta}^{(T)}_{\mathbf{B}} = \boldsymbol{\beta}_0 + \frac{\boldsymbol{\beta}_{\perp}\mathbf{B}^\prime}{T} \end{align} where $\boldsymbol{\alpha}_\perp$ and $\boldsymbol{\beta}_\perp$ are chosen $p\times(p-r_0)$ matrices of rank $p-r_0$ satisfying $\boldsymbol{\alpha}^\prime\boldsymbol{\alpha}_\perp = \boldsymbol{0}_{r_0\times(p-r_0)}$ and $\boldsymbol{\beta}^\prime\boldsymbol{\beta}_\perp = \boldsymbol{0}_{r_0\times(p-r_0)}$, and $\mathbf{C}\in\mathbb{R}^{p-r_0\times p-r_0}$.

With these local alternatives, our inference problem becomes testing the null hypothesis of $\mathbf{C}=\boldsymbol{0}$ (since $r=r_0$ if and only if $\mathbf{C}=\boldsymbol{0}$), while treating $\mathbf{G}$, $\mathbf{A}$ and $\mathbf{B}$ as nuisance parameters.

We incorporate our generalized model into the framework of hallin2016semiparametric. Recalling above that the key difference between our model and theirs is the specification of the trend term, with our model described by equations ((ref))--((ref)), while theirs is described by equation ((ref)). Our case corresponds to theirs when their parameter $\mu = 0$ (and our additional time trend term $\boldsymbol{\mu} + \ptrdt$ can be handled in a similar manner as described in Section (ref)). As a result, all local perturbations associated with $\mu$, specifically their parameters $b$, $d$, and $\mu$ itself, are absent in our model. While HvdAW focus on the parameter of interest $d$, which enjoys a “super-consistency” rate $T^{-3/2}$ and the traditional LAN result, our paper focuses on the parameter $\mathbf{C}$ (corresponding to their $D$), which leads to the LABF result and was left uninvestigated in their work.

Another difference between HvdAW and our present paper is that the former assumes $f$ is elliptical, while we do not make this assumption. The elliptical density assumption in HvdAW is specifically used for their test construction based on the pseudo-Mahalanobis distance-based rank statistic. However, it is not necessary for the limit experiment approach to proceed. Therefore, it is reasonable to conjecture that their limit experiment has the same structure without this distribution shape restriction. In the following, we will translate their results by replacing their $W_{\epsilon}$ and $W_{\phi}$ with our $\boldsymbol{W}_{\boldsymbol{\varepsilon}}$ and $\boldsymbol{W}_{\boldsymbol{\ell}_f}$, respectively, which are both $p$-dimensional Brownian motions that play the role of the limits of the partial-sum processes of the innovations and the score functions.

To save space, we will not present a full exposition of the limit experiment, nor provide a rigorous proof for it (which could be done similarly to Proposition (ref)). Instead, we only list the limiting central sequences with respect to $({\rm vec}\,\mathbf{C})^\prime$ $({\rm vec}\,\mathbf{G})^\prime$, $({\rm vec}\,\mathbf{A})^\prime$ and $({\rm vec}\,\mathbf{B})^\prime$, respectively, as follows (borrowed from hallin2016semiparametric):

align*[align* omitted — 672 chars of source]

where $\mathbf{D} = \mathbf{D}_{\boldsymbol{\Gamma},\boldsymbol{\Pi}} := \boldsymbol{\beta}_\perp\left(\boldsymbol{\alpha}_\perp^\prime(\mathbf{I}_p - \sum_{j=1}^{k-1}\boldsymbol{\Gamma}_j)\boldsymbol{\beta}_\perp\right)^{-1}\boldsymbol{\alpha}_\perp^\prime$, and $\boldsymbol{W}_1$ and $\boldsymbol{W}_2$ are Brownian motions defined in the same space of $(\boldsymbol{W}_{\boldsymbol{\varepsilon}},\boldsymbol{W}_{\boldsymbol{b}},\boldsymbol{W}_{\boldsymbol{\ell}_f})$.

According to the covariance analysis of HvdAW (as discussed above their Lemma A.1), it can be shown that $\boldsymbol{W}_2$ (referred to as $W_{\Delta X\otimes\phi}$ in HvdAW) is uncorrelated with $\boldsymbol{W}_{\boldsymbol{\varepsilon}}$ and $\boldsymbol{W}_{\boldsymbol{\ell}_f}$ when their $\mu$ is zero. This can be explained by observing that the score associated with $\boldsymbol{W}_2$ is constructed by multiplying $\boldsymbol{\ell}_f(\Delta\mathbf{x}_t)$ with the lagged increments, $\Delta\mathbf{x}_{t-1},\dots,\Delta\mathbf{x}_{t-k}$, which is independent of any function of $\Delta\mathbf{x}_t$, and has mean zero under the null hypothesis.\footnote{It is worth noting that this property is also shared by our univariate counterpart, zhou2019semiparametrically, where $\boldsymbol{W}_2$ here corresponds to the $W_\Gamma$ there, and the latter is uncorrelated with the univariate counterparts of $\boldsymbol{W}_{\boldsymbol{\varepsilon}}$ and $\boldsymbol{W}_{\boldsymbol{\ell}_f}$ as described in (3.8) of their paper.} As a result, $\mathit{\Delta}_\mathbf{C}$ is independent of $\mathit{\Delta}_\mathbf{G}$, which implies that the inference for $\mathbf{C}$ is adaptive with respect to $\mathbf{G}$. In other words, the parameters $\boldsymbol{\Gamma}$ that control the serial correlation in the error term can be treated “as if” they are known, and one can simply replace them with their consistent estimates. Unreported simulation results with ARMA errors can confirm our conjecture (these results are available upon request).

The adaptivity result also applies to $\boldsymbol{\alpha}$ and $\boldsymbol{\beta}$, as it can be easily shown that $\mathit{\Delta}_\mathbf{C}$ is independent of both $\mathit{\Delta}_\mathbf{A}$ (due to $\boldsymbol{\beta}^\prime\boldsymbol{\beta}_\perp = \boldsymbol{0}$) and $\mathit{\Delta}_\mathbf{B}$ (due to of $\boldsymbol{\alpha}^\prime\boldsymbol{\alpha}_\perp = \boldsymbol{0}$). Hence, by the same reasoning, one can simply substitute these parameters with their consistent estimates to achieve feasible inference for $\mathbf{C}$ without sacrificing asymptotic efficiency. In summary, for this general reduced null hypothesis case, one can follow these steps: (i) estimate $\alpha$ and $\beta$ based on the original data $\mathbf{y}_t$, obtain an estimator $\hat{\alpha}\perp$ for $\alpha\perp$ using the orthogonality of $\alpha$ and $\alpha_\perp$ based on the estimates from step (i), and (iii) apply the our proposed semiparametric optimal tests to the transformed data $\hat{\alpha}_\perp\mathbf{y}_t$. This procedure is similar to the one outlined in boswijk2015improved except for step (iii), except that we use our proposed tests in step (iii).