EconBase
← Back to paper

Nonparametric Tests of Conditional Independence for Time Series

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

42,807 characters · 8 sections · 40 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonparametric Tests of Conditional Independence for Time Series

abstractWe propose consistent nonparametric tests of conditional independence for time series data. Our methods are motivated from the difference between joint conditional cumulative distribution function (CDF) and the product of conditional CDFs. The difference is transformed into a proper conditional moment restriction (CMR), which forms the basis for our testing procedure. Our test statistics are then constructed using the integrated moment restrictions that are equivalent to the CMR. We establish the asymptotic behavior of the test statistics under the null, the alternative, and the sequence of local alternatives converging to conditional independence at the parametric rate. Our tests are implemented with the assistance of a multiplier bootstrap. Monte Carlo simulations are conducted to evaluate the finite sample performance of the proposed tests. We apply our tests to examine the predictability of equity risk premium using variance risk premium for different horizons and find that there exist various degrees of nonlinear predictability at mid-run and long-run horizons.

Keywords: Conditional CDFs; empirical processes; multiplier bootstrap; nonparametric regression; time series.

JEL Classifications: C12; C14; C15.

\thispagestyle{empty}

\pagenumbering{arabic} \setcounter{page}{1} \setcounter{footnote}{0}

Introduction

A variable $Y$ is said to be conditionally independent of $Z$ given $X$ if and only if the conditional density of $Y$ given $Z$ and $X$ equals to the conditional density of $Y$ given $X$, that is, $Z$ does not carry any information about $Y$ once $X$ is given. Following David (1979), we write $Y\bot Z|X$ to denote that $Y$ is independent of $Z$ given $X$. The hypotheses $Y\bot Z|X$ is related to the hypothesis that $Y$ is independent of $Z$, i.e. $Y\bot Z$ (unconditional independence), and the conditional mean independence (CMI), $\mathrm{E} (Y|Z,X) = \mathrm{E} (Y|X)$.

The assumption of conditional independence plays an important role and is a widely imposed one in both statistical and econometric literature. For example, Markov property of a time series process, Granger non-causality, the assumption of missing at random (MAR) and exogeneity all can be formulated as a conditional independence restriction, see hong2017testing for motivating examples of testing the conditional independence hypothesis in economics and econometrics. However, in contrast to the many tests of unconditional independence or CMI proposed in the context of independent and identically distributed (i.i.d.) data, by far, not many tests are available for testing conditional independence assumption with time series data. Nonparametric tests for unconditional independence between random variables and/or vectors, and nonparametric tests for serial independence are abundant, e.g. a nonparametric test of Cram\'{e}r-von Mises type first introduced by hoeffding1948non, the empirical distribution function-based tests of blum1961distribution, skaug1993nonparametric for testing independence of raw data and delgado2000nonparametric or ghoudi2001nonparametric for testing serial independence of time series or regression errors, the empirical characteristic function-based test of csorgHo1985testing, kernel smoothing-based tests like rosenblatt1975quadratic, robinson1991blup, and hong2005asymptotic, and tests based on measures of association and dependence between random variables and/or vectors such as bakirov2006multivariate, szekely2007measuring or diks2007nonparametric. On the other hand, among the available tests for conditional independence, the majority is designed for i.i.d. data, e.g. linton2014testing develop a non-pivotal nonparametric test based on a generalization of the empirical distribution function, song2009testing employs the Rosenblatt transformation to obtain a distribution-free test for a different type of conditional independence, huang2010testing proposes a test based on the maximal nonlinear conditional correlation, and huang2016flexible develop an integrated conditional moment test. Tests suitable for time series data include su2007consistent, su2008nonparametric, su2012conditional, su2014testing, bouezmarni2012nonparametric and wang2018characteristic, all of which are based on kernel smoothing. The exception is su2012conditional, who provide nonparametric tests for conditional independence using local polynomial quantile regression.

In this paper we aim to further fill the gap of the literature and propose consistent nonparametric tests based on an empirical process approach for testing conditional independence which are applicable to time series data. Our approach exploits a proper conditional moment restriction and is in the same spirit with delgado2001significance, which partially circumvents the “curse of dimensionality” problem. Besides, in comparison with the existing tests based on smoothing methods, our new tests are able to detect local alternatives converging to the null at a parametric rate.

We introduce some notations for testing the conditional independence hypothesis in a time series framework. Let $X_t$, $Y_t$ and $Z_t$ be three generic random vectors with dimensions $d_x$, $d_y$ and $d_z$, respectively. The null hypothesis of interest is that $Y_t$ is independent of $Z_t$ conditional on $X_t$, i.e.,

equation*[equation* omitted — 59 chars of source]

for all $(x,y,z)\in\mathbb{R}^{d_x+d_y+d_z}$, where $F_{Y,Z|X}$, $F_{Y|X}$ and $F_{Z|X}$ denote the conditional cumulative distribution functions (CDFs).

The rest of the paper is as follows. In Section 2 we examine the testing problem and provide the test statistics. Section 3 establishes the asymptotic null distributions. Section 4 studies the consistency property of the test and the asymptotic local power of the test under local alternatives. A bootstrap procedure to implement the tests is proposed and formally justified in Section 5. In Section 6, we study the finite sample performance of our tests by means of Monte Carlo simulations. Section 7 presents an empirical example of using variance risk premium to predict equity risk premium. Finally, Section 8 concludes the paper. All proofs are collected in the online Appendix.

The testing procedure

Consider a $\mathbb{R}^{d_x+d_y+d_z}$-valued strictly stationary ergodic time series process $\{(X_t,Y_{t},Z_{t})\}$ defined on the probability space $(\Omega ,\mathcal{F},\mathbb{P})$, which satisfies the Markov's property

equation[equation omitted — 199 chars of source]

where $\mathcal{A}_{t-1}:=\sigma(\{X_{s},Y_{s-1},Z_{s-1}\}_{s=-\infty}^{t})$ with $\sigma(\cdot)$ the smallest sigma algebra,

equation*[equation* omitted — 84 chars of source]

with $0<p<\infty$ an integer, and “$^\top$” denotes transpose. That is, the only relevant information for explaining $(Y_{t},Z_t)$ are the first $m$ lags of $\left(X_{t+1}, Y_{t},Z_{t}\right)$. Note that $W_t$ is a $d_w\times 1$ vector with $d_w=p(d_x+d_y+d_z)$.

We propose a nonparametric test for the hypothesis that $Y_t$ and $Z_t$ are independent given $W_t$, i.e.,

equation[equation omitted — 159 chars of source]

The alternative hypothesis $\text{H}_1$ is the negation of $\text{H}_0$ in (ref). The alternative hypothesis $\text{H}_1$ consists of a broad class of conditional dependence between $Y_t$ and $Z_t$ given $W_t$. That is, conditional on $W_t$, $Y_t$ could depend on $Z_t$ through mean, variance, skewness, kurtosis, or even higher moments. It is possible to have situations where the dependence between $Y_t$ and $Z_t$ in lower moments (e.g. mean or variance) does not exist, but it does exist in higher moments (e.g. skewness or kurtosis). See Section (ref) for some data generating processes under $\text{H}_1$.

It is important to emphasize that the new formulation of testing conditional independence assumption in (ref) is attractive, since it partly circumvents the problem of “curse of dimensionality” by conditioning on only $W_t$ in all three conditional CDFs. wang2018characteristic has also exploited (ref) to propose a test based on the conditional characteristic functions. On the other hand, existing tests are mainly based on testing $F_{Y|W,Z}(y|W_t,Z_t)=F_{Y|W}(y|W_t)$ a.s., which requires estimation of conditional CDFs given both $W_t$ and $Z_t$, see e.g. su2007consistent, su2008nonparametric and bouezmarni2012nonparametric to name only a few.

Note that $F_{Y,Z|W}(y,z|W_t) = \mathrm{E}[1(Y_t\leq y)1(Z_t\leq z)|W_t]$ and $F_{Y|W}(y|W_t) = \mathrm{E}[1(Y_t\leq y)|W_t]$, with $1(\cdot)$ the indicator function. Thus, $\text{H}_0$ in (ref) can be tested by

equation[equation omitted — 174 chars of source]

The problem of testing (ref) is in fact equivalent to testing a conditional moment restriction (CMR) in (ref), which has the advantage of conditioning on only $W_t$. An important feature of the CMR stated in (ref) is that it involves all values of $(y,z)$. Since it has to hold for all $(y,z)$, we have an infinite number of CMRs to be tested.

Fortunately, stinchcombe1998consistent give us a method to convert conditional to unconditional moment restriction in a convenient way: testing (ref) is further equivalent to testing the following infinite number of unconditional moment restrictions,

equation[equation omitted — 226 chars of source]

where $\mathcal{W} \subseteq \mathbb{R}^{d_{w}}$ is a proper chosen set with typical choices $d_{w} = d_x + d_y + d_z$ or $d_{w} = d_x + d_y + d_z + 1$, and $\varphi$ is a generically comprehensively revealing (GCR) or comprehensively revealing (CR) function. Examples of GCR functions include: (1) $\varphi(W_t, w) = \exp (\mathrm{i} w^{\top} W_t)$; (2) $\varphi(W_t, \gamma) = \sin (w^{\top} W_t)$, and examples of CR functions include: (3) $\varphi(W_t, \gamma) = 1 (W_t \leq w)$; (4) $\varphi (W_t, w) = 1 (\beta^{\top} W_t \leq \alpha)$ with $w = (\alpha, \beta)^{\top}$. It is worthy to note that when $\varphi$ is GCR, the deviations from the null hypothesis can be detected by essentially any choice of $w \in \mathcal{W}$, where $\mathcal{W}$ can be chosen as any small compact set with non-empty interior, whereas CR functions may require the set $\mathcal{W}$ to be the whole Euclidean space to ensure the consistency of the associate test. More discussion can be seen in stinchcombe1998consistent and su2012conditional. Hence in the following, we will always assume $\mathcal{W}$ is a bounded space in $\mathbb{R}^{d_w}$. In addition, to avoid the random denominator problem, we propose to test the following modified version of (ref),

equation[equation omitted — 235 chars of source]

where $f_W(W_t)$ is the density of $W_t$. As in delgado2001significance, the density-weighted formulation helps to avoid conveniently the random denominator in the subsequent nonparametric estimation.

Given a sample $\{(W^\top_t,Y^\top_t,Z^\top_t)^\top\}_{t=1}^n$ of size $n$, if $F_{Z|W}(z|W_t)$ and $f_W(W_t)$ were observable, test statistics could be constructed using the (infeasible) empirical process

align*[align* omitted — 134 chars of source]

Then under $\text{H}_0$, $S^{0}_{n}(\cdot,\cdot,\cdot) \rightsquigarrow S^{0}_\infty(\cdot,\cdot,\cdot)$, where $S^{0}_\infty(\cdot,\cdot,\cdot)$ is a zero mean Gaussian process with covariance kernel $\mathrm{E}\left[S^{0}_\infty(w,y,z)S^{0}_\infty(w',y',z')\right]$.

As $F_{Z|W}(z|W_t)$ and $f_W(W_t)$ are unobservable, testing procedures based on $S^0_{n}(w,y,z)$ are not feasible. In this paper, we replace them with $\widehat{F}_{Z|W}(z|W_t)$ and $\widehat{f}_W(W_t)$, where

equation*[equation* omitted — 156 chars of source]
equation*[equation* omitted — 115 chars of source]

with $K(\cdot)$ and $h:=h_n\in\mathbb{R}^+$ the kernel function and bandwidth, respectively. Define the feasible empirical process

align*[align* omitted — 145 chars of source]

which is algebraically equivalent to

equation*[equation* omitted — 176 chars of source]

The above expression is a variant of $U$-processes considered by delgado2001significance in an i.i.d. context. Note that the limiting distribution of $S_{n}(w,y,z)$ will be different from that of $S^0_{n}(w,y,z)$ due to the estimation of $F_{Z|W}(z|W_t)$.

Test statistics are constructed based on suitable continuous functionals of $S_n(w,y,z)$. A test statistic in the spirit of the Cram\'{e}r-von Mises type is

align[align omitted — 108 chars of source]

where $F_n(w,y,z)=n^{-1}\sum_{t=1}^n1(W_t\leq x)1(Y_t\leq y)1(Z_t\leq z)$ is the empirical distribution function of $(X_t,Y_t,Z_t)$. Henceforth, an unspecified integral denotes integration over the whole space. The Kolmogorov-Smirnov type test statistic basing on the sup-norm is

equation[equation omitted — 146 chars of source]

In practice, one can compute $KS_n$ by simply taking the maximum over the observations, i.e., $\widetilde{KS}_{n}=\max_{1\leq t\leq n}\vert S_n(W_t,Y_t,Z_t)\vert$.

Under $\text{H}_0$, test statistics $CvM_n$ and $KS_n$ converge in distribution, while they diverge to infinity under $\text{H}_1$. We reject the null hypothesis of conditional independence whenever they exceed certain “large” values. Since the asymptotic null distributions of $CvM_n$ and $KS_n$ depend on the data generating process in a complicated way, their critical values are not readily available. To circumvent this problem, we propose a bootstrap procedure to obtain the critical values of our tests in Section (ref).

Asymptotic null distributions

In this section, we will establish the asymptotic null distributions of our test statistics $CvM_n$ and $KS_n$. We need to impose the following assumptions, which are attached in Appendix (ref).

Let $\phi_t(y)=1(Y_{t}\leq y)-F_{Y|W}(y|W_t)$, $\epsilon_t(z)=1(Z_{t}\leq z)-F_{Z|W}(z|W_t)$, and $e_t(w,y,z) = \varphi (W_t, w) \phi_t(y)\epsilon_t(z)f_W(W_t)$. The following theorem shows that $S_n(\cdot,\cdot,\cdot)$ converges weakly to a Gaussian process under the null.

theoremSuppose (ref) and Assumption A1 - A7 in Appendix (ref) hold. Then under the null \begin{equation*} S_n (\cdot, \cdot, \cdot) \rightsquigarrow S_{\infty} (\cdot, \cdot, \cdot), \end{equation*} where $S_{\infty} (\cdot, \cdot, \cdot)$ is a zero mean Gaussian process with covariance kernel $\mathrm{E} \big[ S_{\infty} (w, y, z) S_{\infty} (w', y', z') \big]$. Specially, if $(X_t^{\top}, Y_t^{\top} , Z_t^{\top})$ is IID sample, we have \begin{align*} \mathrm{E}\big[S_{\infty}(w,y,z), S_{\infty}(w',y',z')\big]& = \mathrm{E} \big[e_1(w,y,z)e_1(w',y',z')\big]\\ & = \mathrm{E}[\varphi(W_t, w) \varphi(W_t, w') \psi(y,y';W_1)\rho(z,z';W_1)f^2_W(W_1)], \end{align*} where $\psi(y,y';W)=F_{Y|W}(y\wedge y'|W)-F_{Y|W}(y|W)F_{Y|W}(y'|W)$ and $\rho(z,z';W)=F_{Z|W}(z\wedge z'|W)-F_{Z|W}(z|W)F_{Z|W}(z'|W)$.
remarkThe term $F_{Y|W}(y|W_t)$ in the definition of $e_t(w,y,z)$ reflects the cost paid for replacing $F_{Z|W}(z|W_t)$ with $\widehat F_{Z|W}(z|W_t)$ in the infeasible process $S_n^0(w,y,z)$, a phenomenon known as the “parameter estimation error”, leading to a covariance kernel different from that of $S_n^0(w,y,z)$.

The asymptotic null distributions of test statistics $CvM_n$ and $KS_n$ are given in the corollary below.

corollarySuppose (ref) and Assumption A1 - A7 in Appendix (ref) hold. Then under the null, \begin{align*} CvM_n \rightsquigarrow CvM_{\infty}:=\int_{\mathcal{W} \times \mathbb{R}^{d_y} \times \mathbb{R}^{d_z}} S^2_{\infty}(w,y,z)\,dF_{W,Y,Z}(w,y,z), \end{align*} \begin{equation*} KS_n \rightsquigarrow KS_{\infty}:=\sup_{(w,y,z) \in \mathcal{W} \times \mathbb{R}^{d_y} \times \mathbb{R}^{d_z}}\left\vert S_{\infty}(w,y,z)\right\vert, \end{equation*} where $S_{\infty}(\cdot,\cdot,\cdot)$ is the Gaussian process defined in Theorem (ref).

Consistency and asymptotic local power

We investigate the consistency and asymptotic local power properties of test statistics $CvM_n$ and $KS_n$ based on $S_n(w,y,z)$ under $\text{H}_1$ and under a sequence of local alternatives converging to $\text{H}_0$ at a parametric rate $n^{-1/2}$.

The asymptotic behavior of $S_n(w,y,z)$ under $\text{H}_1$ is stated in the next theorem.

theoremSuppose (ref) and Assumptions A1 - A7 in Appendix (ref) hold.. Then under the alternative, for each $(w,y,z) \in \mathcal{W} \times \mathbb{R}^{d_y} \times \mathbb{R}^{d_z}$, \begin{equation*} n^{-1/2}S_n(w,y,z) \xrightarrow{P} \mathrm{E}\big[\varphi(W_t, w)\big(F_{Y,Z|W}(y,z|W_t)-F_{Y|W}(y|W_t)F_{Z|W}(z|W_t)\big)f_W(W_t)\big]. \end{equation*}

Since $\mathrm{E}\big[\varphi(W_t, w)\big(F_{Y,Z|W}(y,z|W_t)-F_{Y|W}(y|W_t)F_{Z|W}(z|W_t)\big)f_W(W_t)\big]$ in a set with a positive Lebesgue measure under $\text{H}_1$, test statistics $CvM_n$ and $KS_n$ will diverge to infinity and have asymptotic power one against $\text{H}_1$.

To investigate the asymptotic local power properties of the tests, we introduce the following sequence of local alternatives,

equation[equation omitted — 181 chars of source]

where $\Delta(\cdot,\cdot,\cdot)$ is a non-constant measurable function, satisfying $\Delta(W_t,y,z)\neq 0$ a.s. for some $(y,z)$ and $\Delta(W_t,\infty,z)=\Delta(W_t,y,\infty)=\Delta(W_t,-\infty,z)=\Delta(W_t,y,-\infty)=0$ a.s. to deliver a valid conditional CDF $F_{Y,Z|W}(y,z|W_t)$ in (ref). The type of local alternatives in (ref) is widely used in studying the asymptotic local power properties of tests based on empirical processes.

In (ref), the term $n^{-1/2}\Delta(W_t,y,z)$ characterizes the departure of the conditional joint CDF from the product of conditional marginal CDFs. In particular, $\Delta(W_t,y,z)$ specifies the direction of the departure, while $n^{-1/2}$ specifies the rate at which the difference between $F_{Y,Z|W}(y,z|W_t)$ and $F_{Y|W}(y|W_t)F_{Z|W}(z|W_t)$ shrinks to zero. Note that $n^{-1/2}$ is the fastest possible rate found in testing conditional independence.

To derive the local power result, we need the extra assumption i.e. Assumption A8 in Appendix (ref). The next theorem states the asymptotic behavior of $S_n(w,y,z)$ under $\text{H}_{1n}$.

theoremSuppose Assumption A1 - A8 in Appendix (ref) hold. Then under the sequence of local alternatives in (ref), \begin{equation*} S_n(\cdot,\cdot,\cdot)\rightsquigarrow S_{\infty}(\cdot,\cdot,\cdot)+G(\cdot,\cdot,\cdot), \end{equation*} where $S_{\infty}(\cdot,\cdot,\cdot)$ is the Gaussian process defined in Theorem (ref).

Provided that the shift function $G(w,y,z):=\int^{w}_{-\infty}\Delta(\bar{w},y,z)f^2_W(\bar{w})\,d\bar{w}\neq 0$ in a set with a positive Lebesgue measure under $\text{H}_{1n}$, test statistics $CvM_n$ and $KS_n$ based on $S_n(x,y,z)$ will have non-trivial local power against $\text{H}_{1n}$ converging to $\text{H}_0$ at a parametric rate $n^{-1/2}$, the best rate known in testing conditional independence.

remarkFor local power analysis, tests based on empirical processes can achieve a parametric rate $n^{-1/2}$, a rate much faster than those obtained from smoothing-based nonparametric tests. For example, su2008nonparametric test only has power against local alternatives at a rate $n^{-1/2}h^{-d/4}$ with $d=d_w+d_y+d_z$, while tests proposed by su2007consistent and bouezmarni2012nonparametric have power against local alternatives at a rate $n^{-1/2}h^{-(d_w+d_z)/4}$. wang2018characteristic test is the only one that has a better rate, which can detect local alternatives at a rate $n^{-1/2}h^{-d_w/4}$, though still slower than $n^{-1/2}$. Nonetheless, tests based on local smoothing are able to detect high frequency local alternatives considered by rosenblatt1975quadratic, while our tests may not detect such type of local alternatives.

The limiting distributions of $CvM_n$ and $KS_n$ under $\text{H}_{1n}$ is stated in the next result, which is a direct consequence of the continuous mapping theorem and Theorem (ref).

corollarySuppose (ref) and Assumption A1 - A8 in Appendix (ref) hold. Then under the sequence of local alternatives in (ref), \begin{align*} CvM_n \rightsquigarrow \int_{\mathcal{W} \times \mathbb{R}^{d_y} \times \mathbb{R}^{d_z}}\big(S_{\infty}(w,y,z)+G(w,y,z)\big)^2\,dF_{W,Y,Z}(w,y,z), \end{align*} \begin{equation*} KS_n \rightsquigarrow \sup_{(w,y,z) \in \mathcal{W} \times \mathbb{R}^{d_y} \times \mathbb{R}^{d_z}}\left\vert S_{\infty}(w,y,z)+G(w,y,z)\right\vert, \end{equation*} where $S_\infty(\cdot,\cdot,\cdot)$ and $G(\cdot,\cdot,\cdot)$ are defined in Theorem (ref).

Corollary (ref) implies that under $\text{H}_{1n}$ in (ref), the limiting distributions of $CvM_n$ and $KS_n$ are no longer the same as in Corollary 1 and they shift in a non-trivial way. Consequently, our tests have non-trivial power against the sequence of local alternatives in (ref) converging to the null at a parametric rate.

Bootstrap

Since the asymptotic null distributions of $CvM_n$ and $KS_n$ depend on the underlying data generating process due to the complicated covariance kernel of the limiting process $S_\infty(w,y,z)$ in Theorem (ref), it is difficult to tabulate the critical values for our tests. We propose a bootstrap procedure, which is in the spirit of the multiplier bootstrap suggested by delgado2001significance. This bootstrap takes full advantage of the asymptotic theory in Theorem 1, and is also easy to implement as it does not have to compute new nonparametric estimates at each bootstrap replication.

Let $\widehat\phi_t(y)=1(Y_{t}\leq y)-\widehat F_{Y|W}(y|W_t)$, $\widehat\epsilon_t(z)=1(Z_{t}\leq z)-\widehat F_{Z|W}(z|W_t)$, and $\widehat e_t(w,y,z) = \varphi(W_t, w) \widehat\phi_t(y)\widehat\epsilon_t(z)\widehat f_W(W_t)$, where $\widehat{F}_{Y|W}(y|W_t)$ is defined with $1(Y_s\leq y)$ replacing $1(Z_s\leq z)$ in $\widehat{F}_{Z|W}(z|W_t)$. The bootstrap empirical process of $S_{n}(w,y,z)$ is given by

align*[align* omitted — 103 chars of source]

where $\{v_t\}_{t=1}^n$ is a sequence of i.i.d. random variables with zero mean, unit variance, bounded support, and is independent of $\{(W^\top_t,Y^\top_t,Z^\top_t)^\top\}_{t=1}^n$. One popular choice due to mammen1993bootstrap is the i.i.d. Bernoulli variates with probability masses given by

equation*[equation* omitted — 200 chars of source]

See also delgado2001significance and escanciano2006generalized for applications of this choice.

In the next theorem the asymptotic validity the bootstrap procedure is justified formally. We show that the bootstrapped process $S^\ast_{n}(\cdot,\cdot,\cdot)$ converges weakly to the Gaussian process $S_{\infty}(\cdot,\cdot,\cdot)$ in Theorem (ref). As a result, we have convergence in distributions of the bootstrapped test statistics $CvM^\ast_n$ and $KS^\ast_n$, which are simply constructed by replacing $S_{n}(w,y,z)$ with $S^\ast_{n}(w,y,z)$ in (ref) and (ref), respectively.

theoremSuppose (ref) and Assumption A1 - A8 in Appendix (ref) hold. Then, under the null, under the alternative, or under the sequence of local alternatives in (ref), \begin{equation*} S_{n}^{\ast}(\cdot,\cdot,\cdot)\underset{\ast}{\overset{P}{\rightarrow}}S_{\infty}(\cdot,\cdot,\cdot), \end{equation*} where $S_{\infty}(\cdot,\cdot,\cdot)$ is the Gaussian process defined in Theorem (ref), and $\underset{\ast}{\overset{P}{\rightarrow}}$ denotes weak convergence in probability under the bootstrap law, i.e., conditional on $\{(W^\top_t,Y^\top_t,Z^\top_t)^\top\}_{t=1}^n$. In addition, $CvM_n^\ast\underset{\ast}{\rightsquigarrow}CvM_\infty$ and $KS_n^\ast\underset{\ast}{\rightsquigarrow}KS_\infty$ with $CvM_\infty$ and $KS_\infty$ defined in Corollary (ref).

Theorem (ref) implies that the limiting behavior of $S_n(w,y,z)$ can be approximated by that of $S^{\ast}_n(w,y,z)$. Thus, the bootstrap assisted tests $CvM_n$ and $KS_n$ have a correct asymptotic level, are consistent against $\text{H}_1$, and are able to detect local alternatives (ref) converging to the null at a parametric rate. In practice, we can obtain the critical values of $CvM_n$ (and similarly for $KS_n$) as accurately as desired by the following algorithm:\\ Step 1. Compute $S_n(W_t,Y_t,Z_t)$, and get $CvM_n$\\ Step 2. Generate $\{v_t\}_{t=1}^n$ independently, compute $S_{n}^{\ast}(W_t,Y_t,Z_t)$, and get $CvM_n^{\ast}$.\\ Step 3. Repeat Step 2 $B$ times to have $\{CvM_{n,b}^{\ast}\}_{b=1}^B$, and compute its empirical $(1-\alpha)$-th sample quantile $CvM_{n}^{\ast\alpha}$ or bootstrapped $p$-value $p_n^{\ast }=B^{-1}\sum_{b=1}^{B}1\left( CvM_{n,b}^{\ast }\geq CvM_n\right)$. Rejects $\text{H}_0$ at the significance level $\alpha$ if $CvM_n>CvM_{n}^{\ast\alpha}$ or if $p_n^{\ast}<\alpha$.

remarkThe above bootstrap assisted procedure applies when (ref) holds. If (ref) is violated, our proposed tests may suffer size distortion and power loss. However, under the general dependence structure in the data, it is possible to extend the block bootstrap (e.g., buhlmann1994blockwise) to our context, the finite sample performance of which will be investigated through simulations in the next section.

Monte Carlo simulations

We carry out a set of Monte Carlo simulations to examine the finite sample performance of the proposed test statistics $CvM_n$ and $KS_n$. To examine the size performance, we consider the following four data generating processes (DGPs):

align*[align* omitted — 421 chars of source]

To examine the power performance, the following seven DGPs are considered:

align*[align* omitted — 685 chars of source]

Here and below, $\varepsilon_{1,t}$, $\varepsilon_{2,t}$ and $\varepsilon_{3,t}$ are i.i.d. $N(0,1)$ and mutually independent, and $Z_t$ in (S2)-(S3) and (P1)-(P6) follows an AR(1) model: $Z_t =0.5Z_{t-1}+\varepsilon_{2,t}$. These DGPs cover a wide range of both linear and nonlinear time series processes. For (S1), $(Y_t,Z_t,X_t)$ is simply an i.i.d. sequence; we test $Y_t\bot Z_t|X_t$ such that $W_t=X_t$. From (S2) to (P7), we test $Y_t\bot Z_{t-1}|Y_{t-1}$ such that $W_t=Y_{t-1}$; that is, we are interested in checking whether $Z_{t-1}$ has any predictive ability in explaining $Y_t$ (in mean, variance, or higher moments) after controlling the first lag of $Y_t$. Note that (S4) and (P7) do not satisfy (ref).

For each DGP, we first generate $n+500$ observations and then discard the first 500 observations to minimize effects of initial values. The number of Monte Carlo simulations is $2000$ and the bootstrap critical values are obtained from $B=1000$ bootstrap replications. Four sample sizes, $n=100$, $200$, $400$ and $800$, are considered. We only report results for the nominal level of 5% and results for other levels are available upon request. We choose the standard normal density $K(x)=(2\pi)^{-1/2}\exp{(-x^2/2)}$ as our kernel. Bandwidth of the form $h=c n^{-1/3.5}$ is used and results with $c=0.5$, $1.0$ and $1.5$ are reported to check the sensitivity of our tests to different bandwidths. How to choose $h$ to maximize the performance in our testing framework is beyond the scope of this paper and is left for future research. In addition to the fixed bandwidth $h$, in the supplemental appendix we provide more simulation results with a data-driven one $h=1.06\cdot Std(X_t)\cdot n^{-1/3.5}$.

Table (ref) and (ref) report respectively the empirical rejection rates of $CvM_n$ and $KS_n$ under (S1)-(P7) using critical values obtained with the multiplier bootstrap method proposed in Section (ref). Both $CvM_n$ and $KS_n$ have acceptable empirical sizes in the cases of moderate sample sizes for (S1)-(S3), with $KS_n$ having more accurate sizes than $CvM_n$. When (ref) fails in (S4), $CvM_n$ exhibits a severe size distortion for $n=100$ and $n=200$, while $KS_n$ has less distortion; when sample size increases to 800, the two tests show improved sizes especially for $c=0.5$. For the empirical power performance, when sample size is as small as $n=100$, both $CvM_n$ and $KS_n$ are not very powerful against (P2)-(P7), with (P7) the most difficult DGP to detect. They gain power rapidly as sample size increases. Note that in (P7) (ref) fails, which might be the reason for low power in small samples. In summary, our simulation results indicate that $KS_n$ preserves sizes better than $CvM_n$, while $CvM_n$ is more powerful than $KS_n$; in addition, whether (ref) holds will affect the performance. The simulation results also show that the empirical sizes and powers of $CvM_n$ and $KS_n$ are somehow sensitive to bandwidth choices in very small samples; a general pattern is that larger $c$ tends to deliver a higher power but produces a bigger size distortion. When $n=800$, the overall performance is satisfactory.

center[center omitted — 147 chars of source]

To investigate if we have improved performance when we take into account the general dependence structure in the data, we also study the following block multiplier bootstrap:

equation*[equation* omitted — 121 chars of source]

where $\{\zeta_t\}_{t=1}^{n-L+1}$ are i.i.d. $N(0,L^{-1})$ and independent of the sample $\{(W_t^\top,Y_t^\top,Z_t^\top)^\top\}_{t=1}^n$, with $L:=L_n$ the block length diverging to infinity at a slower rate than $n$ as $n\to\infty$. Note that when $L=1$, $S_n^{\ast block}(w,y,z)$ reduces to the multiplier bootstrap introduced in Section (ref), the validity of which replies on the restriction (ref).

Table (ref) and (ref) report respectively the empirical rejection rates of $CvM^{block}_n$ and $KS^{block}_n$ based on $S_n^{\ast block}(w,y,z)$. It is expected that $CvM^{block}_n$ and $KS^{block}_n$ should work when (ref) does not hold and consequently the dependence structure in the data should not be ignored. In implementing the block bootstrap, we have used $L=\lfloor an^{1/4} \rfloor$ with $a=1$, $2$ and $4$ to check the sensitivity of $CvM^{block}_n$ and $KS^{block}_n$ with respect to different block lengths. To save space, we only report the results for bandwidth $h=cn^{-1/3.5}$ with $c=1.0$. The results in Table (ref) and (ref) are based on 1000 simulations and 200 bootstrap replications. It is observed that $CvM^{block}_n$ and $KS^{block}_n$ have less distorted empirical sizes especially for (S4) than $CvM_n$ and $KS_n$ do. At the same time, $CvM^{block}_n$ and $KS^{block}_n$ are not as powerful as $CvM_n$ and $KS_n$ even for (P7). Like before, $KS^{block}_n$ preserves sizes better than $CvM^{block}_n$, while $CvM^{block}_n$ is more powerful than $KS^{block}_n$. Lastly, for given $h$, the tests also depend on the choice of $L$; smaller $L$ tends to deliver a higher power but produces oversized tests.

center[center omitted — 149 chars of source]

An empirical study

We examine whether there exists nonlinear predictability of equity risk premium using variance risk premium. The variance risk premium is defined as the difference between the risk-neutral and objective expectations of realized variance, where the risk-neutral expectation of variance is measured as the end-of-month Volatility Index-squared de-annualized and the realized variance is the sum of squared 5-minute log returns of the S&P 500 index over the month.

There is a rich literature on the predictive power of variance risk premium for the aggregate stock market returns, bond returns or exchange rate returns. For example, Bollerslev et al. (2009) first discover that variance risk premium is able to explain a nontrivial fraction of the time series variation in post 1990 aggregate stock market returns, with high (low) premia predicting high (low) future returns; Wang et al. (2013) find the empirical evidence suggesting that the firm-level variance risk premium has a prominent explanatory power for credit spreads in the presence of market- and firm-level control variables; by defining a “global” variance risk premium, Bollerslev et al. (2013) uncover stronger predictability of aggregate stock market returns using variance risk premium across countries; while Della Corte et al. (2013) investigate the predictive information content in foreign exchange volatility risk premia for exchange rate returns and find that a portfolio that sells currencies with high insurance costs and buys currencies with low insurance costs generates sizeable out-of-sample returns and Sharpe ratios.

We use monthly aggregate S&P 500 composite index over the period January 1996 to September 2008. Our empirical analysis is based on the logarithmic return on the S&P 500 in excess of the 3-month T-bill rate. Let $RP_{t+\tau}$ be the risk premium $\tau$ months ahead and $VRP_{t}$ be the variance risk premium at time $t$. In this empirical study, we take $\tau=$1, 3, 6, and 9 months. We shall examine if the variance risk premium $VRP_t$ explains in a linear or nonlinear way the risk premium $RP_{t+\tau}$ given the information $RP_{t}$, which is equivalent to stating whether $VRP_t$ Granger causes $RP_t$ by setting the lag order to $\tau$. To test for the presence of (nonlinear) predictability of $VRP_t$, we consider to test

equation*[equation* omitted — 122 chars of source]

against $\text{H}_1: \mathrm{P}\left\{F(RP_{t+\tau},VRP_{t}|RP_{t})=F(RP_{t+\tau}|RP_{t})F(VRP_t|RP_{t})\right\}<1$. That is, for a given horizon $\tau$, we test the conditional independence of $RP_{t+\tau}$ and $VRP_t$ given $RP_{t}$, i.e. $RP_{t+\tau}\bot VRP_{t}|RP_{t}$.

For the purpose of comparison, we also perform the popular linear causality analysis in the literature. To this end, we consider the following linear regression model:

equation[equation omitted — 112 chars of source]

The hypothesis of interest is that VRP does not Granger cause RP for $\tau$ months ahead in a linear way, i.e. testing the null hypothesis $\text{H}_0:\alpha_{\tau}=0$ against the alternative hypothesis $\text{H}_1:\alpha_{\tau}\neq 0$. To test $\text{H}_0$, standard $t$-statistic given by $t_{\hat{\alpha}_{\tau}}=\hat{\alpha}_{\tau}/\hat{\sigma}_{\hat{\alpha}_{\tau}}$ will be calculated, where $\hat{\alpha}_{\tau}$ is the least squares estimator of $\alpha_{\tau}$ and $\hat{\sigma}_{\hat{\alpha}_{\tau}}$ is the estimator of its standard error $\sigma_{\hat{\alpha}_{\tau}}$. Moreover, to avoid the impact of possible dependence in the residual terms $\hat{\varepsilon}_{t+\tau}$ on our inference, $\hat{\sigma}_{\hat{\alpha}_{\tau}}$ is calculated using the commonly used heteroscedasticity autocorrelation consistent (HAC) robust variance estimator suggested by Newey and West (1987).

Table (ref) reports the testing results for Granger causality (e.g. nonlinear predictability) from variance risk premium to risk premium, at four different horizons, using our proposed tests $CvM_n$ and $KS_n$ as well as the linear test. To check the robustness of our empirical findings, we also include the block multiplier tests $CvM_n^{block}$ and $KS_n^{block}$, where the block length is set to be $L=\lfloor an^{1/4} \rfloor$ with $a=2$. Results for $a=1$ and $4$ are similar and hence are omitted. The implementation of all bootstrap based tests is as demonstrated in the Monte Carlo simulations part with the number of bootstrap replications $B=10,000$. We have obtained encouraging findings from Table (ref) that are relevant to both empirical and theoretical studies of variance risk premium as a suitable predictor for risk premium. Specifically, results from linear regression based test in (ref) clearly fail to reject the null hypothesis of no linear predictability for one to six months (short-run) horizons; they only show weak evidence of linear predictability until at the nine months (long-run) horizon at the 5% significance level. On the other hand, using our tests, we find convincing evidence that risk premium can be well predicted using variance risk premium at both mid-run and long-run horizons. Testing results are not sensitive to the bandwidth choices for six and nine months horizon, nor does the block type tests. For the one month horizon, we may conclude that variance risk premium has no predictive power, while the findings for three months horizon tend to indicate the existence of predictability. Overall, we find that there is a very high degree of predictability at horizons more than one-month which could be attributed to the nonlinear predictive effect. Our empirical evidence also indicates that caution is needed when interpreting results based on the linear regression.

center[center omitted — 134 chars of source]

Conclusion

This paper proposes new consistent nonparametric tests of conditional independence for time series data based on the empirical process method. The asymptotic properties of the proposed tests under the null, the alternative, and the sequence of local alternatives are investigated. To implement the test in practice, a multiplier bootstrap procedure is suggested and its asymptotic validity is formally justified. The test can be applied to testing for conditional independence in a wide variety of nonparametric models. Using the proposed test, we also study whether there exists some nonlinear predictability of equity risk premium using variance risk premium.