The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
42,804 characters
Nonparametric Tests of Conditional Independence for Time Series
\title{Nonparametric Tests of Conditional Independence for Time Series}
\author{Xiaojun Song \ \thanks{Department of Business Statistics and Econometrics, Guanghua School of
Management and Center for Statistical Science, Peking University, Beijing, 100871, China. E-mail: [email removed]. Financial support from the National Natural Science Foundation of China (Grant No. 71532001) is acknowledged.}\\Peking University
\and
Haoyu Wei \thanks{Guanghua School of Management, Peking University, Beijing, China. Email: [email removed]. }\\Peking University
}
\maketitle
\begin{abstract}
We propose consistent nonparametric tests of conditional independence for time series data. Our methods are motivated from the difference between joint conditional cumulative distribution function (CDF) and the product of conditional CDFs. The difference is transformed into a proper conditional moment restriction (CMR), which forms the basis for our testing procedure. Our test statistics are then constructed using the integrated moment restrictions that are equivalent to the CMR. We establish the asymptotic behavior of the test statistics under the null, the alternative, and the sequence of local alternatives converging to conditional independence at the parametric rate. Our tests are implemented with the assistance of a multiplier bootstrap. Monte Carlo simulations are conducted to evaluate the finite sample performance of the proposed tests. We apply our tests to examine the predictability of equity risk premium using variance risk premium for different horizons and find that there exist various degrees of nonlinear predictability at mid-run and long-run horizons.
\end{abstract}
\noindent \textbf{Keywords}: Conditional CDFs; empirical processes; multiplier bootstrap; nonparametric regression; time series.
\noindent \textbf{JEL Classifications:} C12; C14; C15.
\thispagestyle{empty}
\newpage \pagenumbering{arabic} \setcounter{page}{1} \setcounter{footnote}{0}
\pagebreak
\section{Introduction}
A variable $Y$ is said to be conditionally independent of $Z$ given $X$ if and only if the conditional density of $Y$ given $Z$ and $X$ equals to the conditional density of $Y$ given $X$, that is, $Z$ does not carry any information about $Y$ once $X$ is given. Following David (1979), we write $Y\bot Z|X$ to denote that $Y$ is independent of $Z$ given $X$. The hypotheses $Y\bot Z|X$ is related to the hypothesis that $Y$ is independent of $Z$, i.e. $Y\bot Z$ (unconditional independence), and the conditional mean independence (CMI), $\mathrm{E} (Y|Z,X) = \mathrm{E} (Y|X)$.
The assumption of conditional independence plays an important role and is a widely imposed one in both statistical and econometric literature. For example, Markov property of a time series process, Granger non-causality, the assumption of missing at random (MAR) and exogeneity all can be formulated as a conditional independence restriction, see \cite{hong2017testing} for motivating examples of testing the conditional independence hypothesis in economics and econometrics. However, in contrast to the many tests of unconditional independence or CMI proposed in the context of independent and identically distributed (i.i.d.) data, by far, not many tests are available for testing conditional independence assumption with time series data. Nonparametric tests for unconditional independence between random variables and/or vectors, and nonparametric tests for serial independence are abundant, e.g. a nonparametric test of Cram\'{e}r-von Mises type first introduced by \cite{hoeffding1948non}, the empirical distribution function-based tests of \cite{blum1961distribution}, \cite{skaug1993nonparametric} for testing independence of raw data and \cite{delgado2000nonparametric} or \cite{ghoudi2001nonparametric} for testing serial independence of time series or regression errors, the empirical characteristic function-based test of \cite{csorgHo1985testing}, kernel smoothing-based tests like \cite{rosenblatt1975quadratic}, \cite{robinson1991blup}, and \cite{hong2005asymptotic}, and tests based on measures of association and dependence between random variables and/or vectors such as \cite{bakirov2006multivariate}, \cite{szekely2007measuring} or \cite{diks2007nonparametric}. On the other hand, among the available tests for conditional independence, the majority is designed for i.i.d. data, e.g. \cite{linton2014testing} develop a non-pivotal nonparametric test based on a generalization of the empirical distribution function, \cite{song2009testing} employs the Rosenblatt transformation to obtain a distribution-free test for a different type of conditional independence, \cite{huang2010testing} proposes a test based on the maximal nonlinear conditional correlation, and \cite{huang2016flexible} develop an integrated conditional moment test. Tests suitable for time series data include \cite{su2007consistent, su2008nonparametric, su2012conditional, su2014testing}, \cite{bouezmarni2012nonparametric} and \cite{wang2018characteristic}, all of which are based on kernel smoothing. The exception is \cite{su2012conditional}, who provide nonparametric tests for conditional independence using local polynomial quantile regression.
In this paper we aim to further fill the gap of the literature and propose consistent nonparametric tests based on an empirical process approach for testing conditional independence which are applicable to time series data. Our approach exploits a proper conditional moment restriction and is in the same spirit with \cite{delgado2001significance}, which partially circumvents the ``curse of dimensionality'' problem. Besides, in comparison with the existing tests based on smoothing methods, our new tests are able to detect local alternatives converging to the null at a parametric rate.
We introduce some notations for testing the conditional independence hypothesis in a time series framework. Let $X_t$, $Y_t$ and $Z_t$ be three generic random vectors with dimensions $d_x$, $d_y$ and $d_z$, respectively. The null hypothesis of interest is that $Y_t$ is independent of $Z_t$ conditional on $X_t$, i.e.,
\begin{equation*}
F_{Y,Z|X}(y,z|x)=F_{Y|X}(y|x)F_{Z|X}(z|x),
\end{equation*}
for all $(x,y,z)\in\mathbb{R}^{d_x+d_y+d_z}$, where $F_{Y,Z|X}$, $F_{Y|X}$ and $F_{Z|X}$ denote the conditional cumulative distribution functions (CDFs).
The rest of the paper is as follows. In Section 2 we examine the testing problem and provide the test statistics. Section 3 establishes the asymptotic null distributions. Section 4 studies the consistency property of the test and the asymptotic local power of the test under local alternatives. A bootstrap procedure to implement the tests is proposed and formally justified in Section 5. In Section 6, we study the finite sample performance of our tests by means of Monte Carlo simulations. Section 7 presents an empirical example of using variance risk premium to predict equity risk premium. Finally, Section 8 concludes the paper. All proofs are collected in the online Appendix.
\section{The testing procedure}\label{section2}
Consider a $\mathbb{R}^{d_x+d_y+d_z}$-valued strictly stationary ergodic time series process $\{(X_t,Y_{t},Z_{t})\}$ defined on the probability space $(\Omega ,\mathcal{F},\mathbb{P})$, which satisfies the Markov's property
\begin{equation}
\mathrm{P} (Y_{t}\leq y,Z_t\leq z\vert\mathcal{A}_{t-1}) = \mathrm{P} (Y_{t}\leq y,Z_t\leq z\vert W_{t})\,\,\,\text{a.s.}\,\,\,\forall(y,z)\in \mathbb{R}^{d_y+d_z}, \label{markov}
\end{equation}
where $\mathcal{A}_{t-1}:=\sigma(\{X_{s},Y_{s-1},Z_{s-1}\}_{s=-\infty}^{t})$ with $\sigma(\cdot)$ the smallest sigma algebra,
\begin{equation*}
W_{t}=\{(X^\top_s,Y^\top_{s-1},Z^\top_{s-1})^\top\}_{s=t-p+1}^t
\end{equation*}
with $0<p<\infty$ an integer, and ``$^\top$'' denotes transpose. That is, the only relevant information for explaining $(Y_{t},Z_t)$ are the first $m$ lags of $\left(X_{t+1}, Y_{t},Z_{t}\right)$. Note that $W_t$ is a $d_w\times 1$ vector with $d_w=p(d_x+d_y+d_z)$.
We propose a nonparametric test for the hypothesis that $Y_t$ and $Z_t$ are independent given $W_t$, i.e.,
\begin{equation}
\text{H}_0: F_{Y,Z|W}(y,z|W_t)=F_{Y|W}(y|W_t)F_{Z|W}(z|W_t)\quad\text{a.s.} \quad \forall \, (y,z)\in\mathbb{R}^{d_y+d_z}.\label{eq:null1}
\end{equation}
The alternative hypothesis $\text{H}_1$ is the negation of $\text{H}_0$ in \eqref{eq:null1}. The alternative hypothesis $\text{H}_1$ consists of a broad class of conditional dependence between $Y_t$ and $Z_t$ given $W_t$. That is, conditional on $W_t$, $Y_t$ could depend on $Z_t$ through mean, variance, skewness, kurtosis, or even higher moments. It is possible to have situations where the dependence between $Y_t$ and $Z_t$ in lower moments (e.g. mean or variance) does not exist, but it does exist in higher moments (e.g. skewness or kurtosis). See Section \ref{mc} for some data generating processes under $\text{H}_1$.
It is important to emphasize that the new formulation of testing conditional independence assumption in \eqref{eq:null1} is attractive, since it partly circumvents the problem of ``curse of dimensionality'' by conditioning on only $W_t$ in all three conditional CDFs. \cite{wang2018characteristic} has also exploited \eqref{eq:null1} to propose a test based on the conditional characteristic functions. On the other hand, existing tests are mainly based on testing $F_{Y|W,Z}(y|W_t,Z_t)=F_{Y|W}(y|W_t)$ a.s., which requires estimation of conditional CDFs given both $W_t$ and $Z_t$, see e.g. \cite{su2007consistent, su2008nonparametric} and \cite{bouezmarni2012nonparametric} to name only a few.
Note that $F_{Y,Z|W}(y,z|W_t) = \mathrm{E}[1(Y_t\leq y)1(Z_t\leq z)|W_t]$ and $F_{Y|W}(y|W_t) = \mathrm{E}[1(Y_t\leq y)|W_t]$, with $1(\cdot)$ the indicator function. Thus, $\text{H}_0$ in \eqref{eq:null1} can be tested by
\begin{equation}
\text{H}_0: \mathrm{E} [1(Y_t\leq y)(1(Z_t\leq z)-F_{Z|W}(z|W_t))|W_t]=0 \quad \text{a.s.} \quad \forall \, (y,z)\in\mathbb{R}^{d_y+d_z}. \label{eq:null3}
\end{equation}
The problem of testing \eqref{eq:null1} is in fact equivalent to testing a conditional moment restriction (CMR) in \eqref{eq:null3}, which has the advantage of conditioning on only $W_t$. An important feature of the CMR stated in \eqref{eq:null3} is that it involves all values of $(y,z)$. Since it has to hold for all $(y,z)$, we have an infinite number of CMRs to be tested.
Fortunately, \cite{stinchcombe1998consistent} give us a method to convert conditional to unconditional moment restriction in a convenient way: testing \eqref{eq:null3} is further equivalent to testing the following infinite number of unconditional moment restrictions,
\begin{equation}\label{eq:null4}
\text{H}_0: \mathrm{E} \big[ \varphi(W_t, w) 1 (Y_t \leq y) \big( 1 (Z_t \leq z) - F_{Z | W} (z | W_t)\big) \big] = 0, \quad \forall \, (w, y, z) \in \mathcal{W} \times \mathbb{R}^{d_y+d_z},
\end{equation}
where $\mathcal{W} \subseteq \mathbb{R}^{d_{w}}$ is a proper chosen set with typical choices $d_{w} = d_x + d_y + d_z$ or $d_{w} = d_x + d_y + d_z + 1$, and $\varphi$
is a generically comprehensively revealing (GCR) or comprehensively revealing (CR) function. Examples of GCR functions include: (1) $\varphi(W_t, w) = \exp (\mathrm{i} w^{\top} W_t)$; (2) $\varphi(W_t, \gamma) = \sin (w^{\top} W_t)$, and examples of CR functions include: (3) $\varphi(W_t, \gamma) = 1 (W_t \leq w)$; (4) $\varphi (W_t, w) = 1 (\beta^{\top} W_t \leq \alpha)$ with $w = (\alpha, \beta)^{\top}$. It is worthy to note that when $\varphi$ is GCR, the deviations from the null hypothesis can be detected by essentially any choice of $w \in \mathcal{W}$, where $\mathcal{W}$ can be chosen as any small compact set with non-empty interior, whereas CR functions may require the set $\mathcal{W}$ to be the whole Euclidean space to ensure the consistency of the associate test. More discussion can be seen in \cite{stinchcombe1998consistent} and \cite{su2012conditional}. Hence in the following, we will always assume $\mathcal{W}$ is a bounded space in $\mathbb{R}^{d_w}$. In addition, to avoid the random denominator problem, we propose to test the following modified version of \eqref{eq:null4},
\begin{equation}\label{eq:null5}
\text{H}_0: \mathrm{E} \big[ \varphi(W_t, w) 1 (Y_t \leq y) \big( 1 (Z_t \leq z) - F_{Z | W} (z | W_t)\big) f_W (W_t)\big] = 0, \quad \forall \, (w, y, z) \in \mathcal{W} \times \mathbb{R}^{d_y+d_z},
\end{equation}
where $f_W(W_t)$ is the density of $W_t$. As in \cite{delgado2001significance}, the density-weighted formulation helps to avoid conveniently the random denominator in the subsequent nonparametric estimation.
Given a sample $\{(W^\top_t,Y^\top_t,Z^\top_t)^\top\}_{t=1}^n$ of size $n$, if $F_{Z|W}(z|W_t)$ and $f_W(W_t)$ were observable, test statistics could be constructed using the (infeasible) empirical process
\begin{align*}
S^{0}_{n}(w, y, z)=\frac{1}{\sqrt{n}}\sum_{t=1}^n \varphi(W_t, w) 1(Y_t\leq y)(1(Z_t\leq z)-F_{Z|W}(z|W_t))f_W(W_t).
\end{align*}
Then under $\text{H}_0$, $S^{0}_{n}(\cdot,\cdot,\cdot) \rightsquigarrow S^{0}_\infty(\cdot,\cdot,\cdot)$, where $S^{0}_\infty(\cdot,\cdot,\cdot)$ is a zero mean Gaussian process with covariance kernel $\mathrm{E}\left[S^{0}_\infty(w,y,z)S^{0}_\infty(w',y',z')\right]$.
As $F_{Z|W}(z|W_t)$ and $f_W(W_t)$ are unobservable, testing procedures based on $S^0_{n}(w,y,z)$ are not feasible. In this paper, we replace them with $\widehat{F}_{Z|W}(z|W_t)$ and $\widehat{f}_W(W_t)$, where
\begin{equation*}
\widehat{F}_{Z|W}(z|W_t)=\frac{\frac{1}{(n-1)h^{d_w}}\sum_{s=1, s\neq t}^n1(Z_s\leq z)K\left(\frac{W_t-W_s}{h}\right)}{\hat{f}_W(W_t)},
\end{equation*}
\begin{equation*}
\widehat{f}_W(W_t)=\frac{1}{(n-1)h^{d_w}}\sum_{s=1, s\neq t}^nK\left(\frac{W_t-W_s}{h}\right),
\end{equation*}
with $K(\cdot)$ and $h:=h_n\in\mathbb{R}^+$ the kernel function and bandwidth, respectively. Define the feasible empirical process
\begin{align*}
S_{n}(w,y,z)=\frac{1}{\sqrt n}\sum_{t=1}^n \varphi(W_t, w) 1(Y_t\leq y)(1(Z_t\leq z)-\widehat F_{Z|W}(z|W_t))\widehat f_W(W_t),
\end{align*}
which is algebraically equivalent to
\begin{equation*}
\frac{1}{\sqrt n(n-1)h^{d_w}}\sum_{t=1}^n\sum_{s=1, s\neq t}^nK\left(\frac{W_{t}-W_{s}}{h}\right) \varphi (W_t, w) 1(Y_t\leq y)(1(Z_t\leq z)-1(Z_s\leq z)).
\end{equation*}
The above expression is a variant of $U$-processes considered by \cite{delgado2001significance} in an i.i.d. context. Note that the limiting distribution of $S_{n}(w,y,z)$ will be different from that of $S^0_{n}(w,y,z)$ due to the estimation of $F_{Z|W}(z|W_t)$.
Test statistics are constructed based on suitable continuous functionals of $S_n(w,y,z)$. A test statistic in the spirit of the Cram\'{e}r-von Mises type is
\begin{align}
CvM_n=\int S^2_n(w,y,z)\,dF_{n}(w,y,z)=n^{-1}\sum_{t=1}^{n}S_n^2(W_t,Y_t,Z_t), \label{cvmn}
\end{align}
where $F_n(w,y,z)=n^{-1}\sum_{t=1}^n1(W_t\leq x)1(Y_t\leq y)1(Z_t\leq z)$ is the empirical distribution function of $(X_t,Y_t,Z_t)$. Henceforth, an unspecified integral denotes integration over the whole space. The Kolmogorov-Smirnov type test statistic basing on the sup-norm is
\begin{equation}
KS_n=\sup_{(w,y,z) \in \mathcal{W} \times \mathbb{R}^{d_y} \times \mathbb{R}^{d_z}}\left\vert S_{n}(w,y,z)\right\vert. \label{ksn}
\end{equation}
In practice, one can compute $KS_n$ by simply taking the maximum over the observations, i.e., $\widetilde{KS}_{n}=\max_{1\leq t\leq n}\vert S_n(W_t,Y_t,Z_t)\vert$.
Under $\text{H}_0$, test statistics $CvM_n$ and $KS_n$ converge in distribution, while they diverge to infinity under $\text{H}_1$. We reject the null hypothesis of conditional independence whenever they exceed certain ``large'' values. Since the asymptotic null distributions of $CvM_n$ and $KS_n$ depend on the data generating process in a complicated way, their critical values are not readily available. To circumvent this problem, we propose a bootstrap procedure to obtain the critical values of our tests in Section \ref{boot}.
\section{Asymptotic null distributions}\label{section3}
In this section, we will establish the asymptotic null distributions of our test statistics $CvM_n$ and $KS_n$. We need to impose the following assumptions, which are attached in Appendix \ref{assumptions}.
Let $\phi_t(y)=1(Y_{t}\leq y)-F_{Y|W}(y|W_t)$, $\epsilon_t(z)=1(Z_{t}\leq z)-F_{Z|W}(z|W_t)$, and $e_t(w,y,z) = \varphi (W_t, w) \phi_t(y)\epsilon_t(z)f_W(W_t)$. The following theorem shows that $S_n(\cdot,\cdot,\cdot)$ converges weakly to a Gaussian process under the null.
\begin{theorem}\label{thm1}
Suppose \eqref{markov} and Assumption A1 - A7 in Appendix \ref{assumptions} hold. Then under the null
\begin{equation*}
S_n (\cdot, \cdot, \cdot) \rightsquigarrow S_{\infty} (\cdot, \cdot, \cdot),
\end{equation*}
where $S_{\infty} (\cdot, \cdot, \cdot)$ is a zero mean Gaussian process with covariance kernel $\mathrm{E} \big[ S_{\infty} (w, y, z) S_{\infty} (w', y', z') \big]$. Specially, if $(X_t^{\top}, Y_t^{\top} , Z_t^{\top})$ is IID sample, we have
\begin{align*}
\mathrm{E}\big[S_{\infty}(w,y,z), S_{\infty}(w',y',z')\big]& = \mathrm{E} \big[e_1(w,y,z)e_1(w',y',z')\big]\\
& = \mathrm{E}[\varphi(W_t, w) \varphi(W_t, w') \psi(y,y';W_1)\rho(z,z';W_1)f^2_W(W_1)],
\end{align*}
where $\psi(y,y';W)=F_{Y|W}(y\wedge y'|W)-F_{Y|W}(y|W)F_{Y|W}(y'|W)$ and $\rho(z,z';W)=F_{Z|W}(z\wedge z'|W)-F_{Z|W}(z|W)F_{Z|W}(z'|W)$.
\end{theorem}
\begin{remark}
The term $F_{Y|W}(y|W_t)$ in the definition of $e_t(w,y,z)$ reflects the cost paid for replacing $F_{Z|W}(z|W_t)$ with $\widehat F_{Z|W}(z|W_t)$ in the infeasible process $S_n^0(w,y,z)$, a phenomenon known as the ``parameter estimation error'', leading to a covariance kernel different from that of $S_n^0(w,y,z)$.
\end{remark}
The asymptotic null distributions of test statistics $CvM_n$ and $KS_n$ are given in the corollary below.
\begin{corollary}\label{cor1}
Suppose \eqref{markov} and Assumption A1 - A7 in Appendix \ref{assumptions} hold. Then under the null,
\begin{align*}
CvM_n \rightsquigarrow CvM_{\infty}:=\int_{\mathcal{W} \times \mathbb{R}^{d_y} \times \mathbb{R}^{d_z}} S^2_{\infty}(w,y,z)\,dF_{W,Y,Z}(w,y,z),
\end{align*}
\begin{equation*}
KS_n \rightsquigarrow KS_{\infty}:=\sup_{(w,y,z) \in \mathcal{W} \times \mathbb{R}^{d_y} \times \mathbb{R}^{d_z}}\left\vert S_{\infty}(w,y,z)\right\vert,
\end{equation*}
where $S_{\infty}(\cdot,\cdot,\cdot)$ is the Gaussian process defined in Theorem \ref{thm1}.
\end{corollary}
\section{Consistency and asymptotic local power}
We investigate the consistency and asymptotic local power properties of test statistics $CvM_n$ and $KS_n$ based on $S_n(w,y,z)$ under $\text{H}_1$ and under a sequence of local alternatives converging to $\text{H}_0$ at a parametric rate $n^{-1/2}$.
The asymptotic behavior of $S_n(w,y,z)$ under $\text{H}_1$ is stated in the next theorem.
\begin{theorem}\label{thm2}
Suppose \eqref{markov} and Assumptions A1 - A7 in Appendix \ref{assumptions} hold.. Then under the alternative, for each $(w,y,z) \in \mathcal{W} \times \mathbb{R}^{d_y} \times \mathbb{R}^{d_z}$,
\begin{equation*}
n^{-1/2}S_n(w,y,z) \xrightarrow{P} \mathrm{E}\big[\varphi(W_t, w)\big(F_{Y,Z|W}(y,z|W_t)-F_{Y|W}(y|W_t)F_{Z|W}(z|W_t)\big)f_W(W_t)\big].
\end{equation*}
\end{theorem}
Since $\mathrm{E}\big[\varphi(W_t, w)\big(F_{Y,Z|W}(y,z|W_t)-F_{Y|W}(y|W_t)F_{Z|W}(z|W_t)\big)f_W(W_t)\big]$ in a set with a positive Lebesgue measure under $\text{H}_1$, test statistics $CvM_n$ and $KS_n$ will diverge to infinity and have asymptotic power one against $\text{H}_1$.
To investigate the asymptotic local power properties of the tests, we introduce the following sequence of local alternatives,
\begin{equation}
\text{H}_{1n}: F_{Y,Z|W}(y,z|W_t)=F_{Y|W}(y|W_t)F_{Z|W}(z|W_t)+n^{-1/2}\Delta(W_t,y,z)\,\,\,\text{a.s.}\,\,\,\forall(y,z)\in\mathbb{R}^{d_y+d_z}, \label{alternative}
\end{equation}
where $\Delta(\cdot,\cdot,\cdot)$ is a non-constant measurable function, satisfying $\Delta(W_t,y,z)\neq 0$ a.s. for some $(y,z)$ and $\Delta(W_t,\infty,z)=\Delta(W_t,y,\infty)=\Delta(W_t,-\infty,z)=\Delta(W_t,y,-\infty)=0$ a.s. to deliver a valid conditional CDF $F_{Y,Z|W}(y,z|W_t)$ in \eqref{alternative}. The type of local alternatives in \eqref{alternative} is widely used in studying the asymptotic local power properties of tests based on empirical processes.
In \eqref{alternative}, the term $n^{-1/2}\Delta(W_t,y,z)$ characterizes the departure of the conditional joint CDF from the product of conditional marginal CDFs. In particular, $\Delta(W_t,y,z)$ specifies the direction of the departure, while $n^{-1/2}$ specifies the rate at which the difference between $F_{Y,Z|W}(y,z|W_t)$ and $F_{Y|W}(y|W_t)F_{Z|W}(z|W_t)$ shrinks to zero. Note that $n^{-1/2}$ is the fastest possible rate found in testing conditional independence.
To derive the local power result, we need the extra assumption i.e. Assumption A8 in Appendix \ref{assumptions}. The next theorem states the asymptotic behavior of $S_n(w,y,z)$ under $\text{H}_{1n}$.
\begin{theorem}\label{thm3}
Suppose Assumption A1 - A8 in Appendix \ref{assumptions} hold. Then under the sequence of local alternatives in \eqref{alternative},
\begin{equation*}
S_n(\cdot,\cdot,\cdot)\rightsquigarrow S_{\infty}(\cdot,\cdot,\cdot)+G(\cdot,\cdot,\cdot),
\end{equation*}
where $S_{\infty}(\cdot,\cdot,\cdot)$ is the Gaussian process defined in Theorem \ref{thm1}.
\end{theorem}
Provided that the shift function $G(w,y,z):=\int^{w}_{-\infty}\Delta(\bar{w},y,z)f^2_W(\bar{w})\,d\bar{w}\neq 0$ in a set with a positive Lebesgue measure under $\text{H}_{1n}$, test statistics $CvM_n$ and $KS_n$ based on $S_n(x,y,z)$ will have non-trivial local power against $\text{H}_{1n}$ converging to $\text{H}_0$ at a parametric rate $n^{-1/2}$, the best rate known in testing conditional independence.
\begin{remark}
For local power analysis, tests based on empirical processes can achieve a parametric rate $n^{-1/2}$, a rate much faster than those obtained from smoothing-based nonparametric tests. For example, \cite{su2008nonparametric} test only has power against local alternatives at a rate $n^{-1/2}h^{-d/4}$ with $d=d_w+d_y+d_z$, while tests proposed by \cite{su2007consistent} and \cite{bouezmarni2012nonparametric} have power against local alternatives at a rate $n^{-1/2}h^{-(d_w+d_z)/4}$. \cite{wang2018characteristic} test is the only one that has a better rate, which can detect local alternatives at a rate $n^{-1/2}h^{-d_w/4}$, though still slower than $n^{-1/2}$. Nonetheless, tests based on local smoothing are able to detect high frequency local alternatives considered by \cite{rosenblatt1975quadratic}, while our tests may not detect such type of local alternatives.
\end{remark}
The limiting distributions of $CvM_n$ and $KS_n$ under $\text{H}_{1n}$ is stated in the next result, which is a direct consequence of the continuous mapping theorem and Theorem \ref{thm3}.
\begin{corollary}\label{cor2}
Suppose \eqref{markov} and Assumption A1 - A8 in Appendix \ref{assumptions} hold. Then under the sequence of local alternatives in \eqref{alternative},
\begin{align*}
CvM_n \rightsquigarrow \int_{\mathcal{W} \times \mathbb{R}^{d_y} \times \mathbb{R}^{d_z}}\big(S_{\infty}(w,y,z)+G(w,y,z)\big)^2\,dF_{W,Y,Z}(w,y,z),
\end{align*}
\begin{equation*}
KS_n \rightsquigarrow \sup_{(w,y,z) \in \mathcal{W} \times \mathbb{R}^{d_y} \times \mathbb{R}^{d_z}}\left\vert S_{\infty}(w,y,z)+G(w,y,z)\right\vert,
\end{equation*}
where $S_\infty(\cdot,\cdot,\cdot)$ and $G(\cdot,\cdot,\cdot)$ are defined in Theorem \ref{thm3}.
\end{corollary}
Corollary \ref{cor2} implies that under $\text{H}_{1n}$ in \eqref{alternative}, the limiting distributions of $CvM_n$ and $KS_n$ are no longer the same as in Corollary 1 and they shift in a non-trivial way. Consequently, our tests have non-trivial power against the sequence of local alternatives in \eqref{alternative} converging to the null at a parametric rate.
\section{Bootstrap}\label{boot}
Since the asymptotic null distributions of $CvM_n$ and $KS_n$ depend on the underlying data generating process due to the complicated covariance kernel of the limiting process $S_\infty(w,y,z)$ in Theorem \ref{thm1}, it is difficult to tabulate the critical values for our tests. We propose a bootstrap procedure, which is in the spirit of the multiplier bootstrap suggested by \cite{delgado2001significance}. This bootstrap takes full advantage of the asymptotic theory in Theorem 1, and is also easy to implement as it does not have to compute new nonparametric estimates at each bootstrap replication.
Let $\widehat\phi_t(y)=1(Y_{t}\leq y)-\widehat F_{Y|W}(y|W_t)$, $\widehat\epsilon_t(z)=1(Z_{t}\leq z)-\widehat F_{Z|W}(z|W_t)$, and $\widehat e_t(w,y,z) = \varphi(W_t, w) \widehat\phi_t(y)\widehat\epsilon_t(z)\widehat f_W(W_t)$, where $\widehat{F}_{Y|W}(y|W_t)$ is defined with $1(Y_s\leq y)$ replacing $1(Z_s\leq z)$ in $\widehat{F}_{Z|W}(z|W_t)$. The bootstrap empirical process of $S_{n}(w,y,z)$ is given by
\begin{align*}
S_{n}^{\ast}(w,y,z)=\frac{1}{\sqrt n}\sum_{t=1}^n\widehat e_t(w,y,z)v_t,\label{rstar}
\end{align*}
where $\{v_t\}_{t=1}^n$ is a sequence of i.i.d. random variables with zero mean, unit variance, bounded support, and is independent of $\{(W^\top_t,Y^\top_t,Z^\top_t)^\top\}_{t=1}^n$. One popular choice due to \cite{mammen1993bootstrap} is the i.i.d. Bernoulli variates with probability masses given by
\begin{equation*}
\mathrm{P} \left(v_t=\frac{1-\sqrt{5}}{2}\right)=\frac{1+\sqrt{5}}{2\sqrt{5}}\,\,\,\text{and}\,\,\, \mathrm{P} \left(v_t=\frac{1+\sqrt{5}}{2}\right)=\frac{-1+\sqrt{5}}{2\sqrt{5}}.
\end{equation*}
See also \cite{delgado2001significance} and \cite{escanciano2006generalized} for applications of this choice.
In the next theorem the asymptotic validity the bootstrap procedure is justified formally. We show that the bootstrapped process $S^\ast_{n}(\cdot,\cdot,\cdot)$ converges weakly to the Gaussian process $S_{\infty}(\cdot,\cdot,\cdot)$ in Theorem \ref{thm1}. As a result, we have convergence in distributions of the bootstrapped test statistics $CvM^\ast_n$ and $KS^\ast_n$, which are simply constructed by replacing $S_{n}(w,y,z)$ with $S^\ast_{n}(w,y,z)$ in \eqref{cvmn} and \eqref{ksn}, respectively.
\begin{theorem}\label{thm4}
Suppose \eqref{markov} and Assumption A1 - A8 in Appendix \ref{assumptions} hold. Then, under the null, under the alternative, or under the sequence of local alternatives in \eqref{alternative},
\begin{equation*}
S_{n}^{\ast}(\cdot,\cdot,\cdot)\underset{\ast}{\overset{P}{\rightarrow}}S_{\infty}(\cdot,\cdot,\cdot),
\end{equation*}
where $S_{\infty}(\cdot,\cdot,\cdot)$ is the Gaussian process defined in Theorem \ref{thm1}, and $\underset{\ast}{\overset{P}{\rightarrow}}$ denotes weak convergence in probability under the bootstrap law, i.e., conditional on $\{(W^\top_t,Y^\top_t,Z^\top_t)^\top\}_{t=1}^n$. In addition, $CvM_n^\ast\underset{\ast}{\rightsquigarrow}CvM_\infty$ and $KS_n^\ast\underset{\ast}{\rightsquigarrow}KS_\infty$ with $CvM_\infty$ and $KS_\infty$ defined in Corollary \ref{cor1}.
\end{theorem}
Theorem \ref{thm4} implies that the limiting behavior of $S_n(w,y,z)$ can be approximated by that of $S^{\ast}_n(w,y,z)$. Thus, the bootstrap assisted tests $CvM_n$ and $KS_n$ have a correct asymptotic level, are consistent against $\text{H}_1$, and are able to detect local alternatives \eqref{alternative} converging to the null at a parametric rate. In practice, we can obtain the critical values of $CvM_n$ (and similarly for $KS_n$) as accurately as desired by the following algorithm:\\
\indent \textbf{Step 1.} Compute $S_n(W_t,Y_t,Z_t)$, and get $CvM_n$\\
\indent \textbf{Step 2.} Generate $\{v_t\}_{t=1}^n$ independently, compute $S_{n}^{\ast}(W_t,Y_t,Z_t)$, and get $CvM_n^{\ast}$.\\
\indent \textbf{Step 3.} Repeat \textbf{Step 2} $B$ times to have $\{CvM_{n,b}^{\ast}\}_{b=1}^B$, and compute its empirical $(1-\alpha)$-th sample quantile $CvM_{n}^{\ast\alpha}$ or bootstrapped $p$-value $p_n^{\ast }=B^{-1}\sum_{b=1}^{B}1\left( CvM_{n,b}^{\ast }\geq CvM_n\right)$. Rejects $\text{H}_0$ at the significance level $\alpha$ if $CvM_n>CvM_{n}^{\ast\alpha}$ or if $p_n^{\ast}<\alpha$.
\begin{remark}
The above bootstrap assisted procedure applies when \eqref{markov} holds. If \eqref{markov} is violated, our proposed tests may suffer size distortion and power loss. However, under the general dependence structure in the data, it is possible to extend the block bootstrap (e.g., \cite{buhlmann1994blockwise}) to our context, the finite sample performance of which will be investigated through simulations in the next section.
\end{remark}
\section{Monte Carlo simulations}\label{mc}
We carry out a set of Monte Carlo simulations to examine the finite sample performance of the proposed test statistics $CvM_n$ and $KS_n$. To examine the size performance, we consider the following four data generating processes (DGPs):
\begin{align*}
\text{(S1):} \quad Y_t &=\varepsilon_{1,t}, Z_t=\varepsilon_{2,t}, X_t=\varepsilon_{3,t}.\\
\text{(S2):} \quad Y_t &=0.5Y_{t-1}+\varepsilon_{1,t}.\\
\text{(S3):} \quad Y_t &=0.5Y_{t-1}\exp{(-0.5Y^2_{t-1})}+\varepsilon_{1,t}.\\
\text{(S4):} \quad Y_t &=\sqrt{h_{1,t}}\varepsilon_{1,t}, h_{1,t}=0.01+0.9h_{1,t-1}+0.05Y_{t-1}^2,\\
Z_t &=\sqrt{h_{2,t}}\varepsilon_{2,t}, h_{2,t}=0.01+0.9h_{2,t-1}+0.05Z_{t-1}^2.
\end{align*}
To examine the power performance, the following seven DGPs are considered:
\begin{align*}
\text{(P1):} \quad Y_t &=0.5Y_{t-1}+0.5Z_{t-1}+\varepsilon_{1,t}.\\
\text{(P2):} \quad Y_t &=0.5Y_{t-1}+0.5Z_{t-1}^2+\varepsilon_{1,t}.\\
\text{(P3):} \quad Y_t &=0.5Y_{t-1}Z_{t-1}+\varepsilon_{1,t}.\\
\text{(P4):} \quad Y_t &=0.3+0.2\log(h_t)+\sqrt{h_{t}}\varepsilon_{1,t}, h_{t}=0.01+0.5Y_{t-1}^2+0.3Z_{t-1}^2.\\
\text{(P5):} \quad Y_t &=0.5Y_{t-1}+0.5Z_{t-1}\varepsilon_{1,t}.\\
\text{(P6):} \quad Y_t &=\sqrt{h_{t}}\varepsilon_{1,t}, h_{t}=0.01+0.5Y_{t-1}^2+0.25Z_{t-1}^2.\\
\text{(P7):} \quad Y_t &=\sqrt{h_{1,t}}\varepsilon_{1,t}, h_{1,t}=0.01+0.1h_{1,t-1}+0.4Y_{t-1}^2+0.5Z_{t-1}^2.\\
Z_t &=\sqrt{h_{2,t}}\varepsilon_{2,t}, h_{2,t}=0.01+0.9h_{2,t-1}+0.05Z_{t-1}^2.
\end{align*}
Here and below, $\varepsilon_{1,t}$, $\varepsilon_{2,t}$ and $\varepsilon_{3,t}$ are i.i.d. $N(0,1)$ and mutually independent, and $Z_t$ in (S2)-(S3) and (P1)-(P6) follows an AR(1) model: $Z_t =0.5Z_{t-1}+\varepsilon_{2,t}$. These DGPs cover a wide range of both linear and nonlinear time series processes. For (S1), $(Y_t,Z_t,X_t)$ is simply an i.i.d. sequence; we test $Y_t\bot Z_t|X_t$ such that $W_t=X_t$. From (S2) to (P7), we test $Y_t\bot Z_{t-1}|Y_{t-1}$ such that $W_t=Y_{t-1}$; that is, we are interested in checking whether $Z_{t-1}$ has any predictive ability in explaining $Y_t$ (in mean, variance, or higher moments) after controlling the first lag of $Y_t$. Note that (S4) and (P7) do not satisfy \eqref{markov}.
For each DGP, we first generate $n+500$ observations and then discard the first 500 observations to minimize effects of initial values. The number of Monte Carlo simulations is $2000$ and the bootstrap critical values are obtained from $B=1000$ bootstrap replications. Four sample sizes, $n=100$, $200$, $400$ and $800$, are considered. We only report results for the nominal level of 5\% and results for other levels are available upon request. We choose the standard normal density $K(x)=(2\pi)^{-1/2}\exp{(-x^2/2)}$ as our kernel. Bandwidth of the form $h=c n^{-1/3.5}$ is used and results with $c=0.5$, $1.0$ and $1.5$ are reported to check the sensitivity of our tests to different bandwidths. How to choose $h$ to maximize the performance in our testing framework is beyond the scope of this paper and is left for future research. In addition to the fixed bandwidth $h$, in the supplemental appendix we provide more simulation results with a data-driven one $h=1.06\cdot Std(X_t)\cdot n^{-1/3.5}$.
Table \eqref{cvm_mb} and \eqref{ks_mb} report respectively the empirical rejection rates of $CvM_n$ and $KS_n$ under (S1)-(P7) using critical values obtained with the multiplier bootstrap method proposed in Section \ref{boot}. Both $CvM_n$ and $KS_n$ have acceptable empirical sizes in the cases of moderate sample sizes for (S1)-(S3), with $KS_n$ having more accurate sizes than $CvM_n$. When \eqref{markov} fails in (S4), $CvM_n$ exhibits a severe size distortion for $n=100$ and $n=200$, while $KS_n$ has less distortion; when sample size increases to 800, the two tests show improved sizes especially for $c=0.5$. For the empirical power performance, when sample size is as small as $n=100$, both $CvM_n$ and $KS_n$ are not very powerful against (P2)-(P7), with (P7) the most difficult DGP to detect. They gain power rapidly as sample size increases. Note that in (P7) \eqref{markov} fails, which might be the reason for low power in small samples. In summary, our simulation results indicate that $KS_n$ preserves sizes better than $CvM_n$, while $CvM_n$ is more powerful than $KS_n$; in addition, whether \eqref{markov} holds will affect the performance. The simulation results also show that the empirical sizes and powers of $CvM_n$ and $KS_n$ are somehow sensitive to bandwidth choices in very small samples; a general pattern is that larger $c$ tends to deliver a higher power but produces a bigger size distortion. When $n=800$, the overall performance is satisfactory.
\begin{center}
------------------------------------------
Tables \ref{cvm_mb} \& \ref{ks_mb} about here
------------------------------------------
\end{center}
To investigate if we have improved performance when we take into account the general dependence structure in the data, we also study the following block multiplier bootstrap:
\begin{equation*}
S_n^{\ast block}(w,y,z)=\frac{1}{\sqrt n}\sum_{t=1}^{n-L+1}\zeta_t\sum_{s=t}^{t+L-1}\widehat e_t(w,y,z),
\end{equation*}
where $\{\zeta_t\}_{t=1}^{n-L+1}$ are i.i.d. $N(0,L^{-1})$ and independent of the sample $\{(W_t^\top,Y_t^\top,Z_t^\top)^\top\}_{t=1}^n$, with $L:=L_n$ the block length diverging to infinity at a slower rate than $n$ as $n\to\infty$. Note that when $L=1$, $S_n^{\ast block}(w,y,z)$ reduces to the multiplier bootstrap introduced in Section \ref{boot}, the validity of which replies on the restriction \eqref{markov}.
Table \eqref{cvm_bmb} and \eqref{ks_bmb} report respectively the empirical rejection rates of $CvM^{block}_n$ and $KS^{block}_n$ based on $S_n^{\ast block}(w,y,z)$. It is expected that $CvM^{block}_n$ and $KS^{block}_n$ should work when \eqref{markov} does not hold and consequently the dependence structure in the data should not be ignored. In implementing the block bootstrap, we have used $L=\lfloor an^{1/4} \rfloor$ with $a=1$, $2$ and $4$ to check the sensitivity of $CvM^{block}_n$ and $KS^{block}_n$ with respect to different block lengths. To save space, we only report the results for bandwidth $h=cn^{-1/3.5}$ with $c=1.0$. The results in Table \eqref{cvm_bmb} and \eqref{ks_bmb} are based on 1000 simulations and 200 bootstrap replications. It is observed that $CvM^{block}_n$ and $KS^{block}_n$ have less distorted empirical sizes especially for (S4) than $CvM_n$ and $KS_n$ do. At the same time, $CvM^{block}_n$ and $KS^{block}_n$ are not as powerful as $CvM_n$ and $KS_n$ even for (P7). Like before, $KS^{block}_n$ preserves sizes better than $CvM^{block}_n$, while $CvM^{block}_n$ is more powerful than $KS^{block}_n$. Lastly, for given $h$, the tests also depend on the choice of $L$; smaller $L$ tends to deliver a higher power but produces oversized tests.
\begin{center}
------------------------------------------
Tables \ref{cvm_bmb} \& \ref{ks_bmb} about here
------------------------------------------
\end{center}
\section{An empirical study}
We examine whether there exists nonlinear predictability of equity risk premium using variance risk premium. The variance risk premium is defined as the difference between the risk-neutral and objective expectations of realized variance, where the risk-neutral expectation of variance is measured as the end-of-month Volatility Index-squared de-annualized and the realized variance is the sum of squared 5-minute log returns of the S\&P 500 index over the month.
There is a rich literature on the predictive power of variance risk premium for the aggregate stock market returns, bond returns or exchange rate returns. For example, Bollerslev et al. (2009) first discover that variance risk premium is able to explain a nontrivial fraction of the time series variation in post 1990 aggregate stock market returns, with high (low) premia predicting high (low) future returns; Wang et al. (2013) find the empirical evidence suggesting that the firm-level variance risk premium has a prominent explanatory power for credit spreads in the presence of market- and firm-level control variables; by defining a ``global'' variance risk premium, Bollerslev et al. (2013) uncover stronger predictability of aggregate stock market returns using variance risk premium across countries; while Della Corte et al. (2013) investigate the predictive information content in foreign exchange volatility risk premia for exchange rate returns and find that a portfolio that sells currencies with high insurance costs and buys currencies with low insurance costs generates sizeable out-of-sample returns and Sharpe ratios.
We use monthly aggregate S\&P 500 composite index over the period January 1996 to September 2008. Our empirical analysis is based on the logarithmic return on the S\&P 500 in excess of the 3-month T-bill rate. Let $RP_{t+\tau}$ be the risk premium $\tau$ months ahead and $VRP_{t}$ be the variance risk premium at time $t$. In this empirical study, we take $\tau=$1, 3, 6, and 9 months. We shall examine if the variance risk premium $VRP_t$ explains in a linear or nonlinear way the risk premium $RP_{t+\tau}$ given the information $RP_{t}$, which is equivalent to stating whether $VRP_t$ Granger causes $RP_t$ by setting the lag order to $\tau$. To test for the presence of (nonlinear) predictability of $VRP_t$, we consider to test
\begin{equation*}
\text{H}_0: \mathrm{P}\left\{F(RP_{t+\tau},VRP_{t}|RP_{t})=F(RP_{t+\tau}|RP_{t})F(VRP_t|RP_{t})\right\}=1
\end{equation*}
against $\text{H}_1: \mathrm{P}\left\{F(RP_{t+\tau},VRP_{t}|RP_{t})=F(RP_{t+\tau}|RP_{t})F(VRP_t|RP_{t})\right\}<1$. That is, for a given horizon $\tau$, we test the conditional independence of $RP_{t+\tau}$ and $VRP_t$ given $RP_{t}$, i.e. $RP_{t+\tau}\bot VRP_{t}|RP_{t}$.
For the purpose of comparison, we also perform the popular linear causality analysis in the literature. To this end, we consider the following linear regression model:
\begin{equation}
RP_{t+\tau}=\mu_{\tau}+\beta_{\tau}RP_{t}+\alpha_{\tau}VRP_{t}+\varepsilon_{t+\tau}. \label{lin}
\end{equation}
The hypothesis of interest is that VRP does not Granger cause RP for $\tau$ months ahead in a linear way, i.e. testing the null hypothesis $\text{H}_0:\alpha_{\tau}=0$ against the alternative
hypothesis $\text{H}_1:\alpha_{\tau}\neq 0$. To test $\text{H}_0$, standard $t$-statistic given by $t_{\hat{\alpha}_{\tau}}=\hat{\alpha}_{\tau}/\hat{\sigma}_{\hat{\alpha}_{\tau}}$ will be calculated, where $\hat{\alpha}_{\tau}$ is the least squares estimator of $\alpha_{\tau}$ and $\hat{\sigma}_{\hat{\alpha}_{\tau}}$ is the estimator of its standard error $\sigma_{\hat{\alpha}_{\tau}}$. Moreover, to avoid the impact of possible dependence in the residual terms $\hat{\varepsilon}_{t+\tau}$ on our inference, $\hat{\sigma}_{\hat{\alpha}_{\tau}}$ is calculated using the commonly used heteroscedasticity autocorrelation consistent (HAC) robust variance estimator suggested by Newey and West (1987).
Table \ref{empirical} reports the testing results for Granger causality (e.g. nonlinear predictability) from variance risk premium to risk premium, at four different horizons, using our proposed tests $CvM_n$ and $KS_n$ as well as the linear test. To check the robustness of our empirical findings, we also include the block multiplier tests $CvM_n^{block}$ and $KS_n^{block}$, where the block length is set to be $L=\lfloor an^{1/4} \rfloor$ with $a=2$. Results for $a=1$ and $4$ are similar and hence are omitted. The implementation of all bootstrap based tests is as demonstrated in the Monte Carlo simulations part with the number of bootstrap replications $B=10,000$. We have obtained encouraging findings from Table \ref{empirical} that are relevant to both empirical and theoretical studies of variance risk premium as a suitable predictor for risk premium. Specifically, results from linear regression based test in \eqref{lin} clearly fail to reject the null hypothesis of no linear predictability for one to six months (short-run) horizons; they only show weak evidence of linear predictability until at the nine months (long-run) horizon at the 5\% significance level. On the other hand, using our tests, we find convincing evidence that risk premium can be well predicted using variance risk premium at both mid-run and long-run horizons. Testing results are not sensitive to the bandwidth choices for six and nine months horizon, nor does the block type tests. For the one month horizon, we may conclude that variance risk premium has no predictive power, while the findings for three months horizon tend to indicate the existence of predictability. Overall, we find that there is a very high degree of predictability at horizons more than one-month which could be attributed to the nonlinear predictive effect. Our empirical evidence also indicates that caution is needed when interpreting results based on the linear regression.
\begin{center}
------------------------------------------
Table \ref{empirical} about here
------------------------------------------
\end{center}
\section{Conclusion}
This paper proposes new consistent nonparametric tests of conditional independence for time series data based on the empirical process method. The asymptotic properties of the proposed tests under the null, the alternative, and the sequence of local alternatives are investigated. To implement the test in practice, a multiplier bootstrap procedure is suggested and its asymptotic validity is formally justified. The test can be applied to testing for conditional independence in a wide variety of nonparametric models. Using the proposed test, we also study whether there exists some nonlinear predictability of equity risk premium using variance risk premium.
\bibliographystyle{chicago}
\bibliography{ref}
\newpage