Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
57,536 characters · 12 sections · 38 citation commands
\def\spacingset#1{ {#1}} \spacingset{1}
\spacingset{1.9}
The family of vector autoregressive (VAR) models and the family of multivariate (G)ARCH models are among some of the most popular frameworks for modelling dynamic interactions of multiple variables. The VAR family usually captures the dynamic by imposing structures on the time series itself, while the (G)ARCH family imposes restrictions on the conditional second moments. We acknowledge the vast literature of both families, and have no intention to exhaust all relevant studies in this paper for the sake of space. We refer interested readers to stock2001vector and bauwens2006multivariate for excellent review on both families.
Although both families have rich literature on their own, to the best of the authors' knowledge not many works have been done to bridge them. Among limited attempts (e.g., ling2003asymptotic,bardet2009asymptotic), most (if not all) of these studies rely on the stationarity assumption. While the stationarity assumption comes in handy when deriving asymptotic properties, it may not be very realistic in practice (preuss2015detection,chen2021inference). For example, economic and financial data always include different macro shocks, as a consequence the behaviour can be quite volatile; the climate data may contain certain time trend which recently has attracted lots of attention due to greenhouse emission; etc. Anyway, certain nonstationarity may always occur.
To account for nonstationarity, locally stationary processes have received considerable attention since the seminal work of dahlhaus1996kullback, dette2011measure, zhang2012inference, truquet2017parameter, dahlhaus2019towards, among others. In contrast to the unit root process, the locally stationary process nicely balances stationarity and nonstationarity by allowing for the simultaneous presence of both types of behaviours in one time series process. In a very recent paper, karmakar2021simultaneous consider simultaneous inference for a general class of univariate $p$-Markov processes with time-varying coefficients, which covers several time-varying versions of the classical univariate models (e.g., AR, ARCH, AR-ARCH) as special cases. Despite its generality, their study still rules out the time-varying versions of some widely used models (e.g., ARMA, GARCH, ARMA-GARCH). Also, it is worth mentioning this line of research heavily focuses on univariate time series, which somewhat limits the popularity of locally stationary processes.
That said, it is reasonable to call for a framework which can marry the VAR family and the (G)ARCH family while allowing for nonstationarity. To provide a concrete example, consider a time-varying multivariate GARCH model, which can model the co-movements of financial returns. Detailed investigation on such a model can help answer research questions like (i). Is the volatility of a market leading the volatility of other markets? (ii) Whether the correlations between asset returns change over time? (iii). Are they increasing in the long run, perhaps because of the globalization of financial markets? These are of great practical importance for both investors and policymakers (bauwens2006multivariate,diebold2009measuring).
To allow for flexibility as much as possible from the modelling perspective, we consider a class of multivariate causal processes as follows:
where $\tau_t= t/T$, $\bm{\mu}\left(\cdot\right)$ is an $m$-dimensional random vector, $\mathbf{H}\left(\cdot\right)$ is an $m\times m$-dimensional random matrix, $\bm{\theta}(\tau)$ is a $d\times 1$ time-varying parameter of interest with each element belonging to $C^3[0,1]$, and $\{\bm{\varepsilon}_t\}$ is a sequence of independent and identically distributed (i.i.d.) random vectors. Note that the value of $d$ usually depends on the value of $m$, and the connection becomes clear once a specific model is considered. As far as we are concerned, both of $m$ and $d$ are fixed throughout the paper. Notably, both $\bm{\mu}(\cdot)$ and $ \mathbf{H}(\cdot)$ are known, and share the same unknown parameter $\bm{\theta}(\cdot)$. The setting for $t\le 0$ regulates the time series for the periods that we do not observe, which is commonly adopted when certain nonstationarity gets involved (e.g., vogt2012nonparametric). Essentially, it requires the initial time period does not have a diverging behaviour.
Before proceeding further, we provide two examples to briefly illustrate the rationality behind (ref), and leave the detailed investigation on these examples to Section (ref). We refer interested readers to ling2003adaptive, ling2003asymptotic and bardet2009asymptotic for extensive investigation on the parametric counterparts of these examples.
Example 1: Consider the time-varying VARMA($p,q$) model
It is not hard to show that (ref) admits a presentation in the form of (ref), and
where $\bm{\Omega}(\cdot):=\bm{\omega}(\cdot)\bm{\omega}^\top(\cdot)$.
Example 2: Consider the time-varying multivariate GARCH($p,q$) model
where $h_{j,t}$ stands for the $j^{th}$ element of $\mathbf{h}_t$, and $\bm{\eta}_t = \bm{\Omega}^{1/2}(\tau_t) \bm{\varepsilon}_t$. The model (ref) generalizes the models of bollerslev1990modelling and jeantheau1998strong. Similar to Example 1, we show that (ref) admits a representation in the form of (ref), and
In view of the development of Example 1 and Example 2 in Section (ref), one may further show the time-varying counterparts of the parametric models mentioned in bardet2009asymptotic are also covered by (ref). To this end, we argue that (ref) does not only allows for nonstationarity and conditional heteroskedasticity, but also provides sufficient flexibility to cover many well adopted models in the literature.
In this paper, our contributions are in the following four-fold: (1). we consider a wide class of time-varying multivariate causal processes which nests many classic and new examples as special cases; (2). we prove the existence of a weakly dependent stationary approximation for the model (ref) at any given time of interest (i.e., $\forall\tau\in[0,1]$), which is the foundation in order to establish asymptotic properties associated with the model; (3). we establish the estimation theory, and provide both point-wise and simultaneous inferences on the coefficient functions of which both are important for practical works (zhou2010simultaneous); (4). we demonstrate the theoretical findings through both simulated and real data examples.
The paper is organized as follows. Section (ref) presents the theoretical findings associated with the stationary approximation, estimation and inferences. In Section (ref), we conduct extensive simulation studies to examine the theoretical findings, and further investigate the time-varying conditional correlations between the Chinese and U.S. Stock market. Section (ref) concludes. Due to space limit, we give the proofs of the main results to the online appendices of the paper.
Before proceeding further, it is convenient to introduce some notation: the symbol $|\cdot|$ denotes the Euclidean norm of a vector or the spectral norm for a matrix; $\|\mathbf{v}\|_q:=\left(E|\mathbf{v}|^q\right)^{1/q}$ and $\|\cdot\|:=\|\cdot\|_2$ for short; $\otimes$ denotes the Kronecker product; $\odot$ denotes the Hadamard product; $\mathbf{I}_a$ stands for an $a\times a$ identity matrix; $\mathbf{0}_{a\times b}$ stands for an $a\times b$ matrix of zeros, and we write $\mathbf{0}_a$ for short when $a=b$; for a function $g(w)$, let $g^{(j)}(w)$ be the $j^{th}$ derivative of $g(w)$, where $j\ge 0$ and $g^{(0)}(w) \equiv g(w)$; $K_h(\cdot) =K(\cdot/h)/h$, where $K(\cdot)$ and $h$ stand for a nonparametric kernel function and a bandwidth respectively; let $\tilde{c}_k =\int_{-1}^{1} u^k K(u) \mathrm{d}u$ and $\tilde{v}_k= \int_{-1}^{1} u^k K^2(u) \mathrm{d}u$ for integer $k\ge 0$; $\mathrm{diag}(\mathbf{a})$ is a diagonal matrix with the vector $\mathbf{a}$ on its main diagonal, while $\mathrm{diag}(\mathbf{A})$ creates a vector from the diagonal of matrix $\mathbf{A}$; finally, let $\to_P$ and $\to_D$ denote convergence in probability and convergence in distribution, respectively.
In this section, we first prove the existence of a weakly dependent stationary approximation for the model (ref) in Section (ref); we then provide the estimation approach using the local linear quasi-maximum-likelihood estimation and establish the asymptotic properties of the proposed estimator in Section (ref); Section (ref) provides results on both point-wise and simultaneous inferences; Section (ref) gives some detailed examples to justify the usefulness of our study.
To study (ref), the first challenge lies in the fact that the model may not be stationary. Therefore, for $\forall\tau\in [0,1]$, we initial our analysis by finding a stationary approximation for each $\mathbf{x}_t$ with $t\ge 1$. By doing so, we are able to measure the weak dependence of $\{\mathbf{x}_t \}$ using the nonlinear system theory in wu2005nonlinear, which then provides us a framework to derive the asymptotic properties accordingly.
To be clear on the dependence measure, consider an example in which $\mathbf{e}_t$ is a stationary process, and admits a causal representation $\mathbf{e}_t = \mathbf{J}(\bm{\varepsilon}_t,\bm{\varepsilon}_{t-1},\ldots)$ with $\mathbf{J}(\cdot)$ being a measurable function. See tong1990non for discussion on nonlinear time series of this kind. For $k\geq 0$, we define the following dependence measure:
where $\bm{\varepsilon}_0^*$ is an independent copy of $\{\bm{\varepsilon}_j\}$. Being able to measure the time series dependence such as (ref) is the starting point for time series analyses.
We now introduce some basic assumptions.
Assumption (ref).1 is standard when studying dynamic time series model (lutkepohl2005new). In Assumption (ref).2, $\bm\vartheta$ is a generic $d\times 1$ vector, and has the same length as $\bm \theta(\cdot)$. This assumption imposes Lipschitz-type conditions on $\bm{\mu}(\cdot)$ and $\mathbf{H}(\cdot)$, which are rather minor, and can be easily fulfilled by a variety of models such as those mentioned in Section (ref). See Propositions (ref)-(ref) below for details. Assumption (ref).3 does not only guarantee a stationary approximation for each $\mathbf{x}_t$, but also ensures the approximated process has some proper moments. Similar conditions have also been adopted in bardet2009asymptotic.
With these conditions in hand, we present the following proposition which facilitates the development in what follows.
It is worth mentioning that for a univariate $p$-Markov process
karmakar2021simultaneous show that there exists $0 < \rho < 1$ such that $\sup_{\tau \in[0,1]}\delta_{r}^{\widetilde{x}_p(\tau)}(k) =O(\rho^k)$ based on the development of wu2004limit. From a methodological viewpoint, we give a set of new proofs which allow us to measure the dependence of multivariate causal processes with infinity memory. The term $\sum_{j=p+1}^{\infty}\left[\alpha_j(\bm{\theta}(\tau)) + \beta_j(\bm{\theta}(\tau))\right]$ in the second result of Proposition (ref) arises due to the infinity memory structure of $\widetilde{\mathbf{x}}_t(\tau)$. Thus, the dependence $\delta_{r}^{\widetilde{\mathbf{x}}(\tau)}(k)$ relies on the choice of $p$ and the decay rates of the coefficients $\alpha_j(\bm{\theta}(\tau))$ and $\beta_j(\bm{\theta}(\tau))$.
To ensure $\widetilde{\mathbf{x}}_t(\tau) $ can approximate $\mathbf{x}_t$ reasonably well, we impose more structure below.
Assumption (ref).1 imposes another Lipschitz-type condition with respect to the parameter space. Assumption (ref).2 further restricts the decay rates of $\alpha_j(\bm{\theta}(\tau))$ and $\beta_j(\bm{\theta}(\tau))$.
Using Assumptions (ref)--(ref), we can measure the distance between $\widetilde{\mathbf{x}}_t(\tau) $ and $\mathbf{x}_t$ as follows.
We can consider Proposition (ref) as the stochastic version of the H\"older continuity. Having established the stationary approximation in Proposition (ref), we move on to investigate the estimation theory in the next subsection.
We point out a few facts to facilitate the setup of the likelihood function. First, let $\mathbf{z}_{t} =(\mathbf{x}_t,\mathbf{x}_{t-1},\ldots)$ include all the information of $\mathbf{x}_t$ up to the time period $t$. However, in practice, our observation on $\mathbf{x}_t$ only starting from $t=1$, so we have to work with the truncated version of $\mathbf{z}_{t}$ for each $t\ge 1$:
Second, we note that when $\tau_t$ is sufficiently close to $\tau$,
Therefore, we are able to parametrize $\bm{\theta}(\cdot)$, and consider the maximum-likelihood estimation for each given $\tau$. Finally, since $\bm{\varepsilon}_t$ may not be normally distributed, we consider the local linear quasi-maximum-likelihood estimation (QMLE) method.
Thus, our likelihood function is specified as follows:
where
Accordingly, for $\forall \tau$, $(\bm{\theta}(\tau),h\bm{\theta}^{(1)}(\tau))$ is estimated by
where $\mathbf{E}_T(r) = \bm{\Theta}_r\times(h\cdot\bm{\Theta}^{(1)})$ and $\bm{\Theta}^{(1)}$ is a compact set.
We impose more structures in order to derive the asymptotic distribution.
Assumption (ref).1 ensures the positive definiteness of the covariance matrix of the likelihood function, and is widely adopted when studying the multivariate time series (e.g., page 2736 of bardet2009asymptotic). In fact, the validity of this assumption is easy to justify in view of (ref) and (ref) for Example 1 and Example 2 below. Assumption (ref).2 imposes an standard identification condition in the literature of M-estimation (e.g., Proposition 3.4 of jeantheau1998strong). It is noteworthy that the current form of Assumption (ref) accommodates the flexibility of the model (ref), which is in fact unnecessary if we have a detailed model in practice. See Section (ref) for example.
Assumption (ref) imposes the Lipschitz-type conditions on the first and second order derivatives of $\bm{\mu}(\cdot)$ and $\mathbf{H}(\cdot)$ to ensure the smoothness of their functional components.
Assumption (ref) is a set of regular conditions on the kernel function and the bandwidth.
With these conditions in hand, we summarize the first theorem of this paper below.
After deriving the asymptotic distribution, we will establish both the point-wise inference and the simultaneous inference in the following.
In this section, we first discuss how to conduct point-wise inference, and then move on to derive the asymptotic results associated with the simultaneous inference. Specifically, for some preassigned significance level $\alpha \in (0,1)$, we shall construct a $100(1-\alpha)\%$ asymptotic simultaneous confidence band (SCB) $\{ \Upsilon(\tau), 0\leq \tau\leq 1 \}$ for $\bm{\theta}(\cdot)$ in the sense that $$ \lim_{T\to\infty} \Pr\left(\bm{\theta}(\tau) \in \Upsilon(\tau), 0\leq \tau\leq 1\right) =1-\alpha. $$ Notably, the simultaneous inference nests the traditional constancy test as a special case. It does not only allow one to examine whether a time-varying model should be preferred to its parametric counterpart, but also allows one to test any particular functional form of interest. For example, if a horizontal line can be embedded in the SCB $\{ \Upsilon(\tau)\}$, then we accept the hypothesis that some elements of $\bm{\theta}(\tau)$ are constant.
Point-wise Inference: First, we construct a bias-corrected estimator in order to remove the asymptotic bias of Theorem (ref). Specifically, we let
where $\widehat{\bm{\theta}}_{h/\sqrt{2}}(\tau)$ is defined in the same way as $\widehat{\bm{\theta}} (\tau)$ but using the bandwidth $h/\sqrt{2}$.
After tedious development (Lemma B.7 of Appendix B), we have uniformly over $\tau \in[h,1-h]$
where $\widetilde{K}(x)=2\sqrt{2}K(\sqrt{2}x)-K(x)$ that is essentially a fourth-order kernel. It then infers that under the conditions of Theorem (ref),
where $v_0 = \int_{-1}^{1}\widetilde{K}^2(u)\mathrm{d}u$.
It is noteworthy that the construction of (ref) is different from directly using the fourth-order kernel in the regression. In terms of bandwidth selection, the traditional methods (e.g., cross-validation) still remain valid for (ref) (richter2019cross). However, if one directly employs the fourth-order kernel in the regression, it remains unclear how to select the optimal bandwidth in practice.
Now we discuss how to estimate $\bm{\Sigma}_{\bm{\theta}}(\tau) $ which is constructed by $\bm{\Sigma}(\tau)$ and $\bm{\Omega}(\tau)$. Intuitively, we consider the following estimator
where
Note that we consider a local constant estimator in (ref) rather than a local linear one, that is to avoid an implementation issue for finite sample studies (i.e., nonpositive definite covariance may occur when the local linear approach is employed). Such a numerical problem has been well explained and investigated in the literature. See chen2015local for example.
The following corollary summarizes the asymptotic property of (ref).
Simultaneous Inference: We now consider the simultaneous inference. To allow for flexibility, we first introduce a selection matrix $\mathbf{C}$ with full row rank, which selects the parameters of interest as follows:
Accordingly, the estimator and the corresponding asymptotic covariance matrix become
In Theorem (ref), $\nu$ is slightly smaller than $1/2$ as we only require $r$ to be slightly larger than $6$. Hence, the usual optimal bandwidth $h_{opt} = O(T^{-1/5})$ satisfies the conditions $(\log T)^4/(T^{\nu}h)\to 0$ and $Th^7\log T \to 0$.
As shown in Theorem (ref), the convergence rate of the simultaneous confidence intervals for $\bm{\theta}_{\mathbf{C}}(\cdot)$ is of logarithmic rate and is therefore slow. In order to improve the rate, we consider a bootstrap method which shows a much better finite sample performance. We summarize the result in the following corollary.
By Corollary (ref), we propose the following numerical procedure to construct the SCB of $\bm{\theta}_{\mathbf{C}}(\tau)$:
Below, we demonstrate the usefulness of the aforementioned results by considering Example 1 and Example 2 of Section (ref).
Example 1 (Cont.) --- For $\forall\tau\in[0,1]$, simple algebra shows that the approximated stationary process is defined by
where $\bm{\theta}(\tau)$ has been defined in (ref), and
Additionally, in (ref), $\bm{\Gamma}_j(\tau)$ is yielded as follows:
where $\mathbf{A}_{\tau}(L):=\mathbf{I}_m - \mathbf{A}_1(\tau)L-\cdots-\mathbf{A}_p(\tau)L^p$ and $\mathbf{B}_{\tau}(L):=\mathbf{I}_m + \mathbf{B}_1(\tau)L+\cdots+\mathbf{B}_q(\tau)L^q$.
Then we are able to present the following proposition.
We note that the detailed identification conditions required for VARMA processes (e.g., the final equations form or echelon form) can be found in lutkepohl2005new. We no longer discuss them here in order not to derivative from our main goal.
Example 2 (Cont.) --- We further let
For $\forall\tau\in[0,1]$, the corresponding approximated stationary process is defined as
where
Note that $ \mathbf{\Psi}_j(\tau)$ is generated as follows:
where $\mathbf{C}_{\tau}(L):=\mathbf{C}_1(\tau)L+\cdots+\mathbf{C}_p(\tau)L^p$ and $\mathbf{D}_{\tau}(L):=\mathbf{I}_m - \mathbf{D}_1(\tau)L-\cdots-\mathbf{D}_q(\tau)L^q$.
Consequently, we can present the following proposition.
For the identification conditions of the GARCH process, we refer readers to Proposition 3.4 of jeantheau1998strong, who proves that assuming the minimal representation is enough for ensuring Assumption (ref) holds.
In the following section, we conduct numerical studies using both simulated and real data to evaluate the finite-sample performance of the proposed estimation and inferential methods.
In this section, we first present the details of the numerical implementations in Section (ref), and then conduct extensive simulations in Section (ref). Section (ref) presents a real data example on the conditional correlations between the Chinese and U.S. stock markets.
Throughout the numerical studies, the Epanechnikov kernel $K(u) = 0.75(1-u^2)I(|u|\leq1)$ is adopted. Following zhou2010simultaneous, we use $\widetilde{h} = 2\widehat{h}$ for the biased corrected estimator, where $\widehat{h}$ is the bandwidth selected by the cross-validation method of richter2019cross.
Specifically, define the leave-one-out local linear QMLE
where $$ \mathcal{L}_{T,-t}^c(\tau,\bm{\eta}_1,\bm{\eta}_2)=\frac{1}{T}\sum_{s=1,\neq t}^{T}\mathcal{l}(\mathbf{x}_s,\mathbf{z}_{s-1}^c;\bm{\eta}_1+\bm{\eta}_2\cdot(\tau_s-\tau)/h)K_h(\tau_s-\tau). $$ Then, the bandwidth is chosen by
As shown in richter2019cross, this cross validation method works well as long as $\gradient\mathcal{l}$ is uncorrelated, which implies that this desirable property should hold in our case.
Notably, when considering some specific models, the implementation may be further simplified. We provide more discussions along this line in Appendix B.4.
In the simulation studies, we examine the empirical coverage probabilities of simultaneous confidence intervals for nominal levels $\alpha =90\%,\ 95\%$. We consider the time-varying VARMA($2,1$) and multivariate GARCH($1,1$) model as follows:
Let the sample size be $T \in\{500,1000\}$ ($T \in\{1000,2000,4000\}$) for the VARMA model (the GARCH model). We conduct $1000$ replications for each choice of $T$. Several different bandwidths close to $\widetilde{h}$ are reported to check the sensitivity of bandwidth selection.
We present the empirical coverage probabilities associated with the SCB in Tables (ref)--(ref). For the vector- or matrix-valued unknown coefficients, we take an average across the elements. A few facts emerge from the tables. First, the finite sample coverage probabilities are smaller than their nominal level when $T = 500$ ($T = 1000,2000$) for the VARMA model (the GARCH model), but are fairly close to their nominal level as $T = 1000$ ($T=4000$) for the VARMA model (the GARCH model). Second, the behaviour of the estimated simultaneous confidence intervals is not sensitive to the choices of bandwidths. Third, the GARCH model requires more data to reach a reasonable finite sample performance.
In this subsection, we investigate the time-varying conditional correlations between the Chinese and U.S. stock markets using the time-varying multivariate GARCH model. Recently, there is a growing literature to study the relationship of the two stock markets (e.g., zhang2014has,pan2022modeling), as the Chinese stock market has become the world's second largest stock market after 2009. Understanding the interactions among different financial markets is important for investors and policymakers diebold2009measuring,bensaida2019good. For example, high equity market interdependence implies poor diversification benefits from portfolios, but highlights the possibility of better hedging benefits.
Previous research documents a strong positive link between the degree of globalization and equity market interdependence baele2005volatility. Along this line of research, one important question is that whether the interdependence between the Chinese and U.S. stock markets has increased over time due to globalization so that estimates from historical data are unreliable for modern policy analysis, asset pricing and risk management. The existing results present many discrepancies, which may be due to the fact that the relationship evolves with time. Apparently, the results also indicate that one should use time-varying GARCH model to accommodate potential nonstationarity inherited in these financial variables. In addition, as pointed out by caporin2013ten, dynamic conditional correlation (DCC) GARCH model represents the dynamic conditional covariances of the standardized residuals, and hence does not yield dynamic conditional correlations; DCC yields inconsistent two step estimators; DCC has no asymptotic properties. In what follows, we address these issues using the newly proposed approach. The estimation is conducted in exactly the same way as in Section (ref), so we no longer repeat the details.
We calculate the Chinese and U.S. stock returns based on weekly Shanghai Stock Exchange (SSE) Composite Index and S&P 500 Index as they are the most comprehensive and diversified stock indices. The sample employed in this study spanning from January 2000 to February 2022 provides $1119$ observations\footnote{The data are collected from Yahoo Finance at \url{https://finance.yahoo.com/}.}. Figure (ref) plots the two weekly returns as well as sample autocorrelation functions of squared data, which shows the typical “volatility clustering” phenomenon.
We next fit the data to a time-varying multivariate GARCH(1,1) model and are particularly interested in the estimates of time-varying conditional correlations, i.e., $$ E\left(x_{1,t}x_{2,t}\mid \mathcal{F}_{t-1}\right)/\sqrt{E\left(x_{1,t}^2\mid \mathcal{F}_{t-1}\right)E\left(x_{2,t}^2\mid \mathcal{F}_{t-1}\right)} = \rho_{1,2}(\tau_t), $$ where $\rho_{1,2}(\cdot)$ is defined in (ref). Figure (ref) plots the estimates (black solid line) of time-varying conditional correlations between the two stock markets as well as 95% simultaneous confidence intervals (red dashed line) and 95% pointwise confidence intervals (black dashed line). Based on the simultaneous confidence intervals, apparently, the conditional correlations vary with respect to time. Moreover, as clearly presented in Figure (ref), the interdependence between the two stock markets is increasing over time. By examining the pointwise confidence intervals, we can conclude that the two stock markets are not significantly correlated before 2005, but the relationship has been greatly enhanced in recent years. These results have important implications for investment and risk management. For example, it implies that the Chinese and U.S. investors who use cross-country portfolio strategies to eliminate country specific risks may be benefit from hedging. However, all types of investors should be cautious since the relations between the Chinese and U.S. stock markets are time-varying.
In this paper, we consider a wide class of time-varying multivariate causal processes which nests many classic and new examples as special cases. We first prove the existence of a weakly dependent stationary approximation for the model (ref) which is the foundation to establish the corresponding asymptotic properties. Afterwards, we consider the QMLE estimation approach, and provide both point-wise and simultaneous inferences on the coefficient functions. In addition, we demonstrate the theoretical findings through both simulated and real data examples. In particular, we show the empirical relevance of our study using an application to evaluate the conditional correlations between the stock markets of China and U.S. We find that the interdependence between the two stock markets is increasing over time.
There are several directions for possible extensions. The first one is to consider quantile regression methods for such locally stationary multivariate causal processes. The second one is to propose a more powerful $L_2$ test based on the weighted integrated squared errors for testing whether some coefficients are time-invariant. We wish to leave such issues for future study.
The authors of this paper would like to thank George Athanasopoulos, David Frazier and Gael Martin for their constructive comments on earlier versions of this paper. Thanks also go to seminar participants for their insightful suggestions. Gao and Peng would also like to acknowledge the Australian Research Council Discovery Projects Program for its financial support under Grant Numbers: DP170104421 & DP210100476.
{
}
{
The file includes Appendix A and Appendix B. We first present some technical tools in Appendix (ref), which will be repeatedly used in the development. We then provide the proofs of main results in Appendix (ref). We provide several preliminary lemmas in Appendix (ref) as well as some secondary lemmas in Appendix (ref), and then present the proofs of preliminary lemmas in Appendix (ref). Appendix (ref) discusses several computational issues of the local linear ML estimation.
In what follows, $M$ and $O(1)$ always stand for some bounded constants, and may be different at each appearance.
{