Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,770 characters · 16 sections · 39 citation commands
Vector error-correction models (VECM) have been widely used in practice to study the short-run dynamics and long-run equilibrium of multiple non-stationary time series (e.g., barigozzi2022inference). These models have proven to be useful in forecasting non-stationary time series and testing the validity of various hypotheses and theories, such as the rational expectations hypothesis of the term structure (bauer2020interest), money demand theory (benati2021international), real business-cycle theory (king1991stochastic), and more. However, it is important to note that all of these studies are based on the assumption of constant parameters. In practice, parameters may evolve over time, and failure to take these changes into account leads to incorrect policy implications and predictions (fan2003nonlinear).
To model the changes over time, various parametric VECM models (e.g., hansen2002testing,hansen2003structural,bergamelli2019combining) have been proposed to allow for abrupt structural breaks. However, for this line of research, the corresponding estimation and inferential theory has not been well established, which can undermine the reliability of these models, perhaps because the relevant empirical process theory does not typically apply to this case when the data generating process exhibits some uncertain information (e.g., unknown structural breaks and dates) and involves different time series behaviours (e.g., unit-root and piecewise stationary processes). For instance, hansen2002testing propose a method to test for the presence of a threshold effect in the VECM model with a single cointegration vector, but do not establish the corresponding estimation theory. Similarly, hansen2003structural assumes the change points and the number of cointegration relations are given, while bergamelli2019combining consider testing the structural breaks of a VECM model with unknown break dates but do not provide estimation and inferential theories for model parameters. What is more, model misspecification and parameter instability may undermine the performance of parametric time-varying VECM models. hansen2001new notes that it is unlikely that a structural break would be immediate and that it might be more reasonable to allow for a structural change to take a period of time to take effect. Therefore, models that allow for smooth changes over time may provide a more accurate representation of the dynamics in reality.
Having said that, we consider the time-varying vector error-correction model (VECM):
for $1\leq t\leq T$, where $\bm{\Pi}(\tau) = \bm{\alpha}(\tau)\bm{\beta}^\top$, $\bm{\alpha}(\tau)$ is the $d\times r_0$ adjustment coefficients, $\bm{\beta}$ is the $d\times r_0$ cointegration matrix, and $r_0$ is the cointegration rank to be determined by the dataset. Here, $\bm{\omega}(\cdot)$ governs the time-varying dynamics of the covariance matrix of $\{\mathbf{u}_t\}$, so it characterizes the permanent changes in unconditional volatility. Here, $\{ \mathbf{y}_t \}$ and $\{\Delta\mathbf{y}_t \}$ naturally connect two types of nonstationarity using a time-varying setup. Note that we assume the cointegration matrix $\bm{\beta}$ to be time-invariant, since there are ample empirical evidence showing that the short-run dynamics should be time-varying, while the long-run relationship between economic variables are quite stale (e.g., the expectations hypothesis of the term structure in hansen2003structural; long-run money demand theory in benati2021international). In some specific cases, such as testing the present-value theory for stock returns (e.g., campbell1987cointegration), the cointegration matrix is even known a priori. Mathematically, the decomposition of $\bm{\Pi}(\tau) = \bm{\alpha}(\tau)\bm{\beta}^\top$ is simply an identification restriction. Without any structure, it is impossible to identify the elements in the decomposition of $\bm{\Pi}(\tau)$ due to the multiplication form, so we go along with the existing literature by retaining that $\bm{\beta}$ is a time--invariant matrix.
Having presented our model in (ref), we comment on a challenge from the methodological perspective. In the extant literature of time-varying dynamic models, one often adopts the so-called “stationary approximation” technique in order to find a weakly dependent stationary approximation for time-varying dynamical processes before being able to establish asymptotics (e.g., dahlhaus1996kullback,chandler2012mode,zhang2012inference on time-varying AR models; dahlhaus2009empirical on time-varying ARMA models; dahlhaus2006statistical,truquet2017parameter on time-varying ARCH models; karmakar2021simultaneous on time-varying AR-ARCH models; gao2022time on time-varying VARMA-GARCH models). However, for (ref), $\Delta\mathbf{y}_t$ is expressed in terms of iterated time-varying functions, so the properties of $\Delta\mathbf{y}_t$ involve the integrated processes $\{\mathbf{y}_s\}_{s<t}$ and also depend on infinity past points $\{s/T\}_{s<t}$. As a consequence, how to approximate $\{\Delta\mathbf{y}_t\}$ and $\{ \mathbf{y}_t\}$ becomes challenging and the extant empirical process theories for stationary or locally stationary processes do not apply in this case, to the best of our knowledge.
In view of the aforementioned issues, in this study our contributions to the literature are as follows:
In our empirical analysis, we utilize the newly developed framework to investigate the rational expectations hypothesis of the interest rate term structure. Our results indicate that (1). the predictability of the term structure exhibits significant time-varying behaviour, and (2). the expectations hypothesis of the term structure holds periodically, particularly during periods of unusually high inflation. These findings lend support to the work of andreasen2021yield, who propose a macro-finance model of the term structure to explain variations in bond return predictability and identify the role of Federal Reserve monetary policy in stabilizing inflation as a key factor.
The rest of the paper is organized as follows. Section (ref) presents the estimation methodology and theory. In Section (ref), we conduct extensive simulation studies to examine the theoretical findings. Section (ref) investigates the rational expectations hypothesis of the U.S. term structure. Section (ref) concludes. All proofs are collected in the online supplementary Appendix A of the paper.
Before proceeding further, it is convenient to introduce some notations: $|\cdot|$ denotes the absolute value of a scalar or the spectral norm of a matrix; for a random variable $\mathbf{v}$, let $\|\mathbf{v}\|_q= (E|\mathbf{v}|^q )^{1/q}$ for $q\ge 1$; $\otimes$ denotes the Kronecker product; $\mathbf{I}_a$ stands for an $a\times a$ identity matrix; $\mathbf{0}_{a\times b}$ stands for an $a\times b$ matrix of zeros, and we write $\mathbf{0}_a$ for short when $a=b$; for a function $g(w)$, let $g^{(j)}(w)$ be the $j^{th}$ derivative of $g(w)$, where $j\ge 0$ and $g^{(0)}(w) \equiv g(w)$; let $\tilde{c}_k =\int_{-1}^{1} u^k K(u) \mathrm{d}u$ and $\tilde{v}_k= \int_{-1}^{1} u^k K^2(u) \mathrm{d}u$ for integer $k\ge 0$; $\mathrm{vec}(\cdot)$ stacks the elements of an $m\times n$ matrix as an $mn \times 1$ vector; $\mathbf{A}^{+}$ denotes the Moore-Penrose (MP) inverse of matrix $\mathbf{A}$; for a matrix $\mathbf{A}$ with full column rank, we let $\mathbf{P}_{\mathbf{A}} = \mathbf{A}(\mathbf{A}^\top \mathbf{A})^{-1}\mathbf{A}^\top$ and $\overline{\mathbf{A}} = \mathbf{A}(\mathbf{A}^\top\mathbf{A})^{-1}$; for $m\geq n$, we denote by $\mathbf{M}_{\perp}$ an orthogonal matrix complement of the $m\times n$ matrix $\mathbf{M}$ with $\mathrm{Rank}(\mathbf{M}) = n$; $\to_P$, $\to_D$ and $\Rightarrow$ denote convergence in probability, convergence in distribution, weak convergence with respect to the uniform metric; let $\mathbf{W}_d(\cdot,\bm{\Sigma})$ be a $d$-dimensional Brownian motion with covariance matrix $\bm{\Sigma}$.
In what follows, Section (ref) studies the properties of $\{\mathbf{y}_t \}$ and $\{\Delta\mathbf{y}_t \}$, and provide (integrated) time-varying VMA($\infty$) approximations for both processes. In Section (ref), we assume $r_0$ and $p_0$ are known for simplicity, and consider the estimation of $\bm{\alpha}(\cdot)$ and $\bm{\beta}$. Section (ref) proposes an information criterion to estimate the lag length (i.e., $p_0$), and a singular-value ratio test to determine the cointegration rank (i.e., $r_0$). Finally, Section (ref) gives a parameter stability test, which provides statistical evidence to support the necessity of time-varying VECM models practically.
To investigate (ref), the first obstacle lies in the complexity of the dependence structure of $\Delta\mathbf{y}_t$, which is expressed in terms of iterated time-varying functions with infinity memory. This implies that the population properites of $\Delta\mathbf{y}_t$ involve the integrated processes $\{\mathbf{y}_s\}_{s<t}$ and also depend on infinity past points $\{s/T\}_{s<t}$. As a result, the extant empirical process theory does not apply. To address this issue, we initiate our analysis by seeking local (not necessarily stationary) approximations for each $\Delta\mathbf{y}_t$ and $\mathbf{y}_t$, and then establish some necessary asymptotic properties for these processes.
We now introduce some necessary assumptions.
Assumption (ref).1 explicitly excludes explosive processes, and ensures that $\mathbf{y}_t$ is an integrated process of order one with $d-r_0$ common unit root components, as well as $r_0$ cointegration relationships. In Assumption (ref).1.c, $\bm{\alpha}_{\perp}^\top(\tau) [\mathbf{I}_d - \sum_{i=1}^{p_0-1}\bm{\Gamma}_i(\tau) ]\bm{\beta}_{\perp}$ being nonsingular ensures the existence of a time-varying Granger Representation Theorem for $\mathbf{y}_t$, which facilitates the asymptotic development later on.
Assumption (ref).2.a allows the underlying data generating process to evolve over time in a smooth manner. Assumption (ref).2.b regulates the behaviour of $\mathbf{y}_t$ for $t\leq 0$, which is standard in the literature of locally stationary models (e.g., vogt2012nonparametric) and unit-root processes (e.g., li2019kernel).
Assumption (ref).3 imposes conditions on the error innovations, which are standard in the time series literature (e.g., lutkepohl2005new).
With these conditions in hand, we are able to provide local approximations for each $\Delta\mathbf{y}_t$ and $\mathbf{y}_t$ with $t \geq 1$. For notational simplicity, we denote
in which $\mathbf{J} = [\bm{\beta},\overline{\bm{\beta}}_{\perp}]^\top$ and $\bm{\Gamma}_{\tau}(L)=\mathbf{I}_d - \sum_{i=1}^{p_0-1}\bm{\Gamma}_i(\tau)L^i$.
Lemma (ref) presents a local approximation of $\{\Delta \mathbf{y}_t\}$, which is the foundation of our asymptotic developments. Using Lemma (ref) we are able to establish basic properties of $\{\Delta \mathbf{y}_t\}$ and $\{ \mathbf{y}_t\}$. To put it in a nutshell, studying $\{\Delta \mathbf{y}_t\}$ and $\{ \mathbf{y}_t\}$ directly is technically challenging, we therefore consider their local approximations, which can help us understand $\{\Delta \mathbf{y}_t\}$ and $\{ \mathbf{y}_t\}$ in every small neighbourhood. In addition, Lemma (ref) indicates that $\widetilde{\mathbf{y}}_t(\tau)$ admits a time-varying version of Granger Representation Theorem (johansen1995likelihood, Theorem 4.2), so that $\Delta \widetilde{\mathbf{y}}_t(\tau)$ has a time-varying VMA($\infty$) representation. In the online supplementary Appendix A, we provide a comprehensive study on time-varying VMA($\infty$) processes (such as the Nagaev-type inequality, Gaussian approximation and the limit theorem for quadratic forms), which allows us to further derive estimation and inferential properties for (ref) in the following subsections.
We assume $p_0$ and $r_0$ are known in this subsection, and shall come back to work on their estimation in Section (ref).
The local linear estimator of $[\bm{\Pi}(\tau),\bm{\Gamma}(\tau)]$ with $\bm{\Gamma}(\tau) = [\bm{\Gamma}_1(\tau),\ldots,\bm{\Gamma}_{p_0-1}(\tau)]$ is given by
where $\mathbf{h}_t = \left[\mathbf{y}_{t}^\top,\Delta \mathbf{x}_t^\top\right]^\top$,
Here, we use MP inverse since $\mathbf{S}_{T,l}(\tau)$ is asymptotically singular (cf., Lemma (ref) of the supplementary Appendix A), which is referred to as “kernel-induced degeneracy” in the unit-root literature (phillips2017estimating).
Accordingly, the local linear estimator of $\bm{\Omega}(\tau)$ is defined as $$ \widehat{\bm{\Omega}}(\tau) = \frac{1}{T}\sum_{t=1}^{T}\widehat{\mathbf{u}}_t\widehat{\mathbf{u}}_t^\top w_t(\tau), $$ where $\widehat{\mathbf{u}}_t = \Delta \mathbf{y}_t - \widehat{\bm{\Pi}}(\tau_t)\mathbf{y}_{t-1} - \widehat{\bm{\Gamma}}(\tau_t)\Delta \mathbf{x}_{t-1}$, $\Delta \mathbf{x}_{t} = [\Delta \mathbf{y}_{t}^\top,\ldots,\Delta \mathbf{y}_{t-p_0+2}^\top]^\top$,
To proceed, we require the following conditions to hold.
Assumption (ref).1 imposes a set of regular conditions on the kernel function and the bandwidth. On top of Assumption (ref).3, Assumption (ref).2 imposes more structure on $\{\bm{\varepsilon}_t\}$, and is used to derive Gaussian approximation for the sum of time-varying VECM process. This condition can be relaxed if we impose more dependence structure (e.g., nonlinear system theory as in wu2005nonlinear). Provided $\delta>5$, the usual optimal bandwidth $h_{opt}=O (T^{-1/5} )$ satisfies the condition $T^{1-\frac{4}{\delta}}h/(\log T)^{1-\frac{4}{\delta}} \to \infty$. In addition, Gaussian approximation together with $T^{1-\frac{4}{\delta}}h/(\log T)^{1-\frac{4}{\delta}} \to \infty$ is used to derive the uniform convergence of our nonparametric estimators, which is further used for asymptotic covariance estimation, semiparametric estimation and model specification testing in the below.
With these conditions in hand, we further let
and summarize the asymptotic properties of the local linear estimators in the following theorem.
Theorem (ref) establishes the asymptotic distribution yielded by $[\widehat{\bm{\Pi}}(\tau), \widehat{\bm{\Gamma}}(\tau)]$, and also gives the rates of uniform convergence which will be extensively used in the following development.
In what follows, we consider the estimation of cointegration matrix by utilizing the reduced rank structure of $\bm{\Gamma}(\cdot)$ and using the profile likelihood method (e.g., fan2005profile). To ensure a unique cointegration matrix, we assume
where $\bm{\beta}^*$ is a $(d-r_0)\times r_0$ matrix. Using (ref), $\widehat{\bm{\alpha}}(\tau)$ is the first $r_0$ columns of $\widehat{\bm{\Pi}}(\tau)$, i.e., $\widehat{\bm{\Pi}}(\tau) = [\widehat{\bm{\alpha}}(\tau) , \widehat{\bm{\Pi}}_2(\tau)],$ where the definition of $\widehat{\bm{\Pi}}_2(\tau)$ is obvious.
Specifically, given $\bm{\Pi}(\tau)$, we can estimate the short-run time-varying parameters $\bm{\Gamma}(\tau)$ by
where $\Delta\mathbf{x}_{t}^{*} = \Delta\mathbf{x}_{t}\otimes\left[1,\frac{\tau_{t+1}-\tau}{h}\right]^\top$. In connection with (ref), we can write $$ \widetilde{\mathbf{r}}_t(\bm{\alpha}) = \widetilde{\mathbf{R}}_t^\top(\bm{\alpha}) \mathrm{vec}\left(\bm{\beta}^{*,\top}\right) + \mathbf{u}_t^*, $$ where $\mathbf{u}_t^* = \mathbf{u}_t + \left[\bm{\Gamma}(\tau_t)-\widehat{\bm{\Gamma}}(\tau_t,\bm{\Pi})\right]\Delta \mathbf{x}_{t-1}$, $\mathbf{r}_t(\bm{\alpha}) = \Delta \mathbf{y}_t - \bm{\alpha}(\tau_t)\mathbf{y}_{t-1}^{(1)}$, $\mathbf{y}_{t}^{(1)}$ contains the first $r_0$ elements of $\mathbf{y}_{t}$, $\mathbf{r}^\top(\bm{\alpha}) = [\mathbf{r}_1(\bm{\alpha}),\ldots,\mathbf{r}_T(\bm{\alpha})]$, $$ \widetilde{\mathbf{r}}_t(\bm{\alpha}) = \mathbf{r}_t(\bm{\alpha}) - \mathbf{r}^\top(\bm{\alpha}) \mathbf{K}(\tau_t) \Delta \mathbf{x}^*\left(\Delta \mathbf{x}^{*,\top}\mathbf{k}(\tau_t) \Delta \mathbf{x}^*\right)^{-1}[\mathbf{I}_{d(p_0-1)},\mathbf{0}_{d(p_0-1)}]^\top\Delta\mathbf{x}_{t-1}, $$ $\mathbf{k}(\tau)= \mathrm{diag}\left[K\left(\frac{\tau_1-\tau}{h}\right),\ldots,K\left(\frac{\tau_T-\tau}{h}\right)\right]$, $\Delta \mathbf{x}^{*,\top} = \left[\Delta \mathbf{x}_0^*,\ldots,\Delta \mathbf{x}_{T-1}^* \right]$, $\mathbf{R}_t(\bm{\alpha}) = \mathbf{y}_{t-1}^{(2)} \otimes \bm{\alpha}^\top(\tau_t)$, $\mathbf{y}_{t}^{(2)}$ contains the last $d-r_0$ elements of $\mathbf{y}_{t}$, $$ \widetilde{\mathbf{R}}_t(\bm{\alpha}) = \mathbf{R}_t(\bm{\alpha}) -\mathbf{R}^\top(\bm{\alpha})\mathbf{K}(\tau_t)\Delta \mathbf{X}^{*}\left(\Delta \mathbf{X}^{*,\top}\mathbf{K}(\tau_t) \Delta \mathbf{X}^*\right)^{-1}\left[\mathbf{I}_{d^2(p_0-1)},\mathbf{0}_{d^2(p_0-1)} \right]^\top \Delta\mathbf{X}_{t-1}, $$ $\mathbf{R}^\top(\bm{\alpha}) = [\mathbf{R}_1(\bm{\alpha}),\ldots,\mathbf{R}_T(\bm{\alpha})]$, $\Delta\mathbf{X}^{*,\top}=\left[\Delta \mathbf{X}_0^*,\ldots,\Delta \mathbf{X}_{T-1}^*\right]$, $\Delta\mathbf{X}_t = \Delta\mathbf{x}_t \otimes \mathbf{I}_d$, $\Delta\mathbf{X}_t^* = \Delta\mathbf{x}_t^* \otimes \mathbf{I}_d$ and $\mathbf{K}(\tau) = \mathbf{k}(\tau)\otimes \mathbf{I}_d$. Replacing $\bm{\alpha}(\tau)$ with $\widehat{\bm{\alpha}}(\tau)$, the weighted least squares (WLS) estimator of $\bm{\beta}^*$ is given by
The next theorem summaries the asymptotic distribution associated with $\widehat{\bm{\beta}}^{*}$.
The first result shows that the cointegration matrix can be estimated at a super consistent rate $T$, while the second result concerning asymptotic normality of a properly standardized from of $\widehat{\bm{\beta}}^* - \bm{\beta}^*$ indicates how to construct confidence interval practically.
We now consider the estimation of $p_0$ and $r_0$. Specifically, we first propose an information criterion that can estimate the lag length ($p_0$), and then consider the estimation of cointegration rank ($r_0$).
To estimate $p_0$, we minimize an information criterion as follows:
where $ \text{IC}(p) = \log \left\{\text{RSS}(p)\right\}+p\cdot\chi_T$, $\text{RSS}(p)=\frac{1}{T}\sum_{t=1}^{T}\widehat{\mathbf{u}}_{p,t}^\top \widehat{\mathbf{u}}_{p,t}$, $\chi_T$ is the penalty term, $\widehat{\mathbf{u}}_{p,t}$ is the value of $\widehat{\mathbf{u}}_{t}$ by letting the number of lagged differences be $p-1$, and $P$ is a sufficiently large fixed positive integer. The following proposition summaries the asymptotic property of (ref).
It is noteworthy that Theorem (ref) does not require any knowledge of cointegration rank. In view of the conditions on $\chi_T$, a natural choice is $$ \chi_T = \frac{\log(\log(Th))}{3} \left(h^4 +h^2\left(\frac{\log T}{Th}\right)^{1/2} + \frac{\log T}{Th}\right). $$
We next consider the estimation of cointegration rank $r_0$. The basic principle of our method is to separate the $r_0$ relevant singular values of $\bm{\Pi}(\tau)$ from the zero ones, while the number of nonzero ones corresponds to the cointegration rank. Before proceeding further, it should be pointed out that the choice of lag length $p_0$ is irrelevant with the determination of cointegration rank since we can set the lag length $P$ as sufficiently large but fixed (i.e., $p_0\leq P$), in which case $\bm{\Gamma}_j(\tau) = \mathbf{0}_d$ for $j=p_0,...,P$.
The method is based on the QR decomposition with column-pivoting of $\int_{0}^{1}\bm{\Pi}^\top(\tau)\mathrm{d}\tau$, i.e., $\int_{0}^{1}\bm{\Pi}^\top(\tau)\mathrm{d}\tau = \bm{\beta}\int_{0}^{1}\bm{\alpha}^\top(\tau)\mathrm{d}\tau = \mathbf{S}\mathbf{R}$, where $\mathbf{S}^\top\mathbf{S}=\mathbf{I}_d$, and $\mathbf{R}$ is an upper triangular matrix with the diagonal elements being nonincreasing. The estimator of $\int_{0}^{1}\bm{\Pi}(\tau)\mathrm{d}\tau$ is naturally given by $\int_{0}^{1}\widehat{\bm{\Pi}}(\tau)\mathrm{d}\tau$ and its QR decomposition with column-pivoting is defined as
where $\widehat{\mathbf{S}}$ is $d\times d$ orthonormal, and the partition of the second step should be obvious so we omit the descriptions for each block. By Theorem (ref), $\int_{0}^{1}\widehat{\bm{\Pi}}(\tau)\mathrm{d}\tau\to_P \int_{0}^{1}\bm{\Pi}(\tau)\mathrm{d}\tau .$ Therefore, $\widehat{\mathbf{R}}_{22}$ is expected to be small, which motivates the use of following procedure. Let $\widehat{\mu}_k = \sqrt{\sum_{j=k}^{d} \widehat{\mathbf{R}}^2(k,j)}$, where $\widehat{\mathbf{R}}(k,j)$ denotes the element of $k^{th}$ row and $j^{th}$ column of $\widehat{\mathbf{R}}$.
We consider the following singular value ratio test, taking a suggestion from the literature (e.g., lam2012factor,zhang2019identifying):
where $\widehat{\mu}_0 = \widehat{\mu}_1 + w_T$ is the “mock” singular value since $\frac{\widehat{\mu}_r}{\widehat{\mu}_{r+1}}$ is not defined for $r=0$, $w_T =\frac{\log T}{Th} \log(\log (Th)) $ and $ \widehat{\mu}_1\ge\cdots \ge\widehat{\mu}_d$. In addition, the indicator function is used to ensure that the estimator $\widehat{r}$ is consistent. Note that similar to the case of eigenvalue ratio test for the factor model, both the numerator and denominator converge to zeros at the same rate when $r>r_0$. Therefore, if $\widehat{\mu}_r$ is “small”, we take it as a sign of $r > r_0$ and set $\widehat{\mu}_r/\widehat{\mu}_{r+1}$ to one.
The following theorem summaries the asymptotic property of (ref).
If $r_0 = 0$, $\mathbf{y}_t$ is a pure unit root process with time-varying vector autoregressive errors. Hence, our procedure is also able to test the existence of cointegration relationship.
Practically, it is necessary to test whether the coefficients of (ref) are time-varying before applying the aforementioned framework. Formally, we consider a hypothesis test of the form:
where $\mathbf{b}(\tau) = \mathrm{vec} (\bm{\alpha}(\tau),\bm{\Gamma}(\tau) )$, $\mathbf{C}$ is a selection matrix of full row rank, and $s$ is the number of restrictions. The choice of $\mathbf{C}$ and $\mathbf{c}$ should be theory/data driven. For example, one can let $\mathbf{C} =\left[\mathbf{I}_{r_0},\mathbf{0}_{r_0\times (d-1)r_0+d^2(p_0-1)}\right]$ and $\mathbf{c} = \mathbf{0}$ to test whether there exists an error-correction term for $\Delta y_{1,t}$ over the long-run.
The test statistic is constructed based on the weighted integrated squared errors:
where $\widehat{\mathbf{b}}(\tau) = \mathrm{vec} (\widehat{\bm{\alpha}}(\tau),\widehat{\bm{\Gamma}}(\tau) )$ should be obvious, and $\widehat{\mathbf{c}} = \frac{1}{T}\sum_{t=1}^{T}\mathbf{C}\widehat{\mathbf{b}}(\tau_t)$ is the semiparametric estimator of $\mathbf{c}$. In (ref), $\mathbf{H}(\cdot)$ is an $s\times s$ positive definite weighting matrix, and is typically set as the precision matrix associated with $\widehat{\mathbf{b}}(\cdot)$. We present the asymptotic distribution of the semiparametric estimator $\widehat{\mathbf{c}}$ and the proposed test in the following theorem.
In Theorem (ref).1, the bias term $\frac{1}{2}h^2\widetilde{c}_2\int_{0}^{1}\mathbf{C}\mathbf{b}^{(2)}(\tau)\mathrm{d}\tau$ vanishes under $\mathbb{H}_0$, and thus the parametric component in the corresponding semiparametric model can have a $\sqrt{T}$-consistent estimate $\widehat{\mathbf{c}}$. Theorem (ref).2 states that the test statistic converges to a normal distribution and is asymptotically pivotal. The bias term $s\widetilde{v}_0$ can easily be calculated for any given kernel function, and it arises due to the quadratic form of the test statistic.
To close our theoretical investigation, we note further that the online supplementary Appendix (ref) provides a local alternative of the parameter stability test, and also gives a simulation-assisted testing procedure to improve the finite sample performance of the test.
In this section, we first provide some details of the numerical implementation in Section (ref), and then respectively examine the estimation and hypothesis testing in Sections (ref) and (ref).
Throughout the numerical studies, Epanechnikov kernel (i.e., $K(u)=0.75(1-u^2)I(|u|\leq 1)$) is adopted. The optimal lag length and the cointegration rank are estimated based on (ref) and (ref) respectively. For each given $p$ of (ref), the bandwidth $\widehat{h}_{cv}$ is always chosen by minimizing the following leave-one-out cross-validation criterion function:
where $\widehat{\bm{\Pi}}_{-t}(\cdot)$ and $\widehat{\bm{\Gamma}}_{j,-t}(\cdot)$ are obtained based on the local linear estimator of Section (ref) but leaving the $t^{th}$ observation out. Once $\widehat{p}$, $\widehat{r}$ and $\widehat{h}_{cv}$ are obtained, the estimation procedure is relatively straightforward. As shown in richter2019cross, the leave-one-out cross validation method works well as long as the error terms are uncorrelated, which implies that this desirable property should hold in our case.
The data generating process (DGP) is as follows:
where $T\in \{200, 400, 800\}$, $\bm{\varepsilon}_t$'s are i.i.d. draws from $N(\mathbf{0}_{2\times 1}, \mathbf{I}_2)$, and
In this case, we have $p_0 = 2$. To test the null hypothesis of no cointegration relations, we consider two sets of $\bm{\alpha}(\tau)$ and $\bm{\beta}$:
For each generated dataset, we carry on the methodologies documented in Section (ref), and conduct 1000 replications.
First, we evaluate the performance of the lag length selection procedure (i.e., (ref)), and report the percentages of $\widehat{p} < 2$, $\widehat{p} = 2$, and $\widehat{p} > 2$ respectively based on 1000 replications. Table 1 shows that the information criterion (ref) performs reasonably well, as the percentages associated with $\widehat{p}=2$ are sufficiently close to 1 except for the case with $T=200$. In addition, this information criterion works well in both time-varying VECM (DGP 1) and time-varying VAR models (DGP 2).
Next, we evaluate the performance of the conintegration rank estimator (ref) and report the percentages of $\widehat{r} = 0$ and $\widehat{r} = 1$ respectively based on 1000 replications. Table 2 shows that the singular value ratio method of (ref) performs reasonably well. However, when $r_0=0$, the estimator (ref) tends to identify a false cointegration relationship for small sample size (i.e., $T=200$).
Finally, we evaluate the estimates of $\bm{\alpha}(\tau)$, $\bm{\beta}$ and $\bm{\Gamma}_1(\tau)$ for DGP 1, and calculate the root mean square error (RMSE) as follows
for $\bm{\theta}(\cdot)\in\left\{\bm{\alpha}(\cdot),\bm{\Gamma}(\cdot)\right\}$, where $\widehat{\bm{\theta}}^{(n)}(\tau)$ is the estimate of $\bm{\theta}(\tau)$ for the $n$-th replication. Of interest, we also examine the finite sample coverage probabilities of the confidence intervals based on our asymptotic theories. In the following, we compute the average of coverage probabilities for grid points in $\{\tau_t,t=1,\ldots,T\}$, and then further take an average across the elements of $\bm{\theta}(\cdot)$. The RMSEs and empirical coverage probabilities are reported in Table 3, which reveals several notable points. First, the RMSE decreases as the sample size goes up. Second, the RMSE of $\bm{\beta} $ is much smaller than those of $\bm{\alpha}(\tau)$ and $\bm{\Gamma}_1(\tau)$, which should be expected. Third, the finite sample coverage probabilities are smaller than their nominal level (95%) for small $T$, but are fairly close to 95% as $T$ increases.
To evaluate the size and local power of the proposed test statistic, we consider the following DGP:
where $\bm{\beta}$, $\bm{\Gamma}_1(\cdot)$ and $\mathbf{u}_t$ are generated in the same way as DGP 1 in Section (ref), and $$ \bm{\alpha}(\tau)=\left[
\right]+b\times d_T\times \left[
\right], $$ in which $d_T = T^{-1/2}h^{-1/4}$ and $b$ is set to be $0$, $1$ or $2$ in order to investigate the size and local power of the proposed test. We use the proposed testing procedure to test whether the coefficient $\bm{\alpha}(\cdot)$ is time-varying. Again, we let $T\in \{200,400,800\}$ and conduct $1000$ replications for each choice of $T$. We use the simulation-assisted testing procedure of Appendix \ref{App0} to get the empirical critical value $\widehat{q}_{1-\alpha}$ after $1000$ bootstrap replications. We consider a sequence of bandwidths to check the robustness of the proposed test with respect to a sequence of bandwidths:
Table 4 reports the rejection rates at the $5\%$ and $10\%$ nominal levels. A few facts emerge. First, our test has reasonable sizes using the empirical critical values obtained by the bootstrap procedure if the sample size is not so small. Second, the size behaviour of our test is not sensitive to the choices of bandwidths. As discussed in gao2008bandwidth, the estimation-based optimal bandwidths may also be optimal for testing purposes, so for simplicity we use the cross-validation based bandwidth or the rule-of-thumb bandwidth in practice. Third, the local power of our test increases rapidly as $b$ increases.
In this section, we assess the time-varying predictability of the term structure (i.e., the yield curve) of interest rates, and investigate whether the expectations hypothesis of the term structure holds periodically by using the proposed time-varying VECM model, which allows for shifts in the predictability of the term structure.
We now briefly review the literature. The term structure is crucial to both monetary policy analysis and private individuals. According to the rational expectations hypothesis (campbell1987cointegration), the term structure (or term spread) should provide information on the future changes in both short-term and long-term interest rates. For example, if a long bond yield exceeds a short yield, the long rate subsequently tends to rise, which generates expected capital losses on the long bond and thus offsets the current yield advantage. This also implies that bond returns are predictable from the yield spread, and the expected bond returns and the yield spread should be negatively correlated.
The literature on the term structure and bond return predictability is enormous and it continues to expand (e.g., campbell1987cointegration,bauer2020interest,andreasen2021yield,vayanos2021preferred,he2022treasury). However, the existing results on the expectations hypothesis of the term structure present many discrepancies, which may be due to the fact that the relationship evolves with time. For example, borup2021predicting document that bond return predictability depends on the economic states, while linear forecasting models yield little evidence of unconditional predictability. Along this line of research, one important question is that whether the U.S. monetary and financial system has changed over time so that estimates from historical data are unreliable for modern policy analysis and investment activities. Although the literature has begun to explore whether the predictability of the term structure depends on the economy states (e.g., andreasen2021yield,borup2021predicting), few studies aim to quantify the varying predictability of the term structure over time. In what follows, we address this issue using the newly proposed framework. The estimation procedure is conducted in exactly the same way as in Section (ref), so we no longer repeat the details.
To study the time-varying predictability of the term structure, we consider the following time-varying bivariate VECM($p_0-1$) model: $$ \Delta \mathbf{y}_t = \bm{\alpha}(\tau_t)\bm{\beta}^\top\mathbf{y}_{t-1} + \sum_{j=1}^{p_0-1}\bm{\Gamma}_j(\tau_t)\Delta \mathbf{y}_{t-j} + \mathbf{u}_t,\quad \mathbf{u}_t = \bm{\omega}(\tau_t)\bm{\varepsilon}_t, $$ where $\mathbf{y}_t^\top = [l_t, s_t]$, $s_T$ is the interest rate on a one-period bond and $l_t$ is the interest rate on a multi-period bond. According to the expectations hypothesis of the term structure, $l_t$ and $s_t$ should be integrated with $\bm{\beta}^\top = [1,-1]$ (and thus $\beta^* = -1$ under the identification condition), while the error-correction term is the term spread. {In addition, the adjustment coefficient $\bm{\alpha}(\tau)$ measures the predictability of the term structure and the sign of the elements in $\bm{\alpha}(\tau)$ should be positively significant according to the theory. }
We use a selection of bond rates with maturities ranging from 1 to 5 years. The interest rates are estimated from the prices of U.S. Treasury securities and correspond to zero-coupon bonds. The data are monthly observations from 1961:M6 to 2022:M12, which are collected from Nasdaq Data Link at \url{https://data.nasdaq.com}. Figure (ref) plots these five variables.
We begin by presenting the results using a constant parameter VECM model in order to get some intuition, which is also served as a benchmark for evaluating our time-varying VECM model. Recall that one of the economic implications of the expectations hypothesis of the term structure is that the long and short rates should be cointegrated with $\bm{\beta}=[1,-1]^\top$. To empirically test this implication, we first detect the presence of cointegration, using the Johansen's likelihood ratio test. The testing results are reported in Table 5. From Table 5, for all possible bivariate pairs and lag lengths considered, the tests strongly reject the null of no integration, but do not reject the null of a single cointegration relationship at the 5% significance level.
We then report the parameter estimates for constant VECM models, while the optimal lag $p$ is set to be $2$ according to the literature (e.g., hansen2003structural). Table 6 reports the parameter estimates as well as their 95% confidence intervals, in which $\alpha_i$ denotes the $i^{th}$ elements of adjustment coefficient $\bm{\alpha}$. From Table 6, we can see that for all considered bivariate pairs, the estimated cointegration parameter $\widehat{\beta}^*$ is quite close to unity, which is consistent with the theory. {However, for all bivariate pairs considered, the estimates of $\widehat{\alpha}_1$ and $\widehat{\alpha}_2$ are not positively significant (or even negative), which contradicts to the economic implications of the term structure theory. Note that according to the expectations hypothesis, the sign of adjustment coefficients $\bm{\alpha}$ should be positive. These results are in line with borup2021predicting, who find that linear forecasting models yield little evidence of unconditional bond return predictability.}
In order to solve the puzzle raised by constant parameter VECM models, we then investigate the possibility that the expectations hypothesis of the term structure holds periodically. This is mainly motivated by the following facts. First, several studies have shown that the bond return predictability depends on the economic states (e.g., andreasen2021yield,borup2021predicting). Second the rational expectations hypothesis also implies that the future bond returns should be negatively correlated with the current term spread (i.e., $l_t-s_t$). The above two facts mean that the predictability of term structure shifts over time and the above puzzling contradictions may be explained by using time-varying VECM models.
Table 7 report the estimation and testing results of time-varying VECM models, as well as some robustness checks. For the time-varying VECM($p-1$) model, the optimal lag is $\widehat{p} = 2$ or $\widehat{p} = 3$ by our approach, which is consistent with the literature (e.g., hansen2002testing). We further check whether the model coefficients are indeed time-varying. We employ the proposed test statistic to examine the constancy of model coefficients. The associated $p$-value for these considered bivariate pairs are around 0.000--0.001, which suggest that we should choose the time-varying VECM model over a constant one. Certainly, one may examine each element of these coefficient matrices. However, it will lead to a quite lengthy presentation. In order not to deviate from our main goal, we no longer conduct more testing along this line. These testing results also suggest that the predictability of term structure are time-varying for a range of short and long rates. We then conduct robustness check to see whether the error innovations $\{\bm{\varepsilon}_t\}$ exhibit serial correlation. We use the multivariate version of Breusch-Godfrey LM test (godfrey1978testing) to test the serial-correlations in the innovations $\{\bm{\varepsilon}_t\}$, in which the null hypothesis is $H_0: E(\bm{\varepsilon}_t\bm{\varepsilon}_{t+1}^\top) = \mathbf{0}$. Based on the estimates $\widehat{\bm{\varepsilon}}_t = \widehat{\bm{\Omega}}^{-1/2}(\tau_t)\widehat{\mathbf{u}}_t$, the corresponding $p$-value ranges from 0.559 to 0.945, suggesting that the time-varying VECM model fits the data quite well for all considered bivariate pairs.
We then apply the singular ratio test to detect the presence of cointegration and then check whether $\beta^* = -1$. Table 7 show that the estimates of cointegration rank are $1$ (i.e., $\widehat{r} = 1$) for all considered bivariate pairs. Based on these singular ratio tests, we find strong evidence of the existence of cointegration, which is consistent with Figure (ref) and the results of constant parameter VECM models. In addition, this result, i.e., $\widehat{r} = 1$, is robust to different choices of $p$ ranging from $2$--$6$. We also report the point estimates of $\beta^*$ and their 95% confidence intervals in Table 7. According to Table 7, we find a long-run relationship between the long rate and the short rate, and we cannot reject the null that $H_0: \beta^*=-1$ at a 5% significance level for most considered bivariate pairs. Interestingly, we find that the estimates of $\beta^*$ between these two models are almost identical, while the time-varying VECM models yields quite narrow confidence bands (i.e., smaller standard errors).
Finally, we investigate whether the term structure (or bond returns) is predictable over the long-run by using the term spread as a predictor, i.e., $\bm{\alpha}(\tau) = \mathbf{0}$, and we also examine the sign of $\bm{\alpha}(\tau)$, which should be positive according to the expectations hypothesis. {In order to confirm whether the error--correction component is significant, We test the null hypothesis $\mathbb{H}_0:\ \bm{\alpha}(\tau) = \mathbf{0}$. The testing results are reported in the last column of Table 7, which indicates that we should reject the null at all conventional levels. Therefore, the term structure is predictable in the long-run, at least in some local periods, while the constant parameter VECM model yields little evidence of the predictability of the term structure.} We also investigate the time-varying pattern of the term structure predictability. Figure (ref) plots the estimates of $\alpha_1(\cdot)$ and $\alpha_2(\cdot)$ and their 95% point-wise confidence intervals, as well as the U.S. core inflation. Here, the core inflation data are monthly observations from 1967:M1 to 2022:M12, collected from the Federal Reserve Bank of St. Louis economic database.
{In summary, our finding of great interest is that the estimated error-correction effects for the short and long rates are only positively significant in the 1980s and in recent years, and vary significantly over time for these bivariate pairs considered.} What is more is that the time-varying patterns of $\alpha_1(\cdot)$ and $\alpha_2(\cdot)$ are almost identical, and are similar to the pattern of time-variations in U.S. core inflation rate. Our findings suggest that the expectations hypothesis holds periodically, especially in the period of unusual high inflation. Importantly, our results also provide us with evidence for the economic implications of the theoretical macro-finance term structure model proposed by andreasen2021yield, who find that the monetary policy decisions by the Federal Reserve with respect to stabilizing inflation is a key driver of this switch in bond return predictability.
In this paper, we propose a time-varying vector error-correction model that allows for different time series behaviours (e.g., unit-root and locally stationary processes) interacting with each other to co-exist. From practical perspectives, this framework can be used to estimate shifts in the predictability of non-stationary variables, and test whether economic theories hold periodically. We first develop a time-varying Granger Representation Theorem, which facilitates establishing asymptotic properties, and then propose estimation and inferential theories for both short-run and long-run coefficients. We also propose an information criterion to estimate the lag length, a singular-value ratio test to determine the cointegration rank, and a hypothesis test to examine the parameter stability. To validate the theoretical findings, we conduct extensive simulations. Finally, we demonstrate the empirical relevance by applying the framework to investigate the rational expectations hypothesis of the U.S. term structure. We conclude that the predictability of the term structure vary significantly over time and the expectations hypothesis of the term structure holds periodically, especially in the period of unusual high inflation.
Gao and Peng acknowledge financial support from the Australian Research Council Discovery Grants Program under Grant Numbers: DP200102769 & DP210100476, respectively. Yan acknowledges the financial support by Fundamental Research Funds for the Central Universities (Grant Numbers: 2022110877 & 2023110099).
{
}
\setcounter{page}{1} \linespread{1.1}
This file is organised as follows. In Appendix (ref), we present some extra results on parameter stability test, which include the study on a sequence of local alternatives, and a simulation-assisted numerical procedure to improve finite sample performance. Appendix (ref) includes the notation and mathematical symbols, which will be repeatedly used throughout the development. In Appendix (ref), we provide the preliminary lemmas to facilitate the development of the main theorems of the paper. Appendix (ref) includes the proofs of the main results, while Appendix (ref) covers the proofs of the preliminary lemmas.