Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
74,404 characters · 13 sections · 0 citation commands
Adaptive information-based methods for determining the co-integration rank in heteroskedastic VAR models
\linespread{1.1}
\thispagestyle{empty}
\linespread{1.3}
It is well-known that standard methods for determining the co-integration rank of vector autoregressive (VAR) systems of variables integrated of order one are affected by the presence of heteroskedasticity. In particular, sequential procedures based on (pseudo-) likelihood ratio [PLR] test as developed by Johansen (1996) can be significantly over-sized, even in large samples, when the volatility process displays non-stationary variation (so called non-stationary unconditional volatility) and, moreover, the finite sample power of these tests can vary enormously depending on the pattern of heteroskedasticity present; see, in particular, Cavaliere, Rahbek and Taylor (2010). This is an important issue in practice because time-varying behaviour in unconditional volatility appears to be a common feature in many key macroeconomic and financial time series; see, among many others, McConnell and Perez Quiros (2000), Sensier and van Dijk (2004), and Cavaliere and Taylor (2008); see also McAleer (2005, 2009), Asai et al., (2006) and McAleer and Medeiros (2008).
In a series of recent papers, Cavaliere, Rahbek and Taylor (2010, 2014) show that a solution to the size problems induced by non-stationary volatility is obtained by using wild bootstrap based implementations of the standard PLR tests. In particular, Cavaliere et al.\ (2010) show that the sequential procedure based on wild bootstrap PLR tests leads to consistent co-integration rank determination in the presence of non-stationary unconditional volatility. As alternative solution to the use of wild bootstrap PLR tests is considered by Cavaliere, De Angelis, Rahbek and Taylor (2015, 2018) who show that methods based on information criteria can also be used to consistently determine the co-integration rank in the presence of non-stationary volatility. In particular, they show that popular information criteria such as the Bayesian information criterion [BIC] (Schwarz, 1978) and the Hannan-Quinn information criterion [HQC] (Hannan and Quinn, 1979) provide a useful complement to the wild bootstrap sequential procedures.
The wild bootstrap PLR tests are correctly sized in the presence of non-stationary volatility and attain the same asymptotic local power functions as infeasible size-corrected versions of the standard PLR tests. As such they can therefore display very low power properties for some patterns of non-stationary volatility. Indeed, other things equal, their asymptotic local power functions are reduced, relative to the unconditionally homoskedastic case, under non-stationary volatility. Similarly, the ability of the standard information criteria-based methods discussed above to select the correct co-integration rank can also be greatly reduced under non-stationary volatility. In particular, none of these methods exploits the potential efficiency gains that could be provided by using inference methods which adapt to the volatility process. Adaptive methods, where the covariance matrix process is estimated non-parametrically, have the potential to be particularly useful in this context.
Under the assumption of a known autoregressive lag length, Boswijk and Zu (2022) develop an procedure based on adaptive PLR tests for determining the co-integration rank in possibly heteroskedastic VAR models. Specifically, they propose a procedure where the volatility process is estimated using a non-parametric kernel estimator, with this estimate then used in the adaptive PLR test procedure. Under suitable conditions, they establish that the non-parametric volatility estimator is consistent and that the resulting adaptive PLR co-integration rank tests have the same asymptotic local power functions as for infeasible tests based on the assumption that the volatility process is known. The asymptotic null distribution of their proposed statistics are, however, non-standard and depend on the realisation of the volatility process. As such, asymptotic $p$-values for the adaptive PLR tests need to be obtained using bootstrap methods.
The assumption of a known of autoregressive lag order is problematic in practice. It is well-known that an incorrect lag length choice can significantly impact on the efficacy of both information criteria and PLR tests, in particular where a lag order smaller than the true order is used; see, among others, Boswijk and Franses (1992), Cheung and Lai (1993), Haug (1996), L\"{u}tkepohl and Saikkonen (1999), and Cavaliere et al.\ (2018). In practice the autoregressive lag length will need to be estimated along with the co-integration rank. To that end, the practitioner can use either a sequential procedure, where the lag length is consistently estimated in a first step and then subsequently employed in the second step in a procedure such as either the adaptive PLR test approach of Boswijk and Zu (2022) or an information criterion for determining the co-integration rank, or a joint information criteria-based approach can be used whereby the lag length and co-integration rank are determined simultaneously. Cavaliere et al.\ (2018) show that both joint and sequential procedures based on standard information criteria consistently determine both the lag length and the co-integration rank in the presence of non-stationary unconditional volatility, provided standard conditions hold on the penalty term. They also show the asymptotic validity of a sequential procedure based on wild bootstrap PLR tests with the autoregressive lag length chosen by an information criterion.
The contribution of this paper is to develop adaptive information criteria methods, based around a (non-parametric) estimation of the volatility process, for jointly selecting the co-integration rank and autoregressive lag order. We show that these adaptive information criteria-based methods are weakly consistent for the co-integration rank and autoregressive lag order under the precisely the same conditions on the penalty function are as required for the consistency of standard (non-adaptive) information criteria under non-stationary volatility of the form considered in this paper. We also establish the asymptotic validity of a sequential procedure selecting the autoregressive lag length by an adaptive information criterion [ALS-IC] in the first step and then determining the co-integration rank using again an ALS-IC in the second step based on the first step estimate of the lag length. Because the co-integration rank is determined by minimising an adaptive information criterion over all possible values of the co-integration rank from zero up to the dimension of the system, the practitioner does not therefore need to obtain $p$-values by bootstrap methods, making the procedure considerably less time consuming than the Boswijk and Zu (2022) procedure based on adaptive PLR tests. We also establish the asymptotic validity of a sequential procedure selecting the autoregressive lag length by an ALS-IC in the first step and then using the adaptive PLR test-based approach of Boswijk and Zu (2022) in the second step based on the first step estimate of the lag length.
The remainder of the paper is organised as follows. Section (ref) details our reference heteroskedastic co-integrated VAR model. Section (ref) outlines adaptive information criteria-based methods for determining the co-integration rank and the autoregressive lag length. The large sample properties of these procedures are detailed in Section (ref). Monte Carlo simulation experiments reported in Section (ref) are used to explore the finite sample performance of the ALS-IC methods relative to standard methods such as those based on standard information criteria-based procedures. These results highlight the potential gains that can be achieved by using adaptive methods. Section (ref) provides an empirical application of the methods discussed in this paper to the term structure of interest rates in the US. Section (ref) concludes. Proofs of our main results are contained in the Appendix (ref).
Consider the $p$-dimensional process $\{ X_{t} \} $ which satisfies the $k$ -th order reduced rank VAR model:
where $X_t := ( X_{1t} ,\ldots, X_{pt})^\prime $ and the initial values, $ X_{1-k},\ldots,X_0 $, are taken to be fixed in the statistical analysis. Let $ k_0$ denote the true value of the autoregressive lag length $k$ in (ref) . In the context of (ref) we assume that the standard `I$(1,r_0)$ conditions' hold, where $r_0 \in \{ 0, \ldots, p \} $ denotes the true co-integration rank of the system (see also Cavaliere, Rahbek and Taylor, 2012); that is, the characteristic polynomial associated with (ref) has $p-r_0$ roots equal to 1 with all other roots lying outside the unit circle, and where $\alpha$ and $\beta$ have full column rank $r_0$.
The deterministic variables in (ref) are taken to satisfy one of the following cases (see, e.g., Johansen, 1996): (i) $D_{t}=0$, $d_{t}=0$ (no deterministic); (ii) $D_{t}=1$, $d_{t}=0$ (restricted constant); or (iii) $ D_{t}=t$, $d_{t}=1$ (restricted linear trend).
The innovation process $\varepsilon_t := (\varepsilon_{1t} ,\ldots, \varepsilon_{pt} )^\prime$ in (ref) is taken to satisfy the following set of conditions collectively labelled Assumption (ref).
In this section we discuss adaptive information-based methods for determining the co-integration rank and the autoregressive lag length in the context of (ref). In particular, we first derive the log-likelihood function in Section (ref) and the nonparametric estimator of the volatility matrix in Section (ref). We then outline the adaptive information criterion for the joint determination of the co-integration rank and the lag length in Section (ref) and we discuss how to sequentially estimate the lag length and the co-integration rank using adaptive methods in Section (ref).
Define $\Psi:=[\Gamma_1 : \ldots : \Gamma_{k-1}]$ and $Z_t^{(k)}:=(\Delta X_{t-1}^\prime, \ldots, \Delta X_{t-k+1}^\prime)^\prime$, such that the model in (ref) with no deterministic components (case (i)) can be rewritten more compactly as
Suppose for the present that $\{\sigma_t\}$ is known, and that $z_t$ is Gaussian; i.e., $z_t \sim \mathrm{i.i.d.}\ N(0, I_p)$. Then under Assumption 1 we have that $\varepsilon_t | \mathcal{F}_{t-1} \sim N(0, \Sigma_{t})$, where $\mathcal{F}_{t-1}:=\{X_{t-1}, \ldots, X_1, X_0, \linebreak \ldots, X_{1-k} \}$, and the log-likelihood function is given by (see Boswijk and Zu, 2022):
Maximum likelihood estimation of the parameters $(\alpha, \beta, \Psi)$ can be achieved by using the so-called generalised reduced rank regression procedure (Boswijk, 1995; Hansen, 2002, 2003), which uses a switching algorithm in order to circumvent the issue of the lack of a closed-form expression for the maximum likelihood estimator (MLE). In particular, because the MLE of $(\alpha, \Psi)$ for fixed $\beta$ and the MLE of $\beta$ for fixed $(\alpha, \Psi)$ have closed-form expressions, the maximisation of (ref) can be achieved, starting from an initial guess, by switching between maximisation over $(\alpha, \Psi)$ and $\beta$; see Boswijk and Zu (2022) for further details.
In this paper we focus on the two-sided smoothing nonparametric estimator of the volatility matrix adopted by Boswijk and Zu (2022). This estimator is a multivariate extension of Hansen (1995)'s nonparametric volatility filter based on leads and lags of the outer product of the residual vector. A similar approach to adaptive estimation has also been considered by Xu and Phillips (2008) and Patilea and Ra\"{i}ssi (2012), among others.
Let $K(\cdot)$ denote some kernel function and define $K_h (x) := K(x/h)/h$ with $h>0$ a {\it window width}. The kernel estimator for $\Sigma_t$ that we will consider is then defined as,
where $\hat{e}_t$ is the residual vector obtained by estimating an unrestricted VAR model of order $K$ in the levels of $X_t$, i.e.\ $\hat{e}_t = X_t - \sum_{i=1}^K \hat{A}_i X_{t-i}$, where $A_i$, $i=1,\ldots,K$, are $p \times p$ coefficient matrices. The value $K $ denotes the maximum autoregressive lag order we will allow for which, unless otherwise stated, is assumed in the following to be at least as large as the true lag order, $k_0$ in (ref).
The kernel function in (ref) is implemented with two-sided smoothing, so that $\hat{\Sigma}_t$ is based on leads and lags of $\hat{e}_t \hat{e}_t^\prime$, as outlined in Assumption 3 in Boswijk and Zu (2022). In their Lemma 2, Boswijk and Zu (2022) show that the volatility matrix process implied by the $T$ nonparametrically estimated covariance matrices is uniformly consistent over the compact interval $[0,1]$, which, in turn, implies uniform consistency of the nonparametric estimator $\hat{\Sigma}_t$ in (ref) over $t=1,\ldots,T$. Therefore, these consistent estimators can be used to replace $\Sigma_t$ in the log-likelihood function in (ref), thereby allowing for a feasible version of the generalised reduced rank regression procedure and the computation of the adaptive information criteria and the adaptive bootstrap PLR tests.
In implementing the nonparametric estimator of $ \Sigma_t $ in (ref), we will select the window width $h$ by minimising the quantity
where $||\cdot||$ denotes the Euclidean matrix norm, and where $\hat{\Sigma}_t^{-t}(h)$ is given by (ref), but with $K(0)$ replaced by 0, so that $\hat{e}_{t}\hat{e}_{t}^\prime$ does not enter the expression for $\hat{\Sigma}_t^{-t}(h)$. This leave-one-out cross-validation technique is implemented in Boswijk and Zu (2016) and Patilea and Ra\"{i}ssi (2012), and satisfies the requirement that $h$ decreases with the sample size at a certain rate; see Lemma 2 of Boswijk and Zu (2022) and Section (ref) below.
The maximised pseudo log-likelihood function (ref) associated with (ref) under lag order $k$ and co-integration rank $r$, say $\hat{\ell}_T^{(k,r)} (\alpha, \beta, \Psi)$, in conjunction with the volatility estimator in (ref) substituted for $\Sigma_t$ in (ref) can then be used to construct a feasible adaptive information criterion of the following generic form
where the term $c_T$ may depend on the sample size $T$ (see below) and where $\pi(k,r)$ denotes the number of parameters in the estimated model.\footnote{The number of parameters which defines the penalty term in (ref) depends on the deterministic components included in the model (ref) as follows: (i) in the case of no deterministic component ($D_t=0$, $d_t=0$ in (ref)), $\pi(k,r) = r(2p-r)+p^2(k-1)$; (ii) for the restricted constant case ($D_t=1$, $d_t=0$ in (ref)), $\pi(k,r) = r(2p-r+1)+p^2(k-1)$, and (iii) for the case of a restricted trend ($ D_t=1$, $d_t=1$ in (ref)), $\pi(k,r) = r(2p-r+1)+p+p^2(k-1)$. } The autoregressive lag order and the co-integration rank can then be jointly estimated by minimising the information criterion in (ref) jointly over both all possible lag lengths, $k = 1,\ldots,K$, and over all possible co-integration ranks, $r = 0,\ldots,p $; that is,
Different values of the coefficient $c_T$ yield different adaptive information criteria. In the standard (non-adaptive) case, which can be obtained as a special case of the adaptive information criterion in (ref) by restricting $\Sigma_t = I_p$ in the likelihood function (ref), the most widely used information criteria are the Akaike information criterion [AIC] (Akaike, 1974), the Bayes information criterion [BIC] (Schwarz, 1978), and the Hannan-Quinn information criterion [HQC] (Hannan and Quinn, 1979), which obtain setting $c_T=2$, $\log T$, and $2 \log \log T$, respectively. We will denote the generic standard information criterion in this case as $\text{IC}(k,r)$ and the resulting estimate in (ref) as $ ( \tilde{k}_{\rm IC}, \tilde{r}_{\rm IC} ) $. In the context of (ref), we will refer to the adaptive information criteria based on the AIC, BIC and HQ choices of $c_T$ as ALS-AIC, ALS-BIC, and ALS-HQC, respectively.
Because the lag length $k$ in (ref) is in general unknown and needs to be estimated prior to estimating the co-integration rank, practitioners often use a two-step procedure, whereby the autoregressive lag length is estimated in the first step and then subsequently employed as if it were the known lag length in a second step for determining the co-integration rank, such as a sequential procedure based on PLR tests or an information criterion. In particular, L\"{u}tkepohl and Saikkonen (1999) and Nielsen (2006), inter alia, show that the lag length in nonstationary VAR models can be consistently estimated from the levels of the data using an information criterion. Therefore, the lag length could be selected in the first step of the sequential procedure according to a (standard) information criterion where we do not impose a reduced rank structure on $\Pi:=\alpha\beta^\prime$ in (ref), that is by imposing $r=p$; see, among others, Cavaliere et al.\ (2018).
As for the joint determination of the lag length and the co-integration rank considered in Section (ref), an adaptive version of the information criterion for determining the lag length can also be considered. In particular, the lag length may be selected using an adaptive information criterion of the generic form
where $\hat{\ell}_T^{(k,p)}(\Pi, I_p, \Psi)$ is the maximised pseudo likelihood (ref) associated with (ref) where we do not impose a reduced rank structure on $\Pi=\alpha \beta^\prime$ under lag length $k$ and $\Sigma_t$ in (ref) is substituted with the volatility estimator in (ref). Again, the choice of the $c_T$ term identifies different information criteria as outlined above and, in this case, $\pi(k,p)=p(pk+i)$ with $i = 0$ when no deterministic component is involved, $i = 1$ in the case of restricted constant, and $i = 2$ for the restricted trend. The resulting adaptive information criterion-based lag length estimator is then given by
We note again that the generic standard information criterion, which we will denote by $\text{IC}(k,p)$, can be obtained as a special case of the adaptive information criterion in (ref), by restricting $\Sigma_t = I_p$ in the likelihood function (ref), with the resulting lag length estimator denoted by $\hat{k}_{\text{IC}}$. In the simulation experiments discussed in Section (ref), we will consider both the standard and the adaptive versions of the information criterion for determining the lag length in the first step of the two step sequential procedure. The selected lag length, either $\hat{k}_{\text{IC}}$ or $\hat{k}_{\text{ALS-IC}}$ generically denoted by $\hat{k}$ for the remainder of this section, is then used as if it were the true lag length in the second step for determining the co-integration rank. The second step could be based on either the sequential procedure of Boswijk and Zu (2022) based on adaptive bootstrap PLR tests or an adaptive information criterion for selecting the co-integration rank. We now outline these two possibilities.
The adaptive PLR test-based procedure of Boswijk and Zu (2022). Boswijk and Zu (2022) introduce the adaptive PLR statistic for testing the null hypothesis that the true co-integration rank is (no more than) $r$, $0 \leq r \leq p-1 $,
where $\hat{\varepsilon}_{r,t}$ and $\hat{\varepsilon}_{p,t}$ denote the residuals from the restricted and unrestricted VAR model in (ref), respectively. For the case where the autoregressive lag length is known ($k=k_0$), they demonstrate that the limiting distribution of (ref) depends on the unknown volatility process. Consequently, bootstrap methods are required to approximate the critical values from this distribution. In order to do so, a bootstrap sample $\{X_{r,t}^\ast \}_{t=1}^T$ is generated recursively from
initialised at $X_{r,j}^\ast = X_{j}$, for $j=1-k, \ldots, 0$, where $\hat{\alpha}^{(r)}$, $\hat{\beta}^{(r)}$, and $\hat{\Gamma}_i^{(r)}$ are the estimated parameter matrices from the model (ref) obtained using conventional reduced rank regression under the rank $r$ imposed by the null hypothesis. The adaptive PLR test statistic based on the bootstrap sample is then computed as
where $\hat{\varepsilon}_{r,t}^\ast$ and $\hat{\varepsilon}_{p,t}^\ast$ denote the (bootstrap) residuals from the restricted and unrestricted models, respectively. Following Boswijk and Zu (2022), we consider the following two bootstrap implementations: (i) the variance bootstrap, $\varepsilon_{r,t}^\ast := \hat{\Sigma}_t^{1/2} z_t^\ast$, where $\hat{\Sigma}_t^{1/2}$ is any square root of $\hat{\Sigma}_t$ and $z_t^\ast \sim \mathrm{i.i.d.} N(0, I_p)$; (ii) the wild bootstrap, $\varepsilon_{r,t}^\ast := \hat{\varepsilon}_{r,t} w_t^\ast$, where $w_t^\ast$ is a scalar i.i.d.\ N(0,1) sequence; see Section 4.2 of Boswijk and Zu (2022) for more details. As is typically done in practice, the unknown lag length $k$ in (ref) and (ref) is replaced by the lag length estimated in the first step of the sequential procedure, say $\hat{k}$, in order to compute the bootstrap statistic $Q_{r,\hat{k},T}^\ast$ using the bootstrap sample in (ref) based on $\hat{k}$. The corresponding $p$-value is then computed as $p^\ast_{r,\hat{k},T}:=1-G^\ast_{r,\hat{k},T}(Q_{r,\hat{k},T}^\ast)$, where $G^\ast_{r,\hat{k},T}(\cdot)$ denotes the conditional (on the original data) cdf of $Q_{r,\hat{k},T}^\ast$. Starting from $r=0$, the bootstrap algorithm is repeated as long as $p^\ast_{r,\hat{k},T}$ exceeds the significance level $\eta$, thus yielding $\hat{r}^\ast(\hat{k})=r$. If the null is not rejected for $r=p-1$, then $\hat{r}^\ast(\hat{k})=p$.
The asymptotic validity of the two bootstrap procedures outlined above is established in Theorem 3 of Boswijk and Zu (2022) with the implication that, for the case where the autoregressive lag length is known ($k=k_0$), the variance and wild bootstrap adaptive PLR test-based procedures, $\hat{r}^\ast(k_0)$, are asymptotically accurately capped estimator of the co-integration rank $r_0$.\footnote{The sequential rank determination procedure of Johansen (1996) is asymptotically accurately capped in that if each PLR (or bootstrap PLR) test in the sequence is run with nominal (asymptotic) significance level $ \eta $, then the limiting probability of selecting a rank smaller than, equal to, and greater than the true rank will be 0, $1- \eta $ and $\eta $, respectively, when $r_0 < p $ and 0, 1 and 0, respectively, when $r_0 = p $.} In Section (ref) we will generalise these results to the case where the lag length is unknown and estimated in the first step of the sequential procedure.
The adaptive IC-based procedure. Alternatively, to determine the co-integration rank in the second step of a sequential procedure based on ALS-IC, $k$ in the generic form (ref) can be replaced by the lag length estimated in the first step, thus yielding
The resulting adaptive information criterion-based co-integration rank estimator is then given by
In this section we establish the large sample properties of the adaptive methods for determining the co-integration rank and autoregressive lag length outlined in Sections (ref) and (ref).
Lemma 2 of Boswijk and Zu (2022) establishes that the nonparameteric estimate of the volatility matrix process defined as $\hat{\Sigma}_T(u):=\sum_{t=1}^T \hat{\Sigma}_t 1_{[(t-1)/T,t/T]} (u)$ is uniformly consistent over the compact interval $[0,1]$. This result is a basic building block needed to demonstrate weak consistency\footnote{ An estimator $T_{n}$ is defined to be weakly consistent if it converges in probability to the true value of the unknown parameter $ \theta$; that is, $T_{n} \overset{p}{\rightarrow } \theta$.} for the adaptive information criteria in (ref) and (ref) and so for completeness we first reproduce that result below as Result 1.
Using Result 1, we first show in Lemma (ref) that the adaptive information criterion in (ref) is weakly consistent for the co-integration rank, regardless of the autoregressive lag length used, provided standard conditions hold on the penalty term, $ c_T$. Then second in Lemma (ref) we show that for the true co-integration rank, $ r_0$, the adaptive information criterion in (ref) is weakly consistent for the autoregressive lag length.
Using the results in Lemmas (ref) and (ref), we are now in a position to establish the weak consistency of the joint procedure. This is now given in Theorem (ref).
To conclude this section we now detail the large sample behaviour of the two-step sequential procedures outlined in Section (ref) where in the first step we select the autoregressive lag and then in the second step an adaptive procedure based on this estimated lag length is used to determine the co-integration rank.
First, in Lemma (ref), we generalise the results in Lemma 3 of Cavaliere et al.\ (2018), which show the sufficient conditions on the term $c_T$ that ensure weak consistency for an information criterion of the form given in (ref), to the case of its adaptive analogue, ALS-IC($k,p$). In particular, we derive the conditions under which minimising an adaptive information criterion consistently selects the true lag order, $k_0$, in the first step when we do not impose a reduced rank structure, so that we set $r=p$.
The results in Lemma (ref) imply that $\hat{k}_{\text{ALS-IC}} \overset{p}{\to} k_0$, again provided $\frac{c_T}{T}+\frac{1}{c_T} {\rightarrow} 0$, as $ T \rightarrow \infty $. Using the results in Lemmas (ref) and (ref), we are now in a position in Theorem (ref) to establish the large sample properties of the bootstrap adaptive PLR test-based estimator of the co-integration rank using the lag length estimated by an information criterion as in (ref) at the first step, $\hat{r}^\ast(\hat{k}_{\text{ALS-IC}})$.
Finally, in Theorem (ref) we generalise the results in Theorem 2 of Cavaliere et al.\ (2018) by establishing the large sample properties of the adaptive IC-based estimator of the co-integration rank as in (ref) using the lag length estimated by an information criterion as in (ref) at the first step, $\hat{r}_{\text{ALS-IC}}(\hat{k}_{\text{ALS-IC}})$.
In this section we use Monte Carlo simulation methods to investigate the finite sample performance of the joint and sequential adaptive methods for determining the co-integration rank and autoregressive lag length outlined in Sections (ref) and (ref) and compare these with their standard (non-adaptive) counterparts. The results from these Monte Carlo experiments are reported in Tables 1-6.
We will consider the following second-order VAR model of dimension $p=2$ as our simulation DGP:
with $t=1-K,\ldots,T$, $X_{-K}=\Delta X_{-K}=0$, where $K$ denotes the maximum lag order. In order to allow for true co-integration ranks, $ r_0$, of 0, 1 or 2, we set the parameters $a$ and $b$ in the long-run parameter vector $\alpha$ in (ref) as follows: $a=b=0$ for $r_0=0$, $a=-0.4$ and $b=0$ for $r_0=1$, and $a=b=-0.4$ for $r_0=2$ (full rank). Moreover, we set $\Gamma_{1}:=\gamma I_{2}$ with $\gamma \in \{0, 0.1, 0.5, 0.9\}$.\footnote{For the simulation DGP in (ref), it suffices that $ ( a, b, \gamma ) \in (-2,0]^2 \times [0,1)$ in order to satisfy the I($1,r$) conditions.}
We will consider three cases for the the innovation vector, $\varepsilon_{t}$ in (ref). The first case is that $ \varepsilon_{t} \sim \mathrm{i.i.d.}\ N \left( 0, I_2\right) $ so that $\varepsilon_{t}$ is homoskedastic. This case will provide a useful benchmark to investigate the effects of using adaptive methods when they are not needed. The second case considers conditionally heteroskedastic innovation processes, where the individual components of $\varepsilon_{t}$ follow the first-order AR stochastic volatility [SV] model sets as $\varepsilon_{it}=v_{it}\exp {(h_{it})},$ $ h_{it}=\lambda h_{it-1}+0.5\xi _{it}$, with $(\xi _{it},v_{it})^{\prime }\sim \mathrm{i.i.d.}\ N(0,\mathrm{diag}(\sigma _{\xi }^{2},1))$, independent across $i=1,2$. Results are reported for $\lambda =0.951$, $\sigma _{\xi }=0.314$. This case constitutes a well-known conditionally heteroskedastic model for the innovations which has been used with the same parameter configuration in many other Monte Carlo experiments such as Gon\c{c}alves and Kilian (2004), Cavaliere et al.\ (2010), and Cavaliere et al.\ (2015, 2018). The third case we consider sets $\varepsilon_{t}$ to be a non-stationary, unconditionally heteroskedastic independent sequence of Gaussian variates, characterised by a late positive variance shift. Specifically,
In order to evaluate the behaviour of the adaptive and corresponding standard procedures in practically relevant sample sizes we report results for $T=50$ and 100. All experiments are run over 1,000 Monte Carlo replications and were programmed using MATLAB. Our experiments are based on the no deterministic component case. In all of our simulation experiments we set $K=4$ as the maximum lag length considered. Results for the joint information-based estimates of the co-integration rank and lag length from Section (ref) are reported first in Table 1, while results relating to the sequential procedures from Section (ref) are reported in Tables 2 and 3 for the IC-based approaches in the case of SV innovations and single volatility break, respectively, and in Tables 4 and 5 for the sequential bootstrap-based procedures, again for the SV and single volatility break cases, respectively. Finally, for comparison purposes, Table 6 reports the results for the joint information-based approaches in the homoskedastic case.
Consider first Table 1 which reports results for determining the co-integration rank $r$ (left two panels of Table 1) and the lag order $k$ (right two panels of Table 1) using the joint ALS-IC-based procedures detailed in Section (ref) together with their corresponding standard information criteria-based counterparts. In particular, Table 1 reports the empirical frequencies with which $ \tilde{r} $ and $ \tilde{k} $ from the joint information-based estimator defined in (ref) select the values $ r=0,1,2$ and $ k=1,2, 3, 4 $, respectively, for each of the adaptive criteria ALS-HQC and ALS-BIC, and the corresponding standard criteria, HQC and BIC.\footnote{We do not consider the ALS-AIC estimator nor its standard counterpart in the Monte Carlo experiments because the poor performance of AIC-based approaches in finite samples is documented in many contributions in the literature (see e.g., Kapetanios, 2004; Wang and Bessler, 2005; Cavaliere et al., 2015; Cavaliere et al., 2016). Additional simulations show that, also in the case of adaptive estimation, this criterion tends to overestimate both the true co-integration rank and the lag length. Nevertheless, the adaptation with respect to the variance matrix profile considerably improves the finite sample performance of the AIC-based approach. These results are available on request.}
A number of observations can be made from the results reported in Table 1. Consider first the estimators of the co-integration rank.
The following observations can also be made concerning the behaviour of the estimators of the autoregressive lag length seen in Table 1.
Let us now turn our attention to a discussion of the results in Tables 2-5 which relate to the sequential estimates from Section (ref).
We first focus attention on the results reported for the two-step IC-based procedures in Tables 2 and 3 for the cases of SV innovations and a single volatility break, respectively. In particular, we report the empirical frequencies with which both standard and adaptive IC-based procedures select the lag length, $k$, at the first step (`Step I' in the tables) and those with which they select a co-integration rank, $r$, of zero, one or two at the second step (`Step II' in the tables), using the lag length estimated at the first step by each standard and adaptive information criterion, IC($k,p$) and ALS-IC($k,p$).
The results for where the co-integration rank is determined using the same information criterion at both steps of the sequential procedure are overall similar to the results for the corresponding joint IC-based approaches discussed above. As an example, the joint ALS-BIC estimate of $r$ in Table 1 selects the correct co-integration rank 77.6% (91.3%) of the time when $r_0=1$, $\gamma=0.5$ and $T=50$ ($T=100$) in the SV case, while the corresponding sequential procedure based on ALS-BIC estimate at both steps, i.e.\ ALS-BIC($\hat{k}_{\text{ALS-BIC}},r$), selects the true rank 77.1% (90.9%) of the time. Moreover, all of the approaches considered appear to be fairly robust to the choice of whether to use an adaptive or standard information criterion in the first step of the sequential procedure as the results for the co-integration rank determination appear very similar using either ALS-IC($k,p$) or IC($k,p$).
We now turn to a discussion of the results for the wild bootstrap PLR procedure [denoted PLR-WB], together with the adaptive PLR procedures of Boswijk and Zu (2022) implemented with either a variance bootstrap [denoted ALR-VB] or a wild bootstrap [denoted ALR-WB] in Tables 4 and 5 for the cases of SV innovations and a single break in volatility, respectively. For each of these we report the empirical frequencies with which they select a co-integration rank, $r$, of zero, one or two. We report results for three case for the lag length used in these procedures. The first is an infeasible version based on knowledge of the true lag length, i.e.\ we set $k=k_0$. The other two select the lag length in the first step of the two-step sequential procedure using either standard BIC, $k=\hat{k}_{\text{BIC}}$, or its adaptive counterpart, $k=\hat{k}_{\text{ALS-BIC}}$.\footnote{In Tables 4 and 5 we focus on BIC-based approaches for the selection of $k$ because these provide the best overall performance, see e.g.\ Cavaliere et al.\ (2018). Moreover, we only report the results for the lag length determination for the case of $r_0=1$. The results for $r_0=0$ and 2 are very similar and thus, in the interest of space, are not reported.}
A number of observations can be made from the results reported in Tables 4 and 5.
Finally, we investigate the potential losses of efficacy seen when using the adaptive methods in the benchmark case of homoskedastic innovations by comparing the results reported in Table 6 for the adaptive IC-based methods with those of their standard counterparts. These results suggest that the performance of the joint adaptive IC-based procedures do not deteriorate to any significant degree when the shocks are homoskedastic, such that the use of adaptive methods is unnecessary. Indeed, when either $r_0=1$ or 2, the performance of the ALS-IC-based approaches is similar and sometimes even better than the results for their corresponding standard counterparts. Conversely, in the case of no co-integration, $r_0=0$, standard BIC and HQC-based approaches outperform their adaptive counterparts. However, as pointed out above, this is mainly an artefact of the overall tendency of the standard criteria, especially BIC, to under-fit the true co-integration rank.
To conclude this section, we compare the finite sample behaviour of the adaptive information criteria-based methods with that of the adaptive PLR test-based approaches. By comparing the results reported in Tables 1, 2 and 3 with those in Tables 4 and 5, we observe that, for the co-integrated case ($r_0=1$), the finite sample performance of either joint or sequential ALS-IC is similar to that of adaptive PLR test-based procedures. Conversely, when $r_0=0$ the PLR test-based procedures outperform the adaptive information criteria-based approaches, while this behaviour is reversed when $r_0=2$ and $T=50$. In the case of full rank and $T=100$, the performance of the methods considered are similar. Finally, by comparing the results in Tables 1, 2 and 3 for the joint and the sequential information criteria-based approaches for selecting the lag length, we note that the ability of these methods to determine $k$ are very similar.
In this section we provide an empirical application of the adaptive information criteria-based approaches to the term structure of interest rates in the US. In particular, we analyse the time series $X_t = (X_{1t}, \ldots, X_{5t})^\prime$ of monthly zero yields from January 1970 to December 2012, for maturities equal to 3 months ($X_{1t}$), 1 year ($X_{2t}$), 3 years ($X_{3t}$), 5 years ($X_{4t}$), and 10 years ($X_{5t}$).
The co-integration analysis of $X_t$ has already been considered by Boswijk et al.\ (2016) and Boswijk and Zu (2022). In particular, in order to account for the unconditional heteroskedasticity present in the data, sequential procedures where the lag length is selected at the first step according to (standard) HQC($k,p$), and then the co-integration rank of the system is determined using either PLR-WB (Boswijk et al., 2016) or adaptive PLR tests (Boswijk and Zu, 2022) were adopted. Here we apply the adaptive information-based methods to estimate the co-integration rank and autoregressive lag order of the system and compare these results with those obtained in the two previous analyses cited above. In what follows, the VAR models are fitted with a restricted trend and, for all methods, the maximum number of lags considered is $K=4$. The number of bootstrap samples used in the bootstrap algorithms is $B=999$.
We first focus on the joint determination of the co-integration rank and lag length using adaptive joint information criterion-based procedures as outlined in Section (ref) and the standard counterparts. These results are reported in Table 7. The results in Table 7 show that all of the joint information criteria, both adaptive and non-adaptive, agree on selecting a lag length of $\tilde{k}=2$. Moreover, both standard and adaptive versions of the joint BIC-based approach delivers the same estimate of the co-integration rank, namely $\tilde{r}_{\text{BIC}}=\tilde{r}_{\text{ALS-BIC}}=2$. Conversely, the joint HQC-based approaches select a higher co-integration rank. Specifically, the co-integration rank selected using the (standard) joint HQC-based approach is 3, i.e., $\tilde{r}_{\text{HQC}}= 3$, whereas $\tilde{r}_{\text{ALS-HQC}}=4$ is obtained using the adaptive version.
We now consider in Table 8 the results obtained using the sequential procedures for determining the lag length and then the co-integration rank. In particular, the upper panel of Table 8 shows the results for the selection of $k$ in the first step of the sequential procedure, whereas the results for the determination of $r$ at the second step using information criteria and PLR tests are reported in the middle and lower panels of Table 8, respectively. Note that the results reported in the lower panel of Table 8 for the case of $\hat{k}=2$ reproduce those in Boswijk et al.\ (2016) and Boswijk and Zu (2022) who use standard HQC to select the lag length and therefore they set $\hat{k}=2$. The results for the first step of the sequential procedure show that all but the standard BIC information criteria agree on a choice of $\hat{k}=2$; standard BIC chooses $\hat{k}_{\text{BIC}}=1$. Therefore, on balance, we would recommend a VAR model of order 2.
Let us next focus on the second step of the sequential procedure and, in particular, on the determination of the co-integration rank obtained by the PLR tests (see the lower panel of Table 8). For $\hat{k}=2$, the results for the adaptive and non-adaptive bootstrap-based PLR test procedures vary according to the nominal significance level considered. In particular, at a standard 5% level we select $\hat{r}=2$ using the (non-adaptive) PLR-WB procedure, whereas the two adaptive PLR methods yield $\hat{r}=4$, again replicating the results in Boswijk et al.\ (2016) and Boswijk and Zu (2022), respectively. Using a 1% significance level, we still select a co-integration rank of 2 using the (standard) PLR-WB but we would now select $\hat{r}=3$ using the two adaptive PLR test-based procedures. The results for the information criteria used in the second step of the sequential procedure show that, using $\hat{k}=2$, HQC-based approaches in both adaptive and non-adaptive form agree with the selection of $\hat{r}=4$ also made at the 5% level made by the adaptive PLR test-based procedures. The co-integration rank of $\hat{r}=2$ selected using both the adaptive and non-adaptive BIC-based approaches matches that chosen by the (non-adaptive) PLR-WB test procedure. It is worth noting that, when setting $\hat{k}=1$ as suggested by the (standard) BIC($k,p$), the results for both information criteria and PLR tests in step 2 of the sequential procedure are much more variable across the methods with the rank selected anywhere between 2 and 5. Therefore, we would not recommend the conclusions based on $\hat{k}=1$. In particular, because BIC uses a stricter penalty term than HQC, we would expect, other things equal, that BIC-based approaches will often select a lower lag length and/or co-integration rank than HQC-based approaches. Moreover, this tendency of standard BIC might be exacerbated by the presence of heteroskedasticity in the data, thus allowing the adaptation with the respect to the volatility process to deliver more reliable results in small samples.
In summary, overall our results seem strongly in favour of a selection of an autoregressive lag length of 2. However, the selected co-integration rank varies according to the method used. In particular, the joint and sequential (for $\hat{k}=2$) HQC-based approaches select a co-integration rank of 4, while the joint and sequential (for $\hat{k}=2$) BIC-based approaches select rank 2. The sequential procedures based on PLR tests and $\hat{k}=2$ select $\hat{r}=2$ in non-adaptive form, $\hat{r}=3$ when using a 1% significance level and $\hat{r}=4$ when using a 5% significance level. This is in some ways consistent with the findings for BIC and HQC-based methods since decreasing the significance level is qualitatively the same as using a stricter penalty in the information criterion. Finally, it is worth noting that the choice of a rank equal to 4 implies the presence of a single stochastic trend driving the five yields and is in line with the (weak-form) expectation hypothesis of interest rates (see, for example, Campbell and Shiller, 1987), which implies that the (long-term) level factor - but not the slope nor the curvature - of the interest rate yield curve is a random walk process, so that $\beta^\prime X_t$ consists of spreads $X_{it}-X_{1t}$ for $i =2, 3, 4, 5$.
In this paper we have proposed new methods for determining the co-integration rank and the lag order in heteroskedastic VAR models which exploit the time variation in the unconditional error variance matrix. In particular, we have proposed adaptive information criteria-based approaches to jointly determine the co-integration rank and the autoregressive lag length. Provided standard conditions hold on the penalty term hold, these methods are proved to be weakly consistent for co-integration rank and lag order determination. We have also demonstrated that the adaptive PLR rank determination procedure of Boswijk and Zu (2022), originally developed under the assumption of a known autoregressive lag length, remains asymptotically valid when a consistent lag length estimate, such as that provided by an adaptive information criterion, is used. Monte Carlo experiments reported indicate that the adaptive information criteria-based approaches generally outperform standard methods in finite samples when non-stationary volatility is present in the data.
This paper is dedicated to the memory of our dear friend and colleague Mike McAleer. The authors wish to thank Anders Rahbek and Yang Zu for their helpful comments and suggestions. The authors also thank participants at the Computation and Financial Econometrics in London (December 2015 and 2017), the Bootstrap Workshop at the Amsterdam School of Economics (November 2015), and the RCEA Time Series Econometrics Workshop in Rimini (June 2016). This research was supported by the Danish Council for Independent Research (DSF Grant 015-00028B) and the Italian Ministry of University and Research (PRIN 2017 Grant 2017TA7TYC).
\setcounter{section}{0}