Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
65,158 characters · 8 sections · 96 citation commands
Oracle Efficient Estimation of Structural Breaks in Cointegrating Regressions
\sloppy
\singlespacing
\thispagestyle{empty}
\onehalfspacing
In this paper, we consider modelling cointegration relationships where the long-run equilibrium may differ for subsamples, thereby allowing for (multiple) structural breaks in the cointegrating regression. We assume that cointegration holds over some (fairly long) period of time, but then shifts to a new `long-run' relationship. The number of breaks and their location are unknown to the researcher. Although coefficients of long-run equilibrium equations are relatively persistent by definition, accounting for the possibility of structural breaks is crucial in cointegration analysis, which usually involves long sample periods. On the one hand, long time series are needed to study the long-run behaviour of economic systems, on the other hand, employing long time series increases the likelihood of encountering structural change during the sample period. It is widely known that structural breaks, when present, can mask cointegrating relationships and render cointegration tests uninformative Campos1996, GregoryNasonWatt1996, Qu2007. Hence, we propose a two-step approach to detect (multiple) structural breaks in cointegrating regressions using penalized regression techniques.
Since time series used for economic analyses have become very long in some instances, detecting (multiple) structural breaks has emerged as an important problem in the econometrics literature. For comprehensive surveys on structural breaks in time series models (`change-point' detection in the statistics literature or `pattern recognition' in the context of signal processing), see, for example, Perron2006, AueHorvath2013 and NiuHaoZhang2015. While classical structural break models for linear regressions attempt to detect one unknown break via a grid search procedure Andrews1993, it is not feasible to use grid searches for the detection of multiple breaks because the computational cost increases exponentially with the presumed number of breaks (needing least squares operations of order $O(T^m)$ for $m$ breaks). Addressing this issue, BaiPerron1998, BaiPerron2003 use dynamic programming techniques (henceforth Bai-Perron algorithm), requiring at most least-squares operations of order $O(T^2)$ for any number of breaks, to add breaks sequentially to the model. Recently, several approaches have been proposed that reframe the task of detecting and estimating structural breaks as a model selection problem employing penalized regressions and related model selection techniques Davis2006, Harchaoui2010, JinShiWu2013, ChanYauZhang2014, Ciuperca2014, JinWuShi2016, QianJia2016, QianSu2016, Behrendt2020. Instead of grid search procedures which augment linear regression models with parameter changes, model selection procedures take a top-down approach and try to shrink the set of all possible breakpoint candidates to contain only the true breakpoints. These approaches benefit from high computational efficiency and detect structural breaks with high accuracy.
The theory for (multiple) structural breaks in cointegrating regressions is not nearly as developed as the theory for change-points in the statistics and signal processing literature. Most studies are concerned with cointegration testing in the presence of structural instability. One of the most popular cointegration tests with an unknown breakpoint is the one proposed by GregoryHansen1996, GregoryHansen1996b in which the location of the break can be estimated via grid search at the minimum of the individual cointegration test statistics. Hatemi-J2008 extends the test to account for two breaks and Schweikert2019 allows for the possibility of nonlinear adjustment to the long-run equilibrium. Maki2012 employs a hybrid procedure, detecting $m-1$ breaks by minimizing the sum of squared residuals among all possible sample splits and finally determining the last break by minimizing a cointegration test statistic. Unfortunately, these tests are non-informative about the location of breaks. They optimize the model specification to provide evidence against the null hypothesis of no cointegration, thereby not necessarily finding those breakpoints which optimize the model fit. Maki2012 determines the first $m-1$ breaks based on improving the model fit but does not do so for the last breakpoint. Hence, the set of estimated breakpoints is not completely informative. Other studies with a strong focus on cointegration testing which conduct breakpoint estimation as a by-product are Carrion-i-Silvestre2006 and AraiKurozumi2007. They propose a CUSUM-based approach to test the null hypothesis of cointegration with a structural break against the alternative hypothesis of no cointegration. Qu2007 considers a cointegrated system allowing the cointegrating rank to change during different subsamples so that it is possible to detect cointegrating relationships that exist only in some subsamples. WesterlundEdgerton2006 design LM-based test statistics invariant to structural breaks to test the null of no cointegration and DavidsonMonticini2010 use subsample procedures to account for structural breaks in their cointegration tests.
In contrast, few studies are primarily focussed on modelling structural change in cointegrated systems. KejriwalPerron2008, KejriwalPerron2010 propose to estimate the number and location of structural breaks in cointegrating equations by applying the Bai-Perron algorithm. Inference on breakpoints is studied in, among others, BaiLumsdaineStock1998, QuPerron2007, KejriwalPerron2008, KejriwalPerron2010, LiPerron2017 and OkaPerron2018. Using penalized regression approaches to account for structural breaks in cointegrating regression has not been explored yet in great detail. A similar idea has been proposed by SchmidtSchweikert2019 but their procedure is limited to bivariate cointegrating regressions using a modified adaptive lasso estimator. Here, we extend their methodology to cointegrating regressions with multiple regressors and provide a rigorous proof that the adaptive group lasso estimator is oracle efficient in settings with an unknown number of breaks and a diverging number of breakpoint candidates.
The proposed estimation method in this paper consists of two main steps: in the first step, we apply the group lasso estimator to a cointegration model with a diverging number of breakpoint candidates. We allow that breaks can occur at any point in time except for some lateral trimming which is mostly needed to identify the baseline coefficients in the first regime. We prove that the group lasso estimator consistently estimates parameter changes. However, it is well-known that lasso estimators are not simultaneously parameter estimation consistent and model selection consistent in situations where the restricted eigenvalue condition or related conditions such as the strong irrepresentable condition do not hold ChanYauZhang2014. Under these conditions, we show that the number of selected breaks is greater than the true number of breaks almost surely, but their estimated location is sufficiently close to their true location. In the second step, we use the first step group lasso estimates as weights for the adaptive group lasso. We provide a rigorous proof that the adaptive group lasso has the oracle properties if the first step algorithm assumes a maximum number of breaks and the distance between breaks depends on the sample size ensuring that the breakpoint candidates for the second step estimation are sufficiently distinct. The number of breaks is then estimated as the number of non-zero groups obtained after adaptive group lasso optimization.
The paper is organized as follows. (ref) describes the proposed adaptive group lasso procedure to estimate structural breaks in cointegrating regressions. (ref) is devoted to the Monte Carlo simulation study. (ref) reports the results of an empirical application of our methodology to the US money demand function, and (ref) concludes. Proofs of all theorems in the paper are provided in the Mathematical Appendix.
In the following, we specify a cointegrated system with multiple structural breaks at which it attains new equilibrium states. The cointegrated system does not deviate persistently from each equilibrium until the next break occurs and a new equilibrium is maintained.
Let $\lbrace y_t \rbrace^{\infty}_{t=1}$ denote a scalar process generated by
where $t_j, j \in \lbrace 0, 1, \dots m + 1 \rbrace$ denote the breakpoints $1 = t_0 < t_1 < \dots < t_{m+1} = T + 1$. $\mu$ is the intercept, $\boldsymbol{\beta}'_j = (\beta_{j1}, \beta_{j2}, \dots, \beta_{jN})$ are regime-dependent coefficients and $\lbrace \boldsymbol{X}_t \rbrace^{\infty}_{t=1}$, where $\boldsymbol{X}_t = (X_{1t}, X_{2t}, \dots, X_{Nt})'$, follows an $N$-vector integrated process\footnote{Note that this specification of the process rules out integrated regressors with a deterministic drift component. Relaxing this assumption would be relatively straightforward.}
where $\boldsymbol{X}_0 = 0$. $\lbrace u_t \rbrace^{\infty}_{t=1}$ and $\lbrace v_t \rbrace^{\infty}_{t=1}$ are mean-zero weakly stationary error processes. For expositional simplicity, we restrict our analysis to cointegrating regressions with a constant intercept across regimes.\footnote{Our main results generally hold if the coefficients of included deterministic components, e.g. linear and quadratic trend terms, do not change over the sampling period. A discussion of breaks in the intercept is included in the supplementary material to this paper.} We make the following assumptions about the vector process $w_t = (u_t, v'_t)'$:
While the first three conditions of Assumption (ref) are standard in cointegration analysis, assuming that $\Omega_v$ is positive definite implies that $\boldsymbol{X}_t$ is non-cointegrated. We denote the number of structural breaks by $m$. While the number of true structural breaks $m_0$ is unknown, we assume that the maximum number of structural breaks $m^*$ is known to the researcher. The estimated number of breakpoints is denoted by $\hat{m}$. The locations of breakpoints relative to sample size, so-called break fractions, are denoted by $\tau_{j} = t_{j} / T, j \in \lbrace 0, 1, \dots m + 1 \rbrace$.
Throughout this paper, we use the following notation to present our main results: let $y_T = (y_1, y_2, \dots, y_T)'$ denote the vector containing $T$ observations of our response variable and $u_T = (u_1, u_2, \dots, u_T)'$ denotes the error term vector. The vector of $T$ observations for the $N$-dimensional variable $\boldsymbol{X}_t$ is denoted by $\boldsymbol{X} = (\boldsymbol{X}_1, \dots, \boldsymbol{X}_T)'$. Our design matrix $\boldsymbol{Z}_T$ is an $T \times TN$ matrix defined by
and we define the Gram matrix $\boldsymbol{\Sigma} = \boldsymbol{Z}_T' \boldsymbol{Z}_T / T^2$. Adjacent columns of $\boldsymbol{Z}_T$ differ only by one entry which means that the columns are almost identical for $T \to \infty$. Consequently, $\boldsymbol{\Sigma}$ does not converge to a positive definite asymptotic counterpart. It follows that the restricted eigenvalue condition BickelRitovTsybakov2009 does not hold and we cannot establish our consistency proofs based on this assumption. See ChanYauZhang2014 for a thorough discussion of this issue.
We set $\boldsymbol{\theta}_1 = \boldsymbol{\beta}_1$ and
for $i = 2, \dots, T$. For the remainder of this article, $\boldsymbol{\theta}_i = \boldsymbol{0}$ means that $\boldsymbol{\theta}_i$ has all entries equaling zero and $\boldsymbol{\theta}_i \neq \boldsymbol{0}$ means that $\boldsymbol{\theta}_i$ has at least one non-zero entry. The coefficient vector $\boldsymbol{\theta}(T) = (\boldsymbol{\theta}_1, \boldsymbol{\theta}_2, \dots, \boldsymbol{\theta}_T)'$ is of length $TN$ and contains all time-specific parameter changes. Because we treat structural breaks as rare events and assume that parameter changes persist for some time, the number of non-zero elements in $\boldsymbol{\theta}(T)$ is assumed to be small, i.e.\ smaller than $m^* + 1$ groups of size $N$.
We denote the true value of a parameter with a $0$ superscript. $\lbrace \tau^0_j, j = 1, \dots, m_0 \rbrace$ denotes the set of true break fractions and $\boldsymbol{\beta}^0_j$, $j = 1, \dots, m_0 + 1$ defines the true coefficient of the $j$-th regime. For technical reasons, we additionally set $\boldsymbol{\beta}^0_0 = 0$. We define the index sets $\bar{\mathcal{A}} = \lbrace 1 \leq i \leq T: \boldsymbol{\theta}^0_i \neq \boldsymbol{0} \rbrace$ denoting the indices of truly non-zero coefficients (including the baseline coefficient) and $\mathcal{A} = \lbrace i \geq 2: \boldsymbol{\theta}^0_i \neq \boldsymbol{0} \rbrace$ denoting the non-zero parameter changes. The index set obtained from our first step estimation belonging to all estimated non-zero parameter changes is denoted by $\mathcal{A}_T = \lbrace i \geq 2: \tilde{\boldsymbol{\theta}}_i \neq \boldsymbol{0} \rbrace$. We note that the first regime's coefficient (before the first breakpoint) is not allowed to be zero.\footnote{While the cointegrating vector $(1,\boldsymbol{0}')'$ in principle ensures that $y_t = \mu + u_t$ is stationary under our assumptions, we exclude this case to simplify our exposition. In the following, we need a clear distinction between zero and non-zero coefficients to decide whether their indices belong into the sets $\bar{\mathcal{A}}$ or $\bar{\mathcal{A}}^c$. Allowing zero baseline coefficients would require several case-by-case considerations.} Since we indicate breakpoints with non-zero coefficients in our penalized regression approach, the set $\mathcal{A} = \lbrace t^0_1, t^0_2, \dots, t^0_{m_0} \rbrace$ is also used to denote true breakpoints. Similarly, the set $\mathcal{A}_T = \lbrace \hat{t}_1, \hat{t}_2, \dots, \hat{t}_{\hat{m}} \rbrace$ denotes estimated breakpoints, i.e., indices of those coefficients which are estimated to be non-zero. $|\mathcal{A}|$ denotes the cardinality of the set $\mathcal{A}$ and $\mathcal{A}^c$ denotes the complementary set. We use those sets to index rows and columns of vectors and matrices. For example, let $\boldsymbol{Z}_{T, \mathcal{A}}$, $\boldsymbol{Z}_{T,\mathcal{A}^c}$ contain the columns of $\boldsymbol{Z}_T$ and $\boldsymbol{\theta}_{\mathcal{A}}(T)$, $\boldsymbol{\theta}_{\mathcal{A}^c}(T)$ contain the rows of $\boldsymbol{\theta}(T)$ associated with active and inactive breakpoints, respectively.
For notational convenience, we use `$\Rightarrow$' to signify weak convergence of the associated probability measures and $\overset{p}{\to}$ to denote convergence in probability. Continuous stochastic processes such as a Brownian motion $B(s)$ on [0,1] are simply written as $B$ if no confusion is caused. We also write integrals with respect to the Lebesgue measure such as $\int\limits_{0}^{1} B(s)ds$ simply as $\int\limits_{0}^{1} B$. Throughout the paper, several (distinct) large constants are all denoted with $C$, while small constants are denoted by $\epsilon$.
Using these definitions, our cointegration model described in Equation (ref) can be expressed as a high-dimensional regression model in matrix form
Since only $m_0 + 1$ groups within $\boldsymbol{\theta}(T)$ are truly non-zero, we need to obtain a sparse solution to the high-dimensional regression problem in Equation (ref). This means we frame the detection of structural breaks as a model selection problem and use available methods from this strand of the literature. To reduce the dimensionality of the estimation problem, we assume that breaks occur for all coefficients simultaneously. This allows us to treat all regressors at each point in time as one group. We can therefore apply the group lasso estimator proposed by YuanLin2006 to achieve a sparse solution. As our first step, we minimize the objective function,
to obtain the group lasso estimator for $\boldsymbol{\theta}(T)$ which is henceforth denoted by $\tilde{\boldsymbol{\theta}}(T) = \arg \min_{\boldsymbol{\theta}(T)} \, Q^*$. $\lambda_T$ is the tuning parameter and $\Vert \cdot \Vert$ denotes the $L_2$-norm. Unfortunately, the group lasso estimator inherits the same problems, namely estimation inefficiency and model selection inconsistency, as the plain lasso estimator. Similar to the idea first presented in Zou2006, we reestimate the objective function with individual coefficient weights to alleviate this problem and to try to reduce the number of falsely detected breaks. The statistical properties of adaptive group lasso estimators for a fixed number of groups are investigated in WangLeng2008. Since we have a diverging set of breakpoint candidates, least squares estimation of the full model is not feasible. However, we show that group lasso is a consistent estimator for non-zero parameter changes giving us appropriate weights for a second step adaptive group lasso estimation. This approach is similar to the ideas put forth in WeiHuang2010, HorowitzHuang2013, SchmidtSchweikert2019, and Behrendt2020.
As will be demonstrated later, the group lasso estimator only slightly overselects breaks under the right tuning. The algorithm employed to estimate $\tilde{\boldsymbol{\theta}}(T)$ allows to pre-specify the maximum number of breakpoint candidates $M$, i.e.\ the maximum number of non-zero groups in $\tilde{\boldsymbol{\theta}}(T)$, and the minimum distance between breaks. Since the group lasso overselects breaks in the first step, $M$ should be set large enough to encompass all true breakpoints and some additional falsely selected non-zero groups. This condition guarantees that $\tilde{\boldsymbol{\theta}}(T)$ always contains $MN$ elements. In turn, $TN - MN$ columns of $\boldsymbol{Z}_T$ corresponding to zero coefficients are eliminated during the first step to result in the $T \times MN$ design matrix $\boldsymbol{Z}_S$. Hence, for given $M \ll T$, the column size of the new design matrix is substantially smaller than the original size $TN$ and does not longer depend on the sample size. This allows us to further assume that all eigenvalues of $\boldsymbol{\Sigma}_S = \boldsymbol{Z}_S' \boldsymbol{Z}_S / T^2$ are contained in the interval $[c_*, c^*]$, where $c_*$ and $c^*$ are two positive constants. This means that we can relate to a restricted eigenvalue condition similar to BickelRitovTsybakov2009 for the second step estimation. While the restricted eigenvalue condition in general does not hold for change-point settings, the dimension reduction of the first step allows us to postulate this assumption for our reduced design matrix. It should be noted that our assumption for the second step estimation is not restrictive for empirical applications because the notion of a long-run equilibrium relationship implies a maximum number of breaks and a minimum regime length. A minimum regime length is further justified by the minimum subsample size needed to precisely estimate parameter changes. Consequently, $M$ should be chosen so that the average regime length in case of equidistantly-spaced breaks still guarantees enough observations per regime to estimate all coefficient changes.
We follow WangLeng2008 and define the adaptive group lasso objective function
where $\gamma > 0$ and $w_i$ are the group-specific weights assigned as follows
and set $0 \times \infty = 0$. $\tilde{\boldsymbol{\theta}}_{S,i}$, $i = 1, \dots, |\mathcal{A}_T| + 1$ denotes the non-zero group lasso coefficient estimates obtained from optimizing the objective function in Equation (ref). The remaining $M - |\mathcal{A}_T| - 1$ group elements of $\tilde{\boldsymbol{\theta}}_S$ can be filled with zero groups as long as their selected indices lead to $\boldsymbol{\Sigma}_S$ being a positive definite matrix for all $T$.
We denote the estimator minimizing $Q(\boldsymbol{\theta}_S)$ with $\hat{\boldsymbol{\theta}}_S = \arg \min_{\boldsymbol{\theta}_S} \, Q$. The weight of the first coefficient is usually set to zero to ensure that the system is cointegrated with a cointegrating vector different from $(1,\boldsymbol{0}')'$ if no structural break occurs. Eliminating columns from the initial design matrix requires a mapping of our second step indices to recover the original indices. For notational convenience, we use the mapping $g: \mathbb{N} \to \mathbb{N}, \; i \mapsto g(i) = t_i$, where $t_i$ is the breakpoint corresponding to the index $i$, for this purpose and define the index set $\bar{\mathcal{A}}^*$ ($\mathcal{A}^*$) to pick out the elements that correspond to truly non-zero coefficients (parameter changes).
We note that the major computing cost comes from the first step group lasso estimation considering a large number of observations as potential breakpoints. The second step represents a marginal addition to the total computing time if the first step estimation was sufficiently successful in eliminating inactive breakpoint candidates. The interested reader may consult ChanYauZhang2014 for a detailed discussion of computational complexity in this context.\footnote{While the Bai-Perron algorithm needs at most $O(T^2)$ operations, the group LARS algorithm used to solve Equation (ref) has computational burden of order $O(M^3 + MT)$. Hence, the group LARS algorithm has a stronger dependence on the maximum number of breaks, whereas the Bai-Perron algorithm only depends on the number of observations. This implies that the Bai-Perron algorithm is better suited for small to moderate samples with a potentially large number of breaks, often found in linear regressions modelling short-run relationships. Instead, the group LARS algorithm is well-suited for large sample sizes and a small to moderate number of structural breaks which is often found for long-run relationships in the presence of structural change.}
In the following, we study the asymptotic properties of our adaptive group lasso estimator. To discuss asymptotic properties, we need to impose some further assumptions about the location and magnitude of active breakpoints.
Assumption (ref)(i) requires that the length of the regimes between breaks increases with the sample size and in the same proportions to each other. This allows us to consistently detect and estimate the true break fractions as it makes the break dates asymptotically distinct Perron2006. The first inequality of Assumption (ref)(ii) is a necessary condition to ensure that a structural break occurs at $t_j^0$. We do not consider small breaks with local-to-zero behaviour in this setting (see BaiLumsdaineStock1998 for assumptions used in this context). This assumption is not believed to be restrictive for the intended empirical applications where applied researchers aim to estimate the long-run equilibrium to obtain the error correction term, i.e., the cointegration residuals, for their follow-up analysis. Essentially, they need optimal in-sample forecasts in terms of mean squared error of the cointegrating regression under structural instability to consistently estimate these residuals. BootPick2019 show that in-sample forecasts are largely unaffected by local-to-zero breaks. The second part excludes the possibility of infinitely large parameter changes. Assumption (ref)(iii) implies that the number of active breaks is less than the number of observations and the smallest eigenvalue of $\boldsymbol{\Sigma}_{\bar{\mathcal{A}}}$ is greater than or equal to $C$ by letting $\boldsymbol{m}_j = 0$ for $j \in \bar{\mathcal{A}}^c$. Consequently, Assumption (ref)(iii) ensures that $\boldsymbol{\Sigma}_{\bar{\mathcal{A}}}$ is positive definite for all $T$. This is only then the case if $\boldsymbol{Z}_{T, \bar{\mathcal{A}}}$ contains columns which are sufficiently distinct. This in turn means that the intervals between breaks need to be sufficiently large for all $T$. It is important to note that Assumption (ref)(i) can be deduced as an implication of Assumption (ref)(iii) and we need Assumption (ref)(iii) exclusively for the first step estimation. Our second step estimation requires only Assumption (ref)(i) and (ii) as long as consistent weights are available.
First, we need to show that the initial estimator provides consistent weights for the second step adaptive lasso procedure HuangMaZhang2008. The following theorem provides a consistency result for the group lasso estimator in cointegrating regressions with (possibly) multiple structural breaks.
Theorem (ref) shows that it is crucial to let the tuning parameter $\lambda_T$ grow at the right rate. However, this rate provides only limited practical guidance towards the choice of $\lambda_T$. We follow Kock2016, QianSu2016 and SchmidtSchweikert2019 and propose to select $\lambda_T$ by minimizing an information criterion in the form of
where $SSR$ is the sum of squared residuals resulting from the group lasso estimation of Equation (ref) and $|\mathcal{A}_T|$ gives the number of non-zero breakpoint candidates. The penalty function $\rho_T$ allows for different choices. While Kock2016 suggests to use the BIC for potentially nonstationary autoregressive models which corresponds to $\rho_T = \log(T)/T$, QianSu2016 propose to use $\rho_T = 1/\sqrt{T}$ for the estimation of structural breaks in stationary time series regressions. In this paper, we follow SchmidtSchweikert2019 and employ a modified BIC according to WangLiLeng2009 which incorporates the additional factor $\log \log d^*_T$ where $d^*_T$ denotes the total amount of coefficients in the full model. This modification of the BIC accounts for the fact that the true model must be found in situations where the number of coefficients diverges.
For the next theorem, we temporarily assume that the exact number of breaks is known. This assumption will help us to provide an important consistency result for the estimated location of breakpoints. We note that this temporary assumption will be relaxed for our main results.
The previous result is an important building block for our main results. Next, we prove that the group lasso estimator yields a set of estimated breakpoints for which the number of selected breaks is greater than the true number of breaks almost surely when the exact number of breakpoints is unknown. Further, we evaluate the consistency of estimated breakpoints using the Hausdorff distance between the set of estimated breakpoints and the set of true breakpoints. We follow Boysen2009 and define $d_H(A, B) = \underset{b \in B}{\max} \, \underset{a \in A}{\min} |b - a|$ with $d_H(A, \emptyset) = d_H(\emptyset, B) = 1$, where $\emptyset$ is the empty set. The following theorem shows that the set of estimated breakpoints converges to the set of true breakpoints under the Hausdorff distance.
Finally, we consider the asymptotic properties of the adaptive group lasso estimator with weights obtained from our first step estimation. We note that Theorem (ref) allows us to bound the number of breakpoint candidates by a constant. Hence, the dimensionality of the model selection problem no longer depends on the sample size.
In this section, we conduct simulation experiments to assess the adequacy of our technical results in (ref). We investigate the finite sample performance of our adaptive group lasso procedure with respect to the accuracy in finding the exact number of breaks, their location and the magnitude of parameter changes. We consider model specifications with one, two and four breakpoints, respectively. The following DGP is employed to model a multivariate cointegrated system with multiple structural breaks,
where $\boldsymbol{X}_t = (X_{1t}, X_{2t}, \dots, X_{Nt})'$ and $\Sigma = diag(\sigma^2_{\omega})$, i.e.\ the innovations of our generated random walk processes have identical normal distributions. $\mu$ is a non-zero intercept and $\boldsymbol{\beta}_t = (\beta_{1t}, \beta_{2t}, \dots, \beta_{Nt})$ is a time-varying slope coefficient vector with non-zero baseline value and a finite number of breaks. We note that $cov(\vartheta_t, \omega_t) = 0$, i.e.\ our regressors are strictly exogenous and the asymptotic bias reported in Theorem (ref) is non-existent.
Naturally, the ability of all structural break estimators to detect breaks depends on the overall signal strength. NiuHaoZhang2015 define signal strength in change-point models by $S = m_{\beta}^2 I_{\min}$, where $I_{\min} = \underset{1 \leq j \leq m_0+1}{\min} \vert t_j - t_{j-1} \vert$ is the minimum distance between breaks and $m_{\beta} = \underset{1 \leq j \leq m_0 + 1}{\min} \Vert \boldsymbol{\beta}_j - \boldsymbol{\beta}_{j-1} \Vert$ is the minimum jump size. For our main simulations concerned with consistency of the adaptive group lasso estimator, we use equal jump sizes for multiple breaks and locate the breaks with equidistant spacing between them. Hence, overall signal strength is a linear function of the sample size in our simulations. We use a baseline value of two and a jump size of two which is equal to the standard deviation of the regression error term. Simulations with a better signal-to-noise ratio yield more precise estimates for all sample sizes.
In (ref), we report our results for $N = 2$ regressors. We specify our model for one break located at $\tau = 0.5$, two breaks at $\tau = (0.33, 0.67)$ and four breaks at $\tau = (0.2, 0.4, 0.6, 0.8)$ to have an equidistant spacing on the unit interval. We first compute the percentages of correct estimation (pce) of the number of breaks $m$ and measure the accuracy of the break date estimation conditional on the correct estimation of $m$. For this matter, we compute the average Hausdorff distance and divide it by $T$ (hd/T) to compare the values across different sample sizes. The corresponding figures in our tables are reported in percentages. As $T$ grows larger, the number of breaks is detected with increasing precision and the distance between estimated breakpoints and true breakpoints declines to nearly zero. Parameter estimates are already very accurate at small sample sizes. As expected, the parameter changes of models with fewer breakpoints can be estimated more precisely than those of models with a larger number of breakpoints, as indicated by larger standard deviations obtained for the latter at all sample sizes.\footnote{Results for $N = 1$ reported in SchmidtSchweikert2019 show a very similar pattern. We find that it is slightly more difficult to detect the correct number of breaks in regressions with multiple regressors although the jump size measured as the Euclidean distance is equal for both settings.} Comparing these results with those obtained for the Bai-Perron algorithm\footnote{KejriwalPerron2008, KejriwalPerron2010 obtain estimates of the parameters using the dynamic programming algorithm of BaiPerron2003 with no modification since the algorithm itself is valid irrespective of the nature of the regressors and errors given that it detects break dates that minimize the global sum of squared residuals in a regression.}, where the number of breaks is determined via the BIC, we find that both approaches perform similarly well. The results are reported in (ref). While the Bai-Perron algorithm estimates the true break fractions slightly more accurately, parameter changes on average have larger standard deviations at all samples sizes. The number of structural breaks is estimated with identical accuracy.\footnote{ We can confirm the theoretical claims made about the computational complexity of the Bai-Perron algorithm in comparisons to the group LARS algorithm in (ref). We obtain the following computational times (in seconds) for both algorithms. First, using the simulation set-up for (ref) and a sample size of $T = 1000$, we have $M = 1$: (gLARS: 2.61, BP: 11.03), $M = 2$: (gLARS: 9.24, BP: 11.41), $M = 4$: (gLARS: 21.27, BP: 15.36). Here, we find that the Bai-Perron algorithm is more robust to a larger number of breaks in terms of computational time. Second, we increase the sample size to $T = 10000$ and record the following times, $M = 1$: (gLARS: 4.04, BP: 1013.15), $M = 2$: (gLARS: 27.35, BP: 1317.45), $M = 4$: (gLARS: 35.23, BP: 1563.55). In this case, we can confirm that the group LARS algorithm is much more computationally efficient for large sample sizes. All simulations are computed on a computer with an Intel i5-6500 CPU at 3.20GHz and 16GB RAM.}
Next, we investigate if dynamic augmentation according to Saikkonen1991 and StockWatson1993 yields consistent coefficient estimates if the strict exogeneity condition of our main results is violated. To do so, we follow KejriwalPerron2008 and draw the vector $(\vartheta_t, \omega_{1t}, \omega_{2t})'$ jointly from a multivariate normal distribution with zero mean and covariance matrix
Using this configuration, the strict exogeneity condition is violated for both regressors but the regressors are still generated by independent processes. If we attempt to detect and estimate structural breaks without dynamic augmentation, we still detect breakpoints precisely but obtain strongly biased coefficient estimates. In (ref), we find the corresponding results after the inclusion of $l=1$ and $l=2$ leads and lags. Now, we can recover the number, location and magnitude of all breakpoints with similar accuracy compared to our simulations under strict exogeneity.
In (ref), we consider partial breaks in the cointegrating vector. We use a model specification according to the DGP in Equation (ref) with $N = 2$ regressors and induce partial structural breaks through $\beta_{1t}$ only. Our estimator is applied estimating a full structural change model without prior knowledge that $\beta_{2t}$ is constant over the sampling period. Again, we observe that the number of breaks, their timing and their magnitude is consistently estimated.\footnote{Comparing the results with those obtained for the Bai-Perron algorithm (not reported), we again find that both approaches detect the number of structural breaks with identical accuracy. The Bai-Perron algorithm estimates the true break fractions slightly more accurately, but parameter changes have larger standard deviations at for samples sizes $T=200$ and $T=400$.} The distance between the set of estimated breakpoints and true breakpoints is larger than in the full break setting in (ref). This result is not surprising considering that the break magnitudes for partial breaks are smaller making it more difficult for the adaptive group lasso procedure to detect the true location of the breaks. Consequently, these results also help us to assess how the break magnitude influences the detection rates. Reducing the Euclidean distance from 2 to $\sqrt{2}$, roughly doubles the average Hausdorff distance. The convergence rates for zero parameter changes in $\beta_{2t}$ is almost identical to the convergence rate observed for the non-zero parameter changes in $\beta_{1t}$. This is naturally driven by the joint evaluation of all regressors in each group. Unlike bi-level estimators proposed in HuangMaXieZhang2009 and BrehenyHuang2009, the adaptive lasso procedure is not able to shrink coefficients within active groups to zero. Hence, the usual convergence rate for non-zero coefficients apply. In these cases, the convergence rate for $\beta_{2t}$ could in principle be increased if our procedure was extended to feature bi-level shrinkage. However, this is beyond the scope of this paper and is not investigated further at this point.
Finally, we investigate how sensitive our procedure is to break fractions located near the boundary of the unit interval. While the properties of tests for structural changes in the literature depend strongly on the trimming parameter BaiPerron2006, our method to recover breaks should be more robust in this regard. We only need some lateral trimming to ensure that the first and last regimes identified by our adaptive lasso procedure comprise a sufficiently large number of observations to estimate regime-dependent coefficients.\footnote{For our main results, reported in (ref) and (ref), we follow KejriwalPerron2008 and use a 15% lateral trimming to compare both methods.} The results for breaks near the boundary are summarized in (ref). The first and second panel considers one break located at $\tau = 0.1$ and $\tau = 0.9$, respectively. The pce and average Hausdorff distances over all sample sizes clearly show that a break located close to the beginning of the sample is more difficult to detect than a break located at the end of the sample. GregoryHansen1996 and Schweikert2019 report similar findings for their grid search algorithms. To investigate this further, we consider two breaks located at $\tau = (0.1, 0.9)$ in panel three of (ref). Here, we find that the pce is quite low compared to our main results with equidistant spacing of breakpoints. The first break is estimated less accurately than the second break which can be explained by the fact that parameter changes are measured from one regime to the next and that only a relatively small number of observations is available to estimate the break at $\tau = 0.1$.
The results of our first series of boundary experiments imply that it might be possible to relax our trimming restrictions and assume an asymmetric lateral trimming where the first regime must contain sufficiently many observation, say 5% of the sample, while the end of the sample does not necessarily have to be excluded. We apply a 0.05/0 trimming and estimate breaks located at $\tau = (0.1, 0.95)$. The results for this trimming strategy are presented in panel four of (ref). The break at $\tau = 0.95$ can still be accurately detected, however the standard errors of the parameter changes increase due to the smaller number of observations in the last regime. We conclude that trimming is not necessary to detect breaks located at the end of the sample. Still, we suggest to set a minimum number of observations per regime to ensure that parameter changes are estimated precisely.
In this section, we apply our proposed methodology to the US money demand function. Particularly, we estimate a long-run money demand specification and investigate the presence of long-run instabilities in a cointegrating framework. Juselius2006 considers the condition $M/P = L(Y, R)$ for equilibrium in the money market, which relates $M/P$, the ratio of nominal money balances to price levels, to real income $Y$ and the short term nominal interest rate $R$. Two competing empirical specifications are considered in the literature, namely, the semi-log and the log-log specification. The latter is given by $L(Y, R) = \alpha Y^{\beta_1} R^{\beta_2}$, where $\alpha$ is a constant, $\beta_1$ is the income-elasticity assumed to be unity and $\beta_2 < 0$ is the interest-elasticity.\footnote{Recall that in general, coefficients of log-transformed variables in cointegrating regressions should not be interpreted as elasticities Johansen2005. Only in the special case when those variables are strongly exogeneous, it is allowed to use a ceteris paribus interpretation for the corresponding coefficients.} For our empirical application, we choose a log-log specification which has been found to fit quite well to US data Lucas2000, Bae2007, Ireland2009, MoglianiUrga2018. We extend the dataset used by Maki2012 to span the period from January 1959 to December 2018. Monthly data are obtained from the Federal Reserve Bank of St. Louis. We consider the empirical US money demand function,
where $m^*_t$ and $y_t$ denote the natural logarithm of the ratio of nominal money balances to price levels, and the natural logarithm of real income, respectively. According to the log-log specification, we employ the natural logarithm of the short term nominal interest rate, denoted by $r_t$. $u_t$ denotes the equilibrium error of the money demand function if the system is cointegrated. We use M2 as nominal money, the consumer price index as prices, and the index of industrial production as real income. For the interest rate, we use the 6-month Treasury bill rate. All time series are tested for a unit root using the Dickey-Fuller test. The results, which are not reported, support the assumption that all variables are integrated of order one and we can continue our cointegration analysis.
First, we assume constancy of the parameters and ignore potential structural breaks. Estimation of the long-run equilibrium equation yields coefficients $\hat{\mu} = -0.05$, $\hat{\beta}_1 = 0.80$ and $\hat{\beta}_2 = -0.08$. Dynamic augmentation of the cointegrating regression with two leads and lags each, does not change the coefficient values. The Engle-Granger test based on an ADF regression yields the t-ratio $-0.063$ which does not lead to a rejection of the null hypothesis at the 10% level. Similar results can be obtained for the Phillips-Ouliaris test and the Johansen test. Although it is implausible from a theoretical standpoint that the system is not cointegrated, at least our estimated coefficients have the expected sign and magnitude for post-war data. The estimated income-elasticity measured by $\beta_1$ is slightly below the theoretically expected value. The interest-elasticity of money demand, measured by $\beta_2$ is expected to be negative. Lucas2000 considers $-0.3$, $-0.5$ and $-0.7$ as values of $\beta_2$ and finds that $\beta_2 = -0.5$ gives the best fit for US data. Meltzer1963, Lucas1988, HoffmanRasche1991 and StockWatson1993 find empirical evidence consistent with the theoretical expectation that income-elasticity of money demand is unity and interest-elasticity is relatively high. Ball2001 studies subperiods from 1903 to 1994 and argues against a stable long-run money demand. Further empirical studies have pointed out the presence of structural instability in US money demand for sample periods including data from the 1990s and 2000s TelesZhou2005, Wang2011, LucasNicolini2015. Potential nonlinearities in the functional form are investigated, for example, by ChenWu2005 and JawadiSousa2013. However, we take the perspective that the linear cointegrating regression in Equation (ref) approximates the data well if we simultaneously account for (multiple) parameter changes during the sample period.
A three-dimensional scatterplot of the data in (ref) reveals that the relationship between $r_t$, $y_t$ and $m^*_t$ has changed during the sampling period. We observe at least three two-dimensional surfaces which correspond to distinct long-run levels from which $m^*_t$ does not persistently deviate. However, if we consider linear cointegration without the possibility of structural breaks, we infer from (ref) that the residual series exhibits a clear trend during the latter half of the sample. We note that the presence of structural breaks might mask the cointegrating relationship. Next, we compare several previously mentioned structural break models with our model selection approach. The GregoryHansen1996 test indicates a breakpoint at 2008 m06 but does not reject the null hypothesis at the 10% level. Because the GH-test does not model structural breaks under the null hypothesis, this means that the timing of the indicated breakpoint is not informative. The Hatemi-J2008 test indicates two breakpoints at 1992 m01 and 2008 m06. The null hypothesis of no cointegration can be rejected at the 5% level if these breakpoints are taken into account. The maximum number of breaks chosen for the Maki2012 test is five. It selects the breakpoints at 1986 m05, 1992 m04, 2004 m05, 2008 m11, 2014 m03 and rejects the null hypothesis of no cointegration at the 1% level. We initially also start with a maximum of five breakpoints for our adaptive group lasso procedure. However, imposing a minimum regime length of one year to precisely estimate the parameter changes and dynamically augmenting the cointegrating regression results in a model specification with three breakpoints. The final estimates yield break dates 1992 m07, 2005 m12, and 2015 m11.\footnote{The corresponding breakpoint estimates using the Bai-Perron algorithm are almost identically located at 1991 m10, 2004 m07, and 2014 m06.}
The income-elasticity from 1959 m01 to 1992 m07 is estimated to be $0.95$ and the interest-elasticity amounts to $-0.10$ for the same period. These estimates correspond to the theoretical predictions formulated in Juselius2006 and to the results reported in empirical papers considering this sample period Lucas1988, StockWatson1993, Lucas2000. The first breakpoint leads to an income-elasticity reduction from $0.95$ to $0.89$ while the interest-elasticity remains largely unchanged. A partial decoupling of money demand from income might be explained by the begin of the costly Gulf War and a sharp increase in US debt. In turn, the second breakpoint at 2005 m12 has a negligible effect on the income-elasticity ($0.89$ to $0.90$) but results in a larger reduction of the interest-elasticity from $-0.10$ to $-0.07$. This breakpoint can be related to the beginning Global Financial Crisis of 2007-2008. It must be emphasized at this point that estimated break dates might be affected by the usual lead and lag effects, since parameter changes are representative for the following regime. In the aftermath of the Global Financial Crisis, the Federal Reserve implemented a zero interest rate policy. Consequently, the variation in the interest rate for this period approached zero which naturally reduced the interest-elasticity of money demand. After 2015 m11, the expected interest-elasticity does no longer achieve a good fit to the data and increases to $0.01$. In contrast, the income-elasticity is very close to unity ($0.97$).
Accounting for structural breaks, as indicated by the adaptive group lasso procedure, yields a residual series which much more resembles being generated by a stationary process than the original OLS residual. (ref) illustrates that the residual series does not exhibit a visible trend. The speed of adjustment after equilibrium errors is now $-0.097$ which means that roughly 10% of long-run deviations are corrected each period.
In this paper, we propose a penalized regression approach to the problem of detecting an unknown number of structural breaks and their location in cointegrating regressions. Our estimator eliminates irrelevant breakpoints from a set of candidate breakpoints and, hence, follows a top-down approach regarding the estimation of structural breaks. Practitioners should apply this new methodology in complement to the Bai-Perron algorithm which follows a bottom-up approach, i.e.\ sequentially increasing the number of breaks. Due to the importance of finding the right model specification with respect to the number and location of structural breaks, either approach can serve as a valuable robustness check of the model specification chosen by the other approach. Ideally both approaches should indicate the same breakpoints which would mean that the chosen model specification is sufficiently sparse (bottom-up) and does not ignore important breaks (top-down).
We can show the important theoretical result that the adaptive group lasso estimator has nonstandard oracle properties in settings with a diverging number of breakpoint candidates. This means that the estimator determines the true number of non-zero parameter changes with probability tending to one and consistently estimates their location. The corresponding parameter changes are estimated with the same convergence rate that least squares estimators would have under full information of the number and location of breaks.
The present paper does not consider cointegration testing. It is unclear how optimal cointegration test can be constructed from the proposed penalized regression approach. An attempt to design such cointegration tests has been made by SchmidtSchweikert2019 for a single regressor. Our results depend critically on the stationarity assumption about the error term. Hence, it is required to establish the existence of a cointegration relationship before the penalized regression is estimated. Practitioners should employ cointegration tests which are robust to the presumed number of breaks during the sample period.
Further extensions include the use of bi-level selection via the group fused lasso HuangMaXieZhang2009, BrehenyHuang2009 to estimate partial breaks more efficiently, and the possibility to detect structural breaks in system-based approaches with multiple equilibria BaiLumsdaineStock1998, Qu2007.
I thank Florian Stark, Alexander Schmidt, Markus M��ler, Timo Dimitriadis and the participants of the Doctoral Seminar in Econometrics in T�bingen, German Statistical Week in Trier, ZU Methodenkolloquium in Friedrichshafen, THE Christmas Workshop in Stuttgart, Seminar at Maastricht University, and the 2nd CSL Symposium in Stuttgart for valuable comments and suggestions. Further, I thank Maike Becker and Manuel Huth for excellent research assistance.