Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,176 characters · 12 sections · 86 citation commands
1515Testing for Common Breaks in a Multiple Equations System
\thispagestyle{empty}\setcounter{page}{0}\baselineskip=18.0pt
\pagestyle{plain}
Issues related to structural change have been extensively studied in the statistics and econometrics literature CH1997Book, Perron2006HB. In the last twenty years or so, substantial advances have been made in the econometrics literature to cover models at a level of generality that makes them relevant across time-series applications in the context of unknown change points. For example, Bai1994JTSA, Bai1997REStat studies the least squares estimation of a single change point in regressions involving stationary and/or trending regressors. BaiPerron1998Emtca, BaiPerron2003JAE extend the testing and estimation analysis to the case of multiple structural changes and present an efficient algorithm. Hansen1992JBES and KejriwalPerron2008JoE consider regressions with integrated variables. Andrews1993Emtca and HallSen1999JBES consider nonlinear models estimated by generalized method of moments. Bai1995ET, Bai1998JSPI studies structural changes in least absolute deviation regressions, while Qu2008JoE, SuXiao2008SPL and OkaQu2011JoE analyze structural changes in regression quantiles. HallHanBoldea2012JoE and PerronYamamoto2014, PerronYamamoto2015JAE consider structural changes in linear models with endogenous regressors. Studies about structural changes in panel data models include Bai2010JoE, Kim2011JoE, BaltagiFengKao2015JoE and QianSu2016JoE for linear panel data models and BreitungEickmeier2011JoE, ChenEtAl2016RES, CorradiSwanson2014JoE, HanInoue2015ET and YamamotoTanaka2015JoE for factor models.
The literature on structural breaks in a multiple equations system includes BaiLumsdaineStock1998, Bai2000AEF and QuPerron2007Emtca, among others. Their analysis relies on a common breaks assumption, under which breaks in different basic parameters (regression coefficients and elements of the covariance matrix of the errors) occur at a common location or are separated by some positive fraction of the sample size (i.e., asymptotically distinct).\footnote{ The concept of common breaks here is quite distinct from the notion of co-breaking or co-trending HatanakaYamada2003Book, HendryMizon1998EE. In this literature, the focus is on whether some linear combination of series with breaks do not have a break, a concept akin to that of cointegration. } BaiLumsdaineStock1998 assume a single common break across equations for a multivariate system with stationary regressors and trends as well as for cointegrated systems. For the case of multiple common breaks, Bai2000AEF analyzes vector autoregressive models for stationary variables and QuPerron2007Emtca cover multiple system equations, allowing for more general stationary regressors and arbitrary restrictions across parameters. Under the framework of QuPerron2007Emtca, KurozumiTu2011JoE propose model selection procedures for a system of equations with multiple common breaks and EoMorley2015QE consider a confidence set for the common break date based on inverting the likelihood ratio test. In this literature, it has been documented that common breaks allow more precise estimates of the break dates in multivariate systems. Given unknown break dates, however, an issue of interest for most applications concerns the validity of the assumption of common breaks.\footnote{ The common breaks assumption is also used in the literature on panel data Bai2010JoE, Kim2011JoE, BaltagiFengKao2015JoE. In this paper, we consider a multiple equations system in which the number of equations are relatively small, and thus panel data models are outside our scope. However, testing for common breaks in a system with a large number of equations is an interesting avenue for future research. } To our knowledge, no test has been proposed to address this issue.
Our paper addresses three outstanding issues about testing for common breaks. First, we propose a quasi-likelihood ratio test under a very general framework.\footnote{ One may also consider other type of tests, such as LM-type tests. The literature on structural breaks, however, documents that even though LM-type tests have simple asymptotic representations, they tend to exhibit poor finite sample properties with respect to power. Thus, this paper focuses on the LR test DP2008JoE, KimPerron2009JoE, PerronYamamoto2016ER. } We consider a multiple equations system under a likelihood framework with normal errors, though the limit distribution of the proposed test remains valid with non-normal, serially dependent and heteroskedastic errors. Our framework allows integrated regressors and trends as well as stationary regressors as in BaiLumsdaineStock1998 and also accommodates multiple breaks and arbitrary restrictions across parameters as in QuPerron2007Emtca. Thus, our results apply for general systems of multiple equations considered in existing studies. A case not covered in our framework is when the regressors depend on the break date. This occurs when considering joint segmented trends and this issue was analyzed in Kim2017WP.
Second, we propose a test for common breaks not only across equations within a multivariate system, but also within an equation. As in BaiLumsdaineStock1998, the issue of common breaks is often associated with breaks occurring across equations, whereas one may want to test for common breaks in the parameters within a regression equation, whether a single equation or a system of multiple equations are considered. More precisely, the null hypothesis of interest is that some subsets of the basic parameters share one or more common break dates, so that each regime is separated by some positive fraction of the sample size. Under the alternative hypothesis, the break dates are not the same and also need not be separated by a positive fraction of the sample size, or be asymptotically distinct.
Third, we derive the asymptotic properties of the quasi-likelihood and the parameter estimates, allowing for the possibility that the break dates associated with different basic parameters may not be asymptotically distinct. This poses an additional layer of difficulty, since existing studies establish the consistency and rate of convergence of estimators only when the break dates are assumed to either have a common location or be asymptotically distinct, at least under the level of generality adopted here. Moreover, we establish the results in the presence of integrated regressors and trends as well as stationary regressors. This is by itself a noteworthy contribution. These asymptotic results will allow us to derive the limit distribution of our test statistic under the null hypothesis and also facilitate asymptotic power analyses under fixed and local alternatives. We can show that our test is consistent under fixed alternatives and also has non-trivial local power.
There is one additional layer of difficulty compared to Bai and Perron (1998) or Qu and Perron (2007). In their analysis, it is possible to transform the limit distribution so that it can be evaluated using a closed form solution and thus critical values can be tabulated. Here, no such solution is available and we need to obtain critical values for each case through simulations. This involves simulating the Wiener processes with consistent parameter estimates and evaluating each realization of the limit distribution with and without the restriction of common breaks. While it is conceptually straightforward and quick enough to be feasible for common applications, the procedure needs to be repeated many times to obtain the relevant quantities and can be quite computationally intensive. This is because we need to search over many possible combinations of all the permutations of the break locations for each replication of the simulations. To reduce the computational burden, we propose an alternative procedure based on the particle swarm optimization method developed by Eberhart1995 with the Karhunen-Lo{\`e}ve representation of stochastic processes. Our simulation results suggest that the test proposed has reasonably good size and power performance even in small samples under both computation procedures. Also, we apply our test to inflation series, following the work of Clark2006JAE to illustrate its usefulness.
The remainder of the paper is as follows. Section 2 introduces the models with and without the common breaks assumption and describes the estimation methods under the quasi-likelihood framework. Section 3 presents the assumptions and asymptotic results including the asymptotic null distribution and asymptotic power analyses. Section 4 examines the finite sample properties of our procedure via Monte Carlo simulations. Section 5 presents an empirical application and Section 6 concludes. An appendix contains all the proofs.
In this section, we first introduce models for a multiple equations system with and without common breaks. Subsequently, we describe the quasi-likelihood estimation method assuming normal errors and then propose the quasi-likelihood ratio test for common breaks. For illustration purpose, we also discuss some examples.
As a matter of notation, \textquotedblleft $\overset{p}{\rightarrow }$ \textquotedblright $\ $denotes convergence in probability, \textquotedblleft $\overset{d}{\rightarrow }$\textquotedblright $\ $convergence in distribution and \textquotedblleft $\Rightarrow $\textquotedblright\ weak convergence in the space $D[0,\infty)\ $under the Skorohod topology. We use $\mathbb{R}$, $\mathbb{Z}$ and $\mathbb{N}$ to denote the set of all real numbers, all integers and all positive integers, respectively. For a vector $x$, we use $\| \cdot \|$ to denote the Euclidean norm (i.e., $\|x\|= \sqrt{x' x}$), while for a matrix $A$, we use the vector-induced norm (i.e., $\|A\|= \sup_{x \not = 0}\|A x\|/\|x\|$). Define the $L_{r}$-norm of a random matrix $X$ as $\left\Vert X\right\Vert _{r}=(\sum_{i}\sum_{j}E\left\vert X_{ij}\right\vert ^{r})^{1/r}$ for $r\geq 1 $. Also, $a \wedge b=\min \{a,b\}$ and $a \vee b=\max \{a,b\}$ for any $a, b \in \mathbb{R}$. Let $\circ$ denote the Hadamard product (entry-wise product) and let $\otimes$ denote the Kronecker product. Define $\mathbbm{1}_{ \{\cdot \}}$ as the indicator function taking value one when its argument is true, and zero otherwise and $e_{i}$ as a unit vector having 1 at the $i^{th}$ entry and 0 for the others. We use the operator $\mathrm{vec}(\cdot)$ to convert a matrix into a column vector by stacking the columns of the matrix and the operator $\mathrm{tr}(\cdot)$ to denote the trace of a matrix. The largest integer not greater than $a \in \mathbb{R}$ is denoted by $[a]$ and the sign function is defined as $\mathrm{sgn}(a)= -1, 0, 1$ if $a >0$, $a=0$ or $a <0$, respectively.
Let the data consist of observations $\{(y_{t}, x_{tT})\}_{ t=1}^{T}$, where $y_{t}$ is an $n \times 1$ vector of dependent variables and $x_{tT}$ is a $q \times 1$ vector of explanatory variables for $n, q \in \mathbb{N}$ with a subscript $t$ indexing a temporal observation and $T$ denoting the sample size. We allow the regressors $x_{tT}$ to include stationary variables, time trends and integrated processes, while scaling by the sample size $T$ so that the order of all components is the same. In what follows, we consider
Here, $z_{t}$, $\varphi(t/T)$ and $w_{t}$ respectively denote vectors of stationary, trending and integrated variables with sizes being $q_{z}{\times} 1$, $q_{\varphi}{\times} 1$ and $q_{w}{\times} 1$, so that $q\equiv q_{z}+q_{\varphi}+q_{w}$.\footnote{ The normalization is simply a theoretical device to reduce notational burden. Without it, we would need to handle different convergence rates of the estimates by introducing additional notations. } Also,
where $w_{0}\ $is assumed, for simplicity, to be either $O_{p}(1)$ random variables or fixed finite constants, and $u_{wt}$ is a vector of unobserved random variables with zero means. We label the variables $z_{t}$ as $I(0)$ if the partial sums of the associated noise components satisfy a functional central limit theorem, while we label a variable as $I(1)$ if it is the accumulation of an $I(0)$ process. We discuss in more details the specific conditions in Section 3.
We first explain the case of common breaks through a model in which all of the parameters including those of the covariance matrix of the errors change, i.e., a pure structural change model. The model of interest is a multiple equations system with $n$ equations and $T$ time periods, excluding the initial conditions if lagged dependent variables are used as regressors. We denote the break dates in the system by $T_{1}, \dots, T_{m}$ with $m$ denoting the total number of structural changes and we use the convention that $T_{0}=0$ and $T_{m+1}=T$.
With a subscript $j$ indexing a regime for $j=1,...,m+1$, the model is given by
where $I_{n}$ is an $n \times n$ identity matrix, $S$ is an $n q \times p$ selection matrix with full column rank, $\beta_{j}$ is a $p \times 1$ vector of unknown coefficients, and $u_{t}$ is an $n \times 1$ vector of errors having zero means and covariance matrix $\Sigma_{j}$.\footnote{ An example of models involving stationary and integrated variables is the dynamic ordinary least squares method to estimate cointegrating vectors Saikkonen1991ET, SW1993Emtca. } The selection matrix $S$ usually consists of elements that are $0$ or $1$ and, hence, specifies which regressors appear in each equation, although in principle it is allowed to have entries that are arbitrary constants. To ease notation, define the $n \times p$ matrix $X_{tT}:=$ $S' (x_{tT}\otimes I_n)$ so that ((ref)) becomes, for $j=1,...,m+1$,
The set of basic parameters in the $j^{th}$ regime consists of the coefficients $\beta_{j}$ and the elements of the covariance matrix $\Sigma _{j}$, and we denote it by $\theta_{j}:=(\beta_{j}, \Sigma_{j})$ for each regime $j = 1, \dots, m+1$. We use $ \Theta_{j} \subset \mathbb{R}^{p} \times \mathbb{R}^{n\times n} $ to denote a parameter space for $\theta_{j}$ and we also define a product space $\Theta:= \Theta_{1} \times \dots \times \Theta_{m+1}$ for $\theta:=(\theta_{1}, \dots, \theta_{m+1})$. In model ((ref)), we allow for the imposition of a set of $r$ restrictions through a function $R: \Theta \to \mathbb{R}^{r}$, given by
Note that the equation in ((ref)) can impose restrictions both within and across equations and regimes. Thus the model in ((ref)) with some restrictions of the form ((ref)) can accommodate structural break models other than a pure structural change model, such as partial structural change models in which a part of the basic parameters are constant across regimes. For a discussion of how general the framework is, see Qu and Perron (2007).
Next, we consider a pure structural change model allowing for the possibility that the break dates are not necessarily common across basic parameters. In the equations system with the $p \times 1$ vector of coefficients, we can assign each coefficient an index from $1$ to $p$ and we then group the $p$ indices into disjoint subsets $\mathcal{G}_{1}, \dots, \mathcal{G}_{G} \subset \{1, \dots, p\}$ with $G$ standing for the total number of groups, such that coefficients indexed by elements of $\mathcal{G}_{g}$ share the same break dates for each group $g=1,\dots, G$ and $\cup _{g=1}^{G}\mathcal{G}_{g} = \{1,...,p\}$. Given a collection $\{\mathcal{G}_{g}\}_{g=1}^{G}$, we define, for $(g, j) \in \{1, \dots, G\} {\times} \{1,...,m+1\}$,
Without loss of generality, we assume that the elements of the covariance matrix $\Sigma_{j}$ have break dates that are common to those in the last group $G$. If none of the regression coefficients change at the same time as the elements of the covariance matrix $\Sigma_{j}$, then $\mathcal{G}_{G}$ is simply an empty set.\footnote{ We assume that the different elements of the covariance matrix of the errors change at the same time. The results can be extended to the case where different parameters have distinct break dates, although additional notations would be needed. For the sake of notational simplicity, we only consider the case where the break dates are common within all elements of the covariance matrix. } Here, we introduce groups of basic parameters to accommodate a wide range of empirical applications under our framework. Sometimes, researchers have economic models of interest or empirical knowledge that suggest specific parameter groups having common breaks. Even when one has no knowledge to form parameter groups, our analysis can be applied by considering all basic parameters as separate groups.
To denote the break date for regime $j$ and group $g$, we use $k_{gj}$ for $(g, j) \in \{1, \dots, G\} \times \{1, \dots, m\}$ with the convention that $k_{g0}=0$ and $k_{g,m+1}=T$ for any $g= 1, \dots, G$. Also, define a collection of break dates as,
The regression model can be expressed as one depending on time-varying basic parameters according to the collection $\mathcal{K}$:
where $\beta_{t,\mathcal{K}} := \sum_{g=1}^{G}\beta _{g, t, \mathcal{K}} $ and $E[u_{t}u_{t}^{\prime }] = \Sigma_{t,\mathcal{K}}$ with
for $(g, j) \in \{1, \dots, G\} {\times} \{1,...,m+1\}$. We also use $\theta_{t, \mathcal{K}}:= (\beta_{t, \mathcal{K}}, \Sigma_{t, \mathcal{K}}) $ to denote time-varying basic parameters depending on the collection of break dates $\mathcal{K}$. Thus the restrictions ((ref)) can be imposed on the system ((ref)) to accommodate more general models with structural breaks as in the one with common breaks.
In model ((ref)), the basic parameters, break dates and the number of breaks are unknown and have to be estimated. To select the total number of structural changes, we can apply existing sequential testing procedures or information criteria. For example, if the breaks are common within each equation under both null and alternative hypotheses, but may differ across equations (see Example 1 below), sequential testing procedures proposed by BaiPerron1998Emtca can be used to select the number of structural changes in each equation of a system BaiPerron1998Emtca. In a similar way, the sequential testing procedure in QuPerron2007Emtca can be applied for sets of equations of a system separately. In order to handle more complex cases, we can alternatively use the Bayesian information criterion or the minimum description length principle as in KurozumiTu2011JoE, Lee2000JASA and aue2011. Because we use the likelihood framework, a likelihood function with a relevant penalty can be computed with the use of genetic algorithms Davis2001, which consistently selects the number of structural breaks, as in Lee2000JASA and aue2011. Thus, our analysis in what follows focuses on unknown basic parameters and breaks dates, given a total number of structural changes.
We use a $0$ superscript to denote the true values of the parameters in both ((ref)) and ((ref)). Thus, the true basic parameters and break dates in ((ref)) are denoted by $\{(\beta_{j}^{0},\Sigma_j^0)\}_{j=1}^{m+1}$ and $\{T_{j}^{0}\}_{j=1}^{m}$, respectively, with the convention that $T_{0}^{0}=0$ and $T_{m+1}^{0}=T$, whereas the ones in ((ref)) are denoted by $\{\beta_{1j}^{0}, \dots, \beta_{Gj}^{0}, \Sigma_{j}^{0}\}_{j=1}^{m+1}$ and $\mathcal{K}_{g}^{0}:=(k_{g1}^{0},\dots, k_{gm}^{0})$ with $k_{g0}^{0}=0$ and $k_{g,m+1}^{0}=T$ for $g =1, \dots, G$. Also let $\mathcal{K}^{0}:=\{\mathcal{K}_{1}^{0}, \dots, \mathcal{K}_{G}^{0} \}$. Given a collection of break dates $\mathcal{K}$, let $\theta_{t, \mathcal{K}}^{0}:= (\beta_{t, \mathcal{K}}^{0}, \Sigma_{t, \mathcal{K}}^{0}) $ with a $0$ superscript to denote time-varying true basic parameters $\theta^{0}$, where $\theta^{0}:=(\theta_{1}^{0}, \dots, \theta_{m+1}^{0})$ with $\theta_{j}^{0} := (\beta_{j}^{0},\Sigma_j^0) $ for $j=1, \dots, m+1$.
We consider the quasi-maximum likelihood estimation method with serially uncorrelated Gaussian errors for model ((ref)) with restrictions given by ((ref)).\footnote{ Our framework includes OLS-based estimation by setting the covariance matrix to be an identity matrix. } Given the collection of break dates $\mathcal{K}$ and the basic parameters $\theta$, the Gaussian quasi-likelihood function is defined as
where
To obtain maximum likelihood estimators, we impose a restriction on the set of permissible partitions with a trimming parameter $\nu>0$ as follows\footnote{ For the asymptotic analysis, the trimming value $\nu$ can be an arbitrary small constant such that a positive fraction of the sample size $T\nu$ diverges at rate $T$. }:
This set of permissible partitions ensures that there are enough observations between any break dates within the same group $\mathcal{K}_{g}$, while it accommodates the possibility that the break dates across different groups are not separated by a positive fraction of the sample size.
We propose a test for common breaks under the quasi-likelihood framework. The null hypothesis of common breaks in model ((ref)) can be stated as
and the alternative hypothesis is
The set of permissible partitions under the null hypothesis can be expressed as
The test considered is simply the quasi-likelihood ratio test that compares the values of the likelihood function with and without the common breaks restrictions. The quasi-maximum likelihood estimates under the null hypothesis, denoted by $(\tilde{\mathcal{K}}, \tilde{\theta} )$, can be obtained from the following maximization problem with a restricted set of candidate break dates:
where $\tilde{\mathcal{K}}:= (\tilde{\mathcal{K}}_{1}, \dots, \tilde{\mathcal{K}}_{G})$ with $\tilde{\mathcal{K}}_{g}:=(\tilde{k}_{1}, \dots, \tilde{k}_{m})$ for all $g=1,\dots, G$, $\tilde{\theta}:=(\tilde{\beta}, \tilde{\Sigma})$ with $\tilde{\beta}:=(\tilde{\beta}_{1}, \dots, \tilde{\beta}_{m+1})$ and $\tilde{\Sigma}:=(\tilde{\Sigma}_{1}, \dots, \tilde{\Sigma}_{m+1})$. Also, the quasi-maximum likelihood estimates under the alternative, denoted by $(\hat{\mathcal{K}}, \hat{\theta} )$, are obtained from the following problem:
where $\hat{\mathcal{K}}:= (\hat{\mathcal{K}}_{1}, \dots, \hat{\mathcal{K}}_{G})$ with $\hat{\mathcal{K}}_{g}:=(\hat{k}_{g1}, \dots, \hat{k}_{gm})$ for $g=1,\dots, G$, $\hat{\theta}:=(\hat{\beta}, \hat{\Sigma})$ with $\hat{\beta}:=(\hat{\beta}_{1}, \dots, \hat{\beta}_{m+1})$ and $\hat{\Sigma}:=(\hat{\Sigma}_{1}, \dots, \hat{\Sigma}_{m+1})$. Using the estimates $\hat{\theta}$, we can define $\hat{\beta}_{gj}$ as in ((ref)) and $\hat{\theta}_{t, \mathcal{K}}:= (\hat{\beta}_{t, \mathcal{K}}, \hat{\Sigma}_{t, \mathcal{K}}) $ as in ((ref)) given a collection of break dates $\mathcal{K}$.
We define the quasi-likelihood ratio test for common breaks as
For the asymptotic analysis, it is useful to employ a normalization by using the log-likelihood function evaluated at the true parameters $(\mathcal{K}^{0}, \theta^{0})$ and we consider
where $ \ell_{T}(\mathcal{K}, \theta) := \log L_{T}(\mathcal{K},\theta) - \log L_{T}(\mathcal{K}^{0},\theta^{0}) $ for any $(\mathcal{K}, \theta) \in \Xi_{\nu} \times \Theta$. The common break test $CB_{T}$ depends on two log-likelihoods with and without the common breaks assumption. The break date estimates $\tilde{\mathcal{K}}$ under the null hypothesis are required to either have common locations or be separated by a positive fraction of the sample size. Without common breaks restrictions, however, the break date estimates $\hat{\mathcal{K}}$ are simply allowed to be distinct but not necessarily separated by a positive fraction of the sample size across groups. This will be important since the setup of Bai2000AEF and QuPerron2007Emtca requires the maximization to be taken over asymptotically distinct elements and their proof for the convergence rate of the estimates relies on this premise. Hence, we will need to provide a detailed proof of the convergence rate under this less restrictive maximization problem (see Section 3).
Given that the notation is rather complex, it is useful to illustrate the framework explained in the preceding subsection via examples.
Example 1 (changes in intercepts): We consider a two-equations system of autoregressions with structural changes in intercepts, for $j=1,2$,
where $(u_{1t}, u_{2t})'$ have a covariance matrix $\Sigma$. In this model, the basic parameters except the intercepts are assumed to be constant and the intercepts change at a common break date $T_{1}$. In equation ((ref)), we have $x_{tT}=(1, y_{1,t-1}, y_{2, t-1})'$, $\beta_{j}=(\mu_{1j}, \alpha_{1j}, \mu_{2j}, \alpha_{2j})'$ and $E[u_{t} u_{t}'] = \Sigma_{j}$. The selection matrix $S=<s_{ij}>$ is a $6\times 4$ matrix taking value 1 at the entries $s_{11}$,$s_{22}$, $s_{33}$ and $s_{64}$ and 0 elsewhere. Also, by setting $R(\theta) = \big(\alpha_{11}-\alpha_{12}, \alpha_{21}-\alpha_{22}, \mathrm{vec}(\Sigma_{1})- \mathrm{vec}(\Sigma_{2}) \big )' = 0$ in ((ref)), we impose restrictions on the basic parameters so that a partial structural change model is considered with no changes in the autoregressive parameters and the covariance matrix of the errors. On the other hand, when we allow the possibility that break dates can differ across the two equations as in the model ((ref)), we consider the following system, for $j=1,2$,
Here, we separate $\beta_{j}$ into $\beta_{1j}=(\mu_{1j}, \alpha_{1j}, 0, 0)'$ and $\beta_{2j}=(0, 0, \mu_{2j}, \alpha_{2j})'$, so that we can set $\mathcal{G}_{1}=\{1, 2\}$ and $\mathcal{G}_{2}=\{3, 4\}$. We have two possibly distinct break dates $k_{11}$ and $k_{21}$ for the parameter groups $\{\beta_{1j}\}_{j=1}^{2}$ and $\{(\beta_{2j}, \Sigma_{j})\}_{j=1}^{2}$, respectively. We address the issue of testing the null hypothesis $H_{0}: k_{11} = k_{21}$ against the alternative hypothesis $H_{1}: k_{11} \not = k_{21}$.
Example 2 (a single equation model): Consider a single equation model:
for $T_{j-1} +1 \le t \le T_{j}$ with $j = 1, 2, 3$, where $u_{1t}$ denotes the error term with $E[u_{1t}] = 0$ and $E[u_{1t}^{2}] = \sigma_{j}^{2}$. In this example, the basic parameters other than the intercepts have two structural changes. Under model ((ref)) with break dates $T_{1}$ and $T_{2}$, we have $x_{tT}=(1, z_{1t}, t/T, T^{-1/2}w_{1t})'$, $S= I_{4}$, $\beta_{j}=(\mu_{j}, \alpha_{j}, \gamma_{j}, \rho_{j})'$. Restrictions of the form ((ref)) are imposed by the function $R(\theta) = (\mu_{1} - \mu_{2}, \mu_{2} - \mu_{3} )' = 0$. We consider a test for common breaks against the alternative that all coefficients change at distinct break dates, while the coefficient $\rho_{j}$ and the variance $\sigma_{j}^{2}$ change at the same break dates. In this case, we separate $\beta_{j}$ into three vectors $\beta_{1j}=(\mu_{j}, \alpha_{j}, 0, 0)$, $\beta_{2j}=(0, 0, \gamma_{j}, 0)$ and $\beta_{3j}=(0, 0, 0, \rho_{j})$. For these parameters groups, we assign a set of break dates $\mathcal{K}_{g}=(k_{g1}, k_{g2})$ for $g=1,\dots, 3$ and we set $\mathcal{G}_{1}=\{1, 2\}$, $\mathcal{G}_{2}=\{3\}$ and $\mathcal{G}_{3}=\{4\}$. The break dates for the last group, $\mathcal{K}_{3}$, are also the ones for the variance. This example shows that our framework can accommodate common breaks not only across equations in a system but also within an equation.
This section presents the relevant asymptotic results. We first provide the convergence rates of the estimates of the break dates and the basic parameters, allowing for the possibility that the break dates of different basic parameters may not be asymptotically distinct. This condition is substantially less restrictive than the ones usually assumed in the existing literature and particularly includes the assumption of common breaks as a special case. Next, we provide the limiting distribution of the quasi-likelihood ratio test for common breaks under the null hypothesis. Finally, we provide asymptotic power analyses of the test under a fixed alternative as well as a local one. Our result shows non-trivial asymptotic power.
We consider the case where we obtain the quasi-likelihood estimates $(\hat{\mathcal{K}}, \hat{\theta})$ as in ((ref)), using the observations $\{(y_t , x_{tT} )\}_{t=1}^{T}$ generated by model ((ref)) with collections of true parameter values $(\mathcal{K}^{0}, \theta^{0})$. The results presented in this subsection can apply for the estimates obtained from the model under the null hypothesis since it is a special case of the setup adopted. To obtain the asymptotic results, the following assumptions are imposed.
Assumptions:
Assumption A1 ensures that there is no local collinearity problem so that a standard invertibility requirement holds if the number of observations in some sub-sample is greater than $k_{0}$, not depending on $T$. Assumption A2 determines the dependence structure of $\{\zeta_{t} \otimes \eta_{t}\}$, $\{\eta_{t}\eta_{t}' - I_{n}\}$ and $\{w_{0} \otimes \eta_{t}\}$ to guarantee that they are short memory processes and have bounded fourth moments. The assumptions are imposed to obtain a functional central limit theorem and a generalized HajekRenyi1955AMH type inequality that allow us to derive the relevant convergence rates. Assumption A2 also specifies that the stationary regressors are contemporaneously uncorrelated with the errors and that a constant term is included in $z_{t}$. The former is a standard requirement to obtain consistent estimates and the latter is for notational simplicity since the results reported below are the same without a constant term.\footnote{ One can use the usual ordinary least squares framework to simply estimate the break dates and test for structural change even in the presence of the correlation between the stationary regressors and the errors PerronYamamoto2015JAE. One may also use a two-stage least squares method if relevant instrumental variables are available HallHanBoldea2012JoE, PerronYamamoto2014. }\footnote{ When a constant term is not included in $z_{t}$, in contrast to Assumption A2, one additionally needs to assume that the sequence $\{\eta_{t}\}_{t \in \mathbb{Z}}$ satisfies the same mixing and moment conditions as in Assumption A2(a). } It is important to note that no assumption is imposed on the correlation between the innovations to the $I(1)$ regressors and the errors. Hence, we allow endogenous $I(1)$ regressors. Assumption A3 ensures that $\lambda_{gj}^{0} -\lambda_{g,j-1}^{0} > \nu$ holds for every pair of group and regime $(g, j)$ and thus implies asymptotically distinct breaks within each parameter group, but not necessarily across groups. Assumption A4 implies a shrinking shifts asymptotic framework whereby the magnitudes of the shifts converge to zero as the sample size increases. This condition is necessary to develop a limit distribution theory for the estimates of the break dates that does not depend on the exact distributions of the regressors and the errors, as commonly used in the literature Bai1997REStat, BaiPerron1998Emtca, BaiLumsdaineStock1998. Assumption A5 implies that the data are generated by a model with a finite conditional mean and innovations having a non-degenerate covariance matrix.
As stated above, the break dates are estimated from a set $\Xi_{\nu}$, which requires candidate break dates to be separated by some fraction of the sample size only within parameter groups. Thus, we cannot appeal to the results in Bai2000AEF and Qu and Perron (2007) about the rate of convergence of the estimates, and more general results are needed. The following theorem presents results about the convergence rates of the estimates.
This theorem establishes the convergence rates obtained in BaiPerron1998Emtca, BaiLumsdaineStock1998, Bai2000AEF and QuPerron2007Emtca, while assuming less restrictive conditions regarding the optimization problem and the time-series properties of the regressors.
The importance of these results is that they will allow us to analyze the properties of our test under compact sets for the parameters, namely, for some $M>0$,
We also have a result that expresses the restricted likelihood in two parts: one that involves only the break dates and the true values of the coefficients; the other involving the true values of the break dates, the basic parameters and the restrictions. Thus, asymptotically the estimates of the break dates are not affected by the restrictions imposed on the coefficients, while the limiting distributions of these estimates are influenced by the restrictions.
The result in Theorem 2 implies that when analyzing the asymptotic properties of the break date estimates, one can ignore the restrictions in ((ref)). This will prove especially convenient to obtain the limit distribution of our test. Since the quasi-likelihood ratio test can be expressed as a difference of two normalized log likelihoods evaluated at different break dates, the second term on the right-hand side of ((ref)) is canceled out in the test statistic. The result in Theorem 2 has been obtained in Bai2000AEF for vector autoregressive models and QuPerron2007Emtca for more general stationary regressors, when break dates are assumed to either have a common location or be asymptotically distinct. We establish the results, allowing for the possibility that the break dates associated with different basic parameters may not be asymptotically distinct, and thus expand the scope of prior work such as BaiLumsdaineStock1998, Bai2000AEF and QuPerron2007Emtca.
We now establish the limit distribution of the quasi-likelihood ratio test under the null hypothesis of common breaks in ((ref)). To this end, let the data consist of the observations $\{(y_t , x_{tT} )\}_{t=1}^{T}$ from model ((ref)) with true basic parameters $\theta^{0}=(\beta^{0}, \Sigma^{0})$ and true break dates $\mathcal{T}^{0}$ consisting of $T_{1}^{0}, \dots, T_{m}^{0}$. Theorem (ref)(a) shows that, uniformly in $(g, j) \in \{1, \dots, G\} \times \{1, \dots, m \} $, there exists a sufficiently large $M$ such that $|\hat{k}_{gj} - T_{j}^{0}| \le M v_{T}^{-2}$ and $|\tilde{k}_{j} - T_{j}^{0}| \le M v_{T}^{-2}$ with probability approaching 1. This implies that we can restrict our analysis to an interval centered at the true break $T_{j}^{0}$ with length $2 M v_{T}^{-2}$ for each regime $j \in \{1,\dots ,m\}$. More precisely, given a sufficiently large $M$, we have that $\theta_{t, \hat{\mathcal{K}}}^{0} = \theta_{t, \mathcal{T}^{0}}^{0} $ and $\theta_{t, \tilde{\mathcal{K}}}^{0} = \theta_{t, \mathcal{T}^{0}}^{0} $ for all $t \not \in \cup_{j=1}^{m}[T_{j}^{0}-Mv_{T}^{-2}, T_{j}^{0}+Mv_{T}^{-2}]$, with probability approaching 1. This follows since the break dates estimates are asymptotically in neighborhoods of the true break dates; hence that there are some miss-classification of regimes around the neighborhoods, while the regimes are correctly classified outside of the neighborhoods. This together with Theorem (ref) yields that, under the null hypothesis specified by ((ref)),
where $\overline{k}_{j}:=\max \{k_{1j},\dots, k_{Gj}, T_{j}^{0}\}$, $\underline{k}_{j}:=\min \{k_{1j},\dots, k_{Gj}, T_{j}^{0}\}$, and $\Xi_{M, H_{0}} = \Xi_{M} \cap \Xi_{\eta, H_{0}}$. Under the null hypothesis, the true break dates $T_{1}^{0}, \dots, T_{m}^{0}$ are separated by some positive fraction of the sample size and we can obtain the limit distribution of the common break test by separately analysing terms of the test for each neighborhood of the true break date. We consider a shrinking framework under which the break date estimates $\hat{k}_{gj}$ and $\tilde{k}_{j}$ diverge to $\infty$ as $v_{T}$ decreases and thus an application of a Functional Central Limit Theorem for each neighborhood yields a limit distribution of the test which does not depend on the exact distributions. To derive the limit distribution, we make the following additional assumptions.
Assumptions:
Assumption A6 rules out trending variables in the stationary regressors $z_{t}$. Assumption A7 is mild in the sense that the conditions allow for substantial conditional heteroskedasticity and autocorrelation. It can be shown to apply to a large class of linear processes including those generated by all stationary and invertible ARMA models. This assumption is useful to describe the asymptotic behavior of the test and in particular to characterize the limit distribution. Here, we introduce some processes used later. For each $j=1, \dots, m$, let $\mathbb{V}_{z \eta ,j}^{(1)}(\cdot)$ and $\mathbb{V}_{z \eta ,j}^{(2)}(\cdot)$ be Brownian motions defined on the space $D[0,\infty)^{nq}$ with zero means and covariance functions given by, for $l = 1, 2$ and for $s_{1},s_{2}>0$,
where $ \bar{V}_{T, z\eta,j}^{(1)} := (\Delta T_{j}^{0})^{-1/2} \sum_{t=T_{j-1}^{0}+1}^{T_{j}^{0} } (z_{t}\otimes \eta_{t}) $ and $ \bar{V}_{T, z\eta,j}^{(2)} := (\Delta T_{j+1}^{0})^{-1/2} \sum_{t=T_{j}^{0}+1}^{T_{j+1}^{0}} (z_{t}\otimes \eta_{t}) $. Similarly, define $\mathbb{V}_{\eta\eta ,j}^{(1)}(\cdot)$ and $\mathbb{V}_{\eta\eta ,j}^{(2)}(\cdot)$ as Brownian motions defined on the space $D[0,\infty)^{n^2}$ with zero means and covariance functions given by, for $l = 1, 2$ and for $s_{1}, s_{2} >0$,
where $ \bar{V}_{T, \eta\eta,j}^{(1)} := (\Delta T_{j}^{0})^{-1/2} \sum_{t=T_{j-1}^{0}+1}^{T_{j}^{0}} (\eta_{t}\eta_{t}' {-} I_{n}) $ and $ \bar{V}_{T, \eta\eta,j}^{(2)} := (\Delta T_{j+1}^{0})^{-1/2} \sum_{t=T_{j}^{0}+1}^{T_{j+1}^{0}} (\eta_{t}\eta_{t}' {-} I_{n}) $. We define the following two-sided Brownian motions
Under Assumption A2, $z_{t}$ is assumed to include a constant term and the process $\mathbb{V}_{z\eta ,j}^{(l)} (\cdot)$ includes some process depending purely on $\{\eta_{t}\}$. We denote it by $\mathbb{V}_{\eta ,j}^{(l)} (\cdot)$ for each $l=1,2$ and also define a two-sided Brownian motion, denoted by $\mathbb{V}_{\eta,j}(\cdot)$, as before.
Assumption A8 requires the integrated regressors to follow a homogeneous distribution throughout the sample. Allowing for heterogeneity in the distribution of the errors underlying the $I(1)$ regressors would be considerably more difficult, since we would, instead of having the limit distribution in terms of standard Wiener processes, have time-deformed Wiener processes according to the variance profile of the errors through time; see, e.g., CT2007. This would lead to important complications given that, as shown below, the limit distribution of the estimates of the break dates depends on the whole time profile of the limit Wiener processes. It is possible to allow for trends in the $I(1)$ regressors. The limiting distributions of the test to be derived will remain valid under different Wiener processes Hansen1992JBES. The positive definiteness of the matrix $\Omega _{w}$ rules out cointegration among the $I(1)$ regressors and is needed to ensure a set of regressors that has a positive definite limit.
Assumption A9 is quiet mild and is sufficient but not necessary to obtain a manageable limit distribution of the test. It requires the independence of most Wiener processes described above. Condition (a) ensures that the autocovariance structure of the $I(0)$ regressors and the errors are uncorrelated with the $I(1)$ variables. This guarantees that $\mathbb{V}_{z\eta,j}(\cdot)$ and $\mathbb{V}_{w,j}(\cdot )$ are uncorrelated and thus independent because of Gaussianity. Without these conditions, the analysis would be much more complex. Similarly, the conditions (b) and (c) imply the independence between $\mathbb{V}_{z\eta,j}(\cdot )$ and $\mathbb{V}_{\eta \eta}(\cdot )$. See KejriwalPerron2008JoE for more details.
In order to characterize the limit distribution of $ CB_{T}$ it is useful to first state some preliminary results about the limit distribution of some quantities. For $s \in \mathbb{R}$ and for $j = 1, \dots, m$, let $\overline{T}_{j}(s):=\max \{T_{j}(s),T_{j}^{0}\}$ and $\underline{T}_{j}(s):=\min\{T_{j}(s),T_{j}^{0}\}$ where $T_{j}(s):= T_{j}^{0} + [s v_{T}^{-2}]$. For $s, r \in \mathbb{R}$, we define $ B_{T,j}(s, r) {:=} v_{T}^{2} \sum_{t= \underline{T}_{j}^{0}(s)+ 1}^{ \overline{T}_{j}(s)} X_{tT} (\Sigma_{j + \mathbbm{1}_{ \{T_{j}(r) < t\} }}^{0} )^{-1} X_{tT}' $ and $ W_{T,j}(s, r) {:=} v_{T} \sum_{t= \underline{T}_{j}^{0}(s)+ 1}^{ \overline{T}_{j}(s)} X_{tT} (\Sigma_{j + \mathbbm{1}_{ \{T_{j}(r) < t\} }}^{0} )^{-1} u_{t} $ for $j \in \{1, \dots, m\}$.
The theorem below presents the main result of the paper concerning the limit distribution of the test statistic, which can be expressed as the difference of the maxima of a limit process with and without restrictions implied by the assumption of common breaks.
The limit distribution in Theorem (ref) is quite complex and depends on nuisance parameters. However, they can be consistently estimated and it is easy to show that the coverage rates will be asymptotically valid provided $\sqrt{T}$-consistent estimates are used instead of the true values. The various quantities can be estimated as follows: for $\Delta \tilde{k}_{j}:= \tilde{k}_{j}-\tilde{k}_{j-1}$, we can use $\tilde{Q}_{zz,j} =(\Delta \tilde{k}_{j})^{-1}\sum_{t=\tilde{k}_{j-1} + 1}^{\tilde{k}_{j}}z_{t}z_{t}^{\prime } $, $\tilde{\mu}_{z,j} = (\Delta \tilde{k}_{j})^{-1}\sum_{t=\tilde{k}_{j-1}+1}^{\tilde{k}_{j}}z_{t}$, $\Delta \tilde{\beta}_{j}:= \tilde{\beta}_{j}-\tilde{\beta}_{j-1}$ and $\tilde{\Sigma}_{j} = (\Delta \tilde{k}_{j})^{-1}\sum_{t=\tilde{k}_{j-1} + 1}^{\tilde{k}_{j}}\tilde{u}_{t}\tilde{u}_{t}^{\prime } $, $\tilde{\Delta}_{gj}:= \big \{ \|\Delta \tilde{\beta}_{j}\|^2 + \mathrm{tr}\big ( ( \Delta \tilde{\Sigma}_{j})^2\big) \big \}^{-1/2} \sum_{l \in \mathcal{G}_{g}} e_{l}\circ \Delta \tilde{\beta}_{j+1} $ and $\tilde{\Upsilon}_{j}:= \big \{ \|\Delta \tilde{\beta}_{j}\|^2 + \mathrm{tr}\big ( ( \Delta \tilde{\Sigma}_{j})^2\big) \big \}^{-1/2} \Delta \tilde{\Sigma}_{j} $, where $\Delta \tilde{\beta}_{j}:= \tilde{\beta}_{j}-\tilde{\beta}_{j-1}$ and $\Delta \tilde{\Sigma}_{j}:= \tilde{\Sigma}_{j}-\tilde{\Sigma}_{j-1}$. Also, the estimates of the long run variances of $\{z_{t}\otimes \eta_{t}\}$ and $\{\eta_{t}\eta_{t} - I_{n}\}$ can be constructed using a method based on a weighted sum of sample autocovariances of the relevant quantities, as discussed in Andrews1991Emtca, for instance. Though only $\sqrt{T}$-consistent estimates of $(\beta ,\Sigma )$ are needed, it is likely that more precise estimates of these parameters will lead to better finite sample coverage rates. Hence, it is recommended to use the estimates obtained imposing the restrictions in ((ref)) even though imposing restrictions does not have a first-order effect on the limiting distribution of the estimates of the break dates.
In some cases, the limit distribution of the common breaks test can be derived and expressed in a simpler manner. For illustration purpose, our supplemental material states the limit distribution of the test under the setup of Examples 1 and 2. When the covariance matrix is constant over time (i.e., $\Sigma_{j}^{0} = \Sigma^{0}$ for $j=1, \dots, m+1$), the limit distribution above can be further simplified as stated in the following corollary.
As another immediate corollary to Theorem (ref), when no integrated variables are present, the limit distribution of the test for a common break date only involves the pre and post break date regimes, as is the case for the limit distribution of the estimates when multiple breaks are present BaiPerron1998Emtca. Also, the above result can be easily extended to test the hypothesis of common break dates for a part of the parameter groups, while the break dates of the other groups are not necessarily common. We illustrate the application of the test for common breaks in ((ref)) and its variant through an application in Section 5.
As discussed in Section 1, there is one additional layer of difficulty compared to Bai and Perron (1998) or Qu and Perron (2007). In their analysis, the limit distribution can be evaluated using a closed form solution after some transformation, while no such solution is available here and thus we need to resort simulations to obtain the critical values. This involves first simulating the Wiener processes appearing in the various Brownian motion processes by partial sums of $i.i.d.$ normal random vectors (independent of each others given Assumption A9). One can then evaluate one realization of the limit distribution by replacing unknown values by their estimates as stated above. The procedure is then repeated many times to obtain the relevant quantiles. While conceptually straightforward, this procedure is nevertheless computationally intensive. The reason is that for each replication we need to search over many possible combinations of all the permutations of the locations of the break dates. The procedure suggested is nevertheless quick enough to be feasible for common applications involving testing for few common break dates but the computational burden increases exponentially with the number of common breaks being tested. In Section 4, we propose an alternative approach to alleviate this issue and examine its performance.
In this subsection, we provide an asymptotic power analysis of the test statistic $CB_{T}$ when using a critical value $c_{\alpha }^{\ast}$ at the significance level $\alpha$ from the asymptotic null distribution $CB_{\infty}$. As a fixed alternative hypothesis, we consider, for some $\delta>0$
Given that $k_{gj}^{0} = [T\lambda_{gj}^{0}]$ for $(g, j) \in \{1, \dots, G\} {\times} \{1,...,m\}$ under Assumption A3, the above condition is asymptotically equivalent to $ \max_{1 \le g_{1}, g_{2} \le G} |\lambda_{g_{1},j}^{0} - \lambda_{g_{2},j}^{0}| \ge \delta $ for some $j = 1, \dots, m$, and thus can be considered as a fixed alternative hypothesis in term of break fractions. As a local alternative hypothesis, we consider
for some constant $M>0$, where $v_{T}$ satisfies the condition in Assumption A4. We can also express ((ref)) as $ \max_{1 \le g_{1}, g_{2} \le G} | \lambda_{g_{1},j}^{0} - \lambda_{g_{2},j}^{0} | \ge M (\sqrt{T}v_{T})^{-2} $ for some $j = 1, \dots, m$. The following theorem shows that the proposed test statistic is consistent against fixed alternatives and also has non-trivial local power against local alternatives.
This section provides simulation results about the finite sample performance of the test in terms of size and power. We first consider a direct simulation-based approach to obtain the critical values and then a more computationally efficient algorithm. As a data generating process (DGP), we adopt a similar setup to the one used in BaiLumsdaineStock1998, namely a bivariate autoregressive system with a single break in intercepts as in Example 1. Hence, only the intercepts are allowed to change at some dates $k_{i1}$ for equation $i \in \{1,2\}$. We test the null hypothesis $H_0: k_{11} = k_{21}$ against the alternative hypothesis $H_1: k_{11} \not= k_{21}$. The number of observations is set to $T=100$, and we use $500$ replications. Results are reported for autoregressive parameters $\alpha \in \{0.0,0.4, 0.8\}$. We set $\mu_{i1}=1$ and let $\delta _{i}:=\mu_{i2}-\mu_{i1}$, the magnitude of the mean shift, take values $\{0.50, 0.75, 1.00, 1.25, 1.50\}$.
A direct simulation-based approach: We first present results when we resort direct simulations to obtain the critical values, which involves simulating the Wiener processes by partial sums of i.i.d. normal random vectors and searching over all possible combinations of the break dates. Given the computational cost, we choose a simple setup and focus on limited cases. To examine the empirical sizes and power, we here consider the errors $(u_{1t},u_{2t})^{\prime }$ following $i.i.d.$ $N(0,I_{2})$ and we use 3,000 repetitions to generate the critical values.
We first examine the empirical rejection frequencies under the null hypothesis that $k_{11} = k_{21} = 50$ with a trimming parameter $\nu=0.15$. The results are reported in Table 1 for nominal sizes of 10%, 5% and 1%. First, when the autoregressive process has no or moderate dependency ($\alpha=0.0$ or $\alpha=0.4$), the empirical size of the test is either slightly conservative or close to the nominal size. Given the small sample size, this size property is satisfactory. When the autoregressive parameter is close to the boundary of the non-stationary region, e.g. $\alpha =0.8$, as expected there are some liberal size distortions. When the magnitudes of the breaks are small, the test tends to over-reject the null hypothesis. This is due to the fact that for very small breaks the break date estimates are quite imprecise and are more likely to be affected by the highly dependent series than the break sizes themselves, so that the test depends on the log likelihoods evaluated outside neighborhoods of the true break dates. When the magnitude of the break sizes increases, the size of the test quickly approaches the nominal level. These results are encouraging given the small sample size.
To analyze power, we also set $\mu_{i1}=1$, while we consider values $\{0.50, 1.00, 1.50\}$ for the magnitude of the mean shift. The break date in the first equation is kept fixed at $k_{1}=35$, while the break date in the second equation takes values $k_{2}=35,40,45,50,55$. The power is a function of the difference between the break dates, $k_{2}-k_{1}$. The results are presented in Figure 1, where the horizontal axis in each box represents the difference $k_{2}-k_{1}$ and the vertical axis shows the empirical rejection frequency. As before, when the magnitudes of the breaks are small, the data are not informative enough to reject the common breaks null hypothesis and the test has little power. However, when the magnitudes of the changes reach 1, the power increases rapidly as the distance between the break dates increases. The results are qualitatively similar for all values of $\alpha $ considered.
An alternative approach: The direct simulation-based procedure involves a combinatorial optimization problem and the computational burden increases exponentially with the number of common breaks being tested. Such a procedure may be feasible for a small number of breaks in a parsimonious system. However, in more general cases, it may be prohibitive. Hence, we also propose an alternative approach that solves this problem, using heuristic algorithms that find approximate, if not optimal, solutions. Because heuristic algorithms have mainly been developed to optimize functions having explicit forms, we use the Karhunen-Lo{\`e}ve (KL) representation of stochastic processes, which expresses a Brownian motion as an infinite sum of sine functions with independent Gaussian random multipliers Bosq2012. A truncated series of the KL representation was used to obtain critical values by Durbin1970 and Krivyakov1978, among others. Similarly, we use a truncated series with 500 terms and apply a change of variables to approximately obtain an explicit form of the objects being maximized in the limit distribution of the common breaks test. Also, we use the particle swarm optimization method, which is an evolutionary computation algorithm developed by Eberhart1995.\footnote{ For our simulations, we use the particle swarm algorithm “particleswarm” of the Matlab Global Optimization Toolbox. We also tried the genetic algorithm “ga” from Matlab and found that the two algorithms yield very similar, frequently the same, critical values, while the particle swarm algorithm is faster. }
We examine the performance of the common breaks test using the alternative algorithm under various setups in order to show that similar good finite sample properties are obtained compared to the direct optimization method. In addition to the setup used above, we consider a trimming value $\nu = 0.10$, a pair of break dates (35, 35) and normal errors with correlation coefficient being 0.5 across equations. Columns (1)-(4) of Table 2 present empirical rejection frequencies under the null hypothesis for a nominal size of 5%. Whether the errors are correlated or not, the empirical size of the test is either conservative or close to the nominal size in cases of moderate dependency ($\alpha=0.0$ or $\alpha=0.4$). Also the trimming parameter has little impact. With uncorrelated errors, there are size distortions in cases of high dependency ($\alpha =0.8$) and small break sizes. When the errors are correlated, however, the empirical sizes get closer to the nominal level in all cases. This is likely due to efficiency gains from using a SUR estimation method. Columns (5)-(6) of Table 2 report the empirical power for the case $(k_{1}, k_{2}) = (35, 50)$ and the results show satisfactory power, comparable to the direct method.
In this section, we apply the common breaks test to inflation series, following Clark2006JAE. He analyzes the persistence of a number of disaggregated inflation series based on the sum of the autoregressive (AR) coefficients in an AR model, and documents that the persistence is very high and close to one without allowing for a mean shift, whereas the persistence declines substantially when allowing for one. Although such features have been documented theoretically in the literature Perron1990JBES, he finds that the decline in persistence is more pronounced amongst disaggregated measures compared to various aggregate measures. The issue of importance is that Clark2006JAE assumes a common mean shift for all series, following BaiLumsdaineStock1998, but the validity of this assumption is not established.
We consider a subset of the series analyzed in Clark2006JAE, namely the inflation measures for durables, nondurables and services. These are taken from the NIPA accounts and cover the period 1984-2002 at the quarterly frequency; see Clark2006JAE for more details. Let $\{(y_{1t}, y_{2t}, y_{3t})\}_{t=1}^{T}$ denote the inflation series of durables, nondurables and services and consider an AR model allowing for a mean shift for each series $i= 1, 2, 3$:
where $\mu_{i}$ is an intercept parameter, $\delta_{i}$ is the magnitude of the mean shift with $k_{i}$ being a break date. The parameters, $\alpha_{i}^{(1)}, \dots, \alpha_{i}^{(p_{i})}$, are AR coefficients with $p_{i}$ denoting the lag length and $u_{it}$ is an error term. The persistence of each series is measured by the sum $\alpha_{i}^{(1)} + \cdots + \alpha_{i}^{(p_{i})}$ for $i=1,2, 3$. Clark2006JAE uses the Akaike information criterion (AIC) to select the AR lag length such that $(p_{1}, p_{2}, p_{3}) = (4, 5, 3)$ and also presents some evidence to support a mean shift in the AR models by applying break tests for each series and for groups.
We present our empirical results in Table 3. We first replicate a part of the results in Clark2006JAE. We find that when not allowing for a mean shift, the persistence measure is indeed quite high ranging from 0.855 to 0.921. Also, the persistence measure decreases to a large extent for non-durables and services but not so much for durables when a common break is imposed for the intercept at the break date 1993:Q1, which is not estimated but treated as known in Clark2006JAE. When we use the Seemingly Unrelated Regressions (SUR) method with an unknown common break date, following BaiLumsdaineStock1998, the point estimates are similar expect that the break date is estimated at 1992:Q1.
We now use our test to assess the validity of the common breaks specification. In Table 3, we report values of the test statistic for several null hypotheses as well as critical values corresponding to a 5% significance level, obtained through the computationally efficient algorithm described in Section 4 with 3,000 repetitions. First, we consider the null hypothesis of common breaks in the three inflation series, i.e., $H_{0}: k_{1} = k_{2} = k_{3}$. The value of the test statistic is 9.015 and the critical value is 5.242, so that the test rejects the null hypothesis of common breaks at the 5% significance level. Next, we test for common breaks in two inflation series within the full system of the three inflation series, separately. That is, we separately calculate the test statistic for $H_{0}: k_{1} = k_{2}$, $H_{0}: k_{1} = k_{3}$, and $H_{0}: k_{2} = k_{3}$. The values of the test statistic are 9.735 and 7.684 with corresponding critical values 3.473 and 3.259 for $H_{0}: k_{1} = k_{2}$ and $H_{0}: k_{1} = k_{3}$, respectively, and thus both hypotheses are rejected at the 5% significance level. On the other hand, the value of the statistic for $H_{0}: k_{2} = k_{3}$ is 0.749 with a critical value of 2.501. Thus, we cannot reject the null hypothesis of common breaks in the nondurables and service series.
We then estimate a system with the three inflation series imposing a common break only in the nondurables and service series (i.e., $k_{2} = k_{3}$), estimated at 1992:Q1, which is the same as when allowing for an unknown common break date in all series (the parameter estimates are also broadly similar). Things are quite different for the durables series. In this case, the estimate of the break date is 1995:Q1. What is interesting is that with this break date the decrease in persistence is very important with an estimate of 0.324 compared to 0.805 obtained assuming a common break date across the three series. Hence, allowing for different break dates for durables and the other series, we document a substantial decline in the persistence measure across all three series. Moreover, we report the 95% confidence intervals for the estimated break dates: [1994:Q2, 1995:Q4] for durables and [1991:Q3, 1992:Q3] for the others. These non-overlapping intervals are consistent with our results.
This paper provides a procedure to test for common breaks across or within equations. Our framework is very general and allows integrated regressors and trends as well as stationary regressors. The test considered is the quasi-likelihood ratio test assuming normal errors, though as usual the limit distribution of the test remains valid with non-normal errors. Of independent interest, we provide results about the rate of convergence when searching over all possible partitions subject only to the requirement that each regime contains at least as many observations as some positive fraction of the sample size, allowing break dates not separated by a positive fraction of the sample size across equations. We propose two approaches to obtain critical values. Simulations show that the test has good finite sample properties. We also provide an application to issues related to level shifts and persistence for various measures of inflation to illustrate its usefulness.
\setstretch{0.3} {8pt}
\baselineskip=15pt \setstretch{0.3}