Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
90,339 characters · 13 sections · 50 citation commands
Estimation of a Structural Break Point in Linear Regression Models
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} Change point, Parameter instability, Structural change
\spacingset{1.45}
Researchers in many economic fields extensively address parameter instability in models, which is a common empirical problem in macroeconomics and finance, such as the decrease in output growth volatility in the 1980s, known as “the Great Moderation," oil-price shocks, labor productivity changes, inflation uncertainty, and stock-return prediction models. It is often reasonable to assume that a change occurs over a long period of time or that some historical event affects the dynamics of a structural model. Hence, the interpretation of structural model dynamics or prediction models relies heavily on the estimation and testing of parameter instability. In econometrics literature, these changes in the underlying data generating process (DGP) of time-series are referenced as a structural break. The timing of the break, as a fraction of the sample size, is called the break point.
Researchers have used estimation methods in the structural break literature to analyze threshold effects and tipping points. Studies on policy change, income inequality dynamics, and social interaction models have used structural break estimation methods. Card2008 estimate a tipping point of segregation arising in neighborhoods with white preferences. \citet*{Gonzalez2012} explore the effect of child custody law reforms and Child Support Enforcement on U.S. divorce rates using the method developed by \citeauthor*{Bai1998a} (Bai1998a, Bai2003).
Extensive literature describes structural break estimation methods, starting with maximum likelihood estimators (MLE) on break points. \citet*{Hinkley1970}, \citet*{Bhattacharya1987} and \citet*{Yao1987} provide an asymptotic theory of the MLE of the break point in a sequence of independent and identically distributed random variables. The asymptotic theory of least-squares (LS) estimation of a one-time break in a linear regression model has been developed by \citeauthor*{Bai1994} (Bai1994, Bai1997), with extension to multiple breaks in \citet*{Bai1998a} and Bai1998b. The main problem with the LS estimation of the break point is that its finite sample behavior depends on the size of the parameter shift. In many cases, break magnitudes that are empirically relevant are “small" in a statistical sense. For instance, the quarterly U.S. real gross domestic product (GDP) growth rate from 1970Q1 to 2018Q2 has a mean of 0.68 and a standard deviation of 0.8 percent. A break that decreases the quarterly mean growth rate by 0.25 percentage points is less than half a standard deviation change but is equivalent to a 1 percentage-point decrease in annual growth, which is a significant event for the economy.
In asymptotic analysis, tests have local power against breaks with a magnitude of order $O(T^{-1/2})$\footnote{Asymptotic analysis under a DGP with a drifting sequence of parameters can be considered as a form of weak identification asymptotics. Literature on estimation and inference with a restricted parameter space under weak identification includes Andrews2012 (Andrews2012, Andrews2013, Andrews2014), and \citet*{Han2019}. Their results do not cover an abrupt structural change model, which is our model of interest.}. The magnitude represents a small break, shrinking with sample size $T$, so that structural breaks tests have asymptotic power strictly less than one (Elliott2007 Elliott2007). In the presence of small but detectable breaks, the LS estimator of the break point has a finite sample distribution that exhibits tri-modality with one mode at the true value and two modes at zero and one. Break points at zero or one do not provide any information about a structural break, nor are they likely to be true in practice. Therefore, inference in practical applications based on LS estimation of structural breaks would seem unreliable. Surprisingly, although the methodology is used widely, there are few alternatives for estimating the location of a structural break. Recent literature such as Casini2019 (Casini2019, Casini2020) suggest a Laplace-based procedure to provide an estimator of the break point, which is defined by an integration, rather than an optimization-based method.
This study provides an estimator of the structural break point, which is a generalization of LS estimation, and hence, easy to implement in practice. The new estimator resolves the finite sample issue of LS estimation; it has a finite sample distribution with a unique mode at the true break and flat tails. This is achieved by imposing weights on the LS objective function. Under small breaks, the LS estimator picks boundaries with high probability due to the functional form of the objective. I construct a weight function of the break point and impose it on the LS objective function to incorporate different estimation uncertainties across potential break points. I provide conditions on the weight function that ensure consistency of the break point estimator. I also suggest a representative weight function for empirical researchers to use.
The break point estimator is consistent with the same rate of convergence as the LS estimator (Bai1997 Bai1997) under regularity conditions on the weight functional form in a linear regression model with a structural break on a subset (or all) coefficients. The limit distribution of the break point estimator is derived when the break magnitude is small, under an in-fill asymptotic framework, following the approach of Jiang2017 (Jiang2017, Jiang2018). Monte Carlo simulations show that the break point estimator has smaller root mean squared error (RMSE) than the LS estimator in a finite sample for all break point values considered.
This study provides two empirical applications: estimation of structural breaks on post-war U.S. real GDP growth rate and the U.S. and UK stock return prediction models. For the quarterly U.S. real GDP growth rate under different sample periods, the new method estimates a break in the early 1970s, whereas the LS estimates vary from the 1970s to 1952 or 2000, which are near boundaries of the sample. The break date estimate in the early 1970s is matched with the “productivity growth slowdown" suggested in literature, such as \citet*{Perron1989} and \citet*{Hansen2001}. Thus, my estimation method yields reasonable break point estimates compared to LS estimates, which is sensitive to trimming the sample.
The remainder of this paper proceeds as follows: Section (ref) constructs the break point estimator for a mean shift in a linear process. Section (ref) provides a generalized linear regression model with multiple regressors, and proves the consistency of the break point estimator. Section (ref) presents the in-fill asymptotic theory for stationary and local-to-unit root processes. Monte Carlo simulation results are in Section (ref) and Section (ref) provides three empirical applications of the new structural break estimation method. We provide concluding remarks in Section (ref). Additional theoretical results and proofs are in the Appendix.
In this section, I consider the simplest regression model with a constant term to provide an intuitive explanation of the construction of the break point estimator. I provide theoretical results in Section (ref) under a general linear regression model with multiple regressors. Suppose a single break occurs at time $k_0 = [\rho_0 T]$, where $\rho_0 \in (0,1)$, $[\cdot]$ is the greatest smaller integer function, and $\textbf{1}\{t>k_0\}$ is an indicator function that equals one if $t > k_0$ and zero otherwise.
The disturbances $\{\varepsilon_{t}\}$ are independent and identically distributed (i.i.d.) with mean zero and $E\varepsilon_t^2 = \sigma^2$. The pre-break mean $y_t$ is $\mu$ and the post-break mean is $\mu + \delta$. Assume we know a one-time break occurs, but the break point $\rho_0$ and parameters $(\mu, \delta, \sigma^2)$ are unknown.
The conventional estimation method of the break location in the literature is least-squares. One obtains the LS estimator by finding a value $k$ that minimizes the objective function $S_{T}(k)^2$, which is the sum of squared residuals (SSR) under the assumption that $k$ is the break date, $S_{T}(k)^2 =\sum_{t=1}^{k}(y_t - \bar{y}_{k})^2 + \sum_{t=k+1}^{T}(y_t - \bar{y}_{k}^{*})^2$, where $\bar{y}_{k} = k^{-1}\sum_{j=1}^{k}y_j$ and $\bar{y}_{k}^{*} = (T-k)^{-1}\sum_{j=k+1}^{T}y_j$ are pre- and post-break LS estimates under break date $k$, respectively. Following Bai1994's Bai1994 expression, I use the identity $\sum_{t=1}^{T}(y_t - \bar{y})^2 = S_{T}(k)^2 + TV_{T}(k)^{2}$ (Amemiya1985 Amemiya1985), where $V_{T}(k)^2 = k/T(1-k/T)\left(\bar{y}_{k}^{*} - \bar{y}_{k}\right)^2$, to substitute for the SSR. Then the LS estimator of the break date is equivalent to
Denote $\rho = k/T$ and $\rho_0 = k_0/T$. An issue with the LS estimator $\hat{\rho}_{LS}$ is that under a small magnitude $\vert \delta \vert$, $\hat{\rho}_{LS}$ has a finite distribution that is tri-modal with two modes at the ends of the unit interval and one mode at the true break point $\rho_0$.
A break magnitude that is statistically small is not necessarily small in an economic sense. For example, quarterly U.S. real GDP growth rate from 1970Q1 to 2018Q2 has a mean of 0.68 and a standard deviation of around 0.8 percent. A break that decreases the mean quarterly growth rate by 0.3 percentage points (a 1.2 percentage point decrease in yearly growth) is a significant event for the economy. Suppose model ((ref)) has parameter values similar to the U.S. real GDP growth rate; assume $\rho_0 = 0.3$, the pre-break mean is $\mu = 0.88$ percent and the shift in the mean of growth rate is $\delta = -0.29$. The expectation of $y_t$ is $\mu+(1-\rho_0)\delta = 0.68$, which matches the quarterly U.S. real GDP growth rate. Suppose we have $T=100$ observations and Gaussian disturbances $\varepsilon_t \overset{i.i.d.}{\sim} N(0,0.8^2)$. The left plot of Figure (ref) shows the finite sample distribution of the LS estimator of $\rho$ from a Monte Carlo simulation with 2,000 replications. The LS estimator fails to accurately detect the break that occurs in the constant term of a univariate linear regression model. Thus, we expect that in practice, structural breaks that are economically important are not large enough for the LS estimator to detect in many cases.
This study focuses on such empirically relevant breaks that are not “large" enough. I follow the approach of \citet*{Elliott2007} to provide an asymptotic approximation to finite sample properties under this small break magnitude. The break magnitude has the same order as sampling uncertainty, $\delta = T^{-1/2}d$, where $d$ is fixed. These asymptotics reflect an important feature of finite sample properties under moderate breaks, because the $p$ values of tests for breaks are typically significant, but not zero. See Elliott2007 (Elliott2007, Elliott2014) for details on the justification of this break magnitude.
Importantly, in literature, it is standard to trim the boundaries of the optimization space so that $\hat{k}_{LS}$ in ((ref)) is the argmax function across $k = [\alpha T], \ldots, [(1-\alpha) T]$ for some $0 < \alpha < 1/2$. Trimming the optimization space may help reduce the build-up mass at the boundaries of the finite sample distribution; however, this has its own drawbacks. Figure (ref) shows the finite sample distribution of the LS estimator of $\rho$ with various fractions of trimming, $\alpha \in \{0, 0.1, 0.15, 0.2\}$, from a Monte Carlo simulation with 2,000 replications. Under small break magnitudes $\delta = T^{-1/2}d$, $d \in \{2, 4\}$, the modes at the boundaries remain even after trimming. It is unclear if there is a trade-off between the break size and how large a trimming is needed. With larger trimming ($\alpha = 0.2$), the mass at the boundaries accumulate even more. In addition, there is no reason to believe a break occurs in a restricted period. Thus, we need an alternative method to resolve this issue.
The finite sample distribution of the LS estimator has build-up mass at boundaries under small break magnitudes because of how the objective function is constructed. For each potential break date $k$, the objective function is constructed by partitioning the sample into two sub-samples, before and after $k$. Each sub-sample is used to estimate two different means, $\bar{y}_k$ and $\bar{y}_k^*$. If $k$ is near 1, the pre-break sub-sample size $k$ is small; similarly, if $k$ is near $T-1$, the post-break sub-sample size $T-k$ is small\footnote{Estimation theory does not require $\rho$ to be bounded away from zero and one, provided that a change point is assumed to exist. However, identification in a finite sample typically needs more than one observation pre- and post-break. In Section (ref), I assume the break point exists; see Assumption (ref)$(i)$.}. Hence, when the potential break date $k$ of $\left\vert V_{T}(k) \right\vert$ is near the boundaries, the estimates of pre- or post-break mean are imprecise because of the small sub-sample size. Estimation uncertainty at boundaries distorts picking up the true break location if the break magnitude is small relative to sampling variability.
Because the issue arises from the large variance of the objective function at the boundaries, one can think of shrinking the variance accordingly. Suppose there are non-negative “weights" $\omega_k$ imposed on the LS objective function so that $k$ with a large estimation error has smaller weights than $k$ with a small estimation error. When $k=1$ and $T-1$, weights near zero are imposed, which implies the variance of the weighted objective function $\omega_k \vert V_T(k)\vert$ would shrink toward zero. If the sample period is normalized into a unit interval, the weights are represented by a continuous function $\omega(\rho)$ on $\rho \in [0, 1]$, which is zero at $\rho \in \{0,1\}$, and has positive values otherwise. A continuous function with such properties would look like an inverse U-shaped (or concave downward) function on the unit interval.
I propose a new break point estimator to maximize the value of the objective function $\left\vert Q_{T}(k) \right\vert$, equal to weights $\omega_k$ multiplied by the LS objective $\left\vert V_{T}(k) \right\vert$.
The weight function shrinks the variance of the LS objective function when $k$ is near the boundaries, and thus, the maximizing value $\hat{k}$ is less likely to pick either end. The right plot of Figure (ref) shows the finite sample distribution of the break point estimator ((ref)) under the DGP calibrated from the U.S. real GDP growth rate. As expected, the break point estimator has flat tails at boundaries with a mode at the true break point $\rho_0 = 0.3$, whereas the LS estimator has modes at $0.01$ and $0.99$. Section (ref) provides additional Monte Carlo simulations comparing the two estimators.
The break point estimator ((ref)) is easy to implement, as we simply modify the objective function by multiplying the weight function. It is a generalization of LS estimation because the LS estimator is a special case, when $\omega(\rho) = 1$. Moreover, by employing the weight function, we no longer need to trim the search grid, because the boundaries have zero weight. Section (ref) provides a set of conditions on the weight function that ensures consistency of the break point estimator under a general linear regression model.
I also suggest a representative weight function $\omega(\rho) = (\rho(1-\rho))^{1/2}$ under model ((ref)) (see Section (ref) for its analogue under a model with multiple regressors). The weight function has two interpretations. First, it is related to the weighting function from \citet*{Anderson1954}, which tests whether the sample is drawn from a particular distribution. A non-negative weight function is chosen to accentuate the boundaries of the sample space, where the test is desired to have sensitivity. If the cumulative distribution function (cdf) under the null hypothesis is $F(\cdot)$, the weight function is $\left[F(x)(1-F(x))\right]^{-1}$, which increases as $x$ approaches the boundaries of the sample space. In contrast, I want to down-weight the boundaries of the parameter space $\rho \in [0, 1]$. If $\rho$ is a random variable with cdf $F(\rho)$, the weight would be the reciprocal of the weight function in \citet*{Anderson1954}, $F(\rho)(1-F(\rho))$. We further assume that $\rho$ is uniformly distributed on the unit interval, $F(\rho) = \rho$. We obtain the weight function $\rho(1-\rho)$, which down-weights the variance of the LS objective function $V_T(k)^2$ near the boundaries.
Second, from a Bayesian perspective, the weight function $\omega(\rho)$ can be interpreted as a prior on parameters $\delta$ and $\rho$. We assume the Gaussian disturbances in ((ref)), $\omega(\rho) = (\rho(1-\rho))^{1/2}$ are equivalent to the square root of the Fisher information up to a constant. The Fisher information is interpreted as a way to measure the amount of information the data gives us about the unknown parameter $\delta$, given $\rho$. A prior distribution based on the Fisher information reflects our belief that a structural break is less likely to occur near the boundaries. See Appendix A for details.
This section provides consistency of the break point estimator under a general linear regression model with multiple regressors. The model incorporates a partial break in coefficients and assumes that a one-time break occurs at an unknown date $k_0 = [\rho_0 T]$ with $\rho_0 \in (0,1)$. I follow the notations of \citet*{Bai1997} by denoting the vector of variables associated with a stable coefficient as $w_t$ and the variables associated with coefficients under a break as $z_t$. Let $x_{t} = (w_t', z_t')'$ be a $(p \times 1)$ vector and $z_{t}$ is a $(q\times 1)$ vector with $q \leq p$,
where $\varepsilon_t$ is a mean zero error term. In general, $z_{t}$ can be expressed as a linear function of $x_{t}$ so that $z_{t} = R'x_{t}$, where $R$ is a $(p\times q)$ matrix with full column rank. Let $Y = (y_1,\ldots,y_T)'$ and define $X_{k} := (0,\ldots,0,x_{k+1},\ldots,x_{T})'$ and $X_{0} :=(0,\ldots,0,x_{k_0 +1},$ $\ldots,x_{T})'$. Define $Z_k$ and $Z_0$ analogously so that $Z_k = X_k R$ and $Z_0 = X_0 R$. Let $M := I-X(X'X)^{-1}X'$ and use the maximal invariant to eliminate the nuisance parameter $\beta$.
The subscript on $\delta_T$ shows that the break magnitude may depend on the sample size. We assume the break magnitude is outside the local $T^{-1/2}$ neighborhood of zero. This is because the break point is not consistently estimable if the break magnitude is in the local $T^{-1/2}$ neighborhood of zero. This corresponds to the case of small break magnitudes discussed previously, $\delta_T = O(T^{-1/2})$, in which structural break tests have asymptotic power strictly less than one. In this section, I proceed by assuming that $\delta_T$ is fixed or it converges to zero at a rate slower than $T^{-1/2}$ so that the power of the structural break tests converge to one (Assumption (ref)).
Let $\bar{S} =Y'MY$, and denote $S_{T}(k)^2$ as the SSR regressing $MY$ on $MZ_k$. The LS estimator of break date $\hat{k}_{LS}$ is the value that minimizes $S_{T}(k)^2$, and thus, maximizes $V_{T}(k)^2$ from the identity $\bar{S} = S_{T}(k)^2 + V_{T}(k)^2$ (Amemiya1985 Amemiya1985),
where $\hat{\delta}_{k}$ is the LS estimate of $\delta_T$ by regressing $MY$ on $MZ_k$. Note that $V_{T}(k)^2$ is non-negative from the inner product of the vector $(Z_{k}'MZ_{k})^{1/2}\hat{\delta}_{k}$. The LS objective function is modified by multiplying a $(q \times q)$ positive definite weight matrix $\Omega_k$, which is a generalization of $\omega_k$ in Section (ref) for linear regression models with multiple regressors. Decompose the weight matrix so that $\Omega_k = \Omega_k^{1/2\prime}\Omega_k^{1/2}$ and multiply $\Omega_k^{1/2}$ to the vector $(Z_{k}'MZ_{k})^{1/2}\hat{\delta}_{k}$. Take the inner product and obtain the objective function $Q_T(k)^2 := \hat{\delta}_{k}'(Z_{k}'MZ_{k})^{1/2}\Omega_k (Z_{k}'MZ_{k})^{1/2}\hat{\delta}_{k}$. Then the estimator of the break point is
An example of the weight matrix is $\Omega_k = T^{-1}Z_k'MZ_k$, which is equal to the square of the representative weight function $\omega_k^2 = k/T(1-k/T)$ in model ((ref)) if $R=I$ and $X$ is a $(T\times 1)$ vector of ones. Similarly, the matrix $T^{-1}Z_k'MZ_k$ “decreases" as $k$ approaches either end of the sample from the following rearrangement of terms:
I prove the consistency of the break point estimator $\hat{\rho}$ in ((ref)) under regularity conditions on model ((ref)) and weight matrix $\Omega_k$. The notation $\left\lVert\cdot\right\rVert$ denotes the Euclidean norm $\left\lVertx\right\rVert = \left(\sum_{i=1}^{p}x_i^2\right)^{1/2}$ for $x \in \mathbb{R}^p$. For a matrix $A$, $\left\lVertA\right\rVert$ represents the vector induced norm $\left\lVertA\right\rVert = \sup_x \left\lVertAx\right\rVert/\left\lVertx\right\rVert$ for $x \in \mathbb{R}^p$ and $A \in \mathbb{R}^{p \times p}$.
The conditions of assumption (ref) are similar to assumptions A1 to A6 in \citet*{Bai1997}, with additional restrictions $(iv)$ and $(vi)$. Assumption (ref)$(vi)$ allows for general serial correlation in disturbances and requires $x_t$ to be strictly exogeneous. This is because $\Omega_k$ depends on the moments of regressors and we want to impose zero weights on the boundaries of the unit interval. For instance, if the second moments of $z_t$ changes at $\rho_0$, the boundaries of the unit interval may have positive weights that depend on the distribution of $z_t$. These cases are avoided under strict exogeneity because $\Omega_k$ converges in probability to a nonrandom matrix that varies across $\rho$ only. Note that if $\Omega_k$ is a non-stochastic matrix that satisfies the norm inequality in Assumption (ref), consistency holds under weakly exogeneous regressors (see Assumption (ref)).
Assumption (ref) guarantees that the matrix
is positive definite, and hence, $\left\lVertA_{T}(k)\right\rVert \geq \lambda_{\min}(A_{T}(k)) > 0$, where $\lambda_{\min}$ denotes the minimum eigenvalue of $A_{T}(k)$. The condition can be interpreted as follows: for simplicity, consider the univariate model ((ref)). Assumption (ref) is equivalent to $\vert \omega'(\rho)/ \omega(\rho) \vert < (2\rho(1-\rho))^{-1}$ for all $\rho$, where $\omega'(\rho) = \partial \omega(x)/\partial x \vert_{x=\rho}$. The slope magnitude of the logarithm of $\omega(\rho)$ has an upper bound that increases as $\rho$ approaches zero or one. A sufficient condition is the function $\omega(\rho) = (\rho(1-\rho))^{\gamma}$, with $-1/2 \leq \gamma \leq 1/2$.\footnote{The function $\omega(\rho) = (\rho(1-\rho))^{\gamma}$, with $-1/2 \leq \gamma \leq 1/2$, allows a convex function that has a large weight on the boundaries. This is because I assume “large" break magnitudes (Assumption (ref)) for consistency of the estimator. That is, if the break magnitude is large enough, we no longer have build-up mass at the boundaries and imposing large weights does not matter for consistency.} Note that this is sufficient under Assumption (ref)$(i)$. If $\alpha$ is arbitrarily close to zero, it may restrict the functional of $\omega(\cdot)$. The weight matrix $\Omega_k$ may be close to a singular matrix in a finite sample if $\alpha$ is extremely close to zero under model ((ref)). Under Assumption (ref)$(iv)$, the weight matrix converges in probability to a function of $\rho$ and $\Sigma_x$ as $T$ increases. Because $\bar{\Omega}(\rho)$ is a differentiable function of $\rho$ element-wise, $\left\lVert\Omega_k - \Omega_{k_0}\right\rVert \leq b\vert k-k_0 \vert/T$ for some finite $b > 0$ and all $k$.
The consistency of the break point estimator is proved by showing that if $\delta_T \neq 0$, then with high probability, $Q_T(k)^2$ can only be maximized near the true break $k_0$. The objective function $Q_T(k)^2$ is defined in ((ref)) and $\hat{\delta}_k$ is the LS estimator of the break magnitude, assuming that $k$ is the break date: $\hat{\delta}_k = (Z_k'MZ_k)^{-1}(Z_k'MZ_0)\delta_T + (Z_k'MZ_k)^{-1}Z_k'M\varepsilon$. If $k = k_0$, then $\hat{\delta}_{k_0} = \delta_T + (Z_0'MZ_0)^{-1}Z_0'M\varepsilon$.
See Appendix B for proof of Theorem (ref). For weakly exogenous regressors, the break point estimator is consistent with the same rate of convergence in Theorem (ref), under the following conditions that substitute Assumptions (ref) and (ref).
The proof of Theorem (ref) is similar to the proof of Theorem (ref); hence we have omitted it. Under Assumptions (ref)$(v)$, (ref)$(i)$, and (ref)$(ii)$, the strong law of large numbers holds for $x_t \varepsilon_t$, because the conditions in \citet*{Hansen1991} are satisfied. The weight matrix $\Omega_k$ in Assumption (ref)$(iii)$ depends on $k/T$ but not on the data $\{x_t,\varepsilon_t\}$. Thus, by setting $\rho = k/T$, $\Omega_k$ is a function of $\rho$, which is assumed to be differentiable with respect to $\rho$. Then, for some finite $c>0$, the bound $\left\lVert\Omega_{k_1} - \Omega_{k_2}\right\rVert \leq c \vert k_1-k_2\vert/T$ holds for any $k_1$ and $k_2$. Using these properties, proving the consistency of the estimator under Assumption (ref) follows the same process as in the proof under Assumptions (ref) and (ref).
Given the consistency of the break point estimator from Theorem (ref) or (ref), the estimator of the break magnitude corresponding to $\hat{k}$ is consistent and asymptotically normally distributed. Let $\hat{\delta}(\hat{\rho}) = \hat{\delta}_{\hat{k}}$, then the following results hold. The proof is provided in the Appendix.
\citet*{Bai1997} provides the limit distribution of the LS estimator assuming large breaks ($\delta = O(T^{-1/2+\epsilon})$ with $0<\epsilon< 1/2$). The asymptotic distribution is symmetric at the true break point, if the second moment of variables associated with coefficients under break ($z_t$ in Section (ref)) do not change before and after break. I am interested in small breaks in which the asymptotic distribution depends on the parameters in a complicated manner (Elliott2007 Elliott2007).
Jiang2017 (Jiang2017, Jiang2018) and \citet*{Casini2019} employed a continuous record asymptotic framework to derive the limit distribution of the break point estimator. By assuming that a continuous record is available, a continuous time approximation to the discrete time model is constructed and an in-fill asymptotic distribution is developed. In contrast to the long-span asymptotic, where the time span of the data increases, the in-fill asymptotic assumes a fixed time span with shrinking sampling intervals. The in-fill asymptotic distribution is asymmetric, tri-modal, and dependent on the initial condition. However, the long-span asymptotic distribution of the LS estimator under local-to-unity processes do not depend on the initial condition. See Chong2001, Pang2014 and Pang2018 on the long-span asymptotic distribution of the LS estimator under different settings of the AR root before and after the break. I follow the approach of Jiang2017 (Jiang2017, Jiang2018) to derive the limit distribution of the break point estimator under a stationary and local-to-unity autoregressive process.
Consider the linear regression model ((ref)) with continuous time process $\{W_s,Z_s,\mathcal{E}_s\}_{s\geq 0}$ defined on a filtered probability space $(\Omega,\mathcal{F},(\mathcal{F}_s)_{s\geq 0},P)$, where $s$ can be interpreted as a continuous time index. Assume we observe at discrete points of time so that $\{Y_{th}, W_{th}, Z_{th}: t=0,1,\ldots,T = N/h\}$, where $N$ is the time span. We normalize the time span $N=1$ for simplicity. We denote the increment of processes as $\Delta_h Y_t := Y_{th}-Y_{(t-1)h}$. Let $X_{th} = (W_{th}', Z_{th}')'$ so that $Z_{th} = R'X_{th}$. The model ((ref)) can be expressed as
Divide both sides by $\sqrt{h}$ so that the error term variance is $O(1)$. The parameters $\beta_h$ and $\delta_h$ may depend on the sampling interval, denoted by subscript $h$. Let $\varepsilon_{t} := \Delta_h \mathcal{E}_t/\sqrt{h}$, $y_t := \Delta_h Y_t/\sqrt{h}$, $x_t := \Delta_h X_t/\sqrt{h}$, $z_t := \Delta_h Z_t/\sqrt{h} = R'x_t$,
Notations from Section (ref) are used for model ((ref)): $MY = MZ_0 \delta_h + M\varepsilon$, where $\varepsilon = (\varepsilon_1,\ldots,\varepsilon_T)'$ and $M = I-X(X'X)^{-1}X'$. The objective function of the estimator $\hat{k}$ in ((ref)) is restated as follows:
The in-fill asymptotic distribution is derived for the two different magnitudes of $\delta_h$ in Assumption (ref). Theorem (ref) provides the limit distribution under (ref)$(i)$, which represents small breaks. For proof, see Appendix B.
An equivalent representation of the in-fill asymptotic distribution is (let $\rho = \rho_0+u$)
where $\widetilde{W}(\cdot)$ is defined in Theorem (ref).
Next, consider the case of Assumption (ref)$(ii)$. The proof of Theorem (ref) is in Appendix B.
If the weight matrix is $\Omega_k = I_q$, the estimator is equivalent to the LS estimator and the limiting distribution reduces to the distribution in Proposition 3 of \citet*{Bai1997}. The term $A_u$ shows how the weight matrix down-weights break points near the boundaries. Suppose $\Omega_k = T^{-1}Z_k'MZ_k$ and $\rho_0 > 0.5$ so that $\bar{\Omega}(\rho) = \rho(1-\rho)\Sigma_z$ and $\nabla \bar{\Omega}_0 < 0$. If $u > 0$ increases in a positive direction toward the boundary (i.e., $\rho > \rho_0 > 0.5$), then $A_u$ increases and the term multiplied to the Wiener process with drift, $(d_0'\Sigma_z A_u d_0)^{-1}$ decreases. In contrast, if $u<0$ decreases such that $\rho$ shifts toward the median, then $A_u$ decreases and $(d_0'\Sigma_z A_u d_0)^{-1}$ increases. The result is opposite if $\rho_0 < 0.5$. That is, there is larger weight on $\rho$ near the median $0.5$ and less weight near the boundaries.
In this section, I derive the in-fill asymptotic distribution of an autoregressive (AR) model with a structural break in its lag coefficient, using a deterministic weight function $\omega(\cdot)$. As mentioned in Section (ref), Assumption (ref) excludes lagged dependent variables, due to the dependence of the weight function on regressors. This condition is relaxed to allow weakly exogeneous regressors by assuming non-stochastic weights. Consider a discrete model closely related to the Ornstein-Uhlenbeck process with a break in the drift function:
where $t \in [0,1]$ and $B(\cdot)$ denote a standard Brownian motion. The discrete time model has the form
where $\beta_1 = \exp\{-\mu/T\}$ and $\beta_2 = \exp\{-(\mu+\delta)/T\}$ are the AR roots before and after the break. We denote $y_t = x_t/\sqrt{h}$ so that the order of errors is $O_p(1)$ as in model ((ref)). Then, I have for $t = 1,\ldots,T$,
The initial condition of $y_t$ in ((ref)) diverges at rate $T^{1/2}$; thus, the in-fill asymptotic distribution will depend explicitly on the initial value $x_0$. The break size is $\beta_2-\beta_1 = O(T^{-1})$, whereas the literature on long-span asymptotics assumes $O(T^{-\gamma})$ with $0<\gamma <1$. The model ((ref)) is a local-to-unit root process: $\beta_1 = \exp\{-\mu /T\} \rightarrow 1$ and $\beta_2 = \exp \{-(\mu+\delta)/T\} \rightarrow 1$, as $T \rightarrow \infty$ for any finite $(\mu,\delta)$. In contrast, the long-span asymptotic theory incorporates stationary AR(1) processes, where $\vert \beta_1 \vert < 1$ and $\vert \beta_2 \vert < 1$. \citet*{Chong2001} derives the long-span distribution under $\vert\beta_2-\beta_1\vert = O(T^{-1/2+\gamma})$ with $0 < \gamma < 1/2$. Jiang2017 provides simulation results that the in-fill asymptotic theory works well even when $\beta_1$ and/or $\beta_2$ are distant from unity in the finite sample.
The break point estimator and the LS estimator in model ((ref)) takes the form
where $\hat{\beta}_1(k) = \sum_{t=1}^k y_t y_{t-1}/ \sum_{t=1}^k y_{t-1}^2$ and $\hat{\beta}_2(k) = \sum_{t=k+1}^T y_t y_{t-1}/ \sum_{t=k+1}^T y_{t-1}^2$ are LS estimates of $\beta_1$ and $\beta_2$ under break at $k$, respectively.
The results of Theorem (ref) derive from applying the continuous mapping theorem to the limit distribution $S(k)^2$ in Theorem 4.1 from Jiang2017. See Appendix B for the proof. The difference between the asymptotic distributions of the two estimators is the weight function multiplied by the stochastic process in the argmax function. Both estimators are asymmetrically distributed around the true point and biased when $\rho_0 \neq 1/2$.
This section compares finite sample distributions of the new estimator and the LS estimator using Monte Carlo simulation. It considers two different models; a break in the mean of a univariate regression model and a break in the lag coefficient of the AR(1) process. I compare the root mean squared error (RMSE), bias, and standard errors of the two estimators in the finite sample and in-fill asymptotics.
The first model is when a structural break occurs in model ((ref)), where $x_t = z_t = 1$ for all $t$. The break magnitude $\delta_T = T^{-1/2}d_0$ is in the local $T^{-1/2}$ neighborhood of zero to represent small break magnitudes.
where $\varepsilon_t \overset{i.i.d.}{\sim} N(0,\sigma^2)$ and $\sigma = 1$. Parameter values are $\rho_0 \in \{0.15, 0.3, 0.5, 0.7, 0.85\}$, $\mu = 4$, $d_0 \in \{1, 2, 4 \}$, and $T=100$ with 5,000 replications. The weight function is $\omega_k = (k/T(1-k/T))^{1/2}$, which is the representative weight function motivated in Section (ref)\footnote{If the weight function is $\omega_k = (k/T(1-k/T))^{\gamma}$, Assumption (ref) is satisfied if $-1/2\leq \gamma \leq 1/2$ for an arbitrary small $\alpha$ in Assumption (ref)$(i)$. For $\gamma \in \{1/8, 1/4, 3/8\}$, the results (omitted due to space constraints) do not change qualitatively; the probability at the boundaries decrease compared to the finite sample distribution of LS. Because $\omega(\rho) = (\rho(1-\rho))^\gamma \rightarrow 1$ as $\gamma \rightarrow 0$, the difference between the two estimators finite sample behavior shrinks when $\gamma$ is close to zero.}. The break point estimator $\hat{\rho}_{NEW}$ is defined in ((ref)) and the LS estimator $\hat{\rho}_{LS}$ in ((ref)). Although it is unnecessary to trim under the simple model ((ref)), I trim the optimization space by fraction $\alpha=0.1$ on both ends, following the common practice in the literature.
Table (ref) provides the RMSE, the bias, and the standard error for the finite sample distribution. For all $\rho_0$ and $d_0$ values considered, the RMSE of the estimator $\hat{\rho}_{NEW}$ is smaller than that of $\hat{\rho}_{LS}$ in the finite sample. A comparison of asymptotic RMSE shows the same results qualitatively (see Appendix C). A trade-off emerges of slightly larger bias but a large decrease in standard error for $\hat{\rho}_{NEW}$ compared to $\hat{\rho}_{LS}$, which leads to a decrease in RMSE. When $\rho_0 = 0.5$, both bias and standard error of the new estimator is smaller than the LS estimator.
Figures (ref) and (ref) show the finite sample distribution of the two estimators under $\rho_0 = 0.30$, and 0.85. Under $\rho_0 = 0.30$, the finite sample distribution of $\hat{\rho}_{LS}$ is tri-modal, whereas the $\hat{\rho}_{NEW}$ has an unique mode at $\rho_0$ for all $d_0$ values considered. When the true break point is near the boundaries of the optimization space and the break magnitude is small ($\rho_0 = 0.85$ and $d_0 = 1$), the LS estimator performs particularly worse. The tri-modal LS estimator distribution becomes bi-modal with modes at $\alpha$ and $1-\alpha$. Under these parameter values, the new estimator has its drawbacks; the finite sample distribution is relatively flat. This is partly due to trimming the optimization space, which is not necessary for our estimation method. Without trimming ($\alpha = 0$), the finite sample distribution of our break point estimator has a unique mode at $\rho_0$ for all parameter values considered.
In short, the break point estimator is preferable than the LS estimator in terms of RMSE, under small break magnitudes. The weight function biases the estimator toward the median in trade-off to a significant decrease in standard error. When a break occurs near the boundaries, we need to be careful, because the new estimator can also be problematic. I suggest minimizing the trimming fraction $\alpha$ and using the new break point estimator.
For the AR(1) process, I replicate two experiments from Jiang2017. The first experiment is a break in the lag coefficient so that the stationary process changes to another stationary AR(1) process. The second case is a change from a local-to-unit root to a stationary AR(1) process. Each experiment is generated from model ((ref)) with $h=1/200$ ($T=200$), $\sigma = 1$, $\varepsilon_t \overset{i.i.d.}{\sim} N(0,1)$, $\rho_0 \in \{0.3, 0.5, 0.7\}$, and different combinations of $\mu$ and $\delta$ with $\beta_1 =\exp(-\mu/T)$ and $\beta_2 = \exp(-(\mu+\delta)/T)$.
The stochastic integrals of in-fill asymptotic distributions are approximated over a grid size $h = 0.005$. The optimization space is trimmed by fraction $\alpha=0.1$. The break point estimator $\hat{\rho}_{NEW}$ of the AR(1) model is defined in ((ref)) and its asymptotic distribution is stated in Theorem (ref). The in-fill asymptotic distribution of the LS estimator $\hat{\rho}_{LS}$ is stated in Theorem 4.1 from Jiang2017.
Table (ref) provides the RMSE, bias, and the standard error of $\hat{\rho}_{NEW}$ and $\hat{\rho}_{LS}$ for the finite sample, respectively (see Table (ref) in Appendix C for the asymptotic distribution). Similar to results in section (ref), the RMSE of $\hat{\rho}_{NEW}$ is smaller than that of $\hat{\rho}_{LS}$ for all parameter values $(\beta_1,\beta_2,\rho_0)$ considered. This also holds in the limit. A decrease in RMSE of $\hat{\rho}_{NEW}$ emerges from the trade-off of a relatively large decrease in variance compared to the increase in the squared bias.
Figures (ref) and (ref) are finite sample distributions of the break point in the two experiments. For the stationary to another stationary process change, the LS estimator $\hat{\rho}_{LS}$ mode at the true break point is almost negligible, unless it is the median $\rho_0 = 0.5$. In contrast, the estimator $\hat{\rho}_{NEW}$ has a unique mode at the true break point for all $\rho_0 \in \{0.3, 0.5, 0.7\}$. For the local-to-unit root to a stationary AR(1) change, both estimators have a higher probability at the true break point. However, the LS estimator continues to exhibit tri-modality with modes at the ends, whereas the new estimator has a unique mode at $\rho_0$.
In this section, I estimate the structural breaks in two empirical applications. I analyze the performance of the break point estimator by comparing it with the LS estimator and historical events documented in the literature. Furthermore, I show that the estimator is robust to trimming the sample period, whereas the LS estimator varies significantly depending on the trimmed sample. The first application is about the structural break in postwar U.S. real GDP growth rate. The second application is estimating the break date on the U.S. and UK stock returns using the return prediction model from \citet*{Paye2006}.
In the macroeconomics literature, shocks that affect the mean growth rate are often modeled as a one-time structural break because of their rare occurrence. However, existing estimation methods fail to capture the graphical evidence of postwar European and U.S. growth, slowing sometime in the 1970s. This is known as the “productivity growth slowdown," which is widely hypothesized in macroeconomics literature. For instance, \citet*{Bai1998b} show that for the U.S., most test statistics reject the no-break hypothesis; however, the estimated confidence interval does not contain the slowdown in the 1970s.
I estimate the structural break of an autoregressive model using the postwar quarterly U.S. real GDP growth rate. I use real GDP in chained dollars (base year 2012) data from the Bureau of Economic Analysis (BEA) website for the sample period 1947Q1-2018Q2, seasonally adjusted at annual rates. Annualized quarterly growth rates are calculated as 400 times the first differences of the natural logarithms of the levels data. I assume that log output has a stochastic trend with a drift and a finite-order representation. Following the \citet*{Eo2015} approach, I use Kurozumi2011's Kurozumi2011 modified Bayesian information criterion (BIC) for lag selection to account for structural breaks. The highest lag order selected is 1 for output growth, given an upper bound of four lags and four breaks. The AR(1) model ((ref)) is estimated under three cases. The first case is a break only in the drift term ($\gamma \neq 0$, $\delta = 0$), the second case is a break only in the coefficient of lag, the “propagation term" ($\gamma = 0$, $\delta \neq 0$), and finally, a break in constant and coefficient.
I assume the error term $\{ \varepsilon_{t} \}$ is serially uncorrelated with mean zero disturbances. If a structural break occurs in constant and lag coefficients, the long-run growth rate of log output will change from $E[\Delta y_{t}] = \beta/(1-\phi)$ to $(\beta +\gamma)/(1-\phi-\delta)$ and the volatility of the growth rate will change from $Var[\Delta y_{t}] = \sigma^{2}/(1-\phi^{2})$ to $\sigma^{2}/(1-(\phi+\delta)^{2})$ at time $k_0$.
Using the notations of model ((ref)), we have $x_t = (1,\, \Delta y_{t-1})'$, the dependent variable is $\Delta y_t$, and for each model $z_t = R'x_t$ is as follows:
The break point estimator $\hat{k}_{NEW}$ in ((ref)) is obtained using the weight $\Omega_k = \omega_k^2 I_q$, where $\omega_k = (k/T(1-k/T))^{1/2}$ and $q = $ dim($z_t$)\footnote{For consistency of the break point estimator in a AR(1) model, we use $\Omega_k = \omega_k^2 I_q$, where $\omega_k$ is a function of $k/T$ only.}. I use the full sample 1947Q3-2018Q2 ($T=284$) to estimate the structural break date, then a shorter sub-sample to see if the break date estimates change. The optimization space of $k$ is the sample trimmed by fraction $\alpha = 0.1$ at both ends; the grid starts at 1954Q2 and ends at 2011Q1. The second and third columns of Table (ref) show the break date estimates for the full sample. For M1 and M3, the two estimates are extremely different from each other, $\hat{k}_{NEW}$ is 1973Q1, whereas $\hat{k}_{LS}$ is 2000Q2. The break point estimates on the unit interval are approximately 0.36 and 0.75, respectively. Without any knowledge of historical events, one might think the finite sample properties of $\hat{\rho}_{LS}$ do not appear here, because it is not close to the boundaries of $0.1$ or $0.9$.
However, the LS estimate switches to the boundary if we consider a sub-sample that is one decade shorter. Consider a sub-sample that ends at 2007Q1, with starting date 1947Q3, so that the search grid includes $\hat{k}_{LS}$ from all models. The LS estimate of M1 changes drastically to 1953Q1, which is the end of the search grid, $\hat{\rho}_{LS} = 0.1$. In contrast, my estimator under M1 provides the same break date estimate $\hat{k}_{NEW} = $1973Q1. For M3, both estimates change, so $\hat{k}_{NEW} = $1966Q1 and $\hat{k}_{LS} = $1958Q1. Compared to the full sample estimate, the change in the new estimate is 7 years, whereas it is over 40 years for the LS estimate.
The break date estimate $\hat{k}_{NEW}=$1973Q1 under M1 corresponds to the productivity growth slowdown in the early 1970s. U.S. labor productivity experienced a slowdown in growth after the oil shock in 1973 (see \citet*{Perron1989} and \citet*{Hansen2001}). None of the models estimate a break date in the 1980s, which is known as “the Great Moderation," referencing an empirical fact of a large reduction in the volatility of U.S. real GDP growth in 1984Q1, established by \citet*{Kim1999} and \citet*{McConnell2000}. I focus on events that affect the mean rather than the volatility of growth rate because the change in volatility is not a linear function of the change in the lag coefficient in model ((ref)).
I check the sensitivity of the estimators to trimming the sample, by computing the two estimators for a total of 54 sub-samples, which end at different dates from 2005Q1 to 2018Q2 (the samples used to obtain estimates differ, not the fraction $\alpha$ of the optimization space within the sample). Because the sample is trimmed by one quarter each time, switching to a different estimate, which is farther, implies that the estimator is sensitive to trimming, rather than suggesting multiple breaks\footnote{It is likely that multiple structural breaks exist in the output growth rate because we are considering a sample that is over 70 years. Estimates that vary depending on the sub-sample could be evidence of more than one break in the sample period. The break date estimate $\hat{k}_{LS} =$ 2000Q2 can align with the tech bubble, also known as the dot-com crash in 2000. In relation to business cycles, $\hat{k}_{LS} = $1953Q1 and 1958Q1 are both recession in 1953, with the end of the Korean war. Under M2, both estimates from the full sample are 1966Q1, and the closest historical event that is likely to affect the output growth rate is the Vietnam war.}. Estimating multiple structural breaks using the weighting scheme is beyond the scope of this study. For LS estimation of multiple breaks see \citet*{Bai1998a} and Bai1998b.
Table (ref) provides the number of sub-samples that have the same break date estimates; the entries are fractions of the number of sub-samples out of 54 sub-samples. For M1 and M3, there are sub-samples in which the LS estimates are at the boundaries $\hat{\rho}_{LS} = 0.1$ or $0.9$. In contrast our break point estimates are either mid 1960s or early 1970s, which are in the fraction interval $\hat{\rho} \in [0.2, 0.5]$.
In short, estimating a structural break of postwar U.S. real GDP growth rate using our estimation method, provides evidence of a break occurring in 1973Q1, which corresponds to the productivity growth slowdown period. However, the LS estimates a break occurs in 2000 or 1953, depending on the time interval. Break date estimates are obtained for sub-samples with end dates 2005Q1 to 2018Q2 for both methods; the LS estimates vary considerably, with $\hat{\rho}_{LS}$ near 0.1 and 0.9 for almost 40% of the sub-samples considered under M1. In contrast, my estimates are 1966Q1 or 1973Q1 for all sub-samples and models. This suggests that the difference in LS estimates, depending on the sample period, are due to their finite sample behavior (tri-modality) rather than evidence of multiple structural breaks.
\citet*{Paye2006} studied the instability in models of ex-post predictable components in stock returns by examining structural breaks in the coefficients of state variables. The regression model ((ref)) is specified with four state variables as follows: the lagged dividend yield, short-term interest rate, term spread, and default premium. The model allows for all coefficients to change because no strong reason exists to believe that the coefficient on any of the regressors should be immune from shifts. The multivariate model with a one-time structural break at $k$ with $t=1,\ldots,T$ is
where $Ret_t$ represents the excess return for the international index in question during month $t$, $Div_{t-1}$ is the lagged dividend yield, $Tbill_{t-1}$ is the lagged local country short interest rate, $Spread_{t-1}$ is the lagged local country spread, and $Def_{t-1}$ is the lagged US default premium. From the notation of model ((ref)), $y_t = Ret_t$ and for the multivariate model, $x_t = z_t = (1,Div_{t-1},Tbill_{t-1},$ $Spread_{t-1},Def_{t-1})$. For the univariate model with dividend yield $x_t = z_t = (1,Div_{t-1})$, which is defined analogously for other univariate models. The weight matrix is $w_k = T^{-1}Z_k'MZ_k$, where $Z_k = (0,\ldots,0,z_{k+1},\ldots,z_{T})'$ and $M = I-X(X'X)^{-1}X'$. Following the approach of \citet*{Paye2006}, I examine univariate models to facilitate the interpretation of coefficients, in addition to the multivariate model ((ref)).
I collected data from Global Financial Data and Federal Reserve Economic Data (FRED). The indices of the total return and dividend yield series are the S&P 500 for the U.S. and the Financial Times Stock Exchange (FTSE) All-share for the UK. The dividend yield is expressed as an annual rate and constructed as the sum of dividends over the preceding 12 months, divided by the current price. For both countries, the three-month Treasury bill (T-bill) rate is used as a measure of the short-term interest rate and the 20-year government bond yield is the measure of the long-term interest rate. Excess returns are the total return stocks in the local currency less the total return on T-bills. The term spread is constructed as the difference between the long- and the short-term local country interest rate. The U.S. default premium is the differences in yields between Moody's Baa and Aaa rated bonds. The search grid is obtained by trimming each sample period by fraction $\alpha=0.15$ (which is equivalent to the trimming window of \citet*{Paye2006}). For the full sample, the search grid is 1960:2-1996:3, and for the sub-sample it is 1975:1-1998:10.
Under the univariate model, with the lagged dividend yield as a single forecasting regressor, the LS estimate of the break point for the S&P 500 is close to the boundary of the search grid. \citet*{Paye2006} note that the NYSE or S&P 500 indices have the same estimated break date when the trimming window is shortened, and thus, the discrepancy is not the sole explanation for the timing of the break. However, it is likely that estimates are near the boundaries because of the finite sample behavior of the LS estimator. I check whether the new estimator provides a different break date estimate under model ((ref)), using data similar to the first dataset from \citet*{Paye2006}, which is monthly data on the U.S. and the UK stock returns from 1952:7 to 2003:12. For comparison, I also estimate the break using a shorter period 1970:1-2003:12, which is equivalent to the sample period of their second dataset.
Table (ref) provides estimates of the two samples using the S&P 500 index. One notable feature is that the LS estimates that a break occurred in December 1994, with break point $\hat{\rho}_{LS} = 0.83$, whereas my method estimates a break in the mid-1980s and $\hat{\rho}_{NEW} = 0.62$. Although the LS estimate is close to the boundaries of the grid, it gives the same break estimate in the sub-sample. This suggests that a break may have occurred multiple times. \citet*{Paye2006} use the \citet*{Bai1998a} method and find that two structural breaks occur in the return model ((ref)), using the S&P 500, where each break occurs at 1987:7 and 1995:3. They note that the break in 1987 appears to be an isolated break, not appearing in other international markets. These two break date estimates are similar to estimates in Table (ref), which assume a one-time structural break.
An alternative explanation of the break in the early 1980s is that the estimation method captures a change in the individual state variable itself rather than the coefficient of the prediction model ((ref)), because the noisy nature of stock market returns makes it extremely difficult to detect a break. For instance, the estimate $\hat{\beta}_1$ could be capturing noise caused by the movement in $Div_{t-1}$ (see Figure (ref) in Appendix C).
For UK stock returns, both methods obtain a break date estimate that is (or close to) 1975:1 under all models and sample periods. This is different from the result using the S&P 500 index series, because the excess return for the FTSE All-share index increases nearly 10 standard deviations from 1975:1 to 1975:2 \footnote{For the sample period 1952:7-2003:12, the mean excess return of the FTSE All-share index is 0.5949 and the standard deviation is 5.4890. At $t = $1975:1, the excess return $Ret_t = 0.4556$ and at $t = $1975:2 we have $Ret_t = 53.2187$; therefore, the change is approximately 9.6 standard deviations.}. Hence, the change in excess returns is large enough for the LS to detect the break point appropriately. \citet*{Paye2006} relate the break in the mid-1970s to the large macroeconomic shocks reflecting large oil price increases; breaks in the underlying economic fundamentals process can explain breaks in financial return models. If this is the case, then the break magnitude is large enough for both methods to accurately estimate the break date 1975:1.
This study provides an estimation method of the structural break point in multivariate linear regression models, when a one-time break occurs in a subset of (or all) coefficients. In particular, this study focuses on break magnitudes that are empirically relevant. In practice, it is likely that the shift in parameters is small in a statistical sense. The LS estimation widely used in the literature fails to accurately estimate the break point under small break magnitudes, which motivates us to construct the estimation method in this study.
I construct a weight function on the sample period normalized to a unit interval, which imposes small weights on the LS objective for potential break points with large estimation uncertainty. The break point estimator is the argmax of the objective function that is equal to the LS objective multiplied by a weight function. The break point estimator is consistent under regularity conditions on a general weight function, with the same rate of convergence as the LS estimator from \citet*{Bai1997}. The limit distribution under a small break magnitude derives under an in-fill asymptotic framework, following the approach by Jiang2017 (Jiang2017, Jiang2018). For a structural break in a stationary linear process with a small break magnitude (inside the local $T^{-1/2}$ neighborhood of zero), the asymptotic distribution of the new estimator explicitly depends on the weight function. The limit distribution is also derived for a break in a local-to-unit root process, assuming the break magnitude is $O(T^{-1})$. Monte Carlo simulation results show that for a small break, the break point estimator reduces the RMSE compared to the LS estimator for all parameter values considered.
This study provides two empirical applications as follows: structural breaks on the U.S. real GDP growth and the U.S. and UK stock return prediction models. My break point estimator is robust to trimming of the sample, in contrast to the LS. In particular, my method estimates the break date 1973Q1 in U.S. real GDP growth rates, which LS estimation has failed to confirm. In macroeconomics literature, the “productivity growth slowdown" in the early 1970s is a widely known empirical fact.
In short, this study provides an alternative estimation method that estimates the timing of a structural break in linear regression models under empirically relevant break magnitudes. My estimator shows a uni-modal finite sample distribution under statistically small break magnitudes. To my knowledge, this is the first study to widen the class of break point estimators by generalizing least-squares. I provide theoretical results of the consistency of the estimator and an asymptotic distribution that represents finite sample behavior. If the break magnitude is small, my estimator outperforms the LS estimation in terms of RMSE. Thus, under statistically small but empirically relevant breaks, the estimator described in this study provides reliable inferences of the change point in models. The estimation method can be generalized to estimate multiple structural breaks, which is a topic for future research.