Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
123,036 characters · 19 sections · 141 citation commands
Generalized Laplace Inference in Multiple Change-Points Models
\setcounter{page}{0}
{\bf{JEL Classification}}: C12, C13, C22\\ {\bf{Keywords}}: Asymptotic Distribution, Bet-Proof, Break Date, Change-point, Generalized Laplace Inference, Highest Density Region, Quasi-Bayes.
\onehalfspacing \thispagestyle{empty} \allowdisplaybreaks
In the context of the multiple change-points model analyzed in bai/perron:98, we develop inference methods for the change-point dates for a class of Generalized Laplace (GL) estimators using a classical long-span asymptotic framework. They are defined by an integration rather than an optimization-based method, the latter typically characterizing classical extremum estimators. The idea traces back to laplace:74, who first suggested to interpret transformations of a least-squares criterion function as a statistical belief over a parameter of interest. Hence, a Laplace estimator is defined similarly to a Bayesian estimator although the former relies on a statistical criterion function rather than a parametric likelihood function. As a consequence, the GL estimator is interpreted as a classical (non-Bayesian) estimator and the inference methods proposed retain a frequentist interpretation such that the GL estimators are constructed as a function of integral transformations of the least-squares criterion. In a first step, we use the approach of bai/perron:98 to evaluate the least-squares criterion function at all candidate break dates. We then apply a transformation to obtain a proper distribution over the parameters of interest, referred to as the Quasi-posterior. For a given choice of a loss function and (possibly) a prior density, the estimator is then defined either explicitly as, for example, the mean or median of the (weighted) Quasi-posterior or implicitly as the minimizer of a smooth convex optimization problem.
The underlying asymptotic framework considered is the long-span shrinkage asymptotics of bai:97RES, bai/perron:98 and also perron/qu:06 who considerably relaxed some conditions, where the magnitude of the parameter shift is sample-size dependent and approaches zero as the sample size increases. Early contributions to this approach are hinkley:71, bhattacharya:87, and yao:87 for estimating break points. For testing for structural breaks, see hawkins:77, picard:85, kim/siegmund:89, andrews:93, horvath:93 and andrews/ploberger:94. See also the reviews of csorgo/horvath:97, perron:06, casini/perron_Oxford_Survey and references therein.
One of our goals is to develop GL estimates with better small-sample properties compared to least-squares estimates, namely lower Mean Absolute and Root-Mean Squared Errors, and confidence sets with accurate coverage probabilities and relatively short lengths for a wide range of break sizes, whether small or large; existing methods work well for either small or large breaks, but not for both. A second goal is to establish theoretical results that support the reported finite-sample properties about inference.
The asymptotic distribution of the GL estimator is derived via a local parameter related to a normalized deviation from the true fractional break date. The normalization factor corresponds to the rate of convergence of the original (extremum) least-squares estimator as established by bai/perron:98. The asymptotic distribution of the GL estimator then depends on a sample-size dependent smoothing parameter sequence applied to the least-squares criterion function. We derive two distinct limiting distributions corresponding to different smoothing sequences of the criterion function {[}cf. jun/pinkse/wan:15 for a related application in the context of the cube-root asymptotics of kim/pollard:90{]}. In one case, the estimator displays the same limit law as the asymptotic distribution of the least-squares estimator derived in bai/perron:98 {[}see also hinkley:71, picard:85 and yao:87{]}. In a second case, the limiting distribution is characterized by a ratio of integrals over functions of Gaussian processes and resembles the limiting distribution of Bayesian change-point estimators. The latter is exploited for the purpose of constructing confidence sets for the break dates. We use the concept of highest density regions (HDR) introduced by casini/perron_CR_Single_Break for structural change problems, which best summarizes the properties of the probability distribution of interest. The HDR are common in Bayesian analysis where they are applied to a posterior distribution {[}see, e.g., box/tiao:1973{]}. kendall/stuart:1961 discussed the difference between frequentist confidence intervals and Bayesian approaches in relation to the existence of a sufficient statistic. Our procedure yields confidence sets for the break date which, in finite samples, better account for the uncertainty over the parameter space in finite-samples because it effectively incorporates a statistical measure of the uncertainty in the least-squares criterion function. As noted in the literature on likelihood-based inference in some classes of nongranular problems {[}see e.g., chernozhukov/hong:03, ghosal/ghosh/samanta:95, hirano/porter:03 and ibragimov/has:81{]}, the Maximum Likelihood Estimator (MLE) is generally not an asymptotically sufficient statistic in these models and so the likelihood contains more information asymptotically than the MLE. Hence, likelihood-based procedures are generally not functions of the MLE even asymptotically. This incompleteness property motivated the study of the entire likelihood rather than just the MLE. Likewise, our method exploits the entire behavior of the objective function.
Laplace's seminal insight has been applied successfully in many disciplines. In econometrics, chernozhukov/hong:03 introduced Laplace-type estimators as an alternative to classical (regular) extremum estimators in several problems such as censored median regression and nonlinear instrumental variable; see also forneron/ng:17 for a review and comparisons. Their main motivation was to solve the curse of dimensionality inherent to the computation of such estimators. In contrast, the class of GL estimators in structural change models serves distinct multiple purposes. First, inference about the break dates presents several challenges, in particular to provide methods with a satisfactory performance uniformly over different data-generating mechanisms and break magnitudes. The GL inference proves to be reliable and accurate in finite-samples. Second, it leads to inference methods that have both frequentist and credibility properties which is not shared by the other popular methods.
Turning to the problem of constructing confidence sets for a single break date, the standard asymptotic method for the linear regression model was proposed in bai:97RES, while elliott/mueller:07 proposed to invert the locally best invariant test of nyblom:89, and eo/morley:15 suggested to invert the likelihood-ratio statistic of qu/perron:07. The latter were mainly motivated by finite-sample results indicating that the exact coverage rates of the confidence intervals obtained from Bai's (1997) method are often below the nominal level when the magnitude of the break is small. It has been shown that the method of elliott/mueller:07 delivers the most accurate coverage rates but the average length of the confidence sets is significantly larger than with other methods. The confidence sets for the break dates constructed from the GL inference that we develop result in exact coverage rates close to the nominal level and short length of the confidence sets. This holds true whether the magnitude of the break is small or large. In fact, we show that GL inference is bet-proof, a measure of “reasonableness” of frequentist inference in non-regular problems {[}see, e.g., buehler:59{]}.
The GL inference developed in this paper has been applied by casini/perron_Lap_CR_Single_Inf to achieve finite-sample improvements under the continuous record asymptotic framework of casini/perron_CR_Single_Break. The latter proposed an alternative asymptotic framework to explain the non-standard features of the finite-sample distribution of the least-squares estimator.
The paper is organized as follows. We first focus on the single change-point case. Section (ref) presents the statistical setting. We develop the asymptotic theory in Section (ref) and the inference methods in Section (ref). Results for multiple change-points models are given in Section (ref) while Section (ref) discusses some theoretical properties of GL inference. Section (ref) presents simulation results about the finite-sample performance. Section (ref) concludes. All proofs are included in an online supplement {[}casini/perron_SC_BP_Lap_Supp{]}.
This section introduces the structural change model with a single break, reviews the least-squares estimation method for the break date, and presents the relevant assumptions. We start with introducing the formal setup for our analysis. The following notation is used throughout. We denote the transpose of a matrix $A$ by $A'$. We use $\left\Vert \cdot\right\Vert $ to denote the Euclidean norm of a linear space, i.e., $\left\Vert x\right\Vert =\left(\sum_{i=1}^{p}x_{i}^{2}\right)^{1/2}$ for $x\in\mathbb{R}^{p}.$ For a matrix $A$, we use the vector-induced norm, i.e., $\left\Vert A\right\Vert =\sup_{x\neq0}\left\Vert Ax\right\Vert /\left\Vert x\right\Vert .$ All vectors are column vectors. For two vectors $a$ and $b$, we write $a\leq b$ if the inequality holds component-wise. We use $\left\lfloor \cdot\right\rfloor $ to denote the largest smaller integer function. We use $\overset{\mathbb{P}}{\rightarrow}$ and $\overset{d}{\rightarrow}$ to denote convergence in probability and convergence in distribution, respectively. $\mathbb{C}_{b}\left(\mathbf{E}\right)$ {[}$\mathbb{D}_{b}\left(\mathbf{E}\right)${]} is the collection of bounded continuous {[}c�dl�g{]} functions from some specified set $\mathbf{E}$ to $\mathbb{R}$. Weak convergence on either $\mathbb{C}_{b}\left(\mathbf{E}\right)$ or $\mathbb{D}_{b}\left(\mathbf{E}\right)$ is denoted by $\Rightarrow$. The symbol “$\triangleq$” stands for definitional equivalence.
We consider a sample of observations $\left\{ \left(y_{t},\,w_{t},\,z_{t}\right):\,t=1,\ldots,\,T\right\} ,$ defined on a filtered probability space $\left(\Omega,\,\mathscr{F},\,\mathbb{P}\right)$, on which all of the random elements introduced in what follows are defined. The model is
where $y_{t}$ is a scalar dependent variable, $w_{t}$ and $z_{t}$ are regressors of dimensions, $p$ and $q,$ respectively, and $e_{t}$ is an unobserved error term. The true parameter vectors $\phi^{0},\,\delta_{1}^{0}$ and $\delta_{2}^{0}$ are unknown and we define $\delta^{0}\triangleq\delta_{2}^{0}-\delta_{1}^{0},$ with $\delta^{0}\neq0$ so that a structural change occurs at date $T_{b}^{0}$. It is useful to re-parametrize the model. Letting $x_{t}\triangleq\left(w'_{t},\,z'_{t}\right)'$ and $\beta^{0}\triangleq\left(\left(\phi^{0}\right)',\,\left(\delta_{1}^{0}\right)'\right)',$ we have
More generally, we can define $z_{t}\triangleq D'x_{t}$, where $D$ is a $\left(p+q\right)\times q$ matrix with full column rank. A pure structural change model in which all regression parameters are subject to change corresponds to $D=I_{\left(p+q\right)\times\left(p+q\right)}$, whereas a partial structural change model arises when $D=\left(0_{q\times p},\,I_{q\times q}\right)'.$ In order to facilitate the derivations, we reformulate model (ref) in matrix format. Let $Y=\left(y_{1},\,\ldots,\,y_{T}\right)',\,X=\left(x_{1},\,\ldots,\,x_{T}\right)'$, $e=\left(e_{1},\,\ldots,\,e_{T}\right)',$ $X_{1}=\left(x_{1},\,\ldots,\,x_{T_{b}},\,0,\,\ldots,\,0\right)'$, $X_{2}=(0,\,\ldots,\,0,\,$ $x_{T_{b}+1},\ldots,\,x_{T})'$ and $X_{0}=(0,\,\ldots,\,0,\,x_{T_{b}^{0}+1},\ldots,\,x_{T})'$. Further, define $Z_{1},\,Z_{2}$ and $Z_{0}$ in a similar way: $Z_{1}=X_{1}D,\,Z_{2}=X_{2}D$ and $Z_{0}=X_{0}D$. We omit the dependence of the matrices $X_{i}$ and $Z_{i}$ ($i=1,\,2$) on $T_{b}$. Then, (ref) is equivalent to
Let $\theta^{0}\triangleq\left(\left(\phi^{0}\right)',\,\left(\delta_{1}^{0}\right)',\,\left(\delta^{0}\right)'\right)'$ denote the true value of the parameter vector $\theta\triangleq\left(\phi,\,\delta_{1},\,\delta\right).$ The break date least-squares (LS) estimator $\widehat{T}_{b}^{\mathrm{LS}}$ is the minimizer of the sum of squared residuals {[}denoted $S_{T}\left(\theta,\,T_{b}\right)${]} from (ref). The parameter $\theta$ can be concentrated out resulting in a criterion function depending only on $T_{b}=T\lambda_{b}$, i.e., $\widehat{T}_{b}^{\mathrm{LS}}=\arg\min_{1\leq T_{b}\leq T}S_{T}(\widehat{\theta}^{\mathrm{LS}}(T_{b}),\,T_{b})$ where $\widehat{\theta}^{\mathrm{LS}}(T_{b})=\arg\min_{\theta}S_{T}(\theta,\,T_{b})$ with $S_{T}(\theta,\,T_{b})=\sum_{t=1}^{T_{b}}\left(y_{t}-\phi'w_{t}-\delta'_{1}z_{t}\right)^{2}+\sum_{t=T_{b}+1}^{T}\left(y_{t}-\phi'w_{t}-\delta'z_{t}\right)^{2}$. Also,
where $M_{X}\triangleq I-X\left(X'X\right)^{-1}X'$, $\widehat{\delta}^{\mathrm{LS}}\left(\lambda_{b}\right)$ is the least-squares estimator of $\delta^{0}$ obtained by regressing $Y$ on $X$ and $Z_{2}$ and the statistic $Q_{T}\left(\widehat{\delta}^{\mathrm{LS}}\left(\lambda_{b}\right),\,\lambda_{b}\right)$ is the numerator of the sup-Wald statistic. The Laplace-type inference builds on the least-squares criterion function $Q_{T}\left(\delta\left(\lambda_{b}\right),\,\lambda_{b}\right),$ where $\delta\left(\lambda_{b}\right)$ stands for $\widehat{\delta}^{\mathrm{LS}}\left(\lambda_{b}\right)$ to minimize notational burden.
These assumptions are standard and similar to those in perron/qu:06. It is well-known that only the fractional break date $\lambda_{b}^{0}$ (not $T_{b}^{0}$) can be consistently estimated, with $\widehat{\lambda}_{b}^{\mathrm{LS}}$ having a $T$-rate of convergence. The corresponding result for the break date estimator $\widehat{T}_{b}^{\mathrm{LS}}$ states that, as $T$ increases, $\widehat{T}_{b}^{\mathrm{LS}}$ remains within a bounded distance from $T_{b}^{0}$. However, this does not affect the estimation problem of the regression coefficients $\theta^{0}$, for which $\widehat{\theta}^{\mathrm{LS}}$ is a regular estimator; i.e., $\sqrt{T}$-consistent and asymptotically normally distributed, since the estimation of the regression parameters is asymptotically independent from the estimation of the change-point. Hence, the regression parameters are essentially estimated as if the change-point was known. More complex is the derivation of the asymptotic distribution of $\widehat{\lambda}_{b}^{\mathrm{LS}}$; e.g., hinkley:71 for an i.i.d. Gaussian process with a mean change. Therefore, to make progress it is necessary to consider a shrinkage asymptotic setting in which the size of the shift converges to zero as $T\rightarrow\infty$; see picard:85 and yao:87 and extended by bai:97RES to general linear models.
We define the GL estimator in Section (ref) and discuss its usefulness in Section (ref). Section (ref) describes the asymptotic framework under which we derive the limiting distribution with the results presented in Section (ref).
The class of GL estimators relies on the original least-squares criterion function $Q_{T}\left(\delta\left(\lambda_{b}\right),\,\lambda_{b}\right)$, with the parameter of interest being $\lambda_{b}^{0}=T_{b}^{0}/T$. The Quasi-posterior $p_{T}\left(\lambda_{b}\right)$ is defined by the exponential transformation,
where $\pi\left(\cdot\right)$ is a density function. Note that $p_{T}\left(\lambda_{b}\right)$ defines a proper distribution over the parameter space $\varGamma^{0}$. The $\mathscr{\mathscr{L}}\left(\theta,\,T_{b}\right)$-class of estimators are the solutions of smooth convex optimization problems for a given loss function, restricting attention to convex loss functions $l_{T}\left(\cdot\right)$. Examples include (a) $l_{T}\left(r\right)=a_{T}^{m}\left|r\right|^{m},$ the polynomial loss function (the squared loss function is obtained when $m=2$ and the absolute deviation loss function when $m=1$); (b) $l_{T}\left(r\right)=a_{T}\left(\tau-\mathbf{1}\left(r\leq0\right)\right)r,$ the check loss function; where $a_{T}$ is a divergent sequence. We define the Expected Risk function, under the density $p_{T}\left(\cdot\right)$ and the loss $l_{T}\left(\cdot\right)$ as $\mathcal{R}_{l,T}\left(s\right)\triangleq\mathbb{E}_{p_{T}}\left[l_{T}\left(s-\widetilde{\lambda}_{b}\right)\right],$ where $\widetilde{\lambda}_{b}$ is a random variable with distribution $p_{T}$ and $\mathbb{E}_{p_{T}}$ denotes expectation taken under $p_{T}.$ Using (ref) we have,
The Laplace-type estimator $\widehat{\lambda}_{b}^{\mathrm{GL}}$ shall be interpreted as a decision rule that, given the information contained in the Quasi-posterior $p_{T}$, is least unfavorable according to the loss function $l_{T}$ and the prior density $\pi$. Then $\widehat{\lambda}_{b}^{\mathrm{GL}}$ is the minimizer of the expected risk function (ref), i.e., $\widehat{\lambda}_{b}^{\mathrm{GL}}\triangleq\arg\min_{s\in\varGamma^{0}}\left[\mathcal{R}_{l,T}\left(s\right)\right].$ Observe that the GL estimator $\widehat{\lambda}_{b}^{\mathrm{GL}}$ results in the mean (median) of the Quasi-posterior upon choosing the squared (absolute deviation) loss function. The choice of the loss and of the prior density functions hinges on the statistical problem addressed. In the structural change problem, a natural choice for the Quasi-prior $\pi$ is the density of the asymptotic distribution of $\widehat{\lambda}_{b}^{\mathrm{LS}}$. This requires to replace the population quantities appearing in that distribution by consistent plug-in estimates\textemdash cf. bai/perron:98\textemdash and derive its density via simulations as in casini/perron_CR_Single_Break. The attractiveness of the Quasi-posterior (ref) is that it provides additional information about the parameter of interest $\lambda_{b}^{0}$ beyond what is already included in the point estimate $\widehat{\lambda}_{b}^{\mathrm{LS}}$ and its distribution (see Section (ref)). This approach will result in more accurate inference in finite-samples even in cases with high uncertainty in the data as we shall document in Section (ref). This is supported in Section (ref) showing that the GL inference is bet-proof which is a desirable theoretical property in non-regular problems.
Assumption (ref) is similar to those in bickel/yahav:69, ibragimov/has:81 and chernozhukov/hong:03. The convexity assumption on $l_{T}\left(\cdot\right)$ is guided by practical considerations. The dominant restriction in part (iii) is conventional and implicitly assumes that the loss function has been scaled by some constant. What is important is that the growth of the function $l_{T}\left(r\right)$ as $\left|r\right|\rightarrow\infty$ is slower than $\exp\left(\epsilon\left|r\right|\right)$ for any $\epsilon>0$. Assumption (ref) on the prior is satisfied for any reasonable choice. For priors that have a peak at $\lambda_{b}^{0}$ one can apply some basic smoothing techniques to make it differentiable locally {[}e.g., mean smoothing, Gaussian smoothing and Savitzky-Golay filter{]}. We did not find any particular difference in the empirical results and so we used the mean smoothing. The assumption on the differentiability of the kernel can be relaxed at the expense of one more step in the proof. chernozhukov/hong:03 assumed differentiability of the prior; we also keep the same assumption and applied the smoothing. The large-sample properties of the $\mathscr{\mathscr{L}}\left(\theta,\,T_{b}\right)$-class are studied under the shrinkage asymptotic setting of bai:97RES and bai/perron:98. Thus, we need the following assumption.
We omit the superscript 0 from $\delta_{T}^{0}$ for notational convenience since it should not cause any confusion. Assumption (ref) requires the magnitude of the break to shrink to zero at any slower rate than $T^{-1/2}$. The specific rates allowed differ from those in bai:97RES and bai/perron:98, since they require $\vartheta\in\left(0,\,1/2\right)$. The reason is merely technical; the asymptotics of the Laplace-type estimator involve smoothing the criterion function, and thus one needs to guarantee that $\widehat{\lambda}_{b}$ approaches $\lambda_{b}^{0}$ at a sufficiently fast rate. Under the shrinkage asymptotics, Proposition 1 and Corollary 1 in bai:97RES state that $T\left\Vert \delta_{T}\right\Vert ^{2}\left(\widehat{\lambda}_{b}^{\mathrm{LS}}-\lambda_{b}^{0}\right)=O_{\mathbb{P}}\left(1\right)$ and $\widehat{\delta}_{T}^{\mathrm{LS}}-\delta_{T}=o_{\mathbb{P}}\left(1\right)$.
We use Figure (ref)-(ref) to illustrate the main idea behind the usefulness of the GL method. They present plots of the density of the distribution of $\widehat{T}_{b}^{\mathrm{LS}}$ and $\widehat{T}_{b}^{\mathrm{GL}}$ for the simple model $y_{t}=\phi^{0}+z_{t}\left(\delta_{1}^{0}+\delta^{0}\mathbf{1}\left\{ t>T_{b}^{0}\right\} \right)+e_{t}$ where $\left\{ z_{t}\right\} $ follows an ARMA(1,1) process and $e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right)$. The distributions presented are the exact finite-sample distributions of the LS and GL estimators, bai:97RES\textquoteright s (1997) classical large-$N$ limit distribution, and the asymptotic distribution of the GL estimator. Noteworthy are the non-standard features of the finite-sample distribution of the LS estimator when the break magnitude is small, which include multi-modality, fat tails and asymmetry. The central mode is near $\widehat{T}_{b}^{\mathrm{LS}}$ while the other two modes are in the tails near the start and end of the sample period; when the break magnitude is small $\widehat{T}_{b}^{\mathrm{LS}}$ tends to locate the break in the tails since the evidence of a break is weak. It is evident that the classical large-$N$ asymptotic distribution provides a poor approximation especially for small break sizes. Some of these features have been found in other works {[}see, e.g., perron/zhu:05, deng/perron:06, Jiang, et al. jiang/wang/yu:16,jiang/wang/yu:17, and casini/perron_CR_Single_Break{]}. Turning to the densities of the GL estimators, some of the nonstandard features appear also for the GL estimator although to a much lesser extent. In particular, the densities of the GL estimators are less spread out than the corresponding densities for the LS estimator. For small breaks, the finite-sample distributions of the LS and GL estimators are quite different, which suggests that standard measures of accuracy (e.g., MAE and RMSE) can be expected to differ substantially. For $\lambda_{0}=0.5$ the GL estimator exhibits much less variability and more precision. The figures also show that the asymptotic distribution of the GL estimator provides an accurate approximation for large breaks while for small breaks the approximation is less accurate. However, it captures the fat-tails of the finite-sample distribution which suggests that it does not underestimate uncertainty about the break location unlike bai:97RES\textquoteright s (1997) distribution.
The GL method is useful because it weights the information from the least-squares criterion function with the information from the prior density\textemdash which, here, is the density of the asymptotic distribution of $\widehat{T}_{b}^{\mathrm{LS}}$. Note that the least-squares objective function is quite flat when the magnitude of the break is small and so $\widehat{T}_{b}^{\mathrm{LS}}$ is imprecise. The resulting Quasi-posterior, or, e.g., its median, is likely to lead to better estimates in finite-samples, because it takes into account the overall shape of the objective function which weighted by the prior becomes more informative about the uncertainty of the break date.
In order to develop the asymptotic results, we introduce a smoothing sequence $\left\{ \gamma_{T}\right\} $ whose properties are specified below and work with a normalized version of $\mathcal{R}_{l,T}\left(s\right)$ in order to be able to derive the relevant limit results. We assume that $\lambda_{b}^{0}\in\varGamma^{0}\subset\left(0,\,1\right)$ is the unknown extremum of $\widetilde{Q}\left(\theta^{0},\,\lambda_{b}\right)=\mathbb{E}\left[Q_{T}\left(\theta^{0},\,\lambda_{b}\right)\right]$ and that $\theta^{0}\triangleq\left(\left(\phi^{0}\right)',\,\left(\delta_{1}^{0}\right)',\,\left(\delta^{0}\right)'\right)'\in\mathbf{S}\subset\mathbb{R}^{p}\times\mathbb{R}^{q}\times\mathbb{R}^{q}$. Our analysis is within a vanishing neighborhood of $\theta^{0}$. For any $\theta\in\mathbf{S}$, let $\lambda_{b}^{0}\left(\theta\right)$ be an arbitrary element of $\varGamma^{0}\left(\theta\right)\triangleq\left\{ \lambda_{b}\in\varGamma^{0}:\,\widetilde{Q}\left(\theta,\,\lambda_{b}\right)=\sup_{\widetilde{\lambda}_{b}\in\mathcal{\varGamma}^{0}}\widetilde{Q}\left(\theta,\,\widetilde{\lambda}_{b}\right)\right\} $. Provided a uniqueness condition is assumed (see Assumption (ref)), $\varGamma^{0}\left(\theta\right)$ contains a single element, $\lambda_{b}^{0}$. Further, let $\overline{Q}_{T}\left(\theta,\,\lambda_{b}\right)\triangleq Q_{T}\left(\theta,\,\lambda_{b}\right)-Q_{T}\left(\theta,\,\lambda_{b}^{0}\right),$ $Q_{T}^{0}\left(\theta,\,\lambda_{b}\right)\triangleq\mathbb{E}\left[Q_{T}\left(\theta,\,\lambda_{b}\right)-Q_{T}\left(\theta,\,\lambda_{b}^{0}\right)|\,X\right],$ and $G_{T}\left(\theta,\,\lambda_{b}\right)\triangleq\overline{Q}_{T}\left(\theta,\,\lambda_{b}\right)-Q_{T}^{0}\left(\theta,\,\lambda_{b}\right).$ These expressions are given by $G_{T}\left(\theta,\,\lambda_{b}\right)=g_{e}\left(\theta,\,\lambda_{b}\right)$, $Q_{T}^{0}=g_{d}\left(\theta,\,\lambda_{b}\right)$ and $\overline{Q}_{T}=g_{d}\left(\theta,\,\lambda_{b}\right)+g_{e}\left(\theta,\,\lambda_{b}\right)$, where
and
They are derived in Section (ref). For the purpose of developing the asymptotic theory, the GL estimator $\widehat{\lambda}_{b}^{\mathrm{GL}}\left(\theta\right)$ is defined as the minimizer of a normalized version of $\mathcal{R}_{l,T}\left(s\right)$:
Note that, under Condition (ref) below, this is equivalent to the minimizer of $\mathcal{R}_{l,T}\left(s\right)$ since $\overline{Q}_{T}\left(\theta,\,\lambda_{b}\right)$ can always be normalized without affecting its maximization. Different choices of $\left\{ \gamma_{T}\right\} $ give rise to GL estimators with different limiting distributions. Using $\delta_{T}$ or any consistent estimate (e.g., $\widehat{\delta}_{T}^{\mathrm{LS}}$) in the factor $\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)$ is irrelevant because they are asymptotically equivalent. Our analysis is local in nature and thus we write $\widehat{\lambda}_{b}^{\mathrm{GL}}(\widehat{\theta})\triangleq\widehat{\lambda}_{b}^{\mathrm{GL,}*}(r_{T}(\widehat{\theta}-\theta^{0}),\,r_{T}(\widehat{\theta}-\theta^{0})),$ where $r_{T}$ is the convergence rate of $\widehat{\theta}-\theta^{0}$. Note that $G_{T}\left(\cdot,\,\cdot\right)$ and $Q_{T}^{0}\left(\cdot,\,\cdot\right)$ constitute the stochastic and the deterministic part of the objective function, respectively. Both depend on $r_{T}(\widehat{\theta}-\theta^{0})$ and our proof proceeds in conditioning first on the effect of $r_{T}(\widehat{\theta}-\theta^{0})$ on the deterministic part to obtain weak convergence of the stochastic part to a limit process that does not depend on this conditioning. See below for more details. Hence, it is required to introduce two indices $\widetilde{v}$ and $v$, such that we define $\widehat{\lambda}_{b}^{\mathrm{GL}}(\widehat{\theta})=\widehat{\lambda}_{b}^{\mathrm{GL,*}}(\widetilde{v},\,v)$ as the minimizer of
For each $v,$ we show weak convergence as a function of $\widetilde{v}$ to a limit process that does not depend on $v.$ In a second step, we use the monotonicity in $v$ of $Q_{T}^{0}$ which, relying on the argument in jureckova:77, allows us to achieve weak convergence uniformly in $v$. We first show the consistency and rate of convergence of $\widehat{\lambda}_{b}^{\mathrm{GL}}$. These results imply that $\theta^{0}$ is estimated as if $T_{b}^{0}$ were known. Thus, $\widehat{\theta}$ is $\sqrt{T}$-consistent and asymptotically normal so that we set $r_{T}=\sqrt{T}$ hereafter. We first show, for each pair $\left(v,\,\widetilde{v}\right)$ with $v,\,\widetilde{v}\in\mathbf{V}$, the convergence of the marginal distributions of the sample function $\Psi_{l,T}\left(s;\,v,\,\widetilde{v}\right)$ to the marginal distributions of the random function
where
and $W_{1},\,W_{2}$ are independent standard Wiener processes defined on $[0,\,\infty)$. The limit process $\Psi_{l}^{0}\left(s\right)$ does not depend on $v$ nor $\widetilde{v}$. Next, we show that the family of probability measures in $\mathbb{C}_{b}\left(\mathbf{K}\right)$, with $\mathbf{K}\triangleq\left\{ s\in\mathbb{R}:\,\left|s\right|\leq K\textrm{ and }K<\infty\right\} $, generated by the contractions of $\Psi_{l,T}\left(s;\,\widetilde{v},\,v\right)$ on $\mathbf{K}$ is dense uniformly in $\left(v,\,\widetilde{v}\right)$. Finally, we examine the oscillations of the minimizers of the sample criterion $\Psi_{l,T}\left(s;\,v,\,\widetilde{v}\right)$.
It is important to note that the results derived in this section are more general than what is required for the structural change model. The reason is that the change-point model is recovered as a special case corresponding to $\Psi_{l,T}\left(s\right)=\Psi_{l,T}\left(s;\,0,\,0\right)$. That is, defining the GL estimator in a $1/r_{T}$-neighborhood of the slope parameter vector $\theta^{0}$ is not strictly necessary and one can essentially develop the same analysis with $\theta$ fixed at its true value $\theta^{0}$. This relies on the properties of (orthogonal) least-squares projections and would not apply, for example, to the least absolute deviation (LAD) estimator of the break date {[}cf. baiLAD:95{]} for which $\Psi_{l,T}\left(s;\,\widetilde{v},\,v\right)$ should instead be considered. The same issue is present when estimating structural changes in the quantile regression model {[}cf. oka/qu:10{]} and in using instrumental variables models {[}cf. hall/han/boldea:10 and Perron and Yamamoto perron/yamamoto:14,perron/yanamoto:15{]}. We establish theoretical results under this more general setting since they may be useful for future work.
Let $\lambda_{b,T}^{0}\left(v\right)=\lambda_{b,T}^{0}\left(\theta^{0}+v/r_{T}\right)$. Introduce the local parameter $u=\psi_{T}\left(\lambda_{b}-\lambda_{b,T}^{0}\left(v\right)\right)$ and let $\pi_{T,v}\left(u\right)\triangleq\pi\left(\lambda_{b,T}^{0}\left(v\right)+u/\psi_{T}\right)$, $Q_{T,v}\left(u\right)\triangleq Q_{T}^{0}(\theta^{0}+v/r_{T},\,\lambda_{b,T}^{0}$ $\left(v\right)+u/\psi_{T}$), and $\widetilde{G}_{T,v}\left(u,\,\widetilde{v}\right)\triangleq G_{T}\left(\theta^{0}+\widetilde{v}/r_{T},\,\lambda_{b,T}^{0}\left(v\right)+u/\psi_{T}\right)$, where the sequence $\left\{ \psi_{T}\right\} $ depends on the results on consistency and rate of convergence of $\widehat{\lambda}_{b}^{\mathrm{GL}}$ in Proposition (ref). Apply a simple substitution in (ref) to yield,
where $\Gamma_{T}\triangleq\left\{ u\in\mathbb{R}:\,\lambda_{b}^{0}+u/\psi_{T}\in\varGamma^{0}\right\} $.
Assumptions (ref)-(ref) are equivalent to A9 in bai:97RES and A7 in bai/perron:98. More specifically, Assumption (ref) requires that, within each regime, an Invariance Principle holds for $\left\{ z_{t}e_{t}\right\} .$ Let $\zeta_{t}\triangleq z_{t}e_{t}$. For $u\leq0$ let $g\left(\zeta_{t};\,u\right)\triangleq\left(\delta^{0}\right)'\sum_{t=T_{b}^{0}+\left\lfloor u/v_{T}^{2}\right\rfloor }^{T_{b}^{0}}\zeta_{t}$ and $\widetilde{g}\left(\zeta_{t};\,u,\,\widetilde{v},\,v;\,\psi_{T},\,r_{T}\right)\triangleq\sqrt{\psi_{T}}\left(\delta^{0}+\widetilde{v}/r_{T}\right)'\sum_{t=T\lambda_{b}^{0}\left(\theta^{0}+v/r_{T}\right)+\left\lfloor u/\psi_{T}\right\rfloor }^{T\lambda_{b}^{0}\left(\theta^{0}+v/r_{T}\right)}\zeta_{t}.$ Define analogously $g\left(\zeta_{t};\,u\right)$ and $\widetilde{g}\left(\zeta_{t};\,u,\,\widetilde{v},\,v;\,\psi_{T},\,r_{T}\right)$ for the case $T_{b}>T_{b}^{0}$. We now present some technical assumptions that are necessary for the derivation of the asymptotic results for the GL estimate.
Part (i) of Assumption (ref) is an identification condition. Assumption (ref)-(ii) holds whenever $\widehat{\lambda}_{b}$ is consistent. With Assumption (ref)-(ii) we fully characterize the Gaussian component of the limit process $\mathscr{V}\left(\cdot\right);$ it implies that $\varSigma\left(\cdot,\,\cdot\right)$ is strictly positive and that
where the second implication requires some simple but tedious manipulations. Finally, the following assumption is automatically satisfied if $l\left(\cdot\right)$ is a convex function with a unique minimum.
We first show the consistency and rate of convergence of the GL estimator. The latter allows us to characterize the rate of $\psi_{T}$ and proceed with the asymptotic analysis in a neighborhood of $\lambda_{b}^{0}$. In practice, the squared loss function is often employed. Hence, it is useful to first present in Theorem (ref) the theoretical results for this case for which the GL estimator is $\widehat{\lambda}_{b}^{\mathrm{GL}}=\int_{\varGamma^{0}}\lambda_{b}p_{T}\left(\lambda_{b}\right)d\lambda_{b},$ i.e., the Quasi-posterior mean. This allows us to keep the theoretical results tractable and provide the main intuition without the need of complex notation. This case is also instructive since we can compare our results with corresponding ones for the least-squares and Bayesian change-point estimators. Corresponding results for general loss functions are given in Theorem (ref).
The rate of convergence is similar to that of the LS estimator; the difference being that $\psi_{T}=T^{1-2\vartheta}$ with $\vartheta\in\left(0,\,1/2\right)$ for the LS estimator and $\vartheta\in\left(0,\,1/4\right)$ for the GL estimator.
For the squared loss function $\widehat{\lambda}_{b}^{\mathrm{GL}}\left(\widehat{\theta}\right)\triangleq\widehat{\lambda}_{b}^{\mathrm{GL,}*}\left(\widetilde{v},\,v\right),$ where
and $v,$ $\widetilde{v}$ each belong to some compact set $\mathbf{V}\subset\mathbb{R}^{p+2q}$. For each $v\in\mathbf{V},$ we consider $\widehat{\lambda}_{b}^{\mathrm{GL,*}}\left(\cdot,\,v\right)$ as a random process with paths in $\mathbb{D}_{b}\left(\mathbf{V}\right)$. We focus on the weak convergence of $\widehat{\lambda}_{b}^{\mathrm{GL},*}\left(\cdot,\,v\right)$ for fixed $v$ since the limit process is independent of $v$ and constant as a function of $\widetilde{v}$; we then exploit monotonicity in $v.$ More precisely, we will show that for $\lambda_{b,T}^{0}\left(v\right)=\lambda_{b,T}^{0}\left(\theta^{0}+v/r_{T}\right)$ and diverging sequences $\left\{ \gamma_{T}\right\} $ and $\left\{ r_{T}\right\} $, the sequence $a_{T}\left(\widehat{\lambda}_{b}^{\mathrm{GL},*}\left(\widetilde{v},\,v\right)-\lambda_{b,T}^{0}\left(v\right)\right)$ converges in distribution in $\mathbb{D}_{b}\left(\mathbf{V}\right)$ for each $v$ to a limit process not depending on $v$ nor $\widetilde{v}$. Since it is monotonic in $v$, we do not need to show uniform convergence directly. Introduce the local parameter $u=\psi_{T}\left(\lambda_{b}-\lambda_{b,T}^{0}\left(v\right)\right)$; a simple substitution in (ref) yields,
where again we have used the notation $\pi_{T,v}\left(u\right)=\pi\left(\lambda_{b,T}^{0}\left(v\right)+u/\psi_{T}\right),$ $Q_{T,v}\left(u\right)=Q_{T}^{0}\left(\theta^{0}+v/r_{T}\right.$ $\left.\lambda_{b,T}^{0}\left(v\right)+u/\psi_{T}\right)$ and $\widetilde{G}_{T,v}\left(u,\,\widetilde{v}\right)=G_{T}\left(\theta^{0}+\widetilde{v}/r_{T},\,\lambda_{b,T}^{0}\left(v\right)+u/\psi_{T}\right)$. The limit of the GL estimator depends on the limit of the process $\left(\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)\right)\left(\widetilde{G}_{T,v}\left(u,\,\widetilde{v}\right)+Q_{T,v}\left(u\right)\right)$. As part of the proof of Theorem (ref), we show that the sequence of processes $\left\{ \widetilde{G}_{T,v}\left(u,\,\widetilde{v}\right),\,T\geq1\right\} $ converges weakly in $\mathbb{D}_{b}\left(\mathbb{R}\times\mathbf{V}\right)$ to a Gaussian process $\mathscr{W}$ not varying with $v$, whereas $Q_{T,v}\left(\cdot\right)$ is approximated by a (deterministic) drift process taking negative values, and is monotonic in $v$ and flat in $\widetilde{v}$. We show that this implies that $\widehat{\lambda}_{b}^{\mathrm{GL},*}\left(\widetilde{v},\,v\right)-\lambda_{b,T}^{0}\left(v\right)$ is monotonic in $v$ which then leads to uniform convergence in $v$ following the argument of jureckova:77.
In anticipation of the results, we make a few comments about the notation for the weak convergence of processes on the space of bounded c�dl�g functions $\mathbb{D}_{b}$. Let $\mathbf{V}\subset\mathbb{R}^{p+2q}$ be a compact set. Let $W_{T}\left(u,\,\widetilde{v},\,v\right)$ denote an arbitrary sample process with bounded c�dl�g paths evaluated at the local parameters $u\in\mathbb{R},$ and $v,\,\widetilde{v}\in\mathbf{V}$. For each fixed $v\in\mathbf{V}$, we shall write $W_{T}\left(u,\,\widetilde{v},\,v\right)\Rightarrow\mathscr{W}\left(u,\,\widetilde{v},\,v\right)$ in $\mathbb{D}_{b}\left(\mathbb{R}\times\mathbf{V}\right)$ whenever the process $W_{T}\left(\cdot,\,\cdot,\,v\right)$ converges weakly to $\mathscr{W}\left(\cdot,\,\cdot,\,v\right)$, where $\mathscr{W}\left(\cdot,\,\cdot,\,v\right)$ also belongs to $\mathbb{D}_{b}\left(\mathbb{R}\times\mathbf{V}\right)$. As a shorthand, we shall omit the argument $u\,\left(\widetilde{v}\right)$ if the limit process does not depend on $u\,\left(\widetilde{v}\right)$. The same notational conventions are used for the case when $W_{T}$ is only a function of $\left(\widetilde{v},\,v\right).$ In Theorem (ref) the convergence holds for every $v\in\mathbf{V}$, stated as convergence in $\mathbb{D}_{b}$.
Theorem (ref) states that the asymptotic distribution of the GL estimate is a ratio of integrals of functions of tight Gaussian processes. We shall compare this result with the limiting distribution of the Bayesian change-point estimator of ibragimov/has:81. They considered a simple diffusion process with a change-point in the deterministic drift {[}see their eq. (2.17) on pp. 338{]}. The limiting distribution of the GL estimate from Theorem (ref) for the case of a break in the mean for model (ref) is essentially the same (and exactly so in the i.i.d. case with stationary regimes) as theirs. Hence, while the GL estimator has a classical (frequentist) interpretation, it is first-order equivalent in law to a corresponding Bayes-type estimator.
We now present a result about the dual nature of the limiting distribution of the GL estimator. The following proposition shows that, under different conditions on the smoothing sequence parameter $\left\{ \gamma_{T}\right\} $, the GL estimator achieves different limiting distributions.
Corollary (ref) and Proposition (ref) show that with enough smoothing applied, the GL estimator is (first-order) asymptotically equivalent to the least-squares or MLE {[}cf. bai:97RES and yao:87, respectively{]}. The intuition is that when the criterion function is sufficiently smoothed, the Quasi-posterior probability density converges to the generalized dirac probability measure concentrated at the argmax of the limit criterion function. This is analogous to a well-known result {[}cf. Corollary 5.11 in robert/casella:04{]}, stating that in a parametric statistical experiment indexed by a parameter $\theta\in\Theta,$ the MLE $\widehat{\theta}_{T}^{\mathrm{ML}}$ is the limit of a Bayes estimator as the smoothing parameter $\gamma\rightarrow\infty$, i.e., using obvious notation:
For general loss functions satisfying Assumption (ref), Theorem (ref) shows that $T\left\Vert \delta_{T}\right\Vert ^{2}\left(\widehat{\lambda}_{b}^{\mathrm{GL}}-\lambda_{b}^{0}\right)$ is (first-order) asymptotically equivalent to $\xi_{l}^{0}$ defined by
The existence and uniqueness of $\xi_{l}^{0}$ follow from Assumption (ref). If one interprets $p_{0}^{*}\left(u\right)$ as a true posterior density function, then $\xi_{l}^{0}$ would naturally be viewed as a Bayesian estimator for the loss function $l_{T}\left(\cdot\right).$ In particular, in analogy to the above comparison with the Bayesian estimator of ibragimov/has:81, one can interpret the GL estimator as a Quasi-Bayesian estimator. While this is by itself a theoretically interesting result, we actually exploit it to construct more reliable inference methods about the date of a structural change. Under the least-absolute deviation loss, the GL estimator converges in distribution to the median of $p_{0}^{*}\left(u\right)$. We shall use the results in Theorem (ref)-(ref) but not Proposition (ref) since the latter implies the same confidence intervals as in bai:97RES and bai/perron:98. GL inference based on the Bayes-type limiting distribution provides a more accurate description of the uncertainty over the parameter space than the inference based on the density of $\arg\max_{s\in\mathbb{R}}\mathscr{V}\left(s\right)$ which underestimates uncertainty as shown by confidence intervals with empirical coverage rates below the nominal level particularly when the magnitude of the break is small (see Section (ref)). After some investigation, we found that both estimation and inference under the least-absolute loss works well and this is what will be used in our simulation study.
In this section, we discuss inference procedures for the break date based on the large-sample results of the previous section. Inference under general loss functions based on Theorem (ref) is what we recommend to use in practice, in particular with an absolute loss function.
Since the limiting distribution from Theorem (ref) involves certain unknown quantities, we begin by assuming that they can be replaced by consistent estimates. They are easy to construct {[}cf. bai:97RES and bai/perron:98; see also Section (ref){]}.
The first part and (i) of the second part follow from consistency of $\widehat{\lambda}_{b,T}$ and from an Invariance Principle {[}cf. Assumption (ref){]}. Part (ii)-(iii) are implied by Assumption (ref)-(ii) and consistency of $\widehat{\lambda}_{b,T}$. Let $\left\{ \widehat{\mathscr{W}}_{T}\right\} $ be a (sample-size dependent) sequence of two-sided zero-mean Gaussian processes with covariance $\widehat{\varSigma}_{T}.$ Construct the process $\widehat{\mathscr{V}}_{T}$ by replacing the population quantities in $\mathscr{V}$ by their corresponding estimates from the first part of Assumption (ref) and further, replace $\mathscr{W}$ by $\widehat{\mathscr{W}}_{T}$. Assumption (ref)-(i) basically implies that the finite-dimensional limit law of $\left\{ \widehat{\mathscr{W}}_{T}\right\} $ is the same as the finite-dimensional laws of $\mathscr{W}$ while parts (ii)-(iii) are needed for the integrability of the transform $\exp\left(\widehat{\mathscr{V}}_{T}\left(\cdot\right)\right)$. Part (iv) is needed for the proof of the asymptotic stochastic equicontinuity of $\left\{ \widehat{\mathscr{W}}_{T}\right\} $. Let $\widehat{\xi}_{T}$ be defined as the sample analogue of $\xi_{l}^{0}$ that uses $\widehat{\mathscr{V}}_{T}\left(v\right)$ in place of $\mathscr{V}_{T}\left(v\right)$ in (ref). The distribution of $\widehat{\xi}_{T}$ can be evaluated numerically.
The asymptotic distribution theory of the GL estimator may be exploited in several ways to conduct inference about the break date. The finite-sample distribution of the LS nd GL estimate of the break date displays significant non-standard features (cf. Figure (ref)). Hence, a conventional two-sided confidence interval may not result in a confidence set with reliable properties across all break magnitudes and break locations. Thus, as in casini/perron_CR_Single_Break, we use the concept of Highest Quasi-posterior Density (HQPD) regions, defined analogously to the Highest Density Region (HDR); cf. hyndman:96. See also samworth/wand:10 and mason/polonik:09 (2008, 2009)\nocite{mason/polonik:08} for more recent developments. For an illustrative example on the properties of the HDR see the discussion of Figure 11 in casini/perron_CR_Single_Break.
For $s=T\left\Vert \delta_{T}\right\Vert ^{2}\left(\widehat{\lambda}_{b}^{\mathrm{LS}}-\lambda_{b}^{0}\right)$, the asymptotic distribution theory of bai:97RES suggests a belief $\pi\left(s\right)$ over $s\in\mathbb{R}.$ This belief function can be used as a Quasi-prior for $\lambda_{b}$ in the definition of the Quasi-posterior $p_{T}\left(\lambda_{b}\right).$ Let $\mu\left(\lambda_{b}\right)$ denote some density function defined by the Radon-Nikodym equation $\mu\left(\lambda_{b}\right)=dp_{T}\left(\lambda_{b}\right)/d\lambda_{\mathrm{L}},$ where $\lambda_{\mathrm{L}}$ denotes the Lebesgue measure. The following algorithm describes how to construct a confidence set for $T_{b}^{0}$.
In principle, any Quasi-prior $\pi\left(\lambda_{b}\right)$ satisfying Assumption (ref) can be used. Note that $C_{\mathrm{HQPD}}\left(\mathrm{cv}_{\alpha}\right)$ retains a frequentist interpretation, since no parametric likelihood function of the data is required.
Following bai/perron:98, the multiple linear regression model with $m$ change-points is
for $j=1,\ldots,\,m+1$, where by convention $T_{0}^{0}=0$ and $T_{m+1}^{0}=T$. There are $m$ unknown break points $\left(T_{1}^{0},\ldots,\,T_{m}^{0}\right)$ and consequently $m+1$ regimes each corresponding to a distinct parameter value $\delta_{j}^{0}$. The purpose is to estimate the unknown regression coefficients together with the break points when $T$ observations on $\left(y_{t},\,w_{t},\,z_{t}\right)$ are available. Many of the theoretical results follow directly from the single break case; the break points are asymptotically distinct and thus, given the mixing conditions, our results for the single break date extend readily to multiple breaks. More complicated is the computation of the estimates of the break dates which has been addressed by bai/perron:03 who proposed an efficient algorithm based on the principle of dynamic programming; see also hawkins:76.
Let $T_{i}\triangleq\left\lfloor T\lambda_{i}\right\rfloor $ and $\theta\triangleq\left(\phi',\,\delta'_{1},\,\Delta_{1}^{\prime}\ldots,\,\Delta_{m}^{\prime}\right)'$ where $\Delta_{i}=\delta_{i+1}-\delta_{i}$, $i=1,\ldots,\,m$. The class $\mathscr{\mathscr{L}}\left(\theta,\,T_{i};\,1\leq i\leq m\right)$ of GL estimators in multiple change-points models relies on the least-squares criterion function $Q_{T}\left(\delta\left(\boldsymbol{\lambda}_{b}\right),\,\boldsymbol{\lambda}_{b}\right)=\sum_{i=1}^{m+1}\sum_{t=T_{i}-1}^{T_{i}}\left(y_{t}-w'_{t}\phi-z'_{t}\delta_{i}\right)^{2}$, with $\boldsymbol{\lambda}_{b}\triangleq\left(\lambda_{i};\,1\leq i\leq m\right)$. In order to state the large-sample properties, we need to introduce the shrinkage theoretical framework of bai/perron:98.
Next, for $i=1,\ldots,\,m$, define $\Xi_{Z,i}=\left(\Delta_{i}^{0}\right)'V_{i+1}\Delta_{i}^{0}/\left(\Delta_{i}^{0}\right)'V_{i}\Delta_{i}^{0},$ $\Xi_{e,i}^{2}=\left(\Delta_{i}^{0}\right)'\Sigma_{i+1}\Delta_{i}^{0}/\left(\Delta_{i}^{0}\right)'\Sigma_{i}\Delta_{i}^{0}$, and let $W_{1}^{\left(i\right)}\left(s\right)$ and $W_{2}^{\left(i\right)}\left(s\right)$ be independent Wiener processes defined on $[0,\,\infty)$, starting at 0 when $s=0$; $W_{1}^{\left(i\right)}\left(s\right)$ and $W_{2}^{\left(i\right)}\left(s\right)$ are also independent over $i.$ Finally, define
We now extend the notation of Section (ref) to the present context. By redefining the Quasi-posterior $p\left(\boldsymbol{\lambda}_{b}\right)$ in terms of $\boldsymbol{\lambda}_{b},$ the GL estimator as the minimizer of the associated risk function {[}recall (ref){]}, $\widehat{\boldsymbol{\lambda}}_{b}^{\mathrm{GL}}=\arg\min_{s\in\varGamma^{0}}\left[\mathcal{R}_{l,T}\left(s\right)\right],$ where now $\varGamma^{0}=\mathbf{B}_{1}\times\ldots\times\mathbf{B}_{m}$, with $\mathbf{B}_{i}$ a compact subset of $\left(0,\,1\right)$. The sets $\mathbf{B}_{i}$ are disjoint and satisfy $\sup_{\lambda\in\mathbf{B}_{i}}<\inf_{\lambda\in\mathbf{B}_{i+1}}$ for all $i.$
Assumption (ref) implies that $\xi_{l,i}^{0}\triangleq\xi\left(\lambda_{i}^{0}\right)$ is uniquely defined by $\Psi_{l}\left(\xi_{l,i}^{0}\right)\triangleq\inf_{s}\Psi_{l,i}\left(s\right)=\inf_{s}\int_{\mathbb{R}}l\left(s-u\right)\left(\exp\left(\mathscr{V}^{\left(i\right)}\left(u\right)\right)/\left(\int_{\mathbb{R}}\exp\left(\mathscr{V}^{\left(i\right)}\left(w\right)\right)dw\right)\right)du$. The GL estimator is defined as the minimizer of
The analysis is now in terms of the $m\times1$ local parameter $u$ with components $u_{i}=T\left\Vert \Delta_{T,i}\right\Vert ^{2}(\lambda_{i}-$ $\lambda_{i,T}^{0}\left(v\right))$, with $\lambda_{i,T}^{0}\left(v\right)=\lambda_{i,T}^{0}\left(\theta^{0}+v/r_{T}\right)$.
Theorem (ref)-(ref) extend corresponding results from Theorem (ref)-(ref), respectively, to multiple change-points. The fast rate of convergence implies that asymptotically the behavior of the GL estimator only matters in a small neighborhood of each $T_{i}^{0}$. Since each such neighborhood increases at rate $1/v_{T}$ while $T\rightarrow\infty$ at a faster rate, given the mixing conditions, these are asymptotically distinct and the limiting distribution is then similar to that in the single break case. This is the same argument underlying the analysis of bai/perron:98 and of ibragimov/has:81. The same comments as those in Section (ref) apply.
Turning to the general case of loss functions satisfying Assumption (ref), Theorem (ref) shows that the random quantity $T\left\Vert \delta_{T}\right\Vert ^{2}\left(\widehat{\lambda}_{i}^{\mathrm{GL}}-\lambda_{i}^{0}\right)$ is (first-order) asymptotically equivalent to the random variable $\xi_{l,i}^{0}$ determined by
A direct consequence of the results of this section is that statistical inference for the break dates $T_{i}^{0}$ $\left(i=1,\ldots,\,m\right)$ can be carried out using the same methods for the single break case.
This section shows that the GL-HPDR confidence sets are bet-proof. The betting framework and the notion of bet-proofness are useful to study the properties of frequentist inference in non-regular problems. The literature concluded that frequentist confidence sets may exhibit undesirable properties in non-regular problems {[}e.g., buehler:59, cornfield:69, cox:58, mueller/norest:16 pierce:73, robinson:77 and wallace:59{]}. For example, the confidence sets can be too short or empty with positive probability. This arises because frequentist procedures often have the property that, conditional on a sample point lying in some subset of the sample space, the conditional confidence level is less than the unconditional confidence level uniformly in the parameters.
We use the same betting framework as in buehler:59. Let $\mathbb{P}\left(\cdot|\,\lambda_{b}\right)$ denote the likelihood of the data $Y\in\mathcal{Y}$ conditional on $\lambda_{b}\in\varGamma^{0}$. Assume $\mathbb{P}\left(\cdot|\,\lambda_{b}\right)$ has density $p\left(\cdot|\,\lambda_{b}\right)$ with respect to a finite measure $\zeta.$ We define a $1-\alpha$ confidence set by a rejection probability rule $\varphi:\,\varGamma^{0}\times\mathcal{Y}\mapsto\left[0,\,1\right]$ satisfying $\int\left[1-\varphi\left(\lambda_{b},\,y\right)\right]p\left(y|\,\lambda_{b}\right)d\zeta\left(y\right)\geq1-\alpha$, with $\varphi\left(\lambda_{b},\,y\right)$ the probability that $\lambda_{b}$ is not included in the set when $y$ is observed. For any realization of the data $Y=y$, an inspector can choose to object to the confidence set $\varphi$. The inspector\textquoteright s objection $\widetilde{b}:\mathcal{Y}\mapsto\left[0,\,1\right]$ takes value 1 if there is an objection. Denote by $\mathbf{B}$ the set of all measurable strategies $\widetilde{b}$. When $\widetilde{b}=1$ the inspector receives 1 if $\varphi$ does not contain $\lambda_{b}$, and she loses $\alpha/\left(1-\alpha\right)$ otherwise. For a given parameter $\lambda_{b}$ and betting strategy $\widetilde{b}$, the inspector's expected loss is,
A confidence set $\varphi$ is said to be bet-proof at level $1-\alpha$ if for each $\widetilde{b}\in\mathbf{B}$, $L_{\alpha}\left(\varphi,\,\widetilde{b},\,\lambda_{b}\right)\geq0$ for some $\lambda_{b}\in\varGamma^{0}$. If there exists a strategy $\widetilde{b}$ such that $L_{\alpha}\left(\varphi,\,\widetilde{b},\,\lambda_{b}\right)<0$ for all $\lambda_{b}\in\varGamma^{0}$, then the inspector would be right on average and would make positive expected profits. Hence, such $\varphi$ would be an “unreasonable” confidence set. Without loss of substance, we restrict our attention to a change in the mean of a sequence of i.i.d. Gaussian variables. Let $y_{t}=\delta_{T}\mathbf{1}\left\{ t>T_{b}^{0}\right\} +e_{t},$ where $e_{t}\sim i.i.d.\,\mathrm{\mathscr{N}\left(0,\,1\right)}.$ The result below can also be shown to hold for fixed shifts $\delta_{T}=\delta^{0}$. For ease of exposition, we assume $\delta^{0}$ known. The general case leads to similar results, with more lengthy derivations without any gain in intuition.
Recall that $\varphi$ is such that the Quasi-posterior probability $p_{T}\left(\lambda_{b}|\,y\right)=p_{T}\left(\lambda_{b}\right)$ of excluding $\lambda_{b}$ is less than or equal to $\alpha,$
Part of the proof shows that the Quasi-posterior is asymptotically equivalent (in total variation distance) to the Bayesian posterior. Given the conservativeness allowed by Definition (ref), the GL confidence interval is asymptotically a superset of a Bayesian credible interval. Bet-proofness is a useful criterion in change-point models where popular inference methods face some difficulties, as shown in the next section. Proposition (ref) suggests that GL inference should not suffer from these issues; the simulations in the next section will confirm that this is indeed the case.
The purpose of this section is twofold. Section (ref) assesses the accuracy of the GL estimate of the change-point while Section (ref) evaluates the small-sample properties of the proposed method to construct confidence sets. We consider DGPs that take the form:
with a sample size $T=100.$ Three versions of (ref) are investigated: M1 involves a break in mean: $Z_{t}=1$, $D_{t}$ absent, and $e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right)$; M2 is similar to M1 but with $e_{t}=0.3e_{t-1}+u_{t}$, $u_{t}\sim\mathscr{N}\left(0,\,1\right)$; M3 is a dynamic model with $D_{t}=y_{t-1}$, $Z_{t}=1$, $e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,0.5\right)$ and $\alpha^{0}=0.6$. We set $\beta^{0}=1$ in M1-M2 and $\beta^{0}=0$ in M3. We consider $\lambda_{0}=0.3$ and 0.5, and break magnitudes $\delta^{0}=0.3,\,0.4,\,0.6$ and $1$. Additional simulations are presented in the supplement.
We consider the following estimators of $T_{b}^{0}$: the least-squares estimator (OLS), the GL estimator under a least-absolute loss function (GL-LN); the GL estimator under a least-absolute loss function with a uniform prior (GL-Uni). We compare the mean absolute error (MAE), standard deviation (Std), root-mean-squared error (RMSE), and the 25% and 75% quantiles. We set the trimming parameter $\epsilon$ equal to 0.05. As explained in casini/perron_Lap_CR_Single_Inf, the trimming $\epsilon$ should not be chosen too high because otherwise the estimate might tend to overestimate (resp. underestimate) the break date if it is in the first (resp. second) half of the sample. They found that $\epsilon=0.05$ performs well for different locations of the break date and this is also confirmed in the simulations in this section. See Section 1 and 5 in casini/perron_Lap_CR_Single_Inf for more discussion.
Tables (ref)-(ref) present the results. When the magnitude of the break is small, the OLS estimator displays quite large MAE, which increases as the change-point point moves toward the tails. In contrast, the GL estimator shows substantially lower MAE uniformly over break magnitudes and break locations. In addition, the GL estimator has smaller variance as well as lower RMSE compared to the OLS estimator. Notably, the distribution of GL-LN concentrates a higher fraction of the mass around the mid-sample relative to the finite-sample distribution of the OLS estimate. This is mainly due to the fact that the Quasi-posterior essentially does not share the marked trimodality of the finite-sample distribution {[}cf. casini/perron_CR_Single_Break{]}. When the break magnitude is small, the objective function is quite flat with a small peak at the OLS estimate. The Quasi-posterior has higher mass close to the OLS estimate\textemdash which corresponds to the middle mode\textemdash and accordingly lower mass in the tails. The GL estimator that uses the uniform prior (GL-Uni) is also more precise than the OLS estimator, though the margin is smaller. The latter is due to the fact that the GL estimate uses information only from the OLS objective function. We have not reported the bias. However, here is a summary of its behavior which can also be learned from Figures (ref)-(ref). When $\lambda_{b}^{0}=0.5$, the bias is small and close to zero because the finite-sample distributions of the estimators are symmetric. When $\lambda_{b}^{0}<0.5$, the bias is positive which means that the break date estimators tend to be on the right of $\lambda_{b}^{0}.$ The opposite hold for $\lambda_{b}^{0}>0.5$.
We now assess the performance of the suggested inference procedure for the break date. We compare it with the following existing methods: Bai's (1997) approach, elliott/mueller:07\textquoteright s (2007) approach based on\textcolor{red}{ }inverting a sequence of locally best invariant tests using nyblom:89's (1989) statistic, the inverted likelihood-ratio (ILR) method of eo/morley:15 which inverts the likelihood-ratio test of qu/perron:07 and the HDR method proposed in casini/perron_CR_Single_Break based on continuous record asymptotics, labelled OLS-CR. These methods have been discussed in detail in casini/perron_CR_Single_Break and in chang/perron:18. We can summarize their properties as follows. The confidence intervals obtained from Bai's (1997) method display empirical coverage rates often below the nominal level when the size of the break is small. In general, elliott/mueller:07\textquoteright s (2007) approach achieves the most accurate coverage rates but the average length of the confidence sets is always substantially larger relative to other methods.\footnote{This problem is more severe when the errors are serially correlated or the model includes lagged dependent variables.\nocite{casini_hac} Regarding the former, this in part may be due to issues with newey/west:87 HAC-type estimators when there are breaks {[}see casini_CR_Test_Inst_Forecast (2018, 2019), \nocite{casini_hac} casini/perron_Low_Frequency_Contam_Nonstat:2020, casini/perron_Oxford_Survey, chang/perron:18, crainiceanu/vogelsang:07, deng/perron:06, fossati:17, juhl/xiao:09, kim/perron:09, martins/perron:16, perron/yamamoto:18 and vogeslang:99{]}.} In addition, this approach breaks down in models with serially correlated errors or lagged dependent variables, whereby the length of the confidence set approaches the whole sample as the magnitude of the break increases. The ILR has coverage rates often above the nominal level and an average length significantly longer than with the OLS-CR method when the magnitude of the shift is small. Here, we shall show that the GL inference performs well in terms of coverage probability compared with the other methods and is characterized by shorter lengths of the confidence sets.
When the errors are uncorrelated (i.e., M1 and M3) we simply estimate variances rather than long-run variances. The least-squares estimation method is employed with a trimming parameter $\epsilon=0.15$ and we use the required degrees of freedom adjustment for the statistic $\widehat{\textrm{U}}_{T}$ of elliott/mueller:07. To construct the OLS-CR method, we follow the steps outlined in casini/perron_CR_Single_Break. To implement Bai's (1997) method we use the usual steps described in bai:97RES and bai/perron:98. We implement the GL estimator using a least-absolute loss with the prior from Corollary (ref). For model M2, the estimate of the long-run variance is the pre-whitened heteroskedasticity and autocorrelation (HAC) estimator of andrews/monahan:92. We consider the version $\widehat{\textrm{U}}_{T}$ proposed by elliott/mueller:07 that allows for heterogeneity across regimes; using the restricted version when applicable leads to similar results. Finally, the last row of each panel includes the rejection probability of the 5%-level sup-Wald test using the asymptotic critical value of andrews:93; it serves as a statistical measure of the magnitude of the break.
Overall, the results in Table (ref)-(ref) confirm previous findings about the performance of existing methods. Bai's (1997) method has a coverage rate below the nominal level when the size of the break is small. For example, in model M2, with $\lambda_{0}=0.5$ and $\delta^{0}=0.8$, it has a coverage probability below 82% even though the Sup-Wald test rejects roughly for 70% of the samples. With smaller break sizes, it systematically fails to cover the true break date with correct probability. In contrast, the method of elliott/mueller:07 yields very accurate empirical coverage rates. However, the average length of the confidence intervals obtained is systematically much larger than those from all other methods across all DGPs, break sizes and break locations. For large break sizes, Bai's (1997) method delivers good coverage rates and the shortest average length among all methods.
The GL method displays good coverage rates across different break magnitudes and tends to have the shortest lengths among all methods for all break magnitudes, except for $\delta^{0}=1.6$ in model M2 for which Bai's (1997) confidence interval is slightly shorter. In Model M3, the coverage rates of OLS-CR are more accurate than those with the GL method although the difference is not large. Thus the GL method strikes a good balance between adequate coverage probability and short average lengths, thus confirming the theoretical results on bet-proofness. This is also consistent with Figures (ref)-(ref) which show that the asymptotic distribution of the GL estimator does not underestimate uncertainty about the break location even when the break magnitude is small thereby yielding good coverage rates also in this case. In model M1 the GL method leads to shorter lengths than Bai's even for large breaks. This is not in contradiction with Figure (ref) because model M1 is a simpler model than that reported in the figure which shows that the density of the asymptotic distribution of the GL estimator is more spread out than that from bai:97RES.
Non-reported simulations show that the GL method is robust to heteroskedastic errors $e_{t}=\left|z_{t}\right|u_{t}$ and non-normal errors. The case of multiple breaks is not considered since they are expected to be similar as in the single break case by virtue of the assumption that the break dates are sufficiently separated. Finally, in the supplement we compare the GL method above with its continuous record counterparts developed in casini/perron_Lap_CR_Single_Inf. Overall, we find that both estimation and confidence intervals based on GL-LN perform well relative to the continuous record counterparts, where significant gains appear to occur when there is high serial correlation in the errors. See the additional results reported in the supplement.
We developed large-sample results for a class of Generalized Laplace estimators in multiple change-points models where popular methods face some challenges due to the non-regularities of the problem. The GL method exploits the insight of Laplace who proposed to generate a density from taking an exponential transformation of a least-squares criterion. The class of GL estimators exhibits a dual limiting distribution; namely, the classical shrinkage asymptotic distribution of bai/perron:98, or a Bayes-type asymptotic distribution {[}cf. ibragimov/has:81{]}. Simulations show that the GL estimator is more accurate than OLS. Similarly, inference has superior finite-sample properties relative to popular methods and these properties are shown to be supported by theoretical results. Since the issues about the finite-sample performance of OLS especially for small breaks continue to hold in more complex structural change models, we believe that our method can be usefully extended to those models. For example, the GL approach can be immediately applied to nonlinear models (e.g., instrumental variable models, linear model with restrictions, nonlinear regression models, etc.) even though particular attention to the appropriate choice of the prior should be given in each context. We believe that our approach can also be relevant for high-dimensional regression with structural changes although this would require a careful consideration of additional aspects related to the growing number of regressors.
Casini, A. and P. Perron (2020c). Supplement to \textquotedblleft Generalized Laplace Inference in Multiple Change-points Models\textquotedblright , Econometric Theory Supplementary Material. To view, please visit: {[}{[}doi will be inserted here by typesetter{]}{]}\nocite{casini/perron_CR_Single_Break,casini/perron_Lap_CR_Single_Inf,casini/perron_Low_Frequency_Contam_Nonstat:2020,casini/perron_PrewhitedHAC,casini/perron_SC_BP_Lap,casini_diss,casini_hac}
\setcounter{page}{1}
\setcounter{page}{1}
\addcontentsline{toc}{part}{Supplemental Material}
\thispagestyle{empty}
\setcounter{page}{1} \setcounter{section}{1}