EconBase
← Back to paper

Generalized Laplace Inference in Multiple Change-Points Models

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

123,036 characters

Generalized Laplace Inference in Multiple Change-Points Models


\setcounter{page}{0}
\title{\textbf{Generalized Laplace Inference in Multiple Change-Points Models}\thanks{This paper is based on the fourth chapter of the first author's doctoral
dissertation at Boston University. We thank the Editor and a Co-Editor
for guiding the review, and three anonymous referees for constructive
comments. We also thank Zhongjun Qu for useful comments. }\textbf{ }}
\maketitle
\begin{abstract}
{\footnotesize{}Under the classical long-span asymptotic framework
we develop a class of Generalized Laplace (GL) inference methods for
the change-point dates in a linear time series regression model with
multiple structural changes analyzed in, e.g., \citet{bai/perron:98}.
The GL estimator is defined by an integration rather than optimization-based
method and relies on the least-squares criterion function. It is interpreted
as a classical (non-Bayesian) estimator and the inference methods
proposed retain a frequentist interpretation. This approach provides
a better approximation about the uncertainty in the data of the change-points
relative to existing methods. On the theoretical side, depending on
some input (smoothing) parameter, the class of GL estimators exhibits
a dual limiting distribution; namely, the classical shrinkage asymptotic
distribution, or a Bayes-type asymptotic distribution. We propose
an inference method based on Highest Density Regions using the latter
distribution. We show that it has attractive theoretical properties
not shared by the other popular alternatives, i.e., it is bet-proof.
Simulations confirm that these theoretical properties translate to
good finite-sample performance.}{\footnotesize\par}
\end{abstract}
\indent {\bf{JEL Classification}}: C12, C13, C22\\
\noindent {\bf{Keywords}}: Asymptotic Distribution, Bet-Proof, Break Date, Change-point,   Generalized Laplace Inference, Highest Density Region, Quasi-Bayes.

\onehalfspacing
\thispagestyle{empty}
\allowdisplaybreaks

\vfill{}
\pagebreak{}

\section{\label{Section Introduction Lap_BP}Introduction}

In the context of the multiple change-points model analyzed in \citet{bai/perron:98},
we develop inference methods for the change-point dates for a class
of Generalized Laplace (GL) estimators using a classical long-span
asymptotic framework. They are defined by an integration rather than
an optimization-based method, the latter typically characterizing
classical extremum estimators. The idea traces back to \citet{laplace:74},
who first suggested to interpret transformations of a least-squares
criterion function as a statistical belief over a parameter of interest.
Hence, a Laplace estimator is defined similarly to a Bayesian estimator
although the former relies on a statistical criterion function rather
than a parametric likelihood function. As a consequence, the GL estimator
is interpreted as a classical (non-Bayesian) estimator and the inference
methods proposed retain a frequentist interpretation such that the
GL estimators are constructed as a function of integral transformations
of the least-squares criterion. In a first step, we use the approach
of \citet{bai/perron:98} to evaluate the least-squares criterion
function at all candidate break dates. We then apply a transformation
to obtain a proper distribution over the parameters of interest,
referred to as the Quasi-posterior. For a given choice of a loss function
and (possibly) a prior density, the estimator is then defined either
explicitly as, for example, the mean or median of the (weighted) Quasi-posterior
or implicitly as the minimizer of a smooth convex optimization problem.

The underlying asymptotic framework considered  is the long-span
shrinkage asymptotics of \citet{bai:97RES}, \citet{bai/perron:98}
and also \citet{perron/qu:06} who considerably relaxed some conditions,
where the magnitude of the parameter shift is sample-size dependent
and approaches zero as the sample size increases. Early contributions
to this approach are \citet{hinkley:71}, \citet{bhattacharya:87},
and \citet{yao:87} for estimating break points. For testing for structural
breaks, see \citet{hawkins:77}, \citet{picard:85}, \citet{kim/siegmund:89},
\citet{andrews:93}, \citet{horvath:93} and \citet{andrews/ploberger:94}.
See also the reviews of \citet{csorgo/horvath:97}, \citet{perron:06},
\citet{casini/perron_Oxford_Survey} and references therein.

One of our goals is to develop GL estimates with better small-sample
properties compared to least-squares estimates, namely lower Mean
Absolute and Root-Mean Squared Errors, and confidence sets with accurate
coverage probabilities and relatively short lengths for a wide range
of break sizes, whether small or large; existing methods work well
for either small or large breaks, but not for both. A second goal
is to establish theoretical results that support the reported finite-sample
properties about inference.

The asymptotic distribution of the GL estimator is derived via a local
parameter related to a normalized deviation from the true fractional
break date. The normalization factor corresponds to the rate of convergence
of the original (extremum) least-squares estimator as established
by \citet{bai/perron:98}.  The asymptotic distribution of the GL
estimator then depends on a sample-size dependent smoothing parameter
sequence applied to the least-squares criterion function. We derive
two distinct limiting distributions corresponding to different smoothing
sequences of the criterion function {[}cf. \citet{jun/pinkse/wan:15}
for a related application in the context of the cube-root asymptotics
of \citet{kim/pollard:90}{]}. In one case, the estimator displays
the same limit law as the asymptotic distribution of the least-squares
estimator derived in \citet{bai/perron:98} {[}see also \citet{hinkley:71},
\citet{picard:85} and \citet{yao:87}{]}. In a second case, the
limiting distribution is characterized by a ratio of integrals over
functions of Gaussian processes and resembles the limiting distribution
of Bayesian change-point estimators. The latter is exploited for
the purpose of constructing confidence sets for the break dates.
We use the concept of highest density regions (HDR) introduced by
\citet{casini/perron_CR_Single_Break} for structural change problems,
which best summarizes the properties of the probability distribution
of interest. The HDR are common in Bayesian analysis where they are
applied to a posterior distribution {[}see, e.g., \citeauthor{box/tiao:1973}{]}.
\citeauthor{kendall/stuart:1961} discussed the difference between
frequentist confidence intervals and Bayesian approaches in relation
to the existence of a sufficient statistic. Our procedure yields confidence
sets for the break date which, in finite samples, better account for
the uncertainty over the parameter space in finite-samples because
it effectively incorporates a statistical measure of the uncertainty
in the least-squares criterion function. As noted in the literature
on likelihood-based inference in some classes of nongranular problems
{[}see e.g., \citet{chernozhukov/hong:03}, \citet{ghosal/ghosh/samanta:95},
\citet{hirano/porter:03} and \citet{ibragimov/has:81}{]}, the Maximum
Likelihood Estimator (MLE) is generally not an asymptotically sufficient
statistic in these models and so the likelihood contains more information
asymptotically than the MLE. Hence, likelihood-based procedures are
generally not functions of the MLE even asymptotically. This incompleteness
property motivated the study of the entire likelihood rather than
just the MLE. Likewise, our method exploits the entire behavior of
the objective function.

Laplace's seminal insight has been applied successfully in many disciplines.
In econometrics, \citet{chernozhukov/hong:03} introduced Laplace-type
estimators as an alternative to classical (regular) extremum estimators
in several problems such as censored median regression and  nonlinear
instrumental variable; see also \citet{forneron/ng:17} for a review
and comparisons. Their main motivation was to solve the curse of dimensionality
inherent to the computation of such estimators. In contrast, the class
of GL estimators in structural change models serves distinct multiple
purposes. First,  inference about the break dates  presents several
challenges, in particular to provide methods with a satisfactory performance
uniformly over different data-generating mechanisms and break magnitudes.
The GL inference proves to be reliable and accurate in finite-samples.
Second, it leads to inference methods that have both frequentist and
credibility properties which is not shared by the other popular methods.

Turning to the problem of constructing confidence sets for a single
break date, the standard asymptotic method for the linear regression
model was proposed in \citet{bai:97RES}, while \citet{elliott/mueller:07}
proposed to invert the locally best invariant test of \citet{nyblom:89},
and \citet{eo/morley:15} suggested to invert the likelihood-ratio
statistic of \citet{qu/perron:07}. The latter were mainly motivated
by finite-sample results indicating that the exact coverage rates
of the confidence intervals obtained from Bai's (1997) method are
often below the nominal level when the magnitude of the break is small.
It has been shown that the method of \citet{elliott/mueller:07}  delivers
the most accurate coverage rates but the average length of the confidence
sets is significantly larger than with other methods. The confidence
sets for the break dates constructed from the GL inference that we
develop result in exact coverage rates close to the nominal level
and short length of the confidence sets. This holds true whether the
magnitude of the break is small or large. In fact, we show that
GL inference is bet-proof, a measure of ``reasonableness'' of frequentist
inference in non-regular problems {[}see, e.g., \citet{buehler:59}{]}.

The GL inference developed in this paper has been applied by \citet{casini/perron_Lap_CR_Single_Inf}
to achieve finite-sample improvements under the continuous record
asymptotic framework of \citet{casini/perron_CR_Single_Break}. The
latter proposed an alternative asymptotic framework to explain the
non-standard features of the finite-sample distribution of the least-squares
estimator.

The paper is organized as follows. We first focus on the single change-point
case. Section \ref{Section, Model and its Assumptions, Bai Lap} presents
the statistical setting. We develop the asymptotic theory in Section
\ref{Section Asymptotic Results LapBai97} and the inference methods
in Section \ref{Section Inference Methods}. Results for multiple
change-points models are given in Section \ref{Section Models with Multiple Breaks}
while Section \ref{Section Theoretical-Properties-of GL Inference}
discusses some theoretical properties of GL inference. Section \ref{Section Monte-Carlo-Simulation}
presents simulation results about the finite-sample performance. Section
\ref{Section Conclusions} concludes. All proofs are included in an
online supplement {[}\citet{casini/perron_SC_BP_Lap_Supp}{]}.

\section{\label{Section, Model and its Assumptions, Bai Lap}The Model and
the Assumptions}

This section introduces the structural change model with a single
break, reviews the least-squares estimation method for the break date,
and presents the relevant assumptions. We start with introducing the
formal setup for our analysis. The following notation is used throughout.
We denote the transpose of a matrix $A$ by $A'$.  We use $\left\Vert \cdot\right\Vert $
to denote the Euclidean norm of a linear space, i.e., $\left\Vert x\right\Vert =\left(\sum_{i=1}^{p}x_{i}^{2}\right)^{1/2}$
for $x\in\mathbb{R}^{p}.$ For a matrix $A$, we use the vector-induced
norm, i.e., $\left\Vert A\right\Vert =\sup_{x\neq0}\left\Vert Ax\right\Vert /\left\Vert x\right\Vert .$
All vectors are column vectors. For two vectors $a$ and $b$, we
write $a\leq b$ if the inequality holds component-wise. We use $\left\lfloor \cdot\right\rfloor $
to denote the largest smaller integer function. We use $\overset{\mathbb{P}}{\rightarrow}$
and $\overset{d}{\rightarrow}$ to denote convergence in probability
and convergence in distribution, respectively. $\mathbb{C}_{b}\left(\mathbf{E}\right)$
{[}$\mathbb{D}_{b}\left(\mathbf{E}\right)${]} is the collection of
bounded continuous {[}c�dl�g{]} functions from some specified set
$\mathbf{E}$ to $\mathbb{R}$. Weak convergence on either $\mathbb{C}_{b}\left(\mathbf{E}\right)$
or $\mathbb{D}_{b}\left(\mathbf{E}\right)$ is denoted by $\Rightarrow$.
The symbol ``$\triangleq$'' stands for definitional equivalence.

We consider a sample of observations $\left\{ \left(y_{t},\,w_{t},\,z_{t}\right):\,t=1,\ldots,\,T\right\} ,$
defined on a filtered probability space $\left(\Omega,\,\mathscr{F},\,\mathbb{P}\right)$,
on which all of the random elements introduced in what follows are
defined. The model is
\begin{align}
\begin{split}y_{t} & =w_{t}'\phi^{0}+z_{t}'\delta_{1}^{0}+e_{t},\quad\left(t=1,\ldots,\,T_{b}^{0}\right)\qquad y_{t}=w_{t}'\phi^{0}+z_{t}'\delta_{2}^{0}+e_{t},\quad\left(t=T_{b}^{0}+1,\ldots,\,T\right)\end{split}
\label{Eq. the Single Break Model}
\end{align}
 where $y_{t}$ is a scalar dependent variable, $w_{t}$ and $z_{t}$
are regressors of dimensions, $p$ and $q,$ respectively, and $e_{t}$
is an unobserved error term. The true parameter vectors $\phi^{0},\,\delta_{1}^{0}$
and $\delta_{2}^{0}$ are unknown and we define $\delta^{0}\triangleq\delta_{2}^{0}-\delta_{1}^{0},$
with $\delta^{0}\neq0$ so that a structural change occurs at date
$T_{b}^{0}$. It is useful to re-parametrize the model. Letting $x_{t}\triangleq\left(w'_{t},\,z'_{t}\right)'$
and $\beta^{0}\triangleq\left(\left(\phi^{0}\right)',\,\left(\delta_{1}^{0}\right)'\right)',$
we have
\begin{align}
\begin{split}y_{t} & =x_{t}'\beta^{0}+e_{t},\quad\left(t=1,\ldots,\,T_{b}^{0}\right)\qquad y_{t}=x_{t}'\beta^{0}+z_{t}'\delta^{0}+e_{t},\quad\left(t=T_{b}^{0}+1,\ldots,\,T\right).\end{split}
\label{Eq. Reparameterize SC model}
\end{align}
 More generally, we can define $z_{t}\triangleq D'x_{t}$, where $D$
is a $\left(p+q\right)\times q$ matrix with full column rank. A pure
structural change model in which all regression parameters are subject
to change corresponds to $D=I_{\left(p+q\right)\times\left(p+q\right)}$,
whereas a partial structural change model arises when $D=\left(0_{q\times p},\,I_{q\times q}\right)'.$
In order to facilitate the derivations, we reformulate model \eqref{Eq. Reparameterize SC model}
in matrix format. Let $Y=\left(y_{1},\,\ldots,\,y_{T}\right)',\,X=\left(x_{1},\,\ldots,\,x_{T}\right)'$,
$e=\left(e_{1},\,\ldots,\,e_{T}\right)',$ $X_{1}=\left(x_{1},\,\ldots,\,x_{T_{b}},\,0,\,\ldots,\,0\right)'$,
$X_{2}=(0,\,\ldots,\,0,\,$ $x_{T_{b}+1},\ldots,\,x_{T})'$ and $X_{0}=(0,\,\ldots,\,0,\,x_{T_{b}^{0}+1},\ldots,\,x_{T})'$.
Further, define $Z_{1},\,Z_{2}$ and $Z_{0}$ in a similar way: $Z_{1}=X_{1}D,\,Z_{2}=X_{2}D$
and $Z_{0}=X_{0}D$. We omit the dependence of the matrices $X_{i}$
and $Z_{i}$ ($i=1,\,2$) on $T_{b}$. Then, \eqref{Eq. Reparameterize SC model}
is equivalent to
\begin{align}
Y & =X\beta+Z_{0}\delta+e.\label{Eq Matrix Format of the model}
\end{align}

Let $\theta^{0}\triangleq\left(\left(\phi^{0}\right)',\,\left(\delta_{1}^{0}\right)',\,\left(\delta^{0}\right)'\right)'$
denote the true value of the parameter vector $\theta\triangleq\left(\phi,\,\delta_{1},\,\delta\right).$
The break date least-squares (LS) estimator $\widehat{T}_{b}^{\mathrm{LS}}$
is the minimizer of the sum of squared residuals {[}denoted $S_{T}\left(\theta,\,T_{b}\right)${]}
from \eqref{Eq Matrix Format of the model}. The parameter $\theta$
can be concentrated out resulting in a criterion function depending
only on $T_{b}=T\lambda_{b}$, i.e.,~$\widehat{T}_{b}^{\mathrm{LS}}=\arg\min_{1\leq T_{b}\leq T}S_{T}(\widehat{\theta}^{\mathrm{LS}}(T_{b}),\,T_{b})$
where $\widehat{\theta}^{\mathrm{LS}}(T_{b})=\arg\min_{\theta}S_{T}(\theta,\,T_{b})$
with $S_{T}(\theta,\,T_{b})=\sum_{t=1}^{T_{b}}\left(y_{t}-\phi'w_{t}-\delta'_{1}z_{t}\right)^{2}+\sum_{t=T_{b}+1}^{T}\left(y_{t}-\phi'w_{t}-\delta'z_{t}\right)^{2}$.
Also,
\begin{align}
\arg\min_{1\leq T_{b}\leq T}S_{T}(\widehat{\theta}^{\mathrm{LS}}\left(T_{b}\right),\,T_{b}) & =\arg\max_{T_{b}}\widehat{\delta}^{\mathrm{LS\prime}}(T_{b})(Z_{2}'M_{X}Z_{2})\widehat{\delta}^{\mathrm{LS}}(T_{b})\label{Eq. Relationship Sup Wald and SSR}\\
 & \triangleq\arg\max_{\lambda_{b}}Q_{T}(\widehat{\delta}^{\mathrm{LS}}(\lambda_{b}),\,\lambda_{b}),\nonumber
\end{align}
 where $M_{X}\triangleq I-X\left(X'X\right)^{-1}X'$, $\widehat{\delta}^{\mathrm{LS}}\left(\lambda_{b}\right)$
is the least-squares estimator of $\delta^{0}$ obtained by regressing
$Y$ on $X$ and $Z_{2}$ and the statistic $Q_{T}\left(\widehat{\delta}^{\mathrm{LS}}\left(\lambda_{b}\right),\,\lambda_{b}\right)$
is the numerator of the sup-Wald statistic. The Laplace-type inference
builds on the least-squares criterion function $Q_{T}\left(\delta\left(\lambda_{b}\right),\,\lambda_{b}\right),$
where $\delta\left(\lambda_{b}\right)$ stands for $\widehat{\delta}^{\mathrm{LS}}\left(\lambda_{b}\right)$
to minimize notational burden.
\begin{assumption}
\label{Assumption A Bai97}$T_{b}^{0}=\left\lfloor T\lambda_{b}^{0}\right\rfloor ,$
where $\lambda_{b}^{0}\in\varGamma^{0}\subset\left(0,\,1\right).$
\end{assumption}

\begin{assumption}
\label{Assumption A5 Bai97 and A4 BP98}With $\left\{ \mathscr{F}_{t},\,t=1,\,2,\ldots\right\} $
a sequence of increasing $\sigma$-fields, $\left\{ z_{t}e_{t},\,\mathscr{F}_{t}\right\} $
forms an $L^{r}$-mixingale sequence with $r=2+\nu$ for some $\nu>0$.
That is, there exist nonnegative constants $\left\{ \varrho_{1,t}\right\} _{t\geq1}$
and $\left\{ \varrho_{2,j}\right\} _{j\geq0}$ such that $\varrho_{2,j}\rightarrow0$
as $j\rightarrow\infty$, and for all $t\geq1$, $j\geq0$ and $r\geq1$,
(i) $\left\Vert \mathbb{E}\left(z_{t}e_{t}|\,\mathscr{F}_{t-j}\right)\right\Vert _{r}\leq\varrho_{1,t}\varrho_{2,j}$,
(ii) $\left\Vert z_{t}e_{t}-\mathbb{E}\left(z_{t}e_{t}|\,\mathscr{F}_{t+j}\right)\right\Vert _{r}\leq\varrho_{1,t}\varrho_{2,j+1}$.
In addition, (iii) $\max_{t}\varrho_{1,t}<C_{1}<\infty$ and (iv)
$\sum_{j=0}^{\infty}j^{1+\nu}\varrho_{2,j}<\infty$ for some $\nu>0$,
(v) $\left\Vert z_{t}\right\Vert _{2r}<C_{2}<\infty$ and $\left\Vert e_{t}\right\Vert _{2r}<C_{3}<\infty$
for some $C_{1},\,C_{2},\,C_{3}>0$.
\end{assumption}

\begin{assumption}
\label{Assumption A3 Bai97}There exists an $l_{0}>0$ such that for
all $l>l_{0},$ the minimum eigenvalues of $H_{l}^{*}=\left(1/l\right)\sum_{T_{b}^{0}-l+1}^{T_{b}^{0}}x_{t}x'_{t}$
and $H_{l}^{**}=\left(1/l\right)\sum_{T_{b}^{0}+1}^{T_{b}^{0}+l}x_{t}x'_{t}$
are bounded away from zero. These matrices are invertible when $l\geq p+q$
and have stochastically bounded norms uniformly in $l$.
\end{assumption}

\begin{assumption}
\label{Assumption A4 Bai97}$T^{-1}X'X\overset{\mathbb{P}}{\rightarrow}\Sigma_{XX},$
where $\Sigma_{XX}$, a positive definite matrix.
\end{assumption}
These assumptions are standard and similar to those in \citet{perron/qu:06}.
It is well-known that only the fractional break date $\lambda_{b}^{0}$
(not $T_{b}^{0}$) can be consistently estimated, with $\widehat{\lambda}_{b}^{\mathrm{LS}}$
having a $T$-rate of convergence. The corresponding result for
the break date estimator $\widehat{T}_{b}^{\mathrm{LS}}$ states
that, as $T$ increases, $\widehat{T}_{b}^{\mathrm{LS}}$ remains
within a bounded distance from $T_{b}^{0}$. However, this does not
affect the estimation problem of the regression coefficients $\theta^{0}$,
for which $\widehat{\theta}^{\mathrm{LS}}$ is a regular estimator;
i.e., $\sqrt{T}$-consistent and asymptotically normally distributed,
since the estimation of the regression parameters is asymptotically
independent from the estimation of the change-point. Hence, the regression
parameters are essentially estimated as if the change-point was known.
More complex is the derivation of the asymptotic distribution of $\widehat{\lambda}_{b}^{\mathrm{LS}}$;
e.g., \citet{hinkley:71} for an i.i.d. Gaussian process with a mean
change. Therefore, to make progress it is necessary to consider a
shrinkage asymptotic setting in which the size of the shift converges
to zero as $T\rightarrow\infty$; see \citet{picard:85} and \citet{yao:87}
and extended by \citet{bai:97RES} to general linear models.

\section{\label{Section Asymptotic Results LapBai97}Generalized Laplace Estimation}

We define the GL estimator in Section \ref{subsec:The-Setting-and}
and discuss its usefulness in Section \ref{subsection Discussion-on-GL}.
Section \ref{subsec:Normalized-Version-of} describes the asymptotic
framework under which we derive the limiting distribution with the
results presented in Section \ref{Sub section Theorems Laplace in SC}.


\subsection{\label{subsec:The-Setting-and}The Class of Laplace Estimators}

The class of GL estimators relies on the original least-squares criterion
function $Q_{T}\left(\delta\left(\lambda_{b}\right),\,\lambda_{b}\right)$,
with the parameter of interest being $\lambda_{b}^{0}=T_{b}^{0}/T$.
The Quasi-posterior $p_{T}\left(\lambda_{b}\right)$ is defined by
the exponential transformation,
\begin{align}
p_{T}\left(\lambda_{b}\right) & \triangleq\frac{\exp\left(Q_{T}\left(\delta\left(\lambda_{b}\right),\,\lambda_{b}\right)\right)\pi\left(\lambda_{b}\right)}{\int_{\varGamma^{0}}\exp\left(Q_{T}\left(\delta\left(\lambda_{b}\right),\,\lambda_{b}\right)\right)\pi\left(\lambda_{b}\right)d\lambda_{b}},\label{Eq. Quasi-posterior-1}
\end{align}
where $\pi\left(\cdot\right)$ is a density function. Note that $p_{T}\left(\lambda_{b}\right)$
defines a proper distribution over the parameter space $\varGamma^{0}$.
The $\mathscr{\mathscr{L}}\left(\theta,\,T_{b}\right)$-class of estimators
are the solutions of smooth convex optimization problems for a given
loss function, restricting attention to convex loss functions $l_{T}\left(\cdot\right)$.
Examples include (a) $l_{T}\left(r\right)=a_{T}^{m}\left|r\right|^{m},$
the polynomial loss function (the squared loss function is obtained
when $m=2$ and the absolute deviation loss function when $m=1$);
(b) $l_{T}\left(r\right)=a_{T}\left(\tau-\mathbf{1}\left(r\leq0\right)\right)r,$
the check loss function; where $a_{T}$ is a divergent sequence. We
define the Expected Risk function, under the density $p_{T}\left(\cdot\right)$
and the loss $l_{T}\left(\cdot\right)$ as $\mathcal{R}_{l,T}\left(s\right)\triangleq\mathbb{E}_{p_{T}}\left[l_{T}\left(s-\widetilde{\lambda}_{b}\right)\right],$
where $\widetilde{\lambda}_{b}$ is a random variable with distribution
$p_{T}$ and $\mathbb{E}_{p_{T}}$ denotes expectation taken under
$p_{T}.$ Using \eqref{Eq. Quasi-posterior-1} we have,
\begin{align}
\mathcal{R}_{l,T}\left(s\right) & \triangleq\int_{\varGamma^{0}}l_{T}\left(s-\lambda_{b}\right)p_{T}\left(\lambda_{b}\right)d\lambda_{b}.\label{Eq. Expected Risk function-1}
\end{align}
 The Laplace-type estimator $\widehat{\lambda}_{b}^{\mathrm{GL}}$
shall be interpreted as a decision rule that, given the information
contained in the Quasi-posterior $p_{T}$, is least unfavorable according
to the loss function $l_{T}$ and the prior density $\pi$. Then $\widehat{\lambda}_{b}^{\mathrm{GL}}$
is the minimizer of the expected risk function \eqref{Eq. Expected Risk function-1},
i.e., $\widehat{\lambda}_{b}^{\mathrm{GL}}\triangleq\arg\min_{s\in\varGamma^{0}}\left[\mathcal{R}_{l,T}\left(s\right)\right].$
Observe that the GL estimator $\widehat{\lambda}_{b}^{\mathrm{GL}}$
results in the mean (median) of the Quasi-posterior upon choosing
 the squared (absolute deviation) loss function. The choice of the
loss and of the prior density functions hinges on the statistical
problem addressed. In the structural change problem, a natural choice
for the Quasi-prior $\pi$ is the density of the asymptotic distribution
of $\widehat{\lambda}_{b}^{\mathrm{LS}}$. This requires to replace
the population quantities appearing in that distribution by consistent
plug-in estimates\textemdash cf. \citet{bai/perron:98}\textemdash and
derive its density via simulations as in \citet{casini/perron_CR_Single_Break}.
The attractiveness of the Quasi-posterior \eqref{Eq. Quasi-posterior-1}
is that it provides additional information about the parameter of
interest $\lambda_{b}^{0}$ beyond what is already included in the
point estimate $\widehat{\lambda}_{b}^{\mathrm{LS}}$ and its distribution
(see Section \ref{subsection Discussion-on-GL}). This approach will
result in more accurate inference in finite-samples even in cases
with high uncertainty in the data as we shall document in Section
\ref{Section Monte-Carlo-Simulation}. This is supported in Section
\ref{Section Theoretical-Properties-of GL Inference} showing that
the GL inference is bet-proof which is a desirable theoretical property
in non-regular problems.
\begin{assumption}
\label{Assumption The-loss-function LapBai97}Let $l_{T}\left(r\right)\triangleq l\left(a_{T}r\right)$,
with $a_{T}$ a positive divergent sequence. $\boldsymbol{L}$ denotes
the set of functions $l:\,\mathbb{R}\rightarrow\mathbb{R}_{+}$ that
satisfy (i) $l\left(r\right)$ is defined on $\mathbb{R}$, with $l\left(r\right)\geq0$
and $l\left(r\right)=0$ if and only if $r=0$; (ii) $l\left(r\right)$
is continuous at $r=0$; (iii) $l\left(\cdot\right)$ is convex and
$l\left(r\right)\leq1+\left|r\right|^{m}$ for some $m>0$.
\end{assumption}

\begin{assumption}
\label{Assumption Prior LapBai97}$\pi:\,\mathbb{R}\rightarrow\mathbb{R}_{+}$
is a continuous, uniformly positive density function satisfying $\pi^{0}\triangleq\pi\left(\lambda_{b}^{0}\right)>0,$
and for some finite $C_{\pi}<\infty,$ $\pi^{0}<C_{\pi}$. Also, $\pi\left(\lambda_{b}\right)=0$
for all $\lambda_{b}\notin\varGamma^{0}$, and $\pi$ is twice continuously
differentiable with respect to $\lambda_{b}$ at $\lambda_{b}^{0}$.
\end{assumption}

Assumption \ref{Assumption The-loss-function LapBai97} is similar
to those in \citet{bickel/yahav:69}, \citet{ibragimov/has:81} and
\citet{chernozhukov/hong:03}. The convexity assumption on $l_{T}\left(\cdot\right)$
is guided by practical considerations. The dominant restriction in
part (iii) is conventional and implicitly assumes that the loss function
has been scaled by some constant. What is important is that the growth
of the function $l_{T}\left(r\right)$ as $\left|r\right|\rightarrow\infty$
is slower than $\exp\left(\epsilon\left|r\right|\right)$ for any
$\epsilon>0$. Assumption \ref{Assumption Prior LapBai97} on the
prior is satisfied for any reasonable choice. For priors that have
a peak at $\lambda_{b}^{0}$ one can apply some basic smoothing techniques
to make it differentiable locally {[}e.g., mean smoothing, Gaussian
smoothing and Savitzky-Golay filter{]}. We did not find any particular
difference in the empirical results and so we used the mean smoothing.
The assumption on the differentiability of the kernel can be relaxed
at the expense of one more step in the proof. \citet{chernozhukov/hong:03}
assumed differentiability of the prior; we also keep the same assumption
and applied the smoothing. The large-sample properties of the $\mathscr{\mathscr{L}}\left(\theta,\,T_{b}\right)$-class
are studied under the shrinkage asymptotic setting of \citet{bai:97RES}
and \citet{bai/perron:98}. Thus, we need the following assumption.
\begin{assumption}
\label{Assumption Small Shift BP}Let $\delta_{T}\triangleq\delta_{T}^{0}\triangleq v_{T}\delta^{0}$
where $v_{T}>0$ is a scalar satisfying $v_{T}\rightarrow0$ as $T\rightarrow\infty$
and $T^{1/2-\vartheta}v_{T}\rightarrow\infty$ for some $\vartheta\in\left(0,\,1/4\right)$.
\end{assumption}
We omit the superscript 0 from $\delta_{T}^{0}$ for notational convenience
since it should not cause any confusion. Assumption \ref{Assumption Small Shift BP}
requires the magnitude of the break to shrink to zero at any slower
rate than $T^{-1/2}$. The specific rates allowed differ from those
in \citet{bai:97RES} and \citet{bai/perron:98}, since they require
$\vartheta\in\left(0,\,1/2\right)$. The reason is merely technical;
the asymptotics of the Laplace-type estimator involve smoothing the
criterion function, and thus one needs to guarantee that $\widehat{\lambda}_{b}$
approaches $\lambda_{b}^{0}$ at a sufficiently fast rate. Under the
shrinkage asymptotics, Proposition 1 and Corollary 1 in \citet{bai:97RES}
state that $T\left\Vert \delta_{T}\right\Vert ^{2}\left(\widehat{\lambda}_{b}^{\mathrm{LS}}-\lambda_{b}^{0}\right)=O_{\mathbb{P}}\left(1\right)$
and $\widehat{\delta}_{T}^{\mathrm{LS}}-\delta_{T}=o_{\mathbb{P}}\left(1\right)$.

\subsection{\label{subsection Discussion-on-GL}Discussion about the GL Approach}

We use Figure \ref{Fig1}-\ref{Fig2} to illustrate the main idea
behind the usefulness of the GL method. They present plots of the
density of the distribution of $\widehat{T}_{b}^{\mathrm{LS}}$ and
$\widehat{T}_{b}^{\mathrm{GL}}$ for the simple model $y_{t}=\phi^{0}+z_{t}\left(\delta_{1}^{0}+\delta^{0}\mathbf{1}\left\{ t>T_{b}^{0}\right\} \right)+e_{t}$
where $\left\{ z_{t}\right\} $ follows an ARMA(1,1) process and $e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right)$.
The distributions presented are the exact finite-sample distributions
of the LS and GL estimators, \citeauthor{bai:97RES}\textquoteright s
(1997) classical large-$N$ limit distribution, and the asymptotic
distribution of the GL estimator. Noteworthy are the non-standard
features of the finite-sample distribution of the LS estimator when
the break magnitude is small, which include multi-modality, fat tails
and asymmetry. The central mode is near $\widehat{T}_{b}^{\mathrm{LS}}$
while the other two modes are in the tails near the start and end
of the sample period; when the break magnitude is small $\widehat{T}_{b}^{\mathrm{LS}}$
tends to locate the break in the tails since the evidence of a break
is weak. It is evident that the classical large-$N$ asymptotic distribution
provides a poor approximation especially for small break sizes. Some
of these features have been found in other works {[}see, e.g., \citet{perron/zhu:05},
\citet{deng/perron:06}, Jiang, et al. \citeyearpar{jiang/wang/yu:16,jiang/wang/yu:17},
and \citet{casini/perron_CR_Single_Break}{]}. Turning to the densities
of the GL estimators, some of the nonstandard features appear also
for the GL estimator although to a much lesser extent. In particular,
the densities of the GL estimators are less spread out than the corresponding
densities for the LS estimator. For small breaks, the finite-sample
distributions of the LS and GL estimators are quite different, which
suggests that standard measures of accuracy (e.g., MAE and RMSE) can
be expected to differ substantially. For $\lambda_{0}=0.5$ the GL
estimator exhibits much less variability and more precision. The figures
also show that the asymptotic distribution of the GL estimator provides
an accurate approximation for large breaks while for small breaks
the approximation is less accurate. However, it captures the fat-tails
of the finite-sample distribution which suggests that it does not
underestimate uncertainty about the break location unlike \citeauthor{bai:97RES}\textquoteright s
(1997) distribution.

The GL method is useful because it weights the information from the
least-squares criterion function with the information from the prior
density\textemdash which, here, is the density of the asymptotic distribution
of $\widehat{T}_{b}^{\mathrm{LS}}$. Note that the least-squares objective
function is quite flat when the magnitude of the break is small and
so $\widehat{T}_{b}^{\mathrm{LS}}$ is imprecise. The resulting Quasi-posterior,
or, e.g., its  median, is likely to lead to better estimates in finite-samples,
because it takes into account the overall shape of the objective function
which weighted by the prior becomes more informative about the uncertainty
of the break date.

\subsection{\label{subsec:Normalized-Version-of}Normalized Version of $\mathcal{R}_{l,T}\left(s\right)$}

In order to develop the asymptotic results, we introduce a smoothing
sequence $\left\{ \gamma_{T}\right\} $ whose properties are specified
below and work with a normalized version of $\mathcal{R}_{l,T}\left(s\right)$
in order to be able to derive the relevant limit results. We assume
that $\lambda_{b}^{0}\in\varGamma^{0}\subset\left(0,\,1\right)$ is
the unknown extremum of $\widetilde{Q}\left(\theta^{0},\,\lambda_{b}\right)=\mathbb{E}\left[Q_{T}\left(\theta^{0},\,\lambda_{b}\right)\right]$
and that $\theta^{0}\triangleq\left(\left(\phi^{0}\right)',\,\left(\delta_{1}^{0}\right)',\,\left(\delta^{0}\right)'\right)'\in\mathbf{S}\subset\mathbb{R}^{p}\times\mathbb{R}^{q}\times\mathbb{R}^{q}$.
Our analysis is within a vanishing neighborhood of $\theta^{0}$.
For any $\theta\in\mathbf{S}$, let $\lambda_{b}^{0}\left(\theta\right)$
be an arbitrary element of $\varGamma^{0}\left(\theta\right)\triangleq\left\{ \lambda_{b}\in\varGamma^{0}:\,\widetilde{Q}\left(\theta,\,\lambda_{b}\right)=\sup_{\widetilde{\lambda}_{b}\in\mathcal{\varGamma}^{0}}\widetilde{Q}\left(\theta,\,\widetilde{\lambda}_{b}\right)\right\} $.
Provided a uniqueness condition is assumed (see Assumption \ref{Assumption Gaussian Process for Lap LapBai97}),
$\varGamma^{0}\left(\theta\right)$ contains a single element, $\lambda_{b}^{0}$.
Further, let $\overline{Q}_{T}\left(\theta,\,\lambda_{b}\right)\triangleq Q_{T}\left(\theta,\,\lambda_{b}\right)-Q_{T}\left(\theta,\,\lambda_{b}^{0}\right),$
$Q_{T}^{0}\left(\theta,\,\lambda_{b}\right)\triangleq\mathbb{E}\left[Q_{T}\left(\theta,\,\lambda_{b}\right)-Q_{T}\left(\theta,\,\lambda_{b}^{0}\right)|\,X\right],$
and $G_{T}\left(\theta,\,\lambda_{b}\right)\triangleq\overline{Q}_{T}\left(\theta,\,\lambda_{b}\right)-Q_{T}^{0}\left(\theta,\,\lambda_{b}\right).$
These expressions are given by $G_{T}\left(\theta,\,\lambda_{b}\right)=g_{e}\left(\theta,\,\lambda_{b}\right)$,
$Q_{T}^{0}=g_{d}\left(\theta,\,\lambda_{b}\right)$ and $\overline{Q}_{T}=g_{d}\left(\theta,\,\lambda_{b}\right)+g_{e}\left(\theta,\,\lambda_{b}\right)$,
where

\begin{align}
g_{d}\left(\theta,\,\lambda_{b}\right) & =\delta'_{T}\left\{ \left(Z'_{0}MZ_{2}\right)\left(Z'_{2}MZ_{2}\right)^{-1}\left(Z'_{2}MZ_{0}\right)-Z_{0}'MZ_{0}\right\} \delta_{T},\label{eq. A.2.1-1}
\end{align}
and
\begin{align}
g_{e} & \left(\theta,\,\lambda_{b}\right)=2\delta_{T}'\left(Z'_{0}MZ_{2}\right)\left(Z'_{2}MZ_{2}\right)^{-1}Z_{2}Me-2\delta_{T}'\left(Z'_{0}Me\right)\label{eq. A.2.3-1}\\
 & \quad+e'MZ_{2}\left(Z'_{2}MZ_{2}\right)^{-1}Z_{2}Me-e'MZ_{0}\left(Z'_{0}MZ_{0}\right)^{-1}Z'_{0}Me.\nonumber
\end{align}
They are derived in Section \ref{subsection: Preliminary Lemmas}.
For the purpose of developing the asymptotic theory, the GL estimator
$\widehat{\lambda}_{b}^{\mathrm{GL}}\left(\theta\right)$ is defined
as the minimizer of a normalized version of $\mathcal{R}_{l,T}\left(s\right)$:
\vfill{}

\begin{align}
\Psi_{l,T}\left(s;\,\theta\right) & =\int_{\varGamma^{0}}l\left(s-\lambda_{b}\right)\frac{\exp\left(\left(\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)\right)\overline{Q}_{T}\left(\theta,\,\lambda_{b}\right)\right)\pi\left(\lambda_{b}\right)}{\int_{\varGamma^{0}}\exp\left(\left(\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)\right)\overline{Q}_{T}\left(\theta,\,\lambda_{b}\right)\right)\pi\left(\lambda_{b}\right)d\lambda_{b}}d\lambda_{b}\label{eq. Definition GLE, Psi(s, theta) eq. 5}\\
 & =\int_{\varGamma^{0}}l\left(s-\lambda_{b}\right)\frac{\exp\left(\left(\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)\right)\left(G_{T}\left(\theta,\,\lambda_{b}\right)+Q_{T}^{0}\left(\theta,\,\lambda_{b}\right)\right)\right)\pi\left(\lambda_{b}\right)}{\int_{\varGamma^{0}}\exp\left(\left(\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)\right)\left(G_{T}\left(\theta,\,\lambda_{b}\right)+Q_{T}^{0}\left(\theta,\,\lambda_{b}\right)\right)\right)\pi\left(\lambda_{b}\right)d\lambda_{b}}d\lambda_{b}.\nonumber
\end{align}
Note that, under Condition \ref{Condition 1 LapBai97} below, this
is equivalent to the minimizer of $\mathcal{R}_{l,T}\left(s\right)$
since $\overline{Q}_{T}\left(\theta,\,\lambda_{b}\right)$ can always
be normalized without affecting its maximization. Different choices
of $\left\{ \gamma_{T}\right\} $ give rise to GL estimators with
different limiting distributions. Using $\delta_{T}$ or any consistent
estimate (e.g., $\widehat{\delta}_{T}^{\mathrm{LS}}$) in the factor
$\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)$
is irrelevant because they are asymptotically equivalent. Our analysis
is local in nature and thus we write $\widehat{\lambda}_{b}^{\mathrm{GL}}(\widehat{\theta})\triangleq\widehat{\lambda}_{b}^{\mathrm{GL,}*}(r_{T}(\widehat{\theta}-\theta^{0}),\,r_{T}(\widehat{\theta}-\theta^{0})),$
where $r_{T}$ is the convergence rate of $\widehat{\theta}-\theta^{0}$.
Note that $G_{T}\left(\cdot,\,\cdot\right)$ and $Q_{T}^{0}\left(\cdot,\,\cdot\right)$
constitute the stochastic and the deterministic part of the objective
function, respectively. Both depend on $r_{T}(\widehat{\theta}-\theta^{0})$
and our proof proceeds in conditioning first on the effect of $r_{T}(\widehat{\theta}-\theta^{0})$
on the deterministic part to obtain weak convergence of the stochastic
part to a limit process that does not depend on this conditioning.
See below for more details. Hence, it is required to introduce two
indices $\widetilde{v}$ and $v$, such that we define $\widehat{\lambda}_{b}^{\mathrm{GL}}(\widehat{\theta})=\widehat{\lambda}_{b}^{\mathrm{GL,*}}(\widetilde{v},\,v)$
as the minimizer of
\begin{align}
\Psi_{l,T}(s;\,\widetilde{v},\,v) & \triangleq\int_{\varGamma^{0}}l(s-\lambda_{b})\times\label{Eq. (14) - Definition of Psi(s,v,v)}\\
 & \quad\frac{\exp\left(\left(\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)\right)\left(G_{T}\left(\theta^{0}+\widetilde{v}/r_{T},\,\lambda_{b}\right)+Q_{T}^{0}\left(\theta^{0}+v/r_{T},\,\lambda_{b}\right)\right)\right)\pi\left(\lambda_{b}\right)}{\int_{\varGamma^{0}}\exp\left(\left(\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)\right)\left(G_{T}\left(\theta^{0}+\widetilde{v}/r_{T},\,\lambda_{b}\right)+Q_{T}^{0}\left(\theta^{0}+v/r_{T},\,\lambda_{b}\right)\right)\right)\pi\left(\lambda_{b}\right)d\lambda_{b}}d\lambda_{b}.\nonumber
\end{align}
 For each $v,$ we show weak convergence as a function of $\widetilde{v}$
to a limit process that does not depend on $v.$ In a second step,
we use the monotonicity in $v$ of $Q_{T}^{0}$ which, relying on
the argument in \citet{jureckova:77}, allows us to achieve weak convergence
uniformly in $v$. We first show the consistency and rate of convergence
of $\widehat{\lambda}_{b}^{\mathrm{GL}}$. These results imply that
$\theta^{0}$ is estimated as if $T_{b}^{0}$ were known. Thus, $\widehat{\theta}$
is $\sqrt{T}$-consistent and asymptotically normal so that we set
$r_{T}=\sqrt{T}$ hereafter. We first show, for each pair $\left(v,\,\widetilde{v}\right)$
with $v,\,\widetilde{v}\in\mathbf{V}$, the convergence of the marginal
distributions of the sample function $\Psi_{l,T}\left(s;\,v,\,\widetilde{v}\right)$
to the marginal distributions of the random function
\begin{align*}
\Psi_{l}^{0}\left(s\right) & =\int_{\mathbb{R}}l\left(s-u\right)\left(\mathscr{V}\left(u\right)/\int_{\mathbb{R}}\mathscr{V}\left(v\right)dv\right)du,
\end{align*}
 where
\begin{align}
\mathscr{V}\left(s\right)\triangleq\mathscr{W}\left(s\right)-\varLambda^{0}\left(s\right) & \triangleq\begin{cases}
2\left(\left(\delta^{0}\right)'\Sigma_{1}\delta^{0}\right)^{1/2}W_{1}\left(-s\right)-\left|s\right|\left(\delta^{0}\right)'V_{1}\delta^{0}, & \textrm{if }s\leq0\\
2\left(\left(\delta^{0}\right)'\Sigma_{2}\delta^{0}\right)^{1/2}W_{2}\left(s\right)-s\left(\delta^{0}\right)'V_{2}\delta^{0}, & \textrm{if }s>0,
\end{cases}\label{eq. V(s) Limit Process Theorem Post Mean LapBai97}
\end{align}
 and $W_{1},\,W_{2}$ are independent standard Wiener processes defined
on $[0,\,\infty)$. The limit process $\Psi_{l}^{0}\left(s\right)$
does not depend on $v$ nor $\widetilde{v}$. Next, we show that
the family of probability measures in $\mathbb{C}_{b}\left(\mathbf{K}\right)$,
with $\mathbf{K}\triangleq\left\{ s\in\mathbb{R}:\,\left|s\right|\leq K\textrm{ and }K<\infty\right\} $,
generated by the contractions of $\Psi_{l,T}\left(s;\,\widetilde{v},\,v\right)$
on $\mathbf{K}$ is dense uniformly in $\left(v,\,\widetilde{v}\right)$.
Finally, we examine the oscillations of the minimizers of the sample
criterion $\Psi_{l,T}\left(s;\,v,\,\widetilde{v}\right)$.

It is important to note that the results derived in this section are
more general than what is required for the structural change model.
The reason is that the change-point model is recovered as a special
case corresponding to $\Psi_{l,T}\left(s\right)=\Psi_{l,T}\left(s;\,0,\,0\right)$.
That is, defining the GL estimator in a $1/r_{T}$-neighborhood of
the slope parameter vector $\theta^{0}$ is not strictly necessary
and one can essentially develop the same analysis with $\theta$ fixed
at its true value $\theta^{0}$. This relies on the properties of
(orthogonal) least-squares projections and would not apply, for example,
to the least absolute deviation (LAD) estimator of the break date
{[}cf. \citet{baiLAD:95}{]} for which $\Psi_{l,T}\left(s;\,\widetilde{v},\,v\right)$
should instead be considered. The same issue is present when estimating
structural changes in the quantile regression model {[}cf. \citet{oka/qu:10}{]}
and in using instrumental variables models {[}cf. \citet{hall/han/boldea:10}
and Perron and Yamamoto \citeyearpar{perron/yamamoto:14,perron/yanamoto:15}{]}.
We establish theoretical results under this more general setting
since they may be useful for future work.

Let $\lambda_{b,T}^{0}\left(v\right)=\lambda_{b,T}^{0}\left(\theta^{0}+v/r_{T}\right)$.
 Introduce the local parameter $u=\psi_{T}\left(\lambda_{b}-\lambda_{b,T}^{0}\left(v\right)\right)$
and let $\pi_{T,v}\left(u\right)\triangleq\pi\left(\lambda_{b,T}^{0}\left(v\right)+u/\psi_{T}\right)$,
$Q_{T,v}\left(u\right)\triangleq Q_{T}^{0}(\theta^{0}+v/r_{T},\,\lambda_{b,T}^{0}$
$\left(v\right)+u/\psi_{T}$), and $\widetilde{G}_{T,v}\left(u,\,\widetilde{v}\right)\triangleq G_{T}\left(\theta^{0}+\widetilde{v}/r_{T},\,\lambda_{b,T}^{0}\left(v\right)+u/\psi_{T}\right)$,
where the sequence $\left\{ \psi_{T}\right\} $ depends on the results
on consistency and rate of convergence of $\widehat{\lambda}_{b}^{\mathrm{GL}}$
in Proposition \ref{Proposition: Consistency and Rate of Convergence}.
Apply a simple substitution in \eqref{Eq. (14) - Definition of Psi(s,v,v)}
to yield,
\begin{align}
\Psi_{l,T}\left(s;\,\widetilde{v},\,v\right) & =\int_{\Gamma_{T}}l\left(s-u\right)\frac{\exp\left(\left(\gamma_{T}/T\left\Vert \delta_{T}\right\Vert ^{2}\right)\left(\widetilde{G}_{T,v}\left(u,\,\widetilde{v}\right)+Q_{T,v}\left(u\right)\right)\right)\pi_{T,v}\left(u\right)du}{\int_{\Gamma_{T}}\exp\left(\left(\gamma_{T}/T\left\Vert \delta_{T}\right\Vert ^{2}\right)\left(\widetilde{G}_{T,v}\left(w,\,\widetilde{v}\right)+Q_{T,v}\left(w\right)\right)\right)\pi_{T,v}\left(w\right)dw},\label{eq. (17)-1-1}
\end{align}
 where $\Gamma_{T}\triangleq\left\{ u\in\mathbb{R}:\,\lambda_{b}^{0}+u/\psi_{T}\in\varGamma^{0}\right\} $.

\begin{assumption}
\label{A.9a Bai 97}$\left\{ \left(z_{t},\,e_{t}\right)\right\} $
is second-order stationary within each regime such that $\mathbb{E}\left(z_{t}z'_{t}\right)=V_{1}$
and $\mathbb{E}\left(e_{t}^{2}\right)=\sigma_{1}^{2}$ for $t\leq T_{b}^{0}$
and $\mathbb{E}\left(z_{t}z'_{t}\right)=V_{2}$ and $\mathbb{E}\left(e_{t}^{2}\right)=\sigma_{2}^{2}$
for $t>T_{b}^{0}$.
\end{assumption}

\begin{assumption}
\label{Assumption A.9b Bai 97, LapBai97}For $r\in\left[0,\,1\right],$
$\left(T_{b}^{0}\right)^{-1/2}\sum_{t=1}^{\left\lfloor rT_{b}^{0}\right\rfloor }z_{t}e_{t}\Rightarrow\mathscr{G}_{1}\left(r\right)$
and $\left(T-T_{b}^{0}\right)^{-1/2}\sum_{t=T_{b}^{0}+1}^{T_{b}^{0}+\left\lfloor r\left(T-T_{b}^{0}\right)\right\rfloor }$
$z_{t}e_{t}\Rightarrow\mathscr{G}_{2}\left(r\right)$, where $\mathscr{G}_{i}\left(\cdot\right)$
is a multivariate Gaussian process on $\left[0,\,1\right]$ with zero
mean and covariance $\mathbb{E}\left[\mathscr{G}_{i}\left(u\right),\,\mathscr{G}_{i}\left(s\right)\right]=\min\left\{ u,\,s\right\} \Sigma_{i}$
$\left(i=1,\,2\right)$, and $\Sigma_{1}\triangleq\lim_{T\rightarrow\infty}\mathbb{E}\left[\left(T_{b}^{0}\right)^{-1/2}\sum_{t=1}^{T_{b}^{0}}z_{t}e_{t}\right]^{2}$,
$\Sigma_{2}\triangleq\lim_{T\rightarrow\infty}\mathbb{E}\left[\left(T-T_{b}^{0}\right)^{-1/2}\sum_{t=T_{b}^{0}+1}^{T}z_{t}e_{t}\right]^{2}$.
Furthermore, for any $0<r_{0}<1$ with $r_{0}<\lambda_{0}$, $T^{-1}\sum_{t=\left\lfloor r_{0}T\right\rfloor +1}^{\left\lfloor \lambda_{0}T\right\rfloor }z_{t}z'_{t}\overset{\mathbb{P}}{\rightarrow}\left(\lambda_{0}-r_{0}\right)V_{1},$
and with $\lambda_{0}<r_{0}$ $T^{-1}\sum_{t=\left\lfloor \lambda_{0}T\right\rfloor +1}^{\left\lfloor r_{0}T\right\rfloor }$
$z_{t}z'_{t}\overset{\mathbb{P}}{\rightarrow}\left(r_{0}-\lambda_{0}\right)V_{2}$
so that $\lambda_{-}$ and $\lambda_{+}$ (the minimum and maximum
of the eigenvalues of the last two matrices) satisfy $0<\lambda_{-}\leq\lambda_{+}<\infty.$
\end{assumption}
Assumptions \ref{A.9a Bai 97}-\ref{Assumption A.9b Bai 97, LapBai97}
are equivalent to A9 in \citet{bai:97RES} and A7 in \citet{bai/perron:98}.
More specifically, Assumption \ref{Assumption A.9b Bai 97, LapBai97}
requires that, within each regime, an Invariance Principle holds for
$\left\{ z_{t}e_{t}\right\} .$ Let $\zeta_{t}\triangleq z_{t}e_{t}$.
For $u\leq0$ let $g\left(\zeta_{t};\,u\right)\triangleq\left(\delta^{0}\right)'\sum_{t=T_{b}^{0}+\left\lfloor u/v_{T}^{2}\right\rfloor }^{T_{b}^{0}}\zeta_{t}$
and $\widetilde{g}\left(\zeta_{t};\,u,\,\widetilde{v},\,v;\,\psi_{T},\,r_{T}\right)\triangleq\sqrt{\psi_{T}}\left(\delta^{0}+\widetilde{v}/r_{T}\right)'\sum_{t=T\lambda_{b}^{0}\left(\theta^{0}+v/r_{T}\right)+\left\lfloor u/\psi_{T}\right\rfloor }^{T\lambda_{b}^{0}\left(\theta^{0}+v/r_{T}\right)}\zeta_{t}.$
Define analogously $g\left(\zeta_{t};\,u\right)$ and $\widetilde{g}\left(\zeta_{t};\,u,\,\widetilde{v},\,v;\,\psi_{T},\,r_{T}\right)$
for the case $T_{b}>T_{b}^{0}$. We now present some technical assumptions
that are necessary for the derivation of the asymptotic results for
the GL estimate.
\begin{assumption}
\label{Assumption Gaussian Process for Lap LapBai97}For some neighborhood
 $\Theta^{0}\subset\mathbf{S}$ of $\theta^{0},$ (i) for all $\lambda_{b}\neq\lambda_{b}^{0},\,\widetilde{Q}\left(\theta^{0},\,\lambda_{b}\right)<\widetilde{Q}\left(\theta^{0},\,\lambda_{b}^{0}\right)$;
(ii) for any $v,\,\widetilde{v}_{1},\,\widetilde{v}_{2}\in\mathbf{V}$
and $u,\,s\in\mathbb{R},$
\[
\varSigma\left(u,\,s\right)\triangleq\lim_{T\rightarrow\infty}\mathbb{E}\left[\widetilde{g}\left(\zeta_{t};\,u,\,\widetilde{v}_{1},\,v;\,\psi_{T},\,r_{T}\right)\widetilde{g}\left(\zeta_{t};\,s,\,\widetilde{v}_{2},\,v;\,\psi_{T},\,r_{T}\right)'\right],
\]
does not depend on $v,\,\widetilde{v}_{1},\,\widetilde{v}_{2}\in\mathbf{V}$.
\end{assumption}
Part (i) of Assumption \ref{Assumption Gaussian Process for Lap LapBai97}
is an identification condition. Assumption \ref{Assumption Gaussian Process for Lap LapBai97}-(ii)
holds whenever $\widehat{\lambda}_{b}$ is consistent. With Assumption
\ref{Assumption Gaussian Process for Lap LapBai97}-(ii) we fully
characterize the Gaussian component of the limit process $\mathscr{V}\left(\cdot\right);$
it implies that $\varSigma\left(\cdot,\,\cdot\right)$ is strictly
positive and that
\begin{align}
\forall u,\,s\in\mathbb{R}: & \begin{cases}
\forall c>0:\,\varSigma\left(cu,\,cs\right)=c\varSigma\left(u,\,s\right),\\
\varSigma\left(u,\,u\right)+\varSigma\left(s,\,s\right)-2\varSigma\left(u,\,s\right)=\varSigma\left(u-s,\,u-s\right),
\end{cases}\label{Eq. (18) Covariance of Limit Gaussina process}
\end{align}
 where the second implication requires some simple but tedious manipulations.
Finally, the following assumption is automatically satisfied if $l\left(\cdot\right)$
is a convex function with a unique minimum.
\begin{assumption}
\label{Assumption Uniquess loss Fucntion LapBai97}$\xi_{l}^{0}\triangleq\xi\left(\lambda_{b}^{0}\right)$
is uniquely defined by
\begin{align*}
\Psi_{l}\left(\xi_{l}^{0}\right)\triangleq\inf_{s}\Psi_{l}\left(s\right)=\inf_{s}\int_{\mathbb{R}}l\left(s-u\right)\left(\exp\left(\mathscr{V}\left(u\right)\right)/\left(\int_{\mathbb{R}}\exp\left(\mathscr{V}\left(w\right)\right)dw\right)\right)du & .
\end{align*}

\end{assumption}

\subsection{\label{Sub section Theorems Laplace in SC}Asymptotic Results for
the GL Estimate}

We first show the consistency and rate of convergence of the GL estimator.
The latter allows us to characterize the rate of $\psi_{T}$ and proceed
with the asymptotic analysis in a neighborhood of $\lambda_{b}^{0}$.
In practice, the squared loss function is often employed. Hence, it
is useful to first present in Theorem \ref{Theorem Posterior Mean Lap Estimation in Bai 97}
the theoretical results for this case for which the GL estimator is
$\widehat{\lambda}_{b}^{\mathrm{GL}}=\int_{\varGamma^{0}}\lambda_{b}p_{T}\left(\lambda_{b}\right)d\lambda_{b},$
i.e., the Quasi-posterior mean. This allows us to keep the theoretical
results tractable and provide the main intuition without the need
of  complex notation. This case is also instructive since we can
compare our results with corresponding ones for the least-squares
and Bayesian change-point estimators. Corresponding results for general
loss functions  are given in Theorem \ref{Theorem Geneal Laplace Estimator LapBai97}.


\subsubsection{Consistency and Rate of Convergence}

The rate of convergence is similar to that of the LS estimator; the
difference being that $\psi_{T}=T^{1-2\vartheta}$ with $\vartheta\in\left(0,\,1/2\right)$
for the LS estimator and $\vartheta\in\left(0,\,1/4\right)$ for the
GL estimator.
\begin{prop}
\label{Proposition: Consistency and Rate of Convergence}Under Assumptions
\ref{Assumption A Bai97}-\ref{Assumption A4 Bai97}, \ref{Assumption The-loss-function LapBai97}-\ref{Assumption Small Shift BP}
and \ref{Assumption Gaussian Process for Lap LapBai97}-(i): (i) $\widehat{\lambda}_{b}^{\mathrm{GL}}=\lambda_{b}^{0}+o_{\mathbb{P}}\left(1\right)$;
(ii) $\widehat{\lambda}_{b}^{\mathrm{GL}}=\lambda_{b}^{0}+O_{\mathbb{P}}\left(\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)^{-1}\right)$.
\end{prop}

\subsubsection{The Asymptotic Distribution of the Quasi-posterior Mean}

For the squared loss function $\widehat{\lambda}_{b}^{\mathrm{GL}}\left(\widehat{\theta}\right)\triangleq\widehat{\lambda}_{b}^{\mathrm{GL,}*}\left(\widetilde{v},\,v\right),$
where
\begin{align}
\widehat{\lambda}_{b}^{\mathrm{GL},*}\left(\widetilde{v},\,v\right) & \triangleq\frac{\int_{\varGamma^{0}}\lambda_{b}\exp\left(\left(\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)\right)\left(G_{T}\left(\theta^{0}+\widetilde{v}/r_{T},\,\lambda_{b}\right)+Q_{T}^{0}\left(\theta^{0}+v/r_{T},\,\lambda_{b}\right)\right)\right)\pi\left(\lambda_{b}\right)d\lambda_{b}}{\int_{\varGamma^{0}}\exp\left(\left(\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)\right)\left(G_{T}\left(\theta^{0}+\widetilde{v}/r_{T},\,\lambda_{b}\right)+Q_{T}^{0}\left(\theta^{0}+v/r_{T},\,\lambda_{b}\right)\right)\right)\pi\left(\lambda_{b}\right)d\lambda_{b}},\label{Eq. (14) - Definition of lambda*(v,v)}
\end{align}
 and $v,$ $\widetilde{v}$ each belong to some compact set $\mathbf{V}\subset\mathbb{R}^{p+2q}$.
For each $v\in\mathbf{V},$ we consider $\widehat{\lambda}_{b}^{\mathrm{GL,*}}\left(\cdot,\,v\right)$
as a random process with paths in $\mathbb{D}_{b}\left(\mathbf{V}\right)$.
We focus on the weak convergence of $\widehat{\lambda}_{b}^{\mathrm{GL},*}\left(\cdot,\,v\right)$
for fixed $v$ since the limit process is independent of $v$ and
constant as a function of $\widetilde{v}$; we then exploit monotonicity
in $v.$  More precisely, we will show that for $\lambda_{b,T}^{0}\left(v\right)=\lambda_{b,T}^{0}\left(\theta^{0}+v/r_{T}\right)$
and diverging sequences $\left\{ \gamma_{T}\right\} $ and $\left\{ r_{T}\right\} $,
the sequence $a_{T}\left(\widehat{\lambda}_{b}^{\mathrm{GL},*}\left(\widetilde{v},\,v\right)-\lambda_{b,T}^{0}\left(v\right)\right)$
converges in distribution in $\mathbb{D}_{b}\left(\mathbf{V}\right)$
for each $v$ to a limit process not depending on $v$ nor $\widetilde{v}$.
Since it is monotonic in $v$, we do not need to show uniform convergence
directly.  Introduce the local parameter $u=\psi_{T}\left(\lambda_{b}-\lambda_{b,T}^{0}\left(v\right)\right)$;
a simple substitution in \eqref{Eq. (14) - Definition of lambda*(v,v)}
yields,
\begin{align}
\psi_{T}\left(\widehat{\lambda}_{b}^{\mathrm{GL},*}\left(\widetilde{v},\,v\right)-\lambda_{b,T}^{0}\left(v\right)\right) & =\frac{\int_{\mathbb{R}}u\exp\left(\left(\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)\right)\left(\widetilde{G}_{T,v}\left(u,\,\widetilde{v}\right)+Q_{T,v}\left(u\right)\right)\right)\pi_{T,v}\left(u\right)du}{\int_{\mathbb{R}}\exp\left(\left(\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)\right)\left(\widetilde{G}_{T,v}\left(u,\,\widetilde{v}\right)+Q_{T,v}\left(u\right)\right)\right)\pi_{T,v}\left(u\right)du},\label{eq. (17)-1}
\end{align}
where again we have used the notation $\pi_{T,v}\left(u\right)=\pi\left(\lambda_{b,T}^{0}\left(v\right)+u/\psi_{T}\right),$
$Q_{T,v}\left(u\right)=Q_{T}^{0}\left(\theta^{0}+v/r_{T}\right.$
$\left.\lambda_{b,T}^{0}\left(v\right)+u/\psi_{T}\right)$ and $\widetilde{G}_{T,v}\left(u,\,\widetilde{v}\right)=G_{T}\left(\theta^{0}+\widetilde{v}/r_{T},\,\lambda_{b,T}^{0}\left(v\right)+u/\psi_{T}\right)$.
The limit of the GL estimator depends on the limit of the process
$\left(\gamma_{T}/\left(T\left\Vert \delta_{T}\right\Vert ^{2}\right)\right)\left(\widetilde{G}_{T,v}\left(u,\,\widetilde{v}\right)+Q_{T,v}\left(u\right)\right)$.
As part of the proof of Theorem \ref{Theorem Posterior Mean Lap Estimation in Bai 97},
we show that the sequence of processes $\left\{ \widetilde{G}_{T,v}\left(u,\,\widetilde{v}\right),\,T\geq1\right\} $
converges weakly in $\mathbb{D}_{b}\left(\mathbb{R}\times\mathbf{V}\right)$
to a Gaussian process $\mathscr{W}$ not varying with $v$, whereas
$Q_{T,v}\left(\cdot\right)$ is approximated by a (deterministic)
drift process taking negative values, and is monotonic in $v$ and
flat in $\widetilde{v}$.  We show that this implies that $\widehat{\lambda}_{b}^{\mathrm{GL},*}\left(\widetilde{v},\,v\right)-\lambda_{b,T}^{0}\left(v\right)$
is monotonic in $v$ which then leads to uniform convergence in $v$
following the argument of \citet{jureckova:77}.

In anticipation of the results, we make a few comments about the notation
for the weak convergence of processes on the space of bounded c�dl�g
functions $\mathbb{D}_{b}$. Let $\mathbf{V}\subset\mathbb{R}^{p+2q}$
be a compact set. Let $W_{T}\left(u,\,\widetilde{v},\,v\right)$ denote
an arbitrary sample process with bounded c�dl�g paths evaluated at
the local parameters $u\in\mathbb{R},$ and $v,\,\widetilde{v}\in\mathbf{V}$.
For each fixed $v\in\mathbf{V}$, we shall write $W_{T}\left(u,\,\widetilde{v},\,v\right)\Rightarrow\mathscr{W}\left(u,\,\widetilde{v},\,v\right)$
in $\mathbb{D}_{b}\left(\mathbb{R}\times\mathbf{V}\right)$ whenever
the process $W_{T}\left(\cdot,\,\cdot,\,v\right)$ converges weakly
to $\mathscr{W}\left(\cdot,\,\cdot,\,v\right)$, where $\mathscr{W}\left(\cdot,\,\cdot,\,v\right)$
also belongs to $\mathbb{D}_{b}\left(\mathbb{R}\times\mathbf{V}\right)$.
As a shorthand, we shall omit the argument $u\,\left(\widetilde{v}\right)$
if the limit process does not depend on $u\,\left(\widetilde{v}\right)$.
The same notational conventions are used for the case when $W_{T}$
is only a function of $\left(\widetilde{v},\,v\right).$ In Theorem
\ref{Theorem Posterior Mean Lap Estimation in Bai 97} the convergence
holds for every $v\in\mathbf{V}$, stated as convergence in $\mathbb{D}_{b}$.
\begin{condition}
\label{Condition 1 LapBai97}As $T\rightarrow\infty$ there exist
a positive finite number $\kappa_{\gamma}$ such that $\gamma_{T}/T\left\Vert \delta_{T}\right\Vert ^{2}\rightarrow\kappa_{\gamma}$.
\end{condition}
\begin{thm}
\label{Theorem Posterior Mean Lap Estimation in Bai 97}Assume $l\left(\cdot\right)$
is the squared loss function. Under Assumptions \ref{Assumption A Bai97}-\ref{Assumption A4 Bai97}
and \ref{Assumption The-loss-function LapBai97}-\ref{Assumption Uniquess loss Fucntion LapBai97},
and Condition \ref{Condition 1 LapBai97}, then in $\mathbb{D}_{b}$,
\begin{align}
T\left\Vert \delta_{T}\right\Vert ^{2}\left(\widehat{\lambda}_{b}^{\mathrm{GL}}-\lambda_{b}^{0}\right) & \Rightarrow\frac{\int u\exp\left(\mathscr{W}\left(u\right)-\varLambda^{0}\left(u\right)\right)du}{\int\exp\left(\mathscr{W}\left(u\right)-\varLambda^{0}\left(u\right)\right)du}\triangleq\int up_{0}^{*}\left(u\right)du,\label{eq. Asy Dist Laplace Bai97 part (ii)}
\end{align}
where $\mathscr{W}\left(\cdot\right)$ and $\varLambda^{0}\left(\cdot\right)$
are defined in \eqref{eq. V(s) Limit Process Theorem Post Mean LapBai97}.
\end{thm}
Theorem \ref{Theorem Posterior Mean Lap Estimation in Bai 97} states
that the asymptotic distribution of the GL estimate is a ratio of
integrals of functions of tight Gaussian processes. We shall compare
this result with the limiting distribution of the Bayesian change-point
estimator of \citet{ibragimov/has:81}. They considered a simple diffusion
process with a change-point in the deterministic drift {[}see their
eq. (2.17) on pp. 338{]}. The limiting distribution of the GL estimate
from Theorem \ref{Theorem Posterior Mean Lap Estimation in Bai 97}
for the case of a break in the mean for model \eqref{Eq. the Single Break Model}
is essentially the same (and exactly so in the i.i.d. case with stationary
regimes) as theirs. Hence, while the GL estimator has a classical
(frequentist) interpretation, it is first-order equivalent in law
to a corresponding Bayes-type estimator.

We now present a result about the dual nature of the limiting distribution
of the GL estimator. The following proposition shows that, under different
conditions on the smoothing sequence parameter $\left\{ \gamma_{T}\right\} $,
the GL estimator achieves different limiting distributions.
\begin{condition}
\label{Condition 2 LapBai97}As $T\rightarrow\infty$, $T\left\Vert \delta_{T}\right\Vert ^{2}/\gamma_{T}=o\left(1\right)$.
\end{condition}
\begin{prop}
\label{Proposition}Assume $l\left(\cdot\right)$ is the squared loss
function. Under Assumptions \ref{Assumption A Bai97}-\ref{Assumption A4 Bai97}
and \ref{Assumption The-loss-function LapBai97}-\ref{Assumption Uniquess loss Fucntion LapBai97},
and Condition \ref{Condition 2 LapBai97}, $T\left\Vert \delta_{T}\right\Vert ^{2}\left(\widehat{\lambda}_{b}^{\mathrm{GL}}-\lambda_{b}^{0}\right)\Rightarrow\arg\max_{s\in\mathbb{R}}\mathscr{V}\left(s\right)$
in $\mathbb{D}_{b}$, with $\mathscr{V}\left(\cdot\right)$ defined
in \eqref{eq. V(s) Limit Process Theorem Post Mean LapBai97}.
\end{prop}
\begin{cor}
\label{Corollary part (i) of Theorem LapBai97}Define $\Xi_{e}\triangleq\left(\delta^{0}\right)'\Sigma_{2}\delta^{0}/\left(\delta^{0}\right)'\Sigma_{1}\delta^{0}$
and $\Xi_{Z}\triangleq\left(\delta^{0}\right)'V_{2}\delta^{0}/\left(\delta^{0}\right)'V_{1}\delta^{0}$.
Under Assumptions \ref{Assumption A Bai97}-\ref{Assumption A4 Bai97}
and \ref{Assumption The-loss-function LapBai97}-\ref{Assumption Uniquess loss Fucntion LapBai97},
and Condition \ref{Condition 2 LapBai97}, $\left(\left(\delta'_{T}V_{1}\delta{}_{T}\right)^{2}/\delta'_{T}\Sigma_{1}\delta_{T}\right)\left(\widehat{T}_{b}^{\mathrm{GL}}-T_{b,T}^{0}\right)\overset{d}{\rightarrow}\arg\max_{s\in\mathbb{R}}\mathscr{V}^{*}\left(s\right)$
in $\mathbb{D}_{b}$ where
\begin{align*}
\mathscr{V}^{*}\left(s\right)=W_{1}\left(-s\right)-\left|s\right|/2\,\mathrm{\,if\,}\,s\leq0 & ;\,\,\mathscr{V}^{*}\left(s\right)=\Xi_{e}^{1/2}W_{2}\left(s\right)-\Xi_{Z}s/2\,\,\mathrm{if}\,\,s>0.
\end{align*}
\end{cor}
Corollary \ref{Corollary part (i) of Theorem LapBai97} and Proposition
\ref{Proposition} show that with enough smoothing applied, the GL
estimator is (first-order) asymptotically equivalent to the least-squares
or MLE {[}cf. \citet{bai:97RES} and \citet{yao:87}, respectively{]}.
The intuition is that when the criterion function is sufficiently
smoothed, the Quasi-posterior probability density converges to the
generalized dirac probability measure concentrated at the argmax of
the limit criterion function. This is analogous to a well-known result
{[}cf. Corollary 5.11 in \citet{robert/casella:04}{]}, stating that
in a parametric statistical experiment indexed by a parameter $\theta\in\Theta,$
the MLE $\widehat{\theta}_{T}^{\mathrm{ML}}$ is the limit of a Bayes
estimator as the smoothing parameter $\gamma\rightarrow\infty$, i.e.,
using obvious notation:
\begin{align*}
\widehat{\theta}_{T}^{\mathrm{ML}} & =\arg\max_{\theta\in\Theta}L_{T}\left(\theta\right)=\lim_{\gamma\rightarrow\infty}\frac{\int_{\Theta}\theta\exp\left(\gamma L_{T}\left(\theta\right)\right)\pi\left(\theta\right)d\theta}{\int_{\Theta}\exp\left(\gamma L_{T}\left(\theta\right)\right)\pi\left(\theta\right)d\theta}.
\end{align*}


\subsubsection{The Asymptotic Distribution for General Loss Functions}

For general loss functions satisfying Assumption \ref{Assumption The-loss-function LapBai97},
Theorem \ref{Theorem Geneal Laplace Estimator LapBai97} shows that
$T\left\Vert \delta_{T}\right\Vert ^{2}\left(\widehat{\lambda}_{b}^{\mathrm{GL}}-\lambda_{b}^{0}\right)$
is (first-order) asymptotically equivalent to $\xi_{l}^{0}$ defined
by
\begin{align}
\Psi_{l}\left(\xi_{l}^{0}\right) & \triangleq\inf_{r}\Psi_{l}\left(r\right)=\inf_{r\in\mathbb{R}}\left\{ \int_{\mathbb{R}}l\left(r-u\right)p_{0}^{*}\left(u\right)du\right\} .\label{Eq. Distribution of Bayes Estimator-1}
\end{align}

\begin{thm}
\label{Theorem Geneal Laplace Estimator LapBai97}Under Assumptions
\ref{Assumption A Bai97}-\ref{Assumption A4 Bai97} and \ref{Assumption The-loss-function LapBai97}-\ref{Assumption Uniquess loss Fucntion LapBai97},
and Condition \ref{Condition 1 LapBai97}, for $l\in\boldsymbol{L},$
$T\left\Vert \delta_{T}\right\Vert ^{2}(\widehat{\lambda}_{b}^{\mathrm{GL}}-\lambda_{b}^{0})$
$\Rightarrow\xi_{l}^{0}$ as defined by \eqref{Eq. Distribution of Bayes Estimator-1}.
\end{thm}
The existence and uniqueness of $\xi_{l}^{0}$ follow from Assumption
\ref{Assumption Uniquess loss Fucntion LapBai97}. If one interprets
$p_{0}^{*}\left(u\right)$ as a true posterior density function, then
$\xi_{l}^{0}$ would naturally be viewed as a Bayesian estimator for
the loss function $l_{T}\left(\cdot\right).$ In particular, in analogy
to the above comparison with the Bayesian estimator of \citet{ibragimov/has:81},
one can interpret the GL estimator as a Quasi-Bayesian estimator.
While this is by itself a theoretically interesting result, we actually
exploit it to construct more reliable inference methods about the
date of a structural change. Under the least-absolute deviation loss,
the GL estimator converges in distribution to the median of $p_{0}^{*}\left(u\right)$.
We shall use the results in Theorem \ref{Theorem Posterior Mean Lap Estimation in Bai 97}-\ref{Theorem Geneal Laplace Estimator LapBai97}
but not Proposition \ref{Proposition} since the latter implies the
same confidence intervals as in \citet{bai:97RES} and \citet{bai/perron:98}.
GL inference based on the Bayes-type limiting distribution provides
a more accurate description of the uncertainty over the parameter
space than the inference based on the density of $\arg\max_{s\in\mathbb{R}}\mathscr{V}\left(s\right)$
which underestimates uncertainty as shown by confidence intervals
with empirical coverage rates below the nominal level particularly
when the magnitude of the break is small (see Section \ref{Section Monte-Carlo-Simulation}).
After some investigation, we found that both estimation and inference
under the least-absolute loss works well and this is what will be
used in our simulation study.

\section{\label{Section Inference Methods}Confidence Sets Based on the GL
Estimator}

In this section, we discuss inference procedures for the break date
based on the large-sample results of the previous section. Inference
under general loss functions based on Theorem \ref{Theorem Geneal Laplace Estimator LapBai97}
is what we recommend to use in practice, in particular with an absolute
loss function.

Since the limiting distribution from Theorem \ref{Theorem Geneal Laplace Estimator LapBai97}
involves certain unknown quantities, we begin by assuming that they
can be replaced by consistent estimates. They are easy to construct
{[}cf. \citet{bai:97RES} and \citet{bai/perron:98}; see also Section
\ref{Section Monte-Carlo-Simulation}{]}.
\begin{assumption}
\label{Assumption Laplace Inference LapBai97}There exist sequences
of estimators $\widehat{\lambda}_{b,T},\,\widehat{\delta}_{T},\,\widehat{\Xi}_{Z,T},$
and $\widehat{\Xi}_{e,T}$ such that $\widehat{\lambda}_{b,T}=\lambda_{0}+o_{\mathbb{P}}\left(1\right)$,
$\widehat{\delta}_{T}=\delta_{T}+o_{\mathbb{P}}\left(1\right)$, $\widehat{\Xi}_{Z,T}=\Xi_{Z}+o_{\mathbb{P}}\left(1\right)$
and $\widehat{\Xi}_{e,T}=\Xi_{e}+o_{\mathbb{P}}\left(1\right)$.
Furthermore, for all $u,\,s\in\mathbb{R}$ and any $c>0$, there exist
covariation processes $\widehat{\varSigma}_{i,T}\left(\cdot\right)$
$\left(i=1,\,2\right)$ that satisfy (i) $\widehat{\varSigma}_{i,T}\left(u,\,s\right)=\varSigma_{i}^{0}\left(u,\,s\right)+o_{\mathbb{P}}\left(1\right)$,
(ii) $\widehat{\varSigma}_{i,T}\left(u-s,\,u-s\right)=\widehat{\varSigma}_{i,T}\left(u,\,u\right)+\widehat{\varSigma}_{i,T}\left(s,\,s\right)-2\widehat{\varSigma}_{i,T}\left(s,\,u\right)$,
(iii) $\widehat{\varSigma}_{i,T}\left(cu,\,cu\right)=c\widehat{\varSigma}_{i,T}\left(u,\,u\right)$,
(iv) $\mathbb{E}\left\{ \sup_{\left\Vert u\right\Vert =1}\widehat{\varSigma}_{i,T}^{2}\left(u,\,u\right)\right\} =O\left(1\right)$.
\end{assumption}
The first part and (i) of the second part follow from consistency
of $\widehat{\lambda}_{b,T}$ and from an Invariance Principle {[}cf.
Assumption \ref{Assumption A.9b Bai 97, LapBai97}{]}. Part (ii)-(iii)
are implied by Assumption \ref{Assumption Gaussian Process for Lap LapBai97}-(ii)
and consistency of $\widehat{\lambda}_{b,T}$. Let $\left\{ \widehat{\mathscr{W}}_{T}\right\} $
be a (sample-size dependent) sequence of two-sided zero-mean Gaussian
processes with covariance $\widehat{\varSigma}_{T}.$ Construct the
process $\widehat{\mathscr{V}}_{T}$ by replacing the population quantities
in $\mathscr{V}$ by their corresponding estimates from the first
part of Assumption \ref{Assumption Laplace Inference LapBai97} and
further, replace $\mathscr{W}$ by $\widehat{\mathscr{W}}_{T}$. Assumption
\ref{Assumption Laplace Inference LapBai97}-(i) basically implies
that the finite-dimensional limit law of $\left\{ \widehat{\mathscr{W}}_{T}\right\} $
is the same as the finite-dimensional laws of $\mathscr{W}$ while
parts (ii)-(iii) are needed for the integrability of the transform
$\exp\left(\widehat{\mathscr{V}}_{T}\left(\cdot\right)\right)$. Part
(iv) is needed for the proof of the asymptotic stochastic equicontinuity
of $\left\{ \widehat{\mathscr{W}}_{T}\right\} $. Let $\widehat{\xi}_{T}$
be defined as the sample analogue of $\xi_{l}^{0}$ that uses $\widehat{\mathscr{V}}_{T}\left(v\right)$
in place of $\mathscr{V}_{T}\left(v\right)$ in \eqref{Eq. Distribution of Bayes Estimator-1}.
The distribution of $\widehat{\xi}_{T}$ can be evaluated numerically.
\begin{prop}
\label{Proposition Inference Limi Distr LapBai97}Let $l\in\boldsymbol{L}$
be continuous. Under Assumption \ref{Assumption Laplace Inference LapBai97},
$\widehat{\xi}_{T}$ converges in distribution to the limiting distribution
in Theorem \ref{Theorem Geneal Laplace Estimator LapBai97}.
\end{prop}
The asymptotic distribution theory of the GL estimator may be exploited
in several ways to conduct inference about the break date. The finite-sample
distribution of the LS nd GL estimate of the break date displays significant
non-standard features (cf. Figure \ref{Fig1}). Hence, a conventional
two-sided confidence interval may not result in a confidence set with
reliable properties across all break magnitudes and break locations.
Thus, as in \citet{casini/perron_CR_Single_Break}, we use the concept
of Highest Quasi-posterior Density (HQPD) regions, defined analogously
to the Highest Density Region (HDR); cf. \citet{hyndman:96}. See
also \citet{samworth/wand:10} and \citeauthor{mason/polonik:09}
(2008, 2009)\nocite{mason/polonik:08} for more recent developments.
For an illustrative example on the properties of the HDR see the discussion
of Figure 11 in \citet{casini/perron_CR_Single_Break}.
\begin{defn}
\label{def:Highest-Density-Region}Highest Density Region: Let the
density function $f_{Y}\left(y\right)$ of a random variable $Y$
defined on a probability space $\left(\Omega_{Y},\,\mathscr{F}_{Y},\,\mathbb{P}_{Y}\right)$
and taking values on the measurable space $\left(\mathcal{Y},\,\mathscr{Y}\right)$
be continuous and bounded. The $\left(1-\alpha\right)100\%$ Highest
Density Region is a subset $\mathbf{S}\left(\kappa_{\alpha}\right)$
of $\mathcal{Y}$ defined as $\mathbf{S}\left(\kappa_{\alpha}\right)=\left\{ y:\,f_{Y}\left(y\right)\geq\kappa_{\alpha}\right\} $
where $\kappa_{\alpha}$ is the largest constant that satisfies $\mathbb{P}_{Y}\left(Y\in\mathbf{S}\left(\kappa_{\alpha}\right)\right)\geq1-\alpha$.
\end{defn}
For $s=T\left\Vert \delta_{T}\right\Vert ^{2}\left(\widehat{\lambda}_{b}^{\mathrm{LS}}-\lambda_{b}^{0}\right)$,
the asymptotic distribution theory of \citet{bai:97RES} suggests
a belief $\pi\left(s\right)$ over $s\in\mathbb{R}.$ This belief
function can be used as a Quasi-prior for $\lambda_{b}$ in the definition
of the Quasi-posterior $p_{T}\left(\lambda_{b}\right).$ Let $\mu\left(\lambda_{b}\right)$
denote some density function defined by the Radon-Nikodym equation
$\mu\left(\lambda_{b}\right)=dp_{T}\left(\lambda_{b}\right)/d\lambda_{\mathrm{L}},$
where $\lambda_{\mathrm{L}}$ denotes the Lebesgue measure. The following
algorithm describes how to construct a confidence set for $T_{b}^{0}$.
\begin{lyxalgorithm}
\label{Alg: HQPD}GL HQDR-based Confidence Sets for $T_{b}^{0}$:\\
(1) Estimate by least-squares the break date and the regression coefficients
from model \eqref{Eq Matrix Format of the model};\\
(2) Set the Quasi-prior $\pi\left(\lambda_{b}\right)$ equal to the
probability density of the limiting distribution from Corollary \ref{Corollary part (i) of Theorem LapBai97};
\\
(3) Construct the Quasi-posterior given in \eqref{Eq. Quasi-posterior-1};
\\
(4) Obtain numerically the density $\mu\left(\lambda_{b}\right)$
as explained above and label it by $\widehat{\mu}\left(\lambda_{b}\right)$;\\
(5) Compute the Highest Quasi-Posterior Density (HQPD) region of the
probability distribution $\widehat{p}_{T}\left(\lambda_{b}\right)$
and include the point $T_{b}$ in the level $\left(1-\alpha\right)\%$
confidence set $C_{\mathrm{HQPD}}\left(\mathrm{cv}_{\alpha}\right)$
if $T_{b}$ satisfies Definition \ref{def:Highest-Density-Region}.

If a general Quasi-prior $\pi\left(\lambda_{b}\right)$ is used, one
begins directly with step 3.
\end{lyxalgorithm}
In principle, any Quasi-prior $\pi\left(\lambda_{b}\right)$ satisfying
Assumption \ref{Assumption Prior LapBai97} can be used. Note that
$C_{\mathrm{HQPD}}\left(\mathrm{cv}_{\alpha}\right)$ retains a frequentist
interpretation, since no parametric likelihood function of the data
is required.

\section{\label{Section Models with Multiple Breaks}Models with Multiple
Change-Points}

Following \citet{bai/perron:98}, the multiple linear regression model
with $m$ change-points is
\begin{align*}
y_{t} & =w_{t}'\phi^{0}+z'_{t}\delta_{j}^{0}+e_{t},\qquad\qquad\qquad\left(t=T_{j-1}^{0}+1,\ldots,\,T_{j}^{0}\right)
\end{align*}
for $j=1,\ldots,\,m+1$, where by convention $T_{0}^{0}=0$ and $T_{m+1}^{0}=T$.
There are $m$ unknown break points $\left(T_{1}^{0},\ldots,\,T_{m}^{0}\right)$
and consequently $m+1$ regimes each corresponding to a distinct parameter
value $\delta_{j}^{0}$. The purpose is to estimate the unknown regression
coefficients together with the break points when $T$ observations
on $\left(y_{t},\,w_{t},\,z_{t}\right)$ are available. Many of the
theoretical results follow directly from the single break case; the
break points are asymptotically distinct and thus, given the mixing
conditions, our results for the single break date extend readily to
multiple breaks. More complicated is the computation of the estimates
of the break dates which has been addressed by \citet{bai/perron:03}
who proposed an efficient algorithm based on the principle of dynamic
programming; see also \citet{hawkins:76}.

Let $T_{i}\triangleq\left\lfloor T\lambda_{i}\right\rfloor $ and
$\theta\triangleq\left(\phi',\,\delta'_{1},\,\Delta_{1}^{\prime}\ldots,\,\Delta_{m}^{\prime}\right)'$
where $\Delta_{i}=\delta_{i+1}-\delta_{i}$, $i=1,\ldots,\,m$. The
class $\mathscr{\mathscr{L}}\left(\theta,\,T_{i};\,1\leq i\leq m\right)$
of GL estimators in multiple change-points models relies on the least-squares
criterion function $Q_{T}\left(\delta\left(\boldsymbol{\lambda}_{b}\right),\,\boldsymbol{\lambda}_{b}\right)=\sum_{i=1}^{m+1}\sum_{t=T_{i}-1}^{T_{i}}\left(y_{t}-w'_{t}\phi-z'_{t}\delta_{i}\right)^{2}$,
with $\boldsymbol{\lambda}_{b}\triangleq\left(\lambda_{i};\,1\leq i\leq m\right)$.
In order to state the large-sample properties, we need to introduce
the shrinkage theoretical framework of \citet{bai/perron:98}.
\begin{assumption}
\label{Assumption A1-A6 BP98}(i) Let $x_{t}=\left(w'_{t},\,z'_{t}\right)'$,
$X=\left(x_{1},\ldots x_{T}\right)'$ and $\overline{X}_{0}=\textrm{diag}\left(X_{1}^{0},\ldots,\,X_{m+1}^{0}\right)$
be the diagonal partition of $X$ at $\left(T_{1}^{0},\ldots,T_{m}^{0}\right)$.
For each $i=1,\ldots,m+1$ $\left(X_{1}^{0}\right)'X_{1}^{0}/\left(T_{i}^{0}-T_{i-1}^{0}\right)$
converges to a non-random positive definite matrix not necessarily
the same for all $i$. (ii) Assumption \ref{Assumption A3 Bai97}
holds. (iii) The matrix $\sum_{t=k}^{l}z_{t}z'_{t}$ is invertible
for $l-k\geq q$. (iv) $T_{i}^{0}=\left\lfloor T\lambda_{i}^{0}\right\rfloor $,
where $0<\lambda_{1}^{0}<\cdots<\lambda_{m}^{0}<1$. (v) Let $\Delta_{T,i}=v_{T}\Delta_{i}^{0}$
where $v_{T}>0$ is a scalar satisfying $v_{T}\rightarrow0$ and $T^{1/2-\vartheta}v_{T}\rightarrow\infty$
for some $\vartheta\in\left(0,\,1/4\right)$, and $\mathbb{E}\left\Vert z_{t}\right\Vert ^{2}<C,\,$
$\mathbb{E}\left\Vert e_{t}\right\Vert ^{2/\vartheta}<C$ for some
$C<\infty$ and all $t$.
\end{assumption}

\begin{assumption}
\label{Ass A7 BP98}Let $\Delta T_{i}^{0}=T_{i}^{0}-T_{i-1}^{0}$.
For $i=1,\ldots,\,m+1$, uniformly in $s\in\left[0,\,1\right]$, (a)
$\left(\Delta T_{i}^{0}\right)^{-1}\sum_{t=T_{i-1}^{0}+1}^{T_{i-1}^{0}+\left\lfloor s\Delta T_{i}^{0}\right\rfloor }z_{t}z'_{t}\overset{\mathbb{P}}{\rightarrow}sV_{i}$,
$\left(\Delta T_{i}^{0}\right)^{-1}\sum_{t=T_{i-1}^{0}+1}^{T_{i-1}^{0}+\left\lfloor s\Delta T_{i}^{0}\right\rfloor }e_{t}^{2}\overset{\mathbb{P}}{\rightarrow}s\sigma_{i}^{2}$,
and
\begin{align*}
\left(\Delta T_{i}^{0}\right)^{-1}\sum_{t=T_{i-1}^{0}+1}^{T_{i-1}^{0}+\left\lfloor s\Delta T_{i}^{0}\right\rfloor }\sum_{r=T_{i-1}^{0}+1}^{T_{i-1}^{0}+\left\lfloor s\Delta T_{i}^{0}\right\rfloor }\mathbb{E}\left(z_{t}z'_{r}u_{t}u_{r}\right) & \overset{\mathbb{P}}{\rightarrow}s\Sigma_{i};
\end{align*}

(b) $\left(\Delta T_{i}^{0}\right)^{-1/2}\sum_{t=T_{i-1}^{0}+1}^{T_{i-1}^{0}+\left\lfloor s\Delta T_{i}^{0}\right\rfloor }z_{t}u_{t}\overset{\mathbb{P}}{\rightarrow}\mathscr{G}_{i}\left(s\right)$
where $\mathscr{G}_{i}\left(s\right)$ is a multivariate Gaussian
process on $\left[0,\,1\right]$ with mean zero and covariance $\mathbb{E}\left[\mathscr{G}_{i}\left(s\right)\mathscr{G}_{i}\left(u\right)\right]=\min\left\{ s,\,u\right\} \Sigma_{i}$.
\end{assumption}
Next, for $i=1,\ldots,\,m$, define $\Xi_{Z,i}=\left(\Delta_{i}^{0}\right)'V_{i+1}\Delta_{i}^{0}/\left(\Delta_{i}^{0}\right)'V_{i}\Delta_{i}^{0},$
$\Xi_{e,i}^{2}=\left(\Delta_{i}^{0}\right)'\Sigma_{i+1}\Delta_{i}^{0}/\left(\Delta_{i}^{0}\right)'\Sigma_{i}\Delta_{i}^{0}$,
and let $W_{1}^{\left(i\right)}\left(s\right)$ and $W_{2}^{\left(i\right)}\left(s\right)$
be independent Wiener processes defined on $[0,\,\infty)$, starting
at 0 when $s=0$; $W_{1}^{\left(i\right)}\left(s\right)$ and $W_{2}^{\left(i\right)}\left(s\right)$
are also independent over $i.$ Finally, define
\begin{align}
\mathscr{V}^{\left(i\right)}\left(s\right)\triangleq\mathscr{W}^{\left(i\right)}\left(s\right)-\varLambda_{i}^{0}\left(s\right) & \triangleq\begin{cases}
2\left(\left(\Delta_{i}^{0}\right)'\Sigma_{i}\Delta_{i}\right)^{1/2}W_{1}^{\left(i\right)}\left(-s\right)-\left|s\right|\left(\Delta_{i}^{0}\right)'V_{i}\Delta_{i}, & \textrm{if }s\leq0\\
2\left(\left(\Delta_{i}^{0}\right)'\Sigma_{i+1}\delta^{0}\right)^{1/2}W_{2}^{\left(i\right)}\left(s\right)-s\left(\Delta_{i}^{0}\right)'V_{i+1}\Delta_{i}, & \textrm{if }s>0.
\end{cases}\label{eq. V(s) Limit Process Theorem Post Mean BP98}
\end{align}
 We now extend the notation of Section \ref{Section Asymptotic Results LapBai97}
to the present context. By redefining the Quasi-posterior $p\left(\boldsymbol{\lambda}_{b}\right)$
in terms of $\boldsymbol{\lambda}_{b},$ the GL estimator as the minimizer
of the associated risk function {[}recall \eqref{Eq. Expected Risk function-1}{]},
$\widehat{\boldsymbol{\lambda}}_{b}^{\mathrm{GL}}=\arg\min_{s\in\varGamma^{0}}\left[\mathcal{R}_{l,T}\left(s\right)\right],$
where now $\varGamma^{0}=\mathbf{B}_{1}\times\ldots\times\mathbf{B}_{m}$,
with $\mathbf{B}_{i}$ a compact subset of $\left(0,\,1\right)$.
The sets $\mathbf{B}_{i}$ are disjoint and satisfy $\sup_{\lambda\in\mathbf{B}_{i}}<\inf_{\lambda\in\mathbf{B}_{i+1}}$
for all $i.$
\begin{assumption}
\label{Ass Gaussian BP98}Assumptions \ref{Assumption The-loss-function LapBai97}-\ref{Assumption Prior LapBai97}
hold with obvious modifications to allow for the multidimensional
parameter $\boldsymbol{\lambda}_{b}\in\varGamma^{0}.$ Assumption
\ref{Assumption Gaussian Process for Lap LapBai97} holds where now
in part (i) $\boldsymbol{\lambda}_{b}$ replaces $\lambda_{b}$, and
in part (ii) $\varSigma^{\left(i\right)}\left(\cdot,\,\cdot\right)$
($1\leq i\leq m+1$) replaces $\varSigma\left(\cdot,\,\cdot\right)$
and is defined analogously for each regime.
\end{assumption}
Assumption \ref{Assumption Uniquess loss Fucntion LapBai97} implies
that $\xi_{l,i}^{0}\triangleq\xi\left(\lambda_{i}^{0}\right)$ is
uniquely defined by $\Psi_{l}\left(\xi_{l,i}^{0}\right)\triangleq\inf_{s}\Psi_{l,i}\left(s\right)=\inf_{s}\int_{\mathbb{R}}l\left(s-u\right)\left(\exp\left(\mathscr{V}^{\left(i\right)}\left(u\right)\right)/\left(\int_{\mathbb{R}}\exp\left(\mathscr{V}^{\left(i\right)}\left(w\right)\right)dw\right)\right)du$.
 The GL estimator is defined as the minimizer of
\begin{align*}
\mathcal{R}_{l,T} & \triangleq\int_{\varGamma^{0}}l\left(s-\boldsymbol{\lambda}_{b}\right)\frac{\exp\left(-Q_{T}\left(\delta\left(\boldsymbol{\lambda}_{b}\right),\,\boldsymbol{\lambda}_{b}\right)\right)\pi\left(\boldsymbol{\lambda}_{b}\right)}{\int_{\varGamma^{0}}\exp\left(-Q_{T}\left(\delta\left(\boldsymbol{\lambda}_{b}\right),\,\boldsymbol{\lambda}_{b}\right)\right)\pi\left(\boldsymbol{\lambda}_{b}\right)d\boldsymbol{\lambda}_{b}}d\boldsymbol{\lambda}_{b}.
\end{align*}
 The analysis is now in terms of the $m\times1$ local parameter $u$
with components $u_{i}=T\left\Vert \Delta_{T,i}\right\Vert ^{2}(\lambda_{i}-$
$\lambda_{i,T}^{0}\left(v\right))$, with $\lambda_{i,T}^{0}\left(v\right)=\lambda_{i,T}^{0}\left(\theta^{0}+v/r_{T}\right)$.

Theorem \ref{Theorem Posterior Mean Lap Estimation in BP98}-\ref{Theorem Geneal Laplace Estimator LapBP98}
extend corresponding results from Theorem \ref{Theorem Posterior Mean Lap Estimation in Bai 97}-\ref{Theorem Geneal Laplace Estimator LapBai97},
respectively, to multiple change-points. The fast rate of convergence
implies that asymptotically the behavior of the GL estimator only
matters in a small neighborhood of each $T_{i}^{0}$. Since each such
neighborhood increases at rate $1/v_{T}$ while $T\rightarrow\infty$
at a faster rate, given the mixing conditions, these are asymptotically
distinct and the limiting distribution is then similar to that in
the single break case. This is the same argument underlying the analysis
of \citet{bai/perron:98} and of \citet{ibragimov/has:81}. The same
comments as those in Section \ref{Section Asymptotic Results LapBai97}
apply.
\begin{condition}
\label{Condition 1 LapBP98}For $1\leq i\leq m$ there exist positive
finite numbers $\kappa_{\gamma,i}$ such that $\gamma_{T}/T\left\Vert \Delta_{T,i}\right\Vert ^{2}\rightarrow\kappa_{\gamma,i}$.
\end{condition}
\begin{thm}
\label{Theorem Posterior Mean Lap Estimation in BP98}Assume $l\left(\cdot\right)$
is the squared loss function. Under Assumption \ref{Assumption A1-A6 BP98}-\ref{Ass Gaussian BP98}
and Condition \ref{Condition 1 LapBP98}, we have in $\mathbb{D}_{b}$,
\begin{align}
T\left\Vert \Delta_{T,i}\right\Vert ^{2}\left(\widehat{\lambda}_{i}^{\mathrm{GL}}-\lambda_{i}^{0}\right) & \Rightarrow\frac{\int u\exp\left(\mathscr{W}^{\left(i\right)}\left(u\right)-\varLambda_{i}^{0}\left(u\right)\right)du}{\int\exp\left(\mathscr{W}^{\left(i\right)}\left(u\right)-\varLambda_{i}^{0}\left(u\right)\right)du}.\label{eq. Asy Dist Laplace Bai97 part (ii)-1}
\end{align}
\end{thm}
Turning to the general case of loss functions satisfying Assumption
\ref{Assumption The-loss-function LapBai97}, Theorem \ref{Theorem Geneal Laplace Estimator LapBP98}
shows that the random quantity $T\left\Vert \delta_{T}\right\Vert ^{2}\left(\widehat{\lambda}_{i}^{\mathrm{GL}}-\lambda_{i}^{0}\right)$
is (first-order) asymptotically equivalent to the random variable
$\xi_{l,i}^{0}$ determined by
\begin{align}
\Psi_{l}\left(\xi_{l,i}^{0}\right) & \triangleq\inf_{r}\Psi_{l,i}\left(r\right)=\inf_{r\in\mathbb{R}}\left\{ \int_{\mathbb{R}}l\left(r-u\right)\frac{\exp\left(\mathscr{W}^{\left(i\right)}\left(u\right)-\varLambda_{i}^{0}\left(u\right)\right)}{\int\exp\left(\mathscr{W}^{\left(i\right)}\left(u\right)-\varLambda_{i}^{0}\left(u\right)\right)du}du\right\} .\label{Eq. Distribution of Bayes Estimator-1-1}
\end{align}

\begin{thm}
\label{Theorem Geneal Laplace Estimator LapBP98}Under Assumptions
\ref{Assumption A1-A6 BP98}-\ref{Ass Gaussian BP98} and Condition
\ref{Condition 1 LapBP98}, for $l\in\boldsymbol{L},$ $T\left\Vert \Delta_{T,i}\right\Vert ^{2}\left(\widehat{\lambda}_{i}^{\mathrm{GL}}-\lambda_{i}^{0}\right)\Rightarrow\xi_{l,i}^{0},$
as defined by \eqref{Eq. Distribution of Bayes Estimator-1-1}.
\end{thm}
A direct consequence of the results of this section is that statistical
inference for the break dates $T_{i}^{0}$ $\left(i=1,\ldots,\,m\right)$
can be carried out using the same methods for the single break case.


\section{\label{Section Theoretical-Properties-of GL Inference}Theoretical
Properties of GL Inference}

This section shows that the GL-HPDR confidence sets are bet-proof.
The betting framework and the notion of bet-proofness are useful to
study the properties of frequentist inference in non-regular problems.
The literature concluded that frequentist confidence sets may exhibit
undesirable properties in non-regular problems {[}e.g., \citet{buehler:59},
\citet{cornfield:69}, \citet{cox:58}, \citet{mueller/norest:16}
\citet{pierce:73}, \citet{robinson:77} and \citet{wallace:59}{]}.
For example, the confidence sets can be too short or empty with positive
probability. This arises because frequentist procedures often have
the property that, conditional on a sample point lying in some subset
of the sample space, the conditional confidence level  is less than
the unconditional confidence level uniformly in the parameters.

We use the same betting framework as in \citet{buehler:59}. Let $\mathbb{P}\left(\cdot|\,\lambda_{b}\right)$
denote the likelihood of the data $Y\in\mathcal{Y}$ conditional on
$\lambda_{b}\in\varGamma^{0}$. Assume $\mathbb{P}\left(\cdot|\,\lambda_{b}\right)$
has density $p\left(\cdot|\,\lambda_{b}\right)$ with respect to a
finite measure $\zeta.$ We define a $1-\alpha$ confidence set by
a rejection probability rule $\varphi:\,\varGamma^{0}\times\mathcal{Y}\mapsto\left[0,\,1\right]$
satisfying $\int\left[1-\varphi\left(\lambda_{b},\,y\right)\right]p\left(y|\,\lambda_{b}\right)d\zeta\left(y\right)\geq1-\alpha$,
with $\varphi\left(\lambda_{b},\,y\right)$ the probability that $\lambda_{b}$
is not included in the set when $y$ is observed. For any realization
of the data $Y=y$, an inspector can choose to object to the confidence
set $\varphi$. The inspector\textquoteright s objection $\widetilde{b}:\mathcal{Y}\mapsto\left[0,\,1\right]$
takes value 1 if there is an objection. Denote by $\mathbf{B}$ the
set of all measurable strategies $\widetilde{b}$. When $\widetilde{b}=1$
the inspector receives 1 if $\varphi$ does not contain $\lambda_{b}$,
and she loses $\alpha/\left(1-\alpha\right)$ otherwise. For a given
parameter $\lambda_{b}$ and betting strategy $\widetilde{b}$, the
inspector's expected loss is,
\begin{align*}
L_{\alpha}\left(\varphi,\,\widetilde{b},\,\lambda_{b}\right) & =\frac{1}{1-\alpha}\int\left[\alpha-\varphi\left(\lambda_{b},\,y\right)\right]\widetilde{b}\left(y\right)p\left(y|\lambda_{b}\right)d\zeta\left(y\right).
\end{align*}
 A confidence set $\varphi$ is said to be bet-proof at level $1-\alpha$
if for each $\widetilde{b}\in\mathbf{B}$, $L_{\alpha}\left(\varphi,\,\widetilde{b},\,\lambda_{b}\right)\geq0$
for some $\lambda_{b}\in\varGamma^{0}$. If there exists a strategy
$\widetilde{b}$ such that $L_{\alpha}\left(\varphi,\,\widetilde{b},\,\lambda_{b}\right)<0$
for all $\lambda_{b}\in\varGamma^{0}$, then the inspector would be
right on average and would make positive expected profits. Hence,
such $\varphi$ would be an ``unreasonable'' confidence set.
Without loss of substance, we restrict our attention to a change in
the mean of a sequence of i.i.d. Gaussian variables. Let $y_{t}=\delta_{T}\mathbf{1}\left\{ t>T_{b}^{0}\right\} +e_{t},$
where $e_{t}\sim i.i.d.\,\mathrm{\mathscr{N}\left(0,\,1\right)}.$
The result below can also be shown to hold for fixed shifts $\delta_{T}=\delta^{0}$.
For ease of exposition, we assume $\delta^{0}$ known. The general
case leads to similar results, with more lengthy derivations without
any gain in intuition.

Recall that $\varphi$ is such that the Quasi-posterior probability
$p_{T}\left(\lambda_{b}|\,y\right)=p_{T}\left(\lambda_{b}\right)$
of excluding $\lambda_{b}$ is less than or equal to $\alpha,$
\begin{align}
\int\varphi\left(\lambda_{b},\,y\right)p_{T}\left(\lambda_{b}|\,y\right)d\lambda_{b} & \leq\alpha\quad\mathrm{for\,all\,}y\in\mathcal{Y}.\label{eq (def Quasi-credible set)}
\end{align}

\begin{prop}
\label{Proposition Bet Proof}Under Assumptions \ref{Assumption A Bai97}-\ref{Assumption A4 Bai97}
and \ref{Assumption The-loss-function LapBai97}-\ref{Assumption Uniquess loss Fucntion LapBai97},
and Condition \ref{Condition 1 LapBai97}, for $l\in\boldsymbol{L}:$
(i) $\varphi$ is bet-proof at level $1-\alpha$; (ii) If \eqref{eq (def Quasi-credible set)}
holds with equality, then $\varphi$ is the shortest confidence set
in the class of level $1-\alpha$ confidence sets, i.e., there cannot
exist a level $1-\alpha$ confidence set $\varphi'$ with the property
that, for all $y\in\mathcal{Y}$ $\int\varphi'\left(\lambda_{b},\,y\right)d\lambda_{b}\geq\int\varphi\left(\lambda_{b},\,y\right)d\lambda_{b}$,
and for all $y\in\mathcal{Y}_{0}$ with $\zeta\left(\mathcal{Y}_{0}\right)>0,$
$\int\varphi'\left(\lambda_{b},\,y\right)d\lambda_{b}>\int\varphi\left(\lambda_{b},\,y\right)d\lambda_{b}$.
\end{prop}
Part of the proof shows that the Quasi-posterior is asymptotically
equivalent (in total variation distance) to the Bayesian posterior.
Given the conservativeness allowed by Definition \ref{def:Highest-Density-Region},
the GL confidence interval is asymptotically a superset of a Bayesian
credible interval. Bet-proofness is a useful criterion in change-point
models where popular inference methods face some difficulties, as
shown in the next section. Proposition \ref{Proposition Bet Proof}
suggests that GL inference should not suffer from these issues; the
simulations in the next section will confirm that this is indeed the
case.

\section{\label{Section Monte-Carlo-Simulation}Finite-Sample Evaluations}

The purpose of this section is twofold. Section \ref{subsec:Precision-of-the}
assesses the accuracy of the GL estimate of the change-point while
Section \ref{subsec:Properties-of-the} evaluates the small-sample
properties of the proposed method to construct confidence sets. We
consider DGPs that take the form:
\begin{align}
y_{t}=D{}_{t}\alpha^{0}+Z{}_{t}\beta^{0}+Z{}_{t}\delta^{0}\boldsymbol{1}\left\{ t>T_{b}^{0}\right\} +e_{t}, & \qquad\qquad t=1,\ldots,\,T,\label{Eq. DGP Simulation Study}
\end{align}
with a sample size $T=100.$ Three versions of \eqref{Eq. DGP Simulation Study}
are investigated: M1 involves a break in mean: $Z_{t}=1$, $D_{t}$
absent, and $e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right)$; M2
is similar to M1 but with $e_{t}=0.3e_{t-1}+u_{t}$, $u_{t}\sim\mathscr{N}\left(0,\,1\right)$;
M3 is a dynamic model with $D_{t}=y_{t-1}$, $Z_{t}=1$, $e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,0.5\right)$
and $\alpha^{0}=0.6$. We set $\beta^{0}=1$ in M1-M2 and $\beta^{0}=0$
in M3. We consider $\lambda_{0}=0.3$ and 0.5, and break magnitudes
$\delta^{0}=0.3,\,0.4,\,0.6$ and $1$. Additional simulations are
presented in the supplement.

\subsection{\label{subsec:Precision-of-the}Precision of the Change-point Estimate}

We consider the following estimators of $T_{b}^{0}$: the least-squares
estimator (OLS), the GL estimator under a least-absolute loss function
 (GL-LN); the GL estimator under a least-absolute loss function
with a uniform prior (GL-Uni). We compare the mean absolute error
(MAE), standard deviation (Std), root-mean-squared error (RMSE), and
the 25\% and 75\% quantiles. We set the trimming parameter $\epsilon$
equal to 0.05. As explained in \citet{casini/perron_Lap_CR_Single_Inf},
the trimming $\epsilon$ should not be chosen too high because otherwise
the estimate might tend to overestimate (resp. underestimate) the
break date if it is in the first (resp. second) half of the sample.
They found that $\epsilon=0.05$ performs well for different locations
of the break date and this is also confirmed in the simulations in
this section. See Section 1 and 5 in \citet{casini/perron_Lap_CR_Single_Inf}
for more discussion.

Tables \ref{Table M0 Bias}-\ref{Table M5 Bias} present the results.
When the magnitude of the break is small, the OLS estimator displays
quite large MAE, which increases as the change-point point moves
toward the tails. In contrast, the GL estimator shows substantially
lower MAE uniformly over break magnitudes and break locations.
In addition, the GL estimator has smaller variance as well as lower
RMSE compared to the OLS estimator. Notably, the distribution of
GL-LN concentrates a higher fraction of the mass around the mid-sample
relative to the finite-sample distribution of the OLS estimate. This
is mainly due to the fact that the Quasi-posterior essentially does
not share the marked trimodality of the finite-sample distribution
{[}cf. \citet{casini/perron_CR_Single_Break}{]}. When the break magnitude
is small, the objective function is quite flat with a small peak at
the OLS estimate. The Quasi-posterior has higher mass close to the
OLS estimate\textemdash which corresponds to the middle mode\textemdash and
accordingly lower mass in the tails. The GL estimator that uses the
uniform prior (GL-Uni) is also more precise than the OLS estimator,
though the margin is smaller. The latter is due to the fact that
the GL estimate uses information only from the OLS objective function.
We have not reported the bias. However, here is a summary of its behavior
which can also be learned from Figures \ref{Fig1}-\ref{Fig2}. When
$\lambda_{b}^{0}=0.5$, the bias is small and close to zero because
the finite-sample distributions of the estimators are symmetric. When
$\lambda_{b}^{0}<0.5$, the bias is positive which means that the
break date estimators tend to be on the right of $\lambda_{b}^{0}.$
The opposite hold for $\lambda_{b}^{0}>0.5$.\textbf{ }

\subsection{\label{subsec:Properties-of-the}Properties of the GL Confidence
Sets}

We now assess the performance of the suggested inference procedure
for the break date. We compare it with the following existing methods:
Bai's (1997) approach, \citeauthor{elliott/mueller:07}\textquoteright s
(2007) approach based on\textcolor{red}{{} }inverting a sequence of
locally best invariant tests using \citeauthor{nyblom:89}'s (1989)
statistic, the inverted likelihood-ratio (ILR) method of \citet{eo/morley:15}
which inverts the likelihood-ratio test of \citet{qu/perron:07} and
the HDR method proposed in \citet{casini/perron_CR_Single_Break}
based on continuous record asymptotics, labelled OLS-CR. These methods
have been discussed in detail in \citet{casini/perron_CR_Single_Break}
and in \citet{chang/perron:18}. We can summarize their properties
as follows. The confidence intervals obtained from Bai's (1997) method
display empirical coverage rates often below the nominal level when
the size of the break is small. In general, \citeauthor{elliott/mueller:07}\textquoteright s
(2007) approach achieves the most accurate coverage rates but the
average length of the confidence sets is always substantially larger
relative to other methods.\footnote{This problem is more severe when the errors are serially correlated
or the model includes lagged dependent variables.\nocite{casini_hac}
Regarding the former, this in part may be due to issues with \citeauthor{newey/west:87}
HAC-type estimators when there are breaks {[}see \citeauthor{casini_CR_Test_Inst_Forecast}
(2018, 2019), \nocite{casini_hac} \citet{casini/perron_Low_Frequency_Contam_Nonstat:2020},
\citet{casini/perron_Oxford_Survey}, \citet{chang/perron:18}, \citet{crainiceanu/vogelsang:07},
\citet{deng/perron:06}, \citet{fossati:17}, \citet{juhl/xiao:09},
\citet{kim/perron:09}, \citet{martins/perron:16}, \citet{perron/yamamoto:18}
and \citet{vogeslang:99}{]}.} In addition, this approach breaks down in models with serially correlated
errors or lagged dependent variables, whereby the length of the confidence
set approaches the whole sample as the magnitude of the break increases.
The ILR has coverage rates often above the nominal level and an average
length significantly longer than with the OLS-CR method when the magnitude
of the shift is small. Here, we shall show that the GL inference performs
well in terms of coverage probability compared with the other methods
and is characterized by shorter lengths of the confidence sets.

When the errors are uncorrelated (i.e., M1 and M3) we simply estimate
variances rather than long-run variances. The least-squares estimation
method is employed with a trimming parameter $\epsilon=0.15$ and
we use the required degrees of freedom adjustment for the statistic
$\widehat{\textrm{U}}_{T}$ of \citet{elliott/mueller:07}. To construct
the OLS-CR method, we follow the steps outlined in \citet{casini/perron_CR_Single_Break}.
To implement Bai's (1997) method we use the usual steps described
in \citet{bai:97RES} and \citet{bai/perron:98}. We implement the
GL estimator using a least-absolute loss with the prior from Corollary
\ref{Corollary part (i) of Theorem LapBai97}. For model M2, the
estimate of the long-run variance is the pre-whitened heteroskedasticity
and autocorrelation (HAC) estimator of \citet{andrews/monahan:92}.
We consider the version $\widehat{\textrm{U}}_{T}$ proposed by \citet{elliott/mueller:07}
that allows for heterogeneity across regimes; using the restricted
version when applicable leads to similar results. Finally, the last
row of each panel includes the rejection probability of the 5\%-level
sup-Wald test using the asymptotic critical value of \citet{andrews:93};
it serves as a statistical measure of the magnitude of the break.

Overall, the results in Table \ref{Table M0}-\ref{Table M2} confirm
previous findings about the performance of existing methods. Bai's
(1997) method has a coverage rate below the nominal level when the
size of the break is small. For example, in model M2, with $\lambda_{0}=0.5$
and $\delta^{0}=0.8$, it has a coverage probability below 82\% even
though the Sup-Wald test rejects roughly for 70\% of the samples.
With smaller break sizes, it systematically fails to cover the true
break date with correct probability. In contrast, the method of \citet{elliott/mueller:07}
yields very accurate empirical coverage rates. However, the average
length of the confidence intervals obtained is systematically much
larger than those from all other methods across all DGPs, break sizes
and break locations. For large break sizes, Bai's (1997) method delivers
good coverage rates and the shortest average length among all methods.

The GL method displays good coverage rates across different break
magnitudes and tends to have the shortest lengths among all methods
for all break magnitudes, except for $\delta^{0}=1.6$ in model M2
for which Bai's (1997) confidence interval is slightly shorter.
In Model M3, the coverage rates of OLS-CR are more accurate than those
with the GL method although the difference is not large. Thus the
GL method strikes a good balance between adequate coverage probability
and short average lengths, thus confirming the theoretical results
on bet-proofness. This is also consistent with Figures \ref{Fig1}-\ref{Fig2}
which show that the asymptotic distribution of the GL estimator does
not underestimate uncertainty about the break location even when the
break magnitude is small thereby yielding good coverage rates also
in this case. In model M1 the GL method leads to shorter lengths than
Bai's even for large breaks. This is not in contradiction with Figure
\ref{Fig2} because model M1 is a simpler model than that reported
in the figure which shows that the density of the asymptotic distribution
of the GL estimator is more spread out than that from \citet{bai:97RES}.

Non-reported simulations show that the GL method is robust to heteroskedastic
errors $e_{t}=\left|z_{t}\right|u_{t}$ and non-normal errors. The
case of multiple breaks is not considered since they are expected
to be similar as in the single break case by virtue of the assumption
that the break dates are sufficiently separated. Finally, in the supplement
we compare the GL method above with its continuous record counterparts
developed in \citet{casini/perron_Lap_CR_Single_Inf}. Overall, we
find that both estimation and confidence intervals based on GL-LN
perform well relative to the continuous record counterparts, where
significant gains appear to occur when there is high serial correlation
in the errors. See the additional results reported in the supplement.

\section{\label{Section Conclusions}Conclusions}

We developed large-sample results for a class of Generalized Laplace
estimators in multiple change-points models where popular methods
face some challenges due to the non-regularities of the problem. The
GL method exploits the insight of Laplace who proposed to generate
a density from taking an exponential transformation of a least-squares
criterion. The class of GL estimators exhibits a dual limiting distribution;
namely, the classical shrinkage asymptotic distribution of \citet{bai/perron:98},
or a Bayes-type asymptotic distribution {[}cf. \citet{ibragimov/has:81}{]}.
Simulations show that the GL estimator is more accurate than OLS.
Similarly, inference has superior finite-sample properties relative
to popular methods and these properties are shown to be supported
by theoretical results. Since the issues about the finite-sample performance
of OLS especially for small breaks continue to hold in more complex
structural change models, we believe that our method can be usefully
extended to those models. For example, the GL approach can be immediately
applied to nonlinear models (e.g., instrumental variable models, linear
model with restrictions, nonlinear regression models, etc.) even though
particular attention to the appropriate choice of the prior should
be given in each context. We believe that our approach can also be
relevant for high-dimensional regression with structural changes although
this would require a careful consideration of additional aspects related
to the growing number of regressors.

\newpage{}

\section{Supplementary Material}

Casini, A. and P. Perron (2020c). Supplement to \textquotedblleft Generalized
Laplace Inference in Multiple Change-points Models\textquotedblright ,
Econometric Theory Supplementary Material. To view, please visit:
{[}{[}doi will be inserted here by typesetter{]}{]}\nocite{casini/perron_CR_Single_Break,casini/perron_Lap_CR_Single_Inf,casini/perron_Low_Frequency_Contam_Nonstat:2020,casini/perron_PrewhitedHAC,casini/perron_SC_BP_Lap,casini_diss,casini_hac}

\begin{singlespace}

\bibliographystyle{chicago}
\bibliography{References_JoE}

\addcontentsline{toc}{section}{References}

\end{singlespace}

\newpage{}

\setcounter{page}{1} 
\begin{singlespace}

\begin{center}
\begin{figure}[h]
\includegraphics[width=18cm,height=6.5cm]{Fig1_Bai_Lap_ET}

{\footnotesize{}\caption{{\scriptsize{}\label{Fig1}The probability density of the LS estimator
for the model $y_{t}=\mu^{0}+Z_{t}\delta_{1}^{0}+Z_{t}\delta_{2}^{0}\mathbf{1}\left\{ t>\left\lfloor T\lambda_{0}\right\rfloor \right\} +e_{t},\,Z_{t}=0.3Z_{t-1}+u_{t}-0.1u_{t-1},\,u_{t}\sim\textrm{i.i.d.}\mathscr{N}\left(0,\,1\right),\,e_{t}\sim\textrm{i.i.d.}\mathscr{N}\left(0,\,1\right),$
$\left\{ u_{t}\right\} $ independent from $\left\{ e_{t}\right\} ,$
$T=100$ with $\delta^{0}=0.3$ and $\lambda_{0}=0.25$ and $0.5$
(the left and right panel, respectively). The black dotted line is
the density of the asymptotic distribution from Bai (1997), the red
broken line break is the density of the finite-sample distribution
of the LS estimator, the green broken line is the density of the finite-sample
distribution of the GL estimator, and the blue broken line is the
density of the asymptotic distribution of the GL estimator. }}
}{\footnotesize\par}
\end{figure}
\end{center}

\setlength{\belowcaptionskip}{-0pt}

\begin{center}
\begin{figure}[h]
\includegraphics[width=18cm,height=6.5cm]{Fig2_Bai_Lap_ET}

{\footnotesize{}\caption{{\scriptsize{}\label{Fig2}The descriptions and comments given in
Figure \ref{Fig1} apply but with a break magnitude $\delta^{0}=1.5$. }}
}{\footnotesize\par}
\end{figure}
\end{center}

\end{singlespace}

\newpage{}

\setcounter{page}{1} \begin{center}
\begin{table}[H]
\caption{\label{Table M0 Bias}Small-sample accuracy of the estimate of the
break point $T_{b}^{0}$ for model M1}

\begin{singlespace}
\begin{centering}
{\footnotesize{}}
\begin{tabular}{ccccccc|ccccc}
\hline
 &  & {\footnotesize{}MAE} & {\footnotesize{}Std} & {\footnotesize{}$\textrm{RMSE}$} & {\footnotesize{}$Q_{0.25}$} & \multicolumn{1}{c}{{\footnotesize{}$Q_{0.75}$}} & {\footnotesize{}MAE} & {\footnotesize{}Std} & {\footnotesize{}$\textrm{RMSE}$} & {\footnotesize{}$Q_{0.25}$} & {\footnotesize{}$Q_{0.75}$}\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
 &  & \multicolumn{5}{c|}{{\footnotesize{}$\lambda_{0}=0.3$}} & \multicolumn{5}{c}{{\footnotesize{}$\lambda_{0}=0.5$}}\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
{\footnotesize{}$\delta^{0}=0.3$} & {\footnotesize{}OLS} & 21.99 & 27.51 & 30.53 & 24 & 66 & 21.51 & 26.85 & 26.79 & 34 & 71\tabularnewline
 & {\footnotesize{}GL-LN} & 13.44 & 15.03 & 18.99 & 28 & 54 & 11.85 & 14.51 & 14.93 & 38 & 60\tabularnewline
 & {\footnotesize{}GL-Uni} & 17.56 & 22.88 & 25.51 & 26 & 56 & 16.90 & 22.03 & 22.13 & 38 & 61\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
{\footnotesize{}$\delta^{0}=0.4$} & {\footnotesize{}OLS} & 20.48 & 26.30 & 28.51 & 23 & 57 & 15.64 & 21.79 & 21.23 & 40 & 61\tabularnewline
 & {\footnotesize{}GL-LN} & 13.02 & 15.52 & 18.29 & 29 & 51 & 9.46 & 11.84 & 12.30 & 44 & 56\tabularnewline
 & {\footnotesize{}GL-Uni} & 17.68 & 22.30 & 24.64 & 27 & 54 & 12.38 & 17.69 & 17.15 & 42 & 57\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
{\footnotesize{}$\delta^{0}=0.6$} & {\footnotesize{}OLS} & 13.04 & 20.82 & 15.92 & 28 & 41 & 11.06 & 16.05 & 16.89 & 45 & 55\tabularnewline
 & {\footnotesize{}GL-LN} & 9.20 & 13.67 & 13.67 & 28 & 40 & 7.04 & 9.92 & 10.46 & 47 & 53\tabularnewline
 & {\footnotesize{}GL-Uni} & 11.49 & 18.59 & 14.23 & 27 & 39 & 9.11 & 13.92 & 13.48 & 45 & 55\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
{\footnotesize{}$\delta^{0}=1$} & {\footnotesize{}OLS} & 3.49 & 4.61 & 4.61 & 28 & 32 & 2.92 & 5.24 & 5.23 & 48 & 52\tabularnewline
 & {\footnotesize{}GL-LN} & 3.41 & 4.53 & 4.52 & 28 & 32 & 2.89 & 5.44 & 5.20 & 49 & 51\tabularnewline
 & {\footnotesize{}GL-Uni} & 3.63 & 4.56 & 4.61 & 28 & 32 & 2.90 & 5.21 & 5.22 & 48 & 52\tabularnewline
\hline
\end{tabular}{\footnotesize\par}
\par\end{centering}
\end{singlespace}
\noindent\begin{minipage}[t]{1\columnwidth}
{\scriptsize{}The model is $y_{t}=\delta^{0}\mathbf{1}\left\{ t>\left\lfloor T\lambda_{0}\right\rfloor \right\} +e_{t},\,e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right),\,T=100$.
The columns refer to Mean Absolute Error (MAE), standard deviation
(Std), Root Mean Squared Error (RMSE) and the 25\% and 75\% empirical
quantiles. OLS is the least-squares estimator; GL-LN is the GL estimator
under a least-absolute loss function with the density of the long-span
asymptotic distribution as the prior; GL-Uni is the GL estimator under
a least-absolute loss function with a uniform prior. The number of
simulations is 3,000.}
\end{minipage}
\end{table}
\par\end{center}

\begin{center}

\par\end{center}\begin{center}
\begin{table}[H]
\caption{\label{Table M1 Bias}Small-sample accuracy of the estimates of the
break point $T_{b}^{0}$ for model M2}

\begin{singlespace}
\begin{centering}
{\footnotesize{}}
\begin{tabular}{ccccccc|ccccc}
\hline
 &  & {\footnotesize{}MAE} & {\footnotesize{}Std} & {\footnotesize{}$\textrm{RMSE}$} & {\footnotesize{}$Q_{0.25}$} & \multicolumn{1}{c}{{\footnotesize{}$Q_{0.75}$}} & {\footnotesize{}MAE} & {\footnotesize{}Std} & {\footnotesize{}$\textrm{RMSE}$} & {\footnotesize{}$Q_{0.25}$} & {\footnotesize{}$Q_{0.75}$}\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
 &  & \multicolumn{5}{c|}{{\footnotesize{}$\lambda_{0}=0.3$}} & \multicolumn{5}{c}{{\footnotesize{}$\lambda_{0}=0.5$}}\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
{\footnotesize{}$\delta^{0}=0.3$} & {\footnotesize{}OLS} & 26.61 & 22.85 & 33.03 & 23 & 76 & 24.09 & 28.29 & 28.08 & 23 & 73\tabularnewline
 & {\footnotesize{}GL-LN} & 19.33 & 10.17 & 24.87 & 29 & 61 & 16.01 & 18.78 & 19.81 & 29 & 62\tabularnewline
 & {\footnotesize{}GL-Uni} & 24.76 & 21.05 & 31.34 & 26 & 70 & 20.93 & 25.37 & 25.39 & 28 & 65\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
{\footnotesize{}$\delta^{0}=0.4$} & {\footnotesize{}OLS} & 23.10 & 27.99 & 30.85 & 21 & 68 & 20.47 & 25.55 & 25.54 & 33 & 70\tabularnewline
 & {\footnotesize{}GL-LN} & 16.59 & 18.59 & 22.75 & 29 & 60 & 13.68 & 17.06 & 17.12 & 38 & 61\tabularnewline
 & {\footnotesize{}GL-Uni} & 21.51 & 25.87 & 28.83 & 24 & 61 & 17.91 & 22.94 & 22.91 & 37 & 62\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
{\footnotesize{}$\delta^{0}=0.6$} & {\footnotesize{}OLS} & 17.64 & 23.51 & 25.01 & 24 & 50 & 15.51 & 20.93 & 20.91 & 41 & 59\tabularnewline
 & {\footnotesize{}GL-LN} & 13.42 & 16.63 & 18.63 & 28 & 47 & 11.06 & 14.90 & 14.38 & 46 & 54\tabularnewline
 & {\footnotesize{}GL-Uni} & 16.01 & 21.54 & 22.75 & 25 & 47 & 13.92 & 19.11 & 19.91 & 40 & 58\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
{\footnotesize{}$\delta^{0}=1$} & {\footnotesize{}OLS} & 8.71 & 15.87 & 15.79 & 27 & 34 & 7.24 & 10.73 & 10.72 & 47 & 54\tabularnewline
 & {\footnotesize{}GL-LN} & 8.25 & 15.27 & 15.61 & 27 & 34 & 6.88 & 9.21 & 9.19 & 47 & 52\tabularnewline
 & {\footnotesize{}GL-Uni} & 8.65 & 14.96 & 15.21 & 27 & 33 & 6.96 & 10.44 & 10.45 & 46 & 53\tabularnewline
\hline
\end{tabular}{\footnotesize\par}
\par\end{centering}
\end{singlespace}
\noindent\begin{minipage}[t]{1\columnwidth}
{\scriptsize{}The model is $y_{t}=\delta^{0}\mathbf{1}\left\{ t>\left\lfloor T\lambda_{0}\right\rfloor \right\} +e_{t},\,e_{t}=0.3e_{t-1}+u_{t},\,u_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right),\,T=100$.
The notes of Table \ref{Table M0 Bias} apply.}
\end{minipage}
\end{table}
\par\end{center}

\begin{center}

\par\end{center}\begin{center}
\begin{table}[H]
\caption{\label{Table M5 Bias}Small-sample accuracy of the estimates of the
break point $T_{b}^{0}$ for model M3}

\begin{centering}
{\footnotesize{}}
\begin{tabular}{ccccccc|ccccc}
\hline
 &  & {\footnotesize{}MAE} & {\footnotesize{}Std} & {\footnotesize{}$\textrm{RMSE}$} & {\footnotesize{}$Q_{0.25}$} & \multicolumn{1}{c}{{\footnotesize{}$Q_{0.75}$}} & {\footnotesize{}MAE} & {\footnotesize{}Std} & {\footnotesize{}$\textrm{RMSE}$} & {\footnotesize{}$Q_{0.25}$} & {\footnotesize{}$Q_{0.75}$}\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
 &  & \multicolumn{5}{c|}{{\footnotesize{}$\lambda_{0}=0.3$}} & \multicolumn{5}{c}{{\footnotesize{}$\lambda_{0}=0.5$}}\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
{\footnotesize{}$\delta^{0}=0.3$} & {\footnotesize{}OLS} & 23.66 & 28.14 & 31.32 & 22 & 69 & 22.01 & 26.61 & 26.59 & 33 & 72\tabularnewline
 & {\footnotesize{}GL-LN} & 19.31 & 19.22 & 26.28 & 30 & 57 & 14.89 & 18.18 & 19.08 & 39 & 61\tabularnewline
 & {\footnotesize{}GL-Uni} & 21.38 & 24.12 & 28.08 & 24 & 64 & 18.76 & 22.88 & 22.01 & 31 & 66\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
{\footnotesize{}$\delta^{0}=0.4$} & {\footnotesize{}OLS} & 19.31 & 25.76 & 27.71 & 23 & 57 & 18.14 & 23.43 & 23.44 & 38 & 60\tabularnewline
 & {\footnotesize{}GL-LN} & 15.04 & 17.64 & 21.29 & 29 & 51 & 12.36 & 16.43 & 16.52 & 40 & 60\tabularnewline
 & {\footnotesize{}GL-Uni} & 18.46 & 22.74 & 25.18 & 25 & 58 & 15.91 & 20.42 & 20.42 & 37 & 62\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
{\footnotesize{}$\delta^{0}=0.6$} & {\footnotesize{}OLS} & 12.02 & 19.02 & 19.82 & 25 & 37 & 10.28 & 15.51 & 15.58 & 45 & 55\tabularnewline
 & {\footnotesize{}GL-LN} & 9.29 & 12.86 & 14.61 & 29 & 40 & 8.46 & 11.84 & 11.86 & 45 & 55\tabularnewline
 & {\footnotesize{}GL-Uni} & 12.33 & 18.43 & 19.54 & 27 & 41 & 8.90 & 14.54 & 14.53 & 45 & 55\tabularnewline
\cline{3-12} \cline{4-12} \cline{5-12} \cline{6-12} \cline{7-12} \cline{8-12} \cline{9-12} \cline{10-12} \cline{11-12} \cline{12-12}
{\footnotesize{}$\delta^{0}=1$} & {\footnotesize{}OLS} & 3.72 & 6.88 & 6.89 & 28 & 32 & 3.85 & 6.98 & 6.98 & 48 & 52\tabularnewline
 & {\footnotesize{}GL-LN} & 3.49 & 6.44 & 6.57 & 28 & 32 & 3.45 & 6.09 & 6.10 & 48 & 52\tabularnewline
 & {\footnotesize{}GL-Uni} & 4.37 & 8.12 & 8.24 & 28 & 32 & 3.86 & 6.97 & 6.96 & 48 & 52\tabularnewline
\hline
\end{tabular}{\footnotesize\par}
\par\end{centering}
\noindent\begin{minipage}[t]{1\columnwidth}
{\scriptsize{}The model is $y_{t}=\delta^{0}\mathbf{1}\left\{ t>\left\lfloor T\lambda_{0}\right\rfloor \right\} +\alpha^{0}y_{t-1}+e_{t},\,e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,0.5\right),\,\alpha^{0}=0.6,\,T=100$.
The notes of Table \ref{Table M0 Bias} apply.}
\end{minipage}
\end{table}
\par\end{center}

\begin{center}

\par\end{center}\begin{center}
\begin{table}[H]
\caption{\label{Table M0}Small-sample coverage rates and lengths of the confidence
sets for model M1}

\begin{centering}
{\footnotesize{}}
\begin{tabular}{cccccccccc}
\hline
 &  & \multicolumn{2}{c}{{\footnotesize{}$\delta^{0}=0.4$}} & \multicolumn{2}{c}{{\footnotesize{}$\delta^{0}=0.8$}} & \multicolumn{2}{c}{{\footnotesize{}$\delta^{0}=1.2$}} & \multicolumn{2}{c}{{\footnotesize{}$\delta^{0}=1.6$}}\tabularnewline
 &  & {\footnotesize{}$\textrm{Cov.}$} & {\footnotesize{}$\textrm{Lgth.}$} & {\footnotesize{}$\textrm{Cov.}$} & {\footnotesize{}$\textrm{Lgth.}$} & {\footnotesize{}$\textrm{Cov.}$} & {\footnotesize{}$\textrm{Lgth.}$} & {\footnotesize{}$\textrm{Cov.}$} & {\footnotesize{}$\textrm{Lgth.}$}\tabularnewline
\cline{3-10} \cline{4-10} \cline{5-10} \cline{6-10} \cline{7-10} \cline{8-10} \cline{9-10} \cline{10-10}
{\footnotesize{}$\lambda_{0}=0.5$} & {\footnotesize{}OLS-CR} & 0.922 & 77.52 & 0.934 & 49.46 & 0.946 & 22.51 & 0.938 & 10.48\tabularnewline
 & {\footnotesize{}Bai (1997)} & 0.812 & 58.12 & 0.862 & 28.75 & 0.928 & 13.78 & 0.928 & 8.16\tabularnewline
 & {\footnotesize{}$\widehat{U}_{T}\left(T_{\textrm{m}}\right).\textrm{neq}$} & 0.950 & 75.45 & 0.950 & 41.68 & 0.950 & 21.78 & 0.950 & 14.79\tabularnewline
 & {\footnotesize{}ILR} & 0.959 & 76.14 & 0.973 & 35.79 & 0.976 & 14.44 & 0.977 & 7.15\tabularnewline
 & {\footnotesize{}GL-LN} & 0.942 & 49.76 & 0.948 & 22.45 & 0.958 & 10.47 & 0.965 & 5.15\tabularnewline
 & {\footnotesize{}sup-W} & \multicolumn{2}{c}{0.384} & \multicolumn{2}{c}{0.916} & \multicolumn{2}{c}{1.000} & \multicolumn{2}{c}{1.000}\tabularnewline
\cline{3-10} \cline{4-10} \cline{5-10} \cline{6-10} \cline{7-10} \cline{8-10} \cline{9-10} \cline{10-10}
{\footnotesize{}$\lambda_{0}=0.3$} & {\footnotesize{}OLS-CR} & 0.928 & 74.95 & 0.928 & 46.68 & 0.930 & 21.47 & 0.958 & 10.22\tabularnewline
 & {\footnotesize{}Bai (1997)} & 0.830 & 56.64 & 0.870 & 28.72 & 0.904 & 13.89 & 0.962 & 8.27\tabularnewline
 & {\footnotesize{}$\widehat{U}_{T}\left(T_{\textrm{m}}\right).\textrm{neq}$} & 0.952 & 77.51 & 0.952 & 44.72 & 0.952 & 22.51 & 0.952 & 14.21\tabularnewline
 & {\footnotesize{}ILR} & 0.952 & 78.28 & 0.966 & 39.78 & 0.969 & 31.29 & 0.968 & 18.23\tabularnewline
 & {\footnotesize{}GL-LN} & 0.942 & 49.60 & 0.948 & 23.89 & 0.958 & 11.14 & 0.980 & 5.60\tabularnewline
 & {\footnotesize{}sup-W} & \multicolumn{2}{c}{0.316} & \multicolumn{2}{c}{0.866} & \multicolumn{2}{c}{0.992} & \multicolumn{2}{c}{1.000}\tabularnewline
\hline
\end{tabular}{\footnotesize\par}
\par\end{centering}
\noindent\begin{minipage}[t]{1\columnwidth}
{\scriptsize{}The model is $y_{t}=\delta^{0}\mathbf{1}_{\left\{ t>\left\lfloor T\lambda_{0}\right\rfloor \right\} }+e_{t},\,e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right),\,T=100$.
Cov. and Lgth. refer to the coverage probability and the average length
of the confidence set (i.e., the average number of dates in the confidence
set). sup-W refers to the rejection probability of the sup-Wald test
using a 5\% asymptotic critical value. The number of simulations is
3,000.}
\end{minipage}
\end{table}
\par\end{center}

\begin{center}

\par\end{center}\begin{center}
\begin{table}[H]
\caption{\label{Table M1}Small-sample coverage rates and lengths of the confidence
sets for model M2}

\begin{centering}
{\footnotesize{}}
\begin{tabular}{cccccccccc}
\hline
 &  & \multicolumn{2}{c}{{\footnotesize{}$\delta^{0}=0.4$}} & \multicolumn{2}{c}{{\footnotesize{}$\delta^{0}=0.8$}} & \multicolumn{2}{c}{{\footnotesize{}$\delta^{0}=1.2$}} & \multicolumn{2}{c}{{\footnotesize{}$\delta^{0}=1.6$}}\tabularnewline
 &  & {\footnotesize{}$\textrm{Cov.}$} & {\footnotesize{}$\textrm{Lgth.}$} & {\footnotesize{}$\textrm{Cov.}$} & {\footnotesize{}$\textrm{Lgth.}$} & {\footnotesize{}$\textrm{Cov.}$} & {\footnotesize{}$\textrm{Lgth.}$} & {\footnotesize{}$\textrm{Cov.}$} & {\footnotesize{}$\textrm{Lgth.}$}\tabularnewline
\cline{3-10} \cline{4-10} \cline{5-10} \cline{6-10} \cline{7-10} \cline{8-10} \cline{9-10} \cline{10-10}
{\footnotesize{}$\lambda_{0}=0.5$} & {\footnotesize{}OLS-CR} & 0.952 & 80.29 & 0.954 & 57.70 & 0.957 & 30.04 & 0.963 & 15.10\tabularnewline
 & {\footnotesize{}Bai (1997)} & 0.804 & 64.64 & 0.824 & 43.53 & 0.907 & 13.03 & 0.930 & 7.81\tabularnewline
 & {\footnotesize{}$\widehat{U}_{T}\left(T_{\textrm{m}}\right).\textrm{neq}$} & 0.967 & 87.30 & 0.967 & 72.70 & 0.957 & 36.70 & 0.957 & 30.20\tabularnewline
 & {\footnotesize{}ILR} & 0.937 & 81.88 & 0.945 & 57.43 & 0.972 & 21.99 & 0.972 & 18.96\tabularnewline
 & {\footnotesize{}GL-LN} & 0.933 & 55.13 & 0.912 & 32.97 & 0.935 & 20.03 & 0.961 & 10.62\tabularnewline
 & {\footnotesize{}sup-W} & \multicolumn{2}{c}{0.316} & \multicolumn{2}{c}{0.699} & \multicolumn{2}{c}{1.000} & \multicolumn{2}{c}{1.000}\tabularnewline
\cline{3-10} \cline{4-10} \cline{5-10} \cline{6-10} \cline{7-10} \cline{8-10} \cline{9-10} \cline{10-10}
{\footnotesize{}$\lambda_{0}=0.3$} & {\footnotesize{}OLS-CR} & 0.945 & 79.25 & 0.957 & 54.93 & 0.962 & 29.91 & 0.970 & 15.37\tabularnewline
 & {\footnotesize{}Bai (1997)} & 0.823 & 63.79 & 0.851 & 26.33 & 0895 & 13.07 & 0.946 & 7.87\tabularnewline
 & {\footnotesize{}$\widehat{U}_{T}\left(T_{\textrm{m}}\right).\textrm{neq}$} & 0.966 & 88.23 & 0.953 & 59.66 & 0.950 & 39.65 & 0.951 & 32.39\tabularnewline
 & {\footnotesize{}ILR} & 0.945 & 84.37 & 0.945 & 62.97 & 0.971 & 33.74 & 0.987 & 17.92\tabularnewline
 & {\footnotesize{}GL-LN} & 0.945 & 53.79 & 0.923 & 34.75 & 0.934 & 19.92 & 0.944 & 10.04\tabularnewline
 & {\footnotesize{}sup-W} & \multicolumn{2}{c}{0.314} & \multicolumn{2}{c}{{\footnotesize{}0.881}} & \multicolumn{2}{c}{0.999} & \multicolumn{2}{c}{1.000}\tabularnewline
\hline
\end{tabular}{\footnotesize\par}
\par\end{centering}
\noindent\begin{minipage}[t]{1\columnwidth}
{\scriptsize{}The model is $y_{t}=\delta^{0}\mathbf{1}\left\{ t>\left\lfloor T\lambda_{0}\right\rfloor \right\} +e_{t},\,e_{t}=0.3e_{t-1}+u_{t},\,u_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right),\,T=100$.
The notes of Table \ref{Table M0} apply.}
\end{minipage}
\end{table}
\par\end{center}

\begin{center}

\par\end{center}\begin{center}
\begin{table}[H]
\caption{\label{Table M2}Small-sample coverage rates and lengths of the confidence
sets for model M3}

\begin{centering}
{\footnotesize{}}
\begin{tabular}{cccccccccc}
\hline
 &  & \multicolumn{2}{c}{{\footnotesize{}$\delta^{0}=0.4$}} & \multicolumn{2}{c}{{\footnotesize{}$\delta^{0}=0.8$}} & \multicolumn{2}{c}{{\footnotesize{}$\delta^{0}=1.2$}} & \multicolumn{2}{c}{{\footnotesize{}$\delta^{0}=1.6$}}\tabularnewline
 &  & {\footnotesize{}$\textrm{Cov.}$} & {\footnotesize{}$\textrm{Lgth.}$} & {\footnotesize{}$\textrm{Cov.}$} & {\footnotesize{}$\textrm{Lgth.}$} & {\footnotesize{}$\textrm{Cov.}$} & {\footnotesize{}$\textrm{Lgth.}$} & {\footnotesize{}$\textrm{Cov.}$} & {\footnotesize{}$\textrm{Lgth.}$}\tabularnewline
\cline{3-10} \cline{4-10} \cline{5-10} \cline{6-10} \cline{7-10} \cline{8-10} \cline{9-10} \cline{10-10}
{\footnotesize{}$\lambda_{0}=0.5$} & {\footnotesize{}OLS-CR} & 0.954 & 80.29 & 0.952 & 57.23 & 0.957 & 30.21 & 0.963 & 15.20\tabularnewline
 & {\footnotesize{}Bai (1997)} & 0.781 & 55.85 & 0.845 & 26.23 & 0.902 & 13.03 & 0.932 & 7.81\tabularnewline
 & {\footnotesize{}$\widehat{U}_{T}\left(T_{\textrm{m}}\right).\textrm{neq}$} & 0.958 & 81.28 & 0.959 & 55.34 & 0.957 & 36.71 & 0.957 & 30.20\tabularnewline
 & {\footnotesize{}ILR} & 0.934 & 65.96 & 0.956 & 33.73 & 0.975 & 21.96 & 0.984 & 17.45\tabularnewline
 & {\footnotesize{}GL-LN} & 0.912 & 60.90 & 0.925 & 32.93 & 0.964 & 19.23 & 0.971 & 9.23\tabularnewline
 & {\footnotesize{}sup-W} & \multicolumn{2}{c}{0.407} & \multicolumn{2}{c}{0.931} & \multicolumn{2}{c}{1.000} & \multicolumn{2}{c}{1.000}\tabularnewline
\cline{3-10} \cline{4-10} \cline{5-10} \cline{6-10} \cline{7-10} \cline{8-10} \cline{9-10} \cline{10-10}
{\footnotesize{}$\lambda_{0}=0.3$} & {\footnotesize{}OLS-CR} & 0.968 & 83.69 & 0.951 & 54.13 & 0.962 & 29.31 & 0.970 & 15.37\tabularnewline
 & {\footnotesize{}Bai (1997)} & 0.795 & 64.06 & 0.853 & 26.33 & 0.896 & 13.07 & 0.946 & 7.85\tabularnewline
 & {\footnotesize{}$\widehat{U}_{T}\left(T_{\textrm{m}}\right).\textrm{neq}$} & 0.960 & 86.42 & 0.953 & 59.13 & 0.950 & 39.65 & 0.951 & 32.28\tabularnewline
 & {\footnotesize{}ILR} & 0.934 & 67.73 & 0.964 & 35.30 & 0.971 & 33.74 & 0.987 & 17.92\tabularnewline
 & {\footnotesize{}GL-LN} & 0.912 & 60.28 & 0.945 & 36.08 & 0.974 & 22.72 & 0.975 & 12.71\tabularnewline
 & {\footnotesize{}sup-W} & \multicolumn{2}{c}{0.232} & \multicolumn{2}{c}{0.884} & \multicolumn{2}{c}{0.999} & \multicolumn{2}{c}{1.000}\tabularnewline
\hline
\end{tabular}{\footnotesize\par}
\par\end{centering}
\noindent\begin{minipage}[t]{1\columnwidth}
{\scriptsize{}The model is $y_{t}=\delta^{0}\mathbf{1}\left\{ t>\left\lfloor T\lambda_{0}\right\rfloor \right\} +\alpha^{0}y_{t-1}+e_{t},\,e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,0.5\right),\,\alpha^{0}=0.6,\,T=100$.
The notes of Table \ref{Table M0} apply.}
\end{minipage}
\end{table}
\par\end{center}

\pagebreak{}

\section*{}
\addcontentsline{toc}{part}{Supplemental Material}
\begin{center}
\Large{\uline{Supplemental Material} to}
\end{center}

\begin{center}
\title{\textbf{\Large{Generalized Laplace Inference in Multiple Change-Points Models}}}
\maketitle
\end{center}
\medskip{}
\medskip{}
\medskip{}
\thispagestyle{empty}

\begin{center}
$\qquad$ \textsc{\textcolor{MyBlue}{Alessandro Casini}} $\qquad$ \textsc{\textcolor{MyBlue}{Pierre Perron}}\\
\small{University of Rome Tor Vergata} $\quad$ \small{Boston University}
\\
\medskip{}
\medskip{}
\medskip{}
\medskip{}
\date{\small{\today} \\
}
\medskip{}
\medskip{}
\medskip{}
\end{center}
\begin{abstract}
{\footnotesize{}This supplemental material is structured as follows.
Section \ref{Supp Math App} contains the Mathematical Appendix which
includes all proofs of the results in the paper. Section \ref{Section Comparison-to}
includes further simulation results comparing the GL-LN method to
the GL estimators proposed in \citet{casini/perron_Lap_CR_Single_Inf}.}{\footnotesize\par}
\end{abstract}
\setcounter{page}{1}
\setcounter{section}{1}

\newpage{}

\begin{singlespace}
\noindent
\small