EconBase
← Back to paper

Continuous Record Asymptotics for Change-Points Models

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

87,840 characters

Continuous Record Asymptotics for Change-Point Models


\setcounter{page}{0}
\title{\textbf{Continuous Record Asymptotics for Change-Point Models}\thanks{We wish to thank Mark Podolskij for helpful suggestions. We also thank
Christian Gouri\'{e}roux, Hashem Pesaran, Myung Hwan Seo and Viktor
Todorov for related discussions on some aspects of this project. We
are grateful to Iv\'{a}n Fern\'{a}ndez-Val, Hiro Kaido, Zhongjun
Qu as well as seminar participants at Boston University and participants
at 2018 NBER-NSF Time Series Conference and 11th Annual SoFiE Conference
for valuable comments and suggestions. Seong Yeon Chang, Andres Sagner
and Yohei Yamamoto have provided generous help with computer programming.
We thank Yunjong Eo and James Morley for sharing their programs. Casini
gratefully acknowledges partial financial support from the Bank of
Italy.}}
\maketitle
\begin{abstract}
{\footnotesize{}In the context of a linear regression model with a
single break point, we develop a continuous record asymptotic framework
to build inference methods for the break date. We have $T$ observations
with a sampling frequency $h$ over a fixed time horizon $\left[0,\,N\right],$
and let $T\rightarrow\infty$ with $h\downarrow0$ while keeping the
time span $N$ fixed. We consider the least-squares estimate of the
break date and establish consistency and convergence rate. We provide
a limit theory for shrinking magnitudes of shifts and locally increasing
variances. The asymptotic distribution corresponds to the location
of the extremum of a function of the quadratic variation of the regressors
and of a Gaussian centered martingale process over a certain time
interval. We can account for the asymmetric informational content
provided by the pre- and post-break regimes and show how the location
of the break and shift magnitude are key ingredients in shaping the
distribution. We consider a feasible version based on plug-in estimates,
which provides a very good approximation to the finite sample distribution.
We use the concept of Highest Density Region to construct confidence
sets. Overall, our method is reliable and delivers accurate coverage
probabilities and relatively short average length of the confidence
sets. Importantly, it does so irrespective of the size of the break.}{\footnotesize\par}
\end{abstract}
\indent {\bf{JEL Classification}}: C10, C12, C22\\
\noindent {\bf{Keywords}}: Asymptotic distribution, break date, change-point, highest density region, semimartingale.

\onehalfspacing
\thispagestyle{empty}
\allowdisplaybreaks

\pagebreak{}

\section{Introduction}

In the context of a linear regression model with a single break point,
we develop a continuous record asymptotic framework and inference
methods for the break date. Our model is specified in continuous time
but estimated with discrete-time observations using a least-squares
method. We have $T$ observations with a sampling frequency $h$ over
a fixed time horizon $\left[0,\,N\right],$ where $N=Th$ denotes
the time span of the data. We consider a continuous record asymptotic
framework whereby $T$ increases by shrinking the time interval $h$
to zero while keeping time span $N$ fixed. We impose very mild conditions
on an underlying continuous-time model assumed to generate the data,
basically continuous It� semimartingales.

An extensive amount of research addressed change-point problems under
the classical large-$N$ asymptotics. Early contributions are \citet{hinkley:71},
\citet{bhattacharya:87}, and \citet{yao:87}, who adopted a Maximum
Likelihood (ML) approach, and for linear regression models, \citet{bai:97RES}
and \citet{bai/perron:98}.  See the reviews of \citet{csorgo/horvath:97},
\citet{aue/Horvath:13}, \citet{casini/perron_Oxford_Survey} and
references therein. In this literature, the resulting large-$N$
limit theory for the estimate of the break date depends on the exact
distributions of the regressors and disturbances. Therefore, a so-called
shrinkage asymptotic theory was adopted whereby the magnitude of
the shift, say $\delta_{T},$ converges to zero which leads to an
invariant limit distribution.

We study a general change-point problem under a continuous record
asymptotic framework and develop inference procedures based on the
derived asymptotic distribution. We establish consistency at rate-$T$
convergence for the least-squares estimate of the break date, assumed
to occur at time $N_{b}^{0}$. Given the fast rate of convergence,
we introduce a limit theory with shrinking magnitudes of shifts and
increasing variance of the residual process local to the change-point.
The asymptotic distribution corresponds to the location of the extremum
of a function of the (quadratic) variation of the regressors and of
a Gaussian centered martingale process over some time interval. It
is characterized by some notable aspects. With the time horizon $\left[0,\,N\right]$
fixed, we can account for the asymmetric informational content provided
by the pre- and post-break observations, i.e., the time span and the
position of the break date $N_{b}^{0}$ convey useful information
about the finite-sample distribution. In contrast, this is not achievable
under the large-$N$ shrinkage asymptotic framework because both pre-
and post-break segments expand proportionately as $T$ increases and,
given the mixing assumptions imposed, only the neighborhood around
the break date remains relevant. Further, the domain of the extremum
depends on the position of the break $N_{b}^{0}$ relative to $N$,
or total span, and thus the distribution is asymmetric, in general.
The degree of asymmetry increases as the true break point moves away
from mid-sample. This holds unless the magnitude of the break is large,
in which case the density is symmetric irrespective of the location
of the break. This accords with simulation evidence which documents
that the break point estimate is less precise and the coverage rates
of the confidence intervals less reliable when the break is not at
mid-sample {[}see, e.g., \citet{chang/perron:18}{]}. When the shift
magnitude is small, the  density displays three modes. As the shift
magnitude increases, this tri-modality vanishes. We show empirically
that all of these features are shared by the finite-sample distribution
of the least-squares estimator of the break date. Hence, the continuous
record asymptotics theory provides an accurate approximation to the
finite-sample distribution of the break date estimator. In contrast,
we show that the large-$N$ shrinkage asymptotic distribution of \citet{yao:87}
and \citet{bai:97RES} provides a poor approximation to the finite-sample
distribution and does not share any of those features as we discuss.

Our asymptotics can be seen as intermediate between the shrinkage
asymptotics and more recent approaches relying on weak identification
{[}see e.g., \citet{elliott/mueller:07} and \citet{elliott/mueller/watson:15}{]}.
On the one hand, using the usual shrinking condition of \citet{yao:87}
and \citet{bai:97RES} for which the break magnitude, say $\delta_{T},$
goes to zero at a rate slower than $O(T^{-1/2})$ leads to underestimation
of the uncertainty about the break date. On the other hand, the weak
identification condition of \citet{elliott/mueller:07} for which
$\delta_{T}$ goes to zero at a fast rate (i.e., $\delta_{T}=O(T^{-1/2})$,
so that the change-point cannot be consistently estimated) leads to
overstating the uncertainty. This has opposite consequences for the
confidence intervals of the break date. Confidence sets have poor
coverage probabilities when the break is small under \citeauthor{bai:97RES}'s
framework while they can be too wide under that of \citet{elliott/mueller:07}.
In this paper, the key is not to focus our asymptotic experiment on
shrinking condition on $\delta_{T}$ but to make assumptions on the
signal-to-noise ratio $\delta_{T}/\sigma_{t}$ instead, where $\sigma_{t}$
is the volatility of the errors. We require $\delta_{T}$ to go to
zero at a slower rate than that of \citet{elliott/mueller:07}---to
guarantee strong identification---and require $\sigma_{t}$ to increase
without bound when $t$ approaches the break date $T_{b}^{0}$. This
offers a new characterization of the uncertainty without compromising
strong identification and consistency of the model parameters needed
to conduct inference.

Despite the effort devoted to the construction of confidence intervals
for the break date {[}see e.g., \citet{bai/perron:98}, \citet{elliott/mueller:07}
and \citet{eo/morley:15}{]}, what is still missing is a method that,
for both large and small breaks, achieves both accurate coverage rates
and satisfactory average lengths. The most popular method is \citet{bai/perron:98}
which yields confidence intervals that are relatively short but have
good coverage only when the magnitude of the break is not small. However,
both small and large breaks are relevant for empirical work; breaks
that are statistically small can still be practically relevant.


Given the peculiar properties (e.g., multi-modality and asymmetry)
of the continuous record asymptotic distribution, we propose a non-standard
inference related to Bayesian analyses. We use the concept of Highest
Density Region to construct confidence sets for the break date. Our
method has good coverage and length across all break magnitudes. This
has important implications for empirical work because the user can
be confident that our confidence interval includes the true value
across all break sizes. For small breaks, the length of the confidence
intervals from any method can be quite large for some models. However,
our confidence interval is still informative because it reveals that
there is high uncertainty about the change-point. The same information
cannot be provided by existing methods either because they do not
have good coverage unless the break is not small {[}e.g., \citet{bai/perron:98}{]}
or because they have a large length even when the break is not small
{[}e.g., \citet{elliott/mueller:07}{]}.

We use the continuous record asymptotics to provide an alternative
approximation to the finite-sample distribution of the least-squares
estimator based on discrete-time data. This creates no contradiction
since asymptotic theory is intended as a thought experiment used to
obtain approximations to the distribution of estimators or test statistics.
The continuous record asymptotics has proven to be useful in other
discrete-time settings such as in the context of unit roots {[}cf.
\citet{phillips:87a} and \citet{perron:1991ecma}{]} and nonparametric
regression {[}cf. \citet{brown/low:1996}{]}. \citet{brown/low:1996}
showed the asymptotic equivalence between a nonparametric regression
problem and a white noise with drift problem. Our results show the
asymptotic equivalence between discrete-time and continuous-time regression
models with a change-point.

Recent work in change-point analysis has focused on estimation when
the number of change-points is allowed to increase with the sample
size {[}e.g., \citet{fryzlewicz:14}{]} and when the change-point
is allowed to approach the start and end sample point. A growing literature
has also considered change-points in a high-dimensional setting {[}e.g.,
\citet{lee/seo/shin:16}, \citet{leonardi/buhlmann:16}, \citet{wang/lin/willett:19}
and \citet{wang/yu/rinaldo/willett:19}{]}. This work is mainly concerned
with consistent estimation of the change-point dates and development
of corresponding computational algorithms.  Our focus is on asymptotic
theory and inference within the classical change-point model with
a single break. Our results can also have useful implications for
the growing literature on inference in high-dimensional change-point
analysis and for the literature on threshold regression {[}see, e.g.,
\citet{hansen:00ecma} and \citet{hildago/lee/seo:19}{]}.

This paper relates to other work by the authors, namely \citeauthor{casini/perron_SC_BP_Lap}
(\citeyear{casini/perron_Lap_CR_Single_Inf}; \citeyear{casini/perron_SC_BP_Lap})\nocite{casini/perron_Lap_CR_Single_Inf}.
\citet{casini/perron_Lap_CR_Single_Inf} used the asymptotic results
developed in this paper and proposed a new Generalized Laplace estimator
of the break date under a continuous record asymptotic framework.
 \citet{casini/perron_SC_BP_Lap} analyzed the Generalized Laplace
method under classical asymptotics and focused on the theoretical
relationship between the asymptotic distribution of frequentist and
Bayesian estimators of the break point. Finally, \citet{chambers/taylor:19}
considered both deterministic one-time and continuous stochastic parameter
change in a continuous-time autoregressive model while \citet{casini_CR_Test_Inst_Forecast}
introduced continuous-time asymptotics to test for forecast failure.
Recently, \citet{casini:change-point-spectra} considered testing
and estimating change-points in a locally stationary process using
frequency-domain methods.

The paper is organized as follows. Section \ref{Section: Model-and-Assumptions}
introduces the model and the estimation method. Section \ref{Section Consistency-and-Rate }
contains results about the consistency and rate of convergence for
fixed shifts. Section \ref{Section Asymptotic Distribution: Continuous Case}
develops the asymptotic theory. We compare our limit theory with the
finite-sample distribution in Section \ref{Section Approximation-to-the}.
Section \ref{Section Inference Methods} describes how to construct
the confidence sets, with simulation results  reported in Section
\ref{Section Small-Sample-Effectiveness-of}. Section \ref{Section Conclusions}
provides brief concluding remarks. The Supplement {[}\citet{casini/perron_CR_Single_Break_Supp}{]}
contains the proofs as well as additional material.

\section{Model and Assumptions \label{Section: Model-and-Assumptions}}

We denote the transpose of a matrix $A$ by $A'$ and the $\left(i,\,j\right)$
elements of $A$ by $A^{\left(i,j\right)}$. We use $\left\Vert \cdot\right\Vert $
to denote the Euclidean norm of a linear space, i.e., $\left\Vert x\right\Vert =\left(\sum_{i=1}^{p}x_{i}^{2}\right)^{1/2}$
for $x\in\mathbb{R}^{p}.$ We use $\left\lfloor \cdot\right\rfloor $
to denote the largest smaller integer function. A sequence $\left\{ u_{kh}\right\} _{k=1}^{T}$
is $i.i.d.$ (resp., $i.n.d$) if the $u_{kh}$ are independent and
identically (resp., non-identically) distributed. We use $\overset{P}{\rightarrow},\,\Rightarrow,$
and $\overset{\mathcal{L}-\mathrm{s}}{\Rightarrow}$ to denote convergence
in probability, weak convergence and stable convergence in law, respectively.
For semimartingales $\left\{ S_{t}\right\} _{t\geq0}$ and $\left\{ R_{t}\right\} _{t\geq0}$,
we denote their covariation process by $\left[S,\,R\right]_{t}$ and
their predictable counterpart by $\left\langle S,\,R\right\rangle _{t}$.
The symbol ``$\triangleq$'' denotes definitional equivalence.

Consider a change-point model with a single break point:
\begin{align}
Y_{t} & =D_{t}'\nu^{0}+Z_{t}'\delta_{Z,1}^{0}+e_{t},\quad(t=0,\,1,\ldots,\,T_{b}^{0})\label{Original SC Model}\\
Y_{t} & =D_{t}'\nu^{0}+Z_{t}'\delta_{Z,2}^{0}+e_{t},\quad(t=T_{b}^{0}+1,\ldots,\,T),\nonumber
\end{align}
 where $Y_{t}$ is the dependent variable, $D_{t}$ and $Z_{t}$ are,
respectively, $q\times1$ and $p\times1$ vectors of regressors and
$e_{t}$ is an unobservable disturbance. The vector-valued parameters
$\nu^{0},\,\delta_{Z,1}^{0}$ and $\delta_{Z,2}^{0}$ are unknown
with $\delta_{Z,1}^{0}\neq\delta_{Z,2}^{0}$. Our main purpose is
to develop inference methods for the unknown change-point date $T_{b}^{0}$
when $T+1$ observations on $\left(Y_{t},\,D_{t},\,Z_{t}\right)$
are available. Before moving to the re-parametrization of the model,
we discuss the underlying continuous-time model assumed to generate
the data. The discrete-time variables are assumed to be generated
from the continuous-time processes $\{D_{s},\,Z_{s},\,e_{s}\}_{s\geq0}$
defined on a filtered probability space $(\Omega,\,\mathscr{F},\,(\mathscr{F}_{s})_{s\geq0},\,P).$
We observe realizations of $\{Y_{s},\,D_{s},\,Z_{s}\}$ at discrete
points of time.

The sampling occurs at regularly spaced time intervals of length $h$
within a fixed time horizon $\left[0,\,N\right]$ where $N$ denotes
the span of the data. We observe $\left\{ _{h}Y_{kh},\,{}_{h}D_{kh},\,{}_{h}Z_{kh};\,k=0,\,1,\ldots,\,T=N/h\right\} $.
$_{h}D_{kh}\in\mathbb{R}^{q}$ and $_{h}Z_{kh}\in\mathbb{R}^{p}$
are random vector step functions which jump only at times $0,\,h,\ldots,\,Th$.
We shall allow $_{h}D_{kh}$ and $_{h}Z_{kh}$ to include both predictable
processes and locally-integrable semimartingles, though the case with
predictable regressors is more delicate and discussed in the supplement.
The discretized processes $_{h}D_{kh}$ and $_{h}Z_{kh}$ are assumed
to be adapted to the increasing and right-continuous filtration $\left\{ \mathscr{F}_{t}\right\} _{t\geq0}$.
For any process $X$ we denote its ``increments'' by $\Delta_{h}X_{k}=X_{kh}-X_{\left(k-1\right)h}$.
For $k=1,\ldots,T$, let $\Delta_{h}D_{k}\triangleq\mu_{D,k}h+\Delta_{h}M_{D,k}$
and $\Delta_{h}Z_{k}\triangleq\mu_{Z,k}h+\Delta_{h}M_{Z,k}$ where
the ``drifts'' $\mu_{D,t}\in\mathbb{R}^{q},\,\mu_{Z,t}\in\mathbb{R}^{p}$
are $\mathscr{F}_{t-h}$-measurable (exact assumptions will be given
below), and $M_{D,k}\in\mathbb{R}^{q},\,M_{Z,k}\in\mathbb{R}^{p}$
are continuous local martingales with finite conditional covariance
matrix $P$-a.s., $\mathbb{E}(\Delta_{h}M_{D,t}\Delta_{h}M_{D,t}'|\,\mathscr{F}_{t-h})=\Sigma_{D,t-h}\Delta t$
and $\mathbb{E}(\Delta_{h}M_{Z,t}\Delta_{h}M'_{Z,t}|\,\mathscr{F}_{t-h})=\Sigma_{Z,t-h}\Delta t$
($\Delta t$ and $h$ are used interchangeably). Let $\lambda_{0}\in\left(0,\,1\right)$
denote the fractional break date (i.e., $T_{b}^{0}=\left\lfloor T\lambda_{0}\right\rfloor $).
Via the Doob-Meyer Decomposition, model \eqref{Original SC Model}
can be expressed as
\begin{align}
\Delta_{h}Y_{k} & \triangleq\begin{cases}
\left(\Delta_{h}D_{k}\right)'\nu^{0}+\left(\Delta_{h}Z_{k}\right)'\delta_{Z,1}^{0}+\Delta_{h}e_{k}^{*}, & \left(k=1,\ldots,\left\lfloor T\lambda_{0}\right\rfloor \right)\\
\left(\Delta_{h}D_{k}\right)'\nu^{0}+\left(\Delta_{h}Z_{k}\right)'\delta_{Z,2}^{0}+\Delta_{h}e_{k}^{*}, & \left(k=\left\lfloor T\lambda_{0}\right\rfloor +1,\ldots,\,T\right),
\end{cases}\label{Eq. Model 1, CT}
\end{align}
 where the error process $\left\{ \Delta_{h}e_{t}^{*},\,\mathscr{F}_{t}\right\} $
is a continuous local martingale difference sequence with conditional
variance $\mathbb{E}[\left(\Delta_{h}e_{t}^{*}\right)^{2}|\,\mathscr{F}_{t-h}]=\sigma_{e,t-h}^{2}\Delta t$
$P$-a.s. finite. The underlying continuous-time data-generating process
can thus be represented (up to $P$-null sets) in integral equation
form as
\begin{align}
D_{t}=D_{0}+\int_{0}^{t}\mu_{D,s}ds+\int_{0}^{t}\sigma_{D,s}dW_{D,s},\quad & Z_{t}=Z_{0}+\int_{0}^{t}\mu_{Z,s}ds+\int_{0}^{t}\sigma_{Z,s}dW_{Z,s},\label{Model Regressors Integral Form}
\end{align}
where $\sigma_{D,t}$ and $\sigma_{Z,t}$ are the instantaneous covariance
processes taking values in $\mathcal{M}_{q}^{\textrm{c�dl�g}}$ and
$\mathcal{M}_{p}^{\textrm{c�dl�g}}$ {[}the space of $p\times p$
positive definite real-valued matrices whose elements are c�dl�g{]};
$W_{D}$ (resp., $W_{Z}$) is a $q$ (resp., $p$)-dimensional standard
Wiener process; $e^{*}=\left\{ e_{t}^{*}\right\} _{t\geq0}$ is a
continuous local martingale which is orthogonal (in a martingale sense)
to $\left\{ D_{t}\right\} _{t\geq0}$ and $\left\{ Z_{t}\right\} _{t\geq0}$;
and $D_{0}$ and $Z_{0}$ are $\mathscr{F}_{0}$-measurable random
vectors. In \eqref{Model Regressors Integral Form}, $\int_{0}^{t}\mu_{D,s}ds$
is a continuous adapted process with finite variation paths and $\int_{0}^{t}\sigma_{D,s}dW_{D,s}$
corresponds to a continuous local martingale.
\begin{assumption}
\label{Assumption 1, CT}(i) $\mu_{D,t},\,\mu_{Z,t},\,\sigma_{D,t}$
and $\sigma_{Z,t}$ satisfy $P$-a.s., $\sup_{\omega\in\Omega,\,0<t\leq\tau_{T}}\left\Vert \mu_{D,t}\left(\omega\right)\right\Vert <\infty$,
$\sup_{\omega\in\Omega,}$ $_{0<t\leq\tau_{T}}\left\Vert \mu_{Z,t}\left(\omega\right)\right\Vert <\infty$,
$\sup_{\omega\in\Omega,\,0<t\leq\tau_{T}}\left\Vert \sigma_{D,t}\left(\omega\right)\right\Vert <\infty$
and $\sup_{\omega\in\Omega,\,0<t\leq\tau_{T}}\left\Vert \sigma_{Z,t}\left(\omega\right)\right\Vert <\infty$
for some localizing sequence $\left\{ \tau_{T}\right\} $ of stopping
times. Also, $\sigma_{D,\,s}$ and $\sigma_{Z,s}$ are c�dl�g; (ii)
$\int_{0}^{t}\mu_{D,s}ds$ and \textup{$\int_{0}^{t}\mu_{Z,s}ds$
}belong to the class of continuous adapted finite variation processes;
(iii) $\int_{0}^{t}\sigma_{D,s}dW_{D,s}$ and $\int_{0}^{t}\sigma_{Z,s}dW_{Z,s}$
are continuous local martingales with $P$-a.s. finite positive definite
conditional variances (or spot covariances) defined by $\Sigma_{D,t}=\sigma_{D,t}\sigma'_{D,t}$
and $\Sigma_{Z,t},=\sigma_{Z,t}\sigma'_{Z,t}$, which for all $t<\infty$
satisfy $\int_{0}^{t}\Sigma_{D,s}^{\left(j,j\right)}ds<\infty$ $\left(j=1,\ldots,\,q\right)$
and $\int_{0}^{t}\Sigma_{Z,s}^{\left(j,j\right)}ds<\infty$ $\left(j=1,\ldots,\,p\right)$.
Furthermore, for every $j=1,\ldots,\,q,$ $r=1,\ldots,\,p$, and $k=1,\ldots,\,T$,
$h^{-1}\int_{\left(k-1\right)h}^{kh}\Sigma_{D,s}^{\left(j,j\right)}ds$
and $h^{-1}\int_{\left(k-1\right)h}^{kh}\Sigma_{Z,s}^{\left(r,r\right)}ds$
are bounded away from zero and infinity, uniformly in $k$ and $h$;
(iv) $e_{t}^{*}$ is such that $e_{t}^{*}\triangleq\int_{0}^{t}\sigma_{e,s}dW_{e,s}$
with $0<\sigma_{e,t}^{2}<\infty$, where $W_{e}$ is a one-dimensional
standard Wiener process. Furthermore, $\left\langle e,\,D\right\rangle _{t}=\left\langle e,\,Z\right\rangle _{t}=0$
identically for all $t\geq0$.
\end{assumption}
Part (i) restricts the processes to be locally bounded and part (ii)
requires the drifts to be adapted finite variation processes. These
are standard regularity conditions in the high-frequency statistics
literature {[}cf. \citet{barndorff/shephard:04}{]}. Part (iii) requires
the regressors to have finite integrated covariance. We rule out jump
processes; hence, our results are not expected to provide good approximations
for high-frequency data but for data sampled at lower frequencies.
In principle, ultra high frequency data would be used which, however,
are essentially available only for financial variables, which involve
a host of issues that we cannot handle (e.g., market-microstructure,
bid-ask spread, volatility jumps, non-continuous sampling, etc.).
\begin{assumption}
\label{Assumption 2}$D,\,Z,\,e$ and $\Sigma^{0}\triangleq\left\{ \Sigma_{\cdot,t},\,\sigma_{e,t}\right\} _{t\geq0}$
have $P$-a.s. continuous sample paths.
\end{assumption}
An interesting issue is whether the theoretical results to be derived
for model \eqref{Eq. Model 1, CT} are applicable to classical structural
change models for which an increasing span of data is assumed. This
requires establishing a connection between the assumptions imposed
on the stochastic processes in both settings. Roughly, the classical
long-span setting uses approximation results valid for weakly dependent
data; e.g., ergodic and mixing procesess. Such assumptions are not
needed under our fixed-span asymptotics. Nonetheless, we can impose
restrictions on the probabilistic properties of the latent volatility
processes in our model and thereby guarantee that ergodic and mixing
properties are inherited by the corresponding observed processes.
This follows from Theorem 3.1 in \citet*{genon-catalot/jeantheau/laredo:00}
together with Proposition 4 in \citet{carrasco/chen:02}. For example,
these results imply that the observations $\left\{ Z_{kh}\right\} _{k\geq1}$
(with fixed $h$) can be viewed (under certain conditions) as a hidden
Markov model which inherits the ergodic and mixing properties of $\left\{ \sigma_{Z,t}\right\} _{t\geq0}$.
Hence, our model encompasses those considered in the structural change
literature that uses a long-span asymptotic setting. We shall extend
model \eqref{Eq. Model 1, CT} to allow for predictable processes
(e.g., a constant and/or lagged dependent variable) in the supplement.
\begin{assumption}
\label{Assumption 3 Break Date} $N_{b}^{0}=N\lambda_{0}$ for some
$\lambda_{0}\in\left(0,\,1\right)$.
\end{assumption}
It is useful to re-parametrize model \eqref{Eq. Model 1, CT}. Let
$y_{kh}=\Delta_{h}Y_{k},$ $x_{kh}=(\Delta_{h}D'_{k},\,\Delta_{h}Z'_{k})'$,
$z_{kh}=\Delta_{h}Z_{k}$, $e_{kh}=\Delta_{h}e_{k}^{*},$ $\beta^{0}=((\pi^{0}),\,(\delta_{Z,1}^{0})')'$
and $\delta_{Z}^{0}=\delta_{Z,2}^{0}-\delta_{Z,1}^{0}$. \eqref{Eq. Model 1, CT}
can be expressed as:
\begin{align}
y_{kh} & =x_{kh}'\beta^{0}+e_{kh},\qquad & (k=1,\ldots,\,T_{b}^{0})\label{Model (4), scalar format, CT}\\
y_{kh} & =x_{kh}'\beta^{0}+z_{kh}'\delta_{Z}^{0}+e_{kh},\qquad & (k=T_{b}^{0}+1,\ldots,\,T),\nonumber
\end{align}
where the true parameter $\theta^{0}=((\beta^{0})',\,(\delta_{Z}^{0})')'$
takes value in a compact space $\Theta\subset\mathbb{R}^{\textrm{dim}\left(\theta\right)}$.
Also, define $z_{kh}=R'x_{kh}$, where $R$ is a $\left(q+p\right)\times p$
known matrix with full column rank. We consider a partial structural
change model for which $R=\left(0,\,I\right)'$ with $I$ an identity
matrix.

Finally, we write the model in matrix format which will be useful
for the derivations. Let $Y=(y_{h},\,\ldots,\,y_{Th})',\,X=(x_{h},\,\ldots,\,x_{Th})'$,
$e=(e_{h},\,\ldots,\,e_{Th})',$ $X_{1}=(x_{h},\,\ldots,\,x_{T_{b}h},\,0,\,\ldots,\,0)'$,
$X_{2}=(0,\,\ldots,\,0,\,x_{\left(T_{b}+1\right)h},\ldots,\,x_{Th})'$
and $X_{0}=(0,\,\ldots,\,0,\,x_{\left(T_{b}^{0}+1\right)h},\ldots,\,x_{Th})'$.
Note that the difference between $X_{0}$ and $X_{2}$ is that the
latter uses $T_{b}$ rather than $T_{b}^{0}$. Define $Z_{1}=X_{1}R,$
$Z_{2}=X_{2}R$ and $Z_{0}=XR$. \eqref{Model (4), scalar format, CT}
in matrix format is: $Y=X\beta^{0}+Z_{0}\delta_{Z}^{0}+e$. We consider
the least-squares estimator of $T_{b}$, i.e., the minimizer of $S_{T}\left(T_{b}\right)$,
the sum of squared residuals when regressing $Y$ on $X$ and $Z_{2}$
over all possible partitions, namely: $\widehat{T}_{b}^{\textrm{LS}}=\mathrm{argmin}_{p+q\leq T_{b}\leq T}S_{T}\left(T_{b}\right)$.
It is straightforward to show that $\widehat{T}_{b}^{\textrm{LS}}=\mathrm{\mathrm{argmin}}_{p+q\leq T_{b}\leq T}Q_{T}\left(T_{b}\right)$
where $Q_{T}\left(T_{b}\right)\triangleq\widehat{\delta}'_{T_{b}}\left(Z_{2}'MZ_{2}\right)\widehat{\delta}_{T_{b}}$,
$\widehat{\delta}_{T_{b}}$ is the least-squares estimator of $\delta_{Z}^{0}$
when regressing $Y$ on $X$ and $Z_{2}$, and $M=I-X\left(X'X\right)^{-1}X'$.
For brevity, we will write $\widehat{T}_{b}^{\textrm{}}$ for $\widehat{T}_{b}^{\textrm{LS}}$
with the understanding that $\widehat{T}_{b}$ is a sequence indexed
by $T$. Let $\widehat{\delta}=\widehat{\delta}{}_{\widehat{T}_{b}}$.
The estimate of the break fraction is then $\widehat{\lambda}_{b}=\widehat{T}_{b}/T$.
Both in practice and for theoretical analyses, a trimming parameter
$\pi\in\left(0,\,1/2\right)$ is  applied to restrict the minimization
over the interval $\left[T\pi,\,\left(1-\pi\right)T\right]$.

\section{Consistency and Convergence Rate under Fixed Shifts\label{Section Consistency-and-Rate } }

We now establish the consistency and convergence rate of the least-squares
estimator under fixed shifts. Under the classical large-$N$ asymptotics,
related results have been established by \citet{bai:97RES} and \citet{bai/perron:98}.
Early important results for a mean-shift appeared in \citet{yao:87}
and \citet{bhattacharya:87} for an $\mathit{i.i.d.}$ series, \citet{bai:94a}
for linear processes and \citet{picard:85} for a Gaussian autoregressive
model.
\begin{assumption}
\label{Assumption 4 Eigenvalue}There exists an $l_{0}$ such that
for all $l>l_{0},$ the matrices $\left(lh\right)^{-1}\sum_{k=1}^{l}x_{kh}x'_{kh},$
$\left(lh\right)^{-1}\sum_{k=T-l+1}^{T}x_{kh}x'_{kh},$  $\left(lh\right)^{-1}\sum_{k=T_{b}^{0}-l+1}^{T_{b}^{0}}x_{kh}x'_{kh},$
and $\left(lh\right)^{-1}\sum_{k=T_{b}^{0}+1}^{T_{b}^{0}+l}x_{kh}x'_{kh},$
have minimum eigenvalues bounded away from zero in probability.
\end{assumption}

\begin{assumption}
\label{Assumption 5 Identification}Let $Q_{0}\left(T_{b},\,\theta^{0}\right)\triangleq\mathbb{E}\left[Q_{T}\left(T_{b},\,\theta^{0}\right)-Q_{T}\left(T_{b}^{0},\,\theta^{0}\right)\right]$.
There exists a $T_{b}^{0}$ such that $Q_{0}\left(T_{b}^{0},\,\theta^{0}\right)>\sup_{\left(T_{b},\,\theta^{0}\right)\notin\mathbf{B}}Q_{0}\left(T_{b},\,\theta^{0}\right),$
for every open set $\mathbf{B}$ that contains $\left(T_{b}^{0},\,\theta^{0}\right)$.
\end{assumption}
Assumption \ref{Assumption 4 Eigenvalue} is similar to A2 in \citet{bai/perron:98}
and requires enough variation around the break point and at the beginning
and end of the sample. The factor $h^{-1}$ normalizes the observations
so that the assumption is implied by a weak law of large numbers.
Assumption \ref{Assumption 5 Identification} is a standard uniqueness
identification condition. We then have the following results.
\begin{prop}
\label{Proposition (1) - Consistency}Under Assumption \ref{Assumption 1, CT}-\ref{Assumption 3 Break Date}
and \ref{Assumption 4 Eigenvalue}-\ref{Assumption 5 Identification},
$\widehat{\lambda}_{b}\overset{P}{\rightarrow}\lambda_{0}$.
\end{prop}

\begin{prop}
\label{Proposition 2, (Rate of Convergence)}Under Assumption \ref{Assumption 1, CT}-\ref{Assumption 3 Break Date}
and \ref{Assumption 4 Eigenvalue}-\ref{Assumption 5 Identification}
for any $\varepsilon>0$, there exists a $K>0$ such that for all
large $T$, $P(T\left|\widehat{\lambda}_{b}-\lambda_{0}\right|>K)<\varepsilon.$
\end{prop}
We have the same $T$-convergence rate as under large-$N$ asymptotics.
Let $\theta^{0}=((\beta^{0})',\,(\delta_{1}^{0})',$ $(\delta_{2}^{0})')'$.
The fast $T$-rate of convergence implies that the least-squares estimate
of $\theta^{0}$ is the same as when $\lambda_{0}$ is known. A natural
estimator for $\theta^{0}$ is $\mathrm{\mathrm{\mathrm{argmin}}_{\beta\in\mathbb{R}^{\mathit{p+q}},\delta\in\mathbb{R}^{\mathit{p}}}}||Y-X\beta-\widehat{Z}_{2}\delta||^{2}$,
where we use $\widehat{T}_{b}$ instead of $T_{b}$ in the construction
of $\widehat{Z}_{2}$. Then we have the following result, akin to
an extension of corresponding results in Section 3 of \citet{barndorff/shephard:04}.
As a matter of notation, let $\Sigma^{*}\triangleq\left\{ \mu_{\cdot,t},\,\Sigma_{\cdot,t},\,\sigma_{e,t}\right\} _{t\geq0}$
and denote expectation taken with respect to $\Sigma^{*}$ by $\mathbb{E}^{*}$.
\begin{prop}
\label{Proposition OLS Asymtptoc Distribu}Under Assumption \ref{Assumption 1, CT}-\ref{Assumption 3 Break Date}
and \ref{Assumption 4 Eigenvalue}-\ref{Assumption 5 Identification},
we have as $T\rightarrow\infty$ ($N$ fixed), conditionally on $\Sigma^{*},$
$(\sqrt{T/N}(\widehat{\beta}-\beta^{0}),\,\sqrt{T/N}(\widehat{\delta}-\delta_{Z}^{0}))'\overset{d}{\rightarrow}\mathscr{MN}\left(0,\,V\right)$
where $\mathscr{MN}$ denotes a mixed Gaussian distribution, with
\begin{align*}
V & \triangleq\overline{V}^{-1}\underset{T\rightarrow\infty}{\lim}T\begin{bmatrix}\sum_{k=1}^{T}\mathbb{E}^{*}\left(x_{kh}x'_{kh}e_{kh}^{2}\right) & \sum_{k=T_{b}^{0}}^{T}\mathbb{E}^{*}\left(x_{kh}z'_{kh}e_{kh}^{2}\right)\\
\sum_{k=T_{b}^{0}}^{T}\mathbb{E}^{*}\left(x_{kh}z'_{kh}e_{kh}^{2}\right) & \sum_{k=T_{b}^{0}}^{T}\mathbb{E}^{*}\left(z_{kh}z'_{kh}e_{kh}^{2}\right)
\end{bmatrix}\overline{V}^{-1},
\end{align*}
and
\begin{align*}
\overline{V} & \triangleq\underset{T\rightarrow\infty}{\lim}\begin{bmatrix}\sum_{k=1}^{T}\mathbb{E}^{*}\left(x_{kh}x'_{kh}\right) & \sum_{k=T_{b}^{0}}^{T}\mathbb{E}^{*}\left(x_{kh}z'_{kh}\right)\\
\sum_{k=T_{b}^{0}}^{T}\mathbb{E}^{*}\left(x_{kh}x'_{kh}\right) & \sum_{k=T_{b}^{0}}^{T}\mathbb{E}^{*}\left(z_{kh}z'_{kh}\right)
\end{bmatrix}.
\end{align*}
\end{prop}
 $V$ can be random because   $\Sigma_{\cdot,t}$ and $\sigma_{e,t}$
can be stochastic. Under fixed shifts, Proposition \ref{Proposition (1) - Consistency}-\ref{Proposition OLS Asymtptoc Distribu}
shows the asymptotic equivalence of discrete and continuous-time regression
models with a change-point, a result corresponding to \citet{brown/low:1996}
for nonparametric regression.

\section{\label{Section Asymptotic Distribution: Continuous Case}Asymptotic
Distribution under a Continuous Record}

We now present results about the limiting distribution of the least-squares
estimate of the break date under a continuous record framework. As
in the classical large-$N$ asymptotics, it depends on the exact distribution
of the data and the errors for fixed break sizes {[}c.f., \citet{hinkley:71}{]}.
This has forced researchers to consider a shrinkage asymptotic theory
where the size of the shift is made local to zero as $T$ increases,
an approach developed by \citet{picard:85} and \citet{yao:87}. We
continue with this avenue. Given the consistency result, we know that
there exists some $h^{*}$ such that for all $h<h^{*}$ with high
probability $\eta Th\leq\widehat{N}_{b}\leq\left(1-\eta\right)Th$,
for $\eta>0$ such that $\lambda_{0}\in\left(\eta,\,1-\eta\right)$.
By Proposition \ref{Proposition 2, (Rate of Convergence)}, $\widehat{N}_{b}-N_{b}^{0}=O_{p}\left(T^{-1}\right)$,
i.e., $\widehat{N}_{b}$ is in a shrinking neighborhood of $N_{b}^{0}$.
With a certain rescaling of the objective function one can first obtain
the shrinkage asymptotic distribution of \citet{bai:97RES}. However,
this is unsatisfactory for two reasons. First, as we show below {[}see
also Casini and Perron (\citeyear{casini/perron_SC_BP_Lap}; \citeyear{casini/perron_Lap_CR_Single_Inf}){]},
the shrinkage asymptotic distribution provides a poor approximation
to the finite-sample distribution of the least-squares estimator.
Second, the latter point also explains the poor coverage properties
of the confidence intervals derived from the shrinkage asymptotic
distribution when the magnitude of the break is not large. Some related
results were obtained by Jiang et al. \citeyearpar{jiang/wang/yu:16}
for a simple location model. Their approach is, however, quite restrictive
and no feasible inference procedure suggested. See the supplement
for a more complete discussion.

We begin with the following assumption which specifies that i) we
use a shrinking condition on $\delta_{Z}^{0}$; ii) we introduce a
locally increasing variance condition on the residual process. The
first is similarly used under classical large-$N$ asymptotics, while
the second is new and useful in our context in order to accurately
capture the relevant uncertainty in the change-point problem. We
do not impose restrictions only on $\delta_{Z}^{0}$ but also on the
ratio $\delta_{Z}^{0}/\sigma_{t}$ when $t$ is close to $T_{b}^{0}$.
We refer to $\delta_{Z}^{0}/\sigma_{t}$ as the signal-to-noise ratio.
Controlling this ratio rather than just $\delta_{Z}^{0}$ allows for
an alternative characterization of the uncertainty about the change-point
date in order to obtain an asymptotic distribution which provides
a better approximation of the finite-sample distribution of the estimator.
To emphasize that $\delta_{Z}^{0}$ depends on the sample-size we
denote it by $\delta_{h}$.
\begin{assumption}
\label{Assumption 6 - Small Shifts}Let $\delta_{h}=\delta^{0}h^{1/4}$,
$\delta^{0}\in\mathbb{R}^{p}$ and assume that for all $t\in\left(N_{b}^{0}-\epsilon,\,N_{b}^{0}+\epsilon\right),$
with $\epsilon\downarrow0$ and $T^{1-\kappa}\epsilon\rightarrow B<\infty$,
$0<\kappa<1/2$, $\mathbb{E}[\left(\Delta_{h}e_{t}^{*}\right)^{2}|\,\mathscr{F}_{t-h}]=\sigma_{h,t-h}^{2}\Delta t$
$P$-a.s., where $\sigma_{h,t}\triangleq\sigma_{h}\sigma_{e,t}$,
$\sigma_{h}\triangleq\overline{\sigma}h^{-1/4}$ and $\overline{\sigma}\triangleq\int_{0}^{N}\sigma_{e,s}^{2}ds$.
\end{assumption}
Note that the localization parameter $\delta^{0}$ in the definition
of $\delta_{h}$ is different from the fixed parameter $\delta_{Z}^{0}$
since $h\rightarrow0$. The rate $1/4$ in the conditions $\delta_{h}=O(h^{1/4})$
and $\sigma_{h}=O(h^{-1/4})$ is for tractability. One can show that
consistency also holds for a rate faster than $1/4$, though slower
than $\kappa.$ However, for the derivation of the limiting distribution
one needs $\delta_{h}/\sigma_{h}=O(h^{1/2})$ and $O(\delta_{h})=O(\sigma_{h}^{-1})$
with $\kappa<1/2.$ The vector of scaled true parameters is $\theta{}_{h}\triangleq((\beta^{0})',\,\delta'_{h})'$.
Define
\begin{align}
\Delta_{h}\widetilde{e}_{t} & \triangleq\begin{cases}
\Delta_{h}e_{t}^{*}, & t\notin\left(N_{b}^{0}-\epsilon,\,N_{b}^{0}+\epsilon\right)\\
h^{1/4}\Delta_{h}e_{t}^{*}, & t\in\left(N_{b}^{0}-\epsilon,\,N_{b}^{0}+\epsilon\right)
\end{cases}.\label{Eq. eps WN}
\end{align}
We shall refer to $\left\{ \Delta_{h}\widetilde{e}_{t},\,\mathscr{F}_{t}\right\} $
as the normalized residual process. Under this framework, the rate
of convergence of $\widehat{N}_{b}$ is now $T^{1-\kappa}$ with $0<\kappa<1/2$.
Due to the fast rate of convergence of the change-point estimator,
the objective function oscillates too rapidly as $h\downarrow0$.
By scaling up the volatility of the errors around the change-point,
we make the objective function behave as if it were a function of
a standard diffusion process. The neighborhood in which the errors
have relatively higher variance is shrinking at a rate $1/T^{1-\kappa}$,
the rate of convergence of $\widehat{N}_{b}.$ Hence, in a neighborhood
of $N_{b}^{0}$ in which we study the limiting behavior of the break
point estimator, the rescaled criterion function is regular enough
so that a feasible limit theory can be developed. The rate of convergence
$T^{1-\kappa}$ is still sufficiently fast to guarantee a $\sqrt{T}$-consistent
estimation of the slope parameters, as stated in the following proposition.
Let $\left\langle Z_{\Delta},\,Z_{\Delta}\right\rangle \left(v\right)$
be the predictable quadratic variation process of $Z_{\Delta}$. The
process $\mathscr{W}\left(v\right)$ is, conditionally on $\mathscr{F}$,
a two-sided centered Gaussian martingale with independent increments
and variances given in Section \ref{subsection Description Limiting Process}
of the supplement.
\begin{prop}
\label{Prop 3 Asym}Under Assumption \ref{Assumption 1, CT}-\ref{Assumption 3 Break Date},
\ref{Assumption 4 Eigenvalue}-\ref{Assumption 5 Identification}
and \ref{Assumption 6 - Small Shifts}, (i) $\widehat{\lambda}_{b}\overset{P}{\rightarrow}\lambda_{0}$;
(ii) for every $\varepsilon>0$ there exists a $K>0$ such that for
all large $T,$ $P(T^{1-\kappa}\left|\widehat{\lambda}_{b}-\lambda_{0}\right|>K||\delta^{0}||^{-2}\overline{\sigma}^{2})<\varepsilon$;
and (iii) for $\kappa\in(0,\,1/4],$ $(\sqrt{T/N}(\widehat{\beta}-\beta^{0}),\,\sqrt{T/N}(\widehat{\delta}-\delta_{h}))'\overset{d}{\rightarrow}\mathscr{MN}\left(0,\,V\right)$
as $T\rightarrow\infty,$ with $V$ given in Proposition \ref{Proposition OLS Asymtptoc Distribu}.
\end{prop}
We first present a general result which shows that under Assumption
\ref{Assumption 6 - Small Shifts} one can obtain a shrinkage asymptotic
distribution similar to \citet{bai:97RES}. The latter exploits the
consistency of $\widehat{\lambda}_{b}$ and the fact that mixing conditions
implies that the regimes before and after $\lambda_{0}$ are asymptotically
independent.  Let $Z_{\Delta}\triangleq(0,\ldots,\,0,\,z_{\left(T_{b}+1\right)h},\ldots,\,z_{T_{b}^{0}h},\,0,\ldots,\,0)$
if $T_{b}<T_{b}^{0}$ and $Z_{\Delta}\triangleq(0,\ldots,\,0,\,z_{\left(T_{b}^{0}+1\right)h},\ldots,\,z_{T_{b}h},$
$0,\ldots,\,0)$ if $T_{b}>T_{b}^{0}$.
\begin{prop}
\label{Prop. Slow Time Scale}Under Assumption \ref{Assumption 1, CT}-\ref{Assumption 3 Break Date},
\ref{Assumption 4 Eigenvalue}-\ref{Assumption 5 Identification}
and \ref{Assumption 6 - Small Shifts},
\begin{align}
T^{1-\kappa}(\widehat{\lambda}_{b}-\lambda_{0}) & \overset{\mathcal{L}\mathrm{-}\mathrm{s}}{\Rightarrow}\underset{v\in\left(-\infty,\,\infty\right)}{\mathrm{argmax}}2(\delta^{0})'\mathscr{W}\left(v\right).\label{CR Asym Dist =00003D Bai 97}
\end{align}
\end{prop}
The distribution in Proposition \ref{Prop. Slow Time Scale} is different
from \citet{bai:97RES}. One can show that his distribution can be
obtained under a continuous record if Assumption \ref{Assumption 6 - Small Shifts}
is modified as follows: $\delta_{h}=\delta^{0}h^{\kappa/2}$, $T^{1-\kappa}\epsilon\rightarrow B<\infty$,
$0<\kappa\leq1/2$ and $\sigma_{h}\triangleq\overline{\sigma}h^{-\kappa/2}$.
This would result in,
\begin{align}
T^{1-\kappa} & \left(\widehat{\lambda}_{b}-\lambda_{0}\right)\overset{\mathcal{L}\mathrm{-}\mathrm{s}}{\Rightarrow}\label{CR Asym Dist =00003D Bai 97-1}\\
 & \underset{v\in\left(-\infty,\,\infty\right)}{\mathrm{argmax}}\left\{ -\left(\delta^{0}\right)'\left\langle Z_{\Delta},\,Z_{\Delta}\right\rangle \left(v\right)\delta^{0}+2\left(\delta^{0}\right)'\mathscr{W}\left(v\right)\right\} .\nonumber
\end{align}
The difference between \eqref{CR Asym Dist =00003D Bai 97} and \eqref{CR Asym Dist =00003D Bai 97-1}
is the presence of the drift (or deterministic) part $-\left(\delta^{0}\right)'$
$\left\langle Z_{\Delta},\,Z_{\Delta}\right\rangle \left(v\right)\delta^{0}$.
Without relating the magnitude of the break to the local variance
condition, the order of the stochastic part dominates that of the
deterministic part and so the latter vanishes asymptotically. The
distributions in \eqref{CR Asym Dist =00003D Bai 97}-\eqref{CR Asym Dist =00003D Bai 97-1}
share the same issues as Bai's and so they do not add any particular
insight. We therefore move to discuss how to obtain a more useful
continuous record asymptotic distribution.

Consider the set $\mathscr{\mathcal{D}}\left(C\right)\triangleq\left\{ N_{b}:\,N_{b}\in\left\{ N_{b}^{0}+Ch^{1-\kappa}\right\} ,\,\left|C\right|<\infty\right\} $,
on the original time scale. Let $\psi_{h}\triangleq h^{1-k}$. Here
we use the same device as in \citeauthor{foster/nelson:96} (\citeyear{nelson/foster:94};
\citeyear{foster/nelson:96})\nocite{nelson/foster:94}. Different
scaling factors applied to an objective function can lead to different
asymptotic distributions. We normalize $Q_{T}(T_{b})$ by $\psi_{h}$,
where $\psi_{h}$ corresponds to the rate of convergence in Proposition
\ref{Prop 3 Asym}. The rate of convergence implicitly describes the
order of the terms in the expansion of $Q_{T}\left(T_{b}\right)-Q_{T}\left(T_{b}^{0}\right)$.

\begin{lem}
\label{Lemma 1}Under Assumption \ref{Assumption 1, CT}-\ref{Assumption 3 Break Date},
\ref{Assumption 4 Eigenvalue}-\ref{Assumption 5 Identification}
and \ref{Assumption 6 - Small Shifts},
\begin{align}
( & Q_{T}\left(T_{b}\right)-Q_{T}\left(T_{b}^{0}\right))/\psi_{h}\label{Eq. (1), Lemma 1}\\
 & =-\delta'_{h}\left(Z_{\Delta}'Z_{\Delta}/\psi_{h}\right)\delta_{h}+2\delta_{h}'\left(Z'_{\Delta}e/\psi_{h}\right)\mathrm{sgn}\left(T_{b}^{0}-T_{b}\right)+o_{p}\left(h^{1/2}\right).\nonumber
\end{align}
\end{lem}
For brevity, we use the notation $\pm$ in place of $\mathrm{sgn}\left(T_{b}^{0}-T_{b}\right)$,
henceforth. The conditional first moment of the centered criterion
function $Q_{T}\left(T_{b}\right)-Q_{T}\left(T_{b}^{0}\right)$ is
of order $O\left(h^{1-\kappa}\right)$, i.e., it ``oscillates''
rapidly as $h\downarrow0$. Hence, in order to approximate the behavior
of $\{\widehat{T}_{b}-T_{b}^{0}\}$, we proceed as in Section 3 in
\citet{nelson/foster:94} and rescale ``time''. For any $C>0$,
let $L_{C}\triangleq N_{b}^{0}-Ch^{1-\kappa}$ and $R_{C}\triangleq N_{b}^{0}+Ch^{1-\kappa}$,
where $L_{C}$ and $R_{C}$ are the left and right boundary points
of $\mathcal{D}\left(C\right)$, respectively. We then have $\left|R_{C}-L_{C}\right|=O\left(Ch^{1-\kappa}\right)$.
Now, take the vanishingly small interval $\left[L_{C},\,R_{C}\right]$
on the original time scale, and stretch it into a time interval $\left[T^{1-\kappa}L_{C},\,T^{1-\kappa}R_{C}\right]$
on a new ``fast time scale''. Changing time scale simply means that
we rescale the objective function in such a way that it is of higher
order as $h\downarrow0$, i.e., it fluctuates less. This leads to
an asymptotic distribution that accounts for higher uncertainty. Yet,
under our framework it is still possible to consistently estimate
the break fraction and the regression coefficients so that inference
is feasible.

Since the criterion function is scaled by $\psi_{h}^{-1}$, all scaled
processes are $O_{p}\left(1\right)$. Now, let $N_{b}\left(v\right)=N_{b}^{0}-vh^{1-\kappa},\,v\in\left[-C,\,C\right]$.
Using Lemma \ref{Lemma 1} and Assumption \ref{Assumption 6 - Small Shifts}
(see the appendix),
\begin{align*}
\psi_{h}^{-1} & \left(Q_{T}\left(T_{b}\left(v\right)\right)-Q_{T}\left(T_{b}^{0}\right)\right)=\\
 & -\delta'_{h}\left(\sum_{k=T_{b}\left(v\right)+1}^{T_{b}^{0}}\frac{z_{kh}}{\sqrt{\psi_{h}}}\frac{z'_{kh}}{\sqrt{\psi_{h}}}\right)\delta_{h}\pm2\left(\delta^{0}\right)'\sum_{k=T_{b}\left(v\right)+1}^{T_{b}^{0}}\frac{z_{kh}}{\sqrt{\psi_{h}}}\frac{\widetilde{e}{}_{kh}}{\sqrt{\psi_{h}}}+o_{p}\left(h^{1/2}\right).
\end{align*}
where $\widetilde{e}_{kh}\triangleq h^{1/4}e_{kh}.$ In addition,
in view of \eqref{Model Regressors Integral Form}, we let $dZ_{\psi,s}=\psi_{h}^{-1/2}\sigma_{Z,s}dW_{Z,s}$
for $s\in\left[N_{b}^{0}-vh^{1-\kappa},\,N_{b}^{0}+vh^{1-\kappa}\right]$.
Applying the time scale change $s\rightarrow t\triangleq\psi_{h}^{-1}s$
to all processes including $\Sigma^{0}$, we have $dZ_{\psi,t}=\sigma_{Z,t}dW_{Z,t}$
with $t\in\mathcal{T}\left(C\right)$, where $\mathcal{T}\left(C\right)\triangleq\{t:\,t\in\left[N_{b}^{0}+v\left\Vert \delta^{0}\right\Vert ^{2}/\overline{\sigma}^{2}\right],\,\left|v\right|\leq C\}$.
Therefore,
\begin{align*}
\psi_{h}^{-1} & \left(Q_{T}\left(T_{b}\left(v\right)\right)-Q_{T}\left(T_{b}^{0}\right)\right)\\
 & =-\delta'_{h}\left(\sum_{k=T_{b}\left(v\right)+1}^{T_{b}^{0}}z_{\psi,kh}z'_{\psi,kh}\right)\delta_{h}\pm2\left(\delta^{0}\right)'\sum_{k=T_{b}\left(v\right)+1}^{T_{b}^{0}}z_{\psi,kh}\widetilde{e}_{\psi,kh}+o_{p}\left(h^{1/2}\right),
\end{align*}
with $NT_{b}\left(v\right)/T=N_{b}\left(v\right)=N_{b}^{0}+v$, where
$z_{\psi,kh}\triangleq z_{kh}/\sqrt{\psi_{h}}$ and $\widetilde{e}_{\psi,kh}\triangleq\widetilde{e}_{kh}/\sqrt{\psi_{h}}$.
Because of the change of time scale, all processes in the last display
are scaled up to be $O_{p}\left(1\right)$ and thus behave as diffusion-like
processes. On this new ``fast time scale'', we have $T^{1-\kappa}R_{C}-T^{1-\kappa}L_{C}=O\left(1\right)$
and $Q_{T}\left(T_{b}\left(v\right)\right)-Q_{T}\left(T_{b}^{0}\right)$
is restored to be $O_{p}\left(1\right)$. Observe that changing the
time scale does not affect any statistic which depends on observations
from $k=1$ to $k=\left\lfloor L_{C}/h\right\rfloor $ or from $k=\left\lfloor R_{C}/h\right\rfloor $
to $k=T$ (since these involve a positive fraction of data). However,
it does affect quantities which include observations that fall in
$\left[T_{b}h,\,T_{b}^{0}h\right]$ (assuming $T_{b}<T_{b}^{0}$).
In particular, on the original time scale, the processes $\left\{ D_{t}\right\} ,\,\left\{ Z_{t}\right\} $
and $\left\{ e_{t}\right\} $ are well-defined and scaled to be $O_{p}\left(1\right)$
while $Q_{T}\left(T_{b}\right)-Q_{T}\left(T_{b}^{0}\right)$ (asymptotically)
oscillates more rapidly than a simple diffusion-type process. On the
new ``fast time scale'', $\left\{ D_{t}\right\} ,\,\left\{ Z_{t}\right\} $
and $\left\{ e_{t}\right\} $ are not affected since they have the
same order in $\left[T^{1-\kappa}L_{C},\,T^{1-\kappa}R_{C}\right]$
as $h\downarrow0$. That is, the first conditional moments are $O\left(h\right)$
while the corresponding moments for $Q_{T}\left(T_{b}\right)-Q_{T}\left(T_{b}^{0}\right)$
on $\mathcal{T}\left(C\right)$ are restored to be $O\left(h\right)$.
As the continuous-time limit is approached, the rescaled criterion
function $\left(Q_{T}\left(T_{b}\left(v\right)\right)-Q_{T}\left(T_{b}^{0}\right)\right)/h^{1/2}$\textit{\textcolor{red}{{}
}}operates on a ``fast time scale'' on $\mathcal{T}\left(C\right)$.

Our analysis is local; we examine the limiting behavior of the centered
and rescaled criterion function process in a neighborhood $\mathcal{T}\left(C\right)$
of the the true break date $N_{b}^{0}$ defined on a new time scale.
We first obtain the weak convergence results for the statistic $\left(Q_{T}\left(T_{b}\left(v\right)\right)-Q_{T}\left(T_{b}^{0}\right)\right)/h^{1/2}$
and then apply a continuous mapping theorem for the argmax functional.
However, it is convenient to work with a re-parametrized objective
function. Proposition \ref{Prop 3 Asym} allows us to use
\begin{align*}
\overline{Q}_{T}\left(\theta^{*}\right) & =\left(Q_{T}\left(\theta_{h},\,T_{b}\left(v\right)\right)-Q_{T}\left(\theta^{0},\,T_{b}^{0}\right)\right)/h^{1/2},
\end{align*}
where $\theta^{*}\triangleq\left(\theta'_{h},\,v\right)'$ with $T_{b}\left(v\right)\triangleq T_{b}^{0}+\left\lfloor v/h\right\rfloor $
and $T_{b}\left(v\right)$ is the time index on the ``fast time scale''.
The normalizing factor $\psi_{h}h^{1/2}$ allows us to change the
time scale and obtain an alternative asymptotic distribution. When
$v$ varies, $T_{b}\left(v\right)$ potentially visits all integers
between $1$ and $T$. Thus, on the new time scale, we need to introduce
the trimming parameter $\pi\in\left(0,\,1\right)$ which determines
the region where $T_{b}\left(v\right)$ can vary. We have the normalizations
$T_{b}\left(v\right)=T\pi$ if $T_{b}\left(v\right)\leq T\pi$ and
$T_{b}\left(v\right)=T\left(1-\pi\right)$ if $T_{b}\left(v\right)\geq T\left(1-\pi\right)$.
On the old time scale, $N_{b}\left(u\right)=N_{b}^{0}+u$ with $v\rightarrow\psi_{h}^{-1}u$,
so that $N_{b}\left(u\right)$ is in a vanishing neighborhood of $N_{b}^{0}$.
On $\mathcal{T}\left(C\right)$, we index the process $Q_{T}\left(\theta_{h},\,T_{b}\left(v\right)\right)-Q_{T}\left(\theta^{0},\,T_{b}^{0}\right)$
by two time subscripts: one referring to the time $T_{b}$ on the
original time scale and one referring to the time elapsed since $T_{b}h$
on the ``fast time scale''. For simplicity, we omit the former;
since the limiting distribution of the least-squares estimator will
now depend on the trimming we use the notation $\widehat{T}_{b,\pi}=T\widehat{\lambda}_{b,\pi}$
where $\widehat{\lambda}_{b,\pi}$ is the least-squares estimator
of the fractional break date associated to the fast time scale (i.e.,
associated to the  factor $\psi_{h}h^{1/2}$).

The optimization problem is not affected by the change of time scale.
In fact, by Proposition \ref{Prop 3 Asym}, $u=Th(\widehat{\lambda}_{b}-\lambda_{0})=KO_{p}\left(h^{1-\kappa}\right)$
on the old time scale; whereas on the new ``fast time scale'',
$v=Th(\widehat{\lambda}_{b,\pi}-\lambda_{0})=O_{p}\left(1\right)$.
The maximization problem is not changed because $v/h$ can take any
value in $\mathbb{R}$. The process $Q_{T}\left(\theta_{h},\,T_{b}\left(v\right)\right)-Q_{T}\left(\theta^{0},\,T_{b}^{0}\right)$
is thus analyzed on a fixed horizon since $v$ now varies over $[(N\pi-N_{b}^{0})/(||\delta^{0}||^{-2}\overline{\sigma}^{2}),\,(N\left(1-\pi\right)-N_{b}^{0})/(\left\Vert \delta^{0}\right\Vert ^{-2}\overline{\sigma}^{2})]$.
Define  the modification to the set $\mathcal{D}\left(C\right)$
applicable to the new time scale by
\begin{align*}
\mathcal{D}^{*}\left(C\right) & =\biggl\{\left(\beta^{0},\,\delta_{h},\,v\right):\,\left\Vert \theta^{0}\right\Vert \leq C;\,T_{b}\left(v\right)=T_{b}^{0}+vN^{-1}\left\Vert \delta^{0}\right\Vert ^{-2}\overline{\sigma}^{2};\\
 & \quad\frac{\left(N\pi-N_{b}^{0}\right)}{\left\Vert \delta^{0}\right\Vert ^{-2}\overline{\sigma}^{2}}\leq v\leq\frac{N\left(1-\pi\right)-N_{b}^{0}}{\left\Vert \delta^{0}\right\Vert ^{-2}\overline{\sigma}^{2}}\biggr\}.
\end{align*}
Let $\mathbb{D}\left(\mathcal{D}^{*}\left(C\right),\,\mathbb{R}\right)$
denote the space of all \textit{c\`{a}dl\`{a}g} functions from $\mathcal{D}^{*}\left(C\right)$
into $\mathbb{R}.$ Endow this space with the Skorokhod topology.
 Under a continuous record, we can apply limit theorems for statistics
involving (co)variation between regressors and errors. This enables
us to deduce the limiting process for $\overline{Q}_{T}\left(\theta^{*}\right)$,
 relying upon the work of Jacod \citeyearpar{jacod:94,jacod:97}
 and \citet{jacod/protter:98}.

To guide intuition, note that under the new re-parametrization, the
limit law of $\overline{Q}_{T}\left(\theta^{*}\right)$ is, according
to Lemma \ref{Lemma 1}, the same as the limit law of
\begin{align*}
-h^{-1/2} & \delta_{h}'\left(Z_{\Delta}'Z_{\Delta}\right)\delta_{h}\pm2h^{-1/2}\delta_{h}'\left(Z'_{\Delta}e\right)\\
 & \overset{d}{\equiv}-\left(\delta^{0}\right)'\left(Z_{\Delta}'Z_{\Delta}\right)\delta^{0}\pm2h^{-1/2}\left(\delta^{0}\right)'h^{1/4}\left(Z'_{\Delta}h^{-1/4}\widetilde{e}\right),
\end{align*}
where $\overset{d}{\equiv}$ denotes (first order) equivalence in
law, and since (approximately) $e_{kh}\sim i.n.d.\,\mathscr{N}(0,$
$\sigma_{h,k-1}^{2}h),\,\sigma_{h,k}=\sigma_{h}\sigma_{e,k}$ then
$\widetilde{e}_{kh}\sim i.n.d.\,\mathscr{N}\left(0,\,\sigma_{e,k-1}^{2}h\right)$.
Hence, the limit law of $\overline{Q}_{T}\left(\theta^{*}\right)$
is, to first-order, equivalent to the law of
\begin{align}
-\left(\delta^{0}\right)'\left(Z_{\Delta}'Z_{\Delta}\right)\delta^{0}\pm2\left(\delta^{0}\right)'\left(h^{-1/2}Z'_{\Delta}\widetilde{e}\right) & .\label{Eq Intuition Limit Law}
\end{align}
 We apply a law of large numbers to the first term and a stable convergence
in law under the Skorokhod topology to the second. Assumption \ref{Assumption 6 - Small Shifts}
combined with the normalizing factor $h^{-1/2}$ in $\overline{Q}_{T}\left(\theta^{*}\right)$
account for the discrepancy between the deterministic and stochastic
component in \eqref{Eq Intuition Limit Law}.

Having outlined the main steps in the arguments used to derive the
continuous records limit distribution of the break date estimate,
we now state the main result of this section. The limiting process
is realized on a extension of the original probability space and we
relegate this description to Section \ref{subsection Description Limiting Process}
in the supplement.
\begin{thm}
\label{Theorem 1}Under Assumption \ref{Assumption 1, CT}-\ref{Assumption 3 Break Date},
\ref{Assumption 4 Eigenvalue}-\ref{Assumption 5 Identification}
and \ref{Assumption 6 - Small Shifts},
\begin{align}
N & \left(\widehat{\lambda}_{b,\pi}-\lambda_{0}\right)\overset{\mathcal{L}\mathrm{-}\mathrm{s}}{\Rightarrow}\underset{v\in\mathcal{A}}{\mathrm{argmax}}\left\{ -\left(\delta^{0}\right)'\left\langle Z_{\Delta},\,Z_{\Delta}\right\rangle \left(v\right)\delta^{0}+2\left(\delta^{0}\right)'\mathscr{W}\left(v\right)\right\} ,\label{CR Asymptotic Distribution}
\end{align}
where
\begin{align*}
\mathcal{A} & \triangleq\left[\frac{N\pi-N_{b}^{0}}{\left\Vert \delta^{0}\right\Vert ^{-2}\overline{\sigma}^{2}},\,\frac{N\left(1-\pi\right)-N_{b}^{0}}{\left\Vert \delta^{0}\right\Vert ^{-2}\overline{\sigma}^{2}}\right].
\end{align*}
\end{thm}
Note the differences between the results in Theorem \ref{Theorem 1}
and in Proposition \ref{Prop. Slow Time Scale}. First, on the fast
time scale, $\widehat{\lambda}_{b,\pi}$ behaves as an inconsistent
estimator for $\lambda_{0}$ for $N$ fixed, but it is consistent
as $N\rightarrow\infty$. On the original time scale $\widehat{\lambda}_{b}$
is not only consistent for $\lambda_{0}$ but it also enjoys a similar
asymptotic distribution as in \citet{bai:97RES}. Second, the asymptotic
distribution of $\widehat{\lambda}_{b,\pi}$ depends on the span of
the data and consequently on the trimming $\pi.$ Proposition \ref{Prop. Slow Time Scale},
in contrast, suggests that the span, the trimming and the location
of the break are irrelevant for the limiting behavior of the estimator.
This intuitively follows from the fact that under the original time
scale the break date estimator is consistent. We will show that indeed
the span of the data and the location of the break influence the finite-sample
properties of the least-squares estimator, and that Theorem \ref{Theorem 1}
provides a more useful approximation. An important implication of
Theorem \ref{Theorem 1} is that the precision of the estimator depends
more on the span $N$ than to the number of observations $T$.

Unlike Bai's distribution, the distribution in Theorem \ref{Theorem 1}
involves the location of the maximum of a function of the (quadratic)
variation of the regressors and of a two-sided centered Gaussian martingale
process over the interval $[(N\pi-N_{b}^{0})/(\left\Vert \delta^{0}\right\Vert ^{-2}\overline{\sigma}^{2}),\,(N\left(1-\pi\right)-N_{b}^{0})/(||\delta^{0}||^{-2}\overline{\sigma}^{2})]$.
Notably, this domain depends on the true value of $N_{b}^{0}$ and
therefore the limit distribution is asymmetric, in general. The degree
of asymmetry increases as the true break point moves away from mid-sample.
This holds even when the distributions of the errors and regressors
are the same in the pre- and post-break regimes. The presence of the
trimming confirms that the span of the (trimmed) data affects the
 limit distribution. It is well-known that the least-squares estimator
of the break date can be sensitive to trimming {[}see \citet{bai/perron:03}
for some recommendations on the trimming choice{]}. Our asymptotic
theory accommodates this property of the least-squares estimator while
others do not.

Additional relevant remarks follow; more details are provided in the
supplement. The magnitude of the break plays a key role in determining
the density of the asymptotic distribution. More precisely, the density
displays interesting properties which change when the signal-to-noise
ratio as well as other parameters of the model change. Moreover, the
distribution in Theorem \ref{Theorem 1} is able to reproduce important
features of the small-sample results obtained via simulations {[}e.g.,
\citet{bai/perron:06}{]}. First, the second moments of the regressors
impact the asymptotic mean as well as the second-order behavior of
the break point estimator (e.g., the persistence of the regressors
influences the finite-sample performance of the estimator). Second,
the continuous record setting manages to preserve information about
the time span $N$ of the data, a clear advantage since the location
of the true break point matters for the small-sample distribution
of the estimator. It has been shown via simulations that in small-samples
the break point estimator tends to be imprecise if the break size
is small, and some bias arises if the break point is not at mid-sample.
In our framework, the (trimmed) time horizon $\left[N\pi,\,N\left(1-\pi\right)\right]$
is fixed and thus we can distinguish between the statistical content
of the segments $\left[N\pi,\,N_{b}^{0}\right]$ and $\left[N_{b}^{0},\,N\left(1-\pi\right)\right]$.
In contrast, this is not feasible under the classical shrinkage large-$N$
asymptotics because both the pre- and post-break segments increase
 proportionately and mixing conditions are imposed so that the only
relevant information is a neighborhood around the true break date.
Details on how to simulate the limiting distribution in Theorem \ref{Theorem 1}
are given in Section \ref{subsec:Simulation-of-the LD in Theorem 4.1}
of the supplement.

We further characterize the asymptotic distribution by exploiting
the ($\mathscr{F}$-conditionally) Gaussian property of the limit
process. The analysis also holds unconditionally if we assume that
the volatility processes are non-stochastic. Thus, as in the classical
setting, we begin with a second-order stationarity assumption within
each regime. The following assumption guarantees that the results
below remain valid without the need to condition on $\mathscr{F}.$
\begin{assumption}
\label{Assumtpion - Regimes}The process $\Sigma^{0}$ is (possibly
time-varying) deterministic; $\left\{ z_{kh},\,e_{kh}\right\} $ is
second-order stationary within each regime. For $k=1,\ldots,\,T_{b}^{0}$,
$\mathbb{E}(z_{kh}z'_{kh}|\,\mathscr{F}_{\left(k-1\right)h})=\Sigma_{Z,1}h$,
$\mathbb{E}(\widetilde{e}_{kh}^{2}|$ $\mathscr{F}_{\left(k-1\right)h})=\sigma_{e,1}^{2}h$
and $\mathbb{E}(z_{kh}z'_{kh}\widetilde{e}_{kh}^{2}|\,\mathscr{F}_{\left(k-1\right)h})=\Omega_{\mathscr{W},1}h^{2}$
while for $k=T_{b}^{0}+1,\ldots,\,T$, $\mathbb{E}(z_{kh}z'_{kh}|$
$\mathscr{F}_{\left(k-1\right)h})=\Sigma_{Z,2}h$, $\mathbb{E}(\widetilde{e}_{kh}^{2}|\,\mathscr{F}_{\left(k-1\right)h})=\sigma_{e,2}^{2}h$
and $\mathbb{E}(z_{kh}z'_{kh}\widetilde{e}_{kh}^{2}|\,\mathscr{F}_{\left(k-1\right)h})=\Omega_{\mathscr{W},2}h^{2}$.
\end{assumption}
Let $W_{i}^{*},$ $i=1,\,2,$ be two independent standard Wiener processes
defined on $[0,\,\infty),$ starting at the origin when $s=0.$ Let
\begin{align*}
\mathscr{V}\left(s\right) & =\begin{cases}
-\frac{\left|s\right|}{2}+W_{1}^{*}\left(s\right), & \textrm{if }s<0\\
-\frac{\left(\delta^{0}\right)'\Sigma_{Z,2}\delta^{0}}{\left(\delta^{0}\right)'\Sigma_{Z,1}\delta^{0}}\frac{\left|s\right|}{2}+\left(\frac{\left(\delta^{0}\right)'\Omega_{\mathscr{W},2}\left(\delta^{0}\right)}{\left(\delta^{0}\right)'\Omega_{\mathscr{W},1}\left(\delta^{0}\right)}\right)^{1/2}W_{2}^{*}\left(s\right), & \textrm{if }s\geq0.
\end{cases}
\end{align*}

\begin{thm}
\label{Theorem 2, Asymptotic Distribution immediate Stationary Regimes}Under
Assumption \ref{Assumption 1, CT}-\ref{Assumption 3 Break Date},
\ref{Assumption 4 Eigenvalue}-\ref{Assumption 5 Identification}
and \ref{Assumption 6 - Small Shifts}-\ref{Assumtpion - Regimes},
\begin{align}
\frac{\left(\left(\delta^{0}\right)'\left\langle Z,\,Z\right\rangle _{1}\delta^{0}\right)^{2}}{\left(\delta^{0}\right)'\Omega_{\mathscr{W},1}\delta^{0}}N\left(\widehat{\lambda}_{b,\pi}-\lambda_{0}\right) & \Rightarrow\underset{s\in\mathcal{A}^{*}}{\mathrm{argmax}}\mathscr{V}\left(s\right),\label{Equation (2) Asymptotic Distribution}
\end{align}
 where
\begin{align*}
\mathcal{A}^{*} & \triangleq\left[\frac{N\pi-N_{b}^{0}}{\left\Vert \delta^{0}\right\Vert ^{-2}\overline{\sigma}^{2}}\frac{\left(\left(\delta^{0}\right)'\left\langle Z,\,Z\right\rangle _{1}\delta^{0}\right)^{2}}{\left(\delta^{0}\right)'\Omega_{\mathscr{W},1}\left(\delta^{0}\right)},\,\frac{N\left(1-\pi\right)-N_{b}^{0}}{\left\Vert \delta^{0}\right\Vert ^{-2}\overline{\sigma}^{2}}\frac{\left(\left(\delta^{0}\right)'\left\langle Z,\,Z\right\rangle _{1}\delta^{0}\right)^{2}}{\left(\delta^{0}\right)'\Omega_{\mathscr{W},1}\delta^{0}}\right].
\end{align*}
\end{thm}
Unlike the asymptotic distribution derived under classical large-$N$
asymptotics, the probability density  in \eqref{Equation (2) Asymptotic Distribution}
is not available in closed form. Furthermore, the limiting distribution
depends on unknown quantities. In the next section we explain how
one can derive a feasible counterpart. This will be useful to characterize
the main features of interest that will guide us in devising methods
to construct confidence sets for $T_{b}^{0}$.

\section{Feasible Approximations to the Finite-Sample Distributions\label{Section Approximation-to-the}}

In Section \ref{subsec:Evaluation-of-the} we propose a feasible version
of our limit theory and compare it with the finite-sample distribution.
In Section \ref{Subsec:Comparison-with Bai and EM} we discuss some
differences between our approach and others. Let
\begin{align*}
\rho=\frac{\left(\left(\delta^{0}\right)'\left\langle Z,\,Z\right\rangle _{1}\delta^{0}\right)^{2}}{\left(\left(\delta^{0}\right)'\Omega_{\mathscr{W},1}\delta^{0}\right)},\quad\xi_{1}=\frac{\left(\delta^{0}\right)'\left\langle Z,\,Z\right\rangle _{2}\delta^{0}}{\left(\delta^{0}\right)'\left\langle Z,\,Z\right\rangle _{1}\delta^{0}},\quad & \xi_{2}=\frac{\left(\delta^{0}\right)'\Omega_{\mathscr{W},2}\delta^{0}}{\left(\delta^{0}\right)'\Omega_{\mathscr{W},1}\delta^{0}}.
\end{align*}


\subsection{\label{subsec:Evaluation-of-the}A Feasible Version of the Limit
Distribution}

In order to use the continuous record asymptotic distribution in practice
one needs consistent estimates of the unknown quantities. In this
section, we compare the finite-sample distribution of the least-squares
estimator of the change-point date with a feasible version of the
continuous record asymptotic distribution obtained with plug-in estimates.
We obtain the finite-sample distribution of $\rho(\widehat{T}_{b,\pi}-T_{b}^{0})$
based on 100,000 simulations from the following model:
\begin{align}
Y_{t}=D'_{t}\nu^{0}+Z'_{t}\beta^{0}+Z'_{t}\delta_{Z}^{0}\boldsymbol{1}_{\left\{ t>T_{b}^{0}\right\} }+e_{t}, & \qquad t=1,\ldots,\,T,\label{Casin i-  Model for FS vs Asy Dist}
\end{align}
 where $Z_{t}=0.5Z_{t-1}+u_{t}$ with $u_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right)$
independent of $e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,\sigma_{e}^{2}\right)$,
$\sigma_{e}^{2}=1$, $\nu^{0}=1$, $Z_{0}=0$, $D_{t}=1$ for all
$t$, and $T=100.$ We set $\pi=0.05$, $T_{b}^{0}=\left\lfloor T\lambda_{0}\right\rfloor $
with $\lambda_{0}=0.3,\,0.5,\,0.7$ and consider different break sizes
$\delta_{Z}^{0}=0.2,\,0.3,\,0.5,\,1$. The infeasible continuous record
asymptotic distribution is computed assuming knowledge of the data
generating process (DGP) as well as of the model parameters, i.e.,
using Theorem \ref{Theorem 2, Asymptotic Distribution immediate Stationary Regimes}
where we set $N_{b}^{0}$ equal to its true value, $||\delta^{0}||^{-2}\overline{\sigma}^{2}=||\delta_{Z}^{0}||^{-2}\sigma_{e}^{2},$
and $\xi_{1},\,\xi_{2}$ and $\rho$ equal to their true values, respectively,
with $\delta_{Z}^{0}$ in place of $\delta^{0}$. Note that the scaling
$h^{\kappa/2}$ and $h^{-1/4}$ in the definition of $\delta_{h}$
and $\sigma_{h}$ respectively, cancel using the fact that they appear
in both numerator and denominator and applying a change in variables.
The feasible counterparts are constructed with plug-in estimates of
$\xi_{1},\,\xi_{2},\,\rho$ and $(N_{b}^{0}\left\Vert \delta^{0}\right\Vert ^{2}/\overline{\sigma}^{2})\rho$.
In practice we need to use a normalization for $N$. A common choice
is $N=1$. Then $\widehat{\lambda}_{b}^{\mathrm{}}=\widehat{T}_{b}/T$
 is a natural estimate of $\lambda_{0}$, using the consistency result
of $\widehat{\lambda}_{b}^{\mathrm{}}$  that holds in the setting
of Theorem \ref{Theorem 1} which can also be rationalized for large
$N$ under the conditions of Theorem \ref{Theorem 1}. In practice
this means that we approximate the distribution of the estimator $\widehat{\lambda}_{b,\pi}$
where $\pi$ is chosen by the researcher and we plug-in the estimator
$\widehat{\lambda}_{b}$ which can be based on any  trimming because
of the consistency property. Here, we set $\widehat{\lambda}_{b}$
equal to the least-squares estimator based on a trimming $0.15$,
which is also used for the other plug-in estimates. The estimates
of $\xi_{1}$ and $\xi_{2}$ are given, respectively, by
\begin{align*}
\widehat{\xi}_{1}=\frac{\widehat{\delta}'\left(T-\widehat{T}_{b}^{\mathrm{}}\right)^{-1}\sum_{k=\widehat{T}_{b}^{\mathrm{}}+1}^{T}z_{kh}z'_{kh}\widehat{\delta}}{\widehat{\delta}'\left(\widehat{T}_{b}^{\mathrm{}}\right)^{-1}\sum_{k=1}^{\widehat{T}_{b}^{\mathrm{}}}z_{kh}z'_{kh}\widehat{\delta}} & ,\qquad\widehat{\xi}_{2}=\frac{\widehat{\delta}'\left(T-\widehat{T}_{b}^{\mathrm{}}\right)^{-1}\sum_{k=\widehat{T}_{b}^{\mathrm{}}+1}^{T}\widehat{e}_{kh}^{2}z_{kh}z'_{kh}\widehat{\delta}}{\widehat{\delta}'\left(\widehat{T}_{b}^{\mathrm{}}\right)^{-1}\sum_{k=1}^{\widehat{T}_{b}^{\mathrm{}}}\widehat{e}_{kh}^{2}z_{kh}z'_{kh}\widehat{\delta}},
\end{align*}
where $\widehat{\delta}$ is the least-squares estimator of $\delta_{h}$
and $\widehat{e}_{kh}$ are the least-squares residuals. Note that
in $\widehat{\xi}_{1}$ and $\widehat{\xi}_{2}$, the estimate $\widehat{\delta}$
appears in both numerator and denominator so that the scaling $h^{\kappa/2}$
in the definition of $\delta_{h}$ cancels. Use is made of the fact
that $\left\langle Z,\,Z\right\rangle _{1}$ is consistently estimated
by $\sum_{k=1}^{\widehat{T}_{b}^{\mathrm{}}}z_{kh}z'_{kh}/\widehat{\lambda}_{b}^{\mathrm{}}$
while $\Omega_{\mathscr{W},1}$ is consistently estimated by $T\sum_{k=1}^{\widehat{T}_{b}^{\mathrm{}}}\widehat{e}_{kh}^{2}z_{kh}z'_{kh}/\widehat{\lambda}_{b}^{\mathrm{}}$.
The method to estimate $\lambda_{0}\left\Vert \delta^{0}\right\Vert ^{2}\overline{\sigma}^{-2}\rho$
is less immediate because it involves manipulating the scaling of
each of the three estimates. Let $\vartheta=\left\Vert \delta^{0}\right\Vert ^{2}\overline{\sigma}^{-2}\rho$.
We use the following estimates for $\vartheta$ and $\rho$, respectively,
\begin{align*}
\widehat{\vartheta}= & \widehat{\rho}\left\Vert \widehat{\delta}\right\Vert ^{2}\left(T^{-1}\sum_{k=1}^{T}\widehat{e}_{kh}^{2}\right)^{-1},\qquad\widehat{\rho}=\frac{\left(\widehat{\delta}'\left(\widehat{T}_{b}^{\mathrm{}}\right)^{-1}\sum_{k=1}^{\widehat{T}_{b}^{\mathrm{}}}z_{kh}z'_{kh}\widehat{\delta}\right)^{2}}{\widehat{\delta}'\left(\widehat{T}_{b}^{\mathrm{}}\right)^{-1}\sum_{k=1}^{\widehat{T}_{b}^{\mathrm{}}}\widehat{e}_{kh}^{2}z_{kh}z'_{kh}\widehat{\delta}},
\end{align*}
Whereas we have $\widehat{\xi}_{i}\overset{p}{\rightarrow}\xi_{i}$
$\left(i=1,\,2\right)$, the corresponding approximations for $\widehat{\rho}$
and $\widehat{\vartheta}$ are given by $\widehat{\rho}/h^{\kappa}\overset{p}{\rightarrow}\rho$
and $\widehat{\vartheta}/h^{2\kappa}\overset{p}{\rightarrow}\vartheta$.
 However, before letting $T\rightarrow\infty$ we can apply a change
in variable using the fact that $\widehat{\lambda}_{b}-\lambda_{b}^{0}=O\left(h^{1-\kappa}\right)$
which result in the extra factor $h^{2\kappa}$ canceling.
\begin{prop}
\label{Proposition CI}Under the conditions of Theorem \ref{Theorem 2, Asymptotic Distribution immediate Stationary Regimes},
\eqref{Equation (2) Asymptotic Distribution} holds when using $\widehat{\xi}_{1},\,\widehat{\xi}_{2},\,\widehat{\rho}$
and $\widehat{\vartheta}$ in place of $\xi_{1},\,\xi_{2},\,\rho$
and $\vartheta$, respectively.
\end{prop}
The proposition implies that the limiting distribution can be simulated
using plug-in estimates. This allows feasible inference about the
break date. The results are presented in Figure \ref{Fig19}-\ref{Fig23}
which also plot the asymptotic distribution from \citet{bai:97RES}
and the infeasible distribution from Theorem \ref{Theorem 2, Asymptotic Distribution immediate Stationary Regimes}.
Here by signal-to-noise ratio we mean $\delta_{Z}^{0}/\sigma_{e}$
which, given $\sigma_{e}^{2}=1$, equals the break size $\delta_{Z}^{0}$.

Several interesting observations appear at the outset. The density
of the large-$N$ shrinkage asymptotic distribution does not depend
on the location of the break, and thus it is always unimodal and symmetric
about the origin. None of these features are shared by the density
derived under a continuous record. When the true break is at mid-sample
$\left(\lambda_{0}=0.5\right)$, the density function is symmetric
and centered at zero. However, when the signal-to-noise ratio is low,
the density features three modes. This tri-modality vanishes as the
signal-to-noise ratio increases. When $\delta_{Z}^{0}$ is low and
the break is not at mid-sample the density is asymmetric; for values
of $\lambda_{0}$ less (larger) than 0.5, the density is right (left)
skewed. When the signal is low and $\lambda_{0}$ is less (larger)
than 0.5, the density has highest mode at some value near $\widehat{\lambda}_{b}$
being close to the starting (end) sample point than centered at $\lambda_{0}.$
However, as in the case of $\lambda_{0}=0.5,$ when the signal-to-noise
ratio increases the highest mode is centered at a value which corresponds
to $\widehat{\lambda}_{b}$ being close to $\lambda_{0}$. Asymmetry
and multi-modality of the finite-sample distribution of the break
point estimator were also found by \citet{perron/zhu:05} and \citet{deng/perron:06}
in models with a trend.

The interpretation of these features are straightforward. For example,
asymmetry reflects the fact that the span of the data and the actual
location of the break play a crucial role on the behavior of the estimator.
If the break occurs early in the sample there is a tendency to overestimate
the break date and vice-versa if the break occurs late in the sample.
The marked changes in the shape of the density as we raise $\delta_{Z}^{0}$
confirms that the magnitude of the shift matters a great deal as well.
The tri-modality of the density when the shift size is small reflects
the uncertainty in the data as to whether a structural change is present
at all; i.e., the least-squares estimator finds it easier to locate
the break at either the beginning or the end of the sample. Unlike
the shrinkage asymptotic distribution, the density of the feasible
version of the continuous record distribution provides a remarkably
good approximation to the infeasible one and thus also to the finite-sample
distribution. The extended working paper \citet{casini/perron_CR_Single_Break_Extended}
shows that the quality of the approximation is good for a variety
of models.

\subsection{\label{Subsec:Comparison-with Bai and EM}Comparison with Other Approaches}

The figures reported above have shown that there is a high degree
of uncertainty when the break magnitude is not large. The classical
shrinkage asymptotics of \citet{bai:97RES} with $\delta_{T}$ required
to convergence to zero at a rate slower than $O(T^{-1/2})$ clearly
underestimates that degree of uncertainty and, as the figures show,
provides a poor approximation to the finite-sample behavior of the
least-squares estimator. In Section \ref{Section Small-Sample-Effectiveness-of}
we show that this issue is responsible for the poor coverage probabilities
of the confidence intervals introduced in \citet{bai:97RES} when
the break magnitude is small. On the other hand, \citet{elliott/mueller:07}
and \citet{elliott/mueller/watson:15} require $\delta_{T}$ to go
to zero at the fast rate $O(T^{-1/2})$ leading to weak identification.
The latter implies that the relevant quantities in the model become
inconsistent. This can be problematic for inference and indeed, their
inference often suffers from the opposite problem in that confidence
intervals for $\widehat{T}_{b}$ can be too large {[}Casini and Perron
(\citeyear{casini/perron_Oxford_Survey}; \citeyear{casini/perron_Lap_CR_Single_Inf})
and \citet{chang/perron:18}{]}.

We impose conditions on the signal-to-noise ratio $\delta/\sigma$
rather than just on $\delta.$ Consider a simple location model with
a change $\delta$ in the mean and independent errors. What describes
the uncertainty about the break in this model is the ratio $\delta/\sigma$
where $\sigma$ is the volatility of the errors. We let $\delta$
go to zero at a not too fast rate while letting $\sigma$ increase
to infinity in a neighborhood of $T_{b}^{0}$. That is $\left(\delta_{T}/\sigma_{t}\right)\rightarrow0$
at rate $O(T^{-1/2})$ in a neighborhood of $T_{b}^{0}$. Interestingly,
this is the same rate Elliott and M{\"u}ller used for $\delta_{T}\rightarrow0.$
Away from $T_{b}^{0}$, we require $\left(\delta_{T}/\sigma_{t}\right)\rightarrow0$
at slower rate---similar to \citet{yao:87} and \citet{bai:97RES}.
The difference now is that we do not lose identification and all the
parameters in the model remain consistent.  Under continuous-time,
the variance of the processes is proportional to the sampling interval.
This allows us to trade-off the rate of convergence at which $\widehat{\lambda}_{b}$
approaches $\lambda_{0}$ with the variance of the errors in a neighborhood
of $T_{b}^{0}$ by letting $\sigma_{t}$ become large when $t$ is
close to $T_{b}^{0}$ {[}i.e., a change of time scale as in Foster
and Nelson (\citeyear{nelson/foster:94}, \citeyear{foster/nelson:96}){]}.
This offers a new characterization of higher uncertainty without losing
identification.

\section{Highest Density Region-based Confidence Sets \label{Section Inference Methods}}

The features of the limit and finite-sample distributions suggest
that standard methods to construct confidence intervals may be inappropriate;
e.g., two-sided intervals around the estimated break date based on
the standard deviations of the estimate. Our suggested approach is
rather non-standard and relates to Bayesian methods. In our context,
the Highest Density Region (HDR) seems the most appropriate in light
of the asymmetry and, especially, the multi-modality of the distribution
for small break sizes. All that is needed to implement the procedure
is an estimate of the density function, using plug-in estimates as
explained in Section \ref{Section Approximation-to-the}. Choose
some significance level $0<\alpha<1$ and let $\widehat{P}_{T_{b}}$
denote the empirical counterpart of the probability distribution of
$\rho N(\widehat{\lambda}_{b,\pi}-\lambda_{b}^{0})$ as defined in
Theorem \ref{Theorem 2, Asymptotic Distribution immediate Stationary Regimes}.
Further, let $\widehat{p}_{T_{b}}$ denote the density function defined
by the Radon-Nikodym equation $\widehat{p}_{T_{b}}^{\mathrm{}}=d\widehat{P}_{T_{b}}^{\mathrm{}}/d\lambda_{\mathrm{L}},$
where $\lambda_{\mathrm{L}}$ denotes the Lebesgue measure.
\begin{defn}
\textbf{Highest Density Region:} Assume that the density function
$f_{Y}\left(y\right)$ of some random variable $Y$ defined on a probability
space $(\Omega_{Y},\,\mathscr{F}_{Y},\,\mathbb{P}_{Y})$ and taking
values on the measurable space $\left(\mathcal{Y},\,\mathscr{Y}\right)$
is continuous and bounded. Then the $\left(1-\alpha\right)100\%$
Highest Density Region is a subset $\mathbf{S}(\kappa_{\alpha})$
of $\mathcal{Y}$ defined as $\mathbf{S}(\kappa_{\alpha})=\{y:\,f_{Y}\left(y\right)>\kappa_{\alpha}\}$
where $\kappa_{\alpha}$ the largest constant that satisfies $\mathbb{P}_{Y}(Y\in\mathbf{S}(\kappa_{\alpha}))\geq1-\alpha$.
\end{defn}
The concept of HDR and of its estimation has an established literature
in statistics. The definition reported here is from \citet{hyndman:96};
see also \citet{samworth/wand:10} and \citeauthor{mason/polonik:09}
(\citeyear{mason/polonik:08}; \citeyear{mason/polonik:09}) for more
recent developments.\nocite{mason/polonik:08}
\begin{defn}
\textbf{Confidence Sets for $T_{b}^{0}$ under a Continuous Record:}\label{Def. Confidence Sets for Break Date}
Under Assumption \ref{Assumption 1, CT}-\ref{Assumption 3 Break Date},
\ref{Assumption 4 Eigenvalue}-\ref{Assumption 5 Identification}
and \ref{Assumption 6 - Small Shifts}-\ref{Assumtpion - Regimes},
a $\left(1-\alpha\right)100\%$ confidence set for $T_{b}^{0}$ is
a subset of $\left\{ 1,\ldots,\,T\right\} $ given by $C\left(\textrm{cv}_{\alpha}\right)=\left\{ T_{b}\in\left\{ 1,\ldots,\,T\right\} :\,T_{b}\in\mathbf{S}\left(\textrm{cv}_{\alpha}\right)\right\} ,$
where $\mathbf{S}\left(\textrm{cv}_{\alpha}\right)=\left\{ T_{b}:\,\widehat{p}_{T_{b}}>\textrm{cv}_{\alpha}\right\} $
and $\textrm{cv}_{\alpha}$ satisfies $\sup_{\textrm{cv}_{\alpha}\in\mathbb{R}_{+}}\widehat{P}_{T_{b}}\left(T_{b}\in\mathbf{S}\left(\textrm{cv}_{\alpha}\right)\right)\geq1-\alpha$.
\end{defn}
The confidence set $C(\textrm{cv}_{\alpha})$ has a frequentist interpretation
even though the concept of HDR is often encountered in Bayesian analyses
since it associates naturally to the derived posterior distribution,
especially when the latter is multi-modal. A feature of the confidence
set $C(\textrm{cv}_{\alpha})$ under our context is that, at least
when the size of the shift is small, it consists of the union of several
disjoint intervals. The appeal of using HDR is that one can directly
deal with such features. As the break size increases and the distribution
becomes unimodal, the HDR becomes equivalent to the standard way of
constructing confidence sets. In practice, one can proceed as follows.
\begin{lyxalgorithm}
\textbf{$\mathrm{\mathbf{Confidence\,sets\,for\,}}T_{b}^{0}\mathbf{:}$}\label{algorithmn Confidence-sets-for}
1) Estimate by least-squares the break point and the regression coefficients
from model \eqref{Model (4), scalar format, CT}; 2) Replace quantities
appearing in \eqref{Equation (2) Asymptotic Distribution} by consistent
estimators as explained in Section \ref{Section Approximation-to-the};
3) Simulate the limiting distribution $\widehat{P}_{T_{b}}^{\mathrm{}}$
from Theorem \ref{Theorem 2, Asymptotic Distribution immediate Stationary Regimes};
4) Compute the HDR of the empirical distribution $\widehat{P}_{T_{b}}^{\mathrm{}}$
and include the point $T_{b}$ in the level $1-\alpha$ confidence
set $C\left(\textrm{cv}_{\alpha}\right)$ if $T_{b}$ satisfies the
conditions in Definition \ref{Def. Confidence Sets for Break Date}.
\textcolor{white}{\uline{d}}
\end{lyxalgorithm}
This procedure will not deliver contiguous confidence sets when the
size of the break is small. Indeed, we find that in such cases, the
overall confidence set for $T_{b}^{0}$ consists in general of the
union of disjoint intervals if $\widehat{T}_{b}$ is not near the
tails of the sample. One is located around the estimate of the break
date, while the others are in the pre- and post-break regimes. To
provide an illustration, we consider a simple example involving a
single draw from a simulation experiment. Figure \ref{Fig30} reports
the HDR of the feasible limiting distribution of $\rho(\widehat{T}_{b,\pi}-T_{b}^{0})$
for a random draw from the model in \eqref{Casin i-  Model for FS vs Asy Dist}
with parameters $\nu^{0}=1$, $\beta^{0}=0$, unit variance and autoregressive
coefficient 0.6 for $Z_{t}$ and $\sigma_{e}^{2}=1.2$. We set $\lambda_{0}=0.35,\,0.5$
and $\delta_{Z}^{0}=0.3,\,0.8,\,1.5$. We use a trimming 0.15 for
the plug-in estimator $\widehat{T}_{b}$ and $\pi=0.05$ for $\widehat{T}_{b,\pi}$.
As explained in Section \ref{subsec:Evaluation-of-the}, we could
use any other trimming in place of 0.15. The results remain unchanged.
We set $T=100$ and the significance level is $\alpha=0.05$. Note
that the origin is at the estimated break date. The point on the horizontal
axis corresponds to the true break date. The black intervals on the
horizontal axis correspond to regions of high density. The resulting
confidence set is their union. Once a confidence region for $\rho(\widehat{T}_{b,\pi}-T_{b}^{0})$
is computed, it is straightforward to derive a 95\% confidence set
for $T_{b}^{0}.$ The top panel (left plot) reports results for the
case $\delta_{Z}^{0}=0.3$ and $\lambda_{0}=0.35$ and shows that
the HDR is composed of two disjoint intervals. The estimated break
date is $\widehat{T}_{b}=70$ and the implied 95\% confidence set
for $T_{b}^{0}$ is given by $C(\textrm{cv}_{0.05})=\left\{ 1,\ldots,12\right\} \cup\left\{ 18,\ldots100\right\} $.
This includes  $T_{b}^{0}$ and the overall length is 95 observations.
Table \ref{Table 1 HDR} reports for various methods whether $T_{b}^{0}$
is covered or not and the length of the confidence sets for this example.
The length of \citeauthor{bai:97RES}\textquoteright s (1997) confidence
interval is 55 but does not include $T_{b}^{0}$. \citeauthor{elliott/mueller:07}\textquoteright s
(2007) confidence set, denoted by $\widehat{U}_{T}.\textrm{eq}$ in
Table \ref{Table 1 HDR}, also does not include the true break date
at the 90\% confidence level, but does so at the 95\% and its length
is 95. Our method covers $T_{b}^{0}$ and has a relatively short length
across different $\delta_{Z}^{0}$.


\section{Small-Sample Properties of the HDR Confidence Sets\label{Section Small-Sample-Effectiveness-of}}

We now assess via simulations the finite-sample performance of the
method proposed to construct confidence sets for the break date. We
also make comparisons with alternative methods in the literature:
\citeauthor{bai:97RES}'s \citeyearpar{bai:97RES} approach based
on the large-$N$ shrinkage asymptotics; \citeauthor{elliott/mueller:07}\textquoteright s
\citeyearpar{elliott/mueller:07}, hereafter EM, method on\textcolor{red}{{}
}inverting \citeauthor{nyblom:89}'s \citeyearpar{nyblom:89} statistic;\textcolor{red}{{}
}the Inverted Likelihood Ratio (ILR) approach of \citet{eo/morley:15}.
We omit the technical details of these methods and refer to the original
sources or \citet{chang/perron:18} for a review and comparisons.
We consider two DGPs: M1 is $y_{t}=\beta^{0}+\delta_{Z}^{0}\boldsymbol{1}_{\left\{ t>T_{b}^{0}\right\} }+e_{t}$
with $\beta^{0}=1$ and $e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right)$;
M2 is $y_{t}=\delta_{Z}^{0}\left(1-\nu^{0}\right)\boldsymbol{1}_{\left\{ t>T_{b}^{0}\right\} }+\nu^{0}y_{t-1}+e_{t}$
with $\nu^{0}=0.8$ and $e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,0.04\right)$.
Our companion paper \citet{casini/perron_CR_Single_Break_Extended}
includes extensive simulation results. We set the significance level
at $\alpha=0.05$, and the break occurs at date $\left\lfloor T\lambda_{0}\right\rfloor $,
where $\lambda_{0}=0.2,\,0.35,\,0.5$ and $T=200$ for M1 and $T=100$
for M2. The results are presented in Table \ref{Table M1}-\ref{Table M7}.
The last row in each table includes the rejection probability of a
5\%-level sup-Wald test using the asymptotic critical value in \citet{andrews:93},
which provides a measure of the magnitude of the break relative to
the noise. For models with predictable processes we use the two-step
procedure described in Section \ref{subsection Asymptotic-Results- Pre}.

Overall, the simulation results confirm previous findings about the
performance of existing methods. \citeauthor{bai:97RES}\textquoteright s
\citeyearpar{bai:97RES} method has a coverage rate below the nominal
level when the size of the break is small. Overall, our HDR method
and that of EM show accurate empirical coverage rates for all DGP
considered. However, EM\textquoteright s method almost always displays
confidence sets which are larger than those from the other approaches.
Over all DGPs considered, the average length of the HDR confidence
sets are 40\% to 70\% shorter than those obtained with EM\textquoteright s
approach when the size of the shift is moderate to high. The results
for M2, a change in mean with a lagged dependent variable and strong
correlation, are quite revealing. EM\textquoteright s method yields
confidence intervals that are very wide, increasing with the size
of the break and for large breaks covering nearly the entire sample.
This does not occur with the other methods. For instance, when $\lambda_{0}=0.5$
and $\delta_{Z}^{0}=2$, the average length from the HDR method is
8.34 compared to 93.71 with EM\textquoteright s. This concurs with
the results in \citet{chang/perron:18}.

In summary, the small-sample simulation results suggest that our continuous
record HDR-based inference provides accurate coverage probabilities
close to the nominal level and average lengths of the confidence sets
shorter relative to existing methods. It is also valid and reliable
under a wider range of DGPs including long-memory processes. Specifically
noteworthy is the fact that it performs well for all break sizes,
whether small or large.

\section{\label{Section Conclusions}Conclusions}

We examined a change-point model under a continuous record asymptotics.
With the time horizon $\left[0,\,N\right]$ fixed, we can account
for the asymmetric informational content provided by the pre- and
post-break samples. We derived a feasible counterpart of the continuous
record asymptotic distribution of the change-point estimator using
consistent plug-in estimates and showed that it provides accurate
approximations to the finite-sample distributions.  We used our limit
theory to construct confidence sets for the change-point date based
on the concept of Highest Density Region. Overall, it delivers accurate
coverage probabilities and relatively short average lengths of the
confidence sets. Importantly, it does so irrespective of the magnitude
of the break, whether large or small, a notoriously difficult problem
in the literature.

\newpage{}

\bibliographystyle{elsarticle-harv}
\bibliography{References_JoE}
\addcontentsline{toc}{section}{References}

\clearpage{}

\newpage{}

\clearpage
\pagenumbering{arabic}