Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
87,840 characters · 10 sections · 109 citation commands
Continuous Record Asymptotics for Change-Point Models
\setcounter{page}{0}
{\bf{JEL Classification}}: C10, C12, C22\\ {\bf{Keywords}}: Asymptotic distribution, break date, change-point, highest density region, semimartingale.
\onehalfspacing \thispagestyle{empty} \allowdisplaybreaks
In the context of a linear regression model with a single break point, we develop a continuous record asymptotic framework and inference methods for the break date. Our model is specified in continuous time but estimated with discrete-time observations using a least-squares method. We have $T$ observations with a sampling frequency $h$ over a fixed time horizon $\left[0,\,N\right],$ where $N=Th$ denotes the time span of the data. We consider a continuous record asymptotic framework whereby $T$ increases by shrinking the time interval $h$ to zero while keeping time span $N$ fixed. We impose very mild conditions on an underlying continuous-time model assumed to generate the data, basically continuous It� semimartingales.
An extensive amount of research addressed change-point problems under the classical large-$N$ asymptotics. Early contributions are hinkley:71, bhattacharya:87, and yao:87, who adopted a Maximum Likelihood (ML) approach, and for linear regression models, bai:97RES and bai/perron:98. See the reviews of csorgo/horvath:97, aue/Horvath:13, casini/perron_Oxford_Survey and references therein. In this literature, the resulting large-$N$ limit theory for the estimate of the break date depends on the exact distributions of the regressors and disturbances. Therefore, a so-called shrinkage asymptotic theory was adopted whereby the magnitude of the shift, say $\delta_{T},$ converges to zero which leads to an invariant limit distribution.
We study a general change-point problem under a continuous record asymptotic framework and develop inference procedures based on the derived asymptotic distribution. We establish consistency at rate-$T$ convergence for the least-squares estimate of the break date, assumed to occur at time $N_{b}^{0}$. Given the fast rate of convergence, we introduce a limit theory with shrinking magnitudes of shifts and increasing variance of the residual process local to the change-point. The asymptotic distribution corresponds to the location of the extremum of a function of the (quadratic) variation of the regressors and of a Gaussian centered martingale process over some time interval. It is characterized by some notable aspects. With the time horizon $\left[0,\,N\right]$ fixed, we can account for the asymmetric informational content provided by the pre- and post-break observations, i.e., the time span and the position of the break date $N_{b}^{0}$ convey useful information about the finite-sample distribution. In contrast, this is not achievable under the large-$N$ shrinkage asymptotic framework because both pre- and post-break segments expand proportionately as $T$ increases and, given the mixing assumptions imposed, only the neighborhood around the break date remains relevant. Further, the domain of the extremum depends on the position of the break $N_{b}^{0}$ relative to $N$, or total span, and thus the distribution is asymmetric, in general. The degree of asymmetry increases as the true break point moves away from mid-sample. This holds unless the magnitude of the break is large, in which case the density is symmetric irrespective of the location of the break. This accords with simulation evidence which documents that the break point estimate is less precise and the coverage rates of the confidence intervals less reliable when the break is not at mid-sample {[}see, e.g., chang/perron:18{]}. When the shift magnitude is small, the density displays three modes. As the shift magnitude increases, this tri-modality vanishes. We show empirically that all of these features are shared by the finite-sample distribution of the least-squares estimator of the break date. Hence, the continuous record asymptotics theory provides an accurate approximation to the finite-sample distribution of the break date estimator. In contrast, we show that the large-$N$ shrinkage asymptotic distribution of yao:87 and bai:97RES provides a poor approximation to the finite-sample distribution and does not share any of those features as we discuss.
Our asymptotics can be seen as intermediate between the shrinkage asymptotics and more recent approaches relying on weak identification {[}see e.g., elliott/mueller:07 and elliott/mueller/watson:15{]}. On the one hand, using the usual shrinking condition of yao:87 and bai:97RES for which the break magnitude, say $\delta_{T},$ goes to zero at a rate slower than $O(T^{-1/2})$ leads to underestimation of the uncertainty about the break date. On the other hand, the weak identification condition of elliott/mueller:07 for which $\delta_{T}$ goes to zero at a fast rate (i.e., $\delta_{T}=O(T^{-1/2})$, so that the change-point cannot be consistently estimated) leads to overstating the uncertainty. This has opposite consequences for the confidence intervals of the break date. Confidence sets have poor coverage probabilities when the break is small under bai:97RES's framework while they can be too wide under that of elliott/mueller:07. In this paper, the key is not to focus our asymptotic experiment on shrinking condition on $\delta_{T}$ but to make assumptions on the signal-to-noise ratio $\delta_{T}/\sigma_{t}$ instead, where $\sigma_{t}$ is the volatility of the errors. We require $\delta_{T}$ to go to zero at a slower rate than that of elliott/mueller:07---to guarantee strong identification---and require $\sigma_{t}$ to increase without bound when $t$ approaches the break date $T_{b}^{0}$. This offers a new characterization of the uncertainty without compromising strong identification and consistency of the model parameters needed to conduct inference.
Despite the effort devoted to the construction of confidence intervals for the break date {[}see e.g., bai/perron:98, elliott/mueller:07 and eo/morley:15{]}, what is still missing is a method that, for both large and small breaks, achieves both accurate coverage rates and satisfactory average lengths. The most popular method is bai/perron:98 which yields confidence intervals that are relatively short but have good coverage only when the magnitude of the break is not small. However, both small and large breaks are relevant for empirical work; breaks that are statistically small can still be practically relevant.
Given the peculiar properties (e.g., multi-modality and asymmetry) of the continuous record asymptotic distribution, we propose a non-standard inference related to Bayesian analyses. We use the concept of Highest Density Region to construct confidence sets for the break date. Our method has good coverage and length across all break magnitudes. This has important implications for empirical work because the user can be confident that our confidence interval includes the true value across all break sizes. For small breaks, the length of the confidence intervals from any method can be quite large for some models. However, our confidence interval is still informative because it reveals that there is high uncertainty about the change-point. The same information cannot be provided by existing methods either because they do not have good coverage unless the break is not small {[}e.g., bai/perron:98{]} or because they have a large length even when the break is not small {[}e.g., elliott/mueller:07{]}.
We use the continuous record asymptotics to provide an alternative approximation to the finite-sample distribution of the least-squares estimator based on discrete-time data. This creates no contradiction since asymptotic theory is intended as a thought experiment used to obtain approximations to the distribution of estimators or test statistics. The continuous record asymptotics has proven to be useful in other discrete-time settings such as in the context of unit roots {[}cf. phillips:87a and perron:1991ecma{]} and nonparametric regression {[}cf. brown/low:1996{]}. brown/low:1996 showed the asymptotic equivalence between a nonparametric regression problem and a white noise with drift problem. Our results show the asymptotic equivalence between discrete-time and continuous-time regression models with a change-point.
Recent work in change-point analysis has focused on estimation when the number of change-points is allowed to increase with the sample size {[}e.g., fryzlewicz:14{]} and when the change-point is allowed to approach the start and end sample point. A growing literature has also considered change-points in a high-dimensional setting {[}e.g., lee/seo/shin:16, leonardi/buhlmann:16, wang/lin/willett:19 and wang/yu/rinaldo/willett:19{]}. This work is mainly concerned with consistent estimation of the change-point dates and development of corresponding computational algorithms. Our focus is on asymptotic theory and inference within the classical change-point model with a single break. Our results can also have useful implications for the growing literature on inference in high-dimensional change-point analysis and for the literature on threshold regression {[}see, e.g., hansen:00ecma and hildago/lee/seo:19{]}.
This paper relates to other work by the authors, namely casini/perron_SC_BP_Lap (casini/perron_Lap_CR_Single_Inf; casini/perron_SC_BP_Lap)\nocite{casini/perron_Lap_CR_Single_Inf}. casini/perron_Lap_CR_Single_Inf used the asymptotic results developed in this paper and proposed a new Generalized Laplace estimator of the break date under a continuous record asymptotic framework. casini/perron_SC_BP_Lap analyzed the Generalized Laplace method under classical asymptotics and focused on the theoretical relationship between the asymptotic distribution of frequentist and Bayesian estimators of the break point. Finally, chambers/taylor:19 considered both deterministic one-time and continuous stochastic parameter change in a continuous-time autoregressive model while casini_CR_Test_Inst_Forecast introduced continuous-time asymptotics to test for forecast failure. Recently, casini:change-point-spectra considered testing and estimating change-points in a locally stationary process using frequency-domain methods.
The paper is organized as follows. Section (ref) introduces the model and the estimation method. Section (ref) contains results about the consistency and rate of convergence for fixed shifts. Section (ref) develops the asymptotic theory. We compare our limit theory with the finite-sample distribution in Section (ref). Section (ref) describes how to construct the confidence sets, with simulation results reported in Section (ref). Section (ref) provides brief concluding remarks. The Supplement {[}casini/perron_CR_Single_Break_Supp{]} contains the proofs as well as additional material.
We denote the transpose of a matrix $A$ by $A'$ and the $\left(i,\,j\right)$ elements of $A$ by $A^{\left(i,j\right)}$. We use $\left\Vert \cdot\right\Vert $ to denote the Euclidean norm of a linear space, i.e., $\left\Vert x\right\Vert =\left(\sum_{i=1}^{p}x_{i}^{2}\right)^{1/2}$ for $x\in\mathbb{R}^{p}.$ We use $\left\lfloor \cdot\right\rfloor $ to denote the largest smaller integer function. A sequence $\left\{ u_{kh}\right\} _{k=1}^{T}$ is $i.i.d.$ (resp., $i.n.d$) if the $u_{kh}$ are independent and identically (resp., non-identically) distributed. We use $\overset{P}{\rightarrow},\,\Rightarrow,$ and $\overset{\mathcal{L}-\mathrm{s}}{\Rightarrow}$ to denote convergence in probability, weak convergence and stable convergence in law, respectively. For semimartingales $\left\{ S_{t}\right\} _{t\geq0}$ and $\left\{ R_{t}\right\} _{t\geq0}$, we denote their covariation process by $\left[S,\,R\right]_{t}$ and their predictable counterpart by $\left\langle S,\,R\right\rangle _{t}$. The symbol “$\triangleq$” denotes definitional equivalence.
Consider a change-point model with a single break point:
where $Y_{t}$ is the dependent variable, $D_{t}$ and $Z_{t}$ are, respectively, $q\times1$ and $p\times1$ vectors of regressors and $e_{t}$ is an unobservable disturbance. The vector-valued parameters $\nu^{0},\,\delta_{Z,1}^{0}$ and $\delta_{Z,2}^{0}$ are unknown with $\delta_{Z,1}^{0}\neq\delta_{Z,2}^{0}$. Our main purpose is to develop inference methods for the unknown change-point date $T_{b}^{0}$ when $T+1$ observations on $\left(Y_{t},\,D_{t},\,Z_{t}\right)$ are available. Before moving to the re-parametrization of the model, we discuss the underlying continuous-time model assumed to generate the data. The discrete-time variables are assumed to be generated from the continuous-time processes $\{D_{s},\,Z_{s},\,e_{s}\}_{s\geq0}$ defined on a filtered probability space $(\Omega,\,\mathscr{F},\,(\mathscr{F}_{s})_{s\geq0},\,P).$ We observe realizations of $\{Y_{s},\,D_{s},\,Z_{s}\}$ at discrete points of time.
The sampling occurs at regularly spaced time intervals of length $h$ within a fixed time horizon $\left[0,\,N\right]$ where $N$ denotes the span of the data. We observe $\left\{ _{h}Y_{kh},\,{}_{h}D_{kh},\,{}_{h}Z_{kh};\,k=0,\,1,\ldots,\,T=N/h\right\} $. $_{h}D_{kh}\in\mathbb{R}^{q}$ and $_{h}Z_{kh}\in\mathbb{R}^{p}$ are random vector step functions which jump only at times $0,\,h,\ldots,\,Th$. We shall allow $_{h}D_{kh}$ and $_{h}Z_{kh}$ to include both predictable processes and locally-integrable semimartingles, though the case with predictable regressors is more delicate and discussed in the supplement. The discretized processes $_{h}D_{kh}$ and $_{h}Z_{kh}$ are assumed to be adapted to the increasing and right-continuous filtration $\left\{ \mathscr{F}_{t}\right\} _{t\geq0}$. For any process $X$ we denote its “increments” by $\Delta_{h}X_{k}=X_{kh}-X_{\left(k-1\right)h}$. For $k=1,\ldots,T$, let $\Delta_{h}D_{k}\triangleq\mu_{D,k}h+\Delta_{h}M_{D,k}$ and $\Delta_{h}Z_{k}\triangleq\mu_{Z,k}h+\Delta_{h}M_{Z,k}$ where the “drifts” $\mu_{D,t}\in\mathbb{R}^{q},\,\mu_{Z,t}\in\mathbb{R}^{p}$ are $\mathscr{F}_{t-h}$-measurable (exact assumptions will be given below), and $M_{D,k}\in\mathbb{R}^{q},\,M_{Z,k}\in\mathbb{R}^{p}$ are continuous local martingales with finite conditional covariance matrix $P$-a.s., $\mathbb{E}(\Delta_{h}M_{D,t}\Delta_{h}M_{D,t}'|\,\mathscr{F}_{t-h})=\Sigma_{D,t-h}\Delta t$ and $\mathbb{E}(\Delta_{h}M_{Z,t}\Delta_{h}M'_{Z,t}|\,\mathscr{F}_{t-h})=\Sigma_{Z,t-h}\Delta t$ ($\Delta t$ and $h$ are used interchangeably). Let $\lambda_{0}\in\left(0,\,1\right)$ denote the fractional break date (i.e., $T_{b}^{0}=\left\lfloor T\lambda_{0}\right\rfloor $). Via the Doob-Meyer Decomposition, model (ref) can be expressed as
where the error process $\left\{ \Delta_{h}e_{t}^{*},\,\mathscr{F}_{t}\right\} $ is a continuous local martingale difference sequence with conditional variance $\mathbb{E}[\left(\Delta_{h}e_{t}^{*}\right)^{2}|\,\mathscr{F}_{t-h}]=\sigma_{e,t-h}^{2}\Delta t$ $P$-a.s. finite. The underlying continuous-time data-generating process can thus be represented (up to $P$-null sets) in integral equation form as
where $\sigma_{D,t}$ and $\sigma_{Z,t}$ are the instantaneous covariance processes taking values in $\mathcal{M}_{q}^{\textrm{c�dl�g}}$ and $\mathcal{M}_{p}^{\textrm{c�dl�g}}$ {[}the space of $p\times p$ positive definite real-valued matrices whose elements are c�dl�g{]}; $W_{D}$ (resp., $W_{Z}$) is a $q$ (resp., $p$)-dimensional standard Wiener process; $e^{*}=\left\{ e_{t}^{*}\right\} _{t\geq0}$ is a continuous local martingale which is orthogonal (in a martingale sense) to $\left\{ D_{t}\right\} _{t\geq0}$ and $\left\{ Z_{t}\right\} _{t\geq0}$; and $D_{0}$ and $Z_{0}$ are $\mathscr{F}_{0}$-measurable random vectors. In (ref), $\int_{0}^{t}\mu_{D,s}ds$ is a continuous adapted process with finite variation paths and $\int_{0}^{t}\sigma_{D,s}dW_{D,s}$ corresponds to a continuous local martingale.
Part (i) restricts the processes to be locally bounded and part (ii) requires the drifts to be adapted finite variation processes. These are standard regularity conditions in the high-frequency statistics literature {[}cf. barndorff/shephard:04{]}. Part (iii) requires the regressors to have finite integrated covariance. We rule out jump processes; hence, our results are not expected to provide good approximations for high-frequency data but for data sampled at lower frequencies. In principle, ultra high frequency data would be used which, however, are essentially available only for financial variables, which involve a host of issues that we cannot handle (e.g., market-microstructure, bid-ask spread, volatility jumps, non-continuous sampling, etc.).
An interesting issue is whether the theoretical results to be derived for model (ref) are applicable to classical structural change models for which an increasing span of data is assumed. This requires establishing a connection between the assumptions imposed on the stochastic processes in both settings. Roughly, the classical long-span setting uses approximation results valid for weakly dependent data; e.g., ergodic and mixing procesess. Such assumptions are not needed under our fixed-span asymptotics. Nonetheless, we can impose restrictions on the probabilistic properties of the latent volatility processes in our model and thereby guarantee that ergodic and mixing properties are inherited by the corresponding observed processes. This follows from Theorem 3.1 in \citet*{genon-catalot/jeantheau/laredo:00} together with Proposition 4 in carrasco/chen:02. For example, these results imply that the observations $\left\{ Z_{kh}\right\} _{k\geq1}$ (with fixed $h$) can be viewed (under certain conditions) as a hidden Markov model which inherits the ergodic and mixing properties of $\left\{ \sigma_{Z,t}\right\} _{t\geq0}$. Hence, our model encompasses those considered in the structural change literature that uses a long-span asymptotic setting. We shall extend model (ref) to allow for predictable processes (e.g., a constant and/or lagged dependent variable) in the supplement.
It is useful to re-parametrize model (ref). Let $y_{kh}=\Delta_{h}Y_{k},$ $x_{kh}=(\Delta_{h}D'_{k},\,\Delta_{h}Z'_{k})'$, $z_{kh}=\Delta_{h}Z_{k}$, $e_{kh}=\Delta_{h}e_{k}^{*},$ $\beta^{0}=((\pi^{0}),\,(\delta_{Z,1}^{0})')'$ and $\delta_{Z}^{0}=\delta_{Z,2}^{0}-\delta_{Z,1}^{0}$. (ref) can be expressed as:
where the true parameter $\theta^{0}=((\beta^{0})',\,(\delta_{Z}^{0})')'$ takes value in a compact space $\Theta\subset\mathbb{R}^{\textrm{dim}\left(\theta\right)}$. Also, define $z_{kh}=R'x_{kh}$, where $R$ is a $\left(q+p\right)\times p$ known matrix with full column rank. We consider a partial structural change model for which $R=\left(0,\,I\right)'$ with $I$ an identity matrix.
Finally, we write the model in matrix format which will be useful for the derivations. Let $Y=(y_{h},\,\ldots,\,y_{Th})',\,X=(x_{h},\,\ldots,\,x_{Th})'$, $e=(e_{h},\,\ldots,\,e_{Th})',$ $X_{1}=(x_{h},\,\ldots,\,x_{T_{b}h},\,0,\,\ldots,\,0)'$, $X_{2}=(0,\,\ldots,\,0,\,x_{\left(T_{b}+1\right)h},\ldots,\,x_{Th})'$ and $X_{0}=(0,\,\ldots,\,0,\,x_{\left(T_{b}^{0}+1\right)h},\ldots,\,x_{Th})'$. Note that the difference between $X_{0}$ and $X_{2}$ is that the latter uses $T_{b}$ rather than $T_{b}^{0}$. Define $Z_{1}=X_{1}R,$ $Z_{2}=X_{2}R$ and $Z_{0}=XR$. (ref) in matrix format is: $Y=X\beta^{0}+Z_{0}\delta_{Z}^{0}+e$. We consider the least-squares estimator of $T_{b}$, i.e., the minimizer of $S_{T}\left(T_{b}\right)$, the sum of squared residuals when regressing $Y$ on $X$ and $Z_{2}$ over all possible partitions, namely: $\widehat{T}_{b}^{\textrm{LS}}=\mathrm{argmin}_{p+q\leq T_{b}\leq T}S_{T}\left(T_{b}\right)$. It is straightforward to show that $\widehat{T}_{b}^{\textrm{LS}}=\mathrm{\mathrm{argmin}}_{p+q\leq T_{b}\leq T}Q_{T}\left(T_{b}\right)$ where $Q_{T}\left(T_{b}\right)\triangleq\widehat{\delta}'_{T_{b}}\left(Z_{2}'MZ_{2}\right)\widehat{\delta}_{T_{b}}$, $\widehat{\delta}_{T_{b}}$ is the least-squares estimator of $\delta_{Z}^{0}$ when regressing $Y$ on $X$ and $Z_{2}$, and $M=I-X\left(X'X\right)^{-1}X'$. For brevity, we will write $\widehat{T}_{b}^{\textrm{}}$ for $\widehat{T}_{b}^{\textrm{LS}}$ with the understanding that $\widehat{T}_{b}$ is a sequence indexed by $T$. Let $\widehat{\delta}=\widehat{\delta}{}_{\widehat{T}_{b}}$. The estimate of the break fraction is then $\widehat{\lambda}_{b}=\widehat{T}_{b}/T$. Both in practice and for theoretical analyses, a trimming parameter $\pi\in\left(0,\,1/2\right)$ is applied to restrict the minimization over the interval $\left[T\pi,\,\left(1-\pi\right)T\right]$.
We now establish the consistency and convergence rate of the least-squares estimator under fixed shifts. Under the classical large-$N$ asymptotics, related results have been established by bai:97RES and bai/perron:98. Early important results for a mean-shift appeared in yao:87 and bhattacharya:87 for an $\mathit{i.i.d.}$ series, bai:94a for linear processes and picard:85 for a Gaussian autoregressive model.
Assumption (ref) is similar to A2 in bai/perron:98 and requires enough variation around the break point and at the beginning and end of the sample. The factor $h^{-1}$ normalizes the observations so that the assumption is implied by a weak law of large numbers. Assumption (ref) is a standard uniqueness identification condition. We then have the following results.
We have the same $T$-convergence rate as under large-$N$ asymptotics. Let $\theta^{0}=((\beta^{0})',\,(\delta_{1}^{0})',$ $(\delta_{2}^{0})')'$. The fast $T$-rate of convergence implies that the least-squares estimate of $\theta^{0}$ is the same as when $\lambda_{0}$ is known. A natural estimator for $\theta^{0}$ is $\mathrm{\mathrm{\mathrm{argmin}}_{\beta\in\mathbb{R}^{\mathit{p+q}},\delta\in\mathbb{R}^{\mathit{p}}}}||Y-X\beta-\widehat{Z}_{2}\delta||^{2}$, where we use $\widehat{T}_{b}$ instead of $T_{b}$ in the construction of $\widehat{Z}_{2}$. Then we have the following result, akin to an extension of corresponding results in Section 3 of barndorff/shephard:04. As a matter of notation, let $\Sigma^{*}\triangleq\left\{ \mu_{\cdot,t},\,\Sigma_{\cdot,t},\,\sigma_{e,t}\right\} _{t\geq0}$ and denote expectation taken with respect to $\Sigma^{*}$ by $\mathbb{E}^{*}$.
$V$ can be random because $\Sigma_{\cdot,t}$ and $\sigma_{e,t}$ can be stochastic. Under fixed shifts, Proposition (ref)-(ref) shows the asymptotic equivalence of discrete and continuous-time regression models with a change-point, a result corresponding to brown/low:1996 for nonparametric regression.
We now present results about the limiting distribution of the least-squares estimate of the break date under a continuous record framework. As in the classical large-$N$ asymptotics, it depends on the exact distribution of the data and the errors for fixed break sizes {[}c.f., hinkley:71{]}. This has forced researchers to consider a shrinkage asymptotic theory where the size of the shift is made local to zero as $T$ increases, an approach developed by picard:85 and yao:87. We continue with this avenue. Given the consistency result, we know that there exists some $h^{*}$ such that for all $h<h^{*}$ with high probability $\eta Th\leq\widehat{N}_{b}\leq\left(1-\eta\right)Th$, for $\eta>0$ such that $\lambda_{0}\in\left(\eta,\,1-\eta\right)$. By Proposition (ref), $\widehat{N}_{b}-N_{b}^{0}=O_{p}\left(T^{-1}\right)$, i.e., $\widehat{N}_{b}$ is in a shrinking neighborhood of $N_{b}^{0}$. With a certain rescaling of the objective function one can first obtain the shrinkage asymptotic distribution of bai:97RES. However, this is unsatisfactory for two reasons. First, as we show below {[}see also Casini and Perron (casini/perron_SC_BP_Lap; casini/perron_Lap_CR_Single_Inf){]}, the shrinkage asymptotic distribution provides a poor approximation to the finite-sample distribution of the least-squares estimator. Second, the latter point also explains the poor coverage properties of the confidence intervals derived from the shrinkage asymptotic distribution when the magnitude of the break is not large. Some related results were obtained by Jiang et al. jiang/wang/yu:16 for a simple location model. Their approach is, however, quite restrictive and no feasible inference procedure suggested. See the supplement for a more complete discussion.
We begin with the following assumption which specifies that i) we use a shrinking condition on $\delta_{Z}^{0}$; ii) we introduce a locally increasing variance condition on the residual process. The first is similarly used under classical large-$N$ asymptotics, while the second is new and useful in our context in order to accurately capture the relevant uncertainty in the change-point problem. We do not impose restrictions only on $\delta_{Z}^{0}$ but also on the ratio $\delta_{Z}^{0}/\sigma_{t}$ when $t$ is close to $T_{b}^{0}$. We refer to $\delta_{Z}^{0}/\sigma_{t}$ as the signal-to-noise ratio. Controlling this ratio rather than just $\delta_{Z}^{0}$ allows for an alternative characterization of the uncertainty about the change-point date in order to obtain an asymptotic distribution which provides a better approximation of the finite-sample distribution of the estimator. To emphasize that $\delta_{Z}^{0}$ depends on the sample-size we denote it by $\delta_{h}$.
Note that the localization parameter $\delta^{0}$ in the definition of $\delta_{h}$ is different from the fixed parameter $\delta_{Z}^{0}$ since $h\rightarrow0$. The rate $1/4$ in the conditions $\delta_{h}=O(h^{1/4})$ and $\sigma_{h}=O(h^{-1/4})$ is for tractability. One can show that consistency also holds for a rate faster than $1/4$, though slower than $\kappa.$ However, for the derivation of the limiting distribution one needs $\delta_{h}/\sigma_{h}=O(h^{1/2})$ and $O(\delta_{h})=O(\sigma_{h}^{-1})$ with $\kappa<1/2.$ The vector of scaled true parameters is $\theta{}_{h}\triangleq((\beta^{0})',\,\delta'_{h})'$. Define
We shall refer to $\left\{ \Delta_{h}\widetilde{e}_{t},\,\mathscr{F}_{t}\right\} $ as the normalized residual process. Under this framework, the rate of convergence of $\widehat{N}_{b}$ is now $T^{1-\kappa}$ with $0<\kappa<1/2$. Due to the fast rate of convergence of the change-point estimator, the objective function oscillates too rapidly as $h\downarrow0$. By scaling up the volatility of the errors around the change-point, we make the objective function behave as if it were a function of a standard diffusion process. The neighborhood in which the errors have relatively higher variance is shrinking at a rate $1/T^{1-\kappa}$, the rate of convergence of $\widehat{N}_{b}.$ Hence, in a neighborhood of $N_{b}^{0}$ in which we study the limiting behavior of the break point estimator, the rescaled criterion function is regular enough so that a feasible limit theory can be developed. The rate of convergence $T^{1-\kappa}$ is still sufficiently fast to guarantee a $\sqrt{T}$-consistent estimation of the slope parameters, as stated in the following proposition. Let $\left\langle Z_{\Delta},\,Z_{\Delta}\right\rangle \left(v\right)$ be the predictable quadratic variation process of $Z_{\Delta}$. The process $\mathscr{W}\left(v\right)$ is, conditionally on $\mathscr{F}$, a two-sided centered Gaussian martingale with independent increments and variances given in Section (ref) of the supplement.
We first present a general result which shows that under Assumption (ref) one can obtain a shrinkage asymptotic distribution similar to bai:97RES. The latter exploits the consistency of $\widehat{\lambda}_{b}$ and the fact that mixing conditions implies that the regimes before and after $\lambda_{0}$ are asymptotically independent. Let $Z_{\Delta}\triangleq(0,\ldots,\,0,\,z_{\left(T_{b}+1\right)h},\ldots,\,z_{T_{b}^{0}h},\,0,\ldots,\,0)$ if $T_{b}<T_{b}^{0}$ and $Z_{\Delta}\triangleq(0,\ldots,\,0,\,z_{\left(T_{b}^{0}+1\right)h},\ldots,\,z_{T_{b}h},$ $0,\ldots,\,0)$ if $T_{b}>T_{b}^{0}$.
The distribution in Proposition (ref) is different from bai:97RES. One can show that his distribution can be obtained under a continuous record if Assumption (ref) is modified as follows: $\delta_{h}=\delta^{0}h^{\kappa/2}$, $T^{1-\kappa}\epsilon\rightarrow B<\infty$, $0<\kappa\leq1/2$ and $\sigma_{h}\triangleq\overline{\sigma}h^{-\kappa/2}$. This would result in,
The difference between (ref) and (ref) is the presence of the drift (or deterministic) part $-\left(\delta^{0}\right)'$ $\left\langle Z_{\Delta},\,Z_{\Delta}\right\rangle \left(v\right)\delta^{0}$. Without relating the magnitude of the break to the local variance condition, the order of the stochastic part dominates that of the deterministic part and so the latter vanishes asymptotically. The distributions in (ref)-(ref) share the same issues as Bai's and so they do not add any particular insight. We therefore move to discuss how to obtain a more useful continuous record asymptotic distribution.
Consider the set $\mathscr{\mathcal{D}}\left(C\right)\triangleq\left\{ N_{b}:\,N_{b}\in\left\{ N_{b}^{0}+Ch^{1-\kappa}\right\} ,\,\left|C\right|<\infty\right\} $, on the original time scale. Let $\psi_{h}\triangleq h^{1-k}$. Here we use the same device as in foster/nelson:96 (nelson/foster:94; foster/nelson:96)\nocite{nelson/foster:94}. Different scaling factors applied to an objective function can lead to different asymptotic distributions. We normalize $Q_{T}(T_{b})$ by $\psi_{h}$, where $\psi_{h}$ corresponds to the rate of convergence in Proposition (ref). The rate of convergence implicitly describes the order of the terms in the expansion of $Q_{T}\left(T_{b}\right)-Q_{T}\left(T_{b}^{0}\right)$.
For brevity, we use the notation $\pm$ in place of $\mathrm{sgn}\left(T_{b}^{0}-T_{b}\right)$, henceforth. The conditional first moment of the centered criterion function $Q_{T}\left(T_{b}\right)-Q_{T}\left(T_{b}^{0}\right)$ is of order $O\left(h^{1-\kappa}\right)$, i.e., it “oscillates” rapidly as $h\downarrow0$. Hence, in order to approximate the behavior of $\{\widehat{T}_{b}-T_{b}^{0}\}$, we proceed as in Section 3 in nelson/foster:94 and rescale “time”. For any $C>0$, let $L_{C}\triangleq N_{b}^{0}-Ch^{1-\kappa}$ and $R_{C}\triangleq N_{b}^{0}+Ch^{1-\kappa}$, where $L_{C}$ and $R_{C}$ are the left and right boundary points of $\mathcal{D}\left(C\right)$, respectively. We then have $\left|R_{C}-L_{C}\right|=O\left(Ch^{1-\kappa}\right)$. Now, take the vanishingly small interval $\left[L_{C},\,R_{C}\right]$ on the original time scale, and stretch it into a time interval $\left[T^{1-\kappa}L_{C},\,T^{1-\kappa}R_{C}\right]$ on a new “fast time scale”. Changing time scale simply means that we rescale the objective function in such a way that it is of higher order as $h\downarrow0$, i.e., it fluctuates less. This leads to an asymptotic distribution that accounts for higher uncertainty. Yet, under our framework it is still possible to consistently estimate the break fraction and the regression coefficients so that inference is feasible.
Since the criterion function is scaled by $\psi_{h}^{-1}$, all scaled processes are $O_{p}\left(1\right)$. Now, let $N_{b}\left(v\right)=N_{b}^{0}-vh^{1-\kappa},\,v\in\left[-C,\,C\right]$. Using Lemma (ref) and Assumption (ref) (see the appendix),
where $\widetilde{e}_{kh}\triangleq h^{1/4}e_{kh}.$ In addition, in view of (ref), we let $dZ_{\psi,s}=\psi_{h}^{-1/2}\sigma_{Z,s}dW_{Z,s}$ for $s\in\left[N_{b}^{0}-vh^{1-\kappa},\,N_{b}^{0}+vh^{1-\kappa}\right]$. Applying the time scale change $s\rightarrow t\triangleq\psi_{h}^{-1}s$ to all processes including $\Sigma^{0}$, we have $dZ_{\psi,t}=\sigma_{Z,t}dW_{Z,t}$ with $t\in\mathcal{T}\left(C\right)$, where $\mathcal{T}\left(C\right)\triangleq\{t:\,t\in\left[N_{b}^{0}+v\left\Vert \delta^{0}\right\Vert ^{2}/\overline{\sigma}^{2}\right],\,\left|v\right|\leq C\}$. Therefore,
with $NT_{b}\left(v\right)/T=N_{b}\left(v\right)=N_{b}^{0}+v$, where $z_{\psi,kh}\triangleq z_{kh}/\sqrt{\psi_{h}}$ and $\widetilde{e}_{\psi,kh}\triangleq\widetilde{e}_{kh}/\sqrt{\psi_{h}}$. Because of the change of time scale, all processes in the last display are scaled up to be $O_{p}\left(1\right)$ and thus behave as diffusion-like processes. On this new “fast time scale”, we have $T^{1-\kappa}R_{C}-T^{1-\kappa}L_{C}=O\left(1\right)$ and $Q_{T}\left(T_{b}\left(v\right)\right)-Q_{T}\left(T_{b}^{0}\right)$ is restored to be $O_{p}\left(1\right)$. Observe that changing the time scale does not affect any statistic which depends on observations from $k=1$ to $k=\left\lfloor L_{C}/h\right\rfloor $ or from $k=\left\lfloor R_{C}/h\right\rfloor $ to $k=T$ (since these involve a positive fraction of data). However, it does affect quantities which include observations that fall in $\left[T_{b}h,\,T_{b}^{0}h\right]$ (assuming $T_{b}<T_{b}^{0}$). In particular, on the original time scale, the processes $\left\{ D_{t}\right\} ,\,\left\{ Z_{t}\right\} $ and $\left\{ e_{t}\right\} $ are well-defined and scaled to be $O_{p}\left(1\right)$ while $Q_{T}\left(T_{b}\right)-Q_{T}\left(T_{b}^{0}\right)$ (asymptotically) oscillates more rapidly than a simple diffusion-type process. On the new “fast time scale”, $\left\{ D_{t}\right\} ,\,\left\{ Z_{t}\right\} $ and $\left\{ e_{t}\right\} $ are not affected since they have the same order in $\left[T^{1-\kappa}L_{C},\,T^{1-\kappa}R_{C}\right]$ as $h\downarrow0$. That is, the first conditional moments are $O\left(h\right)$ while the corresponding moments for $Q_{T}\left(T_{b}\right)-Q_{T}\left(T_{b}^{0}\right)$ on $\mathcal{T}\left(C\right)$ are restored to be $O\left(h\right)$. As the continuous-time limit is approached, the rescaled criterion function $\left(Q_{T}\left(T_{b}\left(v\right)\right)-Q_{T}\left(T_{b}^{0}\right)\right)/h^{1/2}$\textcolor{red}{ }operates on a “fast time scale” on $\mathcal{T}\left(C\right)$.
Our analysis is local; we examine the limiting behavior of the centered and rescaled criterion function process in a neighborhood $\mathcal{T}\left(C\right)$ of the the true break date $N_{b}^{0}$ defined on a new time scale. We first obtain the weak convergence results for the statistic $\left(Q_{T}\left(T_{b}\left(v\right)\right)-Q_{T}\left(T_{b}^{0}\right)\right)/h^{1/2}$ and then apply a continuous mapping theorem for the argmax functional. However, it is convenient to work with a re-parametrized objective function. Proposition (ref) allows us to use
where $\theta^{*}\triangleq\left(\theta'_{h},\,v\right)'$ with $T_{b}\left(v\right)\triangleq T_{b}^{0}+\left\lfloor v/h\right\rfloor $ and $T_{b}\left(v\right)$ is the time index on the “fast time scale”. The normalizing factor $\psi_{h}h^{1/2}$ allows us to change the time scale and obtain an alternative asymptotic distribution. When $v$ varies, $T_{b}\left(v\right)$ potentially visits all integers between $1$ and $T$. Thus, on the new time scale, we need to introduce the trimming parameter $\pi\in\left(0,\,1\right)$ which determines the region where $T_{b}\left(v\right)$ can vary. We have the normalizations $T_{b}\left(v\right)=T\pi$ if $T_{b}\left(v\right)\leq T\pi$ and $T_{b}\left(v\right)=T\left(1-\pi\right)$ if $T_{b}\left(v\right)\geq T\left(1-\pi\right)$. On the old time scale, $N_{b}\left(u\right)=N_{b}^{0}+u$ with $v\rightarrow\psi_{h}^{-1}u$, so that $N_{b}\left(u\right)$ is in a vanishing neighborhood of $N_{b}^{0}$. On $\mathcal{T}\left(C\right)$, we index the process $Q_{T}\left(\theta_{h},\,T_{b}\left(v\right)\right)-Q_{T}\left(\theta^{0},\,T_{b}^{0}\right)$ by two time subscripts: one referring to the time $T_{b}$ on the original time scale and one referring to the time elapsed since $T_{b}h$ on the “fast time scale”. For simplicity, we omit the former; since the limiting distribution of the least-squares estimator will now depend on the trimming we use the notation $\widehat{T}_{b,\pi}=T\widehat{\lambda}_{b,\pi}$ where $\widehat{\lambda}_{b,\pi}$ is the least-squares estimator of the fractional break date associated to the fast time scale (i.e., associated to the factor $\psi_{h}h^{1/2}$).
The optimization problem is not affected by the change of time scale. In fact, by Proposition (ref), $u=Th(\widehat{\lambda}_{b}-\lambda_{0})=KO_{p}\left(h^{1-\kappa}\right)$ on the old time scale; whereas on the new “fast time scale”, $v=Th(\widehat{\lambda}_{b,\pi}-\lambda_{0})=O_{p}\left(1\right)$. The maximization problem is not changed because $v/h$ can take any value in $\mathbb{R}$. The process $Q_{T}\left(\theta_{h},\,T_{b}\left(v\right)\right)-Q_{T}\left(\theta^{0},\,T_{b}^{0}\right)$ is thus analyzed on a fixed horizon since $v$ now varies over $[(N\pi-N_{b}^{0})/(||\delta^{0}||^{-2}\overline{\sigma}^{2}),\,(N\left(1-\pi\right)-N_{b}^{0})/(\left\Vert \delta^{0}\right\Vert ^{-2}\overline{\sigma}^{2})]$. Define the modification to the set $\mathcal{D}\left(C\right)$ applicable to the new time scale by
Let $\mathbb{D}\left(\mathcal{D}^{*}\left(C\right),\,\mathbb{R}\right)$ denote the space of all c\`{a}dl\`{a}g functions from $\mathcal{D}^{*}\left(C\right)$ into $\mathbb{R}.$ Endow this space with the Skorokhod topology. Under a continuous record, we can apply limit theorems for statistics involving (co)variation between regressors and errors. This enables us to deduce the limiting process for $\overline{Q}_{T}\left(\theta^{*}\right)$, relying upon the work of Jacod jacod:94,jacod:97 and jacod/protter:98.
To guide intuition, note that under the new re-parametrization, the limit law of $\overline{Q}_{T}\left(\theta^{*}\right)$ is, according to Lemma (ref), the same as the limit law of
where $\overset{d}{\equiv}$ denotes (first order) equivalence in law, and since (approximately) $e_{kh}\sim i.n.d.\,\mathscr{N}(0,$ $\sigma_{h,k-1}^{2}h),\,\sigma_{h,k}=\sigma_{h}\sigma_{e,k}$ then $\widetilde{e}_{kh}\sim i.n.d.\,\mathscr{N}\left(0,\,\sigma_{e,k-1}^{2}h\right)$. Hence, the limit law of $\overline{Q}_{T}\left(\theta^{*}\right)$ is, to first-order, equivalent to the law of
We apply a law of large numbers to the first term and a stable convergence in law under the Skorokhod topology to the second. Assumption (ref) combined with the normalizing factor $h^{-1/2}$ in $\overline{Q}_{T}\left(\theta^{*}\right)$ account for the discrepancy between the deterministic and stochastic component in (ref).
Having outlined the main steps in the arguments used to derive the continuous records limit distribution of the break date estimate, we now state the main result of this section. The limiting process is realized on a extension of the original probability space and we relegate this description to Section (ref) in the supplement.
Note the differences between the results in Theorem (ref) and in Proposition (ref). First, on the fast time scale, $\widehat{\lambda}_{b,\pi}$ behaves as an inconsistent estimator for $\lambda_{0}$ for $N$ fixed, but it is consistent as $N\rightarrow\infty$. On the original time scale $\widehat{\lambda}_{b}$ is not only consistent for $\lambda_{0}$ but it also enjoys a similar asymptotic distribution as in bai:97RES. Second, the asymptotic distribution of $\widehat{\lambda}_{b,\pi}$ depends on the span of the data and consequently on the trimming $\pi.$ Proposition (ref), in contrast, suggests that the span, the trimming and the location of the break are irrelevant for the limiting behavior of the estimator. This intuitively follows from the fact that under the original time scale the break date estimator is consistent. We will show that indeed the span of the data and the location of the break influence the finite-sample properties of the least-squares estimator, and that Theorem (ref) provides a more useful approximation. An important implication of Theorem (ref) is that the precision of the estimator depends more on the span $N$ than to the number of observations $T$.
Unlike Bai's distribution, the distribution in Theorem (ref) involves the location of the maximum of a function of the (quadratic) variation of the regressors and of a two-sided centered Gaussian martingale process over the interval $[(N\pi-N_{b}^{0})/(\left\Vert \delta^{0}\right\Vert ^{-2}\overline{\sigma}^{2}),\,(N\left(1-\pi\right)-N_{b}^{0})/(||\delta^{0}||^{-2}\overline{\sigma}^{2})]$. Notably, this domain depends on the true value of $N_{b}^{0}$ and therefore the limit distribution is asymmetric, in general. The degree of asymmetry increases as the true break point moves away from mid-sample. This holds even when the distributions of the errors and regressors are the same in the pre- and post-break regimes. The presence of the trimming confirms that the span of the (trimmed) data affects the limit distribution. It is well-known that the least-squares estimator of the break date can be sensitive to trimming {[}see bai/perron:03 for some recommendations on the trimming choice{]}. Our asymptotic theory accommodates this property of the least-squares estimator while others do not.
Additional relevant remarks follow; more details are provided in the supplement. The magnitude of the break plays a key role in determining the density of the asymptotic distribution. More precisely, the density displays interesting properties which change when the signal-to-noise ratio as well as other parameters of the model change. Moreover, the distribution in Theorem (ref) is able to reproduce important features of the small-sample results obtained via simulations {[}e.g., bai/perron:06{]}. First, the second moments of the regressors impact the asymptotic mean as well as the second-order behavior of the break point estimator (e.g., the persistence of the regressors influences the finite-sample performance of the estimator). Second, the continuous record setting manages to preserve information about the time span $N$ of the data, a clear advantage since the location of the true break point matters for the small-sample distribution of the estimator. It has been shown via simulations that in small-samples the break point estimator tends to be imprecise if the break size is small, and some bias arises if the break point is not at mid-sample. In our framework, the (trimmed) time horizon $\left[N\pi,\,N\left(1-\pi\right)\right]$ is fixed and thus we can distinguish between the statistical content of the segments $\left[N\pi,\,N_{b}^{0}\right]$ and $\left[N_{b}^{0},\,N\left(1-\pi\right)\right]$. In contrast, this is not feasible under the classical shrinkage large-$N$ asymptotics because both the pre- and post-break segments increase proportionately and mixing conditions are imposed so that the only relevant information is a neighborhood around the true break date. Details on how to simulate the limiting distribution in Theorem (ref) are given in Section (ref) of the supplement.
We further characterize the asymptotic distribution by exploiting the ($\mathscr{F}$-conditionally) Gaussian property of the limit process. The analysis also holds unconditionally if we assume that the volatility processes are non-stochastic. Thus, as in the classical setting, we begin with a second-order stationarity assumption within each regime. The following assumption guarantees that the results below remain valid without the need to condition on $\mathscr{F}.$
Let $W_{i}^{*},$ $i=1,\,2,$ be two independent standard Wiener processes defined on $[0,\,\infty),$ starting at the origin when $s=0.$ Let
Unlike the asymptotic distribution derived under classical large-$N$ asymptotics, the probability density in (ref) is not available in closed form. Furthermore, the limiting distribution depends on unknown quantities. In the next section we explain how one can derive a feasible counterpart. This will be useful to characterize the main features of interest that will guide us in devising methods to construct confidence sets for $T_{b}^{0}$.
In Section (ref) we propose a feasible version of our limit theory and compare it with the finite-sample distribution. In Section (ref) we discuss some differences between our approach and others. Let
In order to use the continuous record asymptotic distribution in practice one needs consistent estimates of the unknown quantities. In this section, we compare the finite-sample distribution of the least-squares estimator of the change-point date with a feasible version of the continuous record asymptotic distribution obtained with plug-in estimates. We obtain the finite-sample distribution of $\rho(\widehat{T}_{b,\pi}-T_{b}^{0})$ based on 100,000 simulations from the following model:
where $Z_{t}=0.5Z_{t-1}+u_{t}$ with $u_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right)$ independent of $e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,\sigma_{e}^{2}\right)$, $\sigma_{e}^{2}=1$, $\nu^{0}=1$, $Z_{0}=0$, $D_{t}=1$ for all $t$, and $T=100.$ We set $\pi=0.05$, $T_{b}^{0}=\left\lfloor T\lambda_{0}\right\rfloor $ with $\lambda_{0}=0.3,\,0.5,\,0.7$ and consider different break sizes $\delta_{Z}^{0}=0.2,\,0.3,\,0.5,\,1$. The infeasible continuous record asymptotic distribution is computed assuming knowledge of the data generating process (DGP) as well as of the model parameters, i.e., using Theorem (ref) where we set $N_{b}^{0}$ equal to its true value, $||\delta^{0}||^{-2}\overline{\sigma}^{2}=||\delta_{Z}^{0}||^{-2}\sigma_{e}^{2},$ and $\xi_{1},\,\xi_{2}$ and $\rho$ equal to their true values, respectively, with $\delta_{Z}^{0}$ in place of $\delta^{0}$. Note that the scaling $h^{\kappa/2}$ and $h^{-1/4}$ in the definition of $\delta_{h}$ and $\sigma_{h}$ respectively, cancel using the fact that they appear in both numerator and denominator and applying a change in variables. The feasible counterparts are constructed with plug-in estimates of $\xi_{1},\,\xi_{2},\,\rho$ and $(N_{b}^{0}\left\Vert \delta^{0}\right\Vert ^{2}/\overline{\sigma}^{2})\rho$. In practice we need to use a normalization for $N$. A common choice is $N=1$. Then $\widehat{\lambda}_{b}^{\mathrm{}}=\widehat{T}_{b}/T$ is a natural estimate of $\lambda_{0}$, using the consistency result of $\widehat{\lambda}_{b}^{\mathrm{}}$ that holds in the setting of Theorem (ref) which can also be rationalized for large $N$ under the conditions of Theorem (ref). In practice this means that we approximate the distribution of the estimator $\widehat{\lambda}_{b,\pi}$ where $\pi$ is chosen by the researcher and we plug-in the estimator $\widehat{\lambda}_{b}$ which can be based on any trimming because of the consistency property. Here, we set $\widehat{\lambda}_{b}$ equal to the least-squares estimator based on a trimming $0.15$, which is also used for the other plug-in estimates. The estimates of $\xi_{1}$ and $\xi_{2}$ are given, respectively, by
where $\widehat{\delta}$ is the least-squares estimator of $\delta_{h}$ and $\widehat{e}_{kh}$ are the least-squares residuals. Note that in $\widehat{\xi}_{1}$ and $\widehat{\xi}_{2}$, the estimate $\widehat{\delta}$ appears in both numerator and denominator so that the scaling $h^{\kappa/2}$ in the definition of $\delta_{h}$ cancels. Use is made of the fact that $\left\langle Z,\,Z\right\rangle _{1}$ is consistently estimated by $\sum_{k=1}^{\widehat{T}_{b}^{\mathrm{}}}z_{kh}z'_{kh}/\widehat{\lambda}_{b}^{\mathrm{}}$ while $\Omega_{\mathscr{W},1}$ is consistently estimated by $T\sum_{k=1}^{\widehat{T}_{b}^{\mathrm{}}}\widehat{e}_{kh}^{2}z_{kh}z'_{kh}/\widehat{\lambda}_{b}^{\mathrm{}}$. The method to estimate $\lambda_{0}\left\Vert \delta^{0}\right\Vert ^{2}\overline{\sigma}^{-2}\rho$ is less immediate because it involves manipulating the scaling of each of the three estimates. Let $\vartheta=\left\Vert \delta^{0}\right\Vert ^{2}\overline{\sigma}^{-2}\rho$. We use the following estimates for $\vartheta$ and $\rho$, respectively,
Whereas we have $\widehat{\xi}_{i}\overset{p}{\rightarrow}\xi_{i}$ $\left(i=1,\,2\right)$, the corresponding approximations for $\widehat{\rho}$ and $\widehat{\vartheta}$ are given by $\widehat{\rho}/h^{\kappa}\overset{p}{\rightarrow}\rho$ and $\widehat{\vartheta}/h^{2\kappa}\overset{p}{\rightarrow}\vartheta$. However, before letting $T\rightarrow\infty$ we can apply a change in variable using the fact that $\widehat{\lambda}_{b}-\lambda_{b}^{0}=O\left(h^{1-\kappa}\right)$ which result in the extra factor $h^{2\kappa}$ canceling.
The proposition implies that the limiting distribution can be simulated using plug-in estimates. This allows feasible inference about the break date. The results are presented in Figure (ref)-(ref) which also plot the asymptotic distribution from bai:97RES and the infeasible distribution from Theorem (ref). Here by signal-to-noise ratio we mean $\delta_{Z}^{0}/\sigma_{e}$ which, given $\sigma_{e}^{2}=1$, equals the break size $\delta_{Z}^{0}$.
Several interesting observations appear at the outset. The density of the large-$N$ shrinkage asymptotic distribution does not depend on the location of the break, and thus it is always unimodal and symmetric about the origin. None of these features are shared by the density derived under a continuous record. When the true break is at mid-sample $\left(\lambda_{0}=0.5\right)$, the density function is symmetric and centered at zero. However, when the signal-to-noise ratio is low, the density features three modes. This tri-modality vanishes as the signal-to-noise ratio increases. When $\delta_{Z}^{0}$ is low and the break is not at mid-sample the density is asymmetric; for values of $\lambda_{0}$ less (larger) than 0.5, the density is right (left) skewed. When the signal is low and $\lambda_{0}$ is less (larger) than 0.5, the density has highest mode at some value near $\widehat{\lambda}_{b}$ being close to the starting (end) sample point than centered at $\lambda_{0}.$ However, as in the case of $\lambda_{0}=0.5,$ when the signal-to-noise ratio increases the highest mode is centered at a value which corresponds to $\widehat{\lambda}_{b}$ being close to $\lambda_{0}$. Asymmetry and multi-modality of the finite-sample distribution of the break point estimator were also found by perron/zhu:05 and deng/perron:06 in models with a trend.
The interpretation of these features are straightforward. For example, asymmetry reflects the fact that the span of the data and the actual location of the break play a crucial role on the behavior of the estimator. If the break occurs early in the sample there is a tendency to overestimate the break date and vice-versa if the break occurs late in the sample. The marked changes in the shape of the density as we raise $\delta_{Z}^{0}$ confirms that the magnitude of the shift matters a great deal as well. The tri-modality of the density when the shift size is small reflects the uncertainty in the data as to whether a structural change is present at all; i.e., the least-squares estimator finds it easier to locate the break at either the beginning or the end of the sample. Unlike the shrinkage asymptotic distribution, the density of the feasible version of the continuous record distribution provides a remarkably good approximation to the infeasible one and thus also to the finite-sample distribution. The extended working paper casini/perron_CR_Single_Break_Extended shows that the quality of the approximation is good for a variety of models.
The figures reported above have shown that there is a high degree of uncertainty when the break magnitude is not large. The classical shrinkage asymptotics of bai:97RES with $\delta_{T}$ required to convergence to zero at a rate slower than $O(T^{-1/2})$ clearly underestimates that degree of uncertainty and, as the figures show, provides a poor approximation to the finite-sample behavior of the least-squares estimator. In Section (ref) we show that this issue is responsible for the poor coverage probabilities of the confidence intervals introduced in bai:97RES when the break magnitude is small. On the other hand, elliott/mueller:07 and elliott/mueller/watson:15 require $\delta_{T}$ to go to zero at the fast rate $O(T^{-1/2})$ leading to weak identification. The latter implies that the relevant quantities in the model become inconsistent. This can be problematic for inference and indeed, their inference often suffers from the opposite problem in that confidence intervals for $\widehat{T}_{b}$ can be too large {[}Casini and Perron (casini/perron_Oxford_Survey; casini/perron_Lap_CR_Single_Inf) and chang/perron:18{]}.
We impose conditions on the signal-to-noise ratio $\delta/\sigma$ rather than just on $\delta.$ Consider a simple location model with a change $\delta$ in the mean and independent errors. What describes the uncertainty about the break in this model is the ratio $\delta/\sigma$ where $\sigma$ is the volatility of the errors. We let $\delta$ go to zero at a not too fast rate while letting $\sigma$ increase to infinity in a neighborhood of $T_{b}^{0}$. That is $\left(\delta_{T}/\sigma_{t}\right)\rightarrow0$ at rate $O(T^{-1/2})$ in a neighborhood of $T_{b}^{0}$. Interestingly, this is the same rate Elliott and M{\"u}ller used for $\delta_{T}\rightarrow0.$ Away from $T_{b}^{0}$, we require $\left(\delta_{T}/\sigma_{t}\right)\rightarrow0$ at slower rate---similar to yao:87 and bai:97RES. The difference now is that we do not lose identification and all the parameters in the model remain consistent. Under continuous-time, the variance of the processes is proportional to the sampling interval. This allows us to trade-off the rate of convergence at which $\widehat{\lambda}_{b}$ approaches $\lambda_{0}$ with the variance of the errors in a neighborhood of $T_{b}^{0}$ by letting $\sigma_{t}$ become large when $t$ is close to $T_{b}^{0}$ {[}i.e., a change of time scale as in Foster and Nelson (nelson/foster:94, foster/nelson:96){]}. This offers a new characterization of higher uncertainty without losing identification.
The features of the limit and finite-sample distributions suggest that standard methods to construct confidence intervals may be inappropriate; e.g., two-sided intervals around the estimated break date based on the standard deviations of the estimate. Our suggested approach is rather non-standard and relates to Bayesian methods. In our context, the Highest Density Region (HDR) seems the most appropriate in light of the asymmetry and, especially, the multi-modality of the distribution for small break sizes. All that is needed to implement the procedure is an estimate of the density function, using plug-in estimates as explained in Section (ref). Choose some significance level $0<\alpha<1$ and let $\widehat{P}_{T_{b}}$ denote the empirical counterpart of the probability distribution of $\rho N(\widehat{\lambda}_{b,\pi}-\lambda_{b}^{0})$ as defined in Theorem (ref). Further, let $\widehat{p}_{T_{b}}$ denote the density function defined by the Radon-Nikodym equation $\widehat{p}_{T_{b}}^{\mathrm{}}=d\widehat{P}_{T_{b}}^{\mathrm{}}/d\lambda_{\mathrm{L}},$ where $\lambda_{\mathrm{L}}$ denotes the Lebesgue measure.
The concept of HDR and of its estimation has an established literature in statistics. The definition reported here is from hyndman:96; see also samworth/wand:10 and mason/polonik:09 (mason/polonik:08; mason/polonik:09) for more recent developments.\nocite{mason/polonik:08}
The confidence set $C(\textrm{cv}_{\alpha})$ has a frequentist interpretation even though the concept of HDR is often encountered in Bayesian analyses since it associates naturally to the derived posterior distribution, especially when the latter is multi-modal. A feature of the confidence set $C(\textrm{cv}_{\alpha})$ under our context is that, at least when the size of the shift is small, it consists of the union of several disjoint intervals. The appeal of using HDR is that one can directly deal with such features. As the break size increases and the distribution becomes unimodal, the HDR becomes equivalent to the standard way of constructing confidence sets. In practice, one can proceed as follows.
This procedure will not deliver contiguous confidence sets when the size of the break is small. Indeed, we find that in such cases, the overall confidence set for $T_{b}^{0}$ consists in general of the union of disjoint intervals if $\widehat{T}_{b}$ is not near the tails of the sample. One is located around the estimate of the break date, while the others are in the pre- and post-break regimes. To provide an illustration, we consider a simple example involving a single draw from a simulation experiment. Figure (ref) reports the HDR of the feasible limiting distribution of $\rho(\widehat{T}_{b,\pi}-T_{b}^{0})$ for a random draw from the model in (ref) with parameters $\nu^{0}=1$, $\beta^{0}=0$, unit variance and autoregressive coefficient 0.6 for $Z_{t}$ and $\sigma_{e}^{2}=1.2$. We set $\lambda_{0}=0.35,\,0.5$ and $\delta_{Z}^{0}=0.3,\,0.8,\,1.5$. We use a trimming 0.15 for the plug-in estimator $\widehat{T}_{b}$ and $\pi=0.05$ for $\widehat{T}_{b,\pi}$. As explained in Section (ref), we could use any other trimming in place of 0.15. The results remain unchanged. We set $T=100$ and the significance level is $\alpha=0.05$. Note that the origin is at the estimated break date. The point on the horizontal axis corresponds to the true break date. The black intervals on the horizontal axis correspond to regions of high density. The resulting confidence set is their union. Once a confidence region for $\rho(\widehat{T}_{b,\pi}-T_{b}^{0})$ is computed, it is straightforward to derive a 95% confidence set for $T_{b}^{0}.$ The top panel (left plot) reports results for the case $\delta_{Z}^{0}=0.3$ and $\lambda_{0}=0.35$ and shows that the HDR is composed of two disjoint intervals. The estimated break date is $\widehat{T}_{b}=70$ and the implied 95% confidence set for $T_{b}^{0}$ is given by $C(\textrm{cv}_{0.05})=\left\{ 1,\ldots,12\right\} \cup\left\{ 18,\ldots100\right\} $. This includes $T_{b}^{0}$ and the overall length is 95 observations. Table (ref) reports for various methods whether $T_{b}^{0}$ is covered or not and the length of the confidence sets for this example. The length of bai:97RES\textquoteright s (1997) confidence interval is 55 but does not include $T_{b}^{0}$. elliott/mueller:07\textquoteright s (2007) confidence set, denoted by $\widehat{U}_{T}.\textrm{eq}$ in Table (ref), also does not include the true break date at the 90% confidence level, but does so at the 95% and its length is 95. Our method covers $T_{b}^{0}$ and has a relatively short length across different $\delta_{Z}^{0}$.
We now assess via simulations the finite-sample performance of the method proposed to construct confidence sets for the break date. We also make comparisons with alternative methods in the literature: bai:97RES's bai:97RES approach based on the large-$N$ shrinkage asymptotics; elliott/mueller:07\textquoteright s elliott/mueller:07, hereafter EM, method on\textcolor{red}{ }inverting nyblom:89's nyblom:89 statistic;\textcolor{red}{ }the Inverted Likelihood Ratio (ILR) approach of eo/morley:15. We omit the technical details of these methods and refer to the original sources or chang/perron:18 for a review and comparisons. We consider two DGPs: M1 is $y_{t}=\beta^{0}+\delta_{Z}^{0}\boldsymbol{1}_{\left\{ t>T_{b}^{0}\right\} }+e_{t}$ with $\beta^{0}=1$ and $e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right)$; M2 is $y_{t}=\delta_{Z}^{0}\left(1-\nu^{0}\right)\boldsymbol{1}_{\left\{ t>T_{b}^{0}\right\} }+\nu^{0}y_{t-1}+e_{t}$ with $\nu^{0}=0.8$ and $e_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,0.04\right)$. Our companion paper casini/perron_CR_Single_Break_Extended includes extensive simulation results. We set the significance level at $\alpha=0.05$, and the break occurs at date $\left\lfloor T\lambda_{0}\right\rfloor $, where $\lambda_{0}=0.2,\,0.35,\,0.5$ and $T=200$ for M1 and $T=100$ for M2. The results are presented in Table (ref)-(ref). The last row in each table includes the rejection probability of a 5%-level sup-Wald test using the asymptotic critical value in andrews:93, which provides a measure of the magnitude of the break relative to the noise. For models with predictable processes we use the two-step procedure described in Section (ref).
Overall, the simulation results confirm previous findings about the performance of existing methods. bai:97RES\textquoteright s bai:97RES method has a coverage rate below the nominal level when the size of the break is small. Overall, our HDR method and that of EM show accurate empirical coverage rates for all DGP considered. However, EM\textquoteright s method almost always displays confidence sets which are larger than those from the other approaches. Over all DGPs considered, the average length of the HDR confidence sets are 40% to 70% shorter than those obtained with EM\textquoteright s approach when the size of the shift is moderate to high. The results for M2, a change in mean with a lagged dependent variable and strong correlation, are quite revealing. EM\textquoteright s method yields confidence intervals that are very wide, increasing with the size of the break and for large breaks covering nearly the entire sample. This does not occur with the other methods. For instance, when $\lambda_{0}=0.5$ and $\delta_{Z}^{0}=2$, the average length from the HDR method is 8.34 compared to 93.71 with EM\textquoteright s. This concurs with the results in chang/perron:18.
In summary, the small-sample simulation results suggest that our continuous record HDR-based inference provides accurate coverage probabilities close to the nominal level and average lengths of the confidence sets shorter relative to existing methods. It is also valid and reliable under a wider range of DGPs including long-memory processes. Specifically noteworthy is the fact that it performs well for all break sizes, whether small or large.
We examined a change-point model under a continuous record asymptotics. With the time horizon $\left[0,\,N\right]$ fixed, we can account for the asymmetric informational content provided by the pre- and post-break samples. We derived a feasible counterpart of the continuous record asymptotic distribution of the change-point estimator using consistent plug-in estimates and showed that it provides accurate approximations to the finite-sample distributions. We used our limit theory to construct confidence sets for the change-point date based on the concept of Highest Density Region. Overall, it delivers accurate coverage probabilities and relatively short average lengths of the confidence sets. Importantly, it does so irrespective of the magnitude of the break, whether large or small, a notoriously difficult problem in the literature.
\addcontentsline{toc}{section}{References}
\pagenumbering{arabic}