The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
110,213 characters
Change-Point Analysis of Time Series with Evolutionary Spectra
\pagebreak{}
\setcounter{page}{0}
\raggedbottom
\title{\textbf{Change-Point Analysis of Time Series with Evolutionary Spectra}\thanks{We thank an Associate Editor and two referees for helpful comments.
We also thank Federico Belotti, Leopoldo Catania and Stefano Grassi
for generous help with computer programming. }}
\maketitle
\begin{abstract}
This paper develops change-point methods for the spectrum of a locally
stationary time series. We focus on series with a bounded spectral
density that change smoothly under the null hypothesis but exhibits
change-points or becomes less smooth under the alternative. We address
two local problems. The first is the detection of discontinuities
(or breaks) in the spectrum at unknown dates and frequencies. The
second involves abrupt yet continuous changes in the spectrum over
a short time period at an unknown frequency without signifying a break.
Both problems can be cast into changes in the degree of smoothness
of the spectral density over time. We consider estimation and minimax-optimal
testing. We determine the optimal rate for the minimax distinguishable
boundary, i.e., the minimum break magnitude such that we are able
to uniformly control type I and type II errors. We propose a novel
procedure for the estimation of the change-points based on a wild
sequential top-down algorithm and show its consistency under shrinking
shifts and possibly growing number of change-points.
\end{abstract}
\indent {\bf JEL Classification:} C12, C13, C22. \\
\noindent{\bf Keywords:} Change-point, Locally stationary, Frequency domain, Minimax-optimal test.
\onehalfspacing
\thispagestyle{empty}
\allowdisplaybreaks
\pagebreak{}
\section{Introduction}
Classical change-point theory focuses on detecting and estimating
structural breaks in the mean or regression coefficients. Early contributions
include, among others, \citet{hinkley:71}, \citet{yao:87}, \citet{andrews:93},
\citet{horvath:93} and \citet{bai/perron:98}, who assume the presence
of a single or multiple change-points in the parameters of an otherwise
stationary time series model; see the reviews of \citet{aue/Horvath:13}
and \citet{casini/perron_Oxford_Survey} for more details. More recently
there has been a growing interest about functional and time-varying
parameter models where the latter are characterized by infinite-dimensional
parameters which change continuously over time {[}see, e.g., \citet{dahlhaus:96},
\citet{neumann/von_sachs:1997}, \citet{hormann/kokoszka:2010}, \citet{dette/preus/vetter:2011},
\citet{zhang/wu:12}, \citet{panaretos/tavakoli:13}, \citet{aue/hormann:2015}
and \citet{vandelft/eichler:2018}{]}. Several authors have extended
the stationarity tests originally introduced by \citet{priesley/rao:1969},
and further developed by, e.g., \citet{dwivedi/rao:2010}, \citet{jentsch/subbarao:2015}
and \citet{bandyopadhyay/carsten/rao}, to these settings. In the
context of local stationarity, \citet{paparoditis:2009} proposed
a test based on comparing a local estimate of the spectral density
to a global estimate and \citet{preuss/vetter/dette:2013} proposed
a test for stationarity using empirical process theory. In the context
of functional time series, tests for stationarity were considered
by \citet{horvath/kokoszka/rice:2014} and \citet{aue/rice/sonmez}
using time-domain methods, and by \citet{aue/vandelft:2020} and \citet{vandelft/Characiejus/dette:2018}
using frequency-domain methods.
There is wide agreement in empirical work that besides breaks in the
mean the detection of breaks in the variance or the correlation structure
of a time series is of importance. For example, the discrimination
between regimes of high and low asset volatility is of central interest
in finance and the detection of changes in the parameters of an autoregressive
process is important to build superior forecasting procedures. In
addition, discerning the type of the changes (continuous or abrupt
changes) is useful in applied work. On the one hand, if one assumes
local stationarity but the true data-generating process involves structural
breaks then parameter estimates may be severely biased and inference
may be misleading. On the other hand, if one assumes a structural
break model but the parameters actually change gradually then similar
issues may arise. A general approach that allows for both continuous
changes as well as breaks is needed to avoid these issues. Therefore,
the detection of breaks in an otherwise locally stationary time series
is important.
We develop inference methods about the changes in the degree of smoothness
of the spectrum of a locally stationary time series, and hence, about
change-points in the spectrum as a special case. The key parameter
is the regularity exponent that governs how smooth the path of the
local spectral density is over time. We address two local problems.
The first is the detection of discontinuities (or breaks) in the spectrum
at some unknown date and frequency. The second involves the detection
of abrupt yet continuous changes in the spectrum over a short time
period at an unknown frequency without signifying a break. For example,
the spectrum becomes rougher over a short time period, meaning that
the paths are less smooth as quantified by the regularity exponent.
This can occur for a stationary process whose parameters start to
evolve smoothly according to Lipschitz continuity, or for a locally
stationary process with Lipschitz parameters that change to continuous
but non-differentiable functional parameters. For example, the volatility
of high-frequency stock prices or of other macroeconomic variables
is known to become rougher (i.e., less smooth) without signifying
a structural break after central banks' official announcements, especially
in periods of high market uncertainty. In seismology, earthquakes
are made up of several seismic waves that arrive at different times
and so changes in the smoothness properties of each wave is important
for locating the epicenter and identifying what materials the waves
have passed through. We consider minimax-optimal testing and estimation
for both problems, following \citet{ingster:93}. We determine the
optimal rate for the minimax distinguishable boundary, i.e., the minimum
break magnitude such that we are still able to uniformly control type
I and type II errors. These results are different from the developments
on minimax optimality obtained recently in the statistics literature
for the classical change-point problem where the mean of the series
is piecewise constant {[}see, e.g., \citet{liu/gao/samworth} and
\citet{verzelen/fromont/lerasle/reynaud-bouret:2020}{]}.
The problem of discriminating discontinuities from a continuous evolution
in a nonparametric framework has received relatively less attention
than the classical change-point problem with a few exceptions {[}\citet{muller:92},
\citet{spokoiny:98}, \citet{muller/stadtmuller:99}, \citet{wu/zhao:07}
and \citet{bibinger/jirak/vetter:16}{]}. These focused on nonparametric
regression and high-frequency volatility, and considered time-domain
methods while we consider frequency-domain methods. This adds a difficulty
in that, e.g., the search for a break or smooth change has to run
over two dimensions, the time and frequency indices. Our test statistics
are the maximum of local two-sample $t$-tests based on the local
smoothed periodogram. We construct statistics that allow the researcher
to test for a change-point in the spectrum at a prespecified frequency
and others that allow to detect a break in the spectrum without prior
knowledge about the frequencies. These test statistics can detect
both discontinuous and smooth changes, and therefore are useful for
both inference problems discussed above. The asymptotic null distribution
follows an extreme value distribution. In order to derive this result,
we first establish several asymptotic results, including bounds for
higher-order cumulants and spectra of locally stationary processes.
These results are complementary to some in \citet{dahlhaus:96}, \citet{paparoditis:2009},
\citet{panaretos/tavakoli:13}, \citet{aue/vandelft:2020} and \citet{casini_hac},
and extend some classical frequency-domain results for stationary
processes {[}e.g., \citet{brillinger:75}{]} to locally stationary
processes.
Change-point problems have also been studied in the frequency-domain
in several fields, though with less generality. \citet{adak:98} investigated
the detection of change-points in piecewise stationary time series
by looking at the difference in the power spectral density for two
adjacent regimes. She compared several distance metrics such as the
Kolmogorov-Smirnov, Cr\'{a}mer-Von Mises and CUSUM-type distance
proposed by \citet{coates/diggle}. \citet{last/shumway:08} focused
on detecting change-points in piecewise locally stationary series.
They exploited some of the results in \citet{Kakizawa/shumway/taniguchi}
and \citet{huang/ombao/stoffer:2004} to propose a Kullback-Liebler
discrimination information but did not derive the null distribution
of the test statistic. \citet{dette/wu/zhou:2019} considered testing
for change-points in the autocorrelation coefficient at some pre-specified
lag $k$ while allowing for local stationarity under the null hypothesis.
They proposed a time-domain method based on a CUSUM-type test. To
the extent that the autocorrelation at a given lag is only one of
the features contained in the spectrum, our setting is more general.
\citet{zhang:16} considered testing for and estimating change-points
in the mean of a piecewise locally stationary time series. This is
a different problem from ours since the spectrum involves the second-order
properties which are complex to study. \citet{preuss/puchstein/dette:2015}
considered the detection and estimation of change-points in the autocovariance
function but they required stationarity under the null hypothesis.
Several authors considered methods based on segmenting the wavelet
spectrum for piecewise stationary time series {[}see, e.g., \citet{barigozzi/cho/fryzlewicz:2018}
and \citeauthor{cho/fryzlewicz:2012} (\citeyear{cho/fryzlewicz:2012};
\citeyear{cho/fryzlewicz:2017}){]}. Outside of the wavelet context
other contributions are \citet{kirch/muhsal/ombao:2015} and \citet{Schrhoder/ombao:2019}.
\citet{vogt/dette:2015} investigated the detection of gradual changes
in a locally stationary time series. Our contribution is different
since we provide a general change-point analysis about the time-varying
spectrum of a time series and establish the relevant asymptotic theory
of the proposed test statistics under both the null and alternative
hypotheses.
We also address the problem of estimating the change-points, allowing
their number to increase with the sample size and the distance between
change-points to shrink to zero. We propose a procedure based on
a wild sequential top-down algorithm that exploits the idea of bisection
combining it with a wild resampling technique similar to the one proposed
by \citet{fryzlewicz:14}. We establish the consistency of the procedure
for the number of change-points and their locations. We compare the
rate of convergence with that of standard change-point estimators
under the classical setting {[}e.g., \citet{yao:87}, \citet{bai:94a},
\citet{casini/perron_CR_Single_Break}, \citet{casini/perron_Lap_CR_Single_Inf}
and \citet{casini/perron_SC_BP_Lap}{]}. We verify the performance
of our methods via simulations which show their benefits. The advantage
of using frequency-domain methods to detect change-points is that
they do not require to make assumptions about the data-generating
process under the null hypothesis beyond the fact that the spectrum
is bounded. Furthermore, the method allows for a broader range of
alternative hypotheses than time-domain methods which usually have
good power only against some specific alternatives. For example, tests
for changes in the volatility do not have power for changes in the
dependence and vice versa. Our methods are readily available for use
in many fields such as speech processing, biomedical signal processing,
seismology, failure detection, economics and finance. It can also
be used as a pre-test before constructing the recently introduced
double kernel long-run variance estimator that accounts more flexibly
for nonstationarity {[}cf. \citet{casini_hac}, \citet{casini/perron_PrewhitedHAC}
and \citet{casini/perron_Low_Frequency_Contam_Nonstat:2020}{]}.
The rest of the paper is organized as follows. Section \ref{Section Hypotheses Testing Problem}
introduces the statistical setting and the hypothesis testing problems.
Section \ref{Section Change-Point-Tests} presents the test statistics
and states their null limit distributions. Section \ref{Section Consistency-and-Minimax}
addresses the consistency of the tests and their minimax optimality.
Section \ref{Section Estimation-of-the Breaks} discusses the estimation
of the change-points while Section \ref{Section Implementation} provides
details for the implementation of the methods. The results of some
Monte Carlo simulations are presented in Section \ref{Section Monte Carlo}.
An empirical application is presented in Section \ref{Section Empirical Application}.
Section \ref{Section Conclusions} reports brief concluding comments.
An online supplement {[}cf. \citet{casini/perron:change-point-spectra_Supp}{]}
contains additional theoretical and empirical results, and all mathematical
proofs.
\section{\label{Section Hypotheses Testing Problem}Statistical Environment
and the Testing Problems}
Section \ref{Subsec Segmented-Locally-Stationary} introduces the
statistical setting and Section \ref{subsec The hypotheses testin problem}
presents the hypotheses testing problems. We work in the frequency-domain
under the locally stationary framework introduced by \citet{dahlhaus:96}
who formalized the ideas of \citet{priestley:65}. \citet{casini_hac}
extended his framework to allow for discontinuities in the spectrum
which then results in a segmented locally stationary process. This
corresponds to the relevant process under the alternative hypothesis
of breaks in the spectrum. Since local stationarity is a special case
of segmented local stationarity we begin with the latter. We use an
infill asymptotic setting whereby we rescale the original discrete
time horizon $\left[1,\,T\right]$ by dividing each $t$ by $T.$
\subsection{\label{Subsec Segmented-Locally-Stationary}Segmented Locally Stationary
Processes}
Suppose $\left\{ X_{t}\right\} _{t=1}^{T}$ is defined on an abstract
probability space $\left(\Omega,\,\mathscr{F},\,\mathbb{P}\right)$,
where $\Omega$ is the sample space, $\mathscr{F}$ is the $\sigma$-algebra
and $\mathbb{P}$ is a probability measure. Let $i\triangleq\sqrt{-1}$.
We use the notation $\overline{A}$ for the complex conjugate of $A\in\mathbb{C}$.
\begin{defn}
\label{Definition Segmented-Locally-Stationary}A sequence of stochastic
processes $\{X_{t,T}\}_{t=1}^{T}$ is called segmented locally stationary
(SLS)\textbf{ }with $m_{0}+1$ regimes, transfer function $A^{0}$
and trend $\mu$ if there exists a representation
\begin{align}
X_{t,T} & =\mu_{j}\left(t/T\right)+\int_{-\pi}^{\pi}\exp\left(i\omega t\right)A_{j,t,T}^{0}\left(\omega\right)d\xi\left(\omega\right),\qquad\qquad\left(t=T_{j-1}^{0}+1,\ldots,\,T_{j}^{0}\right),\label{Eq. Spectral Rep of SLS}
\end{align}
for $j=1,\ldots,\,m_{0}+1$, where by convention $T_{0}^{0}=0$ and
$T_{m_{0}+1}^{0}=T$ ($\mathcal{T}\triangleq\{T_{1}^{0},\,\ldots,\,T_{m_{0}}^{0}\}$),
and the following holds:
(i) $\xi\left(\omega\right)$ is a stochastic process on $\left[-\pi,\,\pi\right]$
with $\overline{\xi\left(\omega\right)}=\xi\left(-\omega\right)$
and
\begin{align*}
\mathrm{cum}\left\{ d\xi\left(\omega_{1}\right),\ldots,\,d\xi\left(\omega_{r}\right)\right\} & =\varphi\left(\sum_{j=1}^{r}\omega_{j}\right)g_{r}\left(\omega_{1},\ldots,\,\omega_{r-1}\right)d\omega_{1}\cdots d\omega_{r},
\end{align*}
where $\mathrm{cum}\left\{ \cdot\right\} $ is the cumulant of
$r$th order, $g_{1}=0,\,g_{2}\left(\omega\right)=1$, $\left|g_{r}\left(\omega_{1},\ldots,\,\omega_{r-1}\right)\right|\leq M_{r}<\infty$
and $\varphi\left(\omega\right)=\sum_{j=-\infty}^{\infty}\delta\left(\omega+2\pi j\right)$
is the period $2\pi$ extension of the Dirac delta function $\delta\left(\cdot\right)$.
(ii) There exists a constant $K>0$ and a piecewise continuous function
$A:\,\left[0,\,1\right]\times\mathbb{R}\rightarrow\mathbb{C}$ such
that, for each $j=1,\ldots,\,m_{0}+1$, there exists a $2\pi$-periodic
function $A_{j}:\,(\lambda_{j-1}^{0},\,\lambda_{j}^{0}]\times\mathbb{R}\rightarrow\mathbb{C}$
with $A_{j}\left(u,\,-\omega\right)=\overline{A_{j}\left(u,\,\omega\right)}$,
$\lambda_{j}^{0}\triangleq T_{j}^{0}/T$ and, for all $T,$
\begin{align}
A\left(u,\,\omega\right)=A_{j}\left(u,\,\omega\right) & \,\mathrm{\,for\,}\,\lambda_{j-1}^{0}<u\leq\lambda_{j}^{0},\label{Eq A(u) =00003D Ai}\\
\sup_{1\leq j\leq m_{0}+1}\sup_{T_{j-1}^{0}<t\leq T_{j}^{0}}\sup_{\omega\in\left[-\pi,\,\pi\right]}\left|A_{j,t,T}^{0}\left(\omega\right)-A_{j}\left(t/T,\,\omega\right)\right| & \leq K\,T^{-1}.\label{Eq. 2.4 Smothenss Assumption on A}
\end{align}
(iii) $\mu_{j}\left(t/T\right)$ is piecewise continuous.
\end{defn}
The smoothness properties of $A$ in $u$ guarantees that $X_{t,T}$
has a piecewise locally stationary behavior. This means that $X_{t,T}$
is locally stationary in each segment where the notion of local stationarity
is as introduced by \citet{dahlhaus:96}. We refer to \citet{casini_hac}
for several theoretical properties of SLS processes. \citet{zhou:2013}
considered piecewise locally stationary processes in a time-domain
setting but his notion is less general. In particular, \citet{casini_hac}
also defined (and worked with) the covariance between observations
belonging to different segments whereas previous works considered
only the covariance between observations belonging to the same segment,
thereby using smoothness which simplifies the analysis.
\subsection{\label{subsec The hypotheses testin problem}The Testing Problems}
We focus on time-varying spectra that are bounded, thereby excluding
unit root, and long memory processes. Unit roots or trending processes
can be handled by, for example, taking first differences or using
other de-trending techniques. For some $D<\infty$, we consider the
following class of time-varying spectra under the null hypothesis,
\begin{align}
\boldsymbol{F}\left(\theta,\,D\right) & =\left\{ \left\{ f\left(u,\,\omega\right)\right\} _{u\in\left[0,\,1\right],\,\omega\in\left[-\pi,\,\pi\right]}:\,\sup_{\omega\in\left[-\pi,\,\pi\right],\,}\sup_{u,\,v\in\left[0,\,1\right],\,\left|v-u\right|<h}\left|f\left(u,\,\omega\right)-f\left(v,\,\omega\right)\right|\leq Dh^{\theta}\right\} .\label{eq (2) BJV}
\end{align}
At all continuity points $u$, $f\left(u,\,\omega\right)$ is related
to $A\left(u,\,\omega\right)$ by the relationship $f\left(u,\,\omega\right)=|A\left(u,\,\omega\right)|^{2}$.
The key parameter of the testing problem under the null hypotheses
is $\theta>0$. This is the regularity exponent of $f$ in the time
dimension. For $\theta>1$, $f$ is constant in $u$ and reduces to
the spectral density of a stationary process. For $\theta=1$, $f$
is Lipschitz continuous in $u$. For $\theta<1$, $f$ is $\theta$-H\"older
continuous. Local stationarity corresponds to $\theta>0$ and $f$
being differentiable {[}see \citet{dahlhaus:96b}{]}. The latter is
the setting that we consider under the null hypothesis. To avoid redundancy,
we do not require differentiability directly for the functions in
$\boldsymbol{F}\left(\theta,\,D\right)$ since below we assume that
the transfer function $A\left(u,\,\omega\right)$ is differentiable
in $u$ which in turn implies that $f$ is differentiable in $u$.
Since most of the applied work concerning local stationarity relies
on Lipschitz continuity (i.e., $\theta=1$), our specification of
the null is more general.
We now discuss features that are relevant under the alternative hypothesis.
Our focus is on (i) discontinuities of $f$ in $u$ which correspond
to $\theta=0$ and (ii) decreases in the smoothness of the trajectory
$u\mapsto f(u,\,\omega)$ for each $\omega$ which correspond to a
decrease in $\theta$. We shall refer to a series affected by the
changes of case (ii) as ``becoming more rough or less smooth''.
Both cases refer to the properties of the spectral density and thus
to the second-order properties of $X_{t,T}$.
Case (i) involves a break in the spectrum, i.e., there exits $\lambda_{b}^{0}\in\left(0,\,1\right)$
such that $\Delta f\left(\lambda_{b}^{0},\,\omega\right)\triangleq(f(\lambda_{b}^{0},\,\omega)-\lim_{u\downarrow\lambda_{b}^{0}}f(u,\,\omega))\neq0$
for some $\omega\in\left[-\pi,\,\pi\right]$.
Case (ii) involves a fall in the regularity exponent from $\theta>0$
to $\theta'\in\left(0,\,\theta\right)$ after some $\lambda_{b}^{0}$
for some period of time and some $\omega$; i.e., the spectrum becomes
rougher after some $\lambda_{b}^{0}\in\left(0,\,1\right)$ for some
time period before returning to $\theta$-smoothness. The case of
an increase in $\theta$ is technically more complex to handle (see
Section \ref{Section Consistency-and-Minimax} for details). As an
example, consider a locally stationary AR(1),
\begin{align*}
X_{t,T} & =a\left(t/T\right)X_{t-1,T}+\sigma\left(t/T\right)e_{t},\qquad t=1,\ldots,\,T,
\end{align*}
where $a:\,\left[0,\,1\right]\rightarrow\left(-1,\,1\right)$ and
$\sigma:\,\left[0,\,1\right]\rightarrow\mathbb{R}_{+}$ are functional
parameters satisfying a Lipschitz condition and $\left\{ e_{t}\right\} $
is an i.i.d. mean-zero sequence. Additionally, if $\sup_{u\in\left[0\,1\right]}|\sigma\left(u\right)|<\infty$
and the initial condition satisfies some regularity condition, then
$f\left(u,\,\omega\right)$ is uniformly bounded and $\theta=1$.
Problem (ii) refers to either $a\left(\cdot\right)$ or $\sigma\left(\cdot\right)$,
or both, becoming less smooth, i.e., we have a change from $\theta=1$
(since the parameters satisfy a Lipschitz condition) to some $\theta'\in\left(0,\,1\right)$.
In this case, $X_{t,T}$ becomes an AR(1) process with functional
parameters that are still continuous but less smooth and may not be
differentiable.
Case (i) has received most attention so far in the time series literature
although under much stronger assumptions {[}e.g., $f\left(u,\,\omega\right)=f\left(\omega\right)${]}.
Case (ii) is a new testing problem and can be of considerable interest
in several fields even though it requires larger sample sizes than
problem (i). For example, the volatility dynamics of financial or
macroeconomic variables vary over time. \citet{bibinger/jirak/vetter:16}
provided evidence that the volatility of stock prices can change substantially
its path properties after a press release following a meeting of the
Federal Open Market Committee, in particular, its path may become
more rough. Since $f\left(u,\,\omega\right)$ is a smooth function
of the parameters of the data-generating process, if $\sigma\left(u\right)$
becomes more rough, then also $f\left(u,\,\omega\right)$ becomes
more rough as $u$ varies. Case (ii) can also be relevant in seismology
since the study of the path properties of the seismic waves is important
for locating the epicenter of an earthquake. We show below that
our tests are consistent and have minimax optimality properties for
both cases (i) and (ii). Note that case (ii) is a local problem. In
this paper, we do not consider more global problems where for example
the spectrum is such that a fall in $\theta$ to $\theta'\in\left(0,\,\theta\right)$
occurs on $(\lambda_{b}^{0},\,1]$. This represents a continuous change
in the smoothness of the spectrum that persists until the end of
the time interval. Different test statistics are needed for this case,
as will be discussed later.
As discussed by \citet{last/shumway:08}, an important question is
which magnitude of the discontinuity in the time-varying spectrum\textcolor{red}{{}
}can be detected. Or equivalently, how much the spectrum can change
over a short time without indicating a break. We introduce the
quantity $b_{T}$, called the detection boundary or simply ``rate'',
which is defined as the minimum break magnitude $\Delta f\left(\lambda_{b}^{0},\,\omega\right)$
such that we are still able to uniformly control the type I and type
II errors as indicated below. To address the minimax-optimal testing
{[}cf. \citet{ingster:93}{]}, we now introduce the testing problems
(i) and (ii) and defer a more general treatment to Section \ref{Section Consistency-and-Minimax}.
\subsubsection*{Testing Problem for Case (i)}
Given the discussion above, for some fractional break point $\lambda_{b}^{0}\in\left(0,\,1\right)$
and frequency $\omega_{0}$, and a decreasing sequence $b_{T}$, we
consider the following class of alternative hypotheses:
\begin{align}
\boldsymbol{F}_{1,\lambda_{b}^{0},\omega_{0}} & \left(\theta,\,b_{T},\,D\right)\label{Eq. (3) BJV - H1}\\
& =\{\left\{ f\left(u,\,\omega\right)\right\} _{u\in\left[0,\,1\right],\,\omega\in\left[-\pi,\,\pi\right]}:\,\left(f\left(u,\,\omega\right)-\Delta f\left(u,\,\omega\right)\right)_{u\in\left[0,\,1\right]}\in\boldsymbol{F}\left(\theta,\,D\right);\nonumber \\
& \quad|\Delta f\left(\lambda_{b}^{0},\,\omega_{0}\right)|\geq b_{T}\}.\nonumber
\end{align}
We can then present first the hypothesis testing problem that we wish
to address:
\begin{align}
\mathcal{H}_{0} & :\,\left\{ f\left(u,\,\omega\right)\right\} _{u\in\left[0,\,1\right],\,\omega\in\left[-\pi,\,\pi\right]}\in\boldsymbol{F}\left(\theta,\,D\right)\label{eq (4) BJV - Testing Problem}\\
\ensuremath{\mathcal{H}_{1}^{\mathrm{B}}} & :\,\exists\lambda_{b}^{0}\in\left(0,\,1\right)\,\mathrm{and}\,\omega_{0}\in\left[-\pi,\,\pi\right]\,\mathrm{with}\,\left\{ f\left(u,\,\omega\right)\right\} _{u\in\left[0,\,1\right],\,\omega\in\left[-\pi,\,\pi\right]}\in\boldsymbol{F}_{1,\lambda_{b}^{0},\omega_{0}}\left(\theta,\,b_{T},\,D\right).\nonumber
\end{align}
Observe that $\mathcal{H}_{1}^{\mathrm{B}}$ requires at least one
break but allows for multiple breaks even across different $\omega$.
For the testing problem \eqref{eq (4) BJV - Testing Problem}, we
establish the minimax-optimal rate of convergence of the tests suggested
{[}see Ch. 2 in \citet{ingster/suslina:03} for an introduction{]}.
A conventional definition is the following. For a nonrandomized test
$\psi$ that maps a sample $\left\{ X_{t,T}\right\} $ to zero or
one, we consider the maximal type I error
\begin{align*}
\alpha_{\psi}\left(\theta\right) & =\sup_{\left\{ f\left(u,\,\omega\right)\right\} _{u\in\left[0,\,1\right],\,\omega\in\left[-\pi,\,\pi\right]}\in\boldsymbol{F}\left(\theta,\,D\right)}\mathbb{P}_{f}\left(\psi=1\right),
\end{align*}
and the maximal type II error
\begin{align*}
\beta_{\psi}\left(\theta,\,b_{T}\right) & =\sup_{\lambda_{b}^{0}\in\left(0,\,1\right),\,\omega_{0}\in\left[-\pi,\,\pi\right]}\sup_{\left\{ f\left(u,\,\omega\right)\right\} _{u\in\left[0,\,1\right],\,\omega\in\left[-\pi,\,\pi\right]}\in\boldsymbol{F}_{1,\lambda_{b}^{0},\omega_{0}}\left(\theta,\,b_{T},\,D\right)}\mathbb{P}_{f}\left(\psi=0\right),
\end{align*}
and define the total testing error as $\gamma_{\psi}(\theta,\,b_{T})=\alpha_{\psi}(\theta)+\beta_{\psi}(\theta,\,b_{T}).$
The notion of asymptotic minimax-optimality is as follows. We want
to find sequences of tests and rates $b_{T}$ such that $\gamma_{\psi}(\theta,\,b_{T})\rightarrow0$
as $T\rightarrow\infty$. The larger is $b_{T}$ the easier it is
to distinguish between $\mathcal{H}_{0}$ and $\mathcal{H}_{1}^{\mathrm{B}}$
but we may incur at the same time a larger type II error $\beta_{\psi}\left(\theta,\,b_{T}\right)$.
The optimal value $b_{T}^{\mathrm{opt}}$, named the minimax distinguishable
rate, is the minimum value of $b_{T}>0$ such that $\lim_{T\rightarrow\infty}\inf_{\psi}\gamma_{\psi}\left(\theta,\,b_{T}\right)=0$.
A sequence of tests $\psi_{T}$ that satisfies the latter relation
for all $b_{T}\geq b_{T}^{\mathrm{opt}}$ is called minimax-optimal.
Minimax-optimality has been considered in other change-point problems.
\citet{loader:96} and \citet{spokoiny:98} considered the nonparametric
estimation of a regression function with break size fixed. \citet{bibinger/jirak/vetter:16}
considered breaks in the volatility of semimartingales under high-frequency
asymptotics while we focus on breaks in the spectral density and thus
we work in the frequency-domain. Another difference from previous
work is that we do not deal with i.i.d. observations; hence, we cannot
use the same approach to derive the minimax bound as in \citet{bibinger/jirak/vetter:16}
because their information-theoretic reductions exploit independence.
We need to rely on approximation theorems {[}cf. \citet{berkes/philipp:79}{]}
to establish that our statistical experiment is asymptotically equivalent
in a strong Le Cam sense to a high dimensional signal detection problem.
This allows us to derive the minimax bound using classical arguments
based on the results in \citet{ingster/suslina:03}, Ch. 8. The relevant
results are stated in Section \ref{Section Consistency-and-Minimax}.
\subsubsection*{Testing Problem for Case (ii)}
We consider alternative hypotheses where $f$ is less smooth than
under $\mathcal{H}_{0}$, including the case of breaks as a special
case. Suppose that under $\mathcal{H}_{0}$ the spectrum $f\left(u,\,\omega\right)$
is differentiable in both arguments and behaves until time $T\lambda_{b}^{0}$
as specified in $\boldsymbol{F}\left(\theta,\,D\right)$ for some
$\theta>0$ and $D<\infty$. After $T\lambda_{b}^{0}$, the regularity
exponent $\theta$ drops to some $\theta'$ with $0<\theta'<\theta$
for some non-trivial period of time. That is, since $\boldsymbol{F}\left(\theta,\,D\right)\subset\boldsymbol{F}\left(\theta',\,D\right)$,
we need that $f$ behaves as $\theta'$-regular for some period of
time such that there exists a $\omega$ with $\left\{ f\left(u,\,\omega\right)\right\} _{u\in\left[0,\,1\right]}\notin\boldsymbol{F}\left(\theta,\,D\right).$
This guarantees that $\mathcal{H}_{0}$ and $\mathcal{H}_{1}^{\mathrm{S}}$
(to be defined below) are well-separated. To this end, define for
some function $g_{u}$ with $u\in\left[0,\,1\right]$, $\Delta_{h}^{\theta'}g_{u}=\left(g_{u+h}-g_{u}\right)/\left|h\right|^{\theta'}$
for $h\in\left[-u,\,1-u\right].$ The set of possible alternatives
is then defined as
\begin{align*}
\boldsymbol{F}'_{1,\lambda_{b}^{0},\omega_{0}}\left(\theta,\,\theta',\,b_{T},\,D\right) & =\biggl\{\left\{ f\left(u,\,\omega\right)\right\} _{u\in\left[0,\,1\right],\,\omega\in\left[-\pi,\,\pi\right]}\in\boldsymbol{F}\left(\theta',\,D\right):\\
& \quad\quad\,\inf_{\left|h\right|\leq2m_{T}/T}\Delta_{h}^{\theta'}f\left(\lambda_{b}^{0},\,\omega_{0}\right)\geq b_{T}\quad\mathrm{or}\quad\sup_{\left|h\right|\leq2m_{T}/T}\Delta_{h}^{\theta'}f\left(\lambda_{b}^{0},\,\omega_{0}\right)\leq-b_{T}\biggr\},
\end{align*}
where $m_{T}\rightarrow\infty$ as $T\rightarrow\infty$ such that
$m_{T}^{-1}T^{\epsilon}\rightarrow0$ with $\epsilon>0.$ This leads
to the following testing problem,
\begin{align}
\mathcal{H}_{0} & :\,\left\{ f\left(u,\,\omega\right)\right\} _{u\in\left[0,\,1\right]}\in\boldsymbol{F}\left(\theta,\,D\right)\label{eq (29) BJV - Testing Problem}\\
\mathcal{H}_{1}^{\mathrm{S}} & :\,\exists\lambda_{b}^{0}\in\left(0,\,1\right)\,\mathrm{and}\,\omega_{0}\in\left[-\pi,\,\pi\right]\,\mathrm{with}\,\left\{ \left\{ f\left(u,\,\omega\right)\right\} _{u\in\left[0,\,1\right],\,\omega\in\left[-\pi,\,\pi\right]}\right\} \in\boldsymbol{F}'_{1,\lambda_{b}^{0},\omega_{0}}\left(\theta,\,\theta',\,b_{T},\,D\right).\nonumber
\end{align}
Note that $\mathcal{H}_{1}^{\mathrm{S}}$ allows for multiple changes.
$\mathcal{H}_{1}^{\mathrm{B}}$ is a special case of $\mathcal{H}_{1}^{\mathrm{S}}$
since it can be seen as the limiting case of $\mathcal{H}_{1}^{\mathrm{S}}$
as $\theta'\rightarrow0.$ In the context of infinite-dimensional
parameter problems one faces the issue of distinguishability between
the null and the alternative hypotheses. It is evident that one cannot
test $f\in\boldsymbol{F}\left(\theta,\,D\right)$ versus $f\in\boldsymbol{F}\left(\theta',\,D\right)$
for $\theta>\theta'$. First, since $\boldsymbol{F}\left(\theta,\,D\right)\subset\boldsymbol{F}\left(\theta',\,D\right)$,
one has at least to remove the set of functions in $\boldsymbol{F}\left(\theta,\,D\right)$
from those in $\boldsymbol{F}\left(\theta',\,D\right)$. Still, as
discussed by \citet{ingster/suslina:03}, this would not be enough
since the two hypotheses are still too close. That explains why we
focus on spectral densities $f$ that belong to $\boldsymbol{F}'_{1,\lambda_{b}^{0},\omega_{0}}\left(\theta,\,\theta',\,b_{T},\,D\right)$
under $\mathcal{H}_{1}^{\mathrm{S}}$. These are rough enough so
as not to be close to functions in $\boldsymbol{F}\left(\theta,\,D\right)$.
This is captured by the requirement that the difference quotient $\Delta_{h}^{\theta'}f$
exceeds the so-called rate $b_{T}$. As $T\rightarrow\infty$ the
requirement becomes less stringent since $b_{T}\rightarrow0$. See
\citet{hoffman/nick:11} and \citet{bibinger/jirak/vetter:16} for
similar discussions in different contexts.
\section{\label{Section Change-Point-Tests}Tests for Changes in the Spectrum
and Their Limiting Distributions }
Section \ref{subsec: Test-Statistics} introduces the test statistics
while Section \ref{subsec:The-Limiting-Distribution} presents the
results concerning their asymptotic distributions under the null hypothesis.
These results apply also to the case of smooth alternatives which
we discuss formally in Section \ref{Section Consistency-and-Minimax}.
\subsection{\label{subsec: Test-Statistics}The Test Statistics}
We first define the quantities needed to define the tests. Let $h:\,\mathbb{R}\rightarrow\mathbb{R}$
be a data taper with $h\left(x\right)=0$ for $x\notin[0,\,1)$,
\[
H_{k,T}\left(\omega\right)=\sum_{s=0}^{T-1}h\left(s/T\right)^{k}\exp\left(-i\omega s\right),
\]
and (for $n_{T}$ even),
\begin{align*}
d_{L,h,T}\left(u,\,\omega\right)\triangleq\sum_{s=0}^{n_{T}-1}h\left(\frac{s}{n_{T}}\right)X_{\left\lfloor Tu\right\rfloor -n_{T}+s+1,T}\exp\left(-i\omega s\right), & \quad I_{L,h,T}\left(u,\,\omega\right)\triangleq\frac{1}{2\pi H_{2,n_{T}}\left(0\right)}\left|d_{L,h,T}\left(u,\,\omega\right)\right|^{2},\\
d_{R,h,T}\left(u,\,\omega\right)\triangleq\sum_{s=0}^{n_{T}-1}h\left(\frac{s}{n_{T}}\right)X_{\left\lfloor Tu\right\rfloor +n_{T}-s,T}\exp\left(-i\omega s\right), & \quad I_{R,h,T}\left(u,\,\omega\right)\triangleq\frac{1}{2\pi H_{2,n_{T}}\left(0\right)}\left|d_{R,h,T}\left(u,\,\omega\right)\right|^{2},
\end{align*}
where $I_{L,h,T}\left(u,\,\omega\right)$ (resp., $I_{R,h,T}\left(u,\,\omega\right)$)
is the local periodogram over a segment of length $n_{T}\rightarrow\infty$
that uses observations to the left (resp. right) of $\left\lfloor Tu\right\rfloor $.
$I_{L,h,T}\left(u,\,\omega\right)$ (resp., $I_{R,h,T}\left(u,\,\omega\right)$)
is a near unbiased estimator of $f\left(u-n_{T}/T,\,\omega\right)$
(resp., $f\left(u+n_{T}/T,\,\omega\right)$). We allow for a data
taper since one may want to put more weight on observations that are
closer to $u$. The smoothed local periodogram is defined as
\begin{align*}
f_{L,h,T}\left(u,\,\omega\right) & =\frac{2\pi}{n_{T}}\sum_{s=1}^{n_{T}-1}W_{T}\left(\omega-\frac{2\pi s}{n_{T}}\right)I_{L,h,T}\left(u,\,\frac{2\pi s}{n_{T}}\right),
\end{align*}
with $f_{R,h,T}\left(u,\,\omega\right)$ defined similarly to $f_{L,h,T}\left(u,\,\omega\right)$
but with $I_{R,h,T}\left(u,\,\omega\right)$ in place of $I_{L,h,T}\left(u,\,\omega\right)$,
where $W_{T}\left(\omega\right)$ $(-\infty<\omega<\infty)$ is a
family of weight functions of period $2\pi$,
\begin{align*}
W_{T}\left(\omega\right) & =\sum_{j=-\infty}^{\infty}b_{W,T}^{-1}W\left(b_{W,T}^{-1}\left(\omega+2\pi j\right)\right),
\end{align*}
with $b_{W,T}$ a bandwidth and $W\left(\beta\right)$ $(-\infty<\beta<\infty)$
a fixed function. We define
\begin{align*}
\widetilde{f}_{L,r,T}\left(\omega\right)=M_{S,T}^{-1}\sum_{j\in\mathbf{S}_{r}}f_{L,h,T}\left(j/T,\,\omega\right) & \quad\mathrm{and}\quad\widetilde{f}_{R,r,T}\left(\omega\right)=M_{S,T}^{-1}\sum_{j\in\mathbf{S}_{r}}f_{R,h,T}\left(j/T,\,\omega\right),
\end{align*}
where
\begin{align*}
\mathbf{S}_{r} & =\{rm_{T}-m_{T}/2+\left\lfloor n_{T}/2\right\rfloor +1,\,rm_{T}-m_{T}/2+\left\lfloor n_{T}/2\right\rfloor +1+m_{S,T},\\
& \qquad\ldots,\,rm_{T}+\left\lfloor n_{T}/2\right\rfloor +1+m_{S,T}M_{S,T}/2\},
\end{align*}
with $m_{S,T}=\left\lfloor m_{T}^{1/2}\right\rfloor $ and $M_{S,T}=\left\lfloor m_{T}/m_{S,T}\right\rfloor $.
$\widetilde{f}_{a,r,T}\left(\omega\right)$ ($a=L,\,R$) denotes the
average local spectral density around time $rm_{T}$ computed using
$f_{a,h,T}\left(j/T,\,\omega\right)$ where $r=1,\ldots,\,M_{T}=\left\lfloor T/m_{T}\right\rfloor -1$.
We do not use all the $m_{T}$ local spectral densities $f_{a,\,h,T}\left(j/T,\,\omega\right)$
($a=L,\,R$) in the block $r$ but only those separated by $m_{S,T}$
points. Thus, $\mathbf{S}_{r}$ is a subset of the indices in the
block $r$. We need to consider a sub-sample of the $f_{a,h,T}\left(j/T,\,\omega\right)$'s
($a=L,\,R$) because there is strong dependence among the adjacent
terms, e.g., $f_{a,h,T}\left(j/T,\,\omega\right)$ and $f_{a,h,T}\left((j+1)/T,\,\omega\right)$
($a=L,\,R$). A large deviation between $\widetilde{f}_{L,r,T}\left(\omega\right)$
and $\widetilde{f}_{R,r+1,T}\left(\omega\right)$ suggests the presence
of a break in the spectrum close to time $\left(r+1\right)m_{T}$
at frequency $\omega$. Note that the latest observation used in the
construction of $\widetilde{f}_{L,r,T}$ is $X_{rm_{T}+m_{T}/2+\left\lfloor n_{T}/2\right\rfloor +1}$
while the earliest observation used in the construction of $\widetilde{f}_{R,r+1,T}$
is $X_{rm_{T}+m_{T}/2+\left\lfloor n_{T}/2\right\rfloor +2}$. This
shows that there is no overlapping in the time points used in the
construction of $\widetilde{f}_{L,r,T}$ and $\widetilde{f}_{R,r+1,T}$.
The reason behind this is that in order to maximize power $\widetilde{f}_{L,r,T}$
and $\widetilde{f}_{R,r+1,T}$ should not use common observations,
otherwise the effect of the common observations to the difference
in the averages would offset the effect of the change-point.
Let
\begin{align*}
\mathbf{S}_{r,+}\left(j\right) & =\{\{rm_{T}-m_{T}/2+\left\lfloor n_{T}/2\right\rfloor +1,\,rm_{T}-m_{T}/2+\left\lfloor n_{T}/2\right\rfloor +1+\widetilde{m}_{S,T},\\
& \qquad\ldots,\,rm_{T}+\left\lfloor n_{T}/2\right\rfloor +1+\widetilde{m}_{S,T}\widetilde{M}_{S,T}/2\}/\,\{\ldots,\,rm_{T}-m_{T}/2+1+\widetilde{m}_{S,T}\left(j-1\right)\}\},\\
\mathbf{S}_{r,-}\left(j\right) & =\{\{rm_{T}-m_{T}/2+\left\lfloor n_{T}/2\right\rfloor +1,\,rm_{T}-m_{T}/2+\left\lfloor n_{T}/2\right\rfloor +1+\widetilde{m}_{S,T},\\
& \qquad\ldots,\,rm_{T}+\left\lfloor n_{T}/2\right\rfloor +1+\widetilde{m}_{S,T}\widetilde{M}_{S,T}/2\}/\,\{\ldots,\,rm_{T}-m_{T}/2+1+\widetilde{m}_{S,T}\left(-j+1\right)\}\},
\end{align*}
where $\widetilde{m}_{S,T}=\left\lfloor m_{T}^{1/3}\right\rfloor $
and $\widetilde{M}_{S,T}=\left\lfloor m_{T}/\widetilde{m}_{S,T}\right\rfloor $.
Define $\widehat{\sigma}_{L,r}^{2}\left(\omega\right)=\sum_{j=-\widetilde{M}_{S,T}+1}^{\widetilde{M}_{S,T}-1}K_{1}\left(b_{1,T}j\right)\widehat{\Gamma}_{r}\left(j\right)$
where
\begin{align*}
\widehat{\Gamma}_{r}\left(j\right) & =\begin{cases}
\widetilde{M}_{S,T}^{-1}\sum_{t\in\mathbf{S}_{r,+}\left(j\right)}\widehat{f}_{L,h,T}\left(t/T,\,\omega\right)\widehat{f}_{L,h,T}\left(\left(t-j\widetilde{m}_{S,T}\right)/T,\,\omega\right), & j\geq0\\
\widetilde{M}_{S,T}^{-1}\sum_{t\in\mathbf{S}_{r,-}\left(j\right)}\widehat{f}_{L,h,T}\left(t/T,\,\omega\right)\widehat{f}_{L,h,T}\left(\left(t+j\widetilde{m}_{S,T}\right)/T,\,\omega\right), & j<0
\end{cases},
\end{align*}
and $\widehat{f}_{L,h,T}\left(j/T,\,\omega\right)=f_{L,h,T}\left(j/T,\,\omega\right)-\widetilde{f}_{L,r,T}\left(\omega\right)$
for $j\in\mathbf{S}_{r}$. The quantity $\widehat{\sigma}_{L,r}^{2}\left(\omega\right)$
is a local long-run variance estimator where $K_{1}$ is a kernel
and $b_{1,T}$ is the associated bandwidth.
We first present a test statistic for the detection of a change-point
in the spectrum $f\left(\cdot,\,\omega\right)$ for a given frequency
$\omega.$ A second test statistic that we consider detects change-points
in $u\in(0,\,1)$ occurring at any frequency $\omega\in\left[-\pi,\,\pi\right]$.
The latter is arguably more useful in practice because often the practitioner
does not know a priori at which frequency the spectrum is discontinuous.
We begin with the following test statistic,
\begin{align}
\mathrm{S}_{\max,T}\left(\omega\right)\triangleq\max_{r=1,\ldots,\,M_{T}-2}\left|\frac{\widetilde{f}_{L,r,T}\left(\omega\right)-\widetilde{f}_{R,r+1,T}\left(\omega\right)}{\widehat{\sigma}_{L,r}\left(\omega\right)}\right| & ,\qquad\omega\in\left[-\pi,\,\pi\right].\label{Eq. (14) Smax Stat}
\end{align}
Test statistics of the form of \eqref{Eq. (14) Smax Stat} were also
used in the time-domain in the context of nonparametric change-point
analysis under a less general framework {[}cf. \citet{eichinger/kirch:2018},
\citet{bibinger/jirak/vetter:16} and \citet{wu/zhao:07}{]} and forecasting
{[}cf. \citet{casini_CR_Test_Inst_Forecast}{]}. The derivation of
the null distribution uses a (strong) invariance principle for nonstationary
processes {[}see, e.g., \nocite{wu:07} and \citet{wu/zhou:11}{]}.
The test statistic $\mathrm{S}_{\mathrm{max},T}\left(\omega\right)$
aims at detecting a break in the spectrum at some given frequency
$\omega$. An alternative would be to consider a double-sup statistic
which takes the maximum over $\omega\in\left[-\pi,\,\pi\right]$.
Theorem \ref{Theorem 5.2.8 Brillinger} in the supplement shows that
$I_{h,L,T}\left(u,\,\omega_{j}\right)$ and $I_{h,L,T}\left(u,\,\omega_{k}\right)$
are asymptotically independent if $2\omega_{j},\,\omega_{j}\pm\omega_{k}\not\equiv0\,(\mathrm{mod\,}2\pi)$.\footnote{The notation $2\omega_{j},\,\omega_{j}\pm\omega_{k}\not\equiv0\,(\mathrm{mod\,}2\pi)$
means $2\omega_{j}\not\equiv0\,(\mathrm{mod\,}2\pi)$ and $\omega_{j}\pm\omega_{k}\not\equiv0\,(\mathrm{mod\,}2\pi).$ } However, the smoothing over frequencies introduces short-range dependence
over $\omega$. Due to this short-range dependence, we cannot consider
the maximum over all frequencies in $\Pi$ because the statistics
would not be independent. Thus, we specify a framework based on an
infill procedure over the frequency-domain $\left[-\pi,\,\pi\right]$
by assuming that there are $n_{\omega}$ frequencies $\omega_{1},\ldots,\,\omega_{n_{\omega}},$
with $\omega_{1}=-\pi$ and $\omega_{n_{\omega}}=\pi-\epsilon,\,\epsilon>0$,
and $\left|\omega_{j}-\omega_{j+1}\right|=O\left(n_{\omega}^{-1}\right)$
for $j=1,\ldots,\,n_{\omega}-2$. Assume that $n_{\omega}\rightarrow\infty$
as $T\rightarrow\infty$. Let $\Pi\triangleq\left\{ \omega_{1},\ldots,\,\omega_{n_{\omega}}\right\} $.
The maximum is taken over the following set of frequencies
\[
\Pi'\triangleq\{\omega_{1},\,\omega_{2+\left\lfloor n_{T}b_{W,T}\right\rfloor },\ldots,\omega_{n_{\omega}-\left\lfloor n_{T}b_{W,T}\right\rfloor -1},\,\omega_{n_{\omega}}\}.
\]
Let $n'_{\omega}=\left\lfloor n_{\omega}/\left(\left\lfloor n_{T}b_{W,T}\right\rfloor +1\right)\right\rfloor $
with $n'_{\omega}\rightarrow\infty$. Note that $\Pi'\subset\Pi$.
This then leads to the double-sup statistic,
\begin{equation}
\mathrm{S}_{\mathrm{Dmax},T}\triangleq\max_{\omega_{k}\in\Pi'}\sqrt{\log\left(M_{T}\right)}(M_{S,T}^{1/2}\mathrm{S}_{\mathrm{max},T}\left(\omega_{k}\right)-\gamma_{M_{T}})-\log\left(n'_{\omega}\right),\label{Eq. (SDmax)}
\end{equation}
where $\gamma_{M_{T}}=\left[4\log\left(M_{T}\right)-2\log\left(\log\left(M_{T}\right)\right)\right]^{1/2}$.
This double-sup form is a new feature for change-point testing under
the frequency-domain.
Next, we consider alternative test statistics that are self-normalized
such that one does not need to estimate $\sigma_{L,r}^{2}\left(\omega\right)$.
We consider the following test statistic,
\begin{align}
\mathrm{R}_{\max,T}\left(\omega\right)\triangleq\max_{r=1,\ldots,\,M_{T}-2}\left|\frac{\widetilde{f}_{L,r,T}\left(\omega\right)}{\widetilde{f}_{R,r+1,T}\left(\omega\right)}-1\right| & ,\label{Eq. (14) Smax Stat-1}
\end{align}
where $\omega\in\left[-\pi,\,\pi\right].$ We can define a test statistic
corresponding to $\mathrm{S}_{\mathrm{Dmax},T}$ by
\[
\mathrm{R}_{\mathrm{Dmax},T}\triangleq\max_{\omega_{k}\in\Pi'}\sqrt{\log\left(M_{T}\right)}(M_{S,T}^{1/2}\mathrm{R}_{\mathrm{max},T}\left(\omega_{k}\right)-\gamma_{M_{T}})-\log\left(n'_{\omega}\right).
\]
\subsection{\label{subsec:The-Limiting-Distribution}The Limiting Distribution
Under the Null Hypothesis}
Let $\mathbf{X}_{t,T}=(X_{t,T}^{\left(a_{1}\right)},\ldots,\,X_{t,T}^{\left(a_{p}\right)})$
with finite $p\geq1$. Denote by $\kappa_{\mathbf{X},t}^{\left(a_{1},\ldots,a_{r}\right)}\left(k_{1},\ldots,\,k_{r-1}\right)$
the time-$t$ cumulant of order $r$ of $(X_{t+k_{1},T}^{\left(a_{1}\right)},\ldots,X_{t+k_{r-1},T}^{\left(a_{r-1}\right)},\,X_{t,T}^{\left(a_{r}\right)})$
with $r\leq p$.
\begin{assumption}
\label{Assumption Locally Stationary}(i) $\left\{ \mathbf{X}_{t,T}\right\} $
is a mean-zero locally stationary process (i.e., $m_{0}=0$); (ii)
for all $j=1,\ldots,\,p$, $A^{\left(a_{j}\right)}\left(u,\,\omega\right)$
is \textcolor{red}{ } $2\pi$-periodic in $\omega$ and the periodic
extensions are differentiable in $u$ and $\omega$ with uniformly
bounded derivative $\left(\partial/\partial u\right)\left(\partial/\partial\omega\right)A\left(u,\,\omega\right)$;
(iii) $g_{4}$ is continuous.
\end{assumption}
Assumption \ref{Assumption Locally Stationary} requires $\left\{ \mathbf{X}_{t,T}\right\} $
to be locally stationary. Without loss of generality, we assume that
$\left\{ \mathbf{X}_{t,T}\right\} $ has zero mean. All results go
through when the mean is non-zero or when using demeaned series. The
differentiability of $A\left(u,\,\omega\right)$ implies that $f\left(u,\,\omega\right)$
is also differentiable. This means that under the null hypothesis
we require $f\left(u,\,\omega\right)$ to be differentiable in $u$
and to have some regularity exponent $\theta>0.$ The differentiability
of $A\left(u,\,\omega\right)$ in $u$ can be relaxed at the expense
of more complex proofs to establish the results in the supplement
on high-order cumulants. Without differentiability, for any $\theta>0$
the test statistics above follow the same asymptotic distribution
as when differentiability holds, though we do not discuss this case
formally.
We need to impose some conditions on the temporal dependence. Let
$\left\{ e_{t}\right\} _{t\in\mathbb{Z}}$ be a sequence of i.i.d.
random variables and $\left\{ e'_{t}\right\} _{t\in\mathbb{Z}}$ be
an independent copy of $\left\{ e_{t}\right\} _{t\in\mathbb{Z}}.$
Assume $X_{t,T}=H_{T}\left(t/T,\,\mathscr{F}_{t}\right)$ where $\mathscr{F}_{t}\triangleq\left\{ \ldots,\,e_{t-1},\,e_{t}\right\} $
and $H_{T}:\,\left[0,\,1\right]\times\mathbb{R}^{\infty}\mapsto\mathbb{R}$
is a measurable function. We use the dependence measure introduced
by \citeauthor{wu:05} (\citeyear{wu:05}, \citeyear{wu:07}) for
stationary processes and extended to nonstationary processes by \citet{wu/zhou:11}.
Let $\mathscr{L}^{q}$ denote the space generated by the $q$-norm,
$q>0$. For all $t$, assume $X_{t,T}\in\mathscr{L}^{q}$. For $w\geq0$
define the dependence measure,
\begin{align}
\phi_{w,q} & =\sup_{t}\left\Vert X_{t,T}-X_{t,T,\left\{ w\right\} }\right\Vert _{q}\label{Eq. (2.1) WZ Delta}\\
& =\sup_{t}\left\Vert H_{T}\left(t/T,\,\mathscr{F}_{t}\right)-H_{T}\left(t/T,\,\mathscr{F}_{t,\left\{ w\right\} }\right)\right\Vert _{q},\nonumber
\end{align}
where $\mathscr{F}_{t,\left\{ w\right\} }$ is a coupled version of
$\mathscr{F}_{t}$ with $e_{w}$ replaced by an i.i.d. copy $e'_{w}$.
For $\{X_{t,T}\}$ locally stationary, $H_{T}\left(\cdot,\,\mathscr{F}_{t}\right)$
is stochastic $\theta$-H\"older continuous in the sense that there
exists $C_{T,H}<\infty$ such that
\begin{align}
\sup_{0\leq u<u'\leq1} & \frac{\left\Vert H_{T}\left(u,\,\mathscr{F}_{t}\right)-H_{T}\left(u',\,\mathscr{F}_{t}\right)\right\Vert }{|u-u'|^{\theta}}\leq C_{T,H}.\label{Eq. Stochastic Lip cont}
\end{align}
Assume $\Upsilon_{n,q}=\sum_{j=n}^{\infty}\phi_{j,q}<\infty$ for
some $n\in\mathbb{Z}$. Let $\tau_{T}=T^{\vartheta_{1}}\left(\log\left(T\right)\right)^{\vartheta_{2}}$
where $\vartheta_{1}=\left(1/2-1/q+\gamma/q\right)$ $/\left(1/2-1/q+\gamma\right)$
and $\vartheta_{2}=\left(\gamma+\gamma/q\right)/\left(1/2-1/q+\gamma\right)$
for some $\gamma>0$.
\begin{assumption}
\label{Assumption WZ 2011}For $q\geq2$, $\sum_{n=0}^{\infty}n^{l+q-1}(\sum_{j=n}^{\infty}\phi_{j,q}^{2})^{1/2}<\infty$
where $l\geq0$.
\end{assumption}
Assumption \ref{Assumption WZ 2011} was also used by \citet{shao/wu:2007ET}
who showed that it is satisfied for many nonlinear time series processes.
Using Assumption \ref{Assumption WZ 2011} we can give a sufficient
condition for the summability of the joint cumulant up to a certain
order. The latter is a common assumption in spectral analysis and
we use it to establish results on the high-order cumulants and spectra
in Section \ref{Section Cumulant and Spectra}. These are used to
obtain the null limiting distribution of the test statistics.
\begin{lem}
\label{Lemma Theorem 4.1 in Shao and Wu (2007)}Let Assumption \ref{Assumption WZ 2011}
hold and $X_{t,T}\in\mathscr{L}^{r}$, $r\geq2$. We have
\begin{align}
\sum_{k_{1},\ldots,\,k_{r-1}=-\infty}^{\infty}\left(1+\left|k_{j}\right|^{l}\right)\sup_{1\leq t\leq T}\left|\kappa_{\mathbf{X},t}^{\left(a_{1},\ldots,a_{r}\right)}\left(k_{1},\ldots,\,k_{r-1}\right)\right| & <\infty,\label{Eq. (2.6.6) Condition on cumulant}
\end{align}
for $l\geq0$, $j=1,\ldots,\,r-1$, and any $r$-tuple $a_{1},\ldots,a_{r}$.
\end{lem}
\citet{shao/wu:2007ET} proved \eqref{Eq. (2.6.6) Condition on cumulant}
for $l=0$ and $\{X_{t,T}\}$ a stationary process.
\begin{assumption}
\label{Taper, nT, W and K1, b1}(i) The data taper $h:\,\mathbb{R}\rightarrow\mathbb{R}$
with $h\left(x\right)=0$ for $x\notin[0,\,1)$ is bounded and of
bounded variation; (ii) The sequence $\left\{ n_{T}\right\} $ satisfies
$n_{T}\rightarrow\infty$ as $T\rightarrow\infty$ with $n_{T}/T\rightarrow0$;
(iii) $W\left(\beta\right)$ $(-\infty<\beta<\infty)$ is real-valued,
even, of bounded variation, and satisfies $\int_{-\infty}^{\infty}W\left(\beta\right)d\beta=1$;
(iv) $b_{1,T}\rightarrow0$ such that $Tb_{1,T}\rightarrow\infty$
and $(\widetilde{M}_{S,T}b_{1,T})^{-1/2}M_{S,T}^{1/2}\log T\rightarrow0$
and $K_{1}\left(\cdot\right):\,\mathbb{R}\rightarrow\left[-1,\,1\right],$
$K_{1}\left(0\right)=1,\,K_{1}\left(x\right)=K_{1}\left(-x\right),\,\forall x\in\mathbb{R}$
and ${\textstyle \int\nolimits _{-\infty}^{\infty}}K_{1}^{2}\left(x\right)dx<\infty$.
\end{assumption}
Assumption \ref{Taper, nT, W and K1, b1}-(i,ii) are standard in the
nonparametric estimation literature�while Assumption \ref{Taper, nT, W and K1, b1}-(iii)
is also used for spectral density estimation under stationarity {[}e.g.,
\citet{brillinger:75}{]}. The conditions on $b_{1,T}$ in Assumption
\ref{Taper, nT, W and K1, b1}-(iv) are necessary for the consistency
of the long-run variance estimator. The class of kernels allowed
by Assumption \ref{Taper, nT, W and K1, b1}-(iv) includes popular
kernels such as the Truncated, Bartlett, Parzen, Quadratic Spectral
(QS) and Tukey-Hanning kernels. For technical reasons inherent to
the proofs, we need to assume that the spectral density is strictly
positive. Theorem \ref{Theorem 5.6.4 Brillinger} in the supplement
shows that the variance of $f_{L,h,T}\left(u,\,\omega\right)$ depends
on $f\left(u,\,\omega\right)$. Thus, the denominator of the test
statistic depends on $f\left(u,\,\omega\right)$. Assumption \ref{Assumption Minimum f}
requires the latter to be bounded away from zero. In practice, if
one suspects that at some frequencies $f\left(u,\,\omega\right)$
can be close to zero, then one can add a small number $\epsilon_{f}>0$
to the denominator of the test statistic to guarantee numerical stability.
\begin{assumption}
\label{Assumption Minimum f}$f_{-}=\min_{u\in\left[0,\,1\right],\,\omega\in\left[-\pi,\,\pi\right]}f\left(u,\,\omega\right)>0$.
\end{assumption}
The next assumption ensures that the local spectral density estimates
are asymptotically independent when evaluated at some given frequencies
(see Theorem \ref{Theorem 5.6.4 Brillinger}). It is used to derive
the asymptotic null distribution of the double-sup test statistics
$\mathrm{S}_{\mathrm{Dmax},T}$ and $\mathrm{R}_{\mathrm{Dmax},T}$.
\begin{assumption}
\label{Assumption Big PI}Assume that $2\omega_{j},\,\omega_{j}\pm\omega_{k}\not\equiv0\,(\mathrm{mod\,}2\pi)$
for $\omega_{j},\,\omega_{k}\in\Pi$.
\end{assumption}
\begin{condition}
\label{Condition n_T h_T}(i) The sequence $\left\{ m_{T}\right\} $
satisfies $m_{T}\rightarrow\infty$ as $T\rightarrow\infty,$ and
\begin{align}
M_{S,T}^{1/2} & m_{T}^{\theta}T^{-\theta}\left(\log\left(M_{T}\right)\right)^{1/2}+\tau_{T}^{2}\log\left(M_{T}\right)M_{S,T}^{-1}\label{eq. Condition auxiliary nT}\\
& +M_{S,T}n_{T}^{4}\log\left(M_{T}\right)T^{-4}+M_{S,T}\left(\log\left(n_{T}\right)\right)^{2}\log\left(M_{T}\right)n_{T}^{-2}\rightarrow0;\nonumber
\end{align}
(ii) $b_{W,T}\rightarrow0$ such that $Tb_{W,T}\rightarrow\infty$,
$\log\left(M_{T}\right)M_{S,T}b_{W,T}^{4}\rightarrow0$ and $\log\left(M_{T}\right)M_{S,T}(n_{T}b_{W,T})^{-1}\rightarrow0.$
\end{condition}
Part (i) imposes lower and upper bounds on the growth condition of
the sequence $\left\{ m_{T}\right\} $. The upper bound relates to
the smoothness of $A\left(u,\,\omega\right)$, the value of $\theta$
under the null hypothesis, $n_{T}$ and the number of summands $M_{S,T}$
in $\widetilde{f}_{L,r,T}^{*}\left(\omega\right)$.
Let $\mathscr{V}$ denote a random variable with an extreme value
distribution defined by $\mathbb{P}\left(\mathscr{V}\leq v\right)=\exp($
$-\pi^{-1/2}\exp\left(-v\right)).$
\begin{thm}
\label{Theorem Asymptotic H0 Distrbution Smax}Let Assumption \ref{Assumption Locally Stationary}-\ref{Assumption Minimum f}
and Condition \ref{Condition n_T h_T} hold. Under $\mathcal{H}_{0}$,
$\sqrt{\log\left(M_{T}\right)}(M_{S,T}^{1/2}\mathrm{S}_{\mathrm{max},T}\left(\omega\right)$
$-\gamma_{M_{T}})\Rightarrow\mathscr{V}$ for any $\omega\in\left[-\pi,\,\pi\right]$.
\end{thm}
Theorem \ref{Theorem Asymptotic H0 Distrbution Smax} shows that the
asymptotic null distribution follows an extreme value distribution.
The derivation of the null distribution uses a (strong) invariance
principle for nonstationary processes {[}see, e.g., \nocite{wu:07}
and \citet{wu/zhou:11}{]}. The following theorems shows that the
asymptotic null distribution of the remaining tests $\mathrm{S}_{\mathrm{Dmax},T}$,
$\mathrm{R}_{\mathrm{max},T}\left(\omega\right)$ and $\mathrm{R}_{\mathrm{Dmax},T}$
also follows an extreme value distribution, though the additional
Assumption \ref{Assumption Big PI} and the extra factor $\log\left(n'_{\omega}\right)$
are needed for $\mathrm{S}_{\mathrm{Dmax},T}$ and $\mathrm{R}_{\mathrm{Dmax},T}$.
\begin{thm}
\label{Theorem Asymptotic H0 Distrbution SDmax}Let Assumption \ref{Assumption Locally Stationary}-\ref{Assumption Big PI}
and Condition \ref{Condition n_T h_T} hold. Under $\mathcal{H}_{0}$,
$\mathrm{S}_{\mathrm{Dmax},T}\overset{}{\Rightarrow}\mathscr{V}$.
\end{thm}
\begin{thm}
\label{Theorem Asymptotic H0 Distrbution Rmax and RDmax}Let Assumption
\ref{Assumption Locally Stationary}-\ref{Assumption Minimum f} and
Condition \ref{Condition n_T h_T} hold. Under $\mathcal{H}_{0}$,
$\sqrt{\log\left(M_{T}\right)}(M_{S,T}^{1/2}\mathrm{R}_{\mathrm{max},T}\left(\omega_{k}\right)$
$-\gamma_{M_{T}})\Rightarrow\mathscr{V}$ and, in addition if Assumption
\ref{Assumption Big PI} holds, then $\mathrm{R}_{\mathrm{Dmax},T}\overset{}{\Rightarrow}\mathscr{V}$.
\end{thm}
The limiting distributions in Theorem \ref{Theorem Asymptotic H0 Distrbution Smax}-\ref{Theorem Asymptotic H0 Distrbution Rmax and RDmax}
are pivotal, so that critical values can be obtained immediately
without the need to rely on simulations. The tests have power against
both breaks (alternative hypothesis (i)) and changes in the smoothness
(alternative hypothesis (ii)). Unfortunately, discerning between the
two types of alternative hypotheses is a quite hard technical problem.
Although knowing whether the rejection of the null hypothesis is due
to a break or a change in the smoothness would be useful, in many
cases knowing that there has been some change in the data-generating
process is sufficient to modify the estimation and/or inference strategy
to account for the change. If the change involves a break, a reasonable
approach would be to apply some sample-splitting method for estimation.
For example, in the context of long-run variance estimation the existence
of a break suggests to modify the double kernel HAC (DK-HAC) estimator
to avoid mixing two different regimes {[}cf. \citet{casini_hac}{]}.
On the other hand, for a change in the smoothness a modification of
a standard nonparametric kernel smoothing could be enough. However,
using a sample-splitting technique even when there is a change in
the smoothness would be robust to the change and would result in better
estimation and inference. In general, treating a change as a break
and applying some sample-splitting method would be technically valid
even when the rejection was due to a change in the smoothness (though
it may not be efficient).
The property of being robust to different alternative hypotheses is
not specific to our method. This property is shared by most of the
existing structural break tests. For example, \citet{andrews:93}
showed that structural break tests also have some power against some
forms of smoothly-varying parameters. This property was seen as a
positive feature in the structural break literature. Following the
same reasoning, the property of being robust to breaks as well as
changes in the smoothness can actually be seen as a virtue of our
method.
\section{\label{Section Consistency-and-Minimax}Consistency and Minimax Optimal
Rate of Convergence}
In this section, we discuss the consistency and minimax-optimal lower
bound for the testing problem \eqref{eq (29) BJV - Testing Problem}
(i.e., case (ii)). The discussion also covers the testing problem
(i) since $\mathcal{H}_{1}^{\mathrm{B}}$ can be seen as the limiting
case of $\mathcal{H}_{1}^{\mathrm{S}}$ as $\theta'\rightarrow0.$
We assume that $X_{t,T}$ is segmented locally stationary with transfer
function $A\left(u,\,\omega\right)$ satisfying the following smoothness
properties.
\begin{assumption}
\label{Assumption Smothness of A}(i) $\left\{ X_{t,T}\right\} $
is a mean-zero segmented locally stationary process; (ii) $A\left(u,\,\omega\right)$
is \textcolor{red}{}twice continuously differentiable in $u$ at
all $u\neq\lambda_{j}^{0}$ $(j=1,\ldots,\,m_{0}+1)$ with uniformly
bounded derivatives $\left(\partial/\partial u\right)A\left(u,\,\cdot\right)$
and $\left(\partial^{2}/\partial u^{2}\right)A\left(u,\,\cdot\right)$;
(iii) $A\left(u,\,\omega\right)$ is twice left-differentiable in
$u$ at $u=\lambda_{j}^{0}$ $(j=1,\ldots,\,m_{0}+1)$ with uniformly
bounded derivatives $\left(\partial/\partial_{-}u\right)A\left(u,\,\cdot\right)$
and $\left(\partial^{2}/\partial_{-}u^{2}\right)A\left(u,\,\cdot\right)$.
\end{assumption}
\begin{assumption}
\label{Assumption: (i) smothness in omega; (ii) g4 is continuous}(i)
$A\left(u,\,\omega\right)$ is twice differentiable in $\omega$
with uniformly bounded derivatives $\left(\partial/\partial\omega\right)$
$A\left(\cdot,\,\omega\right)$ and $\left(\partial^{2}/\partial\omega^{2}\right)A\left(\cdot,\,\omega\right)$;
(ii) $g_{4}\left(\omega_{1},\,\omega_{2},\,\omega_{3}\right)$ is
continuous in its arguments.
\end{assumption}
We now move to the derivation of the minimax lower bound. As explained
before, we restrict attention to a strictly positive spectral density
in the frequency dimension at which the null hypotheses is violated.
That is, $f_{-}\left(\omega_{0}\right)=\inf_{u\in\left[0,\,1\right]}f\left(u,\,\omega_{0}\right)>0$.
Such restriction is not imposed on $f\left(u,\,\omega\right)$ for
$\omega\neq\omega_{0}$.
\begin{thm}
\label{Theorem 4.1 BJV}Let Assumption \ref{Assumption WZ 2011}-\ref{Taper, nT, W and K1, b1},
\ref{Assumption Smothness of A}-\ref{Assumption: (i) smothness in omega; (ii) g4 is continuous}
and $f_{-}\left(\omega_{0}\right)>0$ hold. Consider either set of
hypotheses $\{\mathcal{H}_{0},\,\mathcal{H}_{1}^{\mathrm{B}}\}$ with
$\theta'=0$ or $\{\mathcal{H}_{0},\,\ensuremath{\mathcal{H}_{1}^{\mathrm{S}}}\}$
with $0<\theta'<\theta$. Then, for
\begin{align*}
b_{T} & \leq\left(T/\log\left(M_{T}\right)\right)^{-\frac{\theta-\theta'}{2\theta+1}}D^{-\frac{2\theta'+1}{2\theta+1}}f_{-}\left(\omega_{0}\right),
\end{align*}
we have $\lim_{T\rightarrow\infty}\inf_{\psi}\gamma_{\psi}\left(\theta,\,b_{T}\right)=1.$
\end{thm}
The theorem implies the need for
\[
b_{T}^{\mathrm{opt}}>\left(T/\log\left(M_{T}\right)\right)^{-\frac{\theta-\theta'}{2\theta+1}}D^{-\frac{2\theta'+1}{2\theta+1}}f_{-}\left(\omega_{0}\right),
\]
otherwise there cannot exist a minimax-optimal test yielding $\lim_{T\rightarrow\infty}\inf_{\psi}\gamma_{\psi}\left(\theta,\,b_{T}\right)=0.$
Note that the lower bound does not depend on $\omega$. In Theorem
\ref{Theorem 4.2 BJV 17} we establish a corresponding upper bound.
From the lower and upper bounds we deduce the optimal rate for the
minimax distinguishable boundary. We can also derive tests based on
$b_{T}^{\mathrm{opt}}$. For example, using the test statistic \eqref{Eq. (14) Smax Stat}
for $\{\mathcal{H}_{0},\,\mathcal{H}_{1}^{\mathrm{B}}\}$ we obtain
the following test $\psi^{*}:$ $\psi^{*}(\left\{ X_{t,T}\right\} )=1$
if $\mathrm{S}_{\mathrm{max},T}\left(\omega\right)\geq2D^{*}\sqrt{\log\left(M_{T}^{*}\right)/m_{T}^{*}}$
for $\omega\in\left[-\pi,\,\pi\right]$ where $D^{*}>2$, $m_{T}^{*}=(\sqrt{\log\left(M_{T}^{*}\right)}T^{\theta}/D)^{\frac{2}{2\theta+1}}$
and $M_{T}^{*}=\left\lfloor T/m_{T}^{*}\right\rfloor $. Hence,
in order to construct such a test we need knowledge of $\theta$ under
$\mathcal{H}_{0}$. We discuss this in Section \ref{Section Implementation}.
Next, we establish the optimal rate for minimax distinguishability.
Note that either alternatives $\mathcal{H}_{1}^{\mathrm{B}}$ or
$\mathcal{H}_{1}^{\mathrm{S}}$ allows for multiple breaks. The
following results require further restrictions on the relation between
$n_{T}$ and $m_{T}$.
\begin{thm}
\label{Theorem 4.2 BJV 17}Let Assumption \ref{Assumption WZ 2011}-\ref{Taper, nT, W and K1, b1},
\ref{Assumption Smothness of A}-\ref{Assumption: (i) smothness in omega; (ii) g4 is continuous}
hold. Consider either alternative hypotheses $\ensuremath{\mathcal{H}_{1}^{\mathrm{B}}}$
with $\theta'=0$ and $\lambda_{j}^{0}<\lambda_{j+1}^{0}$ for $j=1,\ldots,\,m_{0}$,
or $\ensuremath{\mathcal{H}_{1}^{\mathrm{S}}}$ with $0<\theta'<\theta$.
If
\begin{align}
\left(\sqrt{\log\left(M_{T}^{*}\right)/m_{T}^{*}}\right)^{-1}\left(\left(m_{T}^{*}/T\right)^{\theta}+\left(n_{T}/T\right)^{2}+\log\left(n_{T}\right)/n_{T}+b_{W,T}^{2}\right) & \rightarrow0,\label{Eq. Cond for Sec. 5}
\end{align}
and
\begin{align}
b_{T}^{*} & >\left(4D^{*}\sup_{u\in\left[0,\,1\right]}f\left(u,\,\omega_{0}\right)+2\right)^{-\frac{\theta-\theta'}{2\theta+1}}\left(T/\log\left(M_{T}\right)\right)^{-\frac{\theta-\theta'}{2\theta+1}}D^{-\frac{2\theta'+1}{2\theta+1}},\label{Eq. (34) BJV 17}
\end{align}
then $\lim_{T\rightarrow\infty}\gamma_{\psi^{*}}\left(\theta,\,b_{T}^{*}\right)=0$
and $b_{T}^{\mathrm{opt}}\propto\left(T/\log\left(M_{T}\right)\right)^{-\frac{\theta-\theta'}{2\theta+1}}.$
\end{thm}
The theorem shows that a smooth change in the regularity exponent
$\theta$ cannot be distinguished from a break of magnitude smaller
than $b_{T}^{\mathrm{opt}}$ because the change from $\theta$ to
$\theta'$ has to persist for some time. This is also indicated by
the restriction $\theta'>0.$ The minimax bound is similar to the
one established by \citet{bibinger/jirak/vetter:16} for the volatility
of a It\^o semimartingale. The theorem suggests that knowledge
of the frequency $\omega_{0}$ at which the spectrum changes regularity
is irrelevant for the determination of the bound. However, we conjecture
that if the spectrum exhibits a break or smooth change of the form
discussed above simultaneously across multiple frequencies then the
lower bound may be further decreased as one can pool additional information
from inspection of the spectrum for the set of frequencies subject
to the change. The key assumption would be that the change occurs
at the same time $\lambda_{b}^{0}$ for a given set of frequencies
$\omega.$ This may be of interest for economic and financial time
series since they often exhibit a break simultaneously at high and
low frequencies. We leave this to future research.
\section{Estimation of the Change-Points\label{Section Estimation-of-the Breaks}}
We now discuss the estimation of the break locations for the case
of discontinuities in the spectrum (i.e., $\mathcal{H}_{1}^{\mathrm{B},m_{0}}$
where $m_{0}$ is the number of breaks, recall Definition \ref{Definition Segmented-Locally-Stationary}).
The same estimator is valid for the locations of the smooth changes
as under $\mathcal{H}_{1}^{\mathrm{S}}$. For the latter case we later
provide intuitive remarks about the consistency result and the conditions
needed for it. We first consider the case of a single break (i.e.,
$\mathcal{H}_{1}^{\mathrm{B},1}$) and then present the results for
the case of multiple breaks (i.e., $\mathcal{H}_{1}^{\mathrm{B},m_{0}}$).
\subsection{Single Break Alternatives $\mathcal{H}_{1}^{\mathrm{B},1}$ }
Let
\begin{align*}
\mathrm{D}_{r,T}\left(\omega\right)\triangleq M_{S,T}^{-1/2}\left|\sum_{j\in\mathbf{S}_{L,r}}f_{L,h,T}\left(j/T,\,\omega\right)-\sum_{j\in\mathbf{S}_{R,r}}f_{R,h,T}\left(j/T,\,\omega\right)\right|, & \qquad\omega\in\left[-\pi,\,\pi\right].
\end{align*}
where
\begin{align*}
\mathbf{S}_{L,r} & =\{r-m_{T}+1,\,r-m_{T}+1+m_{S,T},\ldots,\,r-m_{T}+1+m_{S,T}M_{S,T}\},\\
\mathbf{S}_{R,r} & =\{r+1,\,r+1+m_{S,T},\ldots,\,r+1+m_{S,T}M_{S,T}\},
\end{align*}
and $r=2m_{T},\,3m_{T},\ldots$ with $r<\left(M_{T}-1\right)m_{T}-n_{T}$.
Note that the maximum of the statistics $\mathrm{D}_{r,T}\left(\omega\right)$
is a version of $\mathrm{S}_{\max,T}$ that does not involve the normalization.
The change-point estimator is defined as
\[
T\widehat{\lambda}_{b,T}=\underset{r=2m_{T},\,3m_{T},\ldots}{\mathrm{argmax}}\max_{\omega\in\left[-\pi,\,\pi\right]}\mathrm{D}_{r,T}\left(\omega\right).
\]
Recall that we consider the following alternative hypothesis:
\begin{align*}
\mathcal{H}_{1}^{\mathrm{B},1}: & \,\left\{ f\left(T_{b}^{0}/T,\,\omega_{0}\right)-\lim_{s\downarrow T_{b}^{0}}f\left(s/T,\,\omega_{0}\right)=\delta_{T}\neq0,\quad\omega_{0}\in\left[-\pi,\,\pi\right]\right\} .
\end{align*}
Note that a break does not need to occur simultaneously at all frequencies
for the procedure to work. The break magnitude can be either fixed
or converge to zero as specified by the following assumption.
\begin{assumption}
\label{Assumption small shifts}$\delta_{T}=\delta\neq0$ is fixed
or $\delta_{T}\rightarrow0$ and $\delta_{T}M_{S,T}^{1/2}/\sqrt{\log\left(T\right)}\rightarrow(0,\,\infty]$.
\end{assumption}
\begin{prop}
\label{Prop 4.5 BJV Consitency Change-point}Let Assumption \ref{Assumption WZ 2011}-\ref{Taper, nT, W and K1, b1}-(i-iii),
\ref{Assumption Smothness of A} with $m_{0}=1$ and Condition \ref{Condition n_T h_T}
hold. Under $\mathcal{H}_{1}^{\mathrm{B},1},$ if $\delta_{T}$ satisfies
Assumption \ref{Assumption small shifts}, we have $\widehat{\lambda}_{b,T}-\lambda_{b}^{0}=O_{\mathbb{P}}(m_{S,T}\sqrt{M_{S,T}\log(T)}/(T\delta_{T}))$.
\end{prop}
We compare the rate of convergence in Proposition \ref{Prop 4.5 BJV Consitency Change-point}
with that of classical change-point estimators in the piecewise constant
mean model. For fixed shifts, the latter rate of convergence is $O_{\mathbb{P}}(T^{-1})$
while for shrinking shifts it is $O_{\mathbb{P}}((T\delta_{T}^{2})^{-1})$
where $\delta_{T}\rightarrow0$ with $\delta_{T}T^{1/2-\vartheta}$
for some $\vartheta\in\left(0,\,1/2\right)$ {[}cf. \citet{yao:87}{]}.\footnote{See also \citet{verzelen/fromont/lerasle/reynaud-bouret:2020} for
recent developments on nimimax optimality for change-point estimation
in the piecewise constant mean model. They considered as a significance
of a change-point a measure that depends on both the break magnitude
and the location of the break. They named it the \textit{energy} of
the change-point. They established the uniform detection threshold
for the energy.} Unlike the classical change-point problem where the mean is piecewise
constant, our problem involves a spectrum that can vary smoothly.
The latter represents a local problem that cannot be addressed by
standard sample-splitting methods. Our method is local in nature and
so it is sub-optimal for the classical change-point problem but it
is valid for the more general case of a piecewise smooth spectrum.
Hence, for fixed shifts, the rate of convergence in our problem is
slower. The smallest break magnitude allowed by Proposition \ref{Prop 4.5 BJV Consitency Change-point}
is $\delta_{T}=O(\sqrt{\log\left(T\right)}/M_{S,T}^{1/2})$. Under
this condition the convergence rate for the classical change-point
estimator is $O_{\mathbb{P}}(M_{S,T}(T\log\left(T\right))^{-1})$
which is faster by a factor $O(m_{S,T}\sqrt{\log\left(T\right)})$
than the one suggested by Proposition \ref{Prop 4.5 BJV Consitency Change-point}.
In addition, in classical change-point setting $\delta_{T}\rightarrow0$
is allowed at a faster rate. This is obvious since in our setting
a small break can be confounded with a smooth local change.
Under the smooth alternative $\mathcal{H}_{1}^{\mathrm{S}}$ the
estimator is consistent when $\theta$-regularity is violated only
once in the sample and also when the violation occurs in a small interval
around $\lambda_{b}^{0}$ which does not exceeds $O(m_{S,T}\sqrt{M_{S,T}\log\left(T\right)}/T\delta_{T})$.
If that interval is longer then this becomes a global problem which
cannot be addressed by the estimation method considered in this section.
This also relates to the discussion in Section \ref{Section Consistency-and-Minimax}
that one cannot perfectly separate functions with $\theta$-smoothness
from functions with $\theta'$-smoothness such that $\theta'<\theta.$
\subsection{Multiple Breaks Alternatives $\mathcal{H}_{1}^{\mathrm{B},m_{0}}$ }
Let us assume that there are $m_{0}>1$ break points in $f\left(u,\,\omega\right)$.
Let $0<\lambda_{1}^{0}<\ldots<\lambda_{m_{0}}^{0}<1$. We consider
the following class of alternative hypotheses:
\begin{align*}
\mathcal{H}_{1}^{\mathrm{B},m_{0}}: & \,\left\{ f\left(T_{l}^{0}/T,\,\omega_{l}\right)-\lim_{s\downarrow T_{l}^{0}}f\left(s/T,\,\omega_{l}\right)=\delta_{l,T}\neq0,\quad\omega_{l}\in\left[-\pi,\,\pi\right]\,\mathrm{for\,}1\leq l\leq m_{0}\right\} .
\end{align*}
We provide a consistency result for both $m_{0}$ and the actual locations
of the breaks $\lambda_{l}^{0}$ $(1\leq l\leq m_{0})$. Let $j^{*}$
be the largest integer such that $\left(M_{T}-j^{*}\right)m_{T}\leq\left(M_{T}-1\right)m_{T}-n_{T}$
and $\mathcal{I}\subseteq\left\{ 2m_{T},\,3m_{T},\ldots,\,\left(M_{T}-j^{*}\right)m_{T}\right\} $
denote a generic index set. One can test for a break at some time
index in $\mathcal{I}$ by using the test $\psi^{*}(\{X_{r}\}_{r\in\mathcal{I}})$
based on $\max_{\omega_{k}\in\Pi'}\mathrm{S}_{\mathrm{max},T}\left(\omega_{k}\right)$
and if the test rejects one can estimate the break location using
\begin{align}
T\widehat{\lambda}_{T}\left(\mathcal{I}\right) & =\underset{r\in\mathcal{I}}{\mathrm{argmax}}\max_{\omega\in\left[-\pi,\,\pi\right]}\mathrm{D}_{r,T}\left(\omega\right).\label{eq. (38) BJV - Obj Func Estimation}
\end{align}
We can then update the set $\mathcal{I}$ by excluding a $v_{T}$-neighborhood
of $T\widehat{\lambda}_{T}$ and repeat the above steps. This is a
sequential top-down algorithm exploiting the classical idea of bisection.
However, this procedure may not be efficient. For example, consider
the first step of the algorithm in which we test for the first break;
this is associated with the largest break magnitude ($\delta_{1,T}>\delta_{l,T}$
for all $l=2,\ldots,\,m_{0}$). If the true break date $T_{1}^{0}$
falls in between two indices in $\mathcal{I}$, say $r_{1}$ and $r_{2}=r_{1}+m_{T}$,
then this does not maximize either power or precision of the location
estimate because one would need to compare two adjacent blocks exactly
separated at $T_{1}^{0}$ but $T_{1}^{0}\notin\mathcal{I}$ since
$T_{1}^{0}\in\left(r_{1},\,r_{2}\right).$ Hence, we introduce a
wild sequential top-down algorithm.
Continuing with the above example, we draw randomly without replacement
$K\geq1$ separation points $r^{\diamond}$ from the interval $\left(r_{1},\,r_{2}\right)$
and for each separation point compute $\mathrm{D}_{r^{\diamond},T}\left(\omega\right)$
where $r^{\diamond}\in\left(r_{1},\,r_{2}\right).$ We take the maximum
value. Then, we update $\mathcal{I}$ by removing $r_{1}$ and adding
$r^{\diamond}$. We repeat this for all indices in $\mathcal{I}$.
Because the $K$ separation points are drawn randomly, there is always
some probability to pick up the separation point that guarantees the
highest power. A natural question is why not take all integers between
$r_{1}$ and $r_{2}$ and compute $\mathrm{D}_{r^{\diamond},T}\left(\omega\right)$
for each. The reason is that in applications involving high frequency
data (e.g., weakly, daily, and so on) that would be highly computationally
intensive especially with multiple breaks as one wishes to change
$m_{T}$ when searching for an additional break. This procedure exploits
idea of bisection and combines it with a wild resampling technique
similar to the one in \citet{fryzlewicz:14}. The latter is characterized
by using binary segmentation and drawing a large number of random
intervals. Here the idea of using draws of random intervals is applied
to the sequential top-down algorithm.
We are now ready to present the algorithm. Guidance as to a suitable
choice of $K$ will be given below. Let $v_{T}\rightarrow\infty$
with $v_{T}/T\rightarrow0$ and $m_{T}/v_{T}\rightarrow0.$ Consider
the test $\psi(\left\{ X_{t,T}\right\} ,\,\mathcal{I})=1$ if $\mathrm{S}_{\mathrm{Dmax},T}\left(\mathcal{I}\right)\geq2D^{*}\sqrt{\log\left(M_{T}^{*}\right)/m_{T}^{*}}$
where
\begin{align*}
\mathrm{S}_{\mathrm{Dmax},T}\left(\mathcal{I}\right)\triangleq\max_{r\in\mathcal{I}}\max_{\omega_{k}\in\Pi'}\left|\frac{\sum_{j\in\mathbf{S}_{L,r}}f_{L,h,T}\left(j/T,\,\omega_{k}\right)-\sum_{j\in\mathbf{S}_{R,r}}f_{R,h,T}\left(j/T,\,\omega_{k}\right)}{\widehat{\sigma}_{L,r}\left(\omega_{k}\right)}\right| & ,
\end{align*}
with $D^{*}$, $m_{T}^{*}$ and $M_{T}^{*}$ as defined in Section
\ref{Section Consistency-and-Minimax}.
\begin{lyxalgorithm}
\label{Algorithm 1}Set $\widehat{\mathcal{I}}=\{2m_{T},\,3m_{T},\ldots,\,\left(M_{T}-j^{*}\right)m_{T}\}$
and $\widehat{\mathcal{T}}=\emptyset$. \\
(1) For $r\in\widehat{\mathcal{I}}\backslash\{2m_{T}\}$ uniformly
draw (without replacement) $K\in\{1,\ldots,\,m_{T}\}$ points $r_{k}^{\diamond}$
from $\mathbf{I}\left(r\right)=\{r-m_{T}+1,\ldots,r\}$ and compute
$\overline{r}^{\diamond}=\arg\max_{k=1,\ldots,\,K}\max_{\omega\in\left[-\pi,\,\pi\right]}\mathrm{D}_{r_{k}^{\diamond},T}\left(\omega\right)$;
set $\widehat{\mathcal{I}}=(\widehat{\mathcal{I}}\backslash\left\{ r\right\} )\cup\{\overline{r}^{\diamond}\}$.
\\
(2) If $\psi(\left\{ X_{t,T}\right\} ,\,\widehat{\mathcal{I}})=0$
return $\widehat{\mathcal{T}}=\emptyset$. Otherwise proceed with
step (3).\\
(3) Estimate the change-point $T\widehat{\lambda}_{T}(\widehat{\mathcal{I}})$
via \eqref{eq. (38) BJV - Obj Func Estimation} using $\widehat{\mathcal{I}}$.
\\
(4) Set $\widehat{\mathcal{I}}=\widehat{\mathcal{I}}\backslash\{T\widehat{\lambda}_{T}(\mathcal{\widehat{I}})-v_{T},\ldots,\,T\widehat{\lambda}_{T}(\mathcal{\widehat{I}})+v_{T}\}$
and $\widehat{\mathcal{T}}=\widehat{\mathcal{T}}\cup\{T\widehat{\lambda}_{T}(\mathcal{\widehat{I}})\}$.
Return to step (1).
\end{lyxalgorithm}
Finally, arrange the estimated change-points $\widehat{\lambda}_{l,T}$
in $\widehat{\mathcal{T}}$ in chronological order and use the symbol
$\left|\mathcal{S}\right|$ for the cardinality of a set $\mathcal{S}$.
To each $\widehat{\lambda}_{l,T}$ the procedure can return the frequency
$\widehat{\omega}_{l}$ at which the break is found.
\begin{assumption}
\label{Assumption small shift Multiple breaks}$\delta_{l,T}=\delta_{l}\neq0$
is fixed or $\delta_{l,T}\rightarrow0$ with $\inf_{1\leq l\leq m_{0}}\delta_{l,T}\geq2D^{*}M_{S,T}^{-1/2}(\log(T))^{2/3}$.
For $\nu_{T}\rightarrow\infty$ with $\nu_{T}=o\left(T/v_{T}\right),$
it holds that $\inf_{1\leq l\leq m_{0}-1}|\lambda_{l+1}^{0}-\lambda_{l}^{0}|\geq\nu_{T}^{-1}.$
\end{assumption}
Assumption \ref{Assumption small shift Multiple breaks} allows for
shrinking shifts and a possibly growing number of change-points as
long as $m_{0}/\nu_{T}\rightarrow0.$ The following proposition presents
the consistency result for the number of change-points $m_{0}$ and
for the change-point locations $\lambda_{l}^{0}$ $\left(l=1,\ldots,\,m_{0}\right)$,
and the rate of convergence of their estimates.
\begin{prop}
\label{Prop 4.9 BJV Consitency Change-point}Let Assumption \ref{Assumption WZ 2011}-\ref{Taper, nT, W and K1, b1},
\ref{Assumption Smothness of A} with $m_{0}=1$ and Condition \ref{Condition n_T h_T}
hold. Then, under $\mathcal{H}_{1}^{\mathrm{B},m_{0}}$ we have (i)
$\mathbb{P}(|\widehat{\mathcal{T}}-m_{0}|>\epsilon_{2})\rightarrow0$
for any $\epsilon_{2}>0$ and $\sup_{1\leq l\leq m_{0}}|\widehat{\lambda}_{l,T}-\lambda_{l}^{0}|=o_{\mathbb{P}}\left(1\right)$,
and (ii) $\sup_{1\leq l\leq m_{0}}|\widehat{\lambda}_{l,T}-\lambda_{l}^{0}|=O_{\mathbb{P}}(m_{S,T}\sqrt{M_{S,T}\log\left(T\right)}/(T\inf_{1\leq l\leq m_{0}}\delta_{l,T}))$.
Furthermore, if $K=O(a_{T}m_{T})$ with $a_{T}\in(0,\,1]$ such
that $a_{T}\rightarrow1,$ then the breaks are detected in decreasing
order of magnitude.
\end{prop}
The number of draws $K$ may be fixed or increase with the sample
size. However, the algorithm can return the change-point dates in
decreasing order of the break magnitudes only if $K$ is sufficiently
large. Note that at each loop of the algorithm it is not possible
to know to which $\lambda_{l}^{0}$ $\left(l=1,\ldots,\,m_{0}\right)$
the estimate $\widehat{\lambda}_{l,T}$ is consistent for. Only after
all breaks are detected and we rearrange the estimated change-points
in $\widehat{\mathcal{T}}$ in chronological order, we can learn such
information. The same procedure can be applied for the case of multiple
smooth local changes, though the notation becomes cumbersome and so
we omit it.
\section{Implementation\label{Section Implementation}}
In this section we explain how to choose the tuning parameters. The
choice of $m_{T}$, $n_{T}$ and $b_{W,T}$ could be based on a mean-squared
error (MSE) criterion or cross-validation exploiting results derived
for locally stationary series {[}e.g., data-dependent methods for
bandwidths in the context of locally stationary processes were investigated
by, among others, \citet{casini_hac}, \citet{dahlhaus:12}, \citet{Dahlhaus/Giraitis:98}
and \citet{ritcher/dahlhaus:19}{]}. The optimal amount of smoothing
depends on the regularity exponent $\theta$, on the boundness of
the moments and on the extent of the dependence in $\left\{ X_{t,T}\right\} $.
Here we choose the order of the bandwidths, neglecting the constants,
by following the restrictions in Condition \ref{Condition n_T h_T}.
In particular, we choose the largest possible values allowed by Condition
\ref{Condition n_T h_T} in order to ensure the highest possible power.
We conduct a sensitivity analysis based on simulations in the supplement.
We relegate to future work a more detailed analysis of data-dependent
methods for this problem with multiple smoothing directions.
For spectral densities satisfying Lipschitz continuity, $\theta=1$
so that $m_{T}\propto T^{2/3-\epsilon}$ while for $\theta=1/2$ we
have $m_{T}\propto T^{1/2-\epsilon}$ where in both cases $\epsilon>0.$
In applied work, it is common to work under stationarity ($\theta>1$)
or local stationarity with Lipschitz smoothness ($\theta=1$). Hence,
we use the bandwidths corresponding to $\theta=1$ which works for
both cases. Of course, if one has prior knowledge about the smoothness
of the parameters of the data-generating process, one can choose a
suitable $\theta$. Assuming $q=8$ and $\gamma$ large enough we
have $\tau_{T}\propto T^{1/4}$ and so values that satisfy Condition
\ref{Condition n_T h_T} are $m_{T}=T^{0.66}$, $n_{T}=T^{0.62}$
and $b_{W,T}=n_{T}^{-1/6}$. The scaling is normalized to 1, as
our simulations show this to provide good finite-sample properties,
see Section \ref{Section Monte Carlo}.
As for the tapering function $h\left(\cdot\right)$ and weight function
$W\left(\cdot\right)$ we use a rectangular taper (i.e., $h\left(u\right)=1$
for all $u$) and a rectangular kernel. The rectangular kernel is
known as the Daniell kernel with parameter $n_{T}b_{W,T}$ (it is
a centered moving average which creates a smoothed value at time $Tu$
by averaging all values between $Tu-n_{T}b_{W,T}$ and $Tu+n_{T}b_{W,T}$).
These are the simplest choices for $h\left(\cdot\right)$ and $W\left(\cdot\right)$.
As for the bandwidth and kernel of the estimator $\widehat{\sigma}_{L,r}\left(\omega\right),$
we follow the results in \citeauthor{casini_comment_andrews91} (\citeyear{casini_comment_andrews91},
\citeyear{casini_hac}) that suggest $b_{1,T}=M_{S,T}^{-1/3}.$ This
corresponds to the MSE-optimal bandwidth when $K_{1}\left(\cdot\right)$
is the Bartlett kernel. For the choice of the number and values
of the frequencies, the theory does not suggest particular values.
Thus, we tried several values $n_{\omega}=15,\,11,\,7,\,5.$ Our default
choice is $n_{\omega}=7$. A sensitivity analysis suggests that different
choices for $n_{\omega}$ lead to negligible differences in the results.
For the selection of the set of frequencies we use the function linspace
that generates a linearly spaced sequence, e.g., in \textsc{Matlab}
we used the command $\mathrm{\mathsf{linspace(-\pi+1e-3,\,0,\,(\mathit{n}_{\omega}+1)/2)}.}$
The regularity exponent $\theta$ also affects the test $\psi(\left\{ X_{t,T}\right\} ,\,\mathcal{I})$
in Algorithm \ref{Algorithm 1}. It is possible to get an estimate
of $\theta$ under the null as follows. Compute $\mathrm{S}_{\mathrm{Dmax},T}$
where the maximum is taken among the indices of the blocks such that
the null hypothesis is not violated and label it $s_{\mathrm{Dmax}}^{*}$.
Solve $s_{\mathrm{Dmax}}^{*}=2\sqrt{\log\left(M_{T}^{*}\right)/m_{T}^{*}}$
for $\theta$, where recall that $m_{T}^{*}$ and $M_{T}^{*}$ depend
on $\theta.$ This yields a preliminary estimate of $\theta$ which
can then be used for the test $\psi(\left\{ X_{t,T}\right\} ,\,\mathcal{I})$.
Similarly, $b_{T}^{\mathrm{opt}}$ depends on $\theta$ and $\theta'$.
Using the same approach, for a given $\theta'$ one can solve $s_{\mathrm{Dmax}}^{*}=b_{T}^{\mathrm{opt}}$
for $\theta$ as function of $\theta'.$ If one is interested in the
alternative $\mathcal{H}_{1}^{\mathrm{B}}$, we have $\theta'=0$
and so this immediately yields an estimate for $\theta$. If one is
interested in the alternative $\mathcal{H}_{1}^{\mathrm{S}}$, then
one can try a few values of $\theta'$ in the range $\left(0,\,\theta\right)$.
However, note that in order to use Algorithm \ref{Algorithm 1} only
$\theta$ is needed. The knowledge of $\theta'$ under $\mathcal{H}_{1}^{\mathrm{S}}$
is only needed to obtain $b_{T}^{\mathrm{opt}}.$
We set $v_{T}=T^{0.666}$ which satisfies $m_{T}/v_{T}\rightarrow0.$
Our default recommendation is $K=10.$ Our simulations with different
data-generating processes and sample sizes show that this choice strikes
a good balance between the precision of the change-point estimates
and computing time. For $T>1000$, we recommend setting $K=\left\lfloor m_{T}/3\right\rfloor $.
The test statistics $\mathrm{S}_{\mathrm{max},T}\left(\omega\right)$
and $\mathrm{R}_{\mathrm{max},T}\left(\omega\right)$ depend on $\omega.$
The choice of $\omega$ is, of course, important as it involves different
frequency components and hence different periodicities. If the user
does not have a priori knowledge about the frequency at which the
spectrum has a change-point, our recommendation is to run the tests
for multiple values of $\omega\in\left[0,\,\pi\right]$. Even if the
change-point occurs at some $\omega_{0}$ and one selects a value
of $\omega$ close but not equal to $\omega_{0}$ the tests are still
able reject the null hypothesis given the differentiability of $f\left(u,\,\omega\right)$.
Thus, one can select a few values of $\omega$ evenly spread on $\left[0,\,\pi\right]$.
\section{\label{Section Monte Carlo}Small-Sample Evaluations}
In this section, we conduct a Monte Carlo analysis to evaluate the
properties of the proposed methods. We first discuss the detection
of the change-points and then their localization. We investigate different
types of changes and consider the test statistics $\mathrm{S}_{\max,T}\left(\omega\right),$
$\mathrm{S}_{\mathrm{Dmax},T}$, $\mathrm{R}_{\max,T}\left(\omega\right)$,
$\mathrm{R}_{\mathrm{Dmax},T}$ proposed here and the test statistic
$\widehat{D}$ proposed by \citet{last/shumway:08}. The latter is
included for comparison since it applies to the same problems. We
consider the following data-generating processes where in all models
the innovation $e_{t}$ is a Gaussian white noise $e_{t}\sim\mathrm{i.i.d.}\,\mathscr{N}\left(0,\,1\right)$.
Models M1 involves a stationary AR(1) process $X_{t}=\rho X_{t-1}+e_{t}$
with $\rho=0.3$ and 0.6, while M2 involves a locally stationary AR(1)
$X_{t}=\rho\left(t/T\right)X_{t-1}+e_{t}$ where $\rho\left(t/T\right)=0.4\cos\left(0.8-\cos\left(2t/T\right)\right).$
Note that $\rho\left(t/T\right)$ varies smoothly from 0.1389 to 0.3920.
Model M1 and M2 are used to verify the finite-sample size of the tests.
We verify the power in models M3 and M4 using the specification in
model M1 and M2, respectively, for the first regime and consider two
additional regimes with different specifications. Hence, two breaks
are present. In model M3,
\begin{align*}
X_{t} & =\begin{cases}
0.3X_{t-1}+e_{t}, & 1\leq t\leq\left\lfloor T\lambda_{1}^{0}\right\rfloor \\
0.6X_{t-1}+0.7e_{t}, & \left\lfloor T\lambda_{1}^{0}\right\rfloor +1\leq t\leq\left\lfloor T\lambda_{2}^{0}\right\rfloor \\
0.6X_{t-1}+e_{t}, & \left\lfloor T\lambda_{2}^{0}\right\rfloor +1\leq t\leq T
\end{cases},
\end{align*}
while, for model M4
\begin{align*}
X_{t} & =\begin{cases}
\rho\left(t/T\right)X_{t-1}+0.7e_{t}, & 1\leq t\leq\left\lfloor T\lambda_{1}^{0}\right\rfloor \\
0.8X_{t-1}+e_{t}, & \left\lfloor T\lambda_{1}^{0}\right\rfloor +1\leq t\leq\left\lfloor T\lambda_{2}^{0}\right\rfloor \\
\rho\left(t/T\right)X_{t-1}+0.7e_{t}, & \left\lfloor T\lambda_{2}^{0}\right\rfloor +1\leq t\leq T
\end{cases},
\end{align*}
where $\rho\left(t/T\right)$ is as in model M2. In model M3, the
second regime involves higher serial dependence while in the third
regime the variance doubles relative to the second regime. In model
M4, the second regime involves a stationary autoregressive process
with strong serial dependence while in the third regime $X_{t}$ assumes
the same dynamics as in the first regime. Models M3-M4 feature alternative
hypotheses in the forms of breaks in the spectrum.
We consider the alternative hypothesis of more rough variation without
signifying a break (i.e., $\mathcal{H}_{1}^{\mathrm{S}}$ defined
in Section \ref{Section Consistency-and-Minimax}) in model M5 given
by $X_{t}=\sigma\left(t/T\right)e_{t}$ where $\sigma^{2}\left(t/T\right)=\max\{1.5,\,\overline{\sigma}^{2}+\cos\left(1+\cos\left(10t/T\right)\right)\}$
with $\overline{\sigma}^{2}=1.$ Note that even though $\sigma^{2}\left(\cdot\right)$
is locally stationary, the degree of smoothness alternates throughout
the sample. It starts from $\sigma^{2}\left(\cdot\right)=1.5$ and
maintains this value for some time, then within a short period it
increases slowly to $\sigma^{2}\left(\cdot\right)=2$ and falls slowly
back to $\sigma^{2}\left(\cdot\right)=1.5$. It keeps this value until
the final part of the sample where it increases slowly to $\sigma^{2}\left(\cdot\right)=2$
for a short period. Thus, $\sigma^{2}\left(\cdot\right)$ alternates
between periods where it is constant (i.e., $\theta>1$) and periods
where it becomes non-constant but less smooth (i.e., $\theta=1$).
Importantly, no break occurs; only a change in the smoothness as specified
in $\ensuremath{\mathcal{H}_{1}^{\mathrm{S}}}$. In unreported simulations
we also considered the case where $\theta$ changes from Lipschitz
continuity (i.e., $\theta=1$) to the continuity-path of Wiener processes
(i.e., $\theta\thickapprox1/2$) with results that are similar to
those reported here. For the test statistic $\widehat{D}$ of \citet{last/shumway:08},
we obtain the critical value by simulations. As suggested by the authors
we compute the finite-sample distribution of $\widehat{D}$ by simulating
a white noise under the null hypotheses with a sample size $T=1000$
and then obtain the critical value. We consider the three sample sizes
$T=250,\,500$ and 1000. The significance level is $\alpha=0.05$.
For the test statistics $\mathrm{S}_{\mathrm{max},T}\left(\omega\right)$
and $\mathrm{R}_{\mathrm{max},T}\left(\omega\right)$, we use as a
default value $\omega=0$ given that the interest is often in low
frequency analysis. We set $\lambda_{1}^{0}=0.33$ and $\lambda_{2}^{0}=0.66$
throughout. The number of simulations is 5,000 for all cases.
The results are reported in Table \ref{Table SM1-M2}-\ref{Table Power M3-M5}.
We first discuss the size of the tests. The tests proposed in this
paper have good empirical size for both models and all sample sizes.
The test statistics $\mathrm{S}_{\mathrm{Dmax},T}$ and $\mathrm{R}_{\mathrm{Dmax},T}$
are slightly undersized for $T=250$ but their empirical size improves
for $T=500$ and 1000. The test statistics $\mathrm{S}_{\mathrm{max},T}$
and $\mathrm{R}_{\mathrm{max},T}$ share accurate empirical sizes
in all cases. In contrast, the test statistic $\widehat{D}$ of \citet{last/shumway:08}
is largely oversized for $T=250$ and 500. For $T=1000$ it works
better but it is still oversized. This means that the finite-sample
distribution of $\widehat{D}$ has high variance and changes substantially
across different sample sizes. Since the simulated critical value
is obtained with a sample size $T=1000$ it works better for this
sample size than for the others for which the size control is poor.
Turning to the power of the tests, we note that it is not fair to
compare the proposed tests with the test $\widehat{D}$ when $T=250$
and 500 since the latter is largely oversized in those cases. In model
M3, all the proposed tests have good power which increases with the
sample size. The tests $\mathrm{S}_{\mathrm{Dmax},T}$ and $\mathrm{R}_{\mathrm{max},T}$
have the highest power, followed by $\mathrm{S}_{\mathrm{max},T}$
and lastly $\mathrm{R}_{\mathrm{Dmax},T}$. The power differences
are not large except those involving $\mathrm{R}_{\mathrm{Dmax},T}$
for $T=250$ which has substantially lower power. It is important
to note that for $T=1000$ the proposed tests have higher power than
the test $\widehat{D}$ of \citet{last/shumway:08} even though the
latter is oversized. For $T=250$ and 500, where the $\widehat{D}$
test is largely oversized, the proposed tests only have slightly lower
power. This confirms that the proposed tests have very good power.
Similar comments apply to model M4.
Model M5 involves changes in the smoothness without involving a break.
This constitutes a more challenging alternative hypothesis, and as
expected, the power for each test is lower than in models M3-M4. The
test with the highest power is $\mathrm{S}_{\mathrm{Dmax},T}$. For
$T=1000$, the test with the lowest power is $\widehat{D}$ (for $T=250$
the $\widehat{D}$ test has higher power again due to its oversize
problem). Overall, the results show that the proposed tests have accurate
empirical size even for small sample sizes and have good power against
different forms of breaks or changes in ``roughness''.
Next, we consider the estimation of the number of change-points ($m_{0}$)
and their locations. We consider the following two models, both with
$m_{0}=2$. The model M6 is given by
\begin{align*}
X_{t} & =\begin{cases}
0.7e_{t}, & 1\leq t\leq\left\lfloor T\lambda_{1}^{0}\right\rfloor \\
0.6X_{t-1}+0.7e_{t}, & \left\lfloor T\lambda_{1}^{0}\right\rfloor +1\leq t\leq\left\lfloor T\lambda_{2}^{0}\right\rfloor \\
0.6X_{t-1}+e_{t}, & \left\lfloor T\lambda_{2}^{0}\right\rfloor +1\leq t\leq T
\end{cases},
\end{align*}
while model M7 is the same as model M4. We set $\lambda_{1}^{0}=0.33$
and $\lambda_{2}^{0}=0.66$ and $T=1000$ throughout. Table \ref{Table Estimation}
reports summary statistics for $\widehat{m}-m_{0}$. It displays the
percentage of times with $\widehat{m}=m_{0}$, the median, and the
25\% and 75\% quantile of the distribution of $\widehat{m}$. We only
consider Algorithm \ref{Algorithm 1}. We do not report the results
for the corresponding procedure of \citet{last/shumway:08} because
it is based on $\widehat{D}$ which is oversized and so it finds many
more breaks than $m_{0}$. Table \ref{Table Estimation} shows that
$\widehat{m}=m_{0}$ occurs for about 85\% of the simulations with
model M6 and about 80\% with model M7. This suggests that Algorithm
\ref{Algorithm 1} is quite precise. As expected it performs better
in model M6 since the specification of the alternative is farther
from the null. The quantiles of the empirical distribution also suggest
that the change-point estimates $\widehat{T}_{1}$ and $\widehat{T}_{2}$
are accurate. For example the median is very close to the their respective
true value $T_{1}^{0}=333$ and $T_{2}^{0}=666$. Similar conclusions
arise from different models and sample sizes, in unreported simulations.
\section{\label{Section Empirical Application}Empirical Application}
We demonstrate how to use our change-point methods for studying the
causal effects of monetary policy. A fast growing literature in macroeconomics
uses high-frequency data to identify the effects of monetary policy
on the real economy (i.e., money non-neutrality). The identifying
assumption used by \citet{nakamura/steinsson:2018} is that the volatility
of the daily change in the nominal 2-year Treasury yields, say $\Delta i_{t}$,
is higher during days when the Federal Open Market Committee (FOMC)
meets to make monetary policy announcements relative to regular Tuesdays
and Wednesdays with no announcement. Let $\eta_{t}$ denote a pure
monetary shock and suppose that the policy instrument $\Delta i_{t}$,
which is observed in the data, is governed by both monetary and non-monetary
shocks:
\begin{align*}
\Delta i_{t} & =\mu_{i}+\eta_{t}+\varepsilon_{t},
\end{align*}
where $\varepsilon_{t}$ is a function of all other shocks that affect
$\Delta i_{t}$ and $\mu_{i}$ is a constant. We normalize the impact
of $\eta_{t}$ and $\varepsilon_{t}$ on $\Delta i_{t}$ to one. $\Delta i_{t}$
is a measure of the monetary policy news revealed in the FOMC announcement.
The idea is that changes in the policy instrument during days when
there is a FOMC announcement are dominated by the information about
future monetary policy contained in the announcement. Let $\Delta s_{t}$
denote the change in the outcome variable which is the yield on a
five year zero-coupon Treasury bond. We wish to estimate the effects
of the monetary shock $\eta_{t}$ on the outcome variable $\Delta s_{t}$.
The latter is also affected by both the monetary and non-monetary
shocks:
\begin{align*}
\Delta s_{t} & =\mu_{s}+\beta\eta_{t}+\alpha\varepsilon_{t},
\end{align*}
where $\mu_{s}$ is a constant and $\alpha$ and $\beta$ are two
parameters. The parameter of interest is $\beta$ which represents
the impact of the pure monetary shock $\eta_{t}$ on $\Delta s_{t}$
relative to its impact on $\Delta i_{t}$.
The identifying assumption is that the variance of monetary shocks
increases during days of FOMC announcements, while the variances of
the other shocks are unchanged. Let $T_{P}$ denote the number of
days containing a FOMC announcement, and let $T_{C}$ the number of
days with no such announcements. The subscript ``$P$'' in $T_{P}$
refers to the \textquotedblleft policy or treatment\textquotedblright{}
sample while the subscript ``$C$'' in $T_{C}$ refers to the \textquotedblleft control\textquotedblright{}
sample. The days in the control sample are comparable on other dimensions
since they are all Tuesdays and Wednesdays with no FOMC meeting. The
identifying restriction can then be written as
\begin{align}
\sigma_{\eta,P} & >\sigma_{\eta,C}\qquad\mathrm{and}\qquad\sigma_{\varepsilon,P}=\sigma_{\varepsilon,C},\label{Eq. Vol_P >Vol_C}
\end{align}
where $\sigma_{a,i}$ is the volatility of variable $a=\eta,\,\varepsilon$
in the sample $i=P,\,C$. One can show that
\begin{align}
\beta & =\frac{\mathrm{Cov}_{P}\left(\Delta i_{t},\,\Delta s_{t}\right)-\mathrm{Cov}_{C}\left(\Delta i_{t},\,\Delta s_{t}\right)}{\mathrm{Var}_{P}\left(\Delta i_{t}\right)-\mathrm{Var}_{C}\left(\Delta i_{t}\right)},\label{Eq. beta}
\end{align}
where $\mathrm{Cov}_{i}(\cdot,\,\cdot)$ (resp. $\mathrm{Var}_{i}(\cdot,\,\cdot)$)
denotes the population covariance (resp. variance) in the $i=P,\,C$
sample. The parameter $\beta$ can be identified only if $\mathrm{Var}_{P}\left(\Delta i_{t}\right)\neq\mathrm{Var}_{C}\left(\Delta i_{t}\right)$.
If the volatility of the policy instrument $\Delta i_{t}$ does not
change across treatment and control samples, then $\beta$ is not
identified. If the change in volatility is small, then $\beta$ is
weakly identified.
Subsequent developments in the literature employed robust weak identification
tests to argue that $\beta$ is not strongly identified when using
daily data. In contrast, $\beta$ can be strongly identified if one
uses ultra high-frequency data based on a 30 minute window around
the announcement time. For the case of daily data, we show that the
volatility of $\Delta i_{t}$ in the control sample varies substantially
over time and that several change-points can be detected. Thus, the
regimes in the control sample where the volatility is high contribute
to an average (over the full control sample) volatility that approaches
the average volatility in the treatment sample, thereby violating
$\mathrm{Var}_{P}\left(\Delta i_{t}\right)\neq\mathrm{Var}_{C}\left(\Delta i_{t}\right)$.
This implies that $\beta$ may be weakly identified and its estimates
may be imprecise which supports the recent evidence in the literature.
We obtain the data from Emi Nakamura's webpage. The sample of ``treatment''
days is all regularly scheduled FOMC meeting day from 1/1/2000 to
3/19/2014. The sample of ``control'' days is all Tuesdays and Wednesdays
that are not FOMC meeting days from 1/1/2000 to 12/31/2012. In both
the treatment and control samples, the second half of 2008, the first
half of 2009 and a 10 day period after 9/11/2001 are dropped in \citet{nakamura/steinsson:2018}.
We follow the same practice.
The plot of $\left\{ \Delta i_{t}\right\} $ for the control sample
is reported in Figure \ref{Fig1}. The series displays substantial
changes in volatility and some changes in persistence. The tests $\mathrm{R}_{\mathrm{max},T}$
and $\mathrm{R}_{\mathrm{Dmax},T}$ strongly reject the null hypothesis
of no change-points in the spectrum. Algorithm \ref{Algorithm 1}
detects three change-points.
The first change-point date, denoted $\widehat{T}_{1}$, corresponds
to April 24, 2007. Thus, the first regime $[1,\,\widehat{T}_{1}]$
refers to the period prior to the beginning of the 2007-09 financial
crisis. It is evident that a large change in volatility and possibly
a change in persistence occurred. We note that the change in volatility
does not occur abruptly, it is rather gradual. This means that
the volatility path changes gradually and possibly becomes more rough.
This supports the usefulness of our hypothesis testing framework with
local stationarity under the null hypothesis and with changes in the
smoothness of the parameters that govern the data-generating process
under the alternative hypothesis. The first change-point $\widehat{T}_{1}$
is associated to the volatility path becoming more rough with the
level of the volatility increasing gradually.
The second change-point date, denoted $\widehat{T}_{2}$, corresponds
to July 28, 2009. Thus, the second regime $[\widehat{T}_{1}+1,\,\widehat{T}_{2}]$
includes roughly the 2007-09 financial crisis. In this regime the
volatility of the series is remarkably high. After $\widehat{T}_{2}$
the series shares a pattern similar to that in the first regime $[1,\,\widehat{T}_{1}]$,
both in terms of persistence and volatility. The second change-point
date $\widehat{T}_{2}$ is associated to an abrupt fall in volatility.
The third change-point date, denoted $\widehat{T}_{3}$, corresponds
to February 2, 2011. The third regime $[\widehat{T}_{2}+1,\,\widehat{T}_{3}]$
corresponds to a zero lower bound (ZLB) period when the FOMC announced
the conduct of unconventional monetary policies to stimulate the economy
after the crisis. The ZLB refers to a situation in which the short-term
nominal interest rate is at or near zero, limiting the central bank's
capacity to stimulate growth. Unconventional monetary policies refer
to large-scale asset purchases and active use of communication (i.e.,
``forward guidance'') to shape expectations about future monetary
policies, which can mitigate the limitations imposed by the ZLB {[}see,
e.g., \citet{swanson:2021}{]}.
The last regime, $[\widehat{T}_{3}+1,\,T_{C}]$, corresponds to the
period when the economy was witnessing the first effects of the expansive
monetary policy and of the stability of the unconventional monetary
policies that were introduced previously. This is a regime where initially
the economy started the recovery and then reached stable economic
growth. In this regime, the volatility level of the series is the
lowest of the sample.
Overall, the first change-point date, $\widehat{T}_{1}$, corresponds
to a change in the smoothness while the second change-point date,
$\widehat{T}_{2}$, corresponds to an abrupt break. For the third
change-point date, it is more difficult to tell from the plot whether
this corresponds to an abrupt or smooth break.
We now discuss how the results about the change-points can be useful
for the identification issue based on \eqref{Eq. Vol_P >Vol_C}-\eqref{Eq. beta}.
One should think of $\mathrm{Var}_{C}\left(\Delta i_{t}\right)$ as
the average variance of $\Delta i_{t}$ in the control sample. Our
results show that there is significant time variation in $\mathrm{Var}_{C}\left(\Delta i_{t}\right)$.
In the second regime, $[\widehat{T}_{1}+1,\,\widehat{T}_{2}]$, the
volatility is the highest of the sample. During this period, it is
lower but very close to the average volatility of $\Delta i_{t}$
in the treatment sample, $\mathrm{Var}_{P}\left(\Delta i_{t}\right)$.\footnote{We refer to \citet{nakamura/steinsson:2018} for details about the
series in the treatment sample. We do not test for change-points in
the treatment sample because the sample size in the treatment sample
is relatively small ($T_{P}=74)$ and so we treat it as a single regime.} This contributes to make $\mathrm{Var}_{P}\left(\Delta i_{t}\right)-\mathrm{Var}_{C}\left(\Delta i_{t}\right)$
(over the full sample) closer to zero which then would lead to weak
identification. In fact, \citet{nakamura/steinsson:2018} found the
estimate of $\beta$ to be imprecise and not meaningful from an economic
standpoint. Furthermore, it was quite different from that obtained
using 30 minute data instead of daily data. Our change-point analysis
is useful because it suggests which periods or sub-samples contribute
to this identification problem and which sub-samples can be used to
obtain consistent estimates for $\beta.$
\section{\label{Section Conclusions}Conclusions}
We develop a theoretical framework for inference about the smoothness
of the spectral density over time. We provide frequency-domain statistical
tests for the detection of discontinuities in the spectrum of a segmented
locally stationary time series and for changes in the regularity exponent
of the spectral density over time. The null distribution of the test
follows an extreme value distribution. We rely on the theory on minimax-optimal
testing developed by \citet{ingster:93}. We determine the optimal
rate for the minimax distinguishable boundary, i.e., the minimum break
magnitude such that we are still able to uniformly control type I
and type II errors. We propose a novel procedure to estimate the change-points
based on a wild sequential top-down algorithm and show its consistency
under shrinking shifts and possibly growing number of change-points.
The advantage of using frequency-domain methods to detect change-points
is that it does not require to make assumptions about the data-generating
process under the null hypothesis beyond the fact that the spectrum
is differentiable and bounded. Furthermore, the method allows for
a broader range of alternative hypotheses compared to time-domain
methods which usually have power against a limited set of alternatives.
Overall, our simulations and empirical results show the usefulness
of our method.
\newpage{}
\bibliographystyle{elsarticle-harv}
\bibliography{References_JoE}
\addcontentsline{toc}{section}{References}
\section{Appendix}
\subsection{Tables}
\begin{table}[H]
\caption{\label{Table SM1-M2}Empirical small-sample size for models M1-M2}
\begin{centering}
\begin{tabular}{lccc}
\hline
& \multicolumn{3}{c}{Model M1}\tabularnewline
$\alpha=0.05$ & $T=250$ & $T=500$ & $T=1000$\tabularnewline
\hline
\hline
$\mathrm{S}_{\max,T}\left(0\right)$ & 0.039 & 0.043 & 0.053\tabularnewline
$\mathrm{S}_{\mathrm{Dmax},T}$ & 0.029 & 0.049 & 0.047\tabularnewline
$\mathrm{R}_{\mathrm{max},T}\left(0\right)$ & 0.040 & 0.054 & 0.042\tabularnewline
$\mathrm{R}_{\mathrm{Dmax},T}$ & 0.025 & 0.032 & 0.038\tabularnewline
$\widehat{D}$ statistic & 0.581 & 0.471 & 0.068\tabularnewline
\hline
& \multicolumn{3}{c}{Model M2}\tabularnewline
& $T=250$ & $T=500$ & $T=1000$\tabularnewline
\hline
\hline
$\mathrm{S}_{\max,T}\left(0\right)$ & 0.061 & 0.059 & 0.057\tabularnewline
$\mathrm{S}_{\mathrm{Dmax},T}$ & 0.035 & 0.055 & 0.058\tabularnewline
$\mathrm{R}_{\mathrm{max},T}\left(0\right)$ & 0.036 & 0.035 & 0.039\tabularnewline
$\mathrm{R}_{\mathrm{Dmax},T}$ & 0.025 & 0.032 & 0.035\tabularnewline
$\widehat{D}$ statistic & 0.731 & 0.583 & 0.102\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\end{table}
\begin{table}[H]
\caption{\label{Table Power M3-M5}Empirical small-sample power for models
M3-M5}
\begin{centering}
\begin{tabular}{lccc}
\hline
& \multicolumn{3}{c}{Model M3}\tabularnewline
$\alpha=0.05$ & $T=250$ & $T=500$ & $T=1000$\tabularnewline
\hline
\hline
$\mathrm{S}_{\max,T}\left(0\right)$ & 0.694 & 0.850 & 0.889\tabularnewline
$\mathrm{S}_{\mathrm{Dmax},T}$ & 0.734 & 0.890 & 0.921\tabularnewline
$\mathrm{R}_{\mathrm{max},T}\left(0\right)$ & 0.768 & 0.940 & 0.973\tabularnewline
$\mathrm{R}_{\mathrm{Dmax},T}$ & 0.456 & 0.752 & 0.874\tabularnewline
$\widehat{D}$ statistic & 0.961 & 0.967 & 0.790\tabularnewline
\hline
& \multicolumn{3}{c}{Model M4}\tabularnewline
& $T=250$ & $T=500$ & $T=1000$\tabularnewline
\hline
\hline
$\mathrm{S}_{\max,T}\left(0\right)$ & 0.868 & 0.964 & 0.973\tabularnewline
$\mathrm{S}_{\mathrm{Dmax},T}$ & 0.938 & 0.988 & 0.996\tabularnewline
$\mathrm{R}_{\mathrm{max},T}\left(0\right)$ & 0.927 & 0.997 & 0.999\tabularnewline
$\mathrm{R}_{\mathrm{Dmax},T}$ & 0.775 & 0.983 & 0.998\tabularnewline
$\widehat{D}$ statistic & 1.000 & 1.000 & 1.000\tabularnewline
\hline
& \multicolumn{3}{c}{Model M5}\tabularnewline
& $T=250$ & $T=500$ & $T=1000$\tabularnewline
\hline
\hline
$\mathrm{S}_{\max,T}$ & 0.223 & 0.475 & 0.565\tabularnewline
$\mathrm{S}_{\mathrm{Dmax},T}$ & 0.325 & 0.801 & 0.918\tabularnewline
$\mathrm{R}_{\mathrm{max},T}$ & 0.028 & 0.237 & 0.369\tabularnewline
$\mathrm{R}_{\mathrm{Dmax},T}$ & 0.025 & 0.189 & 0.304\tabularnewline
$\widehat{D}$ statistic & 0.834 & 0.695 & 0.172\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\end{table}
\begin{table}[H]
\caption{\label{Table Estimation}Summary statistics for the empirical distribution
of $\widehat{m}-m_{0}$}
\begin{centering}
\begin{tabular}{ccccc}
Percent time $\widehat{m}=m_{0}$ & & $Q_{0.25}$ & Median & $Q_{0.75}$\tabularnewline
\hline
\hline
\multicolumn{5}{c}{Model M6}\tabularnewline
\hline
85.50 & $\widehat{T}_{1}$ & 299 & 333 & 352\tabularnewline
& $\widehat{T}_{2}$ & 632 & 663 & 688\tabularnewline
\multicolumn{5}{c}{Model M7}\tabularnewline
\hline
80.12 & $\widehat{T}_{1}$ & 317 & 336 & 359\tabularnewline
& $\widehat{T}_{2}$ & 623 & 655 & 685 \tabularnewline
\hline
\end{tabular}
\par\end{centering}
\end{table}
\begin{center}
\begin{figure}[H]
\includegraphics[width=18cm,height=9cm]{fig1}
{\footnotesize{}\caption{{\scriptsize{}\label{Fig1} }{\footnotesize{}Plot of one-day changes
in the nominal Treasury yields ($\Delta i_{t}$) in the control sample.
The sample size is $T_{C}=762$ which corresponds to all Tuesdays
and Wednesdays that are not FOMC meeting days from 1/1/2000 to 12/31/2012.
Following \citet{nakamura/steinsson:2018} we drop the second half
of 2008, the first half of 2009 and a 10 day period after 9/11/2001.
The red dashed lines are change-point dates estimated using Algorithm
\ref{Algorithm 1}.}}
}{\footnotesize\par}
\end{figure}
\end{center}
\newpage{}
\clearpage
\pagenumbering{arabic}