EconBase
← Back to paper

Change-Point Analysis of Time Series with Evolutionary Spectra

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

110,213 characters · 19 sections · 123 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Change-Point Analysis of Time Series with Evolutionary Spectra

\setcounter{page}{0}

\raggedbottom

abstractThis paper develops change-point methods for the spectrum of a locally stationary time series. We focus on series with a bounded spectral density that change smoothly under the null hypothesis but exhibits change-points or becomes less smooth under the alternative. We address two local problems. The first is the detection of discontinuities (or breaks) in the spectrum at unknown dates and frequencies. The second involves abrupt yet continuous changes in the spectrum over a short time period at an unknown frequency without signifying a break. Both problems can be cast into changes in the degree of smoothness of the spectral density over time. We consider estimation and minimax-optimal testing. We determine the optimal rate for the minimax distinguishable boundary, i.e., the minimum break magnitude such that we are able to uniformly control type I and type II errors. We propose a novel procedure for the estimation of the change-points based on a wild sequential top-down algorithm and show its consistency under shrinking shifts and possibly growing number of change-points.

{\bf JEL Classification:} C12, C13, C22. \\ {\bf Keywords:} Change-point, Locally stationary, Frequency domain, Minimax-optimal test.

\onehalfspacing \thispagestyle{empty} \allowdisplaybreaks

Introduction

Classical change-point theory focuses on detecting and estimating structural breaks in the mean or regression coefficients. Early contributions include, among others, hinkley:71, yao:87, andrews:93, horvath:93 and bai/perron:98, who assume the presence of a single or multiple change-points in the parameters of an otherwise stationary time series model; see the reviews of aue/Horvath:13 and casini/perron_Oxford_Survey for more details. More recently there has been a growing interest about functional and time-varying parameter models where the latter are characterized by infinite-dimensional parameters which change continuously over time {[}see, e.g., dahlhaus:96, neumann/von_sachs:1997, hormann/kokoszka:2010, dette/preus/vetter:2011, zhang/wu:12, panaretos/tavakoli:13, aue/hormann:2015 and vandelft/eichler:2018{]}. Several authors have extended the stationarity tests originally introduced by priesley/rao:1969, and further developed by, e.g., dwivedi/rao:2010, jentsch/subbarao:2015 and bandyopadhyay/carsten/rao, to these settings. In the context of local stationarity, paparoditis:2009 proposed a test based on comparing a local estimate of the spectral density to a global estimate and preuss/vetter/dette:2013 proposed a test for stationarity using empirical process theory. In the context of functional time series, tests for stationarity were considered by horvath/kokoszka/rice:2014 and aue/rice/sonmez using time-domain methods, and by aue/vandelft:2020 and vandelft/Characiejus/dette:2018 using frequency-domain methods.

There is wide agreement in empirical work that besides breaks in the mean the detection of breaks in the variance or the correlation structure of a time series is of importance. For example, the discrimination between regimes of high and low asset volatility is of central interest in finance and the detection of changes in the parameters of an autoregressive process is important to build superior forecasting procedures. In addition, discerning the type of the changes (continuous or abrupt changes) is useful in applied work. On the one hand, if one assumes local stationarity but the true data-generating process involves structural breaks then parameter estimates may be severely biased and inference may be misleading. On the other hand, if one assumes a structural break model but the parameters actually change gradually then similar issues may arise. A general approach that allows for both continuous changes as well as breaks is needed to avoid these issues. Therefore, the detection of breaks in an otherwise locally stationary time series is important.

We develop inference methods about the changes in the degree of smoothness of the spectrum of a locally stationary time series, and hence, about change-points in the spectrum as a special case. The key parameter is the regularity exponent that governs how smooth the path of the local spectral density is over time. We address two local problems. The first is the detection of discontinuities (or breaks) in the spectrum at some unknown date and frequency. The second involves the detection of abrupt yet continuous changes in the spectrum over a short time period at an unknown frequency without signifying a break. For example, the spectrum becomes rougher over a short time period, meaning that the paths are less smooth as quantified by the regularity exponent. This can occur for a stationary process whose parameters start to evolve smoothly according to Lipschitz continuity, or for a locally stationary process with Lipschitz parameters that change to continuous but non-differentiable functional parameters. For example, the volatility of high-frequency stock prices or of other macroeconomic variables is known to become rougher (i.e., less smooth) without signifying a structural break after central banks' official announcements, especially in periods of high market uncertainty. In seismology, earthquakes are made up of several seismic waves that arrive at different times and so changes in the smoothness properties of each wave is important for locating the epicenter and identifying what materials the waves have passed through. We consider minimax-optimal testing and estimation for both problems, following ingster:93. We determine the optimal rate for the minimax distinguishable boundary, i.e., the minimum break magnitude such that we are still able to uniformly control type I and type II errors. These results are different from the developments on minimax optimality obtained recently in the statistics literature for the classical change-point problem where the mean of the series is piecewise constant {[}see, e.g., liu/gao/samworth and verzelen/fromont/lerasle/reynaud-bouret:2020{]}.

The problem of discriminating discontinuities from a continuous evolution in a nonparametric framework has received relatively less attention than the classical change-point problem with a few exceptions {[}muller:92, spokoiny:98, muller/stadtmuller:99, wu/zhao:07 and bibinger/jirak/vetter:16{]}. These focused on nonparametric regression and high-frequency volatility, and considered time-domain methods while we consider frequency-domain methods. This adds a difficulty in that, e.g., the search for a break or smooth change has to run over two dimensions, the time and frequency indices. Our test statistics are the maximum of local two-sample $t$-tests based on the local smoothed periodogram. We construct statistics that allow the researcher to test for a change-point in the spectrum at a prespecified frequency and others that allow to detect a break in the spectrum without prior knowledge about the frequencies. These test statistics can detect both discontinuous and smooth changes, and therefore are useful for both inference problems discussed above. The asymptotic null distribution follows an extreme value distribution. In order to derive this result, we first establish several asymptotic results, including bounds for higher-order cumulants and spectra of locally stationary processes. These results are complementary to some in dahlhaus:96, paparoditis:2009, panaretos/tavakoli:13, aue/vandelft:2020 and casini_hac, and extend some classical frequency-domain results for stationary processes {[}e.g., brillinger:75{]} to locally stationary processes.

Change-point problems have also been studied in the frequency-domain in several fields, though with less generality. adak:98 investigated the detection of change-points in piecewise stationary time series by looking at the difference in the power spectral density for two adjacent regimes. She compared several distance metrics such as the Kolmogorov-Smirnov, Cr\'{a}mer-Von Mises and CUSUM-type distance proposed by coates/diggle. last/shumway:08 focused on detecting change-points in piecewise locally stationary series. They exploited some of the results in Kakizawa/shumway/taniguchi and huang/ombao/stoffer:2004 to propose a Kullback-Liebler discrimination information but did not derive the null distribution of the test statistic. dette/wu/zhou:2019 considered testing for change-points in the autocorrelation coefficient at some pre-specified lag $k$ while allowing for local stationarity under the null hypothesis. They proposed a time-domain method based on a CUSUM-type test. To the extent that the autocorrelation at a given lag is only one of the features contained in the spectrum, our setting is more general. zhang:16 considered testing for and estimating change-points in the mean of a piecewise locally stationary time series. This is a different problem from ours since the spectrum involves the second-order properties which are complex to study. preuss/puchstein/dette:2015 considered the detection and estimation of change-points in the autocovariance function but they required stationarity under the null hypothesis. Several authors considered methods based on segmenting the wavelet spectrum for piecewise stationary time series {[}see, e.g., barigozzi/cho/fryzlewicz:2018 and cho/fryzlewicz:2012 (cho/fryzlewicz:2012; cho/fryzlewicz:2017){]}. Outside of the wavelet context other contributions are kirch/muhsal/ombao:2015 and Schrhoder/ombao:2019. vogt/dette:2015 investigated the detection of gradual changes in a locally stationary time series. Our contribution is different since we provide a general change-point analysis about the time-varying spectrum of a time series and establish the relevant asymptotic theory of the proposed test statistics under both the null and alternative hypotheses.

We also address the problem of estimating the change-points, allowing their number to increase with the sample size and the distance between change-points to shrink to zero. We propose a procedure based on a wild sequential top-down algorithm that exploits the idea of bisection combining it with a wild resampling technique similar to the one proposed by fryzlewicz:14. We establish the consistency of the procedure for the number of change-points and their locations. We compare the rate of convergence with that of standard change-point estimators under the classical setting {[}e.g., yao:87, bai:94a, casini/perron_CR_Single_Break, casini/perron_Lap_CR_Single_Inf and casini/perron_SC_BP_Lap{]}. We verify the performance of our methods via simulations which show their benefits. The advantage of using frequency-domain methods to detect change-points is that they do not require to make assumptions about the data-generating process under the null hypothesis beyond the fact that the spectrum is bounded. Furthermore, the method allows for a broader range of alternative hypotheses than time-domain methods which usually have good power only against some specific alternatives. For example, tests for changes in the volatility do not have power for changes in the dependence and vice versa. Our methods are readily available for use in many fields such as speech processing, biomedical signal processing, seismology, failure detection, economics and finance. It can also be used as a pre-test before constructing the recently introduced double kernel long-run variance estimator that accounts more flexibly for nonstationarity {[}cf. casini_hac, casini/perron_PrewhitedHAC and casini/perron_Low_Frequency_Contam_Nonstat:2020{]}.

The rest of the paper is organized as follows. Section (ref) introduces the statistical setting and the hypothesis testing problems. Section (ref) presents the test statistics and states their null limit distributions. Section (ref) addresses the consistency of the tests and their minimax optimality. Section (ref) discusses the estimation of the change-points while Section (ref) provides details for the implementation of the methods. The results of some Monte Carlo simulations are presented in Section (ref). An empirical application is presented in Section (ref). Section (ref) reports brief concluding comments. An online supplement {[}cf. casini/perron:change-point-spectra_Supp{]} contains additional theoretical and empirical results, and all mathematical proofs.

Statistical Environment and the Testing Problems

Section (ref) introduces the statistical setting and Section (ref) presents the hypotheses testing problems. We work in the frequency-domain under the locally stationary framework introduced by dahlhaus:96 who formalized the ideas of priestley:65. casini_hac extended his framework to allow for discontinuities in the spectrum which then results in a segmented locally stationary process. This corresponds to the relevant process under the alternative hypothesis of breaks in the spectrum. Since local stationarity is a special case of segmented local stationarity we begin with the latter. We use an infill asymptotic setting whereby we rescale the original discrete time horizon $\left[1,\,T\right]$ by dividing each $t$ by $T.$

Segmented Locally Stationary Processes

Suppose $\left\{ X_{t}\right\} _{t=1}^{T}$ is defined on an abstract probability space $\left(\Omega,\,\mathscr{F},\,\mathbb{P}\right)$, where $\Omega$ is the sample space, $\mathscr{F}$ is the $\sigma$-algebra and $\mathbb{P}$ is a probability measure. Let $i\triangleq\sqrt{-1}$. We use the notation $\overline{A}$ for the complex conjugate of $A\in\mathbb{C}$.

defnA sequence of stochastic processes $\{X_{t,T}\}_{t=1}^{T}$ is called segmented locally stationary (SLS) with $m_{0}+1$ regimes, transfer function $A^{0}$ and trend $\mu$ if there exists a representation \begin{align} X_{t,T} & =\mu_{j}\left(t/T\right)+\int_{-\pi}^{\pi}\exp\left(i\omega t\right)A_{j,t,T}^{0}\left(\omega\right)d\xi\left(\omega\right),\qquad\qquad\left(t=T_{j-1}^{0}+1,\ldots,\,T_{j}^{0}\right), \end{align} for $j=1,\ldots,\,m_{0}+1$, where by convention $T_{0}^{0}=0$ and $T_{m_{0}+1}^{0}=T$ ($\mathcal{T}\triangleq\{T_{1}^{0},\,\ldots,\,T_{m_{0}}^{0}\}$), and the following holds: (i) $\xi\left(\omega\right)$ is a stochastic process on $\left[-\pi,\,\pi\right]$ with $\overline{\xi\left(\omega\right)}=\xi\left(-\omega\right)$ and \begin{align*} \mathrm{cum}\left\{ d\xi\left(\omega_{1}\right),\ldots,\,d\xi\left(\omega_{r}\right)\right\} & =\varphi\left(\sum_{j=1}^{r}\omega_{j}\right)g_{r}\left(\omega_{1},\ldots,\,\omega_{r-1}\right)d\omega_{1}\cdots d\omega_{r}, \end{align*} where $\mathrm{cum}\left\{ \cdot\right\} $ is the cumulant of $r$th order, $g_{1}=0,\,g_{2}\left(\omega\right)=1$, $\left|g_{r}\left(\omega_{1},\ldots,\,\omega_{r-1}\right)\right|\leq M_{r}<\infty$ and $\varphi\left(\omega\right)=\sum_{j=-\infty}^{\infty}\delta\left(\omega+2\pi j\right)$ is the period $2\pi$ extension of the Dirac delta function $\delta\left(\cdot\right)$. (ii) There exists a constant $K>0$ and a piecewise continuous function $A:\,\left[0,\,1\right]\times\mathbb{R}\rightarrow\mathbb{C}$ such that, for each $j=1,\ldots,\,m_{0}+1$, there exists a $2\pi$-periodic function $A_{j}:\,(\lambda_{j-1}^{0},\,\lambda_{j}^{0}]\times\mathbb{R}\rightarrow\mathbb{C}$ with $A_{j}\left(u,\,-\omega\right)=\overline{A_{j}\left(u,\,\omega\right)}$, $\lambda_{j}^{0}\triangleq T_{j}^{0}/T$ and, for all $T,$ \begin{align} A\left(u,\,\omega\right)=A_{j}\left(u,\,\omega\right) & \,\mathrm{\,for\,}\,\lambda_{j-1}^{0}<u\leq\lambda_{j}^{0},\\ \sup_{1\leq j\leq m_{0}+1}\sup_{T_{j-1}^{0}<t\leq T_{j}^{0}}\sup_{\omega\in\left[-\pi,\,\pi\right]}\left|A_{j,t,T}^{0}\left(\omega\right)-A_{j}\left(t/T,\,\omega\right)\right| & \leq K\,T^{-1}. \end{align} (iii) $\mu_{j}\left(t/T\right)$ is piecewise continuous.

The smoothness properties of $A$ in $u$ guarantees that $X_{t,T}$ has a piecewise locally stationary behavior. This means that $X_{t,T}$ is locally stationary in each segment where the notion of local stationarity is as introduced by dahlhaus:96. We refer to casini_hac for several theoretical properties of SLS processes. zhou:2013 considered piecewise locally stationary processes in a time-domain setting but his notion is less general. In particular, casini_hac also defined (and worked with) the covariance between observations belonging to different segments whereas previous works considered only the covariance between observations belonging to the same segment, thereby using smoothness which simplifies the analysis.

The Testing Problems

We focus on time-varying spectra that are bounded, thereby excluding unit root, and long memory processes. Unit roots or trending processes can be handled by, for example, taking first differences or using other de-trending techniques. For some $D<\infty$, we consider the following class of time-varying spectra under the null hypothesis,

align[align omitted — 368 chars of source]

At all continuity points $u$, $f\left(u,\,\omega\right)$ is related to $A\left(u,\,\omega\right)$ by the relationship $f\left(u,\,\omega\right)=|A\left(u,\,\omega\right)|^{2}$. The key parameter of the testing problem under the null hypotheses is $\theta>0$. This is the regularity exponent of $f$ in the time dimension. For $\theta>1$, $f$ is constant in $u$ and reduces to the spectral density of a stationary process. For $\theta=1$, $f$ is Lipschitz continuous in $u$. For $\theta<1$, $f$ is $\theta$-H\"older continuous. Local stationarity corresponds to $\theta>0$ and $f$ being differentiable {[}see dahlhaus:96b{]}. The latter is the setting that we consider under the null hypothesis. To avoid redundancy, we do not require differentiability directly for the functions in $\boldsymbol{F}\left(\theta,\,D\right)$ since below we assume that the transfer function $A\left(u,\,\omega\right)$ is differentiable in $u$ which in turn implies that $f$ is differentiable in $u$. Since most of the applied work concerning local stationarity relies on Lipschitz continuity (i.e., $\theta=1$), our specification of the null is more general.

We now discuss features that are relevant under the alternative hypothesis. Our focus is on (i) discontinuities of $f$ in $u$ which correspond to $\theta=0$ and (ii) decreases in the smoothness of the trajectory $u\mapsto f(u,\,\omega)$ for each $\omega$ which correspond to a decrease in $\theta$. We shall refer to a series affected by the changes of case (ii) as “becoming more rough or less smooth”. Both cases refer to the properties of the spectral density and thus to the second-order properties of $X_{t,T}$.

Case (i) involves a break in the spectrum, i.e., there exits $\lambda_{b}^{0}\in\left(0,\,1\right)$ such that $\Delta f\left(\lambda_{b}^{0},\,\omega\right)\triangleq(f(\lambda_{b}^{0},\,\omega)-\lim_{u\downarrow\lambda_{b}^{0}}f(u,\,\omega))\neq0$ for some $\omega\in\left[-\pi,\,\pi\right]$.

Case (ii) involves a fall in the regularity exponent from $\theta>0$ to $\theta'\in\left(0,\,\theta\right)$ after some $\lambda_{b}^{0}$ for some period of time and some $\omega$; i.e., the spectrum becomes rougher after some $\lambda_{b}^{0}\in\left(0,\,1\right)$ for some time period before returning to $\theta$-smoothness. The case of an increase in $\theta$ is technically more complex to handle (see Section (ref) for details). As an example, consider a locally stationary AR(1),

align*[align* omitted — 102 chars of source]

where $a:\,\left[0,\,1\right]\rightarrow\left(-1,\,1\right)$ and $\sigma:\,\left[0,\,1\right]\rightarrow\mathbb{R}_{+}$ are functional parameters satisfying a Lipschitz condition and $\left\{ e_{t}\right\} $ is an i.i.d. mean-zero sequence. Additionally, if $\sup_{u\in\left[0\,1\right]}|\sigma\left(u\right)|<\infty$ and the initial condition satisfies some regularity condition, then $f\left(u,\,\omega\right)$ is uniformly bounded and $\theta=1$. Problem (ii) refers to either $a\left(\cdot\right)$ or $\sigma\left(\cdot\right)$, or both, becoming less smooth, i.e., we have a change from $\theta=1$ (since the parameters satisfy a Lipschitz condition) to some $\theta'\in\left(0,\,1\right)$. In this case, $X_{t,T}$ becomes an AR(1) process with functional parameters that are still continuous but less smooth and may not be differentiable.

Case (i) has received most attention so far in the time series literature although under much stronger assumptions {[}e.g., $f\left(u,\,\omega\right)=f\left(\omega\right)${]}. Case (ii) is a new testing problem and can be of considerable interest in several fields even though it requires larger sample sizes than problem (i). For example, the volatility dynamics of financial or macroeconomic variables vary over time. bibinger/jirak/vetter:16 provided evidence that the volatility of stock prices can change substantially its path properties after a press release following a meeting of the Federal Open Market Committee, in particular, its path may become more rough. Since $f\left(u,\,\omega\right)$ is a smooth function of the parameters of the data-generating process, if $\sigma\left(u\right)$ becomes more rough, then also $f\left(u,\,\omega\right)$ becomes more rough as $u$ varies. Case (ii) can also be relevant in seismology since the study of the path properties of the seismic waves is important for locating the epicenter of an earthquake. We show below that our tests are consistent and have minimax optimality properties for both cases (i) and (ii). Note that case (ii) is a local problem. In this paper, we do not consider more global problems where for example the spectrum is such that a fall in $\theta$ to $\theta'\in\left(0,\,\theta\right)$ occurs on $(\lambda_{b}^{0},\,1]$. This represents a continuous change in the smoothness of the spectrum that persists until the end of the time interval. Different test statistics are needed for this case, as will be discussed later.

As discussed by last/shumway:08, an important question is which magnitude of the discontinuity in the time-varying spectrum\textcolor{red}{ }can be detected. Or equivalently, how much the spectrum can change over a short time without indicating a break. We introduce the quantity $b_{T}$, called the detection boundary or simply “rate”, which is defined as the minimum break magnitude $\Delta f\left(\lambda_{b}^{0},\,\omega\right)$ such that we are still able to uniformly control the type I and type II errors as indicated below. To address the minimax-optimal testing {[}cf. ingster:93{]}, we now introduce the testing problems (i) and (ii) and defer a more general treatment to Section (ref).

Testing Problem for Case (i)

Given the discussion above, for some fractional break point $\lambda_{b}^{0}\in\left(0,\,1\right)$ and frequency $\omega_{0}$, and a decreasing sequence $b_{T}$, we consider the following class of alternative hypotheses:

align[align omitted — 458 chars of source]

We can then present first the hypothesis testing problem that we wish to address:

align[align omitted — 563 chars of source]

Observe that $\mathcal{H}_{1}^{\mathrm{B}}$ requires at least one break but allows for multiple breaks even across different $\omega$. For the testing problem (ref), we establish the minimax-optimal rate of convergence of the tests suggested {[}see Ch. 2 in ingster/suslina:03 for an introduction{]}. A conventional definition is the following. For a nonrandomized test $\psi$ that maps a sample $\left\{ X_{t,T}\right\} $ to zero or one, we consider the maximal type I error

align*[align* omitted — 232 chars of source]

and the maximal type II error

align*[align* omitted — 360 chars of source]

and define the total testing error as $\gamma_{\psi}(\theta,\,b_{T})=\alpha_{\psi}(\theta)+\beta_{\psi}(\theta,\,b_{T}).$ The notion of asymptotic minimax-optimality is as follows. We want to find sequences of tests and rates $b_{T}$ such that $\gamma_{\psi}(\theta,\,b_{T})\rightarrow0$ as $T\rightarrow\infty$. The larger is $b_{T}$ the easier it is to distinguish between $\mathcal{H}_{0}$ and $\mathcal{H}_{1}^{\mathrm{B}}$ but we may incur at the same time a larger type II error $\beta_{\psi}\left(\theta,\,b_{T}\right)$. The optimal value $b_{T}^{\mathrm{opt}}$, named the minimax distinguishable rate, is the minimum value of $b_{T}>0$ such that $\lim_{T\rightarrow\infty}\inf_{\psi}\gamma_{\psi}\left(\theta,\,b_{T}\right)=0$. A sequence of tests $\psi_{T}$ that satisfies the latter relation for all $b_{T}\geq b_{T}^{\mathrm{opt}}$ is called minimax-optimal.

Minimax-optimality has been considered in other change-point problems. loader:96 and spokoiny:98 considered the nonparametric estimation of a regression function with break size fixed. bibinger/jirak/vetter:16 considered breaks in the volatility of semimartingales under high-frequency asymptotics while we focus on breaks in the spectral density and thus we work in the frequency-domain. Another difference from previous work is that we do not deal with i.i.d. observations; hence, we cannot use the same approach to derive the minimax bound as in bibinger/jirak/vetter:16 because their information-theoretic reductions exploit independence. We need to rely on approximation theorems {[}cf. berkes/philipp:79{]} to establish that our statistical experiment is asymptotically equivalent in a strong Le Cam sense to a high dimensional signal detection problem. This allows us to derive the minimax bound using classical arguments based on the results in ingster/suslina:03, Ch. 8. The relevant results are stated in Section (ref).

Testing Problem for Case (ii)

We consider alternative hypotheses where $f$ is less smooth than under $\mathcal{H}_{0}$, including the case of breaks as a special case. Suppose that under $\mathcal{H}_{0}$ the spectrum $f\left(u,\,\omega\right)$ is differentiable in both arguments and behaves until time $T\lambda_{b}^{0}$ as specified in $\boldsymbol{F}\left(\theta,\,D\right)$ for some $\theta>0$ and $D<\infty$. After $T\lambda_{b}^{0}$, the regularity exponent $\theta$ drops to some $\theta'$ with $0<\theta'<\theta$ for some non-trivial period of time. That is, since $\boldsymbol{F}\left(\theta,\,D\right)\subset\boldsymbol{F}\left(\theta',\,D\right)$, we need that $f$ behaves as $\theta'$-regular for some period of time such that there exists a $\omega$ with $\left\{ f\left(u,\,\omega\right)\right\} _{u\in\left[0,\,1\right]}\notin\boldsymbol{F}\left(\theta,\,D\right).$ This guarantees that $\mathcal{H}_{0}$ and $\mathcal{H}_{1}^{\mathrm{S}}$ (to be defined below) are well-separated. To this end, define for some function $g_{u}$ with $u\in\left[0,\,1\right]$, $\Delta_{h}^{\theta'}g_{u}=\left(g_{u+h}-g_{u}\right)/\left|h\right|^{\theta'}$ for $h\in\left[-u,\,1-u\right].$ The set of possible alternatives is then defined as

align*[align* omitted — 514 chars of source]

where $m_{T}\rightarrow\infty$ as $T\rightarrow\infty$ such that $m_{T}^{-1}T^{\epsilon}\rightarrow0$ with $\epsilon>0.$ This leads to the following testing problem,

align[align omitted — 544 chars of source]

Note that $\mathcal{H}_{1}^{\mathrm{S}}$ allows for multiple changes. $\mathcal{H}_{1}^{\mathrm{B}}$ is a special case of $\mathcal{H}_{1}^{\mathrm{S}}$ since it can be seen as the limiting case of $\mathcal{H}_{1}^{\mathrm{S}}$ as $\theta'\rightarrow0.$ In the context of infinite-dimensional parameter problems one faces the issue of distinguishability between the null and the alternative hypotheses. It is evident that one cannot test $f\in\boldsymbol{F}\left(\theta,\,D\right)$ versus $f\in\boldsymbol{F}\left(\theta',\,D\right)$ for $\theta>\theta'$. First, since $\boldsymbol{F}\left(\theta,\,D\right)\subset\boldsymbol{F}\left(\theta',\,D\right)$, one has at least to remove the set of functions in $\boldsymbol{F}\left(\theta,\,D\right)$ from those in $\boldsymbol{F}\left(\theta',\,D\right)$. Still, as discussed by ingster/suslina:03, this would not be enough since the two hypotheses are still too close. That explains why we focus on spectral densities $f$ that belong to $\boldsymbol{F}'_{1,\lambda_{b}^{0},\omega_{0}}\left(\theta,\,\theta',\,b_{T},\,D\right)$ under $\mathcal{H}_{1}^{\mathrm{S}}$. These are rough enough so as not to be close to functions in $\boldsymbol{F}\left(\theta,\,D\right)$. This is captured by the requirement that the difference quotient $\Delta_{h}^{\theta'}f$ exceeds the so-called rate $b_{T}$. As $T\rightarrow\infty$ the requirement becomes less stringent since $b_{T}\rightarrow0$. See hoffman/nick:11 and bibinger/jirak/vetter:16 for similar discussions in different contexts.

Tests for Changes in the Spectrum and Their Limiting Distributions

Section (ref) introduces the test statistics while Section (ref) presents the results concerning their asymptotic distributions under the null hypothesis. These results apply also to the case of smooth alternatives which we discuss formally in Section (ref).

The Test Statistics

We first define the quantities needed to define the tests. Let $h:\,\mathbb{R}\rightarrow\mathbb{R}$ be a data taper with $h\left(x\right)=0$ for $x\notin[0,\,1)$, \[ H_{k,T}\left(\omega\right)=\sum_{s=0}^{T-1}h\left(s/T\right)^{k}\exp\left(-i\omega s\right), \] and (for $n_{T}$ even),

align*[align* omitted — 625 chars of source]

where $I_{L,h,T}\left(u,\,\omega\right)$ (resp., $I_{R,h,T}\left(u,\,\omega\right)$) is the local periodogram over a segment of length $n_{T}\rightarrow\infty$ that uses observations to the left (resp. right) of $\left\lfloor Tu\right\rfloor $. $I_{L,h,T}\left(u,\,\omega\right)$ (resp., $I_{R,h,T}\left(u,\,\omega\right)$) is a near unbiased estimator of $f\left(u-n_{T}/T,\,\omega\right)$ (resp., $f\left(u+n_{T}/T,\,\omega\right)$). We allow for a data taper since one may want to put more weight on observations that are closer to $u$. The smoothed local periodogram is defined as

align*[align* omitted — 180 chars of source]

with $f_{R,h,T}\left(u,\,\omega\right)$ defined similarly to $f_{L,h,T}\left(u,\,\omega\right)$ but with $I_{R,h,T}\left(u,\,\omega\right)$ in place of $I_{L,h,T}\left(u,\,\omega\right)$, where $W_{T}\left(\omega\right)$ $(-\infty<\omega<\infty)$ is a family of weight functions of period $2\pi$,

align*[align* omitted — 132 chars of source]

with $b_{W,T}$ a bandwidth and $W\left(\beta\right)$ $(-\infty<\beta<\infty)$ a fixed function. We define

align*[align* omitted — 264 chars of source]

where

align*[align* omitted — 232 chars of source]

with $m_{S,T}=\left\lfloor m_{T}^{1/2}\right\rfloor $ and $M_{S,T}=\left\lfloor m_{T}/m_{S,T}\right\rfloor $. $\widetilde{f}_{a,r,T}\left(\omega\right)$ ($a=L,\,R$) denotes the average local spectral density around time $rm_{T}$ computed using $f_{a,h,T}\left(j/T,\,\omega\right)$ where $r=1,\ldots,\,M_{T}=\left\lfloor T/m_{T}\right\rfloor -1$. We do not use all the $m_{T}$ local spectral densities $f_{a,\,h,T}\left(j/T,\,\omega\right)$ ($a=L,\,R$) in the block $r$ but only those separated by $m_{S,T}$ points. Thus, $\mathbf{S}_{r}$ is a subset of the indices in the block $r$. We need to consider a sub-sample of the $f_{a,h,T}\left(j/T,\,\omega\right)$'s ($a=L,\,R$) because there is strong dependence among the adjacent terms, e.g., $f_{a,h,T}\left(j/T,\,\omega\right)$ and $f_{a,h,T}\left((j+1)/T,\,\omega\right)$ ($a=L,\,R$). A large deviation between $\widetilde{f}_{L,r,T}\left(\omega\right)$ and $\widetilde{f}_{R,r+1,T}\left(\omega\right)$ suggests the presence of a break in the spectrum close to time $\left(r+1\right)m_{T}$ at frequency $\omega$. Note that the latest observation used in the construction of $\widetilde{f}_{L,r,T}$ is $X_{rm_{T}+m_{T}/2+\left\lfloor n_{T}/2\right\rfloor +1}$ while the earliest observation used in the construction of $\widetilde{f}_{R,r+1,T}$ is $X_{rm_{T}+m_{T}/2+\left\lfloor n_{T}/2\right\rfloor +2}$. This shows that there is no overlapping in the time points used in the construction of $\widetilde{f}_{L,r,T}$ and $\widetilde{f}_{R,r+1,T}$. The reason behind this is that in order to maximize power $\widetilde{f}_{L,r,T}$ and $\widetilde{f}_{R,r+1,T}$ should not use common observations, otherwise the effect of the common observations to the difference in the averages would offset the effect of the change-point.

Let

align*[align* omitted — 702 chars of source]

where $\widetilde{m}_{S,T}=\left\lfloor m_{T}^{1/3}\right\rfloor $ and $\widetilde{M}_{S,T}=\left\lfloor m_{T}/\widetilde{m}_{S,T}\right\rfloor $. Define $\widehat{\sigma}_{L,r}^{2}\left(\omega\right)=\sum_{j=-\widetilde{M}_{S,T}+1}^{\widetilde{M}_{S,T}-1}K_{1}\left(b_{1,T}j\right)\widehat{\Gamma}_{r}\left(j\right)$ where

align*[align* omitted — 473 chars of source]

and $\widehat{f}_{L,h,T}\left(j/T,\,\omega\right)=f_{L,h,T}\left(j/T,\,\omega\right)-\widetilde{f}_{L,r,T}\left(\omega\right)$ for $j\in\mathbf{S}_{r}$. The quantity $\widehat{\sigma}_{L,r}^{2}\left(\omega\right)$ is a local long-run variance estimator where $K_{1}$ is a kernel and $b_{1,T}$ is the associated bandwidth.

We first present a test statistic for the detection of a change-point in the spectrum $f\left(\cdot,\,\omega\right)$ for a given frequency $\omega.$ A second test statistic that we consider detects change-points in $u\in(0,\,1)$ occurring at any frequency $\omega\in\left[-\pi,\,\pi\right]$. The latter is arguably more useful in practice because often the practitioner does not know a priori at which frequency the spectrum is discontinuous. We begin with the following test statistic,

align[align omitted — 303 chars of source]

Test statistics of the form of (ref) were also used in the time-domain in the context of nonparametric change-point analysis under a less general framework {[}cf. eichinger/kirch:2018, bibinger/jirak/vetter:16 and wu/zhao:07{]} and forecasting {[}cf. casini_CR_Test_Inst_Forecast{]}. The derivation of the null distribution uses a (strong) invariance principle for nonstationary processes {[}see, e.g., \nocite{wu:07} and wu/zhou:11{]}.

The test statistic $\mathrm{S}_{\mathrm{max},T}\left(\omega\right)$ aims at detecting a break in the spectrum at some given frequency $\omega$. An alternative would be to consider a double-sup statistic which takes the maximum over $\omega\in\left[-\pi,\,\pi\right]$. Theorem (ref) in the supplement shows that $I_{h,L,T}\left(u,\,\omega_{j}\right)$ and $I_{h,L,T}\left(u,\,\omega_{k}\right)$ are asymptotically independent if $2\omega_{j},\,\omega_{j}\pm\omega_{k}\not\equiv0\,(\mathrm{mod\,}2\pi)$.\footnote{The notation $2\omega_{j},\,\omega_{j}\pm\omega_{k}\not\equiv0\,(\mathrm{mod\,}2\pi)$ means $2\omega_{j}\not\equiv0\,(\mathrm{mod\,}2\pi)$ and $\omega_{j}\pm\omega_{k}\not\equiv0\,(\mathrm{mod\,}2\pi).$ } However, the smoothing over frequencies introduces short-range dependence over $\omega$. Due to this short-range dependence, we cannot consider the maximum over all frequencies in $\Pi$ because the statistics would not be independent. Thus, we specify a framework based on an infill procedure over the frequency-domain $\left[-\pi,\,\pi\right]$ by assuming that there are $n_{\omega}$ frequencies $\omega_{1},\ldots,\,\omega_{n_{\omega}},$ with $\omega_{1}=-\pi$ and $\omega_{n_{\omega}}=\pi-\epsilon,\,\epsilon>0$, and $\left|\omega_{j}-\omega_{j+1}\right|=O\left(n_{\omega}^{-1}\right)$ for $j=1,\ldots,\,n_{\omega}-2$. Assume that $n_{\omega}\rightarrow\infty$ as $T\rightarrow\infty$. Let $\Pi\triangleq\left\{ \omega_{1},\ldots,\,\omega_{n_{\omega}}\right\} $. The maximum is taken over the following set of frequencies \[ \Pi'\triangleq\{\omega_{1},\,\omega_{2+\left\lfloor n_{T}b_{W,T}\right\rfloor },\ldots,\omega_{n_{\omega}-\left\lfloor n_{T}b_{W,T}\right\rfloor -1},\,\omega_{n_{\omega}}\}. \] Let $n'_{\omega}=\left\lfloor n_{\omega}/\left(\left\lfloor n_{T}b_{W,T}\right\rfloor +1\right)\right\rfloor $ with $n'_{\omega}\rightarrow\infty$. Note that $\Pi'\subset\Pi$. This then leads to the double-sup statistic,

equation[equation omitted — 236 chars of source]

where $\gamma_{M_{T}}=\left[4\log\left(M_{T}\right)-2\log\left(\log\left(M_{T}\right)\right)\right]^{1/2}$. This double-sup form is a new feature for change-point testing under the frequency-domain.

Next, we consider alternative test statistics that are self-normalized such that one does not need to estimate $\sigma_{L,r}^{2}\left(\omega\right)$. We consider the following test statistic,

align[align omitted — 226 chars of source]

where $\omega\in\left[-\pi,\,\pi\right].$ We can define a test statistic corresponding to $\mathrm{S}_{\mathrm{Dmax},T}$ by \[ \mathrm{R}_{\mathrm{Dmax},T}\triangleq\max_{\omega_{k}\in\Pi'}\sqrt{\log\left(M_{T}\right)}(M_{S,T}^{1/2}\mathrm{R}_{\mathrm{max},T}\left(\omega_{k}\right)-\gamma_{M_{T}})-\log\left(n'_{\omega}\right). \]

The Limiting Distribution Under the Null Hypothesis

Let $\mathbf{X}_{t,T}=(X_{t,T}^{\left(a_{1}\right)},\ldots,\,X_{t,T}^{\left(a_{p}\right)})$ with finite $p\geq1$. Denote by $\kappa_{\mathbf{X},t}^{\left(a_{1},\ldots,a_{r}\right)}\left(k_{1},\ldots,\,k_{r-1}\right)$ the time-$t$ cumulant of order $r$ of $(X_{t+k_{1},T}^{\left(a_{1}\right)},\ldots,X_{t+k_{r-1},T}^{\left(a_{r-1}\right)},\,X_{t,T}^{\left(a_{r}\right)})$ with $r\leq p$.

assumption(i) $\left\{ \mathbf{X}_{t,T}\right\} $ is a mean-zero locally stationary process (i.e., $m_{0}=0$); (ii) for all $j=1,\ldots,\,p$, $A^{\left(a_{j}\right)}\left(u,\,\omega\right)$ is \textcolor{red} $2\pi$-periodic in $\omega$ and the periodic extensions are differentiable in $u$ and $\omega$ with uniformly bounded derivative $\left(\partial/\partial u\right)\left(\partial/\partial\omega\right)A\left(u,\,\omega\right)$; (iii) $g_{4}$ is continuous.

Assumption (ref) requires $\left\{ \mathbf{X}_{t,T}\right\} $ to be locally stationary. Without loss of generality, we assume that $\left\{ \mathbf{X}_{t,T}\right\} $ has zero mean. All results go through when the mean is non-zero or when using demeaned series. The differentiability of $A\left(u,\,\omega\right)$ implies that $f\left(u,\,\omega\right)$ is also differentiable. This means that under the null hypothesis we require $f\left(u,\,\omega\right)$ to be differentiable in $u$ and to have some regularity exponent $\theta>0.$ The differentiability of $A\left(u,\,\omega\right)$ in $u$ can be relaxed at the expense of more complex proofs to establish the results in the supplement on high-order cumulants. Without differentiability, for any $\theta>0$ the test statistics above follow the same asymptotic distribution as when differentiability holds, though we do not discuss this case formally.

We need to impose some conditions on the temporal dependence. Let $\left\{ e_{t}\right\} _{t\in\mathbb{Z}}$ be a sequence of i.i.d. random variables and $\left\{ e'_{t}\right\} _{t\in\mathbb{Z}}$ be an independent copy of $\left\{ e_{t}\right\} _{t\in\mathbb{Z}}.$ Assume $X_{t,T}=H_{T}\left(t/T,\,\mathscr{F}_{t}\right)$ where $\mathscr{F}_{t}\triangleq\left\{ \ldots,\,e_{t-1},\,e_{t}\right\} $ and $H_{T}:\,\left[0,\,1\right]\times\mathbb{R}^{\infty}\mapsto\mathbb{R}$ is a measurable function. We use the dependence measure introduced by wu:05 (wu:05, wu:07) for stationary processes and extended to nonstationary processes by wu/zhou:11. Let $\mathscr{L}^{q}$ denote the space generated by the $q$-norm, $q>0$. For all $t$, assume $X_{t,T}\in\mathscr{L}^{q}$. For $w\geq0$ define the dependence measure,

align[align omitted — 272 chars of source]

where $\mathscr{F}_{t,\left\{ w\right\} }$ is a coupled version of $\mathscr{F}_{t}$ with $e_{w}$ replaced by an i.i.d. copy $e'_{w}$. For $\{X_{t,T}\}$ locally stationary, $H_{T}\left(\cdot,\,\mathscr{F}_{t}\right)$ is stochastic $\theta$-H\"older continuous in the sense that there exists $C_{T,H}<\infty$ such that

align[align omitted — 205 chars of source]

Assume $\Upsilon_{n,q}=\sum_{j=n}^{\infty}\phi_{j,q}<\infty$ for some $n\in\mathbb{Z}$. Let $\tau_{T}=T^{\vartheta_{1}}\left(\log\left(T\right)\right)^{\vartheta_{2}}$ where $\vartheta_{1}=\left(1/2-1/q+\gamma/q\right)$ $/\left(1/2-1/q+\gamma\right)$ and $\vartheta_{2}=\left(\gamma+\gamma/q\right)/\left(1/2-1/q+\gamma\right)$ for some $\gamma>0$.

assumptionFor $q\geq2$, $\sum_{n=0}^{\infty}n^{l+q-1}(\sum_{j=n}^{\infty}\phi_{j,q}^{2})^{1/2}<\infty$ where $l\geq0$.

Assumption (ref) was also used by shao/wu:2007ET who showed that it is satisfied for many nonlinear time series processes. Using Assumption (ref) we can give a sufficient condition for the summability of the joint cumulant up to a certain order. The latter is a common assumption in spectral analysis and we use it to establish results on the high-order cumulants and spectra in Section (ref). These are used to obtain the null limiting distribution of the test statistics.

lemLet Assumption (ref) hold and $X_{t,T}\in\mathscr{L}^{r}$, $r\geq2$. We have \begin{align} \sum_{k_{1},\ldots,\,k_{r-1}=-\infty}^{\infty}\left(1+\left|k_{j}\right|^{l}\right)\sup_{1\leq t\leq T}\left|\kappa_{\mathbf{X},t}^{\left(a_{1},\ldots,a_{r}\right)}\left(k_{1},\ldots,\,k_{r-1}\right)\right| & <\infty, \end{align} for $l\geq0$, $j=1,\ldots,\,r-1$, and any $r$-tuple $a_{1},\ldots,a_{r}$.

shao/wu:2007ET proved (ref) for $l=0$ and $\{X_{t,T}\}$ a stationary process.

assumption(i) The data taper $h:\,\mathbb{R}\rightarrow\mathbb{R}$ with $h\left(x\right)=0$ for $x\notin[0,\,1)$ is bounded and of bounded variation; (ii) The sequence $\left\{ n_{T}\right\} $ satisfies $n_{T}\rightarrow\infty$ as $T\rightarrow\infty$ with $n_{T}/T\rightarrow0$; (iii) $W\left(\beta\right)$ $(-\infty<\beta<\infty)$ is real-valued, even, of bounded variation, and satisfies $\int_{-\infty}^{\infty}W\left(\beta\right)d\beta=1$; (iv) $b_{1,T}\rightarrow0$ such that $Tb_{1,T}\rightarrow\infty$ and $(\widetilde{M}_{S,T}b_{1,T})^{-1/2}M_{S,T}^{1/2}\log T\rightarrow0$ and $K_{1}\left(\cdot\right):\,\mathbb{R}\rightarrow\left[-1,\,1\right],$ $K_{1}\left(0\right)=1,\,K_{1}\left(x\right)=K_{1}\left(-x\right),\,\forall x\in\mathbb{R}$ and ${\textstyle \int\nolimits _{-\infty}^{\infty}}K_{1}^{2}\left(x\right)dx<\infty$.

Assumption (ref)-(i,ii) are standard in the nonparametric estimation literature�while Assumption (ref)-(iii) is also used for spectral density estimation under stationarity {[}e.g., brillinger:75{]}. The conditions on $b_{1,T}$ in Assumption (ref)-(iv) are necessary for the consistency of the long-run variance estimator. The class of kernels allowed by Assumption (ref)-(iv) includes popular kernels such as the Truncated, Bartlett, Parzen, Quadratic Spectral (QS) and Tukey-Hanning kernels. For technical reasons inherent to the proofs, we need to assume that the spectral density is strictly positive. Theorem (ref) in the supplement shows that the variance of $f_{L,h,T}\left(u,\,\omega\right)$ depends on $f\left(u,\,\omega\right)$. Thus, the denominator of the test statistic depends on $f\left(u,\,\omega\right)$. Assumption (ref) requires the latter to be bounded away from zero. In practice, if one suspects that at some frequencies $f\left(u,\,\omega\right)$ can be close to zero, then one can add a small number $\epsilon_{f}>0$ to the denominator of the test statistic to guarantee numerical stability.

assumption$f_{-}=\min_{u\in\left[0,\,1\right],\,\omega\in\left[-\pi,\,\pi\right]}f\left(u,\,\omega\right)>0$.

The next assumption ensures that the local spectral density estimates are asymptotically independent when evaluated at some given frequencies (see Theorem (ref)). It is used to derive the asymptotic null distribution of the double-sup test statistics $\mathrm{S}_{\mathrm{Dmax},T}$ and $\mathrm{R}_{\mathrm{Dmax},T}$.

assumptionAssume that $2\omega_{j},\,\omega_{j}\pm\omega_{k}\not\equiv0\,(\mathrm{mod\,}2\pi)$ for $\omega_{j},\,\omega_{k}\in\Pi$.
condition(i) The sequence $\left\{ m_{T}\right\} $ satisfies $m_{T}\rightarrow\infty$ as $T\rightarrow\infty,$ and \begin{align} M_{S,T}^{1/2} & m_{T}^{\theta}T^{-\theta}\left(\log\left(M_{T}\right)\right)^{1/2}+\tau_{T}^{2}\log\left(M_{T}\right)M_{S,T}^{-1}\\ & +M_{S,T}n_{T}^{4}\log\left(M_{T}\right)T^{-4}+M_{S,T}\left(\log\left(n_{T}\right)\right)^{2}\log\left(M_{T}\right)n_{T}^{-2}\rightarrow0;\nonumber \end{align} (ii) $b_{W,T}\rightarrow0$ such that $Tb_{W,T}\rightarrow\infty$, $\log\left(M_{T}\right)M_{S,T}b_{W,T}^{4}\rightarrow0$ and $\log\left(M_{T}\right)M_{S,T}(n_{T}b_{W,T})^{-1}\rightarrow0.$

Part (i) imposes lower and upper bounds on the growth condition of the sequence $\left\{ m_{T}\right\} $. The upper bound relates to the smoothness of $A\left(u,\,\omega\right)$, the value of $\theta$ under the null hypothesis, $n_{T}$ and the number of summands $M_{S,T}$ in $\widetilde{f}_{L,r,T}^{*}\left(\omega\right)$.

Let $\mathscr{V}$ denote a random variable with an extreme value distribution defined by $\mathbb{P}\left(\mathscr{V}\leq v\right)=\exp($ $-\pi^{-1/2}\exp\left(-v\right)).$

thmLet Assumption (ref)-(ref) and Condition (ref) hold. Under $\mathcal{H}_{0}$, $\sqrt{\log\left(M_{T}\right)}(M_{S,T}^{1/2}\mathrm{S}_{\mathrm{max},T}\left(\omega\right)$ $-\gamma_{M_{T}})\Rightarrow\mathscr{V}$ for any $\omega\in\left[-\pi,\,\pi\right]$.

Theorem (ref) shows that the asymptotic null distribution follows an extreme value distribution. The derivation of the null distribution uses a (strong) invariance principle for nonstationary processes {[}see, e.g., \nocite{wu:07} and wu/zhou:11{]}. The following theorems shows that the asymptotic null distribution of the remaining tests $\mathrm{S}_{\mathrm{Dmax},T}$, $\mathrm{R}_{\mathrm{max},T}\left(\omega\right)$ and $\mathrm{R}_{\mathrm{Dmax},T}$ also follows an extreme value distribution, though the additional Assumption (ref) and the extra factor $\log\left(n'_{\omega}\right)$ are needed for $\mathrm{S}_{\mathrm{Dmax},T}$ and $\mathrm{R}_{\mathrm{Dmax},T}$.

thmLet Assumption (ref)-(ref) and Condition (ref) hold. Under $\mathcal{H}_{0}$, $\mathrm{S}_{\mathrm{Dmax},T}\overset{}{\Rightarrow}\mathscr{V}$.
thmLet Assumption (ref)-(ref) and Condition (ref) hold. Under $\mathcal{H}_{0}$, $\sqrt{\log\left(M_{T}\right)}(M_{S,T}^{1/2}\mathrm{R}_{\mathrm{max},T}\left(\omega_{k}\right)$ $-\gamma_{M_{T}})\Rightarrow\mathscr{V}$ and, in addition if Assumption (ref) holds, then $\mathrm{R}_{\mathrm{Dmax},T}\overset{}{\Rightarrow}\mathscr{V}$.

The limiting distributions in Theorem (ref)-(ref) are pivotal, so that critical values can be obtained immediately without the need to rely on simulations. The tests have power against both breaks (alternative hypothesis (i)) and changes in the smoothness (alternative hypothesis (ii)). Unfortunately, discerning between the two types of alternative hypotheses is a quite hard technical problem. Although knowing whether the rejection of the null hypothesis is due to a break or a change in the smoothness would be useful, in many cases knowing that there has been some change in the data-generating process is sufficient to modify the estimation and/or inference strategy to account for the change. If the change involves a break, a reasonable approach would be to apply some sample-splitting method for estimation. For example, in the context of long-run variance estimation the existence of a break suggests to modify the double kernel HAC (DK-HAC) estimator to avoid mixing two different regimes {[}cf. casini_hac{]}. On the other hand, for a change in the smoothness a modification of a standard nonparametric kernel smoothing could be enough. However, using a sample-splitting technique even when there is a change in the smoothness would be robust to the change and would result in better estimation and inference. In general, treating a change as a break and applying some sample-splitting method would be technically valid even when the rejection was due to a change in the smoothness (though it may not be efficient).

The property of being robust to different alternative hypotheses is not specific to our method. This property is shared by most of the existing structural break tests. For example, andrews:93 showed that structural break tests also have some power against some forms of smoothly-varying parameters. This property was seen as a positive feature in the structural break literature. Following the same reasoning, the property of being robust to breaks as well as changes in the smoothness can actually be seen as a virtue of our method.

Consistency and Minimax Optimal Rate of Convergence

In this section, we discuss the consistency and minimax-optimal lower bound for the testing problem (ref) (i.e., case (ii)). The discussion also covers the testing problem (i) since $\mathcal{H}_{1}^{\mathrm{B}}$ can be seen as the limiting case of $\mathcal{H}_{1}^{\mathrm{S}}$ as $\theta'\rightarrow0.$ We assume that $X_{t,T}$ is segmented locally stationary with transfer function $A\left(u,\,\omega\right)$ satisfying the following smoothness properties.

assumption(i) $\left\{ X_{t,T}\right\} $ is a mean-zero segmented locally stationary process; (ii) $A\left(u,\,\omega\right)$ is \textcolor{red}twice continuously differentiable in $u$ at all $u\neq\lambda_{j}^{0}$ $(j=1,\ldots,\,m_{0}+1)$ with uniformly bounded derivatives $\left(\partial/\partial u\right)A\left(u,\,\cdot\right)$ and $\left(\partial^{2}/\partial u^{2}\right)A\left(u,\,\cdot\right)$; (iii) $A\left(u,\,\omega\right)$ is twice left-differentiable in $u$ at $u=\lambda_{j}^{0}$ $(j=1,\ldots,\,m_{0}+1)$ with uniformly bounded derivatives $\left(\partial/\partial_{-}u\right)A\left(u,\,\cdot\right)$ and $\left(\partial^{2}/\partial_{-}u^{2}\right)A\left(u,\,\cdot\right)$.
assumption(i) $A\left(u,\,\omega\right)$ is twice differentiable in $\omega$ with uniformly bounded derivatives $\left(\partial/\partial\omega\right)$ $A\left(\cdot,\,\omega\right)$ and $\left(\partial^{2}/\partial\omega^{2}\right)A\left(\cdot,\,\omega\right)$; (ii) $g_{4}\left(\omega_{1},\,\omega_{2},\,\omega_{3}\right)$ is continuous in its arguments.

We now move to the derivation of the minimax lower bound. As explained before, we restrict attention to a strictly positive spectral density in the frequency dimension at which the null hypotheses is violated. That is, $f_{-}\left(\omega_{0}\right)=\inf_{u\in\left[0,\,1\right]}f\left(u,\,\omega_{0}\right)>0$. Such restriction is not imposed on $f\left(u,\,\omega\right)$ for $\omega\neq\omega_{0}$.

thmLet Assumption (ref)-(ref), (ref)-(ref) and $f_{-}\left(\omega_{0}\right)>0$ hold. Consider either set of hypotheses $\{\mathcal{H}_{0},\,\mathcal{H}_{1}^{\mathrm{B}}\}$ with $\theta'=0$ or $\{\mathcal{H}_{0},\,\ensuremath{\mathcal{H}_{1}^{\mathrm{S}}}\}$ with $0<\theta'<\theta$. Then, for \begin{align*} b_{T} & \leq\left(T/\log\left(M_{T}\right)\right)^{-\frac{\theta-\theta'}{2\theta+1}}D^{-\frac{2\theta'+1}{2\theta+1}}f_{-}\left(\omega_{0}\right), \end{align*} we have $\lim_{T\rightarrow\infty}\inf_{\psi}\gamma_{\psi}\left(\theta,\,b_{T}\right)=1.$

The theorem implies the need for \[ b_{T}^{\mathrm{opt}}>\left(T/\log\left(M_{T}\right)\right)^{-\frac{\theta-\theta'}{2\theta+1}}D^{-\frac{2\theta'+1}{2\theta+1}}f_{-}\left(\omega_{0}\right), \] otherwise there cannot exist a minimax-optimal test yielding $\lim_{T\rightarrow\infty}\inf_{\psi}\gamma_{\psi}\left(\theta,\,b_{T}\right)=0.$ Note that the lower bound does not depend on $\omega$. In Theorem (ref) we establish a corresponding upper bound. From the lower and upper bounds we deduce the optimal rate for the minimax distinguishable boundary. We can also derive tests based on $b_{T}^{\mathrm{opt}}$. For example, using the test statistic (ref) for $\{\mathcal{H}_{0},\,\mathcal{H}_{1}^{\mathrm{B}}\}$ we obtain the following test $\psi^{*}:$ $\psi^{*}(\left\{ X_{t,T}\right\} )=1$ if $\mathrm{S}_{\mathrm{max},T}\left(\omega\right)\geq2D^{*}\sqrt{\log\left(M_{T}^{*}\right)/m_{T}^{*}}$ for $\omega\in\left[-\pi,\,\pi\right]$ where $D^{*}>2$, $m_{T}^{*}=(\sqrt{\log\left(M_{T}^{*}\right)}T^{\theta}/D)^{\frac{2}{2\theta+1}}$ and $M_{T}^{*}=\left\lfloor T/m_{T}^{*}\right\rfloor $. Hence, in order to construct such a test we need knowledge of $\theta$ under $\mathcal{H}_{0}$. We discuss this in Section (ref).

Next, we establish the optimal rate for minimax distinguishability. Note that either alternatives $\mathcal{H}_{1}^{\mathrm{B}}$ or $\mathcal{H}_{1}^{\mathrm{S}}$ allows for multiple breaks. The following results require further restrictions on the relation between $n_{T}$ and $m_{T}$.

thmLet Assumption (ref)-(ref), (ref)-(ref) hold. Consider either alternative hypotheses $\ensuremath{\mathcal{H}_{1}^{\mathrm{B}}}$ with $\theta'=0$ and $\lambda_{j}^{0}<\lambda_{j+1}^{0}$ for $j=1,\ldots,\,m_{0}$, or $\ensuremath{\mathcal{H}_{1}^{\mathrm{S}}}$ with $0<\theta'<\theta$. If \begin{align} \left(\sqrt{\log\left(M_{T}^{*}\right)/m_{T}^{*}}\right)^{-1}\left(\left(m_{T}^{*}/T\right)^{\theta}+\left(n_{T}/T\right)^{2}+\log\left(n_{T}\right)/n_{T}+b_{W,T}^{2}\right) & \rightarrow0, \end{align} and \begin{align} b_{T}^{*} & >\left(4D^{*}\sup_{u\in\left[0,\,1\right]}f\left(u,\,\omega_{0}\right)+2\right)^{-\frac{\theta-\theta'}{2\theta+1}}\left(T/\log\left(M_{T}\right)\right)^{-\frac{\theta-\theta'}{2\theta+1}}D^{-\frac{2\theta'+1}{2\theta+1}}, \end{align} then $\lim_{T\rightarrow\infty}\gamma_{\psi^{*}}\left(\theta,\,b_{T}^{*}\right)=0$ and $b_{T}^{\mathrm{opt}}\propto\left(T/\log\left(M_{T}\right)\right)^{-\frac{\theta-\theta'}{2\theta+1}}.$

The theorem shows that a smooth change in the regularity exponent $\theta$ cannot be distinguished from a break of magnitude smaller than $b_{T}^{\mathrm{opt}}$ because the change from $\theta$ to $\theta'$ has to persist for some time. This is also indicated by the restriction $\theta'>0.$ The minimax bound is similar to the one established by bibinger/jirak/vetter:16 for the volatility of a It\^o semimartingale. The theorem suggests that knowledge of the frequency $\omega_{0}$ at which the spectrum changes regularity is irrelevant for the determination of the bound. However, we conjecture that if the spectrum exhibits a break or smooth change of the form discussed above simultaneously across multiple frequencies then the lower bound may be further decreased as one can pool additional information from inspection of the spectrum for the set of frequencies subject to the change. The key assumption would be that the change occurs at the same time $\lambda_{b}^{0}$ for a given set of frequencies $\omega.$ This may be of interest for economic and financial time series since they often exhibit a break simultaneously at high and low frequencies. We leave this to future research.

Estimation of the Change-Points

We now discuss the estimation of the break locations for the case of discontinuities in the spectrum (i.e., $\mathcal{H}_{1}^{\mathrm{B},m_{0}}$ where $m_{0}$ is the number of breaks, recall Definition (ref)). The same estimator is valid for the locations of the smooth changes as under $\mathcal{H}_{1}^{\mathrm{S}}$. For the latter case we later provide intuitive remarks about the consistency result and the conditions needed for it. We first consider the case of a single break (i.e., $\mathcal{H}_{1}^{\mathrm{B},1}$) and then present the results for the case of multiple breaks (i.e., $\mathcal{H}_{1}^{\mathrm{B},m_{0}}$).

Single Break Alternatives $\mathcal{H}_{1}^{\mathrm{B},1}$

Let

align*[align* omitted — 253 chars of source]

where

align*[align* omitted — 175 chars of source]

and $r=2m_{T},\,3m_{T},\ldots$ with $r<\left(M_{T}-1\right)m_{T}-n_{T}$. Note that the maximum of the statistics $\mathrm{D}_{r,T}\left(\omega\right)$ is a version of $\mathrm{S}_{\max,T}$ that does not involve the normalization. The change-point estimator is defined as \[ T\widehat{\lambda}_{b,T}=\underset{r=2m_{T},\,3m_{T},\ldots}{\mathrm{argmax}}\max_{\omega\in\left[-\pi,\,\pi\right]}\mathrm{D}_{r,T}\left(\omega\right). \] Recall that we consider the following alternative hypothesis:

align*[align* omitted — 223 chars of source]

Note that a break does not need to occur simultaneously at all frequencies for the procedure to work. The break magnitude can be either fixed or converge to zero as specified by the following assumption.

assumption$\delta_{T}=\delta\neq0$ is fixed or $\delta_{T}\rightarrow0$ and $\delta_{T}M_{S,T}^{1/2}/\sqrt{\log\left(T\right)}\rightarrow(0,\,\infty]$.
propLet Assumption (ref)-(ref)-(i-iii), (ref) with $m_{0}=1$ and Condition (ref) hold. Under $\mathcal{H}_{1}^{\mathrm{B},1},$ if $\delta_{T}$ satisfies Assumption (ref), we have $\widehat{\lambda}_{b,T}-\lambda_{b}^{0}=O_{\mathbb{P}}(m_{S,T}\sqrt{M_{S,T}\log(T)}/(T\delta_{T}))$.

We compare the rate of convergence in Proposition (ref) with that of classical change-point estimators in the piecewise constant mean model. For fixed shifts, the latter rate of convergence is $O_{\mathbb{P}}(T^{-1})$ while for shrinking shifts it is $O_{\mathbb{P}}((T\delta_{T}^{2})^{-1})$ where $\delta_{T}\rightarrow0$ with $\delta_{T}T^{1/2-\vartheta}$ for some $\vartheta\in\left(0,\,1/2\right)$ {[}cf. yao:87{]}.\footnote{See also verzelen/fromont/lerasle/reynaud-bouret:2020 for recent developments on nimimax optimality for change-point estimation in the piecewise constant mean model. They considered as a significance of a change-point a measure that depends on both the break magnitude and the location of the break. They named it the energy of the change-point. They established the uniform detection threshold for the energy.} Unlike the classical change-point problem where the mean is piecewise constant, our problem involves a spectrum that can vary smoothly. The latter represents a local problem that cannot be addressed by standard sample-splitting methods. Our method is local in nature and so it is sub-optimal for the classical change-point problem but it is valid for the more general case of a piecewise smooth spectrum. Hence, for fixed shifts, the rate of convergence in our problem is slower. The smallest break magnitude allowed by Proposition (ref) is $\delta_{T}=O(\sqrt{\log\left(T\right)}/M_{S,T}^{1/2})$. Under this condition the convergence rate for the classical change-point estimator is $O_{\mathbb{P}}(M_{S,T}(T\log\left(T\right))^{-1})$ which is faster by a factor $O(m_{S,T}\sqrt{\log\left(T\right)})$ than the one suggested by Proposition (ref). In addition, in classical change-point setting $\delta_{T}\rightarrow0$ is allowed at a faster rate. This is obvious since in our setting a small break can be confounded with a smooth local change.

Under the smooth alternative $\mathcal{H}_{1}^{\mathrm{S}}$ the estimator is consistent when $\theta$-regularity is violated only once in the sample and also when the violation occurs in a small interval around $\lambda_{b}^{0}$ which does not exceeds $O(m_{S,T}\sqrt{M_{S,T}\log\left(T\right)}/T\delta_{T})$. If that interval is longer then this becomes a global problem which cannot be addressed by the estimation method considered in this section. This also relates to the discussion in Section (ref) that one cannot perfectly separate functions with $\theta$-smoothness from functions with $\theta'$-smoothness such that $\theta'<\theta.$

Multiple Breaks Alternatives $\mathcal{H}_{1}^{\mathrm{B},m_{0}}$

Let us assume that there are $m_{0}>1$ break points in $f\left(u,\,\omega\right)$. Let $0<\lambda_{1}^{0}<\ldots<\lambda_{m_{0}}^{0}<1$. We consider the following class of alternative hypotheses:

align*[align* omitted — 262 chars of source]

We provide a consistency result for both $m_{0}$ and the actual locations of the breaks $\lambda_{l}^{0}$ $(1\leq l\leq m_{0})$. Let $j^{*}$ be the largest integer such that $\left(M_{T}-j^{*}\right)m_{T}\leq\left(M_{T}-1\right)m_{T}-n_{T}$ and $\mathcal{I}\subseteq\left\{ 2m_{T},\,3m_{T},\ldots,\,\left(M_{T}-j^{*}\right)m_{T}\right\} $ denote a generic index set. One can test for a break at some time index in $\mathcal{I}$ by using the test $\psi^{*}(\{X_{r}\}_{r\in\mathcal{I}})$ based on $\max_{\omega_{k}\in\Pi'}\mathrm{S}_{\mathrm{max},T}\left(\omega_{k}\right)$ and if the test rejects one can estimate the break location using

align[align omitted — 223 chars of source]

We can then update the set $\mathcal{I}$ by excluding a $v_{T}$-neighborhood of $T\widehat{\lambda}_{T}$ and repeat the above steps. This is a sequential top-down algorithm exploiting the classical idea of bisection. However, this procedure may not be efficient. For example, consider the first step of the algorithm in which we test for the first break; this is associated with the largest break magnitude ($\delta_{1,T}>\delta_{l,T}$ for all $l=2,\ldots,\,m_{0}$). If the true break date $T_{1}^{0}$ falls in between two indices in $\mathcal{I}$, say $r_{1}$ and $r_{2}=r_{1}+m_{T}$, then this does not maximize either power or precision of the location estimate because one would need to compare two adjacent blocks exactly separated at $T_{1}^{0}$ but $T_{1}^{0}\notin\mathcal{I}$ since $T_{1}^{0}\in\left(r_{1},\,r_{2}\right).$ Hence, we introduce a wild sequential top-down algorithm.

Continuing with the above example, we draw randomly without replacement $K\geq1$ separation points $r^{\diamond}$ from the interval $\left(r_{1},\,r_{2}\right)$ and for each separation point compute $\mathrm{D}_{r^{\diamond},T}\left(\omega\right)$ where $r^{\diamond}\in\left(r_{1},\,r_{2}\right).$ We take the maximum value. Then, we update $\mathcal{I}$ by removing $r_{1}$ and adding $r^{\diamond}$. We repeat this for all indices in $\mathcal{I}$. Because the $K$ separation points are drawn randomly, there is always some probability to pick up the separation point that guarantees the highest power. A natural question is why not take all integers between $r_{1}$ and $r_{2}$ and compute $\mathrm{D}_{r^{\diamond},T}\left(\omega\right)$ for each. The reason is that in applications involving high frequency data (e.g., weakly, daily, and so on) that would be highly computationally intensive especially with multiple breaks as one wishes to change $m_{T}$ when searching for an additional break. This procedure exploits idea of bisection and combines it with a wild resampling technique similar to the one in fryzlewicz:14. The latter is characterized by using binary segmentation and drawing a large number of random intervals. Here the idea of using draws of random intervals is applied to the sequential top-down algorithm.

We are now ready to present the algorithm. Guidance as to a suitable choice of $K$ will be given below. Let $v_{T}\rightarrow\infty$ with $v_{T}/T\rightarrow0$ and $m_{T}/v_{T}\rightarrow0.$ Consider the test $\psi(\left\{ X_{t,T}\right\} ,\,\mathcal{I})=1$ if $\mathrm{S}_{\mathrm{Dmax},T}\left(\mathcal{I}\right)\geq2D^{*}\sqrt{\log\left(M_{T}^{*}\right)/m_{T}^{*}}$ where

align*[align* omitted — 324 chars of source]

with $D^{*}$, $m_{T}^{*}$ and $M_{T}^{*}$ as defined in Section (ref).

lyxalgorithmSet $\widehat{\mathcal{I}}=\{2m_{T},\,3m_{T},\ldots,\,\left(M_{T}-j^{*}\right)m_{T}\}$ and $\widehat{\mathcal{T}}=\emptyset$. \\ (1) For $r\in\widehat{\mathcal{I}}\backslash\{2m_{T}\}$ uniformly draw (without replacement) $K\in\{1,\ldots,\,m_{T}\}$ points $r_{k}^{\diamond}$ from $\mathbf{I}\left(r\right)=\{r-m_{T}+1,\ldots,r\}$ and compute $\overline{r}^{\diamond}=\arg\max_{k=1,\ldots,\,K}\max_{\omega\in\left[-\pi,\,\pi\right]}\mathrm{D}_{r_{k}^{\diamond},T}\left(\omega\right)$; set $\widehat{\mathcal{I}}=(\widehat{\mathcal{I}}\backslash\left\{ r\right\} )\cup\{\overline{r}^{\diamond}\}$. \\ (2) If $\psi(\left\{ X_{t,T}\right\} ,\,\widehat{\mathcal{I}})=0$ return $\widehat{\mathcal{T}}=\emptyset$. Otherwise proceed with step (3).\\ (3) Estimate the change-point $T\widehat{\lambda}_{T}(\widehat{\mathcal{I}})$ via (ref) using $\widehat{\mathcal{I}}$. \\ (4) Set $\widehat{\mathcal{I}}=\widehat{\mathcal{I}}\backslash\{T\widehat{\lambda}_{T}(\mathcal{\widehat{I}})-v_{T},\ldots,\,T\widehat{\lambda}_{T}(\mathcal{\widehat{I}})+v_{T}\}$ and $\widehat{\mathcal{T}}=\widehat{\mathcal{T}}\cup\{T\widehat{\lambda}_{T}(\mathcal{\widehat{I}})\}$. Return to step (1).

Finally, arrange the estimated change-points $\widehat{\lambda}_{l,T}$ in $\widehat{\mathcal{T}}$ in chronological order and use the symbol $\left|\mathcal{S}\right|$ for the cardinality of a set $\mathcal{S}$. To each $\widehat{\lambda}_{l,T}$ the procedure can return the frequency $\widehat{\omega}_{l}$ at which the break is found.

assumption$\delta_{l,T}=\delta_{l}\neq0$ is fixed or $\delta_{l,T}\rightarrow0$ with $\inf_{1\leq l\leq m_{0}}\delta_{l,T}\geq2D^{*}M_{S,T}^{-1/2}(\log(T))^{2/3}$. For $\nu_{T}\rightarrow\infty$ with $\nu_{T}=o\left(T/v_{T}\right),$ it holds that $\inf_{1\leq l\leq m_{0}-1}|\lambda_{l+1}^{0}-\lambda_{l}^{0}|\geq\nu_{T}^{-1}.$

Assumption (ref) allows for shrinking shifts and a possibly growing number of change-points as long as $m_{0}/\nu_{T}\rightarrow0.$ The following proposition presents the consistency result for the number of change-points $m_{0}$ and for the change-point locations $\lambda_{l}^{0}$ $\left(l=1,\ldots,\,m_{0}\right)$, and the rate of convergence of their estimates.

propLet Assumption (ref)-(ref), (ref) with $m_{0}=1$ and Condition (ref) hold. Then, under $\mathcal{H}_{1}^{\mathrm{B},m_{0}}$ we have (i) $\mathbb{P}(|\widehat{\mathcal{T}}-m_{0}|>\epsilon_{2})\rightarrow0$ for any $\epsilon_{2}>0$ and $\sup_{1\leq l\leq m_{0}}|\widehat{\lambda}_{l,T}-\lambda_{l}^{0}|=o_{\mathbb{P}}\left(1\right)$, and (ii) $\sup_{1\leq l\leq m_{0}}|\widehat{\lambda}_{l,T}-\lambda_{l}^{0}|=O_{\mathbb{P}}(m_{S,T}\sqrt{M_{S,T}\log\left(T\right)}/(T\inf_{1\leq l\leq m_{0}}\delta_{l,T}))$. Furthermore, if $K=O(a_{T}m_{T})$ with $a_{T}\in(0,\,1]$ such that $a_{T}\rightarrow1,$ then the breaks are detected in decreasing order of magnitude.

The number of draws $K$ may be fixed or increase with the sample size. However, the algorithm can return the change-point dates in decreasing order of the break magnitudes only if $K$ is sufficiently large. Note that at each loop of the algorithm it is not possible to know to which $\lambda_{l}^{0}$ $\left(l=1,\ldots,\,m_{0}\right)$ the estimate $\widehat{\lambda}_{l,T}$ is consistent for. Only after all breaks are detected and we rearrange the estimated change-points in $\widehat{\mathcal{T}}$ in chronological order, we can learn such information. The same procedure can be applied for the case of multiple smooth local changes, though the notation becomes cumbersome and so we omit it.

Implementation

In this section we explain how to choose the tuning parameters. The choice of $m_{T}$, $n_{T}$ and $b_{W,T}$ could be based on a mean-squared error (MSE) criterion or cross-validation exploiting results derived for locally stationary series {[}e.g., data-dependent methods for bandwidths in the context of locally stationary processes were investigated by, among others, casini_hac, dahlhaus:12, Dahlhaus/Giraitis:98 and ritcher/dahlhaus:19{]}. The optimal amount of smoothing depends on the regularity exponent $\theta$, on the boundness of the moments and on the extent of the dependence in $\left\{ X_{t,T}\right\} $. Here we choose the order of the bandwidths, neglecting the constants, by following the restrictions in Condition (ref). In particular, we choose the largest possible values allowed by Condition (ref) in order to ensure the highest possible power. We conduct a sensitivity analysis based on simulations in the supplement. We relegate to future work a more detailed analysis of data-dependent methods for this problem with multiple smoothing directions.

For spectral densities satisfying Lipschitz continuity, $\theta=1$ so that $m_{T}\propto T^{2/3-\epsilon}$ while for $\theta=1/2$ we have $m_{T}\propto T^{1/2-\epsilon}$ where in both cases $\epsilon>0.$ In applied work, it is common to work under stationarity ($\theta>1$) or local stationarity with Lipschitz smoothness ($\theta=1$). Hence, we use the bandwidths corresponding to $\theta=1$ which works for both cases. Of course, if one has prior knowledge about the smoothness of the parameters of the data-generating process, one can choose a suitable $\theta$. Assuming $q=8$ and $\gamma$ large enough we have $\tau_{T}\propto T^{1/4}$ and so values that satisfy Condition (ref) are $m_{T}=T^{0.66}$, $n_{T}=T^{0.62}$ and $b_{W,T}=n_{T}^{-1/6}$. The scaling is normalized to 1, as our simulations show this to provide good finite-sample properties, see Section (ref).

As for the tapering function $h\left(\cdot\right)$ and weight function $W\left(\cdot\right)$ we use a rectangular taper (i.e., $h\left(u\right)=1$ for all $u$) and a rectangular kernel. The rectangular kernel is known as the Daniell kernel with parameter $n_{T}b_{W,T}$ (it is a centered moving average which creates a smoothed value at time $Tu$ by averaging all values between $Tu-n_{T}b_{W,T}$ and $Tu+n_{T}b_{W,T}$). These are the simplest choices for $h\left(\cdot\right)$ and $W\left(\cdot\right)$. As for the bandwidth and kernel of the estimator $\widehat{\sigma}_{L,r}\left(\omega\right),$ we follow the results in casini_comment_andrews91 (casini_comment_andrews91, casini_hac) that suggest $b_{1,T}=M_{S,T}^{-1/3}.$ This corresponds to the MSE-optimal bandwidth when $K_{1}\left(\cdot\right)$ is the Bartlett kernel. For the choice of the number and values of the frequencies, the theory does not suggest particular values. Thus, we tried several values $n_{\omega}=15,\,11,\,7,\,5.$ Our default choice is $n_{\omega}=7$. A sensitivity analysis suggests that different choices for $n_{\omega}$ lead to negligible differences in the results. For the selection of the set of frequencies we use the function linspace that generates a linearly spaced sequence, e.g., in Matlab we used the command $\mathrm{\mathsf{linspace(-\pi+1e-3,\,0,\,(\mathit{n}_{\omega}+1)/2)}.}$

The regularity exponent $\theta$ also affects the test $\psi(\left\{ X_{t,T}\right\} ,\,\mathcal{I})$ in Algorithm (ref). It is possible to get an estimate of $\theta$ under the null as follows. Compute $\mathrm{S}_{\mathrm{Dmax},T}$ where the maximum is taken among the indices of the blocks such that the null hypothesis is not violated and label it $s_{\mathrm{Dmax}}^{*}$. Solve $s_{\mathrm{Dmax}}^{*}=2\sqrt{\log\left(M_{T}^{*}\right)/m_{T}^{*}}$ for $\theta$, where recall that $m_{T}^{*}$ and $M_{T}^{*}$ depend on $\theta.$ This yields a preliminary estimate of $\theta$ which can then be used for the test $\psi(\left\{ X_{t,T}\right\} ,\,\mathcal{I})$. Similarly, $b_{T}^{\mathrm{opt}}$ depends on $\theta$ and $\theta'$. Using the same approach, for a given $\theta'$ one can solve $s_{\mathrm{Dmax}}^{*}=b_{T}^{\mathrm{opt}}$ for $\theta$ as function of $\theta'.$ If one is interested in the alternative $\mathcal{H}_{1}^{\mathrm{B}}$, we have $\theta'=0$ and so this immediately yields an estimate for $\theta$. If one is interested in the alternative $\mathcal{H}_{1}^{\mathrm{S}}$, then one can try a few values of $\theta'$ in the range $\left(0,\,\theta\right)$. However, note that in order to use Algorithm (ref) only $\theta$ is needed. The knowledge of $\theta'$ under $\mathcal{H}_{1}^{\mathrm{S}}$ is only needed to obtain $b_{T}^{\mathrm{opt}}.$

We set $v_{T}=T^{0.666}$ which satisfies $m_{T}/v_{T}\rightarrow0.$ Our default recommendation is $K=10.$ Our simulations with different data-generating processes and sample sizes show that this choice strikes a good balance between the precision of the change-point estimates and computing time. For $T>1000$, we recommend setting $K=\left\lfloor m_{T}/3\right\rfloor $.

The test statistics $\mathrm{S}_{\mathrm{max},T}\left(\omega\right)$ and $\mathrm{R}_{\mathrm{max},T}\left(\omega\right)$ depend on $\omega.$ The choice of $\omega$ is, of course, important as it involves different frequency components and hence different periodicities. If the user does not have a priori knowledge about the frequency at which the spectrum has a change-point, our recommendation is to run the tests for multiple values of $\omega\in\left[0,\,\pi\right]$. Even if the change-point occurs at some $\omega_{0}$ and one selects a value of $\omega$ close but not equal to $\omega_{0}$ the tests are still able reject the null hypothesis given the differentiability of $f\left(u,\,\omega\right)$. Thus, one can select a few values of $\omega$ evenly spread on $\left[0,\,\pi\right]$.

Small-Sample Evaluations

In this section, we conduct a Monte Carlo analysis to evaluate the properties of the proposed methods. We first discuss the detection of the change-points and then their localization. We investigate different types of changes and consider the test statistics $\mathrm{S}_{\max,T}\left(\omega\right),$ $\mathrm{S}_{\mathrm{Dmax},T}$, $\mathrm{R}_{\max,T}\left(\omega\right)$, $\mathrm{R}_{\mathrm{Dmax},T}$ proposed here and the test statistic $\widehat{D}$ proposed by last/shumway:08. The latter is included for comparison since it applies to the same problems. We consider the following data-generating processes where in all models the innovation $e_{t}$ is a Gaussian white noise $e_{t}\sim\mathrm{i.i.d.}\,\mathscr{N}\left(0,\,1\right)$. Models M1 involves a stationary AR(1) process $X_{t}=\rho X_{t-1}+e_{t}$ with $\rho=0.3$ and 0.6, while M2 involves a locally stationary AR(1) $X_{t}=\rho\left(t/T\right)X_{t-1}+e_{t}$ where $\rho\left(t/T\right)=0.4\cos\left(0.8-\cos\left(2t/T\right)\right).$ Note that $\rho\left(t/T\right)$ varies smoothly from 0.1389 to 0.3920. Model M1 and M2 are used to verify the finite-sample size of the tests. We verify the power in models M3 and M4 using the specification in model M1 and M2, respectively, for the first regime and consider two additional regimes with different specifications. Hence, two breaks are present. In model M3,

align*[align* omitted — 328 chars of source]

while, for model M4

align*[align* omitted — 365 chars of source]

where $\rho\left(t/T\right)$ is as in model M2. In model M3, the second regime involves higher serial dependence while in the third regime the variance doubles relative to the second regime. In model M4, the second regime involves a stationary autoregressive process with strong serial dependence while in the third regime $X_{t}$ assumes the same dynamics as in the first regime. Models M3-M4 feature alternative hypotheses in the forms of breaks in the spectrum.

We consider the alternative hypothesis of more rough variation without signifying a break (i.e., $\mathcal{H}_{1}^{\mathrm{S}}$ defined in Section (ref)) in model M5 given by $X_{t}=\sigma\left(t/T\right)e_{t}$ where $\sigma^{2}\left(t/T\right)=\max\{1.5,\,\overline{\sigma}^{2}+\cos\left(1+\cos\left(10t/T\right)\right)\}$ with $\overline{\sigma}^{2}=1.$ Note that even though $\sigma^{2}\left(\cdot\right)$ is locally stationary, the degree of smoothness alternates throughout the sample. It starts from $\sigma^{2}\left(\cdot\right)=1.5$ and maintains this value for some time, then within a short period it increases slowly to $\sigma^{2}\left(\cdot\right)=2$ and falls slowly back to $\sigma^{2}\left(\cdot\right)=1.5$. It keeps this value until the final part of the sample where it increases slowly to $\sigma^{2}\left(\cdot\right)=2$ for a short period. Thus, $\sigma^{2}\left(\cdot\right)$ alternates between periods where it is constant (i.e., $\theta>1$) and periods where it becomes non-constant but less smooth (i.e., $\theta=1$). Importantly, no break occurs; only a change in the smoothness as specified in $\ensuremath{\mathcal{H}_{1}^{\mathrm{S}}}$. In unreported simulations we also considered the case where $\theta$ changes from Lipschitz continuity (i.e., $\theta=1$) to the continuity-path of Wiener processes (i.e., $\theta\thickapprox1/2$) with results that are similar to those reported here. For the test statistic $\widehat{D}$ of last/shumway:08, we obtain the critical value by simulations. As suggested by the authors we compute the finite-sample distribution of $\widehat{D}$ by simulating a white noise under the null hypotheses with a sample size $T=1000$ and then obtain the critical value. We consider the three sample sizes $T=250,\,500$ and 1000. The significance level is $\alpha=0.05$. For the test statistics $\mathrm{S}_{\mathrm{max},T}\left(\omega\right)$ and $\mathrm{R}_{\mathrm{max},T}\left(\omega\right)$, we use as a default value $\omega=0$ given that the interest is often in low frequency analysis. We set $\lambda_{1}^{0}=0.33$ and $\lambda_{2}^{0}=0.66$ throughout. The number of simulations is 5,000 for all cases.

The results are reported in Table (ref)-(ref). We first discuss the size of the tests. The tests proposed in this paper have good empirical size for both models and all sample sizes. The test statistics $\mathrm{S}_{\mathrm{Dmax},T}$ and $\mathrm{R}_{\mathrm{Dmax},T}$ are slightly undersized for $T=250$ but their empirical size improves for $T=500$ and 1000. The test statistics $\mathrm{S}_{\mathrm{max},T}$ and $\mathrm{R}_{\mathrm{max},T}$ share accurate empirical sizes in all cases. In contrast, the test statistic $\widehat{D}$ of last/shumway:08 is largely oversized for $T=250$ and 500. For $T=1000$ it works better but it is still oversized. This means that the finite-sample distribution of $\widehat{D}$ has high variance and changes substantially across different sample sizes. Since the simulated critical value is obtained with a sample size $T=1000$ it works better for this sample size than for the others for which the size control is poor.

Turning to the power of the tests, we note that it is not fair to compare the proposed tests with the test $\widehat{D}$ when $T=250$ and 500 since the latter is largely oversized in those cases. In model M3, all the proposed tests have good power which increases with the sample size. The tests $\mathrm{S}_{\mathrm{Dmax},T}$ and $\mathrm{R}_{\mathrm{max},T}$ have the highest power, followed by $\mathrm{S}_{\mathrm{max},T}$ and lastly $\mathrm{R}_{\mathrm{Dmax},T}$. The power differences are not large except those involving $\mathrm{R}_{\mathrm{Dmax},T}$ for $T=250$ which has substantially lower power. It is important to note that for $T=1000$ the proposed tests have higher power than the test $\widehat{D}$ of last/shumway:08 even though the latter is oversized. For $T=250$ and 500, where the $\widehat{D}$ test is largely oversized, the proposed tests only have slightly lower power. This confirms that the proposed tests have very good power. Similar comments apply to model M4.

Model M5 involves changes in the smoothness without involving a break. This constitutes a more challenging alternative hypothesis, and as expected, the power for each test is lower than in models M3-M4. The test with the highest power is $\mathrm{S}_{\mathrm{Dmax},T}$. For $T=1000$, the test with the lowest power is $\widehat{D}$ (for $T=250$ the $\widehat{D}$ test has higher power again due to its oversize problem). Overall, the results show that the proposed tests have accurate empirical size even for small sample sizes and have good power against different forms of breaks or changes in “roughness”.

Next, we consider the estimation of the number of change-points ($m_{0}$) and their locations. We consider the following two models, both with $m_{0}=2$. The model M6 is given by

align*[align* omitted — 320 chars of source]

while model M7 is the same as model M4. We set $\lambda_{1}^{0}=0.33$ and $\lambda_{2}^{0}=0.66$ and $T=1000$ throughout. Table (ref) reports summary statistics for $\widehat{m}-m_{0}$. It displays the percentage of times with $\widehat{m}=m_{0}$, the median, and the 25% and 75% quantile of the distribution of $\widehat{m}$. We only consider Algorithm (ref). We do not report the results for the corresponding procedure of last/shumway:08 because it is based on $\widehat{D}$ which is oversized and so it finds many more breaks than $m_{0}$. Table (ref) shows that $\widehat{m}=m_{0}$ occurs for about 85% of the simulations with model M6 and about 80% with model M7. This suggests that Algorithm (ref) is quite precise. As expected it performs better in model M6 since the specification of the alternative is farther from the null. The quantiles of the empirical distribution also suggest that the change-point estimates $\widehat{T}_{1}$ and $\widehat{T}_{2}$ are accurate. For example the median is very close to the their respective true value $T_{1}^{0}=333$ and $T_{2}^{0}=666$. Similar conclusions arise from different models and sample sizes, in unreported simulations.

Empirical Application

We demonstrate how to use our change-point methods for studying the causal effects of monetary policy. A fast growing literature in macroeconomics uses high-frequency data to identify the effects of monetary policy on the real economy (i.e., money non-neutrality). The identifying assumption used by nakamura/steinsson:2018 is that the volatility of the daily change in the nominal 2-year Treasury yields, say $\Delta i_{t}$, is higher during days when the Federal Open Market Committee (FOMC) meets to make monetary policy announcements relative to regular Tuesdays and Wednesdays with no announcement. Let $\eta_{t}$ denote a pure monetary shock and suppose that the policy instrument $\Delta i_{t}$, which is observed in the data, is governed by both monetary and non-monetary shocks:

align*[align* omitted — 63 chars of source]

where $\varepsilon_{t}$ is a function of all other shocks that affect $\Delta i_{t}$ and $\mu_{i}$ is a constant. We normalize the impact of $\eta_{t}$ and $\varepsilon_{t}$ on $\Delta i_{t}$ to one. $\Delta i_{t}$ is a measure of the monetary policy news revealed in the FOMC announcement. The idea is that changes in the policy instrument during days when there is a FOMC announcement are dominated by the information about future monetary policy contained in the announcement. Let $\Delta s_{t}$ denote the change in the outcome variable which is the yield on a five year zero-coupon Treasury bond. We wish to estimate the effects of the monetary shock $\eta_{t}$ on the outcome variable $\Delta s_{t}$. The latter is also affected by both the monetary and non-monetary shocks:

align*[align* omitted — 74 chars of source]

where $\mu_{s}$ is a constant and $\alpha$ and $\beta$ are two parameters. The parameter of interest is $\beta$ which represents the impact of the pure monetary shock $\eta_{t}$ on $\Delta s_{t}$ relative to its impact on $\Delta i_{t}$.

The identifying assumption is that the variance of monetary shocks increases during days of FOMC announcements, while the variances of the other shocks are unchanged. Let $T_{P}$ denote the number of days containing a FOMC announcement, and let $T_{C}$ the number of days with no such announcements. The subscript “$P$” in $T_{P}$ refers to the \textquotedblleft policy or treatment\textquotedblright sample while the subscript “$C$” in $T_{C}$ refers to the \textquotedblleft control\textquotedblright sample. The days in the control sample are comparable on other dimensions since they are all Tuesdays and Wednesdays with no FOMC meeting. The identifying restriction can then be written as

align[align omitted — 141 chars of source]

where $\sigma_{a,i}$ is the volatility of variable $a=\eta,\,\varepsilon$ in the sample $i=P,\,C$. One can show that

align[align omitted — 244 chars of source]

where $\mathrm{Cov}_{i}(\cdot,\,\cdot)$ (resp. $\mathrm{Var}_{i}(\cdot,\,\cdot)$) denotes the population covariance (resp. variance) in the $i=P,\,C$ sample. The parameter $\beta$ can be identified only if $\mathrm{Var}_{P}\left(\Delta i_{t}\right)\neq\mathrm{Var}_{C}\left(\Delta i_{t}\right)$. If the volatility of the policy instrument $\Delta i_{t}$ does not change across treatment and control samples, then $\beta$ is not identified. If the change in volatility is small, then $\beta$ is weakly identified.

Subsequent developments in the literature employed robust weak identification tests to argue that $\beta$ is not strongly identified when using daily data. In contrast, $\beta$ can be strongly identified if one uses ultra high-frequency data based on a 30 minute window around the announcement time. For the case of daily data, we show that the volatility of $\Delta i_{t}$ in the control sample varies substantially over time and that several change-points can be detected. Thus, the regimes in the control sample where the volatility is high contribute to an average (over the full control sample) volatility that approaches the average volatility in the treatment sample, thereby violating $\mathrm{Var}_{P}\left(\Delta i_{t}\right)\neq\mathrm{Var}_{C}\left(\Delta i_{t}\right)$. This implies that $\beta$ may be weakly identified and its estimates may be imprecise which supports the recent evidence in the literature.

We obtain the data from Emi Nakamura's webpage. The sample of “treatment” days is all regularly scheduled FOMC meeting day from 1/1/2000 to 3/19/2014. The sample of “control” days is all Tuesdays and Wednesdays that are not FOMC meeting days from 1/1/2000 to 12/31/2012. In both the treatment and control samples, the second half of 2008, the first half of 2009 and a 10 day period after 9/11/2001 are dropped in nakamura/steinsson:2018. We follow the same practice.

The plot of $\left\{ \Delta i_{t}\right\} $ for the control sample is reported in Figure (ref). The series displays substantial changes in volatility and some changes in persistence. The tests $\mathrm{R}_{\mathrm{max},T}$ and $\mathrm{R}_{\mathrm{Dmax},T}$ strongly reject the null hypothesis of no change-points in the spectrum. Algorithm (ref) detects three change-points.

The first change-point date, denoted $\widehat{T}_{1}$, corresponds to April 24, 2007. Thus, the first regime $[1,\,\widehat{T}_{1}]$ refers to the period prior to the beginning of the 2007-09 financial crisis. It is evident that a large change in volatility and possibly a change in persistence occurred. We note that the change in volatility does not occur abruptly, it is rather gradual. This means that the volatility path changes gradually and possibly becomes more rough. This supports the usefulness of our hypothesis testing framework with local stationarity under the null hypothesis and with changes in the smoothness of the parameters that govern the data-generating process under the alternative hypothesis. The first change-point $\widehat{T}_{1}$ is associated to the volatility path becoming more rough with the level of the volatility increasing gradually.

The second change-point date, denoted $\widehat{T}_{2}$, corresponds to July 28, 2009. Thus, the second regime $[\widehat{T}_{1}+1,\,\widehat{T}_{2}]$ includes roughly the 2007-09 financial crisis. In this regime the volatility of the series is remarkably high. After $\widehat{T}_{2}$ the series shares a pattern similar to that in the first regime $[1,\,\widehat{T}_{1}]$, both in terms of persistence and volatility. The second change-point date $\widehat{T}_{2}$ is associated to an abrupt fall in volatility.

The third change-point date, denoted $\widehat{T}_{3}$, corresponds to February 2, 2011. The third regime $[\widehat{T}_{2}+1,\,\widehat{T}_{3}]$ corresponds to a zero lower bound (ZLB) period when the FOMC announced the conduct of unconventional monetary policies to stimulate the economy after the crisis. The ZLB refers to a situation in which the short-term nominal interest rate is at or near zero, limiting the central bank's capacity to stimulate growth. Unconventional monetary policies refer to large-scale asset purchases and active use of communication (i.e., “forward guidance”) to shape expectations about future monetary policies, which can mitigate the limitations imposed by the ZLB {[}see, e.g., swanson:2021{]}.

The last regime, $[\widehat{T}_{3}+1,\,T_{C}]$, corresponds to the period when the economy was witnessing the first effects of the expansive monetary policy and of the stability of the unconventional monetary policies that were introduced previously. This is a regime where initially the economy started the recovery and then reached stable economic growth. In this regime, the volatility level of the series is the lowest of the sample.

Overall, the first change-point date, $\widehat{T}_{1}$, corresponds to a change in the smoothness while the second change-point date, $\widehat{T}_{2}$, corresponds to an abrupt break. For the third change-point date, it is more difficult to tell from the plot whether this corresponds to an abrupt or smooth break.

We now discuss how the results about the change-points can be useful for the identification issue based on (ref)-(ref). One should think of $\mathrm{Var}_{C}\left(\Delta i_{t}\right)$ as the average variance of $\Delta i_{t}$ in the control sample. Our results show that there is significant time variation in $\mathrm{Var}_{C}\left(\Delta i_{t}\right)$. In the second regime, $[\widehat{T}_{1}+1,\,\widehat{T}_{2}]$, the volatility is the highest of the sample. During this period, it is lower but very close to the average volatility of $\Delta i_{t}$ in the treatment sample, $\mathrm{Var}_{P}\left(\Delta i_{t}\right)$.\footnote{We refer to nakamura/steinsson:2018 for details about the series in the treatment sample. We do not test for change-points in the treatment sample because the sample size in the treatment sample is relatively small ($T_{P}=74)$ and so we treat it as a single regime.} This contributes to make $\mathrm{Var}_{P}\left(\Delta i_{t}\right)-\mathrm{Var}_{C}\left(\Delta i_{t}\right)$ (over the full sample) closer to zero which then would lead to weak identification. In fact, nakamura/steinsson:2018 found the estimate of $\beta$ to be imprecise and not meaningful from an economic standpoint. Furthermore, it was quite different from that obtained using 30 minute data instead of daily data. Our change-point analysis is useful because it suggests which periods or sub-samples contribute to this identification problem and which sub-samples can be used to obtain consistent estimates for $\beta.$

Conclusions

We develop a theoretical framework for inference about the smoothness of the spectral density over time. We provide frequency-domain statistical tests for the detection of discontinuities in the spectrum of a segmented locally stationary time series and for changes in the regularity exponent of the spectral density over time. The null distribution of the test follows an extreme value distribution. We rely on the theory on minimax-optimal testing developed by ingster:93. We determine the optimal rate for the minimax distinguishable boundary, i.e., the minimum break magnitude such that we are still able to uniformly control type I and type II errors. We propose a novel procedure to estimate the change-points based on a wild sequential top-down algorithm and show its consistency under shrinking shifts and possibly growing number of change-points. The advantage of using frequency-domain methods to detect change-points is that it does not require to make assumptions about the data-generating process under the null hypothesis beyond the fact that the spectrum is differentiable and bounded. Furthermore, the method allows for a broader range of alternative hypotheses compared to time-domain methods which usually have power against a limited set of alternatives. Overall, our simulations and empirical results show the usefulness of our method.

\addcontentsline{toc}{section}{References}

Appendix

Tables

table[table omitted — 1,137 chars of source]
table[table omitted — 1,591 chars of source]
table[table omitted — 637 chars of source]
center[center omitted — 642 chars of source]

\pagenumbering{arabic}