The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
127,866 characters
Tests for Forecast Instability and Forecast Failure under a Continuous Record Asymptotic Framework
\begin{titlepage}
\center
\textsc{\Large \textcolor{MyLigthBlack}{research paper and supplemental material}}\\[0.5cm]
\rule{\textwidth}{1.6pt}\vspace*{-\baselineskip}\vspace*{2pt} \rule{\textwidth}{0.4pt}\\[\baselineskip] {\huge \bfseries \textcolor{MyLightRed2!75!white}{TESTS FOR FORECAST INSTABILITY AND FORECAST FAILURE UNDER A CONTINUOUS RECORD ASYMPTOTIC FRAMEWORK}\\[0.2\baselineskip] \rule{\textwidth}{0.4pt}\vspace*{-\baselineskip}\vspace{3.2pt} \rule{\textwidth}{1.6pt}\\[\baselineskip]}
\bigskip
\bigskip
\bigskip
\centering \large \textbf{\textcolor{MyLigthBlue13!70!white}{ALESSANDRO CASINI}}
\\\textcolor{MyLigthBlack}{Department of Economics}
\\\textcolor{MyLigthBlack}{Boston University}
\bigskip
\bigskip
\bigskip
\textcolor{MyLigthBlack}{\normalfont{\today}
\\
\bigskip
\bigskip
\bigskip
\smallskip
}
\vfill
\end{titlepage}
\noindent
\textbf{Correspondence}:
\medskip\\
\begin{minipage}{0.8\textwidth}
\begin{flushleft} \normalfont \textbf{\textcolor{MyLigthBlue13!70!white}{Alessandro Casini}}
\\Department of Economics
\\Boston University
\\270 Bay State Road
\\Boston, MA 02215, US
\\Email:
\texttt{\textcolor{MyLigthBlue13!70!white}{{[email removed]}}}
\\Webpage:
\textcolor{MyLigthBlue13!70!white}{{\url{http://alessandro-casini.com}}}
\end{flushleft}
\end{minipage}
\medskip
\bigskip
\noindent\\
\textbf{Ducument Information}:
This electronic document includes the working paper and supplemental materials. The Supplement starts on pdf page 48. Monte Carlo simulations were conducted in \textsc{Matlab}. The most recent working paper version can be downloaded from the author's webpage. The associated computer package, if available, should be downloaded from the same webpage.
\smallskip
\begin{singlespace}
\noindent\\
\footnotesize{}{
\textbf{Ducument History}}\\
First Version: \today
\medskip
\noindent\\
\textbf{Abstract}:
We develop a novel continuous-time asymptotic framework for inference on whether the predictive ability of a given forecast model remains stable over time. We formally define forecast instability from the economic forecaster's perspective and highlight that the time duration of the instability bears no relationship with stable period. Our approach is applicable in forecasting environment involving low-frequency as well as high-frequency macroeconomic and financial variables. As the sampling interval between observations shrinks to zero the sequence of forecast losses is approximated by a continuous-time stochastic process (i.e., an It� semimartingale) possessing certain pathwise properties. We build an hypotheses testing problem based on the local properties of the continuous-time limit counterpart of the sequence of losses. The null distribution follows an extreme value distribution. While controlling the statistical size well, our class of test statistics feature uniform power over the location of the forecast failure in the sample. The test statistics are designed to have power against general form of insatiability and are robust to common forms of nonstationarity such as heteroschedasticity and serial correlation. The gains in power are substantial relative to extant methods, especially when the instability is short-lasting and when occurs toward the tail of the sample.
\medskip
\noindent\\
\textbf{JEL Classification}: C10, C22\\
\textbf{Keywords:} Asymptotic distribution, break date, continuous-time, forecast failure, forecast instability, infill asymptotics, parameter instability, predicitve ability, semimartingale\\
\medskip
\noindent\\
\textbf{Acknowledgements}: I am indebted to Pierre Perron for introducing me to the issue of structural changes in economics and for sharing insightful discussions on the topic. I am grateful to Raffaella Giacomini for helpful disccusions and to Viktor Todorov for comments on the continous-time approach. I thank Simon Gilchrist for useful remarks on the GZ credit spread index from \citet{gilchrist/zakrajsek:12}. I thank Raffaella Giacomini and Barbara Rossi for making their programs available online. Partial financial support from the Bank of Italy is gratefully acknowledged.
\end{singlespace}
\thispagestyle{empty}
\pagebreak{}
\setcounter{page}{1}
\raggedbottom
\title{\textbf{Tests for Forecast Instability and Forecast Failure under
a Continuous Record Asymptotic Framework}\thanks{I am indebted to Pierre Perron for introducing me to the issue of
structural changes in economics and for sharing insightful discussions
on the topic. I am grateful to Raffaella Giacomini for helpful disccusions
and to Viktor Todorov for comments on the continous-time approach.
I thank Simon Gilchrist for useful remarks on the GZ credit spread
index from \citet{gilchrist/zakrajsek:12}. I thank Raffaella Giacomini
and Barbara Rossi for making their programs available online. Partial
financial support from the Bank of Italy is gratefully acknowledged.}}
\maketitle
\begin{abstract}
We develop a novel continuous-time asymptotic framework for inference
on whether the predictive ability of a given forecast model remains
stable over time. We formally define forecast instability from the
economic forecaster's perspective and highlight that the time duration
of the instability bears no relationship with stable period. Our approach
is applicable in forecasting environment involving low-frequency as
well as high-frequency macroeconomic and financial variables. As the
sampling interval between observations shrinks to zero the sequence
of forecast losses is approximated by a continuous-time stochastic
process (i.e., an It� semimartingale) possessing certain pathwise
properties. We build an hypotheses testing problem based on the local
properties of the continuous-time limit counterpart of the sequence
of losses. The null distribution follows an extreme value distribution.
While controlling the statistical size well, our class of test statistics
feature uniform power over the location of the forecast failure in
the sample. The test statistics are designed to have power against
general form of insatiability and are robust to common forms of non-stationarity
such as heteroskedasticty and serial correlation. The gains in power
are substantial relative to extant methods, especially when the instability
is short-lasting and when occurs toward the tail of the sample.
\end{abstract}
\indent {\bf JEL Classification:} C12, C22, C52, C53 \\
\noindent{\bf Keywords:} Asymptotic distribution, break date, continuous-time, forecast failure, forecast instability, infill asymptotics, parameter instability, predicitve ability, semimartingale
\onehalfspacing
\pagebreak{}
\section{Introduction }
Since the seminal contribution of Klein \citeyearpar{klein:69,klein:71},
economic forecasts had been built upon the presumption that the relationships
between economic variables remain stable over time. However, the last
decades have been subject to many social-economic episodes and technological
advancements that have led economists to reconsider the assumption
of model stability. The resonant empirical evidences documented in,
among others, \citet{perron:89} and \citet{stock/watson:96} {[}see
also the recent survey by \citet{ng/wright:13}{]} have motivated
the development of econometric methods that detect such instabilities\textemdash most
work directed toward structural changes\textemdash and estimate the
actual dates at which economic relationships change. Yet, the issue
of parameter insatiability is not limited to model estimation. In
the forecasting literature, there has been a widespread concordance
that the major issue that prevents good forecasts for economic variables
is parameter instability\textemdash and structural changes as a special
case\textemdash {[}cf. \citet{banerjee/marcellino/masten:08}, \citeauthor{clements/hendry:98}
(1998, 2006)\nocite{clements/hendry:06}, \citet{elliott/timmermann:16},
\citet{giacomini:15}, \citet{giacomini/rossi:15}, \citet{inoue/rossi:11},
\citet{clark/mccracken:05}, \citet{pesaran/pettenuzzo/timmermann:06}
and \citet{rossi:13a}{]}.
This paper develops a statistical setting under infill asymptotics
to address the issue of testing whether the predictive ability of
a given forecast model remains stable over time. \citet{ng/wright:13}
and \citet{stock/watson:03} explain that there has been abundant
evidence for which a predictor that has performed well over a certain
time period may not perform as well during other subsequent periods.
For example, \citet{gilchrist/zakrajsek:12} proposed a new credit
spread index and showed that a residual component labeled as the excess
bond premium\textemdash the credit spread adjusted for expected default
risk\textemdash has considerable predictive content for future economic
activity. They documented that this forecasting ability is stronger
over the subsample 1985-2010 rather than over the full sample starting
from 1973.\footnote{They reported that structural change tests provide some statistical
evidence for a break in a coefficient associated with financial indicators\textemdash more
specifically the coefficient on the federal funds rate. Given the
latter evidence and the well-documented change in the conduct of monetary
policy in the late 1970s and the early 1980s, it seems plausible to
split the sample in 1985 (see p. 1709 and footnote 11 in their paper).} The latter finding can be attributed to a more developed bond market
in the 1985-2010 subsample. Relatedly, \citet{giacomini/rossi:10}
and \citet{ng/wright:13} further examined this finding and found
that indeed the predictive ability of commonly used term and credit
spreads is unstable and somehow episodic. The latter authors suggested
that credit spreads may be more useful predictors of economic activity
in a more highly leveraged economy and that recent developments in
financial markets translate into credit spreads containing more information
than they had previously. We refer to such temporal instability for
a given forecasting method as \textit{forecast instability} or more
specifically, as \textit{forecast failure}. These terminologies are
not new to professional forecasters as they were informally introduced
by \citet{clements/hendry:98} and generalized in econometric terms
by \citet{giacomini/rossi:09} who interpreted \textit{forecast breakdown}
(or \textit{forecast failure}) as a situation in which the out-of-sample
performance of a forecast model significantly deteriorates relative
to its in-sample performance. Our approach is to formally define forecast
instability from the economic forecaster's perspective.\footnote{We use the terminology ``instability'' because not only the deterioration
but also the improvement of the performance of a given forecast model
over time can provide useful information to the forecaster.} We emphasize that a forecast failure may well result from a short
period of instability within the out-of-sample and not necessarily
require that the instability be systematic in the sense of persisting
throughout the whole out-of-sample period. That is, consistency of
a forecast model's performance with expected performance given the
past should hold not only throughout the out-of-sample but also in
any sub-sample of the latter. Indeed, many documented episodes of
forecast failure seemed to arise from parameters nonconstancy data-generating
processes over relatively short time periods compared to the total
sample size. Hence, the desire of focusing on statistical tests being
able to detect short-lasting instabilities is intuitive: if a test
for forecast failure needs the deterioration of the forecasting ability
to last for, say, at least half of the total sample in order to have
sufficiently high power to reject the null hypotheses, then this test
would not perform very well in practice because instability can be
short-lasting. Furthermore, the occurrence of recurrent structural
instabilities or multiple breaks that compensate each other in the
out-of-sample might lead a forecast model to perform, on average,
in a similar fashion as in the in-sample period. However, should a
forecaster know about those recurrent changes she would conceivably
revise its forecast model to adapt to the unstable environment. Hence,
we introduce the following definition.
\begin{defn}
(Forecast Instability)
Forecast Instability refers to a situation of either sustained deterioration
or improvement of the predictive ability of a given forecast model
relative to the historical performance that would had led a forecaster
to revise or reconsider its forecast model if she had known the occurrence
of such instability. The time lengths of these two distinct periods
need not bear any relationship.\footnote{\textit{Forecast Failure} constitutes a special case of the definition\textemdash namely,
a sustained deterioration of predictive ability.}
\end{defn}
Th definition poses at the center the economic forecaster and consequently
it is not merely a statistical definition; rather, it is based on
an equilibrium concept. Since forecasting constitutes a decision theoretic
problem, it should be from the forecaster perspective that a given
forecast model is deemed to have failed. It is implicit from the definition
to distinguish between forecasting method and model. Two forecasters
may share the same forecast model\textemdash the relationship between
the variable of interest and the predictor\textemdash but use different
methods (e.g., recursive scheme versus rolling scheme). Thus, instability
refers to a given method-model pair. The object of the definition
is predictive ability. Since the latter can be measured differently
by different loss functions, then the definition applies to a given
choice of the loss function. A notable aspect of the definition is
the reference to the time span of the historical performance and of
the putative period of instability. They need not be related. Consider
a given forecasting strategy which has performed well during, say,
the Great Moderation (i.e., from mid-1980s up to prior the beginning
of the Great Recession in 2007). Assume that during the years 2007-2012
this method endures a time of poor performance and returns to perform
well thereafter. According to our definition, this episode constitutes
an example of forecast instability. However, if one designs the forecasting
exercise in such a way that half of the sample is used for estimation
and the remaining half for prediction, then this relatively short
period of instability gets ``averaged-out'' from tests which simply
compare the in-sample and out-of-sample averages. Conceivably, such
tests would not reject the null hypotheses of no forecast failure
while it seems that a forecaster would had revised its strategy during
the crisis if she had known about such occurring under-performance
in the present and immediate future period. Finally, detection of
forecast instability does not necessarily mean that a forecast model
should be abandoned. In fact, its performance may have improved over
time. Yet, even if forecast instability is induced by performance
deterioration, a forecaster might not end up switching to a new predictor.
For example, entering a state of high variability might lead to poor
performance even if the forecast model is still correct. Hence, our
definition uses the term \textit{reconsider}. Continuing with the
above example, a forecaster may \textit{reconsider} the choice of
the forecasting window since a longer window may now produce better
forecasts while keeping the \textit{same} forecast model. In other
words, knowledge of forecast instability is important because indicates
that care must be exercised to assess the source of the changes.\footnote{Economists have documented episodes of forecast failure in many areas
of macroeconomics. In the empirical literature on exchange rates a
prominent forecast failure is associated with the Meese and Rogoff's
puzzle {[}cf. \citet{meese/rogoff:83}, \citet{cheung/chinn/garcia:05},
and \citet{rossi:13b} for an up-to-date account{]}. In the context
of inflation forecasting, forecast failures have been reported by
\citet{atkeson/ohanian:01} and \citet{stock/watson:09}. For forecast
instability concerning other macroeconomic variables see the surveys
of \citet{stock/watson:03} and \citet{ng/wright:13}.}
The theoretical implication is that in this paper our tests for forecast
instability shall be based on the local behavior of the sequence of
realized forecast losses. This is opposite to existing tests for forecast
instability\textemdash and classical structural change tests more
generally\textemdash which instead rely on a global and retrospective
methodology merely comparing the average of in-sample losses with
the average of out-of-sample losses. While maintaining approximately
correct nominal size, our class of test statistics achieves substantial
gains in statistical power relative to previous methods. Furthermore,
as the initial timing of the instability moves away from middle sample
toward the tail of the out-of-sample, the gains in power become considerable.
In this paper, we set out a continuous record asymptotic framework
for a forecasting environment where $T$ observations at equidistant
time intervals $h$ are made over a fixed time span $\left[0,\,N\right],$
with $N=Th.$ These observations are realizations from a continuous-time
model for the variable to be forecast and for the predictor. From
these discretely observed realizations we compute a sequence of forecasts
using either a fixed, recursive or rolling scheme. To this sequence
of forecasts there corresponds a continuous-time process which satisfies
mild regularity conditions and that under the null hypotheses possesses
a continuous sample-path. We exploit this pathwise property to base
an hypothesis testing problem on the relative performance of a given
forecast model over time. Under the hypotheses we expect the sequence
of losses to display a smooth and stable path. Any discontinuous or
jump behavior followed by a (possibly short) period of substantial
discrepancy from the same path over the in-sample period provides
evidence against the hypotheses. Our asymptotic theory involves a
continuous record of observations where we let the sample size $T$
grow to infinity by shrinking the sampling interval $h$ to zero with
the time span kept fixed at $N$, thereby approaching the continuous-time
limit.
Our underlying probabilistic model is specified in terms of continuous
It� semimartingales which are standard building blocks for analysis
of macro and financial high-frequency data {[}cf. \citet*{andersen/bollerslev/diebold/labys:01},
\citet{andersen/fusari/todorov:16}, \citet{bandi/reno:16} and \citet{barndorff/shephard:04}{]};
the theoretical methodology is thus related to that of \citet{casini/perron_CR_Single_Break},
\citet{li/todorov/tauchen:17}, \citet{li/xiu:16} and \citet{mykland/zhang:09}.\footnote{Recent work by \citet{li/patton:17} extends standard methods for
testing predictive accuracy of forecasts to a high-frequency financial
setting.} The framework is not only useful for high-frequency data; in particular,
recent work of \citeauthor{casini/perron_CR_Single_Break} (2017a,
2017b)\nocite{casini/perron_Lap_CR_Single_Inf} has adopted this continuous-time
approach for modeling time series regression models with structural
changes fitted to low-frequency data (e.g., macroeconomic data that
are sampled at weekly, monthly, quarterly, annual frequency, etc.).
They have showed that this continuous-time approach delivers a
better approximation to the finite-sample distributions of estimators
in structural change models and inference is more reliable than previous
methods based on classical long-span asymptotics.
The classical approach to economic forecasting for macroeconomic variables
is to formulate models in discrete-time and then base inference on
long-span asymptotics where the sample size increases without bound
and the sampling interval remains fixed {[}cf. \citet{diebold/mariano:95},
\citet{giacomini/white:06} and \citet{west:96}{]}. There are crucial
distinctions between this classical approach and the setting introduced
in this paper. Under long-span asymptotics, identification of parameters
hinges on assumptions on the distributions or moments of the studied
processes {[}cf. the specification of the null hypotheses in \citet{giacomini/rossi:09}{]},
whereas within a continuous-time framework, unknown structural parameters
are identified from the sample paths of the studied processes. Hence,
we only need to assume rather mild pathwise regularity conditions
for the underlying continuous-time model and avoid any ergodic or
weak-dependence assumption. As in \citet{casini/perron_CR_Single_Break},
our framework encompasses any time series regression model allowing
for general forms of non-stationarity such as heteroskedasticty and
serial correlation.
Given a null hypotheses stated in terms of the path properties of
the sequence of losses, we propose a test statistic which compares
the local behavior of the sequence of surprise losses defined as the
difference between the out-of-sample and in-sample losses. More specifically,
our maximum-type statistic examines the smoothness of the sequence
of surprise losses as the continuous-time limit is approached. Under
the hypotheses, the continuous-time analogue of the sequence of losses
follows a continuous motion and any deviation from such smooth path
is interpreted as evidence against the hypotheses. The null distribution
of the test statistic is non-standard and follows an extreme value
distribution. Therefore, our limit theory exploits results from extreme
value theory as elaborated by \citet{bickel/rosenblatt:73} and \citet{galambos:87}.\footnote{In nonparametric change-point testing, related works are \citet{wu/zhao:07}
and \citet{bibinger/jirak/vetter:16}.}
We propose two versions of the test statistic: one that is self-normalized
and one that uses an appropriate estimator of the asymptotic variance.
The test statistic is defined as the maximal deviation between the
average surprise losses over asymptotically vanishing time blocks.
Further, we consider extensions of each of these statistics which
use overlapping rather than non-overlapping blocks. Although they
should be asymptotically equivalent, the statistics based on overlapping
blocks are more powerful in finite-samples. In a framework where one
allows for model misspecification, the problem of non-stationarity
such as heteroskedastcity and serial correlation in the forecast losses
should be taken seriously. Given the block-based form our test statistics
we derive an alternative estimator of the long-run variance of the
forecast losses. This estimator differs from the popular estimators
of \citet{andrews:91} and \citet{newey/west:87} {[}see \citet{muller:07}
for a review{]} and it is of independent interest. Finally, we extend
results to settings that allow for stochastic volatility, and we conduct
a local power analysis and highlight a few differences of our testing
framework from the structural change test of \citet{andrews:93}.
Related aspects, such as estimating the timing of the instability
and covering high-frequency setting with jumps, are being considered
in a companion paper.
The rest of the paper is organized as follows. Section \ref{Section: Statistical Enviromnent}
introduces the statistical setting, the hypotheses of interest and
the test statistics. Section \ref{Section CR Distribution Theory Test Stats}
derives the asymptotic null distribution under a continuous record.
We discuss the estimation of the asymptotic variance in Section \ref{Section Estimation-of-Asymptotic Variance}.
Some extensions and a local power analysis are presented in Section
\ref{Section Ito Vol}. Additional elements that are covered in our
companion paper are briefly described in Section \ref{Section Extensions}.
A simulation study is contained in Section \ref{Section Simulation Study}.
Section \ref{Section Conclusions} concludes the paper. The supplemental
material to this paper contains all mathematical proofs and additional
simulation experiments.
\section{The Statistical Environment\label{Section: Statistical Enviromnent}}
Section \ref{Subsection The-Forecasting-Problem} introduces the statistical
setting with a description of the forecasting problem and the sampling
scheme considered throughout. The underlying continuous-time model
and its assumptions are introduced in Section \ref{Subsection The Continuous-Time Model}.
In Section \ref{Subsection The-Hypotheses-of} we set out the testing
problem and state the relevant null and alternative hypotheses. The
test statistics are presented in Section \ref{Subsection The-Test-Statistics}.
Throughout we adopt the following notational conventions. All limits
are taken as $T\rightarrow\infty$, or equivalently as $h\downarrow0$,
where $T$ is the sample size and $h$ is the sampling interval.
All vectors are column vectors and for two vectors $a$ and $b$,
we write $a\leq b$ if the inequality holds component-wise. For a
sequence of matrices $\left\{ A_{T}\right\} ,$ we write $A_{T}=o_{\mathbb{P}}\left(1\right)$
if each of its elements is $o_{\mathbb{P}}\left(1\right)$ and likewise
for $O_{\mathbb{P}}\left(1\right).$ If $x$ is a non-stochastic vector,
$\left\Vert x\right\Vert $ denotes the its Euclidean norm, whereas
if $x$ is a stochastic vector, the same notation is used for the
$L^{2}$ norm. We use $\left\lfloor \cdot\right\rfloor $ to denote
the largest smaller integer function and for a set $A,$ the indicator
function of $A$ is denoted by $\mathbf{1}_{A}$. A sequence $\left\{ u_{kh}\right\} _{k=1}^{T}$
is $\textrm{i.i.d.}$ if the $u_{kh}$ are independent and identically
distributed. We use $\overset{\mathbb{P}}{\rightarrow},\,\Rightarrow$
to denote convergence in probability and weak convergence, respectively.
$\mathcal{M}_{p}^{\textrm{c�dl�g}}$ is used for the space of $p\times p$
positive define real-valued matrices whose elements are c�dl�g. The
symbol ``$\triangleq$'' is definitional equivalence.
\subsection{\label{Subsection The-Forecasting-Problem}The Forecasting Problem}
The continuous-time stochastic process $Z\triangleq\left(Y,\,X'\right)$
is defined on a filtered probability space $\left(\Omega,\,\mathscr{F},\,\left\{ \mathscr{F}_{t}\right\} _{t\geq0},\,\mathbb{P}\right)$
and takes value in $\mathbf{Z}\subseteq\mathbb{R}^{q+1}$ where $\left\{ Y_{t}\right\} _{t\geq0}$
is the variable to be forecast and $\left\{ X_{t}\right\} _{t\geq0}$
are the predictor variables. The index $t$ is defined as the continuous-time
index and we have $t\in\left[0,\,N\right]$, where $N$ is referred
to as the time span. In this paper, $N$ will remain fixed. That is,
the unobserved process $Z_{t}$ evolves within the fixed time horizon
$\left[0,\,N\right]$ and the econometrician records $T$ of its realizations,
with a sampling interval $h$, at discrete-time points $h,\,2h\ldots,\,Th$,
where accordingly $Th=N.$ A continuous record asymptotic framework
involves letting the sample size $T$ grow to infinity by shrinking
the time interval $h$ to zero at the same rate so that $N$ remains
fixed. The index $k$ is used for the observation (or tick) times
$k=1,\ldots,\,T$.
The objective is to generate a series $\left\{ Y_{\left(k+\tau\right)h}\right\} $
of $\tau$-step ahead forecasts. We shall adopt an out-of-sample
precedure whereby splitting the time span $\left[0,\,N\right]$ into
an \textit{in-sample} and \textit{out-of-sample} window, $\left[0,\,N_{\mathrm{in}}\right]$
and $\left[N_{\mathrm{in}}+h,\,N\right]$, respectively.\footnote{Indeed, $\left[0,\,N_{\mathrm{in}}\right]$ corresponds to the in-sample
window only for the fixed forecasting scheme to be introduced later\textemdash e.g.,
the rolling scheme only uses the most recent span of data of length
$N_{\mathrm{in}}$. A minor and straightforward modification to this
notation should be applied when the recursive and rolling schemes
are considered. However, for all methods $N_{\mathrm{in}}$ indicates
the artificial separation such that $N_{\mathrm{in}}+h$ is the beginning
of the out-of-sample period.} The latter two time horizons are supposed to be fixed and therefore
within the in-sample (or prediction) window a sample of size $T_{m}$
is observed whereas within the out-of-sample (or estimation) window
the sample is of size $T_{n}=T-T_{m}-\tau+1$. We consider a general
framework that allows for the three traditional forecasting schemes:
(1) a \textit{fixed} forecasting scheme with discrete-time observations
$h,\,2h,\ldots,\,\left(T_{m}-1\right)h,\,T_{m}h=N_{\mathrm{in}}$;
(2) a \textit{recursive} forecasting scheme where at time $kh$ the
prediction sample includes observations $h,\ldots,\,\left(k-1\right)h,\,kh$;
(3) a \textit{rolling} forecasting scheme where the time span of the
rolling window is fixed and of the same length as $N_{\mathrm{in}}$
(i.e., at time $kh$ the in-sample window includes observations $kh-T_{m}h+h,\ldots,\,\left(k-1\right)h,\,kh$.\footnote{Equivalently, the observation times within the rolling widow at the
$k$th's observation are $k-T_{m}+1,\ldots,\,k$.}
The forecasts may be based on a parametric model whose time-$kh$
parameter estimates are then collected into the $q\times1$ random
vector $\widehat{\beta}_{k}$. If no parametric assumption is made,
then $\widehat{\beta}_{k}$ represents whatever semiparametric or
nonparametric estimator used for generating the forecasts. The time-$kh$
forecast is denoted by $\widehat{f}_{k}\left(\widehat{\beta}_{k}\right)\triangleq f\left(Z_{kh},\,Z_{\left(k-1\right)h},\ldots,\,Z_{\left(k-m_{f}+1\right)h};\,\widehat{\beta}_{k}\right)$,
where $f$ is some measurable function. The notation indicates that
the $kh$-time forecast is generated from information contained in
a sample of size $m_{f}$.\footnote{$m_{f}$ varies with the forecastis scheme; e.g., for the rolling
scheme we have $m_{f}=T_{m}$ while for the recursive scheme we have
$m_{f}=k$.}
Next, we introduce a loss function $L\left(\cdot\right)$ which serves
for evaluating the performance of a given forecast model. More specifically,
each out-of-sample loss $L_{\left(k+\tau\right)h}\left(\widehat{\beta}_{k}\right)\triangleq L\left(Y_{\left(k+\tau\right)h},\,\widehat{f}_{k}\left(\widehat{\beta}_{k}\right)\right)$
constitutes a statistical measure of accuracy of the $\tau$-step
forecast made at time $kh$. However, given the objective of detecting
potential instability of a certain forecasting method over time, we
need additionally to introduce the in-sample losses $L_{jh}\left(\widehat{\beta}_{k}\right)\triangleq L\left(Y_{jh},\,\widehat{y}_{j}\left(\widehat{\beta}_{k}\right)\right),$
where $\widehat{y}_{j}\left(\widehat{\beta}_{k}\right)$ is an in-sample
fitted value with $j$ varying over the specific in-sample window.
That is, for each time-$kh$ forecast there corresponds a sequence
(indexed by $j$) of in-sample fitted values $\widehat{y}_{j}\left(\widehat{\beta}_{k}\right)$.\footnote{We have $j=\tau+1,\ldots,\,T_{m}$ for the fixed scheme, $j=\tau+1,\ldots,\,k$
for the recursive scheme and $j=k-T_{m}+\tau+1,\ldots,\,k$ for the
rolling scheme. } Then, the testing problem turns into the detection of any ``systematic
difference'' between the sequence of out-of-sample and in-sample
losses; the formal measure of such difference under our context is
provided below.
\subsection{\label{Subsection The Continuous-Time Model}The Underlying Continuous-Time
Model}
The process $Z$ is a $\mathbb{R}^{q+1}$-valued semimartingale on
$\left(\Omega,\,\mathscr{F},\,\left\{ \mathscr{F}_{t}\right\} _{t\geq0},\,\mathbb{P}\right)$
and we further assume that all processes considered in this paper
are c�dl�g adapted and possess a $\mathbb{P}$-a.s. continuous path
on $\left[0,\,N\right]$.\footnote{For accessible treatments of the probabilistic elements used in this
section we refer to \citet{sahalia/jacod:14}, \citet{jacod/shiryaev:03},
\citet{jacod/protter:12}, \citet{karatzas/shreve:96} and \citet{protter:05}.} The continuity property represents a key assumption in our setting
and implies that $Z$ is a continuous It\^o semimartignale. The integral
form for $X_{t}$ is given by,
\begin{align}
X_{t} & =x_{0}+\int_{0}^{t}\mu_{X,s}ds+\int_{0}^{t}\sigma_{X,s}dW_{X,s},\label{Mode for X}
\end{align}
where $\left\{ W_{X,t}\right\} _{t\geq0}$ is a $q\times1$ Wiener
process, $\mu_{X,s}\in\mathbb{R}^{q}$ and $\sigma_{X,s}\in\mathcal{M}_{q}^{\textrm{c�dl�g}}$
are the drift and spot covariance process, respectively, and $x_{0}$
is $\mathscr{F}_{0}$-measurable. We incorporate model misspeficication
into our framework by allowing for a large non-zero drift which adds
to the residual process:
\begin{align}
Y_{t} & \triangleq y_{0}+\left(\beta^{*}\right)'X_{t-}+\int_{0}^{t}\mu_{e,s}h^{-\vartheta}ds+e_{t},\qquad\qquad e_{t}\triangleq\int_{0}^{t}\sigma_{e,s}dW_{e,s}\label{Mode for Y}
\end{align}
where $\beta^{*}\in\mathbb{R}^{q}$, $\left\{ W_{e,t}\right\} _{t\geq0}$
is a standard Wiener process, $\sigma_{e,s}\in\mathbb{R}_{+}$ is
its associated volatility, $\mu_{e,s}\in\mathbb{R}$ and $y_{0}$
is $\mathscr{F}_{0}$-measurable. In \eqref{Mode for Y}, the last
two terms on the right-hand side account for the residual part of
$Y_{t}$ which is not explained by $X_{t-},$ where $X_{t-}=\lim_{s\uparrow t}X_{s}$.
We assume $\vartheta\in[0,\,1/8)$ so that the factor $h^{-\vartheta}$
inflates the infinitesimal mean of the residual component thereby
approximating a setting with arbitrary misspecification.
\begin{rem}
In \eqref{Mode for Y}, misspecification manifests itself in the
form of (time-varying) non-zero conditional mean of the residual process,
and in giving rise to serial dependence in the disturbances which
in turn leads to dependence in the sequence of forecast losses.\footnote{Asymptotically, these features can be dealt with basic arguments used
in the high-frequency financial statistics literature; however, when
$h$ is not too small one needs methods that are robust in finite-samples
to such misspecification-induced properties. More precisely, we will
propose an appropriate estimator of the long-run variance of the sequence
of forecast losses in Section \ref{Section Estimation-of-Asymptotic Variance}.
} Hence, this specification is similar in spirit to the near-diffusion
assumption of \citet{foster/nelson:96} who studied the impact of
misspecification in ARCH models. On the other hand, \citet{casini/perron_CR_Single_Break}
introduced a ``large-drift'' asymptotics with $h^{-1/2}$ to deal
with non-identification of the drift in their context. Technically,
the latter specification implies that as $h$ becomes small the drift
features larger oscillations that add to the local Gaussianity of
the stochastic part. \citet{casini/perron_CR_Single_Break} referred
to this specification as \textit{small-dispersion} assumption. Finally,
note that the presence of $h^{-\vartheta}$ can also be related to
the signal plus small Gaussian noise of \citet{ibragimov/has:81}
if one sets $\varepsilon_{h}=h^{\vartheta}$ in their model in Section
VII.2.
\end{rem}
\begin{assumption}
\label{Assumption 1, CT}We have the following assumptions: (i) The
processes $\left\{ X_{t}\right\} _{t\geq0}$ and $\Sigma^{0}\triangleq\left\{ \sigma_{X,t},\,\sigma_{e,t}\right\} _{t\geq0}$
have $\mathbb{P}$-a.s. continuous sample paths; (ii) The processes
$\left\{ \mu_{X,t}\right\} _{t\geq0},\,\left\{ \mu_{e,t}\right\} _{t\geq0},\,\left\{ \sigma_{X,t}\right\} _{t\geq0}$
and $\left\{ \sigma_{e,t}\right\} _{t\geq0}$ are locally bounded;
(iii) There exists $0<\sigma_{-}<\sigma_{+}<\infty$ such that $\mathbb{P}$-a.s.
$\inf_{t\in\left[0,\,N\right]}\sigma_{V,t}^{2}\geq\sigma_{-}^{2}$
and $\sigma_{+}^{2}\geq\sup_{t\in\left[0,\,N\right]}\sigma_{V,t}^{2}$
with $V=X,\,e$; (iv) $\sigma_{X,t}\in\mathcal{M}_{q}^{\mathrm{c\grave{a}dl\grave{a}g}}$
and $\sigma_{e,t}\in\mathcal{M}_{1}^{\mathrm{c\grave{a}dl\grave{a}g}}$
and the conditional variance (or spot covariance) is defined as $\Sigma_{X,t}=\sigma_{X,t}\sigma'_{X,t}$,
which for all $t<\infty$ satisfies $\int_{0}^{t}\Sigma_{X,s}^{\left(j,j\right)}ds<\infty,\,\left(j=1,\ldots,\,q\right)$
where $\Sigma_{X,t}^{\left(j,r\right)}$ denotes the $\left(j,r\right)$-th
element of $\Sigma_{X,t}$. Furthermore, for every $j=1,\ldots,\,q,$
and $k=1,\,2,\ldots,\,T$, the quantity $h^{-1}\int_{\left(k-1\right)h}^{kh}\Sigma_{X,s}^{\left(j,j\right)}ds$
is bounded away from zero and infinity, uniformly in $k$ and $h$;
(v) The disturbance process $e_{t}$ is orthogonal (in martingale
sense) to $X_{t}:$ $\left\langle e,\,X\right\rangle _{t}=0$ identically
for all $t\geq0$.\footnote{The angle brackets notation $\left\langle \cdot,\cdot\right\rangle $
is used for the predictable quadratic variation process.}
\end{assumption}
Part (i) rules out jump processes from our setting. We relax this
restriction in our companion paper; see Section \ref{Section Extensions}.
Part (ii) restricts those processes to be locally bounded. These should
be viewed as regularity conditions rather than assumptions and are
standard in the financial econometrics literature {[}see \citet{barndorff/shephard:04},
\citet{li/xiu:16} and \citet*{li/todorov/tauchen:17}{]}; recently,
they have been used by \citet{casini/perron_CR_Single_Break} in the
context of structural change models.
The continuous-time model in \eqref{Mode for X}-\eqref{Mode for Y}
is not observable. The econometrician only has access to $T$ realizations
of $Y_{t}$ and $X_{t}$ with a sampling interval $h>0$ over the
horizon $\left[0,\,N\right]$. For each $h>0,$ $Z_{kh}\in\mathbb{R}^{q+1}$
is a random vector step function that jumps only at time $0,\,h,\,2h,\ldots$,
and so on. The discretized processes $Y_{kh}$ and $X_{kh}$ are assumed
to be adapted to the increasing and right-continuous filtration $\left\{ \mathscr{F}_{t}\right\} _{t\geq0}$.
The increments of a process $U$ are denoted by $\Delta_{h}U_{k}=U_{kh}-U_{\left(k-1\right)h}$.
A seminal result known as Doob-Meyer Decomposition {[}cf. the original
sources are \citet{doob:53} and \citet{meyer:67}; see also Section
III.3 in \citet{protter:05}{]} allows us to decompose the semimartingale
process $X_{t}$ into a predictable part and a local martingale part.
Hence, it follows that we can write for $k=1,\ldots,T$, $\Delta_{h}X_{k}\triangleq\mu_{X,kh}\cdot h+\Delta_{h}M_{X,k}$
where the drift $\mu_{X,t}\in\mathbb{R}^{q}$ is $\mathscr{F}_{t-h}$
measurable, and $M_{X,kh}\in\mathbb{R}^{q}$ is a continuous local
martingale with finite conditional covariance matrix $\mathbb{P}$-a.s.
$\mathbb{E}\left(\Delta_{h}M_{X,k}\Delta_{h}M'_{X,k}|\,\mathscr{F}_{\left(k-1\right)h}\right)=\Sigma_{X,\left(k-1\right)h}\cdot h$.
Turning to equation \eqref{Mode for Y}, the error process $\left\{ \Delta_{h}e_{k}^{*},\,\mathscr{F}_{t}\right\} $,
with $\Delta_{h}e_{k}^{*}\triangleq\sigma_{e,\left(k-1\right)h}\Delta_{h}W_{e,k}$,
is then a continuous local martingale difference sequence taking its
values in $\mathbb{R}$ with finite conditional variance $\mathbb{E}\left[\left(\Delta_{h}e_{k}^{*}\right)^{2}|\,\mathscr{F}_{\left(k-1\right)h}\right]=\sigma_{e,\left(k-1\right)h}^{2}\cdot h,\,\mathbb{P}$-a.s.
Therefore, we express the discretized analogue of \eqref{Mode for Y}
as
\begin{align}
\Delta_{h}Y_{k} & =\left(\beta^{*}\right)'\Delta_{h}X_{k-\tau}+\mu_{e,kh}\cdot h^{1-\vartheta}+\Delta_{h}e_{k},\qquad\qquad k=\tau+1,\ldots,T.\label{Discretized Model Y}
\end{align}
\begin{rem}
As explained above, we accommodate possible model misspecification
by adding the component $\mu_{e,k}\cdot h^{1-\vartheta}$. In the
forecasting literature, often one directly imposes restrictions on
the sequence of losses, say, $L\left(e_{k}\right)$ where $e_{k}=Y_{k}-\widehat{f}_{k}\left(\widehat{\beta}_{k}\right)$
is a forecast error. There are two main differences from our approach.
First, in order to facilitate illustrating our novel framework to
the reader, we have chosen, without loss of generality, to express
directly the relationship between $\Delta_{h}Y_{k+\tau}$ and $\Delta_{h}X_{k}$
while at the same time, allowing for misspecification by including
$\mu_{e,kh}\cdot h^{1-\vartheta}$. A second distinction from the
classical approach is that the latter imposes restrictions on the
sequences of losses such as mixing and ergodicity conditions, covariance
stationary and so on. In contrast, our infill asymptotics does not
require us to impose any ergodic or mixing condition {[}cf. \citet{casini/perron_CR_Single_Break}{]}.
\end{rem}
Finally, we have an additional assumption on the path of the volatility
process $\left\{ \sigma_{e,t}^{2}\right\} _{t\geq0}$. This turns
out be important because it partly affects the local behavior of the
forecast losses.
\begin{assumption}
\label{Assumption Lipchtitz cont of Sigma}For small $\eta>0$, define
the modulus of continuity of $\left\{ \sigma_{e,t}\right\} _{t\geq0}$
on the time horizon $\left[0,\,N\right]$ by $\phi_{\sigma,\eta,N}=\sup_{s,t\in\left[0,\,N\right]}\left\{ \left|\sigma_{t}-\sigma_{s}\right|:\,\left|t-s\right|<\eta\right\} .$
We assume that $\phi_{\sigma,\eta,\tau_{h}\wedge N}\leq K_{h}\eta$
for some sequence of stopping times $\tau_{h}\rightarrow\infty$ and
some $\mathbb{P}$-a.s. finite random variable $K_{h}$.
\end{assumption}
The assumption essentially states that $\phi_{\sigma,\eta,N}$ is
locally bounded and $\left\{ \sigma_{e,t}\right\} _{t\geq0}$ is Lipschitz
continuous. Lipschitz volatility is a more than reasonable specification
for the macroeconomic and financial data to which our analysis is
primarily directed. Indeed, the basic case of constant variance $\sigma^{2}$
is easily accommodated by the assumption. Time-varying volatility
is also covered provided $\sigma_{e,t}^{2}$ is sufficiently smooth.
However, the assumption rules out some standard stochastic volatility
models often used in finance. We relax that assumption in Section
\ref{Section Ito Vol}, so that we can extend our results to, for
example, stochastic volatility models driven by a Wiener process.
\subsection{\label{Subsection The-Hypotheses-of}The Hypotheses of Interest}
As time evolves, a forecast model can suffer instability for multiple
reasons. However, incorporating model misspecification into our framework
necessarily implies that the exact form of the instability is unknown
and thus one has to leave it unspecified. This differs from the classical
setting for estimation of structural change models {[}cf. \citet{bai/perron:98}
and \citet{casini/perron_CR_Single_Break}{]} where (i) the break
date is well-defined as it is part of the definition of the econometric
problem, and (ii) the form of the instability is explicitly specified
through a discrete shift in a regression parameter. In contrast, under
our context we remain agnostic regarding both (i) and (ii). There
may be multiple dates at which the forecast model suffers instability
and they might be interrelated in a complicated way. Forecast instability
may manifest itself in several forms, including gradual, smooth or
recurrent changes in the predictive relationship between $Y_{\left(k+\tau\right)h}$
and $X_{kh}$; certainly, there could also be discrete shifts in $\beta^{*}$\textemdash arguably
the most common case in practice\textemdash but this is only a possibility
in our setting and not an assumption as in structural change models.
A forecast failure then reflects the forecaster's failure to recognize
the shift in the predictive power of $X_{kh}$ on $Y_{\left(k+\tau\right)h}$.
On the other hand, even if one can rule out shits in $\beta^{*}$,
a forecast instability may be induced by an increase/decrease in the
uncertainty in the data which might result, for example, from changes
in the unconditional variance of the target variable. In this case,
the predictive ability of $X_{kh}$ on $Y_{\left(k+\tau\right)h}$,
as described for instance by a parameter $\beta,$ remains stable
while due to an increase in the unconditional variance of $Y_{\left(k+\tau\right)h}$
it might become weak and in turn the forecasting power might breakdown.
Tests for forecast failure such as those proposed in this paper and
the ones proposed in \citet{giacomini/rossi:09} are designed to have
power against both of the above hypotheses.\footnote{Recently, \citet{perron/yamamoto:18} proposed to apply modified versions
of classical structural break tests to the forecast failure setting.
However, their testing framework and hence their null hypotheses are
different from ours because they do not fix a model-method pair but
only fix the forecast model under the null.}
\subsubsection{The Null and Alternative Hypotheses on Forecast Instability}
Define at time $\left(k+\tau\right)h$ a \textit{surprise loss} given
by the deviation between the time-$\left(k+\tau\right)h$ out-of-sample
loss and the average in-sample loss: $SL_{\left(k+\tau\right)h}\left(\widehat{\beta}_{k}\right)\triangleq L_{\left(k+\tau\right)h}\left(\widehat{\beta}_{k}\right)-\overline{L}_{kh}\left(\widehat{\beta}_{k}\right),$
for $k=T_{m},\ldots,\,T-\tau$, where $\overline{L}_{kh}\left(\widehat{\beta}_{k}\right)$
is the average in-sample loss computed according to the specific
forecasting scheme. One can then define the average of the out-of-sample
surprise losses
\begin{align}
\overline{SL}{}_{N_{0}}\left(\widehat{\beta}_{k}\right) & \triangleq N_{0}^{-1}\sum_{k=T_{m}}^{T-\tau}SL_{\left(k+\tau\right)h}\left(\widehat{\beta}_{k}\right),\label{eq. SL_bar}
\end{align}
where $N_{0}\triangleq N-N_{\mathrm{in}}-h$ denotes the time span
of the out-of-sample window.\footnote{By definition $N_{0}$ is fixed and should not be confused with $T_{n},$
which indicates the number of observations in the out-of-sample window.
Indeed, $N_{0}=T_{n}h.$ } In the classical discrete-time setting, under the hypotheses of no
forecast instability one would naturally test whether $\overline{SL}{}_{N_{0}}\left(\beta^{*}\right)$
has zero mean, where $\beta^{*}$ is the pseudo-true value of $\beta$.
If the forecasting perfomance remains stable throughout the whole
sample then there should be no systematic surprise losses in the out-of-sample
window and thus $\mathbb{E}\left[N_{0}^{-1}\sum_{k=T_{m}}^{T-\tau}SL_{\left(k+\tau\right)h}\left(\beta^{*}\right)\right]=0.$
This reasoning motivated the forecast breakdown test of \citet{giacomini/rossi:09}.
Therefore, under the classical asymptotic setting one exploits time
series properties of the process $SL_{\left(k+\tau\right)h}\left(\beta^{*}\right)$
such as ergodicity and mixing together with the representation of
the hypotheses by a global moment restriction.\footnote{Global refers to the property that the zero-mean restriction involves
the entire sequence of forecast losses.} By letting the span $N\rightarrow\infty$, this method underlies
the classical approach to statistical inference but does not directly
extend to an infill asymptotic setting. Under continuous-time asymptotics,
identification of parameters is achieved by properties of the paths
of the involved processes and not by moment conditions. This constitutes
the key difference and requires one to recast the above hypotheses
into an infill setting thereby making use of assumptions on an underlying
continuous-time data-generating mechanism which is assumed to govern
the observed data.
We begin with observing that the sequence of losses $\left\{ L_{kh}\left(\cdot\right)\right\} $
can be viewed as realizations from an underlying continuous-time process
$\left\{ \widetilde{L}_{t}\right\} _{t\geq0}$, with $\widetilde{L}_{t}\triangleq\int_{0}^{t}L_{s}\left(Y_{s},\,X_{s-};\,\beta^{*}\right)ds$.
That is, $\widetilde{L}_{t}$ consists of temporally integrated forecast
losses where $L_{t}$ is the loss at time $t$ and is defined by some
transformation of the target variable $Y_{t}$ and of the predictor
$X_{t-}$.\footnote{The definition of $\widetilde{L}_{t}$ uses that so long as the forecast
step $\tau$ is small and finite one can approximate $X_{s-\tau h}$
by $X_{s-}$ for sufficiently small $h>0.$} In order to provide a general theory, we focus on families of loss
functions that depend only on the forecast error.\footnote{The most popular loss functions used in economic forecasting are within
this category {[}see \citet{elliott/timmermann:16} for a recent incisive
account of the literature{]}. Extension to \textit{ad hoc} loss functions
requires specific treatment that might vary from case to case.} We denote this class by $\boldsymbol{L}_{e}$ and we say that the
loss function $L_{\cdot}\left(\cdot,\,\cdot;\,\cdot\right)\in\boldsymbol{L}_{e}$
if $L_{t}\left(Y_{t},\,X_{t-};\,\beta\right)=L_{t}\left(e_{t};\,\beta\right)$
for all $t\in\left[0,\,N\right]$, where $e_{t}=Y_{t}-\widehat{f}_{t}\left(\beta\right)$.
The class $\boldsymbol{L}_{e}$ comprises the vast majority of loss
functions employed in empirical work, including among others the popular
Quadratic loss, Absolute error loss and Linex loss. The following
examples illustrate how these loss functions are constructed under
our setting. For the rest of this section, assume for simplicity $y_{0}=0,\,\mu_{e,\cdot}=0$
and that $X_{t}$ is one-dimensional in \eqref{Mode for Y}.
\begin{example*}
($\mathbf{QL}$: Quadratic Loss)\\
The \textit{Mean Squared Error} or \textit{Quadratic} loss function
is symmetric and is by far the most commonly used by practitioners.
Given \eqref{Mode for Y}, we have $e_{t}=Y_{t}-\beta^{*}X_{t-}$.
Then $L\left(e\right)=ae^{2}$ or $L_{t}\left(Y_{t},\,X_{t-};\,\beta^{*}\right)=ae_{t}^{2}$
with $a>0$.
\end{example*}
\begin{example*}
($\mathbf{LL}$: Linex Loss)\\
The \textit{Linear-exponential} or \textit{Linex }loss was introduced
by \citet{varian:75} and it is an example of asymmetric loss function.
By the same reasoning as in the Quadratic loss case, we have $L\left(e\right)=a_{1}\left(\exp\left(a_{2}e\right)-a_{2}e-1\right)$
or $L_{t}\left(Y_{t},\,X_{t-};\,\beta^{*}\right)=a_{1}\left(\exp\left(a_{2}e_{t}\right)-a_{2}e_{t}-1\right)$
with $a_{1}>0,\,a_{2}\neq0$.
\end{example*}
Below we make very mild pathwise assumptions on the process $Z$ which
imply restrictions on $\left\{ \widetilde{L}_{t}\right\} _{t\geq0}$.
We derive asymptotic results under Lipschitz continuity (in $t$)
of the coefficients of the system of stochastic differential equations
driving the data $\left\{ Z_{t}\right\} _{t\geq0}$. We apply the
techniques of stochastic calculus to formulate our testing problem.
To avoid clutter, we introduce the notation $g\left(Y_{t},\,X_{t-};\,\beta^{*}\right)=L_{t}\left(Y_{t},\,X_{t-};\,\beta^{*}\right)$
and its shorthand $g\left(e_{t};\,\beta^{*}\right)=L_{t}\left(e_{t};\,\beta^{*}\right)$.\footnote{The notation implicitly assumes that the same loss function is used
for estimation and prediction which in turn implies that the subscript
$t$ in $L_{t}\left(e_{t};\,\beta^{*}\right)$ can be omitted since
it can be understood from that of the argument $e_{t}$.} By It\^o Lemma, {[}cf. Section II.7 in \citet{protter:05}{]}, under
smoothness of $g\left(e_{t};\,\beta^{*}\right)$,
\begin{align*}
dL_{t}\left(e_{t};\,\beta^{*}\right) & =\frac{\sigma_{e,t}^{2}}{2}\frac{\partial^{2}g\left(e_{t};\,\beta^{*}\right)}{\partial e^{2}}dt+\sigma_{e,t}\frac{\partial g\left(e_{t};\,\beta^{*}\right)}{\partial e}dW_{e,t}.
\end{align*}
Let $\mathbb{E}_{\sigma}$ denote the expectation conditional on the
path $\left\{ \sigma_{e,t}\right\} _{t\geq0}$. The instantaneous
mean of $dL\left(e_{t};\,\beta^{*}\right)$ is $\mathbb{E}_{\sigma}\left[dL\left(e_{t};\,\beta^{*}\right)/dt\right]=2^{-1}\sigma_{e,t}^{2}\mathbb{E}_{\sigma}\left(\partial^{2}g\left(e_{t};\,\beta^{*}\right)/\partial e^{2}\right)$.
Note that the latter is a symbolic abbreviation for
\begin{align*}
\mathbb{E}_{\sigma}\left[L_{t}\left(e_{t};\,\beta^{*}\right)-L_{s}\left(e_{s};\,\beta^{*}\right)\right] & =\frac{\sigma_{e,t}^{2}}{2}\mathbb{E}_{\sigma}\left[\frac{\partial^{2}g\left(e_{t};\,\beta^{*}\right)}{\partial e^{2}}\right]\left(t-s\right)+o\left(t-s\right),\qquad\qquad\mathrm{as\,}s\uparrow t.
\end{align*}
Since the coefficients of the original system of stochastic equations
are Lipschitz continuous in $t$, one can verify that $\mathbb{E}_{\sigma}\left[dL\left(e_{t};\,\beta^{*}\right)/dt\right]$
is also Lipschitz upon regularity conditions on $g\left(\cdot,\,\beta^{*}\right)$
and time-$t$ information.
We denote by $\boldsymbol{Lip}\left(\left[0,\,N\right]\right)$ the
class of Lipschitz continuous functions on $\left[0,\,N\right]$.
Let $\left\{ c_{t}\right\} _{t\geq0}$ denote a continuous-time stochastic
process that is $\mathbb{P}$-a.s. locally bounded and adapted.
\begin{defn}
\label{def: Lipschitz}The process $\left\{ c_{t}\right\} _{t\geq0}$
belongs to $\boldsymbol{Lip}\left(\left[0,\,N\right]\right)$ if $\sup_{s,t\in\left[0,\,\tau_{h}\wedge N\right],t\neq s}\left|c_{t}-c_{s}\right|<K_{h}\left|t-s\right|$
for some sequence of stopping times $\tau_{h}\rightarrow\infty$ and
some $\mathbb{P}$-a.s. finite random variable $K_{h}$.
\end{defn}
We are in a position to formulate the testing problem in terms of
the pathwise property of $L_{t}\left(e_{t};\,\beta^{*}\right).$ This
implies that the hypotheses are specified in terms of random events
which differs from classical hypotheses testing but it is typical
under continuous-time asymptotics; see \citet{sahalia/jacod:12} (for
many references), \citet{li/todorov/tauchen:16} and \citet{reiss/todorov/tauchen:15}.
We consider the following hypotheses: for any $L\left(\cdot;\,\cdot\right)\in\boldsymbol{L}_{e}$,\footnote{Precise assumptions will be stated below.}
\begin{align}
H_{0}: & \quad\left\{ \lim_{s\uparrow t}\mathbb{E}_{\sigma}\left[L_{t}\left(e_{t};\,\beta^{*}\right)-L_{s}\left(e_{s};\,\beta^{*}\right)\right]\right\} \in\boldsymbol{Lip}\left(\left[N_{\mathrm{in}}+h,\,N\right]\right),\label{eq H0}
\end{align}
which means that we wish to discriminate between the following two
events that divide $\Omega:$
\begin{align*}
\Omega_{0}\triangleq\left\{ \omega\in\Omega:\,\left\{ \lim_{s\uparrow t}\mathbb{E}_{\sigma}\left[L_{t}\left(e_{t}\left(\omega\right);\,\beta^{*}\right)-L_{s}\left(e_{s}\left(\omega\right);\,\beta^{*}\right)\right]\right\} \in\boldsymbol{Lip}\left(\left[N_{\mathrm{in}}+h,\,N\right]\right)\right\} , & \qquad\Omega_{1}\triangleq\Omega\backslash\Omega_{0}.
\end{align*}
The dependence of the hypotheses on $\omega$ is appropriate because
each event $\omega$ generates a certain path of $L\left(e_{t}\left(\cdot\right);\,\beta^{*}\right)$
on $\left[0,\,N\right]$, where $de_{t}\left(\omega\right)=\sigma_{e,t}\left(\omega\right)dW_{e,t}\left(\omega\right)$.
The hypotheses requires a Lipschitz condition to hold on $\left[N_{\mathrm{in}}+h,\,N\right]$,
where $N_{\mathrm{in}}$ is the usual artificial separation date after
which the first forecast is made. $N_{\mathrm{in}}$ is taken as given
here because the testing problem applies to a specific method-model
pair and $N_{\mathrm{in}}$ is part of the chosen forecasting method.
From a practical standpoint, it would be helpful if this separation
date is such that the forecast model is stable on $\left[0,\,N_{\mathrm{in}}\right]$
{[}see \citet{casini/perron_Oxford_Survey} for more details{]}. The
latter property is, however, unknown a priori by the practitioner.
We cover this case in Section \ref{Section Extensions}.
\begin{example*}
($\mathbf{QL}$; cont'd)\\
For the Quadratic loss $L\left(e\right)=ae^{2},$ It\^o Lemma yields
$\mathbb{E}_{\sigma}\left[dL_{t}\left(e_{t};\,\beta^{*}\right)/dt\right]=a\sigma_{e,t}^{2}$.
If $\sigma_{e,t}$ is Lipschitz continuous, then the hypothesis $H_{0}$
holds.
\end{example*}
\begin{example*}
($\mathbf{LL}$; cont'd)\\
From It\^o Lemma, $dL_{t}\left(e_{t};\,\beta^{*}\right)=a_{1}\left\{ a_{2}\left[2^{-1}a_{2}\sigma_{e,t}^{2}\exp\left(a_{2}e_{t}\right)dt+\left(\exp\left(a_{2}e_{t}\right)-1\right)\sigma_{e,t}dW_{e,t}\right]-1\right\} $.
Consequently, by It\^o Isometry {[}cf. Section 3.3.2 in \citet{karatzas/shreve:96}
or Lemma 3.1.5 in \citet{oksendal:00}{]} $\mathbb{E}_{\sigma}\left[dL\left(e_{t};\,\beta^{*}\right)/dt\right]=a_{1}\left(a_{2}^{2}\left(\sigma_{e,t}^{2}/2\right)\right)\exp\left(a_{2}^{2}\left(\int_{0}^{t}\sigma_{e,s}^{2}ds\right)/2\right)$
and hypotheses $H_{0}$ is seen to hold under Lipschitz continuity
of $\sigma_{e,t}$.\footnote{Recall that composition of Lipschitz functions is Lipschitz and that
under our context $\exp\left(a_{2}\left(\int_{0}^{t}\sigma_{e,s}^{2}ds\right)/2\right)$
is Lipschitz because (i) $\sigma_{e,s}^{2}$ is locally bounded and
Lipschitz, and (ii) $t\leq N$ and $N$ remains fixed. }
\end{example*}
We have reduced the forecast instability problem to examination of
the local properties of the path of $L_{t}$. However, we still have
to face the question on how to use the data to test $H_{0}$ in practice.
Even if we could observe $\widetilde{L}_{t}$, it would not be clear
how to formulate a testing problem on the stability of $L_{t}$ by
using path properties of $\widetilde{L}_{t}$. The reason is that
$\widetilde{L}_{t}$ is always absolutely continuous by definition,
and thus it would provide little information on the large deviations
of the forecast error $e_{t}$. In order to study the local behavior
of $L_{t}$ one needs to consider the small increments of $L_{t}$
close to time $t$. Leaving the definition of $\widetilde{L}_{t}$
aside for a moment, observe that $\mathbb{P}$-a.s. continuity of
$Z_{t}$ is equivalent to having the relationship between $Y_{t}$
and $X_{t}$ holding for any infinitesimal interval of time. For the
basic parametric linear model: $dY_{t}=\beta^{*}dX_{t}+de_{t}$. Then,
the forecast loss is $L\left(de_{t}\right)$, which is difficult to
interpret in rigorous probabilistic terms. However, we can consider
its discrete-time analogue. We normalize the forecast error by the
factor $\psi_{h}=h^{1/2}$ and redefine $L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)\triangleq L_{kh}\left(\psi_{h}^{-1}\Delta_{h}e_{k};\,\beta^{*}\right)$.\footnote{Alternatively, $L_{\psi,kh}\left(\Delta_{h}Y_{k},\,\Delta_{h}X_{k};\,\beta^{*}\right)=L_{kh}\left(\psi_{h}^{-1}\left(\Delta_{h}Y_{k}-\beta^{*}\Delta_{h}X_{k}\right)\right)$.}
Then, for all $k$, the mean of $L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)$\textemdash conditional
on $\sigma_{e,kh}$\textemdash depends on the parameters of the model
and its local behavior can be used as a proxy for the local behavior
of the infinitesimal mean of $dL_{t}\left(e_{t};\,\beta^{*}\right)$.
If the corresponding structural parameters of the continuous-time
data-generating process satisfy a Lipschitz continuity in $t$, then\textemdash knowing
$\sigma_{e,kh}$\textemdash also $\mathbb{E}_{\sigma}\left[L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)\right]$
should be Lipschitz in the continuous-time limit. Under the hypotheses
$H_{0}$ there should be no break in $\mathbb{E}_{\sigma}\left[L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)\right]$
and an appropriately defined right local average of $L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)$
should not differ too much from its left local average. That is, one
can test for forecast instability by using a two-sample t-test over
asymptotically vanishing time blocks.
\begin{example*}
($\mathbf{QL}$; cont'd)\\
Conditional on $\left\{ \sigma_{t}\right\} _{t\geq0}$, $\Delta_{h}e_{k}\sim\mathscr{N}\left(0,\,\sigma_{e,\left(k-1\right)h}^{2}\cdot h\right)$.
Thus, $\mathbb{E}_{\sigma}\left[L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)\right]=a\sigma_{e,\left(k-1\right)h}^{2}$.
If $\sigma_{e,t}$ is Lipschitz continuous, then the hypothesis $H_{0}$
holds.
\end{example*}
\begin{example*}
($\mathbf{LL}$; cont'd)\\
Similar to the Quadratic loss case, we have $\mathbb{E}_{\sigma}\left[L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)\right]=a_{1}\left(\exp\left(a_{2}^{2}\sigma_{e,\left(k-1\right)h}^{2}/2\right)-1\right)$.
Again, the hypotheses $H_{0}$ is satisfied if $\sigma_{e,t}$ is
Lipschitz.
\end{example*}
Both examples demonstrate that pathwise assumptions on the data-generating
process implies restrictions on the properties of the sequence of
loss functions. For the QL example, if there is a structural break
at the observation $k=T_{b},$ then this would result in the mean
of $L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)$ shifting
to a new level after time $T_{b}h$. Given that the same reasoning
extends to the sequence of surprise losses, one may consider to construct
a test statistic on the basis of the local behavior of the surprise
losses over time. If there is no instability in the predictive ability
of a certain model, then the sequence of out-of-sample surprise losses
should display a certain stability. Under the framework of \citet{giacomini/rossi:09},
this stability is interpreted in a retrospective and global sense
as a zero-mean restriction on the sequence over the entire out-of-sample.
In contrast, under our continuous-time setting, this stability manifests
itself as a continuity property of the path of the continuous-time
counterpart of the sequence.
\subsection{\label{Subsection The-Test-Statistics}The Test Statistics }
By inspection of the null hypotheses in \eqref{eq H0}, it is evident
that a considerable number of forms of instabilities are allowed.
These may result from discrete shifts in a model's structural parameter
and/or in structural properties of the processes considered such as
conditional and unconditional moments and so on. This first set of
non-stationarities relates to the popular case of structural changes
which are designed to be detected with high probability by the structural
break tests of, among others, \citet{andrews:93} and \citet{andrews/ploberger:94},
\citet{bai/perron:98} and \citet{elliott/mueller:06} in univariate
settings and of \citet{qu/perron:07} in multivariate settings. However,
a forecast instability may be generated by many other forms of non-stationarities
against which such classical tests for structural breaks are not designed
for and consequently they might have little power against. For example,
consider the case of smooth changes in model parameters, or in the
unconditional variance of $Y_{kh}$. Even more serious would be the
presence of recurrent smooth changes in the marginal distribution
of the predictor since in this case the above-mentioned tests are
likely to falsely reject $H_{0}$ too often {[}cf. \citet{hansen:00}{]}.
Thus, the null hypotheses of no forecast instability calls for a new
statistical hypotheses testing framework. Ideally, in this context
one needs a test statistic that retains power against any discontinuity,
jump and recurrent switch at any point in the out-of-sample and for
any magnitude of the shift. We propose a test statistic which aims
asymptotically at distinguishing any discontinuity from a regular
Lipschitz continuous motion. We introduce a sequence of two-sample
t-tests over asymptotically vanishing adjacent time blocks. This should
lead to significant gains in power whenever on fixed time intervals
the out-of-sample losses exhibit instabilities of any form such as
breaks, jumps and relatively large deviations. Such gains are likely
to occur especially when instabilities take place within a small portion
of the sample relative to the whole time span\textemdash a common
case in practice that has characterized many episodes of forecast
failure in economics.
Interestingly, for the Quadratic loss function we can exploit properties
of the local quadratic variation and propose a self-normalized test
statistic. Thus, we separate the discussion on the Quadratic loss
from that on general loss functions. Let $SL_{\psi,\left(k+\tau\right)h}\left(\widehat{\beta}_{k}\right)\triangleq L_{\psi,\left(k+\tau\right)h}\left(\widehat{\beta}_{k}\right)-\overline{L}_{\psi,kh}\left(\widehat{\beta}_{k}\right)$,
$k=T_{m},\ldots,\,T-\tau$. Next, we partition the out-of-sample
into $m_{T}\triangleq\left\lfloor T_{n}/n_{T}\right\rfloor $ blocks
each containing $n_{T}$ observations. Let $B_{h,b}\triangleq n_{T}^{-1}\sum_{j=1}^{n_{T}}SL_{\psi,T_{m}+\tau+bn_{T}+j-1}\left(\widehat{\beta}_{T_{m}+bn_{T}+j-1}\right)$
and $\overline{B}_{h,b}\triangleq n_{T}^{-1}\sum_{j=1}^{n_{T}}L_{\psi,\left(T_{m}+\tau+bn_{T}+j-1\right)h}\left(\widehat{\beta}_{T_{m}+bn_{T}+j-1}\right)$
for $b=0,\ldots,\,,\,\left\lfloor T_{n}/n_{T}\right\rfloor -1$.
\subsubsection{Test Statistics under Quadratic Loss}
We propose the following statistic
\begin{align*}
\mathrm{B}_{\mathrm{max},h}\left(T_{n},\,\tau\right) & \triangleq\max_{b=0,\ldots,\,\left\lfloor T_{n}/n_{T}\right\rfloor -2}\left|\frac{B_{h,b+1}-B_{h,b}}{\overline{B}_{h,b+1}}\right|.
\end{align*}
The quantity $B_{h,b}$ is a local average of the surprise losses
within the block $b$. We have partitioned the out-of-sample window
into $m_{T}$ blocks of asymptotically vanishing length $\left[bn_{T}h,\,\left(b+1\right)n_{T}h\right]$.
We consider an asymptotic experiment in which the number of blocks
$m_{T}$ increases at a controlled rate to infinity while the per-block
sample size $n_{T}$ grows without bound at a slower rate than the
out-of-sample size $T_{n}$. The appeal of the $\mathrm{B}_{\mathrm{max},h}\left(T_{n},\,\tau\right)$
statistic is that a large deviation $B_{h,b+1}-B_{h,b}$ suggests
the existence of either a discontinuity or non-smooth shift in the
surprise losses close to time $bn_{T}h$ and thus it provides evidence
against $H_{0}$. We comment on the nature of the normalization $\overline{B}_{h,b+1}$
in the denominator of $\mathrm{B}_{\mathrm{max},h}$ below, after
we introduce a version of $\mathrm{B}_{\mathrm{max},h}$ statistic
which uses all admissible overlapping blocks of length $n_{T}h$:
\begin{align*}
\mathrm{MB}{}_{\mathrm{max},h}\left(T_{n},\,\tau\right) & \triangleq\max_{i=n_{T},\ldots,\,T_{n}-n_{T}}\left|\left(n_{T}^{-1}\sum_{j=i-n_{T}+1}^{i}SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)\right.\right.\\
& \left.\left.\quad-n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)\right)/\overline{B}_{h,i}\right|,
\end{align*}
where $\overline{B}_{h,i}=n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}L_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)$.
Since under the alternative hypotheses the exact location of the change-point\textendash -or
possibly the locations of the multiple change-points\textemdash within
the block might actually affect the power of the $\mathrm{B}_{\mathrm{max},h}$-based
test in small samples, we indeed find in our simulation study that
the test statistic $\mathrm{MB}{}_{\mathrm{max},h}$ which uses overlapping
blocks is more powerful especially when the instability arises in
forms other than the simple one-time structural change. Thus, the
power of the $\mathrm{B}_{\mathrm{max},h}$ test is slightly sensible
to the actual location of the change-point within the block, with
higher power achieved when the change-point is close to either the
beginning or the end of the block. In contrast, the statistical power
of $\mathrm{MB}_{\mathrm{max},h}$ is uniform over the location of
the change-point in the sample. The latter property is not shared
by the exiting test of \citet{giacomini/rossi:09} given that its
power tends to be substantially lower if the instability is not located
at about middle sample.
An important characteristic of both $\mathrm{B}_{\mathrm{max},h}$
and $\mathrm{MB}_{\mathrm{max},h}$ is that they are self-normalized;
no asymptotic variance appears in their definition. The reason for
why $\overline{B}_{h,b+1}$ appears in the denominator of, for example,
$\mathrm{B}_{\mathrm{max},h}$ is that even though $B_{h,b+1}$ constitutes
a more logical self-normalizing term, it might be close to zero in
some cases. This would occur under Quadratic loss if, for example,
$\sigma_{e,t}=\sigma_{e}$ for all $t\geq0.$ This is not true for
the factor $\overline{B}_{h,b+1}$.
In addition, observe that allowing for misspecification naturally
leads one to deal carefully with artificial serial dependence in the
forecast losses in small samples. Thus, we consider a version of the
statistics $\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{MB}_{\mathrm{max},h}$
that are normalized by their asymptotic variance: $\mathrm{Q}_{\mathrm{max},h}\left(T_{n},\,\tau\right)\triangleq\nu_{L}^{-1}\max_{b=0,\ldots,\,\left\lfloor T_{n}/n_{T}\right\rfloor -2}\left|B_{h,b+1}-B_{h,b}\right|,$
and similarly,
\begin{align*}
\mathrm{MQ}{}_{\mathrm{max},h} & \left(T_{n},\,\tau\right)\\
& \triangleq\nu_{L}^{-1}\max_{i=n_{T},\ldots,\,T_{n}-n_{T}}\left|n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)-n_{T}^{-1}\sum_{j=i-n_{T}+1}^{i}SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)\right|.
\end{align*}
The quantity $\nu_{L}$ standardizes the test statistic so that under
the null hypotheses we obtain a distribution-free limit. This can
be useful because given the fully non-stationary setting together
with the possible consequences of misspecification in finite-samples,
standardization by the square-root of the asymptotic variance $\nu_{L}^{2}$
might lead to a more precise empirical size in small samples. We relegate
theoretical details on $\nu_{L}$ as well as on its estimation to
Section \ref{Section Estimation-of-Asymptotic Variance} where we
also present a discussion about its relation with the choice of the
number of blocks.
\subsubsection{Test Statistics under General Loss Function}
For general loss $L\in\boldsymbol{L}_{e}$, we propose the following
statistic,
\begin{align*}
\mathrm{G}_{\mathrm{max},h}\left(T_{n},\,\tau\right) & \triangleq\max_{b=0,\ldots,\,\left\lfloor T_{n}/n_{T}\right\rfloor -2}\left|\frac{B_{h,b+1}-B_{h,b}}{\sqrt{D_{h,b+1}}}\right|,
\end{align*}
where $B_{h,b},\,B_{h,b+1}$ are defined as in the quadratic case
and
\begin{align*}
D_{h,b+1} & \triangleq n_{T}^{-1}\sum_{j=1}^{n_{T}}\left(L_{\psi,\left(T_{m}+\tau+\left(b+1\right)n_{T}+j-1\right)h}\left(\widehat{\beta}_{T_{m}+\left(b+1\right)n_{T}+j-1}\right)-\overline{L}_{\psi,b+1}\left(\widehat{\beta}\right)\right)^{2},
\end{align*}
with $\overline{L}_{\psi,b}\left(\widehat{\beta}\right)\triangleq n_{T}^{-1}\sum_{j=1}^{n_{T}}L_{\psi,\left(T_{m}+\tau+bn_{T}+j-1\right)h}\left(\widehat{\beta}_{T_{m}+bn_{T}+j-1}\right)$.
The interpretation of $\mathrm{G}_{\mathrm{max},h}$ is essentially
the same as of $\mathrm{B}_{\mathrm{max},h}$, the only difference
arising from the denominator $\sqrt{D_{h,b+1}}$ that estimates the
within-block variance. A version that uses all overlapping blocks
is
\begin{align*}
\mathrm{MG}{}_{\mathrm{max},h} & \left(T_{n},\,\tau\right)\\
& \triangleq\max_{i=n_{T},\ldots,\,T_{n}-n_{T}}\left|\frac{n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)-n_{T}^{-1}\sum_{j=i-n_{T}+1}^{i}SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)}{\sqrt{D_{h,i}}}\right|,
\end{align*}
where $D_{h,i}\triangleq n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}\left(L_{\psi,\left(T_{m}+\tau+j-1\right)h}\left(\widehat{\beta}_{T_{m}+j-1}\right)-\overline{L}_{\psi,i}\left(\widehat{\beta}\right)\right)^{2}$,
with
\begin{align*}
\overline{L}_{\psi,i}\left(\widehat{\beta}\right) & \triangleq n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}L_{\psi,\left(T_{m}+\tau+j-1\right)h}\left(\widehat{\beta}_{T_{m}+j-1}\right).
\end{align*}
As argued above, it is useful to consider versions of the statistic
$\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{MB}_{\mathrm{max},h}$
that are normalized by their asymptotic variance:
\begin{align*}
\mathrm{Q}_{\mathrm{max},h}^{\mathrm{G}}\left(T_{n},\,\tau\right) & \triangleq\nu_{L}^{-1}\max_{b=0,\ldots,\,\left\lfloor T_{n}/n_{T}\right\rfloor -2}\left|B_{h,b+1}-B_{h,b}\right|,
\end{align*}
and similarly,
\begin{align*}
\mathrm{MQ}_{\mathrm{max},h}^{\mathrm{G}} & \left(T_{n},\,\tau\right)\\
& \triangleq\nu_{L}^{-1}\max_{i=n_{T},\ldots,\,T_{n}-n_{T}}\left|n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)-n_{T}^{-1}\sum_{j=i-n_{T}+1}^{i}SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)\right|.
\end{align*}
\section{Continuous Record Distribution Theory for the Test Statistics\label{Section CR Distribution Theory Test Stats}}
\subsection{Asymptotic Distribution under the Null Hypotheses}
We begin with a set of assumptions. Assumption \ref{Assumption Moments of Losses}
below is a finite-moment condition on the sequence of rescaled forecast
losses and and on its first-order derivative. It has a similar scope
to A4 in \citet{giacomini/rossi:09}. Assumption \ref{Assumption A5 in GR: Finite E L'}
is similar to A5 in \Citet{giacomini/rossi:09} and it imposes the
first-order derivative of the forecast losses to be constant over
time. It trivially holds when one employs the same loss function for
estimation and evaluation. Assumption \ref{Assumption Roo-T Consistent beta }
demands the existence of a consistent estimator for $\beta^{*}$ at
the parametric rate $\sqrt{T}$ and it encompasses many estimation
procedures. In model \eqref{Discretized Model Y}, the popular least-squares
method will satisfy the condition {[}cf. \citet{barndorff/shephard:04}
and \citet*{li/todorov/tauchen:17}{]}.
\begin{assumption}
\label{Ass the process}The process $\left\{ Y_{t}-\left(\beta^{*}\right)'X_{t-}\right\} _{t\in\left[0,\,N\right]}$
takes value in an open set $\mathcal{E}\subseteq\mathbb{R}$, and
$\beta^{*}$ takes value in a compact parameter space $\Theta\subset\mathbb{R}^{\mathrm{dim}\left(\beta\right)}$.
\end{assumption}
\begin{assumption}
\label{Assumption Differentiability Loss}For any $L\in\boldsymbol{L}_{e}$
we assume $L:\,\mathcal{E}\times\Theta\mapsto\mathbb{R}$ is a measurable
function and $L\in\boldsymbol{C}^{2,2}$ (i.e., twice continuously
differentiable in both arguments). For every open set $B$ that contains
$\beta^{*}$ there exists $C<\infty$ such that for all $k\geq1$,
$\sup_{\beta\in B}\left\Vert \partial^{2}L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta\right)/\partial\beta\partial\beta'\right\Vert <C$.
\end{assumption}
\begin{assumption}
\label{Assumption Moments of Losses}We have $\sup_{k=1,\ldots,T}\mathbb{E}_{\sigma}\left\Vert \left(L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right),\,\partial L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)/\partial\beta\right)'\right\Vert ^{4+\varpi}<\infty$,
for $\varpi>0$.
\end{assumption}
\begin{assumption}
\label{Assumption A5 in GR: Finite E L'}For all $k\geq1,$ $\mathbb{E}_{\sigma}\left[\partial L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)/\partial\beta\right]=\overline{K}$,
for some $\overline{K}<\infty$.
\end{assumption}
\begin{assumption}
\label{Assumption Bounded on Bounded Sets LX (2-iv)}For all $k\geq1$,
$\left|\partial L_{\psi,kh}\left(e;\,\cdot\right)/\partial e\right|$
is bounded on bounded sets.
\end{assumption}
\begin{assumption}
\label{Assumption Roo-T Consistent beta }There exists a sequence
$\left\{ \widehat{\beta}_{k}\right\} _{k=T_{m}}^{T-\tau}$ such that
$\left\Vert \widehat{\beta}_{k}-\beta^{*}\right\Vert =O_{\mathbb{P}}\left(1/\sqrt{T}\right)$
uniformly over $k=T_{m},\ldots,\,T-\tau$.
\end{assumption}
Our asymptotic results are valid under the following conditions on
the auxiliary sequence $n_{T}$.
\begin{condition}
\label{Cond The-auxiliary-sequence}The sequence $\left\{ n_{T}\right\} $
satisfies for some $\epsilon>0,$
\begin{align}
n_{T}\rightarrow\infty\quad\mathrm{as\quad}T\rightarrow\infty & \qquad\mathrm{and}\qquad T^{\epsilon}n_{T}^{-1}+n_{T}^{3/2}h\sqrt{\log\left(T\right)}\rightarrow0.\label{eq. Condition auxiliary nT}
\end{align}
\end{condition}
Condition \ref{Cond The-auxiliary-sequence} imposes a lower bound
and an upper bound on the growth condition of the sequence $\left\{ n_{T}\right\} $.
The first part of \eqref{eq. Condition auxiliary nT} requires $n_{T}$
to grow to infinity at any faster rate than $T^{\epsilon}$ with $\epsilon>0$,
which we interpret as saying that the number of observations $n_{T}$
in each block cannot be too small. The second part of \eqref{eq. Condition auxiliary nT}
provides an upper bound on the growth of $n_{T}$ and relates to Assumption
\ref{Assumption Lipchtitz cont of Sigma} concerning the smoothness
of $\left\{ \sigma_{e,t}\right\} _{t\geq0}$ thereby ensuring that,
for example, the random oscillations of $\mathrm{B}_{\mathrm{max},h}\left(T_{n},\,\tau\right)$
can be controlled. As we shall explain in the simulation study
of Section \ref{Section Simulation Study}, we recommend to set $n_{T}\propto T_{n}^{2/3-\epsilon}$
for small $\epsilon>0$.
\subsubsection{Asymptotic Distribution Under Quadratic Loss Function}
\begin{thm}
\label{Theoem Asymptotic H0 Distrbution Bmax and Qmax}Let $\gamma_{m_{T}}=\left[4\log\left(m_{T}\right)-2\log\left(\log\left(m_{T}\right)\right)\right]^{1/2}$
and recall $m_{T}=\left\lfloor T_{n}/n_{T}\right\rfloor $. Assume
Assumption \ref{Assumption 1, CT}-\ref{Assumption Lipchtitz cont of Sigma},
\ref{Ass the process}-\ref{Assumption Roo-T Consistent beta }, and
Condition \ref{Cond The-auxiliary-sequence} hold. Let $\mathscr{V}$
denote a random variable defined by $\mathbb{P}\left(\mathscr{V}\leq v\right)=\exp\left(-\pi^{-1/2}\exp\left(-v\right)\right).$
Under $H_{0}$, we have\\
(i) $\sqrt{\log\left(m_{T}\right)}\left(2^{-1/2}n_{T}^{1/2}\mathrm{B}_{\mathrm{max},h}\left(T_{n},\,\tau\right)-\gamma_{m_{T}}\right)\overset{}{\Rightarrow}\mathscr{V};$
\\
(ii) $2^{-1/2}\sqrt{\log\left(m_{T}\right)}n_{T}^{1/2}\mathrm{MB}_{\mathrm{max},h}\left(T_{n},\,\tau\right)-2\log\left(m_{T}\right)-\frac{1}{2}\log\log\left(m_{T}\right)-\log3\overset{}{\Rightarrow}\mathscr{V}.$
\end{thm}
\begin{cor}
\label{Corollary CR Null Distrb Quadratic}Under the same assumptions
of the previous theorem, we have under $H_{0}$, $\sqrt{\log\left(m_{T}\right)}$
~ $\left(n_{T}^{1/2}\nu_{L}^{-1}\mathrm{Q}_{\mathrm{max},h}\left(T_{n},\,\tau\right)-\gamma_{m_{T}}\right)\overset{}{\Rightarrow}\mathscr{V}$
and
\begin{align*}
\sqrt{\log\left(m_{T}\right)}\left(n_{T}^{1/2}\nu_{L}^{-1}\right)\mathrm{MQ}_{\mathrm{max},h}\left(T_{n},\,\tau\right)-2\log\left(m_{T}\right)-\frac{1}{2}\log\log\left(m_{T}\right)-\log3\overset{}{\Rightarrow}\mathscr{V} & ,
\end{align*}
where $\mathscr{V},\,m_{T}$ and $\gamma_{m_{T}}$ are defined as
in the previous theorem.
\end{cor}
Theorem \ref{Theoem Asymptotic H0 Distrbution Bmax and Qmax} shows
that the asymptotic null distribution of our test statistics follows
an extreme vale distribution whose critical values can be computed
directly. In nonparametric change-point analysis, \citet{wu/zhao:07}
and \citet*{bibinger/jirak/vetter:16} have derived an extreme value
null distribution for tests statistics which share a similar form
to ours. As it is stated, the tests statistics are not yet feasible
because the asymptotic variances $\nu_{L}^{2}$ is unknown. However,
we can find statistical consistent estimators which can be used in
place of $\nu_{L}^{2}$ to make the test feasible. We relegate the
treatment of its consistent estimation to Section \ref{Section Estimation-of-Asymptotic Variance}.
\subsubsection{Asymptotic Distribution Under General Loss Function}
\begin{thm}
\label{Theoem Asymptotic H0 Distrbution Gmax QGmax}Under the same
assumptions of the previous theorem and with $\mathscr{V},\,m_{T}$
and $\gamma_{m_{T}}$ defined analogously, we have under $H_{0}$,
\\
(i) $2^{-1/2}\sqrt{\log\left(m_{T}\right)}\left(n_{T}^{1/2}\mathrm{G}_{\mathrm{max},h}\left(T_{n},\,\tau\right)-\gamma_{m_{T}}\right)\overset{}{\Rightarrow}\mathscr{V};$
\\
(ii) $2^{-1/2}\sqrt{\log\left(m_{T}\right)}n_{T}^{1/2}\mathrm{MG}_{\mathrm{max},h}\left(T_{n},\,\tau\right)-2\log\left(m_{T}\right)-\frac{1}{2}\log\log\left(m_{T}\right)-\log3\overset{}{\Rightarrow}\mathscr{V}.$
\end{thm}
\begin{cor}
\label{Corollary CR Null Distrb General}Under the same assumptions
of the previous theorem, we have under $H_{0}$, $\sqrt{\log\left(m_{T}\right)}$
~ $\left(n_{T}^{1/2}\nu_{L}^{-1}\mathrm{Q}_{\mathrm{max},h}^{\mathrm{G}}\left(T_{n},\,\tau\right)-\gamma_{m_{T}}\right)\overset{}{\Rightarrow}\mathscr{V}$
and $\sqrt{\log\left(m_{T}\right)}n_{T}^{1/2}\nu_{L}^{-1}\mathrm{MQ}_{\mathrm{max},h}^{\mathrm{G}}\left(T_{n},\,\tau\right)-2\log\left(m_{T}\right)-\frac{1}{2}\log\log\left(m_{T}\right)-\log3\overset{}{\Rightarrow}\mathscr{V}.$
\end{cor}
\section{\label{Section Estimation-of-Asymptotic Variance}Estimation of Asymptotic
Variance}
The purpose of this section is to show how to construct an asymptotically
valid estimator of the variance $\nu_{L}^{2}$ that enters the definition
of our test statistics. This is an important aspect that together
with the selection of the block length might affect statistical inferences
based on the proposed tests in finite-samples. Allowing for misspecification
is customary in the forecasting literature, and as a consequence this
may result in forecast losses that artificially exhibit heteroskedasticity
and serial dependence in small samples.
\subsection{\label{subsec:Asymptotic-Variance-Estimator}Estimation of the Asymptotic
Variance}
We begin with the case of stationary forecast losses, including constant
$\nu_{L}$ as a special case.
\subsubsection{Stationary Forecast Losses}
Recall that our test statistics are related to a maximum over blocks
of data. Thus, for i.i.d. forecast losses one can use the following
estimator for $\nu_{L}$ in $\mathrm{Q}_{\mathrm{max},h}\left(T_{n},\,\tau\right)$:
\begin{align*}
\widehat{\nu}_{\mathrm{Q}1,b}^{2} & \triangleq\frac{2}{n_{T}}\sum_{j=1}^{n_{T}}\left(SL_{\psi,T_{m}+\tau+bn_{T}+j-1}\left(\widehat{\beta}_{T_{m}+bn_{T}+j-1}\right)-\overline{SL}{}_{\psi,b}\right)^{2},
\end{align*}
where $\overline{SL}{}_{\psi,b}\triangleq n_{T}^{-1}\sum_{j=1}^{n_{T}}SL_{\psi,T_{m}+\tau+bn_{T}+j-1}\left(\widehat{\beta}_{T_{m}+bn_{T}+j-1}\right)$.
The estimator $\widehat{\nu}_{\mathrm{Q}1,b}^{2}$ normalizes the
difference in the out-of-sample forecast losses between the $b+1$
and $b$ blocks. The statistic $\mathrm{Q}_{\mathrm{max},h}\left(T_{n},\,\tau\right)$
then results in $\mathrm{Q}_{\mathrm{max},h}\left(T_{n},\,\tau\right)=\max_{b=0,\ldots,\,\left\lfloor T_{n}/n_{T}\right\rfloor -2}\left|\left(B_{h,b+1}-B_{h,b}\right)/\widehat{\nu}_{\mathrm{Q}1,b+1}\right|$.
For the overlapping blocks case, the estimator is $\widehat{\nu}_{\mathrm{MQ}1,i}^{2}\triangleq2n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}\left(SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)-\overline{SL}{}_{\psi,i}\right)^{2},$
where $\overline{SL}{}_{\psi,i}\triangleq n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)$
so that we can write
\begin{align*}
\mathrm{M} & \mathrm{Q}{}_{\mathrm{max},h}\left(T_{n},\,\tau\right)\\
& \triangleq\max_{i=n_{T},\ldots,\,T_{n}-n_{T}}\left|\widehat{\nu}_{\mathrm{MQ}1,i}^{-1}n_{T}^{-1}\left(\sum_{j=i+1}^{i+n_{T}}SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)-\sum_{j=i-n_{T}+1}^{i}SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)\right)\right|.
\end{align*}
Both $\widehat{\nu}_{\mathrm{Q}1,b+1}^{2}$ and $\widehat{\nu}_{\mathrm{MQ}1,i}^{2}$
apply a natural block-wise normalization in order to guarantee a distribution-free
limit under $H_{0}$. However, it is useful to consider estimators
that use all of the observations in the out-of-sample period. Thus,
one exploits covariance stationarity of the sequence of forecast losses.
Let $\Phi_{0.75}=0.647...$ denote the third quartile of the standard
normal distribution and define
\begin{align*}
\widehat{\nu}_{2,h}\triangleq\frac{\sqrt{\pi n_{T}}}{2\left(m_{T}-1\right)}\sum_{b=1}^{m_{T}-1}\left|B_{h,b}-B_{h,b-1}\right|, & \qquad\qquad\widehat{\nu}_{3,h}\triangleq\frac{\sqrt{n_{T}}}{\sqrt{2\left(m_{T}-1\right)}}\left(\sum_{b=1}^{m_{T}-1}\left|B_{h,b}-B_{h,b-1}\right|^{2}\right)^{1/2},\\
\widehat{\nu}_{4,h}\triangleq\frac{\sqrt{n_{T}}}{\sqrt{2\Phi_{0.75}}}\mathrm{median} & \left(\left|B_{h,b}-B_{h,b-1}\right|\right),\qquad1\leq b\leq m_{T}-1.
\end{align*}
Note that $\widehat{\nu}_{2,h},\,\widehat{\nu}_{3,h}$ and $\widehat{\nu}_{4,h}$
can be used to implement both $\mathrm{Q}{}_{\mathrm{max},h}\left(T_{n},\,\tau\right)$
and $\mathrm{MQ}{}_{\mathrm{max},h}\left(T_{n},\,\tau\right)$.\footnote{They can be also applied to the test statistics $\mathrm{Q}{}_{\mathrm{max},h}^{\mathrm{G}}\left(T_{n},\,\tau\right)$
and $\mathrm{MQ}{}_{\mathrm{max},h}^{\mathrm{G}}\left(T_{n},\,\tau\right)$
with $G_{h,b}$ in place of $B_{h,b}$.} $\widehat{\nu}_{3,h}$ is related to \citeauthor{carlstein:86}'s
(1986) subseries variance estimate in the context of strong mixing
processes and it was also used by \citet{wu/zhao:07}. Each of the
estimators $\widehat{\nu}_{2,h},\,\widehat{\nu}_{3,h}$ and $\widehat{\nu}_{4,h}$
allows for dependence but requires stationarity. The simulation study
in \citet{wu/zhao:07} suggests that $\widehat{\nu}_{4,h}$ is more
robust whereas $\widehat{\nu}_{2,h}$ and $\widehat{\nu}_{3,h}$ are
less precise when there are large instabilities or jumps. For two
sequences $\left\{ a_{k}\right\} $ and $\left\{ b_{k}\right\} $,
we write $a_{k}\asymp b_{k}$ if for some $c\geq1$, $b_{k}/c\leq a_{k}\leq cb_{k}$
for all $T.$ The following theorem is similar to Theorem 3 in \citet{wu/zhao:07}
and in particular, part (ii) states that if $n_{T}\asymp T_{n}^{1/3}$
then $\widehat{\nu}_{3,h}^{2}$ achieves the optimal MSE $O\left(n_{T}^{-2/3}\right)$.
\begin{condition}
\label{Cond The-auxiliary-sequence LR VAR}The sequence $\left\{ n_{T}\right\} $
satisfies
\begin{align}
n_{T}\rightarrow\infty\quad\mathrm{as\quad}T\rightarrow\infty & \qquad\mathrm{and}\qquad\sqrt{T_{n}}n_{T}^{-1}\log\left(T_{n}\right)^{3}+n_{T}T_{n}^{-2/3}\left(\log\left(T\right)\right)^{1/3}\rightarrow0.\label{eq. Condition auxiliary nT-1-1}
\end{align}
\end{condition}
\begin{thm}
\label{Theorem: Var Est 3 in Wu and Zhao}In addition to the assumptions
of Theorem \ref{Theoem Asymptotic H0 Distrbution Bmax and Qmax},
assume that $\mathrm{Cov}\left(L_{\psi,kh}\left(\beta^{*}\right),\,L_{\psi,\left(k-j\right)h}\right.$
~ $\left(\beta^{*}\right)\bigr)$ depends on $j$ but not on $kh$.
Then, under $H_{0}$, (i) Let $n_{T}\asymp T^{5/8}$. Then, $\widehat{\nu}_{2,h},\,\widehat{\nu}_{4,h}=\nu_{L}+O_{\mathbb{P}}\left(T_{n}^{-1/16}\log\left(T_{n}\right)\right)$;
(ii) Let $n_{T}\asymp T^{1/3}$. Then $\mathbb{E}\left(\left[\widehat{\nu}_{3,h}^{2}-\nu_{L}^{2}\right]^{2}\right)=O\left(T_{n}^{-2/3}\right)$.
\end{thm}
Under covariance-stationarity, given Theorem \ref{Theorem: Var Est 3 in Wu and Zhao},
the results of Corollary \ref{Corollary CR Null Distrb Quadratic}-\ref{Corollary CR Null Distrb General}
are applicable after replacing $\nu_{L}$ by an appropriate consistent
estimator.
\subsubsection{Non-Stationary Forecast Losses}
We now consider estimation of the asymptotic variance in the case
the forecast losses are heterogeneous. The estimator $\widehat{\nu}$
that we introduce below depends on the specific loss function and
thus it can be used for replacing $\nu_{L}$ in Corollary \ref{Corollary CR Null Distrb Quadratic}-\ref{Corollary CR Null Distrb General}.
Non-stationarity implies that $\sigma_{e,t}^{2}$ is time-varying
and thus the results of Theorem \ref{Theorem: Var Est 3 in Wu and Zhao}
are not applicable due to the presence of many extra parameters that
account for the time-varying structure. To deal with this issue we
propose a novel block-wise self-normalization technique which simultaneously
addresses two issues. First, the block-wise self-normalization ensures
that the difference in forecast losses between two adjacent blocks
are asymptotically independent across non-adjacent blocks and that
within each block the losses are standardized so that time-varying
variances cancel out. Second, by computing an average\textemdash over
all blocks\textemdash of the self-normalized difference in forecast
losses we account for possible serial dependence. We derive asymptotic
results within a general framework based on the strong invariance
principle for stationary processes developed in \citet{wu:07} and
extended to modulated stationary processes by \citet{zhao/li:13}.
For each block $b=0,\ldots,\,m_{T}-2$, let
\begin{align*}
A_{h,b}\left(\widehat{\beta}\right) & \triangleq n_{T}^{-1}\sum_{j=1}^{n_{T}}\left(L_{\psi,T_{m}+\tau+\left(b+1\right)n_{T}+j-1}\left(\widehat{\beta}_{T_{m}+\left(b+1\right)n_{T}+j-1}\right)\right),\\
V_{h,b}\left(\widehat{\beta}\right) & \triangleq n_{T}^{-1}\sum_{j=1}^{n_{T}}\left(L_{\psi,T_{m}+\tau+\left(b+1\right)n_{T}+j-1}\left(\widehat{\beta}\right)-\overline{L}{}_{\psi,b}\left(\widehat{\beta}\right)\right)^{2},
\end{align*}
where $\overline{L}{}_{\psi,b}\left(\widehat{\beta}\right)=n_{T}^{-1}\sum_{j=1}^{n_{T}}L_{\psi,T_{m}+\tau+\left(b+1\right)n_{T}+j-1}\left(\widehat{\beta}\right)$
and define the statistic
\begin{align*}
\zeta_{h,b}\left(\widehat{\beta}\right)\triangleq & \sqrt{n_{T}}\left(A_{h,b}\left(\widehat{\beta}\right)-A_{h,b-1}\left(\widehat{\beta}\right)\right)/\sqrt{V_{h,b}}.
\end{align*}
Finally, an average\textemdash over all blocks $m_{T}$\textemdash of
the per-block self-normalized statistics $\zeta_{h,b}$'s is used
to define an estimator of the asymptotic variance: $\widehat{\nu}_{L}^{2}\triangleq2^{-1}\left(m_{T}-1\right)^{-1}\sum_{b=0}^{m_{T}-1}\zeta_{h,b}^{2}$.
Let $\sigma_{L,kh}^{2}\triangleq\mathrm{Var}\left(L_{\psi,kh}\left(\beta^{*}\right)\right)$.
We also need to introduce the following quantities,
\begin{align*}
F_{h,b}^{*} & \triangleq\left|\sigma_{L,\left(T_{m}+\tau+\left(b+2\right)n_{T}\right)h}\right|,\qquad\qquad J_{h,b}^{*}\triangleq\sigma_{L,\left(T_{m}+\tau+\left(b+2\right)n_{T}\right)h}^{2},\\
\Sigma_{h,b}^{*} & \triangleq\sum_{j=1}^{n_{T}}\sigma_{L,\left(T_{m}+\tau+\left(b+1\right)n_{T}+j\right)h}^{2},\qquad\qquad\widetilde{\Sigma}_{h,b}^{*}\triangleq\left(\sum_{j=1}^{n_{T}}\sigma_{L,\left(T_{m}+\tau+\left(b+1\right)n_{T}+j\right)h}^{4}\right)^{1/2}.
\end{align*}
\begin{thm}
\label{Theorem Long-Run Variance}Under Condition \ref{Cond The-auxiliary-sequence LR VAR}
we have $\widehat{\nu}_{L}^{2}-\nu_{L}^{2}=O_{\mathbb{P}}\left(r_{h}^{-1}\right)$,
where $r_{h}=O_{\mathbb{P}}\left(T_{n}^{\epsilon}/\left(\log\left(T_{n}\right)\right)^{2}\right)$
with $\epsilon\in\left(0,\,1/4\right)$ such that $r_{h}\rightarrow\infty$.
\end{thm}
The theorem simply states that $\widehat{\nu}_{L}$ is consistent
for $\nu_{L}$ and therefore the asymptotic results of Section \ref{Section CR Distribution Theory Test Stats}
continue to hold when we replace $\nu_{L}$ by $\widehat{\nu}_{L}$.
\section{\label{Section Ito Vol}Continuous Semimartingale Volatility and
Asymptotic Local Power}
\subsection{Asymptotic Results under Continuous Semimartingale Volatility }
In this section we relax the Lipschitz condition on $\sigma_{e,t}$
and extend the results for the quadratic loss case from Theorem \ref{Theoem Asymptotic H0 Distrbution Bmax and Qmax}
to stochastic volatility models driven by a Wiener process. Consequently,
this relaxation enables one to utilize the tests proposed in this
paper in setting involving high-frequency financial variables. More
specifically, we assume that $\sigma_{e,t}$ is an It\^o continuous
semimartingale that is almost surely bounded and strictly positive
adapted process. We replace Assumption \ref{Assumption Lipchtitz cont of Sigma}
by the following.
\begin{assumption}
\label{Assumption Smooth of Sigma}Under $H_{0}$ the process $\left\{ \sigma_{e,t}\right\} _{t\geq0}$
satisfies $\phi_{\sigma,\eta,\tau_{h}\wedge N}\leq K_{h}\eta^{\kappa}$
for some $\kappa>0$, some sequence of stopping times $\tau_{h}\rightarrow\infty$
and some $\mathbb{P}$-a.s. finite random variable $K_{h}$.
\end{assumption}
The assumption implies that $\sigma_{e,t}$ belongs to a rather large
class of volatility processes usually considered in financial econometrics.
The parameter $\kappa$ plays a key role in the testing framework
of this section and we refer to it as the regularity exponent. When
$\kappa=1$ we recover the case of Lipschitz volatility considered
in the previous sections while the standard stochastic volatility
model without jumps correspond to $\kappa=1/2-\epsilon$ for a sufficiently
small $\epsilon>0.$ Next, we have a slightly different version of
Condition \ref{Cond The-auxiliary-sequence}.
\begin{condition}
\label{Cond The-auxiliary-sequence Ito Vol}The sequence $\left\{ n_{T}\right\} $
satisfies for some $\epsilon>0,$
\begin{align}
n_{T}\rightarrow\infty\quad\mathrm{as\quad}T\rightarrow\infty & \qquad\mathrm{and}\qquad T^{\epsilon}n_{T}^{-1}+\sqrt{n_{T}}\left(n_{T}h\right)^{\kappa}\sqrt{\log\left(T\right)}\rightarrow0.\label{eq. Condition auxiliary nT Ito Vol}
\end{align}
\end{condition}
For It\^o continuous semimartingale volatility $\sigma_{e,t}$ the
condition suggests $n_{T}\propto T^{1/2-\epsilon}$ for small $\epsilon>0.$
Let $\varGamma_{t}\triangleq\mathbb{E}_{\sigma}\left[dL_{t}\left(e_{t};\,\beta^{*}\right)/dt\right]$.\footnote{For example, for the quadratic loss with $\mu_{e,t}=0$ the notation
reduces to $\varGamma_{t}=\sigma_{e,t}^{2}.$ } The more general framework considered here requires us to consider
the following null hypotheses: under quadratic loss,
\begin{align}
H_{0}: & \quad\left\{ \varGamma_{t}\right\} _{t\in\left[N_{\mathrm{in}}+h,\,N\right]}\in\boldsymbol{C}\left(\kappa,\,K_{h}\right),\label{eq H0 Ito Vol}
\end{align}
where $\boldsymbol{C}\left(\kappa,\,K_{h}\right)$ is a class of
continuous functions on $\left[N_{\mathrm{in}}+h,\,N\right]$,
\begin{align*}
\boldsymbol{C}\left(\kappa,\,K_{h}\right) & \triangleq\left\{ \left\{ \varGamma_{t}\right\} _{t\in\left[N_{\mathrm{in}}+h,\,N\right]}:\,\sup_{s,t\in\left[N_{\mathrm{in}}+h,\,N\right],\left|t-s\right|<\eta}\left|\varGamma_{t}-\varGamma_{s}\right|\leq K_{h}\eta^{\kappa}\right\} ,
\end{align*}
where $\kappa>0$ and $K_{h}$ is given in Assumption \ref{Assumption Smooth of Sigma}.
Thus, we wish to discriminate between $H_{0}$ and
\begin{align}
H_{1} & :\,\exists\lambda\in\left[N_{\mathrm{in}}+h,\,N\right]\quad\mathrm{with}\quad\left\{ \varGamma_{t}\left(\omega\right)\right\} _{t\in\left[N_{\mathrm{in}}+h,\,N\right]}\in\boldsymbol{J}_{\lambda}\left(\kappa,\,K_{h},\,d_{h}\right),\label{eq. H1 Ito vol}
\end{align}
where $\boldsymbol{J}_{\lambda}\left(\kappa,\,K_{h},\,d_{h}\right)\triangleq\left\{ \left\{ \varGamma_{t}\right\} _{t\in\left[N_{\mathrm{in}}+h,\,N\right]}:\,\left\{ \varGamma_{t}-\Delta\varGamma_{t}\right\} _{t\in\left[N_{\mathrm{in}}+h,\,N\right]}\in\boldsymbol{C}\left(\kappa,\,K_{h}\right);\,\left|\Delta\varGamma_{\lambda}\right|\geq d_{h}\right\} $,
$\Delta\varGamma_{\lambda}=\varGamma_{\lambda}-\lim_{s\uparrow\lambda}\varGamma_{s}$
and $\left\{ d_{h}\right\} $ is a decreasing sequence. The following
theorem extends Theorem \ref{Theoem Asymptotic H0 Distrbution Bmax and Qmax}
to the current setting.
\begin{thm}
\label{Theoem Asymptotic H0 Distrbution Bmax and Qmax Ito Vol}Let
$m_{T},\,\gamma_{m_{T}}$ and $\mathscr{V}$ as defined in Theorem
\ref{Theoem Asymptotic H0 Distrbution Bmax and Qmax}. Assume the
assumptions of Theorem \ref{Theoem Asymptotic H0 Distrbution Bmax and Qmax}
hold with Assumption \ref{Assumption Lipchtitz cont of Sigma} replaced
by Assumption \ref{Assumption Smooth of Sigma}. Under Condition \ref{Cond The-auxiliary-sequence Ito Vol}
and quadratic loss, the same results of Theorem \ref{Theoem Asymptotic H0 Distrbution Bmax and Qmax}
hold.
\end{thm}
\subsection{Asymptotic Local Power}
In this section we consider the behavior of $\mathrm{MQ}_{\mathrm{max},h}$
under a sequence of local alternatives.
\begin{assumption}
\label{Ass Local Power}We have the same assumptions as in Theorem
\ref{Theoem Asymptotic H0 Distrbution Bmax and Qmax Ito Vol} and
assume (i) in model \eqref{Mode for Y} we replace $\beta^{*}$ by
$\beta_{t}=\beta^{*}+\mu_{\beta,t}/\left(\log\left(T_{n}\right)n_{T}\right)^{1/4}$
where $\mu_{\beta,t}\in\mathbb{R}^{q}$ is $\mathbb{P}$-a.s. locally
bounded and adapted process; (ii) we set $\mu_{e,t}=0$ for all $t\geq0$;
(iii) we replace Assumption \ref{Assumption Roo-T Consistent beta }
by $\left\Vert \widehat{\beta}_{k}-\beta^{*}\right\Vert =\mu_{\beta,kh}/\left(\log\left(T_{n}\right)n_{T}\right)^{1/4}+O_{\mathbb{P}}\left(T^{-1/2}\right)$
uniformly in $k.$
\end{assumption}
Part (iii) is a consequence of part (i) as it can be easily verified.
Let
\begin{align*}
\mathrm{\widetilde{MQ}}{}_{\mathrm{max},h}\left(T_{n},\,\tau\right) & \triangleq\nu_{L}^{-1}\max_{i=n_{T},\ldots,\,T_{n}-n_{T}}\left|n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}\left(SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)-2\zeta_{\mu,j,+}\right)\right.\\
& \quad\left.-n_{T}^{-1}\sum_{j=i-n_{T}+1}^{i}\left(SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)-2\zeta_{\mu,j,-}\right)\right|,
\end{align*}
where
\begin{align*}
\zeta_{\mu,j,+} & \triangleq\mu'_{\beta,\left(T_{m}+\tau+j-1\right)h}\Sigma_{X,\left(T_{m}+\tau+i-1\right)h}\mu_{\beta,\left(T_{m}+\tau+j-1\right)h}/\left(\log\left(T_{n}\right)n_{T}\right)^{1/2}\\
\zeta_{\mu,j,-} & \triangleq\mu'_{\beta,\left(T_{m}+\tau+j-1\right)h}\Sigma_{X,\left(T_{m}+\tau+i-n_{T}-1\right)h}\mu_{\beta,\left(T_{m}+\tau+j-1\right)h}/\left(\log\left(T_{n}\right)n_{T}\right)^{1/2}.
\end{align*}
\begin{thm}
\label{Theorem Local Asym Power}Under Assumption \ref{Ass Local Power},
\begin{align*}
\sqrt{\log\left(m_{T}\right)}\left(n_{T}^{1/2}\nu_{L}^{-1}\right)\mathrm{\widetilde{MQ}}_{\mathrm{max},h}\left(T_{n},\,\tau\right)-2\log\left(m_{T}\right)-\frac{1}{2}\log\log\left(m_{T}\right)-\log3 & \overset{}{\Rightarrow}\mathscr{V},
\end{align*}
where $\mathscr{V},$ and $m_{T}$ are defined as in Theorem \ref{Theoem Asymptotic H0 Distrbution Bmax and Qmax}.
\end{thm}
\begin{rem}
(i) The theorem suggests that under the local alternatives $\beta_{t}=\beta^{*}+\mu_{\beta,t}/\left(\log\left(T_{n}\right)n_{T}\right)^{1/4}$
there is a bias term arising from the presence of $\zeta_{\mu}$.
This bias term does not vanish asymptotically and results in shifting
the center of the distribution. Moreover, it depends on the second
moments of the regressors and on the function $\mu_{\beta}$; (ii)
The theorem illustrates the sensitivity of the asymptotic power to
the form of the alternative. We can attempt to compare Theorem \ref{Theorem Local Asym Power}
with the local power result regarding the sup-Wald test of \citet{andrews:93}.
Unlike Theorem 4 in \citet{andrews:93}, our result suggests that
the location of the instability should not play any special role and
the power should not be sensitive to whether the break in predictive
ability occurs at middle sample or toward the tail of the sample.
This follows because of the local nature of our test statistic and
contrasts with classical tests for parameter instability and structural
change since their performance hinges on the location of the break
{[}see \citet{deng/perron:08}, \citet{kim/perron:09} and \citet{perron/yamamoto:16}
for additional results on the power of classical structural break
tests{]}. However, the magnitude of the break\textemdash here shrinking
at rate $\left(\log\left(T_{n}\right)n_{T}\right)^{1/4}$\textemdash under
our specification of the local alternatives is larger than the one
considered by \citet{andrews:93}\textemdash which shrinks at rate
$1/\sqrt{T}$. This implies a trade-off between location and magnitude
of the break, and it is consistent with the evidence provided in our
simulation study; (iii) Although not shown here, the local power
of the tests is the same when a subset of the vector $\beta$ is not
subject to shift.
\end{rem}
Theorem \ref{Theorem Local Asym Power} can be used to show that our
test possesses nontrivial power against alternatives for which the
parameter $\beta_{t}$ is time-varying and non-smooth.
\begin{cor}
\label{Corollary Local Power}Suppose the assumptions of the previous
theorem hold with $\beta_{t}=\beta^{*}+c\mu_{\beta,t}/\left(\log\left(T_{n}\right)n_{T}\right)^{1/4}$,
where $c\in\mathbb{R}$. If $\mu_{\beta,\cdot}$ and/or $\left\{ \mu_{\beta,t}\cdot\sigma_{X,t}\right\} _{t\geq0}$
is non-smooth, we have
\begin{alignat*}{1}
\underset{c\rightarrow\infty}{\lim}\underset{h\downarrow0}{\lim} & \mathbb{P}\left(\sqrt{\log\left(m_{T}\right)n_{T}}\nu_{L}^{-1}\mathrm{MQ}_{\mathrm{max},h}\left(T_{n},\,\tau\right)-2\log\left(m_{T}\right)-\frac{1}{2}\log\log\left(m_{T}\right)-\log3>\mathrm{cv}_{1-\alpha}\right)=1,
\end{alignat*}
where $\mathrm{cv}_{1-\alpha}$ is the level $\left(1-\alpha\right)$
critical value of the distribution of $\mathscr{V}$ and $\alpha\in\left(0,\,1\right)$.
\end{cor}
\section{Extensions\label{Section Extensions}}
A number of extensions is treated in our companion paper \citet{casini_FFb}.
As explained above, it would be useful to ensure that there are no
instabilities in the in-sample period $\left[0,\,N_{\mathrm{in}}\right].$
We propose a procedure that involves a pre-test about instability
on $\left[0,\,N_{\mathrm{in}}\right]$ for a given $N_{\mathrm{in}}$
chosen by the forecaster. Instabilities in the in-sample $\left[0,\,N_{\mathrm{in}}\right]$
are much easier to be detected relative to instabilities in the out-of
sample because they do not face the so-called ``contamination effect''.
The latter arises, for example under the recursive and rolling scheme,
when the instability originally occurring in the out-of-sample eventually
enters the moving in-sample window {[}cf. \citet{casini/perron_Oxford_Survey}
and \citet{perron/yamamoto:18}{]}. The consequence is that existing
tests face substantial power losses. This property is not shared by
our test statistics because of their local nature. Our procedure works
very well and we show through simulations that instabilities occurring
in the in-sample only or occurring both in the in-sample and in the
out-of-sample simultaneously, lead easily to rejection of the null
hypotheses relative to instabilities occurring in the out-of-sample
only\textemdash as we consider here.
A second issue is that, in this paper, we have considered processes
that have a continuous sample path under the null hypotheses. Thus,
it is of interest to extend the results to a setting that involves
jump processes which are important in high-frequency financial data.
This can be achieved by using techniques that are able to separate
the continuous part from the discontinuous part of a It� semimartingale
{[}see e.g. \citet{li/todorov/tauchen:17} and \citet{li/xiu:16}{]}.
Another important issue is the estimation of the time at which forecast
instability occurs. Once the null hypotheses has been rejected, a
forecaster may take into consideration the possibility of revising
the forecasting method and/or model. Hence, it becomes crucial to
learn some information about the timing of the instability. For example,
consider the case of a one-time structural change in a parameter of
the data-generating process at time $T_{b}^{0}=\left\lfloor T\lambda_{b}^{0}\right\rfloor $,
where $T_{b}^{0}$ is the \textit{break point} and $\lambda_{b}^{0}\in\left(0,\,1\right)$
is the \textit{fractional break date}. Once $H_{0}$ is rejected,
a forecaster would benefit from knowing that the forecast method originally
employed is found to statistically either under or over-perform over
part of the sample after $T_{b}^{0}$ relative to the part prior $T_{b}^{0}$.
Then, a forecaster would entertain the possibility of modifying the
forecast model in order to generate future forecasts for $Y_{t}$.
Not only the forecast model might be revised but most importantly,
knowledge of beginning of the instability at $T_{b}^{0}$ can be further
exploited to design the forecasting method for the future forecasts.
It would be inappropriate, for instance, to use a rolling scheme where
the rolling window used to construct the forecast include observations
prior to $T_{b}^{0}$ since those observations provide little informational
content for the purpose of predicting $Y_{t}$ after the change-point
$T_{b}^{0}$. On the other hand, this line of reasoning is justified
by this particular example and indeed in practice many issues arise
when dealing with the timing of the insatiability in our context.
For example, the exact form of the insatiability may be unknown. Under
the latter scenarios, there is no clear-cut break date $T_{b}^{0}$
that can be defined. Thus, it is less obvious how a forecaster should
proceed in those cases. Nonetheless, one can meaningfully think about
the timing of the instability by not just attempting to estimate $T_{b}^{0}$\textemdash which
is not clear how it is defined\textemdash but rather attempting to
detect the initial date in the sample after which the forecasts become
unstable as well as to detect the last date after which the forecasts
remain stable relative to the in-sample period. Since our test statistics
are local in nature, one can introduce a procedure which sequentially
tests the hypotheses $H_{0}$ in regions of the sample where $H_{0}$
has not yet been rejected. One then records the number of times for
which $H_{0}$ is rejected and estimates the corresponding change-point
dates. After ordering these change-point dates, one has finally access
to useful information such has the initial timing of the forecast
instability and the last part of the out-of-sample period which remains
stable. Such information can arguably be advantageous to the forecaster.
\section{\label{Section Simulation Study}Small-Sample Evaluation}
We now examine the empirical size and power of our proposed tests
and compare them to those of \citet{giacomini/rossi:09}, abbreviated
GR (2009). In particular, we consider both the uncorrected and corrected
version of the $t_{T_{m},T_{n},\tau}^{\mathrm{stat}}$ statistic of
GR (2009).\footnote{We use a superscript $\mathrm{c}$ to indicate the corrected version:
$t^{\mathrm{stat,c}}$.} Size and power properties for the Quadratic loss with fixed scheme
are reported in Section \ref{Subsection MC Size} and Section \ref{Subsection MC Power},
respectively. The Supplemental Material includes corresponding simulation
studies for the recursive and rolling schemes and for the Linex loss;
these results are not reported here because they are qualitatively
equivalent. Overall, one can draw the following conclusions from our
simulation study. In terms of size control, the statistics $\mathrm{B}_{\mathrm{max},h}$
and $\mathrm{Q}_{\mathrm{max},h}$ are comparable with the corrected
version $t^{\mathrm{stat,c}}$ proposed by GR (2009).\footnote{As shown by GR (2009), the uncorrected version of $t^{\mathrm{stat}}$
can be oversized for models that induce serial dependence in the forecast
losses. The authors then proposed a finite-sample correction and did
not consider $t^{\mathrm{stat}}$ further in their power analysis.
Similarly, GR (2009) showed that just using classical structural break
tests in this context is not very helpful as they might have statistical
power equals to the size in some cases. Moreover, simulations in
\citet{perron/yamamoto:18} confirmed that, under rolling and recursive
scheme, structural break tests suffer power losses which can be attributed
to a so-called ``contamination effect'' arising when the instability
enters the in-sample window {[}see also \citet{casini/perron_Oxford_Survey}{]}. } Moreover, the test $\mathrm{MQ}_{\mathrm{max},h}$ that uses overlapping
blocks is also comparable in terms of size. The same is not true for
$\mathrm{MB}_{\mathrm{max},h}$ because often it seems to be somewhat
liberal. Turning to the power comparison, each of our test statistics
$\mathrm{B}_{\mathrm{max},h}$, $\mathrm{Q}_{\mathrm{max},h}$ and
$\mathrm{MQ}_{\mathrm{max},h}$ displays significant power gains over
the $t_{T_{m},T_{n},\tau}^{\mathrm{stat}}$ statistics especially
as the period of instability (i) is comparatively short relative to
the total sample size and/or (ii) is not located at middle sample.
In the latter circumstances, the gains in power are, uniformly over
different data-generating processes and over parameter break magnitudes,
on the order of 30-40\%.
Throughout, we restrict attention to one-step ahead forecast horizon
(i.e., $\tau=1$), and we use the same loss function for estimation
and evaluation. We use our asymptotic results as an approximation
for the case where $h=1$ in our theoretical model in \eqref{Discretized Model Y}
and consider discrete-time DGPs. In models with serially correlated
losses (i.e., S2 and S6 below) for the statistics $\mathrm{Q}_{\mathrm{max},h}$
and $\mathrm{MQ}_{\mathrm{max},h}$ we employ the long-run variance
estimator from Theorem \ref{Theorem Long-Run Variance}. With regards
to the tests of GR (2009) we use the appropriate version of $t^{\mathrm{stat}}$
and of $t^{\mathrm{stat,c}}$.\footnote{As recommended by GR (2009) we set the truncation lag of their HAC
estimator equal to $\left\lfloor T_{n}^{1/3}\right\rfloor $; we also
the use the truncation lag $\left\lfloor 0.75T_{n}^{1/3}\right\rfloor $.}
\begin{rem}
Implementation of our tests statistics requires to choose the number
of blocks $m_{T}$\textemdash{} satisfying Condition \ref{Cond The-auxiliary-sequence}.
The finite-sample properties can be sensitive to the choice of $m_{T}$.
This is confirmed in our numerical study, where assigning larger values
to $m_{T}$ than the smallest one allowed by the condition may result
in oversized tests. Therefore, we recommend practitioners to set $m_{T}$
equal to the smallest integer as allowed by Condition \ref{Cond The-auxiliary-sequence}.
This is the strategy we have adopted in the Monte Carlo study of this
section, and as we will show, it results in approximately correct
size and good power across different data-generating mechanisms.
\end{rem}
\subsection{\label{Subsection MC Size}Empirical Size}
We consider discrete-time DGPs of the form
\begin{align}
Y_{t}=\mu+\beta X_{t-1}+e_{t}, & \qquad\qquad t=1,\ldots,\,T,\label{Eq. DGP Simulation Study}
\end{align}
for various in-sample and out-of-sample sizes and with a total sample
size ranging from $T=100$ to $T=500$. Note that \eqref{Eq. DGP Simulation Study}
is a special case of the theoretical model with a sampling interval
$h=1$. We consider six versions of \eqref{Eq. DGP Simulation Study},
where the first and second specification (S1 and S2 below) are calibrated
to the Phillips curve model of U.S. inflation from \citet{staiger/stock/watson:97}:
S1 involves $\mu=2.73$, $\beta=-0.44$, and where $\left\{ X_{t}\right\} $
and $\left\{ e_{t}\right\} $ are independent sequences of zero-mean
i.i.d. Gaussian disturbances with unit variance; S2 is the same as
S1 but with ARCH errors $e_{t}=\sigma_{e,t}u_{t}$, $\sigma_{e,t}=1+0.5e_{t-1}^{2}$
with $u_{t}\sim\mathscr{N}\left(0,\,1\right)$; S3 specifies $\left\{ X_{t}\right\} $
to follow a zero-mean Gaussian AR(1) with autoregressive coefficient
0.4, $\beta=1$ and $e_{t}\sim\mathscr{N}\left(0,\,0.49\right)$ independent
of $X_{t}$; S5 is a model with a lagged dependent variable $X_{t-1}=Y_{t-1}$,
$\mu=0$, $\beta=0.3$ and $e_{t}\sim\mathscr{N}\left(0,\,0.49\right)$;
S6 involves serially correlated disturbances $e_{t}=0.3e_{t-1}+u_{t}$,
$u_{t}\sim\mathscr{N}\left(0,\,1\right)$.
Table \ref{Table S1}-\ref{Table S2} report the rejection rates for
significance levels $\alpha=0.05$ and $0.10$ for model S1-S2. Results
for the other DGPs can be found in Table \ref{Table Size S3}-\ref{Table S6}
in the Supplement. We first focus on i.i.d. forecast losses (i.e.,
models S1 and S3-S5). Both $\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{Q}_{\mathrm{max},h}$
are well-sized. As the sample size increases their performance improves
and we note that their rejection frequencies are closer to the nominal
level when the in-sample size is one half of the total sample. In
model S1, when the in-sample size is $0.25T,$ $\mathrm{B}_{\mathrm{max},h}$
and $\mathrm{Q}_{\mathrm{max},h}$ tend to be slightly conservative
while the opposite occurs when in-sample size is $0.75T$. The version
of $\mathrm{B}_{\mathrm{max},h}$ that uses overlapping blocks $\left(\mathrm{MB}_{\mathrm{max},h}\right)$
can be quite liberal (cf. models S1 and S3). In contrast, $\mathrm{MQ}_{\mathrm{max},h}$
seems to control the size well, though it tends to be slightly liberal
but that depends on the relative size of the in-sample and out-of-sample
windows. We observe that there is no clear pattern in size performance
for our test statistics as we raise the sample size $T$. The reason
is straightforward: as we raise $T$ we also need to adjust the choice
of $m_{T}$ (the number of blocks) in accordance with Condition \ref{Cond The-auxiliary-sequence}.
This explains why, for example, in Table \ref{Table S1}, top panel,
the empirical size of $\mathrm{Q}_{\mathrm{max},h}$ for $\left(T_{m}=100,\,T_{n}=100\right)$
is better than for $\left(T_{m}=150,\,T_{n}=150\right)$. Turning
to the $t^{\mathrm{stat}}$ statistics of \citet{giacomini/rossi:09},
the uncorrected version performs better than the corrected version
since the latter systematically displays an empirical size 2-3\% below
the nominal level. We can conclude that in models with i.i.d. errors
the statistics $\mathrm{B}_{\mathrm{max},h}$, $\mathrm{Q}_{\mathrm{max},h}$,
$\mathrm{MQ}_{\mathrm{max},h}$ and $t^{\mathrm{stat}}$ are comparable
in terms of empirical size, whereas $\mathrm{MB}_{\mathrm{max},h}$
and $t^{\mathrm{stat,c}}$ tend to over-reject and under-reject, respectively.
Let us now turn to models with serially correlated losses. When the
disturbances follow an ARCH process, (cf. model S2, Table \ref{Table S2}),
we observe that both statistics that do not use overlapping blocks,
$\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{Q}_{\mathrm{max},h}$,
show reasonable size control. The same feature applies to $\mathrm{MQ}_{\mathrm{max},h}$
while $\mathrm{MB}_{\mathrm{max},h}$ displays rejection rates that
are systematically above the significance level. It also appears that
the corrected version of the statistic of GR (2009) is now regularly
below the nominal level. In contrast, the uncorrected version $t^{\mathrm{stat}}$
seems to control size well. When the errors follow an autoregressive
process (cf. model S6), Table \ref{Table S6} shows that $t^{\mathrm{stat}}$
and $\mathrm{MB}_{\mathrm{max},h}$ are arbitrarily oversized for
all sample sizes. $\mathrm{MQ}_{\mathrm{max},h}$ and $t^{\mathrm{stat,c}}$
possess rejection rates frequently below the desired nominal level.
The statistic that shows the best empirical sizes across different
$T$ is $\mathrm{Q}_{\mathrm{max},h}$.
Overall, our analysis on the size properties of the tests suggests
that when the DGP involves i.i.d. errors it is fair to compare $\mathrm{B}_{\mathrm{max},h}$,
$\mathrm{Q}_{\mathrm{max},h}$, $\mathrm{MQ}_{\mathrm{max},h}$ and
$t^{\mathrm{stat}}$ whereas the rejection rates of $\mathrm{MB}_{\mathrm{max},h}$
and $t^{\mathrm{stat,c}}$ tend to deviate systematically from the
nominal level. When there are autocorellated errors, it is difficult
to compare $t^{\mathrm{stat}}$ and $\mathrm{MB}_{\mathrm{max},h}$
with the other statistics because the former can be highly oversized.
The statistics that appear to perform better in terms of approximate
size control uniformly over different data-generating mechanisms are
$\mathrm{Q}_{\mathrm{max},h}$ and $\mathrm{MQ}_{\mathrm{max},h}$.
\subsection{\label{Subsection MC Power}Empirical Power}
We report the small sample power of the tests under various sources
of forecast instability. We consider several sample sizes $T$ as
well as several designs varying for the distribution of the total
sample between in-sample and out-of-sample window. The break date\textemdash or
the date of the first change-point when more complicated designs are
used\textemdash is denoted by $T_{b}^{0}=T\lambda_{0}$, where $\lambda_{0}\in\left(0,\,1\right)$
is the fractional break date. We shall bring special attention to
the location of $T_{b}^{0}$ in the sample as well as to the duration
of the instability (i.e., $T-T_{b}^{0}$). We shall see that both
factors are actually important for the performance of the methods
proposed by \citet{giacomini/rossi:09} while our test statistics
being local in nature possess essentially uniform power over distinct
locations $T_{b}^{0}$. Furthermore, our definition of forecast instability
does not demand any relationship between the stable and unstable period
and thus it is useful to examine the differences in power properties
when a one-time change-point is present relative to when short-lasting
instabilities arise.
We consider both discrete shifts\textemdash a structural break\textemdash and
recurrent changes in a parameter: model P1a (break in a regression
coefficient): $Y_{t}=2.73-0.44X_{t-1}+\delta X_{t-1}\mathbf{1}\left\{ t>T_{b}^{0}\right\} +e_{t}$,
where $X_{t-1}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1\right)$ and
$e_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1\right)$; model P1b:
it is the same as model P1a but with $X_{t-1}\sim\mathrm{i.i.d.}\mathscr{N}\left(1,\,1\right)$;
model P2: $Y_{t}=X_{t-1}+\delta X_{t-1}\mathbf{1}\left\{ t>T_{b}^{0}\right\} +e_{t}$,
where $X_{t-1}$ is a Gaussian AR(1) with autoregressive coefficient
0.4 and unit variance, and $e_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,0.49\right)$;
model P3 (recurrent break in mean): $Y_{t}=\beta_{t}+e_{t}$, where
$\beta_{t}$ switches between $\delta$ and $0$ every $p$ periods
and $e_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,0.64\right)$; model
P4 (single break in variance): $Y_{t}=0.5X_{t-1}+\left(1+\delta\mathbf{1}\left\{ t>T_{b}^{0}\right\} \right)e_{t}$
where $X_{t-1}\sim\mathrm{i.i.d.}\mathscr{N}\left(1,\,1\right)$ and
$e_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1\right)$; model P5
(recurrent break in variance): $Y_{t}=\mu+\left(1+\beta_{t}\right)e_{t}$,
where $\beta_{t}$ switches between $\delta$ and $0$ every $p$
periods and $e_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,0.49\right)$;
model P6 (lagged dependent variable): $Y_{t}=\delta\mathbf{1}\left\{ t>T_{b}^{0}\right\} +0.3Y_{t-1}+e_{t}$,
$e_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,0.49\right)$; model
P7 (ARCH disturbances): $Y_{t}=2.73-0.44X_{t-1}+\delta X_{t-1}\mathbf{1}\left\{ t>T_{b}^{0}\right\} +e_{t}$,
where $X_{t-1}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1.5\right)$
and $e_{t}=\sigma_{t}u_{t}$, $\sigma_{t}^{2}=0.5+0.5e_{t-1}^{2}$,
$u_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1\right)$; model P8
(autocorrelated errors): $Y_{t}=1+X_{t-1}+\delta X_{t-1}\mathbf{1}\left\{ t>T_{b}^{0}\right\} +e_{t}$,
where $X_{t-1}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1.4\right)$
and $e_{t}=0.4e_{t-1}+u_{t}$, $u_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1\right)$.
For models that do not involve recurrent changes we also consider
power comparisons when the instability lasts only for some period
of time as opposed to the post-$T_{b}^{0}$ period. This requires
replacing $\mathbf{1}\left\{ t>T_{b}^{0}\right\} $ in models P1-P2,
P4 and P6-P8 with $\mathbf{1}\left\{ T_{b}^{0}<t\leq T_{b}^{0}+p\right\} $
where $p$ is the number of consecutive observations in which the
forecast model is unstable. The value of $p$ depends on the sample
size $T$. For example, when $T=100$ we set $p=10$; when $T=200$
we set $p=20$ and so on.\footnote{Note that for $\left(T_{m}=50,\,T_{n}=50\right)$ the value $p=10$
corresponds to a period of instability lasting for one-fifth of the
out-of-sample; thus, the duration of the instability is nontrivial
and consistent with forecasting applications. See the notes to each
figure for the other values of $p$. The title of a figure corresponding
to a short-lasting instability is labeled ``short-term instability''.} The case of short-term instability is the most prevalent in empirical
work because it is very unlikely that a professional forecaster would
use a poor-performing predictor or forecast model for many consecutive
years (e.g., the whole out-of-sample).
Figure \ref{Fig_P1a_150_078}-\ref{Fig_P4_450_068_st} in the Appendix
plot the power functions for models P1a, P4 and P7. Figure \ref{Fig_P1b_150}-\ref{Fig_P2_1200_078_st}
in the Supplement plot the power functions for the remaining DGPs.
They include several sample sizes ranging from $T=100$ to $T=500,$
several in-sample and out-of-sample sizes as well as different locations
$\lambda_{0}$ of the breaks. We begin with considering general instabilities
first and then move to short-term instabilities. Figure \ref{Fig_P1a_150_078}-\ref{Fig_P1a_230_078}
reports the results for model P1a. When $T=100,\,150$ Figure \ref{Fig_P1a_150_078}
shows that our tests have good power against model P1a while the tests
of GR (2009) seem to be less powerful. For example, when the break
date is at $T_{0}=0.8T$ our tests display reasonable power. However,
both $t^{\mathrm{stat}}$ statistics of GR (2009) perform significantly
worse and the associated power curve is bounded away from one even
for a very large break size $\delta=3$. This feature disappears when
we raise the sample size to $T\geq200$ and maintain the break date
at $T_{0}=0.8T$; see Figure \ref{Fig_P1a_230_078}. The latter figure
also shows that for large sample sizes and instabilities that last
for more than 50\% of the out-of-sample (top panels) all tests have
good power even though the $t^{\mathrm{stat}}$ statistics of GR (2009)
have slightly higher power. The power turns to be essentially the
same when $\lambda_{0}=0.8$ (i.e., the instability only lasts for
40\% of the out-of-sample). For model P2, Figure \ref{Fig_P2_120_078}
plots the power functions for $T=100,\,200$ and $\lambda_{0}=0.7,\,0.8.$
Except for the pair $\left(T=200,\,\lambda_{0}=0.7\right)$ (cf. top-right
panel) for which our tests and the $t^{\mathrm{stat}}$-type tests
display roughly the same power, it is clear that our tests are more
powerful than the $t^{\mathrm{stat}}$ tests (both corrected and uncorrected
version). The power gains are substantial and range from 20\% to 40\%.
Moreover, as for model P1a and P1b when the instability lasts for
less than 50\% of the out-of-sample (cf. $\lambda_{0}=0.8$; bottom-left
panel) the statistics $\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{Q}_{\mathrm{max},h}$
achieve trivial power already for a break magnitude $\delta=1.5$
whereas the $t^{\mathrm{stat}}$ tests of GR (2009) display rejection
rates below 60\% even when $\delta=2$ and yet below 70\% when $\delta=2.5$;
that is, their power function does not attain unit power even for
very large break magnitudes. These properties characterize all models
with i.i.d. errors and extend to model with lagged dependent variables
as predictors (cf. model P7; Figure \ref{Fig_P6_2300_078} in the
Supplement).
Let us now turn to recurrent breaks in the mean. For recurrent breaks
we implement the statistics $\mathrm{MB}_{\mathrm{max},h}$ and $\mathrm{MQ}_{\mathrm{max},h}$
that use overlapping blocks. Figure \ref{Fig_P3_230_056} plots the
power curves for model P3. All tests have power and their performance
is essentially the same. Figure \ref{Fig_P4_230_068} corresponds
to model P4 (single break in the variance) and shows that when the
instability begins in the second half of the out-of-sample (cf. $\lambda_{0}=0.8$;
bottom panels) our tests $\mathrm{MB}_{\mathrm{max},h}$ and $\mathrm{MQ}_{\mathrm{max},h}$
achieves good power while the $t^{\mathrm{stat}}$-type tests have
little power that does not attain unity even for a large break magnitude
$\delta=1.5$. When there are recurrent breaks in the variance as
in model P5, Figure \ref{Fig_P5_230_056} shows that the our tests
$\mathrm{MB}_{\mathrm{max},h}$ and $\mathrm{MQ}_{\mathrm{max},h}$
and the $t^{\mathrm{stat}}$-type tests have all good power and their
performance is analogous.
Let us now consider models with either ARCH errors or autocorrelated
errors. Observe that the latter models both imply that the forecast
losses are serially correlated. Figure \ref{Fig_P7_2300_078} shows
that when the errors follow an ARCH(1) process the statistic $\mathrm{Q}_{\mathrm{max},h}$
based on the asymptotic variance estimator $\widehat{\nu}_{L}^{2}$
performs well in terms of empirical power. In contrast, the $t^{\mathrm{stat}}$-type
tests fail as their power is non-monotonic, never reaches 20\% and
it decreases to zero as the magnitude of the break rises.\footnote{We actually implemented the $t^{\mathrm{stat}}$ by using either \citeauthor{andrews:91}'s
(1991) or \citeauthor{newey/west:87}'s (1987) estiamtor of the long-run
variance. We also experimented different choices for the truncation
lag. The results, however, were unchanged. We suspect that this property
depends on the estimation of the long-run variance in our forecasting
context which can be challenging due to small sample sizes and to
the presence of breaks. The same issues were found in \citet{martins/perron:16}
and \citet{fossati:17}.} We note that the version of $\widehat{\nu}_{L}$ that uses more blocks
is less precises. The same results hold true when the disturbances
are autocorrelated; see Figure \ref{Fig_P8_2300_078}.
Finally, we consider short-term instabilities in Figure \ref{Fig_P1_150_078_st}-\ref{Fig_P4_450_068_st}.
It is straightforward to recognize a general pattern: the tests of
GR (2009) have little power whereas our tests possess good empirical
power against all data-generating processes, break locations and sample
sizes. Furthermore, the small sample power properties are uniform
over the location of the instability and over the relative size of
the in-sample and out-of-sample windows. The latter property is important
in practice because forecast instabilities are frequently short-lived.
To sum up, our test statistics perform well in controlling the size,
even though the versions that use overlapping blocks are somewhat
liberal. For our tests, empirical size being close to the significance
level is a feature that holds over different DGPs and sample sizes.
Turning to power comparison, there is clear evidence that our tests
are reliable in that they have good power against different form of
instabilities. There appears to be substantial power gains relative
to existing methods especially when the instability (i) is short-lasting
and/or (ii) is located toward the tail of the out-of-sample. These
properties characterize both statistics using non-overlapping and
overlapping blocks.
\section{\label{Section Conclusions}Conclusions}
We have formalized the concepts of forecast instability and forecast
failure. Our definition poses at the center the economic forecaster
and emphasizes the importance of the time duration of the instability.
We assume the data arise as an outcome of an underlying system of
stochastic differential equations which then implies that we can approximate
the sequence of forecast losses by a continuous-time stochastic process.
We have built a testing framework based on the local pathwise properties
of that process and have adopted an infill asymptotics to derive the
null distribution of the test statistics. The null distribution follows
an extreme value distribution. Our results can be used to test whether
the predictive ability of a given forecast model changes over time
and can be applied in forecasting exercises involving either low-frequency
as well as high-frequency macroeconomic and financial variables. The
simulation study confirms that there are substantial power gains especially
when the instability (i) is short-lasting and/or (ii) is located toward
the tail of the out-of-sample. Our framework allows for misspecification,
different types of parameter instability and arbitrary forms of non-stationarity
such as heteroskedasticity and serial correlation. Our continuous-time
specification and associated continuous record asymptotic scheme can
provide a promising complementary framework to the classical approach
for forecasting in economics.
\newpage{}
\begin{singlespace}
\bibliographystyle{econometrica}
\bibliography{References}
\addcontentsline{toc}{section}{References}
\end{singlespace}
\newpage{}
\clearpage
\pagenumbering{arabic}