Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
127,865 characters · 25 sections · 119 citation commands
Tests for Forecast Instability and Forecast Failure under a Continuous Record Asymptotic Framework
Correspondence: \\
\\ Ducument Information: This electronic document includes the working paper and supplemental materials. The Supplement starts on pdf page 48. Monte Carlo simulations were conducted in Matlab. The most recent working paper version can be downloaded from the author's webpage. The associated computer package, if available, should be downloaded from the same webpage.
\thispagestyle{empty}
\setcounter{page}{1}
\raggedbottom
{\bf JEL Classification:} C12, C22, C52, C53 \\ {\bf Keywords:} Asymptotic distribution, break date, continuous-time, forecast failure, forecast instability, infill asymptotics, parameter instability, predicitve ability, semimartingale
\onehalfspacing
Since the seminal contribution of Klein klein:69,klein:71, economic forecasts had been built upon the presumption that the relationships between economic variables remain stable over time. However, the last decades have been subject to many social-economic episodes and technological advancements that have led economists to reconsider the assumption of model stability. The resonant empirical evidences documented in, among others, perron:89 and stock/watson:96 {[}see also the recent survey by ng/wright:13{]} have motivated the development of econometric methods that detect such instabilities\textemdash most work directed toward structural changes\textemdash and estimate the actual dates at which economic relationships change. Yet, the issue of parameter insatiability is not limited to model estimation. In the forecasting literature, there has been a widespread concordance that the major issue that prevents good forecasts for economic variables is parameter instability\textemdash and structural changes as a special case\textemdash {[}cf. banerjee/marcellino/masten:08, clements/hendry:98 (1998, 2006)\nocite{clements/hendry:06}, elliott/timmermann:16, giacomini:15, giacomini/rossi:15, inoue/rossi:11, clark/mccracken:05, pesaran/pettenuzzo/timmermann:06 and rossi:13a{]}.
This paper develops a statistical setting under infill asymptotics to address the issue of testing whether the predictive ability of a given forecast model remains stable over time. ng/wright:13 and stock/watson:03 explain that there has been abundant evidence for which a predictor that has performed well over a certain time period may not perform as well during other subsequent periods. For example, gilchrist/zakrajsek:12 proposed a new credit spread index and showed that a residual component labeled as the excess bond premium\textemdash the credit spread adjusted for expected default risk\textemdash has considerable predictive content for future economic activity. They documented that this forecasting ability is stronger over the subsample 1985-2010 rather than over the full sample starting from 1973.\footnote{They reported that structural change tests provide some statistical evidence for a break in a coefficient associated with financial indicators\textemdash more specifically the coefficient on the federal funds rate. Given the latter evidence and the well-documented change in the conduct of monetary policy in the late 1970s and the early 1980s, it seems plausible to split the sample in 1985 (see p. 1709 and footnote 11 in their paper).} The latter finding can be attributed to a more developed bond market in the 1985-2010 subsample. Relatedly, giacomini/rossi:10 and ng/wright:13 further examined this finding and found that indeed the predictive ability of commonly used term and credit spreads is unstable and somehow episodic. The latter authors suggested that credit spreads may be more useful predictors of economic activity in a more highly leveraged economy and that recent developments in financial markets translate into credit spreads containing more information than they had previously. We refer to such temporal instability for a given forecasting method as forecast instability or more specifically, as forecast failure. These terminologies are not new to professional forecasters as they were informally introduced by clements/hendry:98 and generalized in econometric terms by giacomini/rossi:09 who interpreted forecast breakdown (or forecast failure) as a situation in which the out-of-sample performance of a forecast model significantly deteriorates relative to its in-sample performance. Our approach is to formally define forecast instability from the economic forecaster's perspective.\footnote{We use the terminology “instability” because not only the deterioration but also the improvement of the performance of a given forecast model over time can provide useful information to the forecaster.} We emphasize that a forecast failure may well result from a short period of instability within the out-of-sample and not necessarily require that the instability be systematic in the sense of persisting throughout the whole out-of-sample period. That is, consistency of a forecast model's performance with expected performance given the past should hold not only throughout the out-of-sample but also in any sub-sample of the latter. Indeed, many documented episodes of forecast failure seemed to arise from parameters nonconstancy data-generating processes over relatively short time periods compared to the total sample size. Hence, the desire of focusing on statistical tests being able to detect short-lasting instabilities is intuitive: if a test for forecast failure needs the deterioration of the forecasting ability to last for, say, at least half of the total sample in order to have sufficiently high power to reject the null hypotheses, then this test would not perform very well in practice because instability can be short-lasting. Furthermore, the occurrence of recurrent structural instabilities or multiple breaks that compensate each other in the out-of-sample might lead a forecast model to perform, on average, in a similar fashion as in the in-sample period. However, should a forecaster know about those recurrent changes she would conceivably revise its forecast model to adapt to the unstable environment. Hence, we introduce the following definition.
Th definition poses at the center the economic forecaster and consequently it is not merely a statistical definition; rather, it is based on an equilibrium concept. Since forecasting constitutes a decision theoretic problem, it should be from the forecaster perspective that a given forecast model is deemed to have failed. It is implicit from the definition to distinguish between forecasting method and model. Two forecasters may share the same forecast model\textemdash the relationship between the variable of interest and the predictor\textemdash but use different methods (e.g., recursive scheme versus rolling scheme). Thus, instability refers to a given method-model pair. The object of the definition is predictive ability. Since the latter can be measured differently by different loss functions, then the definition applies to a given choice of the loss function. A notable aspect of the definition is the reference to the time span of the historical performance and of the putative period of instability. They need not be related. Consider a given forecasting strategy which has performed well during, say, the Great Moderation (i.e., from mid-1980s up to prior the beginning of the Great Recession in 2007). Assume that during the years 2007-2012 this method endures a time of poor performance and returns to perform well thereafter. According to our definition, this episode constitutes an example of forecast instability. However, if one designs the forecasting exercise in such a way that half of the sample is used for estimation and the remaining half for prediction, then this relatively short period of instability gets “averaged-out” from tests which simply compare the in-sample and out-of-sample averages. Conceivably, such tests would not reject the null hypotheses of no forecast failure while it seems that a forecaster would had revised its strategy during the crisis if she had known about such occurring under-performance in the present and immediate future period. Finally, detection of forecast instability does not necessarily mean that a forecast model should be abandoned. In fact, its performance may have improved over time. Yet, even if forecast instability is induced by performance deterioration, a forecaster might not end up switching to a new predictor. For example, entering a state of high variability might lead to poor performance even if the forecast model is still correct. Hence, our definition uses the term reconsider. Continuing with the above example, a forecaster may reconsider the choice of the forecasting window since a longer window may now produce better forecasts while keeping the same forecast model. In other words, knowledge of forecast instability is important because indicates that care must be exercised to assess the source of the changes.\footnote{Economists have documented episodes of forecast failure in many areas of macroeconomics. In the empirical literature on exchange rates a prominent forecast failure is associated with the Meese and Rogoff's puzzle {[}cf. meese/rogoff:83, cheung/chinn/garcia:05, and rossi:13b for an up-to-date account{]}. In the context of inflation forecasting, forecast failures have been reported by atkeson/ohanian:01 and stock/watson:09. For forecast instability concerning other macroeconomic variables see the surveys of stock/watson:03 and ng/wright:13.}
The theoretical implication is that in this paper our tests for forecast instability shall be based on the local behavior of the sequence of realized forecast losses. This is opposite to existing tests for forecast instability\textemdash and classical structural change tests more generally\textemdash which instead rely on a global and retrospective methodology merely comparing the average of in-sample losses with the average of out-of-sample losses. While maintaining approximately correct nominal size, our class of test statistics achieves substantial gains in statistical power relative to previous methods. Furthermore, as the initial timing of the instability moves away from middle sample toward the tail of the out-of-sample, the gains in power become considerable.
In this paper, we set out a continuous record asymptotic framework for a forecasting environment where $T$ observations at equidistant time intervals $h$ are made over a fixed time span $\left[0,\,N\right],$ with $N=Th.$ These observations are realizations from a continuous-time model for the variable to be forecast and for the predictor. From these discretely observed realizations we compute a sequence of forecasts using either a fixed, recursive or rolling scheme. To this sequence of forecasts there corresponds a continuous-time process which satisfies mild regularity conditions and that under the null hypotheses possesses a continuous sample-path. We exploit this pathwise property to base an hypothesis testing problem on the relative performance of a given forecast model over time. Under the hypotheses we expect the sequence of losses to display a smooth and stable path. Any discontinuous or jump behavior followed by a (possibly short) period of substantial discrepancy from the same path over the in-sample period provides evidence against the hypotheses. Our asymptotic theory involves a continuous record of observations where we let the sample size $T$ grow to infinity by shrinking the sampling interval $h$ to zero with the time span kept fixed at $N$, thereby approaching the continuous-time limit.
Our underlying probabilistic model is specified in terms of continuous It� semimartingales which are standard building blocks for analysis of macro and financial high-frequency data {[}cf. \citet*{andersen/bollerslev/diebold/labys:01}, andersen/fusari/todorov:16, bandi/reno:16 and barndorff/shephard:04{]}; the theoretical methodology is thus related to that of casini/perron_CR_Single_Break, li/todorov/tauchen:17, li/xiu:16 and mykland/zhang:09.\footnote{Recent work by li/patton:17 extends standard methods for testing predictive accuracy of forecasts to a high-frequency financial setting.} The framework is not only useful for high-frequency data; in particular, recent work of casini/perron_CR_Single_Break (2017a, 2017b)\nocite{casini/perron_Lap_CR_Single_Inf} has adopted this continuous-time approach for modeling time series regression models with structural changes fitted to low-frequency data (e.g., macroeconomic data that are sampled at weekly, monthly, quarterly, annual frequency, etc.). They have showed that this continuous-time approach delivers a better approximation to the finite-sample distributions of estimators in structural change models and inference is more reliable than previous methods based on classical long-span asymptotics.
The classical approach to economic forecasting for macroeconomic variables is to formulate models in discrete-time and then base inference on long-span asymptotics where the sample size increases without bound and the sampling interval remains fixed {[}cf. diebold/mariano:95, giacomini/white:06 and west:96{]}. There are crucial distinctions between this classical approach and the setting introduced in this paper. Under long-span asymptotics, identification of parameters hinges on assumptions on the distributions or moments of the studied processes {[}cf. the specification of the null hypotheses in giacomini/rossi:09{]}, whereas within a continuous-time framework, unknown structural parameters are identified from the sample paths of the studied processes. Hence, we only need to assume rather mild pathwise regularity conditions for the underlying continuous-time model and avoid any ergodic or weak-dependence assumption. As in casini/perron_CR_Single_Break, our framework encompasses any time series regression model allowing for general forms of non-stationarity such as heteroskedasticty and serial correlation.
Given a null hypotheses stated in terms of the path properties of the sequence of losses, we propose a test statistic which compares the local behavior of the sequence of surprise losses defined as the difference between the out-of-sample and in-sample losses. More specifically, our maximum-type statistic examines the smoothness of the sequence of surprise losses as the continuous-time limit is approached. Under the hypotheses, the continuous-time analogue of the sequence of losses follows a continuous motion and any deviation from such smooth path is interpreted as evidence against the hypotheses. The null distribution of the test statistic is non-standard and follows an extreme value distribution. Therefore, our limit theory exploits results from extreme value theory as elaborated by bickel/rosenblatt:73 and galambos:87.\footnote{In nonparametric change-point testing, related works are wu/zhao:07 and bibinger/jirak/vetter:16.}
We propose two versions of the test statistic: one that is self-normalized and one that uses an appropriate estimator of the asymptotic variance. The test statistic is defined as the maximal deviation between the average surprise losses over asymptotically vanishing time blocks. Further, we consider extensions of each of these statistics which use overlapping rather than non-overlapping blocks. Although they should be asymptotically equivalent, the statistics based on overlapping blocks are more powerful in finite-samples. In a framework where one allows for model misspecification, the problem of non-stationarity such as heteroskedastcity and serial correlation in the forecast losses should be taken seriously. Given the block-based form our test statistics we derive an alternative estimator of the long-run variance of the forecast losses. This estimator differs from the popular estimators of andrews:91 and newey/west:87 {[}see muller:07 for a review{]} and it is of independent interest. Finally, we extend results to settings that allow for stochastic volatility, and we conduct a local power analysis and highlight a few differences of our testing framework from the structural change test of andrews:93. Related aspects, such as estimating the timing of the instability and covering high-frequency setting with jumps, are being considered in a companion paper.
The rest of the paper is organized as follows. Section (ref) introduces the statistical setting, the hypotheses of interest and the test statistics. Section (ref) derives the asymptotic null distribution under a continuous record. We discuss the estimation of the asymptotic variance in Section (ref). Some extensions and a local power analysis are presented in Section (ref). Additional elements that are covered in our companion paper are briefly described in Section (ref). A simulation study is contained in Section (ref). Section (ref) concludes the paper. The supplemental material to this paper contains all mathematical proofs and additional simulation experiments.
Section (ref) introduces the statistical setting with a description of the forecasting problem and the sampling scheme considered throughout. The underlying continuous-time model and its assumptions are introduced in Section (ref). In Section (ref) we set out the testing problem and state the relevant null and alternative hypotheses. The test statistics are presented in Section (ref). Throughout we adopt the following notational conventions. All limits are taken as $T\rightarrow\infty$, or equivalently as $h\downarrow0$, where $T$ is the sample size and $h$ is the sampling interval. All vectors are column vectors and for two vectors $a$ and $b$, we write $a\leq b$ if the inequality holds component-wise. For a sequence of matrices $\left\{ A_{T}\right\} ,$ we write $A_{T}=o_{\mathbb{P}}\left(1\right)$ if each of its elements is $o_{\mathbb{P}}\left(1\right)$ and likewise for $O_{\mathbb{P}}\left(1\right).$ If $x$ is a non-stochastic vector, $\left\Vert x\right\Vert $ denotes the its Euclidean norm, whereas if $x$ is a stochastic vector, the same notation is used for the $L^{2}$ norm. We use $\left\lfloor \cdot\right\rfloor $ to denote the largest smaller integer function and for a set $A,$ the indicator function of $A$ is denoted by $\mathbf{1}_{A}$. A sequence $\left\{ u_{kh}\right\} _{k=1}^{T}$ is $\textrm{i.i.d.}$ if the $u_{kh}$ are independent and identically distributed. We use $\overset{\mathbb{P}}{\rightarrow},\,\Rightarrow$ to denote convergence in probability and weak convergence, respectively. $\mathcal{M}_{p}^{\textrm{c�dl�g}}$ is used for the space of $p\times p$ positive define real-valued matrices whose elements are c�dl�g. The symbol “$\triangleq$” is definitional equivalence.
The continuous-time stochastic process $Z\triangleq\left(Y,\,X'\right)$ is defined on a filtered probability space $\left(\Omega,\,\mathscr{F},\,\left\{ \mathscr{F}_{t}\right\} _{t\geq0},\,\mathbb{P}\right)$ and takes value in $\mathbf{Z}\subseteq\mathbb{R}^{q+1}$ where $\left\{ Y_{t}\right\} _{t\geq0}$ is the variable to be forecast and $\left\{ X_{t}\right\} _{t\geq0}$ are the predictor variables. The index $t$ is defined as the continuous-time index and we have $t\in\left[0,\,N\right]$, where $N$ is referred to as the time span. In this paper, $N$ will remain fixed. That is, the unobserved process $Z_{t}$ evolves within the fixed time horizon $\left[0,\,N\right]$ and the econometrician records $T$ of its realizations, with a sampling interval $h$, at discrete-time points $h,\,2h\ldots,\,Th$, where accordingly $Th=N.$ A continuous record asymptotic framework involves letting the sample size $T$ grow to infinity by shrinking the time interval $h$ to zero at the same rate so that $N$ remains fixed. The index $k$ is used for the observation (or tick) times $k=1,\ldots,\,T$.
The objective is to generate a series $\left\{ Y_{\left(k+\tau\right)h}\right\} $ of $\tau$-step ahead forecasts. We shall adopt an out-of-sample precedure whereby splitting the time span $\left[0,\,N\right]$ into an in-sample and out-of-sample window, $\left[0,\,N_{\mathrm{in}}\right]$ and $\left[N_{\mathrm{in}}+h,\,N\right]$, respectively.\footnote{Indeed, $\left[0,\,N_{\mathrm{in}}\right]$ corresponds to the in-sample window only for the fixed forecasting scheme to be introduced later\textemdash e.g., the rolling scheme only uses the most recent span of data of length $N_{\mathrm{in}}$. A minor and straightforward modification to this notation should be applied when the recursive and rolling schemes are considered. However, for all methods $N_{\mathrm{in}}$ indicates the artificial separation such that $N_{\mathrm{in}}+h$ is the beginning of the out-of-sample period.} The latter two time horizons are supposed to be fixed and therefore within the in-sample (or prediction) window a sample of size $T_{m}$ is observed whereas within the out-of-sample (or estimation) window the sample is of size $T_{n}=T-T_{m}-\tau+1$. We consider a general framework that allows for the three traditional forecasting schemes: (1) a fixed forecasting scheme with discrete-time observations $h,\,2h,\ldots,\,\left(T_{m}-1\right)h,\,T_{m}h=N_{\mathrm{in}}$; (2) a recursive forecasting scheme where at time $kh$ the prediction sample includes observations $h,\ldots,\,\left(k-1\right)h,\,kh$; (3) a rolling forecasting scheme where the time span of the rolling window is fixed and of the same length as $N_{\mathrm{in}}$ (i.e., at time $kh$ the in-sample window includes observations $kh-T_{m}h+h,\ldots,\,\left(k-1\right)h,\,kh$.\footnote{Equivalently, the observation times within the rolling widow at the $k$th's observation are $k-T_{m}+1,\ldots,\,k$.}
The forecasts may be based on a parametric model whose time-$kh$ parameter estimates are then collected into the $q\times1$ random vector $\widehat{\beta}_{k}$. If no parametric assumption is made, then $\widehat{\beta}_{k}$ represents whatever semiparametric or nonparametric estimator used for generating the forecasts. The time-$kh$ forecast is denoted by $\widehat{f}_{k}\left(\widehat{\beta}_{k}\right)\triangleq f\left(Z_{kh},\,Z_{\left(k-1\right)h},\ldots,\,Z_{\left(k-m_{f}+1\right)h};\,\widehat{\beta}_{k}\right)$, where $f$ is some measurable function. The notation indicates that the $kh$-time forecast is generated from information contained in a sample of size $m_{f}$.\footnote{$m_{f}$ varies with the forecastis scheme; e.g., for the rolling scheme we have $m_{f}=T_{m}$ while for the recursive scheme we have $m_{f}=k$.}
Next, we introduce a loss function $L\left(\cdot\right)$ which serves for evaluating the performance of a given forecast model. More specifically, each out-of-sample loss $L_{\left(k+\tau\right)h}\left(\widehat{\beta}_{k}\right)\triangleq L\left(Y_{\left(k+\tau\right)h},\,\widehat{f}_{k}\left(\widehat{\beta}_{k}\right)\right)$ constitutes a statistical measure of accuracy of the $\tau$-step forecast made at time $kh$. However, given the objective of detecting potential instability of a certain forecasting method over time, we need additionally to introduce the in-sample losses $L_{jh}\left(\widehat{\beta}_{k}\right)\triangleq L\left(Y_{jh},\,\widehat{y}_{j}\left(\widehat{\beta}_{k}\right)\right),$ where $\widehat{y}_{j}\left(\widehat{\beta}_{k}\right)$ is an in-sample fitted value with $j$ varying over the specific in-sample window. That is, for each time-$kh$ forecast there corresponds a sequence (indexed by $j$) of in-sample fitted values $\widehat{y}_{j}\left(\widehat{\beta}_{k}\right)$.\footnote{We have $j=\tau+1,\ldots,\,T_{m}$ for the fixed scheme, $j=\tau+1,\ldots,\,k$ for the recursive scheme and $j=k-T_{m}+\tau+1,\ldots,\,k$ for the rolling scheme. } Then, the testing problem turns into the detection of any “systematic difference” between the sequence of out-of-sample and in-sample losses; the formal measure of such difference under our context is provided below.
The process $Z$ is a $\mathbb{R}^{q+1}$-valued semimartingale on $\left(\Omega,\,\mathscr{F},\,\left\{ \mathscr{F}_{t}\right\} _{t\geq0},\,\mathbb{P}\right)$ and we further assume that all processes considered in this paper are c�dl�g adapted and possess a $\mathbb{P}$-a.s. continuous path on $\left[0,\,N\right]$.\footnote{For accessible treatments of the probabilistic elements used in this section we refer to sahalia/jacod:14, jacod/shiryaev:03, jacod/protter:12, karatzas/shreve:96 and protter:05.} The continuity property represents a key assumption in our setting and implies that $Z$ is a continuous It\^o semimartignale. The integral form for $X_{t}$ is given by,
where $\left\{ W_{X,t}\right\} _{t\geq0}$ is a $q\times1$ Wiener process, $\mu_{X,s}\in\mathbb{R}^{q}$ and $\sigma_{X,s}\in\mathcal{M}_{q}^{\textrm{c�dl�g}}$ are the drift and spot covariance process, respectively, and $x_{0}$ is $\mathscr{F}_{0}$-measurable. We incorporate model misspeficication into our framework by allowing for a large non-zero drift which adds to the residual process:
where $\beta^{*}\in\mathbb{R}^{q}$, $\left\{ W_{e,t}\right\} _{t\geq0}$ is a standard Wiener process, $\sigma_{e,s}\in\mathbb{R}_{+}$ is its associated volatility, $\mu_{e,s}\in\mathbb{R}$ and $y_{0}$ is $\mathscr{F}_{0}$-measurable. In (ref), the last two terms on the right-hand side account for the residual part of $Y_{t}$ which is not explained by $X_{t-},$ where $X_{t-}=\lim_{s\uparrow t}X_{s}$. We assume $\vartheta\in[0,\,1/8)$ so that the factor $h^{-\vartheta}$ inflates the infinitesimal mean of the residual component thereby approximating a setting with arbitrary misspecification.
Part (i) rules out jump processes from our setting. We relax this restriction in our companion paper; see Section (ref). Part (ii) restricts those processes to be locally bounded. These should be viewed as regularity conditions rather than assumptions and are standard in the financial econometrics literature {[}see barndorff/shephard:04, li/xiu:16 and \citet*{li/todorov/tauchen:17}{]}; recently, they have been used by casini/perron_CR_Single_Break in the context of structural change models.
The continuous-time model in (ref)-(ref) is not observable. The econometrician only has access to $T$ realizations of $Y_{t}$ and $X_{t}$ with a sampling interval $h>0$ over the horizon $\left[0,\,N\right]$. For each $h>0,$ $Z_{kh}\in\mathbb{R}^{q+1}$ is a random vector step function that jumps only at time $0,\,h,\,2h,\ldots$, and so on. The discretized processes $Y_{kh}$ and $X_{kh}$ are assumed to be adapted to the increasing and right-continuous filtration $\left\{ \mathscr{F}_{t}\right\} _{t\geq0}$. The increments of a process $U$ are denoted by $\Delta_{h}U_{k}=U_{kh}-U_{\left(k-1\right)h}$. A seminal result known as Doob-Meyer Decomposition {[}cf. the original sources are doob:53 and meyer:67; see also Section III.3 in protter:05{]} allows us to decompose the semimartingale process $X_{t}$ into a predictable part and a local martingale part. Hence, it follows that we can write for $k=1,\ldots,T$, $\Delta_{h}X_{k}\triangleq\mu_{X,kh}\cdot h+\Delta_{h}M_{X,k}$ where the drift $\mu_{X,t}\in\mathbb{R}^{q}$ is $\mathscr{F}_{t-h}$ measurable, and $M_{X,kh}\in\mathbb{R}^{q}$ is a continuous local martingale with finite conditional covariance matrix $\mathbb{P}$-a.s. $\mathbb{E}\left(\Delta_{h}M_{X,k}\Delta_{h}M'_{X,k}|\,\mathscr{F}_{\left(k-1\right)h}\right)=\Sigma_{X,\left(k-1\right)h}\cdot h$. Turning to equation (ref), the error process $\left\{ \Delta_{h}e_{k}^{*},\,\mathscr{F}_{t}\right\} $, with $\Delta_{h}e_{k}^{*}\triangleq\sigma_{e,\left(k-1\right)h}\Delta_{h}W_{e,k}$, is then a continuous local martingale difference sequence taking its values in $\mathbb{R}$ with finite conditional variance $\mathbb{E}\left[\left(\Delta_{h}e_{k}^{*}\right)^{2}|\,\mathscr{F}_{\left(k-1\right)h}\right]=\sigma_{e,\left(k-1\right)h}^{2}\cdot h,\,\mathbb{P}$-a.s. Therefore, we express the discretized analogue of (ref) as
Finally, we have an additional assumption on the path of the volatility process $\left\{ \sigma_{e,t}^{2}\right\} _{t\geq0}$. This turns out be important because it partly affects the local behavior of the forecast losses.
The assumption essentially states that $\phi_{\sigma,\eta,N}$ is locally bounded and $\left\{ \sigma_{e,t}\right\} _{t\geq0}$ is Lipschitz continuous. Lipschitz volatility is a more than reasonable specification for the macroeconomic and financial data to which our analysis is primarily directed. Indeed, the basic case of constant variance $\sigma^{2}$ is easily accommodated by the assumption. Time-varying volatility is also covered provided $\sigma_{e,t}^{2}$ is sufficiently smooth. However, the assumption rules out some standard stochastic volatility models often used in finance. We relax that assumption in Section (ref), so that we can extend our results to, for example, stochastic volatility models driven by a Wiener process.
As time evolves, a forecast model can suffer instability for multiple reasons. However, incorporating model misspecification into our framework necessarily implies that the exact form of the instability is unknown and thus one has to leave it unspecified. This differs from the classical setting for estimation of structural change models {[}cf. bai/perron:98 and casini/perron_CR_Single_Break{]} where (i) the break date is well-defined as it is part of the definition of the econometric problem, and (ii) the form of the instability is explicitly specified through a discrete shift in a regression parameter. In contrast, under our context we remain agnostic regarding both (i) and (ii). There may be multiple dates at which the forecast model suffers instability and they might be interrelated in a complicated way. Forecast instability may manifest itself in several forms, including gradual, smooth or recurrent changes in the predictive relationship between $Y_{\left(k+\tau\right)h}$ and $X_{kh}$; certainly, there could also be discrete shifts in $\beta^{*}$\textemdash arguably the most common case in practice\textemdash but this is only a possibility in our setting and not an assumption as in structural change models. A forecast failure then reflects the forecaster's failure to recognize the shift in the predictive power of $X_{kh}$ on $Y_{\left(k+\tau\right)h}$. On the other hand, even if one can rule out shits in $\beta^{*}$, a forecast instability may be induced by an increase/decrease in the uncertainty in the data which might result, for example, from changes in the unconditional variance of the target variable. In this case, the predictive ability of $X_{kh}$ on $Y_{\left(k+\tau\right)h}$, as described for instance by a parameter $\beta,$ remains stable while due to an increase in the unconditional variance of $Y_{\left(k+\tau\right)h}$ it might become weak and in turn the forecasting power might breakdown. Tests for forecast failure such as those proposed in this paper and the ones proposed in giacomini/rossi:09 are designed to have power against both of the above hypotheses.\footnote{Recently, perron/yamamoto:18 proposed to apply modified versions of classical structural break tests to the forecast failure setting. However, their testing framework and hence their null hypotheses are different from ours because they do not fix a model-method pair but only fix the forecast model under the null.}
Define at time $\left(k+\tau\right)h$ a surprise loss given by the deviation between the time-$\left(k+\tau\right)h$ out-of-sample loss and the average in-sample loss: $SL_{\left(k+\tau\right)h}\left(\widehat{\beta}_{k}\right)\triangleq L_{\left(k+\tau\right)h}\left(\widehat{\beta}_{k}\right)-\overline{L}_{kh}\left(\widehat{\beta}_{k}\right),$ for $k=T_{m},\ldots,\,T-\tau$, where $\overline{L}_{kh}\left(\widehat{\beta}_{k}\right)$ is the average in-sample loss computed according to the specific forecasting scheme. One can then define the average of the out-of-sample surprise losses
where $N_{0}\triangleq N-N_{\mathrm{in}}-h$ denotes the time span of the out-of-sample window.\footnote{By definition $N_{0}$ is fixed and should not be confused with $T_{n},$ which indicates the number of observations in the out-of-sample window. Indeed, $N_{0}=T_{n}h.$ } In the classical discrete-time setting, under the hypotheses of no forecast instability one would naturally test whether $\overline{SL}{}_{N_{0}}\left(\beta^{*}\right)$ has zero mean, where $\beta^{*}$ is the pseudo-true value of $\beta$. If the forecasting perfomance remains stable throughout the whole sample then there should be no systematic surprise losses in the out-of-sample window and thus $\mathbb{E}\left[N_{0}^{-1}\sum_{k=T_{m}}^{T-\tau}SL_{\left(k+\tau\right)h}\left(\beta^{*}\right)\right]=0.$ This reasoning motivated the forecast breakdown test of giacomini/rossi:09. Therefore, under the classical asymptotic setting one exploits time series properties of the process $SL_{\left(k+\tau\right)h}\left(\beta^{*}\right)$ such as ergodicity and mixing together with the representation of the hypotheses by a global moment restriction.\footnote{Global refers to the property that the zero-mean restriction involves the entire sequence of forecast losses.} By letting the span $N\rightarrow\infty$, this method underlies the classical approach to statistical inference but does not directly extend to an infill asymptotic setting. Under continuous-time asymptotics, identification of parameters is achieved by properties of the paths of the involved processes and not by moment conditions. This constitutes the key difference and requires one to recast the above hypotheses into an infill setting thereby making use of assumptions on an underlying continuous-time data-generating mechanism which is assumed to govern the observed data.
We begin with observing that the sequence of losses $\left\{ L_{kh}\left(\cdot\right)\right\} $ can be viewed as realizations from an underlying continuous-time process $\left\{ \widetilde{L}_{t}\right\} _{t\geq0}$, with $\widetilde{L}_{t}\triangleq\int_{0}^{t}L_{s}\left(Y_{s},\,X_{s-};\,\beta^{*}\right)ds$. That is, $\widetilde{L}_{t}$ consists of temporally integrated forecast losses where $L_{t}$ is the loss at time $t$ and is defined by some transformation of the target variable $Y_{t}$ and of the predictor $X_{t-}$.\footnote{The definition of $\widetilde{L}_{t}$ uses that so long as the forecast step $\tau$ is small and finite one can approximate $X_{s-\tau h}$ by $X_{s-}$ for sufficiently small $h>0.$} In order to provide a general theory, we focus on families of loss functions that depend only on the forecast error.\footnote{The most popular loss functions used in economic forecasting are within this category {[}see elliott/timmermann:16 for a recent incisive account of the literature{]}. Extension to ad hoc loss functions requires specific treatment that might vary from case to case.} We denote this class by $\boldsymbol{L}_{e}$ and we say that the loss function $L_{\cdot}\left(\cdot,\,\cdot;\,\cdot\right)\in\boldsymbol{L}_{e}$ if $L_{t}\left(Y_{t},\,X_{t-};\,\beta\right)=L_{t}\left(e_{t};\,\beta\right)$ for all $t\in\left[0,\,N\right]$, where $e_{t}=Y_{t}-\widehat{f}_{t}\left(\beta\right)$. The class $\boldsymbol{L}_{e}$ comprises the vast majority of loss functions employed in empirical work, including among others the popular Quadratic loss, Absolute error loss and Linex loss. The following examples illustrate how these loss functions are constructed under our setting. For the rest of this section, assume for simplicity $y_{0}=0,\,\mu_{e,\cdot}=0$ and that $X_{t}$ is one-dimensional in (ref).
Below we make very mild pathwise assumptions on the process $Z$ which imply restrictions on $\left\{ \widetilde{L}_{t}\right\} _{t\geq0}$. We derive asymptotic results under Lipschitz continuity (in $t$) of the coefficients of the system of stochastic differential equations driving the data $\left\{ Z_{t}\right\} _{t\geq0}$. We apply the techniques of stochastic calculus to formulate our testing problem. To avoid clutter, we introduce the notation $g\left(Y_{t},\,X_{t-};\,\beta^{*}\right)=L_{t}\left(Y_{t},\,X_{t-};\,\beta^{*}\right)$ and its shorthand $g\left(e_{t};\,\beta^{*}\right)=L_{t}\left(e_{t};\,\beta^{*}\right)$.\footnote{The notation implicitly assumes that the same loss function is used for estimation and prediction which in turn implies that the subscript $t$ in $L_{t}\left(e_{t};\,\beta^{*}\right)$ can be omitted since it can be understood from that of the argument $e_{t}$.} By It\^o Lemma, {[}cf. Section II.7 in protter:05{]}, under smoothness of $g\left(e_{t};\,\beta^{*}\right)$,
Let $\mathbb{E}_{\sigma}$ denote the expectation conditional on the path $\left\{ \sigma_{e,t}\right\} _{t\geq0}$. The instantaneous mean of $dL\left(e_{t};\,\beta^{*}\right)$ is $\mathbb{E}_{\sigma}\left[dL\left(e_{t};\,\beta^{*}\right)/dt\right]=2^{-1}\sigma_{e,t}^{2}\mathbb{E}_{\sigma}\left(\partial^{2}g\left(e_{t};\,\beta^{*}\right)/\partial e^{2}\right)$. Note that the latter is a symbolic abbreviation for
Since the coefficients of the original system of stochastic equations are Lipschitz continuous in $t$, one can verify that $\mathbb{E}_{\sigma}\left[dL\left(e_{t};\,\beta^{*}\right)/dt\right]$ is also Lipschitz upon regularity conditions on $g\left(\cdot,\,\beta^{*}\right)$ and time-$t$ information.
We denote by $\boldsymbol{Lip}\left(\left[0,\,N\right]\right)$ the class of Lipschitz continuous functions on $\left[0,\,N\right]$. Let $\left\{ c_{t}\right\} _{t\geq0}$ denote a continuous-time stochastic process that is $\mathbb{P}$-a.s. locally bounded and adapted.
We are in a position to formulate the testing problem in terms of the pathwise property of $L_{t}\left(e_{t};\,\beta^{*}\right).$ This implies that the hypotheses are specified in terms of random events which differs from classical hypotheses testing but it is typical under continuous-time asymptotics; see sahalia/jacod:12 (for many references), li/todorov/tauchen:16 and reiss/todorov/tauchen:15. We consider the following hypotheses: for any $L\left(\cdot;\,\cdot\right)\in\boldsymbol{L}_{e}$,\footnote{Precise assumptions will be stated below.}
which means that we wish to discriminate between the following two events that divide $\Omega:$
The dependence of the hypotheses on $\omega$ is appropriate because each event $\omega$ generates a certain path of $L\left(e_{t}\left(\cdot\right);\,\beta^{*}\right)$ on $\left[0,\,N\right]$, where $de_{t}\left(\omega\right)=\sigma_{e,t}\left(\omega\right)dW_{e,t}\left(\omega\right)$. The hypotheses requires a Lipschitz condition to hold on $\left[N_{\mathrm{in}}+h,\,N\right]$, where $N_{\mathrm{in}}$ is the usual artificial separation date after which the first forecast is made. $N_{\mathrm{in}}$ is taken as given here because the testing problem applies to a specific method-model pair and $N_{\mathrm{in}}$ is part of the chosen forecasting method. From a practical standpoint, it would be helpful if this separation date is such that the forecast model is stable on $\left[0,\,N_{\mathrm{in}}\right]$ {[}see casini/perron_Oxford_Survey for more details{]}. The latter property is, however, unknown a priori by the practitioner. We cover this case in Section (ref).
We have reduced the forecast instability problem to examination of the local properties of the path of $L_{t}$. However, we still have to face the question on how to use the data to test $H_{0}$ in practice. Even if we could observe $\widetilde{L}_{t}$, it would not be clear how to formulate a testing problem on the stability of $L_{t}$ by using path properties of $\widetilde{L}_{t}$. The reason is that $\widetilde{L}_{t}$ is always absolutely continuous by definition, and thus it would provide little information on the large deviations of the forecast error $e_{t}$. In order to study the local behavior of $L_{t}$ one needs to consider the small increments of $L_{t}$ close to time $t$. Leaving the definition of $\widetilde{L}_{t}$ aside for a moment, observe that $\mathbb{P}$-a.s. continuity of $Z_{t}$ is equivalent to having the relationship between $Y_{t}$ and $X_{t}$ holding for any infinitesimal interval of time. For the basic parametric linear model: $dY_{t}=\beta^{*}dX_{t}+de_{t}$. Then, the forecast loss is $L\left(de_{t}\right)$, which is difficult to interpret in rigorous probabilistic terms. However, we can consider its discrete-time analogue. We normalize the forecast error by the factor $\psi_{h}=h^{1/2}$ and redefine $L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)\triangleq L_{kh}\left(\psi_{h}^{-1}\Delta_{h}e_{k};\,\beta^{*}\right)$.\footnote{Alternatively, $L_{\psi,kh}\left(\Delta_{h}Y_{k},\,\Delta_{h}X_{k};\,\beta^{*}\right)=L_{kh}\left(\psi_{h}^{-1}\left(\Delta_{h}Y_{k}-\beta^{*}\Delta_{h}X_{k}\right)\right)$.} Then, for all $k$, the mean of $L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)$\textemdash conditional on $\sigma_{e,kh}$\textemdash depends on the parameters of the model and its local behavior can be used as a proxy for the local behavior of the infinitesimal mean of $dL_{t}\left(e_{t};\,\beta^{*}\right)$. If the corresponding structural parameters of the continuous-time data-generating process satisfy a Lipschitz continuity in $t$, then\textemdash knowing $\sigma_{e,kh}$\textemdash also $\mathbb{E}_{\sigma}\left[L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)\right]$ should be Lipschitz in the continuous-time limit. Under the hypotheses $H_{0}$ there should be no break in $\mathbb{E}_{\sigma}\left[L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)\right]$ and an appropriately defined right local average of $L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)$ should not differ too much from its left local average. That is, one can test for forecast instability by using a two-sample t-test over asymptotically vanishing time blocks.
Both examples demonstrate that pathwise assumptions on the data-generating process implies restrictions on the properties of the sequence of loss functions. For the QL example, if there is a structural break at the observation $k=T_{b},$ then this would result in the mean of $L_{\psi,kh}\left(\Delta_{h}e_{k};\,\beta^{*}\right)$ shifting to a new level after time $T_{b}h$. Given that the same reasoning extends to the sequence of surprise losses, one may consider to construct a test statistic on the basis of the local behavior of the surprise losses over time. If there is no instability in the predictive ability of a certain model, then the sequence of out-of-sample surprise losses should display a certain stability. Under the framework of giacomini/rossi:09, this stability is interpreted in a retrospective and global sense as a zero-mean restriction on the sequence over the entire out-of-sample. In contrast, under our continuous-time setting, this stability manifests itself as a continuity property of the path of the continuous-time counterpart of the sequence.
By inspection of the null hypotheses in (ref), it is evident that a considerable number of forms of instabilities are allowed. These may result from discrete shifts in a model's structural parameter and/or in structural properties of the processes considered such as conditional and unconditional moments and so on. This first set of non-stationarities relates to the popular case of structural changes which are designed to be detected with high probability by the structural break tests of, among others, andrews:93 and andrews/ploberger:94, bai/perron:98 and elliott/mueller:06 in univariate settings and of qu/perron:07 in multivariate settings. However, a forecast instability may be generated by many other forms of non-stationarities against which such classical tests for structural breaks are not designed for and consequently they might have little power against. For example, consider the case of smooth changes in model parameters, or in the unconditional variance of $Y_{kh}$. Even more serious would be the presence of recurrent smooth changes in the marginal distribution of the predictor since in this case the above-mentioned tests are likely to falsely reject $H_{0}$ too often {[}cf. hansen:00{]}. Thus, the null hypotheses of no forecast instability calls for a new statistical hypotheses testing framework. Ideally, in this context one needs a test statistic that retains power against any discontinuity, jump and recurrent switch at any point in the out-of-sample and for any magnitude of the shift. We propose a test statistic which aims asymptotically at distinguishing any discontinuity from a regular Lipschitz continuous motion. We introduce a sequence of two-sample t-tests over asymptotically vanishing adjacent time blocks. This should lead to significant gains in power whenever on fixed time intervals the out-of-sample losses exhibit instabilities of any form such as breaks, jumps and relatively large deviations. Such gains are likely to occur especially when instabilities take place within a small portion of the sample relative to the whole time span\textemdash a common case in practice that has characterized many episodes of forecast failure in economics.
Interestingly, for the Quadratic loss function we can exploit properties of the local quadratic variation and propose a self-normalized test statistic. Thus, we separate the discussion on the Quadratic loss from that on general loss functions. Let $SL_{\psi,\left(k+\tau\right)h}\left(\widehat{\beta}_{k}\right)\triangleq L_{\psi,\left(k+\tau\right)h}\left(\widehat{\beta}_{k}\right)-\overline{L}_{\psi,kh}\left(\widehat{\beta}_{k}\right)$, $k=T_{m},\ldots,\,T-\tau$. Next, we partition the out-of-sample into $m_{T}\triangleq\left\lfloor T_{n}/n_{T}\right\rfloor $ blocks each containing $n_{T}$ observations. Let $B_{h,b}\triangleq n_{T}^{-1}\sum_{j=1}^{n_{T}}SL_{\psi,T_{m}+\tau+bn_{T}+j-1}\left(\widehat{\beta}_{T_{m}+bn_{T}+j-1}\right)$ and $\overline{B}_{h,b}\triangleq n_{T}^{-1}\sum_{j=1}^{n_{T}}L_{\psi,\left(T_{m}+\tau+bn_{T}+j-1\right)h}\left(\widehat{\beta}_{T_{m}+bn_{T}+j-1}\right)$ for $b=0,\ldots,\,,\,\left\lfloor T_{n}/n_{T}\right\rfloor -1$.
We propose the following statistic
The quantity $B_{h,b}$ is a local average of the surprise losses within the block $b$. We have partitioned the out-of-sample window into $m_{T}$ blocks of asymptotically vanishing length $\left[bn_{T}h,\,\left(b+1\right)n_{T}h\right]$. We consider an asymptotic experiment in which the number of blocks $m_{T}$ increases at a controlled rate to infinity while the per-block sample size $n_{T}$ grows without bound at a slower rate than the out-of-sample size $T_{n}$. The appeal of the $\mathrm{B}_{\mathrm{max},h}\left(T_{n},\,\tau\right)$ statistic is that a large deviation $B_{h,b+1}-B_{h,b}$ suggests the existence of either a discontinuity or non-smooth shift in the surprise losses close to time $bn_{T}h$ and thus it provides evidence against $H_{0}$. We comment on the nature of the normalization $\overline{B}_{h,b+1}$ in the denominator of $\mathrm{B}_{\mathrm{max},h}$ below, after we introduce a version of $\mathrm{B}_{\mathrm{max},h}$ statistic which uses all admissible overlapping blocks of length $n_{T}h$:
where $\overline{B}_{h,i}=n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}L_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)$. Since under the alternative hypotheses the exact location of the change-point\textendash -or possibly the locations of the multiple change-points\textemdash within the block might actually affect the power of the $\mathrm{B}_{\mathrm{max},h}$-based test in small samples, we indeed find in our simulation study that the test statistic $\mathrm{MB}{}_{\mathrm{max},h}$ which uses overlapping blocks is more powerful especially when the instability arises in forms other than the simple one-time structural change. Thus, the power of the $\mathrm{B}_{\mathrm{max},h}$ test is slightly sensible to the actual location of the change-point within the block, with higher power achieved when the change-point is close to either the beginning or the end of the block. In contrast, the statistical power of $\mathrm{MB}_{\mathrm{max},h}$ is uniform over the location of the change-point in the sample. The latter property is not shared by the exiting test of giacomini/rossi:09 given that its power tends to be substantially lower if the instability is not located at about middle sample.
An important characteristic of both $\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{MB}_{\mathrm{max},h}$ is that they are self-normalized; no asymptotic variance appears in their definition. The reason for why $\overline{B}_{h,b+1}$ appears in the denominator of, for example, $\mathrm{B}_{\mathrm{max},h}$ is that even though $B_{h,b+1}$ constitutes a more logical self-normalizing term, it might be close to zero in some cases. This would occur under Quadratic loss if, for example, $\sigma_{e,t}=\sigma_{e}$ for all $t\geq0.$ This is not true for the factor $\overline{B}_{h,b+1}$.
In addition, observe that allowing for misspecification naturally leads one to deal carefully with artificial serial dependence in the forecast losses in small samples. Thus, we consider a version of the statistics $\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{MB}_{\mathrm{max},h}$ that are normalized by their asymptotic variance: $\mathrm{Q}_{\mathrm{max},h}\left(T_{n},\,\tau\right)\triangleq\nu_{L}^{-1}\max_{b=0,\ldots,\,\left\lfloor T_{n}/n_{T}\right\rfloor -2}\left|B_{h,b+1}-B_{h,b}\right|,$ and similarly,
The quantity $\nu_{L}$ standardizes the test statistic so that under the null hypotheses we obtain a distribution-free limit. This can be useful because given the fully non-stationary setting together with the possible consequences of misspecification in finite-samples, standardization by the square-root of the asymptotic variance $\nu_{L}^{2}$ might lead to a more precise empirical size in small samples. We relegate theoretical details on $\nu_{L}$ as well as on its estimation to Section (ref) where we also present a discussion about its relation with the choice of the number of blocks.
For general loss $L\in\boldsymbol{L}_{e}$, we propose the following statistic,
where $B_{h,b},\,B_{h,b+1}$ are defined as in the quadratic case and
with $\overline{L}_{\psi,b}\left(\widehat{\beta}\right)\triangleq n_{T}^{-1}\sum_{j=1}^{n_{T}}L_{\psi,\left(T_{m}+\tau+bn_{T}+j-1\right)h}\left(\widehat{\beta}_{T_{m}+bn_{T}+j-1}\right)$. The interpretation of $\mathrm{G}_{\mathrm{max},h}$ is essentially the same as of $\mathrm{B}_{\mathrm{max},h}$, the only difference arising from the denominator $\sqrt{D_{h,b+1}}$ that estimates the within-block variance. A version that uses all overlapping blocks is
where $D_{h,i}\triangleq n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}\left(L_{\psi,\left(T_{m}+\tau+j-1\right)h}\left(\widehat{\beta}_{T_{m}+j-1}\right)-\overline{L}_{\psi,i}\left(\widehat{\beta}\right)\right)^{2}$, with
As argued above, it is useful to consider versions of the statistic $\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{MB}_{\mathrm{max},h}$ that are normalized by their asymptotic variance:
and similarly,
We begin with a set of assumptions. Assumption (ref) below is a finite-moment condition on the sequence of rescaled forecast losses and and on its first-order derivative. It has a similar scope to A4 in giacomini/rossi:09. Assumption (ref) is similar to A5 in \Citet{giacomini/rossi:09} and it imposes the first-order derivative of the forecast losses to be constant over time. It trivially holds when one employs the same loss function for estimation and evaluation. Assumption (ref) demands the existence of a consistent estimator for $\beta^{*}$ at the parametric rate $\sqrt{T}$ and it encompasses many estimation procedures. In model (ref), the popular least-squares method will satisfy the condition {[}cf. barndorff/shephard:04 and \citet*{li/todorov/tauchen:17}{]}.
Our asymptotic results are valid under the following conditions on the auxiliary sequence $n_{T}$.
Condition (ref) imposes a lower bound and an upper bound on the growth condition of the sequence $\left\{ n_{T}\right\} $. The first part of (ref) requires $n_{T}$ to grow to infinity at any faster rate than $T^{\epsilon}$ with $\epsilon>0$, which we interpret as saying that the number of observations $n_{T}$ in each block cannot be too small. The second part of (ref) provides an upper bound on the growth of $n_{T}$ and relates to Assumption (ref) concerning the smoothness of $\left\{ \sigma_{e,t}\right\} _{t\geq0}$ thereby ensuring that, for example, the random oscillations of $\mathrm{B}_{\mathrm{max},h}\left(T_{n},\,\tau\right)$ can be controlled. As we shall explain in the simulation study of Section (ref), we recommend to set $n_{T}\propto T_{n}^{2/3-\epsilon}$ for small $\epsilon>0$.
Theorem (ref) shows that the asymptotic null distribution of our test statistics follows an extreme vale distribution whose critical values can be computed directly. In nonparametric change-point analysis, wu/zhao:07 and \citet*{bibinger/jirak/vetter:16} have derived an extreme value null distribution for tests statistics which share a similar form to ours. As it is stated, the tests statistics are not yet feasible because the asymptotic variances $\nu_{L}^{2}$ is unknown. However, we can find statistical consistent estimators which can be used in place of $\nu_{L}^{2}$ to make the test feasible. We relegate the treatment of its consistent estimation to Section (ref).
The purpose of this section is to show how to construct an asymptotically valid estimator of the variance $\nu_{L}^{2}$ that enters the definition of our test statistics. This is an important aspect that together with the selection of the block length might affect statistical inferences based on the proposed tests in finite-samples. Allowing for misspecification is customary in the forecasting literature, and as a consequence this may result in forecast losses that artificially exhibit heteroskedasticity and serial dependence in small samples.
We begin with the case of stationary forecast losses, including constant $\nu_{L}$ as a special case.
Recall that our test statistics are related to a maximum over blocks of data. Thus, for i.i.d. forecast losses one can use the following estimator for $\nu_{L}$ in $\mathrm{Q}_{\mathrm{max},h}\left(T_{n},\,\tau\right)$:
where $\overline{SL}{}_{\psi,b}\triangleq n_{T}^{-1}\sum_{j=1}^{n_{T}}SL_{\psi,T_{m}+\tau+bn_{T}+j-1}\left(\widehat{\beta}_{T_{m}+bn_{T}+j-1}\right)$. The estimator $\widehat{\nu}_{\mathrm{Q}1,b}^{2}$ normalizes the difference in the out-of-sample forecast losses between the $b+1$ and $b$ blocks. The statistic $\mathrm{Q}_{\mathrm{max},h}\left(T_{n},\,\tau\right)$ then results in $\mathrm{Q}_{\mathrm{max},h}\left(T_{n},\,\tau\right)=\max_{b=0,\ldots,\,\left\lfloor T_{n}/n_{T}\right\rfloor -2}\left|\left(B_{h,b+1}-B_{h,b}\right)/\widehat{\nu}_{\mathrm{Q}1,b+1}\right|$. For the overlapping blocks case, the estimator is $\widehat{\nu}_{\mathrm{MQ}1,i}^{2}\triangleq2n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}\left(SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)-\overline{SL}{}_{\psi,i}\right)^{2},$ where $\overline{SL}{}_{\psi,i}\triangleq n_{T}^{-1}\sum_{j=i+1}^{i+n_{T}}SL_{\psi,T_{m}+\tau+j-1}\left(\widehat{\beta}_{T_{m}+j-1}\right)$ so that we can write
Both $\widehat{\nu}_{\mathrm{Q}1,b+1}^{2}$ and $\widehat{\nu}_{\mathrm{MQ}1,i}^{2}$ apply a natural block-wise normalization in order to guarantee a distribution-free limit under $H_{0}$. However, it is useful to consider estimators that use all of the observations in the out-of-sample period. Thus, one exploits covariance stationarity of the sequence of forecast losses. Let $\Phi_{0.75}=0.647...$ denote the third quartile of the standard normal distribution and define
Note that $\widehat{\nu}_{2,h},\,\widehat{\nu}_{3,h}$ and $\widehat{\nu}_{4,h}$ can be used to implement both $\mathrm{Q}{}_{\mathrm{max},h}\left(T_{n},\,\tau\right)$ and $\mathrm{MQ}{}_{\mathrm{max},h}\left(T_{n},\,\tau\right)$.\footnote{They can be also applied to the test statistics $\mathrm{Q}{}_{\mathrm{max},h}^{\mathrm{G}}\left(T_{n},\,\tau\right)$ and $\mathrm{MQ}{}_{\mathrm{max},h}^{\mathrm{G}}\left(T_{n},\,\tau\right)$ with $G_{h,b}$ in place of $B_{h,b}$.} $\widehat{\nu}_{3,h}$ is related to carlstein:86's (1986) subseries variance estimate in the context of strong mixing processes and it was also used by wu/zhao:07. Each of the estimators $\widehat{\nu}_{2,h},\,\widehat{\nu}_{3,h}$ and $\widehat{\nu}_{4,h}$ allows for dependence but requires stationarity. The simulation study in wu/zhao:07 suggests that $\widehat{\nu}_{4,h}$ is more robust whereas $\widehat{\nu}_{2,h}$ and $\widehat{\nu}_{3,h}$ are less precise when there are large instabilities or jumps. For two sequences $\left\{ a_{k}\right\} $ and $\left\{ b_{k}\right\} $, we write $a_{k}\asymp b_{k}$ if for some $c\geq1$, $b_{k}/c\leq a_{k}\leq cb_{k}$ for all $T.$ The following theorem is similar to Theorem 3 in wu/zhao:07 and in particular, part (ii) states that if $n_{T}\asymp T_{n}^{1/3}$ then $\widehat{\nu}_{3,h}^{2}$ achieves the optimal MSE $O\left(n_{T}^{-2/3}\right)$.
Under covariance-stationarity, given Theorem (ref), the results of Corollary (ref)-(ref) are applicable after replacing $\nu_{L}$ by an appropriate consistent estimator.
We now consider estimation of the asymptotic variance in the case the forecast losses are heterogeneous. The estimator $\widehat{\nu}$ that we introduce below depends on the specific loss function and thus it can be used for replacing $\nu_{L}$ in Corollary (ref)-(ref). Non-stationarity implies that $\sigma_{e,t}^{2}$ is time-varying and thus the results of Theorem (ref) are not applicable due to the presence of many extra parameters that account for the time-varying structure. To deal with this issue we propose a novel block-wise self-normalization technique which simultaneously addresses two issues. First, the block-wise self-normalization ensures that the difference in forecast losses between two adjacent blocks are asymptotically independent across non-adjacent blocks and that within each block the losses are standardized so that time-varying variances cancel out. Second, by computing an average\textemdash over all blocks\textemdash of the self-normalized difference in forecast losses we account for possible serial dependence. We derive asymptotic results within a general framework based on the strong invariance principle for stationary processes developed in wu:07 and extended to modulated stationary processes by zhao/li:13.
For each block $b=0,\ldots,\,m_{T}-2$, let
where $\overline{L}{}_{\psi,b}\left(\widehat{\beta}\right)=n_{T}^{-1}\sum_{j=1}^{n_{T}}L_{\psi,T_{m}+\tau+\left(b+1\right)n_{T}+j-1}\left(\widehat{\beta}\right)$ and define the statistic
Finally, an average\textemdash over all blocks $m_{T}$\textemdash of the per-block self-normalized statistics $\zeta_{h,b}$'s is used to define an estimator of the asymptotic variance: $\widehat{\nu}_{L}^{2}\triangleq2^{-1}\left(m_{T}-1\right)^{-1}\sum_{b=0}^{m_{T}-1}\zeta_{h,b}^{2}$.
Let $\sigma_{L,kh}^{2}\triangleq\mathrm{Var}\left(L_{\psi,kh}\left(\beta^{*}\right)\right)$. We also need to introduce the following quantities,
The theorem simply states that $\widehat{\nu}_{L}$ is consistent for $\nu_{L}$ and therefore the asymptotic results of Section (ref) continue to hold when we replace $\nu_{L}$ by $\widehat{\nu}_{L}$.
In this section we relax the Lipschitz condition on $\sigma_{e,t}$ and extend the results for the quadratic loss case from Theorem (ref) to stochastic volatility models driven by a Wiener process. Consequently, this relaxation enables one to utilize the tests proposed in this paper in setting involving high-frequency financial variables. More specifically, we assume that $\sigma_{e,t}$ is an It\^o continuous semimartingale that is almost surely bounded and strictly positive adapted process. We replace Assumption (ref) by the following.
The assumption implies that $\sigma_{e,t}$ belongs to a rather large class of volatility processes usually considered in financial econometrics. The parameter $\kappa$ plays a key role in the testing framework of this section and we refer to it as the regularity exponent. When $\kappa=1$ we recover the case of Lipschitz volatility considered in the previous sections while the standard stochastic volatility model without jumps correspond to $\kappa=1/2-\epsilon$ for a sufficiently small $\epsilon>0.$ Next, we have a slightly different version of Condition (ref).
For It\^o continuous semimartingale volatility $\sigma_{e,t}$ the condition suggests $n_{T}\propto T^{1/2-\epsilon}$ for small $\epsilon>0.$ Let $\varGamma_{t}\triangleq\mathbb{E}_{\sigma}\left[dL_{t}\left(e_{t};\,\beta^{*}\right)/dt\right]$.\footnote{For example, for the quadratic loss with $\mu_{e,t}=0$ the notation reduces to $\varGamma_{t}=\sigma_{e,t}^{2}.$ } The more general framework considered here requires us to consider the following null hypotheses: under quadratic loss,
where $\boldsymbol{C}\left(\kappa,\,K_{h}\right)$ is a class of continuous functions on $\left[N_{\mathrm{in}}+h,\,N\right]$,
where $\kappa>0$ and $K_{h}$ is given in Assumption (ref). Thus, we wish to discriminate between $H_{0}$ and
where $\boldsymbol{J}_{\lambda}\left(\kappa,\,K_{h},\,d_{h}\right)\triangleq\left\{ \left\{ \varGamma_{t}\right\} _{t\in\left[N_{\mathrm{in}}+h,\,N\right]}:\,\left\{ \varGamma_{t}-\Delta\varGamma_{t}\right\} _{t\in\left[N_{\mathrm{in}}+h,\,N\right]}\in\boldsymbol{C}\left(\kappa,\,K_{h}\right);\,\left|\Delta\varGamma_{\lambda}\right|\geq d_{h}\right\} $, $\Delta\varGamma_{\lambda}=\varGamma_{\lambda}-\lim_{s\uparrow\lambda}\varGamma_{s}$ and $\left\{ d_{h}\right\} $ is a decreasing sequence. The following theorem extends Theorem (ref) to the current setting.
In this section we consider the behavior of $\mathrm{MQ}_{\mathrm{max},h}$ under a sequence of local alternatives.
Part (iii) is a consequence of part (i) as it can be easily verified. Let
where
Theorem (ref) can be used to show that our test possesses nontrivial power against alternatives for which the parameter $\beta_{t}$ is time-varying and non-smooth.
A number of extensions is treated in our companion paper casini_FFb. As explained above, it would be useful to ensure that there are no instabilities in the in-sample period $\left[0,\,N_{\mathrm{in}}\right].$ We propose a procedure that involves a pre-test about instability on $\left[0,\,N_{\mathrm{in}}\right]$ for a given $N_{\mathrm{in}}$ chosen by the forecaster. Instabilities in the in-sample $\left[0,\,N_{\mathrm{in}}\right]$ are much easier to be detected relative to instabilities in the out-of sample because they do not face the so-called “contamination effect”. The latter arises, for example under the recursive and rolling scheme, when the instability originally occurring in the out-of-sample eventually enters the moving in-sample window {[}cf. casini/perron_Oxford_Survey and perron/yamamoto:18{]}. The consequence is that existing tests face substantial power losses. This property is not shared by our test statistics because of their local nature. Our procedure works very well and we show through simulations that instabilities occurring in the in-sample only or occurring both in the in-sample and in the out-of-sample simultaneously, lead easily to rejection of the null hypotheses relative to instabilities occurring in the out-of-sample only\textemdash as we consider here.
A second issue is that, in this paper, we have considered processes that have a continuous sample path under the null hypotheses. Thus, it is of interest to extend the results to a setting that involves jump processes which are important in high-frequency financial data. This can be achieved by using techniques that are able to separate the continuous part from the discontinuous part of a It� semimartingale {[}see e.g. li/todorov/tauchen:17 and li/xiu:16{]}.
Another important issue is the estimation of the time at which forecast instability occurs. Once the null hypotheses has been rejected, a forecaster may take into consideration the possibility of revising the forecasting method and/or model. Hence, it becomes crucial to learn some information about the timing of the instability. For example, consider the case of a one-time structural change in a parameter of the data-generating process at time $T_{b}^{0}=\left\lfloor T\lambda_{b}^{0}\right\rfloor $, where $T_{b}^{0}$ is the break point and $\lambda_{b}^{0}\in\left(0,\,1\right)$ is the fractional break date. Once $H_{0}$ is rejected, a forecaster would benefit from knowing that the forecast method originally employed is found to statistically either under or over-perform over part of the sample after $T_{b}^{0}$ relative to the part prior $T_{b}^{0}$. Then, a forecaster would entertain the possibility of modifying the forecast model in order to generate future forecasts for $Y_{t}$. Not only the forecast model might be revised but most importantly, knowledge of beginning of the instability at $T_{b}^{0}$ can be further exploited to design the forecasting method for the future forecasts. It would be inappropriate, for instance, to use a rolling scheme where the rolling window used to construct the forecast include observations prior to $T_{b}^{0}$ since those observations provide little informational content for the purpose of predicting $Y_{t}$ after the change-point $T_{b}^{0}$. On the other hand, this line of reasoning is justified by this particular example and indeed in practice many issues arise when dealing with the timing of the insatiability in our context. For example, the exact form of the insatiability may be unknown. Under the latter scenarios, there is no clear-cut break date $T_{b}^{0}$ that can be defined. Thus, it is less obvious how a forecaster should proceed in those cases. Nonetheless, one can meaningfully think about the timing of the instability by not just attempting to estimate $T_{b}^{0}$\textemdash which is not clear how it is defined\textemdash but rather attempting to detect the initial date in the sample after which the forecasts become unstable as well as to detect the last date after which the forecasts remain stable relative to the in-sample period. Since our test statistics are local in nature, one can introduce a procedure which sequentially tests the hypotheses $H_{0}$ in regions of the sample where $H_{0}$ has not yet been rejected. One then records the number of times for which $H_{0}$ is rejected and estimates the corresponding change-point dates. After ordering these change-point dates, one has finally access to useful information such has the initial timing of the forecast instability and the last part of the out-of-sample period which remains stable. Such information can arguably be advantageous to the forecaster.
We now examine the empirical size and power of our proposed tests and compare them to those of giacomini/rossi:09, abbreviated GR (2009). In particular, we consider both the uncorrected and corrected version of the $t_{T_{m},T_{n},\tau}^{\mathrm{stat}}$ statistic of GR (2009).\footnote{We use a superscript $\mathrm{c}$ to indicate the corrected version: $t^{\mathrm{stat,c}}$.} Size and power properties for the Quadratic loss with fixed scheme are reported in Section (ref) and Section (ref), respectively. The Supplemental Material includes corresponding simulation studies for the recursive and rolling schemes and for the Linex loss; these results are not reported here because they are qualitatively equivalent. Overall, one can draw the following conclusions from our simulation study. In terms of size control, the statistics $\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{Q}_{\mathrm{max},h}$ are comparable with the corrected version $t^{\mathrm{stat,c}}$ proposed by GR (2009).\footnote{As shown by GR (2009), the uncorrected version of $t^{\mathrm{stat}}$ can be oversized for models that induce serial dependence in the forecast losses. The authors then proposed a finite-sample correction and did not consider $t^{\mathrm{stat}}$ further in their power analysis. Similarly, GR (2009) showed that just using classical structural break tests in this context is not very helpful as they might have statistical power equals to the size in some cases. Moreover, simulations in perron/yamamoto:18 confirmed that, under rolling and recursive scheme, structural break tests suffer power losses which can be attributed to a so-called “contamination effect” arising when the instability enters the in-sample window {[}see also casini/perron_Oxford_Survey{]}. } Moreover, the test $\mathrm{MQ}_{\mathrm{max},h}$ that uses overlapping blocks is also comparable in terms of size. The same is not true for $\mathrm{MB}_{\mathrm{max},h}$ because often it seems to be somewhat liberal. Turning to the power comparison, each of our test statistics $\mathrm{B}_{\mathrm{max},h}$, $\mathrm{Q}_{\mathrm{max},h}$ and $\mathrm{MQ}_{\mathrm{max},h}$ displays significant power gains over the $t_{T_{m},T_{n},\tau}^{\mathrm{stat}}$ statistics especially as the period of instability (i) is comparatively short relative to the total sample size and/or (ii) is not located at middle sample. In the latter circumstances, the gains in power are, uniformly over different data-generating processes and over parameter break magnitudes, on the order of 30-40%.
Throughout, we restrict attention to one-step ahead forecast horizon (i.e., $\tau=1$), and we use the same loss function for estimation and evaluation. We use our asymptotic results as an approximation for the case where $h=1$ in our theoretical model in (ref) and consider discrete-time DGPs. In models with serially correlated losses (i.e., S2 and S6 below) for the statistics $\mathrm{Q}_{\mathrm{max},h}$ and $\mathrm{MQ}_{\mathrm{max},h}$ we employ the long-run variance estimator from Theorem (ref). With regards to the tests of GR (2009) we use the appropriate version of $t^{\mathrm{stat}}$ and of $t^{\mathrm{stat,c}}$.\footnote{As recommended by GR (2009) we set the truncation lag of their HAC estimator equal to $\left\lfloor T_{n}^{1/3}\right\rfloor $; we also the use the truncation lag $\left\lfloor 0.75T_{n}^{1/3}\right\rfloor $.}
We consider discrete-time DGPs of the form
for various in-sample and out-of-sample sizes and with a total sample size ranging from $T=100$ to $T=500$. Note that (ref) is a special case of the theoretical model with a sampling interval $h=1$. We consider six versions of (ref), where the first and second specification (S1 and S2 below) are calibrated to the Phillips curve model of U.S. inflation from staiger/stock/watson:97: S1 involves $\mu=2.73$, $\beta=-0.44$, and where $\left\{ X_{t}\right\} $ and $\left\{ e_{t}\right\} $ are independent sequences of zero-mean i.i.d. Gaussian disturbances with unit variance; S2 is the same as S1 but with ARCH errors $e_{t}=\sigma_{e,t}u_{t}$, $\sigma_{e,t}=1+0.5e_{t-1}^{2}$ with $u_{t}\sim\mathscr{N}\left(0,\,1\right)$; S3 specifies $\left\{ X_{t}\right\} $ to follow a zero-mean Gaussian AR(1) with autoregressive coefficient 0.4, $\beta=1$ and $e_{t}\sim\mathscr{N}\left(0,\,0.49\right)$ independent of $X_{t}$; S5 is a model with a lagged dependent variable $X_{t-1}=Y_{t-1}$, $\mu=0$, $\beta=0.3$ and $e_{t}\sim\mathscr{N}\left(0,\,0.49\right)$; S6 involves serially correlated disturbances $e_{t}=0.3e_{t-1}+u_{t}$, $u_{t}\sim\mathscr{N}\left(0,\,1\right)$.
Table (ref)-(ref) report the rejection rates for significance levels $\alpha=0.05$ and $0.10$ for model S1-S2. Results for the other DGPs can be found in Table (ref)-(ref) in the Supplement. We first focus on i.i.d. forecast losses (i.e., models S1 and S3-S5). Both $\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{Q}_{\mathrm{max},h}$ are well-sized. As the sample size increases their performance improves and we note that their rejection frequencies are closer to the nominal level when the in-sample size is one half of the total sample. In model S1, when the in-sample size is $0.25T,$ $\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{Q}_{\mathrm{max},h}$ tend to be slightly conservative while the opposite occurs when in-sample size is $0.75T$. The version of $\mathrm{B}_{\mathrm{max},h}$ that uses overlapping blocks $\left(\mathrm{MB}_{\mathrm{max},h}\right)$ can be quite liberal (cf. models S1 and S3). In contrast, $\mathrm{MQ}_{\mathrm{max},h}$ seems to control the size well, though it tends to be slightly liberal but that depends on the relative size of the in-sample and out-of-sample windows. We observe that there is no clear pattern in size performance for our test statistics as we raise the sample size $T$. The reason is straightforward: as we raise $T$ we also need to adjust the choice of $m_{T}$ (the number of blocks) in accordance with Condition (ref). This explains why, for example, in Table (ref), top panel, the empirical size of $\mathrm{Q}_{\mathrm{max},h}$ for $\left(T_{m}=100,\,T_{n}=100\right)$ is better than for $\left(T_{m}=150,\,T_{n}=150\right)$. Turning to the $t^{\mathrm{stat}}$ statistics of giacomini/rossi:09, the uncorrected version performs better than the corrected version since the latter systematically displays an empirical size 2-3% below the nominal level. We can conclude that in models with i.i.d. errors the statistics $\mathrm{B}_{\mathrm{max},h}$, $\mathrm{Q}_{\mathrm{max},h}$, $\mathrm{MQ}_{\mathrm{max},h}$ and $t^{\mathrm{stat}}$ are comparable in terms of empirical size, whereas $\mathrm{MB}_{\mathrm{max},h}$ and $t^{\mathrm{stat,c}}$ tend to over-reject and under-reject, respectively.
Let us now turn to models with serially correlated losses. When the disturbances follow an ARCH process, (cf. model S2, Table (ref)), we observe that both statistics that do not use overlapping blocks, $\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{Q}_{\mathrm{max},h}$, show reasonable size control. The same feature applies to $\mathrm{MQ}_{\mathrm{max},h}$ while $\mathrm{MB}_{\mathrm{max},h}$ displays rejection rates that are systematically above the significance level. It also appears that the corrected version of the statistic of GR (2009) is now regularly below the nominal level. In contrast, the uncorrected version $t^{\mathrm{stat}}$ seems to control size well. When the errors follow an autoregressive process (cf. model S6), Table (ref) shows that $t^{\mathrm{stat}}$ and $\mathrm{MB}_{\mathrm{max},h}$ are arbitrarily oversized for all sample sizes. $\mathrm{MQ}_{\mathrm{max},h}$ and $t^{\mathrm{stat,c}}$ possess rejection rates frequently below the desired nominal level. The statistic that shows the best empirical sizes across different $T$ is $\mathrm{Q}_{\mathrm{max},h}$.
Overall, our analysis on the size properties of the tests suggests that when the DGP involves i.i.d. errors it is fair to compare $\mathrm{B}_{\mathrm{max},h}$, $\mathrm{Q}_{\mathrm{max},h}$, $\mathrm{MQ}_{\mathrm{max},h}$ and $t^{\mathrm{stat}}$ whereas the rejection rates of $\mathrm{MB}_{\mathrm{max},h}$ and $t^{\mathrm{stat,c}}$ tend to deviate systematically from the nominal level. When there are autocorellated errors, it is difficult to compare $t^{\mathrm{stat}}$ and $\mathrm{MB}_{\mathrm{max},h}$ with the other statistics because the former can be highly oversized. The statistics that appear to perform better in terms of approximate size control uniformly over different data-generating mechanisms are $\mathrm{Q}_{\mathrm{max},h}$ and $\mathrm{MQ}_{\mathrm{max},h}$.
We report the small sample power of the tests under various sources of forecast instability. We consider several sample sizes $T$ as well as several designs varying for the distribution of the total sample between in-sample and out-of-sample window. The break date\textemdash or the date of the first change-point when more complicated designs are used\textemdash is denoted by $T_{b}^{0}=T\lambda_{0}$, where $\lambda_{0}\in\left(0,\,1\right)$ is the fractional break date. We shall bring special attention to the location of $T_{b}^{0}$ in the sample as well as to the duration of the instability (i.e., $T-T_{b}^{0}$). We shall see that both factors are actually important for the performance of the methods proposed by giacomini/rossi:09 while our test statistics being local in nature possess essentially uniform power over distinct locations $T_{b}^{0}$. Furthermore, our definition of forecast instability does not demand any relationship between the stable and unstable period and thus it is useful to examine the differences in power properties when a one-time change-point is present relative to when short-lasting instabilities arise.
We consider both discrete shifts\textemdash a structural break\textemdash and recurrent changes in a parameter: model P1a (break in a regression coefficient): $Y_{t}=2.73-0.44X_{t-1}+\delta X_{t-1}\mathbf{1}\left\{ t>T_{b}^{0}\right\} +e_{t}$, where $X_{t-1}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1\right)$ and $e_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1\right)$; model P1b: it is the same as model P1a but with $X_{t-1}\sim\mathrm{i.i.d.}\mathscr{N}\left(1,\,1\right)$; model P2: $Y_{t}=X_{t-1}+\delta X_{t-1}\mathbf{1}\left\{ t>T_{b}^{0}\right\} +e_{t}$, where $X_{t-1}$ is a Gaussian AR(1) with autoregressive coefficient 0.4 and unit variance, and $e_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,0.49\right)$; model P3 (recurrent break in mean): $Y_{t}=\beta_{t}+e_{t}$, where $\beta_{t}$ switches between $\delta$ and $0$ every $p$ periods and $e_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,0.64\right)$; model P4 (single break in variance): $Y_{t}=0.5X_{t-1}+\left(1+\delta\mathbf{1}\left\{ t>T_{b}^{0}\right\} \right)e_{t}$ where $X_{t-1}\sim\mathrm{i.i.d.}\mathscr{N}\left(1,\,1\right)$ and $e_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1\right)$; model P5 (recurrent break in variance): $Y_{t}=\mu+\left(1+\beta_{t}\right)e_{t}$, where $\beta_{t}$ switches between $\delta$ and $0$ every $p$ periods and $e_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,0.49\right)$; model P6 (lagged dependent variable): $Y_{t}=\delta\mathbf{1}\left\{ t>T_{b}^{0}\right\} +0.3Y_{t-1}+e_{t}$, $e_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,0.49\right)$; model P7 (ARCH disturbances): $Y_{t}=2.73-0.44X_{t-1}+\delta X_{t-1}\mathbf{1}\left\{ t>T_{b}^{0}\right\} +e_{t}$, where $X_{t-1}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1.5\right)$ and $e_{t}=\sigma_{t}u_{t}$, $\sigma_{t}^{2}=0.5+0.5e_{t-1}^{2}$, $u_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1\right)$; model P8 (autocorrelated errors): $Y_{t}=1+X_{t-1}+\delta X_{t-1}\mathbf{1}\left\{ t>T_{b}^{0}\right\} +e_{t}$, where $X_{t-1}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1.4\right)$ and $e_{t}=0.4e_{t-1}+u_{t}$, $u_{t}\sim\mathrm{i.i.d.}\mathscr{N}\left(0,\,1\right)$. For models that do not involve recurrent changes we also consider power comparisons when the instability lasts only for some period of time as opposed to the post-$T_{b}^{0}$ period. This requires replacing $\mathbf{1}\left\{ t>T_{b}^{0}\right\} $ in models P1-P2, P4 and P6-P8 with $\mathbf{1}\left\{ T_{b}^{0}<t\leq T_{b}^{0}+p\right\} $ where $p$ is the number of consecutive observations in which the forecast model is unstable. The value of $p$ depends on the sample size $T$. For example, when $T=100$ we set $p=10$; when $T=200$ we set $p=20$ and so on.\footnote{Note that for $\left(T_{m}=50,\,T_{n}=50\right)$ the value $p=10$ corresponds to a period of instability lasting for one-fifth of the out-of-sample; thus, the duration of the instability is nontrivial and consistent with forecasting applications. See the notes to each figure for the other values of $p$. The title of a figure corresponding to a short-lasting instability is labeled “short-term instability”.} The case of short-term instability is the most prevalent in empirical work because it is very unlikely that a professional forecaster would use a poor-performing predictor or forecast model for many consecutive years (e.g., the whole out-of-sample).
Figure (ref)-(ref) in the Appendix plot the power functions for models P1a, P4 and P7. Figure (ref)-(ref) in the Supplement plot the power functions for the remaining DGPs. They include several sample sizes ranging from $T=100$ to $T=500,$ several in-sample and out-of-sample sizes as well as different locations $\lambda_{0}$ of the breaks. We begin with considering general instabilities first and then move to short-term instabilities. Figure (ref)-(ref) reports the results for model P1a. When $T=100,\,150$ Figure (ref) shows that our tests have good power against model P1a while the tests of GR (2009) seem to be less powerful. For example, when the break date is at $T_{0}=0.8T$ our tests display reasonable power. However, both $t^{\mathrm{stat}}$ statistics of GR (2009) perform significantly worse and the associated power curve is bounded away from one even for a very large break size $\delta=3$. This feature disappears when we raise the sample size to $T\geq200$ and maintain the break date at $T_{0}=0.8T$; see Figure (ref). The latter figure also shows that for large sample sizes and instabilities that last for more than 50% of the out-of-sample (top panels) all tests have good power even though the $t^{\mathrm{stat}}$ statistics of GR (2009) have slightly higher power. The power turns to be essentially the same when $\lambda_{0}=0.8$ (i.e., the instability only lasts for 40% of the out-of-sample). For model P2, Figure (ref) plots the power functions for $T=100,\,200$ and $\lambda_{0}=0.7,\,0.8.$ Except for the pair $\left(T=200,\,\lambda_{0}=0.7\right)$ (cf. top-right panel) for which our tests and the $t^{\mathrm{stat}}$-type tests display roughly the same power, it is clear that our tests are more powerful than the $t^{\mathrm{stat}}$ tests (both corrected and uncorrected version). The power gains are substantial and range from 20% to 40%. Moreover, as for model P1a and P1b when the instability lasts for less than 50% of the out-of-sample (cf. $\lambda_{0}=0.8$; bottom-left panel) the statistics $\mathrm{B}_{\mathrm{max},h}$ and $\mathrm{Q}_{\mathrm{max},h}$ achieve trivial power already for a break magnitude $\delta=1.5$ whereas the $t^{\mathrm{stat}}$ tests of GR (2009) display rejection rates below 60% even when $\delta=2$ and yet below 70% when $\delta=2.5$; that is, their power function does not attain unit power even for very large break magnitudes. These properties characterize all models with i.i.d. errors and extend to model with lagged dependent variables as predictors (cf. model P7; Figure (ref) in the Supplement).
Let us now turn to recurrent breaks in the mean. For recurrent breaks we implement the statistics $\mathrm{MB}_{\mathrm{max},h}$ and $\mathrm{MQ}_{\mathrm{max},h}$ that use overlapping blocks. Figure (ref) plots the power curves for model P3. All tests have power and their performance is essentially the same. Figure (ref) corresponds to model P4 (single break in the variance) and shows that when the instability begins in the second half of the out-of-sample (cf. $\lambda_{0}=0.8$; bottom panels) our tests $\mathrm{MB}_{\mathrm{max},h}$ and $\mathrm{MQ}_{\mathrm{max},h}$ achieves good power while the $t^{\mathrm{stat}}$-type tests have little power that does not attain unity even for a large break magnitude $\delta=1.5$. When there are recurrent breaks in the variance as in model P5, Figure (ref) shows that the our tests $\mathrm{MB}_{\mathrm{max},h}$ and $\mathrm{MQ}_{\mathrm{max},h}$ and the $t^{\mathrm{stat}}$-type tests have all good power and their performance is analogous.
Let us now consider models with either ARCH errors or autocorrelated errors. Observe that the latter models both imply that the forecast losses are serially correlated. Figure (ref) shows that when the errors follow an ARCH(1) process the statistic $\mathrm{Q}_{\mathrm{max},h}$ based on the asymptotic variance estimator $\widehat{\nu}_{L}^{2}$ performs well in terms of empirical power. In contrast, the $t^{\mathrm{stat}}$-type tests fail as their power is non-monotonic, never reaches 20% and it decreases to zero as the magnitude of the break rises.\footnote{We actually implemented the $t^{\mathrm{stat}}$ by using either andrews:91's (1991) or newey/west:87's (1987) estiamtor of the long-run variance. We also experimented different choices for the truncation lag. The results, however, were unchanged. We suspect that this property depends on the estimation of the long-run variance in our forecasting context which can be challenging due to small sample sizes and to the presence of breaks. The same issues were found in martins/perron:16 and fossati:17.} We note that the version of $\widehat{\nu}_{L}$ that uses more blocks is less precises. The same results hold true when the disturbances are autocorrelated; see Figure (ref).
Finally, we consider short-term instabilities in Figure (ref)-(ref). It is straightforward to recognize a general pattern: the tests of GR (2009) have little power whereas our tests possess good empirical power against all data-generating processes, break locations and sample sizes. Furthermore, the small sample power properties are uniform over the location of the instability and over the relative size of the in-sample and out-of-sample windows. The latter property is important in practice because forecast instabilities are frequently short-lived.
To sum up, our test statistics perform well in controlling the size, even though the versions that use overlapping blocks are somewhat liberal. For our tests, empirical size being close to the significance level is a feature that holds over different DGPs and sample sizes. Turning to power comparison, there is clear evidence that our tests are reliable in that they have good power against different form of instabilities. There appears to be substantial power gains relative to existing methods especially when the instability (i) is short-lasting and/or (ii) is located toward the tail of the out-of-sample. These properties characterize both statistics using non-overlapping and overlapping blocks.
We have formalized the concepts of forecast instability and forecast failure. Our definition poses at the center the economic forecaster and emphasizes the importance of the time duration of the instability. We assume the data arise as an outcome of an underlying system of stochastic differential equations which then implies that we can approximate the sequence of forecast losses by a continuous-time stochastic process. We have built a testing framework based on the local pathwise properties of that process and have adopted an infill asymptotics to derive the null distribution of the test statistics. The null distribution follows an extreme value distribution. Our results can be used to test whether the predictive ability of a given forecast model changes over time and can be applied in forecasting exercises involving either low-frequency as well as high-frequency macroeconomic and financial variables. The simulation study confirms that there are substantial power gains especially when the instability (i) is short-lasting and/or (ii) is located toward the tail of the out-of-sample. Our framework allows for misspecification, different types of parameter instability and arbitrary forms of non-stationarity such as heteroskedasticity and serial correlation. Our continuous-time specification and associated continuous record asymptotic scheme can provide a promising complementary framework to the classical approach for forecasting in economics.
\pagenumbering{arabic}