EconBase
← Back to paper

A Justification of Conditional Confidence Intervals

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

88,704 characters · 13 sections · 69 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Justification of Conditional Confidence Intervals

titlepage\begin{center} {A Justification of Conditional Confidence Intervals} \normalfont \end{center} \begin{center} Eric Beutner$^*$ Alexander Heinemann$^{**}$ Stephan Smeekes$^{***}$ Department of Quantitative Economics Maastricht University \text{\today} \end{center} \begingroup \footnote{ $^*$e-mail: [email removed]\\ $^{**}$e-mail: [email removed] (corresponding author)\\ $^{***}$e-mail: [email removed] The authors thank Franz Palm, Lenard Lieb, Denis de Crombrugghe, Hanno Reuvers and two anonymous referees for constructive comments and suggestions. This research was financially supported by the Netherlands Organisation for Scientific Research (NWO). } \addtocounter{footnote}{-1} \endgroup \begin{abstract} To quantify uncertainty around point estimates of conditional objects such as conditional means or variances, parameter uncertainty has to be taken into account. Attempts to incorporate parameter uncertainty are typically based on the unrealistic assumption of observing two independent processes, where one is used for parameter estimation, and the other for conditioning upon. Such unrealistic foundation raises the question whether these intervals are theoretically justified in a realistic setting. This paper presents an asymptotic justification for this type of intervals that does not require such an unrealistic assumption, but relies on a sample-split approach instead. By showing that our sample-split intervals coincide asymptotically with the standard intervals, we provide a novel, and realistic, justification for confidence intervals of conditional objects. The analysis is carried out for a rich class of time series models.\\ \\ \textbf{Key words:} Conditional confidence intervals, Parameter uncertainty, Sample-splitting, Prediction, Merging\\ \textbf{JEL codes:} C53, C22, C32, G17 \end{abstract}

\doublespacing

Introduction

One of the open questions in time series is how to quantify uncertainty around point estimates of conditional objects such as conditional means or conditional variances. A fundamental issue arises in the construction of confidence intervals that ought to capture the parameter estimation uncertainty contained in these objects. This fundamental issue stems from the fact that on one hand one must condition on the sample as the past informs about the present, yet on the other hand one must allow the data up to now to be treated as random to account for estimation uncertainty. The issue is well-recognized in the econometric literature, however in practice confidence intervals are commonly constructed by treating the sample simultaneously as fixed and random. Frequently, such approach is motivated by presuming to have two independent processes. Assuming two independent processes with the same stochastic structure, using one for conditioning and one for the estimation of the parameters, bypasses the issue. It is a mathematically convenient assumption as in such case the uncertainty quantification reduces to an ordinary inferential problem. However, practitioners rarely have a replicate, independent of the original series, at hand with the exception of perhaps some experimental settings. As such, the intervals commonly constructed by practitioners lack a satisfactory theoretical justification. Therefore it is the objective of the present paper to develop a realistic justification for such confidence intervals around point estimates of conditional objects.

In the literature the fundamental issue described above is encountered in various ways. In the specific case of a first-order autoregressive (AR) process with Gaussian innovations, phillips1979sampling investigates the statistical dependence between the ordinary least squares (OLS) estimator and the endogenous variable conditioned upon. He obtains an Edgeworth-type expansion for the distribution of the conditional mean and, further, studies forecasting, where the fundamental issue equally arises.\footnote{For prediction intervals some solutions have been discussed. We refer to Section (ref).} lutkepohl2005new (lutkepohl2005new, p.\ 95) explicitly states a two-independent-processes assumption in connection with vector AR models. He postulates that such assumption is asymptotically equivalent to using only data not conditioned upon for estimation. ing2003same clearly distinguish between independent-realization and same-realization settings and study the (unconditional) mean-squared prediction error in the latter for an infinite-order AR process. In a companion paper they also provide a theoretical verification for order selection criteria for same-realization predictions and stress that it can be misleading to assume that the results for independent-realization settings carry over to those for corresponding same-realization cases ing2005order. Other studies investigate parameter uncertainty by using resampling methods, that typically mimic a distribution in which the sample, or at least a subsample, is treated as fixed and random at the same time (cf.\ pascual2004bootstrap, pascual2004bootstrap, pascual2006bootstrap, pan2016abootstrap, pan2016abootstrap, pan2016bbootstrap). Aware of this paradox, kreiss2015discussion points out that conditioning on observing specific in-sample values affects the parameter estimator, but the effect is often erroneously disregarded. Deviating from the various bootstrap approaches, hansen2006interval examines parameter uncertainty in interval forecasts in a classical statistical framework. Similar to a general regression framework, he conditions on an arbitrary fixed out-of-sample value to avoid the issue. However, conditioning on arbitrary fixed out-of-sample values appears incompatible with the usual setup of dynamics in which we condition on the final value(s) of the sample. Acknowledging the issue while avoiding the two independent processes argument bears careful statements as in francq2015risk who write in view of this issue “the delta method ... suggests" (p.\ 162). Similarly, pesaran2015time notices that although such intervals “have been discussed in the econometrics literature, the particular assumptions that underlie them are not fully recognized" (p.\ 389).

This paper provides a novel, and realistic, justification for commonly constructed confidence intervals around point estimates of conditional objects. Our solution is based on a simple sample-split approach and a weak dependence condition, which allows to partition our sample into two asymptotically independent subsamples. For a rich class of time series models we construct asymptotically valid sample-split intervals, without relying on the assumption of observing two independent processes, and show that these intervals coincide asymptotically with the intervals commonly constructed by practitioners. As will be argued below, an appropriate concept to study conditional confidence intervals is merging, a concept that generalizes weak convergence. To the best of our knowledge, except for belyaev2000weakly, this paper is the only one to study merging in the context of conditional distributions. Moreover, our paper seems to be the first to employ merging of conditional distributions in time series. By employing this concept we avoid unnatural assumptions such as observing $X_T=x$ (in dynamic models), losing the time index $T$, and instead explicitly acknowledge that the conditional objects vary over time.

The rest of the paper is organized as follows. Section (ref) specifies the general setup and describes the argument of two independent processes as well as our sample-split approach. In Section (ref) we establish merging among the proposed and the two-independent-processes estimator in probability under mild conditions. Further, we construct asymptotically valid sample-split intervals and show that these coincide asymptotically with the standard intervals. The extension to prediction is discussed in Section (ref). Section (ref) concludes. The main proofs are collected in Appendix A, while Appendix B provides additional proofs of intermediate results.

General Setup

The General Prediction Function

Let $\{X_t\}$ be a real-valued stochastic process defined on the probability space $(\Omega,\mathcal{F},\mathbb{P})$. $\theta$ denotes a generic parameter vector of length $r \in \mathbb{N}$ and $\theta_0$ the true value, unknown to the researcher.\footnote{Generally, in particular throughout Section (ref), we do not distinguish between $\theta$ and $\theta_0$ if there is no cause for confusion. In Section (ref) we explicitly use $\theta_0$ to avoid confusion.} Let $\Theta \subseteq \mathbb{R}^r$ be the corresponding parameter space.

Our general setup involves inference on an object that we call the prediction function, which is a function of both the process $\{X_t\}$ and of the parameter $\theta$. It represents the random object of interest, and will typically express quantities such as a conditional mean or conditional variance (without conditioning on a specific value) as a function of the sample.

definitionThe prediction function $\psi: \mathbb{R}^\infty \times \Theta \rightarrow \mathbb{R}$ is depending both on the parameter $\theta$ and the entire history of the process $\{X_t\}$, such that we can write the prediction of the quantity at time $T+1$, using data up to time $T$, as \begin{equation} \psi_{T+1} := \psi (X_T, X_{T-1}, \ldots ; \theta). \end{equation}

With this setup we can describe most of the possible applications of interest. We now provide three examples to illustrate the prediction function.

exampleSuppose the time series $\{X_t\}$ follows an AR(1) process given by \begin{align} X_t=\beta X_{t-1}+\varepsilon_t\:, \end{align} where $|\beta|<1$ and $\{\varepsilon_t\}$ are independent and identically distributed ($\mbox{\textit{i.i.d.}}$) with $\mathbb{E}[\varepsilon_t]=0$. The conditional mean of $X_{T+1}$ given $X_T$ is given by \begin{align} \mu_{T+1} := \mathbb{E} [X_{T+1} | X_T] = \beta X_T. \end{align} Using the prediction function we can then write $\mu_{T+1} = \psi(X_T, X_{T-1}, \ldots; \theta) = \beta X_T$ with $\theta = \beta$.

A more precarious example, due to its large popularity, is the conditional variance in a generalized autoregressive conditional heteroskedasticity (GARCH) model engle1982autoregressive,bollerslev1986generalized. Whereas in the previous AR($1$) case it suffices to condition on the terminal observation, the subsequent Example 2 is more extreme as the entire sample contains information about the object of interest.

exampleSuppose $\{X_t\}$ follows a GARCH$(1,1)$ process given by $X_t = \sigma_t \varepsilon_{t}$ with \begin{align} \sigma_t^2 = \omega + \alpha X_{t-1}^2+ \beta \sigma_{t-1}^2\:, \end{align} where $\omega>0$, $\alpha \geq 0$, $1>\beta \geq 0$ and $\{\varepsilon_t\}$ are $\mbox{\textit{i.i.d.}}$ with $\mathbb{E}[\varepsilon_t]=0$ and $\mathbb{E}[\varepsilon_t^2]=1$. The model's recursive structure implies \begin{align} \sigma_{T+1}^2=& \frac{\omega}{1-\beta}+\alpha \sum_{k=0}^{\infty} \beta^k X_{T-k}^2. \end{align} It follows directly from (ref) that \begin{equation*} \sigma_{T+1}^2 = \psi (X_T, X_{T-1}, \ldots; \theta) = \frac{\omega}{1-\beta}+\alpha \sum_{k=0}^{\infty} \beta^k X_{T-k}^2, \end{equation*} with $\theta = (\omega, \alpha, \beta)^\prime$, and $\Theta \subset(0,\infty)\times [0,\infty) \times [0,1)$.

The next example shows that for a large class of models the prediction function can be written in the form of Definition (ref).

exampleFollowing boussama2011stationarity, consider a Markov chain of the form \begin{align} S_t=\varphi(S_{t-1},X_t;\theta), \qquad t=1,2,\dots \end{align} where $\varphi$ is some map $\varphi:\mathbb{R}^a\times \mathbb{R} \times \Theta \to \mathbb{R}^a$. Whereas $X_t$ is observable by the researcher at time $t$, $S_t$ may be unobservable or only partially observable. The object of interest $\psi_{T+1}$ is typically a function of the state of the Markov chain $S_T$, such that $\psi_{T+1} = \pi(S_T; \theta)$ for some function $\pi(\cdot)$. Through the recursion in (ref), this is in turn a function of the past of $X_T$, such that we may write \begin{equation*} \psi_{T+1} = \pi(S_T; \theta) = \psi_{T+1} (X_T, X_{T-1}, \ldots; \theta). \end{equation*} Many stochastic processes are in fact Markov processes, including ARMA and GARCH models, several GARCH extensions such as zakoian1994threshold's (zakoian1994threshold) threshold GARCH, and the set of observation driven models considered by blasques2016sample. For further details we refer to beutner2017justification and beutner2017technical.

Note that in many cases, such as the GARCH(1,1) of Example (ref), the prediction function actually depends on the infinite past of the series. In order to express (an approximation of) the prediction function in terms of observable variables only, we would need to replace $X_t$ by $s_t$ for all $t<1$, where $\{s_t\}$ is a sequence of (arbitrary) constants to which we refer as starting or initial values. For a fixed $T$, we accordingly define an approximate prediction function $\psi_{T+1}^s: \mathbb{R}^T \times \Theta \rightarrow \mathbb{R}$ that is only a function of observable variables as

equation[equation omitted — 138 chars of source]

where $\mathbf{X}_{1:T} = (X_1, \ldots, X_T)^\prime$. Note that, given the varying input of the left-hand side in ((ref)), we now actually have a sequence of (varying) functions for $T \in \mathbb{N}$.

In many cases the values far in the past are negligible for a wide range of values for $\{s_t\}$. Consequently, $\psi_{T+1}^s$ will be close to $\psi_{T+1}$. This property can be shown to hold for many different processes including the ones in the examples. We formalize the exact condition we need regarding the negligibility of the starting values in Assumption (ref).b.

Although the prediction function typically represents a conditional object, we have not conditioned on anything yet in the definition. We therefore now extend the analysis by formally conditioning on observing a particular sample. Let $\mathbf{x}_{1:T} = (x_1,\dots,x_T)'$ denote a specific sample path of $\mathbf{X}_{1:T}$. Throughout the paper, we will discriminate between random variables and their realized counterparts by writing the former in capital and the latter in lowercase letters to avoid ambiguity.

As we will consider sample splitting later on, we define notation that also allows for conditioning on only a subsample. For that purpose, let $t_1:t_2$ denote the (sub-)period from $t_1$ up to $t_2$, and correspondingly $\mathbf{X}_{t_1:t_2} = (X_{t_1}, \ldots, X_{t_2})^\prime$ for any integers $1 \leq t_1 \leq t_2 \leq T$, with a corresponding definition for the observed subsample $\mathbf{x}_{t_1: t_2}$. Furthermore, let $\mathbf{X}_{t_1:T}^{c} = (c_1, \ldots, c_{t_1-1}, X_{t_1}, \ldots, X_{T})^\prime$ denote the vector where all non-considered subsamples are replaced by a sequence of constants $\{c_t\}$, in a similar way as we did for the starting values. We can now formally define the conditional prediction function.

definitionThe prediction function conditional on observing $\mathbf{X}_{t_1:T} = \mathbf{x}_{t_1:T}$ is defined as \begin{equation} \psi_{T+1|t_1:T} := \psi_{T+1}^s (\mathbf{x}_{t_1:T}^c; \theta). \end{equation}

Note that we phrase the conditional prediction function directly in terms of the approximate prediction function $\psi_{T+1}^s$ rather than the true prediction function. We take this “shortcut” because we cannot observe $x_{0}, x_{-1}, \ldots$, so we cannot condition on those values anyway. Therefore, the “true” conditional object (which we might represent as $\psi_{T+1| -\infty: T}$), is, from an applied point of view, only the theoretical benchmark.

Example(ref). (continued) For the conditional mean of an AR(1) process, conditioning only on the terminal observation $X_T = x_T$ suffices; that is, for any $t_1 \geq 1$ and any sequence $\{c_t\}$, we have that \begin{equation} \psi_{T+1|t_1:T} = \psi_{T+1}^s (\mathbf{x}_{t_1:T}^c; \theta) = \beta x_T = \psi_{T+1}^s (\mathbf{x}_{T:T}^c; \theta) = \psi_{T+1|T}. \end{equation}
Example(ref). (continued) For objects such as the conditional variance for the GARCH(1,1), the conditioning set and the sequence $\{c_t\}$ make a difference, as \begin{equation} \begin{split} \psi_{T+1|t_1:T} = \psi_{T+1}^s (\mathbf{x}_{t_1:T}^c; \theta) &= \frac{\omega}{1-\beta} + \alpha \sum_{k=0}^{T-t_1} \beta^k x_{T-k}^2 \\ &\quad + \alpha \beta^{T-t_1} \sum_{k=1}^{t_1-1} \beta^k c_{t_1-k}^2 + \alpha \beta^T \sum_{k=0}^{\infty} \beta^k s_{-k}^2, \end{split} \end{equation} which differs depending on the choice of $t_1$. However, as will be shown later, with an appropriate choice of $t_1$, our Assumption (ref).b on the negligibility of the initial condition, also implies that the difference between $\psi_{T+1|t_1:T}$ and $\psi_{T+1|1:T}$ becomes negligible asymptotically.

Before introducing estimators for $\theta$ let us discuss the objects we want to construct inference for. In principle there are two unknown objects one could develop statistical intervals for: $\psi_{T+1}^s (\mathbf{X}_{1:T}; \theta)$ (or slightly more generally $\psi_{T+1}^s (\mathbf{X}_{t_1:T}^c; \theta)$) and $\psi_{T+1}^s (\mathbf{x}_{t_1:T}^c; \theta).$ For a GARCH(1,1), for instance, the first would read as $$ \psi_{T+1}^s (\mathbf{X}_{1:T}; \theta) = \frac{\omega}{1-\beta} + \alpha \sum_{k=0}^{T-1} \beta^k X_{T-k}^2 + \alpha \beta^T \sum_{k=0}^{\infty} \beta^k s_{-k}^2 $$ whereas the second with $t_1=1$ reads as $$ \psi_{T+1}^s (\mathbf{x}_{1:T}; \theta) = \frac{\omega}{1-\beta} + \alpha \sum_{k=0}^{T-1} \beta^k x_{T-k}^2 + \alpha \beta^T \sum_{k=0}^{\infty} \beta^k s_{-k}^2 $$ or more generally, if $t_1$ is not taken to be equal to one, as in ((ref)). While statistical intervals for both objects can be constructed we focus here on conditional inference, i.e. on intervals for $\psi_{T+1|t_1:T}=\psi_{T+1}^s (\mathbf{x}_{t_1:T}^c; \theta)$. In a time series context intervals for $\psi_{T+1|t_1:T}$ are motivated by the relevance property of kabaila1999relevance which postulates that intervals should relate to what actually happened during the sample period opposed to what might have happened. Indeed, intervals for $\psi_{T+1|t_1:T}$ can theoretically be shown to be considerably shorter than the intervals for their unconditional counterparts. While the unconditional objects might lead to conceptually easier analysis, our focus on the conditional objects is therefore not only theoretically but also empirically relevant.

Estimating the Prediction Function

As $\theta$ is unobserved, we need to estimate it. We assume that the estimator is based on a subsample $1:T_E$ (with $1 \leq T_E \leq T$) of the process $\{X_t^{E}\}$ which is potentially a different sample than $\{X_t\}$ that arises in the prediction function. The estimator of $\theta$ based on $\mathbf{X}_{1:T_E}^E=(X_1^E,\dots, X_{T_E}^E)^\prime$ will be denoted by $\hat{\theta} (\mathbf{X}_{1:T_E}^E)$. The introduction of $\{X_t^{E}\}$ serves three purposes: first, using a different process allows us to formulate the two-independent-processes argument where $X_t^E = Y_t$, with $\{Y_t\}$ independent of $\{X_t\}$, $T_E = T$ and an interval is constructed for $\psi_{T+1|1:T}$. Second, it will allow us to discuss the standard approach where $ X_t^E = X_t$, $T_E=T$, and an interval is constructed for $\psi_{T+1|1:T}$. Please note already here that this means that the same variables that arise in the prediction function are also used for estimating $\theta$. Third, it allows us to define the sample splitting approach which we study here. In this approach $ X_t^E = X_t$ and an interval is constructed for $\psi_{T+1|T_P:T}$ with $T_E < T_P$ (with $1 < T_P \leq T$) such that in contrast to the standard approach different subsamples are used for prediction and estimation. \\ Before we illustrate why the standard approach is problematic for constructing and evaluating conditional intervals, we need to define the final building block of prediction function estimation: the conditional prediction function estimator.

definitionLet $1 \leq T_P \leq T$. Define the prediction function estimator conditional on observing $\mathbf{X}_{T_P:T} = \mathbf{x}_{T_P:T}$ as \begin{equation} \widehat{\psi}_{T+1|T_P:T} := \widehat{\psi}_{T+1} (\mathbf{x}_{T_P:T}^{c}, \mathbf{X}_{1:T_E}^E) = \psi_{T+1}^s (\mathbf{x}_{T_P:T}^c, \hat{\theta} (\mathbf{X}_{1:T_E}^E)). \end{equation}

Note that in the above definition we do not condition on the sample $\mathbf{X}_{1:T_E}^E = \mathbf{x}_{1:T_E}^E$ that is used to estimate $\theta$. The reason for not conditioning on $\mathbf{X}_{1:T_E}^E = \mathbf{x}_{1:T_E}^E$ is that the goal is to preserve the randomness in the second argument of $\psi_{T+1}^s$, i.e. in $\hat{\theta} (\mathbf{X}_{1:T_E}^E)$, and consequently in $\widehat{\psi}_{T+1|T_P:T}$. Hence, if this goal is achieved we can use the (non-degenerate) conditional (on $\mathbf{X}_{T_P:T} = \mathbf{x}_{T_P:T}$) distribution of $\widehat{\psi}_{T+1|T_P:T}$ to construct confidence intervals for $\psi_{T+1|T_P:T}$. Having this said let us have a closer look at the standard approach. As mentioned above in the standard approach one has $X_t^E = X_t$, $T_P=1$ and $T_E=T$. Hence, denoting by $\widehat{\psi}_{T+1|1:T}^{STA}$ the “standard” estimator of the prediction function conditional on observing $\mathbf{X}_{1:T} = \mathbf{x}_{1:T}$, it becomes

equation[equation omitted — 182 chars of source]

Notice that there is no capital $\mathbf{X}$ in ((ref)) because there is only one sample and one typically conditions on all values of this sample. Hence, ((ref)) is non-random and thus does not have a distribution that could be used to construct intervals. Instead, to still be able to construct a “standard-looking” interval in practice, researchers typically implicitly rely on the (approximate) quantiles of the estimator

align[align omitted — 176 chars of source]

It is well understood in the literature that considering the sample as random and non-random at the same time as in ((ref)) does not provide a fully satisfactory justification of the intervals used in practice. For the readers not so familiar with the problem just discussed we provide two examples that both illustrate the problem arising from ((ref)). The examples illustrate that the severity of the problem may vary; ranging from only complicating the analysis (Example (ref)) to making the analysis impossible (Example (ref)).

Example(ref). (continued) For the AR(1), we know from ((ref)) that $\psi_{T+1|1:T}=\psi^s_{T+1}(\mathbf{x}_{1:T},\theta)=\beta x_T$. Estimating $\beta$ by OLS, say $\hat{\beta}(\mathbf{X}_{1:T})$, the estimator in ((ref)) becomes \begin{equation} \hat{\psi}_{T+1|1:T}^{STA*} = \psi_{T+1} (\mathbf{x}_{1:T}; \hat{\theta}(\mathbf{X}_{1:T})) = \psi_{T+1}^s (\mathbf{x}_{1:T}, \hat{\beta} (\mathbf{X}_{1:T})) = \hat{\beta} (\mathbf{X}_{1:T}) x_T. \end{equation} Note the discrepancy in treating the terminal observation as random in the estimation sample, yet fixed for the prediction sample. To construct an interval for $\beta x_T$, one uses \begin{align} \sqrt{T}\big(\hat{\beta}(\mathbf{X}_{1:T})-\beta\big)\overset{d}{\to} N(0,\sigma_\beta^2) \end{align} with $\sigma_\beta^2=1-\beta^2$ (cf.\ hamilton1994time, hamilton1994time, p.\ 215) and that one can estimate the variance of this normal distribution by $\hat{\sigma}_\beta^2(\mathbf{X}_{1:T})=1-\hat{\beta}(\mathbf{X}_{1:T})^2$. Then an interval for $\beta x_T$ is typically constructed the following way: \begin{align} \hat{\beta}(\mathbf{X}_{1:T})x_T\pm \Phi^{-1}(\gamma/2)\:x_T\:\hat{\sigma}_\beta(\mathbf{X}_{1:T})/\sqrt{T}\:, \end{align} where $\Phi^{-1}$ denotes the standard normal quantile function. However, the interval in (ref) is hard to interpret as the terminal observation is treated simultaneously as fixed and random. In essence, researchers typically approximate the distribution of $\sqrt{T}(\hat{\beta}(\mathbf{X}_{1:T})-\beta) x_T$ instead of the conditional distribution of $\sqrt{T}(\hat{\beta}(\mathbf{X}_{1:T})-\beta) X_T$ given $X_T=x_T$. The approximation of the latter appears rather cumbersome because even the rather simple condition $X_T=x_T$ has an influence on the whole series $\mathbf{X}_{1:T}$ (kreiss2015discussion, kreiss2015discussion). Despite the challenge, phillips1979sampling obtains such approximation based on Edgeworth expansions in the case of $\varepsilon_t\overset{iid}{\sim}N\big(0,\sigma^2_\varepsilon\big)$.
Example(ref). (continued) For the conditional variance of the GARCH(1,1), the standard estimator of the prediction function conditional on $\mathbf{X}_{1:T} = \mathbf{x}_{1:T}$, is given by \begin{align} \begin{split} \hat{\sigma}_{T+1|1:T}^{2\;STA} = \psi_{T+1}^s \big(\mathbf{x}_{1:T};\hat{\theta}(\mathbf{x}_{1:T})\big) &= \frac{\hat{\omega}(\mathbf{x}_{1:T})}{1-\hat{\beta}(\mathbf{x}_{1:T})}+\hat{\alpha}(\mathbf{x}_{1:T}) \sum_{k=0}^{T-1} \hat{\beta}(\mathbf{x}_{1:T})^k x_{T-k}^2\\ &\quad +\hat{\alpha}(\mathbf{x}_{1:T}) \hat{\beta}(\mathbf{x}_{1:T})^T \sum_{k=0}^{\infty} \hat{\beta}(\mathbf{x}_{1:T})^k s_{-k}^2 \:, \end{split} \end{align} where $\hat{\theta}(\mathbf{x}_{1:T})=\big(\hat{\omega}(\mathbf{x}_{1:T}),\hat{\alpha}(\mathbf{x}_{1:T}),\hat{\beta}(\mathbf{x}_{1:T})\big)'$ is some estimate for $\theta$ depending on $\mathbf{x}_{1:T}$. Clearly, ((ref)) illustrates for the GARCH(1,1) the above mentioned problem that the standard estimator is not random (after conditioning). For the GARCH(1,1) the estimator in ((ref)) whose quantiles are used for an interval reads as \begin{align} \begin{split} \hat{\sigma}_{T+1}^{2\;STA*}=\psi_{T+1}\big(\mathbf{x}_{1:T};\hat{\theta}(\mathbf{X}_{1:T})\big) &= \frac{\hat{\omega}(\mathbf{X}_{1:T})}{1-\hat{\beta}(\mathbf{X}_{1:T})}+\hat{\alpha}(\mathbf{X}_{1:T}) \sum_{k=0}^{T-1} \hat{\beta}(\mathbf{X}_{1:T})^k x_{T-k}^2\\ &\quad +\hat{\alpha}(\mathbf{X}_{1:T}) \hat{\beta}(\mathbf{X}_{1:T})^T \sum_{k=0}^{\infty} \hat{\beta}(\mathbf{X}_{1:T})^k s_{-k}^2 \:. \end{split} \end{align} This quantity exemplifies for the GARCH(1,1) that the complete sample is regarded as random and non-random at the same time. While for the AR(1) this complicated the analysis, yet not made it impossible, the dependence on the complete sample here makes it difficult to use this quantity to make meaningful probabilistic statements.

Argument of Two Independent Processes

The argument of two independent processes can at least be traced back to akaike1969fitting, who studies the prediction of AR time series. It reoccurs in lewis1985prediction (lewis1985prediction, p.\ 394): “...the series used for estimation of parameters and the series used for prediction are generated from two independent processes which have the same stochastic structure." The same argument also appears in lutkepohl2005new (lutkepohl2005new, p.\ 95) and in dufour2010short. Let $\{Y_t\}$ be a process independent of $\{X_t\}$ defined on the same probability space $(\Omega,\mathcal{F},\mathbb{P})$ with $\{Y_t\}$ having the same stochastic structure as $\{X_t\}$. In addition to the sample $\mathbf{X}_{1:T}$ of the process $\{X_t\}$, suppose there is a sample $\mathbf{Y}_{1:T} = (Y_1,\dots,Y_T)'$ of the process $\{Y_t\}$ that we use as estimation sample, that is $\mathbf{X}_{1:T}^E = \mathbf{Y}_{1:T}$. In this situation we denote the conditional prediction function estimator of Definition (ref) by $\widehat{\psi}_{T+1|1:T}^{2IP}$ and it equals

align[align omitted — 189 chars of source]

Notice that ((ref)) does not have the same shortcoming as ((ref)) because even if we consider $\mathbf{x}_{1:T}$ to be known we can nevertheless consider $\hat{\theta}(\mathbf{Y}_{1:T})$ to be random and can hence use its distribution to construct intervals. Throughout this paper, we call (ref), the 2IP (two independent processes) estimator. Then, a conditional interval $I_\gamma^{2IP}(\mathbf{x}_{1:T}, \mathbf{Y}_{1:T})$ can be based on the (approximate) quantiles of $\widehat{\psi}_{T+1|1:T}^{2IP}$. It satisfies

align[align omitted — 198 chars of source]

with the approximate sign indicating asymptotic equivalence. Note that the independence implies that the distribution of $\mathbf{Y}_{1:T}$ in ((ref)) does not depend on the realization $\mathbf{x}_{1:T}$, yet the statement does depend on $\mathbf{x}_{1:T}$ because the interval depends on it (for the AR(1) this can be directly seen from ((ref)) when replacing $\mathbf{X}_{1:T}$ by $\mathbf{Y}_{1:T}$). Although the 2IP approach is statistically sound, it assumes two independent processes with the same stochastic structure. phillips1979sampling points out that the assumption “is quite unrealistic in practical situations" (p.\ 241). Indeed, it is difficult to imagine this assumption to be satisfied in any real-life application beyond experimental settings. Moreover, as only one sample realization is available, to compute the estimate of the interval $I_\gamma^{2IP}(\mathbf{x}_{1:T}, \mathbf{Y}_{1:T})$ it is frequently suggested to take $ \mathbf{Y}_{1:T} = \mathbf{x}_{1:T}$, violating the independence assumption. Thus, the 2IP approach appears to be a rather questionable justification for the usual interval, and as such, in this paper we provide an alternative, realistic, justification of (asymptotically equivalent) intervals based on sample splitting.

Sample-split Estimation

An intuitive motivation for the sample-split approach is the successive decline of the influence of past observations present in a substantial class of time series models. This property permits to split our sample into two (asymptotically) independent subsamples. Consider the end point of the estimation sample, $T_E$, and the starting point of the prediction sample, $T_P$ satisfying $1< T_E < T_P \leq T$, such that the two samples are non-overlapping. In this situation we denote the conditional prediction function estimator of Definition (ref) by $\widehat{\psi}_{T+1|T_P:T}^{SPL}$ and it is given by

align[align omitted — 204 chars of source]

Throughout the paper, we call ((ref)) the SPL estimator (due to SPLitting). Similar to the two sample approach, we can consider the first argument of $\psi_{T+1}^s$ in ((ref)) as given and the second argument as random since the subsamples are non-overlapping. A conditional interval $I_\gamma^{SPL}(\mathbf{x}_{T_P:T}, \mathbf{X}_{1:T_E})$ can be constructed such that

align[align omitted — 210 chars of source]

This statement does make sense as there is still randomness in $\hat{\theta}(\mathbf{X}_{1:T_E})$ since $\mathbf{X}_{1:T_E}$ is not conditioned upon, yet the last $T - T_P+1$ values of $\{X_t\}_{t=1}^T$ are fixed such that their randomness is not taken into account. Similar to ((ref)), the statement in ((ref)) does depend on $\mathbf{x}_{T_P:T}$ and in contrast to $\mathbf{x}_{1:T}$ in ((ref)) the realization $\mathbf{x}_{T_P:T}$ may influence the distribution of $\mathbf{X}_{1:T_E}$. However, as said at the beginning of this subsection the idea of the sample split approach is that this dependence will vanish asymptotically.

remarkIn Section (ref) we will discuss how $T_E$ and $T_P$ should be chosen from an asymptotic perspective to ensure that our regularity conditions are fulfilled. As we only consider sample splitting as a theoretical approach to validate commonly constructed conditional confidence intervals, these asymptotic guidelines are sufficient for our purposes and we do not have to consider how to choose $T_E$ and $T_P$ in practice. Of course, one could use the sample-split approach in practice to construct confidence intervals. While one would gain (near) independence between the two subsamples, this would come at a cost of estimation precision as fewer observations are used for parameter estimation. For the Gaussian AR($1$) setting, phillips1979sampling derives asymptotic expansions for the case where, in our notation, $T_E = T-l$ and $T_P = T$ for some $l \geq 0$, showing that even in this simple case there is indeed a trade-off as described above and the optimal choice of $l$ is unclear. An interesting extension of our analysis would therefore be to investigate the optimal choices of $T_E$ and $T_P$ to achieve the most accurate confidence intervals in small samples. However, this choice is likely to be highly dependent on the specific model and as such would have to be investigated on a case-by-case basis. This is therefore outside the scope of the current paper.

Asymptotic Justification

In this section, we connect the sample-split procedure of Section (ref) with the two-independent-samples approach of Section (ref). First, in Section (ref), we show that the notion of weak convergence is inadequate to study asymptotic closeness for objects that vary over time and discuss the concept of merging. Then, in Section (ref) we link the 2IP and the SPL estimator by proving that their conditional distributions merge in probability (Theorem (ref)). Thereafter, in Section (ref), we construct asymptotically valid intervals (Theorem (ref)) and show that the sample-split intervals coincide asymptotically with the intervals commonly constructed by practitioners (Theorem (ref)). Last, in Section (ref), we state intervals of reduced form and simplified theoretical results under asymptotic normality of the parameter estimator.

Merging

To illustrate the inappropriateness of weak convergence in the context considered here, we revisit Example (ref) for the 2IP approach and the SPL approach, which shows that studying asymptotic closeness between conditional distributions is often complicated by the absence of a limiting distribution.

Example(ref). (continued) For the 2IP approach (ref) implies that $\sqrt{T}\big(\hat{\beta}(\mathbf{Y}_{1:T})-\beta\big)\overset{d}{\to} N(0,\sigma_\beta^2)$ and it entails that $\sqrt{T}(\hat{\beta}(\mathbf{Y}_{1:T})-\beta)x$ converges weakly to $N(0,\sigma_\beta^2x^2)$ for any fixed $x\neq 0$. Further, the result suggests that the conditional distribution of $\sqrt{T}(\hat{\beta}(\mathbf{Y}_{1:T})-\beta)X_T$ given $X_T=x_T$, which is just the distribution of $\sqrt{T}(\hat{\beta}(\mathbf{Y}_{1:T})-\beta)x_T$, is asymptotically close to $N(0,\sigma_\beta^2x_T^2)$. Similarly, for the SPL-approach with $T_E \slash T \rightarrow 1$ we have $\sqrt{T}\big(\hat{\beta}(\mathbf{X}_{1:T_E})-\beta\big)\overset{d}{\to} N(0,\sigma_\beta^2)$ which suggests as well (if the gap between $T_E$ and $T$ is large enough which will be formally specified below) that the conditional distribution of $\sqrt{T_E}(\hat{\beta}(\mathbf{X}_{1:T_E})-\beta)X_T$ given $X_T=x_T$ is also close to $N(0,\sigma_\beta^2x_T^2)$. For both approaches the approximating distribution $N(0,\sigma_\beta^2x_T^2)$ varies with $T$ through the terminal realization $x_T$. Note that the concept of weak convergence is not applicable in this context to characterize this asymptotic closeness, as it requires a (fixed) limiting distribution, which is absent here.

Next, we discuss what closeness means in the absence of a limiting distribution. To do so, first recall that weak convergence of a sequence of cdfs $\{F_T\}$ on $\mathbb{R}^k$ with $k \in \mathbb{N}$, i.e.\ $F_T(x)\to F(x)$ for all continuity points of $F$, can alternatively be defined by $d_{BL}(F_T,F) \rightarrow 0$. Here $d_{BL}$ denotes the bounded Lipschitz metric defined by

align[align omitted — 107 chars of source]

where for any real-valued function $f$ on $\mathbb{R}^k$ one puts $||f||_{BL}=\sup_x \big|f(x)\big|+\sup_{x \neq y}\frac{|f(x)-f(y)|}{||x-y||}$, with $||\cdot||$ denoting the Euclidean norm, i.e.\ $||A||=\sqrt{tr(A'A)}$ for any vector or matrix $A$. Following dudley2002real (see d1988merging and davydov2009asymptotic for related work) we state the following definition.

definition(Merging) Two sequences of cdfs $\{F_T\}$ and $\{G_T\}$ are said to merge if and only if $d_{BL}(F_T,G_T) \to 0$ as $T\to \infty$.

Note that weak convergence can be seen as a special case of merging with $G_T=G$ for all $T \in \mathbb{N}$.\\ While merging is appropriate to capture the asymptotic closeness of the conditional distribution of $\sqrt{T}(\hat{\beta}(\mathbf{Y}_{1:T})-\beta)X_T$ and $N(0,\sigma_\beta^2x_T^2)$ for a given sample $X_T=x_T$, we now extend the concept in a way that allows us to deal with asymptotic closeness when we do not condition on a particular sample. The necessity of this definition can again be exemplified by the AR(1), which also illustrates how we will deal the dependence of the statements in ((ref)) and in ((ref)) on the sample that we mentioned below these equations. For instance, in Example (ref) as described at the beginning of this section, the goal would be to formalize a statement like `when $T$ is large, the probability of all $x_T$ such that the distribution of $\sqrt{T}(\hat{\beta}(\mathbf{X}_{1:T})-\beta)x_T$ merges with that of $\sqrt{T_E}(\hat{\beta}(\mathbf{X}_{1:T_E})-\beta)x_T$ is approximately equal to one'. We now first introduce the conditional distributions of the 2IP and the SPL estimator in the general case and then give the definition capturing what we just illustrated for the AR(1).\\ Let $m_T$ be a sequence of normalizing constants with $m_T \to \infty$ (e.g.\ $m_T=\sqrt{T}$). For any $t_1 \geq 1$, we define the sub $\sigma$-algebra $\mathcal{I}_{t_1:T} = \sigma(X_t: t_1 \leq t \leq T)$. We denote the conditional cdfs of the 2IP and SPL estimator by

align[align omitted — 398 chars of source]

respectively, so that by specifying an event of $\mathcal{I}_{1:T}$ and $\mathcal{I}_{T_P:T}$, we see that ((ref)) and ((ref)) are just the centered and scaled distributions of ((ref)) and ((ref)), respectively. Please note that ((ref)) actually also depends on $c$, see ((ref)), but since our assumptions will ensure that this dependence vanishes asymptotically we prefer to suppress the dependence on $c$ here.

remarkAlthough not explicitly mentioned above we consider ((ref)) and ((ref)) to be regular conditional cdfs, which indicates that we assume that $F_T^{2IP}( \cdot|\mathcal{I}_{1:T})(\omega)$ and $F_T^{SPL}( \cdot | \mathcal{I}_{T_p:T})(\omega)$ are cdfs for every $\omega \in \Omega$; for the exact definition and the existence see dudley2002real (dudley2002real, Section 10.2).

We can now define merging in probability (we do so without explicitly using the conditional cdfs of the 2IP and the SPL approach).

definition(Merging in Probability) Two sequences of conditional cdfs $\{F_T\}$ and $\{G_T\}$ are said to merge in probability if and only if $d_{BL}(F_T,G_T) \overset{p}{\to} 0$ as $T\to \infty$, where “$\overset{p}{\to}$" denotes “convergence in probability".

Merging of 2IP and SPL in Probability

Here, we give conditions such that the conditional cdfs of the 2IP and SPL estimator merge. Clearly, the conditional confidence intervals are functions of these distributions so that their merging is a building block for the study of the conditional confidence intervals based on them. The conditions we give are divided into three parts. Roughly speaking, the first part (general assumptions) makes sure that the function we want to predict is well behaved and that we can estimate the parameter it depends on. The second part (two independent processes) and third part (SPL estimator) guarantee that these assumptions are met by the 2IP and the SPL method. To write the conditions in compact form we employ the usual stochastic order symbols $O_p$ and $o_p$. We assume that $\theta_0$ belongs to the interior of $\Theta$, i.e.\ $\theta_0 \in \mathring{\Theta}$, and we denote the set of all bounded, real-valued Lipschitz functions on $\mathbb{R}^r$ by $BL=\big \{h:\mathbb{R}^r\to \mathbb{R}:||h||_{BL}<\infty\big \}$. We start with the general assumptions.

assumption(General Assumptions) \begin{enumerate}[(ref).a] • (Estimator) $m_T \big(\hat{\theta}(\mathbf{X}_{1:T}) - \theta_0\big)\overset{d}{\to} G_{\infty}$ as $T \to \infty$ for some cdf $G_{\infty}: \mathbb{R}^r \rightarrow [0,1]$; • (Differentiability) $\psi(\:\cdot\:; \theta)$ is continuous on $\Theta$ and twice differentiable on $\mathring{\Theta}$; • (Gradient) $\Big|\Big|\frac{\partial \psi (X_T, X_{T-1}, \ldots ;\theta_0)}{\partial \theta}\Big|\Big|=O_{p}(1)$; • (Hessian) $\sup_{\theta \in \mathscr{V}(\theta_0)}\Big|\Big|\frac{\partial^2 \psi (X_T, X_{T-1}, \ldots ;\theta)}{\partial \theta \partial \theta'}\Big|\Big|=O_{p}(1)$ for some open neighborhood $\mathscr{V}(\theta_0)$ around $\theta_0$; • (Initial Condition) Given sequences $\{s_t\}$ and $\{c_t\}$, we have \begin{align*} m_T\big(\psi_{T+1}^s (\mathbf{X}_{t_1:T}^c; \theta_0) - \psi (X_T, X_{T-1}, \ldots ;\theta_0)\big)=o_{p}(1),\\ \bigg|\bigg|\frac{\partial \psi_{T+1}^s (\mathbf{X}_{t_1:T}^c; \theta_0)}{\partial \theta} - \frac{\partial \psi (X_T, X_{T-1}, \ldots ;\theta_0)}{\partial \theta}\bigg|\bigg|=o_{p}(1),\\ \sup_{\theta \in \mathscr{V}(\theta_0)}\bigg|\bigg|\frac{\partial^2 \psi_{T+1}^s (\mathbf{X}_{t_1:T}^c; \theta)}{\partial \theta \partial \theta'} - \frac{\partial^2 \psi (X_T, X_{T-1}, \ldots ;\theta)}{\partial \theta \partial \theta'}\bigg|\bigg|=o_{p}(1) \end{align*} for any $t_1 \geq 1$ such that $(T-t_1) / l_T \rightarrow \infty$ as $T \to \infty$ and for some model-specific $l_T$ with $l_T \rightarrow \infty$. \end{enumerate}

Assumption (ref).a implies the existence of a limiting distribution for the parameter estimator. The differentiability assumption in (ref).b plus the boundedness Assumptions (ref).c ensure that the scaled prediction function estimators can accurately be approximated by a Taylor expansion; see Lemma (ref) for details. Assumption (ref).e with $t_1 = 1$ ensures the negligibility of the starting values when using the full-sample for prediction, while taking $t_1 = T_P$ ensures that this extends to the case where additionally $X_1,\ldots,X_{T_P-1}$ are replaced by constants, i.e. where only the subsample $(X_{T_P},\ldots,X_{T})$ is used for prediction. This assumption implicitly limits the choice of $T_P$; as replacing past values of $X_t$ for $t < T_P$ by arbitrary constants should have a negligible effect, $T - T_P$ needs to increase faster than some lower bound $l_T$. For models exhibiting an exponential decay in memory, it typically suffices to take $l_T = \log T$ (see e.g. beutner2019technical, beutner2019technical, eq. (4.6)).\\ For the 2IP estimator, we additionally need the two-independent-processes assumption, which is formalized in Assumption (ref).

assumption(Two Independent Processes) \quad \begin{enumerate}[(ref).a] • (Existence) $\{Y_t\}$ is a process defined on $(\Omega,\mathcal{F},\mathbb{P})$, distributed as $\{X_t\}$; • (Independence) $\{Y_t\}$ is independent of $\{X_t\}$. \end{enumerate}

For the SPL estimator we replace the two-independent-processes assumption by a stationarity and a weak dependence condition, which allows to split our sample into two (asymptotically) independent and identical subsamples. In addition we need an assumption on $T_P$ and $T_E$ as functions of $T$, that is $T_E (T)$ and $T_P (T)$.

assumption(SPL Estimator) \begin{enumerate}[(ref).a] • (Rates) The functions $T_P:\mathbb{N} \to \mathbb{N}$ and $T_E:\mathbb{N} \to \mathbb{N}$ satisfy $T_E (T) < T_P(T)$ for all $T$, while $\frac{T - T_P (T)}{l_T} \rightarrow \infty$ and $m_{T_E (T)}/ m_T \rightarrow 1$ as $T \rightarrow \infty$; • (Strict Stationarity) $\{X_t\}$ is a strictly stationary process; • (Weak Dependence) $\{X_t\}$ satisfies for each $h \in BL$ \begin{align*} \int h \:d \Big(G_{T_E}^{SPL}(\cdot|\mathcal{I}_{T_P:T})- G_{T_E}^{SPL}\Big) \overset{p}{\to}0 \qquad as\qquad T \to \infty, \end{align*} where $G_{T_E}^{SPL}$ denotes the unconditional cdf of $m_{T_E} \big(\hat{\theta}(\mathbf{X}_{1:T_E}) - \theta_0\big)$ and $G_{T_E}^{SPL}(\cdot|\mathcal{I}_{T_P:T})$ the corresponding conditional cdf given $\mathcal{I}_{T_P:T}$. \end{enumerate}

The subsample size assumption in (ref).a ensures that the number of observations used for conditioning is increasing, which along with the negligibility of the initial conditions implies that the truncation of the prediction function is negligible. Furthermore, the sample size used for estimation should increase fast enough that the respective scaling of the 2IP and SPL estimators, $m_{T}$ and $m_{T_E}$ respectively, are asymptotically identical. If $m_T$ increases no faster than a polynomial rate, which is generally the case, it is sufficient that $T_E/T \rightarrow 1$ for $m_{T_E}/m_T \rightarrow 1$ to hold.

The stationarity assumption in (ref).b can actually be relaxed; what matters is that the conditions in Assumption (ref) are still true if only a subsample is considered. In particular, we need that $m_{T_E} \big(\hat{\theta}(\mathbf{X}_{1:T_E}) - \theta_0\big)\overset{d}{\to} G_{\infty}$, which - along with the assumptions on gradient and Hessian - is certainly satisfied under stationarity. However, in general the assumption will be far too strict; here we use it simply to have a clear, interpretable assumption rather than a list of high-level assumptions that are difficult to interpret. The weak dependence condition in (ref).c is met by numerous Markov processes. Intuitively, $(X_1,\dots,X_{T_E})$ and $(X_{T_P},\dots,X_T)$ approach independence as their temporal distance $T_P-T_E$ increases. We illustrate a particular case in the Remark (ref).

remarkSuppose $\{X_t\}$ is strong mixing (cf.\ doukhan1994mixing, doukhan1994mixing) and let $\alpha$ denote the strong mixing coefficient. For $h \in BL$ and for all $\epsilon>0$, Markov's and Ibragimov's inequality (cf.\ hall2014martingale, hall2014martingale, Theorem A.5) imply \begin{align*} &\mathbb{P}\bigg[\bigg|\int h\:d\Big( G_{T_E}^{SPL}(\cdot|\mathcal{I}_{T_P:T})- G_{T_E}^{SPL}\Big)\bigg|\geq \epsilon\bigg]\leq \frac{1}{\epsilon}\mathbb{E}\bigg|\int h\:d\Big( G_{T_E}^{SPL}(\cdot|\mathcal{I}_{T_P:T})- G_{T_E}^{SPL}\Big)\bigg|\\ &=\frac{1}{\epsilon} \mathbb{C}ov\bigg[h\Big(m_{T_E} \big(\hat{\theta}(\mathbf{X}_{1:T_E}) - \theta_0\big)\Big),sign\Big\{\int h\:d\Big( G_{T_E}^{SPL}(\cdot|\mathcal{I}_{T_P:T})- G_{T_E}^{SPL}\Big)\Big \}\bigg]\\ &\leq \frac{4||h||_{BL}}{\epsilon}\alpha(T_P - T_E)\:. \end{align*} Taking $T_P - T_E \to \infty$ such that $\alpha(T_P - T_E)\to 0$ verifies Assumption (ref).c.

Assumptions (ref) to (ref) are met by the AR and GARCH processes considered in Examples (ref) and (ref) (with 1.e holding for bounded sequences). A detailed verification of each assumption under mild conditions is provided in beutner2019technical. We state the following theorem.

theorem{(Merging of 2IP and SPL)} Under Assumptions (ref) to (ref), $F_T^{2IP}(\cdot|\mathcal{I}_{1:T})$ and $F_T^{SPL}(\cdot|\mathcal{I}_{T_P:T})$ merge in probability.

Having established asymptotic closeness between the conditional cdfs $F_T^{2IP}(\cdot|\mathcal{I}_{1:T})$ and $F_T^{SPL}(\cdot|\mathcal{I}_{T_P:T})$, we now turn to the construction of asymptotic intervals.

Interval Construction

Henceforth, for any cdf $F$ we write $F^{-1}$ to denote its generalized inverse given by $F^{-1}(u)=\inf\big\{\tau \in \mathbb{R}:F(\tau)\geq u\big\}$. A confidence interval for $\psi_{T+1}$ based on quantiles of (ref) or (ref) is typically infeasible as these cumulative distribution functions are unknown for finite $T$. Here, they are infeasible because roughly they are the distribution functions of some weights which induce merging multiplied by $m_T \big(\hat{\theta}(\mathbf{X}_{1:T}) - \theta_0\big)$ and $m_{T_E} \big(\hat{\theta}(\mathbf{X}_{1:T_E}) - \theta_0\big)$, respectively, where, in general, the distributions of $m_T \big(\hat{\theta}(\mathbf{X}_{1:T}) - \theta_0\big)$ and $m_{T_E} \big(\hat{\theta}(\mathbf{X}_{1:T_E}) - \theta_0\big)$ are unknown in finite samples. Since these are the only unknown distributions an asymptotic approximation can be based on $G_{\infty}$ with merging induced by the non-convergent weights. In general, we also need to estimate $G_{\infty}$; see Examples (ref) and (ref) below for common approaches. We denote estimators of (ref) and (ref) resulting from this approximation by $\widehat{F_T^{2IP}}(\cdot)$ and $\widehat{F_T^{SPL}}(\cdot)$, respectively. In Section (ref), we provide explicit expressions when $G_{\infty}$ is multivariate normal. For the general construction, we refer to relations (ref) and (ref) in Appendix (ref) and the explanations preceding these relations. Based on the 2IP approach, we consider an interval of the form

align[align omitted — 514 chars of source]

where $\gamma_1,\gamma_2 \in [0,1)$ satisfy $\gamma=\gamma_1+\gamma_2$. We typically take $\gamma_1=\gamma_2=\gamma/2$ such that the interval is equal-tailed. Similarly, we construct the following sample split interval:

align[align omitted — 550 chars of source]

To achieve correct coverage, we need that $\widehat{F_T^{2IP}}(\cdot)$ and $F_T^{2IP}(\cdot|\mathcal{I}_{1:T})$ merge in probability and likewise for SPL. A sufficient condition for this in our setting is that we can consistently estimate the asymptotic distribution of the parameter estimator, $G_{\infty}$, by an appropriate estimator. This is formulated in Assumption (ref) below.

assumption{(CDF Estimator)} Let $\widehat{G}_T (\cdot)$ denote a random ($r$-dimensional) cdf as a function of $\mathbf{X}_{1:T}$, used to estimate $G_{\infty}$. Then $\int h \:d \widehat{G}_T (\cdot) \overset{p}{\to}\int h \:d G_{\infty}$ as $T \to \infty$ for all $h \in BL$.

Although we did not explicitly specify in Assumption (ref) the dependence of $\widehat{G}_T$ on $\mathbf{X}_{1:T}$, it should be understood to hold for any subsample of $\mathbf{X}_{1:T}$ whose size goes to infinity. The verification of Assumption (ref) is a standard step in asymptotic analysis. The two examples below provide common methods for verifying Assumption (ref).

exampleSuppose that $G_{\infty}$ belongs to some parametric family $\{G_{\theta,\xi} | \theta \in \Theta, \xi \in \Xi\}$. Then, given some consistent estimators $\hat{\theta}(\mathbf{X}_{1:T})$ and $\hat{\xi}(\mathbf{X}_{1:T})$ for $\theta_0$ and $\xi_0$ respectively, it follows from the continuous mapping theorem that $\widehat{G}_T = G_{\hat{\theta}(\mathbf{X}_{1:T}), \hat{\xi}(\mathbf{X}_{1:T})}$ satisfies Assumption (ref) if $G_{\theta,\xi}$ is continuous in $\theta$ and $\xi$.
exampleIf $\hat{G}_T$ is based on a consistent bootstrap procedure for $G_{\infty}$ then Assumption (ref) clearly holds.

The following theorem states the intervals' asymptotic validity.

theorem{ (Asymptotic Coverage)} \begin{enumerate} • \begin{enumerate}[(a)] • Under Assumption (ref), (ref) and (ref), $F_T^{2IP}(\cdot|\mathcal{I}_{1:T})$ and $\widehat{F_T^{2IP}}(\cdot)$ merge in probability. • If in addition $\widehat{F_T^{2IP}}(\cdot)$ is stochastically uniformly equicontinuous, then \begin{align} \mathbb{P} \Big[ I_\gamma^{2IP}(\mathbf{x}_{1:T},\mathbf{Y}_{1:T}) \ni \psi_{T+1} \Big|\mathcal{I}_{1:T}\Big]\overset{p}{\to} 1 -\gamma\:. \end{align} \end{enumerate} • \begin{enumerate}[(a)] • Under Assumption (ref), (ref) and (ref), $F_T^{SPL}(\cdot|\mathcal{I}_{T_P:T})$ and $\widehat{F_T^{SPL}}(\cdot)$ merge in probability. • If in addition $\widehat{F_T^{SPL}}(\cdot)$ is stochastically uniformly equicontinuous, then \begin{align} \mathbb{P} \Big[ I_\gamma^{SPL}(\mathbf{x}_{T_P:T}^c,\mathbf{X}_{1:T_E}) \ni \psi_{T+1} \Big|\mathcal{I}_{T_P:T}\Big]\overset{p}{\to} 1 -\gamma\:. \end{align} \end{enumerate} \end{enumerate}

However, the standard approach, motivated by $I_\gamma^{2IP}$ as in (ref), computes an interval of the form $I_\gamma^{STA}(\mathbf{x}_{1:T},\mathbf{x}_{1:T}) = I_\gamma^{2IP}(\mathbf{x}_{1:T},\mathbf{x}_{1:T})$ as only one sample realization is available. This, of course, strongly violates the independence assumption of $\{X_t\}$ and $\{Y_t\}$. Specifically, replacing $\mathbf{Y}_{1:T}$ by $\mathbf{X}_{1:T}$ in equation (ref), leads to

align[align omitted — 516 chars of source]

where $\widehat{F_T^{STA}}(\cdot)$ is defined in relation (ref) and the text preceding it. Whereas it is difficult to justify a conditional confidence interval like $I_{\gamma}^{STA}(\mathbf{x}_{1:T},\mathbf{X}_{1:T})$ directly due to the lack of randomness, we can provide a justification by characterizing how closely the interval resembles the SPL interval. We establish the asymptotic equivalence, defined in terms of location and (scaled) length, between the two intervals in the following theorem. Note that, as our characterization of equivalence is probabilistic, we need to introduce the “doubly random” versions of the STA and SPL estimators, where the sample we condition on is considered random. These estimators are denoted as $\hat\psi_{T+1} (\mathbf{X}_{1:T}, \mathbf{X}_{1:T})$ and $\hat\psi_{T+1} (\mathbf{X}_{T_P:T}^c, \mathbf{X}_{1:T_E})$ respectively.

theorem(Asymptotic Equivalence Confidence Intervals) \begin{enumerate} • (Location) If Assumptions 1-3 hold, then $\hat\psi_{T+1} (\mathbf{X}_{1:T}, \mathbf{X}_{1:T}) - \hat\psi_{T+1} (\mathbf{X}_{T_P:T}^c, \mathbf{X}_{1:T_E}) \overset{p}{\to} 0$. • (Length) Under the assumptions of Theorem (ref) and (ref) and $\widehat{F_T^{SPL}}^{-1}(\cdot)$ being stochastically pointwise continuous at $u =\gamma_1,1-\gamma_2$, we have \begin{align} \widehat{F_T^{STA}}^{-1}(u)-\widehat{F_T^{SPL}}^{-1}(u)\overset{p}{\to}0\:. \end{align} \end{enumerate}

The first implication states that the locations of the two intervals coincide asymptotically. The second statement establishes asymptotic closeness of the selected quantiles such that the scaled lengths of the intervals in (ref) and (ref) coincide asymptotically. As such, our sample-split interval coincides asymptotically with the standard interval, meaning that the standard interval can be substituted for an (asymptotically) equivalent interval which has a formal justification in terms of conditional coverage. As such, this provides a justification for the intervals commonly constructed in practice without having to rely on the two-independent-processes assumption.

Interval Construction Under Normality

In this subsection we present intervals of reduced form and simplified theoretical results under asymptotic normality of the parameter estimator.

assumption{(Normality)} Let $G_{\infty}$ be the cdf of the $N(0,\Upsilon_0)$ distribution with $\Upsilon_0=\Upsilon(\theta_0,\xi_0)$ and assume there exist $\hat{\Upsilon}(\mathbf{X}_{1:T})$ converging in probability to $\Upsilon_0$.

Usually, the covariance estimator is obtained by inserting consistent estimators for $\theta_0$ and $\xi_0$ into $\Upsilon_0$. Following the plug-in principle, we estimate $F_T^{2IP}(\cdot|\mathcal{I}_{1:T})$ by a normal distribution with mean $0$ and variance $\hat{\upsilon}_T^{2IP}=\frac{\partial \psi_{T+1}(\mathbf{x}_{1:T};\hat{\theta}(\mathbf{Y}_{1:T}))}{\partial \theta'}\hat{\Upsilon}(\mathbf{Y}_{1:T})\frac{\partial \psi_{T+1}(x_{1:T};\hat{\theta}(\mathbf{Y}_{1:T}))}{\partial \theta}$ such that $\widehat{F_T^{2IP}}(\cdot)=\Phi\big(\cdot/\sqrt{\hat{\upsilon}_T^{2IP}}\big)$. Then, the interval in (ref) simplifies to

align[align omitted — 281 chars of source]

Similarly, for the sample-split approach we consider $\widehat{F_T^{SPL}}(\cdot)= \Phi\big(\cdot / \sqrt{\hat{\upsilon}_T^{SPL}}\big)$ with $\hat{\upsilon}_T^{SPL}=\frac{\partial \psi_{T+1}(\mathbf{x}_{T_P:T}^{c};\hat{\theta}(\mathbf{X}_{1:T_E}))}{\partial \theta'}\hat{\Upsilon}(\mathbf{X}_{1:T_E})\frac{\partial \psi_{T+1}(\mathbf{x}_{T_P:T}^{c};\hat{\theta}(\mathbf{X}_{1:T_E}))}{\partial \theta'}$ such that (ref) reduces to

align[align omitted — 314 chars of source]

In Appendix B we show that if the variance estimator is bounded away from zero in probability, e.g. $1/\hat{\upsilon}_T^{2IP}=O_p(1)$, then $\widehat{F_T^{2IP}}(\cdot)$ is stochastically uniform equicontinuous. Therefore, the asymptotic validity of both intervals can be deduced from Theorem (ref).

corollary{(Asymptotic Coverage under Normality)} \begin{enumerate} • \begin{enumerate}[(a)] • Under Assumption (ref), (ref) and (ref), $F_T^{2IP}(\cdot|\mathcal{I}_{1:T})$ and $\Phi\big(\cdot/ \sqrt{\hat{\upsilon}_T^{2IP}}\big)$ merge in probability. • If in addition $1/\hat{\upsilon}_T^{2IP}=O_p(1)$, $\mathbb{P} \Big[ I_\gamma^{2IP}(\mathbf{x}_{1:T},\mathbf{Y}_{1:T}) \ni \psi_{T+1} \Big|\mathcal{I}_{1:T}\Big]\overset{p}{\to} 1 -\gamma$. \end{enumerate} • \begin{enumerate}[(a)] • Under Assumption (ref), (ref) and (ref), $F_T^{SPL}(\cdot|\mathcal{I}_{T_P:T})$ and $\Phi\big(\cdot/\sqrt{\hat{\upsilon}_T^{SPL}}\big)$ merge in probability. • If in addition $1/\hat{\upsilon}_T^{SPL}=O_p(1)$, $\mathbb{P} \Big[ I_\gamma^{SPL}(\mathbf{x}_{T_P:T}^{c},\mathbf{X}_{1:T_E}) \ni \psi_{T+1} \Big|\mathcal{I}_{T_P:T}\Big]\overset{p}{\to} 1 -\gamma$. \end{enumerate} \end{enumerate}

Bounding the variance estimator away from zero in probability to establish that the conditional coverage probability converges to $1-\gamma$ in probability has intuitive appeal: as $\hat{\upsilon}_T^{2IP}$ approaches zero, $N\big(0,\hat{\upsilon}_T^{2IP}\big)$ becomes degenerate while the interval in (ref) collapses (similar for SPL).

For the standard interval, replacing $\mathbf{Y}_{1:T}$ by $\mathbf{X}_{1:T}$ in (ref) leads to

align[align omitted — 290 chars of source]

with $\hat{\upsilon}_{T}^{STA}=\frac{\partial \psi_{T+1}(x_{1:T};\hat{\theta}(\mathbf{X}_{1:T}))}{\partial \theta'}\hat{\Upsilon}(\mathbf{X}_n)\frac{\partial \psi_{T+1}(\mathbf{x}_{1:T};\hat{\theta}(\mathbf{X}_{1:T}))}{\partial \theta}$. In Appendix B we prove that $\hat{\upsilon}_T^{SPL}$ is bounded in probability, which in turn implies that the quantile function $\widehat{F_T^{SPL}}^{-1}(\cdot)=\sqrt{\hat{\upsilon}_T^{SPL}}\Phi^{-1}(\cdot)$ is stochastically pointwise equicontinuous at any $u \in \mathbb{R}$. Hence, Theorem (ref) applies. Whereas the first statement of the theorem remains unaffected, its second statement reads as follows under normality.

corollary{(Length under Normality)} Under the assumptions of Theorem (ref) and Corollary (ref), we have $\sqrt{\hat{\upsilon}_{T}^{STA}}\Phi^{-1}(u)-\sqrt{\hat{\upsilon}_T^{SPL}}\Phi^{-1}(u)\overset{p}{\to}0$ for $u =\gamma_1,1-\gamma_2$.

Prediction Intervals

The preceding sections have focused purely on the construction of conditional confidence intervals to account for parameter uncertainty. Regarding prediction, a second source of uncertainty arises, that corresponds to the model's innovation process. In this setting, parameter estimation is typically disregarded in textbooks as the stochastic fluctuation stemming from the estimation procedure is generally dominated by the stochastic fluctuation of the innovations. Although the resulting prediction intervals may be asymptotically valid, they are typically characterized by under-coverage in finite samples. In response, pan2016abootstrap introduce the concept of asymptotic pertinence to evaluate distribution approximations that account for the two sources of randomness, innovation and parameter estimation uncertainty, according to their general orders of magnitude. Whereas kunitomo1985properties and samaranayake1988properties study properties of the unconditional law of the forecast error, we focus on its conditional distribution to conform with the relevance property of kabaila1999relevance. The fundamental issue also arises when considering prediction if one attempts to account for parameter uncertainty. To illustrate this point, we revisit the introductory examples and write $*$ to denote the convolution operator.\footnote{For independent variables $X$ and $Y$ with $X\sim F_X$, $Y\sim F_Y$ and $Z=X+Y \sim F_{Z}$, we write $F_{Z}=F_X * F_Y$ to denote $F_{Z}(z) = \int_{-\infty}^zF_X(z-y)dF_Y(y)$.}

Example(ref). (continued) Prediction intervals for the AR are often constructed around the point estimate for the conditional mean. The conditional distribution of the forecast error decomposes into \begin{align} \begin{split} \mathbb{P}\big[X_{T+1}-\hat{\beta}(\mathbf{X}_{1:T})X_T\leq \cdot|X_T=x_T\big] =&\mathbb{P}\big[\beta X_T-\hat{\beta}(\mathbf{X}_{1:T})X_T\leq \cdot|X_T=x_T\big]\\ &*\mathbb{P}\big[\varepsilon_{T+1}\leq \cdot\big]\:, \end{split} \end{align} corresponding to estimation and innovation uncertainty, respectively. As argued above, an approximation of $ \mathbb{P}\big[\beta X_T-\hat{\beta}(\mathbf{X}_{1:T})X_T\leq \cdot|X_T=x_T\big]$ appears rather cumbersome. In the special case of $\varepsilon_t\overset{iid}{\sim}N\big(0,\sigma^2_\varepsilon\big)$, phillips1979sampling (phillips1979sampling, Thm.\ 3) derives an approximation for (ref) based on Edgeworth expansions.
Example(ref). (continued) Suppose we are interested in providing a prediction interval for $X_{T+1}^2$ in the GARCH(1,1). Conditioning on $\mathbf{X}_{1:T}=\mathbf{x}_{1:T}$, a natural estimate of $X_{T+1}^2$ is $\hat{\sigma}_{T+1|1:T}^{2\;STA}$ as defined in (ref), since $\sigma_{T+1|1:T}^2$ is its expected value given information up to time $T$. As \begin{align} \begin{split} \mathbb{P}\big[X_{T+1}^2-\hat{\sigma}_{T+1}^{2\;STA*}\leq \cdot|\mathbf{X}_{1:T}=\mathbf{x}_{1:T}\big] =&\mathbb{P}\big[\sigma_{T+1}^2-\hat{\sigma}_{T+1}^{2\;STA*}\leq \cdot|\mathbf{X}_{1:T}=\mathbf{x}_{1:T}\big]\\ &*\mathbb{P}\big[\sigma_{T+1|1:T}^2(\varepsilon_{T+1}^2-1)\leq \cdot\big]\:, \end{split} \end{align} where $\hat{\sigma}_{T+1}^{2\;STA*}$ is defined in (ref), the desired prediction interval, say $J_\gamma^{STA}$, leads to a sensible probabilistic statement due to variability in $\varepsilon_{T+1}^2$: \begin{align} \mathbb{P}\Big[ X_{T+1}^2 \in J_\gamma^{STA}(\mathbf{X}_{1:T},\mathbf{X}_{1:T}) \Big|\mathbf{X}_{1:T}=\mathbf{x}_{1:T}\Big] \underset{(\approx)}{=} 1 -\gamma\:. \end{align} However, it cannot incorporate parameter uncertainty either, since the conditional distribution $\mathbb{P}\big[\sigma_{T+1}^2-\hat{\sigma}_{T+1}^{2\;STA*}\leq \cdot|\mathbf{X}_{1:T}=\mathbf{x}_{1:T}\big]=\mathbb{P}\big[\sigma_{T+1|1:T}^2-\hat{\sigma}_{T+1|1:T}^{2\;STA}\leq \cdot\big]$ is degenerate.

In his textbook pesaran2015time resorts to a Bayesian-akin approach to avoid the fundamental issue in forecasting. He argues that $\theta$, although “fixed at the estimation stage ... is viewed best as a random variable at the forecasting stage" (p.\ 389). Consequently, he assigns some posterior distribution to $\theta$ motivated by an uninformed prior. Treating $\theta$ not fixed but random, the fundamental issue does not arise, however combining a frequentist view with a Bayesian method does not seem to be coherent.

barndorff1996prediction require the existence of a transitive statistic $U=U(\mathbf{X}_{1:T})$ of fixed low dimension to establish conditional independence between the sample $\mathbf{X}_{1:T}$ and their considered future random variable given $U=u$. vidoni2004improved (vidoni2004improved, vidoni2009improved, vidoni2009simple, vidoni2016improved), kabaila1999relevance, kabaila2004adjustment, and kabaila2008improved (kabaila2008improved, kabaila2010asymptotic) extend their approach and derive improved prediction intervals. Although these methods absorb an additional $O(T^{-1})$ term in the associated conditional coverage probability, there are several drawbacks associated with them: the innovation distribution needs typically be specified (e.g.\ Gaussian), the results apply only to a limited set of estimators (e.g.\ maximum likelihood) and their framework can only incorporate finite autoregressive components (e.g.\ AR$(p)$).

Assuming two independent processes with the same stochastic structure, using one for prediction and one for the estimation of the parameters, alleviates the fundamental issue faced in the continued Examples (ref) and (ref). As the conditional distributions of the 2IP and SPL estimators merge in probability by Theorem (ref), the 2IP assumption can be avoided by following a sample-split approach as described in Section (ref).

Conclusion

In the paper at hand, we study the construction of confidence intervals for conditional objects such as conditional means or conditional variances, focusing on the conceptual issue that arises in the process of taking parameter uncertainty into account. It stems from the fact that on one hand one must condition on the sample as the past informs about the present and future, yet on the other hand one must allow the data up to now to be treated as random to account for estimation uncertainty. Assuming two independent processes with the same stochastic structure, where one is used for conditioning and one for the estimation of the parameters, bypasses this issue, but the assumption itself can generally not be justified in applications. To avoid this assumption, we propose a solution based on a simple sample-split approach, that requires a much more realistic weak dependence condition instead. To acknowledge that the conditional quantities vary over time, we employ a merging concept generalizing the notion of weak convergence. The conditional distributions of the sample-split estimator and the estimator based on the two-independent-processes assumption are shown to merge in probability under mild conditions. The corresponding sample-split intervals are shown to coincide asymptotically with the intervals commonly constructed by practitioners, which provides a novel and theoretically satisfactory justification for commonly constructed confidence intervals for conditional objects, applicable to a wide class of time series models, including ARMA and GARCH-type models.

One limitation to our approach is that we restrict ourselves to univariate time series and objects of interests. At the expense of more involved notation this could be readily extended to multivariate time series and objects of interests. A second, and more restrictive, limitation is our weak dependence assumption needed to achieve asymptotic independence between the two subsamples, which for instance rules out application to integrated processes. Given the fundamental role of this assumption in our setup, it appears difficult to generalize this. However, this also casts further doubt on the two-independent-processes assumption as validation for confidence intervals constructed for such persistent processes. A case-by-case treatment, as for instance done by gospodinov2002median for near unit root processes and samaranayake1988properties for explosive processes, appears to be necessary in such cases, and standard confidence intervals should be treated with caution.

thebibliography\bibitem[\citeauthoryear{Akaike}{Akaike}{1969}]{akaike1969fitting} Akaike, H. (1969). \newblock Fitting autoregressive models for prediction. \newblock {\em Annals of the Institute of Statistical Mathematics\/} {\em 21\/}(1), 243--247. \bibitem[\citeauthoryear{Barndorff-Nielsen and Cox}{Barndorff-Nielsen and Cox}{1996}]{barndorff1996prediction} Barndorff-Nielsen, O. E. and D. R. Cox (1996). \newblock Prediction and asymptotics. \newblock {\em Bernoulli\/} {\em 2\/}(4), 319--340. \bibitem[\citeauthoryear{Belyaev and Sj\"{o}stedt-De Luna}{Belyaev and Sj\"{o}stedt-De Luna}{2000}]{belyaev2000weakly} Belyaev, Y. and S. Sj\"{o}stedt-De Luna (2000). \newblock Weakly approaching sequences of random distributions. \newblock {\em Journal of Applied Probability\/} {\em 37\/}(3), 807--822. \bibitem[\citeauthoryear{Beutner, Heinemann, and Smeekes}{Beutner et al.}{2017a}]{beutner2017justification} Beutner, E., A. Heinemann, and S. Smeekes (2017a). \newblock A justification of conditional confidence intervals. \newblock GSBE Research Memorandum RM/17/023, Maastricht University. \bibitem[\citeauthoryear{Beutner, Heinemann, and Smeekes}{Beutner et al.}{2017b}]{beutner2017technical} Beutner, E., A. Heinemann, and S. Smeekes (2017b). \newblock Technical report on “a justification of conditional confidence intervals”. \newblock Technical report, Maastricht University, \url{http://researchers-sbe.unimaas.nl/stephansmeekes/research/}. \bibitem[\citeauthoryear{Beutner, Heinemann, and Smeekes}{Beutner et al.}{2019}]{beutner2019technical} Beutner, E., A. Heinemann, and S. Smeekes (2019). \newblock A general framework for prediction in time series models. \newblock Working paper, Maastricht University, \url{http://researchers-sbe.unimaas.nl/stephansmeekes/research/}. \bibitem[\citeauthoryear{Billingsley}{Billingsley}{1986}]{billingsley1986probability} Billingsley, P. (1986). \newblock {\em Probability and Measure}. \newblock New York: John Wiley & Sons. \bibitem[\citeauthoryear{Blasques, Koopman, {\L}asak, and Lucas}{Blasques et al.}{2016}]{blasques2016sample} Blasques, F., S. J. Koopman, K. {\L}asak, and A. Lucas (2016). \newblock In-sample confidence bands and out-of-sample forecast bands for time-varying parameters in observation-driven models. \newblock {\em International Journal of Forecasting\/} {\em 32\/}(3), 875--887. \bibitem[\citeauthoryear{Bollerslev}{Bollerslev}{1986}]{bollerslev1986generalized} Bollerslev, T. (1986). \newblock Generalized autoregressive conditional heteroskedasticity. \newblock {\em Journal of Econometrics\/} {\em 31\/}(3), 307--327. \bibitem[\citeauthoryear{Boussama, Fuchs, and Stelzer}{Boussama et al.}{2011}]{boussama2011stationarity} Boussama, F., F. Fuchs, and R. Stelzer (2011). \newblock Stationarity and geometric ergodicity of BEKK multivariate Garch models. \newblock {\em Stochastic Processes and Their Applications\/} {\em 121\/}(10), 2331--2360. \bibitem[\citeauthoryear{Castillo and Rousseau}{Castillo and Rousseau}{2015}]{CastilloAoS15} Castillo, I. and J. Rousseau (2015). \newblock A Bernstein–von Mises theorem for smooth functionals in semiparametric models. \newblock {\em Annals of Statistics\/} {\em 43\/}(6), 2353--2383. \bibitem[\citeauthoryear{Cavaliere, Georgiev, and Taylor}{Cavaliere et al.}{2013}]{cavaliere2013wild} Cavaliere, G., I. Georgiev, and A. M. R. Taylor (2013). \newblock Wild bootstrap of the sample mean in the infinite variance case. \newblock {\em Econometric Reviews\/} {\em 32\/}(2), 204--219. \bibitem[\citeauthoryear{D'Aristotile, Diaconis, and Freedman}{D'Aristotile et al.}{1988}]{d1988merging} D'Aristotile, A., P. Diaconis, and D. Freedman (1988). \newblock On merging of probabilities. \newblock {\em Sankhy{\=a}: The Indian Journal of Statistics, Series A\/} {\em 50\/}(3), 363--380. \bibitem[\citeauthoryear{Davydov and Rotar}{Davydov and Rotar}{2009}]{davydov2009asymptotic} Davydov, Y. and V. Rotar (2009). \newblock On asymptotic proximity of distributions. \newblock {\em Journal of Theoretical Probability\/} {\em 22\/}(1), 82--98. \bibitem[\citeauthoryear{Doukhan}{Doukhan}{1994}]{doukhan1994mixing} Doukhan, P. (1994). \newblock {\em Mixing: Properties and Examples}. \newblock New York: Springer. \bibitem[\citeauthoryear{Dudley}{Dudley}{2002}]{dudley2002real} Dudley, R. M. (2002). \newblock {\em Real Analysis and Probability}. \newblock Cambridge: Cambridge University Press. \bibitem[\citeauthoryear{Dufour and Taamouti}{Dufour and Taamouti}{2010}]{dufour2010short} Dufour, J.-M. and A. Taamouti (2010). \newblock Short and long run causality measures: theory and inference. \newblock {\em Journal of Econometrics\/} {\em 154\/}(1), 42--58. \bibitem[\citeauthoryear{Engle}{Engle}{1982}]{engle1982autoregressive} Engle, R. F. (1982). \newblock Autoregressive conditional heteroscedasticity with estimates of the variance of United Kingdom inflation. \newblock {\em Econometrica\/} {\em 50\/}(4), 987--1007. \bibitem[\citeauthoryear{Francq and Zako{\"\i}an}{Francq and Zako{\"\i}an}{2015}]{francq2015risk} Francq, C. and J.-M. Zako{\"\i}an (2015). \newblock Risk-parameter estimation in volatility models. \newblock {\em Journal of Econometrics\/} {\em 184\/}(1), 158--173. \bibitem[\citeauthoryear{Gospodinov}{Gospodinov}{2002}]{gospodinov2002median} Gospodinov, N. (2002). \newblock Median unbiased forecasts for highly persistent autoregressive processes. \newblock {\em Journal of Econometrics\/} {\em 111\/}(1), 85--101. \bibitem[\citeauthoryear{Hall and Heyde}{Hall and Heyde}{1980}]{hall2014martingale} Hall, P. and C. C. Heyde (1980). \newblock {\em Martingale Limit Theory and Its Application}. \newblock London: Academic Press. \bibitem[\citeauthoryear{Hamilton}{Hamilton}{1994}]{hamilton1994time} Hamilton, J. D. (1994). \newblock {\em Time Series Analysis}. \newblock Princeton: Princeton University Press. \bibitem[\citeauthoryear{Hansen}{Hansen}{2006}]{hansen2006interval} Hansen, B. E. (2006). \newblock Interval forecasts and parameter uncertainty. \newblock {\em Journal of Econometrics\/} {\em 135\/}(1), 377--398. \bibitem[\citeauthoryear{Huber}{Huber}{2009}]{huber2009robust} Huber, P. J. (2009). \newblock {\em Robust Statistics\/} (2nd ed.). \newblock New York: John Wiley & Sons. \bibitem[\citeauthoryear{Ing and Wei}{Ing and Wei}{2003}]{ing2003same} Ing, C.-K. and C.-Z. Wei (2003). \newblock On same-realization prediction in an infinite-order autoregressive process. \newblock {\em Journal of Multivariate Analysis\/} {\em 85\/}(1), 130--155. \bibitem[\citeauthoryear{Ing and Wei}{Ing and Wei}{2005}]{ing2005order} Ing, C.-K. and C.-Z. Wei (2005). \newblock Order selection for same-realization predictions in autoregressive processes. \newblock {\em The Annals of Statistics\/} {\em 33\/}(5), 2423--2474. \bibitem[\citeauthoryear{Kabaila}{Kabaila}{1999}]{kabaila1999relevance} Kabaila, P. (1999). \newblock The relevance property for prediction intervals. \newblock {\em Journal of Time Series Analysis\/} {\em 20\/}(6), 655--662. \bibitem[\citeauthoryear{Kabaila and He}{Kabaila and He}{2004}]{kabaila2004adjustment} Kabaila, P. and Z. He (2004). \newblock The adjustment of prediction intervals to account for errors in parameter estimation. \newblock {\em Journal of Time Series Analysis\/} {\em 25\/}(3), 351--358. \bibitem[\citeauthoryear{Kabaila and Syuhada}{Kabaila and Syuhada}{2008}]{kabaila2008improved} Kabaila, P. and K. Syuhada (2008). \newblock Improved prediction limits for \uppercase{AR}($p$) and \uppercase{ARCH}($p$) processes. \newblock {\em Journal of Time Series Analysis\/} {\em 29\/}(2), 213--223. \bibitem[\citeauthoryear{Kabaila and Syuhada}{Kabaila and Syuhada}{2010}]{kabaila2010asymptotic} Kabaila, P. and K. Syuhada (2010). \newblock The asymptotic efficiency of improved prediction intervals. \newblock {\em Statistics & Probability Letters\/} {\em 80\/}(17), 1348--1353. \bibitem[\citeauthoryear{Kreiss}{Kreiss}{2016}]{kreiss2015discussion} Kreiss, J.-P. (2016). \newblock Discussion: bootstrap prediction intervals for linear, nonlinear and nonparametric autoregressions. \newblock {\em Journal of Statistical Planning and Inference\/} {\em 177}, 28--30. \bibitem[\citeauthoryear{Kunitomo and Yamamoto}{Kunitomo and Yamamoto}{1985}]{kunitomo1985properties} Kunitomo, N. and T. Yamamoto (1985). \newblock Properties of predictors in misspecified autoregressive time series models. \newblock {\em Journal of the American Statistical Association\/} {\em 80\/}(392), 941--950. \bibitem[\citeauthoryear{Lewis and Reinsel}{Lewis and Reinsel}{1985}]{lewis1985prediction} Lewis, R. and G. C. Reinsel (1985). \newblock Prediction of multivariate time series by autoregressive model fitting. \newblock {\em Journal of Multivariate Analysis\/} {\em 16\/}(3), 393--411. \bibitem[\citeauthoryear{L{\"u}tkepohl}{L{\"u}tkepohl}{2005}]{lutkepohl2005new} L{\"u}tkepohl, H. (2005). \newblock {\em New Introduction to Multiple Time Series Analysis}. \newblock Berlin: Springer. \bibitem[\citeauthoryear{Pan and Politis}{Pan and Politis}{2016a}]{pan2016abootstrap} Pan, L. and D. N. Politis (2016a). \newblock Bootstrap prediction intervals for linear, nonlinear and nonparametric autoregressions. \newblock {\em Journal of Statistical Planning and Inference\/} {\em 177}, 1--27. \bibitem[\citeauthoryear{Pan and Politis}{Pan and Politis}{2016b}]{pan2016bbootstrap} Pan, L. and D. N. Politis (2016b). \newblock Bootstrap prediction intervals for \uppercase{M}arkov processes. \newblock {\em Computational Statistics & Data Analysis\/} {\em 100}, 467--494. \bibitem[\citeauthoryear{Pascual, Romo, and Ruiz}{Pascual et al.}{2004}]{pascual2004bootstrap} Pascual, L., J. Romo, and E. Ruiz (2004). \newblock Bootstrap predictive inference for \uppercase{ARIMA} processes. \newblock {\em Journal of Time Series Analysis\/} {\em 25\/}(4), 449--465. \bibitem[\citeauthoryear{Pascual, Romo, and Ruiz}{Pascual et al.}{2006}]{pascual2006bootstrap} Pascual, L., J. Romo, and E. Ruiz (2006). \newblock Bootstrap prediction for returns and volatilities in \uppercase{garch} models. \newblock {\em Computational Statistics & Data Analysis\/} {\em 50\/}(9), 2293--2312. \bibitem[\citeauthoryear{Pesaran}{Pesaran}{2015}]{pesaran2015time} Pesaran, M. H. (2015). \newblock {\em Time Series and Panel Data Econometrics}. \newblock Oxford: Oxford University Press. \bibitem[\citeauthoryear{Phillips}{Phillips}{1979}]{phillips1979sampling} Phillips, P. C. B. (1979). \newblock The sampling distribution of forecasts from a first-order autoregression. \newblock {\em Journal of Econometrics\/} {\em 9\/}(3), 241--261. \bibitem[\citeauthoryear{Samaranayake and Hasza}{Samaranayake and Hasza}{1988}]{samaranayake1988properties} Samaranayake, V. A. and D. P. Hasza (1988). \newblock Properties of predictors for multivariate autoregressive models with estimated parameters. \newblock {\em Journal of Time Series Analysis\/} {\em 9\/}(4), 361--383. \bibitem[\citeauthoryear{van der Vaart}{van der Vaart}{2000}]{van2000asymptotic} van der Vaart, A. W. (2000). \newblock {\em Asymptotic Statistics}. \newblock Cambridge: Cambridge University Press. \bibitem[\citeauthoryear{Vidoni}{Vidoni}{2004}]{vidoni2004improved} Vidoni, P. (2004). \newblock Improved prediction intervals for stochastic process models. \newblock {\em Journal of Time Series Analysis\/} {\em 25\/}(1), 137--154. \bibitem[\citeauthoryear{Vidoni}{Vidoni}{2009a}]{vidoni2009improved} Vidoni, P. (2009a). \newblock Improved prediction intervals and distribution functions. \newblock {\em Scandinavian Journal of Statistics\/} {\em 36\/}(4), 735--748. \bibitem[\citeauthoryear{Vidoni}{Vidoni}{2009b}]{vidoni2009simple} Vidoni, P. (2009b). \newblock A simple procedure for computing improved prediction intervals for autoregressive models. \newblock {\em Journal of Time Series Analysis\/} {\em 30\/}(6), 577--590. \bibitem[\citeauthoryear{Vidoni}{Vidoni}{2017}]{vidoni2016improved} Vidoni, P. (2017). \newblock Improved multivariate prediction regions for \uppercase{M}arkov process models. \newblock {\em Statistical Methods & Applications\/} {\em 26\/}(1), 1--18. \bibitem[\citeauthoryear{Xiong and Li}{Xiong and Li}{2008}]{xiong2008some} Xiong, S. and G. Li (2008). \newblock Some results on the convergence of conditional distributions. \newblock {\em Statistics & Probability Letters\/} {\em 78\/}(18), 3249--3253. \bibitem[\citeauthoryear{Zako{\"\i}an}{Zako{\"\i}an}{1994}]{zakoian1994threshold} Zako{\"\i}an, J.-M. (1994). \newblock Threshold heteroskedastic models. \newblock {\em Journal of Economic Dynamics and Control\/} {\em 18\/}(5), 931--955.