Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
88,704 characters · 13 sections · 69 citation commands
A Justification of Conditional Confidence Intervals
\doublespacing
One of the open questions in time series is how to quantify uncertainty around point estimates of conditional objects such as conditional means or conditional variances. A fundamental issue arises in the construction of confidence intervals that ought to capture the parameter estimation uncertainty contained in these objects. This fundamental issue stems from the fact that on one hand one must condition on the sample as the past informs about the present, yet on the other hand one must allow the data up to now to be treated as random to account for estimation uncertainty. The issue is well-recognized in the econometric literature, however in practice confidence intervals are commonly constructed by treating the sample simultaneously as fixed and random. Frequently, such approach is motivated by presuming to have two independent processes. Assuming two independent processes with the same stochastic structure, using one for conditioning and one for the estimation of the parameters, bypasses the issue. It is a mathematically convenient assumption as in such case the uncertainty quantification reduces to an ordinary inferential problem. However, practitioners rarely have a replicate, independent of the original series, at hand with the exception of perhaps some experimental settings. As such, the intervals commonly constructed by practitioners lack a satisfactory theoretical justification. Therefore it is the objective of the present paper to develop a realistic justification for such confidence intervals around point estimates of conditional objects.
In the literature the fundamental issue described above is encountered in various ways. In the specific case of a first-order autoregressive (AR) process with Gaussian innovations, phillips1979sampling investigates the statistical dependence between the ordinary least squares (OLS) estimator and the endogenous variable conditioned upon. He obtains an Edgeworth-type expansion for the distribution of the conditional mean and, further, studies forecasting, where the fundamental issue equally arises.\footnote{For prediction intervals some solutions have been discussed. We refer to Section (ref).} lutkepohl2005new (lutkepohl2005new, p.\ 95) explicitly states a two-independent-processes assumption in connection with vector AR models. He postulates that such assumption is asymptotically equivalent to using only data not conditioned upon for estimation. ing2003same clearly distinguish between independent-realization and same-realization settings and study the (unconditional) mean-squared prediction error in the latter for an infinite-order AR process. In a companion paper they also provide a theoretical verification for order selection criteria for same-realization predictions and stress that it can be misleading to assume that the results for independent-realization settings carry over to those for corresponding same-realization cases ing2005order. Other studies investigate parameter uncertainty by using resampling methods, that typically mimic a distribution in which the sample, or at least a subsample, is treated as fixed and random at the same time (cf.\ pascual2004bootstrap, pascual2004bootstrap, pascual2006bootstrap, pan2016abootstrap, pan2016abootstrap, pan2016bbootstrap). Aware of this paradox, kreiss2015discussion points out that conditioning on observing specific in-sample values affects the parameter estimator, but the effect is often erroneously disregarded. Deviating from the various bootstrap approaches, hansen2006interval examines parameter uncertainty in interval forecasts in a classical statistical framework. Similar to a general regression framework, he conditions on an arbitrary fixed out-of-sample value to avoid the issue. However, conditioning on arbitrary fixed out-of-sample values appears incompatible with the usual setup of dynamics in which we condition on the final value(s) of the sample. Acknowledging the issue while avoiding the two independent processes argument bears careful statements as in francq2015risk who write in view of this issue “the delta method ... suggests" (p.\ 162). Similarly, pesaran2015time notices that although such intervals “have been discussed in the econometrics literature, the particular assumptions that underlie them are not fully recognized" (p.\ 389).
This paper provides a novel, and realistic, justification for commonly constructed confidence intervals around point estimates of conditional objects. Our solution is based on a simple sample-split approach and a weak dependence condition, which allows to partition our sample into two asymptotically independent subsamples. For a rich class of time series models we construct asymptotically valid sample-split intervals, without relying on the assumption of observing two independent processes, and show that these intervals coincide asymptotically with the intervals commonly constructed by practitioners. As will be argued below, an appropriate concept to study conditional confidence intervals is merging, a concept that generalizes weak convergence. To the best of our knowledge, except for belyaev2000weakly, this paper is the only one to study merging in the context of conditional distributions. Moreover, our paper seems to be the first to employ merging of conditional distributions in time series. By employing this concept we avoid unnatural assumptions such as observing $X_T=x$ (in dynamic models), losing the time index $T$, and instead explicitly acknowledge that the conditional objects vary over time.
The rest of the paper is organized as follows. Section (ref) specifies the general setup and describes the argument of two independent processes as well as our sample-split approach. In Section (ref) we establish merging among the proposed and the two-independent-processes estimator in probability under mild conditions. Further, we construct asymptotically valid sample-split intervals and show that these coincide asymptotically with the standard intervals. The extension to prediction is discussed in Section (ref). Section (ref) concludes. The main proofs are collected in Appendix A, while Appendix B provides additional proofs of intermediate results.
Let $\{X_t\}$ be a real-valued stochastic process defined on the probability space $(\Omega,\mathcal{F},\mathbb{P})$. $\theta$ denotes a generic parameter vector of length $r \in \mathbb{N}$ and $\theta_0$ the true value, unknown to the researcher.\footnote{Generally, in particular throughout Section (ref), we do not distinguish between $\theta$ and $\theta_0$ if there is no cause for confusion. In Section (ref) we explicitly use $\theta_0$ to avoid confusion.} Let $\Theta \subseteq \mathbb{R}^r$ be the corresponding parameter space.
Our general setup involves inference on an object that we call the prediction function, which is a function of both the process $\{X_t\}$ and of the parameter $\theta$. It represents the random object of interest, and will typically express quantities such as a conditional mean or conditional variance (without conditioning on a specific value) as a function of the sample.
With this setup we can describe most of the possible applications of interest. We now provide three examples to illustrate the prediction function.
A more precarious example, due to its large popularity, is the conditional variance in a generalized autoregressive conditional heteroskedasticity (GARCH) model engle1982autoregressive,bollerslev1986generalized. Whereas in the previous AR($1$) case it suffices to condition on the terminal observation, the subsequent Example 2 is more extreme as the entire sample contains information about the object of interest.
The next example shows that for a large class of models the prediction function can be written in the form of Definition (ref).
Note that in many cases, such as the GARCH(1,1) of Example (ref), the prediction function actually depends on the infinite past of the series. In order to express (an approximation of) the prediction function in terms of observable variables only, we would need to replace $X_t$ by $s_t$ for all $t<1$, where $\{s_t\}$ is a sequence of (arbitrary) constants to which we refer as starting or initial values. For a fixed $T$, we accordingly define an approximate prediction function $\psi_{T+1}^s: \mathbb{R}^T \times \Theta \rightarrow \mathbb{R}$ that is only a function of observable variables as
where $\mathbf{X}_{1:T} = (X_1, \ldots, X_T)^\prime$. Note that, given the varying input of the left-hand side in ((ref)), we now actually have a sequence of (varying) functions for $T \in \mathbb{N}$.
In many cases the values far in the past are negligible for a wide range of values for $\{s_t\}$. Consequently, $\psi_{T+1}^s$ will be close to $\psi_{T+1}$. This property can be shown to hold for many different processes including the ones in the examples. We formalize the exact condition we need regarding the negligibility of the starting values in Assumption (ref).b.
Although the prediction function typically represents a conditional object, we have not conditioned on anything yet in the definition. We therefore now extend the analysis by formally conditioning on observing a particular sample. Let $\mathbf{x}_{1:T} = (x_1,\dots,x_T)'$ denote a specific sample path of $\mathbf{X}_{1:T}$. Throughout the paper, we will discriminate between random variables and their realized counterparts by writing the former in capital and the latter in lowercase letters to avoid ambiguity.
As we will consider sample splitting later on, we define notation that also allows for conditioning on only a subsample. For that purpose, let $t_1:t_2$ denote the (sub-)period from $t_1$ up to $t_2$, and correspondingly $\mathbf{X}_{t_1:t_2} = (X_{t_1}, \ldots, X_{t_2})^\prime$ for any integers $1 \leq t_1 \leq t_2 \leq T$, with a corresponding definition for the observed subsample $\mathbf{x}_{t_1: t_2}$. Furthermore, let $\mathbf{X}_{t_1:T}^{c} = (c_1, \ldots, c_{t_1-1}, X_{t_1}, \ldots, X_{T})^\prime$ denote the vector where all non-considered subsamples are replaced by a sequence of constants $\{c_t\}$, in a similar way as we did for the starting values. We can now formally define the conditional prediction function.
Note that we phrase the conditional prediction function directly in terms of the approximate prediction function $\psi_{T+1}^s$ rather than the true prediction function. We take this “shortcut” because we cannot observe $x_{0}, x_{-1}, \ldots$, so we cannot condition on those values anyway. Therefore, the “true” conditional object (which we might represent as $\psi_{T+1| -\infty: T}$), is, from an applied point of view, only the theoretical benchmark.
Before introducing estimators for $\theta$ let us discuss the objects we want to construct inference for. In principle there are two unknown objects one could develop statistical intervals for: $\psi_{T+1}^s (\mathbf{X}_{1:T}; \theta)$ (or slightly more generally $\psi_{T+1}^s (\mathbf{X}_{t_1:T}^c; \theta)$) and $\psi_{T+1}^s (\mathbf{x}_{t_1:T}^c; \theta).$ For a GARCH(1,1), for instance, the first would read as $$ \psi_{T+1}^s (\mathbf{X}_{1:T}; \theta) = \frac{\omega}{1-\beta} + \alpha \sum_{k=0}^{T-1} \beta^k X_{T-k}^2 + \alpha \beta^T \sum_{k=0}^{\infty} \beta^k s_{-k}^2 $$ whereas the second with $t_1=1$ reads as $$ \psi_{T+1}^s (\mathbf{x}_{1:T}; \theta) = \frac{\omega}{1-\beta} + \alpha \sum_{k=0}^{T-1} \beta^k x_{T-k}^2 + \alpha \beta^T \sum_{k=0}^{\infty} \beta^k s_{-k}^2 $$ or more generally, if $t_1$ is not taken to be equal to one, as in ((ref)). While statistical intervals for both objects can be constructed we focus here on conditional inference, i.e. on intervals for $\psi_{T+1|t_1:T}=\psi_{T+1}^s (\mathbf{x}_{t_1:T}^c; \theta)$. In a time series context intervals for $\psi_{T+1|t_1:T}$ are motivated by the relevance property of kabaila1999relevance which postulates that intervals should relate to what actually happened during the sample period opposed to what might have happened. Indeed, intervals for $\psi_{T+1|t_1:T}$ can theoretically be shown to be considerably shorter than the intervals for their unconditional counterparts. While the unconditional objects might lead to conceptually easier analysis, our focus on the conditional objects is therefore not only theoretically but also empirically relevant.
As $\theta$ is unobserved, we need to estimate it. We assume that the estimator is based on a subsample $1:T_E$ (with $1 \leq T_E \leq T$) of the process $\{X_t^{E}\}$ which is potentially a different sample than $\{X_t\}$ that arises in the prediction function. The estimator of $\theta$ based on $\mathbf{X}_{1:T_E}^E=(X_1^E,\dots, X_{T_E}^E)^\prime$ will be denoted by $\hat{\theta} (\mathbf{X}_{1:T_E}^E)$. The introduction of $\{X_t^{E}\}$ serves three purposes: first, using a different process allows us to formulate the two-independent-processes argument where $X_t^E = Y_t$, with $\{Y_t\}$ independent of $\{X_t\}$, $T_E = T$ and an interval is constructed for $\psi_{T+1|1:T}$. Second, it will allow us to discuss the standard approach where $ X_t^E = X_t$, $T_E=T$, and an interval is constructed for $\psi_{T+1|1:T}$. Please note already here that this means that the same variables that arise in the prediction function are also used for estimating $\theta$. Third, it allows us to define the sample splitting approach which we study here. In this approach $ X_t^E = X_t$ and an interval is constructed for $\psi_{T+1|T_P:T}$ with $T_E < T_P$ (with $1 < T_P \leq T$) such that in contrast to the standard approach different subsamples are used for prediction and estimation. \\ Before we illustrate why the standard approach is problematic for constructing and evaluating conditional intervals, we need to define the final building block of prediction function estimation: the conditional prediction function estimator.
Note that in the above definition we do not condition on the sample $\mathbf{X}_{1:T_E}^E = \mathbf{x}_{1:T_E}^E$ that is used to estimate $\theta$. The reason for not conditioning on $\mathbf{X}_{1:T_E}^E = \mathbf{x}_{1:T_E}^E$ is that the goal is to preserve the randomness in the second argument of $\psi_{T+1}^s$, i.e. in $\hat{\theta} (\mathbf{X}_{1:T_E}^E)$, and consequently in $\widehat{\psi}_{T+1|T_P:T}$. Hence, if this goal is achieved we can use the (non-degenerate) conditional (on $\mathbf{X}_{T_P:T} = \mathbf{x}_{T_P:T}$) distribution of $\widehat{\psi}_{T+1|T_P:T}$ to construct confidence intervals for $\psi_{T+1|T_P:T}$. Having this said let us have a closer look at the standard approach. As mentioned above in the standard approach one has $X_t^E = X_t$, $T_P=1$ and $T_E=T$. Hence, denoting by $\widehat{\psi}_{T+1|1:T}^{STA}$ the “standard” estimator of the prediction function conditional on observing $\mathbf{X}_{1:T} = \mathbf{x}_{1:T}$, it becomes
Notice that there is no capital $\mathbf{X}$ in ((ref)) because there is only one sample and one typically conditions on all values of this sample. Hence, ((ref)) is non-random and thus does not have a distribution that could be used to construct intervals. Instead, to still be able to construct a “standard-looking” interval in practice, researchers typically implicitly rely on the (approximate) quantiles of the estimator
It is well understood in the literature that considering the sample as random and non-random at the same time as in ((ref)) does not provide a fully satisfactory justification of the intervals used in practice. For the readers not so familiar with the problem just discussed we provide two examples that both illustrate the problem arising from ((ref)). The examples illustrate that the severity of the problem may vary; ranging from only complicating the analysis (Example (ref)) to making the analysis impossible (Example (ref)).
The argument of two independent processes can at least be traced back to akaike1969fitting, who studies the prediction of AR time series. It reoccurs in lewis1985prediction (lewis1985prediction, p.\ 394): “...the series used for estimation of parameters and the series used for prediction are generated from two independent processes which have the same stochastic structure." The same argument also appears in lutkepohl2005new (lutkepohl2005new, p.\ 95) and in dufour2010short. Let $\{Y_t\}$ be a process independent of $\{X_t\}$ defined on the same probability space $(\Omega,\mathcal{F},\mathbb{P})$ with $\{Y_t\}$ having the same stochastic structure as $\{X_t\}$. In addition to the sample $\mathbf{X}_{1:T}$ of the process $\{X_t\}$, suppose there is a sample $\mathbf{Y}_{1:T} = (Y_1,\dots,Y_T)'$ of the process $\{Y_t\}$ that we use as estimation sample, that is $\mathbf{X}_{1:T}^E = \mathbf{Y}_{1:T}$. In this situation we denote the conditional prediction function estimator of Definition (ref) by $\widehat{\psi}_{T+1|1:T}^{2IP}$ and it equals
Notice that ((ref)) does not have the same shortcoming as ((ref)) because even if we consider $\mathbf{x}_{1:T}$ to be known we can nevertheless consider $\hat{\theta}(\mathbf{Y}_{1:T})$ to be random and can hence use its distribution to construct intervals. Throughout this paper, we call (ref), the 2IP (two independent processes) estimator. Then, a conditional interval $I_\gamma^{2IP}(\mathbf{x}_{1:T}, \mathbf{Y}_{1:T})$ can be based on the (approximate) quantiles of $\widehat{\psi}_{T+1|1:T}^{2IP}$. It satisfies
with the approximate sign indicating asymptotic equivalence. Note that the independence implies that the distribution of $\mathbf{Y}_{1:T}$ in ((ref)) does not depend on the realization $\mathbf{x}_{1:T}$, yet the statement does depend on $\mathbf{x}_{1:T}$ because the interval depends on it (for the AR(1) this can be directly seen from ((ref)) when replacing $\mathbf{X}_{1:T}$ by $\mathbf{Y}_{1:T}$). Although the 2IP approach is statistically sound, it assumes two independent processes with the same stochastic structure. phillips1979sampling points out that the assumption “is quite unrealistic in practical situations" (p.\ 241). Indeed, it is difficult to imagine this assumption to be satisfied in any real-life application beyond experimental settings. Moreover, as only one sample realization is available, to compute the estimate of the interval $I_\gamma^{2IP}(\mathbf{x}_{1:T}, \mathbf{Y}_{1:T})$ it is frequently suggested to take $ \mathbf{Y}_{1:T} = \mathbf{x}_{1:T}$, violating the independence assumption. Thus, the 2IP approach appears to be a rather questionable justification for the usual interval, and as such, in this paper we provide an alternative, realistic, justification of (asymptotically equivalent) intervals based on sample splitting.
An intuitive motivation for the sample-split approach is the successive decline of the influence of past observations present in a substantial class of time series models. This property permits to split our sample into two (asymptotically) independent subsamples. Consider the end point of the estimation sample, $T_E$, and the starting point of the prediction sample, $T_P$ satisfying $1< T_E < T_P \leq T$, such that the two samples are non-overlapping. In this situation we denote the conditional prediction function estimator of Definition (ref) by $\widehat{\psi}_{T+1|T_P:T}^{SPL}$ and it is given by
Throughout the paper, we call ((ref)) the SPL estimator (due to SPLitting). Similar to the two sample approach, we can consider the first argument of $\psi_{T+1}^s$ in ((ref)) as given and the second argument as random since the subsamples are non-overlapping. A conditional interval $I_\gamma^{SPL}(\mathbf{x}_{T_P:T}, \mathbf{X}_{1:T_E})$ can be constructed such that
This statement does make sense as there is still randomness in $\hat{\theta}(\mathbf{X}_{1:T_E})$ since $\mathbf{X}_{1:T_E}$ is not conditioned upon, yet the last $T - T_P+1$ values of $\{X_t\}_{t=1}^T$ are fixed such that their randomness is not taken into account. Similar to ((ref)), the statement in ((ref)) does depend on $\mathbf{x}_{T_P:T}$ and in contrast to $\mathbf{x}_{1:T}$ in ((ref)) the realization $\mathbf{x}_{T_P:T}$ may influence the distribution of $\mathbf{X}_{1:T_E}$. However, as said at the beginning of this subsection the idea of the sample split approach is that this dependence will vanish asymptotically.
In this section, we connect the sample-split procedure of Section (ref) with the two-independent-samples approach of Section (ref). First, in Section (ref), we show that the notion of weak convergence is inadequate to study asymptotic closeness for objects that vary over time and discuss the concept of merging. Then, in Section (ref) we link the 2IP and the SPL estimator by proving that their conditional distributions merge in probability (Theorem (ref)). Thereafter, in Section (ref), we construct asymptotically valid intervals (Theorem (ref)) and show that the sample-split intervals coincide asymptotically with the intervals commonly constructed by practitioners (Theorem (ref)). Last, in Section (ref), we state intervals of reduced form and simplified theoretical results under asymptotic normality of the parameter estimator.
To illustrate the inappropriateness of weak convergence in the context considered here, we revisit Example (ref) for the 2IP approach and the SPL approach, which shows that studying asymptotic closeness between conditional distributions is often complicated by the absence of a limiting distribution.
Next, we discuss what closeness means in the absence of a limiting distribution. To do so, first recall that weak convergence of a sequence of cdfs $\{F_T\}$ on $\mathbb{R}^k$ with $k \in \mathbb{N}$, i.e.\ $F_T(x)\to F(x)$ for all continuity points of $F$, can alternatively be defined by $d_{BL}(F_T,F) \rightarrow 0$. Here $d_{BL}$ denotes the bounded Lipschitz metric defined by
where for any real-valued function $f$ on $\mathbb{R}^k$ one puts $||f||_{BL}=\sup_x \big|f(x)\big|+\sup_{x \neq y}\frac{|f(x)-f(y)|}{||x-y||}$, with $||\cdot||$ denoting the Euclidean norm, i.e.\ $||A||=\sqrt{tr(A'A)}$ for any vector or matrix $A$. Following dudley2002real (see d1988merging and davydov2009asymptotic for related work) we state the following definition.
Note that weak convergence can be seen as a special case of merging with $G_T=G$ for all $T \in \mathbb{N}$.\\ While merging is appropriate to capture the asymptotic closeness of the conditional distribution of $\sqrt{T}(\hat{\beta}(\mathbf{Y}_{1:T})-\beta)X_T$ and $N(0,\sigma_\beta^2x_T^2)$ for a given sample $X_T=x_T$, we now extend the concept in a way that allows us to deal with asymptotic closeness when we do not condition on a particular sample. The necessity of this definition can again be exemplified by the AR(1), which also illustrates how we will deal the dependence of the statements in ((ref)) and in ((ref)) on the sample that we mentioned below these equations. For instance, in Example (ref) as described at the beginning of this section, the goal would be to formalize a statement like `when $T$ is large, the probability of all $x_T$ such that the distribution of $\sqrt{T}(\hat{\beta}(\mathbf{X}_{1:T})-\beta)x_T$ merges with that of $\sqrt{T_E}(\hat{\beta}(\mathbf{X}_{1:T_E})-\beta)x_T$ is approximately equal to one'. We now first introduce the conditional distributions of the 2IP and the SPL estimator in the general case and then give the definition capturing what we just illustrated for the AR(1).\\ Let $m_T$ be a sequence of normalizing constants with $m_T \to \infty$ (e.g.\ $m_T=\sqrt{T}$). For any $t_1 \geq 1$, we define the sub $\sigma$-algebra $\mathcal{I}_{t_1:T} = \sigma(X_t: t_1 \leq t \leq T)$. We denote the conditional cdfs of the 2IP and SPL estimator by
respectively, so that by specifying an event of $\mathcal{I}_{1:T}$ and $\mathcal{I}_{T_P:T}$, we see that ((ref)) and ((ref)) are just the centered and scaled distributions of ((ref)) and ((ref)), respectively. Please note that ((ref)) actually also depends on $c$, see ((ref)), but since our assumptions will ensure that this dependence vanishes asymptotically we prefer to suppress the dependence on $c$ here.
We can now define merging in probability (we do so without explicitly using the conditional cdfs of the 2IP and the SPL approach).
Here, we give conditions such that the conditional cdfs of the 2IP and SPL estimator merge. Clearly, the conditional confidence intervals are functions of these distributions so that their merging is a building block for the study of the conditional confidence intervals based on them. The conditions we give are divided into three parts. Roughly speaking, the first part (general assumptions) makes sure that the function we want to predict is well behaved and that we can estimate the parameter it depends on. The second part (two independent processes) and third part (SPL estimator) guarantee that these assumptions are met by the 2IP and the SPL method. To write the conditions in compact form we employ the usual stochastic order symbols $O_p$ and $o_p$. We assume that $\theta_0$ belongs to the interior of $\Theta$, i.e.\ $\theta_0 \in \mathring{\Theta}$, and we denote the set of all bounded, real-valued Lipschitz functions on $\mathbb{R}^r$ by $BL=\big \{h:\mathbb{R}^r\to \mathbb{R}:||h||_{BL}<\infty\big \}$. We start with the general assumptions.
Assumption (ref).a implies the existence of a limiting distribution for the parameter estimator. The differentiability assumption in (ref).b plus the boundedness Assumptions (ref).c ensure that the scaled prediction function estimators can accurately be approximated by a Taylor expansion; see Lemma (ref) for details. Assumption (ref).e with $t_1 = 1$ ensures the negligibility of the starting values when using the full-sample for prediction, while taking $t_1 = T_P$ ensures that this extends to the case where additionally $X_1,\ldots,X_{T_P-1}$ are replaced by constants, i.e. where only the subsample $(X_{T_P},\ldots,X_{T})$ is used for prediction. This assumption implicitly limits the choice of $T_P$; as replacing past values of $X_t$ for $t < T_P$ by arbitrary constants should have a negligible effect, $T - T_P$ needs to increase faster than some lower bound $l_T$. For models exhibiting an exponential decay in memory, it typically suffices to take $l_T = \log T$ (see e.g. beutner2019technical, beutner2019technical, eq. (4.6)).\\ For the 2IP estimator, we additionally need the two-independent-processes assumption, which is formalized in Assumption (ref).
For the SPL estimator we replace the two-independent-processes assumption by a stationarity and a weak dependence condition, which allows to split our sample into two (asymptotically) independent and identical subsamples. In addition we need an assumption on $T_P$ and $T_E$ as functions of $T$, that is $T_E (T)$ and $T_P (T)$.
The subsample size assumption in (ref).a ensures that the number of observations used for conditioning is increasing, which along with the negligibility of the initial conditions implies that the truncation of the prediction function is negligible. Furthermore, the sample size used for estimation should increase fast enough that the respective scaling of the 2IP and SPL estimators, $m_{T}$ and $m_{T_E}$ respectively, are asymptotically identical. If $m_T$ increases no faster than a polynomial rate, which is generally the case, it is sufficient that $T_E/T \rightarrow 1$ for $m_{T_E}/m_T \rightarrow 1$ to hold.
The stationarity assumption in (ref).b can actually be relaxed; what matters is that the conditions in Assumption (ref) are still true if only a subsample is considered. In particular, we need that $m_{T_E} \big(\hat{\theta}(\mathbf{X}_{1:T_E}) - \theta_0\big)\overset{d}{\to} G_{\infty}$, which - along with the assumptions on gradient and Hessian - is certainly satisfied under stationarity. However, in general the assumption will be far too strict; here we use it simply to have a clear, interpretable assumption rather than a list of high-level assumptions that are difficult to interpret. The weak dependence condition in (ref).c is met by numerous Markov processes. Intuitively, $(X_1,\dots,X_{T_E})$ and $(X_{T_P},\dots,X_T)$ approach independence as their temporal distance $T_P-T_E$ increases. We illustrate a particular case in the Remark (ref).
Assumptions (ref) to (ref) are met by the AR and GARCH processes considered in Examples (ref) and (ref) (with 1.e holding for bounded sequences). A detailed verification of each assumption under mild conditions is provided in beutner2019technical. We state the following theorem.
Having established asymptotic closeness between the conditional cdfs $F_T^{2IP}(\cdot|\mathcal{I}_{1:T})$ and $F_T^{SPL}(\cdot|\mathcal{I}_{T_P:T})$, we now turn to the construction of asymptotic intervals.
Henceforth, for any cdf $F$ we write $F^{-1}$ to denote its generalized inverse given by $F^{-1}(u)=\inf\big\{\tau \in \mathbb{R}:F(\tau)\geq u\big\}$. A confidence interval for $\psi_{T+1}$ based on quantiles of (ref) or (ref) is typically infeasible as these cumulative distribution functions are unknown for finite $T$. Here, they are infeasible because roughly they are the distribution functions of some weights which induce merging multiplied by $m_T \big(\hat{\theta}(\mathbf{X}_{1:T}) - \theta_0\big)$ and $m_{T_E} \big(\hat{\theta}(\mathbf{X}_{1:T_E}) - \theta_0\big)$, respectively, where, in general, the distributions of $m_T \big(\hat{\theta}(\mathbf{X}_{1:T}) - \theta_0\big)$ and $m_{T_E} \big(\hat{\theta}(\mathbf{X}_{1:T_E}) - \theta_0\big)$ are unknown in finite samples. Since these are the only unknown distributions an asymptotic approximation can be based on $G_{\infty}$ with merging induced by the non-convergent weights. In general, we also need to estimate $G_{\infty}$; see Examples (ref) and (ref) below for common approaches. We denote estimators of (ref) and (ref) resulting from this approximation by $\widehat{F_T^{2IP}}(\cdot)$ and $\widehat{F_T^{SPL}}(\cdot)$, respectively. In Section (ref), we provide explicit expressions when $G_{\infty}$ is multivariate normal. For the general construction, we refer to relations (ref) and (ref) in Appendix (ref) and the explanations preceding these relations. Based on the 2IP approach, we consider an interval of the form
where $\gamma_1,\gamma_2 \in [0,1)$ satisfy $\gamma=\gamma_1+\gamma_2$. We typically take $\gamma_1=\gamma_2=\gamma/2$ such that the interval is equal-tailed. Similarly, we construct the following sample split interval:
To achieve correct coverage, we need that $\widehat{F_T^{2IP}}(\cdot)$ and $F_T^{2IP}(\cdot|\mathcal{I}_{1:T})$ merge in probability and likewise for SPL. A sufficient condition for this in our setting is that we can consistently estimate the asymptotic distribution of the parameter estimator, $G_{\infty}$, by an appropriate estimator. This is formulated in Assumption (ref) below.
Although we did not explicitly specify in Assumption (ref) the dependence of $\widehat{G}_T$ on $\mathbf{X}_{1:T}$, it should be understood to hold for any subsample of $\mathbf{X}_{1:T}$ whose size goes to infinity. The verification of Assumption (ref) is a standard step in asymptotic analysis. The two examples below provide common methods for verifying Assumption (ref).
The following theorem states the intervals' asymptotic validity.
However, the standard approach, motivated by $I_\gamma^{2IP}$ as in (ref), computes an interval of the form $I_\gamma^{STA}(\mathbf{x}_{1:T},\mathbf{x}_{1:T}) = I_\gamma^{2IP}(\mathbf{x}_{1:T},\mathbf{x}_{1:T})$ as only one sample realization is available. This, of course, strongly violates the independence assumption of $\{X_t\}$ and $\{Y_t\}$. Specifically, replacing $\mathbf{Y}_{1:T}$ by $\mathbf{X}_{1:T}$ in equation (ref), leads to
where $\widehat{F_T^{STA}}(\cdot)$ is defined in relation (ref) and the text preceding it. Whereas it is difficult to justify a conditional confidence interval like $I_{\gamma}^{STA}(\mathbf{x}_{1:T},\mathbf{X}_{1:T})$ directly due to the lack of randomness, we can provide a justification by characterizing how closely the interval resembles the SPL interval. We establish the asymptotic equivalence, defined in terms of location and (scaled) length, between the two intervals in the following theorem. Note that, as our characterization of equivalence is probabilistic, we need to introduce the “doubly random” versions of the STA and SPL estimators, where the sample we condition on is considered random. These estimators are denoted as $\hat\psi_{T+1} (\mathbf{X}_{1:T}, \mathbf{X}_{1:T})$ and $\hat\psi_{T+1} (\mathbf{X}_{T_P:T}^c, \mathbf{X}_{1:T_E})$ respectively.
The first implication states that the locations of the two intervals coincide asymptotically. The second statement establishes asymptotic closeness of the selected quantiles such that the scaled lengths of the intervals in (ref) and (ref) coincide asymptotically. As such, our sample-split interval coincides asymptotically with the standard interval, meaning that the standard interval can be substituted for an (asymptotically) equivalent interval which has a formal justification in terms of conditional coverage. As such, this provides a justification for the intervals commonly constructed in practice without having to rely on the two-independent-processes assumption.
In this subsection we present intervals of reduced form and simplified theoretical results under asymptotic normality of the parameter estimator.
Usually, the covariance estimator is obtained by inserting consistent estimators for $\theta_0$ and $\xi_0$ into $\Upsilon_0$. Following the plug-in principle, we estimate $F_T^{2IP}(\cdot|\mathcal{I}_{1:T})$ by a normal distribution with mean $0$ and variance $\hat{\upsilon}_T^{2IP}=\frac{\partial \psi_{T+1}(\mathbf{x}_{1:T};\hat{\theta}(\mathbf{Y}_{1:T}))}{\partial \theta'}\hat{\Upsilon}(\mathbf{Y}_{1:T})\frac{\partial \psi_{T+1}(x_{1:T};\hat{\theta}(\mathbf{Y}_{1:T}))}{\partial \theta}$ such that $\widehat{F_T^{2IP}}(\cdot)=\Phi\big(\cdot/\sqrt{\hat{\upsilon}_T^{2IP}}\big)$. Then, the interval in (ref) simplifies to
Similarly, for the sample-split approach we consider $\widehat{F_T^{SPL}}(\cdot)= \Phi\big(\cdot / \sqrt{\hat{\upsilon}_T^{SPL}}\big)$ with $\hat{\upsilon}_T^{SPL}=\frac{\partial \psi_{T+1}(\mathbf{x}_{T_P:T}^{c};\hat{\theta}(\mathbf{X}_{1:T_E}))}{\partial \theta'}\hat{\Upsilon}(\mathbf{X}_{1:T_E})\frac{\partial \psi_{T+1}(\mathbf{x}_{T_P:T}^{c};\hat{\theta}(\mathbf{X}_{1:T_E}))}{\partial \theta'}$ such that (ref) reduces to
In Appendix B we show that if the variance estimator is bounded away from zero in probability, e.g. $1/\hat{\upsilon}_T^{2IP}=O_p(1)$, then $\widehat{F_T^{2IP}}(\cdot)$ is stochastically uniform equicontinuous. Therefore, the asymptotic validity of both intervals can be deduced from Theorem (ref).
Bounding the variance estimator away from zero in probability to establish that the conditional coverage probability converges to $1-\gamma$ in probability has intuitive appeal: as $\hat{\upsilon}_T^{2IP}$ approaches zero, $N\big(0,\hat{\upsilon}_T^{2IP}\big)$ becomes degenerate while the interval in (ref) collapses (similar for SPL).
For the standard interval, replacing $\mathbf{Y}_{1:T}$ by $\mathbf{X}_{1:T}$ in (ref) leads to
with $\hat{\upsilon}_{T}^{STA}=\frac{\partial \psi_{T+1}(x_{1:T};\hat{\theta}(\mathbf{X}_{1:T}))}{\partial \theta'}\hat{\Upsilon}(\mathbf{X}_n)\frac{\partial \psi_{T+1}(\mathbf{x}_{1:T};\hat{\theta}(\mathbf{X}_{1:T}))}{\partial \theta}$. In Appendix B we prove that $\hat{\upsilon}_T^{SPL}$ is bounded in probability, which in turn implies that the quantile function $\widehat{F_T^{SPL}}^{-1}(\cdot)=\sqrt{\hat{\upsilon}_T^{SPL}}\Phi^{-1}(\cdot)$ is stochastically pointwise equicontinuous at any $u \in \mathbb{R}$. Hence, Theorem (ref) applies. Whereas the first statement of the theorem remains unaffected, its second statement reads as follows under normality.
The preceding sections have focused purely on the construction of conditional confidence intervals to account for parameter uncertainty. Regarding prediction, a second source of uncertainty arises, that corresponds to the model's innovation process. In this setting, parameter estimation is typically disregarded in textbooks as the stochastic fluctuation stemming from the estimation procedure is generally dominated by the stochastic fluctuation of the innovations. Although the resulting prediction intervals may be asymptotically valid, they are typically characterized by under-coverage in finite samples. In response, pan2016abootstrap introduce the concept of asymptotic pertinence to evaluate distribution approximations that account for the two sources of randomness, innovation and parameter estimation uncertainty, according to their general orders of magnitude. Whereas kunitomo1985properties and samaranayake1988properties study properties of the unconditional law of the forecast error, we focus on its conditional distribution to conform with the relevance property of kabaila1999relevance. The fundamental issue also arises when considering prediction if one attempts to account for parameter uncertainty. To illustrate this point, we revisit the introductory examples and write $*$ to denote the convolution operator.\footnote{For independent variables $X$ and $Y$ with $X\sim F_X$, $Y\sim F_Y$ and $Z=X+Y \sim F_{Z}$, we write $F_{Z}=F_X * F_Y$ to denote $F_{Z}(z) = \int_{-\infty}^zF_X(z-y)dF_Y(y)$.}
In his textbook pesaran2015time resorts to a Bayesian-akin approach to avoid the fundamental issue in forecasting. He argues that $\theta$, although “fixed at the estimation stage ... is viewed best as a random variable at the forecasting stage" (p.\ 389). Consequently, he assigns some posterior distribution to $\theta$ motivated by an uninformed prior. Treating $\theta$ not fixed but random, the fundamental issue does not arise, however combining a frequentist view with a Bayesian method does not seem to be coherent.
barndorff1996prediction require the existence of a transitive statistic $U=U(\mathbf{X}_{1:T})$ of fixed low dimension to establish conditional independence between the sample $\mathbf{X}_{1:T}$ and their considered future random variable given $U=u$. vidoni2004improved (vidoni2004improved, vidoni2009improved, vidoni2009simple, vidoni2016improved), kabaila1999relevance, kabaila2004adjustment, and kabaila2008improved (kabaila2008improved, kabaila2010asymptotic) extend their approach and derive improved prediction intervals. Although these methods absorb an additional $O(T^{-1})$ term in the associated conditional coverage probability, there are several drawbacks associated with them: the innovation distribution needs typically be specified (e.g.\ Gaussian), the results apply only to a limited set of estimators (e.g.\ maximum likelihood) and their framework can only incorporate finite autoregressive components (e.g.\ AR$(p)$).
Assuming two independent processes with the same stochastic structure, using one for prediction and one for the estimation of the parameters, alleviates the fundamental issue faced in the continued Examples (ref) and (ref). As the conditional distributions of the 2IP and SPL estimators merge in probability by Theorem (ref), the 2IP assumption can be avoided by following a sample-split approach as described in Section (ref).
In the paper at hand, we study the construction of confidence intervals for conditional objects such as conditional means or conditional variances, focusing on the conceptual issue that arises in the process of taking parameter uncertainty into account. It stems from the fact that on one hand one must condition on the sample as the past informs about the present and future, yet on the other hand one must allow the data up to now to be treated as random to account for estimation uncertainty. Assuming two independent processes with the same stochastic structure, where one is used for conditioning and one for the estimation of the parameters, bypasses this issue, but the assumption itself can generally not be justified in applications. To avoid this assumption, we propose a solution based on a simple sample-split approach, that requires a much more realistic weak dependence condition instead. To acknowledge that the conditional quantities vary over time, we employ a merging concept generalizing the notion of weak convergence. The conditional distributions of the sample-split estimator and the estimator based on the two-independent-processes assumption are shown to merge in probability under mild conditions. The corresponding sample-split intervals are shown to coincide asymptotically with the intervals commonly constructed by practitioners, which provides a novel and theoretically satisfactory justification for commonly constructed confidence intervals for conditional objects, applicable to a wide class of time series models, including ARMA and GARCH-type models.
One limitation to our approach is that we restrict ourselves to univariate time series and objects of interests. At the expense of more involved notation this could be readily extended to multivariate time series and objects of interests. A second, and more restrictive, limitation is our weak dependence assumption needed to achieve asymptotic independence between the two subsamples, which for instance rules out application to integrated processes. Given the fundamental role of this assumption in our setup, it appears difficult to generalize this. However, this also casts further doubt on the two-independent-processes assumption as validation for confidence intervals constructed for such persistent processes. A case-by-case treatment, as for instance done by gospodinov2002median for near unit root processes and samaranayake1988properties for explosive processes, appears to be necessary in such cases, and standard confidence intervals should be treated with caution.