EconBase
← Back to paper

A multivariate extension of the Misspecification-Resistant Information Criterion

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

64,569 characters · 15 sections · 23 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A multivariate extension of the Misspecification-Resistant Information Criterion.

\def\spacingset#1{ {#1}} \spacingset{1}

\if00 \fi

\if10 {

center[center omitted — 112 chars of source]

} \fi

abstractThe Misspecification-Resistant Information Criterion (MRIC) proposed in [H.-L. Hsu, C.-K. Ing, H. Tong: On model selection from a finite family of possibly misspecified time series models. The Annals of Statistics. 47 (2), 1061--1087 (2019)] is a model selection criterion for univariate parametric time series that enjoys both the property of consistency and asymptotic efficiency. In this article we extend the MRIC to the case where the response is a multivariate time series and the predictor is univariate. The extension requires novel derivations based upon random matrix theory. We obtain an asymptotic expression for the mean squared prediction error matrix, the vectorial MRIC and prove the consistency of its method-of-moments estimator. Moreover, we prove its asymptotic efficiency. Finally, we show with an example that, in presence of misspecification, the vectorial MRIC identifies the best predictive model whereas traditional information criteria like AIC or BIC fail to achieve the task.

{\it Keywords:} information criteria, model selection,multivariate time series, Mean Square Prediction Error.

\spacingset{1.45}

Introduction

The model selection step is a fundamental task in statistical modelling and its implementation typically depends upon the objective of the exercise. In the time series framework the focus is on either forecasting future values or describing/controlling the process that has generated the data (DGP). A good model selection criterion must feature a good ability to identify the model with the “best” fit to future values, in a specified sense. In particular, in the parametric time series framework, we can identify two main properties. The first one is consistency, i.e., the ability to select the true DGP with probability one as the sample size diverges. This assumes that a true model exists and is among the set of candidate models. If either the set of candidate models does not contain the true DGP, or, for some reason, a true model cannot be postulated, then a selection criterion should be asymptotically efficient, for instance, in the mean square sense, i.e. it minimizes the mean squared prediction error as the sample size diverges. Starting from the seminal work of Akaike, AKA1973 a plethora of model selection criteria has been proposed. These include Akaike's AIC AKA1973, AKA1974, Schwarz's Bayesian Information Criterion (BIC) SCH1978, and Rissanen's Minimum Description Length (MDL) RIS1978. Such criteria paved the way for various extensions dealing with different unsolved issues. For instance, the AIC is efficient but not consistent (i.e. it leads to select overfitting models), whereas the BIC is consistent but not efficient, see HSU2019 for a discussion. A recent development for model selection in possibly misspecified parametric time series models in the fixed-dimensionality setting is given by the Misspecification-Resistant Information Criterion (hereafter $\operatorname{MRIC}$) HSU2019. Fixed-dimensionality means that the number of observations increases to infinity while the number of ‘true’ parameters is finite. In this respect, the MRIC provides a solution to the original research question of Akaike: it enjoys both consistency, in case the true model is included as a candidate, and asymptotic efficiency when a true model either cannot be assumed or is not included. Moreover, when the number of variables in the model grows with the sample size, the $\operatorname{MRIC}$ can achieve asymptotic efficiency, without the need for additional criteria. Finally, in the high-dimensional setting, the $\operatorname{MRIC}$ can be used together with appropriate model selection criteria to identify the best predictive models. The $\operatorname{MRIC}$ is based upon the additive decomposition of the mean squared prediction error in a term that depends upon the misspecification level and a term that measures the sampling variability of the predictor. The idea is to select the model with smallest variability among those that minimize the misspecification index. The appealing properties of the $\operatorname{MRIC}$ make it an ideal tool for omnibus time series model selection but, to date, only the univariate response case has been studied HSU2019. In this work we extend the $\operatorname{MRIC}$ to multivariate time series with a single regressor as to obtain the vectorial MRIC (hereafter $\operatorname{VMRIC}$). As it will be clear, such an extension does not easily derive from the univariate case since it requires dealing with the dependence structure within the components of the vector of forecasting error and hence relies upon random matrix theory. Such multivariate extension can be used in all those models where many time series depend upon a single regressor, like for instance, in econometrics, where many interest rates depend upon a single macroeconomic indicator, such as inflation. Other possible applications include dimension reduction and hedging, which is intimately connected to the problem of model selection BES16. The rest of the paper is organized as follows: in Section (ref) we introduce the notation and in Section (ref) summarize the available results for the univariate case; in Section (ref) we extend the MRIC approach to multivariate time series with a single regressor. In particular, in Section (ref) we obtain the asymptotic decomposition of the Mean Squared Prediction Error (hereafter MSPE) matrix into two parts: the first one is linked to the goodness of fit of the model and the second one depends upon the prediction variance. In Section (ref) we present the $\operatorname{VMRIC}$ and derive a consistent estimator for it, whereas in Section (ref), we prove the asymptotic efficiency of the $\operatorname{VMRIC}$. Section (ref) presents an example to assess the effect of misspecification in the $\operatorname{VMRIC}$ framework. All the proofs are detailed in Section (ref). (ref) contains auxiliary technical lemmas.

Notation and preliminaries

For each $t$, let $\{\mathbf{x}_t\}$ and $\{\mathbf{y}_t\}$, with $\mathbf{x}_t=(x_{t,1},\dots,x_{t,m})^\top$ and $\mathbf{y}_t=(y_{t,1},\dots,y_{t,w})^\top$, be two weakly stationary stochastic processes defined over the probability space $\left(\Omega, \mathcal{F}, \mathbb{P}\right)$. When $m=1$ ($w=1$, respectively) we write $x_t$ ($y_t$). Given a vector $\mathbf{v}$ and a matrix $\mathbf{M}$, we use $\left\lVert\mathbf{v}\right\rVert$ and $\left\lVert\mathbf{M}\right\rVert$ to refer to the $\mathcal{L}_2$ vectorial norm and the matrix norm induced by the Euclidean norm, respectively. We write $o(1)$ ($o_p(1)$) to indicate a sequence that converges (in probability) to zero and $O(1)$ ($O_p(1)$) to indicate a sequence that is bounded (in probability). Moreover, let $\{c_n\}$ be a sequence of scalar random variables whereas $\{\mathbf{v}_n\}$ and $\{\mathbf{M}_n\}$ are sequences of random vectors and random matrices, respectively. We adopt the following notation: $\mathbf{v}_n =o_p(c_n)$ if $\left\lVert\mathbf{v}_n\right\rVert/c_n=o_p(1)$; $\mathbf{v}_n=O_p(c_n)$, if $\left\lVert\mathbf{v}_n\right\rVert/c_n=O_p(1)$, $\mathbf{M}_n =o_p(c_n)$ if $\left\lVert\mathbf{M}_n\right\rVert/c_n=o_p(1)$; $\mathbf{M}_n = O_p(c_n)$ if $\left\lVert\mathbf{M}_n\right\rVert/c_n=O_p(1)$. For further details on matrix algebra see SEB2008, HOR2013, ODE2018, for multivariate time series see REI1993, LUT2005, TSA2013, and for asymptotic tools for vector and matrices, see JIA2010. Let $\{(\mathbf{x}_t,\mathbf{y}_t),t \in \{1,\dots,n\}\}$ be the observed sample, and divide the interval $[1,n]$ into the training set $[1,N]$ and the test set $[N+1,N+h]$, with $h$ being the forecasting horizon. Note that $\mathbf{x}_t$ can contain both endogenous and exogenous variables, therefore, Model ((ref)) encompasses many different models including, inter alia, VAR and VARX models. We denote $\bar{\mathbf{x}}=n^{-1}\sum_{t=1}^{n}\mathbf{x}_t$ and $\bar{\mathbf{y}}=n^{-1}\sum_{t=1}^{n}\mathbf{y}_t$, i.e. the two sample means. Without loss of generality assume $\operatorname{E}[\mathbf{x}_t]=\operatorname{E}[\mathbf{y}_t]=\boldsymbol 0$. In order to forecast $\mathbf{y}_{n+h}$, $h\geq 1$, we adopt the following $h$-step ahead forecasting Model:

equation[equation omitted — 113 chars of source]

where $\mathbf{B}_h$ is a $(w\times m)$ matrix and $\boldsymbol{\varepsilon}_t^{(h)}$ is the vector containing the $w$ $h$-step ahead forecast errors; as before, if $w=1$ we write $\varepsilon_t^{(h)}$.

remarkSince the model can possibly be misspecified, the prediction error vector $\boldsymbol{\varepsilon}_t^{(h)}$ can be serially correlated, and also correlated with $\mathbf{x}_s$, $s\neq t$. Moreover, the multivariate framework differs from HSU2019 in different key aspects. For instance, $(i)$ the components of the error vector can be cross-correlated, and $(ii)$ $\mathbf{x}_t \boldsymbol{\varepsilon}_t^{(h)}$ and $\mathbf{x}_k \boldsymbol{\varepsilon}_k^{(h)}$, for $t\neq k$, can also be both serially and cross correlated.

Define

equation[equation omitted — 187 chars of source]

Then, the ordinary least squares estimator (hereafter OLS) of $\mathbf{B}_h$ results:

equation[equation omitted — 145 chars of source]

When $w=1$, $\mathbf{R}$ and $\mathbf{B}$ become $R$ and $\boldsymbol\beta$, respectively. The prediction of $\mathbf{y}_{n+h}$, $h\geq 1$, is given by

equation[equation omitted — 90 chars of source]

and the corresponding Mean Squared Prediction Error matrix is

equation[equation omitted — 181 chars of source]

The MRIC for parametric univariate time series models

In HSU2019, the authors focused on the case $w=1$ and $m\geq1$. Under appropriate conditions, they obtained the following asymptotic decomposition of $\operatorname{MSPE}$:

align[align omitted — 454 chars of source]

where $\mathbf{C}_{h,s} = \operatorname{E}\left[\mathbf{x}_1\mathbf{x}_{1+s}^\top \varepsilon_1^{(h)} \varepsilon_{1+s}^{(h)}\right]$, $s\geq0$, is the cross-covariance matrix between the regressors and the $h$-step ahead prediction error at lag $s$.

remarkThe first part of Eq. ((ref)) is the Misspecification Index ($\operatorname{MI}$), linked to the goodness-of-fit of the model and coincides with the $h$-step ahead prediction error variance. The second component is the Variability Index ($\operatorname{VI}$), which depends upon the variance of the $h$-step ahead predictor, $\hat{y}_{n+h}= \hat{\boldsymbol{\beta}}_n^\top (h) \mathbf{x}_n$, and is also linked to the bias of the estimator of $\boldsymbol{\beta}_h$.

Based upon the above decomposition, the $\operatorname{MRIC}$ is defined as follows:

align[align omitted — 129 chars of source]

with $\hat{\operatorname{MI}}_h$ and $\hat{\operatorname{VI}}_h$ being the estimators of $\operatorname{MI}_h$ and $\operatorname{VI}_h$ respectively, i.e.: \[ \hat{\operatorname{MI}}_h = N^{-1}\sum_{t=1}^{N}\left(\hat{\varepsilon}_t^{(h)}\right)^2,\quad \hat{\operatorname{VI}}_h = \operatorname{tr}\left(\hat R^{-1}\hat{\mathbf{C}}_{h,0}\right) + 2 \sum_{s=1}^{h-1} \operatorname{tr}\left(\hat R^{-1} \hat{\mathbf{C}}_{h,s}\right), \] where $\hat{\mathbf{C}}_{h,s} = (N-s)^{-1} \sum_{t=1}^{N-s}\mathbf{x}_t\mathbf{x}_{t+s}^\top \hat{\varepsilon}_t^{(h)}\hat{\varepsilon}_{t+s}^{(h)}$ and $\hat{\varepsilon}_t^{(h)}=y_{t+h}-\hat{\boldsymbol{\beta}}_n(h)\mathbf{x}_t$ is the estimated forecast error; $\alpha_n$ is a penalization term sequence such that, as $n$ increases:

align[align omitted — 115 chars of source]

It is shown that $\hat{\operatorname{MI}}_h$ and $\hat{\operatorname{VI}}_h$ are consistent estimators of ${\operatorname{MI}}_h$ and ${\operatorname{VI}}_h$, moreover the asymptotic efficiency of the $\operatorname{MRIC}$ is proved. By minimizing this criterion, the model which minimizes $\operatorname{VI}$ among those with minimum $\operatorname{MI}$ is selected. Among other features, the $\operatorname{MRIC}$ is particularly helpful in situations where competing models present the same goodness-of-fit and the same number of parameters.

remarkThe type of penalty considered in HSU2019 is similar to that used in SHI1989 for the correctly specified case.

A multivariate extension of the MRIC framework

In this section we extend the $\operatorname{MRIC}$ approach to the case where the response is a multivariate time series ($w\geq2$) and the predictor is univariate ($m=1$), for a generic $h$-step ahead forecast. Hence, Model ((ref)) reduces to $\mathbf{y}_{t+h}=\boldsymbol{\beta}_h x_t+ \boldsymbol{\varepsilon}_t^{(h)}$, namely:

equation[equation omitted — 252 chars of source]

Asymptotic decomposition of the MSPE matrix

We extend the asymptotic representation of the $\operatorname{MSPE}_h$ defined in ((ref)) which is the key step to derive the $\operatorname{VMRIC}$ in this multivariate framework. We rely upon the following assumptions, which are the natural multivariate extensions of those in HSU2019.

assumptions\begin{align*} (C1) & \quad \exists \ q_1 > 5, 0<K_1<\infty : \ for any 1 \leq n_1 < n_2 \leq n,\\ & \operatorname{E}\left[ \left\lvert \left( n_2-n_1+1 \right)^{-1/2} \sum_{t=n_1}^{n_2} x_{t}^2 - \operatorname{E}\left[ x_{t}^2 \right] \right\rvert^{q_1} \right]\leq K_1.\\ (C2) &\quad 1. \ \mathbf{C}_{h,s} = \ \operatorname{E}\left[ \boldsymbol{\varepsilon}_t^{(h)} x_t \left(\boldsymbol{\varepsilon}_{t+s}^{(h)} x_{t+s}\right)^\top \right] \perp t,\\ &\quad 2. \ \operatorname{E}\left[ x_1 x_n \varepsilon_{1,i}^{(h)} \varepsilon_{n,j}^{(h)} \right] = \ o(n^{-1}) \ \forall \ i,j \in\{ 1, \dots, w\}.\\ (C3) & \quad 1. \sup_{-\infty<t<\infty} \operatorname{E}\left[ | x_t |^{10} \right] < \infty,\\ & \quad 2. \sup_{-\infty<t<\infty} \operatorname{E}\left[ \left\lVert\boldsymbol{\varepsilon}_{t}^{(h)}\right\rVert^{6} \right] < \infty.\\ (C4) & \quad \exists \ 0 < K_2 < \infty : \ for \ 1\leq n_1 < n_2 \leq n, \quad \operatorname{E}\left[\left\lVert\left(n_2-n_1+1\right)^{-\frac{1}{2}} \sum_{t=n_1}^{n_2} \boldsymbol{\varepsilon}_t^{(h)} x_t \right\rVert^5 \right] < K_2.\\ (\text{C}5) & \quad \text{For any } q>0, \ \operatorname{E}\left[ \left|\hat{R}^{-1} \right|^q \right] = O(1).\\ (\text{C}6) & \quad \exists \mathcal{F}_t \subseteq \mathcal{F}, \mathcal{F}_t \ \text{an increasing sequence of } \sigma\text{-fields such that: }\\ & \quad 1. \ x_t \ \text{is} \ \mathcal{F}_t\text{-measurable}\\ & \quad 2. \ \sup_{-\infty<t<\infty} \operatorname{E}\left[ \left|\operatorname{E}\left[ x_t^2 \mid \mathcal{F}_{t-k}\right] - R\right|^3\right] = \ o(1), \text{ as } k \rightarrow\infty,\\ & \quad 3. \ \sup_{-\infty<t<\infty} \operatorname{E}\left[ \left\lVert\operatorname{E}\left[ \boldsymbol{\varepsilon}_{t}^{(h)} x_t \mid \mathcal{F}_{t-k}\right]\right\rVert^3\right] = \ o(1), \text{ as } k \rightarrow\infty. \end{align*}
theoremUnder the regularity conditions (C1) -- (C6), the asymptotic expression of the $\operatorname{\textbf{MSPE}}_h$ defined in ((ref)) results \begin{align} &N \left\{ \operatorname{E}\left[ \left( \mathbf{y}_{n+h} - \hat{\mathbf{y}}_{n+h} \right) \left( \mathbf{y}_{n+h} - \hat{\mathbf{y}}_{n+h} \right)^\top - \operatorname{E}\left[ \boldsymbol{\varepsilon}_{n}^{(h)} \boldsymbol{\varepsilon}_{n}^{(h)^\top} \right] \right] \right\}\\ & = R^{-1} \operatorname{E}\left[ \left(\boldsymbol{\varepsilon}_{1}^{(h)} x_1\right)\left(\boldsymbol{\varepsilon}_{1}^{(h)} x_1\right)^\top \right]\nonumber\\ &+ R^{-1} \operatorname{E}\left[\sum_{s=1}^{h-1} \left\{\left( \boldsymbol{\varepsilon}_1^{(h)} x_1\right) \left(\boldsymbol{\varepsilon}_{s+1}^{(h)} x_{s+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{s+1}^{(h)} x_{s+1}\right) \left( \boldsymbol{\varepsilon}_1^{(h)} x_1 \right)^\top\right\} \right] + o(1).\nonumber \end{align}

VMRIC and its consistent estimation

In this section we introduce the $\operatorname{VMRIC}$. Let $\{\alpha_n\}$ be the penalization term sequence defined as in Eq. ((ref)).

align[align omitted — 672 chars of source]

The $\operatorname{VMRIC}$ can be estimated via the method of moments as to obtain:

align[align omitted — 560 chars of source]

and $\hat{\mathbf{C}}_{h,s} = (N-s)^{-1} \sum_{t=1}^{N-s} x_t x_{t+s} \hat{\boldsymbol{\varepsilon}}_t \hat{\boldsymbol{\varepsilon}}_{t+s}^\top$, and $\hat{\boldsymbol{\varepsilon}}_t=\mathbf{y}_{t+h}-\hat{\boldsymbol{\beta}}_n(h) x_t$ is the estimated forecast error vector. In Theorem (ref) we prove that $\hat{\operatorname{\textbf{MI}}}_h$ and $\hat{\operatorname{\textbf{VI}}}_h$ are consistent estimators of ${\operatorname{\textbf{MI}}}_h$ and ${\operatorname{\textbf{VI}}}_h$, respectively. Theorem (ref) relies upon the following assumptions, that are less restrictive with respect to (C1) -- (C6). For further discussions on the assumptions see HSU2019.

assumptionsFor each $0\leq s \leq h-1$, we assume the following: \begin{align*} (A1) & \quad n^{-1} \sum_{t=1}^{n} \left( \boldsymbol{\varepsilon}_t^{(h)} \boldsymbol{\varepsilon}_t^{(h)\top}\right) = \operatorname{E}\left[ \boldsymbol{\varepsilon}_1^{(h)} \boldsymbol{\varepsilon}_1^{(h)\top} \right] + O_p \left(n^{-1/2}\right)\%, and any 1\leq i,j \leq m,\\ (A2) &\quad n^{-1} \sum_{t=1}^{n} \left(x_t \boldsymbol{\varepsilon}_t^{(h)}\right) \left(x_{t+s}\boldsymbol{\varepsilon}_{t+s}^{(h)}\right)^\top = \boldsymbol{C}_{h,s} + o_p(1),\\ (A3) & \quad n ^{-1/2} \sum_{t=1}^{n} x_t \boldsymbol{\varepsilon}_t^{(h)} = O_p(1).\\ (A4) & \quad n ^{-1} \sum_{t=1}^{n} x_t^{2} = R + o_p(1),\\ (A5) & \quad \sup_{-\infty < t < \infty} \operatorname{E} \left[\left\|\boldsymbol{\varepsilon}_t^{(h)}\right\|^4\right] + \sup_{-\infty<t<\infty} \operatorname{E}\left[\abs {x_t}^4 \right] < \infty. \end{align*}
theoremIf Assumptions (A1) -- (A5) hold, then for the case $w\geq2$, and $m=1$ we obtain: \begin{align*} \hat{\operatorname{MI}}_h&= \operatorname{MI}_h +\, O_p (n^{-1/2}),\\ \hat{\operatorname{VI}}_h&= \operatorname{VI}_h +\, o_p (1). \end{align*}

Asymptotic efficiency

In this section we prove the asymptotic efficiency of the $\operatorname{VMRIC}$ in the fixed dimensionality framework. To this end, let $\mathcal{M}$ be the set of $K$ candidate models; each model is indicated either by $\ell$ or $\kappa$, $1\leq \ell,\kappa\leq K$. Define the subsets $M_1$ and $M_2$ as follows:

equation[equation omitted — 218 chars of source]
equation[equation omitted — 208 chars of source]

In short, for a given forecast horizon $h$, $M_1$ contains the models with the minimum $\operatorname{\textbf{MI}}_h$ whereas in $M_2$ we are minimizing $\operatorname{\textbf{VI}}_h$ among the candidates models in $M_1$. The definition of efficiency used in our framework is the same as that of HSU2019:

definitionGiven a sample of size $n$, a model selection criterion is said to be asymptotically efficient if it selects the model $\hat{\ell}_h$ such that $$\lim_{n\to\infty}\Pr \left( \hat\ell_h \in M_2\right) = 1.$$
remarkAlternative definitions of asymptotic efficiency for model selection are available. For instance, in the framework of linear stationary processes, SHI1980 defines the Mean Efficiency when a criterion attains asymptotically a lower bound for the sum of squared prediction errors. Also, the notion of Approximate Efficiency is given in SHI1984. In LI1987, a criterion that depends upon the ratio between loss functions is introduced. This latter definition is similar to the Loss Efficiency proposed in SHA1997.

The $\operatorname{VMRIC}$ selects the model with the smallest variability index among those that achieve the best goodness of fit. Hence, the selected model $\hat\ell_h$ is such that:

equation[equation omitted — 298 chars of source]

In the next Theorem we show that the $\operatorname{VMRIC}$ is an asymptotic efficient model selection criterion in the sense of Definition (ref).

theoremAssume that for each $ 1 \leq \ell \leq K$, $0\leq s \leq h-1$, Theorem (ref) holds and let $\hat\ell_h$ be the model selected by the $\operatorname{VMRIC}$. Then we have that: \begin{align*} \lim_{n\rightarrow \infty} \Pr \left( \hat{\ell}_h \in M_2\right) &= 1, \end{align*} namely, the $\operatorname{VMRIC}$ is asymptotically efficient in the sense of Definition (ref).

Example: a misspecified bivariate AR(2) model

The aim of this section is twofold. First, we assess the goodness of the theoretical derivations and the finite sample behaviour of the method of moments estimator for the $\operatorname{VMRIC}$. Second, we show that in presence of misspecification the $\operatorname{VMRIC}$ leads to selecting the best predictive model (i.e. is asymptotically efficient) whereas both the $\operatorname*{AIC}$ and the $\operatorname*{BIC}$ fail to do so. In order to achieve the goals we consider a bivariate AR(2) DGP and use two misspecified predictive models for it: in Model 1 there is one omitted lagged predictor, whereas Model 2 uses only one non-informative predictor. We derive theoretically the Mean Square Prediction Error matrix and the $\operatorname{VMRIC}$ for both models and these show that Model 1 is a better predictive model over Model 2. Based on this, we assess the ability of the $\operatorname{VMRIC}$, and of the multivariate versions of the $\operatorname*{AIC}$ and $\operatorname*{BIC}$ to select the best model (Model 1) in finite samples and for different parameterizations. We start by providing the definition of misspecification. Consider an increasing sequence of $\sigma$-fields, $\left\{ \mathcal{G}_t \right\}$ such that $\sigma\left( \mathbf{x}_s, s\leq t \right) \subseteq \mathcal{G}_t \subseteq \mathcal{F} $, where $\left\{ \mathbf{x}_t \right\}$ is an $m$-dimensional weakly stationary process defined over the probability space $\left(\Omega, \mathcal{F}, \mathbb{P} \right)$.

definitionThe $h$-step ahead forecasting model: \begin{equation} \mathbf{y}_{t+h} = \boldsymbol{\beta}_h^\top \mathbf{x}_t + \boldsymbol{\varepsilon}_t^{(h)}, \end{equation} is correctly specified with respect to an increasing sequence of $\sigma$-fields, $\left\{ \mathcal{G}_t \right\}$ if \begin{equation} E\left[\mathbf{y}_{t+h} \mid \mathcal{G}_t \right] = \boldsymbol{\beta}_h^\top \mathbf{x}_t \ a.s., \ \forall \ -\infty < t < \infty. \end{equation} Otherwise, it is misspecified.
remarkThe presence of misspecification implies that: $E\left[ \boldsymbol{\varepsilon}_t^{(h)} x_t \right] = \boldsymbol{0}$, while it is possible to have $E\left[ \boldsymbol{\varepsilon}_t^{(h)} x_s \right] \neq \boldsymbol{0}$, for $s\neq t$, i.e. we have null simultaneous correlation and non-null cross correlation between the forecasting error vector and the regressor.

Consider the following DGP :

equation[equation omitted — 102 chars of source]

where $\boldsymbol{a}\neq \boldsymbol{0}$, $\left\{ \boldsymbol{\varepsilon}_t \right\}$ is a sequence of independent and identically distributed (hereafter i.i.d.) bivariate random vectors with $E\left[ \boldsymbol{\varepsilon}_1 \right] = \boldsymbol{0} $, $E\left[ \boldsymbol{\varepsilon}_1 \boldsymbol{\varepsilon}_1^{\top} \right] >\boldsymbol{0} $ and $w_t$ is the following scalar AR(2) process:

equation[equation omitted — 80 chars of source]

where $\phi_1 \phi_2 \neq 0$ , $\left\{ \delta_t \right\}$ a sequence of i.i.d. random variables independent of $\left\{ \boldsymbol{\varepsilon}_t \right\}$ such that $$E\left[ \delta_1 \right] = 0\quad\text{and}\quad E\left[ \delta_1^2 \right] = 1 - \phi_2^2 - \left\{ \phi_1^2 \frac{1+\phi_2}{1-\phi_2} \right\}.$$ Hence, we obtain $E\left[ w_t^2 \right] \equiv \gamma_w(0) = 1$, where $\gamma_w(j) = E\left[ w_t w_{t+j} \right]$ is the $j$-th lag autocovariance of $w$. We consider the correctly specified $2$-step ahead forecasting model:

align[align omitted — 270 chars of source]

where $\boldsymbol{\varepsilon}_t^{*(2)} = \boldsymbol{\varepsilon}_{t+2} + \boldsymbol{a} \delta_{t+1}$. It can be easily proved that $E\left[ \boldsymbol{\varepsilon}_t^{*(2)} w_{t-j} \right] = \boldsymbol{0}$ for $j \geq 0$. Now, consider the following misspecified model, Model 1:

align*[align* omitted — 275 chars of source]

The forecasting error results:

equation[equation omitted — 190 chars of source]
remarkIn presence of misspecification $E\left[ \boldsymbol{\varepsilon}_{t}^{(2)} w_{t} \right] = \boldsymbol{0}$, whereas $E\left[ \boldsymbol{\varepsilon}_{t}^{(2)} w_{t-j} \right] \neq \boldsymbol{0}$ for $j\neq0$. We show that this occurs in our case: \begin{align} E\left[ \boldsymbol{\varepsilon}_{t}^{(2)} w_{t-j} \right]& = -\boldsymbol{a} \frac{\phi_2}{1-\phi_2} \left\{ \phi_1 E\left[ w_t w_{t-j} \right] - \left( 1- \phi_2 \right) E\left[ w_{t-1} w_{t-j} \right] \right\}\nonumber\\ & = -\boldsymbol{a} \frac{\phi_2}{1-\phi_2} \left\{\gamma_w(j+1) - \gamma_w (j-1)\right\},\nonumber \end{align} which is zero if $j=0$, otherwise this is generally not the case.

We compute the theoretical value of the $\operatorname{VMRIC}$ by using Eq. ((ref)). After some routine algebra, we get:

equation[equation omitted — 322 chars of source]

which highlights how the variance-covariance matrix of the $2$-step ahead forecast vector is equal to the {DGP}'s variance-covariance plus a bias term that depends upon the misspecification considered. Now we focus on the variability index $\operatorname{\textbf{VI}}$. We get

align[align omitted — 320 chars of source]

and

align[align omitted — 232 chars of source]

where

align*[align* omitted — 159 chars of source]

Following Eq. ((ref)), the results from Eq. ((ref)), ((ref)), and ((ref)), deliver the $\operatorname{VMRIC}$ for this case. Now we consider a second misspecified model, Model 2:

equation[equation omitted — 85 chars of source]

where $z_t$ is a weakly stationary linear AR(1) process independent of $w_t$:

equation[equation omitted — 63 chars of source]

with $\psi_1 \in (-1, 1)$, and $\left\{ \upsilon_t \right\}$ is a sequence of i.i.d. random variables independent of both the error terms $\left\{ \delta_t\right\}$ and $\left\{\boldsymbol{\varepsilon}_{t}\right\}$ such that $E[\upsilon_t] = 0$ and $E[\upsilon_t^2] = 1 - \psi_1^2$, delivering $E[z_t] = 0$ and $E[z_t^2]=1$. Thus, $z_t$ is uncorrelated with both $w_t$ and $\mathbf{y}_t$, therefore $\boldsymbol{\rho}=\boldsymbol{0}$. The forecasting error in this case results $\boldsymbol{\eta}_t^{(2)} = \boldsymbol{a} w_{t+1} + \boldsymbol{\varepsilon}_{t+2}$. Following similar arguments as above we obtain ${\operatorname{\textbf{MI}}}$ and ${\operatorname{\textbf{VI}}}$ for Model 2:

align[align omitted — 279 chars of source]

As mentioned above, Model 1 is misspecified since it omits the lagged predictor $w_{t-1}$, while Model 2 only includes the non-informative predictor $z_t$.

Finite sample performance

First, we compare the above theoretical derivations with their sample counterpart. We consider three different parameterizations, presented in Table (ref). Also, $\alpha_n = n^\alpha$ with $\alpha=0.85$. Note that, in order for Eq. ((ref)) to hold, $\alpha$ must range in $(0, 1)$. Further experiments showed that results are fairly robust if reasonable values of $\alpha$ are selected. For an empirical method to determine it, see hsu2019_supp. We take the following variance/covariance matrix for the innovations: $$E[\boldsymbol{\varepsilon}_t\boldsymbol{\varepsilon}_t^{\top}] = \left[

array[array omitted — 37 chars of source]

\right].$$ \par We compute both the {$\operatorname{VMRIC}$} for Model 1 and Model 2, and estimate the $\hat{\operatorname{VMRIC}}$ and $\hat{\operatorname{VMRIC}}$ on a large sample of $n=10^6$ observations. The results are shown in Table~\ref{tab:2} for the two models, where the theoretical $\operatorname{VMRIC}$ (rows 1 and 3) is compared with the estimated one (rows 2 and 4). The results seem to confirm the consistency of the estimator shown in Eq.~(\ref{eqn:VMRIC_hat}). Clearly, the $\operatorname{VMRIC}$ of Model 1 is consistently smaller than that of Model 2 and indicates its superior predictive capability.

table[table omitted — 463 chars of source]
table[table omitted — 708 chars of source]

The finite sample behaviour of the method of moments estimator of the $\operatorname{VMRIC}$ can be further appreciated in Table (ref) where we show their bias and Mean Squared Error (MSE). The results are based upon 1000 Monte Carlo replications and seem to indicate a rate of convergence of the order of $n^{-1}$. In Table (ref), we show the percentages of correct model selection by the $\operatorname{VMRIC}$, compared with the multivariate version of the $\operatorname*{AIC}$ and $\operatorname*{BIC}$ for the three parameterizations of Table (ref). For a sample size of $n=100$, both the $\operatorname*{AIC}$ and $\operatorname*{BIC}$ select the best predictive model in about 50% of the cases and relying upon them is tantamount to tossing a fair coin. In such a case, the $\operatorname{VMRIC}$ selects the correct model in about $80\%$ of the cases and reaches $100\%$ for $n=1000$. On the contrary, for Case 3, both the $\operatorname*{AIC}$ and $\operatorname*{BIC}$ cannot go above $64\%$ for a sample size as large as $n=10000$ observations and this is a general indication of their lack of asymptotic efficiency.

table[table omitted — 1,191 chars of source]
table[table omitted — 780 chars of source]

Proofs

In this section we detail the proofs of the three theorems. Hereafter all the derivations hold for any fixed $h\geq 1$; for the sake of presentation we write $ \boldsymbol{\varepsilon}_{t}$ instead of $ \boldsymbol{\varepsilon}_{t}^{(h)}$. Remember that $\{l_n\}$ indicates an increasing sequence of positive integers such that:

equation[equation omitted — 99 chars of source]

and define $a = n-l_n-h$ and $b=n-l_n-h+1$.

Proof of Theorem (ref)

The proof of Theorem (ref) relies upon four propositions.

propositionUnder assumptions of Theorem (ref), it holds that: \begin{equation} N(\mathrm{I}) = (\mathrm{III}) + o(1), \end{equation} where \begin{align*} (\mathrm{I}) =- \operatorname{E}\left[x_n \hat{R}^{-1} \left( \boldsymbol{\hat{\Sigma}} \boldsymbol{\varepsilon}_{n}^\top + \boldsymbol{\varepsilon}_{n} \boldsymbol{\hat{\Sigma}}^\top \right) \right],\quad (\mathrm{III})= \ - \operatorname{E}\left[x_n R^{-1} \left( \boldsymbol{\hat{\Sigma}}_A \boldsymbol{\varepsilon}_{n}^\top + \boldsymbol{\varepsilon}_{n} \boldsymbol{\hat{\Sigma}}_A^\top\right) \right], \end{align*} with $\boldsymbol{\hat{\Sigma}} = \left( N^{-1} \sum_{t=1}^{N}x_t \boldsymbol{\varepsilon}_{t} \right)$ and $\boldsymbol{\hat{\Sigma}}_A = \sum_{t=1}^{N}\boldsymbol{\varepsilon_{t}} x_t$.
proofLet $\boldsymbol{A}_1 = \sum_{t=1}^{N} \left(\boldsymbol{\varepsilon}_t x_t\right) \boldsymbol{\varepsilon}_n^\top$ and note that \begin{align} \left\lVert(\mathrm{I})-(\mathrm{III}) \right\rVert &= \left\lVert\operatorname{E}\left[x_n \left(\hat{R}^{-1}-R^{-1}\right)\left(\boldsymbol{A}_1+\boldsymbol{A}_1^\top\right)\right] \right\rVert. \end{align} By using standard properties of the norm, ((ref)) follows upon proving that \begin{equation} \left\lVert \operatorname{E}\left[x_n \left(\hat{R}^{-1}-R^{-1}\right)\boldsymbol{A}_1^\top\right]\right\rVert=o(1). \end{equation} Let \begin{equation} \tilde{R} = \left(n-l_n\right)^{-1} \sum_{t=1}^{n-l_n} x_t^2. \end{equation} By adding and subtracting $ \boldsymbol{\varepsilon}_n x_n \left(\tilde{R}^{-1} \left[\sum_{t=1}^{N}\left(\boldsymbol{\varepsilon}_t x_t\right)\right]^\top \right)$, we have \begin{align} \operatorname{E}\left[x_n \left(\hat{R}^{-1}-R^{-1}\right)\boldsymbol{A}_1^\top\right]&=\operatorname{E}\left[\boldsymbol{\varepsilon}_n x_n \left(\hat{R}^{-1}-\tilde{R}^{-1}\right) \sum_{t=1}^{N} \boldsymbol{\varepsilon}_t^\top x_t\right]\nonumber\\ &+ \operatorname{E}\left[\boldsymbol{\varepsilon}_n x_n \left(\tilde{R}^{-1}-R^{-1}\right) \sum_{t=1}^{N} \boldsymbol{\varepsilon}_t^\top x_t\right]\nonumber\\ &= \operatorname{E}\left[\boldsymbol{\varepsilon}_n x_n \left(\hat{R}^{-1}-\tilde{R}^{-1}\right) \left(\sum_{t=1}^{N} \boldsymbol{\varepsilon}_t x_t\right)^\top\right]\\ &+ \operatorname{E}\left[\boldsymbol{\varepsilon}_n x_n \left(\tilde{R}^{-1}-R^{-1}\right) \left(\sum_{t=b}^{N} \boldsymbol{\varepsilon}_t x_t\right)^\top\right]\\ &+ \operatorname{E}\left[\boldsymbol{\varepsilon}_n x_n \left(\tilde{R}^{-1}-R^{-1}\right) \left(\sum_{t=1}^{a} \boldsymbol{\varepsilon}_t x_t\right)^\top\right]. \end{align} We show below that the norms of ((ref)), ((ref)) and ((ref)) are asymptotically negligible. Focus on the first one: by combining conditions (C3), (C4), Lemma (ref), and H\"older's inequality, it follows that $\left\lVert(\ref{eqn:1})\right\rVert$ is bounded by \begin{align*} \operatorname{E} \left[\left\lVert\boldsymbol{\varepsilon}_n x_n \left(\hat{R}^{-1}-\tilde{R}^{-1}\right) \left(\sum_{t=1}^{N} \boldsymbol{\varepsilon}_t x_t\right)^\top\right\rVert\right] &\leq \operatorname{E}\left[\left\lVert\boldsymbol{\varepsilon}_n\right\rVert^6\right]^{\frac{1}{6}} \operatorname{E}\left[\left\lvert x_n \right\rvert^{6} \right]^{\frac{1}{6}} \operatorname{E}\left[\left\lvert \hat{R}^{-1}-\tilde{R}^{-1} \right\rvert^{3} \right]^{\frac{1}{3}}\\ & \times \operatorname{E}\left[\left\lVert N^{\frac{1}{2}}N^{-\frac{1}{2}}\sum_{t=1}^{N} \boldsymbol{\varepsilon}_t x_t \right\rVert ^{3}\right]^{\frac{1}{3}} = O\left(\frac{l_n}{n^{1/2}}\right), \end{align*} which converges to zero due to the definition of $l_n$ in ((ref)). Similarly, we have that $\left\lVert(\ref{eqn:2})\right\rVert$ is bounded by \begin{align*} \operatorname{E}\left[\left\lVert\boldsymbol{\varepsilon}_n\right\rVert^6\right]^\frac{1}{6} \operatorname{E}\left[\left\lvert x_n\right\rvert^6\right]^\frac{1}{6} \operatorname{E}\left[\left\lvert\tilde{R}^{-1}-R^{-1}\right\rvert^3\right]^\frac{1}{3} \operatorname{E}\left[\left\lVert\left(\left(N-b+1\right)^{\frac{1}{2}}\left(N-b+1\right)^{-\frac{1}{2}}\sum_{t=b}^{N} \boldsymbol{\varepsilon}_t x_t\right)^\top\right\rVert^3\right]^\frac{1}{3}. \end{align*} which is an $O\left(n^{-1/2}l_n\right)$ thereby vanishing asymptotically. Lastly, Condition (C6), Lemma (ref), and H\"older's inequality imply that $\left\lVert(\ref{eqn:2})\right\rVert$ is bounded by \begin{align*} \operatorname{E}\left[\left\lVert\operatorname{E}\left[ \boldsymbol{\varepsilon}_t x_t \mid \mathcal{F}_{t-l_n}\right]\right\rVert^3\right]^\frac{1}{3} \operatorname{E}\left[\left\lvert\tilde{R}^{-1}-R^{-1}\right\rvert^{3}\right]^\frac{1}{3} \operatorname{E}\left[\left\lVerta^{\frac{1}{2}} a^{-\frac{1}{2}}\sum_{t=1}^{a} \boldsymbol{\varepsilon}_t^\top x_t\right\rVert^3\right]^\frac{1}{3}= o\left(1\right) \end{align*} and this completes the proof.
propositionUnder assumptions of Theorem (ref), it holds that: \begin{equation} N(\mathrm{II}) = (\mathrm{IV}) + o(1), \end{equation} where \begin{align*} (\mathrm{II}) = \operatorname{E}\left[\hat{R}^{-1} \boldsymbol{\hat{\Sigma}}x_n x_n \boldsymbol{\hat{\Sigma}}^\top \hat{R}^{-1}\right],\quad (\mathrm{IV}) = \ \operatorname{E}\left[ \boldsymbol{\hat{\Sigma}}_B R^{-1} \boldsymbol{\hat{\Sigma}}_B^\top \right], \end{align*} with $\boldsymbol{\hat{\Sigma}}$ being defined in Proposition (ref) and $\boldsymbol{\hat{\Sigma}}_B = N^{-\frac{1}{2}} \sum_{t=1}^{N}\boldsymbol{\varepsilon_{t}} x_t.$
proofLet $\boldsymbol{M}_1 = x_n \left(\hat{R}^{-1} - R^{-1}\right) \boldsymbol{\hat{\Sigma}}_B$ and $\boldsymbol{M}_2 = x_n R^{-1} \boldsymbol{\hat{\Sigma}}_B$. Since \begin{align*} N \left(\mathrm{II}\right)= \operatorname{E}\left[\left(\boldsymbol{M}_1+ \boldsymbol{M}_2\right) \left(\boldsymbol{M}_1+ \boldsymbol{M}_2\right)^\top\right] &= \operatorname{E}\left[\boldsymbol{M}_1 \boldsymbol{M}_1^\top\right] + \operatorname{E}\left[\boldsymbol{M}_2\boldsymbol{M}_2^\top\right]\\ & + \operatorname{E}\left[\boldsymbol{M}_1 \boldsymbol{M}_2^\top\right] + \operatorname{E}\left[\boldsymbol{M}_2 \boldsymbol{M}_1^\top\right] \end{align*} the proof of ((ref)) reduces to show that the following conditions hold: \begin{align} \left\lVert\operatorname{E}\left[ \boldsymbol{M}_1 \boldsymbol{M}_1^\top \right]\right\rVert &= o\left(1\right),\\ \left\lVert\operatorname{E}\left[ \boldsymbol{M}_1 \boldsymbol{M}_2^\top \right]\right\rVert &= o\left(1\right), \\ \left\lVert\operatorname{E}\left[ \boldsymbol{M}_2 \boldsymbol{M}_2^\top \right]- (\mathrm{IV})\right\rVert&= o\left(1\right). \end{align} Conditions ((ref)) and ((ref)) readily follow from Assumptions (C3) and (C4), Lemma (ref), the non singularity of $R$ and H\"older's inequality: \begin{align*} \operatorname{E}\left[\left\lVert\boldsymbol{M}_1 \boldsymbol{M}_1^\top \right\rVert\right] &= \operatorname{E}\left[\left\lVertx_n^2 \left(\hat{R}^{-1} - R^{-1}\right)^2 \boldsymbol{\hat{\Sigma}}_B \boldsymbol{\hat{\Sigma}}_B^\top \right\rVert\right] \leq \left(\operatorname{E}\left[\left\lvert x_n \right\rvert^{10}\right]\right)^\frac{1}{5} \left(\operatorname{E}\left[\left\lvert\hat{R}^{-1}-R^{-1}\right\rvert^5\right]\right)^\frac{2}{5}\\ &\times\left(\operatorname{E}\left[\left\lVert\boldsymbol{\hat{\Sigma}}_B\right\rVert^5\right]\right)^\frac{2}{5}= o\left(1\right);\\ \operatorname{E}\left[ \left\lVert\boldsymbol{M}_1\boldsymbol{M}_2^\top\right\rVert \right] &= \operatorname{E}\left[\left\lVertx_n^2 \left(\hat{R}^{-1} - R^{-1}\right) R^{-1} \boldsymbol{\hat{\Sigma}}_B \boldsymbol{\hat{\Sigma}}_B^\top \right\rVert\right] \leq \left(\operatorname{E}\left[\left\lvert x_n\right\rvert^{10}\right]\right)^\frac{1}{5} \left(\operatorname{E}\left[\left\lvert\hat{R}^{-1}-R^{-1}\right\rvert ^5\right]\right)^\frac{1}{5}\\ &\times \left(\operatorname{E}\left[\left\lvert R^{-1} \right\rvert^5\right]\right)^\frac{1}{5} \left(\operatorname{E}\left[\left\lVert\boldsymbol{\hat{\Sigma}}_B\right\rVert^5\right]\right)^\frac{2}{5}= o\left(1\right). \end{align*} As concerns ((ref)), decompose the vector $\boldsymbol{\hat{\Sigma}}_B$ as follows: \begin{align*} \boldsymbol{\hat{\Sigma}}_B = N^{-\frac{1}{2}} \sum_{t=1}^{N}\boldsymbol{\varepsilon_{t}} x_t = \boldsymbol{u} + \boldsymbol{w}, \quad with \quad \boldsymbol{u} = N^{-\frac{1}{2}} \sum_{t=1}^{a}\boldsymbol{\varepsilon_{t}} x_t \quadand\quad \boldsymbol{w} = N^{-\frac{1}{2}} \sum_{t=b}^{N}\boldsymbol{\varepsilon_{t}} x_t. \end{align*} Hence, we have that \begin{align*} \operatorname{E}\left[\boldsymbol{M}_2\boldsymbol{M}_2^\top\right]- (\mathrm{IV}) &= \operatorname{E}\left[\boldsymbol{u} R^{-1} x_n x_n R^{-1} \boldsymbol{u}^\top\right] - \operatorname{E}\left[\boldsymbol{u} R^{-1} R R^{-1} \boldsymbol{u}^\top\right]\\ & + \operatorname{E}\left[\boldsymbol{u} R^{-1} x_n x_n R^{-1} \boldsymbol{w}^\top\right] - \operatorname{E}\left[\boldsymbol{u} R^{-1} R R^{-1} \boldsymbol{w}^\top\right]\\ & + \operatorname{E}\left[\boldsymbol{w} R^{-1} x_n x_n R^{-1} \boldsymbol{u}^\top\right] - \operatorname{E}\left[\boldsymbol{w} R^{-1} R R^{-1} \boldsymbol{u}^\top\right]\\ & + \operatorname{E}\left[\boldsymbol{w} R^{-1} x_n x_n R^{-1} \boldsymbol{w}^\top\right] - \operatorname{E}\left[\boldsymbol{w} R^{-1} R R^{-1} \boldsymbol{w}^\top\right]. \end{align*} The law of iterated expectations implies that: \begin{align} &\left\lVert\operatorname{E}\left[\boldsymbol{M}_2 \boldsymbol{M}_2^\top\right]-(\mathrm{IV})\right\rVert\nonumber\\ &\leq \left\lVert\operatorname{E}\left[\boldsymbol{u} R^{-1} \left(\operatorname{E}\left[x_n^2 \mid \mathcal{F}_{n-l_n}\right]-R\right) R^{-1} \boldsymbol{u}^\top\right]\right\rVert\\ &+ \left\lVert\operatorname{E}\left[\boldsymbol{u} R^{-1} \left(\operatorname{E}\left[x_n^2 \mid \mathcal{F}_{n-l_n}\right]-R\right) R^{-1} \boldsymbol{w}^\top\right]\right\rVert\\ &+ \left\lVert\operatorname{E}\left[\boldsymbol{w} R^{-1} \left(\operatorname{E}\left[x_n^2 \mid \mathcal{F}_{n-l_n}\right]-R\right) R^{-1} \boldsymbol{u}^\top\right]\right\rVert\\ &+ \left\lVert\operatorname{E}\left[\boldsymbol{w} R^{-1} \left(\operatorname{E}\left[x_n^2 \mid \mathcal{F}_{n-l_n}\right]-R\right) R^{-1} \boldsymbol{w}^\top\right]\right\rVert. \end{align} By using arguments previously developed, it is easy to see that, under Assumptions (C4) and (C6), ((ref)) -- ((ref)) asymptotically vanish. Therefore, conditions ((ref)) -- ((ref)) are fulfilled and the proof is completed.
propositionUnder assumptions of Theorem (ref), it holds that: \begin{equation} (\mathrm{III}) = - (D) + o\left(1\right), \end{equation} where \begin{align*} (D) &= \operatorname{E}\left[ R^{-1} \left[\sum_{j=h}^{N-1} \left\{\left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right) \left( \boldsymbol{\varepsilon}_1 x_1\right)^\top\right\} \right] \right] \end{align*}
proofThe result readily follows upon noting that, under Assumption (C2) and the weakly stationarity of the process $\{x_t\}$, it holds that: \begin{align*} (\mathrm{III}) & = - \sum_{t=1}^{N} \operatorname{E}\left[R^{-1} \left\{ \left(\boldsymbol{\varepsilon}_t x_t\right) \left(\boldsymbol{\varepsilon}_n x_n\right)^\top + \left(\boldsymbol{\varepsilon}_n x_n\right) \left(\boldsymbol{\varepsilon}_t x_t\right)^\top\right\} \right]\\ & = - \ \sum_{j=h}^{n-1} \operatorname{E}\left[R^{-1} \left\{ \left(\boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right) \left(\boldsymbol{\varepsilon}_1 x_1\right)^\top\right\} \right]\\ &= - \operatorname{E}\left[ R^{-1} \left(\sum_{j=h}^{N-1} \left\{\left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right) \left( \boldsymbol{\varepsilon}_1 x_1\right)^\top\right\} \right) \right] +o(1). \end{align*}
propositionUnder assumptions of Theorem (ref), it holds that: \begin{equation} (\mathrm{IV}) = (1) + \left(Q\right) + \left(D\right) + o(1), \end{equation} where \begin{align*} (1) & = N^{-1} \operatorname{E}\left[ R^{-1} \left\{\sum_{t=1}^{N} \left(\boldsymbol{\varepsilon}_t x_t\right)\left(\boldsymbol{\varepsilon}_t x_t\right)^\top \right\} \right], \\ \left(Q\right) &= \operatorname{E}\left[ R^{-1} \left[\sum_{s=1}^{h-1} \left\{\left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{s+1} x_{s+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{s+1} x_{s+1}\right) \left( \boldsymbol{\varepsilon}_1 x_1 \right)^\top\right\} \right] \right]\\ \left(D\right) &= \operatorname{E}\left[ R^{-1} \left[\sum_{j=h}^{N-1} \left\{\left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right) \left( \boldsymbol{\varepsilon}_1 x_1\right)^\top\right\} \right] \right] \end{align*}
proofLet \begin{equation*} (2) = N^{-1} \operatorname{E}\left[ R^{-1} \left\{\sum_{j=1}^{N-1} \sum_{k=j+1}^{N} \left(\boldsymbol{\varepsilon}_j x_j\right) \left(\boldsymbol{\varepsilon}_k x_k\right)^\top\right\} \right], \end{equation*} and note that $(\mathrm{IV})-(1) = (2) + (2)^\top$. Moreover \begin{align} (2) & = N^{-1} \operatorname{E}\left[ R^{-1} \left\{\sum_{j=1}^{N-1} \left( N-j \right) \left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top \right\} \right] \nonumber\\ & = \operatorname{E}\left[ R^{-1} \left\{\sum_{j=1}^{N-1} \left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top \right\} \right] \\ & - N^{-1} \operatorname{E}\left[ R^{-1} \left\{\sum_{j=1}^{N-1} j\left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top \right\} \right]. \end{align} Assumptions (C2) implies that ((ref)) is $o(1)$. Since ((ref)) can be written as \begin{equation*} \operatorname{E}\left[ R^{-1} \left\{\sum_{s=1}^{h-1} \left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{s+1} x_{s+1}\right)^\top \right\} \right] + \operatorname{E}\left[ R^{-1} \left\{\sum_{j=h}^{N-1} \left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top \right\} \right], \end{equation*} then $(\ref{eqn:a})+(\ref{eqn:a})^\top=(Q)+(D)$ and this completes the proof.

Proof of Theorem (ref)

We prove that:

align[align omitted — 823 chars of source]

Since $$\left( \boldsymbol{\hat{\beta}}-\boldsymbol{\beta}\right) = \hat{R}^{-1} \left( N^{-1} \sum_{t=1}^{N} x_t \boldsymbol{y}_{t+h} \right) - \boldsymbol{\beta} = \hat{R}^{-1} \left(N^{-1} \sum_{t=1}^{N} x_t \boldsymbol{\varepsilon}_t \right),$$ routine algebra implies that:

equation[equation omitted — 281 chars of source]

By applying Propositions (ref) -- Propositions (ref), we have:

align*[align* omitted — 374 chars of source]

The proof is completed upon noting that $(1)=(\ref{eq:main1})$ and $(Q)=(\ref{eq:main2})$.

Proof of Theorem (ref)

We start proving that

equation[equation omitted — 114 chars of source]

Note that

align*[align* omitted — 298 chars of source]

hence, it holds that $\hat{\operatorname{\textbf{MI}}}_h - \operatorname{\textbf{MI}}_h$ equals

align[align omitted — 412 chars of source]

Assumption (A1) implies that $(\ref{eq:mi_A})=O_p(n^{-1/2})$ whereas, by combining Assumptions (A3) and (A4) with the non-singularity of $R$ and H\"older's inequality, it can be shown that $(\ref{eq:mi_B})=O_p(n^{-1})$ and hence the proof of ((ref)) is complete. Next, we prove that

equation*[equation* omitted — 92 chars of source]

It suffices to show that

equation[equation omitted — 81 chars of source]

It holds that $\hat{\mathbf{C}}_{h,s}$ is equal to

align[align omitted — 771 chars of source]

We prove that ((ref)) is $o_p(1)$ componentwise. To this end consider: $$\operatorname{E}\left[\left( N-s\right)^{-1}\abs*{\sum_{t=1}^{N-s} x_t^2 x_{t+s} \varepsilon_{t+s,i}}\right],$$ with $\varepsilon_{t+s,i}$ being the $i$-th component of the vector $\boldsymbol{\varepsilon}_{t+s}$. The triangular inequality and H\"older's inequality imply that:

align*[align* omitted — 426 chars of source]

Since $ \hat{\boldsymbol{\beta}}_n(h) - \boldsymbol{\beta}_h = \hat{R}^{-1} \left( N^{-1} \sum_{j=1}^{N} x_j \boldsymbol{\varepsilon}_{j,h} \right)$, by combining Assumptions (A3), (A4) and (A5) with Chebyshev's inequality we obtain that ((ref)) is $o_p(1)$. Similarly, we can verify that ((ref)) and ((ref)) are $o_p(1)$. Lastly, Condition (A2) implies that $(\ref{Cs_A}) = \mathbf{C}_{h,s}+ o_p (1)$, hence ((ref)) is verified and the whole proof is complete.

Proof of Theorem (ref)

By Theorem (ref) the $\operatorname{VMRIC}_h$ defined in ((ref)) can be written as:

equation[equation omitted — 320 chars of source]

Therefore,

equation[equation omitted — 181 chars of source]

and hence

equation[equation omitted — 96 chars of source]

Now, consider two models $\ell_1$ and $\ell_2$ in the candidates set $J_{\ell_1}, J_{\ell_2} \in M_1$ such that $\operatorname{VI}_h(\ell_1)\neq \operatorname{VI}_h(\ell_2)$. We show that

equation[equation omitted — 332 chars of source]

By defining $\operatorname{\textbf{MI}}^*_h$ to be the minimum value of $\operatorname{\textbf{MI}}_h$ over the family of candidate models, we have:

align*[align* omitted — 472 chars of source]

Therefore, for sufficiently large $n$, it holds that:

equation*[equation* omitted — 258 chars of source]

Thus

equation*[equation* omitted — 260 chars of source]

and ((ref)) is verified and implies that

equation[equation omitted — 85 chars of source]

This completes the proof.

Acknowledgments

Greta Goracci acknowledges the support of Libera Università di Bolzano, Grant WW201L (ESAMD).