The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
64,569 characters
A multivariate extension of the Misspecification-Resistant Information Criterion.
\def\spacingset#1{\renewcommand{\baselinestretch}
{#1}\small\normalsize} \spacingset{1}
\if00
{
\title{\bf A multivariate extension of the Misspecification-Resistant Information Criterion.}
\author{Gery Andrés Díaz Rubio\\
Department of Statistical Sciences, University of Bologna, Italy\\
and \\
Simone Giannerini \\
Department of Statistical Sciences, University of Bologna, Italy\\
and \\
Greta Goracci \\
Faculty of Economics and Management, Free University of Bolzano-Bozen, Italy
}
\maketitle
} \fi
\if10
{
\bigskip
\bigskip
\bigskip
\begin{center}
{\LARGE\bf A multivariate extension of the Misspecification-Resistant Information Criterion.}
\end{center}
\medskip
} \fi
\begin{abstract}
The Misspecification-Resistant Information Criterion (MRIC) proposed in [H.-L. Hsu, C.-K. Ing, H. Tong: \emph{On model selection from a finite family of possibly misspecified time series models}. The Annals of Statistics. 47 (2), 1061--1087 (2019)] is a model selection criterion for univariate parametric time series that enjoys both the property of consistency and asymptotic efficiency. In this article we extend the MRIC to the case where the response is a multivariate time series and the predictor is univariate. The extension requires novel derivations based upon random matrix theory. We obtain an asymptotic expression for the mean squared prediction error matrix, the vectorial MRIC and prove the consistency of its method-of-moments estimator. Moreover, we prove its asymptotic efficiency. Finally, we show with an example that, in presence of misspecification, the vectorial MRIC identifies the best predictive model whereas traditional information criteria like AIC or BIC fail to achieve the task.
\end{abstract}
\noindent
{\it Keywords:} information criteria, model selection,multivariate time series, Mean Square Prediction Error.
\vfill
\newpage
\spacingset{1.45}
\section{Introduction}\label{sec1}
The model selection step is a fundamental task in statistical modelling and its implementation typically depends upon the objective of the exercise. In the time series framework the focus is on either forecasting future values or describing/controlling the process that has generated the data (DGP). A good model selection criterion must feature a good ability to identify the model with the ``best'' fit to future values, in a specified sense. In particular, in the parametric time series framework, we can identify two main properties. The first one is consistency, i.e., the ability to select the true DGP with probability one as the sample size diverges. This assumes that a true model exists and is among the set of candidate models. If either the set of candidate models does not contain the true DGP, or, for some reason, a true model cannot be postulated, then a selection criterion should be asymptotically efficient, for instance, in the mean square sense, i.e. it minimizes the mean squared prediction error as the sample size diverges. Starting from the seminal work of Akaike, \cite{AKA1973} a plethora of model selection criteria has been proposed. These include Akaike's AIC \cite{AKA1973, AKA1974}, Schwarz's Bayesian Information Criterion (BIC) \cite{SCH1978}, and Rissanen's Minimum Description Length (MDL) \cite{RIS1978}. Such criteria paved the way for various extensions dealing with different unsolved issues. For instance, the AIC is efficient but not consistent (i.e. it leads to select overfitting models), whereas the BIC is consistent but not efficient, see \cite{HSU2019} for a discussion.
\par
A recent development for model selection in possibly misspecified parametric time series models in the fixed-dimensionality setting is given by the Misspecification-Resistant Information Criterion (hereafter $\operatorname{MRIC}$) \cite{HSU2019}. Fixed-dimensionality means that the number of observations increases to infinity while the number of ‘true’ parameters is finite. In this respect, the MRIC provides a solution to the original research question of Akaike: it enjoys both consistency, in case the true model is included as a candidate, and asymptotic efficiency when a true model either cannot be assumed or is not included. Moreover, when the number of variables in the model grows with the sample size, the $\operatorname{MRIC}$ can achieve asymptotic efficiency, without the need for additional criteria. Finally, in the high-dimensional setting, the $\operatorname{MRIC}$ can be used together with appropriate model selection criteria to identify the best predictive models. The $\operatorname{MRIC}$ is based upon the additive decomposition of the mean squared prediction error in a term that depends upon the misspecification level and a term that measures the sampling variability of the predictor. The idea is to select the model with smallest variability among those that minimize the misspecification index.
\par
The appealing properties of the $\operatorname{MRIC}$ make it an ideal tool for omnibus time series model selection but, to date, only the univariate response case has been studied \cite{HSU2019}. In this work we extend the $\operatorname{MRIC}$ to multivariate time series with a single regressor as to obtain the vectorial MRIC (hereafter $\operatorname{VMRIC}$). As it will be clear, such an extension does not easily derive from the univariate case since it requires dealing with the dependence structure within the components of the vector of forecasting error and hence relies upon random matrix theory. Such multivariate extension can be used in all those models where many time series depend upon a single regressor, like for instance, in econometrics, where many interest rates depend upon a single macroeconomic indicator, such as inflation. Other possible applications include dimension reduction and hedging, which is intimately connected to the problem of model selection \cite{BES16}.
\par
The rest of the paper is organized as follows: in Section~\ref{sec:notation} we introduce the notation and in Section~\ref{subsec:MRIC} summarize the available results for the univariate case; in Section~\ref{sec:VMIRC} we extend the MRIC approach to multivariate time series with a single regressor. In particular, in Section~\ref{subsec:MSPE_dec} we obtain the asymptotic decomposition of the Mean Squared Prediction Error (hereafter MSPE) matrix into two parts: the first one is linked to the goodness of fit of the model and the second one depends upon the prediction variance. In Section~\ref{subsec:consistency} we present the $\operatorname{VMRIC}$ and derive a consistent estimator for it, whereas in Section~\ref{subsec:eff}, we prove the asymptotic efficiency of the $\operatorname{VMRIC}$. Section~\ref{sec:example} presents an example to assess the effect of misspecification in the $\operatorname{VMRIC}$ framework. All the proofs are detailed in Section~\ref{sec:proof}. \ref{appendx} contains auxiliary technical lemmas.
\section{Notation and preliminaries}\label{sec:notation}
For each $t$, let $\{\mathbf{x}_t\}$ and $\{\mathbf{y}_t\}$, with $\mathbf{x}_t=(x_{t,1},\dots,x_{t,m})^\top$ and $\mathbf{y}_t=(y_{t,1},\dots,y_{t,w})^\top$, be two weakly
stationary stochastic processes defined over the probability space $\left(\Omega, \mathcal{F}, \mathbb{P}\right)$. When $m=1$ ($w=1$, respectively) we write $x_t$ ($y_t$). Given a vector $\mathbf{v}$ and a matrix $\mathbf{M}$, we use $\left\lVert\mathbf{v}\right\rVert$ and $\left\lVert\mathbf{M}\right\rVert$ to refer to the $\mathcal{L}_2$ vectorial norm and the matrix norm induced by the Euclidean norm, respectively. We write $o(1)$ ($o_p(1)$) to indicate a sequence that converges (in probability) to zero and $O(1)$ ($O_p(1)$) to indicate a sequence that is bounded (in probability). Moreover, let $\{c_n\}$ be a sequence of scalar random variables whereas $\{\mathbf{v}_n\}$ and $\{\mathbf{M}_n\}$ are sequences of random vectors and random matrices, respectively. We adopt the following notation: $\mathbf{v}_n =o_p(c_n)$ if $\left\lVert\mathbf{v}_n\right\rVert/c_n=o_p(1)$; $\mathbf{v}_n=O_p(c_n)$, if $\left\lVert\mathbf{v}_n\right\rVert/c_n=O_p(1)$,
$\mathbf{M}_n =o_p(c_n)$ if $\left\lVert\mathbf{M}_n\right\rVert/c_n=o_p(1)$; $\mathbf{M}_n = O_p(c_n)$ if $\left\lVert\mathbf{M}_n\right\rVert/c_n=O_p(1)$.
For further details on matrix algebra see \cite{SEB2008, HOR2013, ODE2018}, for multivariate time series see \cite{REI1993, LUT2005, TSA2013}, and for asymptotic tools for vector and matrices, see \cite{JIA2010}.
\par
Let $\{(\mathbf{x}_t,\mathbf{y}_t),t \in \{1,\dots,n\}\}$ be the observed sample, and divide the interval $[1,n]$ into the \emph{training set} $[1,N]$ and the \emph{test set} $[N+1,N+h]$, with $h$ being the forecasting horizon. Note that $\mathbf{x}_t$ can contain both endogenous and exogenous variables, therefore, Model~(\ref{eqn:for_mod}) encompasses many different models including, inter alia, VAR and VARX models. We denote $\bar{\mathbf{x}}=n^{-1}\sum_{t=1}^{n}\mathbf{x}_t$ and $\bar{\mathbf{y}}=n^{-1}\sum_{t=1}^{n}\mathbf{y}_t$, i.e. the two sample means. Without loss of generality assume $\operatorname{E}[\mathbf{x}_t]=\operatorname{E}[\mathbf{y}_t]=\boldsymbol 0$. In order to forecast $\mathbf{y}_{n+h}$, $h\geq 1$, we adopt the following $h$-step ahead forecasting Model:
\begin{equation}\label{eqn:for_mod}
\mathbf{y}_{t+h}= \mathbf{B}_h \mathbf{x}_t+\boldsymbol{\varepsilon}_t^{(h)},
\end{equation}
\par\noindent
where $\mathbf{B}_h$ is a $(w\times m)$ matrix and $\boldsymbol{\varepsilon}_t^{(h)}$ is the vector containing the $w$ $h$-step ahead forecast errors; as before, if $w=1$ we write $\varepsilon_t^{(h)}$.
\begin{remark}
Since the model can possibly be misspecified, the prediction error vector $\boldsymbol{\varepsilon}_t^{(h)}$ can be serially correlated, and also correlated with $\mathbf{x}_s$, $s\neq t$. Moreover, the multivariate framework differs from \cite{HSU2019} in different key aspects. For instance, $(i)$ the components of the error vector can be cross-correlated, and $(ii)$ $\mathbf{x}_t \boldsymbol{\varepsilon}_t^{(h)}$ and $\mathbf{x}_k \boldsymbol{\varepsilon}_k^{(h)}$, for $t\neq k$, can also be both serially and cross correlated.
\end{remark}
Define
\begin{equation}\label{eqn:def_R}
\hat{\mathbf{R}}=N^{-1} \sum_{t=1}^{N}\mathbf{x}_t\mathbf{x}_t^\top\qquad\text{ and }\qquad \mathbf{R} =\operatorname{E}[\mathbf{x}_1 \mathbf{x}_1^\top].
\end{equation}
Then, the ordinary least squares estimator (hereafter OLS) of $\mathbf{B}_h$ results:
\begin{equation}\label{eqn:ols}
\hat{\mathbf{B}}_n(h)= \hat{\mathbf{R}}^{-1}\left(N^{-1}\sum_{t=1}^{N} \mathbf{x}_t \mathbf{y}_{t+h}^\top\right).
\end{equation}
When $w=1$, $\mathbf{R}$ and $\mathbf{B}$ become $R$ and $\boldsymbol\beta$, respectively. The prediction of $\mathbf{y}_{n+h}$, $h\geq 1$, is given by
\begin{equation}\label{eqn:for_y}
\hat{\mathbf{y}}_{n+h}=\hat{\mathbf{B}}_n(h)\mathbf{x}_n
\end{equation}
and the corresponding Mean Squared Prediction Error matrix is
\begin{equation}\label{eqn:MSPE}
\operatorname{\textbf{MSPE}}_h=\operatorname{E}\left[(\mathbf{y}_{n+h}-\hat{\mathbf{y}}_{n+h})(\mathbf{y}_{n+h}-\hat{\mathbf{y}}_{n+h})^\top\right].
\end{equation}
\subsection{The MRIC for parametric univariate time series models}\label{subsec:MRIC}
In \cite{HSU2019}, the authors focused on the case $w=1$ and $m\geq1$. Under appropriate conditions, they obtained the following asymptotic decomposition of $\operatorname{MSPE}$:
\begin{align}
\operatorname{MSPE}_h &= \operatorname{E}\left[(y_{n+h}-\hat{y}_{n+h})^2\right] = \operatorname{MI}_h+n^{-1}(\operatorname{VI}_h+o(1)),\label{eqn:MSPE_dec}\\
\text{with}\quad\operatorname{MI}_h &= \operatorname{E}\left[\left(\varepsilon_n^{(h)}\right)^2\right],\quad
\operatorname{VI}_h = \operatorname{tr}\left(\mathbf{R}^{-1}\mathbf{C}_{h,0}\right) +2\sum_{s=1}^{h-1}\operatorname{tr}\left(\mathbf{R}^{-1}\mathbf{C}_{h,s}\right),\nonumber
\end{align}
\noindent
where $\mathbf{C}_{h,s} = \operatorname{E}\left[\mathbf{x}_1\mathbf{x}_{1+s}^\top \varepsilon_1^{(h)} \varepsilon_{1+s}^{(h)}\right]$, $s\geq0$, is the cross-covariance matrix between the regressors and the $h$-step ahead prediction error at lag $s$.
\begin{remark}
The first part of Eq.~(\ref{eqn:MSPE_dec}) is the Misspecification Index ($\operatorname{MI}$), linked to the goodness-of-fit of the model and coincides with the $h$-step ahead prediction error variance. The second component is the Variability Index ($\operatorname{VI}$), which depends upon the variance of the $h$-step ahead predictor, $\hat{y}_{n+h}= \hat{\boldsymbol{\beta}}_n^\top (h) \mathbf{x}_n$, and is also linked to the bias of the estimator of $\boldsymbol{\beta}_h$.
\end{remark}
\noindent
Based upon the above decomposition, the $\operatorname{MRIC}$ is defined as follows:
\begin{align}
\operatorname{MRIC}_h &= \hat{\operatorname{MI}}_h+\frac{\alpha_n}{n}\hat{\operatorname{VI}}_h,\label{eqn:MRIC_uni}
\end{align}
with $\hat{\operatorname{MI}}_h$ and $\hat{\operatorname{VI}}_h$ being the estimators of $\operatorname{MI}_h$ and $\operatorname{VI}_h$ respectively, i.e.:
\[
\hat{\operatorname{MI}}_h = N^{-1}\sum_{t=1}^{N}\left(\hat{\varepsilon}_t^{(h)}\right)^2,\quad
\hat{\operatorname{VI}}_h = \operatorname{tr}\left(\hat R^{-1}\hat{\mathbf{C}}_{h,0}\right) + 2 \sum_{s=1}^{h-1} \operatorname{tr}\left(\hat R^{-1} \hat{\mathbf{C}}_{h,s}\right),
\]
\noindent
where $\hat{\mathbf{C}}_{h,s} = (N-s)^{-1} \sum_{t=1}^{N-s}\mathbf{x}_t\mathbf{x}_{t+s}^\top \hat{\varepsilon}_t^{(h)}\hat{\varepsilon}_{t+s}^{(h)}$ and
$\hat{\varepsilon}_t^{(h)}=y_{t+h}-\hat{\boldsymbol{\beta}}_n(h)\mathbf{x}_t$ is the estimated forecast error; $\alpha_n$ is a penalization term sequence such that, as $n$ increases:
\begin{align}\label{eqn:penalty}
\frac{\alpha_n}{\sqrt{n}}\to+\infty\qquad\text{ and }\qquad \frac{\alpha_n}{n}\to0.
\end{align}
\noindent
It is shown that $\hat{\operatorname{MI}}_h$ and $\hat{\operatorname{VI}}_h$ are consistent estimators of ${\operatorname{MI}}_h$ and ${\operatorname{VI}}_h$, moreover the asymptotic efficiency of the $\operatorname{MRIC}$ is proved. By minimizing this criterion, the model which minimizes $\operatorname{VI}$ among those with minimum $\operatorname{MI}$ is selected. Among other features, the $\operatorname{MRIC}$ is particularly helpful in situations where competing models present the same goodness-of-fit and the same number of parameters.
\begin{remark}
The type of penalty considered in \cite{HSU2019} is similar to that used in \cite[p. 230]{SHI1989} for the correctly specified case.
\end{remark}
\section{A multivariate extension of the MRIC framework}\label{sec:VMIRC}
In this section we extend the $\operatorname{MRIC}$ approach to the case where the response is a multivariate time series ($w\geq2$) and the predictor is univariate ($m=1$), for a generic $h$-step ahead forecast.
\noindent
Hence, Model~(\ref{eqn:for_mod}) reduces to $\mathbf{y}_{t+h}=\boldsymbol{\beta}_h x_t+ \boldsymbol{\varepsilon}_t^{(h)}$, namely:
\begin{equation}\label{eqn:for_mod_case1}
\begin{cases}
y_{t+h,1}= \beta_{h,1}x_t+\varepsilon_{t,1}^{(h)} \\
y_{t+h,2}= \beta_{h,2}x_t+\varepsilon_{t,2}^{(h)}\\
\vdots\\
y_{t+h,w}= \beta_{h,w}x_t+\varepsilon_{t,w}^{(h)}
\end{cases}
\end{equation}
\subsection{Asymptotic decomposition of the MSPE matrix}\label{subsec:MSPE_dec}
We extend the asymptotic representation of the $\operatorname{MSPE}_h$ defined in (\ref{eqn:MSPE_dec}) which is the key step to derive the $\operatorname{VMRIC}$ in this multivariate framework. We rely upon the following assumptions, which are the natural multivariate extensions of those in \cite{HSU2019}.
\begin{assumptions}
\begin{align*}
(\text{C}1) & \quad \exists \ q_1 > 5, 0<K_1<\infty : \ \text{for any } 1 \leq n_1 < n_2 \leq n,\\
& \operatorname{E}\left[
\left\lvert \left( n_2-n_1+1 \right)^{-1/2}
\sum_{t=n_1}^{n_2} x_{t}^2 - \operatorname{E}\left[ x_{t}^2 \right]
\right\rvert^{q_1}
\right]\leq K_1.\\
(\text{C}2) &\quad 1. \ \mathbf{C}_{h,s} = \ \operatorname{E}\left[ \boldsymbol{\varepsilon}_t^{(h)} x_t \left(\boldsymbol{\varepsilon}_{t+s}^{(h)} x_{t+s}\right)^\top \right] \perp t,\\
&\quad 2. \ \operatorname{E}\left[ x_1 x_n \varepsilon_{1,i}^{(h)} \varepsilon_{n,j}^{(h)} \right] = \ o(n^{-1}) \ \forall \ i,j \in\{ 1, \dots, w\}.\\
(\text{C}3) & \quad 1. \sup_{-\infty<t<\infty} \operatorname{E}\left[ | x_t |^{10} \right] < \infty,\\
& \quad 2. \sup_{-\infty<t<\infty} \operatorname{E}\left[ \left\lVert\boldsymbol{\varepsilon}_{t}^{(h)}\right\rVert^{6} \right] < \infty.\\
(\text{C}4) & \quad \exists \ 0 < K_2 < \infty : \ \text{for} \ 1\leq n_1 < n_2 \leq n,
\quad \operatorname{E}\left[\left\lVert\left(n_2-n_1+1\right)^{-\frac{1}{2}} \sum_{t=n_1}^{n_2} \boldsymbol{\varepsilon}_t^{(h)} x_t \right\rVert^5 \right] < K_2.\\
(\text{C}5) & \quad \text{For any } q>0, \ \operatorname{E}\left[ \left|\hat{R}^{-1} \right|^q \right] = O(1).\\
(\text{C}6) & \quad \exists \mathcal{F}_t \subseteq \mathcal{F}, \mathcal{F}_t \ \text{an increasing sequence of } \sigma\text{-fields such that: }\\
& \quad 1. \
x_t \ \text{is} \ \mathcal{F}_t\text{-measurable}\\
& \quad 2. \ \sup_{-\infty<t<\infty} \operatorname{E}\left[ \left|\operatorname{E}\left[ x_t^2 \mid \mathcal{F}_{t-k}\right] - R\right|^3\right] = \ o(1), \text{ as } k \rightarrow\infty,\\
& \quad 3. \ \sup_{-\infty<t<\infty} \operatorname{E}\left[ \left\lVert\operatorname{E}\left[ \boldsymbol{\varepsilon}_{t}^{(h)} x_t \mid \mathcal{F}_{t-k}\right]\right\rVert^3\right] = \ o(1), \text{ as } k \rightarrow\infty.
\end{align*}
\end{assumptions}
\begin{theorem}\label{thm:MSPE_dec_case1}
Under the regularity conditions (C1) -- (C6), the asymptotic expression of the $\operatorname{\textbf{MSPE}}_h$ defined in (\ref{eqn:MSPE}) results
\begin{align}\label{eqn:MSPE_dec_case1}
&N \left\{ \operatorname{E}\left[ \left( \mathbf{y}_{n+h} - \hat{\mathbf{y}}_{n+h} \right) \left( \mathbf{y}_{n+h} - \hat{\mathbf{y}}_{n+h} \right)^\top - \operatorname{E}\left[
\boldsymbol{\varepsilon}_{n}^{(h)} \boldsymbol{\varepsilon}_{n}^{(h)^\top} \right] \right] \right\}\\
& = R^{-1} \operatorname{E}\left[ \left(\boldsymbol{\varepsilon}_{1}^{(h)} x_1\right)\left(\boldsymbol{\varepsilon}_{1}^{(h)} x_1\right)^\top \right]\nonumber\\
&+ R^{-1} \operatorname{E}\left[\sum_{s=1}^{h-1} \left\{\left( \boldsymbol{\varepsilon}_1^{(h)} x_1\right) \left(\boldsymbol{\varepsilon}_{s+1}^{(h)} x_{s+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{s+1}^{(h)} x_{s+1}\right) \left( \boldsymbol{\varepsilon}_1^{(h)} x_1 \right)^\top\right\} \right] + o(1).\nonumber
\end{align}
\end{theorem}
\noindent
\subsection{VMRIC and its consistent estimation}\label{subsec:consistency}
In this section we introduce the $\operatorname{VMRIC}$. Let $\{\alpha_n\}$ be the penalization term sequence defined as in Eq.~(\ref{eqn:penalty}).
\begin{align}
\operatorname{VMRIC}_h &= \left\lVert{\operatorname{\textbf{MI}}}_h\right\rVert + \left\lVert \frac{\alpha_n}{n} {\operatorname{\textbf{VI}}}_h\right\rVert \label{eqn:VMRIC}\\
\text{where } \operatorname{\textbf{MI}}_h &= \operatorname{E}\left[\left(\boldsymbol{\varepsilon}_t^{(h)}\boldsymbol{\varepsilon}_t^{(h)\top} \right)\right],\quad
\operatorname{\textbf{VI}}_h = R^{-1} \left( \mathbf{C}_{h,0} + \sum_{s=1}^{h-1} \left( \mathbf{C}_{h,s} + \mathbf{C}_{h,s}^\top \right) \right),\nonumber\\
\mathbf{C}_{h,s} &= \operatorname{E}\left[\left(x_t\boldsymbol{\varepsilon}_t^{(h)}\right)\left(x_t\boldsymbol{\varepsilon}_t^{(h)}\right)^\top\right].\nonumber
\end{align}
The $\operatorname{VMRIC}$ can be estimated via the method of moments as to obtain:
\begin{align}
\hat\operatorname{VMRIC}_h &\equiv \left\lVert\hat{\operatorname{\textbf{MI}}}_h\right\rVert + \left\lVert \frac{\alpha_n}{n} \hat{\operatorname{\textbf{VI}}}_h \label{eqn:VMRIC_hat}\right\rVert,\\
\text{where } \hat{\operatorname{\textbf{MI}}}_h &= N^{-1}\sum_{t=1}^{N}\left(\hat{\boldsymbol{\varepsilon}}_t \hat{\boldsymbol{\varepsilon}}_t^\top \right),\quad
\hat{\operatorname{\textbf{VI}}}_h = \hat{R}^{-1} \left[ \hat{\mathbf{C}}_{h,0} + \sum_{s=1}^{h-1} \left( \hat{\mathbf{C}}_{h,s} + \hat{\mathbf{C}}_{h,s}^\top \right) \right],\nonumber
\end{align}
and $\hat{\mathbf{C}}_{h,s} = (N-s)^{-1} \sum_{t=1}^{N-s} x_t x_{t+s} \hat{\boldsymbol{\varepsilon}}_t \hat{\boldsymbol{\varepsilon}}_{t+s}^\top$, and $\hat{\boldsymbol{\varepsilon}}_t=\mathbf{y}_{t+h}-\hat{\boldsymbol{\beta}}_n(h) x_t$ is the estimated forecast error vector.
\par
\noindent In Theorem~\ref{thm:MOME_case1} we prove that $\hat{\operatorname{\textbf{MI}}}_h$ and $\hat{\operatorname{\textbf{VI}}}_h$ are consistent estimators of ${\operatorname{\textbf{MI}}}_h$ and ${\operatorname{\textbf{VI}}}_h$, respectively. Theorem~\ref{thm:MOME_case1} relies upon the following assumptions, that are less restrictive with respect to (C1) -- (C6). For further discussions on the assumptions see \cite[][Remark~1--3, p. 1073]{HSU2019}.
\begin{assumptions}
For each $0\leq s \leq h-1$, we assume the following:
\begin{align*}
(\text{A}1) & \quad n^{-1} \sum_{t=1}^{n} \left( \boldsymbol{\varepsilon}_t^{(h)} \boldsymbol{\varepsilon}_t^{(h)\top}\right) = \operatorname{E}\left[ \boldsymbol{\varepsilon}_1^{(h)} \boldsymbol{\varepsilon}_1^{(h)\top} \right] + O_p \left(n^{-1/2}\right)\\%, \text{and any } 1\leq i,j \leq m,\\
(\text{A}2)
&\quad n^{-1} \sum_{t=1}^{n} \left(x_t \boldsymbol{\varepsilon}_t^{(h)}\right) \left(x_{t+s}\boldsymbol{\varepsilon}_{t+s}^{(h)}\right)^\top = \boldsymbol{C}_{h,s} + o_p(1),\\
(\text{A}3)
& \quad n ^{-1/2} \sum_{t=1}^{n} x_t \boldsymbol{\varepsilon}_t^{(h)} = O_p(1).\\
(\text{A}4)
& \quad n ^{-1} \sum_{t=1}^{n} x_t^{2} = R + o_p(1),\\
(\text{A}5)
& \quad \sup_{-\infty < t < \infty} \operatorname{E} \left[\left\|\boldsymbol{\varepsilon}_t^{(h)}\right\|^4\right] +
\sup_{-\infty<t<\infty} \operatorname{E}\left[\abs {x_t}^4 \right] < \infty.
\end{align*}
\end{assumptions}
\begin{theorem}\label{thm:MOME_case1}
If Assumptions (A1) -- (A5) hold, then for the case $w\geq2$, and $m=1$ we obtain:
\begin{align*}
\hat{\operatorname{\textbf{MI}}}_h&= \operatorname{\textbf{MI}}_h +\, O_p (n^{-1/2}),\\
\hat{\operatorname{\textbf{VI}}}_h&= \operatorname{\textbf{VI}}_h +\, o_p (1).
\end{align*}
\end{theorem}
\subsection{Asymptotic efficiency}\label{subsec:eff}
In this section we prove the asymptotic efficiency of the $\operatorname{VMRIC}$ in the fixed dimensionality framework. To this end, let $\mathcal{M}$ be the set of $K$ candidate models; each model is indicated either by $\ell$ or $\kappa$, $1\leq \ell,\kappa\leq K$. Define the subsets $M_1$ and $M_2$ as follows:
\begin{equation}
M_1 = \left\{
\kappa: 1 \leq \kappa \leq K,
\left\lVert\operatorname{\textbf{MI}}_h(\kappa)\right\rVert = \min_{1\leq \ell \leq K}
\left\lVert\operatorname{\textbf{MI}}_h(\ell)\right\rVert
\right\}
\end{equation}
\begin{equation}
M_2 = \left\{
\kappa: \kappa \in M_1,
\left\lVert\operatorname{\textbf{VI}}_h(\kappa)\right\rVert = \min_{\ell \in M_1}
\left\lVert\operatorname{\textbf{VI}}_h(\ell)\right\rVert
\right\}.
\end{equation}
In short, for a given forecast horizon $h$, $M_1$ contains the models with the minimum $\operatorname{\textbf{MI}}_h$ whereas in $M_2$ we are minimizing $\operatorname{\textbf{VI}}_h$ among the candidates models in $M_1$. The definition of efficiency used in our framework is the same as that of \cite{HSU2019}:
\begin{definition}\label{def:eff}
Given a sample of size $n$, a model selection criterion is said to be asymptotically efficient if it selects the model $\hat{\ell}_h$ such that
$$\lim_{n\to\infty}\Pr \left( \hat\ell_h \in M_2\right) = 1.$$
\end{definition}
\begin{remark}
Alternative definitions of asymptotic efficiency for model selection are available. For instance, in the framework of linear stationary processes, \cite{SHI1980} defines the Mean Efficiency when a criterion attains asymptotically a lower bound for the sum of squared prediction errors. Also, the notion of Approximate Efficiency is given in \cite{SHI1984}. In \cite{LI1987}, a criterion that depends upon the ratio between loss functions is introduced. This latter definition is similar to the Loss Efficiency proposed in \cite{SHA1997}.
\end{remark}
The $\operatorname{VMRIC}$ selects the model with the smallest variability index among those that achieve the best goodness of fit. Hence, the selected model $\hat\ell_h$ is such that:
\begin{equation}\label{eq:vmric}
\operatorname{VMRIC}_h \left( \hat{\ell}_h\right) \equiv
\min_{1 \leq \ell \leq K}
\left\lVert \hat{\boldsymbol{\operatorname{MI}}}_h (\ell) \right\rVert +
\min_{ \ell \in M_1}
\left\lVert \frac{C_n}{n} \hat{\boldsymbol{\operatorname{VI}}}_h (\ell)\right\rVert.
\end{equation}
In the next Theorem we show that the $\operatorname{VMRIC}$ is an asymptotic efficient model selection criterion in the sense of Definition~\ref{def:eff}.
\begin{theorem}\label{thm:EFF_case1}
Assume that for each $ 1 \leq \ell \leq K$, $0\leq s \leq h-1$, Theorem~\ref{thm:MOME_case1} holds and let $\hat\ell_h$ be the model selected by the $\operatorname{VMRIC}$. Then we have that:
\begin{align*}
\label{eq:theo3}
\lim_{n\rightarrow \infty} \Pr \left( \hat{\ell}_h \in M_2\right) &= 1,
\end{align*}
namely, the $\operatorname{VMRIC}$ is asymptotically efficient in the sense of Definition~\ref{def:eff}.
\end{theorem}
\section{Example: a misspecified bivariate AR(2) model}\label{sec:example}
The aim of this section is twofold. First, we assess the goodness of the theoretical derivations and the finite sample behaviour of the method of moments estimator for the $\operatorname{VMRIC}$. Second, we show that in presence of misspecification the $\operatorname{VMRIC}$ leads to selecting the best predictive model (i.e. is asymptotically efficient) whereas both the $\operatorname*{AIC}$ and the $\operatorname*{BIC}$ fail to do so. In order to achieve the goals we consider a bivariate AR(2) DGP and use two misspecified predictive models for it: in Model 1 there is one omitted lagged predictor, whereas Model 2 uses only one non-informative predictor. We derive theoretically the Mean Square Prediction Error matrix and the $\operatorname{VMRIC}$ for both models and these show that Model 1 is a better predictive model over Model 2. Based on this, we assess the ability of the $\operatorname{VMRIC}$, and of the multivariate versions of the $\operatorname*{AIC}$ and $\operatorname*{BIC}$ to select the best model (Model 1) in finite samples and for different parameterizations.
\par
We start by providing the definition of misspecification. Consider an increasing sequence of $\sigma$-fields, $\left\{ \mathcal{G}_t \right\}$ such that $\sigma\left( \mathbf{x}_s, s\leq t \right) \subseteq \mathcal{G}_t \subseteq \mathcal{F} $, where $\left\{ \mathbf{x}_t \right\}$ is an $m$-dimensional weakly stationary process defined over the probability space $\left(\Omega, \mathcal{F}, \mathbb{P} \right)$.
\begin{definition}\label{def:miss}
The $h$-step ahead forecasting model:
\begin{equation}
\mathbf{y}_{t+h} = \boldsymbol{\beta}_h^\top \mathbf{x}_t + \boldsymbol{\varepsilon}_t^{(h)},
\end{equation}
is correctly specified with respect to an increasing sequence of $\sigma$-fields, $\left\{ \mathcal{G}_t \right\}$ if
\begin{equation}
E\left[\mathbf{y}_{t+h} \mid \mathcal{G}_t \right] = \boldsymbol{\beta}_h^\top \mathbf{x}_t \ \text{a.s.}, \ \forall \ -\infty < t < \infty.
\end{equation}
Otherwise, it is misspecified.
\end{definition}
\begin{remark}\label{re:miss}
The presence of misspecification implies that: $E\left[ \boldsymbol{\varepsilon}_t^{(h)} x_t \right] = \boldsymbol{0}$, while it is possible to have
$E\left[ \boldsymbol{\varepsilon}_t^{(h)} x_s \right] \neq \boldsymbol{0}$, for $s\neq t$, i.e. we have null simultaneous correlation and non-null cross correlation between the forecasting error vector and the regressor.
\end{remark}
Consider the following DGP :
\begin{equation}\label{eq:DGP}
\mathbf{y}_{t+1} = \boldsymbol{a} w_t + \boldsymbol{\varepsilon}_{t+1},
\end{equation}
where $\boldsymbol{a}\neq \boldsymbol{0}$, $\left\{ \boldsymbol{\varepsilon}_t \right\}$ is a sequence of independent and identically distributed (hereafter i.i.d.) bivariate random vectors with $E\left[ \boldsymbol{\varepsilon}_1 \right] = \boldsymbol{0} $, $E\left[ \boldsymbol{\varepsilon}_1 \boldsymbol{\varepsilon}_1^{\top} \right] >\boldsymbol{0} $ and $w_t$ is the following scalar AR(2) process:
\begin{equation}\label{eq:AR2}
w_t = \phi_1 w_{t-1} + \phi_2 w_{t-2} + \delta_t,
\end{equation}
where $\phi_1 \phi_2 \neq 0$ , $\left\{ \delta_t \right\}$ a sequence of i.i.d. random variables independent of $\left\{ \boldsymbol{\varepsilon}_t \right\}$ such that
$$E\left[ \delta_1 \right] = 0\quad\text{and}\quad E\left[ \delta_1^2 \right] = 1 - \phi_2^2 - \left\{ \phi_1^2 \frac{1+\phi_2}{1-\phi_2} \right\}.$$
Hence, we obtain $E\left[ w_t^2 \right] \equiv \gamma_w(0) = 1$, where $\gamma_w(j) = E\left[ w_t w_{t+j} \right]$ is the $j$-th lag autocovariance of $w$. \par
We consider the correctly specified $2$-step ahead forecasting model:
\begin{align}\label{eqn:fcst_mod_corr}
\mathbf{y}_{t+2} &= \boldsymbol{a} w_{t+1} + \boldsymbol{\varepsilon}_{t+1}, \text{ which leads to } \nonumber\\
\mathbf{y}_{t+2} &= \boldsymbol{a} \phi_1 w_t + \boldsymbol{a} \phi_2 w_{t-1} + \boldsymbol{\varepsilon}_t^{*(2)},
\end{align}
where $\boldsymbol{\varepsilon}_t^{*(2)} = \boldsymbol{\varepsilon}_{t+2} + \boldsymbol{a} \delta_{t+1}$. It can be easily proved that $E\left[ \boldsymbol{\varepsilon}_t^{*(2)} w_{t-j} \right] = \boldsymbol{0}$ for $j \geq 0$.
\par
Now, consider the following \textit{misspecified} model, Model 1:
\begin{align*}
\mathbf{y}_{t+2}& = \boldsymbol{\beta} w_t + \boldsymbol{\varepsilon}_t^{(2)}, \quad \text{with}\quad\boldsymbol{\beta} = \frac{E\left[ \mathbf{y}_{t+2} w_t \right]}{V\left[ w_T \right]} = \boldsymbol{a} \left( \phi_1 + \frac{\phi_1 \phi_2}{1- \phi_2} \right).
\end{align*}
The forecasting error results:
\begin{equation}\label{eqn:fcst_error_miss}
\boldsymbol{\varepsilon}_t^{(2)} = \boldsymbol{\varepsilon}_t^{*(2)} - \boldsymbol{a} \phi_2 \left[ \frac{\phi_1}{1-\phi_2} w_t - w_{t-1} \right].
\end{equation}
\par\noindent
\begin{remark}
In presence of misspecification $E\left[ \boldsymbol{\varepsilon}_{t}^{(2)} w_{t} \right] = \boldsymbol{0}$, whereas $E\left[ \boldsymbol{\varepsilon}_{t}^{(2)} w_{t-j} \right] \neq \boldsymbol{0}$ for $j\neq0$. We show that this occurs in our case:
\begin{align}
E\left[ \boldsymbol{\varepsilon}_{t}^{(2)} w_{t-j} \right]& = -\boldsymbol{a} \frac{\phi_2}{1-\phi_2} \left\{ \phi_1 E\left[ w_t w_{t-j} \right] - \left( 1- \phi_2 \right) E\left[ w_{t-1} w_{t-j} \right] \right\}\nonumber\\
& = -\boldsymbol{a} \frac{\phi_2}{1-\phi_2} \left\{\gamma_w(j+1) - \gamma_w (j-1)\right\},\nonumber
\end{align}
which is zero if $j=0$, otherwise this is generally not the case.
\end{remark}
We compute the theoretical value of the $\operatorname{VMRIC}$ by using Eq.~(\ref{eqn:VMRIC}).
After some routine algebra, we get:
\begin{equation}\label{eqn:example_mi2}
\operatorname{\textbf{MI}} = E\left[
\boldsymbol{\varepsilon}_{n}^{(2)}
\boldsymbol{\varepsilon}_{n}^{(2)^\top}
\right] = \boldsymbol{\sigma}^2_\varepsilon +
\boldsymbol{a} \boldsymbol{a}^\top
\left[
\sigma^2_{\delta}
+ \phi_2^2
\left(
1 - \gamma_w^2 (1)
\right)
\right],
\end{equation}
which highlights how the variance-covariance matrix of the $2$-step ahead forecast vector is equal to the {DGP}'s variance-covariance plus a bias term that depends upon the misspecification considered.
\par
Now we focus on the variability index $\operatorname{\textbf{VI}}$. We get
\begin{align}\label{eqn:example_c20}
\mathbf{C}_{2,0} = \boldsymbol{\sigma}^2_{\boldsymbol{\varepsilon}} +
\boldsymbol{a} \boldsymbol{a}^\top \left\{
\sigma^2_{\delta}
+ \phi^2_2 \left(
\gamma_w(1)^2 E\left[w_t^4\right]- 2 \gamma_w(1)E\left[w_t^3 w_{t-1}\right] +
E\left[w_t^2 w_{t-1}^2\right]
\right)
\right\}
\end{align}
and
\begin{align}\label{eqn:example_c21}
\mathbf{C}_{2,1} =
\boldsymbol{a} \boldsymbol{a}^\top \gamma_w(1)\left(
b_1 E\left[w_{t-1}^3 w_{t-2}\right] +
b_2 E\left[w_{t-1} w_{t-2}^3 \right] +
b_3 E\left[w_{t-1}^2 w_{t-2}^2\right]
\right),
\end{align}
where
\begin{align*}
b_1= 2\phi_1\phi_2\gamma_w(1) - \phi_2,\qquad
b_2= -\phi_2^2, \qquad
b_3= \phi_2 \left(\phi_2\gamma_w(1) - 2\phi_1 + \gamma_w(1)^{-1} \right).
\end{align*}
Following Eq.~(\ref{eqn:VMRIC}), the results from Eq.~(\ref{eqn:example_mi2}), (\ref{eqn:example_c20}), and (\ref{eqn:example_c21}), deliver the $\operatorname{VMRIC}$ for this case.
\par
Now we consider a second misspecified model, Model 2:
\begin{equation}
\mathbf{y}_{t+2} = \boldsymbol{\rho} z_t + \boldsymbol{\eta}_t^{(2)},
\end{equation}
where $z_t$ is a weakly stationary linear AR(1) process independent of $w_t$:
\begin{equation}\label{eq:AR1}
z_t = \psi_1 z_{t-1} + \upsilon_t
\end{equation}
with $\psi_1 \in (-1, 1)$, and $\left\{ \upsilon_t \right\}$ is a sequence of i.i.d. random variables independent of both the error terms $\left\{ \delta_t\right\}$ and $\left\{\boldsymbol{\varepsilon}_{t}\right\}$ such that $E[\upsilon_t] = 0$ and $E[\upsilon_t^2] = 1 - \psi_1^2$, delivering $E[z_t] = 0$ and $E[z_t^2]=1$. Thus, $z_t$ is uncorrelated with both $w_t$ and $\mathbf{y}_t$, therefore $\boldsymbol{\rho}=\boldsymbol{0}$. The forecasting error in this case results $\boldsymbol{\eta}_t^{(2)} = \boldsymbol{a} w_{t+1} + \boldsymbol{\varepsilon}_{t+2}$. Following similar arguments as above we obtain ${\operatorname{\textbf{MI}}}$ and ${\operatorname{\textbf{VI}}}$ for Model 2:
\begin{align}
{\operatorname{\textbf{MI}}} &= \boldsymbol{\sigma}^2_{\boldsymbol{\varepsilon}} + \boldsymbol{a} \boldsymbol{a}^\top \\
{\operatorname{\textbf{VI}}} &= \boldsymbol{\sigma}^2_{\boldsymbol{\varepsilon}} + \boldsymbol{a} \boldsymbol{a}^\top (1 + 2 \psi_1 \gamma_w(1))
\end{align}
\par
As mentioned above, Model 1 is misspecified since it omits the lagged predictor $w_{t-1}$, while Model 2 only includes the non-informative predictor $z_t$.
\subsection{Finite sample performance}
First, we compare the above theoretical derivations with their sample counterpart. We consider three different parameterizations, presented in Table~\ref{tab:1}. Also, $\alpha_n = n^\alpha$ with $\alpha=0.85$. Note that, in order for Eq.~(\ref{eqn:penalty}) to hold, $\alpha$ must range in $(0, 1)$. Further experiments showed that results are fairly robust if reasonable values of $\alpha$ are selected. For an empirical method to determine it, see \cite[Section 5]{hsu2019_supp}. We take the following variance/covariance matrix for the innovations:
$$E[\boldsymbol{\varepsilon}_t\boldsymbol{\varepsilon}_t^{\top}] =
\left[\begin{array}{cc}
1 & 0.5 \\
0.5 & 1
\end{array}\right].$$
\par
We compute both the {$\operatorname{VMRIC}$} for Model 1 and Model 2, and estimate the $\hat{\operatorname{VMRIC}}$ and $\hat{\operatorname{VMRIC}}$ on a large sample of $n=10^6$ observations. The results are shown in Table~\ref{tab:2} for the two models, where the theoretical $\operatorname{VMRIC}$ (rows 1 and 3) is compared with the estimated one (rows 2 and 4). The results seem to confirm the consistency of the estimator shown in Eq.~(\ref{eqn:VMRIC_hat}). Clearly, the $\operatorname{VMRIC}$ of Model 1 is consistently smaller than that of Model 2 and indicates its superior predictive capability.
\begin{table}[]
\centering \caption{Parameters' combinations for the DGP of Eq.~(\ref{eq:DGP}), (\ref{eq:AR2}), and (\ref{eq:AR1}).}\label{tab:1}
\begin{tabular}{crrrrr}
Case & $\phi_1$ & $\phi_2$ & $a_1$ & $a_2$ & $\psi_1$ \\
\cmidrule(lr){2-6}
1 & 0.4 & -0.75 & 1.50 & -2.00 & 0.80 \\
2 & -0.4 & -0.45 & -0.75 & 1.25 & -0.65 \\
3 & 0.3 & -0.80 & 1.00 & 0.50 & -0.75 \\
\cmidrule(lr){2-6}
\end{tabular}
\end{table}
\begin{table}[]
\centering \caption{Theoretical and estimated {$\operatorname{VMRIC}$} of Models 1 and 2, for the three parameterizations of Table~\ref{tab:1}, computed on a data set of $n=10^6$ observations.}
\label{tab:2}
\begin{tabular}{ccccc}
& \multicolumn{2}{c}{Model 1} & \multicolumn{2}{c}{Model 2} \\
\cmidrule(lr){2-3}\cmidrule(lr){4-5}
Case & $\operatorname{VMRIC}$ & $\hat{\operatorname{VMRIC}}$ & $\operatorname{VMRIC}$ & $\hat{\operatorname{VMRIC}}$ \\
\cmidrule(lr){2-3}\cmidrule(lr){4-5}
1 & 6.671 & 6.636 & 7.914 & 7.902 \\
2 & 2.777 & 2.768 & 3.164 & 3.168 \\
3 & 2.801 & 2.784 & 2.994 & 2.993 \\
\cmidrule(lr){2-3}\cmidrule(lr){4-5}
\end{tabular}
\end{table}
The finite sample behaviour of the method of moments estimator of the $\operatorname{VMRIC}$ can be further appreciated in Table~\ref{tab:3} where we show their bias and Mean Squared Error (MSE).
The results are based upon 1000 Monte Carlo replications and seem to indicate a rate of convergence of the order of $n^{-1}$.
\par
In Table~\ref{tab:4}, we show the percentages of correct model selection by the $\operatorname{VMRIC}$, compared with the multivariate version of the $\operatorname*{AIC}$ and $\operatorname*{BIC}$ for the three parameterizations of Table~\ref{tab:1}. For a sample size of $n=100$, both the $\operatorname*{AIC}$ and $\operatorname*{BIC}$ select the best predictive model in about 50\% of the cases and relying upon them is tantamount to tossing a fair coin. In such a case, the $\operatorname{VMRIC}$ selects the correct model in about $80\%$ of the cases and reaches $100\%$ for $n=1000$. On the contrary, for Case 3, both the $\operatorname*{AIC}$ and $\operatorname*{BIC}$ cannot go above $64\%$ for a sample size as large as $n=10000$ observations and this is a general indication of their lack of asymptotic efficiency.
\begin{table}[]
\centering
\caption{Bias and Mean-Squared Error (MSE) for the (method of moments) estimator of the $\operatorname{VMRIC}$ for the three parameterizations, $\alpha=0.85$ and different sample size $n$. The results are based upon $1000$ Monte Carlo replications.}
\label{tab:3}
\begin{tabular}{rcccccc}
& \multicolumn{2}{c}{Case 1} & \multicolumn{2}{c}{Case 2} & \multicolumn{2}{c}{Case 3} \\
\cmidrule(lr){2-3}\cmidrule(lr){4-5} \cmidrule(lr){6-7}
$n$ & Bias & MSE & Bias & MSE & Bias & MSE \\
\cmidrule(lr){2-3}\cmidrule(lr){4-5} \cmidrule(lr){6-7}
100 & 0.227 & 1.137 & 0.063 & 0.306 & 0.030 & 0.182 \\
250 & 0.117 & 0.455 & 0.022 & 0.107 & 0.032 & 0.076 \\
500 & 0.061 & 0.225 & 0.015 & 0.048 & 0.004 & 0.032 \\
1000 & 0.019 & 0.109 & 0.010 & 0.023 & 0.002 & 0.015 \\
2500 & 0.008 & 0.044 & 0.001 & 0.009 & 0.001 & 0.006 \\
5000 & 0.009 & 0.023 & 0.001 & 0.004 & 0.003 & 0.003 \\
10000 & 0.001 & 0.012 & 0.003 & 0.002 & 0.001 & 0.002 \\
15000 & 0.004 & 0.008 & 0.001 & 0.001 & 0.002 & 0.001 \\
30000 & 0.002 & 0.004 & 0.001 & 0.001 & 0.001 & 0.001 \\
\cmidrule(lr){2-3}\cmidrule(lr){4-5} \cmidrule(lr){6-7}
\end{tabular}
\end{table}
\begin{table}[]
\centering
\caption{Percentages of correctly selected models by the three information criteria for the three parameterizations and varying sample size $n$.}\label{tab:4}
\small
\begin{tabular}{rccccccccc}
&\multicolumn{3}{c}{Case 1} & \multicolumn{3}{c}{Case 2} & \multicolumn{3}{c}{Case 3} \\
\cmidrule(lr){2-4}\cmidrule(lr){5-7} \cmidrule(lr){8-10}
$n$ & VMRIC & AIC & BIC & VMRIC & AIC & BIC &VMRIC & AIC & BIC \\
\cmidrule(lr){2-4}\cmidrule(lr){5-7} \cmidrule(lr){8-10}
100 & 85.9 & 52.5 & 52.5 & 84.6 & 56.2 & 56.2 & 72.1 & 49.0 & 49.0 \\
1000 & 99.9 & 65.6 & 65.6 & 99.9 & 73.7 & 73.7 & 97.0 & 56.8 & 56.8 \\
10000 & 100 & 88.0 & 88.0 & 100 & 97.8 & 97.8 & 100 & 63.8 & 63.8 \\
\cmidrule(lr){2-4}\cmidrule(lr){5-7} \cmidrule(lr){8-10}
\end{tabular}
\end{table}
\section{Proofs}\label{sec:proof}
In this section we detail the proofs of the three theorems. Hereafter all the derivations hold for any fixed $h\geq 1$; for the sake of presentation we write $ \boldsymbol{\varepsilon}_{t}$ instead of $ \boldsymbol{\varepsilon}_{t}^{(h)}$.
Remember that $\{l_n\}$ indicates an increasing sequence of positive integers such that:
\begin{equation}\label{eqn:ln}
l_n \rightarrow \infty,\qquad \dfrac{l_n}{\sqrt{n}}=o\left(1\right)
\end{equation}
and define $a = n-l_n-h$ and $b=n-l_n-h+1$.
\subsection{Proof of Theorem~\ref{thm:MSPE_dec_case1}}\label{sec:A1}
The proof of Theorem~\ref{thm:MSPE_dec_case1} relies upon four propositions.
\begin{proposition}\label{prop:thm1_1}
Under assumptions of Theorem~\ref{thm:MSPE_dec_case1}, it holds that:
\begin{equation}
N(\mathrm{I}) = (\mathrm{III}) + o(1),\label{eqn:thm1_1}
\end{equation}
where
\begin{align*}
(\mathrm{I}) =- \operatorname{E}\left[x_n \hat{R}^{-1} \left( \boldsymbol{\hat{\Sigma}} \boldsymbol{\varepsilon}_{n}^\top + \boldsymbol{\varepsilon}_{n} \boldsymbol{\hat{\Sigma}}^\top \right) \right],\quad
(\mathrm{III})= \ - \operatorname{E}\left[x_n R^{-1} \left( \boldsymbol{\hat{\Sigma}}_A \boldsymbol{\varepsilon}_{n}^\top + \boldsymbol{\varepsilon}_{n} \boldsymbol{\hat{\Sigma}}_A^\top\right) \right],
\end{align*}
with $\boldsymbol{\hat{\Sigma}} = \left( N^{-1} \sum_{t=1}^{N}x_t \boldsymbol{\varepsilon}_{t} \right)$ and $\boldsymbol{\hat{\Sigma}}_A = \sum_{t=1}^{N}\boldsymbol{\varepsilon_{t}} x_t$.
\end{proposition}
\begin{proof}
Let $\boldsymbol{A}_1 = \sum_{t=1}^{N} \left(\boldsymbol{\varepsilon}_t x_t\right) \boldsymbol{\varepsilon}_n^\top$ and note that
\begin{align}
\label{I-IIIdiv} \left\lVert(\mathrm{I})-(\mathrm{III}) \right\rVert
&= \left\lVert\operatorname{E}\left[x_n \left(\hat{R}^{-1}-R^{-1}\right)\left(\boldsymbol{A}_1+\boldsymbol{A}_1^\top\right)\right] \right\rVert.
\end{align}
By using standard properties of the norm, (\ref{eqn:thm1_1}) follows upon proving that
\begin{equation}\label{eqn:thm1_1main}
\left\lVert \operatorname{E}\left[x_n \left(\hat{R}^{-1}-R^{-1}\right)\boldsymbol{A}_1^\top\right]\right\rVert=o(1).
\end{equation}
Let
\begin{equation}\label{eqn:R_tilde}
\tilde{R} = \left(n-l_n\right)^{-1} \sum_{t=1}^{n-l_n} x_t^2.
\end{equation}
By adding and subtracting $ \boldsymbol{\varepsilon}_n x_n \left(\tilde{R}^{-1} \left[\sum_{t=1}^{N}\left(\boldsymbol{\varepsilon}_t x_t\right)\right]^\top \right)$, we have
\begin{align}
\operatorname{E}\left[x_n \left(\hat{R}^{-1}-R^{-1}\right)\boldsymbol{A}_1^\top\right]&=\operatorname{E}\left[\boldsymbol{\varepsilon}_n x_n \left(\hat{R}^{-1}-\tilde{R}^{-1}\right) \sum_{t=1}^{N} \boldsymbol{\varepsilon}_t^\top x_t\right]\nonumber\\
&+ \operatorname{E}\left[\boldsymbol{\varepsilon}_n x_n \left(\tilde{R}^{-1}-R^{-1}\right) \sum_{t=1}^{N} \boldsymbol{\varepsilon}_t^\top x_t\right]\nonumber\\
&= \operatorname{E}\left[\boldsymbol{\varepsilon}_n x_n \left(\hat{R}^{-1}-\tilde{R}^{-1}\right) \left(\sum_{t=1}^{N} \boldsymbol{\varepsilon}_t x_t\right)^\top\right]\label{eqn:1}\\
&+ \operatorname{E}\left[\boldsymbol{\varepsilon}_n x_n \left(\tilde{R}^{-1}-R^{-1}\right) \left(\sum_{t=b}^{N} \boldsymbol{\varepsilon}_t x_t\right)^\top\right]\label{eqn:2}\\
&+ \operatorname{E}\left[\boldsymbol{\varepsilon}_n x_n \left(\tilde{R}^{-1}-R^{-1}\right) \left(\sum_{t=1}^{a} \boldsymbol{\varepsilon}_t x_t\right)^\top\right]\label{eqn:3}.
\end{align}
We show below that the norms of (\ref{eqn:1}), (\ref{eqn:2}) and (\ref{eqn:3}) are asymptotically negligible. Focus on the first one: by combining conditions (C3), (C4), Lemma \ref{lma:technical}, and H\"older's inequality, it follows that $\left\lVert(\ref{eqn:1})\right\rVert$ is bounded by
\begin{align*}
\operatorname{E} \left[\left\lVert\boldsymbol{\varepsilon}_n x_n \left(\hat{R}^{-1}-\tilde{R}^{-1}\right) \left(\sum_{t=1}^{N} \boldsymbol{\varepsilon}_t x_t\right)^\top\right\rVert\right]
&\leq \operatorname{E}\left[\left\lVert\boldsymbol{\varepsilon}_n\right\rVert^6\right]^{\frac{1}{6}} \operatorname{E}\left[\left\lvert x_n \right\rvert^{6} \right]^{\frac{1}{6}} \operatorname{E}\left[\left\lvert \hat{R}^{-1}-\tilde{R}^{-1} \right\rvert^{3} \right]^{\frac{1}{3}}\\
& \times \operatorname{E}\left[\left\lVert N^{\frac{1}{2}}N^{-\frac{1}{2}}\sum_{t=1}^{N} \boldsymbol{\varepsilon}_t x_t \right\rVert ^{3}\right]^{\frac{1}{3}} = O\left(\frac{l_n}{n^{1/2}}\right),
\end{align*}
which converges to zero due to the definition of $l_n$ in (\ref{eqn:ln}). Similarly, we have that $\left\lVert(\ref{eqn:2})\right\rVert$ is bounded by
\begin{align*}
\operatorname{E}\left[\left\lVert\boldsymbol{\varepsilon}_n\right\rVert^6\right]^\frac{1}{6} \operatorname{E}\left[\left\lvert x_n\right\rvert^6\right]^\frac{1}{6} \operatorname{E}\left[\left\lvert\tilde{R}^{-1}-R^{-1}\right\rvert^3\right]^\frac{1}{3} \operatorname{E}\left[\left\lVert\left(\left(N-b+1\right)^{\frac{1}{2}}\left(N-b+1\right)^{-\frac{1}{2}}\sum_{t=b}^{N} \boldsymbol{\varepsilon}_t x_t\right)^\top\right\rVert^3\right]^\frac{1}{3}.
\end{align*}
which is an $O\left(n^{-1/2}l_n\right)$ thereby vanishing asymptotically. Lastly, Condition (C6), Lemma \ref{lma:technical}, and H\"older's inequality imply that $\left\lVert(\ref{eqn:2})\right\rVert$ is bounded by
\begin{align*}
\operatorname{E}\left[\left\lVert\operatorname{E}\left[ \boldsymbol{\varepsilon}_t x_t \mid \mathcal{F}_{t-l_n}\right]\right\rVert^3\right]^\frac{1}{3} \operatorname{E}\left[\left\lvert\tilde{R}^{-1}-R^{-1}\right\rvert^{3}\right]^\frac{1}{3} \operatorname{E}\left[\left\lVerta^{\frac{1}{2}} a^{-\frac{1}{2}}\sum_{t=1}^{a} \boldsymbol{\varepsilon}_t^\top x_t\right\rVert^3\right]^\frac{1}{3}= o\left(1\right)
\end{align*}
and this completes the proof.
\end{proof}
\begin{proposition}\label{prop:thm1_2}
Under assumptions of Theorem~\ref{thm:MSPE_dec_case1}, it holds that:
\begin{equation}\label{eqn:thm1_2}
N(\mathrm{II}) = (\mathrm{IV}) + o(1),
\end{equation}
where
\begin{align*}
(\mathrm{II}) = \operatorname{E}\left[\hat{R}^{-1} \boldsymbol{\hat{\Sigma}}x_n x_n \boldsymbol{\hat{\Sigma}}^\top \hat{R}^{-1}\right],\quad
(\mathrm{IV}) = \ \operatorname{E}\left[ \boldsymbol{\hat{\Sigma}}_B R^{-1} \boldsymbol{\hat{\Sigma}}_B^\top \right],
\end{align*}
with $\boldsymbol{\hat{\Sigma}}$ being defined in Proposition~\ref{prop:thm1_1} and $\boldsymbol{\hat{\Sigma}}_B = N^{-\frac{1}{2}} \sum_{t=1}^{N}\boldsymbol{\varepsilon_{t}} x_t.$
\end{proposition}
\begin{proof}
Let $\boldsymbol{M}_1 = x_n \left(\hat{R}^{-1} - R^{-1}\right) \boldsymbol{\hat{\Sigma}}_B$ and $\boldsymbol{M}_2 = x_n R^{-1} \boldsymbol{\hat{\Sigma}}_B$. Since
\begin{align*}
N \left(\mathrm{II}\right)= \operatorname{E}\left[\left(\boldsymbol{M}_1+ \boldsymbol{M}_2\right) \left(\boldsymbol{M}_1+ \boldsymbol{M}_2\right)^\top\right]
&= \operatorname{E}\left[\boldsymbol{M}_1 \boldsymbol{M}_1^\top\right] + \operatorname{E}\left[\boldsymbol{M}_2\boldsymbol{M}_2^\top\right]\\
& + \operatorname{E}\left[\boldsymbol{M}_1 \boldsymbol{M}_2^\top\right] + \operatorname{E}\left[\boldsymbol{M}_2 \boldsymbol{M}_1^\top\right]
\end{align*}
the proof of (\ref{eqn:thm1_2}) reduces to show that the following conditions hold:
\begin{align}
\left\lVert\operatorname{E}\left[ \boldsymbol{M}_1 \boldsymbol{M}_1^\top \right]\right\rVert &= o\left(1\right),\label{prop:cond1}\\
\left\lVert\operatorname{E}\left[ \boldsymbol{M}_1 \boldsymbol{M}_2^\top \right]\right\rVert &= o\left(1\right),\label{prop:cond2} \\
\left\lVert\operatorname{E}\left[ \boldsymbol{M}_2 \boldsymbol{M}_2^\top \right]- (\mathrm{IV})\right\rVert&= o\left(1\right).\label{prop:cond3}
\end{align}
Conditions (\ref{prop:cond1}) and (\ref{prop:cond2}) readily follow from Assumptions (C3) and (C4), Lemma \ref{lma:technical}, the non singularity of $R$ and H\"older's inequality:
\begin{align*}
\operatorname{E}\left[\left\lVert\boldsymbol{M}_1 \boldsymbol{M}_1^\top \right\rVert\right] &= \operatorname{E}\left[\left\lVertx_n^2 \left(\hat{R}^{-1} - R^{-1}\right)^2 \boldsymbol{\hat{\Sigma}}_B \boldsymbol{\hat{\Sigma}}_B^\top \right\rVert\right]
\leq \left(\operatorname{E}\left[\left\lvert x_n \right\rvert^{10}\right]\right)^\frac{1}{5} \left(\operatorname{E}\left[\left\lvert\hat{R}^{-1}-R^{-1}\right\rvert^5\right]\right)^\frac{2}{5}\\ &\times\left(\operatorname{E}\left[\left\lVert\boldsymbol{\hat{\Sigma}}_B\right\rVert^5\right]\right)^\frac{2}{5}= o\left(1\right);\\
\operatorname{E}\left[ \left\lVert\boldsymbol{M}_1\boldsymbol{M}_2^\top\right\rVert \right] &= \operatorname{E}\left[\left\lVertx_n^2 \left(\hat{R}^{-1} - R^{-1}\right) R^{-1} \boldsymbol{\hat{\Sigma}}_B \boldsymbol{\hat{\Sigma}}_B^\top \right\rVert\right]
\leq \left(\operatorname{E}\left[\left\lvert x_n\right\rvert^{10}\right]\right)^\frac{1}{5} \left(\operatorname{E}\left[\left\lvert\hat{R}^{-1}-R^{-1}\right\rvert ^5\right]\right)^\frac{1}{5}\\
&\times \left(\operatorname{E}\left[\left\lvert R^{-1} \right\rvert^5\right]\right)^\frac{1}{5} \left(\operatorname{E}\left[\left\lVert\boldsymbol{\hat{\Sigma}}_B\right\rVert^5\right]\right)^\frac{2}{5}= o\left(1\right).
\end{align*}
As concerns (\ref{prop:cond3}), decompose the vector $\boldsymbol{\hat{\Sigma}}_B$ as follows:
\begin{align*}
\boldsymbol{\hat{\Sigma}}_B = N^{-\frac{1}{2}} \sum_{t=1}^{N}\boldsymbol{\varepsilon_{t}} x_t = \boldsymbol{u} + \boldsymbol{w}, \quad
\text{ with } \quad \boldsymbol{u} = N^{-\frac{1}{2}} \sum_{t=1}^{a}\boldsymbol{\varepsilon_{t}} x_t \quad\text{and}\quad \boldsymbol{w} = N^{-\frac{1}{2}} \sum_{t=b}^{N}\boldsymbol{\varepsilon_{t}} x_t.
\end{align*}
Hence, we have that
\begin{align*}
\operatorname{E}\left[\boldsymbol{M}_2\boldsymbol{M}_2^\top\right]- (\mathrm{IV}) &= \operatorname{E}\left[\boldsymbol{u} R^{-1} x_n x_n R^{-1} \boldsymbol{u}^\top\right] - \operatorname{E}\left[\boldsymbol{u} R^{-1} R R^{-1} \boldsymbol{u}^\top\right]\\
& + \operatorname{E}\left[\boldsymbol{u} R^{-1} x_n x_n R^{-1} \boldsymbol{w}^\top\right] - \operatorname{E}\left[\boldsymbol{u} R^{-1} R R^{-1} \boldsymbol{w}^\top\right]\\
& + \operatorname{E}\left[\boldsymbol{w} R^{-1} x_n x_n R^{-1} \boldsymbol{u}^\top\right] - \operatorname{E}\left[\boldsymbol{w} R^{-1} R R^{-1} \boldsymbol{u}^\top\right]\\
& + \operatorname{E}\left[\boldsymbol{w} R^{-1} x_n x_n R^{-1} \boldsymbol{w}^\top\right] - \operatorname{E}\left[\boldsymbol{w} R^{-1} R R^{-1} \boldsymbol{w}^\top\right].
\end{align*}
The law of iterated expectations implies that:
\begin{align}
&\left\lVert\operatorname{E}\left[\boldsymbol{M}_2 \boldsymbol{M}_2^\top\right]-(\mathrm{IV})\right\rVert\nonumber\\
&\leq \left\lVert\operatorname{E}\left[\boldsymbol{u} R^{-1} \left(\operatorname{E}\left[x_n^2 \mid \mathcal{F}_{n-l_n}\right]-R\right) R^{-1} \boldsymbol{u}^\top\right]\right\rVert\label{norm1}\\
&+ \left\lVert\operatorname{E}\left[\boldsymbol{u} R^{-1} \left(\operatorname{E}\left[x_n^2 \mid \mathcal{F}_{n-l_n}\right]-R\right) R^{-1} \boldsymbol{w}^\top\right]\right\rVert\label{norm2}\\
&+ \left\lVert\operatorname{E}\left[\boldsymbol{w} R^{-1} \left(\operatorname{E}\left[x_n^2 \mid \mathcal{F}_{n-l_n}\right]-R\right) R^{-1} \boldsymbol{u}^\top\right]\right\rVert\label{norm3}\\
&+ \left\lVert\operatorname{E}\left[\boldsymbol{w} R^{-1} \left(\operatorname{E}\left[x_n^2 \mid \mathcal{F}_{n-l_n}\right]-R\right) R^{-1} \boldsymbol{w}^\top\right]\right\rVert\label{norm4}.
\end{align}
By using arguments previously developed, it is easy to see that, under Assumptions (C4) and (C6), (\ref{norm1}) -- (\ref{norm4}) asymptotically vanish. Therefore, conditions (\ref{prop:cond1}) -- (\ref{prop:cond3}) are fulfilled and the proof is completed.
\end{proof}
\begin{proposition}\label{prop:thm1_3}
Under assumptions of Theorem~\ref{thm:MSPE_dec_case1}, it holds that:
\begin{equation}\label{eqn:prop3}
(\mathrm{III}) = - (D) + o\left(1\right),
\end{equation}
where
\begin{align*}
(D) &= \operatorname{E}\left[ R^{-1} \left[\sum_{j=h}^{N-1} \left\{\left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right) \left( \boldsymbol{\varepsilon}_1 x_1\right)^\top\right\} \right] \right]
\end{align*}
\end{proposition}
\begin{proof}
The result readily follows upon noting that, under Assumption (C2) and the weakly stationarity of the process $\{x_t\}$, it holds that:
\begin{align*}
(\mathrm{III}) & = - \sum_{t=1}^{N} \operatorname{E}\left[R^{-1} \left\{ \left(\boldsymbol{\varepsilon}_t x_t\right) \left(\boldsymbol{\varepsilon}_n x_n\right)^\top + \left(\boldsymbol{\varepsilon}_n x_n\right) \left(\boldsymbol{\varepsilon}_t x_t\right)^\top\right\} \right]\\
& = - \ \sum_{j=h}^{n-1} \operatorname{E}\left[R^{-1} \left\{ \left(\boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right) \left(\boldsymbol{\varepsilon}_1 x_1\right)^\top\right\} \right]\\
&= - \operatorname{E}\left[ R^{-1} \left(\sum_{j=h}^{N-1} \left\{\left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right) \left( \boldsymbol{\varepsilon}_1 x_1\right)^\top\right\} \right) \right] +o(1).
\end{align*}
\end{proof}
\begin{proposition}\label{prop:thm1_4}
Under assumptions of Theorem~\ref{thm:MSPE_dec_case1}, it holds that:
\begin{equation}\label{eqn:prop4}
(\mathrm{IV}) = (1) + \left(Q\right) + \left(D\right) + o(1),
\end{equation}
where
\begin{align*}
(1) & = N^{-1} \operatorname{E}\left[ R^{-1} \left\{\sum_{t=1}^{N} \left(\boldsymbol{\varepsilon}_t x_t\right)\left(\boldsymbol{\varepsilon}_t x_t\right)^\top \right\} \right], \\
\left(Q\right) &= \operatorname{E}\left[ R^{-1} \left[\sum_{s=1}^{h-1} \left\{\left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{s+1} x_{s+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{s+1} x_{s+1}\right) \left( \boldsymbol{\varepsilon}_1 x_1 \right)^\top\right\} \right] \right]\\
\left(D\right) &= \operatorname{E}\left[ R^{-1} \left[\sum_{j=h}^{N-1} \left\{\left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right) \left( \boldsymbol{\varepsilon}_1 x_1\right)^\top\right\} \right] \right]
\end{align*}
\end{proposition}
\begin{proof}
Let
\begin{equation*}
(2) = N^{-1} \operatorname{E}\left[ R^{-1} \left\{\sum_{j=1}^{N-1} \sum_{k=j+1}^{N} \left(\boldsymbol{\varepsilon}_j x_j\right) \left(\boldsymbol{\varepsilon}_k x_k\right)^\top\right\} \right],
\end{equation*}
and note that $(\mathrm{IV})-(1) = (2) + (2)^\top$. Moreover
\begin{align}
(2) & = N^{-1} \operatorname{E}\left[ R^{-1} \left\{\sum_{j=1}^{N-1} \left( N-j \right) \left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top \right\} \right] \nonumber\\
& = \operatorname{E}\left[ R^{-1} \left\{\sum_{j=1}^{N-1} \left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top \right\} \right] \label{eqn:a} \\
& - N^{-1} \operatorname{E}\left[ R^{-1} \left\{\sum_{j=1}^{N-1} j\left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top \right\} \right].\label{eqn:b}
\end{align}
Assumptions (C2) implies that (\ref{eqn:b}) is $o(1)$. Since (\ref{eqn:a}) can be written as
\begin{equation*}
\operatorname{E}\left[ R^{-1} \left\{\sum_{s=1}^{h-1} \left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{s+1} x_{s+1}\right)^\top \right\} \right] + \operatorname{E}\left[ R^{-1} \left\{\sum_{j=h}^{N-1} \left( \boldsymbol{\varepsilon}_1 x_1\right) \left(\boldsymbol{\varepsilon}_{j+1} x_{j+1}\right)^\top \right\} \right],
\end{equation*}
then $(\ref{eqn:a})+(\ref{eqn:a})^\top=(Q)+(D)$ and this completes the proof.
\end{proof}
\subsubsection*{Proof of Theorem~\ref{thm:MSPE_dec_case1}}
We prove that:
\begin{align}
&N \left\{ \operatorname{E}\left[ \left( \mathbf{y}_{n+h} - \hat{\mathbf{y}}_{n+h} \right) \left( \mathbf{y}_{n+h} - \hat{\mathbf{y}}_{n+h} \right)^\top - \operatorname{E}\left[
\boldsymbol{\varepsilon}_{n}^{(h)} \boldsymbol{\varepsilon}_{n}^{(h)^\top} \right] \right] \right\}\nonumber\\
& = R^{-1} \operatorname{E}\left[ \left(\boldsymbol{\varepsilon}_{1}^{(h)} x_1\right)\left(\boldsymbol{\varepsilon}_{1}^{(h)} x_1\right)^\top \right] \label{eq:main1} \\
&+ R^{-1} \operatorname{E}\left[ \sum_{s=1}^{h-1} \left\{\left( \boldsymbol{\varepsilon}_1^{(h)} x_1\right) \left(\boldsymbol{\varepsilon}_{s+1}^{(h)} x_{s+1}\right)^\top + \left(\boldsymbol{\varepsilon}_{s+1}^{(h)} x_{s+1}\right) \left( \boldsymbol{\varepsilon}_1^{(h)} x_1 \right)^\top\right\} \right] \label{eq:main2}\\
&+ o(1). \nonumber
\end{align}
Since $$\left( \boldsymbol{\hat{\beta}}-\boldsymbol{\beta}\right) = \hat{R}^{-1} \left( N^{-1} \sum_{t=1}^{N} x_t \boldsymbol{y}_{t+h} \right) - \boldsymbol{\beta} = \hat{R}^{-1} \left(N^{-1} \sum_{t=1}^{N} x_t \boldsymbol{\varepsilon}_t \right),$$
routine algebra implies that:
\begin{equation}
\operatorname{E}\left[ \left(\mathbf{y}_{n+h} - \mathbf{\hat{y}}_{n+h}\right) \left(\mathbf{y}_{n+h} - \mathbf{\hat{y}}_{n+h}\right)^\top \right]
- \operatorname{E}\left[\boldsymbol{\varepsilon}_n \boldsymbol{\varepsilon}_n^\top\right] =
(\mathrm{I})+(\mathrm{II}).
\end{equation}
By applying Propositions~\ref{prop:thm1_1} -- Propositions~\ref{prop:thm1_4}, we have:
\begin{align*}
N \left\{ \operatorname{E}\left[ \left( \mathbf{y}_{n+h} - \hat{\mathbf{y}}_{n+h} \right) \left( \mathbf{y}_{n+h} - \hat{\mathbf{y}}_{n+h} \right)^\top - \operatorname{E}\left[
\boldsymbol{\varepsilon}_{n}^{(h)} \boldsymbol{\varepsilon}_{n}^{(h)^\top} \right] \right] \right\}= N(\mathrm{I})+N(\mathrm{II})\\
=(\mathrm{III})+(\mathrm{IV})+o(1)
=(1)+(Q)+o(1).
\end{align*}
The proof is completed upon noting that $(1)=(\ref{eq:main1})$ and $(Q)=(\ref{eq:main2})$.
\subsection{Proof of Theorem~\ref{thm:MOME_case1}}\label{sec:A2}
We start proving that
\begin{equation}\label{mi_hat}
\hat{\operatorname{\textbf{MI}}}_h= \operatorname{\textbf{MI}}_h + O_p (n^{-1/2}).
\end{equation}
Note that
\begin{align*}
\hat{\operatorname{\textbf{MI}}}_h=N^{-1} \left(\sum_{t=1}^{N} \boldsymbol{\varepsilon}_{t} \boldsymbol{\varepsilon}_{t}^\top\right)- \left(N^{-1}\sum_{t=1}^{N} x_t \boldsymbol{\varepsilon}_{t}\right) \hat{R}^{-1}\left(N^{-1}\sum_{s=1}^{N} x_s \boldsymbol{\varepsilon}_{s}\right)^\top
\end{align*}
hence, it holds that $\hat{\operatorname{\textbf{MI}}}_h - \operatorname{\textbf{MI}}_h$ equals
\begin{align}
& N^{-1} \left\{\sum_{t=1}^{N} \left( \boldsymbol{\varepsilon}_{t} \boldsymbol{\varepsilon}_{t}^\top - \operatorname{E}\left[ \boldsymbol{\varepsilon}_{1} \boldsymbol{\varepsilon}_{1}^\top\right] \right) \right\}\label{eq:mi_A}\\
& -\left(N^{-1}\sum_{t=1}^{N} x_t \boldsymbol{\varepsilon}_{t}\right)\hat{R}^{-1} \left(N^{-1}\sum_{t=1}^{N} x_t \boldsymbol{\varepsilon}_{t}\right)^\top.\label{eq:mi_B}
\end{align}
Assumption (A1) implies that $(\ref{eq:mi_A})=O_p(n^{-1/2})$ whereas, by combining Assumptions (A3) and (A4) with the non-singularity of $R$ and H\"older's inequality, it can be shown that $(\ref{eq:mi_B})=O_p(n^{-1})$ and hence the proof of (\ref{mi_hat}) is complete.
\par\noindent
Next, we prove that
\begin{equation*}
\hat{\operatorname{\textbf{VI}}}_h= \operatorname{\textbf{VI}}_h + o_p (1).
\end{equation*}
It suffices to show that
\begin{equation}\label{Cs_hat}
\hat{\mathbf{C}}_{h,s} = \mathbf{C}_{h,s}+ o_p (1).
\end{equation}
It holds that $\hat{\mathbf{C}}_{h,s}$ is equal to
\begin{align}
&\phantom{-}\left( N-s\right)^{-1} \sum_{t=1}^{N-s} \left( x_t \boldsymbol{\varepsilon}_{t}^{\top}\right)^{\top} \left( x_{t+s} \boldsymbol{\varepsilon}_{t+s}^{\top}\right)\label{Cs_A}\\
& - \left( N-s\right)^{-1} \sum_{t=1}^{N-s} x_t^2 x_{t+s} \left( \hat{\boldsymbol{\beta}}_n(h) - \boldsymbol{\beta}_h \right) \boldsymbol{\varepsilon}_{t+s}^{\top}\label{Cs_B}\\
& - \left( N-s\right)^{-1} \sum_{t=1}^{N-s} x_t x_{t+s}^2 \boldsymbol{\varepsilon}_{t} \left( \hat{\boldsymbol{\beta}}_n(h) - \boldsymbol{\beta}_h \right)^{\top}\label{Cs_C}\\
& + \left( N-s\right)^{-1} \sum_{t=1}^{N-s} x_t^2 x_{t+s}^2 \left( \hat{\boldsymbol{\beta}}_n(h) - \boldsymbol{\beta}_h \right)\left( \hat{\boldsymbol{\beta}}_n(h) - \boldsymbol{\beta}_h \right)^{\top}.\label{Cs_D}
\end{align}
We prove that (\ref{Cs_B}) is $o_p(1)$ componentwise. To this end consider:
$$\operatorname{E}\left[\left( N-s\right)^{-1}\abs*{\sum_{t=1}^{N-s} x_t^2 x_{t+s} \varepsilon_{t+s,i}}\right],$$
with $\varepsilon_{t+s,i}$ being the $i$-th component of the vector $\boldsymbol{\varepsilon}_{t+s}$. The triangular inequality and H\"older's inequality imply that:
\begin{align*}
\operatorname{E}\left[\left( N-s\right)^{-1}\abs*{\sum_{t=1}^{N-s} x_t^2 x_{t+s} \varepsilon_{t+s,i}}\right]
&\leq \left( N-s\right)^{-1}\sum_{t=1}^{N-s}\operatorname{E}\left[\abs*{x_t^2 x_{t+s} \varepsilon_{t+s,i}}\right]\\
&\leq (N-s)^{-1}\sum_{t=1}^{N-s}\left\{\left(\operatorname{E}\left[x_t^4\right]\right)^{1/2}\left( \operatorname{E}\left[\abs*{x_{t+s}\varepsilon_{t+s,i}}^2 \right]\right)^{1/2}\right\}.
\end{align*}
Since $ \hat{\boldsymbol{\beta}}_n(h) - \boldsymbol{\beta}_h = \hat{R}^{-1} \left( N^{-1} \sum_{j=1}^{N} x_j \boldsymbol{\varepsilon}_{j,h} \right)$, by combining Assumptions (A3), (A4) and (A5) with Chebyshev's inequality we obtain that (\ref{Cs_B}) is $o_p(1)$. Similarly, we can verify that (\ref{Cs_C}) and (\ref{Cs_D}) are $o_p(1)$. Lastly, Condition (A2) implies that $(\ref{Cs_A}) = \mathbf{C}_{h,s}+ o_p (1)$, hence (\ref{Cs_hat}) is verified and the whole proof is complete.
\subsection{Proof of Theorem~\ref{thm:EFF_case1}}\label{sec:A3}
By Theorem \ref{thm:MOME_case1} the $\operatorname{VMRIC}_h$ defined in (\ref{eq:vmric}) can be written as:
\begin{equation}\label{eq:theo3_proof_1}
\operatorname{VMRIC}_h \left( \hat{\ell}_h\right) = \min_{1 \leq \ell \leq K} \left\lVert \operatorname{\textbf{MI}}_h + O_p (n^{-1/2}) \right\rVert + \min_{ \ell \in M_1} \left\lVert\frac{\alpha_n}{n}\operatorname{\textbf{VI}}_h + o_p \left(\frac{\alpha_n}{n}\right)\right\rVert.
\end{equation}
Therefore,
\begin{equation}
\lim\limits_{n\rightarrow \infty} \operatorname{VMRIC}_h \left( \hat{\ell}_h\right) = \min_{1 \leq \ell \leq K}
\left\lVert\operatorname{\textbf{MI}}_h\right\rVert
\end{equation}
and hence
\begin{equation}
\lim\limits_{n\rightarrow +\infty} \Pr \left( \hat{\ell}_h \in M_1 \right) = 1.
\end{equation}
Now, consider two models $\ell_1$ and $\ell_2$ in the candidates set $J_{\ell_1}, J_{\ell_2} \in M_1$ such that $\operatorname{VI}_h(\ell_1)\neq \operatorname{VI}_h(\ell_2)$. We show that
\begin{equation}\label{eqn:eff_main}
\lim_{n\rightarrow \infty}\Pr \left[ \text{sign}\left\{\operatorname{VMRIC}_h(\ell_1) - \operatorname{VMRIC}_h(\ell_2)\right\} = \text{sign}\left\{\left\lVert\boldsymbol{\operatorname{VI}}_h(\ell_1)\right\rVert- \left\lVert\boldsymbol{\operatorname{VI}}_h(\ell_2)\right\rVert\right\}\right] = 1.
\end{equation}
By defining $\operatorname{\textbf{MI}}^*_h$ to be the minimum value of $\operatorname{\textbf{MI}}_h$ over the family of candidate models, we have:
\begin{align*}
\operatorname{VMRIC}_h (\ell_1) &= \left\lVert \operatorname{\textbf{MI}}_h^* + O_p (n^{-1/2}) \right\rVert + \left\lVert \frac{\alpha_n}{n}\operatorname{\textbf{VI}}_h(\ell_1) + o_p\left(\frac{\alpha_n}{n}\right)\right\rVert,\\
\operatorname{VMRIC}_h (\ell_2) &= \left\lVert \operatorname{\textbf{MI}}_h^* + O_p (n^{-1/2}) \right\rVert + \left\lVert \frac{\alpha_n}{n} \operatorname{\textbf{VI}}_h(\ell_2) + o_p\left(\frac{\alpha_n}{n}\right)\right\rVert.
\end{align*}
Therefore, for sufficiently large $n$, it holds that:
\begin{equation*}
\operatorname{VMRIC}_h (\ell_1) - \operatorname{VMRIC}_h (\ell_2) = \left\|\frac{\alpha_n}{n}\right\| \left( \left\lVert \operatorname{\textbf{VI}}_h(\ell_1)\right\rVert - \left\lVert \operatorname{\textbf{VI}}_h(\ell_2)\right\rVert\right).
\end{equation*}
Thus
\begin{equation*}
\text{sign}\left\{ \operatorname{VMRIC}_h(\ell_1) - \operatorname{VMRIC}_h(\ell_2)\right\} = \text{sign}\left\{ \left\lVert\operatorname{\textbf{VI}}_h(\ell_1)\right\rVert - \left\lVert\operatorname{\textbf{VI}}_h(\ell_2)\right\rVert\right\},
\end{equation*}
and (\ref{eqn:eff_main}) is verified and implies that
\begin{equation}
\lim_{n\rightarrow \infty} \Pr\left(\hat{\ell}_h \in M_2\right) = 1.
\end{equation}
This completes the proof.
\section*{Acknowledgments}
Greta Goracci acknowledges the support of Libera Università di Bolzano, Grant WW201L (ESAMD).