EconBase
← Back to paper

Time-Varying Model Averaging of Multi-layer Network Vector Autoregressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

108,488 characters · 16 sections · 52 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Time-Varying Model Averaging of Multi-layer Network Vector Autoregressions

\newtheorem{corollary}{Corollary} \newtheorem{definition}{Definition} \newtheorem{lemma}{Lemma} \newtheorem{proposition}{Proposition} \newtheorem{remark}{Remark} \newtheorem{theorem}{Theorem} \newtheorem{example}{Example} \newtheorem{assumption}{Assumption} \newtheorem{prop}{Proposition}

\numberwithin{corollary}{section} \numberwithin{definition}{section} \numberwithin{equation}{section} \numberwithin{lemma}{section} \numberwithin{proposition}{section} \numberwithin{remark}{section} \numberwithin{theorem}{section}

\onehalfspacing

\newtheorem{Proof}{Proof} \newtheorem{Lemma}{Lemma} \newtheorem{Mth}{Main Theorem} \newtheorem{Res}{\bf Result} \newtheorem{Def}{Definition} \newtheorem{Rem}{\bf Remark} \newtheorem{Qes}{Question} \newtheorem{Aim}{Aim} \newtheorem{Pro}{Proposition} \newtheorem{Lem}{\bf Lemma} \newtheorem{Cor}{\bf Corollary} \newtheorem{Ex}{Example} \newtheorem{Eq}{Equation} \newtheorem{condition}{Assumption} \newtheorem{conditionnew}{Condition}

\def\0{{\bf 0}}

\def\squarebox#1{\hbox to #1{\vbox to #1}}

\allowdisplaybreaks[4]

\thispagestyle{empty}

abstractIn this paper, we introduce a flexible time-varying multi-layer network vector autoregression (VAR) model framework for large-scale time series, allowing agents in dynamic systems to interact through multiple channels and incorporating multiple adjacency matrices to capture network spillover effects. We propose a penalized model averaging method to determine a time-varying optimal combination of multi-layer network VAR candidate models whose number may be divergent. Under some regularity conditions, the asymptotic properties such as asymptotic optimality and convergence rates of the proposed time-varying weight estimation are derived in the contexts of both the in-sample fitting and out-of-sample prediction. In addition, we extend the conformal prediction method to construct prediction bands for locally stationary time series. Monte-Carlo simulation studies and an empirical application to forecast CPI inflation by combining multiple network information are given to illustrate reliable finite-sample estimation and predictive performance of the developed methodology. {\em Keywords}: asymptotic optimality, conformal prediction, model averaging, multi-layer network, time-varying VAR.

\setcounter{page}{1}

Introduction

\setcounter{equation}{0}

The vector autoregression (VAR) has become a standard econometric tool to model multivariate time series since the seminal work by Si80. In the classic setting with a fixed number of time series variables, Lu06 and KL17 provide a comprehensive review of various VAR-based model estimation, inference and forecasting techniques. However, the conventional OLS or MLE estimation and forecasting methods often perform poorly when the number of variables diverges with the time series length. To address this issue in the high-dimensional time series setup, a commonly-used approach is to impose some structural restriction on transition matrices such as the sparsity or reduced-rank assumption and subsequently adopt regularized estimation methods NW11, KC15, DZZ16, BLM19, MPS23. However, these models and methods often neglect potential inter-connection or network structure among a large number of variables, a feature which commonly exists for large-scale dynamic macroeconomic or financial systems. The network VAR introduced by ZPLLW17 whose model formulation is virtually similar to the so-called global VAR PSW04, provides a flexible autoregressive framework for large-scale time series to jointly incorporate the network spillover (cross-lag) effect, momentum (own-lag) effect and nodal effect.

Network models are an effective tool to tackle interconnected systems with applications to various fields such as economics, finance and social networks. There have been extensive studies on estimation and inference of both the static network DY14, S17, DDLY18, BB19 and time-varying dynamic network KSAX10, ZLW10, CLLL25. The existing literature on network VAR models typically assumes that the network structure is available and incorporates an observed adjacency matrix in VAR to capture the network spillover effect ZPLLW17, CFZ23, YSM23, YSM26, LPTW26. They often limit attention to the assumption of a single adjacency matrix to measure linkages between nodes. This assumption may be restrictive for modern social or economic networks as agents in dynamic systems often interact through multiple channels, resulting in the so-called multi-layer network KABGMP14. BCM25 introduce a factor network VAR model and propose a tensor decomposition of multiple adjacency matrices which may vary over time to achieve dimension reduction in the subsequent estimation procedure. In this paper, we propose an alternative approach for multi-layer network VAR, adopting the model averaging method to construct an optimal combination of candidate time-varying multi-layer network VAR models with different linkage structures (to capture network spillover effects), varying lagged orders of dependent variables and node-specific predictors.

Model averaging tackles model uncertainty by averaging over a number of candidate predictive models and serves as an alternative to model selection which aims to select one model according to a proper criterion. Model averaging methods can be broadly classified into two categories: Bayesian model averaging with prior information HMRV99, and frequentist model averaging, which focuses on determining optimal weights for candidate models. Commonly-used frequentist model averaging approaches include the Mallows model averaging H07,LT20 and jackknife model averaging HR12, LZZZ19, YZL25. These classic methods aim to select time-invariant combination weights. Their numerical performance may be unstable for dynamic environments with structural changes. SHLWZ21 is among the first to propose time-varying jackknife model averaging to accommodate smooth structural changes, with subsequent extensions to high-dimensional penalized time-varying model averaging SHWZ23, factor-augmented regression CHL24 and quantile regression TW25, and time-varying VAR SCG25. The aforementioned works, however, do not account for interconnections between time series variables and potential multi-layer network structures. Our paper aims to fill in this gap by studying time-varying model averaging of multi-layer network VAR for large-scale time series with smooth structural changes.

Meanwhile, there has been increasing attention in recent years on time-varying VAR models, accommodating smooth structural changes over a long time span. GPY24 study estimation and inference of time-varying VAR models including the impulse response analysis for finite-dimensional time series. The time-varying VAR is naturally connected to the locally stationary time series model framework, which has been systematically studied in the literature since D97. In the high-dimensional time-varying VAR model setup, DQC17 combine the $\ell_1$-regularised method with kernel smoothing to estimate transition matrices; XCW20 study the multiple break structure and estimate smooth time-varying covariance and precision matrices between the break points; and CLLL25 estimate dual network structures via directed Granger causality and undirected partial correlation linkages. A recent paper by LPTW26 further introduces a general time-varying network VAR model framework with a latent group structure, allowing both structural changes and structural breaks on model coefficients. However, they limit attention to a single-layer network structure to capture the network spillover effect.

In this paper, we propose a penalized time-varying model averaging (PTVMA) method to estimate the optimal weights for candidate multi-layer network VAR models whose time-varying coefficients are estimated via the classic local linear smoothing. The candidate models are constructed via distinct adjacency matrices in network spillover effects, varying lags in momentum effects and different subsets of node-specific exogenous predictors. In particular, we allow all the candidate models to be misspecified, enhancing the applicability of the proposed model and methodology. Under some regularity conditions, we establish the asymptotic optimality and convergence rates of the proposed time-varying weight estimates in the contexts of both the in-sample fitting and out-of-sample prediction. In particular, we show that the PTVMA criterion achieves the so-called “selection consistency" when there exist correctly-specified candidate models, i.e., the sum of the estimated time-varying weights for correctly-specified candidate models approaches one asymptotically.

In addition to point forecasts, we further construct the prediction interval, accounting for forecast uncertainty. Specifically, we propose a PTVMA-based conformal prediction method to construct interval forecasts without requiring the number of candidate models or predictors to be fixed. The conformal prediction method is first introduced by VGS05 as a sequential approach to form prediction intervals and has been extended by LW14 and LGRTW18 to accommodate the nonparametric regression framework. In this paper, we make a further extension by combining the conformal prediction with the developed PTVMA method to construct prediction bands in the context of large-scale and locally stationary network time series. The proposed PTVMA-based conformal prediction algorithm is easy to implement and can guarantee the coverage probability without requiring the asymptotic distribution theory of the model averaging estimation.

Monte Carlo simulations demonstrate that the proposed PTVMA method delivers nearly identical in-sample fit and superior out-of-sample prediction relative to the benchmark, even when all candidate models are misspecified. By assigning time-varying weights to candidate models to accommodate parameter instability, the PTVMA criterion achieves selection consistency when correctly specified candidates exist, and concentrates weight on the best-approximating models when all candidates are misspecified. The proposed conformal prediction intervals exhibit robust finite-sample performance, with empirical coverage probabilities close to the nominal level, particularly as the sample size increases. The developed model and methodology are further applied to analyze the monthly CPI inflation in 36 economies from January 2007 to December 2024. The empirical results reveal multiple channels of CPI inflation transmission across economies, with the production, equity, trade, and policy networks alternately dominating global inflation transmission. Our model and method consistently outperform a broad set of competing models, including the static and time-varying network VAR models and state-of-the-art machine learning methods, highlighting the advantage of the proposed time-varying multi-layer network VAR and PTVMA.

The rest of the paper is organized as follows. Section (ref) introduces the model framework and the PTVMA methodology. Section (ref) gives the asymptotic properties of the developed weight estimation together with technical assumptions. Section (ref) proposes the PTVMA-based conformal prediction interval construction. Sections (ref) and (ref) report the simulation and empirical studies, respectively. Section (ref) concludes the paper. All the theoretical proofs are contained in the appendix.

Model and methodology

\setcounter{equation}{0}

In this section, we introduce the time-varying multi-layer network VAR model framework and propose the PTVMA method to determine a time-varying optimal combination of candidate models.

Time-varying multi-layer network VAR

Consider ${\sf N}$-dimensional vectors of time series observations ${\bf Y}_t=(y_{1,t},y_{2,t},\ldots,y_{{\sf N}, t})^{{^\intercal}}$, $1\leq t\leq {\sf T}$, with ${\sf N}$ being the number of nodes in the large-scale network, and ${\sf Q}$-dimensional node-specific vectors ${\bf X}_i=(x_{i,1},\ldots,x_{i,{\sf Q}})^{{^\intercal}}$, $1\leq i\leq {\sf N}$. Throughout the paper, ${\sf N}$ is allowed to diverge to infinity while ${\sf Q}$ is a fixed positive integer. Let ${\mathbf W}^{(k)}=(\omega_{ij}^{(k)})_{{\sf N}\times {\sf N}}$ denote the $k$-th adjacency matrix, which is observed and non-stochastic, $\omega_{ii}^{(k)}=0$ and $\omega_{ij}^{(k)}=1$ if $i$ follows $j$, $k=1,\ldots,{\sf K}$. These adjacency matrices capture economic linkages, geographic distances, or other forms of connectedness. All the adjacency matrices in this paper can be either directed or undirected. We consider the following data generating process:

eqnarray[eqnarray omitted — 550 chars of source]

where $\widetilde \omega_{ij}^{(k)}=\omega_{ij}^{(k)}/n_i^{(k)}$ with $n_i^{(k)}=\sum_{j\neq i} \omega_{ij}^{(k)}$ denoting the total number of nodes that $i$ follows according the $k$-th linkage, $\boldsymbol{\beta}_t=(\beta_{t1},\ldots,\beta_{t{\sf K}})^{{^\intercal}}$, $\boldsymbol {\alpha}_t= (\alpha_{t1}, \ldots,\alpha_{t{\sf P}} )^{{^\intercal}}$ and ${\boldsymbol\gamma}_t=(\gamma_{t1},\ldots,\gamma_{t{\sf Q}})^{{^\intercal}}$ are unknown time-varying parameter vectors to be estimated, ${\bf Z}_{i,t}$ denotes the predictor vector with dimension ${\sf K}+{\sf P}+{\sf Q}$, i.e.,

eqnarray[eqnarray omitted — 489 chars of source]

${\boldsymbol\theta}_t=(\boldsymbol{\beta}_t^{^\intercal},\boldsymbol {\alpha}_t^{^\intercal}, {\boldsymbol\gamma}_t^{^\intercal})^{^\intercal}$, and $\epsilon_{i,t}=\sigma_i(t/T)u_{i,t}$ with $\sigma_i^2(\cdot)$ being a time-varying variance function and $u_{i,t}$ being independent and identically distributed (i.i.d.) with zero mean and unit variance.

The time-varying VAR in ((ref)) is split into multiple network spillover effects (or cross-lag effects), the momentum effects with order ${\sf P}$ (or own-lag effects) and the nodal effects, all of which are allowed to evolve smoothly over time. Here we only include the first lag of network spillover effects for simplicity and the extension to higher-order cross-lag effects is straightforward.

Penalized time-varying model averaging

Unlike the model selection which aims to select one time-varying network VAR model, to reduce model estimation and prediction uncertainty caused by selection of adjacency matrices, order of lagged dependent variables and subset of node-specific exogenous predictors, we next propose the PTVMA procedure, seeking to construct a dynamic optimal combination of candidate time-varying network VAR models. Consider a sequence of candidate network VAR models indexed by $m=1,\ldots, {\sf M}$, which may be misspecified for the underlying data generating process. Specifically, we formulate the $m$-th candidate model as

eqnarray[eqnarray omitted — 462 chars of source]

where ${\cal I}_{{\sf K}}^{(m)}=\{i_1,\ldots,i_{k_m}\}$ is a subset of $\{1,2,\ldots,{\sf K}\}$, ${\cal I}_{{\sf P}}^{(m)}=\{i_1,\ldots,i_{p_m}\}$ is a subset of $\{1,2,\ldots,{\sf P}\}$, ${\cal I}_{{\sf Q}}^{(m)}=\{i_1,\ldots,i_{q_m}\}$ is a subset of $\{1,2,\ldots,{\sf Q}\}$, $${\boldsymbol\theta}_t^{(m)}=\left(\boldsymbol{\beta}_t^{(m){^\intercal}},\boldsymbol {\alpha}_t^{(m){^\intercal}},{\boldsymbol\gamma}_t^{(m){^\intercal}}\right)^{{^\intercal}}$$ with

eqnarray[eqnarray omitted — 555 chars of source]

and ${\bf Z}_{i,t}^{(m)}$ denotes the predictor vector with dimension $k_m+p_m+q_m$. All the time-varying parameters are defined as smooth functions of scaled times. Let $$\bar{\rho}= ({\sf K}+{\sf P}+{\sf Q})\vee \max \{ \rho_1, \rho_2, \cdots, \rho_{\sf M}\} \ \ \ \text{with}\ \ \rho_m=k_m+p_m+q_m,$$ ${\boldsymbol\mu}_t^{(m)}=(\mu_{1,t}^{(m)},\ldots,\mu_{{\sf N}, t}^{(m)})^{^\intercal}$ and ${\boldsymbol\epsilon}_t^{(m)}=(\epsilon_{1,t}^{(m)},\ldots,\epsilon_{{\sf N}, t}^{(m)})^{^\intercal}$. Throughout the paper, we allow candidate time-varying network VAR models to be non-nested and misspecified, and both $\bar{\rho}$ and ${\sf M}$ may diverge to infinity.

For any given $t_0$, we define the local linear estimate of ${{\boldsymbol\theta}}_{t_0}^{(m)}$ as

equation[equation omitted — 349 chars of source]

where ${\mathbf I}_{\rho_m}$ and ${\mathbf O}_{\rho_m}$ are the identity and zero matrices, respectively, with size $\rho_m \times \rho_m$, \[ \widetilde{\mathbf{Z}}_{t-1}^{(m)}=\left[{\mathbf Z}_{t-1}^{(m){^\intercal}},\ \left(\frac{t-t_0}{Th}\right){\mathbf Z}_{t-1}^{(m){^\intercal}}\right],\quad {\bf Z}_{t}^{(m)}=({\bf Z}_{1,t}^{(m)},\ldots,{\bf Z}_{{\sf N},t}^{(m)}),\quad K_{t,t_0}=K\left(\frac{t-t_0}{Th}\right) \] with $K(\cdot)$ being a kernel function and $h$ being a bandwidth.

Let \[ {\mathscr W}=\left\{{\bf w}=(w^1,w^2,\ldots,w^{\sf M})^{{^\intercal}}\in[0,1]^{\sf M}:\ \sum_{m=1}^{\sf M} w^m=1\right\}. \] With a given vector ${\bf w}=(w^1,w^2,\ldots,w^{\sf M})^{{^\intercal}}\in{\mathscr W}$, we construct the model averaging estimators of time-varying parameters and conditional mean as follows

eqnarray[eqnarray omitted — 422 chars of source]

where ${\boldsymbol \Pi}^{(m)}$ is a matrix of size $\rho_m\times({\sf K}+{\sf P}+{\sf Q})$ mapping ${\bf Z}_{i,t}$ to ${\bf Z}_{i,t}^{(m)}$ and \[ \widehat{\mu}_{i,t}({\bf w})=\sum_{m=1}^{\sf M} w^m{\bf Z}_{i,t-1}^{(m)^{{^\intercal}}}\widehat{{\boldsymbol\theta}}_t^{(m)}. \] We propose the PTVMA criterion:

eqnarray[eqnarray omitted — 251 chars of source]

where $\lambda$ is a user-specified tuning parameter. The optimal weights are determined by

equation[equation omitted — 200 chars of source]

Finally the one-step ahead forecast combination is constructed by

eqnarray[eqnarray omitted — 429 chars of source]

Asymptotic theory

\setcounter{equation}{0}

In this section, we give some regularity conditions and present the asymptotic properties of the PTVMA estimation developed in Section (ref) including the asymptotic optimality and convergence rates. Section (ref) considers the in-sample fitting, i.e., $\widehat{\bf w}_t$ for the interior point $t$ in the sense that $0<t/{\sf T}<1$, whereas Section (ref) tackles the out-of-sample prediction, i.e., $\widehat{\bf w}_{\sf T}$ for any forecast time point ${\sf T}+1$.

In-sample asymptotic theory

To establish the asymptotic optimality of the proposed PTVMA method for in-sample fitting, we define the following infeasible locally weighted (in-sample) quadratic error loss:

equation[equation omitted — 199 chars of source]

where ${\boldsymbol\mu}_s=(\mu_{1,s},\ldots,\mu_{{\sf N}, s})^{^\intercal}$, $ \mu_{i,s} = {\bf Z}_{i,s-1}^{{^\intercal}} {\boldsymbol\theta}_{s} $, $t$ is a fixed interior point such that $0<t/{\sf T}<1$, in which case the local kernel smoothing utilizes the time series sample information on both sides of $t$. Denote $${\boldsymbol\mu}_t^*({\bf w})=\sum_{m=1}^{\sf M} w^m{\boldsymbol\mu}_{t,\ast}^{(m)}=\left[\mu_{1,t}^{*}({\bf w}),\ldots,\mu_{{\sf N},t}^{*}({\bf w})\right]^{^\intercal},$$ where ${\boldsymbol\mu}_{t,\ast}^{(m)}=(\mu_{1,t}^{(m)*},\ldots,\mu_{{\sf N},t}^{(m)*})^{^\intercal}$ and $\mu_{i,t}^{(m)*}={\bf Z}_{i,t-1}^{(m){^\intercal}}{\boldsymbol\theta}_{t,\ast}^{(m)}$ with ${\boldsymbol\theta}_{t,\ast}^{(m)}$ being the pseudo time-varying coefficient vector for the $m$-th candidate model to be defined in Assumption (ref) below. Write

equation[equation omitted — 260 chars of source]

The following assumptions are required to derive the in-sample asymptotic optimality.

condition(i) For the $m$-th candidate model, there exists a limit time-varying vector ${\boldsymbol\theta}_{s,\ast}^{(m)}$ such that $$\left\|\widehat{{\boldsymbol\theta}}_s^{(m)}-{\boldsymbol\theta}_{s,\ast}^{(m)}\right\|=O_P\left(\bar{\rho}^{1/2}\zeta\right)$$ uniformly over $m$ and $s$, where $\zeta$ is a typical nonparametric uniform convergence rate. In addition, elements of ${\boldsymbol\theta}_{t,\ast}^{{(m)}}$ are bounded over $t$ and $m$. (ii) The second moments for each element of ${\bf Z}_{i,t}$ exist and are uniformly bounded. (iii) Both true and pseudo time-varying coefficients $ {\boldsymbol\theta}_s $ and $ {\boldsymbol\theta}_{s,\ast}^{{(m)}} $ are Lipschitz continuous over $ [0, 1] $ and their difference $ \| {\boldsymbol\theta}_s - {\boldsymbol \Pi}^{{(m)}{^\intercal}} {\boldsymbol\theta}_{s,\ast}^{{(m)}} \| = O(\bar{\rho}^{1/2}) $ uniformly over $ m $ and $ s $.
condition(i)\ The kernel $K(\cdot)$ is positive, symmetric and continuous with a bounded support $[-1,1]$. (ii)\ $\bar{\rho}\zeta({\sf M}{\sf N}{\sf T} h)^{1/2}$ is larger than a positive number, and $$\bar{\rho}{\sf M}\zeta \rightarrow0,\quad \left(\lambda+ \bar{\rho}\zeta{\sf M}{\sf N}{\sf T} h\right)\bar{\rho}\xi_t^{-1}=o_P(1).$$

{\bf Remark 3.1}.\ \ Assumption (ref)(i) imposes a high-level condition on uniform convergence of local linear smoothing for misspecified candidate models, which is comparable to those adopted by some recent works on time-varying model averaging SHWZ23, SCG25. With standard techniques for locally stationary time series and nonparametric kernel-weighted smoothing LPTW26, we may derive an explicit rate: $\zeta=\sqrt{\hbox{log} {\sf T}/({\sf N}{\sf T} h)}+h^2$ under some sufficient conditions which require higher moment restriction than Assumption (ref)(ii). With this explicit rate for $\zeta$, assuming that both $\bar\rho$ and ${\sf M}$ are bounded, ${\sf N}{\sf T} h^5=O(1)$ and $\lambda=O(({\sf N} {\sf T} h\cdot\hbox{log} {\sf T})^{1/2})$, we may verify Assumption (ref)(ii) if \[ ({\sf N}{\sf T} h\cdot \hbox{log} {\sf T})^{1/2}/\xi_t=o_P(1). \] The latter indicating that {\em all the candidate models must be misspecified}, otherwise $\xi_t\equiv0$.

Theorem (ref) below establishes the point-wise asymptotic optimality of the PTVMA estimator, similar to Theorem 1 in SHWZ23 and Theorem 3.3 in CHL24.

theoremSuppose that Assumptions (ref) and (ref) are satisfied. For any given time $t$ such that $0<t/{\sf T}<1$, the proposed PTVMA estimator satisfies the asymptotic optimality, i.e., $$\frac{L_{t}(\widehat{{\bf w}}_t)}{\inf_{{\bf w}\in {\mathscr W}}L_{t}({\bf w})}\stackrel{P}\longrightarrow 1.$$

The following extra condition is required to derive an explicit in-sample convergence rate of the PTVMA weight estimation.

condition(i) There exists a positive constant $\underline{\kappa}$ such that $$ \lambda_{\min}\left(\frac{1}{{\sf N}}{\sf E}\left[\widetilde{{\boldsymbol\mu}}^*_{s}\widetilde{{\boldsymbol\mu}}_{s}^{*{^\intercal}}\right]\right)> \underline{\kappa}\quad \text{with}\quad \widetilde{\boldsymbol\mu}_{s}^* = \left({\boldsymbol\mu}_{s,\ast}^{(1)},\ldots,{\boldsymbol\mu}_{s,\ast}^{({\sf M})}\right)^{^\intercal} $$ for any $s$ such that $|s-t|\leq {\sf T} h$. (ii) For any given time point $t$, \[ \left\Vert\sum_{s=1}^{\sf T}\sum_{i=1}^{\sf N} K_{s,t}\left({\bf Z}_{i,s}{\bf Z}_{i,s}^{^\intercal}-{\sf E}\left[{\bf Z}_{i,s}{\bf Z}_{i,s}^{^\intercal}\right]\right)\right\Vert_{\sf F}=O_P\left(\bar\rho({\sf N} {\sf T} h)^{1/2}\right) \] and \[ \sum_{s=1}^{\sf T}\sum_{i=1}^{\sf N} \sigma_{i}^{2}(s/{\sf T})(u_{is}^2-1)K_{s,t}=O_P\left(({\sf N}{\sf T} h)^{1/2}\right). \] (iii) Let the following rate condition hold: $$\left(\lambda+\zeta{\sf M}{\sf N}{\sf T} h\right)\bar{\rho}({\sf N}{\sf T} h)^{-1}+\bar\rho^2({\sf N}{\sf T} h)^{-1/2}\rightarrow0$$

{\bf Remark 3.2.}\ \ Assumption (ref)(i) is comparable to some similar restriction on predictor vectors or the so-called “hat matrix". Assumption (ref)(ii) implies weak temporal dependence as well as cross-sectional correlation for predictors as in LPTW26. It is easy to verify Assumption (ref)(iii) when $\zeta=\sqrt{\hbox{log} {\sf T}/({\sf N}{\sf T} h)}+h^2$, $\lambda=O(({\sf N}{\sf T} h \hbox{log}{\sf T})^{1/2})$ and both $\bar\rho$ and ${\sf M}$ are bounded.

theoremSuppose Assumptions (ref)--(ref) are satisfied. For any given time point $t$ such that $0<t/{\sf T}<1$, $\widehat{{\bf w}}_t$ is a consistent in-sample estimate of ${\bf w}_t^*$, satisfying \begin{equation} \left\|\widehat{{\bf w}}_t-{\bf w}_t^*\right\|=O_P\left((\lambda+\zeta{\sf M}{\sf N}{\sf T} h)^{1/2}\bar{\rho}^{1/2}({\sf N}{\sf T} h)^{-1/2}+\bar\rho({\sf N}{\sf T} h)^{-1/4}\right), \end{equation} where ${\bf w}_t^*=\mbox{argmin}_{{\bf w}\in{\mathscr W}} {\sf E}[L_{t}({\bf w})]$.

{\bf Remark 3.3.}\ \ Theorem (ref) above reveals that the point-wise convergence rate of the in-sample time-varying weight estimation $\widehat{{\bf w}}_t$ relies on the number of candidate models ${\sf M}$, the maximum dimension of potential predictors $\bar\rho$, the tuning parameter in penalization $\lambda$, the bandwidth $h$ as well as $({\sf N}, {\sf T})$. With the sufficient conditions discussed in Remark 3.2 to verify Assumption (ref)(iii), we may simplify the rate in ((ref)) to \[ \left\|\widehat{\mathbf w}_t-{\mathbf w}_t^*\right\|=O_P\left(h+(\hbox{log} {\sf T})^{1/4}({\sf N}{\sf T} h)^{-1/4}\right). \]

We next consider the scenario when there exist correctly-specified candidate time-varying multi-layer network VAR models. Let ${\mathscr D}$ denote the index set of correctly-specified models, i.e., $m\in{\mathscr D}$ if the $m$-th candidate model is correctly specified. Similar to $\xi_t$ in ((ref)), we define $$\widetilde\xi_{t}=\inf_{{\bf w}\in{\mathscr W}({\mathscr D})} L_{t}^*({\bf w}),\quad {\mathscr W}({\mathscr D})=\{{\bf w}=(w^1,w^2,\ldots,w^{\sf M})^{{^\intercal}}\in{\mathscr W}:\ w^m=0, \ m\in{\mathscr D}\}.$$ Theorem (ref) below shows that the proposed PTVMA criterion achieves the in-sample selection consistency (of correctly-specified candidate models) as the conventional model selection.

theoremSuppose Assumptions (ref), (ref) and (ref)(ii) are satisfied but with the condition $\left(\lambda+\zeta{\sf M}{\sf N} Th\right)\bar{\rho}\xi_t^{-1}=o_P(1)$ in Assumption (ref)(ii) replaced by \begin{equation} \left[\bar\rho(\lambda+ \zeta{\sf M}{\sf N} Th)+\bar\rho^2({\sf N} Th)^{1/2}\right]\widetilde\xi_t^{-1}=o_P(1), \end{equation} and ${\mathscr D}$ is not empty. Then, $\sum_{m\in \mathscr{D}} \widehat{w}_t^m\stackrel{P}\longrightarrow 1$ for any interior time point $t$.

Out-of-sample asymptotic theory

We next consider the asymptotic properties of $\widehat{\bf w}_{\sf T}$ which has been adopted for constructing out-of-sample prediction combination. In this case, the local linear smoothing and ${\sf PTVMA}_{T}({\bf w})$ essentially use one-sided local sample information, which would subsequently slow down the weight estimation convergence. Define the out-of-sample prediction risk: \[ R_{{\sf T}+1}({\bf w})={\sf E}\left\{\left[{\bf Y}_{{\sf T}+1}-\widehat{{\bf Y}}_{{\sf T}+1}({\bf w})\right]^{^\intercal}\left[{\bf Y}_{{\sf T}+1}-\widehat{{\bf Y}}_{{\sf T}+1}({\bf w})\right]\right\}-\frac{1}{{\sf T} h}\sum_{i=1}^{\sf N}\sum_{s=1}^{\sf T}\sigma_{i}^{2}(s/{\sf T})K_{s,{\sf T}}, \] where $\widehat{{\bf Y}}_{{\sf T}+1}({\bf w})$ is defined similarly to $\widehat{{\bf Y}}_{{\sf T}+1}(\widehat{\bf w}_{\sf T})$ defined in ((ref)) but with $\widehat{\bf w}_{\sf T}$ replaced by ${\bf w}$. Note that the second term of $R_{{\sf T}+1}({\bf w})$ is independent of ${\bf w}$. In the out-of-sample prediction, we expect the PTVMA estimate $\widehat{\bf w}_{\sf T}$ to asymptotically minimize $R_{{\sf T}+1}({\bf w})$ over ${\bf w}\in{{\mathscr W}}$. Furthermore, we let \[ R_{{\sf T}+1}^*({\bf w})={\sf E}\left\{\left[{\bf Y}_{{\sf T}+1}-{\bf Y}_{{\sf T}+1}^*({\bf w})\right]^{^\intercal}\left[{\bf Y}_{{\sf T}+1}-{\bf Y}_{{\sf T}+1}^*({\bf w})\right]\right\}-\frac{1}{{\sf T} h}\sum_{i=1}^{\sf N}\sum_{s=1}^{\sf T}\sigma_{i}^{2}(s/{\sf T})K_{s,{\sf T}}, \] and \[ \xi_{{\sf T}+1}^*=\inf_{{\bf w}\in {\mathscr W}}R_{{\sf T}+1}^*({\bf w}), \] where ${\bf Y}_{{\sf T}+1}^\ast({\bf w})=\sum_{m=1}^{\sf M} w^m{\bf Y}_{{\sf T}+1}^{(m)\ast}$ and ${\bf Y}_{{\sf T}+1}^{(m)\ast}=(y_{1,{\sf T}+1}^{(m)*},\ldots,y_{{\sf N},{\sf T}+1}^{(m)*})^{^\intercal}$ with $y_{i,{\sf T}+1}^{(m)*}={\bf Z}_{i,{\sf T}}^{(m){^\intercal}}{\boldsymbol\theta}_{{\sf T},\ast}^{(m)}$. Let ${\boldsymbol\epsilon}_{t}^{(m)*}=(\epsilon_{1,t}^{(m)*},\ldots,\epsilon_{{\sf N},t}^{(m)*})^{^\intercal}$ with $\epsilon_{i,t}^{(m)*}=y_{i,t}-\mu_{i,t}^{(m)*}$, $1\leq t\leq {\sf T}$, and ${\boldsymbol\epsilon}_{{\sf T}+1}^{(m)*}=(\epsilon_{1,{\sf T}+1}^{(m)*},\ldots,\epsilon_{{\sf N}, {\sf T}+1}^{(m)*})^{^\intercal}$ with $\epsilon_{i, {\sf T}+1}^{(m)*}=y_{i, {\sf T}+1}-y_{i, {\sf T}+1}^{(m)*}$ for the out-of-sample period.

condition(i) $\sigma_i(\cdot)$, $1\leq i\leq {\sf N}$, are uniformly bounded and Lipschitz continuous. (ii) There exists $\delta_1>0$ such that \begin{equation} \sup_{\sf T}{\sf E}\left[\xi_{{\sf T}+1}^{*-1/2}\left(\max_{1\leq m\leq {\sf M}}\left\|\widehat{{\bf Y}}_{{\sf T}+1}^{(m)}-{\bf Y}_{{\sf T}+1}^{(m)*}\right\|+\max_{1\leq m,m'\leq {\sf M}}\left\|\widehat{{\bf Y}}_{{\sf T}+1}^{(m)}-{\bf Y}_{{\sf T}+1}^{(m)*}\right\|^{1/2} \left\|{\boldsymbol\epsilon}_{{\sf T}+1}^{(m')*}\right\|^{1/2}\right)\right]^{2+\delta_1}<\infty, \end{equation} and in addition, $\sup_{\sf T}\xi_{{\sf T}+1}^*/{\sf N}<\infty$. (iii) Let the following conditions hold: \[ \xi_{{\sf T}+1}^{\ast-1}\bar\rho \left[{\sf M}{\sf N} \zeta+\lambda({\sf T} h)^{-1}\right]\rightarrow0,\quad \xi_{{\sf T}+1}^{\ast-1}\left[\bar{\rho} {\sf N} h+{\sf M}^2{\sf N}^{1/2}({\sf T} h)^{-1/2}\right]\rightarrow0. \] (iv) For any given time point $t$ such that ${\sf T}- {\sf T} h \leq t\leq {\sf T}$, $$ \left\| \frac{1}{{\sf T} h} \sum_{t=1}^{{\sf T}} K_{t,{\sf T}} {\sf E} \left[{\bf Z}_{i,t-1}{\bf Z}_{i,t-1}^{^\intercal}\right] - {\sf E} \left[{\bf Z}_{i,{\sf T}}{\bf Z}_{i,{\sf T}}^{^\intercal}\right] \right\|_{\sf op} = O_P \left(h\right), $$ where $\|\cdot\|_{\sf op}$ denotes the operator norm of a square matrix.

{\bf Remark 3.4.}\ \ Assumption (ref)(i) imposes a mild smoothness restriction on heterogeneous variance functions $\sigma_i^2(\cdot)$ and is automatically satisfied for the case of homoskedasticity. The condition ((ref)) requires the potential best prediction (over all the candidate models) has finite risk in large-scale networks. A similar condition can be found in Assumptions 3(iii) and 5(iii) of ZZ23. In fact, it provides a sufficient condition to derive the uniform integrability (over ${\bf w}$) of $\xi_{{\sf T}+1}^{*-1}\left|\|{\bf Y}_{{\sf T}+1}-\widehat{{\bf Y}}_{{\sf T}+1}({\bf w})\|^2-\|{\bf Y}_{{\sf T}+1}-{\bf Y}^*_{{\sf T}+1}({\bf w})\|^2\right|$, which is required in the proof of Theorem (ref) in the appendix. If, in addition, $\sup_{\sf T}\xi_{{\sf T}+1}^*/{\sf N}$ is strictly larger than a positive constant, we may replace $\xi_{{\sf T}+1}^\ast$ by ${\sf N}$ in ((ref)) and further show that all the candidate models are misspecified. In fact, if the $m_0$-th candidate model is correctly specified, ${\boldsymbol \Pi}^{(m_0){^\intercal}}\widehat{\boldsymbol\theta}_{\sf T}^{(m_0)}$ converges to the true coefficient vector ${\boldsymbol\theta}_{\sf T}$. Consequently, ${\bf Y}_{{\sf T}+1}^{(m_0)*}={\boldsymbol\mu}_{{\sf T}+1}(1+o_P(h))$ and \[ \xi_{{\sf T}+1}^{*}=\inf_{{\bf w}\in{\mathscr W}}R_{{\sf T}+1}^*({\bf w})\leq {\sf E}\left[\left({\bf Y}_{{\sf T}+1}-{\bf Y}_{{\sf T}+1}^{(m_0)*}\right)^{^\intercal}\left({\bf Y}_{{\sf T}+1}-{\bf Y}_{{\sf T}+1}^{(m_0)*}\right)\right]-\frac{1}{{\sf T} h}\sum_{i=1}^{\sf N}\sum_{s=1}^{\sf T}\sigma_{i}^{2}(s/{\sf T})K_{s,{\sf T}}=o({\sf N}), \] which leads to contradiction. Assumption (ref)(iii) is comparable to Assumption (ref)(ii) and can be easily justified by using the sufficient conditions discussed in Remarks 3.1 and 3.2. Assumption (ref)(iv) contains a high-level condition for locally stationary ${\bf Z}_{i,t}$.

theoremSuppose that Assumptions (ref), (ref)(i) and (ref) are satisfied. For any forecast time point ${\sf T}+1$, $$\frac{R_{{\sf T}+1}(\widehat{{\bf w}}_{\sf T})}{\inf_{{\bf w}\in {\mathscr W}}R_{{\sf T}+1}({\bf w})}\stackrel{P}\rightarrow 1.$$

{\bf Remark 3.5.}\ \ Theorem (ref) above shows that the PTVMA estimator is out-of-sample asymptotically optimal in the sense that its out-of-sample prediction risk is asymptotically equivalent to that of the infeasible optimal weight estimator ${\bf w}_{\sf T}^\ast:=\inf_{{\bf w}\in {\mathscr W}}R_{{\sf T}+1}({\bf w})$. A similar property is also derived by ZL23 and ZZ23. A combination of Theorems (ref) and (ref) demonstrates that our proposed PTVMA estimator achieves not only the minimum of in-sample estimation error at the interior time point but also the minimum of out-of-sample prediction risk.

To derive an explicit out-of-sample convergence rate of $\widehat{{\bf w}}_{\sf T}$, we require the following assumption which is analogous to Assumption (ref)(i)(iii) in Section (ref).

condition(i) There exists a positive constant $\underline{\kappa}_\ast$ such that \[ \lambda_{\min}\left(\frac{1}{{\sf N}}{\sf E}\left[\widetilde{{\bf Y}}^*_{{\sf T}+1}\widetilde{{\bf Y}}_{{\sf T}+1}^{*{^\intercal}}\right]\right)> \underline{\kappa}_\ast,\quad \widetilde{\bf Y}_{{\sf T}+1}^* = \left({\bf Y}_{{\sf T}+1,\ast}^{(1)},\ldots,{\bf Y}_{{\sf T}+1,\ast}^{({\sf M})}\right)^{^\intercal}. \] (ii) Let the following condition hold: \[ \left(\lambda+\zeta{\sf M}{\sf N}{\sf T} h\right)\bar{\rho}({\sf N}{\sf T} h)^{-1}+{\sf M}^2({\sf N}{\sf T} h)^{-1/2}\rightarrow0. \]
theoremSuppose that Assumptions (ref), (ref)(i), (ref)(ii)(iii), (ref)(i) and (ref) are satisfied, and ${\bf w}_{\sf T}^*$ uniquely minimizes $R_{{\sf T}+1}({\bf w})$ in ${\mathscr W}$. Then, for any forecast time point ${\sf T}+1$, $\widehat{{\bf w}}_{\sf T}$ is a consistent estimate of ${\bf w}_{\sf T}^\ast$, satisfying \begin{equation} \left\Vert\widehat{{\bf w}}_{\sf T}-{\bf w}_{\sf T}^*\right\Vert=O_P\left( (\lambda+\zeta{\sf M}{\sf N}{\sf T} h)^{1/2}\bar{\rho}^{1/2}({\sf N}{\sf T} h)^{-1/2}+{\sf M}({\sf N}{\sf T} h)^{-1/4}+ (\bar{\rho} h)^{1/2 } \right). \end{equation}

{\bf Remark 3.6.}\ \ As in Theorem (ref), the convergence rate of $\widehat{{\bf w}}_{\sf T}$ relies on ${\sf M}$, $\bar\rho$, $\lambda$, $h$ and $({\sf N}, {\sf T})$. In fact, the rate in ((ref)) is slightly slower than that in ((ref)) due to the extra term $h^{1/2}$, which is caused by the well-known boundary effect of kernel smoothing. Furthermore, as discussed in Remark 3.3, under some sufficient conditions, we may simplify the rate in ((ref)) to \[ \left\|\widehat{\mathbf w}_{\sf T}-{\mathbf w}_{\sf T}^*\right\|=O_P\left(h^{1/2}+(\hbox{log} {\sf T})^{1/4}({\sf N} {\sf T} h)^{-1/4}\right). \]

We finally consider the case when there exist correctly-specified candidate models. Let ${\mathscr D}$ and ${\mathscr W}({\mathscr D})$ be defined as in Section (ref) and $$\widetilde\xi_{{\sf T}+1}^{*}=\inf_{{\bf w}\in{\mathscr W}({\mathscr D})}R_{{\sf T}+1}^*({\bf w}),$$ denoting the minimum out-of-sample prediction risk over misspecified models. The following theorem is an extension of Theorem (ref) from in-sample fitting to out-of-sample prediction.

theoremSuppose that Assumptions (ref), (ref)(i), (ref)(ii)(iii) and (ref)(i) are satisfied. In addition, \begin{equation} \widetilde\xi_{{\sf T}+1}^{\ast-1}\bar\rho \left[{\sf M}{\sf N} \zeta+\lambda({\sf T} h)^{-1}\right]+ \widetilde\xi_{{\sf T}+1}^{\ast-1}\left[\bar{\rho} {\sf N} h+{\sf M}^2{\sf N}^{1/2}({\sf T} h)^{-1/2}\right]=o_P(1), \end{equation} and ${\mathscr D}$ is not empty. Then, for any forecast time point ${\sf T}+1$, we have $\sum_{m\in {{\mathscr D}}} \widehat{w}_{\sf T}^m\stackrel{P}\rightarrow 1$.

PTVMA-based conformal prediction

\setcounter{equation}{0}

Despite the growing literature on frequentist model averaging, construction of prediction intervals using model averaging estimators remain under-explored. A commonly-used approach to constructing interval forecasts often relies on the asymptotic distribution of the model averaging estimator. The relevant literature can be divided into two strands. The first examines the limiting distributions of least squares averaging estimators within a local asymptotic framework, where time-invariant regression coefficients lie in a local ${\sf T}^{-1/2}$ neighborhood of zero HC03, H14,Liu15. However, the realism of this local misspecification assumption has been questioned by RZ03, who argue that it lacks face validity in certain empirical settings. The second strand relies on the asymptotic distribution of model averaging estimators under proper model assumptions such as the (non-)linear regression with fixed parameters ZL18,YLSZH24 and time-varying coefficient regressions SHWZ23,TW25. These results are restricted to the setting where both the number of candidate models and the number of predictors remain fixed. We next propose the PTVMA-based conformal prediction approach to construct interval forecasts without requiring the number of candidate models or predictors to be fixed; see Algorithm (ref). Without loss of generality, we assume that ${\sf T} h$ takes a positive integer value.

algorithm[algorithm omitted — 2,245 chars of source]

To establish a valid coverage property of the proposed conformal prediction procedure, we require the following technical assumptions.

condition(i) $\sigma_i(\cdot)$ is second-order continuously differentiable, and $\underline{\sigma}\leq\sigma_i(\cdot)\leq\overline\sigma$, where $\underline{\sigma}$ and $\overline\sigma$ are finite positive constants. (ii) $\{u_{i,t}\}$ are i.i.d. with mean zero, unit variance and ${\sf E} [u_{i,t}^4]<\infty$. Furthermore, their common cumulative distribution function ${\sf F}$ satisfies a Lipschitz condition, i.e., \[\sup_{x_1,x_2}|{\sf F}(x_1)-{\sf F}(x_2)|\leq M_{\sf F}|x_1-x_2|,\quad 0<M_{\sf F}<\infty. \] (iii) Let $|{\boldsymbol\theta}_{{\sf T}+1} - {\boldsymbol\theta}_{\sf T} | \leq M_\theta/{\sf T}$ with $M_\theta$ being a positive constant.

Letting $\widehat{u}_{i,t}=\sigma_i^{-1}(\tau_t)\widehat{\epsilon}_{i,t}$ with $\tau_t=t/{\sf T}$, we define the following in-sample and out-of-sample performance measures: $$\delta_{i,{\sf T}}^2=\frac{1}{{\sf T} h+1}\sum_{t={\sf T}-{\sf T} h}^{{\sf T}}(|\widehat{u}_{i,t}|-|u_{i,t}|)^2\quad \text{and}\quad \psi_{i,{\sf T}+1}=\big||\widehat{u}_{i,{\sf T}+1}|-|u_{i, {\sf T}+1}|\big|.$$ The following proposition gives the finite-sample upper bound on the gap between the actual coverage probability and the nominal level $1-\alpha$ and further establishes the asymptotic validity of the developed PTVMA-based conformal prediction method.

propositionSuppose that Assumption (ref) is satisfied. There exists a positive constant $c_\dag$ such that \begin{equation} \left|{\sf P}\left(y_{i,{\sf T}+1}\in C_{i}^{\alpha}({\bf Z}_{{\sf T}})\right)-(1-\alpha)\right|\leq c_\dag \left(\left[{\sf E}\left(\delta_{i,{\sf T}}^{2}\right)\right]^{1/4}+\left[{\sf E}\left(\psi_{i,{\sf T}+1}\right)\right]^{1/2}+h^{4/5}+({\sf T} h)^{-1/3}\right) \end{equation} for any fixed $i$, if $\left[{\sf E}\left(\psi_{i,{\sf T}+1}\right)\right]<1$. Furthermore, if $$ \bar{\rho} \widetilde\xi_{{\sf T}+1}^{\ast-1} \left[ \bar\rho {\sf M}{\sf N} \zeta + \bar\rho\lambda({\sf T} h)^{-1} + \bar{\rho} {\sf N} h + {\sf M}^2{\sf N}^{1/2}({\sf T} h)^{-1/2} \right]= o_P(1) $$ and ${\mathscr D}$ is not empty, we have \begin{equation} {\sf P}\left(y_{i,{\sf T}+1}\in C_{i}^{\alpha}({\bf Z}_{{\sf T}})\right)=1-\alpha+o(1). \end{equation}

{\bf Remark 4.1.}\ \ We allow all the candidate models to be misspecified to obtain the finite-sample upper bound for the deviation of the actual coverage probability from the nominal level in ((ref)). To further derive the asymptotic validity of the constructed prediction interval in ((ref)), we need to prove that ${\sf E}(\delta_{i,{\sf T}}^{2})\rightarrow0$ and ${\sf E}(\psi_{i,{\sf T}+1})\rightarrow0$, which requires ${\mathscr D}$ to be not empty, indicating that there exists at least one correctly-specified candidate model so that Theorem (ref) is applicable.

Monte-Carlo simulation

\setcounter{equation}{0}

To demonstrate advantages of the proposed PTVMA procedure, this section creates a sharp challenge for classic approaches in which any single-network specification may be misspecified. In the two simulation settings, model selection over a set of single-network specifications is global misspecified because no candidate model coincides with the true DGP throughout the entire period. By comparing the finite-sample performance of PTVMA with several competing models, our simulation experiment examines whether the proposed PTVMA can adaptively learn relevant network layers and variables over time and deliver accurate predictions despite persistent specification uncertainty.

Evaluation measures on finite-sample performance

To evaluate the in-sample performance, we calculate the in-sample ratio of the root mean squared prediction errors (RMSPE) to compare the proposed model with benchmark models:

align*[align* omitted — 290 chars of source]

where $R=1000$ is the replication number in simulations, $y_{it:r} $ is the true value in the $r$-th replication, $ \widehat{y}_{it:r}^\text{PTVMA}$ is the fitted value from the PTVMA, and $ \widehat{y}_{it:r}^\text{Benchmark} $ is from a benchmark model. We consider two benchmarks. The network vector autoregressive (NAR) model ZPLLW17 includes the same response and predictors as the DGP but restricts the parameters to be time-invariant, estimating them via the OLS. The time-varying NAR (tv-NAR) model LPTW26 adopts the same DGP specification, utilizing a local linear approach to estimate time-varying coefficients. The proposed PTVMA method outperforms the benchmark model when the ${\sf RMSPE}_\text{in}$ ratio is smaller than one.

The out-of-sample predictive performance is assessed by employing the rolling forecasting procedure. Specifically, we compute the out-of-sample ratio of the RMSPE for the $l$-step-ahead out-of-sample rolling prediction of $ \widehat{\bm{Y}}_{{\sf T}+l} $:

align*[align* omitted — 332 chars of source]

where $ \widehat{y}_{i,t+l:r}^\text{PTVMA}$ and $\widehat{y}_{i,t+l:r}^\text{Benchmark}$ denotes the $l$-step-ahead rolling prediction in the $r$-th replication using the proposed PTVMA and benchmark methods, respectively, ${\sf T}_0 = \lfloor 0.4{\sf T} \rfloor $ is the rolling window length. Under the iterated scheme, multi-step-ahead forecasts are obtained by estimating a one-step-ahead model and recursively replacing future lagged values with their forecasts.

The performance of conformal prediction is evaluated by the empirical coverage probability. Specifically, we use the test set to calculate empirical coverage probabilities:

align*[align* omitted — 198 chars of source]

where $C_{i:r}^{\alpha}(\mathbf{Z}_{t-1})$ is the conformal prediction interval constructed for individual $i$ at time $t$ in the $r$-th replication.

One candidate model is correctly specified over half-period.

The DGP 1 simulates a setting in which the response variable $ y_{it} $ is driven by two distinct networks, $\mathbf{W}^{(1)}$ and $\mathbf{W}^{(2)}$, with time-varying impacts $\beta_{t1}$ and $\beta_{t2}$. Consider

align[align omitted — 289 chars of source]

where $ \widetilde{\omega}_{ij}^{(k)} = \omega_{ij}^{(k)} / \sum_{j \neq i} \omega_{ij}^{(k)} $ is the $ (i,j) $-th element of the matrix obtained by row-normalizing the adjacency matrix $ \mathbf{W}^{(k)} $, $ \mathbf{X}_i = \left( x_{i,1}, x_{i,2}, x_{i,3} \right)^{^\intercal} $ is independently drawn from a three-dimensional normal distribution across $ i $ with mean $ \bm{0} $ and covariance $ \bm{\Sigma}_X = (\sigma_{x,q_1q_2})_{3 \times 3} $ with $ \sigma_{x,q_1q_2} = 0.5^{|q_1-q_2|} $ for $ 1 \leq q_1, q_2 \leq 3 $, the error term $ \epsilon_{i,t} = \sigma_i(t/{\sf T}) u_{i,t} $, $ u_{i,t} $ is independently simulated by the standard normal distribution, $ \sigma_i(t/{\sf T}) = (0.25 \sin (t/{\sf T})+0.3)^{0.5} $ for odd $ i $, and $ \sigma_i(t/{\sf T}) = (0.25 \cos (t/{\sf T})+0.3)^{0.5} $ for even $ i $. True time-varying coefficients are set as

eqnarray[eqnarray omitted — 380 chars of source]

The first adjacency matrix $ \mathbf{W}^{(1)} $ is generated by the stochastic block model with two blocks. Each node is randomly assigned for one block label with equal probability and probabilities within and between blocks are generated by: $ {\sf P} (\omega_{ij}^{(1)} = 1) = 0.5 $ if $ i $ and $ j $ are in the same block; $ {\sf P} (\omega_{ij}^{(1)} = 1) = 0.1 $ otherwise. The second adjacency matrix $ \mathbf{W}^{(2)} $ is generated by using the power-law distribution model. The power-law distribution of degrees widely exists in the real world, where most agents have low degrees while a few hubs have very high degrees. We follow ZPLLW17 and define $ \mathbf{W}^{(2)} $ as follows. First, generate for each agent's $ degree_i $ via the discrete power-law distribution $ {\sf P} ({\sf degree}_i = k) \propto k^{-c} $, where $c $ is the exponent parameter and a smaller value implies a heavier distribution tail.\footnote{Here, we set $ c = 2.5 $ and the created networks share a similar density size at around $ 20\% $.} Next, we randomly select ${\sf degree}_i $ agents to be agent $ i $'s followers.

We consider $ {\sf M} = 8 $ candidate models in the proposed PTVMA for DGP 1 and list their respective predictors in Table (ref). Note that $ \beta_{t1} $ is significant in the first half of the simulation period but not in the second, while $ \beta_{t2} $ exhibits the reverse pattern that is insignificant in the first half and significant in the second. Consequently, candidate models 2 and 4 are correctly specified in the first and second half of the sample period, respectively, but misspecified in the other period. We also involve irrelevant predictors $ x_{i4} $, $ x_{i5} $, $ \sum_{j \neq i} \widetilde{\omega}_{ij}^{(3)} y_{j,t-1} $ and $ \sum_{j \neq i} \widetilde{\omega}_{ij}^{(4)} y_{j,t-1} $ into some candidate models to demonstrate that the PTVMA method can dynamically select the correctly specified model if it is in the candidate models. Irrelevant predictors $ x_{i4} $ and $ x_{i5} $ are independently drawn from a normal distribution with mean $ 1 $ and variance $ 1 $. We refer to ZY18 and adopt the left-right matrix as one irrelevant adjacency matrix $ \mathbf{W}^{(3)} $, where each agent interacts only with its left and right neighbors. The other irrelevant adjacency matrix $ \mathbf{W}^{(4)} $ is constructed by a distance-based spatial weight matrix. We first independently generate agent $ i $'s location $ (x_i, y_i) $ from $ \mathrm{U} [0, 1] $, and then compute the pairwise Euclidean distance $ d_{ij} $ between agent $ i $ and $ j $. The spatial weight matrix is constructed using an exponential distance-decay function, with off-diagonal elements $ \omega_{ij}^{(4)} = \mathrm{I} \left( \exp \{ -10 d_{ij} \} < d_0 \right) $ and zero diagonal, where $ d_0 $ is the 25% quantile of $ \{ d_{ij} \} $.

As reported in Table (ref), the PTVMA delivers uniformly better in-sample fit than the NAR benchmark. The RMSPE ratio (PTVMA/NAR) ranges from $ 0.746 $ to $ 0.825 $ across all $ ({\sf N}, {\sf T}) $ designs, indicating significant reductions of in-sample prediction error. Furthermore, PTVMA delivers nearly the same in-sample performance as the correct specification (tv-NAR model), with RMSPE ratios tightly around $ 1.002 $ and decreasing as either $ {\sf N} $ or $ {\sf T} $ grows. Given that the tv-NAR model is correctly specified here, this comparison suggests that the PTVMA approximates the true model specification without priori information.

Table (ref) also reports the average weights of eight candidate models over the $ 1000 $ replications. The estimated weights are highly concentrated on candidate model 2 and 4, each receiving about $ 0.48 $ and jointly accounting for more than $ 95\% $ of the total weight. In contrast, the (averaged) time-varying weight estimates for the other six candidate models are negligible across all $ {\sf N} $ and $ {\sf T} $. This pattern is consistent with the design of DGP 1: the true network is generated from a combination of $ \mathbf{W}^{(1)} $ and $ \mathbf{W}^{(2)} $ while $ \mathbf{W}^{(3)} $ and $ \mathbf{W}^{(4)} $ are irrelevant alternatives. Overall, the evidence indicates that PTVMA places most of the weight on DGP-relevant network components while shrinking irrelevant candidates toward zero, confirming Theorem (ref) in finite samples.

table[table omitted — 3,586 chars of source]
table[table omitted — 2,263 chars of source]

Figure (ref) illustrates a central implication of the proposed PTVMA: when the strength of network dependence is time-varying, the optimal weights are expected to be time-varying as well. The solid lines report the averaged time-varying weights on candidate model 2 and 4, and the dashed lines plot the true coefficient functions: $ \beta_{t1} $ and $ \beta_{t2} $. The plots show that the assigned weights closely track the DGP 1: the PTVMA places most weight on candidate model 2 in the first half of the sample period and rapidly shifts weight toward candidate model 4 in the second half, consistent with the time-varying pattern of true coefficient functions. In addition, the weight reallocation is concentrated around the mid-sample transition where $ \beta_{t1} $ and $ \beta_{t2} $ switch most sharply, indicating PTVMA's strong adaptivity to smooth structural changes.

figure[figure omitted — 2,233 chars of source]

Table (ref) reports out-of-sample RMSPEs of the PTVMA relative to two benchmarks under DGP 1. Nearly all reported ratios fall below unity, confirming that the PTVMA uniformly dominates both the NAR and the tv-NAR models across all forecast horizons and sample sizes. In particular, the PTVMA outperforms the tv-NAR model even though the latter is correctly specified under DGP 1. The proposed PTVMA achieves the forecast gain by adaptively weighting candidate models as network dependence varies over time, and thus delivers higher forecast accuracy than a correctly specified benchmark model. The last column of Table (ref) reports the empirical coverage probabilities of conformal prediction intervals at 90% nominal level. The coverage probabilities are close to 90% in general, suggesting that the conformal intervals are well-calibrated in finite samples.

table[table omitted — 3,263 chars of source]

All the candidate models are misspecified.

In the DGP 2 below, we consider a more challenging setting with all the single-network specifications misspecified over any time period. Define

align[align omitted — 289 chars of source]

where the true time-varying coefficients are set as

eqnarray[eqnarray omitted — 213 chars of source]

indicating that the impact of $ \mathbf{W}^{(1)} $ diminishes and that of $ \mathbf{W}^{(2)} $ grows over the time. The other model settings are the same as those in DGP 1. The predictors of candidate models for DGP 2 are listed in Table (ref).

The in-sample performance of the PTVMA under DGP 2 exhibits a similar pattern to that in DGP 1, demonstrating the PTVMA's advantages in the misspecification scenario. As reported in Table (ref), the PTVMA substantially outperforms the NAR model (RMSPE ranging from $ 0.550 $ to $ 0.570 $) while delivers a near-identical fit to the correctly specified tv-NAR model (RMSPE tightly around $ 1.001 $). This result confirms that even when all the candidate models are misspecified, the proposed PTVMA can adaptively combine them to approximate the response as effectively as a correctly-specified model.

table[table omitted — 2,080 chars of source]

Table (ref) also reports the averaged weight estimates of six candidate models over $1000$ replications to further reveal PTVMA's weight-assignment capacity. Due to incorporating both autoregressive lags, candidate models 3 and 6 receive the majority of weight with approximately $ 0.41 $ and $ 0.31 $, respectively. Note that the other candidate models receiving negligible weights either omit lags or contain irrelevant predictors. Consequently, these specifications are less informative than candidate models 3 and 6 under the DGP 2. The results confirm that the proposed PTVMA successfully allocates significant weights to the candidate models with most relevant predictors whereas effectively shrinks the other irrelevant specifications.

Table (ref) presents out-of-sample RMSPE for both the one- and multiple-step-ahead forecasts. Across all horizons and sample sizes, the PTVMA consistently outperforms two benchmarks and achieves lower forecast errors than the correctly specified model (tv-NAR), again demonstrating that the proposed model averaging procedure delivers forecast gains by adaptively combining information from candidate models even when all of them are misspecified.

table[table omitted — 2,645 chars of source]

An empirical application

\setcounter{equation}{0}

Accurate forecasting of Consumer Price Index (CPI) inflation plays a critical role in the conduct of monetary policy and preservation of financial stability. Inflation expectations serve as a key anchor for policy credibility and exert a substantial influence on asset pricing. Nevertheless, forecasting inflation remains a persistent challenge, even when employing comprehensive datasets and state-of-the-art econometric methods SW07. These difficulties are compounded in open economies, where domestic inflation is increasingly shaped by global shocks and the cross-country transmission of price pressures. A growing body of literature demonstrates that inflation dynamics are increasingly dominated by global determinants, even as domestic conditions remain relevant CM10. Key international transmission channels include global production networks, international financial markets, trade and policy shocks CNST21, MR20, ARW19. These channels subsequently form multi-layer networks that evolve over time, creating substantial model uncertainty and structural instability. Traditional univariate or panel models often fail to capture complex cross-country spillovers and time-varying nature of inflation transmission, leading to potentially misspecified models and degraded forecast performance. This section aims to address this issue by combining information from multiple international networks and applying the proposed PTVMA method to forecast CPI inflation,

Data

The CPI inflation data is obtained from the International Monetary Fund (IMF)\footnote{The data is downloaded from the IMF Data, available at: \url{https://data.imf.org/en}.}. We consider a dataset on monthly CPI inflation in 36 economies ($ {\sf N} = 36 $) from Januaray 2007 to December 2024 with 216 observations ($ {\sf T} = 216 $). This economy set covers the largest and most systemically relevant economies with established equity markets, ensuring representation of advanced economies, major emerging markets, and key commodity producers that jointly account for the overwhelming majority of global output.\footnote{Selected economies include Belgium, Brazil, Canada, Switzerland, Chile, China, Colombia, the Czech Republic, Germany, Denmark, Spain, Finland, France, the United Kingdom, Greece, Hong Kong, Hungary, Indonesia, India, Ireland, Israel, Italy, Japan, South Korea, Mexico, Malaysia, the Netherlands, Norway, Poland, Portugal, Singapore, Sweden, Turkey, the United States, Vietnam, and South Africa. We exclude nine candidate economies due to data limitations during the sample period: Russia and Saudi Arabia are omitted due to missing daily MSCI indices, while the United Arab Emirates, Argentina, Australia, New Zealand, Peru, the Philippines, and Thailand are excluded due to the unavailability of monthly inflation data.} We include two variables as controls in the model: a time-varying trend term, capturing exogenous shocks such as oil price fluctuations; and the monthly period-average nominal exchange rate for economy $ i $ (domestic currency per U.S. dollar), sourced from the IMF Data.

We consider four types of networks to characterize the international transmission channels of CPI inflation, each reflecting a distinct economic mechanism. The economic rationale for each network and the corresponding data sources are discussed below, with detailed network construction procedures presented in Appendix (ref).

itemize• Global production networks. Global production networks transmit cost shocks across borders through input-output linkages, generating price co-movement that extends beyond direct bilateral trade ALS19, CNST21. Since the CPI measures prices of final goods and services consumed by households, we construct the production network layer by using the origin of value added embodied in final demand.\footnote{The data is downloaded from the Trade in Value-Added (TiVA) database, available at: \url{https://www.oecd.org/en/topics/sub-issues/trade-in-value-added.html}.} Within this framework, an inflation increase in economy $ i $ raises the cost of its value added, which subsequently transmits through global value chains and becomes embodied in final goods consumed in economy $ j $, thereby exerting upward pressure on economy $ j $'s CPI. • Global equity networks. The integration of national equity markets affects cross-country inflation dynamics through information diffusion and the formation of inflation expectations MR20, CM10, CCL09. As institutional ownership and algorithmic trading become more prevalent globally, cross-market equity linkages generate forward-looking signals about macroeconomic fundamentals. Equity-market connectedness thus serves as a leading transmission channel through which global financial conditions shape inflation dynamics across countries, often ahead of real-sector and policy channels. We construct the global equity network using daily MSCI country-level equity indices (standard, price-return, U.S. dollar–denominated) and the VAR decomposition framework of DY14.\footnote{Data are obtained from the MSCI database. \url{https://www.msci.com/indexes\#featured-indexes}.} These MSCI indices provide comparable, internationally tradable measures of equity price movements and capture high-frequency financial channels relevant for inflation transmission. • Trade and policy networks. Trade policy acts as a distinct channel of international inflation transmission by affecting import prices and the cost of intermediate inputs. Evidence from the 2018–2019 U.S.–China trade conflicts indicates that tariff costs pass through to consumer prices and propagate along input–output linkages, generating inflationary effects that extend beyond directly targeted goods FGKK20, ARW19. To disentangle market-based trade integration from policy-driven exposure, we construct two types of network layers capturing trade linkages and political connectedness, respectively. The bilateral trade linkages in the trade network layer are measured by monthly FOB (free on board) export values (in U.S. dollars) sourced from the IMF Direction of Trade Statistics (DOTS).\footnote{Data are obtained from the IMF DOTS database (\url{https://data.imf.org/en/datasets/IMF.STA:IMTS}).} We also quantify bilateral political connectedness using ideal-point estimates derived from United Nations General Assembly (UNGA) roll-call votes BSV17, following standard practice in the political economy literature.

Empirical analysis

Consider the $m$-th candidate model specification:

align[align omitted — 256 chars of source]

where $y_{i,t} $ denotes the CPI inflation of economy $ i $ in month $ t $, $ {\sf ER}_{i,t} $ is the period-average nominal exchange rate of economy $ i $ in month $ t $, $ \gamma_{t1}^{(m)} $ is the time-varying trend term, and $ \widetilde{\omega}_{ij}^{(m)} $ is the $ (i,j) $-th element of the matrix constructed from the network layer $ \mathbf{W}^{(m)} $ by setting its diagonal elements to zero and applying row-normalization. We consider $ 4 $ and $ 10 $ network layers for in- and out-of-sample analysis, respectively. Details on network constructions are presented in Appendix (ref).

Figure (ref) plots the estimated time-varying weights of four candidate models over 2007--2024. The weights of candidate models exhibit pronounced time variation. The production network layer is dominant in the early part of the time period, especially around 2009--2012, and then rapidly loses most weight in the later period. The policy network layer becomes prominent in 2013--2019, while the global equity network layer receives substantial weight in the late period (after 2019). In contrast, the trade network layer is assigned only intermittent weight and never attains persistent dominance. Figure (ref) suggests that the transmission mechanism of CPI inflation shifts across different types of network dependence and varies along the time. This finding is consistent with multiple transmission channels for CPI inflation, with no single network maintaining uniformly dominant weights. A single-layer specification would thus be too restrictive when another layer better matches the network dependence structure. It also shows that the candidate weights concentrate on one or two layers at a time and shift as the underlying structure evolves, indicating that the PTVMA adaptively captures locally optimal specifications which may vary over the temporal dimension.

figure[figure omitted — 1,223 chars of source]

Overall, the in-sample results provide evidence that global inflation transmission is governed by a multi-layer, regime-dependent structure in which production, financial, trade, and policy linkages alternately dominate. Therefore, conventional time series models with a single network structure or time-invariant parameters fail to capture the shifting dominance of these transmission channels, underscoring the value of the proposed PTVMA.

Table (ref) reports the out-of-sample prediction performance of the proposed model for CPI inflation. Predictive accuracy is evaluated by the RMSPE, and our model is compared with seven alternative specifications: NAR, time-varying NAR (tv-NAR), VAR(1) estimated by the OLS (VAR), VARX(1) estimated by the Elastic Net (HD-VARX), adaptive LASSO (ALASSO), random forest (RF), and stochastic gradient boosting (SGB), spanning one-, two-, and three-step-ahead forecasts across VAR, penalized regression, and state-of-the-art machine learning methods. Exogenous variables are the nominal exchange rates for all economies. In addition, the ALASSO, the RF and the SGB also involves in- and out-degrees of networks from $ \mathbf{W}^{(1)} $ to $ \mathbf{W}^{(8)} $ and degrees of networks $ \mathbf{W}^{(9)} $ and $ \mathbf{W}^{(10)} $ as controls, see Appendix (ref) for details.

table[table omitted — 1,285 chars of source]

Table (ref) shows a pronounced advantage of the PTVMA in short-run CPI inflation forecasting. At the one-step horizon, our method uniformly outperforms all competitors and attains the lowest RMSPE (with all the ratios strictly smaller than one), indicating a marked improvement in short-run predictive accuracy. At the two-step horizon, the proposed PTVMA performs slightly worse than the NAR and ALASSO but continues to outperform the remaining methods. At the three-step horizon, however, the PTVMA only performs better than the tv-NAR and two VAR-based methods, suggesting that the short-horizon advantage attenuates as the forecast horizon expands. This is because the time-varying coefficient specification amplifies the negative impact of cumulative errors on multi-step-ahead forecasting. It is also worth noting that the proposed PTVMA still dominates the tv-NAR across three horizons, highlighting the advantage of forecast combination via time-varying model averaging.

Conclusion

This paper introduces a dynamic model averaging criterion within a general time-varying network VAR framework which incorporates multiple adjacency matrices to capture network spillover effects. The main aim is to determine a time-varying optimal combination of multi-layer network VAR candidate models whose number may be divergent. In particular, all the candidate models with different linkage structures and varying lagged orders of dependent variables and node-specific predictors are allowed to be misspecified. Both the in-sample and out-of-sample asymptotic properties including the asymptotic optimality and estimation convergence rates are derived under some technical conditions. In particular, the proposed PTVMA achieves the selection consistency (of correctly-specified models) as the conventional model selection criterion. Furthermore, an extended conformal prediction method is proposed to construct prediction intervals for network time series. The Monte Carlo simulation demonstrates that the proposed PTVMA delivers superior out-of-sample predictive performance relative to the benchmark, even when all the candidate models are misspecified. The empirical case study further reveals the time-varying patterns of optimal weight estimation and shows that our model and method consistently outperform other competing models in the short-run CPI inflation prediction.

thebibliography{xx} \harvarditem{Amiti, Redding and Weinstein}{2019}{ARW19} Amiti, M., Redding, S. J. and Weinstein, D. E. (2019). The impact of the 2018 tariffs on prices and welfare. {\em Journal of Economic Perspectives} {\bf 33}, 187--210. \harvarditem{Auer, Levchenko and Saur{\'e}}{2019}{ALS19} Auer, R. A., Levchenko, A. A. and Saur{\'e}, P. (2019). International inflation spillovers through input linkages. {\em The Review of Economics and Statistics} {\bf 101}, 507--521. \harvarditem{Bailey, Strezhnev and Voeten}{2017}{BSV17} Bailey, M. A., Strezhnev, A. and Voeten, E. (2017). {Estimating dynamic state preferences from United Nations voting data}. {\em Journal of Conflict Resolution} {\bf 61}, 430--456. \harvarditem{Barigozzi and Brownlees}{2019}{BB19} Barigozzi, M. and Brownlees, C. (2019). NETS: Network estimation for time series. {\em Journal of Applied Econometrics} {\bf 34}, 347--364. \harvarditem{Barigozzi, Cavaliere and Moramarco}{2025}{BCM25} Barigozzi, M., Cavaliere, G. and Moramarco, G. (2025). Factor network autoregressions. {\em Journal of Business and Economic Statistics} {\bf 43}, 1105--1118. \harvarditem{Basu, Li and Michailidis}{2019}{BLM19} Basu, S., Li, X. and Michailidis, G. (2019). Low rank and structured modeling of high-dimensional vector autoregressions. {\em IEEE Transactions on Signal Processing} {\bf 67}, 1207--1222. \harvarditem{Basu and Michailidis}{2015}{BM15} Basu, S. and Michailidis, G. (2015). Regularized estimation in sparse high-dimensional time series models. {\em The Annals of Statistics} {\bf 43}, 1535--1567. \harvarditem{Cai, Chou and Li}{2009}{CCL09} Cai, Y., Chou, R. Y. and Li, D. (2009). Explaining international stock correlations with CPI fluctuations and market volatility. {\em Journal of Banking & Finance} {\bf 33}, 2026--2035. \harvarditem{Carvalho {\em et al}}{2021}{CNST21} Carvalho, V. M., Nirei, M., Saito, Y. U. and Tahbaz-Salehi, A. (2021). Supply chain disruptions: Evidence from the great east Japan earthquake. {\em The Quarterly Journal of Economics} {\bf 136}, 1255--1321. \harvarditem{Ciccarelli and Mojon}{2010}{CM10} Ciccarelli, M. and Mojon, B. (2010). Global inflation. {\em The Review of Economics and Statistics} {\bf 92}, 524--535. \harvarditem{Chen, Fan and Zhu}{2023}{CFZ23} Chen, E., Fan, J. and Zhu, X. (2023). Community network autoregression for high-dimensional time series. {\em Journal of Econometrics} {\bf 235}, 1239--1256. \harvarditem{Chen {\em et al}}{2025}{CLLL25} Chen, J., Li, D., Li, Y. and Linton, O. (2025). Estimating time-varying networks for high-dimensional time series. {\em Journal of Econometrics} {\bf 249}, 105941. \harvarditem{Chen, Hong and Li}{2024}{CHL24} Chen, Q., Hong, Y. and Li, H. (2024). Time-varying forecast combination for factor-augmented regressions with smooth structural changes. {\em Journal of Econometrics} 240, 105693. \harvarditem{Dahlhaus}{1997}{D97} Dahlhaus, R. (1997). Fitting time series models to nonstationary processes. {\em The Annals of Statistics} {\bf 25}, 1--37. \harvarditem{DasGupta}{2008}{D08} DasGupta, A. (2008). {\em Asymptotic Theory of Statistics and Probability}. Springer. \harvarditem{Davis, Zang and Zheng}{2016}{DZZ16} Davis, R., Zang, P. and Zheng, T. (2016). Sparse vector autoregressive modeling. {\em Journal of Computational and Graphical Statistics} {\bf 25}, 1077--1096. \harvarditem{Demirer {\em et al}}{2018}{DDLY18} Demirer, M., Diebold, F., Liu, L. and Yilmaz, K. (2018). Estimating global bank network connectedness. {\em Journal of Applied Econometrics} {\bf 33}, 1--15. \harvarditem{Diebold and Yilmaz}{2014}{DY14} Diebold, F. and Yilmaz, K. (2014). On the network topology of variance decompositions: Measuring the connectedness of financial firms. {\em Journal of Econometrics} {\bf 182}, 119--134. \harvarditem{Ding, Qiu and Chen}{2017}{DQC17} Ding, X., Qiu, Z. and Chen, X. (2017). Sparse transition matrix estimation for high-dimensional and locally stationary vector autoregressive models. {\em Electronic Journal of Statistics} {\bf 11}, 3871--3902. \harvarditem{Fajgelbaum {\em et al}}{2020}{FGKK20} Fajgelbaum, P. D., Goldberg, P. K., Kennedy, P. J. and Khandelwal, A. K. (2020). The return to protectionism. {\em The Quarterly Journal of Economics} {\bf 135}, 1--55. \harvarditem{Gao, Peng and Yan}{2024}{GPY24} Gao, J., Peng, B. and Yan, Y. (2024). Estimation, inference, and empirical analysis for time-varying VAR models. {\em Journal of Business and Economic Statistics} {\bf 42}, 310--321. \harvarditem{Gao {\em et al}}{2019}{GZWCZ19} Gao, Y., Zhang, X., Wang, S., Chong, T. T. L. and Zou, G. (2019). Frequentist model averaging for threshold models. {\em Annals of the Institute of Statistical Mathematics} {\bf 71}(2), 275--306. \harvarditem{Hansen}{2007}{H07} Hansen, B. E. (2007). Least squares model averaging. {\em Econometrica} 75, 1175--1189. \harvarditem{Hansen}{2014}{H14} Hansen, B. E. (2014). Model averaging, asymptotic risk, and regressor groups. {\em Quantitative Economics} 5, 495--530. \harvarditem{Hansen and Racine}{2012}{HR12} Hansen, B. E. and Racine, J. S. (2012). Jackknife model averaging. {\em Journal of Econometrics} 167, 38--46. \harvarditem{Hjort and Claeskens}{(2003)}{HC03} Hjort, N. L. and Claeskens, G. (2003). Frequentist model average estimators. {\em Journal of the American Statistical Association} 98, 879--899. \harvarditem{Hoeting {\em et al}}{1999}{HMRV99} Hoeting, J. A., Madigan, D., Raftery, A. E. and Volinsky, C. T. (1999). Bayesian model averaging: a tutorial. {\em Statistical Science} 14, 382--417. \harvarditem{Kilian and L\"utkepohl}{2017}{KL17} Kilian, K. and L\"utkepohl, H. (2017). {\em Structural Vector Autoregressive Analysis}. Cambridge University Press. \harvarditem{Kivel\"a {\em et al}}{2014}{KABGMP14} Kivel\"a, M., Arenas, A., Barthelemy, M., Gleeson, J.P., Moreno, Y. and Porter, M.A. (2014). Multilayer networks. {\em Journal of Complex Networks} {\bf 2}, 203--271. \harvarditem{Kock and Callot}{2015}{KC15} Kock, A.B. and Callot, L. (2015). Oracle inequalities for high dimensional vector autoregressions. {\em Journal of Econometrics} {\bf 186}, 325--344. \harvarditem{Kolar {\em et al}}{2010}{KSAX10} Kolar, M., Song, L., Ahmed, A. and Xing, E. (2010). Estimating time-varying networks. {\em The Annals of Applied Statistics} {\bf 4}, 94--123. \harvarditem{Lei {\em et al}}{2018}{LGRTW18} Lei, J., G'Sell, M., Rinaldo, A. and Tibshirani, R.J. and Wasserman, L. (2018). Distribution-free prediction inference for regression. {\em Journal of the American Statistical Association} {\bf 113}, 1094-1111. \harvarditem{Lei and Wasserman}{2014}{LW14} Lei, J. and Wasserman, L. (2014). Distribution-free prediction bands for non-parametric regression. {\em Journal of the Royal Statistical Society, Series B} {\bf 76}, 71--96. \harvarditem{Li {\em et al}}{2026}{LPTW26} Li, D., Peng, B., Tang, S. and Wu, W. (2026). Estimation of grouped time-varying network vector autoregressive models. {\em The Annals of Statistics} {\bf 54}, 621--646. \harvarditem{Liao {\em et al}}{2019}{LZZZ19} Liao, J., Zong, X., Zhang, X. and Zou, G. (2019). Model averaging based on leave-subject-out cross-validation for vector autoregressions. {\em Journal of Econometrics} \textbf{209}, 35--60. \harvarditem{Liao and Tsay}{2020}{LT20} Liao, J.-C. and Tsay, W.-J. (2020). Optimal multistep var forecast averaging. {\em Econometric Theory} \textbf{36}, 1099--1126. \harvarditem{Liu}{2015}{Liu15} Liu, C.-A. (2015). Distribution theory of the least squares averaging estimator. {\em Journal of Econometrics} \textbf{186}, 142--159. \harvarditem{L\"utkepohl}{2006}{Lu06} L\"utkepohl, H. (2006). {\em New Introduction to Multiple Time Series Analysis}. Springer. \harvarditem{Miao, Phillips and Su}{2023}{MPS23} Miao, K., Phillips, P.C.B. and Su, L. (2023). High-dimensional VARs with common factors. {\em Journal of Econometrics} {\bf 233}, 155--183. \harvarditem{Miranda-Agrippino and Rey}{2020}{MR20} Miranda-Agrippino, S. and Rey, H. (2020). US monetary policy and the global financial cycle. {\em The Review of Economic Studies} {\bf 87}, 2754--2776. \harvarditem{Negahban and Wainwright}{2011}{NW11} Negahban, S. and Wainwright, M. J. (2011). Estimation of (near) low-rank matrices with noise and high-dimensional scaling. {\em The Annals of Statistics} {\bf 39}, 1069--1097. \harvarditem{Pesaran, Schuermann and Weiner}{2004}{PSW04} Pesaran, M.H., Schuermann, T. and Weiner, S.M. (2004). Modeling regional interdependencies using a global error correcting macroeconometric model. {\em Journal of Business and Economic Statistics} {\bf 22}, 129--162. \harvarditem{Raftery and Zheng}{2003}{RZ03} Raftery, A. E. and Zheng, Y. (2003). Discussion: Performance of Bayesian model averaging. {\em Journal of the American Statistical Association} {\bf 98}, 931--938. \harvarditem{Scott}{2017}{S17} Scott, J. (2017). {\em Social Network Analysis} (4th Edition). Sage, London. \harvarditem{Sims}{1980}{Si80} Sims, C. A. (1980). Macroeconomics and reality. {\em Econometrica} {\bf 48}, 1--48. \harvarditem{Sun, Chen and Gao}{2025}{SCG25} Sun, Y., Chen, F. and Gao, J. (2025). Model averaging for time-varying vector autoregressions. Working paper, Department of Econometrics and Business Statistics, Monash University. \harvarditem{Sun {\em et al}}{2021}{SHLWZ21} Sun, Y., Hong, Y., Lee, T.-H., Wang, S. and Zhang, X. (2021). Time-varying model averaging. {\em Journal of Econometrics} \textbf{222}, 974--992. \harvarditem{Sun {\em et al}}{2023}{SHWZ23} Sun, Y., Hong, Y., Wang, S. and Zhang, X. (2023). Penalized time-varying model averaging. {\em Journal of Econometrics} \textbf{235}, 1355--1377. \harvarditem{Stock and Watson}{2007}{SW07} Stock, J. H. and Watson, M. W. (2007). Why has US inflation become harder to forecast? {\em Journal of Money, Credit and Banking} {\bf 39}, 3--33. \harvarditem{Tu and Wang}{2025}{TW25} Tu, Y. and Wang, S. (2025). Quantile prediction with factor-augmented regression: Structural instability and model uncertainty. {\em Journal of Econometrics} \textbf{249}, 105999. \harvarditem{Vovk, Gammerman and Shafer}{2005}{VGS05} Vovk, V., Gammerman, A. and Shafer, G. (2005). {\em Algorithmic Learning in a Random World}. Springer, NewYork. \harvarditem{Xu, Chen and Wu}{2020}{XCW20} Xu, M., Chen, X. and Wu, W. (2020). Estimation of dynamic networks for high-dimensional nonstationary time series. {\em Entropy} {\bf 22}, 55. \harvarditem{Yin, Safikhani and Michailidis}{2023}{YSM23} Yin, H., Safikhani, A. and Michailidis, G. (2023). A general modeling framework for network autoregressive processes. {\em Technometrics} {\bf 65}, 579--589. \harvarditem{Yin, Safikhani and Michailidis}{2026}{YSM26} Yin, H., Safikhani, A. and Michailidis, G. (2026). A functional coefficients network autoregressive model. {\em Statistica Sinica} {\bf 36}, 1-25. \harvarditem{Yu {\em et al}}{2024}{YLSZH24} Yu, D., Lian, H., Sun, Y., Zhang, X. and Hong, Y. (2024). Post-averaging inference for optimal model averaging estimator in generalized linear models. {\em Econometric Reviews} \textbf{43}, 98--122. \harvarditem{Yu, Zhang and Liang}{2025}{YZL25} Yu, D., Zhang, X. and Liang, H. (2025). Unified optimal model averaging with a general loss function based on cross-validation. {\em Journal of the American Statistical Association} {\bf 120}, 2697--2708. \harvarditem{Zhang and Liu}{2019}{ZL18} Zhang, X. and Liu, C. A. (2019). Inference after model averaging in linear regression models. {\em Econometric Theory} \textbf{35}, 816--841. \harvarditem{Zhang and Liu}{2023}{ZL23} Zhang, X. and Liu, C. A. (2023). Model averaging prediction by K-fold cross-validation. {\em Journal of Econometrics} \textbf{235}, 280--301. \harvarditem{Zhang and Zhang}{2023}{ZZ23} Zhang, X. and Zhang, X. (2023). Optimal model averaging based on forward-validation. {\em Journal of Econometrics} \textbf{237}, 105295. \harvarditem{Zhang and Yu}{2018}{ZY18} Zhang, X. and Yu, J. (2018). Spatial weights matrix selection and model averaging for spatial autoregressive models. {\em Journal of Econometrics}, {\bf 203}, 1--18. \harvarditem{Zhou, Lafferty and Wasserman}{2010}{ZLW10} Zhou, S., Lafferty, J. and Wasserman, L. (2010). Time varying undirected graphs. {\em Machine Learning} {\bf 80}, 295--319. \harvarditem{Zhu {\em et al}}{2017}{ZPLLW17} Zhu, X., Pan, R., Li, G., Liu, Y. and Wang, H. (2017). Network vector autoregression. {\em The Annals of Statistics} {\bf 45}, 1096--1123.