Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
108,488 characters · 16 sections · 52 citation commands
Time-Varying Model Averaging of Multi-layer Network Vector Autoregressions
\newtheorem{corollary}{Corollary} \newtheorem{definition}{Definition} \newtheorem{lemma}{Lemma} \newtheorem{proposition}{Proposition} \newtheorem{remark}{Remark} \newtheorem{theorem}{Theorem} \newtheorem{example}{Example} \newtheorem{assumption}{Assumption} \newtheorem{prop}{Proposition}
\numberwithin{corollary}{section} \numberwithin{definition}{section} \numberwithin{equation}{section} \numberwithin{lemma}{section} \numberwithin{proposition}{section} \numberwithin{remark}{section} \numberwithin{theorem}{section}
\onehalfspacing
\newtheorem{Proof}{Proof} \newtheorem{Lemma}{Lemma} \newtheorem{Mth}{Main Theorem} \newtheorem{Res}{\bf Result} \newtheorem{Def}{Definition} \newtheorem{Rem}{\bf Remark} \newtheorem{Qes}{Question} \newtheorem{Aim}{Aim} \newtheorem{Pro}{Proposition} \newtheorem{Lem}{\bf Lemma} \newtheorem{Cor}{\bf Corollary} \newtheorem{Ex}{Example} \newtheorem{Eq}{Equation} \newtheorem{condition}{Assumption} \newtheorem{conditionnew}{Condition}
\def\0{{\bf 0}}
\def\squarebox#1{\hbox to #1{\vbox to #1}}
\allowdisplaybreaks[4]
\thispagestyle{empty}
\setcounter{page}{1}
\setcounter{equation}{0}
The vector autoregression (VAR) has become a standard econometric tool to model multivariate time series since the seminal work by Si80. In the classic setting with a fixed number of time series variables, Lu06 and KL17 provide a comprehensive review of various VAR-based model estimation, inference and forecasting techniques. However, the conventional OLS or MLE estimation and forecasting methods often perform poorly when the number of variables diverges with the time series length. To address this issue in the high-dimensional time series setup, a commonly-used approach is to impose some structural restriction on transition matrices such as the sparsity or reduced-rank assumption and subsequently adopt regularized estimation methods NW11, KC15, DZZ16, BLM19, MPS23. However, these models and methods often neglect potential inter-connection or network structure among a large number of variables, a feature which commonly exists for large-scale dynamic macroeconomic or financial systems. The network VAR introduced by ZPLLW17 whose model formulation is virtually similar to the so-called global VAR PSW04, provides a flexible autoregressive framework for large-scale time series to jointly incorporate the network spillover (cross-lag) effect, momentum (own-lag) effect and nodal effect.
Network models are an effective tool to tackle interconnected systems with applications to various fields such as economics, finance and social networks. There have been extensive studies on estimation and inference of both the static network DY14, S17, DDLY18, BB19 and time-varying dynamic network KSAX10, ZLW10, CLLL25. The existing literature on network VAR models typically assumes that the network structure is available and incorporates an observed adjacency matrix in VAR to capture the network spillover effect ZPLLW17, CFZ23, YSM23, YSM26, LPTW26. They often limit attention to the assumption of a single adjacency matrix to measure linkages between nodes. This assumption may be restrictive for modern social or economic networks as agents in dynamic systems often interact through multiple channels, resulting in the so-called multi-layer network KABGMP14. BCM25 introduce a factor network VAR model and propose a tensor decomposition of multiple adjacency matrices which may vary over time to achieve dimension reduction in the subsequent estimation procedure. In this paper, we propose an alternative approach for multi-layer network VAR, adopting the model averaging method to construct an optimal combination of candidate time-varying multi-layer network VAR models with different linkage structures (to capture network spillover effects), varying lagged orders of dependent variables and node-specific predictors.
Model averaging tackles model uncertainty by averaging over a number of candidate predictive models and serves as an alternative to model selection which aims to select one model according to a proper criterion. Model averaging methods can be broadly classified into two categories: Bayesian model averaging with prior information HMRV99, and frequentist model averaging, which focuses on determining optimal weights for candidate models. Commonly-used frequentist model averaging approaches include the Mallows model averaging H07,LT20 and jackknife model averaging HR12, LZZZ19, YZL25. These classic methods aim to select time-invariant combination weights. Their numerical performance may be unstable for dynamic environments with structural changes. SHLWZ21 is among the first to propose time-varying jackknife model averaging to accommodate smooth structural changes, with subsequent extensions to high-dimensional penalized time-varying model averaging SHWZ23, factor-augmented regression CHL24 and quantile regression TW25, and time-varying VAR SCG25. The aforementioned works, however, do not account for interconnections between time series variables and potential multi-layer network structures. Our paper aims to fill in this gap by studying time-varying model averaging of multi-layer network VAR for large-scale time series with smooth structural changes.
Meanwhile, there has been increasing attention in recent years on time-varying VAR models, accommodating smooth structural changes over a long time span. GPY24 study estimation and inference of time-varying VAR models including the impulse response analysis for finite-dimensional time series. The time-varying VAR is naturally connected to the locally stationary time series model framework, which has been systematically studied in the literature since D97. In the high-dimensional time-varying VAR model setup, DQC17 combine the $\ell_1$-regularised method with kernel smoothing to estimate transition matrices; XCW20 study the multiple break structure and estimate smooth time-varying covariance and precision matrices between the break points; and CLLL25 estimate dual network structures via directed Granger causality and undirected partial correlation linkages. A recent paper by LPTW26 further introduces a general time-varying network VAR model framework with a latent group structure, allowing both structural changes and structural breaks on model coefficients. However, they limit attention to a single-layer network structure to capture the network spillover effect.
In this paper, we propose a penalized time-varying model averaging (PTVMA) method to estimate the optimal weights for candidate multi-layer network VAR models whose time-varying coefficients are estimated via the classic local linear smoothing. The candidate models are constructed via distinct adjacency matrices in network spillover effects, varying lags in momentum effects and different subsets of node-specific exogenous predictors. In particular, we allow all the candidate models to be misspecified, enhancing the applicability of the proposed model and methodology. Under some regularity conditions, we establish the asymptotic optimality and convergence rates of the proposed time-varying weight estimates in the contexts of both the in-sample fitting and out-of-sample prediction. In particular, we show that the PTVMA criterion achieves the so-called “selection consistency" when there exist correctly-specified candidate models, i.e., the sum of the estimated time-varying weights for correctly-specified candidate models approaches one asymptotically.
In addition to point forecasts, we further construct the prediction interval, accounting for forecast uncertainty. Specifically, we propose a PTVMA-based conformal prediction method to construct interval forecasts without requiring the number of candidate models or predictors to be fixed. The conformal prediction method is first introduced by VGS05 as a sequential approach to form prediction intervals and has been extended by LW14 and LGRTW18 to accommodate the nonparametric regression framework. In this paper, we make a further extension by combining the conformal prediction with the developed PTVMA method to construct prediction bands in the context of large-scale and locally stationary network time series. The proposed PTVMA-based conformal prediction algorithm is easy to implement and can guarantee the coverage probability without requiring the asymptotic distribution theory of the model averaging estimation.
Monte Carlo simulations demonstrate that the proposed PTVMA method delivers nearly identical in-sample fit and superior out-of-sample prediction relative to the benchmark, even when all candidate models are misspecified. By assigning time-varying weights to candidate models to accommodate parameter instability, the PTVMA criterion achieves selection consistency when correctly specified candidates exist, and concentrates weight on the best-approximating models when all candidates are misspecified. The proposed conformal prediction intervals exhibit robust finite-sample performance, with empirical coverage probabilities close to the nominal level, particularly as the sample size increases. The developed model and methodology are further applied to analyze the monthly CPI inflation in 36 economies from January 2007 to December 2024. The empirical results reveal multiple channels of CPI inflation transmission across economies, with the production, equity, trade, and policy networks alternately dominating global inflation transmission. Our model and method consistently outperform a broad set of competing models, including the static and time-varying network VAR models and state-of-the-art machine learning methods, highlighting the advantage of the proposed time-varying multi-layer network VAR and PTVMA.
The rest of the paper is organized as follows. Section (ref) introduces the model framework and the PTVMA methodology. Section (ref) gives the asymptotic properties of the developed weight estimation together with technical assumptions. Section (ref) proposes the PTVMA-based conformal prediction interval construction. Sections (ref) and (ref) report the simulation and empirical studies, respectively. Section (ref) concludes the paper. All the theoretical proofs are contained in the appendix.
\setcounter{equation}{0}
In this section, we introduce the time-varying multi-layer network VAR model framework and propose the PTVMA method to determine a time-varying optimal combination of candidate models.
Consider ${\sf N}$-dimensional vectors of time series observations ${\bf Y}_t=(y_{1,t},y_{2,t},\ldots,y_{{\sf N}, t})^{{^\intercal}}$, $1\leq t\leq {\sf T}$, with ${\sf N}$ being the number of nodes in the large-scale network, and ${\sf Q}$-dimensional node-specific vectors ${\bf X}_i=(x_{i,1},\ldots,x_{i,{\sf Q}})^{{^\intercal}}$, $1\leq i\leq {\sf N}$. Throughout the paper, ${\sf N}$ is allowed to diverge to infinity while ${\sf Q}$ is a fixed positive integer. Let ${\mathbf W}^{(k)}=(\omega_{ij}^{(k)})_{{\sf N}\times {\sf N}}$ denote the $k$-th adjacency matrix, which is observed and non-stochastic, $\omega_{ii}^{(k)}=0$ and $\omega_{ij}^{(k)}=1$ if $i$ follows $j$, $k=1,\ldots,{\sf K}$. These adjacency matrices capture economic linkages, geographic distances, or other forms of connectedness. All the adjacency matrices in this paper can be either directed or undirected. We consider the following data generating process:
where $\widetilde \omega_{ij}^{(k)}=\omega_{ij}^{(k)}/n_i^{(k)}$ with $n_i^{(k)}=\sum_{j\neq i} \omega_{ij}^{(k)}$ denoting the total number of nodes that $i$ follows according the $k$-th linkage, $\boldsymbol{\beta}_t=(\beta_{t1},\ldots,\beta_{t{\sf K}})^{{^\intercal}}$, $\boldsymbol {\alpha}_t= (\alpha_{t1}, \ldots,\alpha_{t{\sf P}} )^{{^\intercal}}$ and ${\boldsymbol\gamma}_t=(\gamma_{t1},\ldots,\gamma_{t{\sf Q}})^{{^\intercal}}$ are unknown time-varying parameter vectors to be estimated, ${\bf Z}_{i,t}$ denotes the predictor vector with dimension ${\sf K}+{\sf P}+{\sf Q}$, i.e.,
${\boldsymbol\theta}_t=(\boldsymbol{\beta}_t^{^\intercal},\boldsymbol {\alpha}_t^{^\intercal}, {\boldsymbol\gamma}_t^{^\intercal})^{^\intercal}$, and $\epsilon_{i,t}=\sigma_i(t/T)u_{i,t}$ with $\sigma_i^2(\cdot)$ being a time-varying variance function and $u_{i,t}$ being independent and identically distributed (i.i.d.) with zero mean and unit variance.
The time-varying VAR in ((ref)) is split into multiple network spillover effects (or cross-lag effects), the momentum effects with order ${\sf P}$ (or own-lag effects) and the nodal effects, all of which are allowed to evolve smoothly over time. Here we only include the first lag of network spillover effects for simplicity and the extension to higher-order cross-lag effects is straightforward.
Unlike the model selection which aims to select one time-varying network VAR model, to reduce model estimation and prediction uncertainty caused by selection of adjacency matrices, order of lagged dependent variables and subset of node-specific exogenous predictors, we next propose the PTVMA procedure, seeking to construct a dynamic optimal combination of candidate time-varying network VAR models. Consider a sequence of candidate network VAR models indexed by $m=1,\ldots, {\sf M}$, which may be misspecified for the underlying data generating process. Specifically, we formulate the $m$-th candidate model as
where ${\cal I}_{{\sf K}}^{(m)}=\{i_1,\ldots,i_{k_m}\}$ is a subset of $\{1,2,\ldots,{\sf K}\}$, ${\cal I}_{{\sf P}}^{(m)}=\{i_1,\ldots,i_{p_m}\}$ is a subset of $\{1,2,\ldots,{\sf P}\}$, ${\cal I}_{{\sf Q}}^{(m)}=\{i_1,\ldots,i_{q_m}\}$ is a subset of $\{1,2,\ldots,{\sf Q}\}$, $${\boldsymbol\theta}_t^{(m)}=\left(\boldsymbol{\beta}_t^{(m){^\intercal}},\boldsymbol {\alpha}_t^{(m){^\intercal}},{\boldsymbol\gamma}_t^{(m){^\intercal}}\right)^{{^\intercal}}$$ with
and ${\bf Z}_{i,t}^{(m)}$ denotes the predictor vector with dimension $k_m+p_m+q_m$. All the time-varying parameters are defined as smooth functions of scaled times. Let $$\bar{\rho}= ({\sf K}+{\sf P}+{\sf Q})\vee \max \{ \rho_1, \rho_2, \cdots, \rho_{\sf M}\} \ \ \ \text{with}\ \ \rho_m=k_m+p_m+q_m,$$ ${\boldsymbol\mu}_t^{(m)}=(\mu_{1,t}^{(m)},\ldots,\mu_{{\sf N}, t}^{(m)})^{^\intercal}$ and ${\boldsymbol\epsilon}_t^{(m)}=(\epsilon_{1,t}^{(m)},\ldots,\epsilon_{{\sf N}, t}^{(m)})^{^\intercal}$. Throughout the paper, we allow candidate time-varying network VAR models to be non-nested and misspecified, and both $\bar{\rho}$ and ${\sf M}$ may diverge to infinity.
For any given $t_0$, we define the local linear estimate of ${{\boldsymbol\theta}}_{t_0}^{(m)}$ as
where ${\mathbf I}_{\rho_m}$ and ${\mathbf O}_{\rho_m}$ are the identity and zero matrices, respectively, with size $\rho_m \times \rho_m$, \[ \widetilde{\mathbf{Z}}_{t-1}^{(m)}=\left[{\mathbf Z}_{t-1}^{(m){^\intercal}},\ \left(\frac{t-t_0}{Th}\right){\mathbf Z}_{t-1}^{(m){^\intercal}}\right],\quad {\bf Z}_{t}^{(m)}=({\bf Z}_{1,t}^{(m)},\ldots,{\bf Z}_{{\sf N},t}^{(m)}),\quad K_{t,t_0}=K\left(\frac{t-t_0}{Th}\right) \] with $K(\cdot)$ being a kernel function and $h$ being a bandwidth.
Let \[ {\mathscr W}=\left\{{\bf w}=(w^1,w^2,\ldots,w^{\sf M})^{{^\intercal}}\in[0,1]^{\sf M}:\ \sum_{m=1}^{\sf M} w^m=1\right\}. \] With a given vector ${\bf w}=(w^1,w^2,\ldots,w^{\sf M})^{{^\intercal}}\in{\mathscr W}$, we construct the model averaging estimators of time-varying parameters and conditional mean as follows
where ${\boldsymbol \Pi}^{(m)}$ is a matrix of size $\rho_m\times({\sf K}+{\sf P}+{\sf Q})$ mapping ${\bf Z}_{i,t}$ to ${\bf Z}_{i,t}^{(m)}$ and \[ \widehat{\mu}_{i,t}({\bf w})=\sum_{m=1}^{\sf M} w^m{\bf Z}_{i,t-1}^{(m)^{{^\intercal}}}\widehat{{\boldsymbol\theta}}_t^{(m)}. \] We propose the PTVMA criterion:
where $\lambda$ is a user-specified tuning parameter. The optimal weights are determined by
Finally the one-step ahead forecast combination is constructed by
\setcounter{equation}{0}
In this section, we give some regularity conditions and present the asymptotic properties of the PTVMA estimation developed in Section (ref) including the asymptotic optimality and convergence rates. Section (ref) considers the in-sample fitting, i.e., $\widehat{\bf w}_t$ for the interior point $t$ in the sense that $0<t/{\sf T}<1$, whereas Section (ref) tackles the out-of-sample prediction, i.e., $\widehat{\bf w}_{\sf T}$ for any forecast time point ${\sf T}+1$.
To establish the asymptotic optimality of the proposed PTVMA method for in-sample fitting, we define the following infeasible locally weighted (in-sample) quadratic error loss:
where ${\boldsymbol\mu}_s=(\mu_{1,s},\ldots,\mu_{{\sf N}, s})^{^\intercal}$, $ \mu_{i,s} = {\bf Z}_{i,s-1}^{{^\intercal}} {\boldsymbol\theta}_{s} $, $t$ is a fixed interior point such that $0<t/{\sf T}<1$, in which case the local kernel smoothing utilizes the time series sample information on both sides of $t$. Denote $${\boldsymbol\mu}_t^*({\bf w})=\sum_{m=1}^{\sf M} w^m{\boldsymbol\mu}_{t,\ast}^{(m)}=\left[\mu_{1,t}^{*}({\bf w}),\ldots,\mu_{{\sf N},t}^{*}({\bf w})\right]^{^\intercal},$$ where ${\boldsymbol\mu}_{t,\ast}^{(m)}=(\mu_{1,t}^{(m)*},\ldots,\mu_{{\sf N},t}^{(m)*})^{^\intercal}$ and $\mu_{i,t}^{(m)*}={\bf Z}_{i,t-1}^{(m){^\intercal}}{\boldsymbol\theta}_{t,\ast}^{(m)}$ with ${\boldsymbol\theta}_{t,\ast}^{(m)}$ being the pseudo time-varying coefficient vector for the $m$-th candidate model to be defined in Assumption (ref) below. Write
The following assumptions are required to derive the in-sample asymptotic optimality.
{\bf Remark 3.1}.\ \ Assumption (ref)(i) imposes a high-level condition on uniform convergence of local linear smoothing for misspecified candidate models, which is comparable to those adopted by some recent works on time-varying model averaging SHWZ23, SCG25. With standard techniques for locally stationary time series and nonparametric kernel-weighted smoothing LPTW26, we may derive an explicit rate: $\zeta=\sqrt{\hbox{log} {\sf T}/({\sf N}{\sf T} h)}+h^2$ under some sufficient conditions which require higher moment restriction than Assumption (ref)(ii). With this explicit rate for $\zeta$, assuming that both $\bar\rho$ and ${\sf M}$ are bounded, ${\sf N}{\sf T} h^5=O(1)$ and $\lambda=O(({\sf N} {\sf T} h\cdot\hbox{log} {\sf T})^{1/2})$, we may verify Assumption (ref)(ii) if \[ ({\sf N}{\sf T} h\cdot \hbox{log} {\sf T})^{1/2}/\xi_t=o_P(1). \] The latter indicating that {\em all the candidate models must be misspecified}, otherwise $\xi_t\equiv0$.
Theorem (ref) below establishes the point-wise asymptotic optimality of the PTVMA estimator, similar to Theorem 1 in SHWZ23 and Theorem 3.3 in CHL24.
The following extra condition is required to derive an explicit in-sample convergence rate of the PTVMA weight estimation.
{\bf Remark 3.2.}\ \ Assumption (ref)(i) is comparable to some similar restriction on predictor vectors or the so-called “hat matrix". Assumption (ref)(ii) implies weak temporal dependence as well as cross-sectional correlation for predictors as in LPTW26. It is easy to verify Assumption (ref)(iii) when $\zeta=\sqrt{\hbox{log} {\sf T}/({\sf N}{\sf T} h)}+h^2$, $\lambda=O(({\sf N}{\sf T} h \hbox{log}{\sf T})^{1/2})$ and both $\bar\rho$ and ${\sf M}$ are bounded.
{\bf Remark 3.3.}\ \ Theorem (ref) above reveals that the point-wise convergence rate of the in-sample time-varying weight estimation $\widehat{{\bf w}}_t$ relies on the number of candidate models ${\sf M}$, the maximum dimension of potential predictors $\bar\rho$, the tuning parameter in penalization $\lambda$, the bandwidth $h$ as well as $({\sf N}, {\sf T})$. With the sufficient conditions discussed in Remark 3.2 to verify Assumption (ref)(iii), we may simplify the rate in ((ref)) to \[ \left\|\widehat{\mathbf w}_t-{\mathbf w}_t^*\right\|=O_P\left(h+(\hbox{log} {\sf T})^{1/4}({\sf N}{\sf T} h)^{-1/4}\right). \]
We next consider the scenario when there exist correctly-specified candidate time-varying multi-layer network VAR models. Let ${\mathscr D}$ denote the index set of correctly-specified models, i.e., $m\in{\mathscr D}$ if the $m$-th candidate model is correctly specified. Similar to $\xi_t$ in ((ref)), we define $$\widetilde\xi_{t}=\inf_{{\bf w}\in{\mathscr W}({\mathscr D})} L_{t}^*({\bf w}),\quad {\mathscr W}({\mathscr D})=\{{\bf w}=(w^1,w^2,\ldots,w^{\sf M})^{{^\intercal}}\in{\mathscr W}:\ w^m=0, \ m\in{\mathscr D}\}.$$ Theorem (ref) below shows that the proposed PTVMA criterion achieves the in-sample selection consistency (of correctly-specified candidate models) as the conventional model selection.
We next consider the asymptotic properties of $\widehat{\bf w}_{\sf T}$ which has been adopted for constructing out-of-sample prediction combination. In this case, the local linear smoothing and ${\sf PTVMA}_{T}({\bf w})$ essentially use one-sided local sample information, which would subsequently slow down the weight estimation convergence. Define the out-of-sample prediction risk: \[ R_{{\sf T}+1}({\bf w})={\sf E}\left\{\left[{\bf Y}_{{\sf T}+1}-\widehat{{\bf Y}}_{{\sf T}+1}({\bf w})\right]^{^\intercal}\left[{\bf Y}_{{\sf T}+1}-\widehat{{\bf Y}}_{{\sf T}+1}({\bf w})\right]\right\}-\frac{1}{{\sf T} h}\sum_{i=1}^{\sf N}\sum_{s=1}^{\sf T}\sigma_{i}^{2}(s/{\sf T})K_{s,{\sf T}}, \] where $\widehat{{\bf Y}}_{{\sf T}+1}({\bf w})$ is defined similarly to $\widehat{{\bf Y}}_{{\sf T}+1}(\widehat{\bf w}_{\sf T})$ defined in ((ref)) but with $\widehat{\bf w}_{\sf T}$ replaced by ${\bf w}$. Note that the second term of $R_{{\sf T}+1}({\bf w})$ is independent of ${\bf w}$. In the out-of-sample prediction, we expect the PTVMA estimate $\widehat{\bf w}_{\sf T}$ to asymptotically minimize $R_{{\sf T}+1}({\bf w})$ over ${\bf w}\in{{\mathscr W}}$. Furthermore, we let \[ R_{{\sf T}+1}^*({\bf w})={\sf E}\left\{\left[{\bf Y}_{{\sf T}+1}-{\bf Y}_{{\sf T}+1}^*({\bf w})\right]^{^\intercal}\left[{\bf Y}_{{\sf T}+1}-{\bf Y}_{{\sf T}+1}^*({\bf w})\right]\right\}-\frac{1}{{\sf T} h}\sum_{i=1}^{\sf N}\sum_{s=1}^{\sf T}\sigma_{i}^{2}(s/{\sf T})K_{s,{\sf T}}, \] and \[ \xi_{{\sf T}+1}^*=\inf_{{\bf w}\in {\mathscr W}}R_{{\sf T}+1}^*({\bf w}), \] where ${\bf Y}_{{\sf T}+1}^\ast({\bf w})=\sum_{m=1}^{\sf M} w^m{\bf Y}_{{\sf T}+1}^{(m)\ast}$ and ${\bf Y}_{{\sf T}+1}^{(m)\ast}=(y_{1,{\sf T}+1}^{(m)*},\ldots,y_{{\sf N},{\sf T}+1}^{(m)*})^{^\intercal}$ with $y_{i,{\sf T}+1}^{(m)*}={\bf Z}_{i,{\sf T}}^{(m){^\intercal}}{\boldsymbol\theta}_{{\sf T},\ast}^{(m)}$. Let ${\boldsymbol\epsilon}_{t}^{(m)*}=(\epsilon_{1,t}^{(m)*},\ldots,\epsilon_{{\sf N},t}^{(m)*})^{^\intercal}$ with $\epsilon_{i,t}^{(m)*}=y_{i,t}-\mu_{i,t}^{(m)*}$, $1\leq t\leq {\sf T}$, and ${\boldsymbol\epsilon}_{{\sf T}+1}^{(m)*}=(\epsilon_{1,{\sf T}+1}^{(m)*},\ldots,\epsilon_{{\sf N}, {\sf T}+1}^{(m)*})^{^\intercal}$ with $\epsilon_{i, {\sf T}+1}^{(m)*}=y_{i, {\sf T}+1}-y_{i, {\sf T}+1}^{(m)*}$ for the out-of-sample period.
{\bf Remark 3.4.}\ \ Assumption (ref)(i) imposes a mild smoothness restriction on heterogeneous variance functions $\sigma_i^2(\cdot)$ and is automatically satisfied for the case of homoskedasticity. The condition ((ref)) requires the potential best prediction (over all the candidate models) has finite risk in large-scale networks. A similar condition can be found in Assumptions 3(iii) and 5(iii) of ZZ23. In fact, it provides a sufficient condition to derive the uniform integrability (over ${\bf w}$) of $\xi_{{\sf T}+1}^{*-1}\left|\|{\bf Y}_{{\sf T}+1}-\widehat{{\bf Y}}_{{\sf T}+1}({\bf w})\|^2-\|{\bf Y}_{{\sf T}+1}-{\bf Y}^*_{{\sf T}+1}({\bf w})\|^2\right|$, which is required in the proof of Theorem (ref) in the appendix. If, in addition, $\sup_{\sf T}\xi_{{\sf T}+1}^*/{\sf N}$ is strictly larger than a positive constant, we may replace $\xi_{{\sf T}+1}^\ast$ by ${\sf N}$ in ((ref)) and further show that all the candidate models are misspecified. In fact, if the $m_0$-th candidate model is correctly specified, ${\boldsymbol \Pi}^{(m_0){^\intercal}}\widehat{\boldsymbol\theta}_{\sf T}^{(m_0)}$ converges to the true coefficient vector ${\boldsymbol\theta}_{\sf T}$. Consequently, ${\bf Y}_{{\sf T}+1}^{(m_0)*}={\boldsymbol\mu}_{{\sf T}+1}(1+o_P(h))$ and \[ \xi_{{\sf T}+1}^{*}=\inf_{{\bf w}\in{\mathscr W}}R_{{\sf T}+1}^*({\bf w})\leq {\sf E}\left[\left({\bf Y}_{{\sf T}+1}-{\bf Y}_{{\sf T}+1}^{(m_0)*}\right)^{^\intercal}\left({\bf Y}_{{\sf T}+1}-{\bf Y}_{{\sf T}+1}^{(m_0)*}\right)\right]-\frac{1}{{\sf T} h}\sum_{i=1}^{\sf N}\sum_{s=1}^{\sf T}\sigma_{i}^{2}(s/{\sf T})K_{s,{\sf T}}=o({\sf N}), \] which leads to contradiction. Assumption (ref)(iii) is comparable to Assumption (ref)(ii) and can be easily justified by using the sufficient conditions discussed in Remarks 3.1 and 3.2. Assumption (ref)(iv) contains a high-level condition for locally stationary ${\bf Z}_{i,t}$.
{\bf Remark 3.5.}\ \ Theorem (ref) above shows that the PTVMA estimator is out-of-sample asymptotically optimal in the sense that its out-of-sample prediction risk is asymptotically equivalent to that of the infeasible optimal weight estimator ${\bf w}_{\sf T}^\ast:=\inf_{{\bf w}\in {\mathscr W}}R_{{\sf T}+1}({\bf w})$. A similar property is also derived by ZL23 and ZZ23. A combination of Theorems (ref) and (ref) demonstrates that our proposed PTVMA estimator achieves not only the minimum of in-sample estimation error at the interior time point but also the minimum of out-of-sample prediction risk.
To derive an explicit out-of-sample convergence rate of $\widehat{{\bf w}}_{\sf T}$, we require the following assumption which is analogous to Assumption (ref)(i)(iii) in Section (ref).
{\bf Remark 3.6.}\ \ As in Theorem (ref), the convergence rate of $\widehat{{\bf w}}_{\sf T}$ relies on ${\sf M}$, $\bar\rho$, $\lambda$, $h$ and $({\sf N}, {\sf T})$. In fact, the rate in ((ref)) is slightly slower than that in ((ref)) due to the extra term $h^{1/2}$, which is caused by the well-known boundary effect of kernel smoothing. Furthermore, as discussed in Remark 3.3, under some sufficient conditions, we may simplify the rate in ((ref)) to \[ \left\|\widehat{\mathbf w}_{\sf T}-{\mathbf w}_{\sf T}^*\right\|=O_P\left(h^{1/2}+(\hbox{log} {\sf T})^{1/4}({\sf N} {\sf T} h)^{-1/4}\right). \]
We finally consider the case when there exist correctly-specified candidate models. Let ${\mathscr D}$ and ${\mathscr W}({\mathscr D})$ be defined as in Section (ref) and $$\widetilde\xi_{{\sf T}+1}^{*}=\inf_{{\bf w}\in{\mathscr W}({\mathscr D})}R_{{\sf T}+1}^*({\bf w}),$$ denoting the minimum out-of-sample prediction risk over misspecified models. The following theorem is an extension of Theorem (ref) from in-sample fitting to out-of-sample prediction.
\setcounter{equation}{0}
Despite the growing literature on frequentist model averaging, construction of prediction intervals using model averaging estimators remain under-explored. A commonly-used approach to constructing interval forecasts often relies on the asymptotic distribution of the model averaging estimator. The relevant literature can be divided into two strands. The first examines the limiting distributions of least squares averaging estimators within a local asymptotic framework, where time-invariant regression coefficients lie in a local ${\sf T}^{-1/2}$ neighborhood of zero HC03, H14,Liu15. However, the realism of this local misspecification assumption has been questioned by RZ03, who argue that it lacks face validity in certain empirical settings. The second strand relies on the asymptotic distribution of model averaging estimators under proper model assumptions such as the (non-)linear regression with fixed parameters ZL18,YLSZH24 and time-varying coefficient regressions SHWZ23,TW25. These results are restricted to the setting where both the number of candidate models and the number of predictors remain fixed. We next propose the PTVMA-based conformal prediction approach to construct interval forecasts without requiring the number of candidate models or predictors to be fixed; see Algorithm (ref). Without loss of generality, we assume that ${\sf T} h$ takes a positive integer value.
To establish a valid coverage property of the proposed conformal prediction procedure, we require the following technical assumptions.
Letting $\widehat{u}_{i,t}=\sigma_i^{-1}(\tau_t)\widehat{\epsilon}_{i,t}$ with $\tau_t=t/{\sf T}$, we define the following in-sample and out-of-sample performance measures: $$\delta_{i,{\sf T}}^2=\frac{1}{{\sf T} h+1}\sum_{t={\sf T}-{\sf T} h}^{{\sf T}}(|\widehat{u}_{i,t}|-|u_{i,t}|)^2\quad \text{and}\quad \psi_{i,{\sf T}+1}=\big||\widehat{u}_{i,{\sf T}+1}|-|u_{i, {\sf T}+1}|\big|.$$ The following proposition gives the finite-sample upper bound on the gap between the actual coverage probability and the nominal level $1-\alpha$ and further establishes the asymptotic validity of the developed PTVMA-based conformal prediction method.
{\bf Remark 4.1.}\ \ We allow all the candidate models to be misspecified to obtain the finite-sample upper bound for the deviation of the actual coverage probability from the nominal level in ((ref)). To further derive the asymptotic validity of the constructed prediction interval in ((ref)), we need to prove that ${\sf E}(\delta_{i,{\sf T}}^{2})\rightarrow0$ and ${\sf E}(\psi_{i,{\sf T}+1})\rightarrow0$, which requires ${\mathscr D}$ to be not empty, indicating that there exists at least one correctly-specified candidate model so that Theorem (ref) is applicable.
\setcounter{equation}{0}
To demonstrate advantages of the proposed PTVMA procedure, this section creates a sharp challenge for classic approaches in which any single-network specification may be misspecified. In the two simulation settings, model selection over a set of single-network specifications is global misspecified because no candidate model coincides with the true DGP throughout the entire period. By comparing the finite-sample performance of PTVMA with several competing models, our simulation experiment examines whether the proposed PTVMA can adaptively learn relevant network layers and variables over time and deliver accurate predictions despite persistent specification uncertainty.
To evaluate the in-sample performance, we calculate the in-sample ratio of the root mean squared prediction errors (RMSPE) to compare the proposed model with benchmark models:
where $R=1000$ is the replication number in simulations, $y_{it:r} $ is the true value in the $r$-th replication, $ \widehat{y}_{it:r}^\text{PTVMA}$ is the fitted value from the PTVMA, and $ \widehat{y}_{it:r}^\text{Benchmark} $ is from a benchmark model. We consider two benchmarks. The network vector autoregressive (NAR) model ZPLLW17 includes the same response and predictors as the DGP but restricts the parameters to be time-invariant, estimating them via the OLS. The time-varying NAR (tv-NAR) model LPTW26 adopts the same DGP specification, utilizing a local linear approach to estimate time-varying coefficients. The proposed PTVMA method outperforms the benchmark model when the ${\sf RMSPE}_\text{in}$ ratio is smaller than one.
The out-of-sample predictive performance is assessed by employing the rolling forecasting procedure. Specifically, we compute the out-of-sample ratio of the RMSPE for the $l$-step-ahead out-of-sample rolling prediction of $ \widehat{\bm{Y}}_{{\sf T}+l} $:
where $ \widehat{y}_{i,t+l:r}^\text{PTVMA}$ and $\widehat{y}_{i,t+l:r}^\text{Benchmark}$ denotes the $l$-step-ahead rolling prediction in the $r$-th replication using the proposed PTVMA and benchmark methods, respectively, ${\sf T}_0 = \lfloor 0.4{\sf T} \rfloor $ is the rolling window length. Under the iterated scheme, multi-step-ahead forecasts are obtained by estimating a one-step-ahead model and recursively replacing future lagged values with their forecasts.
The performance of conformal prediction is evaluated by the empirical coverage probability. Specifically, we use the test set to calculate empirical coverage probabilities:
where $C_{i:r}^{\alpha}(\mathbf{Z}_{t-1})$ is the conformal prediction interval constructed for individual $i$ at time $t$ in the $r$-th replication.
The DGP 1 simulates a setting in which the response variable $ y_{it} $ is driven by two distinct networks, $\mathbf{W}^{(1)}$ and $\mathbf{W}^{(2)}$, with time-varying impacts $\beta_{t1}$ and $\beta_{t2}$. Consider
where $ \widetilde{\omega}_{ij}^{(k)} = \omega_{ij}^{(k)} / \sum_{j \neq i} \omega_{ij}^{(k)} $ is the $ (i,j) $-th element of the matrix obtained by row-normalizing the adjacency matrix $ \mathbf{W}^{(k)} $, $ \mathbf{X}_i = \left( x_{i,1}, x_{i,2}, x_{i,3} \right)^{^\intercal} $ is independently drawn from a three-dimensional normal distribution across $ i $ with mean $ \bm{0} $ and covariance $ \bm{\Sigma}_X = (\sigma_{x,q_1q_2})_{3 \times 3} $ with $ \sigma_{x,q_1q_2} = 0.5^{|q_1-q_2|} $ for $ 1 \leq q_1, q_2 \leq 3 $, the error term $ \epsilon_{i,t} = \sigma_i(t/{\sf T}) u_{i,t} $, $ u_{i,t} $ is independently simulated by the standard normal distribution, $ \sigma_i(t/{\sf T}) = (0.25 \sin (t/{\sf T})+0.3)^{0.5} $ for odd $ i $, and $ \sigma_i(t/{\sf T}) = (0.25 \cos (t/{\sf T})+0.3)^{0.5} $ for even $ i $. True time-varying coefficients are set as
The first adjacency matrix $ \mathbf{W}^{(1)} $ is generated by the stochastic block model with two blocks. Each node is randomly assigned for one block label with equal probability and probabilities within and between blocks are generated by: $ {\sf P} (\omega_{ij}^{(1)} = 1) = 0.5 $ if $ i $ and $ j $ are in the same block; $ {\sf P} (\omega_{ij}^{(1)} = 1) = 0.1 $ otherwise. The second adjacency matrix $ \mathbf{W}^{(2)} $ is generated by using the power-law distribution model. The power-law distribution of degrees widely exists in the real world, where most agents have low degrees while a few hubs have very high degrees. We follow ZPLLW17 and define $ \mathbf{W}^{(2)} $ as follows. First, generate for each agent's $ degree_i $ via the discrete power-law distribution $ {\sf P} ({\sf degree}_i = k) \propto k^{-c} $, where $c $ is the exponent parameter and a smaller value implies a heavier distribution tail.\footnote{Here, we set $ c = 2.5 $ and the created networks share a similar density size at around $ 20\% $.} Next, we randomly select ${\sf degree}_i $ agents to be agent $ i $'s followers.
We consider $ {\sf M} = 8 $ candidate models in the proposed PTVMA for DGP 1 and list their respective predictors in Table (ref). Note that $ \beta_{t1} $ is significant in the first half of the simulation period but not in the second, while $ \beta_{t2} $ exhibits the reverse pattern that is insignificant in the first half and significant in the second. Consequently, candidate models 2 and 4 are correctly specified in the first and second half of the sample period, respectively, but misspecified in the other period. We also involve irrelevant predictors $ x_{i4} $, $ x_{i5} $, $ \sum_{j \neq i} \widetilde{\omega}_{ij}^{(3)} y_{j,t-1} $ and $ \sum_{j \neq i} \widetilde{\omega}_{ij}^{(4)} y_{j,t-1} $ into some candidate models to demonstrate that the PTVMA method can dynamically select the correctly specified model if it is in the candidate models. Irrelevant predictors $ x_{i4} $ and $ x_{i5} $ are independently drawn from a normal distribution with mean $ 1 $ and variance $ 1 $. We refer to ZY18 and adopt the left-right matrix as one irrelevant adjacency matrix $ \mathbf{W}^{(3)} $, where each agent interacts only with its left and right neighbors. The other irrelevant adjacency matrix $ \mathbf{W}^{(4)} $ is constructed by a distance-based spatial weight matrix. We first independently generate agent $ i $'s location $ (x_i, y_i) $ from $ \mathrm{U} [0, 1] $, and then compute the pairwise Euclidean distance $ d_{ij} $ between agent $ i $ and $ j $. The spatial weight matrix is constructed using an exponential distance-decay function, with off-diagonal elements $ \omega_{ij}^{(4)} = \mathrm{I} \left( \exp \{ -10 d_{ij} \} < d_0 \right) $ and zero diagonal, where $ d_0 $ is the 25% quantile of $ \{ d_{ij} \} $.
As reported in Table (ref), the PTVMA delivers uniformly better in-sample fit than the NAR benchmark. The RMSPE ratio (PTVMA/NAR) ranges from $ 0.746 $ to $ 0.825 $ across all $ ({\sf N}, {\sf T}) $ designs, indicating significant reductions of in-sample prediction error. Furthermore, PTVMA delivers nearly the same in-sample performance as the correct specification (tv-NAR model), with RMSPE ratios tightly around $ 1.002 $ and decreasing as either $ {\sf N} $ or $ {\sf T} $ grows. Given that the tv-NAR model is correctly specified here, this comparison suggests that the PTVMA approximates the true model specification without priori information.
Table (ref) also reports the average weights of eight candidate models over the $ 1000 $ replications. The estimated weights are highly concentrated on candidate model 2 and 4, each receiving about $ 0.48 $ and jointly accounting for more than $ 95\% $ of the total weight. In contrast, the (averaged) time-varying weight estimates for the other six candidate models are negligible across all $ {\sf N} $ and $ {\sf T} $. This pattern is consistent with the design of DGP 1: the true network is generated from a combination of $ \mathbf{W}^{(1)} $ and $ \mathbf{W}^{(2)} $ while $ \mathbf{W}^{(3)} $ and $ \mathbf{W}^{(4)} $ are irrelevant alternatives. Overall, the evidence indicates that PTVMA places most of the weight on DGP-relevant network components while shrinking irrelevant candidates toward zero, confirming Theorem (ref) in finite samples.
Figure (ref) illustrates a central implication of the proposed PTVMA: when the strength of network dependence is time-varying, the optimal weights are expected to be time-varying as well. The solid lines report the averaged time-varying weights on candidate model 2 and 4, and the dashed lines plot the true coefficient functions: $ \beta_{t1} $ and $ \beta_{t2} $. The plots show that the assigned weights closely track the DGP 1: the PTVMA places most weight on candidate model 2 in the first half of the sample period and rapidly shifts weight toward candidate model 4 in the second half, consistent with the time-varying pattern of true coefficient functions. In addition, the weight reallocation is concentrated around the mid-sample transition where $ \beta_{t1} $ and $ \beta_{t2} $ switch most sharply, indicating PTVMA's strong adaptivity to smooth structural changes.
Table (ref) reports out-of-sample RMSPEs of the PTVMA relative to two benchmarks under DGP 1. Nearly all reported ratios fall below unity, confirming that the PTVMA uniformly dominates both the NAR and the tv-NAR models across all forecast horizons and sample sizes. In particular, the PTVMA outperforms the tv-NAR model even though the latter is correctly specified under DGP 1. The proposed PTVMA achieves the forecast gain by adaptively weighting candidate models as network dependence varies over time, and thus delivers higher forecast accuracy than a correctly specified benchmark model. The last column of Table (ref) reports the empirical coverage probabilities of conformal prediction intervals at 90% nominal level. The coverage probabilities are close to 90% in general, suggesting that the conformal intervals are well-calibrated in finite samples.
In the DGP 2 below, we consider a more challenging setting with all the single-network specifications misspecified over any time period. Define
where the true time-varying coefficients are set as
indicating that the impact of $ \mathbf{W}^{(1)} $ diminishes and that of $ \mathbf{W}^{(2)} $ grows over the time. The other model settings are the same as those in DGP 1. The predictors of candidate models for DGP 2 are listed in Table (ref).
The in-sample performance of the PTVMA under DGP 2 exhibits a similar pattern to that in DGP 1, demonstrating the PTVMA's advantages in the misspecification scenario. As reported in Table (ref), the PTVMA substantially outperforms the NAR model (RMSPE ranging from $ 0.550 $ to $ 0.570 $) while delivers a near-identical fit to the correctly specified tv-NAR model (RMSPE tightly around $ 1.001 $). This result confirms that even when all the candidate models are misspecified, the proposed PTVMA can adaptively combine them to approximate the response as effectively as a correctly-specified model.
Table (ref) also reports the averaged weight estimates of six candidate models over $1000$ replications to further reveal PTVMA's weight-assignment capacity. Due to incorporating both autoregressive lags, candidate models 3 and 6 receive the majority of weight with approximately $ 0.41 $ and $ 0.31 $, respectively. Note that the other candidate models receiving negligible weights either omit lags or contain irrelevant predictors. Consequently, these specifications are less informative than candidate models 3 and 6 under the DGP 2. The results confirm that the proposed PTVMA successfully allocates significant weights to the candidate models with most relevant predictors whereas effectively shrinks the other irrelevant specifications.
Table (ref) presents out-of-sample RMSPE for both the one- and multiple-step-ahead forecasts. Across all horizons and sample sizes, the PTVMA consistently outperforms two benchmarks and achieves lower forecast errors than the correctly specified model (tv-NAR), again demonstrating that the proposed model averaging procedure delivers forecast gains by adaptively combining information from candidate models even when all of them are misspecified.
\setcounter{equation}{0}
Accurate forecasting of Consumer Price Index (CPI) inflation plays a critical role in the conduct of monetary policy and preservation of financial stability. Inflation expectations serve as a key anchor for policy credibility and exert a substantial influence on asset pricing. Nevertheless, forecasting inflation remains a persistent challenge, even when employing comprehensive datasets and state-of-the-art econometric methods SW07. These difficulties are compounded in open economies, where domestic inflation is increasingly shaped by global shocks and the cross-country transmission of price pressures. A growing body of literature demonstrates that inflation dynamics are increasingly dominated by global determinants, even as domestic conditions remain relevant CM10. Key international transmission channels include global production networks, international financial markets, trade and policy shocks CNST21, MR20, ARW19. These channels subsequently form multi-layer networks that evolve over time, creating substantial model uncertainty and structural instability. Traditional univariate or panel models often fail to capture complex cross-country spillovers and time-varying nature of inflation transmission, leading to potentially misspecified models and degraded forecast performance. This section aims to address this issue by combining information from multiple international networks and applying the proposed PTVMA method to forecast CPI inflation,
The CPI inflation data is obtained from the International Monetary Fund (IMF)\footnote{The data is downloaded from the IMF Data, available at: \url{https://data.imf.org/en}.}. We consider a dataset on monthly CPI inflation in 36 economies ($ {\sf N} = 36 $) from Januaray 2007 to December 2024 with 216 observations ($ {\sf T} = 216 $). This economy set covers the largest and most systemically relevant economies with established equity markets, ensuring representation of advanced economies, major emerging markets, and key commodity producers that jointly account for the overwhelming majority of global output.\footnote{Selected economies include Belgium, Brazil, Canada, Switzerland, Chile, China, Colombia, the Czech Republic, Germany, Denmark, Spain, Finland, France, the United Kingdom, Greece, Hong Kong, Hungary, Indonesia, India, Ireland, Israel, Italy, Japan, South Korea, Mexico, Malaysia, the Netherlands, Norway, Poland, Portugal, Singapore, Sweden, Turkey, the United States, Vietnam, and South Africa. We exclude nine candidate economies due to data limitations during the sample period: Russia and Saudi Arabia are omitted due to missing daily MSCI indices, while the United Arab Emirates, Argentina, Australia, New Zealand, Peru, the Philippines, and Thailand are excluded due to the unavailability of monthly inflation data.} We include two variables as controls in the model: a time-varying trend term, capturing exogenous shocks such as oil price fluctuations; and the monthly period-average nominal exchange rate for economy $ i $ (domestic currency per U.S. dollar), sourced from the IMF Data.
We consider four types of networks to characterize the international transmission channels of CPI inflation, each reflecting a distinct economic mechanism. The economic rationale for each network and the corresponding data sources are discussed below, with detailed network construction procedures presented in Appendix (ref).
Consider the $m$-th candidate model specification:
where $y_{i,t} $ denotes the CPI inflation of economy $ i $ in month $ t $, $ {\sf ER}_{i,t} $ is the period-average nominal exchange rate of economy $ i $ in month $ t $, $ \gamma_{t1}^{(m)} $ is the time-varying trend term, and $ \widetilde{\omega}_{ij}^{(m)} $ is the $ (i,j) $-th element of the matrix constructed from the network layer $ \mathbf{W}^{(m)} $ by setting its diagonal elements to zero and applying row-normalization. We consider $ 4 $ and $ 10 $ network layers for in- and out-of-sample analysis, respectively. Details on network constructions are presented in Appendix (ref).
Figure (ref) plots the estimated time-varying weights of four candidate models over 2007--2024. The weights of candidate models exhibit pronounced time variation. The production network layer is dominant in the early part of the time period, especially around 2009--2012, and then rapidly loses most weight in the later period. The policy network layer becomes prominent in 2013--2019, while the global equity network layer receives substantial weight in the late period (after 2019). In contrast, the trade network layer is assigned only intermittent weight and never attains persistent dominance. Figure (ref) suggests that the transmission mechanism of CPI inflation shifts across different types of network dependence and varies along the time. This finding is consistent with multiple transmission channels for CPI inflation, with no single network maintaining uniformly dominant weights. A single-layer specification would thus be too restrictive when another layer better matches the network dependence structure. It also shows that the candidate weights concentrate on one or two layers at a time and shift as the underlying structure evolves, indicating that the PTVMA adaptively captures locally optimal specifications which may vary over the temporal dimension.
Overall, the in-sample results provide evidence that global inflation transmission is governed by a multi-layer, regime-dependent structure in which production, financial, trade, and policy linkages alternately dominate. Therefore, conventional time series models with a single network structure or time-invariant parameters fail to capture the shifting dominance of these transmission channels, underscoring the value of the proposed PTVMA.
Table (ref) reports the out-of-sample prediction performance of the proposed model for CPI inflation. Predictive accuracy is evaluated by the RMSPE, and our model is compared with seven alternative specifications: NAR, time-varying NAR (tv-NAR), VAR(1) estimated by the OLS (VAR), VARX(1) estimated by the Elastic Net (HD-VARX), adaptive LASSO (ALASSO), random forest (RF), and stochastic gradient boosting (SGB), spanning one-, two-, and three-step-ahead forecasts across VAR, penalized regression, and state-of-the-art machine learning methods. Exogenous variables are the nominal exchange rates for all economies. In addition, the ALASSO, the RF and the SGB also involves in- and out-degrees of networks from $ \mathbf{W}^{(1)} $ to $ \mathbf{W}^{(8)} $ and degrees of networks $ \mathbf{W}^{(9)} $ and $ \mathbf{W}^{(10)} $ as controls, see Appendix (ref) for details.
Table (ref) shows a pronounced advantage of the PTVMA in short-run CPI inflation forecasting. At the one-step horizon, our method uniformly outperforms all competitors and attains the lowest RMSPE (with all the ratios strictly smaller than one), indicating a marked improvement in short-run predictive accuracy. At the two-step horizon, the proposed PTVMA performs slightly worse than the NAR and ALASSO but continues to outperform the remaining methods. At the three-step horizon, however, the PTVMA only performs better than the tv-NAR and two VAR-based methods, suggesting that the short-horizon advantage attenuates as the forecast horizon expands. This is because the time-varying coefficient specification amplifies the negative impact of cumulative errors on multi-step-ahead forecasting. It is also worth noting that the proposed PTVMA still dominates the tv-NAR across three horizons, highlighting the advantage of forecast combination via time-varying model averaging.
This paper introduces a dynamic model averaging criterion within a general time-varying network VAR framework which incorporates multiple adjacency matrices to capture network spillover effects. The main aim is to determine a time-varying optimal combination of multi-layer network VAR candidate models whose number may be divergent. In particular, all the candidate models with different linkage structures and varying lagged orders of dependent variables and node-specific predictors are allowed to be misspecified. Both the in-sample and out-of-sample asymptotic properties including the asymptotic optimality and estimation convergence rates are derived under some technical conditions. In particular, the proposed PTVMA achieves the selection consistency (of correctly-specified models) as the conventional model selection criterion. Furthermore, an extended conformal prediction method is proposed to construct prediction intervals for network time series. The Monte Carlo simulation demonstrates that the proposed PTVMA delivers superior out-of-sample predictive performance relative to the benchmark, even when all the candidate models are misspecified. The empirical case study further reveals the time-varying patterns of optimal weight estimation and shows that our model and method consistently outperform other competing models in the short-run CPI inflation prediction.