EconBase
← Back to paper

Granger Causality Testing in High-Dimensional VARs: a Post-Double-Selection Procedure

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

92,211 characters · 14 sections · 113 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Granger Causality Testing in High-Dimensional VARs: a Post-Double-Selection Procedure

abstractWe develop an LM test for Granger causality in high-dimensional VAR models based on penalized least squares estimations. To obtain a test retaining the appropriate size after the variable selection done by the lasso, we propose a post-double-selection procedure to partial out effects of nuisance variables and establish its uniform asymptotic validity. We conduct an extensive set of Monte-Carlo simulations that show our tests perform well under different data generating processes, even without sparsity. We apply our testing procedure to find networks of volatility spillovers and we find evidence that causal relationships become clearer in high-dimensional compared to standard low-dimensional VARs. \noindentKeywords: Granger causality, Post-double-selection, vector autoregressive models, high-dimensional inference.\\ JEL codes: C55, C12, C32.

\doublespacing

Introduction

Economics, statistics and finance have seen a rapid increase of applications involving time series in high-dimensional systems. Central to many of these applications is the vector autoregressive (VAR) model that allows for a flexible modelling of dynamic interactions between multiple time series. In this paper we develop a simple method to test for Granger causality in high-dimensional VARs (HD-VARs) with potentially many variables.

Many financial applications consider Granger causality analysis, especially for constructing high-dimensional networks. Networks of financial firms' intedependencies are investigated in basu2015network, gao2017efficient, demirer2018estimating and barigozzi2019nets. Similarly, spillovers and contagion among stock returns are investigated in networks using Granger causality analysis in lin2017regularized, vyrost2015granger and corsi2018measuring.

Most of the econometric literature has traditionally been focused on allowing for high dimensionality in VARs through the use of factor models bernanke2005measuring,chudik2016theory or Bayesian methods banbura2010large. For instance billio2012econometric develops measures of connectedness to assess systemic risk propagation among institutions in the financial system using principal component analysis and Granger causality networks. Recent years have seen an increase in regularized, or penalized, estimation of sparse VARs based on popular methods from statistics such as the lasso tibshirani1996regression and elastic net zou2005regularization, which impose sparsity by setting a (data-driven) selection of the coefficients to zero.

Compared to factor models, such sparsity-seeking methods have often an advantage of interpretability, as in many economic applications, it appears natural to believe that the most important dynamic interactions among a large set of variables can be adequately captured by a relatively small -- but unknown -- number of `key' variables. As such, the use of these methods for estimating HD-VAR models has also increased significantly in recent years, see e.g. nicholson2017varx, basu2019low, billio2019bayesian, wilms2018algorithm, korobilis2019adaptive).

Regularized estimation theory for high-dimensional time series and VAR models is now well established, see among others song2011large, basu2015regularized, kock2015oracle, davis2016sparse, medeiros2016l1, audrino2018oracle and masini2019regularized and wong2020lasso; kock2020penalized provide a recent review. However, performing inference on HD-VARs, such as testing for Granger causality, still remains a non-trivial matter. As is well known, performing inference after model selection (post-selection inference) is complicated as the selection step invalidates `standard' inference where the uncertainty regarding the selection is ignored leeb2005model. Complexities introduced by the temporal and cross-sectional dependencies in the VAR mean that most recently developed post-selection inference methods are not automatically applicable.

Most existing literature on Granger causality testing in HD-VARs therefore has so far not considered post-selection inferential procedures. wilms2016predictive propose a bootstrap Granger causality test in HD-VARs, but do not account for post-selection issues. Similarly, skripnikov2018joint investigate the problem of jointly estimating multiple network Granger causal models in VARs with sparse transition matrices using lasso-type methods, but focus mostly on estimation rather than testing. song2019better focus on statistical procedures for testing indirect/spurious causality in high-dimensional scenarios, but consider factor models rather than regularized regression techniques. lin2017regularized consider high-dimensional multi-block VARs derived from a two-blocks recursive linear dynamical system and use a maximum likelihood (ML) estimator for Gaussian data. In order to obtain the ML estimates for the system transition matrices and the precision matrix, respectively the lasso and graphical lasso on the residuals are iterated until convergence. krampe2018bootstrap develops bootstrap techniques for sparse VAR models combining a model-based bootstrap procedure and the de-sparsified lasso (see van2014asymptotically) to perform inference on the autoregressive parameters. chaudhry2017uncertainty look at de-biased estimators as in javanmard2014confidence, for Gaussian and sub-Gaussian VAR processes with a focus on Granger-causality and control of the false discovery rate.

In this paper we build on the post-double-selection approach proposed by belloni2014inference, to develop a valid post-selection test of Granger causality in HD-VARs. The finite-sample performance depends heavily on the exact implementation of the method. In particular, the tuning parameter selection in the penalized estimation is crucial. We therefore perform an extensive simulation study to investigate the finite-sample performance of the different ways to set up the test in order to be able to give some practical recommendations. In addition, we investigate the construction of networks of realized volatilities using a sample of 30 financial stocks modeled as a vector heterogeneous VAR corsi2009simple. We are able to demonstrate how our approach allows for obtaining much sharper conclusions than standard low-dimensional VAR techniques.

The remainder of the paper is as follows: Section (ref) introduces the high-dimensional VAR model and Granger causality tests. In Section (ref) we propose our estimation and inferential framework. Section (ref) establishes the asymptotic properties of our method and discusses the assumptions required for the theory to hold. Section (ref) reports the results of the Monte Carlo simulations. We apply our method in Section (ref) to construct volatility spillover networks. Section (ref) concludes. Proofs and supplemental results can be found in the appendix.

A few words on notation. For any $n$-dimensional vector $\bm x$, we let ${\left\lVert\bm x\right\rVert}_p = \left(\sum_{i=1}^n {\left\lvertx_i\right\rvert}^p \right)^{1/p}$ denote the $\ell_p$-norm. For any index set $S \subseteq \{1, \ldots, n\}$, let $\bm x_{S}$ denote the sub-vector of $\bm x_t$ containing only those elements $x_i$ such that $i \in S$. ${\left\lvertS\right\rvert}$ denotes the cardinality of the set $S$. We use $\xrightarrow{p}$ and $\xrightarrow{d}$ to denote convergence in probability and distribution, respectively.

High-dimensional Granger causality tests

Loosely speaking, the notion of Granger causality captures predictability given a particular information set granger1969investigating,granger1980testing. If the addition of variable $X$ to the given information set $\Omega$ alters the conditional distribution of another variable $Y$, and both $X$ and $\Omega$ are observed prior to $Y$, then $X$ improves predictability of $Y$, and is said to Granger cause $Y$ with respect to $\Omega$. granger1969investigating originally envisioned the information set $\Omega$ “be all the information in the universe” (p. 428), which is of course not a workable concept. Yet clearly the choice of information set has a major effect on the interpretation of the finding of (non-)Granger causality, as discussed in granger1980testing. In particular, spurious Granger causality from $X$ to $Y$ may be found when both $X$ and $Y$ are Granger caused by $Z$, but $Z$ is omitted from $\Omega$. As such, one might want to include as many potentially relevant variables in the information set as possible in order to avoid finding spurious causality due to omitted variables, thereby moving as much as possible towards the universal information set envisioned by Granger. However, conditioning on so many variables leads to obvious problems of high-dimensionality rendering many standard statistical techniques invalid.

In this paper we focus on testing Granger causality in mean using linear models, in which setup the VAR model is the natural tool to investigate this problem. However, to enlarge the information set means estimating a VAR with an increasing number of variables. The number of parameters in a VAR increases quadratically with the number of time series included; an unrestricted VAR($p$) has $K^2p$ coefficients to be estimated, where $K$ is the number of series and $p$ is the lag-length. As the time series dimension $T$ is typically fairly small for many economic applications, the data do not contain sufficient information to estimate the parameters and consequently standard least squares and maximum likelihood methods suffer from the curse of dimensionality, resulting in estimators with high variance that overfit the data.

Granger causality testing in VAR models

Let $\bm y_1,\ldots,\bm y_T$ be a $K$-dimensional multiple time series process, where $\bm y_t=(y_{1,t},\ldots,y_{K,t})^{'}$ is generated by a VAR($p$) process

equation[equation omitted — 118 chars of source]

where for notational simplicity we assume the variables have zero mean; if not they can be demeaned prior to the analysis, or equivalently a vector of intercepts is added. $\bm A_1,\ldots,\bm A_{p}$ are $K\times K$ parameter matrices and $\bm u_t$ is a martingale difference sequence (mds) of error terms. We consider weakly stationary VAR models, as formalized in Assumption (ref) below.

assumptionThe VAR model in (ref) satisfies: \begin{enumerate}[(a)] • $\{\bm u_t\}_{t=1}^T$ is a weakly stationary mds with respect to $\mathcal{F}_t = \sigma(\bm y_t, \bm y_{t-1}, \bm y_{t-2}, \ldots)$ $\bm u_t$ such that $\mathbb{E} (\bm u_t| \mathcal{F}_{t-1}) = \bm 0$ for all $t$ and $\bm \varSigma_{u} = \mathbb{E} (\bm u_t \bm u_t^\prime)$ is positive definite. • All roots of $\det(\bm I_{K}-\sum_{j=1}^{p} \bm A_j z^j)$ lie outside the unit disc, such that the lag polynomial is invertible. \end{enumerate}

In the VAR model (ref) we are interested in testing whether variables in the set $J$ Granger cause variables in the set $I$ in mean, conditional on all the other variables, where $J, I \subset \{1,\ldots, K\}$ and $J \cap I = \emptyset$. Let $N_I = {\left\lvertI\right\rvert}$ and $N_J = {\left\lvertN_J\right\rvert}$ denote the number of variables in $I$ and $J$ respectively. We describe our proecdure here in general form for testing blocks of variables. For any sets $S_1, S_2 \subseteq \{1, \ldots, K\}$ of variables define the best linear predictor in $L_2$-norm of $\bm y_{S_1,t}$ given $\bm x_{S_2,t-1}^{(p)} = (\bm y_{S_2,t-1}^\prime, \ldots, \bm y_{S_2, t-p}^\prime)^\prime$ as $\mathcal{P} (\bm y_{S_1,t}|\bm x_{S_2,t-1}^{(p)}) = \bm \varGamma^{*} \bm x_{S_2,t-1}^{(p)}$, where $\bm \varGamma^* = \min_{\bm \varGamma} \mathbb{E} \left[\left.{\left\lVert\bm y_{S_1,t} - \bm \varGamma \bm x_{S_2, t-1}\right\rVert}_2^2 \right. \right]$. Then we say that $\bm y_{J,t}$ does not Granger cause $\bm y_{I,t}$ conditionally on $\bm x_{J^c,t}$ if

equation[equation omitted — 128 chars of source]

for any value of $\bm x_{J^c,t}$. In other words, conditional on $\bm x_{J^c,t}$, addition of the lags of $\bm y_{J,t}$ to the information set does not improve predictability of $\bm y_{I,t}$. Note that Granger (non-)causality as defined in (ref) is a property of the population. In the VAR (ref) this means that testing for Granger causality can be done via testing the joint significance of the blocks of coefficients in the matrices $\bm A_1, \ldots, \bm A_p$ corresponding to the impact of variables $J$ on $I$.

To illustrate, consider (ref) with $p=1$ lag, and assume without loss of generality that the variables in $\bm y_t$ are ordered such that $\bm y_t = \left(\bm y_{I,t}^\prime, \bm y_{J, t}^\prime, \bm y_{-(I \cup J), t}^\prime\right)^\prime$, where $-(I \cup J))$ refers to all variables not in $J$ or $I$. Then we can write

equation[equation omitted — 439 chars of source]

where $\bm A$ is partitioned conformably with the blocks in $\bm y_t$. In this case, the best linear predictors in (ref) are given by

equation*[equation* omitted — 430 chars of source]

For any arbitrary value of $\bm y_{t-1}$, these can only coincide if $\bm A_{I, J} = \bm 0$. Hence, the null hypothesis of no Granger causality from $J$ to $I$ in the VAR($1$) model can be formulated in terms of $\bm A_{I, J} = \bm 0$. This is easily extended to $p>1$ by simply testing if the $(I, J)$-block of all $p$ lag matrices is equal to zero.

In the remainder of the paper, we will be working with a stacked representation of (ref) for the variables in $I$. Specifically, let $\bm Y = \left(\bm y_{p+1}, \ldots, \bm y_T\right)^\prime$ and let $\bm y_{I} = \operatorname{vec}\left(\bm Y_{I}\right)$ denote the $N_I \times 1$ stacked vector containing all observations corresponding to the variables in $I$. Similarly, let $\bm u_{I} = \operatorname{vec}(\bm U_{I})$, where $\bm U = \left(\bm u_{p+1}, \ldots, \bm u_T\right)^\prime$. Let $\bm X = \left(\bm x_{p}^{(p)}, \ldots, \bm x_{T-1}^{(p)}\right)^\prime$ and $\bm X^{\otimes} = \bm I_{N_I} \otimes \bm X$, while defining the stacked parameter vector $\bm \beta = \operatorname{vec}((\bm A_1, \ldots, \bm A_p)^\prime)$. Then we can write

equation[equation omitted — 173 chars of source]

where $\bm X_{GC}^{\otimes} = \bm I_{N_I} \otimes \bm X_{GC}$, and $\bm X_{GC} = \left(\bm x_{J, p}^{(p)}, \ldots, \bm x_{J, T-1}^{(p)}\right)^\prime$ contains those columns of $\bm X$ corresponding to the potentially Granger causing variables in $J$; $\bm X_{-GC}$ and $\bm X_{-GC}^{\otimes}$ are then defined similarly but containing the remaining variables.\footnote{Note that if $I = \{i\}$ for one particular value of interest, then (ref) simply corresponds to a single equation from the VAR in (ref).} Testing for no Granger causality is then equivalent to testing $H_0: \bm \beta_{GC}=\bm 0$ against $H_1: \bm \beta_{GC}\neq \bm 0$.

Define $N_J = {\left\lvertJ\right\rvert}$ and $N_I = {\left\lvertI\right\rvert}$. Note that $\bm \beta_{-GC}$ has $\left(K - N_J \right) \times N_I \times p$ elements, which we assume large through having a large number of variables $K$. On the other hand, throughout the paper we assume that $N_J$, $N_I$ and $p$ are small, or more precisely, fixed when sample size increases to infinity. As $\bm \beta_{GC}$ has $N_{GC} = N_J \times N_I \times p$ elements, these are also implied to be fixed. While theoretically it is possible to consider an increasing number of elements in $\bm \beta_{GC}$ (see Remark (ref) for details), it would not be required for typical applications. $J$ and $I$ are under the researcher's control and in most applications it is natural to consider a small number of variables of interest; often both $J$ and $I$ will only consist of a single variable, as in our application.

For $p$ it may appear more restrictive to assume it small. However, large $p$ in univariate regressions or small systems often arise from neglected dynamics with omitted variables hecq2016univariate. As our HD-VAR attempts to include many more variables than typical small systems, we hope to alleviate the omitted variable issue, and thereby also directly making smaller $p$ much more realistic. Of course, $p$ is generally unknown in practice. However, in many applications it is possible to give a reasonable (and small) upper bound on $p$, which is sufficient for our algorithm. If not, $p$ has to be estimated. We discuss two ways in the next section.

remarkOur operational version of Granger causality only considers causality in mean. Additionally, one might argue that considering only linear models is a further restriction on the generality of the concept of Granger causality. However, in our high-dimensional approach linear models are less restrictive as would appear. First, the VAR does not have to be formulated for levels of variables of interest. In fact, in our application we formulate a VAR for (realized) variances, such that we are implicitly testing Granger causality in second moments rather than first moments. Second. the linear VAR model in many cases provides a good approximation to a general nonlinear process via a Wold-type representation argument, see e.g. meyer2015on. Finally, non-linear transformations (such as powers) of the original variables can be added to (ref), by which general functional forms can be approximated (even if one then strictly loses the VAR equivalence). While in small systems this is infeasible as it increases the dimensionality dis-proportionally, our high-dimensional approach can handle this without any conceptual issues. In fact, belloni2014high explicitly motivate their high-dimensional linear approach as an approximation to a general function; their arguments apply here as well.

Inference after selection by the lasso

In this section we introduce our inferential procedure to the Granger causality tests in high-dimensional VARs. We first discuss the lasso, which we use in the initial stage to select relevant variables. Next we discuss how naive use of the lasso introduces post-selection problems for inference, and we propose our algorithm to remedy this.

The lasso estimator

As $\bm \beta$ is high-dimensional when $Kp$ is large relative to $T$, least squares estimation is not appropriate, and a structure must be imposed on $\bm \beta$ to be able to estimate it consistently. We assume sparsity of $\bm \beta$; that is, we assume that $\bm \beta$ can accurately be approximated by a coefficient vector with a (significant) portion of the coefficients equal to zero.

The sparsity assumption validates the use of variable selection methods, thereby reducing the dimensionality of the system without having to sacrifice predictability. For a general $n$-dimensional vector of responses $\bm y$ and $n \times M$-dimensional matrix of covariates $\bm X$, the (weighted) lasso simultaneously performs variable selection and estimation of the parameters by solving

equation[equation omitted — 237 chars of source]

where $\lambda$ is a non-negative tuning parameter determining the strength of the penalty, and $\{w_m\}_{m=1}^{M}$ are non-negative weights corresponding to the parameters in $\bm \beta$. For the standard lasso the weights are either equal to one, or equal to zero (if this parameter should not be penalized). The notation $\hat{\bm \beta}(\lambda)$ highlights that the solution to the minimization problem depends on $\lambda$, which has te be selected as well (see Section (ref)). When no confusion can arise, we simply write $\hat{\bm \beta}$.

One may also consider the adaptive lasso zou2006adaptive with parameter-specific weights $w_j$ in (ref) based on an initial estimation of $\bm \beta$, which is able to delete more irrelevant variables. However, for our purpose such oracle properties are not very relevant; we wish to eliminate the effects of the other “nuisance” variables on the relation between the variables tested for Granger causality, but we do not need to identify which of these nuisance variables matter.

Theoretical properties of lasso estimation in stable VAR models have now been studied extensively. We here non-exhaustively mention some of the key results for our setting; see kock2020penalized for a thorough review. kock2015oracle derive oracle properties of the adaptive lasso for VAR models. basu2015regularized establish restricted eigenvalue conditions for VAR models and show their sufficiency for estimation consistency. medeiros2016l1 relax the Gaussianity assumptions of these papers by considering conditionally heteroskedastic errors, and demonstrate that the adaptive lasso retains oracle properties in time series settings. Finally, masini2019regularized derive bounds on estimation errors in approximately sparse VAR models under very general conditions, allowing for heavy tails and dependence in the error terms. In particular, they show that several commonly used volatility processes in financial research satisfy these assmuptions, thereby formally establishing the suitability of the lasso for many financial applications of VAR models.

Post-Double-Selection Granger causality test

The need for post selection inference

One might be tempted to simply perform the (adaptive) lasso as in (ref) on (ref), setting $w_{GC} = 0$, and then testing whether $\bm \beta_{GC}=0$, potentially after re-estimating the model by OLS on only the selected variables. However, this ignores the fact that the final, selected, model is random and a function of the data. The randomness contained in the selection step means the post-selection estimators do not converge uniformly to a normal distribution, as the potential omitted variable bias from omitting (weakly) relevant variables in the selection step is too large to maintain uniformly valid inference.

In a sequence of papers leeb2005model, Leeb and P\"otscher address these issues, showing that distributions of post-selection estimators only converge point-wise but not uniformly in the parameter space to normal distributions. Therefore, “standard” asymptotics fail to deliver a proper approximation of finite-sample behavior due to the presence of small, hard to detect parameters, whose omitted variable bias is too large to ignore asymptotically. As such, post-selection based on oracle properties is only appropriate if one a priori rules out small parameters conditions (via beta-min conditions, see e.g. van2011adaptive) thus obtaining a sharp separation of non-zero from zero coefficients. This is typically far too strong to be reasonable in applications, and methods explicitly accounting for selection are required.

Several approaches to valid post-selection inference, also referred to as honest inference, have been developed in recent years based on various philosophies, such as simultaneous inference across models berk2013valid, inference conditional on selected models lee2016exact, or debiasing (desparsifying) the lasso estimates van2014asymptotically,zhang2014confidence. We focus on the double selection approach developed by Belloni, Chernozhukov and co-authors; see e.g. belloni2014high for an overview. This approach is tailored for the lasso, easy to implement, and can be extended to dependent data.

belloni2014uniform develop a post-double-selection approach to construct uniform inference for treatment effects in partially linear models with high-dimensional controls using the lasso. Two initial lasso estimations of both the outcome and the treatment variable on all the controls are performed, and a final post-selection least squares estimation is conducted of the outcome variable on the treatment variable and all the controls selected in at least one of the two steps. The double variable selection step substantially diminishes the omitted variable bias and ensures the errors of the final model are (close enough to) orthogonal with respect to the treatment. The authors proved uniform validity of the procedure under a wide range of DGPs, including heteroskedastic and non-Gaussian errors.

chernozhukov2019lasso extend the analysis of estimation and inference for highly-dimensional systems in regressions, allowing for (weak) temporal and cross-sectional dependency. Regularization techniques for dimensionality reduction are applied iteratively in the system and the overall penalty is jointly chosen by a block multiplier bootstrap procedure. Oracle properties and bootstrap consistency of the test procedure are derived. Furthermore, simultaneous valid inference is obtained via algorithms employing least square or least absolute deviation after (double) lasso selection step(s). Although our approach is closely related to that of chernozhukov2019lasso, it differs in a number of ways. Our method is simpler and faster to implement as it does not rely on bootstrap methods. Also, chernozhukov2019lasso focus on general systems of equations and general ways of performing inference, which is different from our specific focus on Granger causality and VAR models. Third, we consider a different set of assumptions to establish the validity of our method, where we specifically focus on the relevance of these assumptions for applications in financial econometrics.

High-dimensional Granger causality test

We here describe how to implement the post-double-selection procedure in a VAR context. Let $\bm x_{GC,j}$, $j=1,\ldots, N_X$, where $N_X = p N_J$, denote the $j$-th column of $\bm X_{GC}$ and consider the partial regressions:

align[align omitted — 201 chars of source]

where $\bm \gamma_j, \; j = 0,\ldots, N_X$, are the best linear prediction coefficients\footnote{Note that Assumption (ref)((ref)) implies that $\left(\mathbb{E} \bm x_{-GC,t-1} \bm x_{-GC,t-1} \right)^{-1}$ and hence $\left(\mathbb{E} \bm X_{-GC,t-1}^{\otimes} \bm X_{-GC,t-1}^{\otimes \prime} \right)^{-1}$ exist.}

equation*[equation* omitted — 624 chars of source]

where $\bm X_{-GC,t}^{\otimes} = \bm I_{N_I} \otimes \bm x_{-GC,t-1}$. As the errors $\bm e_0, \ldots, \bm e_{N_X}$ are orthogonal to $\bm X_{-GC}$, partialling out the effects of these variables would allow for a valid test of Granger causality. Of course, (ref) and (ref) are still high-dimensional and cannot be estimated by least squares. However, we can select the relevant variables from lasso estimation of (ref) and (ref) and collect all these for the final estimation of $\bm y_{I}$ on $\bm X_{GC}^{\otimes}$ plus only those relevant variables.

Intuitively, this works because to cause omitted variable bias on the coefficients of $\bm X_{GC}$, a particular variable in $\bm X_{-GC}$ must have a nonzero coefficient in both (ref) and one of the regressions in (ref). If its coefficient is zero in (ref), it has no effect on $\bm y_{I}$ and is therefore not wrongfully omitted. If it has a zero coefficient in all regressions in (ref), it is not correlated with any variables of interest, and omitting it will not result in a bias. By including all variables that are selected in at least a single of these regressions, we essentially allow for “one free mistake” by the lasso in failing to select a relevant variable. That is, omitted variable bias will only occur if the lasso fails to select a relevant variable in both regressions simultaneously. As the probability of this occurring decreases quadratically, this is sufficient to be negligible asymptotically and allow for uniformly valid inference. We provide a formal justification in Section (ref).

We now state the details of our algorithm which executes the post-double-section along the lines described above, and conclude this section with some remarks.

algorithm[algorithm omitted — 2,874 chars of source]
remarkWe perform the initial regressions in terms of $\bm X_{GC}$ amd $\bm X_{-GC}$ instead of $\bm X_{GC}^{\otimes}$ and $\bm X_{-GC}^{\otimes}$. The two are equivalent, as the Kronecker product essentially just copies the columns of $\bm X$ both in the dependent and explanatory variables. Running the initial regressions in terms of $\bm X^{\otimes}$ therefore essentially means running the same regression $N_I$ times, which is pointless as the selected variables remain the same in terms of the columns of $\bm X_{-GC}$. We therefore perform the regressions just once for each column in $\bm X_{GC}$. The construction of $\hat{S}_X^{\otimes}$ ensures that for any selected column $\bm x_{-GC,m}$, we select every column of $\bm X_{-GC}^{\otimes}$ in which $\bm x_{-GC,m}$ appears.
remarkThe feasible generalized least squares (FGLS) estimation in Step [2] is needed when $N_I > 1$ to account for the correlation between equations of the VAR, and the fact we do not have the same selected regressors in each equation, as those coming from (ref) differ. Note that if $N_I = 1$, FGLS estimation collapses to the familiar form of the LM statistic. In that case one regresses $\hat{\bm \xi}$ by OLS onto the variables retained by the previous regularization steps plus the Granger causality variables, and retain the residuals $\hat{\bm \nu} = \hat{\bm \xi} - \bm X_{\hat{S}\; \cup \;GC}^{\otimes} \hat{\bm \beta}^*$, obtaining $R^2 = 1 - \hat{\bm \nu}^\prime \hat{\bm \nu} / \hat{\bm \xi}^\prime \hat{\bm \xi}$.
remarkOur lasso estimation of an HD-VAR can be interpreted as a general, data-driven, approach to Granger causality testing which encompasses the theory-driven `standard' approach in low-dimensional VARs. In particular, the lasso can be interpreted as imposing (approximate) sparsity over a high-dimensional information set, with the extent and location of the sparsity, or irrelevance, determined in a data-driven way. Conversely, testing Granger causality in a low-dimensional setting can then be interpreted as a priori assuming an extreme degree of sparsity over the same information set; in other words, it amounts to assuming that none of the additional series are relevant.
remarkGiven that we essentially have $N_{GC} = N_J \times N_I \times p$ steps of selection, it would be more appropriate to refer to our method as “post-$N_{GC}$-selection” approach. For expositional simplicity however we stick to the post-double-selection name, as this is the common name for such a procedure, and conveys the essence of our method equally well.
remarkAlthough the lasso regressions can handle increasing $N_J$, $N_I$ or $p$ with any issues, inference becomes more complicated when $N_{GC}$ increases with the sample size as the proposed LM statistic (or similarly a Wald test) will not have a limit distribution anymore. In such a case one could use recently developed Gaussian approximations of maxima of high-dimensional vectors chernozhukov2013gaussian,zhang2017gaussian to base a test statistic on $\max_{m = 1, \ldots, N_{GC}} {\left\lvert\hat{\beta}_{GC,m}^{\textsc{pds}}\right\rvert}$, where $\hat{\bm \beta}_{GC}^{\textsc{pds}}$ are the coefficients of $\bm X_{GC}^\otimes$ in a regression of $\bm y_{I}$ on $\bm X_{GC}$ and $\bm X_{\hat{S}^\otimes}^{\otimes}$ as in (ref). However, the critical values of this test statistic have to be simulated, which complicates the testing. As we argued in Section (ref) that a fixed $N_{GC}$ is a reasonable assumption for typical Granger causality applications, we do not pursue this route.
remarkIn Step [1] we propose not to consider the GC variables in the first regularization and insert them back at Step [2]. Alternatively, the GC variable(s) can be left in the regression, such that, we regress on the full $\bm X^\otimes$ matrix. In this case there are then two further possibilities by either penalizing these variables or not. Simulations for these two alternatives have been carried out and in practice we do not find significant differences among the three in terms of size and power. The approach proposed in Step 1 delivers the best results in terms of size.
remarkWhen the time series length is of same magnitude as the number of covariates, information criteria and time series cross-validation tend to break down and select too many covariates in order to perform a post-selection by OLS. To overcome this issue we propose to place a lower bound on the penalty to ensure that in each selection regression at most $c\, T N_I$ variables are selected, for some $0 < c < 1$. In our simulation and empirical studies we set $c=0.5$. Note that, as we have $N_{GC}$ selection steps, the possibility remains that different variables are selected in each steps, making the number of variables in the union $\hat{s}$ still too large to perform the post-selection OLS, although this problem is likely to occur far less often. This can be addressed by ensuring that fewer than $N_I T/ N_{GC} = T / N_X$ variables are selected in each selection step. We do not impose this stricter bound in general, as it will often be much too strict. Instead, we recommend to only address this issue if it arises in practice by an ad-hoc increase of the lower bound on the penalty.\footnote{Although it happens less often, the theoretical plug-in method for the tuning parameter occasionally also selects too many variables to make the post-OLS estimation infeasible. However, for this method no easy solution is available for bounding the penalty. One could increase the constant in the plug-in expression, thus strengthening the penalty, but this would be a rather ad-hoc adjustment. In particular, imposing the lower bound for the other methods only limits the allowed range of the tuning parameter, forcing the minimization to choose another (local) minimum that can still be far away from the boundary and justified graphically. For the plug-in method it is however difficult to justify the right amount of the increase, as the tuning parameter will be fixed to that value, and thus the chosen increase is rather arbitrary.}
remarkAlthough our Granger causality test has a $\chi^2$ distribution under the null hypothesis asymptotically, in smaller samples the test might still suffer from the usual small-sample approximation error. As such we propose a finite-sample correction to the test in Step [3b], which in our simulation studies improved the size of our test.
remarkInstead of the (adaptive) lasso, other estimators can be used in Step [1] as long as they deliver a sparse coefficient vector. For instance, the elastic net of zou2005regularization that adds an $\ell_2$-penalty in addition to the $\ell_1$-penalty of the lasso can be used. The additional penalty ensures that the elastic net is strictly convex, and as a consequence tends to select highly correlated variables as a group together, whereas the lasso would tend to select only one of these variables zou2005regularization. Given the typically strong correlations between many economic variables, this appears particularly useful for our context. However, we used the elastic net for both the simulations and the empirical application, and in both cases we found that the results are widely comparable to those of lasso. Therefore we chose to omit them from the paper; they are available upon request.
remarkOne can also perform a standard Wald test of Granger causality instead of the LM test, by regressing the variables of interest on $\bm X_{GC}^{\otimes}$ and $\bm X_{\hat{S}^\otimes}^\otimes$, and testing for the significance of the coefficients of $\bm X_{GC}^\otimes$. While asymptotically the LM and Wald tests behave equally, differences might arise in small samples. We investigated the Wald version of the test in simulations as well, with results reported in Appendix (ref), Table (ref). In general, differences between the two methods are negligible. However, for the Wald test, occasionally we run into the problem described in Remark (ref), where even with the imposed lower bound on the penalty, too many variables are selected for performing a post-selection OLS. For this reason we prefer the LM version.

Tuning parameter selection

Appropriate selection of the lasso tuning parameter $\lambda$ in (ref) is crucial to achieve good performance. Many different data-driven methods exist giving wildly varying results. We provide a systematic comparison of several popular methods discussed in the literature in a simulation study. To the best of our knowledge, this is the first such comparison in the context of post-selection inference. We now introduce the methods considered in our study.

One option is to minimize an information criterion (IC) to determine an appropriate data-driven $\lambda$. Let $\hat{S}(\lambda) = \left\{m \in \{1,\ldots,Kp\}: {\left\lvert\hat{\beta}_m (\lambda)\right\rvert} > 0 \right\}$ denote the set of active variables in the lasso solution for a given $\lambda$. For a generic response vector $\bm y$ and predictor matrix $\bm X$, the value $\lambda^{IC}$ is found as

equation*[equation* omitted — 225 chars of source]

where $C_T$ is the penalty specific to each criterion. We consider the Akaike information criterion (AIC) by akaike1974new with $C_T=2$, the Bayesian information criterion (BIC) by schwarz1978estimating with $C_T=\ln(T)$, and the Extended Bayesian information criterion (EBIC) by chen2008extended with $C_T = \ln(T) + 2 \gamma \ln(Kp)$ with $\gamma=0.5$ proposed by chen2012extended who argue that BIC fails to select the correct variables when the number of parameters is larger than the sample size.

An alternative approach is to plug in estimates of theoretically optimal values bickel2009simultaneous,belloni2013least,belloni2011square. The lasso requires that $\lambda\geq c{\left\lVert\bm X'\bm u\right\rVert}_{\infty}/T$ for some constant $c>0$ with “high probability”. The central limit theorem motivates a Gaussian approximation where one chooses $\lambda^{th}=\frac{2c\hat{\sigma}}{\sqrt{T}}\Phi^{-1}\bigg(1-\frac{\alpha}{2N}\bigg)$ for a small $\alpha=o(1)$, where $\Phi^{-1}(\cdot)$ is the inverse of standard Gaussian cumulative distribution function and $\hat{\sigma}$ is an estimate the variance of $\bm u$. In this paper we set $\alpha =0.05/\ln(T) $ and $c = 0.5$, while we follow belloni2012sparse in the estimation of $\sigma$. Specifically, we obtain an initial (conservative) estimate by least squares estimation of $\bm y$ on the five most correlated regressors. This estimate is then updated iteratively, for details see belloni2012sparse.

Perhaps the most popular way to choose the tuning parameter is cross-validation (CV), although CV is not always appropriate in the time series setup without modifications bergmeir2018note. To estimate the tuning parameter with CV in a time series setup (TSCV) we use an expanding window out-of-sample forecasting scheme and minimize its squared forecasting error. The rolling window is set up with $80\%$ of the sample for training and $20\%$ for testing. Cross-validation is appealing since it does not require any plug-in estimates, however, as observed in chetverikov2020cross it typically yields small values of $\lambda$ thus still gaining fast convergence rate but at the price of less variable selection.

remarkAlthough we assume $p$ fixed, in practice it may still need to be estimated if no reasonable value (or upper bound) can be given. As $p$ determines the number of selection regressions to be conducted, it has to be determined a priori and cannot be integrated in the lasso estimation. It can still be determined though by a (separate) lasso-type algorithm. For example, one may estimate (ref) with a large initial lag length $p^*$, and let $p$ be determined as the largest lag for which variables are selected, possibly also varying the lag length over variables. For this approach the hierarchical penalties of nicholson2020high provide a better option than the regular lasso, as the regular lasso tends to select occasional “spurious” high lags, which would have a significant impact on the testing procedure. Alternatively one may marginalize the VAR to a collection of univariate AR($p$) processes, and select the lag length by minimizing an information criterion on the residual covariance matrix. As marginalization increases the lag length, such an approach would yield a simple to compute upper bound on $p$.

Asymptotic Properties

In this section we derive the asymptotic properties of our method. We first present and discuss our general high-level assumption under which the properties are derived, and then state our main results.

assumptionLet $\delta_T$ and $\Delta_T$ denote sequences such $\delta_T, \Delta_T \rightarrow 0$ as $T \rightarrow \infty$. Then assume that the following conditions are satisfied: \begin{enumerate}[(a)] • Population Eigenvalues: Let $\bm e_t = \left(e_{1,t}, \ldots, e_{N_X,t} \right)^\prime$, $\bm E = \left(\bm e_1, \ldots, \bm e_T \right)^\prime$ and $\bm E^{\otimes} = \bm I_{N_I} \otimes \bm E$. Define \begin{align*} \bm \varSigma &= \begin{bmatrix} \bm \varSigma_{GC,GC} & \bm \varSigma_{GC, -GC} \\ \bm \varSigma_{-GC,GC} & \bm \varSigma_{-GC, -GC} \\ \end{bmatrix} = \begin{bmatrix} \mathbb{E} \left(\bm x_{GC,t} \bm x_{GC,t}^\prime \right) & \mathbb{E} \left(\bm x_{GC,t} \bm x_{-GC,t}^\prime \right) \\ \mathbb{E} \left(\bm x_{-GC,t} \bm x_{GC,t}^\prime \right) & \mathbb{E} \left(\bm x_{-GC,t} \bm x_{-GC,t}^\prime \right) \\ \end{bmatrix} \end{align*} Then there exists a constant $c_L > 0$ not depending on $T$ and $k$ such that $\lambda_{\min} (\bm \varSigma) > c_L$, where $\lambda_{\min}(\bm \varSigma)$ denotes the minimum eigenvalue of $\bm \varSigma$. • Limit Behavior: Let \begin{equation*} \begin{split} \bm E^{\otimes\prime} \bm u_{I} / \sqrt{T} &= \operatorname{vec}(\bm E^\prime \bm U_{I})/\sqrt{T} = \frac{1}{\sqrt{T}} \sum_{t=p+1}^T \operatorname{vec}(\bm e_t \bm u_{I,t}^\prime) \xrightarrow{d} N(0, \bm \varOmega),\\ \bm E^\prime \bm E / T &= \frac{1}{T} \sum_{t=p+1}^T \bm e_t \bm e_t^\prime \xrightarrow{p} \bm \varSigma_{GC|-GC} = \bm \varSigma_{GC,GC} - \bm \varSigma_{GC,-GC} \bm \varSigma_{-GC,-GC}^{-1} \bm \varSigma_{-GC,GC},\\ \bm U_{I}^\prime \bm U_{I} / T &\xrightarrow{p} \bm \varSigma_{u,I}, \end{split} \end{equation*} where $\bm \varOmega = \operatorname*{plim}_{T\rightarrow\infty} \left(\bm E^{\otimes\prime} \bm u_{I} \bm u_{I}^\prime \bm E^{\otimes} \right)/T$. • Empirical Process: We have with probability at least $1 - \Delta_T$ that ${\left\lVert\bm X_{-GC}^{\prime} \bm u_i/\sqrt{T}\right\rVert}_{\infty} \leq \gamma_T$ for all $i \in I$ and ${\left\lVert\bm X_{-GC}^\prime \bm e_j/\sqrt{T}\right\rVert}_{\infty} \leq \gamma_T$ for all $j = 1, \ldots, N_X$, with $\bm e_j$ the $j$-th column of $\bm E$, for some deterministic sequence $\gamma_T$ subject to the restrictions in ((ref)). • Boundedness: The (Granger causality) parameters of interest are bounded, that is, there exists a fixed constant $C>0$ such that ${\left\lVert\bm \beta_{GC}\right\rVert}_1 \leq C$. • Consistency: The initial estimators $\hat{\bm \gamma}_{j}$ are consistent in the prediction sense; specifically, with probability at least $1 - \Delta_T$ we have that \begin{equation*} \begin{split} &{\left\lVert\bm X_{-GC}^{\otimes} (\hat{\bm \gamma}_0 - \bm \gamma_0)\right\rVert}_2 / \sqrt{T} \leq \delta_T T^{-1/4}, &\max_{j = 1, \ldots, N_X} {\left\lVert\bm X_{-GC} (\hat{\bm \gamma}_j - \bm \gamma_j)\right\rVert}_2 / \sqrt{T} \leq \delta_T T^{-1/4}. \end{split} \end{equation*} • Sparsity: Let $S_j = \{m: \gamma_{m,j} \neq 0\}$ denote the sets of active variables in (ref) and (ref) and let $s = {\left\lvertS_0\right\rvert} + \sum_{j=1}^{N_X} {\left\lvertS_j\right\rvert}$. Let $\hat{s}$ be as defined in Algorithm 1. Then both the DGP and the initial estimators are sufficiently sparse; in particular, we have that with probability at least $1 - \Delta_T$, $\max(s, \hat{s}) \leq \bar{s}_T$ for a deterministic sequence $\bar{s}_T$ subject to the restrictions in ((ref)). • \textbf{Sparse Eigenvalues:} for any $\bm \gamma \in \mathbb{R}^{(K - N_J)p}$ with ${\left\lVert\bm \gamma\right\rVert}_0 \leq \bar{s}_T$, we have with probability at least $1 - \Delta_T$ that ${\left\lVert\bm \gamma \right\rVert}_2^2 \leq {\left\lVert\bm X_{-GC} \bm \gamma/\sqrt{T}\right\rVert}_2^2 / \phi_{T,\min}^2$, where $\phi_{T,\min} > 0$ is subject to the restrictions in ((ref)). • \textbf{Rate Conditions:} The deterministic sequences $\bar{s}_T, \gamma_T$ and $\phi_{T,\min}$ introduced above satisfy the restriction $\bar{s}_T \gamma_T / \phi_{T,\min} \leq \delta_T T^{1/4}$. \end{enumerate}

Assumption (ref) is a high-level assumption that allows for much flexibility on the underlying DGP and the used estimators in the first step. We now discuss each part in turn. Part (a) assumes that the minimum eigenvalue of $\bm \varSigma$ is bounded. This is required for application of lasso methods, as well as for the inverse covariance matrix $\bm \varSigma^{-1}$ and the projection coefficients in (ref) and (ref) to exist. Part ((ref)) assumes that a central limit theorem and weak law of large numbers hold. Essentially this require that the process is sufficiently well-behaved in terms of moments and dependence allowed. Although for convenience we assume martingale difference errors in Assumption (ref), ((ref)) holds under much weaker conditions such as mixing errors; see e.g. davidson1994stochastic.

Part ((ref)) is closely related to ((ref)), but additionally controls the tail behavior of the empirical process. Results of this kind are standard in the lasso literature and can be derived using a variety of tail bounds depending on the properties of the random variables of interest, see e.g. kock2015oracle and medeiros2016l1 for results relevant to VAR and time series models. Of particular interest for financial applications, masini2019regularized show that this condition is satisfied for VAR models with general weakly dependent erorrs that include many popular multivariate volatility models. The boundedness assumption in (ref) is not very restrictive, and with $N_{GC}$ fixed follows directly if the parameter space of $\bm \beta$ is a compact set.

Part ((ref)) imposes an appropriate consistency rate on the predictions coming from the first-stage estimator. Such prediction consistency is a standard result for lasso estimators; in particular, wong2020lasso obtain it for a very general class of VAR models allowing for conditional heteroskedasticity and dependence in the error terms. adamek2020lasso derive consistency of the lasso under misspecified time series models, and show that their setting covers (among others) the first-step regressions of the relevant predictors in $\bm X_{GC}$ on the other regressions, which are inherently misspecified in a VAR setup due to the missing lags; see their Remark 3 for further details.

Next to consistency, we also require sparsity of the DGP and the estimator, as controlled by part ((ref)). The assumption of exact sparsity in the DGP for the initial regressions can be relaxed to approximate sparsity as in belloni2014inference. For the sake of expositional clarity we do not work under that assumption here but stick to the simpler exact sparsity. Sparsity of the first-stage estimator is needed in our framework as we perform OLS on the selected variables from the first-stage regressions. If the selected variables are not sparse enough, too many variables will be selected for OLS to be feasible. Sparsity of lasso estimators is analysed in belloni2013least, while kock2015oracle and medeiros2016l1 provide results for adaptive lasso for time series. Importantly, we do not require consistent model selection; the selection method used is allowed to make “persistent” mistakes, allowing for both variables to be incorrectly included and relevant variables to be missed, as long as the estimator remains sufficiently sparse and consistency is guaranteed. Unlike belloni2014inference, we allow for the order of sparsity of the estimator to differ from the true sparsity thereby opening the way for conservative selection procedures.

Given the assumptions above, the eigenvalue assumption in ((ref)) becomes almost superfluous, as it is generally needed to establish ((ref)) and ((ref)) for lasso-type estimators; see e.g. belloni2013least and medeiros2016l1 for details. It requires that for sufficiently sparse vectors, the eigenvalues of the subset of the Gram matrix corresponding to their non-zero support do not decrease to zero too fast. Such assumptions are standard in the lasso literature in various guises as restricted eigenvalue conditions, and can typically be derived by making similar conditions on the population covariance matrix $\bm \varSigma_{-GC, -GC}$ coupled with a convergence result of the Gram matrix $\bm X_{-GC}^\prime \bm X_{GC}$ to $\bm \varSigma_{-GC, -GC}$. basu2015network, masini2019regularized and wong2020lasso establish the plausibility of such restricted eigenvalue conditions for various VAR models. We state the condition here explicitly as it is needed directly in the proofs.

Finally, note that the restrictions on tail behavior (via $\gamma_T$), sparsity (via $\bar{s}_T$) and minimum eigenvalues (via $\phi_{T,\min}$) are meaningless if no rates on these sequences are imposed. Part ((ref)) therefore is the key part which connects all assumptions with explicit rates needed for the validity of the PDS method. The restrictions here represents a trade-off between sparsity, thickness of tails and minimum eigenvalues. For example, if, as often assumed $\phi_{T,\min}$ is fixed and $\bm u_t$ is Gaussian, tails are sufficiently thin that $\gamma_T$ can be chosen as roughly the order of $\sqrt{\ln(K^2 p)}$ kock2015oracle, leaving room for either almost exponentially large $K$ relative to $T$, or a fairly non-sparse model. On the other hand, if only $m$ moments of $\bm u_t$ exist, $\gamma_t$ should be taken roughly of the order $(K^2 p)^{2/m}$ masini2019regularized, requiring polynomial growth of $K$ compared to $T$ and sparser models.

The most restrictive and crucial assumption needed on the underlying DGP for satisfying Assumption (ref) is the sparsity of the underlying DGP formulated in part ((ref)). The plausibility of this assumption highly depends on the specific application. In many financial applications sparsity (or its approximate version) is natural, for example in portfolio selection when the number of assets is large and the estimation of high-dimensional volatility matrices in financial risk assessment (see fan2011sparse for an overview), as well as in our investigation of Granger causality in networks of realized volatilities in Section (ref). The volatility of one particular stock is likely to have specific channels of contagion rather than affecting the whole stock market at the same time. Shocks to one asset therefore likely propagate through the system via specific channels, which corresponds to sparse lag polynomials. One might worry about systemic shocks affecting many assets; however, the dense covariance matrix $\bm \varSigma_u$ can accommodate simultaneous common shocks. Moreover, the dynamic of such shocks can generally well be captured through a sparse combination of the most important and most affected assets. Similarly, in macroeconomic applications it has been found that a few important variables can capture the effects of unobserved common factors, leading sparse models to perform as well as common factors demol2008forecasting,smeekes2018macroeconomic.

We are now ready to state our main asymptotic result of this section in Theorem (ref) which establishes the asymptotic normality of the post-lasso (generalized) least squares estimator. Here we slightly deviate from the LM test in Algorithm (ref); after the double selection procedure carried out in Step [1], we regress the transformed outcome variables $\widetilde \bm y_t = (\bm G_T \otimes \bm I_T) \bm y_{I}$ on both the Granger causing $\widetilde\bm X_{GC}^\otimes = (\bm G_T \otimes \bm I_T) \bm X_{GC}^\otimes$ and selected variables $\widetilde \bm X_{\hat{S}^\otimes}^\otimes = (\bm G_T \otimes \bm I_T) \bm X_{\hat{S}^{\otimes}}^{\otimes} $

equation[equation omitted — 231 chars of source]

The transformation by the matrix $\bm G_T$ allows for the GLS esitmation needed in the LM procedure by taking $\bm G_T = \hat\bm \varSigma_{u,I}^{-1/2}$, while OLS is performed with $\bm G_T = \bm I_{N_I}$. In the latter case the theorem provides the foundation for the Wald test discussed in Remark (ref) (minus the required variance estimation for that test). We state this result separately as it is interesting in its own right, and can be used to establish validity of other tests such as the Wald test.

theoremLet $\hat{\bm \beta}_{GC}^{\textsc{pds}}$ denote the OLS estimator of $\bm \beta_{GC}^{\textsc{pds}}$ in (ref). Let $\bm G_T$ be any matrix satisfying with probability at least $1-\Delta_T$ that $0 < c_1 \leq \lambda_{\min}(\bm G_T^\prime \bm G_T) \leq {\left\lVert\bm G_T^\prime\bm G_T\right\rVert}_{\max} \leq c_2 < \infty$, where $c_1, c_2$ are constants not depending on $T$. Then, uniformly in all DGPs that satisfy Assumption (ref), we have as $T \rightarrow \infty$, \begin{equation*} \sqrt{T} (\hat{\bm \beta}_{GC}^{pds} - \bm \beta_{GC})\overset{d}{\to}\mathcal{N}\left(\bm 0, (\bm G^\prime \bm G \otimes \bm \varSigma_{GC|-GC})^{-1} \bm \varOmega_{\bm G} (\bm G^\prime \bm G \otimes \bm \varSigma_{GC|-GC})^{-1} \right), \end{equation*} where $\bm \varOmega_{\bm G} = \operatorname*{plim}_{T\rightarrow\infty} \left[(\bm G_T^\prime \bm G_T \otimes \bm E^\prime) \bm u_{I} \bm u_{I}^\prime (\bm G_T^\prime \bm G_T \otimes \bm E) \right]/T$.

Theorem (ref) establishes the asymptotic normality of the post-double-selection OLS estimators. The statement `uniformly in all DGPs that satisfy Assumption (ref)' should be interpreted as the theorem holding uniformly over a parameter space that is defined such that Assumption (ref) holds for all parameters in that parameter space. Importantly, no beta-min conditions on the smallest magnitude of parameters are required, thus alleviating the post-selection inference problem. We refer to Comments 3.4 and 3.5 in belloni2014inference for further details regarding the uniformity. The limit distribution of the LM test now follows straightforwardly from Theorem (ref), and is stated in the corollary below.

theoremLet $\bm \beta_{GC} = \bm 0$. Then, uniformly in all DGPs that satisfy Assumption (ref) and for which $\bm \varOmega = \bm \varSigma_{u, I} \otimes \bm \varSigma_{GC|-GC}$, we have that \begin{align*} &LM \xrightarrow{d} \chi_{N_{GC}}^2 \qquad as T \rightarrow \infty. \end{align*}

Theorem (ref) establishes the limiting distribution of the PDS-LM test under an additional condition on the (co)variances of the partial regression errors, which is satisfied if the errors are iid. To allow for heteroskedaticity the LM test has to be modified, which would only lead to more cumbersome proofs without adding any novelty specific to the high-dimensional case. Therefore we focus on the homoskedastic case here, although we do consider a heteroskedasticity-robust version of the test in the volatility application in Section (ref).\footnote{Note that this is no different for the Wald test, for which the variance estimation has to be adjusted as well.}

Monte-Carlo Simulations

We now evaluate the finite-sample performance of our proposed Granger causality test. We consider three Data Generating Processes (DGPs) inspired by kock2015oracle:

align*[align* omitted — 712 chars of source]

The diagonal VAR in DGP1 respects the sparsity assumption while in DGP2 the entries are set to decrease exponentially fast in the distance from the main diagonal and hence the sparsity assumption is not met. DGP2 could be empirically motivated by looking e.g. at financial interconnectedness. Financial institution, such as banks, lend to and borrow from one another becoming interconnected through interbank credit exposures. The financial distress experienced by one bank is likely to be most heavily transmitted the closer the connections are as well as less transmitted, the weaker the connections. DGP3 is a block-diagonal system. Such a structure is motivated by e.g. typical quarterly macroeconomic models capturing business cycle dynamic and monetary and fiscal policy effects. One such example is DSGE models, where the dynamic of the economy through time is monitored on quarterly frequency. Note that as written above, DGP1 satisfies the null of no Granger causality from unit 2 to 1, while DGP2 and DGP3 do not. Therefore, we adapt DGP 1 for the power analysis by setting the coefficient in position $(2,1)$ equal to 0.2. Conversely, we set the same coefficient equal to zero for DGP2 and DGP3 for the size analysis.

We choose our series of interest as $I=\{2\}$ and $J = \{1\}$, therby focusing on the case where we have single variables of interest for both elements of the test. Here we consider for simplicity $p=1$ lag, namely the same lag-length as in the DGPs, so $j=1$. The equation of interest can then be written as

equation*[equation* omitted — 100 chars of source]

Hence, for each DGP we test $H_0:\beta_{GC}=0$ against $H_1:\beta_{GC}\neq 0$ using our proposed PDS-LM test.

Table (ref) reports the size and power of the test for 1000 replications by using different combinations of time series length $T=(50,100,200,500)$ and number of variables in the system $K=(10,20,50,100)$ and a fixed lag-length $p=1$. All the rejection frequencies are reported using a burn-in period of fifty observations. For each scenario, AIC, BIC and EBIC are compared with the theoretical choice of the tuning parameter $\lambda^{th}$ and time series cross validation $\lambda^{TSCV}$ as described in Subsection (ref).

Simulations are also reported for different types of covariance matrices of the error terms. We employ a Toepliz-version for calculating the covariance matrix as $\Sigma_{i,j}=\rho^{|i-j|}$ by using two scenarios of correlation: $\rho=(0,0.7)$. The first case corresponds to no correlation, and is equivalent to set $\bm \varSigma = \bm I_{K}$.

In the Appendix we provide some additional simulation results. First, Table (ref) reports the simulation results for all three DGPs using $\Sigma_{i,j}=0.7^{|i-j|}$. Second, we investigate the Wald version of our test in Table (ref). Third, in Table (ref) we investigate the effects of miss-specification of the lag length by estimating the over-specified VAR$(p+1)$ instead of the true-order VAR$(p)$.\footnote{For both the Wald test and the over-specified VAR$(p+1)$ we report the simulations for $\Sigma_{i,j}=0.7^{|i-j|}$ and DGP1 only. Results for the other DGPs are available upon request.} Fourth, in Table (ref) we report the results for the size of a bivariate Granger causality test for a non-sparse DGP when using a standard Wald ($F$) test. This test is obviously sensitive to omitted variable bias, and our goal is to demonstrate its effect. Finally, although all results reported here use the finite sample correction in Step 3b of the algorithm, we also investigated the differences with Step 3a. We comment on these results in the next subsection. All results not reported in this paper are available from the authors upon request.

sidewaystable\begin{threeparttable} \caption{Simulation results for the PDS-LM Granger causality test} \tiny \begin{tabular}{ l S[table-format=3.0] cccc c c c c c c cc cccccccccccc} \toprule DGP&{Size/Power}&$\rho$&$T$ &&&50&&&&&100&&&&&200&&&&&500&&& \\ \cmidrule(l){1-24} &&&K& \multicolumn{5}{c}{AIC \;\;\; BIC \;\;\;EBIC \;\;\;\;$\lambda^{th}$ \; $\lambda^{TSCV}$} & \multicolumn{5}{c}{AIC \;\;\; BIC \;\;\;EBIC \;\;\;\;$\lambda^{th}$ \; $\lambda^{TSCV}$}& \multicolumn{5}{c}{AIC \;\;\; BIC \;\;\;EBIC \;\;\;\;$\lambda^{th}$ \; $\lambda^{TSCV}$}& \multicolumn{5}{c}{AIC \;\;\; BIC \;\;\;EBIC \;\;\;\;$\lambda^{th}$ \; $\lambda^{TSCV}$}\\ \cmidrule(l){4-24} &&& 10 &6.7&6.1&5.3&6.8&6.9&6.7&6.1&5.9&6.5&6.4&5.5&4.5&4.5&4.7&5.4&4.3&4.1&4.3&4.2&4.3 \\ 1&{Size}&0&20 &7.3&6.0&6.4&5.9&6.2&6.0&4.7&4.7&5.5&5.8&4.0&4.7&5.4&4.3&4.5&4.3&4.1&4.3&4.1&4.3\\ &&& 50 &7.0&5.9&5.7&6.4&5.5&7.4&6.4&6.4&6.2&6.6&6.9&5.9&5.9&6.2&6.1&6.8&6.5&6.7&6.8&6.7\\ &&& 100&7.4&7.2&6.6&NA&6.8&7.2&4.9&4.9&5.9&5.4&6.5&4.7&4.9&5.0&5.2&5.8&3.9&4.0&5.0&4.1 \\ \cmidrule(l){1-24} &&&10 &30.0&31.2&33.3&30.4&30.9&58.3&58.9&60.7&58.5&57.4&89.1&89.3&89.6&89.2&89.5&99.9&99.9&99.9&99.9&99.9 \\ 1&{Power}&0&20 &22.8&26.4&30.2&25.8&24.2&52.8&55.1&57.4&53.9&54.5&85.6&88.0&89.0&86.6&86.1&99.9&100&100&99.9&99.9 \\ &&&50 &13.6&24.3&33.3&17.9&18.7&38.7&53.8&59.0&46.4&45.5&78.2&85.8&87.0&81.3&80.2&99.9&99.9&99.9&99.9&99.9\\ &&&100&10.4&20.3&32.1&NA&18.7&20.4&51.9&57.0&31.2&35.8&61.6&85.4&86.7&74.1&73.7&99.5&100&100&99.8&99.7\\ \cmidrule(l){1-24} &&&10 &6.2&5.5&6.1&6.1&5.9&5.0&4.9&5.4&5.0&5.0&4.7&4.6&4.9&4.6&5.2&3.9&3.9&3.9&3.7&3.7 \\ 2&{Size}&0&20 &7.8&5.5&6.0&5.8&6.5&5.4&4.6&4.7&4.8&5.3&5.9&5.9&5.4&5.7&6.0&4.7&3.8&4.3&4.6&4.7\\ &&&50 &7.3&5.9&4.7&6.0&7.1&7.9&6.3&6.5&7.3&6.5&7.3&6.1&6.0&6.6&6.6&6.3&6.6&6.4&6.1&6.0\\ &&&100&6.1&6.9&5.5&NA&6.6&6.8&5.1&6.0&5.3&4.8&5.4&5.5&5.8&4.6&4.9&6.1&4.3&5.0&4.9&4.5\\ \cmidrule(l){1-24} &&&10 &18.0&19.7&21.3&18.2&17.0&37.9&39.3&40.4&38.0&38.4&64.7&64.7&66.9&64.8&65.1&97.4&97.4&97.5&97.4&97.4 \\ 2&{Power}&0&20 &16.0&19.6&24.6&18.8&16.9&35.4&39.8&44.4&37.4&36.7&64.4&67.4&69.7&66.0&65.0&97.2&97.5&97.5&97.5&97.4\\ &&&50 &8.6&15.2&21.7&12.8&13.1&25.0&36.1&43.2&32.6&31.2&57.0&66.4&71.9&61.8&61.1&95.0&96.2&96.8&96.1&95.4\\ &&&100&9.2&14.1&25.1&NA&10.1&15.1&34.9&45.9&25.9&25.8&44.8&65.1&74.7&58.0&57.8&94.7&97.3&97.7&96.3&96.5\\ \cmidrule(l){1-24} &&&10 &5.2&5.0&5.6&5.7&4.6&5.6&4.9&5.7&5.7&6.2&4.0&4.1&6.1&4.1&4.4&4.1&4.0&3.8&3.9&4.1 \\ 3&{Size}&0&20 &4.2&5.1&5.7&5.6&4.9&4.3&4.3&7.2&4.2&3.9&5.2&5.6&9.4&4.7&4.6&4.7&4.4&4.5&4.5&4.8\\ &&&50 &7.5&6.3&7.4&6.9&6.6&6.4&7.0&9.5&5.9&5.5&6.9&6.9&12.4&6.0&6.6&4.9&5.3&6.6&5.3&5.4\\ &&&100&7.1&6.7&8.3&NA&7.0&6.2&5.7&8.3&5.7&6.0&4.7&6.2&10.7&4.2&4.7&4.4&5.1&6.5&5.1&4.7\\ \cmidrule(l){1-24} &&&10 &15.4&20.0&23.4&16.2&16.1&31.5&36.4&44.3&32.6&31.4&58.4&61.3&63.7&58.9&59.3&95.2&95.6&95.7&95.5&95.3 \\ 3&{Power}&0&20 &13.4&18.7&26.0&13.8&13.8&29.5&37.0&47.8&30.0&29.8&56.5&62.0&69.8&56.8&55.5&94.0&94.6&94.8&94.4&94.3\\ &&&50 &11.4&20.7&28.0&12.9&10.6&24.1&39.8&52.3&26.9&27.6&50.3&59.5&73.6&52.1&51.8&91.1&92.7&93.4&92.3&90.9\\ &&&100&7.9&18.7&26.9&NA&14.1&13.7&42.6&55.0&20.2&22.2&41.4&62.0&75.2&44.9&44.8&90.0&94.0&94.8&91.4&90.2\\ \cmidrule(l){1-24} \end{tabular} \begin{tablenotes} \scriptsize • Notes: Size and Power for the different DGPs described in Section (ref) are reported for 1000 replications. $T=(50,100,200,500)$ is the time series length, $K=(10,20,50,100)$ the number of variables in the system, the lag-length is fixed to $p=1$. $\rho$ indicates the correlation employed to simulate the time series with the Toeplitz covariance matrix. The different choices of the tuning parameter $\lambda$ are reported as: AIC, BIC, EBIC for information criteria, $\lambda^{th}$ for the theoretical plug-in and TSCV for time series cross-validation as explained in Section (ref). \end{tablenotes} \end{threeparttable}

Our proposed approach shows a good performance in terms of size and (unadjusted) power for all DGPs considered. Both for the setting of no correlation and high correlation of errors, sizes are in the vicinity of 5% and power is increasing with the sample size $T$.

Only moderate size distortion is visible in large systems for small samples (e.g. $K\ge 50,\; T=50$). As expected, the test procedure works remarkably well for the sparse DGP1 in high dimensions. However, size properties under the non-sparse DGP2 do not deviate much from its sparse counterpart, although for both DGP2 and DGP3 we do observe a slight deterioration of size when the dimension of the system increases.

Interestingly, the three different information criteria show substantially different behavior. EBIC, due to its very stringent nature, tends to perform well only in very large systems, while it is essentially equal to a bivariate Granger causality test in small systems. We have to add though that the good performance of AIC in particular is somewhat inflated by the imposed lower bound on the penalty; unreported simulations show that without the lower bound AIC performs significantly worse, often selecting too many variables rendering the post-OLS estimation infeasible. The one advantage of using EBIC as information criterion to tune $\lambda$ in the $K>>T$ settings when $T$ is small (e.g. $T= 50,100$) is the possibility to avoid the lower bound on the penalty. However, since this comes at a price of more size distortion, we recommend the use of BIC instead, along with the lower bound on the penalty. When comparing the different choices of the tuning parameter we can narrow down the best performing ones (in terms of size and power) to BIC and $\lambda^{th}$. However, in terms of computational time, estimating the tuning parameter using information criteria is considerably faster.

Comparing our test to the bivariate VAR in Table (ref), it is clear that our proposed PDS-LM is very robust to omitted variable bias, unlike the bivariate test, whose size distortions increase with both the sample size and the number of variables, with sizes of 45% observed for the sample sizes we consider in our application in Section (ref). There we will also further elaborate on this difference between our method and the bivariate test. Table (ref) shows that for sample sizes smaller than $T=500$, rarely the power exceeds 90%. However, one must keep in mind that the powers are not size-adjusted, and thus the high reported power of the low-dimensional test is an artefact of the huge size distortions rather than genuine power. It also seems unreasonable to expect that PDS-LM test has vey high power if $T$ is small; we are still considering large systems with many parameters to estimate, and there seems to be no way around this if one desires to test Granger causality in large systems with many (control) variables. In that sense we may fully expect the bivariate test to also have higher size-corrected power; yet with all its disadvantages and sensitivity to omitted variables this is not a good comparison. All in all, we believe our test still has sufficiently adequate power properties to be useful in practice.

The results of robustness to misspecification of the lag length order with $p=2$ instead of $p=1$ are reported in Table (ref) in Appendix (ref). As the size distortions across the range of considered DGPs are only marginally higher for large $K$ and $T$ comparatively small, the test appears to be quite robust to this misspecification. Again, BIC seems to be the best choice for tuning the penalty for all DGPs. Unreported simulations (available upon request) further show that the finite sample adjustment for the test performed in Step 3b of the algorithm is able to substantially reduce size distortions in small samples compared to the asymptotic version of Step 3a.

Networks in Realized Volatilities

Realized Variances

We also compare it to a four-variate VAR (4-VAR) model

where we also consider the variation of nominal interest rates (3-months treasury bill) $\Delta(TB3MS)$ and variation of inflation $\Delta^2\log(CPI)$. Finally, we compare our method with the factor-augmented VAR (FAVAR) as introduced by bernanke2005measuring. Following mccracken2016fred, we estimate the static factors with PCA from the original (standardized) dataset, excluding the series of GDP and M1. The significant factors are selected by means of the $PC_p=\frac{\log(\min(K,T))}{\min(K,T)}$ criterion of bai2002determining. We obtain a total of eight factors ($F_1,\ldots,F_8$); the top three series most correlated with each factor are reported in the Appendix B. We then test for Granger non causality using the FAVAR model where we stacked the eight extracted factors along with M1 and GDP.

For 2-VAR, 4-VAR and FAVAR models we select the lag-length via AIC or BIC. We exclude EBIC since BIC is already parsimonious enough for selecting lags in low-dimensional VARs.

table[table omitted — 1,649 chars of source]

Table (ref) reports the results from the Granger causality tests. We find that our PDS-LM method provides strong evidence of Granger causality from real money to real output when the lasso is appropriately tuned with (E)BIC or the theoretical plug-in method. When AIC or time series cross-validation is used, too many variables are selected, resulting in a counter-intuitive high $p$-value.\footnote{In order to obtain the $p$-value for the PDS-LM test with AIC we had to adjust the penalty lower bound to $0.5(T)-4$.} Granger causality in the opposite direction is instead always clearly rejected.

The two small VAR models provide conflicting evidence on Granger causality from money to GDP, as well as the other direction. The 2-VAR shows a weak significance in the $M1\rightarrow GDP$ causal relation, yet the 4-VAR is very far from rejecting the null hypothesis, although the difference between the two VAR models might here be attributed to differences in lag length selection. In the opposite direction, the differences between the two small VAR models are even much larger, with the 2-VAR having $p$-values below 0.1 and the 4-VAR above 0.8, regardless of lag length selection.

The factor-augmented VAR appears to be very sensitive to the selected lag length. Adding a single lag from $p=1$ to $p=2$ decreases the $p$-value for Granger causality from M1 to GDP by nearly 70%, while in the other direction an increase is observed from highly significant to decisively not significant. It appears that, even though BIC selects one lag, this is insufficient to capture enough dynamics. We also observed that when manually increasing the lag length to four the $p$-values appear to remain stable, and are qualitatively similar as those of the appropriately tuned PDS-LM test. However, we have not investigated the estimation of the number of factors, which is notoriously difficult, instead going with the established choice of mccracken2016fred. It is likely that uncertainty about the number of factors will further increase the variability of the outcome of the test using the FAVAR.

To wrap up, our results show empirically that by increasing the information set by considering a high-dimensional VAR model, and thus allowing for the potential interaction of many other indicators, one is able to reduce the effect of the omitted variables and thereby retrieve a clearer picture of the causal relations between money and outcome. \fi

We first investigate the volatility transmission in stock return prices using the daily realized variances of 30 US assets. \footnote{We would like to thank Marcelo C. Meideiros for providing us with the high frequency data on stock prices that we have used to construct the realized variances. See Table (ref) for the stocks considered. The R package HDGCvar is available on the GitHub page of the corresponding author (\url{https://github.com/Marga8}).} Both the computational simplicity and the theoretical foundations make realized volatility measures (realized variance, bi-power variation, median realized variance, etc.) very attractive among practitioners and academics for modelling time varying volatilities and monitoring financial risk. We have considered 10-minute realized variances

equation[equation omitted — 116 chars of source]

using $j=1,\ldots, M$ intraday 10 minutes stock prices $P_{j,t}$. We consider 10 minute returns as this is the frequency that minimizes for our sample the microstructure noise (mcaleer2008realized).\footnote{To determine the optimal frequency, we computed realized variances using different frequencies of 1, 5, 10, 15, 30, 65 and 130 minutes, in addition to the estimation using daily returns. The latter estimation has the advantage of being unbiased but the drawback of being very noisy (pooter2008predicting). To find an optimal trade-off between bias and variance martens2004estimating), mean, variances and mean squared errors (MSE) were computed for each estimation frequency in a similar way as pooter2008predicting, and it was found that the frequency of 10 minutes minimizes the MSE.} We investigate the period from March 2008 until February 2017 (2236 trading days).

Given the time series of realized volatilities as defined in ((ref)), we employ a multivariate version of the heterogeneous autoregressive model (VHAR) of corsi2009simple to model their joint behavior (see also cubadda201911). To formally define the VHAR model, we log-transform the series and we stack the logarithmic RV into a vector $y_t$. The VHAR specification is given by the following model:

equation*[equation* omitted — 147 chars of source]

where $\bm y_{t}^{(week)} = \frac{1}{5} \sum_{j=0}^{4} \bm y_{t-j}$ and $\bm y_{t}^{(month)} = \frac{1}{22} \sum_{j=0}^{21} \bm y_{t-j}$ are the vectors containing the average volatility over the last 5 (week) and 22 (month) days. Granger causality in this context represents contagion, or spillover, of volatility from one asset to another. To test for the null hypothesis of no Granger causality / no volatility spillovers from $y_{k,t}$ to $y_{i,t}$ against the alternative of spillovers, we test

equation*[equation* omitted — 165 chars of source]

where $\beta_{i,k}^{(1)}$ is the $(i,k)$-th element of $B^{(1)}$. We perform this test for every $(i,k)$-pair to obtain the full $29\times 29$ network of spillover effects. As heteroskedasticity is likely present in these data, we robustify the PDS-LM procedure by implementing the heteroskedasticity-robust LM test such as for example described in wooldridge2015introductory. The full algorithm for the heteroskedasticity-robust PDS-LM test is given in Appendix (ref).\footnote{In the presence of heteroskedasticity, one might prefer the Wald version of the test, as this can be corrected in the standard way by using heteroskedasticty-robust standard errors. Empirically we found hardly any differences between the LM and Wald versions.}

We now report the results of our spillover tests for the volatility network. We use BIC to select the tuning parameter of the lasso, and perform the Granger causality tests with a 1% significance level.\footnote{We do not perform a correction for multiple testing, as this would only qualitatively affect our results. Moreover, our goal is not to identify exactly the set of spillovers, but to get a feeling of the relations between two variables at a time. As such, we believe a multiple testing correction is not needed, though it can be easily implemented.} Figure (ref) reports the transmission networks of volatilities estimated with the high-dimensional HVAR (PDS-LM HVAR), bivariate Granger causality tests (BiHVAR) for each pair of stocks, Granger causality tests from a full-system VAR (FullHVAR). The latter is feasible because of our large time series dimension with $T=2236$. For all methods we consider heteroskedasticity-robust variants.

The drawback of that log transform is that the interpretation of the original series, in our case the volatilities, is lost as the combinations involve nonlinear transforms of both realized variances and covariances. This is not compatible with the aim of this paper.

The second approach uses the log realized volatilities and the correlations separately, as done by for instance oh2016high. The underlying idea, following the DCC model of engle2002dynamic, is to decompose $ RC_{t}^{(d)}=D_{t}^{(d)}R_{t}^{(d)}D_{t}^{(d)}$ with $D_{t}^{(d)}$ a diagonal matrix with the square root of the individual realized variance and $R_{t}^{(d)}$ the realized correlation matrix. oh2016high use the HVAR model structure for each realized volatilities, they consequently assume no Granger causality across volatilities.

We propose something which is, to some extent, in between these two approaches. We look at two separate objects as in the DCC model, but stack the log of the realized variances $\bm y_{1t}^{(d)^{\prime }}$ and $z$-transforms $\bm y_{2t}^{(d)}=\mathrm{arc}\tanh \left( \mathrm{vech}(\bm R_{t}^{(d)})\right) $ of the realized correlations in a larger vector $\bm y_{t}^{(d)}=(\underset{1\times 30}{\underbrace{\bm y_{1t}^{(d)^{\prime }}}}, \underset{1\times 435}{\underbrace{\bm y_{2t}^{(d)^{\prime }}}})^{\prime }$ on which we estimate a VHAR of dimension 465.

In this HVAR each of the 465 equations depends on 1395 dynamic parameters plus the constant. We focus on the 30 equations corresponding to $\bm y_{1t}^{(d)}$ volatilities and consequently the bivariate causalities between these realized volatilities as in the previous section. Figure (ref) reports a total of 113 connections, which is about twice the connections in Figure (ref). In red we highlighted the 31 common connections with Figure (ref). Interestingly, adding more variables therefore allows us to uncover more relations. It seems that this allows us to uncover partial effects that were previously obscured by counteracting effects of the correlations. Importantly, the number of connections is still far less than compared to the BiHVAR in Figure (ref), and the PDS-LM HVAR is still able to deliver a clear picture of the causal connections when the system considered is high-dimensional. While the different connections found here obviously also lead to a different clustering, Figure (ref) shows that the clustering is quite similar, certainly regarding qualitative conclusions.

figure[figure omitted — 373 chars of source]

Conclusion

We propose an LM test in order to test for Granger causality in high-dimensional VAR models. We employ a post-double selection procedure using the lasso to select the set of relevant covariates in the system. The double selection step allows to substantially reduce the omitted variable bias and thereby allowing for valid post-selection inference on the parameters.

We provide an extensive simulation study to evaluate the performance of our method in finite samples, paying particular attention to the tuning of the penalty parameter. We compare different information criteria, time series cross-validation and a plug-in method based on theoretical arguments, and find that generally BIC and the theoretically tuned penalty perform best. However, to use information criteria in systems with a significantly larger number of variables than observations, a lower bound on the penalty parameter is needed to prevent too many variables being selected. The simulations also show that, when properly tuned, our proposed PDS-LM test attains good results both for size and power under different DGPs. Especially, it is shown to be robust both to non-sparse settings as well as to lag-length overspecification.

We also empirically investigate the usefulness of our method in a study where we apply our PDS-LM method to a high-dimensional VHAR process in order to construct a contagion network of volatility spillovers for 30 large capital stocks, also accounting for effects from changing correlations. We find that by increasing the information set through considering a high-dimensional VAR model instead of bivariate models, we are able to obtain more realistic effects than in low-dimensional models. Furthermore, even when the sample size is not large enough to use standard full-system VAR techniques, our method remains reliable and delivers accurate results.

Note that unlike belloni2014high, we do not give a “truly” causal interpretation to the established Granger causalities. In how far Granger causality is a useful concept to study true causality is (and has long been) open to debate, see for example eichler2013causal and the references therein. Moreover, though it appears desirable and in line with granger1969investigating's (granger1969investigating) original intentions to make the information set as large as possible, it is well known in the literature on graphical models eichler2013causal for causality that considering only the full model is not sufficient for establishing true causal relations from Granger causal ones. For instance, one-period Granger causality in systems with more than two variables cannot capture indirect causal chains spanning over multiple periods. However, the analysis of the full model is a necessary ingredient for any study of causality in a graphical framework. It would therefore be an interesting avenue for further research to study how the method proposed here could fit into such a graphical framework.

small\onehalfspacing