Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
59,903 characters · 10 sections · 71 citation commands
Non-linear dependence and Granger causality: A vine copula approach
\setcounter{page}{1}
The notion of Granger causality was seminally introduced by granger1969, and then usefully applied in several scientific fields, including economics, finance, neuroscience, genomics, and climate science shojaie2022, where the characterization of the dependence relations between time series is a relevant issue. It is based on the comparison of the quality of the prediction when forecasting future values of one series with or without the information about the past values of the other time series. In other words, a time series has a Granger causal influence on another time series if its past values facilitate the prediction of the other time series. Therefore, briefly, testing Granger causality means checking if there is an improvement in prediction if we leverage not only the past values of the considered series but also the past values of the other series.\\
In this paper, we propose a novel Granger causality-in-the-mean test for bivariate $k-$Mar\-kov stationary processes, which is based on vine copula models. In general, copula functions are the mathematical tool that describes any type of dependence structure and so they are able to capture non-linear dependencies. In particular, vine copulas are highly flexible models that allow for the representation of a wide array of different joint distributions czado2022 and computationally tractable model-selection and estimation procedures. Formally, exploiting the concept of pair copula construction joe1996, a multivariate copula is constructed using bivariate copulas as building blocks, overcoming the limited and restrictive selection of multivariate copulas. As shown in recent literature brechmannczado2015,beare2015,smith2015,nagler2022, vine copulas are non-linear models that can represent multivariate time series by capturing both their cross-sectional and serial dependence within the same model, giving place to a myriad of applications in the context of time series modelling, one of them being the study of Granger causality. \\
In literature, there are already several tests differing from each other in the models and techniques used for measuring and estimating the correspondent measure of Granger causality. Gran\-ger's original formulation and most of the existing tests are based on linear models such as Vector Autoregressive (VAR) models. Yet, the application of these tools to non-linear systems may not be appropriate ancona2004. Therefore, a few non-linear Granger causality tests have been proposed hiemstra1994,diks2006,diks2016, song2018, kim2020. In economics, the hiemstra1994 test (HJ test hereafter) has been the most popular approach for testing non-linear Granger causality bai2017; however, it has been shown through simulation studies that this test has a tendency to over-reject the null hypothesis. diks2006 proposed a modified version of the HJ test that overcomes the rejection rate issues by changing the test statistics and providing some guidelines for the optimal bandwidth selection based on the sample size. Afterwards, diks2016 extended the setting by diks2006 from a bivariate to a multivariate one. song2018 introduced a model-free measure of non-linear Granger causality that takes the form of a log-difference between restricted and unrestricted mean squared errors, building a test by means of a non-parametric estimate of this measure using kernel estimators of mean squared errors. hiemstra1994, diks2006, diks2016 and song2018 introduce non-parametric methods. More recently, exploiting the measure of song2018, jang2022 proposed a semi-parametric Granger causality test based on vine copulas, which can capture non-linear dependencies with a good statistical performance in terms of size and power if compared to the aforementioned tests. In line with this recent literature, our aim is to propose a test that extends and improves on the family of semi-parametric non-linear Granger causality tests, which statistical properties are not heavily dependent on the choices of bandwidth or smoothing parameters, as it is instead the case for most of the existing non-parametric non-linear Granger causality tests.\\
In this contribution, we start from the test procedure proposed in jang2022 and incorporate some relevant methodological differences. Thus, we build our non-linear Granger causality-in-the-mean test for bivariate $k-$Markov stationary processes, which we call M-vine Granger causality test. By means of a simulation study based on the setting from song2018, we find that the M-vine test has a higher power than the original version from jang2022 in all scenarios considered, while still controlling its size under or considerably close to the predefined significance level. Moreover, for linear specifications and large samples, our test displays a power that resembles the one from a traditional Granger-causality test based on linear models, even if the latter is based on the real data generating process. Yet, for non-linear specifications, the M-vine test has a substantially higher power than the linear Granger-causality test. Hence, we show that our test improves on the statistical properties of the original one in jang2022, and also of other previous methods, and makes it an excellent tool for testing Granger causality in the presence of non-linear dependence structures.\\
Finally, we apply the M-vine Granger causality test to study the pairwise causal relationships between GDP, energy consumption and investment in the U.S. The nexus between energy consumption and economic growth has been studied in various works kraft1978, chen2007, payne2009, ozturk2010. Nonetheless, results appear to be diverse and often conflicting, even when they refer to the same country: for instance, in Appendix (ref), Table (ref), we provide a synthetic recap of the results from different studies that specifically investigated Granger causality between energy consumption and economic growth with a focus on the United States. Here, we provide our contribution by means of the M-vine test, and we compare results with the S-vine test and a more classical linear test based on VARs. At $5\%$ significance level, we find a Granger causality relationship that runs from GDP to energy consumption, which is in line with previous literature kraft1978, fallahi2011, aslan2014). Yet, at the same time, we also find that there is an albeit less significant Granger causality running from energy consuption to economic growth that is still in line with other pieces of economic literature stern1993, stern2000, bowden2009, fallahi2011. Crucially, only our M-vine test is able to catch the two-way Granger causality between energy consumption and GDP. \\
The rest of the paper is organized as follows. Section (ref) introduces the methodology and the setting of the M-vine test, while section (ref) studies the statistical properties of the proposed test by means of a simulation study. In section (ref), we present a macroeconomic application to analyse the pairwise relationships between energy consumption, GDP and investment in the U.S. Conclusions and final remarks are offered in Section (ref).
Measures of Granger causality are classified into three types, according to the criterion chosen to measure their prediction quality: namely, Granger causality in the mean, in the quantiles or in distribution. Here, we focus only on the first, which is the most widely used in literature, while we refer to lee2014 and candelon2016 for the others. Nonetheless, we believe it is important to remark that the methodology that we present can be naturally extended to Granger causality in quantiles taamouti2021,jang2023. \\
Let $(X_t, Y_t)_{t\in\mathbb{Z}}$ be a strictly stationary bivariate stochastic process, that is $$ [(X_t, Y_t),\dots, (X_{t+h}, Y_{t+h})] \stackrel{d}= [(X_s, Y_s),\dots, (X_{s+h}, Y_{s+h})]\qquad\forall s,\,t\in\mathbb{Z},\, h\in\mathbb{N} $$ where the symbol $\stackrel{d}=$ means equality in distribution. Suppose also that $(X_t, Y_t)_{t\in\mathbb{Z}}$ is a Markov process of order $k\in \mathbb{N}\setminus\{0\}$ (more briefly, a $k$-Markov process), that is $$ [X_{t},Y_{t}] \,|\, [(X_s,Y_s):\,s\leq t-1]\stackrel{d}= [X_{t},Y_t] \,|\, [(X_{t-1},Y_{t-1}),\dots, (X_{t-k},Y_{t-k})] \qquad\forall t\,. $$ Set $X=(X_t)_{t\in\mathbb{Z}}$ and $Y=(Y_t)_{t\in\mathbb{Z}}$.
A measure for Granger causality in the mean has been proposed in song2018 as
Note that $GC^{mean}(Y\rightarrow X)$ is a non-negative quantity, which is equal to zero when there is no Granger causality in the mean (i.e. when the two above mean squared errors are equal). Moreover, it is not restricted to a given model for the distributions of the involved random variables. This measure can be used to construct a statistical test in order to check the presence or absence of Granger causality in the mean for two $k$-Markov stationary time series. \\
jang2022 proposed a semi-parametric Granger causality test based on the above measure $GC^{mean}(Y\rightarrow X)$, estimated using a {\em copula approach} (see Appendix (ref)). Copula functions are the mathematical tool that describes any type of dependence structure. Therefore, this test is able to capture non-linear dependence and, moreover, it exhibits good statistical properties when compared to other existing non-linear tests (see jang2022). Hence, it is our starting point. We here propose a test for Granger causality in the mean based on the M-vine copula structure presented in beare2015 for a bivariate $k-$Markov stationary stochastic process. The presented test improves both on the size and the power when compared to the one in jang2022, and it is more suitable for empirical applications since it takes into account possible non-linear dependencies that are commonly encountered when working with real data. Please, refer to Appendix (ref) for some recalls on the notion of copulas and, in particular, on vine copulas.
Let $(X,Y)=(X_t,Y_t)_t$ be a bivariate $k-$Markov square-integrable stationary stochastic process and let $\{(x_t,y_t)\;:\;t=1,\dots,T\}$ be a sample of it. To simplify notation, from now on, we will assume $k=1$, but everything can be naturally extended to any $k\in\mathbb{N}\setminus\{0\}$. \\
We can split the proposed test, which we call the M-vine Granger causality test, into two parts, say part A and part B: the first one regards the evaluation of an estimator of the Granger causality (in the mean) measure, which is used as the test statistics, and the second one concerns the computation of the $p$-value of the test. \\
We start by describing part A (computation of the value of the test statistic). We use the sample version of the above-recalled measure $GC^{mean}(Y\rightarrow X)$, that is
where the quantities $\widehat{\mathbb{E}}[X_t|X_{t-1}=x_{t-1}]$ and $\widehat{\mathbb{E}}[X_t|X_{t-1}=x_{t-1}, Y_{t-1}=y_{t-1}]$ are computed by fitting an M-vine copula model to the data (see Appendix (ref)) and using it to generate i.i.d. observations of $X_t$ given $X_{t-1}=x_{t-1}$ and of $X_t$ given $X_{t-1}=x_{t-1}$ and $Y_{t-1}=y_{t-1}$ and using the corresponding empirical means to estimate the needed conditional expectations. More precisely, we proceed as follows:
Introducing the sample estimates of both conditional expectations in $\widehat{GC}^{mean}_{Y\rightarrow X}$, we obtain
Note that, by the strong law of large numbers, for $N\to +\infty$, the empirical means $\frac{1}{N}\sum_{i=1}^N \tilde{x}^{M_x}_{i,t}$ and $\frac{1}{N}\sum_{i=1}^N \tilde{x}^{M_{xy}}_{i,t}$ are strongly consistent estimators of the conditional mean of $X_t$ given $X_{t-1}=x_{t-1}$ and given $X_{t-1}=x_{t-1},\, Y_{t-1}=y_{t-1}$ respectively, with respect to the model obtained from Step A1. Therefore, the computed quantity $\widehat{GC}^{mean}_{Y\rightarrow X}$ results to be a strongly consistent estimator (for $N\to +\infty$) of the Granger causality measure with respect to the model obtained from Step A1 applied to $M_x$ and $M_{xy}$). \\
Roughly speaking, $\widehat{GC}^{mean}_{Y\rightarrow X}$, measures the log-difference between the prediction error related to the model with an M-vine copula structure fitted only on the information set from $X$, and the prediction error computed with the M-vine copula model fitted to the entire sample form both $X$ and $Y$, that is using the information set of both series. In line with the definition of Granger causality in the mean, if this estimated quantity is significantly higher than zero, we reject the null hypothesis of no Granger causality (in the mean) from $Y$ to $X$: indeed, under the null hypothesis, the measure $GC^{mean}(Y\rightarrow X)$ is zero, and so, the higher is the value of the statistics $\widehat{GC}^{mean}_{Y\rightarrow X}$, the more statistically significant is the evidence of the presence of Granger causality (in the mean) running from $Y$ to $X$. Regarding $T_0\geq k+1$, we can choose it by taking into account the goodness of fit of the models to the data. \\
In order to test if $\widehat{GC}^{mean}_{Y\rightarrow X}$ is statistically greater than zero, we rely, as in jang2022, on a method, which takes inspiration from paparoditis2000. This method relies on the M-vine copula model fitted on the entire sample $M_{xy}=\{(x_t,y_t):t=1,\dots,T\}$ in order to generate independent samples under the null hypothesis of no Granger causality (in the mean) and use them for computing the $p-$value for the test. More precisely, the part B (computation of the $p$-value) of the proposed procedure works as follows:
Therefore, we reject the null hypothesis of no Granger causality (in the mean) when $p<\alpha$, where $\alpha$ is a given significance level.
There are two relevant differences between the test by jang2022 and the one we propose. Firstly, in the above-described Step 1A), we choose the class of the M-vine copulas to fit the data and to compute the value of the test statistic. The same model is then used in the procedure (part B) to compute the $p$-value for the test. On the contrary, jang2022 use a more general class of copula models (that includes the M-vine copulas), i.e. the class of S-vine copulas, for fitting the data and computing the test statistic, but then they restrict to the M-vine copulas for the computation of the $p$-value. The justification of this fact given in jang2022 simply relies on the implementation of this last computation: indeed, to draw the samples in the above Steps B2) and B3), both $c_{X_{t},X_{t+1}}$ and $c_{X_t,Y_t}$ are needed, and these two copulas are guaranteed to be present in the first tree of a given M-vine structure, whilst these two copulas might not appear in an S-vine structure (see Section (ref) for more details). Consequently, in jang2022, the conditional distributions used for generating a sample of the Granger causality measure under the null hypothesis might be different from the ones in the model fitted to the data to estimate the Granger causality measure, i.e. to compute the value of the test statistic. In our procedure, the two parts, A (computation of the test statistic) and B (computation of the $p$-value), of the Granger causality test are aligned: they rely on the same class of vine copulas, i.e. the M-vine copulas. \\
Secondly, in the above Step 1A), we fit the M-vine copula model on the entire sample, i.e., on all the $T$ observations, while jang2022 split the entire sample into a training set, say observations from $t=1$ to $t=T^*$ (usually $T^*=T/2$), and a testing set, say observations from $t=T^*+1$ to $t=T$. Therefore, the fit of the $S$-vine copula model is performed only on the training set, while the testing set is used to compute the Granger causality statistics. The latter choice implies that all the predictions $\{\tilde{x}^{M_x}_{i,t}\}_{i=1}^N$ and $\tilde{x}^{M_{xy}}_{i,t}$ for $t=T^*+1,\dots, t$, used to compute the Granger causality statistics are based on a model fitted only on the observations from $t=1,\dots, T^*$. In real applications, when the time series are not perfectly stationary or the sample is small, this way of proceeding could generate poor estimates of the Granger causality measure. This is especially true in the case of non-linearities in the data generating process. In our case, instead, predictions are based on a model fitted to the entire sample, and this potentially makes our test better at detecting Granger causality, especially when the dependence structure is not linear. Note that fitting the model on more observations does not mean to add more model parameters, hence the risk of “overfitting” the data is not present.\\
In the next subsection, we show that our variant has a better performance in terms of size and power.
In this section, we analyse the statistical performance of our M-vine Granger causality test (briefly, M-vine test) and compare it, in terms of size and power, with the traditional linear Granger-causality test based on VAR models \footnote{The results of the linear test are obtained using the function grangertest from the R package lmtest. For the formulation of the linear test, we refer to Appendix (ref)}, the S-vine Granger causality test (briefly, S-vine test) by jang2022, and with a non-linear extension of the traditional Granger causality test using feed-forward neural networks from Hmamouche2020 (briefly, NlinTS) \footnote{For the simulation study, we decided to compare the M-vine test only with methods for which the code is available, so that we could run the simulations ourselves and ensure the validity of the results.}. We recall that, as shown in jang2022, the S-vine test has a good statistical performance in terms of size and power when compared to other non-linear Granger causality tests diks2006, song2018, kim2020. Specifically, through a simulation study, they show that S-vine test, ST test song2018 and KLH test kim2020 are the ones whose size is closer to the predefined significance level, whereas the size of the DP test diks2006 is considerably smaller than the significance level and this suggests a tendency of this test to under-reject. Indeed, in terms of statistical power, jang2022 evidences that the DP test is the one that performs the worst. On the contrary, the S-vine test shows, in large samples ($T\geq 200$ for Markov processes of order $k=1$), a good power for all the dependence structure, while the KLH test has an excellent power only when working with linear or quasi-linear models and the ST test displays a good power exclusively when the dependence structure is non-linear. In the following, we will show in particular that our M-vine test outperforms the S-vine test in all the considered scenarios, exhibiting a very good power already starting from $T=100$ in the case of Markov order $k=1$. \\
We performed a simulation study based on the assessment models from jang2022. We used their three size assessment models (S1, S2 and S3) and their three power assessment models (P1, P2 and P3). Furthermore, we enhanced the exercise by including other data generating processes: two size assessment models (S4 and S5) and one power assessment model (P5) taken from song2018 and an additional power assessment model (P4) where both series come from non-linear specifications.\\
Size assessment models
Power assessment models
where $(\epsilon_t,\eta_t)$ are i.i.d. from a standard bivariate normal distribution. Recall the size of a test corresponds to the probability of incorrectly rejecting the null hypothesis when it is true, whilst the power of a test is the probability of correctly rejecting the null hypothesis when it is false. Consequently, the size assessment models correspond to data-generating processes for which there is no Granger causality from $Y$ to $X$ (note that S3 exhibits Granger causality from $X$ to $Y$, but not from $Y$ to $X$), whereas the models used for the power assessment of the test are built under the presence of Granger causality from $Y$ to $X$.\\
For the simulation study, we used $S=500$ simulations for each model with sample sizes of $T=50$, $T=100$ and $T=200$, and computed the empirical size and power using a significance level $\alpha=0.05$. Furthermore, for part A of the procedure, we used $N=200$ i.i.d. predictions for each series to estimate the conditional expectations and, for part B, i.e. the computation of the $p$-value, we used a number $B=200$ of generated samples. We used $T_0=T/2$ in the computation of the test statistic in accordance with the choice $T^*=T/2$ in jang2022.\\
Table (ref) shows evidence that the M-vine test and the linear test control the empirical size below or around the significance level of $\alpha=0.05$ better than the S-vine test. Furthermore, the NlinTS test is the one that exhibits higher size distortions; in fact, when compared to the other tests, only the S-vine test by jang2022 exhibits in some cases a size greater than or equal to $0.070$. In terms of power, the NlinTS test is the one with the worst performance. Moreover, the M-vine test outperforms the S-vine test, in every power assessment model, in small and large samples, with the biggest difference when working with series from the data-generating model P3, that exhibit a quadratic dependence relationship. Furthermore, the power of the M-vine test seems to increase as the sample size grows faster than the one of the S-vine test. Notice that the M-vine and the S-vine tests have the same value of the test statistic under the null hypothesis, i.e. $\widehat{GC}^{0\;mean}_{Y\rightarrow X}$, and so the difference in the performance relies upon the computation of $\widehat{GC}^{mean}_{Y\rightarrow X}$, that explains the difference in the obtained $p$-values. When compared to the linear test, the M-vine test tends to have a smaller power across the assessment models P1 and P2 that correspond, respectively, to a linear model and to a non-linear model with a bounded non-linear term; this difference is decreased as we move to larger samples, becoming negligible at $T=200$. On the contrary, the power of M-vine test is substantially bigger than the one of the linear test when working with non-linear model specifications (precisely, with models exhibiting a quadratic dependence), such as P3, P4, and P5, particularly when working with larger samples: indeed, the power of the linear test does not increase as the sample size grows, remaining considerably low. Notably, all the tests considered in the analysis exhibit low power working with the P5 model, in which the weight of the term $Y_{t-1}$ in the dynamics of the process $X$ is lower than the one in the previous models, which implies that $Y$ Granger-causes $X$ to a lesser extent, requiring a larger sample size to achieve a higher power. Hence, the presented analyses show that the M-vine test brings remarkable improvements in the current literature regarding tools for detecting the Granger causality in the presence of non-linear dependencies. \\
Table (ref) and (ref) show the average $p$-value across simulations for each size and power assessment model, respectively, together with their empirical standard deviation. One can see that, on average, the $p$-values for all methods are around the significance level $\alpha=0.05$ in every size assessment model (with the exception of P5 due to the sample size as previously mentioned). However, those of the M-vine approach have the lower standard deviation among the three tests in all size models, both in small and larger samples. On the other hand, from Table (ref), we can see that for the linear (P1) and close to linear (P2) model specifications, the $p$-values of the Linear method are the ones that are, on average, further below the predefined significance level, also displaying the lower standard deviation. When working with non-linear dependence structures (P3, P4 and P5), this no longer holds as the $p$-values for the M-vine procedure are, on average, closer to the significance level, and they exhibit the lower empirical deviation out of the three methods. In Figure (ref) we plot the distributions of the $p$-values obtained with each method in cases P3 and P4 with $T=100$ in order to further support the previous considerations.\\
Thus far, we have assumed that $(X_t,Y_t)_t$ is a first-order Markov process (i.e. $k=1$), and, under this condition, the M-vine test results to have good statistical properties for sample size $T\geq 100$. In order to analyse how these properties change when working with higher-order Markov processes, we repeat the simulation study for the case $k=4$. The selection of this particular Markov order is driven by the empirical application from Section (ref), as it is the only other Markov order, besides $k=1$, obtained by our selection procedure when fitting the M-vine models to the U.S. macroeconomic data (we refer to Section (ref) for a detailed description of the data). To adjust the simulation study for $k=4$, we modified a selection of the previously presented size and power assessment models so that they represent the same dependence structure scenarios, but letting $(X_t,Y_t)_t$ be a bivariate Markov stationary stochastic process of order $k=4$. The modified assessment models are given in Appendix (ref). Table (ref) and Table (ref) show the empirical size and power of each test when applied to a bivariate Markov stationary process of order $k=4$. From them, we can see that overall the results exhibit the same pattern as the case of $k=1$, with the M-vine test still outperforming in terms of power the S-vine test in both linear and non-linear assessment models. Moreover, when compared to the linear test, the M-vine test has a comparable performance when the processes have a linear or almost linear dependence and the sample size is large enough, while it performs better in non-linear cases. However, we note that for the order $k=4$, the power of the M-vine test seems to converge to 1 at a slower pace than in the case of $k=1$. Thus, for $k=4$, we require a greater $T$ to let the M-vine test have a good power in the cases P3 and P4.\\
The nexus between energy consumption and economic growth is important in times when economic and environmental sustainability are at stake. On the one hand, the correlation is positive everywhere in the world. No wealthy country consumes only a little energy, and no poor country consumes a lot of energy. Yet, beyond correlations, we believe it is important to investigate whether there is any Granger causality between the two variables, especially in a period in which we do have concerns about the impact on the environment and the depletion of natural resources that are needed to generate energy. In the latest decades, since kraft1978, the nexus between energy consumption and economic growth has been analysed in a wide array of studies. Nonetheless, results appear to be diverse and often conflicting as some of them find neutrality, whilst others find a causal relationship, often running in different directions. As pointed out in chen2007, the main reason for conflicting results might be the perspective taken from different countries that have their own energy policies, institutions, and cultural factors. However, we can find conflicting conclusions even when they refer to the same country. Indeed, the employment of data sets with different proxy variables for energy consumption or different temporal frequencies, as well as different econometric methods, can certainly contribute to different findings about the energy-growth relationship. Therefore, by now, previous literature payne2009, ozturk2010 has been inconclusive about policy recommendations for either a specific country or across different nations. Most of these studies apply a notion of Granger causality tests based on linear models, such as Vector Auto Regressive models (VAR). Only a small number of studies experiment with non-linear models (huang2008, chiou2008, fallahi2011). In Appendix (ref), Table (ref), we provide a snapshot of the results from different studies that specifically investigated Granger causality between energy consumption and economic growth with a focus on the United States. Finally, we include in our study also the U.S. time series for investment, proxied by using the Gross Fixed Capital Formation, to show how Granger causality can work differently when considering a variable that is not supposed to have a short-term impact on GDP. Crucially, we decide to investigate the U.S. case for two reasons. On the one hand, it is the most investigated case, and we will not find it difficult to compare our results with previous applications. On the other hand, it is easier to find longer time series that are available for our variables of interest.
We assemble our data set from different sources. We obtain real GDP and Gross Fixed Capital Formation (investment hereafter) from the Federal Reserve Economic Data (FRED) as made available by the Federal Reserve Bank of St. Louis \footnote{Accessible through https://fred.stlouisfed.org/}. Energy consumption data are proxied by the so-called total primary energy consumption measured in quadrillion British thermal units (BTU), which is made freely available by the U.S. Energy Information Administration (EIA) \footnote{Accessible through https://www.eia.gov/opendata/}. Crucially, we consider all the time series with a quarterly frequency in the period 1973-2018.\\
A preliminary investigation of the time series for this empirical application is available in Appendix (ref). To proceed, we need stationarity of the time series. Thus, we work with the first differences of original variables after testing that they are indeed stationary by means of the Phillips-Perron test pp1988. First of all, we focus on how the fit to the entire data set of an M-vine model compares to both an S-vine model (on which the test from jang2022 is based) and a traditional linear model, i.e., a VAR. Table (ref) shows the results of the fit of these models in terms of the Akaike Information Criterion (AIC) for each pair of variables considered in the analysis. The Markov order of each bivariate process is selected by iteration, fitting M-vine copulas with Markov orders from $1$ to $4$, thus keeping the one with the better fit according to the AIC. Analogously, the same process is repeated for the case of S-vine models. For the linear model, we select the optimal amount of lags in the bivariate VAR models based on the AIC. That is, we sequentially fit VAR models with an increasing number of lags, and then we pick the one with the minimum AIC \footnote{Please note that we implement the command VARselect from the RStudio package VAR Modelling (vars).}. \\
Table (ref) put in evidence the non-linear dependence structure between every pair of considered variables: indeed, once we allow for non-linear models, the goodness of fit is considerably higher. Furthermore, when comparing the goodness of fit between the two vine copula models, one can notice that the M-vine and the S-vine models exhibit almost the same AIC. In particular, in the first two cases, the model selected is exactly the same for both vine copula models, suggesting that the M-vine is the structure that has the best goodness of fit among all the vine structures that can represent stationary time series (i.e. S-vine copulas).\\
The results we obtain with our M-vine test, the test from jang2022, and the more traditional test based on a linear model are prima facie relatively similar. Please recall that the null hypothesis of these tests is that no Granger causality exists, and it is rejected at a given significance level $\alpha$ if $p<\alpha$. Remarkably, the most important difference in the results is for the pair GDP and energy consumption. In fact, as Table (ref) shows, all three methods find that Granger causality runs from GDP to energy consumption, although the linear test shows a relatively weaker statistical significance at $10\%$ level, in line with the findings from kraft1978, fallahi2011 and aslan2014.
Interestingly, only the M-vine test detects an inverse Granger causality from energy consumption to GDP at the $10\%$ level. In this specific case, findings are in line with stern1993, stern2000, bowden2009, fallahi2011. As shown in Table (ref), the difference we observe when we compare the M-test and the VAR is due to a non-linear dependence structure between the variables. In this case, we argue that the M-vine and the S-vine tests are better than a classical VAR test at catching the underlying functional form. Moreover, as already shown with the simulation study in previous paragraphs, the M-vine test outperforms the S-vine test thanks to the methodological differences we discussed in section (ref). Hence, the differences that we record in the $p$-values of Table (ref).
Notably, our findings point to a two-way Granger causality between energy consumption and GDP that is important from an economic perspective. Inevitably, energy consumption is a relevant expenditure component for consumers and producers of any budget. For this reason, economic growth is reflected necessarily on increasing energy consumption. From another point of view, we can say that the demand for energy is quite rigid for both consumers and producers.
Please note how another important difference arises when we test the pairwise causal relationship between investment and economic growth. In this case, the linear test rejects the null hypothesis, thus finding evidence of a Granger causality that runs from investment to GDP, whilst both non-linear tests find no Granger-causal relationship in either direction. In this regard, please consider that Table (ref) shows the VAR model has absolutely the worst fit when modelling the dependence structure between investment and economic growth. Once again, we believe this is a hint that the results of non-linear tests are more reliable when applied to relationships that can hardly be proxied as linear. From an economic perspective, we know that investment might not have a direct effect on GDP, but it can contribute indirectly to economic growth in the medium term after raising the technological abilities of a country and, thus, its production.
Finally, all three methods find evidence that investment Granger-causes energy consumption growth. In this case, we believe that the tests are catching the intrinsic prediction power of investment plans that always require the use of additional energy resources.\\
Motivated by the recent literature about non-linear Granger causality tests, our aim was to construct a semi-parametric Granger causality-in-the-mean test for bivariate $k$-Markov stationary processes based on vine copulas. Departing from the procedure in jang2022, we added coherence between the two parts of the test, i.e, the estimation of the Granger causality measure and the computation of the $p$-value, by fitting the same M-vine structure in both parts. In addition, we modified the approach by fitting the vine copula model to the entire sample. Therefore, by means of a simulation study with time series from different data generating processes and Markov orders, we showed that we improve on the statistical properties, making the proposed M-vine test an excellent tool for testing Granger causality in the presence of non-linear dependence structures.\\
Finally, in an application, we use the M-vine test to study the relationship between energy consumption, economic growth and investment in the U.S. Crucially, the copula-based M-vine test revealed a two-way Granger causality between energy consumption and GDP that was not detected by the other tests. We argue that, indeed, a quite rigid demand for energy can be the basis for the two-way prediction power in the relationship.
More in general, we conclude that vine copulas are able to model the dependence structure of multiple stochastic processes, better catching intrinsic non-linearities. A possible avenue for future research is the extension of vine copula-based tests to multivariate settings.
Irene Crimaldi is partially supported by the project MOTUS - Automated Analysis and Prediction of Human Movement Qualities (code P2022J8AXY, cup: D53D23017470001), financed by the Italian Ministry of University and Research (PRIN 2022 PNRR)