EconBase
← Back to paper

Dominant Drivers of National Inflation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

55,146 characters · 11 sections · 81 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Dominant Drivers of National Inflation

abstractFor western economies a long-forgotten phenomenon is on the horizon: rising inflation rates. We propose a novel approach christened $D^2ML$ to identify drivers of national inflation. $D^2ML$ combines machine learning for model selection with time dependent data and graphical models to estimate the inverse of the covariance matrix, which is then used to identify dominant drivers. Using a dataset of 33 countries, we find that the US inflation rate and oil prices are dominant drivers of national inflation rates. For a more general framework, we carry out Monte Carlo simulations to show that our estimator correctly identifies dominant drivers. \newline JEL Codes: C22, C23, C55. \newline Keywords: Time Series; Machine Learning; LASSO; High dimensional data; Dominant Units; Inflation \newline

Introduction

The late 1990s marked the start of a period with low inflation rates across the world Rogoff2003. The only exception were the years around the Great Financial crisis. The emergence of the COVID19 pandemic and related supply chain issues brought inflation back to the headlines. The Russian invasion of Ukraine and the resulting increases in energy prices further fuelled inflation, especially in western countries. While it is well understood that national inflation is driven by national factors Auer2019, the effects of spillovers across countries are less researched. Auer2019 find that international input linkages synchronize inflation rates between countries. Bataa2013 present evidence that the Euro Area leads inflation in North America and that inflation rates are more synchronized since the 1980s. The effect of global factors or common factors on national inflation is well understood. Ciccarelli2010 establish that 2/3 of national inflation is due to global inflation. However global inflation is not a stand in for common shocks such as changes in commodity prices and they find that no country is leading global inflation.

To fill this gap in the literature, we use a novel approach to identify dominant drivers influencing national inflation using the GVAR database MehdiRaissi2020. Drivers can be other countries' inflation rate or macroeconomic variables of the same or other countries.\footnote{Dominant drivers are often labelled as dominant series or units. Brownlees2021 call them granular series and KapetaniosPesaranReese2020 call them units with pervasive effects. Throughout the paper we will refer to them as dominant drivers.} We find that the inflation rates in the United States and oil prices changes have a dominant effect on national inflation rates in a set of 33 countries. Our results are robust to different estimation methods and specifications. An advantage of our approach is that it allows to identify dominant drivers even if the number of variables is larger than the number of observations over time.

Our approach relies on the inverse of the covariance matrix of the data and consists of two steps. The first step is based on Meinshausen2006 and Sulaimanov2016 and uses a graphical model to estimate the inverse of the covariance matrix the precision or concentration matrix. In a graphical model nodes (variables) are connected by edges (connections). A zero in the concentration matrix indicates independence between the variables, in a graphical model no connection. A non-zero in the concentration matrix implies dependence or a common edge between two nodes in the graphical model. Hence estimating the linkages between units using a graphical model is informative about the structure or sparseness of the concentration matrix. The entries of the concentration matrix can be represented by partial correlations and are therefore related to estimated regression coefficients Sulaimanov2016. Meinshausen2006 and Sulaimanov2016 propose to use the least absolute shrinkage and selection operator (LASSO) estimator to estimate the graphical model and then use post LASSO OLS to estimate the elements of the concentration matrix. We further extend the approach by Meinshausen2006 and Sulaimanov2016 to time dependent data by combining it with either the rigorous LASSO Bickel2009a,Belloni2016,Chernozhukov2019,AhrensAitkenDitzenEtAl2020 or the adaptive LASSO estimator Zou2006,Medeiros2016a. The advantage is that our approach can be applied to examples where the time dimension is smaller than the number of series or units.

The second step is the selection of the dominant drivers. We use the procedure in Brownlees2021, henceforth BM, to select the dominant drivers from the estimated concentration matrix. In the BM procedure, the column norms of the concentration matrix are ordered by their size and the dominant drivers identified using a criterion similar to the eigenvalue ratio criterion in AhnHorenstein2013. We propose to use the heteroskedastic and autocorrelation robust rigorous LASSO and the adaptive LASSO to estimate the graphical model. Based on the linkages, post-LASSO OLS estimates the entries of the concentration matrix. Monte Carlo simulation results show that our proposed extension correctly identifies the dominant drivers. Since our approach relies on machine learning methods to identify the dominant drivers, we call it Dominant Drivers by Machine Learning, $D^2ML$ .

The theoretical contribution of $D^2ML$ is threefold: First, it can be applied to data which has more variables than observations, a disadvantage of the method by Brownlees2021. Secondly it is computationally more simple than the method in KapetaniosPesaranReese2020 and requires less assumptions. Finally, it can be applied to various types of data as shown in our Monte Carlo Simulations.

Dominant drivers have a strong influence on other units or series and their identification received growing attention in recent years. In an infinite VAR ChudikPesaran2013a suggest to model the dominant driver as a common factor. KapetaniosPesaranReese2020 propose a sequential multiple testing approach to identify drivers with pervasive effects in a large panel model. The underlying idea is to identify the drivers using their error variance, on which the multiple testing approach Bailey2020 is applied to. Parker2016 identify dominant drivers by analysing the residual variances of regressions of principal components on the time series and other principal components. PesaranYang2020 identify dominant drivers in production networks using a criterion similar to the exponent of cross-section dependence. Brownlees2021 define a dominant driver by the means of the column norms of the inverse of the covariance matrix and a selection criteria. Their criteria requires to invert the covariance matrix and is therefore only applicable to datasets with \(N<T\).

Our work extends the literature on inflation and the theoretical literature on the identification of dominant drivers. The results from the empirical application show that the inflation in the US and oil prices act as a dominating series. We do not find evidence that real GDP, equity prices, exchange rates and interest rates have a dominating effect on national inflation.

The remaining part of this paper is structured as follows: the next section describes the theoretical background, followed by a discussion our $D^2ML$ . We provide evidence for our approach using Monte Carlo Simulations and discuss our findings of US and oil prices dominating national inflation rates.

The notation throughout this paper is as follows: matrices are in capital and bold, such as \(\boldsymbol{X}\), vectors are in lowercase letters and bold, \(\boldsymbol{x}\) and scalars are lowercase \(x\). Time indices are denoted by \(t=,1...,T\), unit indices by \(i=1,...,N\) and the number of variables are defined as \(k=1,...,K\). \(||\boldsymbol{X}_i||\) refers to the i-th column norm of matrix \(\boldsymbol{X}\).

The $D^2ML$ approach

This section defines the Dominant Drivers by Machine Learning\ two step approach to identify dominant drivers in large panel models. In the first step we identify the links between the units and in the second step we select the dominant drivers.

In the first step, called the network selection step, henceforth NSS, we recover the network structure using a graphical model and then estimate the concentration matrix. A graphical model links the estimation of a network in the form of nodes connected by edges and the structure of a covariance and concentration matrix, for a summary see Meinshausen2006,Sulaimanov2016,Friedman2008. The link between the graphical model and the inverse of the covariance matrix is that if two series are independent they have a a zero partial correlation and are not in the same edge set.

We want to identify the set \(\Gamma(N_d)\) of \(N_d\) dominant drivers in the \(T \times N\) variable \(\boldsymbol{X}\). We define the entries of the dominant drivers as:

align[align omitted — 233 chars of source]

where \(\beta_{i,j}\) measure how much dominant driver \(j\) influences the non-dominant unit \(i\). A unit cannot influence itself, hence \(\beta_{i,i} = 0\). \(x_{i,t}\) and the iid error component \(u_{i,t}\) are can be serially correlated but stationary over time and can include common factors. The error component \(u_{i,t}\) can also be potentially heteroskedastic. We define the covariance matrix of \(\mathbf{X} = (\boldsymbol{x}_1,...,\boldsymbol{x}_N)'\) by \(\boldsymbol{\Sigma}\) and the concentration matrix is \(\boldsymbol{\kappa} = \boldsymbol{\Sigma}^{-1}\).

Following Meinshausen2006 Sulaimanov2016 the graphical model is recovered based on the optimisation problem:

align[align omitted — 196 chars of source]

where \(\lambda > 0\) is penalty level (or tuning parameter) and \(\psi_j\) is the penalty loading. To estimate Equation (ref) and select the penalty level and loading, we propose two estimator: the rigorous or plug-in LASSO Bickel2009a,Belloni2016,Chernozhukov2019,AhrensAitkenDitzenEtAl2020 and the adaptive LASSO Zou2006,Medeiros2016a. The rigorous LASSO is a data driven method to select \(\lambda\). The penalty loading \(\psi_j\) is estimated and adjusted to the assumptions of the error variances and can account for clustered error variances Belloni2016, heteroskedasticity and autocorrelated errors AhrensAitkenDitzenEtAl2020. The adaptive LASSO is a two step method which allows simultaneous estimation and consistent variable selection by weighting the \(\ell_1\) penalty term. The penalty loading is obtained from an initial regression and the tuning parameter is obtained by cross-validation (CV) or information criteria such as the AIC, BIC, AICC.

Meinshausen2006 show that the inverse of the covariance can be estimated for each node or variable individually. This implies that the problem in Equation (ref) is repeated for each unit or variable which is influenced by the potential dominant driver. Depending on the data, the dominant driver can be a specific cross-sectional unit or a variable.

The solution to Equation (ref) yields the non-zero elements in each row of the concentration matrix, or in different words it informs which units influence the unit in question, unit \(i\). To construct the concentration matrix, the post-LASSO estimates \(\hat{\beta}_{i,j}\) for each cross-section are collected and the \(N\times N\) matrix \(\boldsymbol{\beta}\) constructed:

align[align omitted — 244 chars of source]

Following Sulaimanov2016 we construct the concentration matrix as:

align[align omitted — 178 chars of source]

where \(\boldsymbol{\beta}\) is a \(N\times N\) matrix of the estimated post-LASSO coefficients from (ref) and \(\mathbf{D}\) is a diagonal matrix with the inverse of the error variances of the i-th regression. The matrix \(\boldsymbol{\hat{\beta}}\) will be sparse and the sparseness will carry over to the concentration matrix Sulaimanov2016.

The second step, the dominant driver selection, henceforth DDS, is based on Brownlees2021. The authors show that the entries and therefore the column norm of the concentration matrix will be larger for dominant than non-dominant units.\footnote{The same applies to the row norm, for the remainder of this work we will use the column norm.}

The problem can be divided into two problems: the estimation of the number of dominant drivers and the identification of which units are dominant drivers. Brownlees2021 propose to use the eigenvalue ratio criterion from AhnHorenstein2013 to select the number of dominant drivers applied to ordered column norms of the concentration matrix. The number of dominant drivers is defined as:

align[align omitted — 128 chars of source]

where \(\hat{\kappa}_{(s)}\) is the s-largest column norm of matrix \(\boldsymbol{\hat{\kappa}}\). All units with a larger column norm than column \(N_d\) are considered dominant drivers. Brownlees2021 show that their approach can be easily extended to common factors. The number of dominant drivers is a subset of all units, however no ratio is assumed.

The NSS step requires the assumptions from Meinshausen2006 to ensure oracle properties of the estimator. The oracle properties imply that model selection and estimation of \(\boldsymbol{\beta}\) are unbiased and consistent. The properties are important because otherwise the second step, the DDS, will select falsely non dominant units as dominant drivers. In summary the assumptions for the oracle properties in Meinshausen2006 are that the graph is sparse, independence in the error components, correlations are bounded from below and neighbourhood stability. The aim of $D^2ML$ is to identify dominant drivers in high dimensional datasets where the number of cross-sections or variables is larger than the number of time periods. The dominant units are ordered in a block structure as in Brownlees2021, however we explicitly allow for \(N>T\). The column norms of the dominant drivers are larger than a threshold and larger than those of non-dominant units. This is equivalent to a sparse concentration matrix, in which elements of non-dominant units are close to or exactly zero. The graphical model acts a thresholding method to select only the influential connections between units, which then imply a larger column norm in the concentration matrix.\footnote{Alternative methods for the estimation of the concentration matrix including thresholding the sample covariance matrix are discussed in Sulaimanov2016.} The second assumption is that the sole source of dependence between units is via the dominant drivers and the residuals are cross-sectionally independent. While this assumption is restrictive, our simulations show that it can be relaxed to a certain degree. The last two assumptions are technical and ensure that the entries of the concentration matrix are not going to infinity and that there are no circle connections between two units via a third one.

There are two notable challenges when applying the BM procedure to an estimated sparse concentration matrix. First the procedure can falsely select non-dominant units if the diagonal elements are very large in comparison to the off diagonal elements in a given column. The diagonal elements are the inverse of the residual variance. In combination with the LASSO estimator which minimizes the RSS, the residual variance can become relatively small making the diagonal element in the concentration matrix large in relation to the off diagonal elements. A second challenge is that the BM procedure does not take the number of non-zero elements, or connections, into account. In the extreme, the BM procedure can therefore select a unit with none or a small number connections to other units. To avoid this issue, we restrict the selection in the Monte Carlo Simulation and the empirical exercise.\footnote{KapetaniosPesaranReese2020 discuss this issue and call it the modified BM procedure.} We also assume at least one dominant driver. The advantages of the BM procedure are besides the simple implementation the robustness to common factors, something which is confirmed in our Monte Carlo Simulations.

Finally it should be noted that the the sample covariance matrix can be degenerated if \(N>T\). Bien2011 discuss this case assuming that the selected data is only a subset of the underlying data. The missing units are not connected to the sample. We can ignore those units in our approach since non-connected units will have no influence on the selection of the dominant drivers. Secondly we note that the matrices \(\boldsymbol{\hat{\beta}}\) and \(\boldsymbol{\hat{\kappa}}\) are not symmetric. Meinshausen2006 ensures symmetry by using the AND criterion. Nodes \(i\) and \(j\) are connected if \(\beta_{i,j}\neq0\) and \(\beta_{j,i}\neq0\). Since we are interested in directed networks, \(\boldsymbol{\hat{\beta}}\) and \(\boldsymbol{\hat{\kappa}}\) are required to be non-symmetric. Therefore in the DDS step we can only use the column norm and not the row norm.\footnote{In a footnote BM state that the row norm can be used instead because of the symmetry of the concentration matrix.}

Monte Carlo Simulation

To show the Oracle properties of the proposed estimator, we employ a Monte Carlo Simulation. The simulations aims to shed light on three criteria: 1) model selection; that is if the individual LASSO estimators select the correct units; 2) if the number of dominant drivers is correctly estimated and 3) if the correct dominant drivers are selected. In total we are comparing 5 different specifications. Our data generating process follows KapetaniosPesaranReese2020:

align[align omitted — 333 chars of source]

where \(\boldsymbol{y}_{t,d}\) denotes a \(N_d \times 1\) vector of the \(N_d\) dominant drivers and \(y_{t,nd}\) a \(N_{nd}\times 1\) vector of the \(N_{nd}\) non dominant units. The fixed effects \(\boldsymbol{\mu}_d\) and \(\boldsymbol{\mu}_{nd}\) are drawn from a uniform distribution with \(IIDU(0,1)\).

\(\boldsymbol{\beta}\) is a \(N_{d} \times N_{nd}\) matrix and measures the impact of the dominant drivers on the non-dominant units and the individual elements are generated as:

align*[align* omitted — 173 chars of source]

In the case of \(\alpha = 1\) the dominant drivers affect all non dominant units. In the case of \(\alpha<1\), only a subset is affected by the dominant drivers. For a discussion see KapetaniosPesaranReese2020. The common factors \(f_t\) are uncorrelated across time and the loadings \(\boldsymbol{\gamma}_{d}\) and \(\boldsymbol{\gamma}_{nd}\) are drawn for each unit separately from a \(IIDU(0,1)\) distribution.\footnote{For more details see Section (ref) in the Appendix.} The random noise \(u_{d,t}\) and \(u_{nd,t}\) are allowed to be correlated over time, measured by \(\rho_i\) and generated as a Gaussian process. The weak cross-sectional dependence in \(u_{d,t}\) and \(u_{nd,t}\) is measured by \(\rho_d\) and \(\rho_{nd}\). The dependence structure over time and space is varied between the different specifications.

table[table omitted — 978 chars of source]

Specification 1 is the simplest, with neither autocorrelation or dependence in the normally distributed (Gaussian) errors. The dominant driver affects all units. Specification 2 allows for weakly dependent dominant drivers by changing \(\alpha\) to \(0.5\). Specification 3 is an alternation of specification 1 and allows for autocorrelation in the errors. Specification 4 relaxes the cross-section independence assumption of the errors. Finally, Specification 5 is the same as specification 1 but the number of dominant drivers increases with the number of cross-sections.

The number of time periods and cross-sections varies between \(N\) and \(T=50,100,150\). The number of dominant drivers is fixed to \(5\) and the factors varies between 0, 1 and 5.\footnote{Additional simulations with 0 and 1 dominant factors, specifications with \(\chi^2\) distributed errors and different dependence structures are available in the appendix.} We present results for the HAC robust rigorous LASSO and the adaptive LASSO.\footnote{We also considered the elastic net LASSO in some preliminary simulations. Results were qualitatively worse than the adaptive or rigorous LASSO. We also acknowledge that other LASSO estimation methods such as the graphical LASSO Friedman2008 can be employed to select the model.} To select the hyperparamter \(\lambda\) of the adaptive LASSO we use the AIC, AICC or BIC criterion. The first stage loadings \(\hat{w}\) are calculated \(\hat{w}_j = 1/abs(\hat{\beta}_j)\) where \(\hat{\beta}_j\) are the from an univariate OLS regression of \(y_{i,t}\) on \(y_{j,t}\) Zou2006,Huang2008. We employ all estimators and estimate the \(N \times N\) matrix \(\boldsymbol{\beta}\). In the first step we use the LASSO estimators to select the non zero elements for each cross-section unit (row), then use post-LASSO OLS to estimate the coefficients in the matrix \(\boldsymbol{\beta}\) and finally construct the concentration matrix \(\boldsymbol{\hat{\kappa}}\) following Equation (ref). We then calculate the column norms, order them by size and use the BM procedure based on Equation (ref).

To assess if the LASSO estimators select the correct dominant drivers in the individual estimations we calculate the average number of non-zero elements in each column of \(\kappa\). Additionally we present the True Positive Rate (TPR), the False Positive Rate (FPR) and the False Discovery Rate (FPR) which are calculated as:

align[align omitted — 516 chars of source]

For the dominant drivers we perform the same analysis. We present the average number of estimated dominant drivers and assess if the correct ones are selected by comparing the TPR, FPR and FDR.

The simulations are done in R for the adaptive LASSO using glmnet Friedman2010 and repeated 1000 times. For the rigorous LASSO the Stata command rlasso Ahrens2020 is used with 100 repetitions.\footnote{We use a correction for the Brownlees2021 criterion. If the criterion selects as the largest growth the last possible growth rate, we use the 2nd largest growth rate. This in particular happens if the column norm for specific units turns to zero.}

Results

We start with analysing the NSS step of Specification 1 with 5 dominant drivers. Table (ref) shows the results for the cases with no, one and five common factors. \(\hat{s}\) is the average of non zero column norms and should be equal to \(\left(N_d(N_d-1) + (N-N_d) N_d \right)/ N = [4.9,4.95,4.97]\), as shown in the last block called “Oracle OLS”. The rigorous and adaptive LASSO both overselect the number of non-zero elements in the \(\boldsymbol{\beta}\) matrix. While the adaptive LASSO tends to improve with an increase in \(N\) and \(T\), the rigorous LASSO tends to select more non zero elements. Still the true positive rate is relatively small, indicating that the rigorous LASSO misses out many non-zero elements. However this is not necessarily a disadvantage because the identification of the dominant driver relies on the column norms of the concentration matrix and thus on the size of the estimated coefficients. As long as the falsely selected non dominant drivers obtain a smaller entry in \(\hat\kappa{i,j}\) than the dominant drivers, the column norms of the dominant drivers will be larger than those of the non-dominant units. An advantage of the rigorous LASSO is that it falsely selects an element in the concentration matrix less often than the adaptive LASSO (FPR). If the number of common factors is increased, the adaptive LASSO tends to select better than the rigorous LASSO. This can be seen by the general increase in the TPR.

Next we turn to the estimation of the dominant drivers. While the number of non-zero elements in a given column of \(\kappa\) gives an indication if a unit is dominant or not, the size of the post LASSO coefficient matters more. Table (ref) shows the estimated number of dominant drivers and if the correct units were select for specification 1 with 5 dominant drivers. In the case of no common factors, the approach using the rigorous LASSO and adaptive LASSO, independent of the selection criterion, estimate the number of dominant drivers precisely. This is especially the case for \(T=150\). The TPR is exactly or close to 100% implying that all dominant drivers are correctly identified. A special case for the adaptive LASSO is if \(N=T\), in which the number of dominant drivers is underestimated. A reason for this might be the estimation of the initial loadings, which is done by OLS for each unit separately and relies on 50, respectively 100 observations.\footnote{We tested a specification with \(N=50,100,150\) and \(T=N-5\) and results behave better.} Noteworthy is that for small \(T\), the rigorous LASSO underestimates the number of dominant drivers. Both, the FPR and the FDR are small and converge to zero for both LASSO estimators. For the remainder of the paper we will focus on the results from the estimation of the dominant drivers.

Next we allow the dominant drivers to be weakly dominant, meaning a dominant driver affects only a subset of the non-dominant units. We set \(\alpha=0.5\), implying that only the first 6 (\(N=50\)), 9 (\(N=100\)) and 12 (\(N=150\)) non dominant units are affected. The results are displayed in Table (ref). Again the adaptive LASSO underestimates the number of dominant drivers if \(N=T\). The estimated number of dominant drivers using the rigorous LASSO is slightly downward biased as well, however the bias decreases with \(N\) and \(T\) increasing. Interestingly the rigorous LASSO improves with an increase in the number of factors, while the adaptive LASSO does much worse in comparison to Specification 1. It selects a smaller number of dominant drivers and falsely identifies non dominant drivers as such. A reason for this is that the adaptive LASSO tends to select less non-zero elements in the \(\boldsymbol{\beta}\) matrix and therefore raises the chance to miss out dominant drivers respectively gives more weight to incorrectly selected non-dominant units. Both LASSO approaches outperform the BM criterion, even for the case of \(N<T\).

Specification 3 allows for autocorrelated errors. Results are similar to Specification 1, however distortions when using the adaptive LASSO in the case of \(N=T\) are less pronounced. As expected, both methods control well for autocorrelation in the errors. Noteworthy is that the BM procedure performs well for the case \(N<T\), but is affected when the number of dominant drivers is smaller than the number of common factors.

So far we assumed strong cross-section independence in the random noise. The sole source of dependence were the dominant drivers or the common factors. To relax this assumption, Specification 4, Table (ref), allows for weak dependence in the random noise components. The $D^2ML$ approach is still robust to weak dependence. However the TPR for the cases with more than one dominant driver shrinks, especially for the rigorous LASSO as it underestimates the number of dominant drivers. An increase in T mitigates the problem, but the bias remains.

As a final exercise we return to specification 1 but increase the number of dominant drivers as \(N_d = h N\) with \(h = [0.1,0.5,0.9]\). Table (ref) shows the results for \(N_d = 0.1 N\), implying that \(10\%\) of the cross-sections are dominant drivers. In the case of \(h=0.1\), all methods identify the number of dominant drivers well, but as before the adaptive LASSO underselects if the number of common factors is \(5\). Similarly, the rigorous LASSO does better if the number of common factors increases. Both estimator improve with an increase in \(T\). If the share increases, it is getting harder for the estimator to identify the correct units. A reason for this is that the cut-off point in the BM criterion is less defined. This is in particular the case if \(h=0.9\), implying that \(90\%\) of the units are dominant ones.

In general the results show that the our proposed method reliably estimates the correct number of dominant drivers and identifies the correct drivers. In comparison to the criterion from Brownlees2021 our method can be applied to datasets with \(N>T\). Both LASSO estimator have their strength and weaknesses. While the rigorous LASSO is less affected by common factors, the adaptive LASSO tends to do better in the presence of weakly correlated errors. The case of 5 dominant drivers and 5 common factors is interesting with respect that the method based on the rigorous LASSO identifies the correct dominant drivers and the correct number, while the adaptive LASSO performs poorly. A possible reason might again be the first stage of the adaptive LASSO and difficulties differentiating the common factors and dominant drivers.

Dominant Drivers in National Inflation Rates

In this section we turn back to the question if national inflation rates are exposed to dominant drivers. We use the GVAR database MehdiRaissi2020 with quarterly observations from 1979Q2 to 2019Q4 (\(T=163\)) for \(N_g=33\) countries. The data is in first differences, standardized and demeaned on a country level. Taking first differences is necessary to remove potential non stationary, which ensures that the covariances are not time dependent and that the post-LASSO estimates are unbiased. Standardisation is required to ensure that the coefficients from the sequential regressions of $D^2ML$ are estimated in the same drivers.

The years covered in the GVAR dataset are different to Ciccarelli2010 and more in line with Levin2021,Rogoff2003 who cover the years from 1980 onwards. A key difference to Ciccarelli2010 is that we investigate the effect of global inflation on national inflation. If a variable country combination is selected as a dominant driver, it implies that it is an important driver for the inflation rate in many countries.

We will start by applying the BM procedure to the inflation series of the GVAR dataset. However the procedure can only be applied to a single series and dominant drivers of national inflation might be influenced by further covariates. We therefore apply the $D^2ML$ approach afterwards which overcomes the two limitations of the BM procedure.

Number of Common Factors and BM Procedure

In a first step we estimate the number of common factors using the criteria from Bai2002 and AhnHorenstein2013. The panel criteria from Bai2002 identifies between 4 and 5 common factors and the panel information criteria 1. Both estimator from AhnHorenstein2013 point to 1 common factor. The latter results are in line with the finding in Ciccarelli2010 who find one common factor. In addition testing for strong cross-section dependence PesaranCD2015 confirms the occurrence of strong cross-section dependence. \\

figure[figure omitted — 593 chars of source]

The results imply an underlying common factor structure. To shed more light if the dependence structure is driven by common factors or dominant drivers, we apply the BM procedure next. Therefore we invert the sample covariance matrix to obtain the concentration matrix. Figure (ref) shows the results of the column norms. The dotted line indicates the dominant driver following the BM procedure. We find that the US and Belgium are dominant drivers in the national inflation series for the 33 countries. The two dominant drivers are connected to each other and all other drivers because the concentration matrix is non-sparse. While the US is somewhat expected to be a dominant driver, the finding that Belgium is a dominant drivers is surprising. However this is in line with Ciccarelli2010 who find that a large share of the detrended inflation variance of Belgium is explained by alternative measures of global inflation, pointing that Belgium's inflation rate is highly connected to others'.

A disadvantage of this approach is that the resulting concentration matrix picks up noise which can drive the determination of the dominant drivers. Secondly it is not possible to add any further covariates which might have an influence on national inflation and limiting the effects of the dominant drivers only on inflation. We therefore apply the $D^2ML$ approach next.

Using the $D^2ML$ approach

To allow other variables to have an effect on national inflation, we the $D^2ML$ approach as described in Section (ref). We model inflation as a function of inflation in other units (\(\mathbf{dp}_{-i}\)), real GDP (\(\mathbf{y}_{i}\)), real equity prices (\(\mathbf{ep}_{i}\)) and exchange rates (\(\mathbf{er}_{i}\)) and the nominal short run (\(\mathbf{r}_{i}\)) and long run interest rate (\(\mathbf{lr}_{i}\)). All variables are in first differences to remove potential non stationary and standardized for each country. Standardisation is necessary to ensure that all estimated coefficients are measured in the same units. An advantage of our approach is that if a variable has no influence on the inflation rate, the LASSO estimator will set the respective coefficient and thus influence to zero. In detail, we estimate the following model:

align[align omitted — 433 chars of source]

where \(\mathbf{Y}_{i,k}\) is the i-th element of the k-th variable of \(\mathbf{X}\). \(\mathbf{dp}, \mathbf{y}, \mathbf{dp}, \mathbf{ep}, \mathbf{eq}, \mathbf{lr}, \mathbf{r}\) are \(T \times N\) matrices and \(\boldsymbol{\beta}_{k,i}, k = 1,..,6\) are \(1 \otimes N\). The subscript \(-(i,k)\) denotes that the i-th element in the k-th variable of \(X\) is zero. \(\mathbf{e}_{i}\) is a \(T\times 1\) vector of random noise. \(\boldsymbol{\beta}\) is then a \(NK \times NK\) matrix which is used to calculate the concentration matrix. \(\boldsymbol{\tilde{\beta}}_k\) is a \(N\times NK\) matrix with the coefficients which measure the influence on the k-th variable.

To estimate \(\boldsymbol{\beta}\), the following equation is estimated using the rigorous and adaptive LASSO:

align[align omitted — 236 chars of source]

The concentration matrix \(\boldsymbol{\kappa}\) is calculated following Equation (ref). Since we are only interested in the dominant drivers for inflation, that is in the first \(N\) off diagonal elements of \(\boldsymbol{\beta}\), we place further constraints on \(\boldsymbol{\kappa}\) for the application of the BM procedure. Let's denote \(\boldsymbol{\hat{\tilde{\beta}}^c}\) the constraint version of \(\boldsymbol{\hat{\tilde{\beta}}}\) and define it as:

align[align omitted — 171 chars of source]

Then the estimate of the concentration matrix \(\hat{\boldsymbol{\kappa}}\) becomes:

align[align omitted — 208 chars of source]

The lower \((N-1)K\) rows of \(\boldsymbol{\hat{\tilde{\boldsymbol{\beta}}}^c}\) represent the effects of the covariates on all variables but inflation. Since we are only interested in inflation, this part is set to zero with the only exception of the diagonal. Non-zeros on the diagonal are required to ensure that the columns receive the equal weight for the BM procedure. The diagonal elements are the inverse of the residual variance of the regression representing the respective row. For example \(\hat\sigma_{1,1}\) is the variance of the residuals of inflation of Argentina on the variables selected in the NSS stage by the LASSO estimators.

For the selection of the dominant drivers, we use the modified version of the BM procedure. The selection is restricted to the \(N/2\) most connected units KapetaniosPesaranReese2020. This has the advantage that we filter out units with very large values on the diagonal, but without any connection to other units.

No Common Factors

\paragraph{Rigorous LASSO} First we employ the rigorous LASSO on Equation (ref) which allows for heteroskedasticity and autocorrelation of order 2.\footnote{We varied the bandwidth between 0, 1, 2, 4, and 8, to cover autocorrelation over several quarters. The results remain unchanged and are presented in the Online Appendix.} Column (1) in Table (ref) displays the results. We identify the inflation series of the US as the strongest and of Belgium as the second dominant drivers. Together the dominant drivers account for 57.58% of the column norms and they influence in total 15 other units.\footnote{Diagonal elements are not counted in the column norm shares or the number of non-zero entries.} $D^2ML$ identifies 26 connections between units of which 13 are related to the two dominant drivers. Important to note is that the inflation rate in the US influences the inflation rate of 8 other countries, while the Belgium inflation rate influences 5 others. \\

sidewaystable\begin{tabular}{@{\extracolsep{4pt}}l ll ll ll ll @}\hline \hline & \multicolumn{2}{c}{(1)}& \multicolumn{2}{c}{(2)} & \multicolumn{2}{c}{(3)}& \multicolumn{2}{c}{(4)} \\ & \multicolumn{4}{c}{Rigorous LASSO}& \multicolumn{4}{c}{Adaptive LASSO} \\ \cline{2-5} \cline{6-9} Common Factors & & No & & Yes & & No && Yes\\ \cline{2-3} \cline{4-5} \cline{6-7} \cline{8-9} Number of dom. Units & \multicolumn{8}{l} \\ 1 & dp(bel) & 5 & dp(bel) & 5 & r(chl) & 5 & r(chl) & 5 \\ & & (27.38%) & & (29.16%) & & (23.90%) & & (28.65%) \\ 2 & dp(usa) & 8 & dp(usa) & 7 & dp(usa) & 10 & dp(usa) & 10 \\ & & (30.19%) & & (27.71%) & & (19.64%) & & (23.10%) \\ 3 & & & poil & 8 & & & & \\ & & & & (20.29%) & & & & \\ \hline \multicolumn{9}{l}{Share columnnorms of dom. units.} \\ & & 57.58% & & 77.17% & & 43.54% & & 51.76% \\ \hline \multicolumn{9}{l}{No. Connections} \\ \multicolumn{2}{l}{\ \ Dominant} & 13 & & 20 & & 15 & & 15 \\ \multicolumn{2}{l}{\ \ Total} & 26 & & 26 & & 38 & & 33 \\ \hline\hline \end{tabular} \caption{Column norms of estimated concentration matrix based on Equation (ref). The diagonal elements for calculation of column norm shares and number of connections are removed. Column norm shares for dominant drivers are in parenthesis. dp() denotes national inflation rates, r() national short run interest rates and poil oil prices. Abbreviations for countries are United States (USA), Belgium (bel) and Chile (chl). See Table (ref) for country name definitions and section (ref) for a detailed description.}

An alternative method to display the results is to look at the column norms directly. Panel (a) of Figure (ref) shows the column norms across all country and variable combinations, Panel (b) only the largest 20. Each bar represents a country and a variable. For example the largest bar in Figure (ref) is the column norm of the US inflation. All country-variable combinations to the left of the dashed red line are dominant drivers. \\

figure[figure omitted — 839 chars of source]

Panel (a) shows the distribution of all column norms. We note that inflation has six non-zero column norms, real GDP and real equity prices have one non-zero column norm each. In Panel (b) it becomes evident that the growth rate from the US to the German interest rate is largest and therefore marks the border between dominant and non dominant units.

The lower panel of the figure displays a network graph of the dominant drivers. It is noteworthy that the US influences Belgium but not the other way around. Belgium is a dominant driver for only European countries, while the US influences most of western Europe and Japan and Canada. The effect of Belgium is somewhat surprising. Besides the explanation discussed in the previous section, it is possible that the influence of Eurozone inflation on the national inflation rates not only within the Eurozone but outside of it are picked up by the Belgium inflation. Our finding is in line with argument in Bataa2013 that national inflation in Eurozone countries moves together and Belgium is a proxy for it. Further it extends the finding in Billo2016 that US leads the Eurozone cycle to inflation. \\

\paragraph{Adaptive LASSO}

Next we employ the adaptive LASSO as the selection method of the NSS stage.\footnote{We further restrict the selection of the dominant drivers such that the diagonal is less than 50% of the absolute sum.} The adaptive weights origin from a univariate regression, similar to the Monte Carlo simulation.

The third column in Table (ref) shows the results for this approach. Two dominant drivers are identified, the short run interest rate in Chile and as in the case of the rigorous LASSO the inflation rate in the US. While the US influences a larger number of drivers, the influence of the Chilean short run interest rate is stronger. The number of identified connections is much larger than in the case of the rigorous LASSO. Figure (ref) shows that the inflation series are again picked up most often influencing other inflation series. Panel (b) shows that inflation in Belgium is again influencing strongly other inflation rates, but it is not picked as a dominant driver. \\

figure[figure omitted — 932 chars of source]

Turning to the network graph in Figure (ref) it becomes evident that the short run rate in Chile not only influences the inflation rate in Chile but other countries as well. Among the main drivers for the column norm of the short run rate of Chile is however the affect on the inflation in Chile. The US has a similar widespread influence on inflation rates again. Together with the short run rate of Chile, the two dominant drivers influence 15 out of the 38 national inflation series.

Our results are in contrast to Ciccarelli2010 who find that no country is leading global inflation. Our results strongly suggest that the United States are a dominant factor for national inflation. Depending on the model selection method, Belgium respectively the short run interest rate in Chile are selected as dominant drivers.

Observed Common Factors

It is likely that inflation is not only driven by specific countries, but by other global factors. Examples would be commodity prices, see AasteveitBjornlandThorsrud2015. To investigate the effect of such observed common factor, we add the prices of oil (\(p_{oil}\)), metals (\(p_{met}\)) and materials (\(p_{mat}\)) to the set of variables in matrix \(\mathbf{X}\) in Equation (ref):

align[align omitted — 209 chars of source]

The commodity prices are the same across the 33 countries and they are allowed to be selected as a dominant driver. Results are presented in column (2) and (4) of Table (ref) and in Figure (ref). A comparison between Column (3) and (4) reveals that the adaptive LASSO is not influenced by the additional observed common factors and the results remain similar. We therefore turn directly to the rigorous LASSO. Again the US and Belgium are selected as dominant drivers, however oil prices are selected as a third a dominant driver. In fact, they account for 20.3% of the column norms. The network structure of the two other drivers remain almost unchanged. The US is not directly connected to Switzerland any longer but the shares of the column norm remain stable. As shown in the Monte Carlo simulation, not accounting for common factors does not worsen the identification of dominant drivers using the rigorous LASSO as a selection method.\\

figure[figure omitted — 928 chars of source]

The network graph in Figure (ref) reveals that inflation in Austria, Finland, France, Italy, Switzerland and Thailand is influenced by the global oil prices. Oil prices are also influencing the two other dominant drivers, Belgium and the US, emphasising further their importance.

Our results so far show that the US and oil prices play a crucial role for national inflation rates. In line with the results in Ciccarelli2010 is that some countries are sheltered from the dominant drivers. Our analysis shows that the inflation hardly spills over into countries such as China, India or Norway. Implying that for those countries other factors than the ones covered in our empirical application play an important role.

Unobserved Common Factors

We further account for unobserved common factors.\footnote{Detailed results are available in the Online Appendix.} We approximate potential unobserved common factors using principal components (PCA) or cross-section averages. Both methods are well established in the literature to account for unobserved common factors Pesaran2006,Bai2009. In the previous section we found up to 4 common factors. Assuming that some of those are dominant drivers, we add the first 3 principal components (PCA). Separately the cross-section averages (CSA) of all variables are added to the model. The PCA and CSA are added in the same way as the observed common factors in the previous section.

The findings are similar to the case of adding observed common factors. The nodelasso using the adaptive LASSO identifies Chile as the sole dominant drivers. The rigorous LASSO only identifies the first principal component respectively the cross-section average of inflation as a dominant driver. This implies that the approximation of the unobserved common factors using accounts for too much of the variation and overlays the network structure. The result hints that the strong approximation of the common factor overlays weaker dependence structure, a similar finding than in Juodis2022. Furthermore the long run interest rate in Germany is selected as a dominant driver if the bandwidth is equal to two years and it influences the inflation rate in Switzerland and the United Kingdom.

Conclusion

This paper combines the approach by Meinshausen2006,Sulaimanov2016 to estimate the inverse of a covariance matrix using time dependent data. It then identifies dominant drivers using the procedure from Brownlees2021. Monte Carlo simulations show that $D^2ML$ identifies the correct units as dominant drivers in a panel. We then apply the method to the GVAR dataset and find evidence that inflation in the US and oil prices are dominant factors for national inflation rates. Our results are informative about potential spillovers into national inflation. It can also be used to improve forecasts using a approach as in Bjornland2017.

While $D^2ML$ is simple, it has several limitations. First of all it is computationally expensive, in particular for a large number of cross-sections, respectively variables. Secondly as criticized by Yuan2007,Banerjee2008,Friedman2008 the approach by Meinshausen2006 is only an approximation to the problem. An extension in the spirit of this paper would be to use graphical LASSO Friedman2008 for time dependent data. The selection of an oracle estimator for the NSS step is crucial. It the selected estimator fails, the sequential estimator will fail as well. In our current setting higher order spatial effects and time varying dominant drivers can also not be considered and are left for future research.