EconBase
← Back to paper

Estimating the Effect of Central Bank Independence on Inflation Using Longitudinal Targeted Maximum Likelihood Estimation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

96,091 characters · 26 sections · 89 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimating the Effect of Central Bank Independence on Inflation Using Longitudinal Targeted Maximum Likelihood Estimation

\pagenumbering{roman} \thispagestyle{empty}

abstractThe notion that an independent central bank reduces a country's inflation is a controversial hypothesis. To date, it has not been possible to satisfactorily answer this question because the complex macroeconomic structure that gives rise to the data has not been adequately incorporated into statistical analyses. We develop a causal model that summarizes the economic process of inflation. Based on this causal model and recent data, we discuss and identify the assumptions under which the effect of central bank independence on inflation can be identified and estimated. Given these and alternative assumptions, we estimate this effect using modern doubly robust effect estimators, i.e., longitudinal targeted maximum likelihood estimators. The estimation procedure incorporates machine learning algorithms and is tailored to address the challenges associated with complex longitudinal macroeconomic data. We do not find strong support for the hypothesis that having an independent central bank for a long period of time necessarily lowers inflation. Simulation studies evaluate the sensitivity of the proposed methods in complex settings when certain assumptions are violated and highlight the importance of working with appropriate learning algorithms for estimation.

{\it Keywords:} causal inference, doubly robust, super learning, macroeconomics, monetary policy.

\pagenumbering{arabic}

Introduction

The impact of the institutional design of central banks on real economic outcomes has received considerable attention over the past three decades. Whether central bank independence (CBI) can lower inflation and provide inflation stability in a country is a particularly controversial issue. It has been claimed that more than 9,000 works have been devoted to the investigation of the role of CBI in influencing economic outcomes vuletin2011replacing. After the 2008-09 Global Financial Crisis, the debate on the optimal design of monetary policy authorities has become even more intense.

The statistical and economic literature is rich in studies that evaluate the relationship between CBI and inflation. A common approach is to treat countries as units in a linear regression model where inflation (the percentage change in the consumer price index, CPI) is the outcome and a binary CBI index and several economic and political variables are covariates. While many studies have found that an independent central bank may lower inflation grilli1991political, cukierman1992measuring, alesina1993central, klomp2010inflation, klomp2010central, arnone2013dynamic, other studies that have used a broader range of characteristics of a nation's economy have been unable to find such a relationship cargill1995statistical, fuhrer1997central,oatley1999central. Moreover, there have been studies suggesting that the effect of CBI on inflation can only be seen during specific time periods klomp2010inflation or only in developed countries klomp2010central,neyapti2012monetary,alpanda2014impact.

Numerous articles have pointed out the weaknesses that come with simple cross-sectional regression approaches when evaluating the effect of CBI on inflation. First, the problem at hand is longitudinal in nature, and only an appropriate panel setup may be suitable to estimate the (long-term) effect of CBI on inflation. Second, the question of interest is essentially causal: i.e., what (average) inflation would we observe in 10 years' time, if -- from now on -- each country's monetary institution had an independent central bank compared to the situation in which the central bank were not independent? However, not all cross-sectional regression approaches embed their analyses in a holistic causal framework.

Some more recent work has attempted to overcome at least parts of these problems. For example, crowe2007evolution,crowe2008central use a panel data setup with two time intervals, and klomp2010central work with a random coefficient panel model to answer the question of interest in a longitudinal setup. Other authors, e.g., walsh2005optimal, acknowledge not only that current CBI may cause future inflation but also that current inflation is possibly related to future CBI status. Several authors have thus tried to use instrumental variable approaches to estimate the effect of CBI on inflation within a causal framework, but have been unable to find strong instruments crowe2008central,jacome2008there.

It is clear that evaluating the effect of CBI on inflation requires a longitudinal causal estimation approach. However, it has been shown repeatedly that standard regression approaches are typically not suitable to answer causal questions, particularly when the setup is longitudinal and when the time-dependent confounders of the outcome-intervention relationship are affected by previous intervention decisions (greenland1998introduction, Daniel:2013). There are at least three methods to evaluate the effect of longitudinal (multiple time-point) interventions on an outcome in such complex situations: 1) inverse probability of treatment weighted (IPTW) approaches (Robins:2000); 2) standardization with respect to the time-dependent confounders (i.e., g-formula-type approaches Robins:1986, Bang:2005); and 3) doubly robust methods, such as targeted maximum-likelihood estimation (TMLE, vanderLaan:2011), which can be seen as a combination and generalization of the other two approaches.

Longitudinal targeted maximum likelihood estimation (LTMLE, vanderLaan:2012) is a doubly robust estimation technique that requires iteratively fitting models for the outcome and intervention mechanisms at each time point. With LTMLE, the causal quantity of interest (such as differences in counterfactual outcomes after intervening at multiple time points) is estimated consistently if either the iterated outcome regressions or the intervention mechanisms are estimated consistently. LTMLE, like other doubly robust methods, has an advantage over other approaches in that it can more readily incorporate machine learning methods while retaining valid statistical inference. Recent research has shown that this is important if correct model specification is difficult, such as when dealing with complex longitudinal data, potentially of small sample size, where relationships and interactions are most likely highly nonlinear and where the number of variables is large compared to the sample size (Tran:2019, Schomaker:2019).

Using causal inference in economics has a long history, starting with path analyses and potential outcome language (tinbergen1928supplycurves, wright1934path) and continuing with regression discontinuity analyses (Hahn2008regressiondiscontinuity), instrumental variable designs (imbens2014instrumentalvariables), and propensity score approaches in the context of the potential outcome framework (rosenbaum1983propensity), among many other methods. More recently, there have been works advocating the use of doubly robust techniques in econometrics (Chernozhukov:2018). From the perspective of statistical inference, this is a very promising suggestion because the integration of modern machine learning methods in causal effect estimation is almost inevitable in areas with a large number of covariates and complex data-generating processes (Schomaker:2019).

However, the application of doubly robust effect estimation can be challenging for (macro-)economic data. First, the causal model that summarizes the knowledge about the data-generating process is often more complex for economic than for epidemiological questions, where most successful implementations have been demonstrated thus far (Kreif:2017, Decker:2014, Schnitzer:2014, Schnitzer:2014b, Schnitzer:2016, Tran:2016, Schomaker:2019, Bell-Gorrod:2019). The task of representing the causal model in a directed acyclic graph (DAG) becomes particularly challenging when considering how economic variables interact with each other over time. Thus, to build a DAG, a thorough review of a vast amount of literature is needed, and economic feedback loops need to be incorporated appropriately. Imbens:2019, who discusses different schools of causal inference and their use in statistics and econometrics, as well as different estimation techniques, emphasizes this point:

\parbox{0.95\textwidth}{ {\it{\glqq [...] a major challenge in causal inference is coming up with the causal model.\grqq}}}

Second, even if a causal model has been developed, identification of an estimand has been established and data have been collected, statistical estimation may be nontrivial given the complexity of a particular data set (Schomaker:2019). If the sample size is small, potentially smaller than the number of (time-varying) covariates, recommended estimation techniques can fail, and the development of an appropriate set of learning and screening algorithms is important. The benefits of LTMLE, which is doubly robust effect estimation in conjunction with machine learning to reduce the chance of model misspecification, can be best utilized under a good and broad selection of learners that are tailored to the problem of interest.

Estimating the effect of CBI on inflation is a typical example of a causal inference question that faces all of the challenges described above. Our paper makes five novel contributions to the literature. i) We discuss identification and estimation for our question of interest and estimate the effect of CBI on inflation; ii) we develop a causal model that can be applied to other macroeconomic questions; iii) we demonstrate that it is possible to develop a DAG for economic questions, which is important, as it has been argued that "the lack of adoption in economics is that the DAG literature has not shown much evidence of the benefits for empirical practice in settings that are important in economics." Imbens:2019; iv) we demonstrate how to integrate machine learning into complex causal effect estimation, including how to define a successful learner set when the number of covariates is larger than the sample size and when there is time-dependent confounding with treatment-confounder feedback hernan2017causal; and v) we use simulation studies to study the performance of doubly robust estimation techniques under the challenges described above.

This paper is structured as follows. In the next section, we motivate our question of interest, and this is followed by the general description of our framework. Section (ref) contains the data analysis and describes the doubly robust estimation strategy to estimate the effect of CBI on inflation. In Section (ref), we conduct simulation studies motivated by our data analysis. Section (ref) concludes the paper.

Motivating Question: Central Bank Independence and Inflation

When governments have discretionary control over monetary instruments, typically a short-term interest rate, they can prioritize other policy goals over price stability. For instance, after nominal wages have been negotiated (or nominal bonds purchased), politicians may be tempted to create inflation to boost employment and output (gross domestic product, GDP) or to devalue government debt. This is referred to as the time-inconsistency problem of commitments to price stability. It results in an inflation rate higher than what is socially desirable. To overcome this outcome, the literature discusses a variety of commitment mechanisms (also called “commitment technologies”), ranging from simple rules (such as the imposition of strict rules on the rate of monetary expansion, inflation targeting and nominal exchange rate targeting), contracts between the government and the central bank, reputational forces and, from a practical perspective the best known and implemented mechanism, the delegation of monetary policy-making to an independent central bank. In particular, rogoff1985optimal has proposed delegating monetary policy to an independent and “conservative” central banker to reduce the tendency to produce high inflation. Here, conservative means that the central banker dislikes inflation more than the government, in the sense that they places a greater weight on price stability than the government does. Once central bankers are insulated from political pressures, commitments to price stability can be credible, which helps to maintain low inflation. Rogoff's seminal paper had a twofold effect: stimulating the implementation of central bank reforms on the policy side and creating avenues for the design of indices that are suitable to capture the degree of independence of these institutions on the research side.

Following these ideas, a considerable policy consensus grew around the potential of having independent central banks to promote inflation stability bernhard2002political, kern2019imf. Numerous countries followed this policy advice. Between 1985 and 2012, and excluding the creation of regional central banks, there were 266 reforms to the statutory independence of central banks, 236 of which were being implemented in developing countries. Most of these reforms (77%) strengthened CBI garriga2016central, though some also weakened it. For instance, the law governing the Reserve Bank of Australia was changed in 2002. While previously the governor and board members were appointed by the governor general; in 2002 appointment power was given to the treasurer, which produced a lower independence score. Moreover, whereas board members had been appointed for exactly five years before this amendment, after the amendment the term was specified as not exceeding five years at the discretion of the appointing person dincer2014central.

Despite the broad impact of the policy advice to make central banks more independent, the empirical evidence in support of it remains controversial. We investigate the effect of CBI on inflation with a causal framework that treats countries as units in a longitudinal (panel) setup. The data set we use in our analysis was created specifically for this purpose and extends the data set from BRV_Inflation. To describe and address relevant confounding structures, the crucial question is: What are possible reasons that motivate the decision of a country to adopt a certain degree of CBI? What macroeconomic factors drive central bank independence farvaque.2002, aghion2004endogenous, polillo2005globalization, brumm2006effect, brumm2011inflation, bodea2015price, romelli.2018, masciandaroRomelli.2019? Four arguments stand out:

enumerate[i)] • Political institutions: Federally organized countries with good checks and balances grant their monetary institutions greater autonomy and thus a greater level of central bank independence deHaan.95, Moser.1999, farvaque.2002. • Political instability: Central bank reforms are more likely to follow elections, which lead to political consolidation or to changes in the political orientation of the government romelli.2018. cukierman.1995 find that de facto CBI, as measured by the turnover rate of the central bank governor, is lower in less stable political systems. • Past inflation: crowe2008central show that over the period 1990–2003, greater changes in CBI have occurred in countries originally characterized by lower levels of independence and higher inflation. This finding is strengthened by the research of masciandaroRomelli.2019 where it is shown that countries which experienced long periods of inflation are characterized by a higher inflation aversion, which may cause the government to grant a higher level of CBI. According to wachtel.2020 the arguments in favor of an independent central bank began to crystallize in the 1980s after a decade or more of traumatic inflationary experience that put a spotlight on central bank policymaking and its failures. • International pressure: Binding agreements with international money lenders like the International Monetary Fund or the World Bank often require countries to commit to a particular set of policies blejer2002inflation, gutierrez2003inflation, polillo2005globalization, rodrik2006goodbye, romelli.2018, kern2019imf, reinsberg2020bad. According to dincer2014central, countries with less developed financial markets, more open economies and countries that have participated in IMF programs have more independent central banks. Similarly, romelli.2018 finds that countries receiving an IMF loan or becoming a member of a currency union adopt reforms that increase CBI. Another type of external pressure can come from regional clustering, which is often found to be cohesive of certain types of reform processes such as democratisations and economic liberalisations simmons2004globalization, elhorst2013impact, giuliano2013democracy, acemoglu2019democracy.

Those arguments inform our causal model and estimation strategies in Section (ref).

Methodological Framework

Notation

We consider panel data with $n$ units (i.e., countries in our case) studied over time ($t=0,1,\ldots,T$). At each time point $t$, we observe an outcome $Y_t$, an intervention of interest $A_t$ and several time-dependent covariates $L^j_t$, $j=1,\ldots,q$, collected in a set $\mathbf{L}_t=\{L^1_t,\ldots,L^q_t\}$. Variables measured at the first time point ($t=0$) are denoted as $\mathbf{L_0}=\{L^1_0,\ldots,L^{q_0}_0\}$ and are called “baseline variables”. The intervention and covariate histories of a unit $i$ (up to and including time $t$) are $\bar{A}_{t,i}=(A_{0,i},\ldots,A_{t,i})$ and $\bar{L}^s_{t,i}=(L^s_{0,i},\ldots,L^s_{t,i})$, $s=1,...,q$, respectively, with $q,q_0 \in \mathbb N$.

We are interested in the counterfactual outcome $Y_{t,i}^{\bar{a}_{t}}$ that would have been observed at time $t$ if unit $i \in \{1,\ldots,n\}$ had received, possibly contrary to the fact, the intervention history $\bar{A}_{t,i}=\bar{a}_t$. For a given intervention $\bar{A}_{t,i}=\bar{a}_t$, the counterfactual covariates are denoted as $\bar{\mathbf{L}}_{t,i}^{\bar{a}_{t}}$. If an intervention depends on covariates, it is dynamic. A dynamic intervention ${d}_t(\mathbf{\bar{L}}_{t})=\bar{d}_t$ assigns treatment ${A}_{t,i} \in \{0,1\}$ as a function $\mathbf{\bar{L}}_{t,i}$. If $\mathbf{\bar{L}}_{t,i}$ is the empty set, the treatment $\bar{d}_t$ is static. We use the notation $\bar{A}_t = \bar{d}_t$ to refer to the intervention history up to and including time $t$ for a given rule $\bar{d}_t$. The counterfactual outcome at time $t$ related to a dynamic rule $\bar{d}_t$ is $Y_{t,i}^{\bar{d}_t}$, and the counterfactual covariates at the respective time point are $\bar{\mathbf{L}}_{t,i}^{\bar{d}_t}$. More specific notation concerning the data analysis is given in Section (ref).

Likelihood

If we assume a time ordering of $\mathbf{L}_t \rightarrow A_t$ at each time point, use $Y_T$ as the outcome, and define $Y_t$, $t<T$, to be contained in $\mathbf{L}_t$, the data can be represented as $n$ iid copies of the following longitudinal data structure:

eqnarray*[eqnarray* omitted — 119 chars of source]

Note that in Section (ref), in the data analysis, the ordering of variables is different. However, for the given ordering, we can write the respective likelihood $\mathcal{L}(O)$ as

eqnarray*[eqnarray* omitted — 718 chars of source]

In the above factorization, $p_0(\cdot)$ refers to the density of $P_0$ (with respect to some dominating measure) and $A_{-1} := \mathbf{L}_{-1} := \emptyset$. If an order for $\mathbf{L}_t$ is given, e.g., $L^1_t \rightarrow \ldots \rightarrow L^q_t$, a more refined factorization is possible. In line with the notation of other papers (e.g., Tran:2019), we define the $q$-portion of the likelihood to also contain the outcome: ${q}_{0,\mathbf{L}_t} := \tilde{q}_{0,\mathbf{L}_t} \times p_0(Y_{T,i}|\bar{A}_{T-1,i},\bar{\mathbf{L}}_{T-1,i})$. Similarly, we define $g_0 := \prod_{t=0}^T g_{0,A_t}$ and $q_0 := \prod_{t=0}^{T} {q}_{0,\mathbf{L}_t}$.

On the distinction between the causal and statistical model

Estimating causal effects cannot be established from data alone but requires additional structural (i.e. causal) assumptions about the data\--gen\-erating process. Therefore, any causal analysis comes with both a structural (i.e. causal) and a statistical model. The former can be represented by a directed acyclic graph (DAG), which encodes conditional independence assumptions and is logically equivalent to a (non-parametric) structural equation framework. Ideally, the structural model is supported by knowledge from the literature. The statistical model encodes assumptions about the family of possible observed data distributions associated with the DAG, with the ultimate aim to estimate post-intervention distributions and quantities. With doubly robust effect estimation, any parametric assumptions are typically eschewed to avoid model mis-specification; and to incorporate machine learning while retaining valid inference. In our framework and analyses below, we proceed as follows: for the causal model, we begin with the basic assumption that variables can be affected by the past, but not the future (Section (ref)). In our analysis in Section (ref), we then make more detailed assumptions with respect to the causal model: we encode our structural assumptions in a DAG (Figure (ref)) and support this model with references from the economic literature (Appendix). For the statistical model, we first don't impose any parametric restrictions on the statistical model (Section (ref)). In the analysis (Section (ref)), we then use the above likelihood factorization and targeted maximum likelihood estimation with super learning, to avoid any overly restrictive parametric assumptions.

Statistical Model

In line with the notation of Section (ref), we consider a statistical model $\mathcal{M}=\{P = q \times g: q \in \mathcal{Q}, g \in \mathcal{G}\}$ for the true distribution $P_0$ that requires minimal (parametric) assumptions. In contrast to many medical applications, we do not impose restrictions on this model; that is, $A_t$ and $Y_t$ are not deterministically determined for any given data history. Once an intervention is implemented, it can be stopped at any time point and potentially started again. Similarly, the outcome can be observed at any time point, and we do not assume that censoring is possible.

Causal Model

Causal assumptions about the data-generating process are encoded in the model $\mathcal{M}^{\mathcal{F}}$. This nonparametric (structural equation) model states our assumptions about the time ordering of the data and the causal mechanism that gave rise to the data. Thus far, it relates to

eqnarray*[eqnarray* omitted — 289 chars of source]

where $\mathbf{U}:=(U_{Y_T},\boldsymbol{U_{\boldsymbol{L_{t}}}},U_{A_t})$ are unmeasured variables from some underlying distribution $P_{\boldsymbol{U}}$. For now, we do not make any assumptions regarding $P_{\boldsymbol{U}}$. However, in the data example further below, we need to enforce some restrictions on this distribution. The functions $f_{O}(\cdot)$ are (deterministic) nonparametric structural equations that assume that each variable may be affected only by variables measured in the past and not those that are measured in the future. Section (ref) refines the causal model for the data-generating process of the motivating question and represents any additional assumptions made in a DAG.

Causal Target Parameter and Identifiability

In this paper, we focus on the differences in intervention-specific means, i.e., in target parameters such as

eqnarray[eqnarray omitted — 145 chars of source]

If we set the intervention according to a static or dynamic rule ($\bar{A}_t=\bar{d}_t^l$ $\forall t$) with $l \in \{j,k\}$ in the causal model $\mathcal{M}^{\mathcal{F}}$, we obtain the post-intervention distribution $P_0^{\bar{d}_t^l}$. The counterfactual outcome $Y_{T}^{\bar{d}_t^l}$ is the one that would have been observed had $A_t$ been set deterministically to $0$ or $1$ according to rule $\bar{d}_t^l$. We thus restrict the set of possible interventions to those where the intervention is binary $A_{t,i} \in \{0,1\}$.

It has been shown that target parameters of the form ((ref)) can be identified under the (partly untestable) assumptions of consistency, conditional exchangeability and positivity, which are defined below. Specifically, it follows from the work of Bang:2005 and vanderLaan:2012 that given these three assumptions, using the iterative conditional expectation rule, and for the particular time-ordering as defined in Section (ref), we can write the target parameter as

eqnarray[eqnarray omitted — 693 chars of source]

The assumptions of consistency, conditional exchangeability and positivity have been discussed in the literature in detail (Daniel:2011, Daniel:2013, Robins:2009, Young:2011, Tran:2019). Briefly, consistency is the requirement that $Y^{\bar{d}_t}_T = Y_T$ if $\bar{\mathbf{A}}_{t-1} = \bar{d}_{t-1}$ and $\bar{\mathbf{L}}_{t}^{\bar{d}_{t}}=\bar{\mathbf{L}}_{t}$ if $\bar{\mathbf{A}}_{t-1} = \bar{d}_{t-1}$. Conditional exchangeability requires the counterfactual outcome under the assigned treatment rule to be independent of the observed treatment assignment, given the observed past: $Y^{\bar{d}_t}_T\coprod {\mathbf{A}_{t-1}|\bar{\mathbf{L}}_{t-1}, \bar{\mathbf{A}}_{t-2}}$ $\forall \bar{\mathbf{A}}_t=\bar{d}_t, \bar{\mathbf{L}}_t=\bar{\mathbf{l}}_t, \forall t$, and positivity says that each unit should have a positive probability of continuing to receive the intervention according to the assigned treatment rule, given that this has been done so far, and irrespective of the covariate history: $P(\mathbf{A}_t=\bar{d}_t|\bar{\mathbf{L}}_t=\bar{\mathbf{l}}_t,\bar{\mathbf{A}}_{t-1}=\bar{d}_{t-1})>0$ $\forall t,\bar{d}_t,\bar{\mathbf{l}}_t$ with $P(\bar{\mathbf{L}}_t=\bar{l}_t,\bar{\mathbf{A}}_{t-1}=\bar{d}_{t-1}) \neq 0$.

In principle, (conditional) exchangeability can be evaluated graphically for an assumed structural model represented in a DAG using the back-door criterion (Pearl:2010, Molina:2014); i.e., by closing all back-door paths and by nonconditioning on descendants of the intervention. For multiple time-point interventions, a generalized version of this criterion can be used to verify conditional exchangeability. This requires blocking all back-door paths from $A_t$ to $Y_T$ that do not go through any future treatment node $A_{t+1}$ (hernan2017causal). More generally, it has been suggested to use single-world intervention graphs to verify exchangeability, particularly to evaluate identification for complex dynamic interventions. See Richardson and Robins for details (Richardson:2013).

Effect estimation with Longitudinal TMLE

The longitudinal TMLE estimator (vanderLaan:2012) relies on equation ((ref)). To estimate $\psi_{j,k}$, one can separately evaluate each of the two nested expectation terms and integrate out $\bar{\mathbf{L}}_{T-1}$ with respect to the post-intervention distribution $P_0^{\bar{d}_t^l}$. To improve inference with respect to $\psi_{j,k}$, a targeted estimation step at each time point yields a doubly robust estimator of the desired target quantity (see vanderLaan:2011 or Schnitzer:2017 for details). Specifically, we recur to the following algorithm for $t=T,...,1$:

enumerate• Estimate $\bar{Q}_T = \mathbb{E}(Y_T|\mathbf{\bar{A}}_{T-1}, \mathbf{\bar{L}}_{T-1})$ with an appropriate model (for $t=T$). If $t<T$, use the prediction from step 3d (of iteration $t-1$) as the outcome, and fit the respective model. The estimated model is denoted as $\hat{{Q}}_{0,t}$. • Now, plug in $\mathbf{\bar{A}}_{t-1}={\bar{d}}^l_{t-1}$ based on rule $\bar{d}_t^l$, and use the fitted model from step 1 to predict the outcome at time $t$ (which we denote as $\hat{Q}^{\bar{d}_t^l}_{0,t}$). • To improve estimation with respect to the target parameter, update the initial estimate of step 2 by means of the following regression: \begin{enumerate}[a)] • The outcome refers again to the measured outcome for $t=T$ and to the prediction from item 3d (of iteration $t-1$) if $t<T$. • The offset is the original predicted outcome $\hat{Q}^{\bar{d}_t^l}_{0,t}$ from step 2 (iteration $t$). • The “clever covariate” is defined as: \begin{eqnarray} {H}_{t-1} = \prod_{s=0}^{t-1} \frac{I(\bar{{A}}_s=\bar{d}_{s})}{g_{0,A_t=\bar{d}_s^l}} \end{eqnarray} with $g_{0,A_t=\bar{d}_s^l}=P(A_s=\bar{d}_s^l|\bar{\mathbf{L}}=\bar{\mathbf{l}}_s,\bar{A}_{s-1}=\bar{d}_{s-1}^l)$. The estimate of $g_{0,A_t=\bar{d}_s^l}$ is denoted as $\hat{g}_{A_t=\bar{d}_s^l}$. • predict the updated (nested) outcome, $\hat{Q}_{1,t}^{\bar{d}_t^l}$, based on the model defined through 3a, 3b, and 3c. \end{enumerate} This model contains no intercept. Alternatively, the same model can be fitted with $H_{t-1}$ as a weight rather than a covariate Kreif:2017, Tran:2019. In this case, an intercept is required. We follow the latter approach in our implementations. • The estimate for $\mathbb{E}(Y_{T}^{\bar{d}_t^l})$ is obtained by calculating the mean of the predicted outcome from step 3d (where $t=1$). • Confidence intervals can, for example, be obtained using the vector of the estimated influence curve; see Tran:2018 for a review of adequate choices. • Repeat 1.-5. to estimate $\mathbb{E}(Y_{T}^{\bar{d}_t^j})$ and $\mathbb{E}(Y_{T}^{\bar{d}_t^k})$. Now, $\hat{\psi}_{j,k}$ and its corresponding confidence intervals can be calculated.

Inference and Properties of LTMLE

For an arbitrary distribution $P\in \mathcal{M}$ and a specific intervention rule $g = g(P)$ we consider the statistical model $M(g) = \{P^* \in \mathcal{M} : g(P^*) = g\}$ for the respective treatment rules $g$. With such a model we could estimate $\psi ^*$ with the algorithm described in (ref). For $\psi^*$ it can be shown (e.g. van2018targeted) that $\hat{\psi}^*$ is an asymptotically efficient estimator of $\psi^*$ where

equation[equation omitted — 90 chars of source]

The variance can be estimated with the sample variance of the estimated influence curve. This is essentially because the construction of the covariate in step 3c, guarantees that the estimating equation corresponding to the (efficient) influence curve is solved, which in turn yields desirable (asymptotic) inferential properties. The influence curve emerges from the linear span of the scores (i.e. first derivative) of the logistic loss for the density of the outcome variable (evaluated at zero) for a given value of the clever covariate Schnitzer:2014b. Thus, in the longitudinal case, for interventions rules $\bar{g}_{t}$, these score components can be summed across the points in time which yields the efficient influence curve

equation[equation omitted — 239 chars of source]

Data-Adaptive Estimation for Complex (Macroeconomic) Data

The above estimation procedure is doubly robust, which means that the estimator is consistent as long as either the Q- or the g-models (steps 1 and 3c in the algorithm described above) are estimated consistently (Bang:2005). If both are estimated consistently (at reasonable rates), the estimator is asymptotically efficient because the construction of the covariate in step 3c guarantees that the estimating equation corresponding to the efficient influence curve is solved, which in turn yields desirable (asymptotic) inferential properties (vanderLaan:2011, Schnitzer:2017).

To estimate the conditional expectations in the algorithm, one could use (parametric) regression models. Under the assumption that they are correctly specified, this approach would be valid. However, in the context of complex macroeconomic data, as in our motivating example below, it is challenging to estimate appropriate parametric models because of small sample sizes, a large number of relevant variables and complex nonlinear relationships. Longitudinal TMLE can (in contrast to many competing estimation techniques) incorporate machine learning algorithms while still retaining valid inference to reduce the possibility of model misspecification. However, in the settings presented below, machine learning approaches need to be tailored to the specific problem and address the following challenges:

enumerate[i)] • Complexity: Macroeconomic relationships are often highly nonlinear and have various interactions of higher order, which need to be modeled in a sophisticated manner while taking into account the time ordering of the data. • Dispensable variables: The inclusion of covariates in the estimation procedure that are not required for identification, i.e., do not block any back-door paths, can potentially be harmful even if they are not colliders or mediators schnitzer2016variable; that is, the inclusion of such variables can increase the finite-sample variance and lead to small estimated probabilities of following a particular treatment rule given the past, which may be both incorrectly interpreted as positivity violations and make the updating step in the TMLE algorithm unstable. They may also amplify bias Pearl:2011. • p$\bm{>}$n: For longitudinal macroeconomic data, the number of parameters is often larger than the sample size. This is because for long follow-up, the whole covariate history needs to be considered, interactions may be nonlinear, and different variables may have different scales and features that need to be modeled adequately. Consequently, one needs to either reduce the number of parameters with an appropriate estimation procedure or eliminate variables beforehand using variable screening. It has been argued that screening of variables is inevitable to facilitate estimation with LTMLE in many settings schnitzer2016variable.

Section (ref) recommends possible approaches to tackle these challenges in common macroeconomic settings.

Data Analysis: Estimating the Effect of Central Bank Independence on Inflation

Data

We accessed databases of the World Bank and the International Monetary Fund to collect annual data for economic, political and institutional variables. Our outcome of interest is inflation in 2010 ($Y_{2010}$). All covariates are measured annually at equidistant points in time for $t^{\ast}=1998,\ldots,2010$. The intervention variable is central bank independence at time $t^{\ast}$ (CBI, $A_{t^{\ast}}$), which we define as suggested by dincer2014central: their CBI index measures several dimensions of independence and runs from 0, the lowest level of independence, to 1, the highest level of independence. It contains considerations such as the independence of the chief executive officer (CEO) and limits on his/her reappointment, the bank's independence in terms of policy formulation, its objective or mandate, the stringency of limits on lending money to the public sector, measures of provisions affecting (re)appointment of board members other than the CEO, restrictions on government representation on the board, and intervention of the government in exchange rate policy formulation. This definition implies that our central bank independence index is an intervention that can in principle be modified through legislative amendments, although it is the very nature of an index to represent multiple facets of a phenomenon that cannot not be easily dealt with in an actual experiment. We binarized dincer2014centrals' index at a value of 0.45 by setting countries with a value greater than 0.45 to 1 (independent) for each time point and 0 (dependent) otherwise. We then used the binarized index for estimation. The trajectories of their original indices and our binarized version can be seen in Figure (ref). Our outcome variable is defined as the year-on-year changes (expressed as annual percentages) of average consumer prices measured by a consumer price index (CPI). A CPI measures changes in the prices of goods and services that households consume. To calculate CPIs, government agencies conduct household surveys to identify a basket of commonly purchased items and then track the cost of purchasing this basket over time. The cost of this basket at a given time, expressed relative to a base year, is the CPI, and the percentage change in the CPI over a certain period is referred to as consumer price inflation, the most widely used measure of inflation. Our measured covariates are $\mathbf{L}_{t^{\ast}}=\{L^1_{t^{\ast}},\ldots,L^{18}_{t^{\ast}}\}$ and include a variety of macroeconomic variables such as money supply, energy prices, economic openness, institutional variables such as central bank transparency and monetary policy strategies, and political variables (see Figure (ref), Table (ref) and BRV_Inflation for details.). In line with the notation of Section (ref), we consider $Y_{t^{\ast}}$, $t^{\ast}<T=2010$, to be part of $\mathbf{L}_{t^{\ast}}$, i.e., we define $L^8_{t^{\ast}} := Y_{t^{\ast}}$.

Our aim was to include as many countries as possible in our analysis. This entailed a tradeoff between the number of countries and the completeness of the data set. We were able to collect annual data from 1998 to 2010 for 124 countries for 14 explanatory variables and for the dependent variable $Y_{t^{\ast}}$. We further derived growth rates and other indicators from those measured variables to capture data for all 18 covariates ($\mathbf{L}_{t^{\ast}}$). Some of the data were missing, however. To decide whether the missing data were likely missing not at random (MNAR) and therefore possibly not useful without making additional assumptions, we examined countries' characteristics. We decided that observations for certain variables, countries or groups of countries had to be excluded because they were not available; for instance, sometimes wars, insufficiently developed institutions, social unrest or other reasons made the collection of data impossible. We split the data set according to our assessment of whether the observation was MNAR. Data that we regarded as missing at random (MAR) (2.7% of the data set) were multiply imputed using Amelia II (AmeliaII), taking the time-series cross-sectional structure of the data into account. We did not impute data that were likely MNAR. However, some variables that were categorized as MNAR were used in the analysis (e.g., CBI). As a result, we obtained observations for 60 countries and 13 points in time (i.e., calendar years 1998-2010) for 19 measured variables ($L_{t^{\ast}}^1$,\ldots,$L_{t^{\ast}}^{7}$,$L_{t^{\ast}}^9$,\ldots,$L_{t^{\ast}}^{18}$,$Y_{t^{\ast}} \equiv L^8_{t^{\ast}}$,$A_{t^{\ast}}$). In this final data set, 0.1% of observations were missing and thus imputed.

According to the World Bank's income classification, approximately 20% of the remaining 60 countries are low-income countries, 36% belong to the lower-middle-income category, 27% to the upper-middle-income category and 17% belong to the high-income category. While our sample reflects considerable heterogeneity with respect to countries' development level, it is possible that the included countries are not representative of all countries in the world: as many excluded countries faced periods of violent conflicts or had no well-developed governmental institutions, our sample likely reflects economies of (reasonably) stable countries.

Target Parameters and Interventions

Our target parameters are ATEs as defined in ((ref)). To be more specific, consider the following three interventions, of which two are static and one dynamic, each of them applied to $\forall t^\ast \in \{1998,\ldots,2008\}$:

flalign*\bar{d}_{t^\ast}^{1} \phantom{\quad\,\,\bar{L}^1_{t^{\ast}-1}} &= \,\, \left\{a_{t^\ast}=1\right.&
flalign*\bar{d}_{t^\ast,i}^{2} (\bar{L}^8_{t^\ast-1}) &=\left\{ \begin{array}{cl} a_{t^\ast,i}=1 & \quad if \quad \widehat{median}({L^8_{t^\ast-1,i},\ldots,L^8_{t^\ast-7,i}}) \leq 0 \quad or \quad \widehat{median}({L^8_{t^\ast-1,i},\ldots,L^8_{t^\ast-7,i}}) \geq 5\\ a_{t^\ast,i}=0 & \quad otherwise \end{array} \right. &
flalign*\bar{d}_{t^\ast}^{3} \phantom{\,\,\quad\bar{L}^1_{t^\ast-1}} &= \,\, \left\{a_{t^\ast}=0\right.&

A country's central bank is set to be either independent (i.e., $\bar{d}_{t^\ast}^{1}$) or dependent (i.e., $\bar{d}_{t^\ast}^{3}$) during the whole time period under the first and third intervention above. This means that we intervene on the first 11 (i.e. from 1998-2008) out of 13 points (i.e. from 1998-2010) in time. This is because we assume a two-year lag between the CBI intervention and its effect on inflation. The transmission mechanism of monetary policy is said to exhibit “long and variable” lags friedman.1972,batini_nelson.2001,Goodhart.2001. In line with this view, inflation-targeting central banks have adopted a value between 12 and 24 months as transmission lag (the horizon at which the response of prices becomes the strongest). Theoretical models usually imply transmission lags of similar length taylor.2012. According to a meta analysis of 67 published studies for 30 different countries rusnak.2013 the average transmission lag is 29 months. However, transmission lags are longer in developed economies (25-50 months) than in post-transition economies (10-20 months). Overall, after filtering out effects of misspecifications, the results suggest that prices bottom out approximately two-and-a-half years after a monetary contraction. Given the heterogeneity of our countries, we chose a somewhat shorter lag to take into account countries' differences in their stage of development.

The second (dynamic) intervention sets a country's central bank to be independent if its median inflation rate in the past 7 years was below 0% or greater than 5%. The rationale for this relates to the fact that excessive inflation and deflation over several years are considered to produce harmful effects on a country's economy (see, e.g., 10.2307/1910352, fisher1933debt). To guarantee price stability, which excludes inflation beyond a certain level and deflation, an independent central bank is required. Over the last twenty years, the optimal level of inflation has been associated with approximately 2% mishkin2017rethinking. If a country's inflation is constantly well above this level, in our case 5%, it will change the status of its central bank towards independence. The same holds for an inflation rate systematically falling below a value of zero. Note that for the dynamic intervention $\bar{d}_{t^\ast,i}^{2}$, data prior to 1998 had to be collected and utilized.

We define the following two target parameters:

eqnarray[eqnarray omitted — 220 chars of source]

The first, $\psi_{1,3}$, quantifies the expected difference in inflation two years after the last intervention (i.e., in 2010) if every country had an independent central bank for 11 years in a row compared to a dependent central bank for 11 consecutive years. The second, $\psi_{2,3}$, quantifies the effect that would have been observed if every country's central bank had become independent for time points when the country's median inflation in the preceding 7 years had been outside the range from 0 to 5, compared to a strictly dependent central bank for 11 consecutive years (i.e., for the period 1998-2008).

Statistical and Causal Model (DAG)

We separate the measured variables into blocks. The first block comprises $\mathbf{L}^{A}_{t^{\ast}}:=\{L^{1}_{t^{\ast}},\ldots,L^{7}_{t^{\ast}}, L^{9}_{t^{\ast}},\ldots, L^{15}_{t^{\ast}}\}$, and the second comprises $\mathbf{L}^{B}_{t^{\ast}}:=\{L^{16}_{t^{\ast}},\ldots,L^{18}_{t^{\ast}}\}$. In line with Sections (ref) and (ref), we do not make any overly restrictive assumptions with respect to our statistical model. First, we assume that our data come from a general true distribution $P_0$ and are ordered such that

eqnarray*[eqnarray* omitted — 257 chars of source]

In the context of our application, we do not need to make any deterministic assumptions regarding our intervention assignment: a central bank can, in principle, be independent or dependent at any point in time, irrespective of the country's history -- and thus be intervened upon.

As discussed in Section (ref), we assume that each variable may be affected only by variables measured in the past and not those that are measured in the future. In addition, we make several assumptions regarding the data-generating process, which are summarized in the DAG in Figure (ref). Not all variables listed in $O$ are needed during estimation; see Section (ref).

The DAG contains both measured variables (in grey colour) and unmeasured variables (in white colour). The outcome variable is coloured in green, and the intervention in red.

The DAG summarizes our knowledge of the transmission channels of monetary policy. An arrow $A \rightarrow B$ reflects our belief, corroborated by economic theory, that $A$ may cause $B$, whereas an absence of such an arrow states that we assume no causal relationship between the respective two variables. Figure (ref) has been developed based on economic theory. For example, arrow number 6 describes the causal effect from real GDP (Output) on one component of companies' price setting (Price Markup), which is motivated by the fact that changes in demand (c.p.) in the goods market enable companies to set higher prices in a profit-maximizing environment. Detailed definitions of the considered variables, as well as detailed justification for the assumptions encoded in our DAG, are given in Tables (ref) and {(ref)} in the Appendix as well as in Section (ref).

Identifiability Considerations

The DAG shows the causal pathways through which CBI can affect consumer prices and thus ultimately inflation. We next explain the main paths from the intervention node to consumer prices. An independent central bank sets its policy tools autonomously to achieve its objective(s). Moreover, an independent central bank is less pressured to pursue an overly expansionary monetary policy that would produce only high inflation. Such a central bank is more likely to live up to its word, which increases its credibility (arrow 74). Higher credibility keeps inflation expectations in check (arrow 32). The more contained inflation expectations are, the lower the demands for nominal wage compensation will be (arrow 75), which, in turn, keeps labor costs (arrow 29), production costs (arrow 23) and companies' prices (arrow 3) low. This will ultimately also be reflected in relatively low consumer prices (arrow 2). Another pathway from the intervention to the outcome acts through monetary policy decisions. Following an intervention, monetary policy makers' time preferences are reduced (arrow 69), and this will be taken into account in their monetary policy decisions (arrow 49). Monetary policy decisions are mirrored in money supply (arrow 52), which is tantamount to banks' loan creation (arrow 66) and, as a result, affects firms' investment decisions (arrow 67) and thus output (arrow 11). The final stage affects firms' markups (arrow 6) in their prices with a final effect on consumer prices (arrows 4 and 2).

\newgeometry{ left=0.2cm, right=0.2cm, top=0.2cm, bottom=0.6cm }

landscape\begin{figure} \scalebox{0.53}{ \begin{tikzpicture} \tikzstyle{VertexStyle} = [shape = ellipse, draw,minimum size = 2pt,inner sep=0pt] \SetUpEdge[lw = 1pt, color = black, labelcolor = white] \tikzset{bigbox/.style={draw, inner sep=5pt,label={[align=center,shift={(5ex,0ex)}]south:\llap{#1}}}} \SetUpVertex[FillColor=green!20] \Vertex[Math,L={Consumer\ Prices_{t^\ast}\ (Y_{t^\ast})},x=43,y=0] {Inflation} \SetUpVertex[FillColor=white] \Vertex[Math,L={Consumption\ Tax_{t^\ast}},x=43,y=2] {VAT} \Vertex[Math,L={Pricing\ by\ Companies_{t^\ast}},x=35,y=0] {PC} \Vertex[Math,L={Price\ Mark-Up_{t^\ast-1}},x=35,y=-2] {Mark Up} \Vertex[Math,L={Market\ Power_{t^\ast-1}},x=43,y=-2] {Market Power} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Output_{t^\ast-1}\ (L^{5}_{t^\ast-1})},x=35,y=-8] {GDP} \SetUpVertex[FillColor=white] \Vertex[Math,L={Savings_{t^\ast-1}},x=31,y=-15] {Savings} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Age\ Structure_{t^\ast-1}\ (L^{1}_{t^\ast-1})},x=31,y=-18] {Aging} \SetUpVertex[FillColor=white] \Vertex[Math,L={Nominal\ Wages_{t^\ast-1}},x=35,y=-16] {Last Nominal Wage} \Vertex[Math,L={Investments_{t^\ast-1}},x=25,y=-16.5] {Investments} \Vertex[Math,L={Tobin's\ q_{t^\ast-2}},x=19,y=-16.5] {Tobin} \Vertex[Math,L={Firms'\ net\ worth_{t^\ast-2}},x=19.2,y=-20] {FirmValue} \Vertex[Math,L={AS\ &\ MH_{t^\ast-2}},x=19.1,y=-23] {MarketEfficiency} \Vertex[Math,L={Firms'\ liquid._{t^\ast-2}},x=30,y=-22.4] {Liquidity} \Vertex[Math,L={Asset\ Prices_{t^\ast-2}},x=21,y=-18] {Assets} \Vertex[Math,L={Consumption_{t^\ast-1}},x=35,y=-11] {Consumption} \Vertex[Math,L={Disposable\ Income_{t^\ast-1}},x=35,y=-14] {Disposable Income} \Vertex[Math,L={Taxes\ and\ Social\ Securities_{t^\ast-1}},x=39,y=-17] {Last TSS} \Vertex[Math,L={Wealth_{t^\ast-1}},x=35,y=-12.5] {Wealth} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Public\ Debt_{t^\ast-2}\ (L^{3}_{t^\ast-2})},x=18,y=-14] {PastDebt} \Vertex[Math,L={Public\ Debt_{t^\ast-2}\ (L^{3}_{t^\ast-2})},x=7,y=-22] {PastDebt2} \Vertex[Math,L={Public\ Debt_{t^\ast-1}\ (L^{3}_{t^\ast-1})},x=26.5,y=-12.5] {PublicDebt} \SetUpVertex[FillColor=white] \Vertex[Math,L={Debt\ Management_{t^\ast-2}},x=9,y=-24.5] {DebtManagement} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Prim.\ Balance_{t^\ast-1}\ (L^{2}_{t^\ast-1})},x=26.5,y=-10.5] {BudgetBalance} \SetUpVertex[FillColor=white] \Vertex[Math,L={Fiscal\ Spending_{t^\ast-1}},x=30,y=-9] {Fiscal Spending} \Vertex[Math,L={Fiscal\ Revenue_{t^\ast-1}},x=17.5,y=-10.5] {Fiscal Revenue} \Vertex[Math,L={Taxes\ and\ Social\ Security_{t^\ast-1}},x=17.5,y=-12] {TASS1} \Vertex[Math,L={Consumption\ Tax_{t^\ast-1}},x=17.5,y=-9] {VAT1} \Vertex[Math,L={Net\ Exports_{t^\ast-1}},x=43,y=-16] {Net Exports} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Foreign\ Output_{t^\ast-1}\ (L^{4}_{t^\ast-1})},x=43,y=-12] {Foreign Output} \SetUpVertex[FillColor=white] \Vertex[Math,L={Real\ Exchange\ Rate_{t^\ast-2}},x=43,y=-19] {Past X-Rate} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Trade\ Openness_{t^\ast-2}\ (L^{10}_{t^\ast-2})},x=3,y=-28] {TradeOpenness} \SetUpVertex[FillColor=white] \Vertex[Math,L={Share\ of\ Non-Tradables_{t^\ast-2}},x=14,y=-26] {ShareNonTrade} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Past\ Inflation_{t^\ast-2}\ (L^{9}_{t^\ast-2})},x=0,y=-25.5] {MovingAverage} \Vertex[Math,L={Consumer\ Prices _{t^\ast-2}\ (L^{8}_{t^\ast-2})},x=42,y=-24.5] {Past Prices} \SetUpVertex[FillColor=white] \Vertex[Math,L={Production\ Cost_{t^\ast-1}},x=28,y=0] {Production Cost} \Vertex[Math,L={Non-Labor\ Costs_{t^\ast-1}},x=28,y=-1.5] {NLC} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Energy\ Prices_{t^\ast-1}\ (L^{7}_{t^\ast-1})},x=18,y=-1.5] {Energy Price} \SetUpVertex[FillColor=white] \Vertex[Math,L={Technological\ Progress_{t^\ast-2}},x=28,y=-3] {Machine Productivity} \Vertex[Math,L={Technological\ Progress_{t^\ast-1}},x=28,y=-4] {Machine Productivity1} \Vertex[Math,L={Technological\ Progress_{t^\ast, \dots t+8}},x=30.5,y=-5.5] {Machine ProductivityFuture} \Vertex[Math,L={Labor\ Costs_{t^\ast-1}},x=18,y=0] {LC} \Vertex[Math,L={Taxes\ and\ Social\ Security_{t^\ast-1}},x=18,y=2] {TASS} \Vertex[Math,L={Nominal\ Wages_{t^\ast-1}},x=9,y=0] {Nominal Wages} \Vertex[Math,L={Bargaining\ Power_{t^\ast-2}},x=9,y=-1.5] {EBP} \Vertex[Math,L={Bargaining\ Power_{t^\ast-1}},x=9,y=-2.7] {EBP1} \Vertex[Math,L={Labor\ Productivity_{t^\ast-1}},x=12.5,y=-4.5] {LProductivity} \Vertex[Math,L={Labor\ Productivity_{t^\ast, \dots t+8}},x=12.5,y=-8] {LProductivityFuture} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Output\ Gap_{t^\ast-1}\ (L^{6}_{t^\ast-1})},x=18,y=-4] {BC1} \SetUpVertex[FillColor=white] \Vertex[Math,L={Labor\ Unions_{t^\ast-1}},x=7,y=-4.5] {LUnion} \Vertex[Math,L={Human\ and\ Public\ Capital_{t^\ast-1}},x=21,y=-8] {RAD} \Vertex[Math,L={Human\ and\ Public\ Capital_{t^\ast-10,\dots , t-2}},x=18,y=-6] {RADT2} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Past\ Inflation_{t^\ast-2}\ (L^{9}_{t^\ast-2})},x=0.5,y=-2] {MovingAverage1} \Vertex[Math,L={Inflation\ Expectations_{t^\ast-2}\ (L^{17}_{t^\ast-2})},x=4.5,y=-6] {Inflation Expectation} \SetUpVertex[FillColor=white] \Vertex[Math,L={CB\ Credibility_{t^\ast-2}},x=4.5,y=-10] {CBC} \SetUpVertex[FillColor=white] \Vertex[Math,L={Exchange-Rate\ Regime_{t^\ast-2}},x=10.5,y=-17.5] {ERA} \Vertex[Math,L={Targeting\ Regime_{t^\ast-2}},x=4.5,y=-17] {Targeting} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Money\ Supply_{t^\ast-2}\ (L^{16}_{t^\ast-2})},x=4.5,y=-19] {MoneySupply1} \SetUpVertex[FillColor=white] \SetUpVertex[FillColor=black!20] \Vertex[Math,L={CBT_{t^\ast-2}\ (L^{13}_{t^\ast-2})},x=1,y=-12] {CBT} \SetUpVertex[FillColor=red!20] \Vertex[Math,L={CB\ Independence_{t^\ast-2}\ (A_{t^\ast-2})},x=0,y=-21] {CBI} \SetUpVertex[FillColor=white] \Vertex[Math,L={Time\ Preference_{t^\ast-2}},x=7.5,y=-21] {TimePref} \Vertex[Math,L={Govern.\ Dec.{t^\ast-2}},x=-1.5,y=-18] {Government1} \Vertex[Math,L={Int. Press.{t^\ast-2}},x=0,y=-11] {IntPress} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Pol.\ Instit._{t^\ast-2}\ (L^{14}_{t^\ast-2})},x=0,y=-9] {PolInstitution} \Vertex[Math,L={Pol.\ Instab._{t^\ast-2}\ (L^{15}_{t^\ast-2})},x=1,y=-7] {Instability} \Vertex[Math,L={GDP\ p.c._{t^\ast-2}\ (L^{12}_{t^\ast-2})},x=17.1,y=-18.75] {gdppc} \SetUpVertex[FillColor=white] \Vertex[Math,L={Money\ Demand_{t^\ast-2}},x=19,y=-26.75] {MoneyDemand} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Consumer\ Prices _{t^\ast-2}\ (L^{8}_{t^\ast-2})},x=12,y=-28] {PastPrices1} \Vertex[Math,L={Money\ Supply_{t^\ast-2}\ (L^{16}_{t^\ast-2})},x=25,y=-27] {MoneySupply} \SetUpVertex[FillColor=white] \Vertex[Math,L={Nominal\ Exchange\ Rate_{t^\ast-2}},x=37,y=-21.5] {X-Rate} \Vertex[Math,L={Nominal\ Interest\ Rate_{t^\ast-2}},x=25,y=-21.5] {InterestRate} \Vertex[Math,L={Real\ Interest\ Rate_{t^\ast-2}},x=25,y=-20] {RealInterestRate} \SetUpVertex[FillColor=black!20] \SetUpVertex[FillColor=white] \Vertex[Math,L={MP\ Decision_{t^\ast-2}},x=14,y=-22] {MPDecision} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Capital\ Openness_{t^\ast-2}\ (L^{11}_{t^\ast-2})},x=1.5,y=-24.5] {CapitalOpenness} \SetUpVertex[FillColor=white] \Vertex[Math,L={Currency\ Competition_{t^\ast-2}},x=8.5,y=-23] {CurrencyCompetition} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Output_{t^\ast-2}\ (L^{5}_{t^\ast-2})},x=18.5,y=-28.5] {PastGDP} \SetUpVertex[FillColor=white] \Vertex[Math,L={Nominal\ Interest\ Rate_{t^\ast-3}},x=25.5,y=-28.5] {PastInterestRate} \SetUpVertex[FillColor=black!20] \Vertex[Math,L={Bank\ Loans_{t^\ast-2}\ (L^{18}_{t^\ast-2})},x=32.5,y=-23.5] {BankLoans} \SetUpVertex[FillColor=white] \Vertex[Math,L={Consumption_{t^\ast-1}},x=39.5,y=-23.5] {Consumption1} \tikzset{EdgeStyle/.style={->,dashed}} \tikzset{EdgeStyle/.style={->}} \Edge[label=$1$](VAT)(Inflation) \Edge[label=$2$](PC)(Inflation) \Edge[label=$3$](Production Cost)(PC) \Edge[label=$4$](Mark Up)(PC) \Edge[label=$5$](Market Power)(Mark Up) \Edge[label=$6$](GDP)(Mark Up) \Edge[label=$7$](InterestRate)(Liquidity) \Edge[label=$8$](Net Exports)(GDP) \Edge[label=$9$](Fiscal Spending)(GDP) \Edge[label=$10$](Consumption)(GDP) \Edge[label=$11$](Investments)(GDP) \Edge[label=$12$](Disposable Income)(Savings) \Edge[label=$13$](Fiscal Spending)(RAD) \Edge[label=$14$](Tobin)(Investments) \Edge[label=$15$](Investments)(RAD) \Edge[label=$16$](GDP)(BC1) \tikzset{EdgeStyle/.style={->}} \Edge[label=$17$](RADT2)(Machine Productivity1) \Edge[label=$17$](RAD)(Machine Productivity1) \Edge[label=$17$](RAD)(Machine ProductivityFuture) \tikzset{EdgeStyle/.style={->,relative=false,in=250,out=175}} \Edge[label=$18$](RAD)(LProductivity) \tikzset{EdgeStyle/.style={->}} \Edge[label=$18$](RADT2)(LProductivity) \Edge[label=$18$](RAD)(LProductivityFuture) \Edge[label=$19$](Machine Productivity1)(BC1) \Edge[label=$20$](Machine Productivity)(NLC) \Edge[label=$21$](NLC)(Production Cost) \Edge[label=$22$](Energy Price)(NLC) \Edge[label=$23$](LC)(Production Cost) \Edge[label=$24$](TASS)(LC) \Edge[label=$25$](LProductivity)(EBP1) \Edge[label=$26$](BC1)(EBP1) \Edge[label=$27$](LUnion)(EBP1) \Edge[label=$28$](EBP)(Nominal Wages) \Edge[label=$29$](Nominal Wages)(LC) \tikzset{EdgeStyle/.style={->,relative=false,in=340,out=0}} \Edge[label=$30$](Targeting)(Inflation Expectation) \tikzset{EdgeStyle/.style={->}} \Edge[label=$31$](Fiscal Revenue)(Fiscal Spending) \Edge[label=$32$](CBC)(Inflation Expectation) \Edge[label=$33$](Foreign Output)(Net Exports) \Edge[label=$34$](Disposable Income)(Wealth) \Edge[label=$35$](Last Nominal Wage)(Disposable Income) \Edge[label=$36$](Fiscal Spending)(BudgetBalance) \Edge[label=$37$](BudgetBalance)(PublicDebt) \tikzset{EdgeStyle/.style={->,relative=false,in=220,out=20}} \Edge[label=$38$](DebtManagement)(MPDecision) \tikzset{EdgeStyle/.style={->}} \Edge[label=$39$](Last TSS)(Disposable Income) \Edge[label=$40$](Past X-Rate)(Net Exports) \tikzset{EdgeStyle/.style={->,relative=false,in=270,out=90}} \Edge[label=$41$](Past Prices)(Past X-Rate) \tikzset{EdgeStyle/.style={->}} \Edge[label=$42$](Fiscal Revenue)(BudgetBalance) \Edge[label=$43$](Savings)(Investments) \Edge[label=$44$](MoneySupply)(InterestRate) \tikzset{EdgeStyle/.style={->,relative=false,in=260,out=20}} \Edge[label=$45$](MoneyDemand)(InterestRate) \tikzset{EdgeStyle/.style={->}} \Edge[label=$46$](RealInterestRate)(Investments) \Edge[label=$47$](InterestRate)(X-Rate) \Edge[label=$48$](X-Rate)(Past X-Rate) \Edge[label=$49$](TimePref)(MPDecision) \tikzset{EdgeStyle/.style={->,relative=false,in=270,out=90}} \Edge[label=$50$](Targeting)(CBC) \tikzset{EdgeStyle/.style={->}} \Edge[label=$51$](ERA)(CBC) \Edge[label=$52$](MPDecision)(MoneySupply) \Edge[label=$53$](PastDebt)(PublicDebt) \tikzset{EdgeStyle/.style={->,relative=false,in=325,out=140}} \Edge[label=$54$](MPDecision)(Targeting) \tikzset{EdgeStyle/.style={->}} \Edge[label=$55$](MPDecision)(ERA) \Edge[label=$56$](MoneyDemand)(MPDecision) \Edge[label=$57$](PastInterestRate)(MoneyDemand) \Edge[label=$58$](PastGDP)(MoneyDemand) \tikzset{EdgeStyle/.style={->,relative=false,in=343,out=75}} \Edge[label=$59$](MPDecision)(Inflation Expectation) \tikzset{EdgeStyle/.style={->}} \Edge[label=$60$](Wealth)(Consumption) \tikzset{EdgeStyle/.style={->,relative=false,in=180,out=90}} \Edge[label=$61$](Savings)(Wealth) \tikzset{EdgeStyle/.style={->}} \Edge[label=$62$](RealInterestRate)(Assets) \Edge[label=$63$](Assets)(Tobin) \tikzset{EdgeStyle/.style={->}} \Edge[label=$66$](MoneySupply)(BankLoans) \tikzset{EdgeStyle/.style={->,relative=false,in=290,out=90}} \Edge[label=$67$](BankLoans)(Investments) \tikzset{EdgeStyle/.style={->}} \Edge[label=$68$](CBT)(CBC) \Edge[label=$69$](CBI)(TimePref) \Edge[label=$70$](TradeOpenness)(ShareNonTrade) \Edge[label=$71$](ShareNonTrade)(MPDecision) \Edge[label=$72$](MovingAverage1)(Inflation Expectation) \tikzset{EdgeStyle/.style={->,relative=false,in=120,out=225}} \Edge[label=$101$](MovingAverage1)(Government1) \tikzset{EdgeStyle/.style={->}} \tikzset{EdgeStyle/.style={->,relative=false,in=250,out=0}} \Edge[label=$65$](MovingAverage)(MPDecision) \tikzset{EdgeStyle/.style={->}} \Edge[label=$74$](CBI)(CBC) \tikzset{EdgeStyle/.style={->,relative=false,in=180,out=90}} \Edge[label=$75$](Inflation Expectation)(Nominal Wages) \tikzset{EdgeStyle/.style={->}} \Edge[label=$76$](Assets)(FirmValue) \Edge[label=$77$](FirmValue)(MarketEfficiency) \Edge[label=$78$](MarketEfficiency)(BankLoans) \Edge[label=$79$](CapitalOpenness)(CurrencyCompetition) \Edge[label=$80$](CurrencyCompetition)(MPDecision) \Edge[label=$81$](PolInstitution)(CBC) \Edge[label=$82$](Instability)(CBC) \tikzset{EdgeStyle/.style={->,relative=false,in=110,out=190}} \Edge[label=$59$](Instability)(Government1) \tikzset{EdgeStyle/.style={->,relative=false,in=100,out=210}} \Edge[label=$99$](PolInstitution)(Government1) \tikzset{EdgeStyle/.style={->}} \Edge[label=$100$](IntPress)(PolInstitution) \tikzset{EdgeStyle/.style={->,relative=false,in=260,out=90}} \Edge[label=$86$](Government1)(CBT) \tikzset{EdgeStyle/.style={->,relative=false,in=90,out=270}} \Edge[label=$98$](Government1)(CBI) \tikzset{EdgeStyle/.style={->}} \Edge[label=$83$](gdppc)(MPDecision) \Edge[label=$84$](Aging)(Savings) \Edge[label=$85$](InterestRate)(RealInterestRate) \Edge[label=$87$](Liquidity)(MarketEfficiency) \Edge[label=$88$](PastPrices1)(MoneyDemand) \Edge[label=$89$](TASS1)(Fiscal Revenue) \Edge[label=$90$](VAT1)(Fiscal Revenue) \Edge[label=$91$](BudgetBalance)(Disposable Income) \tikzset{EdgeStyle/.style={->,relative=false,in=340,out=90}} \Edge[label=$92$](ERA)(Inflation Expectation) \tikzset{EdgeStyle/.style={->}} \Edge[label=$93$](ERA)(MoneySupply1) \Edge[label=$94$](Targeting)(MoneySupply1) \Edge[label=$95$](PastDebt2)(MPDecision) \Edge[label=$96$](BankLoans)(Consumption1) \tikzset{EdgeStyle/.style={->,relative=false,in=260,out=90}} \Edge[label=$97$](RealInterestRate)(Savings) \tikzset{EdgeStyle/.style={->}} \tikzset{EdgeStyle/.style={->,relative=false,in=240,out=20}} \Edge[label=$64$](Assets)(Savings) \tikzset{EdgeStyle/.style={->}} \tikzset{EdgeStyle/.style={->,relative=false,in=200,out=170}} \tikzset{EdgeStyle/.style={->}} \end{tikzpicture} } \caption{DAG containing the structural assumptions about the data generating process for a specific time point $t^{^\ast} = 2000,\ldots,2010$. The target quantity is $\psi_{j,k}$ and relates to $Y_{2010}$, which refers to $Consumer\ Prices_{t^\ast}$ colored in green. The intervention rules relate to CBI at time $t^{^\ast}-2$, colored in red. Measured covariates are grey, and unmeasured covariates are white. A justification of the DAG is given in Appendix (ref).} \end{figure}

\restoregeometry

There are several back-door paths from the intervention to the outcome. They all start with arrow 98 because CBI status is ultimately influenced by government decisions, which are in turn affected by past inflation, political institutions and political stability see Section (ref) or a detailed justification. As an example, consider the back-door path that goes through government decisions (arrow 98) and past inflation (arrow 101): the latter affects current monetary policy decisions (arrow 65). Monetary policy will in turn impact the formation of inflation expectations (arrow 59) or the money supply (arrow 52). Along edges 66, 67, 11, 6, 4 and 2, this affects the outcome.

Under the assumption that the DAG as motivated in Appendix (ref) is correct, establishing identification in terms of the (generalized) back-door criterion requires the following considerations: all back-door paths start with arrow 98 and can be blocked by conditioning on the following 4 variables: past inflation ($L_{t^{\ast}}^{9}$), central bank transparency ($L_{t^{\ast}}^{13}$), political institution ($L_{t^{\ast}}^{14}$) and political instability ($L_{t^{\ast}}^{15}$). There are various paths from the intervention to the outcome that start with edges 69, 49 and 52. All those paths contain mediators one should not necessarily condition on in our example because otherwise the effect of CBI on inflation through these paths would be blocked (hernan2017causal). The same considerations apply to the paths starting with edges 74 and 32.

In summary, our DAG suggests that all back-door paths from $A_{t^\ast}$ to the outcome (that do not go through any future treatment node $A_{t^{\ast}+1}$) can be blocked by including $L_{t^{\ast}}^{9}$, $L_{t^{\ast}}^{13}$, $L_{t^{\ast}}^{14}$ and $L_{t^{\ast}}^{15}$ in the analysis. As many other variables lie on a mediating path from the intervention to the outcome (i.e., are descendants of $A_{t^{\ast}}$), they should not be conditioned upon.

We argue that the developed DAG should serve as the basis for identification considerations and estimation strategies. However, in complex macroeconomic situations, violations of this causal model need to be taken into account, and other estimation strategies may also be useful. We now explain how this can be facilitated.

Data-adaptive Estimation with longitudinal TMLE

We can, in principle, follow the algorithm described in Section (ref) to estimate the target quantity of interest. This includes estimation of the (nested) outcome model $\bar{Q}_{t^\ast}$ (step 1) and the intervention model $g_{0,A_{t^\ast}=\bar{d}^l_s}$ (step 3c) for each time point. That is, we can estimate the $g$-model for $t^\ast=1998, \ldots,2008$ and $Q_{t^{\ast}}$ for $t^\ast = 2000,\ldots,2010$. As mentioned above, the DAG assumes a 2-year lag before an independent central bank can potentially affect the outcome. It is thus sufficient to estimate the first Q-model in 2000 given the assumed lag structure in the DAG. We define $Y_T := Y_{2010}$, which corresponds to the value of inflation in 2010, while $\bar{d}_{t^{\ast}}^{1}$, $\bar{d}_{t^{\ast},i}^{2} (\bar{L}^8_{t^{\ast}-1})$ and $\bar{d}_{t^{\ast}}^{3}$ are the interventions targeting CBI as described in Section (ref).

We consider three approaches to covariate inclusion. The first is based on the identifiability considerations related to our DAG, and the other two refine variable inclusion criteria based on the scenario in which some structural causal assumptions in the DAG may be incorrect.

enumerate[i)] • DAG-based approach (PlainDAG): Based on the identifiability arguments from Section (ref), $\mathbf{\bar{L}}_{t^{\ast}}$ contains only the relevant baseline variables from 1998 that were measured prior to the first intervention node, as well as $L_{t^{\ast}}^{9}$, $L_{t^{\ast}}^{13}$, $L_{t^{\ast}}^{14}$ and $L_{t^{\ast}}^{15}$. • Greedy super learning approach (ScreenLearn): This approach contains the full set of measured variables $\mathbf{L}_{t^\ast}$. This approach assumes that each variable could potentially lie on a back-door path but that this was undiscovered due to misspecification of the causal model. For example, a researcher who argues that bank loans directly affect a central bank's independence (i.e., that there is an arrow from bank loans to CBI) would have to consider a back-door path along arrows 67, 11, 6, 4, 2 and thus include public debt in $\mathbf{L}_{t^\ast}$. Similarly, if it is doubted that some variables are not necessarily mediators but rather confounders on a back-door path that exists due to unmeasured variables, e.g., $\text{\textit{CBI}} \leftarrow \text{\textit{unmeasured variable}} \rightarrow \text{\textit{Output}} \rightarrow \ldots \rightarrow \text{\textit{Consumer Prices}}$, then measured variables such as Output (real GDP) would also have to be included in $\mathbf{L}_{t^\ast}$. We suggest that an analysis that includes all measured variables in $\mathbf{L}_{t^\ast}$ can serve as a useful sensitivity analysis to explore the extent to which effect estimates may change under different assumptions. • Economic theory approach (\textit{EconDAG}): A further approach, termed \textit{EconDAG}, includes only variables that are measured during a particular 2-yearly transmission cycle, as defined by our DAG. That is, for the Q-model at $t^{\ast}$, every measured variable between $t^{\ast}-2$ and $t^{\ast}-1$ is included, while for the estimation of the g-model at $t^{\ast}-2$, only variables during the respective cycle are considered. As above, given the assumed time ordering, only variables from the past, and not from the future, are utilized in the respective models.

Given the complexity of the data-generating process, it makes sense to use machine learning techniques to estimate the respective g- and Q-models. For a specified set of learning algorithms and a given set of data, the method minimizing the expected prediction error (as estimated by $k$-fold cross validation) could be chosen. As the best algorithm in terms of prediction error may depend on the given data set, it is often recommended to use super learning instead -- and this is what we use for i), ii) and iii). Super learning van2007super (or “stacking”, Breiman:1996) considers a set of learners; instead of picking the learner with the smallest prediction error, one chooses the convex combination of learners that minimizes the $k$-fold cross validation error (for a given loss function, we use $k=10$). The weights relating to this convex combination can be obtained with non-negative least squares estimation (which is implemented in the $R$-package SuperLearner, Polley:2017). It can be shown that this weighted combination will perform asymptotically at least as well as the best algorithm, if not better, given that no correctly specified parametric model is contained in the set of learners (vanderLaan:2008).

As described in Section (ref), the challenge of model specification, including the choice of appropriate learners and screening algorithms, is to address the complex nonlinear relationships in the data and the $p>n$ problem.

Our strategy is to use the following algorithms: the arithmetic mean of the outcome; generalized linear models (with main terms only and including all two-way interactions); Bayesian generalized linear models with an independent Gaussian prior distribution for the coefficients; classification and regression trees; multivariate adaptive (polynomial) regression splines; generalized additive models; Breimans' random forest; generalized boosted regression modeling; and single-hidden-layer neural networks. The algorithms are carefully chosen to reflect a balance between simple and computationally efficient strategies and more sophisticated approaches that are able to model highly nonlinear relationships and higher-order interactions that may be prevalent in the data. Furthermore, parametric, semiparametric and nonparametric approaches were applied to allow for enough flexibility with respect to committing to parametric assumptions. In particular, tree-based procedures were chosen to handle challenges that frequently come with economic data -- for instance outliers. In addition, since some of the continuous predictors are transformed by the natural logarithm, this strict monotone transformation may affect its variable importance in a regression-based procedure, while trees are not impaired in that respect.

For strategies i)-iii), we use the following learning and screening algorithms:

enumerate[a)] • Screening algorithms: Used only for estimation approach ii) because of the large covariate set compared to the sample size; we used the elastic net elasticNet, the random forest randomForest, Cramer's V (with either 4 or 8 variables selected at a maximum) and the Pearson correlation coefficient. The screening algorithms were chosen such that at least a subset of them could handle both categorical and quasi-continuous variables well. • Learning algorithms: The 11 learning algorithms mentioned above are the same for estimation strategies i) and iii). i) and iii) were thus estimated with 11 algorithms each. In contrast, strategy ii) additionally benefited from the 5 screening algorithms mentioned in a) where each screening algorithms was run prior to each learning algorithm. We omitted generalized boosted regression modeling from the learner set such that $50 = 5 \times (11 - 1)$ algorithms (i.e. $\{\text{Screener},\text{Learner}\}$ tuples) emerged. In addition, learning algorithms that are applicable in the $p>n$ case were added without prior screening to the 50 tuples. As a result, when Breimans’ random forest and single-hidden-layer neural networks were added without screening, 52 algorithms could be used for strategy ii); see also Figure (ref) in the Appendix.

All estimates have been obtained using the ltmle package in $R$ (LTMLEPackage).

Results

Descriptive summaries of the data are given in the appendix, in Figures (ref)-(ref). They show the variables' distribution over time. Between 1998 and 2010 most measured variables show interesting patterns and changes. For example, one can see a continuously aging population in the countries included, as well as increased levels of central bank transparency. There is support in the data for all three treatment strategies, with 23 countries having an independent central bank throughout the whole time period, 27 countries never having an independent central bank for the period considered and 16 countries which experienced periods with a negative median inflation rate or median inflation above 5% in the last seven years during 1998 and 2010, while having legislated an independent central bank during the same time period (Figure (ref)).

A naive analysis comparing the mean reductions in inflation between 2000 and 2010 between those countries that had an independent central bank (from 1998 to 2008) and those that had a dependent central bank led to the following results: the mean reduction was 2.3 percentage points for those with an independent central bank, compared to 1.0 percentage points for those with a dependent central bank. This equates to a difference of 1.3 percentage points (95% CI: -6.1; 3.5). However, such a crude comparison does not allow a causal interpretation and is not an estimate of $\psi_{1,3}$.

The results of the analyses described in Section (ref) are visualized in Figure (ref).

Our main analysis (PlainDAG) suggests that if a country had legislated CBI for every year between 1998 and 2008, it would have had an average increase in inflation of 0.01 (95% confidence interval (CI): -1.48; 1.50) percentage points in 2010. The other two approaches led to slightly different results: -0.44 (95% CI: -2.38; 1.59) for ScreenLearn and 0.01 (95% CI: -1.46; 1.47) for EconDAG.

figure[figure omitted — 204 chars of source]

Similarly, when considering the estimation strategy PlainDAG, we can conclude that if a country had legislated an independent central bank for every year when the median of the past seven years of inflation had been above 5% or below 0% from 1998 to 2008, it would have achieved an average reduction in inflation of 0.07 percentage points (95% CI: -1.29; 1.15) in 2010 compared to a central bank that was independent during the same time span (that is, dichotomized CBI = 0). The other two strategies suggest somewhat stronger inflation reductions.

Our findings can be summarized as follows: First, depending on the degree of structural assumptions imposed, we find that an independent central bank has either a negative or no effect on inflation. Second, as suggested by the confidence intervals, we cannot exclude the possibility of a strong negative or positive ATE. Third, the largest estimated ATE (in absolute terms) amounts to -0.61 percentage points (EconDAG). From a monetary policy perspective this can be considered as substantial, given that our study period covers an era of overall low to moderate inflation (characterized by a median inflation rate of about 4%).

For a sensitivity analysis, we stratified our sample according to the World Bank's income classification into high income (n = 26) and low income (n = 34) countries and reran all analyses. The results are reported in the appendix (cf. Figure (ref) and (ref)). For high-income countries, the ATE (averaged across estimation strategies) is slightly positive. In contrast, for low-income countries, where inflation has typically been higher, we obtain almost no effect for the static treatment strategy (i.e. $\hat{\Psi}_{1,3}$) and an average reduction of about -0.4 percentage points for the dynamic treatment strategy $\hat{\Psi}_{2,3}$. However, due to the small sample sizes, these results need to be interpreted with caution.

The diagnostics for all analyses are given in Table (ref) and Figure (ref) in the Appendix. The cumulative product of inverse probabilities was never below the truncation level of $0.01$, which was re-assuring (Petersen:2012). The maximum value of clever covariates, as defined in ((ref)), was always well below 5, which suggests that the chosen super learning approach worked well. However, the mean clever covariate, which is supposed to be broadly approximately 1, was not ideal for dynamic treatment strategy 2, suggesting that $\psi_{2,3}$ should be interpreted cautiously.

table[table omitted — 2,026 chars of source]

Figure (ref) (Appendix) visualizes the learner weight distribution. In our analysis, a multitude of learners and screening algorithms were important, including neural networks, random forests, regression trees and Bayesian generalized linear models.

Simulations

Motivated by our data analysis, we explore the extent to which model misspecification and choice of learner sets may affect effect estimation with longitudinal maximum likelihood estimation (and competing methods).

Data-Generating Processes

We specified two data-generating processes: a simple one with 3 time points and one time-dependent confounder and a more complex one with up to 6 time points and 10 time-varying variables.

For the first simulation (Simulation 1), we assume the following time ordering:

eqnarray*[eqnarray* omitted — 65 chars of source]

Using the $R$-package simcausal (Sofrygin:2016), we define preintervention distributions as listed in Table (ref) (Appendix).

For the second simulation (Simulation 2), we use the following time ordering:

eqnarray*[eqnarray* omitted — 124 chars of source]

We generated the preintervention data according to the distributions specified in Table (ref) (Appendix).

Target Parameter and Interventions

For both simulations, we were interested in evaluating ATEs between two static interventions. That is, we were interested in

flalign*\bar{d}_{t^{+}}^{Sim1,1} \phantom{\quad\,\,\bar{L}^1_{t^{\ast}-1}} &= \,\, \left\{a_{t^{+}}=1 \quad \forall t^{+} \in \{1,2,3\}\right.&
flalign*\bar{d}_{t^{+}}^{Sim1,0} \phantom{\,\,\quad\bar{L}^1_{t^{\ast}-1}} &= \,\, \left\{a_{t^{+}}=0 \quad \forall t^{+} \in \{1,2,3\}\right.&

and

flalign*\bar{d}_{t^{++}}^{Sim2,1} \phantom{\quad\,\,\bar{L}^1_{t^{\ast}-1}} &= \,\, \left\{a_{t^{++}}=1 \quad \forall t^{++} \in \{1,2,3,4,5,6\}\right.&
flalign*\bar{d}_{t^{++}}^{Sim2,0} \phantom{\,\,\quad\bar{L}^1_{t^{\ast}-1}} &= \,\, \left\{a_{t^{++}}=0 \quad \forall t^{++} \in \{1,2,3,4,5,6\}\right.&

The target parameters of interest are thus

eqnarray[eqnarray omitted — 219 chars of source]

Estimations

In our primary analysis, we used longitudinal targeted maximum likelihood estimation for both simulations. In a secondary analysis, we also evaluated the performance of (longitudinal) inverse probability of treatment weighting (see, e.g., Daniel:2013 and the references therein).

For LTMLE, we considered four different estimation approaches, the first for the first simulation and another three for the second simulation:

enumerate[i)] • Estimation as explained in Section (ref). Q- and g-models were fitted with (generalized) linear models. This is estimation approach GLM. • Estimation as explained in Section (ref). Q- and g-models were fitted with a data-adaptive approach using super learning. There were four candidate learners: the arithmetic mean, GLMs, Bayesian generalized linear models with an independent Gaussian prior distribution for the coefficients, as well as classification and regression trees. No screening of variables was conducted. This is estimation approach L1. • Estimation as explained in Section (ref). Q- and g-models were fitted with a data-adaptive approach using super learning. The same four learners as in L1 are utilized; however, variable screening with Pearson's correlation coefficient was conducted. In addition, four more learners were added: multivariate adaptive (polynomial) regression splines friedman1991MARS, generalized additive models, and generalized linear models including the main effects with all corresponding two-way interactions. These additional four learners included variable screening with the elastic net ($\alpha = 0.75$). This is estimation approach L2. • Estimation as explained in Section (ref). Q- and g-models were fitted with a data-adaptive approach using super learning. The eight learning/screening combinations from L2 were used. In addition, single-hidden-layer neural networks were used, once without variable screening and once with elastic net screening. Finally, the last learner is composed of classification and regression with the random forest. This is estimation approach L3.

We also obtained estimates for the ATE based on IPTW. The estimation of the propensity scores was identical to the estimation of the g-models within LTMLE and is thus also based on the estimation procedures described in i)-iv).

Comparisons

We compared the estimated absolute (abs.) bias and coverage probabilities for the estimated ATEs for the two simulations and for both correctly and incorrectly specified Q-models (see details below).

enumerate[i)] • Simulation 1: The incorrect, misspecified, Q-models omit $\mathbf{L}:=(L_1, L_2, L_3)$ entirely. By contrast, the g-models were specified such that the entire covariate histories are taken into account. As a result, if no screening is applied (estimation strategies GLM and L1), all relevant variables are used for estimation; however, with screening (estimation strategies L2 and L3), some variables might be omitted. • Simulation 2: The incorrect, misspecified, Q-models do not use $\mathbf{L^1} := (L^1_1, L^1_2, L^1_3, L^1_4, L^1_5, L^1_6, L^1_7)$ for estimation. Thus, one relevant back-door path remains unblocked, which leads to time-dependent confounding with treatment-confounder feedback. As in simulation 1, all g-models were specified such that the entire covariate histories are taken into account.

Results

The results after 1000 simulation runs are summarized in Figure (ref).

figure[figure omitted — 312 chars of source]

In simulation 1, LTMLE provides approximately unbiased estimates even under misspecified Q-models. This is because targeted maximum likelihood estimation is a doubly robust estimator, and thus misspecification of either the Q- or g-models can be handled. However, the coverage probabilities are too high. See Tran:2018 for a discussion of this issue.

Under the more complex setup of simulation 2, there is small bias if both the Q- and g-models contain the relevant adjustment variables (Both Correct) and learner set L1 is used (Bias = 0.991). The more sophisticated learner sets L2 and L3 yield much better estimates (Bias = 0.158 and 0.144). With incorrect specification of the Q-model, there is again some bias (Bias = 1.438, 0.639, 0.663). Interestingly, for simulation 2, the most complex estimation approach with the largest learner set L3 does not produce a substantial improvement over L2. This highlights that a simple increase in learners does not necessarily improve the finite sample performance of LTMLE, although sufficient breadth and complexity is certainly always needed, as seen by the inferior performance of the first learner set.

In simulation 1, the confidence intervals have too large coverage probabilities. However, in simulation 2, using L2 and L3 yields (close to) nominal coverage probabilities. Nevertheless, our results highlight the need to develop more reliable variance estimators, such that overall better coverage can be achieved.

Note that while LTMLE may produce approximately unbiased point estimates, IPTW does not seem to benefit from complex estimation procedures for the propensity scores (g-models) in the second simulation. The estimates are rather volatile, with some bias and poor coverage probabilities. These conclusions hold for all learner sets considered (Appendix, Figure (ref)).

Conclusions

We have shown that even for complex macroeconomic questions, it is possible to develop a causal model and implement modern doubly robust longitudinal effect estimators. We believe that this is an important contribution in light of the current debate on the appropriate implementation and use of causal inference for economic questions (Imbens:2019). Our suggestion was to commit to a causal model, motivate it in substantial detail (as in Appendix (ref)), discuss possible violations of it, and ultimately conduct sensitivity analyses that evaluate effect estimates under different (structural) assumptions.

While the statistical literature has emphasized the benefits of doubly robust effect estimation in conjunction with extensive machine learning (vanderLaan:2011), its use in sophisticated longitudinal settings has sometimes been limited due to computational challenges and constraints (Schomaker:2019). We have shown how the use of screening and learning algorithms that are tailored to the question of interest can help to facilitate a successful implementation of this approach.

As stressed by Imbens:2019: “[...] models in econometric papers are often developed with the idea that they are useful on settings beyond the specific application in the paper”. We hope that both our causal model, i.e., the DAG, and our proposed estimation techniques will be useful in applications other than ours.

Our simulation studies suggest that LTMLE with super learning can yield good point estimates compared to competing approaches, even under model misspecification. However, both the coverage of confidence intervals and the appropriate choice of learners are challenges that warrant more investigation. Recent research confirms that the development of more robust variance estimators is urgently needed (Tran:2018) and that learner selection is becoming more diverse (Gehringer:2018).

From a monetary policy point of view, we conclude that based on the estimates from the main analysis there is no strong support for the hypothesis that an independent central bank necessarily affects inflation, although our confidence intervals were wide. Making fewer or different structural assumptions, as in our secondary analyses, leads to an average inflation reduction of up to 0.6 percentage points under central bank independence. An ATE of -0.6 percentage points may be seen as substantial, considering that the period on which our estimations are based is overall characterized by low to moderate inflation. However, a naive use of super learning (as in our “ScreenLearn” secondary analysis) may be potentially dangerous because important collider and mediator structures may be overlooked, which can yield different, possibly incorrect results. A comparison of the point estimates from the main and secondary analysis reflects this consideration. As highlighted throughout this paper, while a sophisticated computational approach can be advantageous for doubly robust causal effect estimation, it can not replace the commitment to well-thought-out structural assumptions about the macroeconomic process under consideration.

\addcontentsline{toc}{section}{References}

{}