EconBase
← Back to paper

Identification, estimation and inference in Panel Vector Autoregressions using external instruments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

43,689 characters · 6 sections · 45 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification, estimation and inference in Panel Vector Autoregressions using external instruments

abstractThis paper proposes an identification inspired from the SVAR-IV literature that uses external instruments to identify PVARs, and discusses associated issues of identification, estimation, and inference. I introduce a form of local average treatment effect - the $\mu$-LATE - which arises when a continuous instrument targets a binary treatment. Under standard assumptions of independence, exclusion, and monotonicity, I show that externally instrumented PVARs estimate the $\mu$-LATE. Monte Carlo simulations illustrate that confidence sets based on the Anderson-Rubin statistics deliver reliable convergence for impulse responses. As an application, I instrument state-level military spending with the state's share of national spending to estimate the dynamic fiscal multiplier. I find multipliers above unity, with effects concentrated in the contemporaneous year and persisting into the following year. {\bf{JEL Classification}}: C32, C33. \\ {\bf{Keywords}}: Vector autoregressive models, panel data, identification, instruments.

Panel vector autoregressions (PVAR) are one of the standard tools for the estimation of dynamic causal effects in macroeconomics. However, their causal interpretation is frequently limited. Sometimes they are described as identifying a temporally related causality, “Granger causality”; sometimes, as discussed in pala2024pvarcontrol, they may posses a contemporaneous causal interpretation under either exogeneity assumptions, or a suitable control type of assumption. However, Granger causality is not contemporaneous causality\footnote{In the words of granger2014forecasting: “A better term might be temporally related, but since it is such a simple term we shall continue to use it .}; exogeneity is hard to satisfy due to the rich endogeneity among macroeconomic variables (Nakamura2018a); and the control type of approach proposed by pala2024pvarcontrol may not be satisfied if some units cannot act as controls for some others. However, when all the previous cases fail, instruments could be utilised to carry causal claims.

This paper shows that it is possible to identify causal effects by the means of instrumental variables in the case of panel vector autoregressions. By implementing the same approach as gertler2015monetary,Mertens2013,Mertens2014,StockWatson2018,Olea2021 - defined as Proxy-SVAR or SVAR-IV- in PVARs it is possible to retrieve a causal estimand that I define as $\mu-\text{LATE}$. Normally, the causal literature has a clear target in mind in the case of a dummy policy and a dummy instrument (angrist1996identification) - the LATE. LATEs can be viewed as a special case of principal stratification for which there are 4 categories of compliance status. In the case of continuous instruments the interpretation becomes troublesome because there are infinitely many possible strata (antonelli2023principal). Contrarily, $\mu$-LATEs seemlessy move from continuous variables to a binary interpretation akin to a LATE. In particular, they compare the value of the outcome and policy under a 1% and 0% instrument assignment.

In this special context, the defiers are (under positive monotonicity) units that do not observe an increase (decrease) in their policy variable residuals when the instrument increases (decreases). Moreover, because usually the residuals of a PVAR are assumed to be continuous and normally distributed, the $\mu$- LATE strongly depends on the instrument and policy continuity. This latter feature is rarely discussed by the causal literature but imposes a strong linearity assumption.\footnote{In fact, most of the IV literature has focused on either binary (angrist1996identification) or multi-valued instruments (heckman2001policy,vytlacil2002independence), but rarely discussed how a LATE could emerge from the comparison of of two predicted values of the outcome variables under a continuous instrument case. In general, semi parametric or non parametric estimators tend to be preferred when the data allows it. } This is, however, necessary. Because macroeconomic data is frequently relatively small in the unit component, making use of a linearity assumption to obtain meaningful estimates is often the only possibility. \footnote{This point, and its many drawbacks, are developed in kolesar2024dynamic.}

On the other hand, inference can be carried by simply translating the conclusions of Olea2021 regarding VAR-IVs to their panel counterpart. This implies that the optimal approach to inference is to make use of the Anderson Rubin test statistic (henceforth defined as AR - see Fieller1944,Anderson1949 for its definition) in the presence of one instrument, and, at the current state of the literature, the conditional likelihood ratio test (henceforth defined as CLR - see andrews2006optimal for its definition) in the presence of multiple instruments.

I give a context to the proposed causal framework through an original application in which I compute the dynamic fiscal military multiplier in the United States. Using local level data can be advantageous for several reasons. First, Nakamura2014 argue that in a monetary union, the central bank (the Federal Reserve) cannot raise interest rates in some states relative to others, and federal tax policy is common across states in the union. This means that the open economy multiplier identifies the measure of fiscal policy that does not depend on monetary movements. Second, the advantage of using military spending is that the DD-350 military procurement forms made available from the US Department of Defense provide a clean, direct, and localised measure of government spending. Unfortunately, data of such quality, that starts from such an early age - 1966 - is not available for any other form of spending.

For this reason, I consider a PVAR that includes military state-level spending as instrumented by the states fraction of total national military spending, with GDP growth as the outcome variable. Such IV approach is generally defined as Bartik instrument and is becoming increasingly popular in social sciences, where the use of local-level data is intensifying (see bartik1991benefits,goldsmith2020bartik). My findings indicate that the fiscal multiplier may be about $\sim1.7$ on impact, a value not dissimilar from the ones previously observed by the state-multiplier literature (chodorow2019geographic). The advantage of using PVARs over simple panel linear regressions becomes clear when considering the usefulness of Impulse Response Functions. In this sense, my findings suggest a multiplier of about $1.5$ one year ahead and negligible after, suggesting the existence of large cumulative multipliers.

This paper contributes to several streams of the literature. One natural contribution is at the intersection of the causal inference and the time series fields. In this sense, it complements pala2024pvarcontrol in the research agenda oriented to give a causal interpretation to panel vector autoregressions and, more broadly, the literature that uses the Rubin causal model to motivate the causal interpretation of time series models (menchettibojinov2022,menchetticipollinimealli2022,RambachanSheppard2021,bojinovshephard2019). Second, it contributes to the econometric literature by extending the identification of structural VARs through an external instrument to panel VARs. Third, it indicates a potentially interesting recipe for the interpretation of any estimator in the case of continuous instruments that is different from the current literature. Fourth, it extends the common knowledge of the good coverage properties of the Anderson-Rubin statistics for the impulse response functions generated by a PVAR-IV, indicating a venue for the construction of reliable confidence sets. Fifth, it contributes to the applied literature for the estimation of the dynamic fiscal multipliers.

The discussion is organized as follows. Section (ref) introduces the potential outcome framework and the $\mu-\text{LATE}$. Section (ref) introduces the assumptions required for the identification of the causal effects. Section (ref) introduces the estimation procedure of the PSVAR-IV. Section (ref) discusses the Anderson Rubin statistics\footnote{In appendix (ref) I provide the main theorems and simulations that showcase the good coverage properties of the statistic.}. Section (ref) discusses the estimation of the dynamic fiscal multiplier using the PSVAR-IV.

Finally, allow me to introduce some useful notation. A set of observations of variable $s$ for state $i$ at time $t$, $x_{s,it}$, is part of a larger matrix of variables $\boldsymbol{x}_{s}$, and $\boldsymbol{x}_{2:S}$ indicates the partition from the second to the $S$ variable of such matrix. $\mathbb{E}(X)$ will indicate the expected value of the variable $X$, $var(X)$ will indicate its variance, $cov(X,Y)$ the covariance between variables $X$ and $Y$, and $\boldsymbol{cov}(X,Y)$ their variance-covariance matrix, which contains their variance on the diagonal and their covariance in the off diagonal. For $R$ being a square matrix, $R_{a,b}$ indicates the entry of $R$ in the row $a$ and column $b$. $X\perp Y$ will indicate independence between $X$ and $Y$.

Identification

Let us say the researcher is interested in the estimation of a dynamic causal effect, represented as the impact of a change in a policy variable $W_{i,t}$ on one or many outcome variables $Y_{j,i,t}$. As such, the potential outcomes of the outcome variables can be represented as follows: \[ Y_{j,i,t}(w,z)=Y_{j,i,t}((W_{i,1:t-1},w,W_{i,t+1:T})(Z_{i,1:t-1},z,Z_{i,t+1:T})). \] Such definition is similar to the one introduced by RambachanSheppard2021 and indicates that the potential outcome of the different outcome variables $j$ depend on the policy assignment $w$ and on the instrumental variable $z$\footnote{Such form was first introduced in LATEs by angrist1996identification.}. In the leading example of this paper, there will only be GDP growth as outcome variable ($j=1$), while $W_{i,t}$ will indicate a fiscal policy variable. Moreover, $Z_{it}$ will be the instrumental variable. Finally, the index $i,t$ will refer to region $i$ at time $t$.\footnote{In this paper, I will solely focus on a clean case of one instrument, one policy. This choice stems from the higher likelihood of researchers to come up with identification strategies that involve a clear one insturment, one instrumented, many outcomes scenarios. It can be extended to multiple instruments and multiple instrumented variables. The analytical extensions are provided in Olea2021.} All variables are assumed to be continuous.

I define a $\mu-\text{LATE}_{j}$ as the following causal effect.

defnA $\mu\text{-LATE}_{j}$ is the LATE of receiving assignment $w$ versus assignment $w'$ for those units that comply with the assignment \[ \mu\text{-LATE}_{j} = \mathbb{E}\bigl[ Y_{j}(w) - Y_{j}(w') \,\big|\, W(z)=w \bigr]. \]

Assumptions

The assumptions required to give a causal interpretation to a LATE estimator are independence, exclusion and monotonicity. Here the assumptions are put on the residuals of the PVAR. In particular, define $\widetilde{W}_{it}=\mathbb{E}[W_{it}|\omega_{it}]$ and $\widetilde{Y}_{j,it}=\mathbb{E}[Y_{j,it}|\omega_{it}]$, where $\omega_{it}=(W_{i,t-1}.^{\prime}.. W_{i,t-p}^{\prime},Y_{1,i,t-1}^{\prime},..,Y_{1,i,t-p}^{\prime},Y_{J,i,t-1}^{\prime},..,Y_{J,i,t-p}$ are the regressors. The PVAR estimates the impact of variations in $\widetilde{W}_{it}$ on $\widetilde{Y}_{j,it}$.

The treatments are assumed to be continuous, so that $\widetilde{w}^{\circ}$ is any value extracted from the distribution of $\widetilde{W}$ and $z^{\circ}$ is any value extracted from the distribution of $Z$.

assumption(Independence) For all $\widetilde{w}^{\circ}\in\widetilde{W}$, all $\widetilde{z}^{\circ}\in Z$, all $t\geq1$, all $i\geq1$, and all $j\geq1$, it holds that \begin{equation} \{\widetilde{Y}_{j,i,t}(\widetilde{w}^{\circ},z^{\circ}),\widetilde{W}_{i,t}(z^{\circ})\}\perp Z_{i,t} \end{equation}

Assumption (ref) establishes that the instruments are as good as randomly assigned with respect to any of the outcome variable's residuals or any of the policy variable's residuals. The researcher will normally assume that the potential outcomes are not affected by $Z_{it}$ if not through $\widetilde{W}_{it}$, i.e.

assumption(Exclusion) For all $\widetilde{w},\widetilde{w}^{\prime}\in\widetilde{W}_{i,t}$, $t\geq1$, and $i\geq1$ it holds that \[ \{\widetilde{Y}_{j,it}(\widetilde{w},z)=Y_{j,it}(\widetilde{w},z^{\prime})\} \] where $z,z^{\prime}\in Z_{i,t}$ are any possible combination of values of $Z_{i,t}$.

Notice that one of the consequences of assumption (ref) is that it implies that the potential outcomes of $\widetilde{Y}_{j,it}$ - the residual of the outcome variable - do not depend on the realized value of the instruments, if not by the means of the policy. Such assumption, coupled with assumption (ref), allows for a seamless transition from conditional expected values and realized outcomes to potential outcomes frameworks. In the empirical example, this would mean that the assignment of military expenditures at the federal-level are independent with respect to the potential outcome process of any GDP growth innovations and with respect to any military spending innovations at the state-level.

Finally, monotonicity as in angrist1996identification is required, so that

assumption(Monotonicity) For all $z,z^{\prime}\in Z_{i,t}$, and all $t\geq1$ and $i\geq1$, it holds that either $\widetilde{W}_{i,t}(z)\geq\widetilde{W}_{i,t}(z^{\prime})$ or $\widetilde{W}_{i,t}(z)\leq\widetilde{W}_{i,t}(z^{\prime})$.

Notice that this type of monotonicity assumption needs to hold for every couple of instrument assignments $z,z^{\prime}$. For example, in the case of fiscal multipliers, it means that if there is an increase in total national military spending, the residuals of state-level military spending are either increasing or decreasing for each unit $i$ at each time $t$. Such assumption, differently from a non parametric case, does not allow any discountinuity of the mapping of $\widetilde{W}$ on $Z$.

rem(Weakness of a fully parametric monotonicity assumption). Let us say that $Z$ is multivalued, such that it can only take three integer values $Z=\{z^{0},z^{1},z^{2}\}$. A parametric estimator needs to assume $W(z^{2})\geq W(z^{1})\geq W(z^{0})$. A non parametric estimator could potentially solve this issue by estimating two separate quantities, one for $z^{1}$ and $z^{2}$, and assuming that $W(z^{1})\geq W(z^{0})$, $W(z^{2})\geq W(z^{0})$, but the ordering of $W(z^{2})\underline{?}W(z^{1})$ does not need to be assumed. This comes at the cost of requiring the data to be dense enough around the quantities.

Finally, the instrument needs to be a predictor of the instrumented variable, a condition frequently defined as instrument relevance or instrument strength.

assumption(Relevance). The instrument satisfies $\mathbb{E}[\widetilde{W}_{i,t},Z_{i,t}]\neq0$.
remFrequently, SVAR-IV are thought to estimate a causal effect only under a relevance and an non correlation condition. In macro economics, the non correlation condition is frequently stated as $\mathbb{E}[\widetilde{Y}_{j,it},\widetilde{W}_{i,t}]=0$\footnote{See, among many others, StockWatson2018,Olea2021,bruns2024testing,brignone2023robust.}. Yet, such imposition would only refer to a statistical relationship among the two, and would ignore the benefits of having a potential outcome representation\footnote{In this sense, the condition usually stated is potentially testable (see bruns2024testing), but does not allow to claim the identification of a LATE.}. In practice, this means that SVAR-IV would not be able to claim that a meaningful causal estimand has been identified unless the econometrician identifies a set of assumptions that maps the estimator to an estimand. Typically, the two conditions that are sought for are insufficient to achieve any causal identification.

Estimation

Consider several known outcome variables and one intervention variable aggregated in a process of the kind \[ x_{i,t}=(W_{i,t}^{\prime},Y_{j=1,i,t}^{\prime},Y_{j=2,i,t}^{\prime},..,Y_{j=J,i,t}^{\prime})'. \] Here $W_{i,t}$ could indicate military procurement spending in region $i$ at time $t$, and $j=1,..,J$ could indicate output and other outcomes of interest in region $i$ at time $t$. PVARs are generally represented as processes that depend on their past, a series of unit-specific characteristics, and some random disturbances, which leads to:

equation[equation omitted — 154 chars of source]

Here $\Phi$ denotes an $m\times m$ matrix of slope coefficients\footnote{This framework could be extended to random effects instead of fixed effects. However, such change would have no impact on the nature of the causal effects estimated.}, $\mu_{i}$ is an $m\times1$ vector of individual-specific effects, $\tilde{x}_{i,t}$ is an $m\times1$ vector of disturbances, and $I_{m}$ denotes the identity matrix of dimension $m\times m$. The model can be extended to include higher lags, but to keep the notation compact I will use a one lag representation.

The focus of the following section will be on the disturbances \[ \tilde{x}_{i,t}=(\widetilde{W}_{i,t}^{\prime},\widetilde{Y}_{j=1,i,t}^{\prime},\widetilde{Y}_{j=2,i,t}^{\prime},..,\widetilde{Y}_{j=J,i,t}^{\prime})', \] where the tilde represents the specific disturbance related to the original variable. Such disturbances are distributed according to $\widetilde{x}_{i,t}\sim\mathcal{N}(\boldsymbol{0}_{J+1},\Sigma)$ where $\mathcal{N}$ is a normal distribution. In the rest of the paper will assume that the policy variable goes first\footnote{This is a standard assumption in the SVAR-IV literature (see StockWatson2018) .}, and the outcome variables follow. For example, in the case of fiscal multipliers this would mean that $\widetilde{W}_{i,t}$ refers to the innovations of military spending growth in each region of the US, and $\widetilde{Y}_{1,i,t}$ refers to the innovations of GDP growth in those same regions. Exactly like vector autoregressions, panel vector autoregressions have a contemporaneous causal representation that is commonly defined as panel structural vector autoregression (PSVAR), as follows

equation[equation omitted — 169 chars of source]

Generally, the estimation of the contemporaneous causal effects is carried on transformations of $\Sigma=R^{-1}R^{-1\prime}$, where $R^{-1}$ is the unique lower-triangular Cholesky factor with non-negative diagonal elements. The reduced-form innovations $\widetilde{x}_{it}$ are related to the SPVAR shocks $\eta_{t}$ by an invertible matrix $H$: \[ \widetilde{x}_{i,t}=H\Gamma\eta_{i,t}=R^{-1}\eta_{i,t},\qquad\eta_{i,t}\sim(0,I_{J+1}),\qquad diag(H)=1, \] where $R^{-1}=H\Gamma$, and $\Gamma$ is a diagonal matrix with variance of the shocks in the diagonal entries. The structural shocks $\eta_{i,t}$ are mean zero with unit variance, serially and mutually uncorrelated. Since the autoregressive parameters $\hat{\Phi}$ can be consistently estimated under regularity conditions, the sample residuals $\hat{\widetilde{x}}_{i,t}$ are consistent estimates of $\widetilde{x}_{i,t}$. The empirical SPVAR problem reduces to finding $R$ from $\hat{\Phi}$. But there are $(J+1)^{2}$ parameters in $R$ and the sample covariance of $\hat{\widetilde{x}}_{i,t}$ only provides $(J+1)((J+1)+1)$ conditions in face of $(J+1)^{2}$ parameters to be estimated. The SPVAR is therefore under-identified as there can be infinitely many solutions that satisfy the covariance restrictions.

The IV procedure for the estimation of the structural matrix in SVARs generally corresponds to either one or two separate steps, depending on whether the economist is interested in a unit normalized shock or a standard shock\footnote{See gertler2015monetary,Mertens2013,Mertens2014,StockWatson2018,Olea2021 for a full description of the procedure.}. First, an IV is estimated with the following first and second stages \[

aligned\widetilde{W}_{i,t}=\delta Z_{i,t}+\eta_{i,t} & first stage\\ \widetilde{Y}_{j,i,t}=\beta_{j}\widetilde{W}_{i,t}+\epsilon_{i,t} & second stage(s)

\]

Then, if the economist is interested in a unit normalized shock, the IV estimator simply becomes $\beta_{j}^{IV}=(\rho_{j}/\delta)$, where $\rho_{j}$ comes from the regression $\widetilde{Y}_{j,i,t}=\rho_{j}Z_{i,t}+\nu_{i,t}$, $R_{1,1}=\delta$ ,and $R_{1,j}=\beta_{j}^{IV}$. In the case in which the economist is interested in a standardized shock, instead, the IV estimator becomes $\beta_{j}^{IV}=c_{j}(\rho_{j}/\delta)$.\footnote{Following gertler2015monetary, consider the partition of the covariance matrix of the residuals $R=[R_{1}R_{2}]=\left[

array[array omitted — 50 chars of source]

\right]$, $Q=\frac{r_{21}}{r_{11}}\Sigma_{11}\frac{r_{21}^{\prime}}{r_{11}}-(\Sigma_{21}\frac{r_{21}^{\prime}}{r_{11}}+\frac{r_{21}}{r_{11}}\Sigma_{21}^{\prime})+\Sigma_{22}$, and $r_{12}r_{12}^{\prime}=(\Sigma_{21}-\frac{r_{21}}{r_{11}}\Sigma_{11})^{\prime}Q^{-1}(\Sigma_{21}-\frac{r_{21}}{r_{11}}\Sigma_{11}).$ Then, it follows that, to obtain the structural form, $c_{j}=\sqrt{r_{12}}$ and $R_{1,1}=\delta\cdot c_{1}$ and $R_{1,k}=\rho_{k}c_{k}$.} In this case, the estimator is plugged in the covariance matrix as an affine transformation that depends on a normalization that allows to move to the reduced form and fully identifies the first column, so that $R_{1,1}=\delta c_{1}$ and $R_{1,j}=\beta_{j}c_{j}$. More simply, in the first case the estimator reduces to $\beta_{j}^{IV}=(\rho_{j}/\delta)$; in the second case the normalization provided by the vector $c_{j}$ returns the modified estimator $\beta_{j}^{IV}=\frac{1}{\sqrt{\sigma_{\widetilde{W}_{it}}}}(\rho_{j}/\delta)$. I will focus on the first case as it provides an interpretation that is akin to the one of a standard IV estimator.

Finally, the following normalizing assumptions will be considered to hold across the rest of the paper.

assumption(Normalizing assumptions):\\ (1) $\widetilde{x}_{it}$ and $Z_{it}$ are stationary,\\ (2) $\widetilde{x}_{it}\sim\mathcal{N}(\boldsymbol{0}_{J+1},\Sigma)$ and $Z_{it}\sim\mathcal{N}(\mu_{Z},\sigma_{Z})$.

Notice that assumption (ref) is required for several different reasons. Part (1) allows to consider the estimator of first stage and second stage(s) without violations of the Wold theorem and mean that $R^{-1}$ is invertible and part (2) allows the derivative interpretation of $\delta$ and $\rho_{j}$.

rem(Why normality?) While normality is not a necessary condition, it has some important properties that are convenient when discussing the estimators. In fact, alternative estimators that make different assumptions about the distribution of $Z_{it}$ and $\widetilde{x}_{it}$ may be considered. For example, in the case in which $Z_{it}$ and $\widetilde{W}_{it}$ are treatment dummy indicators, the theory goes back to the traditional case of angrist1996identification, and in the case in which $Z_{it}$ and $\widetilde{W}_{it}$ are multi-valued, the theory goes back to the cases analyzed by vytlacil2002independence,heckman2001policy.

Under the assumptions laid out in section (ref) it can be shown that

thm(PSVAR-IV estimates a ratio of derivatives). Under assumptions (ref);(ref);(ref); (ref);$\beta^{IV}$ estimates \[ \beta_{j}^{IV}=\frac{\delta\mathbb{E}[\widetilde{Y}_{j}(z^{\circ})]/\delta z^{\circ}}{\delta\mathbb{E}[\widetilde{W}(z^{\circ})]/\delta z^{\circ}}. \]

Then, according to theorem (ref), $\beta_{j}^{IV}$ simply captures the ratio of the effect of moving along different values of $z^{\circ}$ on $\widetilde{Y}_{j}$ and on $\widetilde{W}$. Therefore, a useful property of impulse response function can be established according to the following theorem.

thm(Interpretation of the impulse response functions). The immediate impulse response function of a shock in $\widetilde{W}$ captures \[ \hat{\text{IRF}}_{j}=\mu-\text{LATE}_{j}=\mathbb{E}[\widetilde{Y}_{j}(\widetilde{w})-\widetilde{Y}_{j}(\widetilde{w}^{\prime})|\widetilde{W}(z)=\widetilde{w}] \] representing the difference between the shock being equal to $\widetilde{w}$ and $\widetilde{w}^{\prime}$.

Notice that theorem (ref) implies that, considering two different impulse response functions, such as the difference between a $1\%$ and a $0\%$ shock, results in the $\mu-\text{LATE}$ that captures the difference of GDP growth for those units that complies with the national spending growth. Hence, the impulse response captures $\mathbb{E}[\widetilde{Y}_{j}(1\%)|\widetilde{W}(z)=1\%]$, the impact of a one percent deviation in regional military spending growth on GDP growth for those units that observed a spending increase.

Inference

Instrumental variables are useful only as far as they satisfy the relevance condition (assumption (ref)). In fact, it is easy to see that, being the IV estimator \[ \beta_{j}^{IV}=(\rho_{j}/\delta), \] weak identification could be tested by the means of a null hypothesis $H_{0}:\delta=0$. Hence, for $\delta\rightarrow0$, it must be that either $\beta_{j}^{IV}\rightarrow\infty$ or $\beta_{j}^{IV}\rightarrow-\infty$ depending on the sign of $\rho_{j}$. For some time the general consensus has been to carry two different inferential procedures: one in the first stage, by the means of the Cragg-Donald statistic, frequently defined as first-stage F-statsistic (StockStaiger1997); and one, separately, in the second stage, by the means of standard confidence intervals or bootstrap. Such procedure heavily relied on the idea that standard confidence intervals possess asymptotically good coverage properties under the alternative hypothesis ($H_{1}:\delta\neq0$).

However, such approach does not cover situations in which the instrument is weak but satisfies independence. Indeed, an approach that generates confidence intervals on $\beta_{j}^{IV}$ on the basis of the strength of the instrument, even when the first stage coefficient is near zero, may be preferred to one that may end up discarding interesting research hypothesis on the basis of a weak - but independent - instrument. On the basis of such wisdom, the two step approach may be sub optimal compared to the Anderson-Rubin statistic approach (Anderson1949,Olea2021,AndrewsStockSun2019,stock2002testing,mikusheva2006tests,mikusheva2010robust). mikusheva2010robust introduces a procedure for generating confidence sets for the second stage that have good coverage properties even in the null hypothesis case.\footnote{Hence, for $\delta=0$, the confidence set would be distributed between minus and plus infinity.} Olea2021 extend the result by introducing a confidence set for the Impulse Response function generated by a SVAR-IV by using the AR statistic.

While it is known that the Anderson-Rubin confidence sets are optimal in the case of one instrument, there is yet to form a consensus about which approach may be preferred in the case of two or more instruments (some recent advancements include the CLR test of andrews2006optimal).

Appendix (ref) extends the conventional wisdom present in the SVAR-IV field to the PSVAR-IV case using rotations of the Anderson-Rubin statistic and demonstrates the good coverage properties of the AR statistic.

Estimation of the dynamic fiscal multiplier

The fiscal multiplier is generally defined as the coefficient a regression with gdp as the dependent variable and government spending as the independent variable. Such measure is of great interest because of its policy relevance: a relatively large fiscal multiplier is often times evoked by governments as the reason to increase spending\footnote{For example, Nakamura2014 observes that the American Recovery and Reinvestment Act (ARRA) was justified on the basis of large estimates of the fiscal multiplier.}.

The aggregate fiscal multiplier is generally computed using vector autoregressions (romer2010macroeconomic,Blanchard2002) or local projections (Ramey2018). Normally, the aggregate fiscal multiplier was found to be rarely above one.

A local fiscal multiplier could be preferred to the aggregate fiscal multiplier for several reasons. First, the assumptions required to obtain an unbiased estimand are less restrictive than their aggregate counterpart. Indeed, the computation of a national aggregate fiscal multiplier often poses some credibility issues due to the unreliability of the underlying assumptions. The fiscal policy literature has therefore explored the quantification of the impact of a fiscal expansion on GDP by using more granular and localized data, either at the state or regional-level. Second, the open economy multiplier can be potentially more interesting to central bankers because, by using state heterogeneity, it is essentially independent of monetary policy, as overnight rates are fixed for all states.

The recent emergence of this literature has generated different relevant contributions that seem to indicate a regional fiscal multiplier of about 1.5 (farhi2016fiscal,Nakamura2014,shoag2010impact,chodorow2019geographic). To the best of my knowledge, there is currently little work done in estimating the dynamic regional fiscal multiplier with the same type of external instrument approach that has characterized the aggregate data literature. Perhaps closer to the main idea of this paper is dupor2023regional, which estimates the dynamic regional fiscal multiplier using a model to frame the impact of the ARRA. Yet, the ARRA is representative of a particular context of the US economy of low inflation and low interest rates, and may not be representative of different state dependencies.

Motivated by the lack of empirical evidence at the intersection of the two streams of literature, I estimate a regional fiscal multiplier for the US using aggregate national military spending as an instrument for the innovations of regional military spending. Two cautions are however invited to the reader. First, the fiscal multiplier identified using military spending data is particularly useful, but may not be representative of a generic spending multiplier. It is useful as it is inherently a measure of direct spending of the US government (Nakamura2014); but it is not representative as it does not include all the government spending (koo2023impulseresponseinstrumentalvariables.

In the case of panel vector autoregressions, the dynamic regional fiscal multiplier can be estimated by simply by running a PVAR estimation on the vector \[ x_{it}=(\frac{\text{exp}_{it}-\text{exp}_{it-1}}{\text{gdp}_{it-1}}^{\prime},\frac{\text{gdp}_{it}-\text{gdp}_{it-1}}{\text{gdp}_{it-1}}^{\prime})^{\prime}. \] Here, the issue of endogeneity arises because the contemporaneous innovations of fiscal expanses growth may not be thought as exogenous with respect to the contemporaneous innovations of GDP growth.

The leading assumption for the case of PVAR therefore is that the United States do not embark on military buildups because states that receive a disproportionate amount of military spending are doing more poorly than before relative to other states. To exploit this assumption, I use data from the US extracted from the electronic database of DD-350 military procurement forms available from the US Department of Defense by Nakamura2014, which includes military spending for equipment of 10000\$ or more in the period 1966-1984 and above 100000\$ in the period until 2006.\footnote{Unfortunately, the data is not updated any further.} The rest of the analysis follows Nakamura2014 fairly closely: the data is at a yearly frequency, and region refers to the aggregation of different states that are close and not densely populated, resulting in 10 different macro-regions and 39 different years. Differently from the original paper, however, the main estimation makes use of the variable's growth with respect to the previous year, rather than the previous two years; and the Bartik/shift-share instrument is given a preference over using 10 different instruments (one for each state). The reason I made such choice is that conventionally time series regressions are framed in terms of growth with respect to the previous period, and a one dimensional instrument is known to have an optimal confidence set, whereas the case of multiple instruments may provide less reliable confidence sets.\footnote{In the appendix, I show that alternative formulations using growth with respect to two previous periods may change the results slightly. While the quantities tend to be similar in the impulse response, the mechanism by which fiscal expansions tend to be associated with an increase in output in the following year is by the means of a large output autocorrelation, rather than a direct correlation between output and past expanses.}

The model is estimated as follows. First, I use a one-lag model as suggested by the MAIC, MBIC, MHIQ from table (ref).

table[table omitted — 324 chars of source]

The residuals plotted in Figure (ref).

figure[figure omitted — 181 chars of source]

From the figure, there appears to be no indication of residual autocorrelation. This is confirmed by the regression coefficients obtained by regressing the residuals against their lags. The coefficients, being not statistically different from zero, do not seem to suggest to reject the null hypothesis of a statistically significant relationship between the residuals and their lags. Finally, I turn to the assumption of normality of the model, discussed in assumption (ref). In fact, violations of assumption (ref)(iii) would suggest that non-parametric estimators could be preferred over the ones implied by the 2sls utilized for the IV regression because of the assumption that are required from parametric estimators. The histograms displaying the error's distribution in figure (ref) seem to indicate that the residuals may be normally distributed and therefore continuously differentiable with appropriate weights, indicating the adequateness of the normality assumption.\footnote{Several statistical tests (such as shapiro1965analysis,shapiro1972approximate) with the null hypothesis of normality fail to reject the null, indicating that the residuals may indeed be normally distributed.}

figure[figure omitted — 231 chars of source]

The results from the IV regression are instead displayed in Table (ref). The coefficient from the first stage is statistically significantly different from zero, and the first stage F-statistic is above the commonly advocated threshold level of 10 (see StockStaiger1997). Moreover, the AR statistic is above the $\chi_{1,1-\alpha}$ critical value for $\alpha=.05$, suggesting that the results may be statistically significantly different from zero.

table[table omitted — 1,184 chars of source]

Finally, consider figure (ref), which displays the impulse response functions of a 1% shock in fiscal spending growth. The results are similar to the ones of the literature, suggesting a value of the fiscal multiplier of approximately $\sim1.7$ in the first period\footnote{Notice that, by definition, the IRF on impact is the second stage regression in table (ref).}. However, the dynamic fiscal multiplier displays an interesting feature, as it appears that the impact of a change in the fiscal spending in year $t$ results in a corresponding increase in output growth by $\sim1.5$. To better highlight the mechanism by which such response happens, table (ref) displays the AR coefficients. The high correlation between GDP growth and fiscal policy in the previous period is the main mechanism by which the fiscal multiplier can result in a GDP growth that may last for more than one year. Moreover, fiscal policy tends to not be particularly correlated with past fiscal policy or output, resulting in a response close to zero in the second horizon.

table[table omitted — 975 chars of source]
figure[figure omitted — 393 chars of source]

Conclusions

This paper discussed the causal interpretation of panel vector autoregressions identified by the means of external instruments. The IRF generated by a PVAR can estimate a LATE representing the difference between the outcome variable under a treatment and no treatment status for the compilers. However, such LATE needs to be read differently from the panel linear regression literature, as it refers to the residuals and emerges as a counterfactual assignment of different predictions, such as a 1% shock versus a 0% shock. I have discussed under which assumptions the LATE may be captured: independence, exclusion, and monotonicity. Some drawbacks of the proposed identification scheme include the severity of the parametric linear nature of the monotonicity assumption.

Moreover, I discussed the best approaches to conduct inference in a PVAR identified using external instrument. In appendix (ref) I showcase the good small sample properties of the AR confidence sets calibrating a simulation on the basis of the dataset from the application.

Finally, I have applied these tools to the estimation of a dynamic regional fiscal multiplier for the United States, a quantity that has been rarely targeted by the literature. My empirical findings suggest that the dynamic regional fiscal multiplier may be above one in the second period, indicating some possibly longer term effects of fiscal expansions on GDP growth.

Future researchers are invited to develop two points. First, the $\mu-\text{LATE}$ interpretation of the PSVAR-IV relies on an underlying linearity assumption. Yet, non-parametric estimators, which could potentially alleviate the linearity assumption, are never utilized in the SVAR nor the PSVAR literature. If the data utilised is sufficiently large, such methods could be further explored. Second, because the IV literature is still uncertain about which statistics to use when dealing with multiple instruments, the inference issue of overidentification naturally carry to SVAR-IV and PSVAR-IV. Hence, future researchers should properly discuss the unreliability of confidence sets in such cases and possibly implement novel methodologies with better coverage properties.