EconBase
← Back to paper

The causal interpretation of panel vector autoregressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

68,671 characters · 11 sections · 52 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The causal interpretation of panel vector autoregressions

abstractThis paper discusses the different contemporaneous causal interpretations of Panel Vector Autoregressions (PVAR). I show that the interpretation of PVARs depends on the distribution of the causing variable, and can range from average treatment effects, to average causal responses, to a combination of the two. If the researcher is willing to postulate a no residual autocorrelation assumption, and some units can be thought of as controls, PVAR can identify average treatment effects on the treated. This method complements the toolkits already present in the literature, such as staggered-DiD, or LP-DiD, as it formulates assumptions in the residuals, and not in the outcome variables. Such a method features a notable advantage: it allows units to be “sparsely” treated, capturing the impact of interventions on the innovation component of the outcome variables. I provide an example related to the evaluation of the effects of natural disasters economic activity at the weekly frequency in the US.I conclude by discussing solutions to potential violations of the SUTVA assumption arising from interference.

{\bf{JEL Classification}}: C32, C33. \\ {\bf{Keywords}}: Vector autoregressive models, panel data, identification, difference in differences.

\onehalfspacing \thispagestyle{empty}

Introduction

Panel Vector Autoregressions (PVARs) are widely used across the social sciences because they provide an intuitive framework for estimating impulse response functions in large panel datasets. However, their use for causal inference remains limited, typically confined to notions of Granger causality. Recent advances at the intersection of time-series and causal inference show that traditional estimators can, under suitable conditions, admit causal interpretations in the Rubin causal framework (bojinovshephard2019, menchettibojinov2022).

This paper extends that insight to PVARs. I show that PVARs possess a causal interpretation under the same independence conditions discussed in RambachanSheppard2021, and I characterize how the identified causal estimands, such as an Average Treatment Effect (ATE) and Average Causal Response (ACR) depend on the distribution of the policy variable.

The most important contribution of this paper is to show that Panel Vector Autoregressions (PVARs) can also be used to conduct causal inference in the presence of endogenous policies. By exploiting their panel structure, the estimation of a PVAR can be reframed as a dynamic comparison between treated and control units. In this sense, PVARs can serve as a useful and complementary tool to Difference-in-Differences (DiD) methods (Card1990,CardKrueger), as well as to their recent extensions such as staggered DiD (sunabraham2021) and local-projection DiD (dube2023local).

The key intuition is that if the autoregressive component of the model is capable of absorbing the dynamic effects of treatment, then previously treated units can act as valid controls in subsequent periods. This makes it possible to identify causal effects through the recursive estimation of a PVAR without the need for explicit orthogonality or independence assumptions typically required in other frameworks.

This approach also highlights several advantages over conventional DiD estimators. Standard DiD designs usually assume that once a unit is treated, it remains treated forever, restricting the set of possible control observations. Local-projection DiD estimators allow units to re-enter the control group after a certain time, but they impose a specific treatment duration ex ante.

PVARs offer a more flexible perspective. Because their identification focuses on the residuals rather than on mean outcome differences, it suffices to examine the distribution of the innovations to address residual autocorrelation. If treated and control units are comparable after controlling for the autoregressive component, the resulting contrasts can be interpreted as an Average Treatment effect on the Treated (ATT). Importantly, this ATT does not arise from direct comparisons of observed outcomes but from differences in the innovation components of treated and untreated units.

The paper concludes with an application related to the evaluation of the effects of natural disasters and discussing possible solutions to violations of the SUTVA assumption.

Narrative examples are constantly used through the paper to give a proper meaning to the theoretical results, and the discussion is organized as follows. Section (ref) discusses a causal framework for PVARs. Section (ref) defines different types of causal effects.

Section (ref) shows that PVARs can estimate an ATE if the policy is a dummy, an ACR if the policy is continuous, a mix of the two if it is continuous and non-negative. Section (ref) shows that PVARs can estimate ATT if some units are extracted as controls and there is no residual autocorrelation. Section (ref) provides an empirical example of this last result by discussing the causal effect of natural disasters on weekly economic activity in the US. Section (ref) discusses non-compliance.

Section (ref) concludes.

Introduction to Panel Vector Autoregressions

Consider a number of known interventions and outcomes, respectively indexed by $k=1,...,K$ which could indicate fiscal and monetary policy, and $j=1,..,J$ which could indicate output and employment. A PVAR is generally ran on the aggregation of processes \[ x_{it}=(W_{k=1,it}^{\prime},W_{k=2,it}^{\prime},...,W_{k=K,it}^{\prime},Y_{j=1,it}^{\prime},Y_{j=2,it}^{\prime},..,Y_{j=J,it}^{\prime})', \] where $W_{k,it}$ is a variable that indicates the treatment value for unit $i$ by a treatment of type $k$ at time $t$, $Y_{j,it}$ is the outcome variable $j$ for unit $i$ at time $t$.

Panel Vector Autoregressions are generally represented as a process that depends on its past, a series of unit-specific characteristics, and some random disturbances, hence:

equation[equation omitted — 152 chars of source]

where $\Phi$ denotes an $m\times m$ matrix of slope coefficients\footnote{This framework could be extended to random effects instead of fixed effects. However, this change would have no impact on the nature of the causal effects estimated.}, $\mu_{i}$ is an $m\times1$ vector of individual-specific effects, $\tilde{x}_{it}$ is an $m\times1$ vector of disturbances, and $I_{m}$ denotes the identity matrix of dimension $m\times m$. The model can be extended to include higher lags, but to keep the notation compact I will use a one lag representation. \footnote{Any VAR(p) for $p>1$ can always be represented as a VAR(1).} Moreover, through the paper I will make the implicit assumptions that $x_{it}$ is stationary and that the disturbances are assumed to be multivariate iid distributions with mean zero and not necessarily normal. This is because assuming the normality of the distributions could be in opposition to some of the conclusions that will be drawn on the paper.\footnote{I will show later that the normality of the distribution of the first variable in the policy column of $\widetilde{x}_{it}$ is just one case over many possible.}

The focus of the following section will be on the disturbances \[ \tilde{x}_{it}=(\widetilde{W}_{k=1,it}^{\prime},\widetilde{W}_{k=2,it}^{\prime},...,\widetilde{W}_{k=K,it}^{\prime},\widetilde{Y}_{j=1,it}^{\prime},\widetilde{Y}_{j=2,it}^{\prime},..,\widetilde{Y}_{j=J,it}^{\prime})', \] where the tilde represents the specific disturbance related to the original variable. I will assume that the policy variables go first, and the outcome variables follow. This system is in line with the conventional way of thinking about Vector Autoregressions. For this reason there exists one causal effect that is estimated per $j=1,..,J$ that corresponds to every intervention variable $k=1,..,K.$

remThe results that follow can also be obtained from a different system in which the PVAR only includes the outcome variables and which residuals are regressed against a $W_{k,it}$ which has the same distribution as $\widetilde{W}_{k,it}$. This system may present several advantages in the case of a no autocorrelated intervention which include, but are not limited to: (1) a better capability of testing the distribution of $W_{k,it}$ than $\widetilde{W}_{k,it}$, (2) the ability to avoid an overparametrisation.\footnote{Another way of viewing this issue could be to consider this system as the panel equivalent to the one analysed by bojinovshephard2019 for the case of AR and VAR estimators.}
remThe results that follow will be valid only in the case of recursive identification. This is because other methods of inference, such as sign, short run and long run restrictions, are much more difficult to give a causal interpretation if they only partially identify the system.\footnote{This is a known issue in the macroeconomics literature which is a leading cause to several debates around the interpretability of Impulse Response Functions in VARs and how restrictive the restrictions can and should be. For example, the oil shock literature has long discussed about the appropriateness of the IRF generated by a VAR that is partially identified because it is unclear whether the assumption that oil shocks can only have effects within {[}-1.5,0{]} is reasonable in light of Bayesian inference (KilianOil2009,kilianzhou2019,baumeisterhamilton2019). Moreover, sign restrictions have been shown to assume a boundary that is frequently not discussed by the economist but is implicit in the way the rotations are formalized (baumeisterhamilton2015econometrica).} However, methods such as IV or proxy variables (Mertens2014,Olea2021) have been shown to posses a causal interpretation akin to the LATE one by RambachanSheppard2021. A follow up to this paper includes a full description of identification, estimation and inference in PVARs identified through external instruments (pala2024identificationestimationPVARIV).

Causal effects

The object of the estimation of any VAR is a dynamic causal effect. Such effects arise as the difference between the outcome variables under the assignment path $\widetilde{w}_{i,1:t}\in\mathcal{W}^{t}$ and a counterfactual path $\widetilde{w}_{i,1:t}'\in W^{t}$. Hence, the objective of the estimation is the difference $\widetilde{Y}_{it}(\widetilde{w}_{i,1:t})-\widetilde{Y}_{it}(\widetilde{w}_{i,1:t}')$. A reasonable way to discuss the potential outcome process is to represent it as a function of past assignments, potential contemporaneous assignments, and future assignments (see RambachanSheppard2021): \[ \widetilde{Y}_{it}(\widetilde{w}):=\widetilde{Y}_{it}(\widetilde{W}_{i,1:t-1},\widetilde{w},\widetilde{W}_{i,t+1:T}). \]

When it comes to Panel VARs, the contemporaneous effect is usually identified by imposing restrictions on the covariance matrix of the residuals. The $t+h$ effect, for $h\geq1$, is instead generally computed by the means of the Impulse Response Functions.\footnote{This paper will only refer to the causal interpretation of the former, as the interpretation of the dynamic IRF follows trivially.}

Starting from this definition, I will consider the following causal effects:

defn(Causal effects). For $t\geq1$ , and any fixed $\widetilde{w},\widetilde{w}'\in\mathcal{W}$
enumerate• Average Treatment Effect is $\mathbb{E}[\widetilde{Y}_{it}(\widetilde{w})-\widetilde{Y}_{it}(\widetilde{w}')]$, • Average Treatment effect on the Treated is $\mathbb{E}[\widetilde{Y}_{it}(\widetilde{w})-\widetilde{Y}_{it}(\widetilde{w}^{\prime})|\widetilde{W}_{it}=\widetilde{w}]$, • Average Causal Response is $\frac{\delta\mathbb{E}[\widetilde{Y}_{it}(\lambda)]}{\delta\lambda}$ , • Average Causal Response on the Treated is $\frac{\delta\mathbb{E}[\widetilde{Y}_{it}(\lambda)|\widetilde{W}_{it}=\lambda]}{\delta\lambda}$.

The difference $\widetilde{Y}_{it}(\widetilde{w})-\widetilde{Y}_{it}(\widetilde{w}')$ is the so called Impulse effect. The Average Treatment Effect (ATE) is usually the target of analysis, since it is informative of the effect of receiving treatment $\widetilde{w}$ versus receiving the treatment $\widetilde{w}'$.

The average treatment effect on the treated (ATT) therefore measures its expectation, but only conditionally on the subset of time being one in which there is a treatment. It carries a substantially different meaning compared to ATEs since it includes a selection bias term.\footnote{The selection bias is $\frac{cov(\widetilde{Y}_{j,it}(w_{k}),1\{\widetilde{W}_{k,it}=w_{k}\})}{var(1\{\widetilde{W}_{k,it}=w_{k}\})}$.}

Notice that both of these effects assume a structure associated to the policy variable that is a dummy variable that is equal to $\tilde{w}$ for treated units and to $\tilde{w}^{\prime}$ for the untreated ones.

The definitions of the average causal response (ACR) is instead well discussed by callawaysantannabacon2021. It represents the causal effect of changes in the policy variables on the outcome variable. In this context, the researcher could easily derive the effect of a change of a given size in a policy on the outcome variable. Lastly, the concept of average causal response on the treated (ACRT) represents the same idea, but it includes the conditioning factor on the observed times, indicating the presence of a selection bias component. Both of those causal effects can result from continuous policy variables.

The researcher is usually interested in the effects of changes in one specific policy variable to an outcome variable. Hence, the dynamic causal effects for a particular assignment $k$, defined $\widetilde{w}_{k}\in\mathcal{W}_{k}$ on the outcome $j$, are defined as: \[ \widetilde{Y}_{j,it}(\widetilde{w}_{k}):=\widetilde{Y}_{j,it}(\widetilde{W}_{i,1:t-1},\widetilde{W}_{1:k-1,it},\widetilde{w}_{k},\widetilde{W}_{k+1:K,it},\widetilde{W}_{i,t+1:T}). \]

enumerate• Its associated ATE is $\mathbb{E}[\widetilde{Y}_{j,it}(\widetilde{w}_{k})-\widetilde{Y}_{j,it}(\widetilde{w}_{k}')]$, • its associated ATT is $\mathbb{E}[\widetilde{Y}_{j,it}(\widetilde{w}_{k})-\widetilde{Y}_{j,it}(\widetilde{w}_{k}')|\widetilde{W}_{k,it}=\widetilde{w}_{k}]$, • its associated ACR is $\frac{\delta\mathbb{E}[\widetilde{Y}_{j,it}(\lambda_{k})]}{\delta\lambda_{k}}$, • its associated ACRT is $\frac{\delta\mathbb{E}[\widetilde{Y}_{j,it}(\lambda_{k})|\widetilde{W}_{k,it}=\lambda_{k}]}{\delta\lambda_{k}}$.

Estimation of causal effects of interest of Panel Vector Autoregressions

Until now, the structure of the policy variable has been kept as flexible as possible to highlight the high potential of PVAR to estimate different causal effects depending on the distribution of the policy. In the following sections, I will show that PVAR can estimate several different causal effects depending on the distribution of the policy variable. Figure (ref) represents the different type of causal effects that can be identified according to the policy variable, which will be discussed more in depth in the subparagraphs that follow.

figure[figure omitted — 161 chars of source]

Through the paper, the stable unit value treatment assumption is assumed to hold. In this context, it can be framed as follows

assumption(SUTVA): For each unit $i$ and time $t$, the potential outcome given a contemporaneous assignment $\widetilde{w}$ and past and future assignments of the unit, \[ \widetilde{Y}_{j,it}(\widetilde{w}_{k}) := \widetilde{Y}_{it}\big(\widetilde{W}_{k,i,1:t-1}, \widetilde{w}_{k}, \widetilde{W}_{k,i,t+1:T} \big), \] is invariant to the assignment paths of all other units: \[ \widetilde{Y}_{j,it}\big(\widetilde{W}_{k,i,1:t-1}, \widetilde{w}_{k}, \widetilde{W}_{k,i,t+1:T}; \widetilde{W}_{k,-i,1:T}\big) = \widetilde{Y}_{j,it}\big(\widetilde{W}_{k,i,1:t-1}, \widetilde{w}_{k}, \widetilde{W}_{k,i,t+1:T}\big), \] where $\widetilde{W}_{k,-i,1:T} = \{\widetilde{W}_{k,i,1:T}: -i\neq i\}$ denotes the collection of all other units' assignment paths.

potential violations of the assumption due to interference are discussed in section (ref), which revises the main empirical application explicitly estimating spillover effects.

The policy variable is randomized and homogeneous across units

A researcher is interested in the evaluation of the effects of hurricanes in the economy of some states. All the data they have collected corresponds to unemployment, inflation, and the moments in which all the states have been affected by the disaster shock. This means that they have no control regions available. The researcher decides to make the assumption that, were the regions not affected by a disaster shock, they would have followed a VAR process. In this case, all units are defined as either treated or control in each period.

assumption(Policy variable is homogeneous across units): For all $i=1,..,N$, at a specific time $t$, either $\widetilde{W}_{k,it}=1$ or $\widetilde{W}_{k,it}=0$.
assumption(Policy variable is random): The policy variable is independent from the other policies, its past, future and the potential outcome process, so that: \[ \widetilde{W}_{k,it}\perp(\widetilde{W}_{k,i,1:t-1},\widetilde{W}_{1:k-1,i,t},\widetilde{W}_{k+1:K,i,t},\widetilde{W}_{k,i,t+1:T},\{\widetilde{Y}_{ji,t:T}(1):1\in\mathcal{W}^{t:T}\}) \] for all $k$, $j$, $i$, $t$ .

Assumption (ref) indicates that the treated moments are the ones for which $\widetilde{W}_{k,it}=1$, while control units are the ones for which $\widetilde{W}_{k,it}=0$. This interpretation aligns with common practice in IRF analysis, where shocks are typically expressed as unit shocks or one-standard-deviation shocks. Assumption (ref) indicates that the hurricanes:

enumerate• Must be independent from previous hurricanes (after controlling for their autocorrelation). • Must be independent with respect to future hurricanes. • Must be independent with respect to the potential outcome of inflation and unemployment. • Must be independent with respect to other influencing factors of inflation and unemployment.

If the previous assumptions are satisfied, obtaining the conditions under which the estimated causal effects are ATEs is straightforward. Notice that the estimated effects will be the following:

thm(PVAR: Randomized, dummy, homogeneous policy): Under assumption (ref) and (ref), Panel Vector Autoregressions estimate, for all $j$, all $k$, all $i$ and all $t$, the following causal effect: \[ \gamma_{jk}=\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=1]-\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=0] \] Which means that $\gamma_{jk}$ can be decomposed into the average treatment effect and a selection bias term as follows: \[ \gamma_{jk}=\mathbb{E}[\widetilde{Y}_{j,it}(1)-\widetilde{Y}_{j,it}(0)]+\Delta_{j,k,i,t}(1,0) \] where:

\[ \Delta_{j,k,i,t}(1,0):=\frac{Cov(\widetilde{Y}_{j,it}(1),1\{\widetilde{W}_{k,it}=1\})}{\mathbb{E}[1\{\widetilde{W}_{k,it}=1\}]}-\frac{Cov(\widetilde{Y}_{j,it}(0),1\{\widetilde{W}_{k,it}=0\})}{\mathbb{E}[1\{\widetilde{W}_{k,it}=0\}]} \] This means that the causal effects identified by the IRF of a PVAR will be Average Treatment Effects only if $\Delta_{j,k,i,t}(1,0)=0$, i.e. the selection bias is zero. This can happen if assumption (ref) is satisfied.

thm(PVARs estimate ATEs): Under assumptions (ref), (ref) and (ref) the following is true: \begin{equation} \begin{aligned}Cov(\widetilde{Y}_{j,it}(1),1\{\widetilde{W}_{k,it}=1\})=0,\qquad & Cov(\widetilde{Y}_{j,it}(0),1\{\widetilde{W}_{k,it}=0\})=0\end{aligned} \end{equation} then $\Delta_{j,k,it}(1,0)=0$. Moreover, equation (ref) is satisfied if: \[ \begin{aligned}\widetilde{W}_{k,it}\perp\widetilde{Y}_{j,it}(1), & \text{and} & \widetilde{W}_{k,it}\perp\widetilde{Y}_{j,it}(0)\end{aligned} \] which is in turn implied by: \[ \widetilde{W}_{k,it}\perp\{\widetilde{Y}_{j,it}(1):1\in\mathcal{W}_{k}\} \] which can be satisfied if assumption (ref) is respected. Hence, Panel Vector Autoregressions identify, for all $j$, all $k$, all $t$ and $i$: \[ \gamma_{jk}=\widetilde{\text{ATE}}_{j,k,i,t}=\mathbb{E}[\widetilde{Y}_{j,it}(1)-\widetilde{Y}_{j,it}(0)] \]

This result is similar to the one of RambachanSheppard2021 for VARs. The intuition for the similarity is that the system is purposefully shrinked down to the case in which unit heterogeneity is not utilized by the system. I will focus now in the cases in which the distribution of the policy variable may not be binary. Subsection (ref) discusses the case of a normal distribution of the policy variable's innovations.

The policy variable is continuous and normally distributed

A researcher is interested in evaluating the effects of monetary policy on inflation, similarly to the problem of Sims1980. Since their interest is in a monetary union, they collect data on interest rates, inflation, and other possibly relevant variables from all the states belonging to the monetary union. They estimate a PVAR, and observe that the innovation component of the interest rates is normally distributed. They make the argument that after controlling for past information, monetary policy decisions are as if random, and leverage this assumption by using a recursively identified PVAR. To carry a causal claim, they will need the following assumptions.

assumption(Continuous differentiability): The forecast errors of the outcome variables $\widetilde{Y}_{j,it}$ are continuously differentiable with respect to $\widetilde{W}_{k,it}.$
assumption(Normal distribution): The forecast error of the policy variable is normally distributed $\widetilde{W}_{k,it}\sim\mathcal{N}(0,\sigma_{\widetilde{W}_{k,it}}^{2})$.

Assumption (ref) may appear technical, but indicates that the forecast errors of inflation are not ill defined in some regions. Let us define $g(\widetilde{w}_{k})=\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=\lambda_{k}]$.

thm(PVARS: Randomized, continuous, unbounded Policy): Under assumptions (ref), (ref), and for each $j$, $k$, $i$, and $t$, Panel Vector Autoregressions capture the following estimand: \[ \gamma_{jk}=\int q(\lambda_{k})g'(\lambda_{k})d\lambda_{k}. \] Here $q(\lambda_{k})>0$ and $\int q(\lambda_{k})d\lambda_{k}=1$. The weights are: \[ \begin{aligned}q(\lambda_{k})=\frac{1}{\sigma_{\widetilde{W}_{k,it}}^{2}}\int_{-\infty}^{\infty}(\mathbb{E}[\widetilde{W}_{k,it}]F_{\widetilde{W}_{k,it}}(\lambda_{k})-\theta_{\widetilde{W}_{k,it}}(\lambda_{k})),\end{aligned} \] and $\Theta_{\widetilde{W}_{k,it}}(\lambda_{k})=\int_{-\infty}^{\lambda_{k}}mf_{\widetilde{W}_{k,it}}(m)dm=F_{\widetilde{W}_{k,it}}(\lambda_{k})\mathbb{E}[\widetilde{Y}_{j,it}|\tilde{W}_{k,it}=\lambda_{k}]$.

Here $f_{\widetilde{W}_{k}}$ indicates the marginal density, $F_{\widetilde{W}_{k}}$ the marginal cumulative distribution, $\sigma_{\widetilde{W}_{kit}}^{2}$ indicates the variance of $\widetilde{W}_{k,it}$. This theorem indicates that the causal effect estimated is a set of weights $q(\lambda_{k})$ that depends only on the distribution of interest rates disturbances.

thm(PVARs identify ACRT):Under assumptions (ref)(ref) and (ref), Panel Vector Autoregressions identify, for all $j$, $k$, $i$, $t$: \[ \gamma_{jk}=\widetilde{\text{ACRT}}_{j,k,it}(\lambda_{k}) \]

This result is partially similar to the one of Yitazaki1996. The ACRT is of particular interest since it captures the effect on the forecast error of the outcome variable by moving along the curve of observed forecast errors of the policy. This first derivative interpretation, however, possess a selection bias in $g(\widetilde{w}_{k})$, which depends on the conditioning argument $\widetilde{W}_{k,it}=\lambda_{k}$. This is where the assumption that the monetary policy is fully exogenous with respect to inflation can help obtain a causal effect without a selection bias.

thm(PVARs identify ACR): Under assumptions (ref),(ref),(ref), and (ref), \[ \gamma_{jk}=\widetilde{\text{ACR}}_{jk,it}(\lambda_{k}). \]

Notice that $\widetilde{\text{ACR}}_{jk,it}(\lambda_{k})$ is free of any selection bias, and allows to formulate counterfactual policy scenarios. If one truly believes in the assumption that monetary policy innovations are fully random with respect to inflation, it becomes possible to identify the effects of any impact shock in monetary policy on inflation.

The policy variable is continuous and non negative

A researcher is, once again, interested in the evaluation of the impact of hurricanes in the economy of some states. This time they make the argument that not all hurricanes impact the economy in the same way. Some create more destruction; some simply affect mildly the economy; some are technically registered, but have no effect on the economy. Therefore, they collect the value of all the destroyed property. The periods in which there are no hurricanes are registered as zero, and the periods in which they happen, the destruction measure is instead utilized. Therefore, the treatment variable is either positive or zero. Such a restriction is the same one proposed by callawaysantannabacon2021 of either continuous or multi-valued treatments.

assumption(Continuous differentiability): The forecast errors of the outcome variables $\widetilde{Y}_{j,it}$ are continuously differentiable with respect to $\widetilde{W}_{k,it}$.
assumption(Non-negative $\widetilde{W}_{k,it}$): $\widetilde{W}_{k,it}$ satisfies $0\leq\widetilde{W}_{k,it}<\infty$ for all $i$ and all $t$.
assumption(Strong Parallel Trends): For all $j$, $k$, $i$, $t$ \[ \mathbb{E}[\widetilde{Y}_{j,it}(\lambda_{k})-\widetilde{Y}_{j,it}(0)]=\mathbb{E}[\widetilde{Y}_{j,it}(\lambda_{k})-\widetilde{Y}_{j,it}(0)|\widetilde{W}_{k,it}=\lambda_{k}]. \]

Assumption (ref) is akin to the one in the previous subsection. Assumption (ref) regularizes the limits of the distribution of the treatment variable. Assumption (ref) is instead used by callawaysantannabacon2021 to restrict the treatment variable to have the same potential effect for the treated units and control units. Such an assumption imposes that receiving a treatment of size $\lambda_{k}$ compared to receiving no treatment (a form of ATE) is the same quantity as conditioning on the group of treated units. It is a convenient, but strong, assumption that allows to freely move from ATEs to ATTs. It is required to only hold for $\lambda_{k}$ and zero potential outcomes as this is the source of selection bias.

The causal effect can be computed from the Cholesky decomposition of the estimated residuals, and will represent:

thm(PVARS: Randomized, continuous, non-negative Policy): Under assumption (ref)(ref), Panel Vector Autoregressions capture the following estimand, for all $j$, all $k$, all $t$ and all $i$: \[ \begin{aligned}\gamma_{jk}= & \int_{d_{L}}^{d_{U}}q_{1}(\lambda_{k})\left(\frac{\delta\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=\lambda_{k}]}{d\lambda_{k}}+\frac{\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=\lambda_{k}]-\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=a]}{\delta a}\mid_{a=\lambda_{k}}\right)d\lambda_{k}+\\ & +q_{0}\frac{\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=d_{L}]-\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=0]}{d_{L}}. \end{aligned} \] Here \[ \frac{\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=\lambda_{k}]-\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=a]}{\delta a}\mid_{a=\lambda_{k}} \] is a selection bias term and \[ \begin{aligned}q_{1}(\lambda):=\frac{\mathbb{E}[\widetilde{W}_{k,it}|\widetilde{W}_{k,it}\geq\lambda_{k}]-\mathbb{E}[\widetilde{W}_{k,it}])\mathbb{P}(\widetilde{W}_{k,it}\geq\lambda_{k})}{var(\widetilde{W}_{k,it})}\end{aligned} \] and \[ q_{0}:=\frac{(\mathbb{E}[\widetilde{W}_{k,it}|\widetilde{W}_{k,it}>0]-\mathbb{E}[\widetilde{W}_{k,it}])\mathbb{P}(\widetilde{W}_{k,it}>0)d_{L}}{var(\widetilde{W}_{k,it})} \] and the weights satisfy $\int_{d_{L}}^{d_{U}}q_{1}(\lambda_{k})d\lambda_{k}+q_{0}=1$. Under assumption (ref) it becomes \[ \begin{aligned}\gamma_{jk}= & \int_{d_{L}}^{d_{U}}q_{1}(\lambda_{k})\frac{\delta\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=\lambda_{k}]}{d\lambda_{k}}d\lambda_{k}+q_{0}\frac{\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=d_{L}]-\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=0]}{d_{L}}\end{aligned} \]

This approach decomposes the estimated coefficient in different partitions of a curve of causal effects, so that $\gamma_{jk}$ will be a weighted average of an $\text{ACRT}_{j,k,it}(\lambda)$ and an $\text{ATE}_{j,k,it}(d_{L})$.

thm(PVARs identify a weighted average of ACR and ATE): The result in Theorem (ref) can be expressed as a weighted average of ACR and ATEs, hence Panel Vector Autoregressions identify, for all $j$, $k$, $t$, $h$: \[ \gamma_{jk}=\int_{d_{L}}^{d_{U}}q_{1}(\lambda_{k})\widetilde{\text{ACRT}}_{j,k,it}(\lambda_{k})+q_{0}\frac{\widetilde{\text{ATE}}_{j,k,it}(d_{L})}{d_{L}}, \] which collapses to a different range of causal effects.

This result can be interpreted as a negative one. Even under strong assumptions, the asymmetry of the distribution can result in a mix of causal effects, which may be of hard interpretation. Moreover, the causal interpretation of $\widetilde{\text{ACRT}}(\lambda_{k})$ may still be problematic due to the presence of a selection bias term. A key difference with respect to callawaysantannabacon2021 is that while they consider other possible forms of estimators for quantities that may be informative about the shape of the causal response, the approach of PVARs has an additional downside. The researcher has no control over the distribution of the innovations of the policy variable, making other alternative estimators that appropriately weight the observations difficult to consider.

What does this mean for the researcher that collected information about the negative impact of hurricanes? Intuitively, this results speaks about the exercise of trying to capture different types of information within a unique coefficient. In this case, $\gamma_{jk}$ is being asked to capture both the impact of moving from a non-event to the lowest destructive hurricane, but also the one of moving from the lowest possible hurricane to the second lowest possible and so on. This is not necessarily a unique coefficient. In the case of callawaysantannabacon2021 the consequence is that it becomes possible to move from a parametric to a non-parametric estimation, and capture the different causal effects through the estimation of different coefficients.

This issue opens up several venues for research. While currently PVARs are identified by simply running a recursive ordering based estimation, it could be possible to consider different estimators that select different observations due to their common distribution. In that case, rather than obtaining just one Impulse Response Function, the researcher would have in hand several Impulse Responses, each capturing the impact of an intervention of, say, a treatment 1 versus a treatment 0; or a treatment 2 versus a treatment 0; or a treatment 2 versus a treatment 1; each referring to a different Cholesky decomposition that refers to commonly distributed residuals. Unfortunately such estimators generally require a large amount of data to cover all the different densities, which may be reasonable in the case of weekly or monthly macroeconomic data, but unfeasible in the case of quarterly or yearly data. In general, however, the issue of non-linearity is becoming more and more apparent in many macroeconomic estimators (see kolesar2024dynamic).

The policy variable is a dummy that indicates treated or untreated units

A researcher is, for the last time, interested in the evaluation of the impact of hurricanes to the economy of some states. This time they observe that hurricanes affect only some regions, leaving others untouched. They therefore notice that the regions not affected by hurricanes can be used as a counterfactual for the ones that are impacted. They make the case that, by considering the forecast errors, they are capable of eliminating the non-immediate effect of the disaster. For example, if a region is particularly affected, the autoregressive coefficients will capture this shift, and eliminate the long term impact of hurricanes.

For this reason, if the units can contemporaneously be extracted to be treated/controls, the burden of generating a counterfactual will be both on the forecasts and on the other untreated units. Consider the following structure for the policy:

assumption(Policy is dummy across heterogeneous units): Units can be divided into four subgroups, respectively \[ \begin{cases} \widetilde{W}_{k,it}=1 & \text{ for }i\in I_{P}\text{ at times }t\in T_{P}\\ \widetilde{W}_{k,it}=0 & \text{ for }i\in I_{C}\text{ at times }t\in T_{P}\\ \widetilde{W}_{k,it}=0 & \text{ for }i\in I_{P}\text{ at times }t\in T_{C}\\ \widetilde{W}_{k,it}=0 & \text{ for }i\in I_{C}\text{ at times }t\in T_{C} \end{cases} \]

Here $I_{P}$ represents a subgroup of units which are treated, $I_{C}$ represents a subgroup of units which are not treated, $T_{P}$ represents the subgroup of times that are treated, and $T_{C}$ represents the subgroup of times that are not treated.\footnote{The $P$ notation is used to indicate “policy” times or “policy” affected units.} It follows that $I_{P}+I_{C}=I$ and $T_{P}+T_{C}=T$.

A consequence of assumption (ref) is that there exist at least a control unit for times in which some units are subject to a treatment. It is necessary in order to guarantee that $\widetilde{W}_{k,it}$ is a dummy that also includes non treated units that can act as contemporaneous counterfactuals for the treated group.\footnote{Notice that this assumption also relates deeply to the results of callawaysantannamultipletime2021, which argue against fixed effects estimation specifically in the case in which there are no more control units after a series of treatments in the staggered treatment framework. Assumption (ref) avoids such instances completely by assuming that there are sufficient counterfactuals contemporaneously.}

assumption(Parallel trends): For each $j\geq1$, $k\geq1$,$t\geq1$,$i\geq1$: \[ \mathbb{E}[\widetilde{Y}_{j,it}(0)|t\in T_{P},i\in I_{P}]=\mathbb{E}[\widetilde{Y}_{j,it}(0)|t\in T_{P},i\in I_{C}]. \]
assumption(No anticipation): For each $j\geq1$, $k\geq1$,$t\geq1$$i\geq1$: \[ \mathbb{E}[\widetilde{Y}_{j,it}(0)|t\in T_{C},i\in I_{P}]=\mathbb{E}[\widetilde{Y}_{j,it}(0)|t\in T_{C},i\in I_{C}]. \]
assumption(No residual autocorrelation): For each $j\ge1$, $k\geq1$, $t\geq1$,,$i\geq1$ and for $s\geq1$: \[ corr(\widetilde{x}_{j,t},\widetilde{x}_{1,t-s}),..,corr(\widetilde{x}_{j,t},\widetilde{x}_{J,t-s})=0. \]
assumption(No contamination): For each $k\geq1$, $t\geq1$,$i\geq1$: \[ \widetilde{W}_{k,it} \perp \widetilde{W}_{-k,it} \]

Assumption (ref) implies that the treated and non treated units are comparable in potential outcomes in treated times. This is akin to a parallel trend assumption in the DiD literature, with the difference that it is imposed in the forecast errors. Assumption (ref) means that units belonging to the treatment and control group have identical potential outcomes in non treated times. Both of those assumptions are standard in the DiD literature. Notice that an implication of the two assumptions is that they may be violated in the case of spillover effects.

Assumption (ref) states that there must be no residual autocorrelation in the innovations of the outcome variables. It is required to avoid the contamination of previous causal effects or anticipatory effects. This is deeply rooted with the idea of stationarity and its reliability strictly depends on the nature of the data and the appropriate additions of dummies to control for time series confounding factors in the PVAR regression, lag selection, and the overall nature of the estimated system. This is because the idea of this type of identification procedure is that any unit, at any time, is allowed to be treated.

Finally, assumption (ref) assumes independence across multiple treatments. For example, if the economist is analysing the effects of two types of natural disasters, such as heatwaves and droughts, the assumption implies no sytematic correlation across the two assignments. \footnote{Notice that this assumption does still allow for assignments to be temporally related. Consider the example of usman2024going, which eliminate regions affected by multiple, subsequent, heterogeneous shocks, such as heatwaves and droughts, from the control sample. From the point of view of a panel VAR, as long as the disasters are not contemporaneously correlated, the occurrence of droughts after heatwaves will not contaminate inference as long as it is estimated by the AR coefficients of the two shocks. Rather, the systematic occurrence of disasters at the same time could indeed contaminate inference, as there would be no way to distinguish the two dummy events.}

To understand the difference with respect to traditional causal inference settings in panel data, such as DiD, staggered-DiD and LP-DiD, consider Table (ref).

table[table omitted — 1,835 chars of source]

The first column represents the traditional DiD literature. The economist could compute the difference before and after period $t=3$ between group 1 and group 2.

However, if more units can be included, the DiD becomes a staggered setup like the one in column 4,5, and 6. The DiD literature (see, among others, callawaysantannamultipletime2021,sunabraham2021) often times solves the problem of potentially shifting effects in the treated units in staggered treatment designs by carefully selecting a subgroup of units for which the econometrician knows that they did not receive a treatment yet. Such units are supposed to act as a credible counterfactual in a non-parametric framework. In this case, at time $t=2$ the units in the group 1 receive a treatment, and the others can be used as a control. It is possible to identify the effect of the treatment at time $t=2$ by simply comparing the difference of the units before and after the treatment. The same quantity can be computed using the difference of $t=3$ and $t=1$. However, the $t=4$ observations cannot be used for causal inference.

Clearly, a drawback of this approach is that, for an increasing $t$, less and less counterfactuals are available. This is because the researcher is not allowed to include the units which previously received a temporary treatment in the control group. This is the reason why such methods often limit themselves to the case of staggered treatment and not a fully sparse treatment. Sparse treatments could be used in scenarios in which the treatment becomes obsolete after one period, and allows to use units from any group as counterfactual for the following period.

dube2023local show that, in the case of sparse treatments, it is possible to avoid the underlying assumption of the staggered DiD setup by imposing the existence of a time after which the units can go back to the control group. This case may be more apt for scenarios in which the researcher is aware of a short lasting effect of the intervention.

PVARs take a different approach. PVARs do not make use of the variables directly, but of their residuals. Therefore, they can realistically make the assumption that even treated units can become controls after just one period. This interpretation is engrained in the idea of looking at deviations from the autoregressive component of the treatment. Effectively, as long as the carryover effect is known, it is possible to use units which are still receiving some form of treatment by using their residuals. While dube2023local would exclude units which are still receiving a form of treatment, PVARs allows them to be on the sample. Clearly, if the distribution of the carryover effects estimated by the autoregressive coefficients $\frac{cov(Y_{it},W_{i,t-j})}{W_{i,t-j}}$, is ill-behaved, inference will be contaminated. This may be more credible in cases in which we believe that the autoregressive component is capable of absorbing the effect of the shock, i.e. the treatment does not alter the structural equilibrium of the outcome variables. In this case, it can be reasonable to assume the following:

enumerate• There is a cyclical component which motivates the use of a PVAR, • at least one unit is either treated or control at each time $t$ (assumption (ref)), • units that have received a treatment would have otherwise seen a similar outcome to the ones that did not (assumption (ref)), • the AR coefficients are capable of eliminating the effects of the treatment (assumption (ref)).
thm(PVAR: Randomised, dummy, heterogeneous policy): Under assumption (ref)(ref), (ref) Panel Vector Autoregressions capture the following estimand: \[ \begin{aligned}\gamma_{jk}= & \mathbb{E}[\widetilde{Y}_{j,it}|t\in T_{P},i\in I_{P}]-\mathbb{E}[\widetilde{Y}_{j,it}|t\in T_{P},i\in I_{C}]\\ & -\mathbb{E}[\widetilde{Y}_{j,it}|t\in T_{C},i\in I_{P}]+\mathbb{E}[\widetilde{Y}_{j,it}|t\in T_{C},i\in I_{C}]. \end{aligned} \]

Hence, the causal effects can be easily found to be the ones in the following Theorem:

thm(PVARs identify ATTs): Under assumptions (ref);(ref); (ref); (ref) Panel Vector Autoregressions identify, for all $j$, all $k$, all $t$, and all $i$: \[ \gamma_{jk}=\widetilde{\text{ATT }}_{j,k,it}=\mathbb{E}[\widetilde{Y}_{j,it}(1)-\widetilde{Y}_{j,it}(0)|\widetilde{W}_{k,it}=1]. \]

Notice that it is often times the case that inference with aggregate variables is conducted with methods that do not have direct causal interpretation. Even if in this specific case the assumptions may appear to be too restrictive, it should be noted that the advantage of this approach is the ability to make truly causal claims. The applied literature has been historically welcoming to this assumption by the means of DiD, since it involves a direct counterfactual and the credibility of identification can be evaluated on a case by case scenario. This is not true when dealing with aggregate data, where often times inference is conducted without appropriately discussing the required assumptions to make causal claims in potential outcome frameworks.

remDiD a-là-callawaysantannamultipletime2021 should be preferred if the researcher is in a staggered treatment setup. LP a-là dube2023local may be preferred in cases in which the researcher is confident about a date in which the effect of the policy terminates. PVARs should be preferred if the researcher is confident about the temporary nature of the intervention in the autoregressive residuals.
remTo prove the results I have used a Cholesky decomposition, but no independence assumption was necessary to obtain the ATT.

Application using natural disasters

Due to the potentially detrimental effects of climate change related disasters, much of the scientific literature has focussed on the issue of trying to quantify the effects of disasters on the overall economy, obtaining mixed results. Often times the effects are found to be neutral or positive in real activity (see Strobl2011,linder2013). In the nominal side of the economy, it appears as though there is consensus of a positive effect of disasters on inflation. For example, beirne2022 finds a marked positive effect of natural disasters on inflation in the euro area, parker2018 finds that natural disasters positively affect inflation disproportionally in advanced economies, and faccia2021feeling find that extreme temperature events have a non-negligible impact on prices. Regarding the timing, there seems to be an agreement about the overall long term neutrality of natural disasters (Cavallo2013) but some empirical evidence suggests a medium term effect of extreme climate events (usman2024going).

The application I propose is the one that is at the most granular level. I use the series on economic activity produced by baumeister2024tracking that tracks weekly economic activity at the state level in the United States. The series is produced by considering data from various sources, including data on mobility, labour market conditions, real activity indicators, expectations, financial returns and household spending and salaries. Moreover, I consider data on the natural disasters coming from the NOAA\footnote{https://www.ncei.noaa.gov/access/billions/events/US/1980-2024?disasters{[}{]}=all-disasters}, which collects state level data on natural disasters exceeding the billion dollar mark. Such disasters are classified into droughts, floods, freezes, severe storms, tropical cyclones, wildfires and winter storms. They tend to vary in length and impact, and a more granular analysis of their impact is certainly warranted but outside the focus of this paper.

The classic assumption when evaluating the effect of disaster shocks is to assume them to be exogenous with respect to a wide variety of macroeconomic indicators. However, it may be that economic activity is correlated with periods of high economic growth. This may overestimate the disrupting effects of disasters if such growth is positively correlated with deforestation, environmental degradation, or other geographical features that may increase the damaging impact of the disaster. Many other arguments could be put forth in favour of a violation of the exogeneity condition.

However, if the researcher is willing to believe that territories affected by natural disasters would have otherwise seen a level of changes in economic activity which are the same ones as the one of the territories not affected by disasters, in accordance to the previous discussion, an ATT, rather than an ATE, may still be estimated. In this case, I consider the following model

equation[equation omitted — 97 chars of source]

where $d_{it}$ is an indicator that is equal to one when a state $i$ faces a natural disaster in week $t$, and $\Delta e_{it}$ is the first difference weekly variable produced by baumeister2024tracking\footnote{I use the first difference because the economic indicator is non stationary.}. The data analyzed starts in April 1987 and ends in October 2024, as that is the range of the series of baumeister2024tracking.

The underlying assumptions behind this modelling approach are the following:

enumerate• The potential outcome under no treatment of innovations in economic activity of the units for which $\widetilde{d}_{it}\sim1$ at time $t$ can be credibly represented by the units for which $\widetilde{d}_{it}\sim0$, in accordance to assumption (ref), • The potential outcome under no treatment in the innovations in economic activity of the units for which $\widetilde{d}_{it}\sim0$ at time $t$ is common to all the units, in accordance to assumption (ref), • The system presents no residual autocorrelation, in accordance to assumption (ref).

The PVAR is estimated through GMM, with five lags as it strikes the best balance according to a MBIC test (see Table (ref)). While the other tests may suggest other lags to be optimal, I focus on the case of five lags as it also strikes the best economic intuition: using the five previous weeks corresponds to making use of the previous month data. Moreover, I add a dummy variable to control for the COVID-19 pandemic period.

table[table omitted — 497 chars of source]
figure[figure omitted — 251 chars of source]

To verify that the no residual autocorrelation assumption holds, figure (ref) displays the residuals of each variable plotted against their lags. From the figure and the regressions of the residuals against their lags, no significant correlation emerges.

The resulting Impulse Response Functions are plotted in figure (ref), which indicates a negative effect on economic activity that lasts for two periods. The impulse response function suggests a negative cumulative impact of $\sim.04$ on average coming from the fact that the two weeks after the disaster appear to face a negative $\sim0.19$ deviation from their previous value.

figure[figure omitted — 295 chars of source]

Interference in panel vector autoregressions

According to the stable unit treatment value assumption (SUTVA (ref)), potential outcomes only depend on one's own treatment assignment. In many cases, SUTVA may fail due to an unknown interference structure among neighbours. In the fields of environmental economics, urban economics, labour economics, and in the case of natural disasters, place-based policies and interventions often generate spillover effects. In the case of natural disasters, Cavallo2013 uses a synthetic control approach a-la-Abadie2010 to generate a counterfactual for countries that faced a significant natural disaster. They discuss possible SUTVA violations, and come to the conclusion that they are unlikely to play a significant role in the case of disasters in the medium term. Yet, bacchiocchi2024macroeconomic find statistically significant short-run effects of severe weather shocks on local economic activity and cross-border spillovers operating through economic linkages between US states. Therefore, it seems reasonable to revise this paper's approach to inference and discuss potential violations of the SUTVA assumption and possible solutions.\\

Typically, the approach to dealing with interference through spillover effects is to explicitly assume a network structure (see, among others aronow2017estimating,manski2013identification,vazquez2023causal). I follow their lead in explicitly modifying the potential outcome framework to the following

assumption[Exposure mapping] Define an exposure mapping $\mathcal{S}_{it}=s(\widetilde{W}_{-i,1:T})$ that summarizes the influence of other units' assignment paths on unit $i$. The potential outcome can be written as \[ \widetilde{Y}_{it}\big(\widetilde{w}\;;\;\widetilde{W}_{-i,1:T}\big) = \widetilde{Y}_{it}\big(\widetilde{W}_{i,1:t-1},\widetilde{w},\widetilde{W}_{i,t+1:T}\;;\;\mathcal{S}_{it}\big), \]

Notice that assumption (ref) allows potential spillovers among individual states due to a mapping $S_{it}$ which depends on the $-i$ units. For example, this may mean that if a natural disaster hits the $-i$ state of California, the effect may reach out to $i$ Arizona. In this context, several important causal effects can be discussed:

enumerate• The total average treatment effect on the treated $\widetilde{\text{ATTE}}_{j,k,it} = \mathbb{E}[\widetilde{Y}_{j,it}(1,\mathcal{S}_{it}=s)-\widetilde{Y}_{j,it}(0,0)|\widetilde{W}_{k,it}=1,S_{j,j,it}=s]$ • The average spillover effect on the treated $\widetilde{\text{ASTE}}_{j,k,it} = \mathbb{E}[\widetilde{Y}_{j,it}(0,\mathcal{S}_{it}=s)-\widetilde{Y}_{j,it}(0,0)|\widetilde{W}_{k,it}=1,S_{j,j,it}=s]$

Here, the total average treatment effect on the treated represents the average effect of receiving a treatment and being subject to a form of spillover effect. The average spillover effect on the treated, instead, captures the effect of the spillover effect under a no treatment scenario. One useful way to frame the estimation of causal effects is to identify what $\gamma_{j,k}$ could identify under the assumptions in section (ref):

thm(PVARs identify total effects and spillovers on the treated): Under assumptions (ref); (ref); (ref);(ref) Panel Vector Autoregressions identify, for all $j$, all $k$, all $t$, and all $i$: \[ \gamma_{jk}=\widetilde{\text{ATTET}}_{j,k,it}-\widetilde{\text{ASTE}}_{j,k,it} \]

Essentially, if the SUTVA assumption is violated, the researcher would claim the identification of an ATE, but would be capturing the spillover effect too. This theorem resembles similar findings in different streams of literature. For example, vazquez2023identification finds that experimental designs with spillover effects do not correctly estimate an average treatment effect. forastiere2021identification finds that an estimator that wrongly assumes away interference is biased. butts2021difference,xu2023difference find that DiD designs do not correctly identify an ATT in the case of spillovers. In essence, the theorem shows that assumptions (ref); (ref) are not sufficient to capture a meaningful causal estimand in the context of interference. Rather, $\gamma_{j,k}$ isolates the average total effect on the treated - the impact of the treatment and the spillover - and the average spillover effect on the treated - the impact of spillovers on the counterfactual.

Several papers, however, propose different possible solutions to the issue of spillovers. If one can reasonably restrict the exposure mapping to specific neighbors, it becomes possible to estimate the spillover effect and therefore eliminate the source of bias.

For example, under a correctly specified exposure mapping $ S_{jk,it} $, the set of regressions

equation[equation omitted — 169 chars of source]

can estimate the total average treatment effect (on the treated) and the average spillover effect (on the treated) through $\delta$ and $\rho$, respectively. Typically, $ S_{jk,it} $ is assumed to be a function of neighbour distance. For example butts2021difference revises kline2014local using concentric county rings to measure the spillover exposure in the context of DiD. Another approach is proposed by xu2023difference, which uses a propensity score matching (doubly robust estimator) to measure the spillover exposure.\footnote{xu2023difference in particular does not specify an exact exposure map, but relies on the propensity score match. While this approach is not vulnerable to a misspecification of the exposure map, it does require a robust propensity score matching, whose accuracy depends on the information utilised to generate the matching function. It should also be noted that xu2023difference operates in a finite population framework, which should provide less conservative inferential claims than superpopulation approaches. }

It is possible to recast the assumptions-estimands as follows. If the empirical application suits it, units which are immediately close to the treated ones may be affected by spillover effects, but units far away may still act as credible counterfactuals. For example, while Arizona may be subject to spillover effects if a drought hits California, it seems unrealistic to think that the effect may spill over to other states which are further away.\footnote{Especially in the context of the empirical application of this paper, usually NOAA well classifies the natural disasters and their economic impact. If a drought happens in California, but also negatively affects Arizona, the NOAA will classify the disaster as happening in Arizona too.} The most commonly utilised approach is to assume a known spillover matrix \footnote{For example, vazquez2023causal formally assumes a distribution of compliance types in an instrumental variable framework.}

In this case, the following assumptions need to be satisfied in place of the common parallel trend and no anticipation.

assumption(Modified parallel trends) For each $j\geq1$, $k\geq1$,$t\geq1$: \[ \mathbb{E}[\widetilde{Y}_{jk}(0,s)|t\in T_{p},i\in I_{c};S_{j,k,it}=s]=E[Y(0,0)|t\in T_{p},i\in I_{p};S_{j,k,it}=0] \]
assumption(No lagged spillovers) For each $j\geq1$, $k\geq1$,$t\geq1$: \[ \mathbb{E}[\widetilde{Y}_{jk,it}(0,0)|t\in T_{c},i\in I_{p};S_{j,k,it}=0]=\mathbb{E}[\widetilde{Y}_{j,k,it}(0,0)|t\in T_{c},i\in I_{c};S_{j,k,it}=0] \]

Assumption (ref) implies that the treated units would have seen similar economic developments to "far away" controls. Assumption (ref) is only mildly stronger than (ref). It reinforces the non-anticipation assumption to also require there not to be lagged spillover effects.

Under those assumptions, it is possible to show that the estimator $\delta_{j,k}$ in equation (ref) recovers the average total effect on the treated.

thm(PVARs identify ATTs-spillovers): Under assumptions (ref); (ref); (ref);(ref) equation (ref) identifies, for all $j$, all $k$, all $t$, and all $i$: \[ \delta_{j,k}=\widetilde{\text{ATTET}}_{j,k,it} \]

Once the $ATTE_{j,k,it}$ has been estimated, it can be plugged in as a value for the impulse response function \footnote{Notice that this approach, differently from the Cholesky decomposition, does not fully identify the rotations of the covariance matrix, $\hat{O}\hat{O}^(-1)=\hat{\Sigma}$, but it does provide enough information to generate the impulse response functions of interest. In the equation for $\hat{O}$ in lemma (ref) this means that, for a unitary shock in natural disasters $\hat{o}_{it}^{11}=1$, $\hat{o}_{it}^{12}=\delta$, but $\hat{o}^{22}_{it}$ is not identified. This feature is common in externally identified SVARS, such as gertler2015monetary.}

Under this framework, it is possible to recover the errors from the PVAR estimation in (ref) and estimate a linear regression against the natural disasters innovation and a spillover matrix. In this context, the spillover matrix is zero every period in which there are no disasters, and one in times in which there are disasters for the states that border the treated.\footnote{The matrix is generated based on USSWM from merryman2005.} In this case $S_{it}$ becomes the inverse of how many states, on average, are treated in the neighbour.

The results of the regression are displayed in table (ref). They show an ATTE broadly similar to the ATT estimated by the Cholesky decomposition ($-.0109$), which does not take into account potential SUTVA violations.

table[table omitted — 670 chars of source]

The estimated impulse response is very similar to the original work because its evolution strictly depends on the AR coefficients, which are the same as the specification in section (ref), and is therefore not presented in this text to avoid redundancies.

Conclusion

Panel Vector Autoregressions are a flexible tool for the estimation of causal effects. Depending on the shape of the treatment and the assumptions the researcher is willing to put on their distributions, the causal effects can either be:

enumerate• Average Treatment Effects (ATE) $\mathbb{E}[\widetilde{Y}_{j,it}(1)-\widetilde{Y}_{j,it}(0)]$ if the policy variable is an “homogeneous” dummy, i.e. all the units are either jointly treated or jointly non treated at each time, and is independent on its past, its future, other policy variables, and the outcome variable. • Average Causal Response on the Treated (ACRT) $\frac{\delta\mathbb{E}[\widetilde{Y}_{j,it}(\lambda_{k})|\widetilde{W}_{k,it}=\lambda_{k}]}{\delta\lambda_{k}}$ if the policy variable is heterogeneous among units, continuous, and normally distributed. • A weighted average of Average Causal Response on the Treated (ACRT) and Average Treatment Effects (ATE) $\int_{d_{L}}^{d_{U}}\frac{\delta\mathbb{E}[\widetilde{Y}_{j,it}(\lambda_{k})|\widetilde{W}_{kit}=\widetilde{w}_{k}^{\prime}]}{\delta\lambda_{k}}q_{1}(\lambda_{k})+q_{0}\frac{\mathbb{E}[\widetilde{Y}_{j,it}(d_{L})-\widetilde{Y}_{j,it}(0)]}{d_{L}}$ if the policy variable is heterogeneous among units in a given time, non-negative, continuous, and strong parallel trends hold. • Average Treatment Effects on the Treated (ATT) $\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=1]-\mathbb{E}[\widetilde{Y}_{j,it}|\widetilde{W}_{k,it}=0]$ if the policy variable is a dummy but some units are treated while others are not, the units respect the “parallel trend” and “no anticipation” conditions, and there is no residual autocorrelation.

I showcase the high potential of this last set of assumptions to evaluate the effects of natural disasters on the real economy. To best complement the work, I provide alternative formulations to deal with spillovers.