Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
93,067 characters · 16 sections · 67 citation commands
When and Why State-Dependent Local Projections Work
In macroeconomics, the effect of an observed shock $X_t$ on a future outcome $Y_{t+h}$ is commonly estimated by running a local projection Jorda:05 of the form\footnote{Since this paper only studies asymptotic properties, it abstracts from control variables that are included to improve finite-sample performance. If controls are used for identification, assume that they have already been projected out using the Frisch-Waugh-Lovell theorem.}
To study whether the effect of $X_t$ on $Y_{t+h}$ depends on the initial state of the economy, a state-dependent version of this regression can be estimated:
where $S_{t-1}$ is a lagged, observed state variable which can be continuous or binary.\footnote{Most applied papers seem to use a lagged state, even though some interact with a contemporary state $S_t$ (see Appendix (ref)). Also see Remark (ref) for a discussion of this issue.} If the regression results indicate that the interaction term $\beta_1^h$ is non-zero, the effect of interest is commonly judged to be state-dependent.
State-dependent LPs are popular, but so far it has been unclear whether their common interpretation is valid when the true data generating process is not of the form (ref). I show that state-dependent LPs estimate a causal effect, even if the true structural function does not correspond to the estimating equation. This is relevant since LPs are commonly used when the researcher does not want to commit to a particular structural model, but still has to rely on a parsimonious parametric estimation technique due to small sample sizes of macroeconomic time series. My paper makes three points that should help clarify the scope and limitations of state-dependent LPs.
First, state-dependent LPs estimate weighted averages of conditional marginal effects if the shock is observed and independent of the lagged state. The weights only depend on the distribution of the shock and are identical across state and application. This nonparametric guarantee has already been derived for linear LPs Rambachan:21,Kolesar:24, but I show that it also holds for state-dependent LPs very generally. To estimate more specific causal quantities such as the average response to a shock of size $\delta$, the data generating process has to be substantially restricted. However, this is true for both linear and state-dependent LPs. In this sense, state-dependent LPs are as valid as linear LPs. Moreover, the interpretation remains transparent even when practitioners depart from simple linear interactions. Even if a continuous interaction term is used in (ref) and the relationship between effect and state is not of the form $\beta_0^h + S_{t-1} \beta_1^h$, state-dependent LPs still estimate a best approximation in the familiar MSE sense: A linear regression of the effect at $S_{t-1}$ onto $(1, S_{t-1})$. Therefore, my result covers virtually all specifications of state-dependent LPs used in the applied literature. In addition, the formulas derived here can be easily applied to new functional-form specifications of state-dependent LPs. Researchers can use their well-trained intuition for misspecified linear regressions to interpret the causal estimand implied by any chosen specification. Since at the moment much of the applied literature relies on only a small set of functional forms,\footnote{In particular, many papers interact the shock with a logistic transformation of a continuous state variable, as in Auerbach:13, or with a binary state indicator, as in Ramey:18.} these results provide guidance for exploring new specifications.
Building on this foundation, I next compare state-dependent LPs with their VAR counterparts. In the linear case, those two methods asymptotically yield the same effect estimates Plagborg-Moller:21. Using a simple DSGE model, I show with simulations and analytically that this equivalence breaks down in the state-dependent case. This occurs even in the idealized scenario where the state follows a known, fully exogenous Markov process and the researcher can manually adjust for the future evolution of the state. Therefore, the favorable asymptotic properties of state-dependent LPs derived in this paper do not carry over to state-dependent VARs. As a remedy, I introduce an impulse response estimate constructed from multiple state-dependent VAR models. This estimator is easy to construct and asymptotically matches the state-dependent LP estimand. This allows researchers that prefer VARs over LPs to reap the asymptotic benefits derived in this paper.
Finally, I extend the analysis to the IV setting, which is central in much applied work. State-dependent LPs using instrumental variables (LP-IVs) also estimate a weighted average of marginal effects. However, the weights generally depend on the state. This makes interpretation challenging without additional information on the data generating process: A non-zero interaction term can arise due to differences in the weighting scheme across states, even if the effect of interest is not state-dependent. To interpret state-dependent LP-IVs in the usual way, either the structural relationship between instrument and regressor or between regressor and outcome have to be restricted. This bears many similarities to the microeconometric literature on local average treatment effects Imbens:94. My paper is the first to raise this issue in the context of state-dependent LPs.
\noindentLiterature.---Linear regressions in a non-linear environment have been studied at least since Yitzhaki:96 and Angrist:00. Rambachan:21 first applied results of this literature to local projections and recently Kolesar:24 weakened the required regularity conditions. In a similar framework, Caravello:24 show how to identify sign and size nonlinearities and Casini:25 study high-frequency event studies. My paper is the first thorough treatment of state-dependent LPs in a nonlinear environment.\footnote{Kolesar:24 note that their results generalize to state-dependent LPs with a binary state since interacting with a dummy amounts to running two separate regressions. However, my results go beyond the binary case by covering continuous and multi-dimensional states. This is necessary to cover specifications commonly used in the literature: 19 of the 44 papers surveyed by Goncalves:24 use a continuous state variable (see Appendix (ref)).}
Some papers have studied state-dependent LPs in a parametric setting to obtain specific estimands of interest: Cloyne:23 extend the Kitagawa-Oaxaca-Blinder decomposition to decompose channels of impulse response heterogeneity. Goncalves:24 study state-dependent LPs under the assumption that the data generating process is a state-dependent VAR. Their estimand of interest is the average response to a non-marginal shock of size $\delta>0$ and they demonstrate that state-dependent LPs can fail to estimate this quantity. The goal of this paper is more modest: I show that state-dependent LPs estimate some weighted average of causal effects. The average effect of a shock of size $\delta$ is a special weighted effect that may or may not correspond to the LP estimand, depending on the data generating process.
Lastly, this paper adds to a literature relating LPs and VARs. Plagborg-Moller:21 first showed that both models asymptotically yield the same effect estimates. Recently, Ludwig:24 derived a finite sample version of this equivalence. This paper shows analytically and with simulations that this equivalence breaks in the state-dependent case. As a remedy, I propose a VAR-based estimate that asymptotically matches the state-dependent LP estimand.
\noindentOutline.---Section (ref) sets up the econometric framework and reviews a key result for linear LPs. Section (ref) contains the main approximation result for state-dependent LPs with observed shocks and Section (ref) discusses its implications for specific empirical specifications. Section (ref) studies the relationship between state-dependent SVARs and LPs. Section (ref) covers instrumental variable methods, Section (ref) concludes. Appendix (ref) presents some properties of the applied state-dependent LP papers surveyed by Goncalves:24, which provides additional information about some claims made in this paper.
This section presents an important result for linear LPs that later sections build upon. The notation and required regularity conditions follow Kolesar:24.
\noindentStructural Functions.---We are interested in the response of a scalar outcome variable $Y_{t+h}$ to a change in the scalar $X_t$. For example, think of $Y_{t+h}$ and $X_t$ as output and a fiscal policy shock in period $t+h$ and $t$, respectively. As is common in the applied literature, I assume that the shock $X_t$ is observed without measurement error, which makes a regression of $Y_{t+h}$ on $X_t$ feasible.\footnote{With classical measurement error, attenuation bias will yield a rescaled version of this regression, which leaves the shape of the estimated impulse response intact Plagborg-Moller:21.} Without loss of generality, let $Y_{t+h}$ be determined by the structural function
where $U_{h,t+h}$ is a collection of variables that influence the outcome variable. In most macroeconomic models, $U_{h,t+h}$ would be a collection of shocks, lags of $Y_{t}$ and other macroeconomic variables that affect $Y_{t+h}$. To gain intuition, consider a simple example:
Note that in econometric practice, one often neither knows the functional form of $\psi_h$ nor the variables $U_{h,t+h}$. It will turn out useful to marginalize the structural function over $U_{h,t+h}$ to obtain the average structural function Blundell:03:
\noindentCausal Effects.---In nonlinear time series models, the size of the response of $Y_{t+h}$ to a change from $X_t$ to $X_t + \delta$ depends on the history of past shocks, the baseline shock level $X_t$ and the sign as well as absolute size of $\delta$. Therefore, there are many different causal effects one could possibly consider. For pragmatic reasons, I focus on average marginal effects\footnote{This type of effect is often the only one that can be estimated with reasonable precision, given typical sample sizes of macroeconomic time series Kolesar:24. If $\Psi_h$ is identified, in theory more general impulse response functions could be estimated using nonparametric methods. The few attempts of nonparametric local projections so far include Goncalves:24b and Paranhos:25.} of the form
where $\omega \geq 0$ satisfies $\int \omega(x) dx = 1$ and is therefore a weight function across the baseline values of the shock. If $\omega$ is the shock density, $\theta_h(\omega) = \mathbb{E}[\Psi_h'(X_t)]$, which I will call the population effect.
The main results in this paper build on an important identity popularized by Yitzhaki:96 and Angrist:00, which Rambachan:21 first applied to local projections. It turns out that the LP estimand has a causal interpretation even if the structural function $\psi_h$ is not linear. I present this result using the weakened regularity conditions of Kolesar:24. Throughout the paper, $\perp\!\!\! \!\perp$ denotes statistical independence and $\perp$ uncorrelatedness.
Assumption (ref) is a collection of regularity conditions that ensure that the LP estimand is well defined, the conditional mean function $g_h$ has a derivative almost everywhere and a specific weighted average of the derivative is finite. Assumption (ref) requires the shock $X_t$ and the other variables entering $Y_{t+h}$ to be independent. This ensures that the conditional mean function $g_h$ nonparametrically identifies the average structural function $\Psi_h$ so that the derivative of $g_h$ has a causal interpretation.
The following result is part of Proposition 1 of Kolesar:24:
The weight function $\omega_X$ is non-negative, integrates to one and is peaked around zero. The proof of Lemma (ref) effectively amounts to using the fundamental theorem of calculus and Fubini's theorem. If $\omega_X$ were the density of the shock $X_t$, LPs would estimate the population effect. For shocks that are Normally distributed, this is the case Stein:81. However, this is the only distribution with smooth density function and decaying tails that has this property.
Since commonly used shocks are often far from Gaussian Kolesar:24, LPs generally fail to estimate the population effect. Nevertheless, Lemma (ref) is reassuring: Even with a data generating process that is far from linear, LPs estimate a proper weighted average of causal effects. In particular, if the process has no size or sign nonlinearities in the shock $X_t$, LPs always estimate the unambiguous correct effect.\footnote{In this case, $\Phi'_h(x) \equiv b_h$ does not depend on $x$. Therefore, $\theta(\omega) = \int \omega(x) dx \cdot b_h = b_h$ for every weight function $\omega$. This is the average response of $Y_{t+h}$ of a shock $X_t$ of any size.} The next section shows that this result seamlessly carries over to state-dependent LPs.
This section considers state-dependent local projections of the form
where the data is generated by the structural function (ref), $f: \mathcal{S} \to \mathbb{R}^k$ is a function mapping states to interaction terms and $\beta^h \in \mathbb{R}^k$ is the regression coefficient. For example, in Ramey:18, $S_{t-1}$ is the unemployment rate and $f$ consists of two indicator functions defining a slack and expansionary state, respectively:
More examples will be discussed later on. The results are commonly interpreted as
This interpretation is clearly adequate if the specification (ref) fully captures the nonlinearities in the structural function $\psi_h$. Also, if $f(S_{t-1})$ consists of dummy variables, the logic of running separate regression on split sub-samples can be evoked. However, in many applications a more complex interaction variable is used and misspecification of the LP equation is possible. In general, some caution is required when interpreting higher-order terms in a linear regression. The coefficients of these terms do not correspond to Taylor coefficients of the structural function White:80 and LPs including nonlinear transformations of $X_t$ are not straightforward to interpret in a causal way.\footnote{See Proposition 2 of Kolesar:24 for an example with the regressor $X_t^2$. Caravello:24 more generally show how nonlinear terms in $X_t$ can be used to dis-entangle sign and size nonlinearities of shock effects.} Luckily, for the state-dependent setup considered here, the common interpretation turns out to be appropriate under mild conditions.
\noindentState Variable.---When estimating state-dependent LPs of the form (ref), the researcher is interested in the response of $Y_{t+h}$ to changes in $X_t$ conditional on some state $S_{t-1} \in \mathcal{S}$, where $\mathcal{S}$ is a possibly multi-dimensional state space. The state is allowed to be endogenous in the sense that $X_t$ can affect current and future realizations of the state. However, it will be important that the shock cannot affect past states. Many states of economic interest such as high unemployment states Ramey:18 or ZLB episodes Auerbach:16 fulfill this requirement. Notably, the recession index used in Auerbach:12 does not meet this criterion, since it is a centred moving average of the output growth rate.
\noindentCausal Effects.---Now I define conditional versions of the causal quantities used in Section (ref). First, define the conditional average structural function as
The only difference to the average structural function (ref) is the conditioning on the state level $s$ in addition to the shock level $x$. With slight abuse of notation, I use the same symbol for both functions. Similarly, for a weight function $\omega \geq 0$, $\int \omega(x) dx = 1$, define the conditional average effect
If $\omega$ is the shock density, $\theta_h(s; \omega)$ is equal to $\mathbb{E}[\Psi_h'(X_t,s)]$, which I call the population conditional effect. If $\psi_h$ is smooth, this is equal to $\mathbb{E}[\psi_h'(X_t, U_{h,t+h}) \mid S_{t-1} = s]$.
It will turn out that state-dependent LPs have a causal estimand under marginally stronger conditions than in the linear case. To ease notation, from now on let $f_{t-1}$ denote $f(S_{t-1})$. Also recall that $\perp$ and $\perp\!\!\! \!\perp$ denote uncorrelatedness and independence, respectively.
Assumption (ref) ensures that Lemma (ref) holds for the conditional measure depending on $S_{t-1}$ and Assumption (ref) ensures that the lagged state variable $S_{t-1}$ is independent of the shock $X_t$. Again, note that $X_t$ is allowed to influence current or future realizations of $S_{t-1}$.
The following result shows, that the state-dependent LP estimand is the projection coefficient of the conditional average effect $\theta_h(S_{t-1};\omega_X)$ on $f_{t-1}$:
To numerically verify and illustrate Proposition (ref), in Appendix (ref) I simulated data from a smooth transition VAR model á la Auerbach:12. In this setting, the causal effect of $X_t$ can be computed analytically and compared to the LP estimand.
Proposition (ref) shows that running a state-dependent local projection of the form (ref) yields the same estimand as regressing the unobserved average conditional effect $\theta_h(S_{t-1}; \omega_X)$ on the interaction term $f(S_{t-1})$. I use this insight to derive the causal estimand of common state-dependent LP specifications and propose an LP estimator that accounts for state dependence by re-weighting observations.
One popular specification of state-dependent LPs interacts $X_t$ with a binary state variable $S_{t-1}$. This is equivalent to running two linear LPs on split subsamples of the data and it follows immediately from Lemma (ref) that weighted averages of conditional average effects are estimated. However, in 19 of the 44 studies listed by Goncalves:24, the authors use a continuous state index, so this split-sample logic cannot be evoked. This is where Proposition (ref) comes to shine: It implies that the popular interaction with a logistic term pioneered by Auerbach:13b as well as similar specifications all approximate a conditional average effect. Throughout the subsection, I assume that Assumptions (ref), (ref) and (ref) are all met.
Specification 1: Binary States. Let $S_{t-1} \in \{0,1\}$ and consider a researcher running the regression
It follows from Proposition (ref) that the estimands satisfy
If $\beta_1^h \neq 0$, the effect of $X_t$ on $Y_{t+h}$ is commonly interpreted as depending on the state $S_{t-1}$. This is justified since the interaction term captures the difference between average conditional effects with the same weighting function for both states. In particular, if the effect of $X_t$ is larger in state 1 than in state 0 across all baseline shock levels $x$, the non-negativity of the weights $\omega_X$ ensures that $\beta_1^h > 0$. On the contrary, if $\beta_1^h \neq 0$, at least for some baseline shock levels $x$ the effect of $X_t$ on $Y_{t+h}$ is state-dependent.
Specification 2: Continuous State. Suppose $S_{t-1}$ is scalar, $\tilde f$ is a logistic function and the state-dependent LP
is estimated. This is the popular setup due to Auerbach:13b. The estimand $\beta_1^h$ satisfies
Therefore, if $\beta_1^h = 0$, the state index $\tilde f(S_{t-1})$ and the conditional average effect at $S_{t-1}$ with weights $\omega_X$ are uncorrelated. Note that (ref) does not depend on $\tilde f$ being logistic so it holds for general functions.
Specification 3: Series Expansion. Auer:21 address nonlinearities in the relationship between the state and the conditional effect by interacting $X_t$ with a polynomial basis in the state, i.e.
with some degree $P>0$. Proposition (ref) shows that the estimand satisfies
Therefore, one can use standard series approximation theory to justify $\sum_{p=0}^{P-1} s^p \beta_{p}^h \approx \theta_h(s;\omega_X)$ for sufficiently large $P$. The same logic applies to other choices of basis functions, such as wavelets or splines.
Suppose a researcher is interested in the effect of $X_t$ on $Y_{t+h}$ at some state level $s^* \in \mathcal{S}$, but $S_{t-1}$ is continuously distributed so she cannot take a subset of all observations that satisfy $S_{t-1} = s^*$. This is a common situation: If $S_t$ is a continuous index of the business cycle, effect estimates for a high and low value of $s^*$ are often reported. Usually, some functional form $f(S_{t-1})$ for the dependence of the effect on the state is assumed and $f(s^*)'\hat \beta^h$ is taken as the desired effect estimate. Since the true relationship between effect and state is unknown, misspecification of $f$ is possible. A natural approximation of the split-sample logic is to weight the observations according to some weight function $w: \mathcal{S} \to \mathbb{R}_+$.\footnote{This idea came from a comment of Haoge Chang to a presentation of this project.} This could be $w(s) = K(h^{-1}\lVert s-s^*\rVert)$, where $K$ is a kernel function and $h$ is a tuning parameter. Now weighting can be implemented by running the OLS regression
This regression is not of the form (ref). Expanding the fraction and using independence reveals, however, that
so $\beta^h$ is the re-scaled coefficient from the regression of $Y_{t+h}$ on $w(S_{t-1}) X_t$, which is of the form (ref). Now Proposition (ref) yields
which is the probability limit of a Nadaraya-Watson kernel regression of the conditional average effect $\theta_h(S_{t-1};\omega)$ on the state using weighting kernel $w$. If $\theta_h$ is sufficiently smooth and the bandwidth $h$ is small, $\beta^h \approx \theta_h(s^*; \omega_X)$. Compared to interactions with fixed functions $f$, such a weighted local projection might have the advantage that extrapolation bias from regions of $\mathcal{S}$ that are far away from $s^*$ is minimized. By a similar argument it can be shown that the estimand $\beta_{0}^h$ of the regression
is a locally linear estimator of $\theta_h(s; \omega_X)$. Since a locally linear estimator is known to be preferable to a locally constant estimator in many situations, the specification (ref) might have desirable approximation properties too. To my knowledge, up to now no empirical study has used weighted LPs to estimate state-dependent effects. However, the above discussion shows that such state-weighted LPs approximate a causal quantity and Proposition (ref) can be used to study its asymptotic properties.
State-dependent Vector Autoregressions (VARs) are among the most commonly used nonlinear time series models Granger:93, Auerbach:12. I show with simulations and analytically that the well known asymptotic equivalence between LPs and VARs Plagborg-Moller:21 breaks down in the state-dependent case. State-dependent VARs lack some desirable robustness properties of state-dependent LPs: Even in the absence of sign and size nonlinearities they may not recover the true effect of $X_t$ on $Y_{t+h}$ conditional on $S_{t-1} = s$. As a remedy, I derive an impulse response estimate based on state-dependent VARs that has the same probability limit as state-dependent LPs.
First, define state-dependent VARs as a projection model. Note that this section remains agnostic about the structural function, so the true data generating process might be arbitrarily non-linear.
Begin by stacking the shock $X_t$ and the outcome $Y_t$ in a vector
It simplifies the analysis to assume that the shock is independent of the past:
Next, define $P_s[\bullet|\bullet]$ as the projection operator with respect to the conditional expectation $\mathbb{E}[\bullet|S_{t-1}=s]$, where $S_{t-1}$ is some state variable. For simplicity, $S_{t-1} \in \{0,1\}$ is assumed throughout the section. Similarly, let $P[\bullet|\bullet]$ be the projection with respect to the unconditional expectation $\mathbb{E}[\bullet]$. With a binary state, the coefficients of the state-dependent LP
satisfy
Now the reduced form VAR conditional projection model can be defined via
where $\mathbb{E}[E_t \mathbf{Y}_{t-k} \mid S_{t-1}] = \mathbf{0}$ for all lags $k\geq 1$. From now on, let only the first lag coefficient be non-zero, i.e. $\Pi_k(s) = \mathbf{0}$ for all $k>1$ and write $\Pi(s) := \Pi_1(s)$. This is to ease notation and without much loss of generality due to the companion form. Each result of this section generalizes to the infinite-lag case.\footnote{The main technical detail that has to be added in the infinite-lag case is a square summability condition to ensure the infinite sum of the projection exists.} By applying the common recursive identification scheme, utilizing that $X_t$ is exogenous, there is a structural SVAR representation of $\mathbf{Y}_t$ in terms of projection coefficients:
where $A(S_{t-1})$ is lower triangular and $\mathbb{E}[X_t e_t^\perp \mid S_{t-1}]=0$.\footnote{Formally, denote the elements of the reduced form error as $(X_t,e_t)' = E_t$. Then the $e_t^\perp$ is defined via
Lastly, the contemporaneous slope coefficients are computed as
where $\text{chol}$ denotes the Cholesky decomposition. } Despite looking like a structural model, this representation is defined purely in terms of population moments and exists under minimal regularity conditions. The only economic assumption so far is $X_t$ being independent of the past. The orthogonalized error $e_t^\perp$, however, is allowed to be dependent with $X_t$ and over time.
After estimating the parameters of the projection model, impulse response estimates can be constructed in an iterative way. The most straightforward way to do this is computing
where $f$ stands for fixed state. This is the impulse response estimate used by Auerbach:13. They are aware that this estimate does not account for the possibility that the economy might move out of state $s$ between time $t-1$ and $t+h-1$. Since it is well known that LPs average over future state changes, it is no surprise that $\theta_{\mathit{VAR,h}}^f$ will be different from the LP estimand. An effect estimate that accounts for the possibility of future state changes would be
where $m$ stands for moving state. As derived by Goncalves:24, for a state-dependent VAR model with fully exogenous state and independent error terms this is the response of $Y_{t+h}$ to a shock $X_t$ of arbitrary size.\footnote{See Proposition 3.1 of Goncalves:24. For this data generating process, $\theta^m_{\mathit{VAR,h}}(s)$ is both what they call the conditional average response and the conditional marginal response.} Since this estimate averages over future paths of the state, it is a natural comparison to the LP estimand.
To investigate the relationship between state-dependent VAR and LP based impulse response estimates, recall the structural SVAR representation (ref) and note that by assumption and construction, respectively,
This implies that $(A(s))_{21}$ is a conditional projection coefficient:
so the state-dependent LP and both VAR estimands $\theta^f_{\mathit{VAR,h}}(s)$, $\theta^m_{\mathit{VAR,h}}(s)$ agree on impact.\footnote{This equivalence on impact was already noted by Auerbach:13. For longer horizons $h>0$, however, they focus on differences between LP and VAR due to varying future states or holding them fixed.} For the horizon $h=1$, iterate (ref) forward and write in terms of expected slope coefficients:
The error term $\mathcal{E}_{t+1}^\Pi$ is the forecast error of the parameter at $t+1$ times the projection error of the endogenous variables at $t$. The term $\mathcal{E}_{t+1}^P$ is the one-step projection error of the endogenous variables at $t+1$. If the state $S_t$ is fully exogenous\footnote{If the state can be influenced by current or past values of $X_t$, $\theta^{m}_{\textit{VAR,h}}$ might not be the correct effect estimate even in the favorable case of independent errors Goncalves:24.}, this provides a condition for equivalence between $\theta^m_{\mathit{VAR,h}}(s)$ and the state-dependent LP estimand:
The condition of Proposition (ref) is not necessarily satisfied. Section (ref) presents a case where $\mathbb{E}[\mathcal{E}_{t+1}^\Pi X_t \mid S_{t-1}] \neq 0$ and also the condition $\mathbb{E}[\mathcal{E}_{t+1}^P X_t \mid S_{t-1}] = 0$ can be violated.\footnote{A simple example is $Y_t = S_{t-2}X_{t-1}$. For this process, $e_{t+1} = (S_{t-1}-\mathbb{E}[S_{t-1} \mid S_t])X_t$. One can verify that $\mathbb{E}[\mathcal{E}^P_{t+1} X_t \mid S_{t-1}] = (0,(S_{t-1}-\mathbb{E}[\mathbb{E}[S_{t-1}\mid S_t]\mid S_{t-1}])\mathbb{V}[X_t])' \neq 0$.} The reason for the latter is that orthogonality with respect to $\mathbb{E}[\bullet|S_t]$ does not imply orthogonality with respect to $\mathbb{E}[\bullet | S_{t-1}]$. Therefore, for horizon $h > 0$, $\theta^m_{\mathit{VAR,h}}(s)$ and the LP estimand differ in general---even in the special case of a fully exogenous state $S_t$.
Even though the VAR based estimates $\theta_{\mathit{VAR,h}}^f$ and $\theta_{\mathit{VAR,h}}^m$ both differ from the LP estimand, there is still a connection between both methods. Consider $h+1$ state-dependent VAR models where each successive model shifts the state back one more lag:
The orthognalized projection error is of the form $E_t^{k,\perp} = (X_t, e_t^{k,\perp})'$. These projection models are just as described in (ref) with the difference that for the $k$'th projection model the conditional expectation $\mathbb{E}[\bullet | S_{t-1}=s]$ is replaced with $\mathbb{E}[\bullet| S_{t-k} = s]$. Iterating forward, using the $k$'th model for the $k$'th prediction step\footnote{This iterative combination of multiple different VAR models is similar in spirit to Ludwig:24's (Ludwig:24) VAR-sequence. Using this technique, he is able to prove a finite sample equivalence between linear VARs and LPs. However, he combines linear VAR models with different lag lengths, while I combine state-dependent VAR models that condition on different lags of the states.} gives the representation
See Appendix (ref) for a recursive formula of the parameters in the more general case of infinitely many lags of the endogenous variables. This representation yields a third VAR-based impulse response estimate
where $b$ stands for backshifted state. It turns out that $\theta^b_{\mathit{VAR,h}}(s)$ is identical to the state-dependent LP estimand.
Like the equivalence results of Plagborg-Moller:21 and Ludwig:24, Proposition (ref) is essentially an application of the law of iterated projections. Projecting $\mathbf{Y}_{t+h}$ on $\text{span}\{\mathbf{Y}_{t+h-1},\mathbf{Y}_{t+h-2},...\}$, then on $\text{span} \{\mathbf{Y}_{t+h-2},\mathbf{Y}_{t+h-3},...\}$ and so on yields the same result as directly projecting on the smallest space, $\text{span} \{X_{t}, \mathbf{Y}_{t-1},...\}$. The iterative procedure corresponds to VAR-based methods, the direct procedure to the LP. The law of iterated projections cannot be applied to the impulse response estimates based on a single state-dependent VAR model that are considered in the previous subsection. The reason is that the VAR prediction conditions on a different lag of the state at every iteration: To predict $\mathbf{Y}_{t}$ given previous values condition on $S_{t-1}$, to predict $\mathbf{Y}_{t+1}$ condition on $S_t$, to predict $\mathbf{Y}_{t+2}$ condition on $S_{t+1}$, and so on. As a result, each projection step uses a different inner product so the law of iterated projections does not hold. Using $h+1$ state-dependent VAR models to compute $\theta^b_{\mathit{VAR,h}}$ ensures that each projection step uses the same inner product as the state-dependent LP such that both methods are equivalent again. Note that the equivalence holds regardless of whether the state $S_t$ is exogenous. Figure (ref) visualizes the different prediction steps underlying each method.
Proposition (ref) has useful practical implications: The estimator $\theta^b_{\mathit{VAR,h}}$ is easy to compute, it does not rely on knowledge about the law of movement of the state like the moving state estimator $\theta^m_{\mathit{VAR,h}}$ defined in (ref). But unless $\theta^f_{\textit{VAR,h}}$ defined in (ref) it also does not implicitly assume that the state remains the same between impulse and response. At the same time, $\theta^b_{\mathit{VAR,h}}$ inherits the favorable asymptotic properties of state-dependent LPs that are presented in this paper. Therefore, the estimator $\theta^b_{\textit{VAR,h}}$ might be an attractive option for researchers who prefer to use VARs for convention or finite sample properties while wishing to benefit from the robustness properties of state-dependent LPs. The next section compares state-dependent LPs to the various VAR based estimators using a numerical example.
To evaluate the asymptotic properties of state-dependent VARs and LPs, consider a simple DSGE growth model. Income consists of output produced with an AK-technology and transfers or windfall income:
The state $S_t$ is a binary recession index, $A(s)$ is the productivity in state $s$, $\nu$ is a perturbation parameter and $\nu B(s)$ is the standard deviation of windfall income in $s$. The state is assumed to move exogenously with known Markov transition matrix
Naturally $A(1)<A(0)$, so the economy is more productive in expansions. To close the model, assume there is a representative household with CRRA preferences that owns the capital stock:
Capital depreciates fully, such that
This can be justified by letting one period represent multiple years. Full depreciation is a convenient assumption popularized by Brock:72 to obtain a closed form solution. As $\nu \to 0,$\footnote{This amounts to assuming that agents do not consider future windfall income when making savings decisions.} income evolves as
where $\phi(s)$ is a savings rate that has to be computed numerically. See Appendix (ref) for details. With high enough intertemporal substitution, $\sigma > 1$, the economy will save more in good times and spend more in bad times. Table (ref) displays the parameter choices for the model. It is calibrated in a way that income $Y_t$ experiences periods of endogenous growth and shrinkage but is stationary overall. The resulting savings rates in good and bad times are $\phi(0) \approx 0.86$ and $\phi(1) \approx 0.77$, respectively. This income process is well suited to study the properties of state-dependent LPs and VARs for three reasons: (i) It allows for analytical computation of the true state-dependent effect of $X_t$ on $Y_{t+h}$, (ii) both state-dependent LP and VAR are misspecified when applied to this process, allowing for a fair comparison and (iii) the average structural function $\Psi_h(x,s)$ is linear in $x$. Therefore, the effect of interest is unambiguously defined: It does not depend on the sign or size of the shock. This lets me assess which method estimates the correct effect and which does not without committing to a particular effect of interest.
Figure (ref) shows the true impulse response of the model and compares it to four econometric estimands. The left two panels show impulse responses conditional on the lagged recession state, the right panel shows the unconditional impulse response as comparison. If a shock hits after a recession, $S_{t-1} = 1$, it raises income by more than after an expansion, which is by assumption. However, the effect evaporates more quickly after a recession, since both savings rate and productivity are lower. Local projections estimate the true effect in all three cases. This is as expected given Proposition (ref). The figure also plots the VAR-based estimands $\theta^f_{\mathit{VAR,h}}$, $\theta^m_{\mathit{VAR,h}}$ and $\theta^b_{\mathit{VAR,h}}$ that are defined in (ref), (ref) and (ref), respectively. Of those three, only my novel estimate $\theta^b_{\mathit{VAR,h}}$ recovers the true effect, which verifies Proposition (ref). If the state is held fixed, the VAR exaggerates the difference between effects after recessions and expansions. The reason is that both the true IRF and the LP estimand account for the possibility of switching to the other state after the shock hits, while $\theta^f_{\textit{VAR,h}}$ implicitly assumes the economy remains in the initial state. The difference between $\theta^m_{\textit{VAR,h}}$ and the LP estimand is more novel: Even when (correctly) accounting for the possibility of state changes, the IRF based on a single VAR model asymptotically yields a different effect estimate than the LP.
To understand why $\theta^m_{\textit{VAR,h}}$ is asymptotically different from the LP estimand in this case, consider a slightly simplified version of the income process with $A(0) = A(1)=1$ but $\phi(0) \neq \phi(1)$:\footnote{This has the advantage that the state-dependent VAR only has one non-zero lag, which eases the exposition. Of course, when solving the model with $A(0) = A(1)$, the savings rates would be the same in both states. One can think about the simplification as follows: The productivities in both states changed, but the agent's policy rules did not change (yet).}
The forecast error of the parameters times the reduced form errors is then
This term is not conditionally orthogonal to $X_t$:
Therefore, state-dependent LP and VAR disagree for $h=1$ if the savings rate $\phi(S_t)$ and the impact of windfall income shocks $\nu B(S_t)$ are correlated.
This section considers LPs of the form
where $f(S_{t-1}) Z_t$ is used as an instrument. For example, $X_t$ could be government spending, which has a large endogenous component, and $Z_t$ could be some government spending shock. This is a common setup, 19 out of the 44 studies surveyed by Goncalves:24 use some kind of 2SLS estimator for state-dependent LPs. This section shows that state dependent LP-IV's identify a weighted average of conditional marginal effects. However, the weights now generally depend on the states. To interpret state-dependent LP-IVs in the usual way, the data generating process has to be restricted.
\noindentEconometric Setup.---Again, suppose the outcome $Y_{t+h}$ is determined by the structural functions $\psi_h$ defined in (ref). However, now $X_t$ is not assumed to be a shock, but is more generally determined by
where $Z_t$ is some instrument and $V_t$ is generally related to $U_{h,t+h}$, so the regressor is endogenous. It will turn out useful to marginalize the structural function $\psi_h$ over $U_{h,t+h}$, conditional on some realization $(z,v)$ of $(Z_t,V_t)$. Define the IV average structural function as
Similarly, define the conditional IV average structural function as
These functions define the average value of $Y_{t+h}$ given fixed outcomes of the shock $Z_t$ and the unobserved component $V_t$.
Equipped with the above definition and the chain rule, a causal expression of the linear LP-IV estimand can be derived from Lemma (ref) under mild conditions.
Assumption (ref) is a collection of regularity conditions, Assumption (ref) ensures monotonicity and Assumption (ref) is an exogeneity condition.
Note that in the case of an observed shock, $Z_t = X_t$ and $V_t$ is a constant, so $X'(z,v) \equiv 1$, $\Psi_{IV,h} = \Psi_h$ and (ref) collapses to
so Lemma (ref) generalizes Lemma (ref). The result shows that LP-IV still identifies weighted averages of causal effects. But in addition to the weight $\omega_Z$ that depends on the marginal distribution of $Z_t$, there is now a weight across the $(Z_t,V_t)$ dimension that depends on the joint behavior of $Z_t$ and $X_t$. When the instrument $Z_t$ has a large effect on $X_t$ for a given $(Z_t,V_t)$-pair, the corresponding effect of $X_t$ on $Y_{t+h}$ will receive more weight than when the instrument affects $X_t$ only little.
Before deriving an analogous result to Proposition (ref), some regularity conditions as well as independence of instrument and lagged state have to be assumed. Again, let $f_{t-1}$ denote $f(S_{t-1})$.
This set of assumptions ensures that the LP-IV estimator and all the causal quantities used in Lemma (ref) exist in conditional form. The following result shows that state-dependent LPs estimate a weighted average of conditional effects analogous to (ref):
Proposition (ref) shows that state-dependent LP-IVs estimate the same causal quantity as linear LP-IVs---just in a conditional way. If $f$ is misspecified, this quantity is approximated in a weighted least square sense, where the non-negative weights $\theta_X(s)$ indicate the strength of the instrument in a given state.\footnote{$\theta_X(s)$ is just the conditional average effect used in Section (ref) and Proposition (ref) with $X_t$ being the dependent variable and $Z_t$ the shock. It is the regression coefficient of $X_t$ on $Z_t$ in the sub-sample where $S_{t-1} = s$.} Again, if the interaction term consists of dummy variables, state-dependent LP-IVs directly estimate $\theta_{\textit{IV,h}}(s)$. This estimand is an integral over a product of three components: (i) The effect of interest at a certain instrument and state realization, $\Psi_{\textit{IV,h}}'(z,s;V_t)$, (ii) the weight $\omega_Z$ and (iii) the weight $\kappa(z,V_t):= X'(z,V_t)/\theta_X(s)$ that corresponds to the effect of the instrument on the regressor $X_t$. The first weight $\omega_Z$ only depends on the marginal distribution of $Z_t$ and therefore is identical across states and applications. The second weight $\kappa$, however, depends on the joint distribution of $(Z_t,X_t)$ and can vary across states. This makes it hard to correctly interpret state-dependent LP-IV coefficients: The result $\theta_{\textit{IV,h}}(1) > \theta_{\textit{IV,h}}(0)$ would commonly be interpreted as $X_t$ having a stronger effect on $Y_{t+h}$ in state 1 than in state 0. However, the result could well be driven by differences in the weighting scheme, i.e. state dependence of the effect of $Z_t$ on $X_t$, which is not actually of interest. The next section shows that with certain model restrictions, the common interpretation of LP-IVs is still valid. However, the last example shows that in the absence of such restrictions this common interpretation can easily fail.
If the data generating process features arbitrary nonlinearities, no strong conclusions can be drawn from state-dependent LP-IVs. For this, either the relationship between regressor and outcome or instrument and regressor has to be restricted. The next two examples demonstrate how this works.
Sometimes, one might know more about the relationship between the instrument $Z_t$ and $X_t$ than about the structural function $\psi_h$. Knowledge of the mechanism linking $Z_t$ and $X_t$ can come from the construction of the shock or from investigating validity of the exogeneity assumption.
The preceding examples hinge on either $Y_{t+h}$ being linear in $X_t$ conditionally on $S_{t-1}$ or $X_t$ being linear in $Z_t$. If neither of those holds, the common interpretation of state-dependent LP-IVs can be misleading.
The study of LP-IVs in a nonlinear environment is closely tied to microeconometric work on limited compliance. Unrestricted linearity of the structural function $\psi_h$ effectively corresponds to (unobserved) treatment effect heterogeneity. Having that in mind, the second weight in (ref) can be understood as indicating compliance, i.e. how strong the treatment reacts to the instrument. While in binary treatment settings compliance is an on-off decision, in the continuous case it is itself a continuum. In microeconometrics, the treatment effect weighted by the compliance decision is called the Local Average Treatment Effect (LATE), which corresponds to the IV estimand. Indeed, this seminal result by Imbens:94 is a special case of Lemma (ref).
The three examples in Section (ref) can also be re-interpreted in the language of microeconometrics: It is well known that limited compliance poses no problems, if every individual has the same treatment effect (Example (ref)). In this case, IVs estimate the average treatment effect (ATE), which is equal to every other weighted average of treatment effects. If compliance is independent of the effect size (corresponding to $X_t$ being linear in $Z_t$), IVs have the same estimand as a regression using data where the treatment is perfectly randomized (Example (ref)). Lastly, Example (ref) corresponds to having two populations with the same treatment effect distribution but different compliance decisions: In the first population, which corresponds to the expansion state, compliance is perfect and so the ATE is estimated. In the second population (the recession state), individuals with higher treatment effect are more likely to comply, so the LATE is higher than the ATE. The resulting difference in the IV estimands is not due to differences in the effect distribution of interest but due to compliance.
This paper shows that state-dependent LPs estimate weighted averages of conditional marginal effects. The result holds without making parametric assumptions and the shock of interest is allowed to influence current and future realizations of the state. The weighted average of effects is generally different from the average response to a shock of both marginal and strictly positive size. Unless one commits to specific functional forms, no stronger guarantee holds even for linear LPs. Therefore I conclude that generally state-dependent LPs are just as valid as linear LPs. If the shock of interest is observed, the weights on the causal effects are identical across states and applications. Therefore, a non-zero interaction coefficient implies state dependence of the effect of interest. If the relationship between state and effect is misspecified, state-dependent LPs approximate the weighted average of conditional marginal effects in the familiar MSE sense. Since asymptotic equivalence between VARs and LPs breaks down in the state-dependent case, those favorable properties do not carry over to conventional state-dependent VAR estimates. As a remedy, I propose a VAR-based impulse response estimate that is easy to compute and converges to the state-dependent LP estimand. This should give researchers more freedom to choose between both methods based on finite sample considerations.
My analysis also raises an issue that warrants caution: When using instrumental variables, the weights on the effects depend on the joint distribution of instrument and regressor. If the instrument $Z_t$ affects the regressor $X_t$ strongly in a certain state, the corresponding effect of $X_t$ on $Y_{t+h}$ receives disproportionate weight. As a consequence, non-zero interaction coefficients in state-dependent LP-IVs can be due to differences in the weighting scheme that have nothing to do with the effect of interest. Knowledge about the relationship between instrument and regressor or regressor and outcome can rule out this option.
Another caveat concerns the assumptions: While linear data generating processes usually require orthogonality conditions for identification, papers studying LPs in a nonparametric setting assume that the shock $X_t$ is serially independent and independent of the nuisance variable $U_{h,t+h}$ Rambachan:21, Caravello:24, Kolesar:24. This paper additionally assumes that the shock $X_t$ is independent of the past state $S_{t-1}$. So far, this strengthening of assumptions has not been discussed a lot. However, it might be problematic: While the fact that shocks are not linearly predictable using past information is intimately tied to the notion of a shock and rational expectations econometrics, the same cannot be said about higher-moment dependence. For example in a financial context, the volatilities of excess returns are often clustered and way easier to forecast than its levels. Thus, being agnostic about the functional form of the data generating process comes at a cost. The required independence conditions should be taken seriously and tested empirically.
\singlespacing \setlength\bibsep{0pt}