Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
21,984 characters · 5 sections · 12 citation commands
Nonlinearity in Dynamic Causal Effects: Making the Bad into the Good, and the Good into the Great?
\onehalfspacing We are grateful for the opportunity to discuss this insightful paper by Michal Koles\'{a}r and Mikkel Plagborg-M{\o}ller (KP hereafter). For more than three decades, program evaluation, causal inference, and nonseparable structural equation models have been prominent areas of research in econometrics. These topics have interacted with the “credibility revolution" in applied microeconomics, which has driven methodological advancements and strengthened the credibility of evidence obtained through empirical research. Key features of these approaches are (i) minimal restrictions on the heterogeneity of causal effects and the functional form of causal equations and (ii) providing a transparent interpretation for causal estimands in order to clarify their relevance and usefulness for policy analysis. KP shares these features, and we believe that it offers novel and useful perspectives on causal analysis in macroeconometrics.
To motivate KP's approach, it is useful to contrast it with conventional econometric research, in which a causal model is combined with identification assumptions to yield a statistical estimand: \[ \mbox{(Causal model) + {(Identification)}} \Rightarrow \mbox{{(Statistical estimand)}}. \] A more recent contrasting approach, which we call “reverse econometrics”, begins with a statistical estimand which is then combined with identification assumptions to obtain a causal interpretation: \[ \mbox{{(Statistical estimand)}+{(Identification)}} \Rightarrow \mbox{(Causal parameter)}. \] Research in reverse econometrics has been growing. Influential papers include imbens1994identification which introduces the local average treatment effect interpretation of the two-stage least squares estimand, angrist1995estimating which derives the weighted-average conditional average treatment effect interpretation of the ordinary least squares estimator with additive controls, and de2020two which shows the weighted-average treatment effect interpretation of the two-way fixed effects estimator. A particular merit of the reverse econometrics approach is that it shows the robustness and policy relevance of existing, statistically well-studied estimands in the presence of heterogeneity or nonlinearity. KP performs reverse econometric analysis of the impulse response (IR) estimands that are dominant in linear time-series models. See rambachan2021when, casini2024identification, and Koo_etal_2023 for other reverse econometric analyses of impulse response estimands; see also chen2024potential, which seeks to offer alternative interpretations of linear causal estimands.
In the following, we make three points. First, we argue that the presence of negative weights is not always problematic. We present a novel proposition that shows conditions under which a “bad” estimand can be salvaged by transforming it from a weighted-average of structural derivatives with negative weights into one with positive weights without altering the estimand's value. Second, we question how nonstationary causal effects can be accommodated. Third, we discuss the possibility of elevating the “good” properties of the estimands to “great” by linking the reverse econometrics to policy decision making. Throughout our discussion, we maintain the definitions and notation of KP.
KP focuses on the estimation of the average marginal effect, which can be expressed as
where $g(x)$ represents the average of a structural function of randomly assigned shocks $x$, and $\omega(\cdot)$ is a weight function determined by the associated estimation procedure. To simplify notation, we suppress the dependence of the statistical objects on the lag index $h$ and the time index $t$.
Throughout their paper, KP views nonnegative weights, $\omega(\cdot) \geq 0$, as the gold standard for an estimand to represent a “meaningful causal summary”. Their justification is that if the region where $\omega(\cdot)$ is negative has a positive measure, then there exists a $g'(\cdot)$ that takes sufficiently large positive value therein to render $\beta$ negative even when $g^{\prime}(\cdot)$ is positive everywhere in $\mathcal{X}$. This claim is valid if one is fully agnostic about $g^{\prime}(\cdot)$. What is missing, however, is an assessment of the likelihood that negative weights lead to a reversal of the sign of $\beta$ in typical macroeconomic applications.
As KP note as motivation: “...both macroeconomic theorists and policymakers think nonlinearities are important”. This statement reflects that macroeconomic theorists and policymakers have some ideas about what are admissible shapes and structures of the average marginal effect rather than being fully agnostic about $g'(x)$. Thus, even in cases where $\omega(x)$ can take negative values, there can be some credible restrictions available on the shape of $g'(x)$ under which the negative weights are rendered inconsequential. Since we use macroeconomic theory to guide empirical implementations in macroeconomics, we should also exploit it to assess whether the negative weight is a serious concern or not in a given context.
In this spirit, we show a proposition below to illustrate that shape restrictions on $g^{\prime}(\cdot)$ can resolve the issue of negative weight. In addition, these restrictions can even guarantee the existence of a representation $\beta = \int_{\mathcal{X}} \tilde{\omega}(x) g^{\prime}(x) dx$ with nonnegative weights $\tilde{\omega}(\cdot) \geq 0$, where $\tilde{\omega}(\cdot)$ is a transformation of the original potentially negative weights $\omega(\cdot)$.
To provide a formal statement, let us partition the domain of the shocks $\mathcal{X}$ into regions according to the sign of $\omega(\cdot)$: \[ \mathcal{X}^{-}:=\{x:\omega(x)<0\},\;\mathcal{X}^{+}:=\{x:\omega(x)\geq0\}. \] Our proposition restricts $g'(x)$ to behave similarly in the negative weight region $\mathcal{X}^{-}$ and the positive weight region $\mathcal{X}^{+}$. Given a weight function \(\omega(x)\) and a marginal effect function \(g(x)\), we aim to find an injective link function $Q(\cdot)$, which is continuously first-order differentiable on $\mathcal{X}^-$ (except for a finite number of points) and maps each $x \in \mathcal{X}^{-}$ to some $Q(x) \in \mathcal{X}^{+}$ with an equal value of $g'(\cdot)$, i.e.,
In addition, we require the link function to have the property that the weight at every $x \in \mathcal{X}^{-}$ is dominated by the weight at $Q(x) \in \mathcal{X}^{+}$ in the sense that
Availability of such link function $Q$ satisfying conditions ((ref)) and ((ref)) constrains the shape of $g'(x)$ in such a way that we can shift $g'(\cdot)$ on $\mathcal{X}^{-}$ to match $g'(\cdot)$ on $\mathcal{X}^{+}$ subject to the positive net weight condition ((ref)).
Define $G(\mathcal{X}^{-}) = \{Q(x) : x \in \mathcal{X}^{-}\}$. Assuming that ((ref)) and ((ref)) hold, consider the transformed weights $\tilde{\omega}(\cdot)$ defined as follows:
These transformed weights $\tilde{\omega}(\cdot)$ are nonnegative by construction and supported on $\mathcal{X^{+}}$. Under additional conditions on the Jacobian of $Q(\cdot)$, the next proposition shows that we can express $\beta = \int_{\mathcal{X}} \omega(x) g'(x) dx$ as a weighted average of $g'(x)$ with positive weights $\tilde{\omega}(x)$.
See Appendix (ref) for a proof. This proposition, which is new to the literature to our knowledge, shows that weighted-average estimands with negative weights $\omega(\cdot)$ can be salvaged by finding an alternative representation with positive weights $\tilde{\omega}(\cdot)$. The existence of $\tilde{\omega}(\cdot)$ relies on $g'(x)$ satisfying a set of shape conditions guaranteeing the existence of the link function characterized by ((ref)), ((ref)), and ((ref)). With the additional condition ((ref)) imposed, we can also ensure that the transformed weights $\tilde{\omega}(\cdot)$ sum to the same value as the original weights $\omega(\cdot)$. The following examples show how this proposition can be applied.
Our intention with Proposition (ref) is to highlight an unexplored connection between restrictions on $g'(\cdot)$ and the harmlessness of negative weights. We do not argue that the conditions of this proposition are likely to be met in all common macroeconometric applications, but further research on salvaging negative weight estimands through restrictions on $g'(\cdot)$ may be an interesting direction for future research following the contributions of KP.
To simplify exposition, Proposition (ref) assumes the existence of a differentiable link function $Q(\cdot)$. We can obtain a more general version of this proposition, which does not involve a link function, by imposing a dominance condition on the measures induced by $\omega(\cdot)$ and the mappings between $g^{\prime}(\cdot)$ and $\mathcal{X}^{-}$ and $\mathcal{X}^{+}$. See Proposition (ref) in Appendix (ref) for more details. Additionally, in Example (ref) of Appendix (ref), we also revisit KP's Proposition 2 to show our approach can restore the interpretation of the linear regression coefficient as an average-effect estimator with positive weights when the shock is positive.
One of the fundamental assumptions underlying the analysis of KP is that, while the average structural equation can depend on the time lag between the period in which the outcome is observed and the period in which treatment is given, it does not depend on the time index $t$. That is, the causal effect is assumed to be stationary. However, in reality, nonstationary causal effects can be as great a concern as nonlinear structural equations.
For a binary $X_t$ which satisfies sequential unconfoundedness and has known propensity scores, bojinov2019time shows that the regression coefficient of $Y_t$ on $X_t \in \{0, 1\}$ estimated with inverse propensity score weighting (i.e., the difference in inverse-propensity-score weighted means between $X_t = 1$ and $X_t = 0$ observations) provides an unbiased estimate of
\[ \frac{1}{T} \sum_{t=1}^T \mathbb{E}[\psi_t(1, U_t) - \psi_t(0, U_t) \mid \mathcal{F}_{t-1}], \quad \mathcal{F}_{t-1} = \sigma(Y_{t-1}, X_{t-1}, Y_{t-2}, X_{t-2}, \dots), \] where $\mathcal{F}_{t-1}$ denotes the set of conditioning variables (information set) consisting of observable variables up to $t-1$. Each summand can be viewed as the conditional average treatment effect at $t$ given the history $\mathcal{F}_{t-1}$, with the structural equation $\psi_t(\cdot, \cdot)$ allowed to vary in $t$. In turn, the whole estimand can be viewed as a sample average of conditional average effects over the sampled periods. casini2024identification shows a similar result when $X_t$ is continuous in a high-frequency context with infill asymptotics.
In the case of a continuous $X_t$, can the regression coefficient $\hat{\beta}$ be interpreted as an average of causal effects even if the structural equation is nonstationary? Can we allow instability in the economic environment to the extent that the time series of observables have unit roots? Can we improve the statistical precision of the estimator or the interpretability of the estimand by finding alternative methods that place greater weight on more recent observations? Answering these questions would help clarify whether the unrestricted heterogeneity of causal effects allowed in the cross-sectional setting can be extended to the time-series setting. To our knowledge, little is known about these questions beyond the binary case.
KP argues that the average marginal effect identified by $\beta$ is an estimand of interest as it can be estimated well even when the sample size is limited, as is common in empirical macroeconomic research. This argument is appealing and shares a common motivation with the procedure of selecting estimands based on their efficient asymptotic variances, as proposed in Crump_etal_2009 for the case of treatment effect estimation in microeconometrics.
A counterargument is that, instead of narrowing the choice of estimand based on identifiability or estimation precision, one should choose the estimand of interest based on the ultimate goal of analysis and the type of policy questions that researchers wish to investigate. This raises the question of how $\beta$ can be connected to counterfactual policy analysis and ex ante policy decisions. However, given that the analytical expression of the weights in $\beta$ is driven purely by the characteristics of the data generating processes, we doubt the existence of a direct link between $\beta$ and a planner's policy decision problem.
Making $\beta$ policy-relevant would turn “the good” into “the great”. One way to achieve this is if, based on economic theory or available background knowledge, we can justify the assumption that the sign or magnitude of $\beta$ is informative for current or future optimal policy. Building upon the potential outcome time-series framework angrist2018semiparametric, bojinov2019time, rambachan2021when, kitagawa2022policy formulates a policy choice problem in the time-series setting with a binary $X_t$ and studies ex ante policy choices while accounting for nonstationary causal effects. Our approach relies on the “invariance of welfare ordering”. This assumption states that the ex post welfare ranking of past policies provides reliable information about the ex ante welfare ranking that is relevant to current policy decisions. For the case of binary $X_t$, a sufficient condition for the welfare ordering to be invariant is that the sign of $\beta$ coincides with the sign of the welfare impact of the policy if it were implemented today.
Extending this type of approach to the case of a continuous $X_t$ is an area that remains unexplored. Establishing how an estimand involving negative weights in its weighted-average interpretation interacts with the validity of the invariance of welfare ordering or a similar assumption could be an insightful exercise.
KP have performed a thorough and insightful reverse econometric analysis of the estimands that are commonly used in empirical macroeconomics. We have presented sketch ideas on how the bad can be salvaged (Comment 1) and how the good can be made great (Comment 3), but leave further development and assessment of the usefulness of these ideas to future research.
We want to highlight that KP contributes to econometrics more broadly by bridging microeconometrics and macroeconometrics, highlighting the potential for greater synergies between these two areas and paving the way for future research that benefits from the strengths of both micro and macroeconometric analysis.