Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
131,158 characters · 20 sections · 114 citation commands
Revisiting Event Study Designs: Robust and Efficient Estimation
Event studies are one of the most popular tools in applied economics and policy evaluation. An event study is a difference-in-differences (DiD) design in which a set of units in the panel receive treatment at different points in time. In this paper, we investigate the robustness and efficiency of estimators of causal effects in event studies, with a focus on the role of treatment effect heterogeneity. We first develop a simple econometric framework that delineates the identification assumptions from each other and from the estimation target, defined as some average of heterogeneous causal effects. We then apply this framework in three ways. First, we analyze the conventional practice of implementing event studies via two-way fixed effect Ordinary Least Squares (TWFE OLS) regressions and show how the implicit conflation of different assumptions leads to biases. Second, leveraging event study assumptions in an explicit and principled way allows us to derive the robust and efficient estimator, along with appropriate inference methods and tests. The estimator takes an intuitive “imputation” form when treatment-effect heterogeneity is unrestricted. Finally, we illustrate the practical relevance of our approach in an application estimating the marginal propensity to spend (MPX) out of tax rebates; our MPX estimates are lower than in prior work, implying that fiscal stimulus is less powerful than commonly thought.
Event studies are frequently used to estimate treatment effects when treatment is not randomized, but the researcher has panel data allowing them to compare outcome trajectories before and after the onset of treatment, as well as across units treated at different times. By analogy to conventional DiD designs without staggered rollout, event studies are commonly implemented by two-way fixed effect regressions, such as
where outcome $Y_{it}$ and binary treatment $D_{it}$ are measured in periods $t$ and for units $i$, $\alpha_{i}$ are unit fixed effects (FEs) that allow for different baseline outcomes across units, and $\beta_{t}$ are period fixed effects that accommodate overall trends in the outcome. Specifications like (ref) are meant to isolate a treatment effect $\tau$ from unit- and period-specific confounders. A commonly-used dynamic version of this regression includes “lags” and “leads” of the indicator for the onset of treatment, to capture treatment effects for different “horizons” since the onset of treatment and test for the parallel trajectories of the pre-treatment outcomes.
To understand the problems with conventional two-way fixed effect estimators in event-study designs and provide a principled econometric approach to overcoming these issues, in (ref) we develop a simple framework that makes the estimation targets and underlying assumptions explicit and clearly isolated. We suppose that the researcher chooses a particular weighted average (or weighted sum) of heterogeneous treatment effects they are interested in estimating. We make (and later test) two standard DiD identification assumptions: that potential outcomes without treatment are characterized by parallel trends and that there are no anticipatory effects. We also allow for \textemdash but do not require \textemdash an auxiliary assumption that the treatment effects themselves follow some model that restricts their heterogeneity for a priori specified economic reasons. This explicit approach is in contrast to regression specifications like (ref), both static and dynamic, which implicitly conflate choices of estimation target and identification assumptions. Our framework covers a broad class of empirically relevant estimands beyond the standard average treatment-on-the-treated (ATT), including heterogeneous treatment effects by observed covariates and ATTs at different horizons that hold the composition of units fixed.
Through the lens of this framework, in (ref) we uncover a set of challenges with conventional event-study estimation methods and trace them back to a mismatch between estimation target, identification assumptions, and the flexibility of the regression specification. First, we note that failing to rule out anticipation effects in “fully-dynamic” specifications (with all leads and lags of the event included) leads to an underidentification problem when there are no never-treated units, such that the dynamic path of anticipation and treatment effects over time is not point-identified. We conclude that it is important to separate out testing the assumptions about pre-trends from the estimation of dynamic treatment effects under those assumptions. Second, implicit assumptions of homogeneous treatment effects embedded in static DiD regressions like (ref) may lead to estimands that put negative weights on some long-run treatment effects. With staggered rollout, regression-based estimation leverages comparisons between groups that got treated over a period of time and reference groups which had been treated earlier. We label such cases “forbidden comparisons.” Indeed, these comparisons are only valid when the homogeneity assumption is true; when it is violated, they can substantially distort the weights the estimator places on treatment effects, or even make them negative. Third, in dynamic specifications, implicit assumptions about treatment effect homogeneity across groups first treated at different times lead to the spurious identification of long-run treatment effects for which no DiD comparisons valid under heterogeneous treatment effects are available. The last two challenges highlight the danger of imposing implicit treatment effect homogeneity assumptions instead of allowing for heterogeneity and explicitly specifying the target estimand. We show that these challenges are not resolved by trimming the sample to a fixed window around the event date.
From the above discussion, the reader should not conclude that event study designs are plagued by fundamental problems. On the contrary, these challenges only arise due to a mismatch between treatment effect heterogeneity and specifications which restrict it. We therefore use our framework to circumvent these issues and derive robust and efficient estimators.
In (ref), we first establish a simple characterization for the most efficient linear unbiased estimator of any pre-specified weighted sum of treatment effects, in the baseline case of spherical errors, i.e. homoskedasticity with no serial correlation. This estimator explicitly incorporates the researcher's estimation goal and assumptions about parallel trends, anticipation effects, and restrictions on treatment effect heterogeneity. It is constructed by estimating a flexible high-dimensional regression that differs from conventional event study specifications, and aggregating its coefficients appropriately. While spherical errors are a natural starting point, the principled construction of this estimator more generally ensures unbiasedness and yields attractive efficiency properties, as we later confirm in simulations.
In our leading case where the heterogeneity of treatment effects is not restricted, the efficient robust estimator can be implemented using a transparent “imputation” procedure. First, the unit and period fixed effects $\hat{\alpha}_{i}$ and $\hat{\beta}_{t}$ are fitted by regressions using untreated observations only. Second, these fixed effects are used to impute the untreated potential outcomes and therefore obtain an estimated treatment effect $\hat{\tau}_{it}=Y_{it}-\hat{\alpha}_{i}-\hat{\beta}_{t}$ for each treated observation. Finally, a weighted sum of these treatment effect estimates is taken, with weights corresponding to the estimation target.
To relate our efficient imputation estimator to other unbiased estimators that have been proposed in the literature, we derive two additional results showing the generality of the imputation structure. First, any other linear estimator that is unbiased in our framework with unrestricted causal effects can be represented as an imputation estimator, albeit with an inefficient way of imputing untreated potential outcomes. Second, even when assumptions that restrict treatment effect heterogeneity are imposed, any unbiased estimator can still be understood as an imputation estimator for an adjusted estimand. Together, these two results allow us to characterize estimators of treatment effects in event studies as a combination of how they impute unobserved potential outcomes and which weights they put on treatment effects.
For the efficient estimator in our framework, we provide tools for valid inference. Specifically, we derive conditions under which the estimator is consistent and asymptotically normal and propose standard error estimates. Inference is challenging under arbitrary treatment effect heterogeneity, because causal effects cannot be separated from the error terms. We instead show how asymptotically conservative standard errors can be derived, by attributing some variation in estimated treatment effects to the error terms.\footnote{While the generality of our setting only allows for conservative inference (on any robust estimator, including ours), we obtain asymptotically exact standard errors in the special case that received the most attention in the literature: when units are randomly sampled from a population and the estimand consists of average treatment effects by period\textendash cohort pairs.} Our inference results apply under mild conditions in short panels. Advancing the existing literature on DiD estimation with staggered adoption, we also provide conditions for consistency and inference that extend to panels where the number of time periods grows, as long as growth is not too fast. We also propose a leave-one-out modification to our conservative variance estimates with improved finite-sample performance.
Another important practical advantage of our approach is that it provides a principled way of testing the identifying assumptions of parallel trends and no anticipation effects, based on OLS regressions with untreated observations only. Compared to conventional specifications with leads and lags of treatment that implicitly restrict treatment effects, this approach avoids the contamination of the tests by treatment effect heterogeneity shown by Abraham2018. Moreover, our strategy circumvents the inference problems after pre-testing that were pointed out by Roth2018a, under spherical errors. These attractive properties result from the clear separation of estimation and testing.
It is also useful to point out two limitations of our analysis. First, all event study designs assume a restrictive parametric model for untreated outcomes. We do not evaluate when these assumptions may be applicable, and therefore when the event study design are ex ante appropriate, as Roth2020 do. We similarly do not consider estimation that is robust to violations of parallel-trend type assumptions, as Roth2019 propose, although our framework allows relaxing those assumptions by including unit-specific trends and time-varying covariates. We instead take the standard assumptions of event study designs as given and derive optimal estimators, valid inference, and practical tests to assess whether parallel-trend assumptions hold. Second, we also do not consider event studies as understood in the finance literature, based on high-frequency panel data, which typically do not use period fixed effects MacKinlay1997.
In (ref), we illustrate the practical relevance of our theoretical insights by revisiting the estimation of the marginal propensity to spend out of tax rebates in the event study of Broda2014. First, we show that the choice of a binned specification used by Broda2014 leads to a substantial upward bias in estimated MPXs. Indeed, we find that the binned specification puts a large weight on the effects happening in the first week after the rebate receipt, and negative weights on some longer-run effects, biasing the estimate upwards because the spending response quickly decays over time. Second, we highlight that, due to the implicit extrapolation of treatment effects in specifications restricting treatment effect heterogeneity, some dynamic specifications could be mistakenly interpreted as evidence for a large and persistent increase in spending. Our imputation estimator eliminates unstable patterns found across such specifications. Finally, we illustrate the underidentification problem with the fully-dynamic specification: the dynamic path of estimates is very sensitive to the choice of leads to drop.
Our findings deliver several insights for the macroeconomics literature. While commonly used estimates of the quarterly MPX covering all expenditures range from 50-90% and estimates of the quarterly MPX for nondurable expenditure range from 15-25%,\footnote{Broda2014, Parker2013 and Johnson2006a estimate different versions of the MPX out of tax rebates. Laibson2022, kaplan2020marginal and di2020stock provide recent reviews of the literature on the estimation of the marginal propensity to spend and consume.} our estimates, when appropriately rescaled, are about half as large, at 25\textendash 37% for the MPX one quarter after tax rebate receipt for all expenditures and 8\textendash 11% for nondurables. Using the scaling methodology of Laibson2022, we estimate that the model-consistent, or “notional,” MPC in the quarter following the tax rebate ranges between 7.8% and 11.4%, compared with 15.9% to 23.4% in the original estimation of Broda2014. Furthermore, our preferred estimates are much more short-lived than benchmark estimates, falling to a statistical zero beyond the first month after receiving the tax rebate. Thus, our new estimates imply that fiscal stimulus may be less potent than predicted by leading macroeconomic models targeting benchmark estimates.\footnote{Orchard2022 apply our imputation estimator to the Parker2013 quarterly data, covering the full consumption basket, and also obtain estimates around half as large as in the original study. Our analysis complements their results since, thanks to the high-frequency data, it allows us to investigate the dynamics of the effect and explain the source of the bias of conventional approaches. See also Baker2021 for evidence that using robust event study estimation methods matters in other empirical contexts.}
For convenient application of our results, we supply a Stata command, did_imputation, which implements the imputation estimator and inference for it in a computationally efficient way. Our command handles a variety of practicalities which are also covered by our theoretical results, such as time-varying covariates, triple-difference designs, and repeated cross-sections. We also provide a second command, event_plot, for producing “event study plots” that visualize the estimates with both our estimator and the alternative ones.
Our paper contributes to a growing methodological literature on event studies. To the best of our knowledge, our paper is the first and only one to characterize the underidentification and spurious identification of long-run treatment effects that arise in conventional implementations of event study designs. The negative weighting problem has received more attention. It was first shown by DeChaisemartin2015. The earlier manuscript of our paper Borusyak2017 independently pointed it out and additionally explained how it arises because of forbidden comparisons and why it affects long-run effects in particular, which we now discuss in (ref) below. The issue has since been further investigated by Goodman-Bacon2021, Strezhnev2018, and DeChaisemartin2018, while Abraham2018 have shown similar problems with dynamic specifications. Abraham2018 and Roth2018a have further uncovered problems with conventional pre-trend tests, and Schmidheiny2018 have characterized the problems which arise from binning multiple lags and leads in dynamic specifications. Besides being the first to point out some of these issues, our paper provides a unifying econometric framework which explicitly relates these issues to the conflation of the target estimand and the underlying identification assumptions.
Several papers have proposed ways to address these problems, introducing estimators that remain valid when treatment effects can vary arbitrarily DeChaisemartin2020,Abraham2018,Callaway2018,Marcus2020,Cengiz2019. An important limitation of these robust estimators is that their efficiency properties are not known.\footnote{There are three notable exceptions. Marcus2020 consider a two-stage generalized method of moments (GMM) estimator and establish its semiparametric efficiency under heteroskedasticity in a large-sample framework with a fixed number of periods. However, they find this estimator to be impractical, as it involves many moments, e.g. almost as many as the number of observations in the application they consider. Second, Roth2021 characterize the efficient DiD estimator which leverages random timing of the treatment, rather than a more conventional parallel trends assumption, as we do. Finally, harmon_DiD builds on our framework to characterize the efficiency properties of difference-in-differences estimators when error terms follow a random walk \textemdash the opposite case from our benchmark analysis of efficiency which imposes no serial correlation of errors. In (ref), we generalize our results to intermediate cases, allowing for models of heteroskedasticity and serial correlation.} A key contribution of our paper is to derive a practical, robust, and finite-sample efficient estimator from first principles. We show that this estimator takes a particularly transparent form under unrestricted treatment effect heterogeneity, while our construction also yields efficiency when some restrictions on treatment effects are imposed. By clearly separating the testing of underlying assumptions from the estimation step imposing these assumptions, we simultaneously increase estimation efficiency and avoid problems with inference after pre-testing under spherical errors. Our estimator uses all pre-treatment periods for imputation, as appropriate under the standard DiD assumptions, while alternative estimators use more limited information.\footnote{This efficiency gain relative to DeChaisemartin2020 and Abraham2018 is obtained without stronger assumptions. The Callaway2018 assumptions are also equivalent to ours when there is only one period before any unit is treated and there are no covariates (see Marcus2020).}
In the MPX application, we find large gains of our imputation estimator: the confidence interval is about 50% longer for each week relative to the rebate for DeChaisemartin2020, and 2\textendash 3.5 times longer for Abraham2018 (which without extra controls are equivalent to the two versions of the Callaway2018 estimator). We confirm these gains in a simulation study, finding that the standard deviations of alternative robust estimators are 1.3\textendash 3.6 times higher with spherical errors, and that these gains are generally preserved under heteroskedasticity and serial correlation of errors.
Finally, our paper is related to a nascent literature that develops robust estimators similar to the imputation estimator. To the best of our knowledge, this idea has been first proposed for factor models Gobillon2016a,Xu2017. Athey2018b consider a general class of “matrix-completion” estimators for panel data that first impute untreated potential outcomes by regularized factor- and fixed-effects models and then average over the implied treatment-effect estimates. The imputation idea has been explicitly applied to fixed-effect estimators in event studies by Liu2020a, Gardner2020a, Thakral2020, and thakral2023two. Specifically, the counterfactual estimator of Liu2020a, the two-stage estimator of Gardner2020a, Thakral2020, and thakral2023two, and a version of the matrix-completion estimator from Athey2018b without factors or regularization coincide with the imputation estimator in our model for the specific class of estimands their papers consider. Relative to these papers, we make four contributions: we derive a general imputation estimator from first principles, show its efficiency, provide tools for valid asymptotic inference when unit fixed effects are included, and show its robustness to pre-testing. Subsequently to our work, Wooldridge2021 derives a two-way Mundlak estimator, which is also equivalent to the imputation estimator for a restricted class of estimands in complete panels with controls that are not allowed to change over time (but that may have time-varying effects). The robustness and efficiency properties of our estimator are not limited to those situations.
We consider estimation of causal effects of a binary treatment $D_{it}$ on an outcome $Y_{it}$ in a panel of units $i$ and periods $t$. We focus on “staggered rollout” designs in which being treated is an absorbing state. For each unit there is an event date $E_{i}$ when $D_{it}$ switches from 0 to 1 forever: $D_{it}=\mathbf{1}\left[K_{it}\ge0\right]$, where $K_{it}=t-E_{i}$ is the number of periods since the event date (“horizon”). Some units may never be treated, denoted by $E_{i}=\infty$. Units with the same event date are referred to as a cohort.
We do not make any random sampling assumptions and work with a set of observations $it\in\Omega$ of total size $N$, which may or may not form a complete panel. We similarly view the event date for each unit, and therefore all treatment indicators, as fixed. We define the set of treated observations by $\Omega_{1}=\left\{ it\in\Omega\colon\ D_{it}=1\right\} $ of size $N_{1}$ and the set of untreated (i.e., never-treated and not-yet-treated) observations by $\Omega_{0}=\left\{ it\in\Omega\colon\ D_{it}=0\right\} $ of size $N_{0}$.\footnote{Viewing the set of observations and event times as non-stochastic is not essential. In (ref), we show how this framework can be derived from one in which both are stochastic, by appropriate conditioning. Our conditional framework avoids random sampling assumptions made in other work on DiD designs (e.g. DeChaisemartin2018, Abraham2018, and Callaway2018).}
We denote by $Y_{it}(0)$ the period-$t$ stochastic potential outcome of unit $i$ if it is never treated. Causal effects on the treated observations $it\in\Omega_{1}$ are denoted $\tau_{it}=\expec{Y_{it}-Y_{it}(0)}$. We suppose a researcher is interested in a statistic which sums or averages treatment effects $\tau=\left(\tau_{it}\right)_{it\in\Omega_{1}}$ over the set of treated observations with pre-specified non-stochastic weights $w_{1}=\left(w_{it}\right)_{it\in\Omega_{1}}$ that can depend on treatment assignment and timing, but not on realized outcomes:
For notation brevity, we consider scalar estimands.
Different weights are appropriate for different research questions. The researcher may be interested in the overall ATT, formalized by $w_{it}=1/N_{1}$ for all $it\in\Omega_{1}$. In event study analyses a common estimand is the average effect $h$ periods since treatment for a given horizon $h\ge0$: $w_{it}=\mathbf{1}\left[K_{it}=h\right]/\lvert\Omega_{1,h}\rvert$ for $\Omega_{1,h}=\left\{ it\colon\ K_{it}=h\right\} $. Our approach also allows researchers to specify target estimands that place unequal weights on units within the same cohort-by-horizon cell. For example, one may be interested in weighting units by their size, or in estimating a “balanced” version of horizon-average effects: the ATT at horizon $h$ computed only for the subset of units also observed at horizon $h^{\prime}$, such that the gap between two or more estimates is not confounded by compositional differences. Finally, we do not require the $w_{it}$ to add up to one; for example, a researcher may be interested in the difference between average treatment effects at different horizons or across some groups of units (e.g. women and men), corresponding to $\sum_{it\in\Omega_{1}}w_{it}=0$.\footnote{More broadly, the choice of weights allows for estimation of treatment effect heterogeneity by observed characteristics $R_{it}$. Indeed, the slope of the linear projection of $\tau_{it}$ on some observable $R_{it}$ (which may or may not be time-varying) is a weighted sum of treatment effects, $\sum_{it\in\Omega_{1}}w_{it}\tau_{it}$ for $w_{it}=\left(R_{it}-\bar{R}\right)/\sum_{js\in\Omega_{1}}(R_{js}-\bar{R})^{2}$ and $\bar{R}=\frac{1}{\lvert\Omega_{1}\rvert}\sum_{js\in\Omega_{1}}R_{js}$. The same logic generalizes when $R_{it}$ is a vector, via the Frisch\textendash Waugh\textendash Lowell theorem. This approach also allows for tests of restrictions on treatment effect heterogeneity, e.g. to assess whether ATTs vary across time horizons.}
To identify $\tau_{w}$, we consider three assumptions. We start with the parallel-trends assumption, which imposes a two-way fixed effect (TWFE) model on the untreated potential outcomes.
An equivalent formulation requires $\expec{Y_{it}(0)-Y_{it'}(0)}$ to be the same across units $i$ for all periods $t$ and $t'$ (whenever $it$ and $it'$ are observed).
Parallel trend assumptions are standard in DiD designs, but their details may vary. First, we impose the TWFE model on the entire sample. Although weaker assumptions can be sufficient for identification of $\tau_{w}$ Callaway2021, those alternative restrictions depend on the realized treatment timing. Since parallel trends is an assumption on potential outcomes, we prefer its stronger version which can be made a priori.\footnote{Specifically, Assumption 4 in Callaway2018 requires that the TWFE model only holds for all treated observations ($D_{it}=1$), observations directly preceding the treatment onset ($K_{it}=-1$), and in all periods for never-treated units. Similarly, Goodman-Bacon2021 proposes to impose parallel trends on a “variance-weighted” average of units, as the weakest assumption under which static specifications we discuss in (ref) identify some average of causal effects. While technically weaker, this assumption may be hard to justify ex ante without imposing parallel trends on all units as it is unlikely that non-parallel trends will cancel out by averaging.} Moreover, (ref) can be tested by using pre-treatment data, while minimal assumptions cannot. Second, we impose (ref) at the unit level, while sometimes it is imposed on cohort-level averages. Our approach is in line with the practice of including unit, rather than cohort, FEs in DiD analyses and allows us to avoid biases in incomplete panels where the composition of units changes over time. Moreover, we show in (ref) that, under random sampling and without compositional changes, assumptions on cohort-level averages imply (ref).
Our framework extends immediately to richer models of $Y_{it}(0)$:
\setcounter{assumption}{1} The first term in this model of $Y_{it}(0)$ nests unit FEs, but also allows to interact them with some observed covariates unaffected by the treatment status, e.g. to include unit-specific trends. This term looks similar to a factor model, but differs in that regressors $A_{it}$ are observed. The second term nests period FEs but additionally allows any time-varying covariates, i.e. $X_{it}^{\prime}\delta=\beta_{t}+\tilde{X}_{it}^{\prime}\tilde{\delta}$. In (ref) we clarify that $X_{it}$ have to be unaffected by treatment and strictly exogenous to be included in the specification.
We next rule out anticipation effects, i.e. the causal effects of being treated in the future on current outcomes (e.g. Abbring2003):
(ref) together imply that the observed outcomes $Y_{it}$ for untreated observations follow the TWFE model. It is straightforward to weaken this assumption, e.g. by allowing anticipation for some $k$ periods before treatment: this simply requires redefining event dates to earlier ones. However, some form of this assumption is necessary for DiD identification, as there would be no reference periods for treated units otherwise.
Finally, researchers sometimes impose restrictions on causal effects, explicitly or implicitly. For instance, $\tau_{it}$ may be assumed to be homogeneous for all units and periods, or only depend on the number of periods since treatment (but be otherwise homogeneous across units and calendar periods). We will consider such restrictions as a possible auxiliary assumption:
It will be more convenient for us to work with an equivalent formulation of (ref), based on $N_{1}-M$ free parameters driving treatment effects rather than $M$ restrictions on them:
\setcounter{assumption}{3} (ref) imposes a parametric model of treatment effects. For example, the assumption that treatment effects all be the same, $\tau_{it}\equiv\theta_{1}$, corresponds to $N_{1}-M=1$ and $\Gamma=\left(1,\dots,1\right)^{\prime}$. Conversely, a “null model” $\tau_{it}\equiv\theta_{it}$ that imposes no restrictions is captured by $M=0$ and $\Gamma=\mathbb{I}_{N_{1}}$.
If restrictions on the treatment effects are implied by economic theory, imposing them will increase estimation power. Often, however, such restrictions are implicitly imposed without an ex ante justification, but just because they yield a simple model for the outcome. We will show in (ref) how estimators that rely on this assumption can fail to estimate reasonable averages of treatment effects, let alone the specific estimand $\tau_{w}$, when the assumption is violated.\footnote{We view the null (ref) as a conservative default. We note, however, that this makes the assumptions inherently asymmetric in that they impose restrictive models on potential control outcomes $Y_{it}(0)$ ((ref)), but not on treatment effects $\tau_{it}$. This asymmetry reflects the standard practice in staggered rollout DiD designs and is natural when the structure of treatment effects is ex ante unknown, while our framework also accommodates the case where the researcher is willing to impose structure. Restrictions on treatment effects, when appropriate, are also useful for external validity: unless some structure is imposed on treatment effects, one cannot use estimates from past data to inform future policy, for instance extending a given treatment to currently untreated units. However, one can use our framework without restrictions to learn about the structure of treatment effects, e.g. whether they vary across cohorts for each horizon.}
While we formulated our setting for staggered-adoption DiD designs with binary treatments in panel data, our framework applies without change in many related research designs. In repeated cross-sections, a different random sample of units $i$ (e.g., individuals) from the same groups $g(i)$ (e.g., regions) is observed in each period. Unit FEs are not possible to include but can be replaced with group FEs in (ref): $\expec{Y_{it}(0)}=\alpha_{g(i)}+\beta_{t}$. In triple-differences designs, the data have two dimensions in addition to periods, e.g. $i$ corresponds to a pair of region $j(i)$ and demographic group $g(i)$. (ref) can be specified as $\expec{Y_{it}(0)}=\alpha_{j(i)g(i)}+\alpha_{j(i)t}+\alpha_{g(i)t}$.\footnote{Another variation is when the outcome is measured in a single period but across two cross-sectional dimensions, such as regions $i$ and birth cohorts $g$, with the treatment implemented in a set of regions for the cohorts born after some cutoff period $E_{i}$ (e.g., Hoynes2016). Then one may write $\expec{Y_{ig}(0)}=\alpha_{i}+\beta_{g}$.} With non-binary treatment intensity, our setting applies if each unit is observed untreated before $E_{i}$ and treated with heterogenous intensity $R_{it}\ne0$ (that may or may not vary over time) from period $E_{i}$. (ref) can apply, and the researcher can consider estimands such as the “ATT per unit of intensity”, $\frac{1}{\left|\Omega_{1}\right|}\expec{\sum_{it\in\Omega_{1}}\left(Y_{it}-Y_{it}(0)\right)/R_{it}}$, by setting $w_{it}$ proportionally to $1/R_{it}$. The challenges we describe in (ref) for standard staggered DiDs and the solutions of (ref) directly apply in all of these cases.\footnote{This is also the case of non-staggered DiD designs, in which units receive treatment in a single period or never. Our insights in (ref) and (ref) are still relevant if continuous covariates or unit-specific trends are included (see SantAnna2020 and Wolfers2003 for related ideas).}
In this section, we first introduce the common two-way fixed effects regressions with restricted treatment effect heterogeneity that have traditionally been used in DiD designs. We then discuss several estimation challenges that pertain to these specifications, including underidentification in certain dynamic specifications, negative weighting, and spurious identification of long-run causal effects. We conclude the section by discussing how our framework also relates to other problems that have been pointed out by Roth2018a and Abraham2018.
Causal effects in staggered adoption DiD designs have traditionally been estimated via OLS regressions with two-way fixed effects, using specifications that implicitly restrict treatment effect heterogeneity across units. While details may vary, the following specification covers many studies:
Here $\tilde{\alpha}_{i}$ and $\tilde{\beta}_{t}$ are the unit and period (“two-way”) fixed effects, $a\ge0$ and $b\ge0$ are the numbers of included “leads” and “lags” of the event indicator, respectively, and $\varepsilon_{it}$ is the error term. The first lead, $\mathbf{1}\left[K_{it}=-1\right]$, is often excluded as a normalization, while the coefficients on the other leads (if present) are interpreted as measures of “pre-trends,” and the hypothesis that $\tau_{-a}=\dots=\tau_{-2}=0$ is tested visually or statistically. Conditionally on this test passing, the coefficients on the lags are interpreted as a dynamic path of causal effects: at $h=0,\dots,b-1$ periods after treatment and, in the case of $\tau_{b+}$, at longer horizons binned together. We will refer to this specification as “dynamic” (as long as $a+b>0$) and, more specifically, “fully-dynamic” if it includes all available leads and lags except $h=-1$, or “semi-dynamic” if it includes all lags but no leads.
Viewed through the lens of the (ref) framework, these specifications make implicit assumptions on untreated potential outcomes, anticipation and treatment effects, and the estimand of interest. First, they make (ref) but, for $a>0$, do not fully impose (ref), allowing for anticipation effects for $a$ periods before treatment.\footnote{One can alternatively view this specification as imposing (ref) but making a weaker (ref) which includes some pre-trends into $Y_{it}(0)$. This difference in interpretation is immaterial for our results.} Typically this is done as a means to test (ref) rather than to relax it, but the resulting specification is the same. Second, equation (ref) imposes strong restrictions on causal effect heterogeneity ((ref)), with treatment (and anticipation) effects assumed to only vary by horizon $h$ and not across units and periods otherwise. Most often, this is done without an a priori justification. If the lags are binned into the term with $\tau_{b+}$, the effects are further assumed to be time-invariant once $b$ periods have elapsed since the event. Finally, dynamic specifications do not explicitly define the estimands $\tau_{h}$ as particular averages of heterogeneous causal effects, even though researchers often consider that effects may vary across observations, as evidenced by a literature on the interpretation of OLS estimands going back to at least Angrist1998a and humphreys2009bounds.
Besides dynamic specifications, equation (ref) also nests a very common specification used when a researcher is interested in a single parameter summarizing all causal effects. With $a=b=0$, we have the “static” specification in which a single treatment indicator is included:
In line with our (ref) setting, the static equation imposes the parallel trends and no anticipation (ref). However, it also makes a particularly strong version of (ref) \textemdash that all treatment effects are the same. Moreover, the target estimand is again not written out as an explicit average of potentially heterogeneous causal effects.
In the rest of this section we turn to the challenges associated with OLS estimation of equations (ref) and (ref). We explain how these issues result from the conflation of the target estimand, (ref) and (ref), providing a new and unified perspective on the problems of static and dynamic specifications with restricted treatment effect heterogeneity.
The first problem pertains to fully-dynamic specifications and arises because a strong enough (ref) is not imposed. We show that those specifications are under-identified if there is no never-treated group:
To illustrate this result with a simple example, (ref) plots the outcomes for a simulated dataset with two units (or equal-sized cohorts), one treated at $t=2$ and the other at $t=4$. Both units exhibit linear growth in the outcome, starting from different levels. There are two interpretations of these dynamics. First, treatment could have no impact on the outcome, in which case the level difference corresponds to the unit FEs, while trends are just a common feature of the environment, through period FEs. Alternatively, note that the outcome equals the number of periods since the event for both groups and all time periods: it is zero at the moment of treatment, negative before, and positive after. A possible interpretation is that the outcome is entirely driven by causal effects and anticipation of treatment. Thus, one cannot hope to distinguish between unrestricted dynamic causal effects and a combination of unit effects and time trends.\footnote{Formally, the problem arises because a linear time trend $t$ and a linear term in the cohort $E_{i}$ (subsumed by the unit FEs) can perfectly reproduce a linear term in horizon $K_{it}=t-E_{i}$. Therefore, a complete set of treatment leads and lags, which is equivalent to the horizon FEs, is collinear with the unit and period FEs.}
The problem may be important in practice, as statistical packages may resolve this collinearity by dropping an arbitrary unit or period indicator. Some estimates of $\left\{ \tau_{h}\right\} $ would then be produced, but because of an arbitrary trend in the coefficients they may suggest a violation of parallel trends even when the specification is in fact correct, i.e. (ref) hold and there is no heterogeneity of treatment effects for each horizon ((ref)).
To break the collinearity problem, stronger restrictions on anticipation effects, and thus on $Y_{it}$ for untreated observations, have to be introduced. One could consider imposing minimal restrictions on the specification that would make it identified. In typical cases, only a linear trend in $\left\{ \tau_{h}\right\} $ is not identified in the fully dynamic specification, while nonlinear paths cannot be reproduced with unit and period fixed effects. Therefore, just one additional normalization, e.g. $\tau_{-a}=0$ in addition to $\tau_{-1}=0$, breaks multicollinearity.\footnote{Additional collinearity arises, e.g., when treatment is staggered but happens at periodic intervals.}
However, minimal identified models rely on ad hoc identification assumptions which are a priori unattractive. For instance, just imposing $\tau_{-a}=\tau_{-1}=0$ means that anticipation effects are assumed away $1$ and $a$ periods before treatment, but not in other pre-periods. This assumption therefore depends on the realized event times. Instead, a systematic approach is to impose the assumptions \textemdash some forms of no anticipation effects and parallel trends \textemdash that the researcher has an a priori argument for and which motivated the use of DiD. Such assumptions also give much stronger identification power.\footnote{Our suggestion to impose identification assumptions at the estimation stage does not mean that those assumptions should not also be tested; we discuss testing in detail in (ref).}
We now show how, by imposing (ref) instead of specifying the estimation target, the static TWFE specification does not identify a reasonably-weighted average of heterogeneous treatment effects: the underlying weights may be negative, particularly for the long-run causal effects. The issues we discuss here also arise in dynamic specifications that bin multiple lags together.
First, we note that, if the parallel-trends and no-anticipation assumptions hold, the static specification identifies some weighted average of treatment effects:\footnote{This result was previously stated in Theorem 1 of DeChaisemartin2018 for general designs, and later in Appendix C of Borusyak2017 for staggered adoption designs.}
The underlying weights $w_{it}^{\text{static}}$ can be computed from the data using the Frisch\textendash Waugh\textendash Lovell theorem (see equation (ref) in the proof of (ref)) and only depend on the timing of treatment for each unit and the set of observed units and periods. The static specification's estimand, however, cannot be interpreted as a proper weighted average, as some weights can be negative, which we illustrate with a simple example:
This example illustrates the severe short-run bias of the static specification: the long-run causal effect, corresponding to the early-treated unit $A$ and the late period 3, enters with a negative weight ($-1/2$). Thus, larger long-run effects make the coefficient smaller.
This problem results from what we call “forbidden comparisons” performed by the static specification. Recall that the original idea of DiD estimation is to compare the evolution of outcomes over some time interval for the units which got treated during that interval relative to a reference group of units which didn't, identifying the period FEs. In the (ref) example, such an “admissible” comparison is between units $A$ and $B$ in periods 2 and 1, $\left(Y_{A2}-Y_{A1}\right)-\left(Y_{B2}-Y_{B1}\right)$. However, panels with staggered treatment timing also lend themselves to a second type of comparisons \textemdash which we label “forbidden” \textemdash in which the reference group has been treated throughout the relevant period. For units in this group, the treatment indicator $D_{it}$ does not change over the relevant period, and so the restrictive specification uses them to identify period FEs, too. The comparison between units $B$ and $A$ in periods 3 and 2, $\left(Y_{B3}-Y_{B2}\right)-\left(Y_{A3}-Y_{A2}\right)$, in (ref) is a case in point. While a comparison like this is appropriate and increases efficiency when treatment effects are homogeneous (which the static specification was designed for), forbidden comparisons are problematic under treatment effect heterogeneity. For instance, subtracting $\left(Y_{A3}-Y_{A2}\right)$ not only removes the gap in period FEs, $\beta_{3}-\beta_{2}$, but also deducts the evolution of treatment effects $\tau_{A3}-\tau_{A2}$, placing a negative weight on $\tau_{A3}$. The restrictive specification leverages comparisons of both types and estimates the treatment effect by $\hat{\tau}^{static}=\left(Y_{B2}-Y_{A2}\right)-\frac{1}{2}\left(Y_{B1}-Y_{A1}\right)-\frac{1}{2}\left(Y_{B3}-Y_{A3}\right)$.\footnote{The proof of (ref) shows why long-run effects in particular are subject to the negative weights problem. In general, negative weights arise for the treated observations, for which the residual from an auxiliary regression of $D_{it}$ on the two-way FEs is negative. DeChaisemartin2018 show that, in complete panels, the unit FEs are higher for early-treated units (which are observed treated for a larger shares of periods) and period FEs are higher for later periods (in which a larger shares of units are treated). The early-treated units observed in later periods correspond to the long-run effects.}
Fundamentally, this problem arises because the specification imposes very strong restrictions on treatment effect homogeneity, i.e. (ref), instead of acknowledging the heterogeneity and specifying a particular target estimand (or perhaps a class of estimands that the researcher is indifferent between).
With a large number of never-treated units or a large number of periods before any unit is treated (relative to other units and periods), our setting becomes closer to a classical non-staggered DiD design, and therefore negative weights disappear, as our next result illustrates:
Even when weights are non-negative, they may remain highly unequal and diverge from the estimands that the researcher is interested in. Our preferred strategy is therefore to commit to the estimation target and explicitly allow for treatment effect heterogeneity, except when some form of (ref) is ex ante appropriate.
Another consequence of inappropriately imposing (ref) concerns estimation of long-run causal effects. Conventional dynamic specifications (except those subject to the underidentification problem) yield some estimates for all $\tau_{h}$ coefficients. Yet, for large enough $h$, no averages of treatment effects are identified under (ref) with unrestricted treatment effect heterogeneity. Therefore, estimates from restrictive specifications are fully driven by unwarranted extrapolation of treatment effects across observations and may not be reliable, unless strong ex ante reasons for (ref) exist.
This issue is well illustrated in the example of (ref). To identify the long-run effect $\tau_{A3}$ under (ref), one needs to form an admissible DiD comparison, of the outcome growth over some period between unit $A$ and another unit not yet treated in period 3. However, by period 3 both units have been treated. Mechanically, this problem arises because the period fixed effect $\beta_{3}$ is not identified separately from the treatment effects $\tau_{A3}$ and $\tau_{B3}$ in this example, absent restrictions on treatment effects. Yet, the semi-dynamic specification \[ Y_{it}=\tilde{\alpha}_{i}+\tilde{\beta}_{t}+\tau_{0}\mathbf{1}\left[K_{it}=0\right]+\tau_{1}\mathbf{1}\left[K_{it}=1\right]+\tilde{\varepsilon}_{it} \] will produce an estimate $\hat{\tau}_{1}$ via extrapolation. Specifically, two different parameters, $\tau_{A3}-\tau_{B3}$ and $\tau_{A2}$, are identified by comparing the two units in periods 2 or 3, respectively, with period 1. Therefore, when imposing homogeneity of short-run effects across units, $\tau_{A2}=\tau_{B3}\equiv\tau_{0}$, we estimate the long-run effect $\tau_{A3}\equiv\tau_{1}$ as the sum of $\tau_{1}-\tau_{0}$ and $\tau_{0}$: \[ \hat{\tau}_{1}=\left[\left(Y_{A3}-Y_{A1}\right)-\left(Y_{B3}-Y_{B1}\right)\right]+\left[\left(Y_{A2}-Y_{A1}\right)-\left(Y_{B2}-Y_{B1}\right)\right]. \] However, when $\tau_{A2}\ne\tau_{B3}$, this estimator is biased.
In general, the gap between the earliest and the latest event times observed in the data provides an upper bound on the number of dynamic coefficients that can be identified without extrapolation of treatment effects. This result, which follows by the same logic of non-identification of the later period effects, is formalized by our next proposition:
Robust estimators, including the one we characterize in (ref), can only be computed for identified estimands, never resulting in spurious estimates.
We finally note that the challenges described in (ref) apply even if the sample is “trimmed” to a fixed window around the event time; see (ref).
To overcome the challenges affecting conventional practice, we now derive the robust and efficient estimator and show that it takes a particularly transparent “imputation” form when no restrictions on treatment-effect heterogeneity are imposed. We then perform asymptotic analysis, establishing the conditions for the estimator to be consistent and asymptotically normal, derive conservative standard error estimates for it, and discuss appropriate pre-trend tests.
Throughout, we continue to suppose that the researcher chose the estimation target $\tau_{w}$ and assumed a model of $Y_{it}(0)$ ((ref)) and no anticipation. Some model of treatment effects ((ref)) may also be assumed, although our main focus is on the null model, under which treatment effect heterogeneity is unrestricted. Letting $\varepsilon_{it}=Y_{it}-\expec{Y_{it}}$ for $it\in\Omega$, we thus have under (ref):
We assume throughout that $\tau_{w}=w_{1}'\Gamma\theta$ is identified. (ref) provides conditions for identification. First, we provide a general rank condition on the matrices of unit-specific and other covariates that requires that the covariate space of treated observations is spanned by that of the untreated ones. This assumption allows us to estimate from the untreated observations those nuisance parameters that are necessary to impute control outcomes of the treated observations, thus providing identification. Second, we derive specific conditions for the case where the parameter $\delta$ represents time fixed effects and $A_{it}$ may vary over time but not across units (as with unit FEs and unit-specific linear trends). In this case, we show that there is identification if (i) the $A_{it}$ are not collinear for any relevant unit and (ii) there is at least one untreated unit at the end of the time period of interest.
For our efficiency result, we impose an additional assumption on the error variances:
While this assumption is strong, our efficiency results also apply without change under dependence that is due to unit random effects, i.e. if $\varepsilon_{it}=\eta_{i}+\tilde{\varepsilon}_{it}$ for $\tilde{\varepsilon}_{it}$ that satisfy (ref) and for some $\eta_{i}$. Moreover, these results are straightforward to relax to any known form of heteroskedasticity or mutual dependence.\footnote{For instance, if error terms are uncorrelated and have known variances $\sigma_{it}^{2}$ (up to a scaling factor), efficiency requires the estimation step (Step 1) of (ref) to be performed with weights proportional to $\sigma_{it}^{-2}$. One example for this is when the data are aggregated from $n_{it}$ individuals randomly drawn from group $i$ in period $t$ and spherical individual-level errors, in which case efficiency is obtained with weights proportional to $n_{it}$.} Under (ref) and allowing for restrictions on causal effects, we have:
Under (ref), regression (ref) is correctly specified. Thus, this estimator for $\theta$ is unbiased by construction, and efficiency under spherical error terms is a direct consequence of the Gauss\textendash Markov theorem. Moreover, OLS yields the most efficient estimator for any linear combination of $\theta$, including $\tau_{w}=w_{1}^{\prime}\Gamma\theta$. While assuming spherical errors may be unrealistic in practice, we think of this assumption as a natural conceptual benchmark to decide between the many unbiased estimators of $\tau_{w}$.\footnote{This benchmark appears natural as it parallels the Gauss-Markov theorem which also relies on spherical errors. In Monte Carlo simulations ((ref)), the estimator performs well even under deviations from spherical errors. In (ref) we generalize the results to parametric models of heteroskedasticity and serial correlation, in the spirit of generalized least squares (GLS) and relating to Wooldridge2021.}
In the important special case of unrestricted treatment effect heterogeneity, $\hat{\tau}_{w}^{*}$ has a useful “imputation” representation. The idea is to estimate the model of $Y_{it}(0)$ using the untreated observations $it\in\Omega_{0}$ and leverage it to impute $Y_{it}(0)$ for treated observations $it\in\Omega_{1}$. Then, observation-specific causal effect estimates can be averaged appropriately. Perhaps surprisingly, the estimation and imputation steps are identical regardless of the target estimand. Applying any weights to the imputed causal effects yields the efficient estimator for the corresponding estimand. We have:
The imputation representation offers computational and conceptual benefits. First, it is computationally efficient as it only requires estimating a simple TWFE model, for which fast algorithms are available Guimaraes2010,Correia2017. This is in contrast to the OLS estimator from (ref), as equation (ref) has regressors $\Gamma_{it}D_{it}$ in addition to the fixed effects, which are high-dimensional unless a low-dimensional model of treatment effect heterogeneity is imposed.
Second, the imputation approach is intuitive and transparently links the parallel trends and no-anticipation assumptions to the estimator. Indeed, imbens2015causal write: “At some level, all methods for causal inference can be viewed as imputation methods, although some more explicitly than others” (p. 141). We formalize this statement in the next proposition, which shows that any estimator unbiased for $\tau_{w}$ can be represented in the imputation way, but the way of imputing the $Y_{it}(0)$ may be less explicit and no longer efficient.
This result establishes an imputation representation when treatment effects can vary arbitrarily. (ref) in the appendix establishes that the imputation structure applies even when restrictions $\tau=\Gamma\theta$ are imposed, albeit with an additional step in which the weights $w_{1}$ defining the estimand are adjusted in a way that does not change $\tau_{w}$ under the imposed model.\footnote{As a special case, we can still write the efficient estimator from (ref) as an imputation estimator from (ref) with alternative weights $v_{1}^{*}$ on the imputed treatment effects. (ref) shows that these adjusted weights $v_{1}^{*}$ solve a quadratic variance-minimization problem with a linear constraint that preserves unbiasedness under (ref). We also provide an explicit formula for the resulting weights in (ref).} In this sense, unbiased causal inference is equivalent to imputation in our framework.
Having derived the linear unbiased estimator $\hat{\tau}_{w}^{\ast}$ for $\tau_{w}$ in (ref) that is also efficient under spherical error terms, we now consider its asymptotic properties without imposing that assumption. We study convergence along a sequence of panels indexed by the sample size $N$, where randomness stems from the error terms $\varepsilon_{it}$ only, as in (ref). Our approach applies to asymptotic sequences where both the number of units and the number of time periods may grow, but the assumptions are least restrictive when the number of time periods remains constant or grows slowly, as in short panels.
Instead of assuming that error terms are spherical, we now assume that error terms are clustered by units $i$.
The key role in our results is played by the weights that the (ref) estimator places on each observation. Since the estimator is linear in the observed outcomes $Y_{it}$, we can write it as $\hat{\tau}_{w}^{*}=\sum_{it\in\Omega}v_{it}^{*}Y_{it}$ with non-stochastic weights $v_{it}^{*}$, derived in (ref) in the appendix.
We now formulate high-level conditions on the sequence of weight vectors that ensure consistency, asymptotic normality, and will later allow us to provide valid inference. These results apply to any unbiased linear estimator $\hat{\tau}_{w}=\sum_{it\in\Omega}v_{it}Y_{it}$ of $\tau_{w}$, not just the efficient estimator $\hat{\tau}_{w}^{*}$ from (ref) \textendash that is, if the respective conditions are fulfilled for the weights $v_{it}$, then consistency, asymptotic normality, and valid inference follow as stated. For the specific estimator $\hat{\tau}_{w}^{*}$ introduced above, we then provide sufficient low-level conditions for short panels.
First, we obtain consistency of $\hat{\tau}_{w}$ under a Herfindahl condition on the weights $v$ that takes the clustering structure of error terms into account.
The condition on the clustered Herfindahl index $\|v\|_{\textnormal{H}}^{2}$ states that the sum of squared weights vanishes, where weights are aggregated by units. One can think of the inverse of the sum of squared weights, $n_{H}=\|v\|_{\textnormal{H}}^{-2}$, as a measure of effective sample size, which (ref) requires to grow large along the asymptotic sequence. If it is satisfied, and variances are uniformly bounded, we obtain consistency of $\hat{\tau}_{w}$:\footnote{The Herfindahl condition can be restrictive since it allows for a worst-case correlation of error terms within units. When such correlations are limited, other sufficient conditions may be more appropriate instead, such as $R\left(\sum_{it\in\Omega}v_{it}^{2}\right)\rightarrow0$ with $R=\max_{i}\left(\text{largest eigenvalue of \ensuremath{\Sigma_{i}}}\right)/\bar{\sigma}^{2}$, where $\Sigma_{i}=\left(\cov{\varepsilon_{it},\varepsilon_{is}}\right)_{t,s}$. Here $R$ is a measure of the maximal joint covariation of all observations for one unit. If error terms are uncorrelated, then $R\leq1$, since the maximal eigenvalue of $\Sigma_{i}$ corresponds to the maximal variance of an error term $\varepsilon_{it}$ in this case, which is bounded by $\bar{\sigma}^{2}$. An upper bound for $R$ is the maximal number of periods for which we observe a unit, since the maximal eigenvalue of $\Sigma_{i}$ is bounded by the sum of the variances on its diagonal.}
We note that the large number of unit-specific parameters $\left\{ \lambda_{i}\right\} _{i}$ cannot generally be estimated consistently in panels with a small number of time periods, raising a potential incidental-parameters problem. However, consistency of $\hat{\tau}_{w}$ does not rely on consistency for unit-specific parameters, since our estimator averages over many units.
We next consider the asymptotic distribution of the estimator around the estimand.
This result establishes conditions under which the difference between estimator and estimand is asymptotically normal. Besides regularity, this proposition requires that the estimator variance $\sigma_{w}^{2}$ does not decline faster than $1/n_{H}$. It is violated if the clustered Herfindahl formula is too conservative: for instance, if the number of periods is growing along the asymptotic sequence while the within-unit over-time correlation of error terms remains small. Alternative sufficient conditions for asymptotic normality can be established in such cases, e.g. along the lines of (ref).
So far, we have formulated high-level conditions on the weights $v_{it}$ of any linear unbiased estimator of $\tau_{w}$. (ref) presents low-level sufficient conditions for consistency and asymptotic normality of the imputation estimator $\hat{\tau}_{w}^{*}$ for the benchmark case of a panel with unit and period FEs, a fixed or slowly growing number of periods, and no restrictions on treatment effects. Unlike (ref), these conditions are imposed directly on the weights $w_{1}$ chosen by the researcher, and not on the the implied weights $v_{it}^{\ast}$, such that the researcher can assess more directly whether the asymptotic approximation is likely to be precise. In particular, the estimator achieves consistency and asymptotic normality in the common case where the number of time periods is fixed, the size of all cohorts increases, the weights on treatment effects do not vary within the same period and cohort, and the sum of (absolute) weights is bounded. In addition, the sufficient conditions are also fulfilled when the number of periods grows slowly and when weights differ across observations within the same cohort and period, but not by too much. With covariates other than unit and period FEs, e.g. with unit-specific linear trends, the general weight conditions in (ref) and (ref) can also be used to verify consistency and asymptotic normality. In those cases, the sufficient conditions are typically fulfilled for convex combinations of cohort-average treatment effects whenever the size of cohorts grows sufficiently fast relatively to the number of periods (see (ref)).
We next estimate the variance of $\hat{\tau}_{w}=\sum_{it\in\Omega}v_{it}Y_{it}$, which equals $\sigma_{w}^{2}=\expec{\sum_{i}\left(\sum_{t;it\in\Omega}v_{it}\varepsilon_{it}\right)^{2}}$ with clustered error terms ((ref)). We start with the case where treatment effect heterogeneity is unrestricted (i.e. $\Gamma=\mathbb{I}$). As in DeChaisemartin2018, exact inference becomes infeasible when treatment effects are heterogeneous, but conservative inference is possible. Following (ref), the inference tools we propose apply to a generic linear unbiased estimator but we use them for the efficient estimator $\hat{\tau}_{w}^{\ast}$. Our strategy is to estimate individual error terms by some $\tilde{\varepsilon}_{it}$ and then use a plug-in estimator,
Estimating the error terms presents two challenges, which become apparent when we consider the benchmark choice $\tilde{\varepsilon}_{it}=\hat{\varepsilon}_{it}$ based on the regression residuals $\hat{\varepsilon}_{it}=Y_{it}-A_{it}^{\prime}\hat{\lambda}_{i}^{*}-X_{it}^{\prime}\hat{\delta}^{*}-D_{it}\hat{\tau}_{it}^{*}$ in the regression (ref). The first challenge is the incidental-parameters problem in estimating $\lambda_{i}$. However, by using cluster-robust variance estimates, our inference does not suffer from this problem since the variance estimator $\hat{\sigma}_{w}^{2}$ does not rely on the consistent estimation of $\lambda_{i}$ any more, similar to the insight of Stock2008a.
A second challenge arises from unrestricted treatment-effect heterogeneity. In (ref), treatment effects are estimated by fitting the corresponding outcomes $Y_{it}$ perfectly, with residuals $\hat{\varepsilon}_{it}\equiv0$ for all treated observations. This issue is not specific to our estimation procedure: one generally cannot distinguish between $\tau_{it}$ and $\varepsilon_{it}$ from observations of $Y_{it}=A_{i}^{\prime}\lambda_{i}+X_{it}^{\prime}\delta+\tau_{it}+\varepsilon_{it}$ for treated observations, making it impossible to produce unbiased estimates of $\sigma_{w}^{2}$ (see Lemma 1 in Kline2020 for a similar impossibility result).
While unbiased estimation of $\sigma_{w}^{2}$ is not possible, we show that this variance can be estimated conservatively. Our variance estimator is based on an auxiliary parsimonious model of treatment effects. We do not require this model to be correct, in the sense that inference is weakly asymptotically conservative under misspecification. However, auxiliary models which better approximate $\tau_{it}$ will make confidence intervals tighter and closer to asymptotically exact. In the computation of $\hat{\sigma}_{w}^{2}$ we set $\tilde{\varepsilon}_{it}$ for the treated observations equal to the residuals of the auxiliary model. We require the model to be parsimonious, such that it does not overfit and the residuals include $\varepsilon_{it}$. When the model is incorrect, $\tilde{\varepsilon}_{it}$ also include a component due to the misspecification of $\tau_{it}$, leading to conservative inference.
We formalize the auxiliary model by considering estimators $\tilde{\tau}_{it}$ for each $it\in\Omega_{1}$ which satisfy two properties: (1) $\tilde{\tau}_{it}$ converges to some non-stochastic limit $\bar{\tau}_{it}$ and (2) if the auxiliary model is correct, $\bar{\tau}_{it}=\tau_{it}$. The following theorem presents conditions under which our construction yields asymptotically conservative inference:
The theorem shows that the proposed variance estimate addresses the two challenges laid out above. First, the estimates remain valid even though we may not be able to estimate the unit-specific parameters $\lambda_{i}$ consistently. This is because unit-specific parameter estimates drop out when summing over all observations of one unit in (ref), as shown in the proof. Second, by using estimates $\tilde{\tau}_{it}$ that fulfill the convergence condition of the theorem, we avoid the issue of obtaining trivial residuals for the treated observations. The resulting variance estimates are asymptotically conservative. From these estimates we can also obtain conservative confidence intervals (that asymptotically have coverage that is at least nominal) if the estimator is also asymptotically normal, such as under the sufficient conditions of (ref).
It remains to choose the estimates $\tilde{\tau}_{it}$. We focus on auxiliary models that impose the equality of treatment effects across large groups of treated observations: for a partition $\Omega_{1}=\bigcup_{g}G_{g}$, $\tau_{it}\equiv\tau_{g}$ for all $it\in G_{g}$. The $\tau_{g}$ can then be estimated by some weighted average of $\hat{\tau}_{it}^{*}$ among $it\in G_{g}$. Specifically, we propose averages of the form
In (ref), we show that this choice of weights leads to minimal excess variance $\sigma_{\tau}^{2}$ in the case where there is only a single group $g$, corresponding to a conservative auxiliary model which requires all treatment effects to be the same. The choice of the partition aims to maintain a balance between avoiding overly conservative variance estimates and ensuring consistency. If the sample is large enough, one may want to partition $\Omega_{1}$ into multiple groups of observations such that treatment effect heterogeneity is expected to be smaller within them than across. For instance, with many units, a group may consist of observations corresponding to the same horizon relative to treatment onset. If cohorts are large, one can further partition observations into groups defined by cohort and period, which we use as the default in our Stata command.
While sufficiently large groups in (ref) avoid overfitting asymptotically (under appropriate conditions), in finite samples these $\tilde{\tau}_{it}$ still use $\hat{\tau}_{it}^{*}$ and thus partially overfit to $\varepsilon_{it}$. In (ref) we therefore also consider leave-out versions of these $\tilde{\tau}_{it}$.
We make four final remarks on (ref). First, our strategy for estimating the variance extends directly to conservative estimation of variance-covariance matrices for vector-valued estimands, e.g. for average treatment effects at multiple horizons $h$. Second, the result applies in short panels under the low-level conditions of (ref) (see (ref)). Third, while we have focused here on the case of unrestricted heterogeneity ($\Gamma=\mathbb{I}$), (ref) can be extended to the case with a non-trivial treatment-effect model imposed in (ref).\footnote{By (ref), the general efficient estimator can be represented as an imputation estimator for a modified estimand, i.e. by changing $w_{1}$ to some $v_{1}$. (ref) then yields a conservative variance estimate for it. We note that under sufficiently strong restrictions on treatment effects, asymptotically exact inference may be possible, as the residuals $\hat{\varepsilon}_{it}$ in (ref) may be estimated consistently even for treated observations (except for the inconsequential noise in $\hat{\lambda}_{i}$), alleviating the need for an additional auxiliary model.} Finally, computation of $\hat{\sigma}_{w}^{2}$ for the estimator $\hat{\tau}_{w}^{*}$ from (ref) involves the implied weights $v_{it}^{*}$, which becomes computationally challenging with multiple sets of high-dimensional FEs. In (ref) we develop a computationally efficient algorithm for computing $v_{it}^{*}$ based on the iterative least squares algorithm for conventional regression coefficients Guimaraes2010.
In this section, we discuss testing the (generalized) parallel-trend and no-anticipation assumptions (ref). We propose a testing procedure based on OLS regressions with untreated observations only, departing from both traditional regression-based tests and more recent placebo tests. This procedure is robust to treatment effect heterogeneity and, under spherical errors, has attractive power properties and avoids the problem of inference after pre-testing explained by Roth2018a. We propose:
This test is valid because equation (ref) is implied by (ref) if the null $\gamma=0$ holds.\footnote{There is a natural alternative test of the null $\gamma=0$ in the model (ref), namely the Hausman test based on the difference between the imputation estimator $\hat{\tau}_{w}^{W}$ based on the model in (ref) and the efficient imputation estimator $\hat{\tau}_{w}^{*}$ that is only valid when $\gamma=0$. Like our test, this test only uses untreated observations and avoids the Roth2018a pre-testing problem for spherical errors. The Hausman approach has the advantage of quantifying the magnitude of bias from omitting $W_{it}$, while (ref) has the advantage that it is also informative about violations that cancel out in the Hausman test. The two tests are equivalent for scalar $W_{it}$.}
The test requires choosing $W_{it}$ to parametrize the possible violation of (ref). A natural choice for $W_{it}$, which parallels conventional pre-trend tests, is a set of indicators for observations $1,\dots,k$ periods before the onset of treatment for some $k$, with periods before $E_{i}-k$ serving as the reference group.\footnote{The optimal choice of $k$ is a challenging question. As usual with Wald tests, choosing a $k$ that is too large can lead to low power against many alternatives, in particular those that generate large biases in treatment effect estimates that impose invalid (ref).} This choice is appropriate, for instance, if the researcher's main worry is the possible effects of treatment anticipation, i.e. violations of (ref). This choice of $W_{it}$ also lends itself to making “event study plots,” which combine the ATT estimates by horizon $h\ge0$ with a series of pre-trend coefficients; we supply the event_plot Stata command for this goal. Alternatively, the researcher may focus on possible violations of (ref). For instance, with data spanning many years one could test for the presence of a structural break in unit FEs.
(ref) can be contrasted with two existing strategies to test parallel trends. Traditionally, researchers estimated a dynamic specification including lags and leads of treatment onset, and tested \textemdash visually or statistically \textemdash that the coefficients on leads are equal to zero. More recent papers DeChaisemartin2018,Liu2020a replace it with a placebo strategy: pretend that treatment happened $k$ periods earlier for all eventually treated units, and estimate the average effects $h=0,\dots,k-1$ periods after the placebo treatment using the same estimator as for actual estimation.
Both of these alternatives strategies have drawbacks. Because the traditional regression-based test uses the full sample, including treated observations, and imposes restrictions on treatment effects (which are assumed homogeneous within each horizon), it is not a test for (ref) only. Rather, it is a joint test that is sensitive to violations of the implicit (ref) Abraham2018. Even if a researcher has reasons to impose a non-trivial (ref) in estimation, a robust test for parallel trends and no anticipation per se should avoid those restrictions on treatment effect heterogeneity. With a null (ref), treated observations are not useful for testing, and our test only uses the untreated ones.\footnote{Wooldridge2021 shows that as long as treatment effects are allowed to vary flexibly, tests based on specifications estimated on the full sample do not use treated observations. Therefore, such tests are also not contaminated by treatment effect heterogeneity.}
Tests based on placebo estimates appropriately use untreated observations only and may have intuitive appeal. However, mimicking the estimator does not generally correspond to an efficient test of a class of plausible alternatives. In contrast, (ref) possesses well-known asymptotic efficiency properties when $W_{it}$ is correctly specified. For example, when $\varepsilon_{it}$ are spherical and normal, it is asymptotically equivalent to the homoskedastic $F$-test, which is a uniformly most powerful invariant test lehmann2006testing.
Finally, we show an additional advantage of (ref): if the researcher conditions on the test passing (i.e. does not report the results otherwise), inference on $\hat{\tau}_{w}^{\ast}$ is still asymptotically valid under the null of no violations of (ref) and under spherical errors. This avoids the issue pointed out by Roth2018a in the context of restrictive dynamic event study regressions: that variance estimates which do not take pre-testing into account are inflated, leading to unnecessarily conservative inference.\footnote{Roth2018a points out another issue, which we also avoid under the assumptions of (ref): when pre-trend and treatment effect estimators are correlated, the bias arising from violations of (ref) is affected by pre-testing; it is exacerbated in specific cases Roth2018a.}
Having derived the attractive theoretical properties of the imputation estimator, we now illustrate their practical relevance by revisiting the estimation of the marginal propensity to spend in the event study of Broda2014. We also use this empirical setting to verify the properties of the imputation estimator in a simulation study.
The marginal propensity to spend out of tax rebates is a crucial parameter for economic policy. In the US, the Economic Stimulus Act of 2008 consisted primarily of a 100 billion dollar program that sent tax rebates to approximately 130 million tax filers. Parker2013 and Broda2014 estimate the marginal propensity to spend (MPX) out of the 2008 tax rebates. The rebate was disbursed using two methods: either via direct deposit to a bank account, if known by the IRS, or with a mailed paper check. For each method, the week in which the funds were disbursed depended on the second-to-last digit of the taxpayer's Social Security number (SSN). This number provides a source of quasi-experimental variation because the last four digits of a SSN are assigned sequentially to applicants within geographic areas.
Broda2014 use an event study design to examine the response of nondurable spending to tax rebate receipt, leveraging the quasi-experimental variation in the timing of the receipt. The quasi-random assignment of the last digits of the SSN makes the parallel-trends assumption for expenditures a priori plausible.\footnote{Thakral2020 point out that for the paper check group pre-rebate household characteristics (in levels) are not balanced with respect to the timing of the receipt. While this is problematic for randomization-based approaches to DiD (e.g. Arkhangelsky2019 and Roth2021), parallel trends in expenditures may still hold. Indeed, we fail to reject them with pre-trend tests below.} The no-anticipation assumption may also be expected to hold: although the disbursement schedule was known in advance, households were directly notified by mail only several days before disbursement.
We estimate the performance of various estimators at estimating the impulse response function of nondurable spending to tax rebate receipt using the same data as BP. While earlier work by Parker2013 estimates the impulse responses using quarterly spending data from the Consumer Expenditure Survey, BP leverage more detailed data from the Nielsen Homescan Consumer Panel. The Nielsen dataset tracks transactions at a much higher (in principle, daily) frequency, which is why we choose it for our analysis. The Nielsen data cover expenditures on consumer packaged goods (food, beverages, beauty and health products, household supplies, and general merchandise), representing around 15% of total household expenditures. Our dataset, identical to that of BP, is a complete panel of 21,760 households (including 21,690 with non-missing disbursement method information) observed over 52 weeks of year 2008.
We show how BP's estimates of the MPX suffer from an upward bias in the short-run due to the choice of a binned specification ((ref)) and how they may be spurious in the long-run ((ref)). In (ref) we present our preferred robust estimates and discuss implications for the macroeconomics literature.
We replicate BP's estimates, focusing on the first three months since the receipt, while leaving longer-run effects to (ref). BP estimate conventional dynamic specifications of the form:
where $Y_{it}$ is the dollar amount of spending in calendar week $t$ for household $i$, $\alpha_{i}$ are household FEs, and $\beta_{t}$ are week FEs. In some specifications, week FEs are interacted with the disbursement method $m(i)$ (i.e., $\beta_{m(i)t}$ is included instead of $\beta_{t}$ in (ref)) to leverage the variation in timing only within each disbursement method; we refer to those specifications as “with disbursement method FEs.” The set of $\mathbf{1}\left[K_{it}=h\right]$ are the lead/lag indicator variables tracking the number of weeks $K_{it}=t-E_{i}$ since the week of the tax rebate receipt for the household, $E_{i}$; $b$ is chosen such that all possible lags in the sample are covered; $a$ varies as discussed below. MPXs for each horizon, as well as pre-trend coefficients, are captured by $\tau_{h}$. Regressions are weighted by the Nielsen projection weights.
BP's preferred specification is a binned version of (ref) which constrains $\tau_{h}$ to be constant across four-week periods \textemdash “months” \textemdash around the event, starting with the week of tax rebate receipt: e.g., $\tau_{0}=\dots=\tau_{3}$. This specification also includes one monthly pre-trend coefficient, i.e. $a=4$ with $\tau_{-1}=\dots=\tau_{-4}.$ These estimates, without and with disbursement method FEs, are replicated in (ref), columns 1 and 2, suggesting that tax rebate receipt led to an increase in spending in the contemporaneous month of \$42.6 (s.e. 7.2) in col. 1 to \$47.6 (s.e. 9.2) in col. 2, and a cumulative increase over three months of \$60.5 (s.e. 25.7) in col. 1 to \$94.4 (s.e. 33.5) in col. 2. As we will discuss in (ref), extrapolating these estimates from Nielsen products to all consumption implies very large total MPX.
Next, we show that the MPX estimates are much smaller without binning. In columns 3 and 4 of (ref) we report the estimates from the conventional specification (ref) without binning and with one weekly lead ($a=1$), as in BP's Table 3. We report the coefficients aggregated to the monthly level. Compared to columns 1 and 2, there is a large fall in the cumulative three-month MPX, from \$60.5 (s.e. 25.7) to \$26.8 (s.e. 21.4) without disbursement method FEs and from \$94.4 (s.e. 33.5) to \$9.6 (s.e. 34.4) with these FEs. In columns 5 and 6, we use the robust and efficient imputation estimator to estimate weekly average responses and aggregate them to the monthly level. The point estimates are similar to columns 3 and 4 for the contemporaneous month response, while for the quarterly MPX they are in between the results obtained with a binned specification and without binning.\footnote{(ref) reports the differences between the estimates from OLS with no binning or imputation and the binned OLS specification, with standard errors and p-values. All binned estimates are significantly different from those without binning at the 10% (5%) significance level without (with) disbursement method FEs. The difference between the binned and imputation estimates is only significant at the 10% level for the contemporaneous month with disbursement method FEs. We explain below that the difference in estimates is due to the difference in estimands, rather than statistical noise.}
Could the difference between binned and other estimates indicate a violation of the DiD assumptions? (ref) provides evidence against this possibility, showing that there is no sign of pre-trends.\footnote{This figure reports the imputation estimates for 8 weeks since the rebate along with the pre-trend coefficients from the (ref) test that allows for 8 weeks of anticipation effects. The figure also reports conventional specifications without binning augmented to include $a=8$ weeks of pre-trends, dropping observations more than 8 weeks since the rebate.} The Wald test confirms this finding: the p-value for the null of no pre-trends is 0.185 (0.403) without (with) disbursement method FEs.
We find instead that the higher estimates from binned specifications are explained by the estimand they implicitly choose. Specifically, this estimand places a very large weight on the first weeks after the rebate, when the effects are the largest, and negative weights on other weeks. (ref) shows that the increase in spending after the receipt is concentrated in the first weeks since the rebate. (ref) in turn shows the weights with which the quarterly MPX estimated from the monthly binned specification of (ref) aggregates the MPXs at each weekly horizon. These weights show how the estimand of the binned specification diverges from the true quarterly MPX, which is a simple sum of the effects at each horizon $h=0,\dots,11$ weeks, i.e. with constant weights of one on each week.\footnote{The binned specification's estimand also diverges from the true MPX in how it weights different households for the same weekly horizon (similar to the issues studied theoretically by Abraham2018 for dynamic specifications without binning). We focus on the variation across horizons here because MPXs have a very strong dynamic pattern.} In the specification without disbursement method FEs, the weight placed on the first-week response is three times larger than it would be with an equally weighted sum; it is five times larger with the FEs. Furthermore, within each month the weights become negative for the last weeks of the month. Applying the weights of the binned specification across weeks from (ref) to the estimates without binning (underlying col. 3 of (ref)), we obtain a point estimate of \$42.6 for the contemporaneous month \textemdash indistinguishable from col. 1, instead of \$35.0 in col. 3. Similarly, we get \$60.4 for the quarter, nearly identical to \$60.5 in col. 1, instead of \$26.8 in col. 3. Thus, the short-run biased weighting scheme due to binning explains nearly all the difference between columns 1 and 3 of (ref).\footnote{Short-run biased weighting also explains the majority, although not all, of the difference between the specification with disbursement method FEs in columns 2 and 4 of (ref). Applying the binned specification weights to the specifications without binning, we get an estimate of \$40.0 for the contemporaneous month and \$69.0 for the quarter, thus reducing the discrepancy between columns 2 and 4 by 62% and 70% for the month and quarter, respectively.}
We now examine the long-run dynamics of MPXs obtained with conventional specifications and the imputation estimator. The timing of the tax rebate is such that we simultaneously observe treated and untreated households for at most 13 weeks.\footnote{The first treated households received the rebate during week 17 of 2008 (week ending April 26), while the last treated households received it during week 30 (week ending July 26).} Per (ref), without restrictive assumptions on treatment effect heterogeneity it is not possible to estimate causal effects beyond 12 weeks. Yet conventional dynamic specifications produce estimates for longer horizons via extrapolation. We examine whether the estimates obtained in this way could paint a misleading picture of the long-run dynamics of MPXs.
In (ref), we use the same specifications as in (ref) but we report the full set of dynamic estimates for the treatment effects. Panel A reports the estimates from the binned specification. With disbursement method FEs, the point estimates are large and positive for all nine months following the receipt of the tax rebate. Thus, due to the extrapolation resulting from binning, this specification could be mistakenly interpreted as evidence for a very large and persistent increase in spending. Without these FEs, the estimates tend to hover around zero.
In Panel B, we show estimates with the conventional dynamic specification without binning. Both specifications with and without disbursement method FEs yield point estimates that are almost all negative in the long run. Taken at face value, these estimates could misleadingly suggest that households intertemporally substitute consumption by making purchases at the time of tax rebate that they would have made 20 to 30 weeks later. As in Panel A, these point estimates are noisy but could lend themselves to some economic interpretation.
In contrast, Panel C describes the results from the robust imputation estimator, which does not allow extrapolation in the absence of an explicit control group. This panel shows that, for the horizons for which imputation is possible, there is no evidence of any impact on spending beyond two to four weeks after tax rebate receipt. The patterns are the same both with and without disbursement FEs. These results highlight the practical relevance of the insights from (ref): the imputation estimators avoid extrapolation, thus eliminating seemingly unstable patterns found across conventional specifications.
Finally, in (ref) we illustrate the importance of the insights on the underidentification of fully-dynamic specifications from (ref). Unlike earlier specifications, which only included a small number of treatment leads, here we run the specification (ref) with a full set of weekly leads and lags around tax rebate receipt. We drop two leads since the set of lead and lag coefficients is only identified up to a linear trend, as discussed in (ref). We find that the fully-dynamic estimates change drastically depending on which two leads are dropped. We illustrate this by comparing the MPXs when dropping leads $-1$ and $-2$ or $-3$ and $-4$. This shows another source of instability in conventional practice, which the imputation estimator directly avoids.
We now discuss the implications of our findings for the macroeconomics literature. We proceed in two steps: selecting our preferred MPX estimate from (ref) for the Nielsen products and then extrapolating it to broader consumption baskets, following the strategy of BP.
Our preferred estimate for the average cumulative MPX out of the tax rebate for the Nielsen products is \$30.5, corresponding to the imputation estimator with disbursement method FEs ((ref), column 6) in the first month since the rebate. This constitutes 3.4% of the average rebate amount. We choose the specification with disbursement method FEs because the variation in timing is more plausibly exogenous within disbursement methods. We focus on the first (i.e., contemporaneous) month and impose zero effects for the following months based on the evidence from (ref) that the MPXs rapidly decay to zero, while estimation noise increases.\footnote{Our preferred estimate is robust to the choice of the time window: the cumulative MPX would have been similar (at \$25.7 instead of \$30.5) if we focused on the first two weeks only.} Finally, we choose the imputation estimator over conventional specifications for its robustness properties. In contrast to columns 1\textendash 2 of (ref), it avoids the short-term bias due to binning. Moreover, in contrast to column 3\textendash 4 it avoids extrapolation of long-run effects (the estimates are similar for the contemporaneous month). Robustness to treatment effect heterogeneity is gained without an efficiency loss in this application: the standard errors are similar across columns of (ref).
To obtain MPX estimates covering the full consumption basket, BP propose to rescale the estimates obtained with the Nielsen data. This scaling is done in three different ways: (i) by the ratio of spending per capita in the National Income and Product Account (NIPA) and Nielsen data; (ii) by the ratio of the self-reported change in spending on all goods after the rebate relative to that on Nielsen goods alone; (iii) by a factor based on the relative shares of spending and relative responsiveness to the rebate across subcategories of goods as measured in Consumer Expenditure Survey (CE). Using these three approaches and BP's preferred MPX estimate (reproduced in our (ref), col. 1), they estimate that the tax rebate raised the annualized expenditure growth rate by 1.3\textendash 1.9 percentage points (p.p.) in 2008Q2 and by 0.6\textendash 0.9 p.p. in 2008Q3, depending on the choice of rescaling.
Applying the same scaling methods to our preferred MPX estimate for the contemporaneous month and assuming zero response in the following months paints a very different picture, with an increase in annualized expenditure growth of only 0.8\textendash 1.1p.p. in 2008Q2 and 0.15\textendash 0.22p.p. in 2008Q3. Our estimate implies a 40% smaller response of consumption expenditures in 2008Q2, and 75% smaller in 2008Q3. Correspondingly, while BP conclude that the propensity to spend at the individual level from a tax rebate over three months since the rebate is between 51 and 75 percent, our preferred estimates are half as large, between 25 and 37 percent.\footnote{We obtain these estimates by replicating the first row of Panel A of BP's Table 5 and using our preferred estimates.}
In (ref), we summarize the MPX estimates for the first quarter after tax rebate obtained with BP's and our preferred specification. The first row reports the observed marginal propensity to spend on products included in the Nielsen sample during that quarter, as a fraction of the average rebate amount. The next rows rescale these estimates to extrapolate the marginal propensity to spend to broader samples, i.e. the full consumption basket (second row) and nondurables (third row). For the full consumption basket, we implement the three scaling procedures from BP and report the lower and upper bounds; for nondurables we leverage the scaling method of Laibson2022. The fourth row reports the model-consistent, or “notional,” marginal propensity to consume (MPC) that can be used as a target for macroeconomic models, also following the methodology of Laibson2022.\footnote{Standard macroeconomic models assume a notional consumption flow that does not distinguish between nondurable and durable consumption. Prior to Laibson2022 showing that the notional MPC should be the relevant target, state-of-the-art macroeconomic models targeted nondurable MPX estimates. For instance, Kaplan2014 targeted the estimates from Johnson2006a, which are quantitatively similar to those from BP when rescaled as in (ref), despite using more aggregated data and a different rebate episode.} The estimates based on BP in column 1 are closely in line with the literature: typical estimates of the quarterly MPX for all expenditures range from 50-90%, while estimates of the quarterly MPX for nondurable expenditure range from 15-25%.\footnote{Laibson2022 provide a recent review of the literature. kaplan2020marginal review nondurable MPX, and di2020stock review total MPX.} In contrast, the imputation estimator in column 2 of (ref) delivers estimates that are about half as large in all rows. These smaller MPC estimates imply a lower effectiveness of fiscal stimulus.
{}
Thus, our new estimates for the impact of the 2008 fiscal stimulus on the U.S. economy yield two lessons for the calibration of macroeconomic models: (1) that the targeted MPC should be significantly smaller \textemdash about half as large \textemdash and (2) that it is best to calibrate the model using weekly-level estimates of the MPC, as we report in (ref), rather than monthly or, especially, quarterly MPC estimates, which are much noisier. Indeed, models should reflect that most of the spending response occurs in the very short run, in the first two to four weeks after tax rebate receipt.
Finally, we compare the efficiency of the imputation estimator to the alternative robust estimators of DeChaisemartin2020 and Abraham2018, abbreviated dCDH and SA. We document the in-sample efficiency gains in (ref) by showing the point estimates and confidence intervals for weekly average MPXs based on the imputation estimator and the two alternatives in Panel A. We use the specification without disbursement method FEs.\footnote{We implement the dCDH method using the csdid Stata command developed for the Callaway2018 estimator: the two estimators are identical absent additional controls, and csdid allows for projection weights.} The point estimates are very similar for dCDH and the imputation estimator, but they differ from those of SA, because this estimator uses a much smaller control group (only the households who received the rebate in the latest possible week) and is therefore much noisier.
Panel B zooms in on the efficiency comparison by reporting the lengths of the confidence intervals for SA and dCDH relative to that of the imputation estimator. The differences are large: the confidence interval from dCDH is about 50% longer for all periods, and 2\textendash 3.5 times longer for SA.
In (ref) we confirm these efficiency gains, obtained from a single sample, in a Monte Carlo study based on the BP data, for several data-generating processes. We find that the imputation estimator has sizable efficiency advantages over alternative robust estimators not only with spherical errors but also in presence of heteroskedasticity, serial correlation, or both. Moreover, these gains do not come at a cost of systematically higher sensitivity to parallel trend violations. We also confirm that our analytical standard errors have correct coverage.
In this paper, we provided a unified framework that formalizes an explicit set of goals and assumptions underlying event study designs, reveals and explains challenges with conventional practice, and yields an efficient estimator. In a benchmark case where treatment-effect heterogeneity remains unrestricted, this robust and efficient estimator takes a particularly simple “imputation” form that estimates fixed effects among the untreated observations only, imputes untreated outcomes for treated observations, and then forms treatment-effect estimates as weighted averages over the differences between actual and imputed outcomes. We developed results for asymptotic inference and testing and compared our approach to other estimators. We also highlighted the importance of separating testing of identification assumptions from estimation, which increases estimation efficiency and helps address inference biases due to pre-testing. We demonstrated the practical relevance of these insights in an empirical application documenting that the notional marginal propensity to consume is between 8 and 11 percent in the first quarter, about half as large as benchmark estimates.
\paragraph{Data Availability Statement}
The data underlying this article are publicly available on Zenodo, at \href{https://doi.org/10.5281/zenodo.10037585}{doi.org/10.5281/zenodo.10037585}.
\printbibliography