Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
90,912 characters · 36 sections · 60 citation commands
Treatment-Effect Estimation in Complex Designs under a Parallel-trends Assumption
Most work on difference-in-differences has focused on simple designs, namely classical designs where some units receive a binary treatment at the same time whereas other do not, or staggered adoption designs, where units begin to receive an absorbing binary treatment at different points in time, or heterogeneous adoption designs where all units are initially untreated before the treated units start receiving heterogeneous doses. However, more complex designs are actually prevalent. In previous work de2020difference, we conducted a survey of the 100 papers with the most Google Scholar citations published by the American Economic Review from 2015 to 2019. Of those, 26 use a two-way fixed effects (TWFE) regression to estimate the effect of a treatment on an outcome.\footnote{Counting TWFE regressions is in line with the long-held belief that this regression is the uncontroversial treatment-effect estimator in a DID research design.} Of these 26 papers, only two have a classical design, four have an absorbing and binary treatment with variation in treatment timing, and two have an heterogeneous adoption design. The remaining 18 have a more complex design, with a non-binary and/or non-absorbing treatment, and where units may receive heterogeneous treatment doses even at period one.
The aim of this paper is to study what can be learnt in such designs. We first focus on identification under no-anticipation and parallel-trends assumptions, summarizing and extending results from de2020difference. A first challenge is that the usual parallel-trends condition, where potential-outcome trends without treatment are mean independent of the treatment path, may not have much identification power if many units actually receive a non-zero dose of treatment at period one. We consider another parallel-trends condition, which accommodates this issue. It requires that conditional on their period-one treatment, outcome trends in the status-quo counterfactual where units keep their period-one treatment are mean independent of units' treatment path. Pre-trends tests can be used to assess the plausibility of this parallel-trends assumption. Conditioning on units' period-one treatment is important: without that conditioning, our parallel-trends assumption would actually rule out effects of the lagged treatments on the outcome, once combined with the usual parallel-trends assumption.
Under our parallel-trends condition, to identify the effect of having departed from one's status-quo treatment $\ell$ periods ago, we can form a difference-in-difference estimand comparing units that have first switched treatment $\ell$ periods ago with units that have not switched yet and that had the same treatment at period one. The corresponding event-study effects may be more challenging to interpret than in staggered adoption designs. Nonetheless, we show that they can be useful to do an ex-post cost-benefit analysis of the policies that effectively took place over the study period. Moreover, once properly normalized they estimate weighted averages of marginal effects of the current and lagged treatments on the outcome. Finally, they can be used to test relevant null hypotheses, such as whether lagged treatments affect the outcome or not. The corresponding estimators are computed by the did_multiplegt_dyn Stata, R, and Python commands.
Still, these event-study effects cannot be used to separate the marginal effects of the current treatment and of some specific treatment lags. Moreover, they cannot be used to evaluate alternative policies, or determine the optimal policy. To make progress on these important questions, we consider in the second part of the paper a further restriction, in the form of a random coefficients distributed-lag linear model. This model is restrictive, but allows for heterogeneity in treatment effects that can be correlated with the treatment path. We first decompose the usual distributed-lag TWFE regressions under this model, and show that they usually fail to identify convex combinations of unit-specific effects. Then, we show that expectations of the random coefficients can be identified and estimated simply and without tuning parameters if the treatment variable takes a finite number of values. The corresponding estimators are computed by the dist_lag_het R package, available at \url{https://github.com/chaisemartinPackages/dist_lag_het}.
Finally, we illustrate our theoretical results by studying the effect of newspapers on electoral turnout in the US between 1868 and 1928, revisiting gentzkow2011. We extend their analysis by considering a dynamic rather than static setup, thus allowing the number of past newspapers to potentially affect current turnout. Our event-study effects, relying only on no-anticipation and parallel trends assumptions, show a large positive effect of newspapers on turnout. Though these parameters cannot bring a definitive answer to this question, they also suggest that the current number of newspapers has a larger effect than the lagged number of newspapers, and our estimates of the random coefficients distributed lag model confirm this finding. On the other hand, the distributed-lag TWFE regression leads to the opposite conclusion: the coefficient on lagged newspapers is positive and significant, while the coefficient on current newspapers is small and insignificant. Our decomposition of those coefficients shows that they may be unreliable if effects vary across counties.
This paper contributes to a rapidly growing literature on difference-in-differences (DID). In that literature, most papers have focused on designs with a binary and absorbing treatment dcDH2020,callaway2018,abraham2018,borusyak2020revisiting or on heterogeneous adoption designs where all units are untreated at period one before treated units start receiving heterogeneous doses fricke2017identification,dcDH2020,callaway2021difference,de2022heterogeneousadoption. The first part of the paper summarizes and extends results from de2020difference, that can accommodate complex designs with a non-binary and/or non-absorbing treatment, and where units may receive heterogeneous treatment doses even at period one. We emphasize that under a parallel-trends assumption alone, we can identify weighted averages of the effect of the current and lagged treatments on the outcome, without being able to separately estimate those effects. This motivates our analysis of the random coefficients distributed-lag model in the second part of the paper. There, we show that the work of chamberlain1992efficiency, arellano2012identifying, and graham2012identification can fruitfully be applied to estimate distributed-lag models with heterogeneous effects. Compared to these papers, we highlight the fact that averages of the random coefficients can be estimated simply and without tuning parameters if the treatment variable takes a finite number of values. We also provide identification conditions for the average of each random coefficient, rather than the average of the whole vector.
The paper is organized as follows. Section 2 presents the setup, our main assumptions and the type of designs we consider. Section 3 considers identification and estimation under no-anticipation and parallel-trends assumptions. Section 4 considers identification and estimation when one also imposes a random coefficients distributed-lag linear model. Section 5 applies the results of Sections 3 and 4 to estimate the effect of newspapers on electoral turnout. Proofs are relegated to the appendix.
We consider a panel of $G$ groups observed at $T$ periods, respectively indexed by $g$ and $t$. Groups can be locations, like states, counties, or municipalities, but could also just be individuals or firms. We seek to identify (dynamic) effects of a treatment on a given outcome. Let $D_{g,t}$ denote the treatment of group $g$ at period $t$, with support denoted by $\mathcal{D}_t$, supposed to be independent of $g$. Also, let $\bm{D}_g=(D_{g,1},...,D_{g,T})$ denote the treatment path of group $g$, and let $\mathcal{D}$ denote its support, also assumed to be independent of $g$. We let $F_g$ be the first date at which $g$ switches treatment: $F_g= \min \{t\geq 2: D_{g,t}\ne D_{g,1}\}$. If $g$ never switches, we let $F_g=T+1$.
For all $(d_1,...,d_T)\in \mathcal{D}$, let $Y_{g,t}(d_1,...,d_T)$ be the potential outcome of group $g$ at period $t$ if $g$ has the treatment path $(d_1,...,d_T)$. This potential outcome model, introduced by robins1986new, allows for dynamic effects of lagged treatments on the current outcome, and for anticipation effects of future treatments on the current outcome. However, following most of the literature, we rule out anticipation effects hereafter:\footnote{The notation $Y_{g,t}(d_1,...,d_T)$ also implicitly rules out dependence on treatments that may have occurred before period 1; see Section 1.8 in the web appendix of de2020difference for a discussion on this issue.}
Assumption (ref) requires that a group's current outcome does not depend on its future treatments. It is plausible when treatment's introduction is hard to anticipate. It is less plausible when treatment's introduction is announced saliently ahead of time. Then, researchers sometimes redefine a $(g,t)$ cell as treated if at period $t$, it has been announced that group $g$ will get treated in the future.
The potential outcome notation $Y_{g,t}(d_1,...,d_t)$ implicitly assumes that groups' treatment prior to period one, the first time period in the data, does not affect their outcome, the so-called “initial conditions” assumption. When groups might have been exposed to treatment before period one, this assumption is not innocuous, though very few papers have attempted to relax it in the DID literature.
Given that groups are identically distributed, we omit the index $g$ hereafter in the absence of ambiguity.
We first describe several common designs that are not classical designs or binary and staggered adoption designs.
\paragraph{Non-absorbing binary treatments.} First, social scientists are often interested in the effect of a non-absorbing binary treatment. For instance, burgess2015value study the effect, in Kenya, of sharing the ethnicity of a country's president, on a district's volume of public expenditures. Districts can enter and leave the treatment (sharing the president's ethnicity) twice over the study period. An interesting special case is when groups can join and leave treatment once:
where both $E$ and $F$ are random. When (ref) holds, groups may get treated and leave treatment once, at possibly heterogeneous dates $F$ and $E$.
\paragraph{Absorbing treatments with variation in treatment timing and dose.} Second, social scientists are often interested in the effect of an absorbing treatment with variation in treatment timing and dose:
where both $I$ and $F$ are random ($I>0$). If (ref) holds, treatment is absorbing but there is variation across groups in the period at which they start receiving the treatment, and in the dose they receive. For instance, favara2015credit study the effect, in the US, of financial deregulations conducted during the 1990s, on the volume of credit and housing prices. Their design almost satisfies (ref): US states deregulate at heterogeneous times and with heterogeneous intensities. The only difference is that a small number of states deregulate more than once over the study period, so strictly speaking the treatment is not absorbing.
\paragraph{Treatments that vary at baseline.} Third, social scientists are often interested in the effect of treatments whose intensity varies across groups at all time periods, including at period one:
For instance, gentzkow2011 study the effect, in the US, of the number of newspapers in circulation in a county on turnout in presidential elections in that county. In 1868, the first presidential election used in their analysis, counties' number of newspapers ranges from 0 to 33. Another example is fuest2018higher, who study the effect, in Germany, of the local business tax rate on wages. In 1993, the first period in their data, municipalities have business tax rates ranging from 10 to 37 percentage points.
\paragraph{Restriction on the design.} We seek to consider as general designs as possible, to include in particular the cases above. Nonetheless, we impose Assumption (ref) below.
We thus require that there exists an initial treatment value for which there is heterogeneity in the date at which groups change treatment for the first time. This requirement is natural in difference-in-differences contexts, and there are many applications where it holds. Still, it fails in designs without stayers, where $D_2\ne D_1$ and $F=2$ almost surely, as will for instance be the case if $D_{g,t}$ is the amount of rainfall or the average temperature in location $g$ and year $t$: all locations will experience different precipitations or temperatures in years one and two. This also fails if groups all change treatment for the first time at the same date $t_0$, for instance due to a universal policy affecting them all: $F=t_0$ almost surely. In such cases, if all groups are untreated at period one and receive heterogeneous treatment doses at $t_0$, the design is actually an heterogeneous adoption design and one can then use the estimators considered by de2022heterogeneousadoption.
We consider two versions of the parallel trends assumption. The first is the classical one. Hereafter, we let $\bm{0}_t$ denote the vector of $t$ zeros (similarly, we use below $\bm{1}_t$ to denote the vector of $t$ ones).
$Y_{t}(D_1,...,D_1)$ denotes groups' outcome in the counterfactual where they keep their period-one treatment $D_1$ from period one to $t$, hereafter referred to as their status-quo potential outcome. Assumption (ref) requires that groups' outcome evolutions in the status-quo counterfactual do not depend on their actual treatment path, once we condition on $D_1$. If all groups are untreated at period 1 and $V(F|D_{1}=0)>0$, we have $\mathcal{D}^{\text{r}}_1=\{0\}$. Then, Assumption (ref) is equivalent to Assumption (ref). Note that Assumption (ref) restricts only one potential outcome per group, so Assumption (ref) alone does not restrict groups' treatment effects.
Assumptions (ref) or (ref) may be both seen as strict exogeneity conditions. One may worry that non-absorbing designs arise because those that receive the treatment self-select in and out of it, and self-selection makes such strict exogeneity conditions implausible. However, in most of the aforementioned examples, the multiple treatment changes come from laws that are changed several times or repealed after having been enacted. Therefore, while it is important to document the reasons that led the legislator to further or cancel an initial policy change, parallel-trends assumptions are not by construction less plausible in non-absorbing designs.
We first define our parameters of interest, and to this end we introduce additional notation. First, let $S=\text{sgn}(D_F-D_1)$ if $F\le T$, $S=0$ otherwise. Namely, $S=1$ for groups whose treatment increases when it first switches, $S=-1$ for groups whose treatment decreases, and $S=0$ for groups whose treatment never changes. Second, let us define $$\mathcal{D}^{\text{nc}}_t := \left\{(d_1,...,d_T)\in \mathcal{D}: \;\text{either } \min_{t\ge s>1} d_s \ge d_1 \text{ or } \max_{t\ge s>1} d_s \le d_1\right\}.$$ In words, $\mathcal{D}^{\text{nc}}_t$ is the set of treatment paths $(d_1,...,d_T)$ without “crossing” until $t$: there cannot be treatment values $d_s$ and $d_{s'}$ ($s, s'\le t$) both above and below the initial treatment value ($d_s< d_1<d_{s'}$). Finally, let $\overline{T}_d:=\max\text{Supp}(F-1|D_1=d)$ and $\overline{T}:=\overline{T}_{D_1}$. The random variable $\overline{T}$ represents, for a given group, the last date at which we can find with a positive probability a “control group” with the same initial treatment value and whose treatment has not switched yet.
Now, for any $\ell \in\{1,...,T-1\}$ for which $P(D_1\in\mathcal{D}^{\text{r}}_1, F-1+\ell \le \overline{T}, \bm{D}\in\mathcal{D}^{\text{nc}}_{F-1+\ell})>0$, let $$\text{AVSQ}_\ell=E\left[S\times \left(Y_{F-1+\ell}-Y_{F-1+\ell}(D_1,...,D_1)\right)|D_1\in\mathcal{D}^{\text{r}}_1, F-1+\ell \le \overline{T}, \bm{D}\in\mathcal{D}^{\text{nc}}_{F-1+\ell}\right].$$ To interpret $\text{AVSQ}_\ell$, let us first suppose that $S\ge 0$ and $\bm{D}\in\mathcal{D}^{\text{nc}}_{F-1+\ell}$ almost surely (a.s.). Then, $\text{AVSQ}_\ell$ is an average expected difference between groups' actual outcome and their counterfactual “status quo” outcome if their treatment had always remained equal to their period-one value $D_1$, $\ell$ periods after the first switch occurs. Because of this, we refer to $\text{AVSQ}_\ell$ as an actual-versus-status-quo (AVSQ) event-study effect. The expectation is over groups for which $D_1\in\mathcal{D}^{\text{r}}_1, F-1+\ell \le \overline{T}$. Intuitively, and as shown below, these are the groups for which the expected counterfactual outcome can be identified.
In a binary and staggered adoption design, $S\ge 0$ and $\bm{D}\in\mathcal{D}^{\text{nc}}_{F-1+\ell}$ a.s., $\mathcal{D}^{\text{r}}_1=\{0\}$, and thus, $\text{AVSQ}_\ell$ boils down to $$\text{ATT}_\ell:=E\left[Y_{F-1+\ell}(\bm{0}_{F-1},\bm{1}_\ell)-Y_{F-1+\ell}(\bm{0}_{F-1+\ell})|F-1+\ell \le \overline{T}\right].$$ This is the average effect of having received the treatment for $\ell$ periods of time, among treated groups that become treated early enough to reach $\ell$ periods of exposure to treatment at a time period where there is still an untreated group that can be used as a control (the condition $F-1+\ell \le \overline{T}$). $\text{ATT}_\ell$ is the the event-study effect estimated by callaway2018, borusyak2020revisiting, and abraham2018. Therefore, $\text{AVSQ}_\ell$ generalizes $\text{ATT}_\ell$ to non-binary and/or non-staggered designs.
The interpretation of $\text{AVSQ}_\ell$ is more delicate in non-binary and/or non-staggered designs, even if $S\ge 0$ and $\bm{D}\in\mathcal{D}^{\text{nc}}_{F-1+\ell}$ a.s. For instance, if Eq. (ref) holds, so that $D_t=1\{E \geq t\geq F\}$ for some random $E$ and $F$, $\text{AVSQ}_\ell$ is a weighted average of $$E\left[Y_{F-1+\ell}(\bm{0}_{F-1},\bm{1}_\ell)-Y_{F-1+\ell}(\bm{0}_{F-1+\ell})|F-1+\ell \le \min(E, \overline{T})\right]$$ and $$E\left[Y_{f-1+\ell}(\bm{0}_{f-1},\bm{1}_{e-(f-1)}, \bm{0}_{f-1+\ell-e}) -Y_{f-1+\ell}(\bm{0}_{f-1+\ell})|E=e, F=f\right],$$ for all $(e,f)$ satisfying $e< f-1+\ell\le \overline{T}$. The latter expectation is an effect of having been treated for $e-(f-1)$ periods, $f-1+\ell-e$ periods ago. Thus, not only the date at which the effect is evaluated ($F-1+\ell$), but also the number of treatment periods and the recency of the treatment episode varies across groups, complicating the interpretation of $\text{AVSQ}_\ell$. Similarly, with three periods and $\mathcal{D}=\{(0,0,4),(0,1,2),(0,0,0)\}$, $\text{AVSQ}_1$ is a weighted average of $E(Y_{3}(0,0,4)-Y_{3}(0,0,0)|D_3=4)$ and $E(Y_2(0,1)-Y_2(0,0)|D_2=1)$. Thus, not only the date at which the effect is evaluated but also the magnitude of the treatment increments generating $\text{AVSQ}_\ell$ varies across groups, which again complicates the interpretation of $\text{AVSQ}_\ell$.
We introduce the condition $\bm{D}\in\mathcal{D}^{\text{nc}}_{F-1+\ell}$ in the definition of $\text{AVSQ}_\ell$ to ensure that it satisfies the following “no-sign reversal” property \citep*{Imbens94,small2017instrumental}:\\ If $(d_1,...,d_t)\ge (d'_1,...,d'_t)$ (where the inequality should be understood component-wise) $ \Rightarrow Y_t(d_1,...,d_t)\ge Y_t(d'_1,...,d'_t)$ a.s., then $\text{AVSQ}_\ell\ge 0$.\\ Suppose we do not condition on $\bm{D}\in\mathcal{D}^{\text{nc}}_{F-1+\ell}$ in $\text{AVSQ}_\ell$ and assume that $\mathcal{D}=\{(1,1,1),(1,2,0)\}$. Then,
Hence, $\text{AVSQ}_2$ would weight negatively $E\left[Y_3(1,2,1)-Y_3(1,2,0)|F=2\right]$. As a result, we could have $\text{AVSQ}_2<0$ even if increasing the current treatment almost surely increases the outcome.
Similarly, it is also to ensure that $\text{AVSQ}_\ell$ satisfies the no-sign reversal that we multiply the difference between the actual and the status-quo outcome by $S$ in its definition. Conditional on $\mathcal{D}^{nc}_{F-1+\ell}$, groups such that $D_F>D_1$, hereafter referred to as “switchers-in”, are also such that $(D_F,...,D_{F-1+\ell})\ge(D_1,...,D_1)$: the difference between their actual and status-quo outcome is an effect of having exposed to a weakly larger treatment dose for $\ell$ periods. Similarly, groups such that $D_F<D_1$, hereafter referred to as “switchers-out”, are such that $(D_F,...,D_{F-1+\ell})\le(D_1,...,D_1)$: the difference between their actual and status-quo outcome is an effect of having exposed to a weakly lower dose for $\ell$ periods. Multiplying it by -1 ensures that now, this difference is an effect of having been exposed to a weakly larger dose, which can then be aggregated with switchers-in's AVSQ effects.
The following theorem shows that under the above assumptions and for suitable $\ell\ge 1$, $\text{AVSQ}_\ell$ is identified by a difference-in-difference estimand. To simplify notation, let us define
so that $\{D_1\in\mathcal{D}^{\text{r}}_1, F-1+\ell \le \overline{T},\bm{D}\in\mathcal{D}^{\text{nc}}_{F-1+\ell}\}$ is equivalent to $\bm{D}\in \mathcal{D}^\ell$.
For every switcher, the estimand in Theorem 1 compares, for groups that have first switched $\ell-1$ periods ago, their $F-1$ to $F-1+\ell$ outcome evolution to $\Delta_\ell(D_1,F)$, the average $F-1$ to $F-1+\ell$ outcome evolution of groups with the same $D_1$ as the switcher and whose treatment has not changed yet at period $F-1+\ell$.
This theorem implies that when $D_1$ has a finitely supported distribution,\footnote{See chaisemartin2022continuous for an analysis of the case when $D_1$ has a continuous distribution.} we can form the following simple estimator of $\text{AVSQ}_\ell$: $$\widehat{\text{AVSQ}}_\ell = \frac{1}{\# \mathcal{S}_{\ell}} \sum_{g\in \mathcal{S}_{\ell}} S_g\left(Y_{g,F_g-1+\ell}-Y_{g,F_g-1} -\widehat{\Delta}_{g,\ell}\right),$$ where $\# A$ denotes the cardinality of the set $A$, $$\mathcal{S}_{\ell}:=\big\{g\in\{1,...,G\}: \, F_g-1+\ell\le T,\; \bm{D}_g\in\mathcal{D}^{\text{nc}}_{F_g-1+\ell} \text{ and } \#\mathcal{C}_{g,\ell}>0\big\},$$ $\mathcal{C}_{g,\ell}:=\{g'\in\{1,...,G\}:D_{g',1}=D_{g,1}, F_{g'}>F_g-1+\ell\}$ and $$\widehat{\Delta}_{g,\ell} := \frac{1}{\# \mathcal{C}_{g,\ell}} \sum_{g'\in \mathcal{C}_{g,\ell}} (Y_{g',F_g-1+\ell}-Y_{g',F_g-1}).$$ When $\mathcal{S}_{\ell}=\emptyset$, we simply let $\widehat{\text{AVSQ}}_\ell =0$. The set $\mathcal{C}_{g,\ell}$ includes the control groups for $g$, namely the groups with the same period-one treatment as switcher $g$ and whose treatment has not changed yet at period $F_g-1+\ell$. Hence, $\mathcal{S}_{\ell}$ denotes the subset of switchers $g$ satisfying the no-crossing condition and for which a control group can be found. The condition $g\in\mathcal{S}_{\ell}$ can therefore be seen as the finite-sample counterpart of $\bm{D}_g\in\mathcal{D}^\ell$.
By slightly adapting the proof of Theorem 1 in de2020difference, we can prove that $\widehat{\text{AVSQ}}_\ell$ is consistent and asymptotically normal for $\text{AVSQ}_\ell$, as $G$ tends to infinity, under mild restrictions de2020difference.
An appealing feature of the parallel trend assumption is that it is partly testable. As in classical designs, we can define pre-trend estimands that mimic the estimands identifying the $\text{AVSQ}_\ell$ effects. Specifically, for $\ell\ge 1$, let $$\text{AVSQ}_{-\ell} = E\left[S\times(Y_{F-1-\ell} - Y_{F-1} - \Delta_{-\ell}(D_1,F)) |\bm{D}\in\mathcal{D}^\ell, F-1-\ell\ge 1 \right],$$ where $\Delta_{-\ell}(d_1,f) = E[Y_{f-1-\ell}-Y_{f-1}|D_1=d,F>f-1+\ell]$. Under Assumptions (ref)-(ref) and (ref), $\text{AVSQ}_{-\ell} = 0$, so under the maintained Assumptions (ref)-(ref), we can test for Assumption (ref) by testing this condition. Note that compared to the parameter $\text{AVSQ}_{\ell}$, in $\text{AVSQ}_{-\ell}$ we also condition on $F-1-\ell\ge 1$, to ensure that $Y_{F-1-\ell}$ can be computed. Note that if groups switch early on, we may have $P(\bm{D}\in\mathcal{D}^\ell, F-1-\ell\ge 1)=0$, even with $\ell=1$; then, $\text{AVSQ}_{-1}$ is undefined and we cannot test Assumption (ref). Abstracting from this condition $F-1-\ell\ge 1$, $\text{AVSQ}_{-\ell}$ is computed on the same subpopulation as $\text{AVSQ}_{\ell}$.
The estimator of $\text{AVSQ}_{-\ell}$ is similar to that of $\text{AVSQ}_{\ell}$.
As mentioned above, $\text{AVSQ}_\ell$ may be delicate to interpret, as it averages the effects of various treatment paths. Actually, under no-anticipation and parallel trends, we can identify all disaggregated effects of the kind
for all $(d_1,...,d_T)$ such that $d_1\in\mathcal{D}^{\text{r}}_1$ and $\min\{t:d_t\ne d_1\}\le\max\text{Supp}(F-1|D_1=d_1)$. If the design is such that the number of treatment paths meeting these conditions is low relative to $G$, then one may be able to precisely estimate such path-specific event-study effects. This may be a feasible approach, for instance, if $D_t=1\{E\ge t\geq F\}$. But in more complicated designs, the number of paths may be too large for this solution to be practical, especially as $\ell$ increases. In such instances, we still recommend that researchers report the treatment paths entering in $\text{AVSQ}_\ell$, as well their distribution: this information may be helpful to interpret $\widehat{\text{AVSQ}}_\ell$.
\paragraph{Difference-in-difference estimators. }
In a binary and staggered design, $\widehat{\text{AVSQ}}_1$ is numerically equal to the $\text{DID}_M$ estimator in dcDH2020, and for all $\ell\ge 1$, $\widehat{\text{AVSQ}}_\ell$ is numerically equal to the event-study estimator of the effect of $\ell$ periods of exposure to treatment of callaway2018, using the not-yet treated as controls. Outside of binary and staggered designs, when all groups are untreated at period one, $\widehat{\text{AVSQ}}_\ell$ is numerically equal to the estimator obtained by redefining the treatment as an indicator equal to one if group $g$'s treatment has ever changed at $t$, and then computing the event-study estimator of $\ell$ periods of exposure to treatment of callaway2018 with this binarized and staggerized treatment. This “binarize and staggerize” idea has for instance been used by deryugina2017fiscal or krolikowski2018choosing. When groups' period-one treatment varies, the two estimators are not equal:\footnote{For instance, east2023multi consider designs where groups' period-one treatment varies, and binarize and staggerize the treatment and compute the event-study estimators of callaway2018.} $\widehat{\text{AVSQ}}_\ell$ only compares switchers and not-yet-switchers with the same period-one treatment, whereas the estimator of callaway2018 applied to this binarized and staggerized treatment compares switchers and non-switchers with different period-one treatments. Then, that estimator relies on a different parallel trend assumption, namely
When combined with Assumption (ref), one can show that contrary to Assumption (ref), (ref) rules out effects of lagged treatments and/or time-varying treatment effects. We refer to Section 3.1 in de2020difference for a detailed discussion of this issue.
\paragraph{Imputation estimators.}
Our estimator of $\text{AVSQ}_\ell$ consists in estimating the missing counterfactual outcome $Y_{g,F_g-1+\ell}(D_{g,1},...,D_{g,1})$ by $Y_{g,1} + \widehat{\Delta}_g$. Instead, one could consider an imputation estimator, following borusyak2020revisiting, gardnertwo, and liu2021practical. Their papers focus on the case where $D_1=0$, but when $D_1$ varies we propose to extend their method as follows. For each $d_1\in\mathcal{D}^{\text{r}}_1$, we regress $Y_{g,t}$ on groups and time fixed effects, on the subsample $\{(g,t):D_{g,t}=d_1,t<F_g\}$. Let $\widehat{\alpha}_g$ and $\widehat{\gamma}_{d_1,t}$ denote the corresponding estimated group and time fixed effects. Then, we impute $Y_{g,F_g-1+\ell}(D_{g,1},...,D_{g,1})$ by $\widehat{\alpha}_g + \widehat{\gamma}_{D_{g,1},F_g-1+\ell}$. Given the asymptotic results on the imputation estimator in borusyak2020revisiting when $D_1=0$, we anticipate that this estimator would also be asymptotically normal as $G$ tends to infinity, provided that for each $d_1$, the number of groups for which $\{(g,t):D_{g,t}=d_1,t<F_g\}$ also goes to infinity.
Finally, in principle we could use other identifying restrictions than parallel trends, e.g. those underlying synthetic controls \citep*{abadie2010synthetic}, factor models xu2017generalized, or synthetic difference-in-differences \citep*{arkhangelsky2021synthetic}. Those methods also rely on imputing $Y_{g,t}(D_1,...,D_1)$. However, they may be challenging to apply in complex designs, as they require a large number of periods before any treatment change, and a large number of control groups for each set of switchers with a specific value of $D_1$.
Caution must be used when comparing the different $(\text{AVSQ}_\ell)_{\ell\ge 1}$. Not only the elapsed time since the first switch, but also the average increase in received treatment doses (compared to the status-quo) varies with $\ell$. For this reason, we now consider a normalized version of $\text{AVSQ}_\ell$, which allows one to interpret $\text{AVSQ}_\ell$ as a weighted average of marginal effects. To this end, let $$\text{AVSQ}^D_\ell := E\left[S\sum_{k=0}^{\ell-1} \left(D_{F+k}-D_1\right)|\bm{D}\in\mathcal{D}^\ell\right].$$ Hence, $\text{AVSQ}^D_\ell$ represents the average change (always counted positively) in total treatment doses between the actual treatment paths and the status quo. This average is taken over the same groups as in $\text{AVSQ}_\ell$. Note that in binary staggered adoption designs, we simply have $\text{AVSQ}^D_\ell=\ell$. In designs satisfying (ref) ($D_t=I \times 1\left\{t\ge F\right\}$), $\text{AVSQ}^D_\ell=\ell\times E[I]$.
Then, the normalized event-study effect we consider is $$\text{AVSQ}^n_\ell := \frac{\text{AVSQ}_\ell}{\text{AVSQ}^D_\ell}.$$ We define placebo effects as $\text{AVSQ}^n_{-\ell} := \text{AVSQ}_{-\ell}/\text{AVSQ}^D_\ell$ for $\ell\ge 1$. Note that we simply modify the numerator here, since $\text{AVSQ}^D_{-\ell}=0$ by construction.
To better interpret $\text{AVSQ}^n_\ell$, let us consider more disaggregated effects, where only one specific lag is modified. For $k\in \{0,...,\ell-1\}$, let
be the slope of the potential outcome $\ell-1$ periods after the first switch, when the $k$-th lag (the underlined terms) is switched from its status-quo counterfactual value $D_{1}$ to its actual value $D_{F-1+\ell-k}$, whereas all previous treatments are held at their actual values, and all subsequent treatments are held at their status-quo value.\footnote{Thus, when $k=0$, there are actually no $D_1$ terms after the underlined treatment values in (ref) (and the 0-th treatment lag is simply the group's current treatment). We also use in (ref) the convention 0/0=0.}
Next, for any $k\in \{0,...,\ell-1\}$, let $$\omega^\ell_k=\frac{S\times(D_{F-1+\ell-k}-D_{1})}{\text{AVSQ}^D_\ell}.$$ Because $S\times(D_{F-1+\ell-k}-D_{1})\ge 0$ and $\text{AVSQ}^D_\ell>0$, $\omega^\ell_k\ge 0$. Moreover, by definition of $\text{AVSQ}^D_\ell$, $E\left[\sum_{k=0}^{\ell-1} \omega^\ell_k |\bm{D}\in\mathcal{D}^\ell\right]=1$. Finally, note that $$\sum_{k=0}^{\ell-1} \omega^\ell_k \text{SL}^\ell_k = S\times(Y_{F-1+\ell}(D_1,...,D_{F-1},...,D_{F-1+\ell}) - Y_{F-1+\ell}(D_1,...,D_1)).$$ As a result, we obtain that $\text{AVSQ}^n_{\ell}$ is the expectation of a weighted average of $(\text{SL}^\ell_k)_{k=0,...,\ell-1}$, with positive weights:
The decomposition (ref) simplifies much in designs where groups' treatment can change at most once ($D_{g,t}=D_{g,F_g}$ for all $t\ge F_g$). Then, $\omega^\ell_k=1/\ell$ for all $k=0,...,\ell-1$. As a result, $\text{AVSQ}^n_1$ is an effect of the current treatment on the outcome, $\text{AVSQ}^n_{2}$ is a weighted average of the effect of the current treatment and of the first treatment lag on the outcome with weights 1/2, and so on. When groups' treatment can change more than once, estimating $k \mapsto \Omega^\ell_k:= E[\omega^{\ell}_k|\bm{D}\in\mathcal{D}^\ell]$ helps documenting which lags contribute the most to $\text{AVSQ}^n_{\ell}$.
The shape of $\ell\mapsto \text{AVSQ}^n_{\ell}$ may provide insights as to whether the outcome is more affected by recent or by old treatments. To see this, suppose that $\ell \mapsto \Omega^\ell_k$ is decreasing for all $\ell>k$ and assume a linear model $Y_t(d_1,...,d_t)=\mu_t+\sum_{k=0}^t \delta_k d_{t-k}$, in which case $\text{SL}^\ell_k = \delta_k$ for all $k\in\{0,...,\ell-1\}$ and $\ell\ge 1$ for which $P(\bm{D}\in\mathcal{D}^\ell)>0$. Then, $\text{AVSQ}^n_{\ell}=\sum_{k=0}^{\ell-1} \Omega_k^\ell \delta_k$. In this case, one can show that if $k\mapsto \delta_k$ is decreasing, $\ell\mapsto \text{AVSQ}^n_{\ell}$ is decreasing as well. Hence, a decreasing $\ell\mapsto \text{AVSQ}^n_{\ell}$ provides suggestive evidence that potential outcomes are more affected by recent than by past treatments.
It follows directly from the definition of $\text{AVSQ}^n_\ell$ and the results on $\text{AVSQ}_\ell$ that $\text{AVSQ}^n_\ell$ can be identified by $$\text{AVSQ}^n_\ell = \frac{E\left[S\times(Y_{F-1+\ell}-Y_{F-1} - \Delta_\ell(D_1,F))|\bm{D}\in\mathcal{D}^\ell\right]}{E\left[S\sum_{k=0}^{\ell-1} \left(D_{F+k}-D_1\right)|\bm{D}\in\mathcal{D}^\ell\right]}.$$ Like that of $\text{AVSQ}_\ell$, its plug-in estimator is asymptotically normal as $G$ tends to infinity, under mild restrictions.
We show here that the actual-versus-status-quo event-study effects can be used to perform a cost-benefit analysis. Let us suppose that groups always switch to a weakly larger treatment than their period-one treatment:
We impose (ref) to reduce the notational burden. When it fails, one can just conduct separate cost-benefit analyses for groups such that $S=1$ and groups such that $S=-1$.
Now, let us consider the parameter $$\text{ACE}=\frac{\sum_{\ell=1}^{T-1}P(F-1+\ell\le \overline{T}) \text{AVSQ}_\ell}{\sum_{\ell=1}^{T-1}P(F-1+\ell\le \overline{T})E[D_{F-1+\ell}-D_1|F-1+\ell\le\overline{T}]}.$$
As explained in de2020difference, $\text{ACE}$ corresponds to an average cumulative effect per unit of treatment, whence its name. By definition of $\text{AVSQ}_\ell$ and (ref), we can rewrite the $\text{ACE}$ as $$\text{ACE}:=\frac{E\left[\sum_{\ell=1}^{\overline{T}-(F-1)} Y_{F-1+\ell}-Y_{F-1+\ell}(D_1,...,D_1)\right]}{ E\left[\sum_{\ell=1}^{\overline{T}-(F-1)} (D_{F-1+\ell}-D_1)\right]},$$ using the convention that the sums are 0 if $\overline{T}\le F-1$. Now, let us take the perspective of a planner, seeking to conduct a cost-benefit analysis comparing groups' actual treatments $\bm{D}$ to the counterfactual “status-quo” scenario where they would have always kept their period-one treatment. Assume that the outcome is a measure of output, such as agricultural yields or wages, expressed in monetary units. Assume also that the treatment is costly, with a cost linear in dose, and known to the analyst. Then, let $C_\ell\geq 0$ denote the (random) cost of administering one treatment dose in a given group, $\ell-1$ periods after this group first switches. Assuming that the planner's discount factor is equal to $1$, groups' actual treatments are beneficial in monetary terms relative to the status quo, up to period $\overline{T}$, if and only if $$E\left[\sum_{\ell=1}^{\overline{T}-(F-1)} Y_{F-1+\ell}-Y_{F-1+\ell}(D_1,...,D_1)\right] \ge E\left[\sum_{\ell=1}^{\overline{T}-(F-1)} C_\ell(D_{F-1+\ell}-D_1)\right].$$ Equivalently, $\text{ACE} \ge c$, where $$c:=\frac{E\left[\sum_{\ell=1}^{\overline{T}-(F-1)} C_\ell(D_{F-1+\ell}-D_1)\right]}{E\left[\sum_{\ell=1}^{\overline{T}-(F-1)} (D_{F-1+\ell}-D_1)\right]}$$ is the average treatment cost, across all the incremental treatment doses received with respect to the status-quo counterfactual. Then, comparing the $\text{ACE}$ to $c$ is sufficient to evaluate if changing groups' treatments from their initial treatments to $\bm{D}$ was beneficial.
Two additional remarks on the $\text{ACE}$ are in order. First, in binary and staggered designs, if no group is treated at period one and there are never-treated groups ($\overline{T}=T$), we have $$\text{ACE}=\sum_{t=1}^T \frac{P(D_t=1)}{\sum_{t'=1}^T P(D_{t'}=1)} E\left[Y_t(\bm{0}_{F-1},\bm{1}_{t-(F-1)}) - Y_t(\bm{0}_t)|D_t=1\right].$$ The expectation on the right-hand side may be seen as the ATT at period $t$, so the right-hand side may be seen as the global ATT, over the $T$ periods. Hence, the $\text{ACE}$ generalizes the ATT to non-binary and/or non-staggered designs.
Second, it follows directly from the results on $\text{AVSQ}_\ell$ that we can consistently estimate the $\text{ACE}$ by $$\widehat{\text{ACE}} = \frac{\sum_{\ell=1}^{T-1}\widehat{P}(F-1+\ell\le \overline{T}) \widehat{\text{AVSQ}}_\ell}{ \sum_{\ell=1}^{T-1}\widehat{P}(F-1+\ell\le \overline{T})\widehat{E}[D_{F-1+\ell}-D_1|F-1+\ell\le\overline{T}]}.$$
While allowing for dynamic effects is appealing, doing so reduces the sample that can be used in the estimation, which may come with high costs in terms of external validity and statistical precision \citep*{chaisemartin2022continuous}. We show here that parameters closely related to the $\text{AVSQ}_\ell$ can be used to test the null that there are no dynamic effects, namely:
Not rejecting those tests may suggest we could use estimators ruling out dynamic effects chaisemartin2022continuous, which may lead to external-validity and precision gains with respect to the estimators discussed in this paper.
In designs with a binary treatment and where some groups leave the treatment after having been previously treated, liu2021practical propose a test of Assumption (ref) under the standard parallel-trends condition (Assumption (ref)), which amounts to estimating the average treatment effect across previously treated groups, at time periods where those groups have left the treatment. Under Assumption (ref), this average treatment effect should be equal to zero. Their test is implemented by the fect Stata \citep*{fectStata} and R \citep*{fectR} commands.
We can extend this test to non-binary treatments, under our alternative parallel-trends condition (Assumption (ref)). Specifically, let us define $$\text{AVSQ}^o_\ell:= E\left[S\left(Y_{F-1+\ell}- Y_{F-1+\ell}(D_1,...,D_1)\right)|D_{F-1+\ell}=D_1, \bm{D}\in\mathcal{D}^\ell\right],$$ for all $\ell\ge 2$ for which $P(D_{F-1+\ell}=D_1, \bm{D}\in\mathcal{D}^\ell)>0$ (note that $\text{AVSQ}^o_1$ is undefined since by construction $D_F\ne D_1$). Under Assumption (ref), $Y_{F-1+\ell}=Y_{F-1+\ell}(D_{F-1+\ell})$, and thus, $$\text{AVSQ}^o_\ell=0.$$ This can be tested easily. In particular, the Stata command did_multiplegt_dyn estimates the $(\text{AVSQ}_\ell)_{\ell\ge 1}$. To estimate $\text{AVSQ}^o_\ell$ for a specific $\ell\ge 2$, it suffices to use the command on the subsample of $(g,t)$ cells such that $D_{g,F_g-1+\ell}=D_{g,1}$ or $t<F_g$.
A disadvantage of the previous tests is that they can only be used when some switchers eventually revert to their initial treatment value. This does not occur, for instance, with absorbing treatments, e.g. in staggered adoption designs. In such cases, testing static effects is more challenging because basically, treatment effects remain unconstrained even if effects are static. As explained earlier, under no-anticipation and parallel trends, we can identify all disaggregated effects of the kind
for all $(d_1,...,d_T)$ such that $d_1\in\mathcal{D}^{\text{r}}_1$ and $\min\{t:d_t\ne d_1\}\le\max\text{Supp}(F-1|D_1=d_1)$. Under Assumption (ref), $Y_t - Y_t(d_1,...,d_1)= Y_t(d_t)-Y_t(d_1)$, so the effect in the previous display reduces to
but that effect may still depend on $\bm{D}$ in an unrestricted way, if $P(D_{F-1+\ell}= D_1|F-1+\ell\le \overline{T})=0$. Therefore, no-anticipation, parallel trends and Assumption (ref) do not have any testable implication.
To make progress, we add an additional assumption, and propose a joint test of no-anticipation, parallel trends, Assumption (ref), and that additional assumption. Specifically, on top of Assumption (ref), we also assume the following:
Condition (ref) allows for treatment effects to depend on the treatment path in an unrestricted way, but requires that they are time-invariant. In particular, if $D_t=D_{t+1}$, $$E[Y_{t+1}- Y_{t+1}(D_1)|\bm{D}] = \beta(D_{t+1},D_1,\bm{D})= E[Y_t- Y_t(D_1)|\bm{D}].$$ This point is key to obtain our test, based on Proposition (ref).
Proposition (ref) implies that under Assumptions (ref)-(ref) and (ref), we can construct a simple test of both Assumption (ref) and (ref): we just need to estimate $(\text{AVSQ}^{\text{bal}}_\ell)_{\ell \in \{1,...,L\}}$, and test that all effects are equal. For that purpose, one can use the did_multiplegt_dyn command on the subsample of $(g,t)$ cells such that $D_{g,F_g}=D_{g,F_g+1}=...=D_{g,F_g+L}$ or $t<F_g$, specifying the same_switchers and effects_equal options.
Compared to the test of liu2021practical, the test based on $(\text{AVSQ}^{\text{bal}}_\ell)_{\ell=1,...,L}$ is a joint test, rather than a test of Assumption (ref) alone. This makes a rejection of this test more difficult to interpret. On the other hand, this test is feasible even if $P(D_{F-1+\ell}= D_1|F-1+\ell\le \overline{T})=0$, whereas the previous one is not. Even if $P(D_{F-1+\ell}= D_1|F-1+\ell\le \overline{T})>0$, there may be few groups reverting back to their initial treatment, in which case the test of liu2021practical may have lower power than the test based on $(\text{AVSQ}^{\text{bal}}_\ell)_{\ell=1,...,L}$.
We have seen so far that in complex designs with non-absorbing and/or non-binary treatment, we can still identify reduced-form parameters that may be policy relevant, to evaluate the policies that were conducted with respect to what would have happened without any policy. Identification is obtained under no-anticipation and parallel-trends assumptions, whose plausibility can be assessed via pre-trends tests, as in standard DID estimation.
However, this approach also has limitations. While normalized event-study effects estimate weighted averages of marginal effects of the current and lagged treatments on the outcome, they cannot be used to separately estimate those marginal effects. Moreover, the cost-benefit analysis considered above allows one to compare the actual policy with the status-quo, but cannot evaluate alternative policies, or determine the optimal policy.
We address these issues by considering the following additional restriction:
Assumption (ref) allows for heterogeneous treatment effects across groups, and it does not impose any restriction on the dependence between the random coefficients $B$ and the treatment path $\bm{D}$. On the other hand, Assumption (ref) also imposes some restrictions. First, it assumes that the effects of the current treatment and its lags do not vary over time. This condition is strong, but imposing it seems unavoidable if one wants to use past policy changes to do an ex-ante evaluation of future policies: with time-varying effects, the effects of past policy changes cannot be used to infer the effect of future policies. Second, it assumes that the analyst knows the number of lags up to which past treatments still affect the current outcome. Third, it assumes that the effects of the current treatment and its lags are linear, and that they do not complement or substitute each other.
We first study below what the usual distributed-lag TWFE regressions identify under no anticipation, parallel trends and Assumption (ref). Because of the possible correlation between $B$ and $\bm{D}$, such regressions may fail to identify a causal effect. We then consider how to identify expectations of $B$. Finally, we consider relaxations of some of the restrictions imposed by Assumption (ref).
In general designs, a commonly-used estimator of dynamic treatment effects, for instance discussed in Equation (5.2.6) of angrist2009mostly, is the distributed-lag TWFE estimator. For $k \in \{0,...,K\}$, let $\widehat{\beta}_{k}$ denote the coefficient on $D_{g,t-k}$ in a regression of $Y_{g,t}$ on group and period FEs and $\left(D_{g,t-k}\right)_{k \in \{0,...,K\}}$, in the subsample such that $t\geq K+1$:
In practice, researchers may slightly augment or modify (ref). They may include treatment leads in the regression, to test Assumptions (ref) and (ref). They may define the lagged treatments as equal to 0 at time periods when they are not observed, and estimate the regression in the full sample. They may also estimate the regression in first difference and without group fixed effects. Finally, they may include control variables. Results similar to Theorem (ref) below apply to all those variations on (ref).
Before stating this result, we introduce additional notation. For $g=1,...,G$, let $\eta^k_{g,t}$ denote the sample residual of cell $(g,t)$ ($t\ge K+1$) in the regression of $D_{g,t-k}$ on group and time fixed effects and $(D_{g,t-k'})_{k'\ne k}$. Then, for any $(k,k')\in\{0,...,K\}^2$, let $$W^{k,k'}_g=\frac{\sum_{t\ge K+1} \eta^k_{g,t} D_{g,t-k'}}{(1/G) \sum_{g'=1}^G \sum_{t\ge K+1} \eta^k_{g't}D_{g',t-k}}.$$ This random variable depends on $G$ but we omit the dependence to simplify notation. Since the initial sample is i.i.d., the distribution of $W^{k,k'}_g$ does not depend on $g$, so we omit this dependence below.
The proof of Theorem (ref) is very similar to that of Theorem 2 in de2020two, but we include it in the appendix for completeness. It essentially extends Proposition 3 in abraham2018 to non-binary or non-staggered designs. Theorem (ref) shows that even under Assumption (ref), which is restrictive, the distributed-lag TWFE estimator $\widehat{\beta}_{k}$ does not identify $E\left[\beta_{k}\right]$ in general. Rather, it estimates the sum of $K+1$ terms. The first term is a weighted average of $\beta_{k}$, where the weight variable has expectation one but may be negative. Then, if treatment effects are heterogeneous between groups, we may not have $E[W^{k,k} \beta_k]= E[\beta_k]$. If $P(W^k<0)>0$, we may even have a sign reversal, namely $P(\beta_{k}>0)=1$ yet $E\left[\widehat{\beta}_{k}\right]<0$. Now, (ref) also includes a sum of $K$ terms. These terms include $\beta_{k'}$, multiplied by another “weight” $W^{k,k'}$, which has expectation 0. Thus, $\widehat{\beta}_{k}$, which is supposed to estimate the effect of the $k$th treatment lag, is actually contaminated by the effects of other treatment lags. If $V(\beta_{k'})=0$, these contamination terms disappear as the expectation of the contamination weights is equal to zero. But if $V(\beta_{k'})>0$, the contamination terms may differ from zero and can also lead to sign reversal.
Importantly, the weights $(W_g^{k,k'})_{g=1,...,G}$ can be computed, and this computation can help diagnosing the robustness of the distributed-lag TWFE regression to heterogeneous treatment effects. For instance, a useful diagnostic is to look at the variance of the weights $(W_g^{k,k})_{g=1,...,G}$ around their expectation, equal to one. A large variance indicates that the regression strongly upweights the $\beta_k$ of some groups and strongly downweights or even weights negatively the $\beta_k$ of other groups, and if the weights are correlated with groups' treatment effects, this could bias $\widehat{\beta}_{k}$ far away from $E[\beta_k]$.
Finally, note that Theorem (ref) is derived under Assumption (ref), a parallel-trends assumption on the never-treated outcome, while event-study estimands in the previous section relied on Assumption (ref), a parallel-trends assumption on the status-quo outcome. In Section 1 of their Web Appendix, de2020two provide an example where a TWFE regression is actually “less robust” under a parallel-trends assumption in the spirit of Assumption (ref) than under a parallel-trends assumption in the spirit of Assumption (ref).
Under Assumption (ref), the slope $\text{SL}^\ell_k$ defined in (ref) reduces to $\beta_k$. Then, Theorem (ref) shows that $\text{AVSQ}^n_{\ell}$ identifies a weighted average of expectations of the $\beta_k$s, for $k$ ranging from $0$ to $\ell-1$. Actually, as discussed in a working paper version of de2020difference, under Assumption (ref) one can leverage DID estimators to separately estimate the expectation of each $\beta_k$. First, one estimates group-specific regressions of $$(Y_{g,F_g-1+\ell}-Y_{g,F_g-1} -\widehat{\Delta}_{g,\ell})_{\ell\in \{1,...,\overline{T}-(F_g-1)\}},$$ the building-block DID estimators in the previous section, on $(D_{g,F_g-1+\ell}-D_{g,1})_{\ell\in \{1,...,\overline{T}-(F_g-1)\}}$ and its lags. Then, one averages those group-specific regression coefficients across switchers de2020differencev5. In this paper, we will not use DIDs to estimate the expectation of each $\beta_k$. Instead, we will propose estimators that may be used in more general designs, namely in designs that may not have stayers.
First, we show that when combined with the standard parallel trends assumption (Assumption (ref)), Assumption (ref) leads to a particular case of the random coefficients model considered by chamberlain1992efficiency and arellano2012identifying. To see this, first note that Assumption (ref) is equivalent to $$\Delta Y_{t}(\bm{0}_t) = \gamma_t + \varepsilon_{t},\quad E[\varepsilon_{t}|\bm{D}]=0,$$ where $\Delta$ is the first-difference operator. Next, let $\Delta\bm{Y}=(\Delta Y_{K+2},\dots,\Delta Y_{T})$, $\Gamma=(\gamma_{K+2},\dots,\gamma_T)'$, $\bm{\varepsilon} = (\varepsilon_{K+1},\dots,\varepsilon_{T})'$ and $$\bm{M}=
.$$ Then,
We now introduce an identifying assumption on the design, discussed in details below. For any matrix $A$, let $A^+$ denote its Moore-Penrose inverse, $I$ denote the identity matrix (without specifying its size in the absence of ambiguity) and let $\Pi(A):=I-AA^+$ be the orthogonal projector on the orthocomplement of the image of $A$.
We consider identification of average effects of the kind $E[\beta_k|A]$ for some subpopulation $A$ specified below. First, because $\Pi(\bm{M})\bm{M}=0$, it follows from (ref) that under Assumption (ref), $\Gamma$ is identified by
Then, let $e_k$ denote the $k+1$-th canonical vector in $\mathbb R^{K+1}$ (e.g., $e_0=(1,0,...,0)'$). Informally, we can identify $\overline{\beta}_k:=E[\beta_k |e_k\in \text{Im}(\bm{M}')]$, provided that $P(e_k\in \text{Im}(\bm{M}'))>0$, by
Heuristically, (ref) follows from the fact that when $e_k\in \text{Im}(\bm{M}')$, $e'_k\bm{M}^+\bm{M}=e'_k$ and thus $e'_k\bm{M}^+(\Delta Y -\Gamma)=\beta_k+e_k'\bm{M}^+\bm{\varepsilon}$. This identification argument is informal, though, because as discussed in graham2012identification and chaisemartin2022continuous in related setups, the expectation on the right-hand side of (ref) may fail to exist. Namely, we may have
due to groups for which $\|\bm{M}'{}^+e_k\|$ is arbitrary large. In Appendix (ref), we show that indeed, (ref) holds if $T=2$, $K=0$, $\Delta D_2$ has a positive density around 0 and a weak regularity condition holds on $\varepsilon_2$. Nevertheless, the following theorem shows that (i) we can still identify $\overline{\beta}_k$ by a suitable truncation and a continuity argument; (ii) the concern above does not apply if $D_t$ is finitely supported for all $t$.
This theorem is similar to Proposition 1 in arellano2012identifying, with a few differences. First, Model (ref) is a particular case of their random coefficient model. Second, we consider the identification of averages of each component of $B$, rather than the identification of an average of the full vector $B$. This may be important in practice, because in the former case, we just need to restrict the estimation to groups for which $e_k\in \text{Im}(\bm{M}')$, whereas in the latter case we need to restrict the estimation to groups for which $\bm{M}'\bm{M}$ is invertible, and such groups may fail to exist (this is the case for instance if $T<2(K+1)$). The third difference with Proposition 1 in arellano2012identifying is that we highlight the fact that identification can be achieved without trimming when $D_t$ is finitely supported, but may require trimming otherwise.
With respect to the identification results in Section (ref), Theorem (ref) applies to a broader class of designs. In particular, expectations of $(\beta_k)_{k\in \{0,...,\ell-1\}}$ may be identified in designs that do not have stayers for long enough for $\text{AVSQ}_\ell$ to be identified, especially when $T>2(K+1)$, namely when the $(T-(K+1))\times (K+1)$ matrix $\bm{M}$ has more rows than columns.\footnote{Intuitively, this helps for Assumption (ref) because then, the image of $\bm{M}$ has dimension at most $K+1$, and thus the matrix $\Pi(\bm{M})$ has rank at least $T-2(K+1)>0$.} For instance, assume that $K=1$, $T=5$ and $\mathcal{D}$, the support of $\bm{D}$, satisfies $\mathcal{D}=\{(0,1,0,1,0), (0,0,1,1,0)\}$. Then, one can show that Assumption (ref) holds, and $\overline{\beta}_0$ and $\overline{\beta}_1$ are identified. On the other hand, with such a design, $\text{AVSQ}_2$ is not identified because $P(D_1=D_2=D_3)=0$. More generally, when $T>2(K+1)$ we can expect $E[\Pi(\bm{M})]$ to be invertible without requiring $P(D_1=...=D_{K+2})>0$, which is necessary to identify $\text{AVSQ}_{K+1}$. In fact, simulations suggest that Assumption (ref) holds if $\bm{D}$ is continuous with respect to the Lebesgue measure on $\mathbb R^T$, a case where there is no stayer as $D_2\ne D_1$ and $F=2$ a.s, thus implying that Assumption (ref) fails.
When $T \le 2(K+1)$, $E[\Pi(\bm{M})]$ may not be invertible without groups that keep the same treatment value over consecutive periods, so averages of the random coefficients $(\beta_k)_{k\in \{0,...,\ell-1\}}$ are identified under conditions similar to those under which $\text{AVSQ}_\ell$ is identified. For instance, Assumption (ref) fails if $\bm{D}$ is continuous with respect to the Lebesgue measure on $\mathbb R^T$. Then, $P(\bm{M} \text{ is invertible})=1$ and $\Pi(\bm{M})=0$ a.s. Even with a binary treatment, if $K=1$ and $T=4$, one can show that a necessary condition for invertibility is
a condition under which $\text{AVSQ}_2$ is typically identified.
Theorem (ref) suggests the following simple, plug-in estimator of $\Gamma$: $$\widehat{\Gamma} = \left(\frac{1}{G}\sum_{g=1}^G \Pi(\bm{\bm{M}_g})\right)^{-1}\frac{1}{G}\sum_{g=1}^G \Pi(\bm{\bm{M}_g}) \Delta\bm{Y}_g.$$ This estimator is root-$G$ consistent and asymptotically normal, even when $T=2(K+1)$, as soon as Assumption (ref) holds. This contrasts with Proposition 1.1 in graham2012identification, who consider a closely related random coefficients model: their result indicates that the common parameters ($\Gamma$ in our context, $\bm{\delta}_0$ in theirs) cannot be estimated at the standard, root-$G$ rate when the number of time periods is equal to the number of regressors ($T=2(K+1)$ in our context). An important difference with their setup is that we restrict the random coefficients to be time-invariant. This allows us to identify $\Gamma$ using observations for which $\bm{M}$ is singular.
Following Theorem (ref), we discuss separately the estimation of $\overline{\beta}_k$ depending on whether $D_t$ is finitely supported or not.
\paragraph{Finitely supported treatment.}
In this case, we can consider the following simple plug-in estimator:
where $\mathcal{S}^{\text{rc}}:=\{g\in\{1,...,G\}: e_k\in\text{Im}(\bm{M}_g')\}$. Like that of $\Gamma$, this estimator is root-$G$ consistent and asymptotically normal under Assumption (ref)-(ref) and (ref)-(ref). Its asymptotic variance can be obtained by seeing $(\widehat{\Gamma}, \widehat{\overline{B}})$ as a joint GMM estimator, see, e.g., newey1984method.
\paragraph{Not finitely supported treatment.}
When $D_t$ is not finitely supported, the situation is more delicate. In particular, if (ref) holds, we have to rely on (ref) for identification. Then, we expect root-$G$ estimation of $\overline{B}$ to be impossible. Equation (ref) suggests the following estimator, which is close to an estimator of graham2012identification: $$\widehat{\overline{\beta}}_k = \frac{1}{\# \mathcal{S}_C^{\text{rc}}}\sum_{g\in \mathcal{S}_C^{\text{rc}}} e_k' \bm{M}_g^+(\Delta \bm{Y}_g - \widehat{\Gamma}),$$ for some $C>0$ and where $\mathcal{S}_C^{\text{rc}}:=\{g\in\{1,...,G\}: e_k\in\text{Im}(\bm{M}_g'), \,\|\bm{M}_g'{}^+e_k\| \le C\}$. When choosing $C$, we face a usual trade-off between bias, which is small when $C$ is large, and variance, which is small when $C$ is small because we trim groups for which $e_k' \bm{M}_g^+(\Delta \bm{Y}_g - \widehat{\Gamma})$ has a large variance (due to $\|\bm{M}_g'{}^+e_k\|$ being large). Choosing an appropriate $C$ may, however, be delicate in practice.
In this section, we replace Assumption (ref) by the following condition:
Then, we have $$\Delta Y = \Gamma + \bm{M} B + \bm{\varepsilon},$$ where now $\Delta Y=(\Delta Y_2,...,\Delta Y_T)'$, $\bm{\varepsilon}=(\varepsilon_2,...,\varepsilon_T)'$ $\Gamma = (\gamma_2,...,\gamma_T)'$ and $$\bm{M} =
.$$ Then, identification and estimation under Assumption \ref{hyp:RC_model2} proceeds as in Sections \ref{ssub:identification_rc} and \ref{ssub:estimation_rc}. Remark that $\bm{M}$ is now a $(T-1)\times (T-1)$ matrix. Then, one can show that Assumption \ref{hyp:design_RC} holds if and only if there are some ``never switchers'', namely $P(D_1=...=D_T)>0$.
Assumption (ref) allows all lagged treatments up to period one to affect the outcome. Thus, on that dimension it is less restrictive than Assumption (ref), which only allows the first $K$ lags to have an effect. At the same time, note that under Assumption (ref),
the number of lags allowed to affect the outcome depends on $t$, the period when the outcome is measured. This reflects an implicit assumption embedded in our potential outcome notation, namely that treatments prior to period one have no effect on groups' outcomes from period one to $T$. If all lags can affect the outcome, namely if instead of (ref) we have that $$Y_{t}= Y_{t}(\bm{0}_t)+ \sum_{k=0}^{+\infty} \beta_{k} D_{t-k},$$ then the estimators introduced in this section can still be used if groups' treatments do not change before period 1:
(ref) is a strong condition. Still, it holds by construction when the treatment does not exist before period one. When the treatment is absorbing, it also holds by construction in the subsample of untreated groups at period one. If treatments prior to period one have an effect on groups' outcomes and groups' treatments changed before period one, then the estimators introduced in this section may be misleading. Instead, the estimators proposed in Section (ref) can be used in such cases, at the expense of ruling out effects of lagged treatments beyond the $K$th lag, and dropping the outcomes $(Y_t)_{t=1,...,K}$ from the estimation.
In this section, we replace Assumption (ref) by the following condition, which allows for interaction effects between treatment lags:
Then, (ref) still holds, except that now the matrix $\bm{M}$ is of dimension $(T-K+1)\times \frac{(K+1)(K+2)}{2}$, and includes columns of the kind $\Delta (D_{t-k} D_{t-k'})$ for some fixed $(k,k')$ and where $t$ varies over the column. As a result, identification may fail even in some canonical designs. For instance, with a binary and absorbing treatment, one cannot tease out the effects of $D_{t-1}D_t$ and $D_{t-1}$: as no treated group goes back to being untreated, $D_{t-1}D_t=D_{t-1}$.\footnote{Formally, the identifying condition corresponding to $P(e_k\in\text{Im}(\bm{M}'))>0$ in Theorem (ref) does not hold.} On the other hand, identification may still be achieved with non-absorbing treatments. For instance, if $K=1$, $T=5$ and $\mathcal{D}=\{(0,0,0,0,0),(0,1,1,0,0)\}$, we can identify averages of $\beta_{0,0}$, $\beta_{0,1}$, and $\beta_{1,1}$ on groups for which $\bm{D}=(0,1,1,0,0)$.
In this application, we revisit the effect of newspapers on voters' participation in elections, studied by gentzkow2011. They use a US panel data set at the county $\times$ presidential-election level, with 1,195 counties and from the 1868 to the 1928 election. They seek to test a conjecture in de1850democratie, that newspapers encourage citizens to participate more in democratic institutions. For that purpose, they let $Y_{g,t}$ denote the turnout rate in county $g$ and the presidential election that took place in year $t$, they let $D_{g,t}$ denote the number of newspapers circulating in county $g$ and year $t$, and they run a static first-difference regression of $\Delta Y_{g,t}$ on $\Delta D_{g,t}$, and state-year fixed effects. Hereafter, we consider possibly dynamic effects. We first estimate the AVSQ event-study effects, under parallel trends only. Then, we estimate a standard distributed-lag TWFE regressions, before estimating the random coefficients distributed-lag linear model.
First, note that the design of this application is truly complex: the number of newspapers is a non-binary treatment, it can increase or decrease over time, counties can experience several changes in their number of newspapers, at different points in time. The right panel of Figure (ref) displays the distribution of the number of switches in $D_t$. It shows that over the sixteen years of presidential elections, counties experiencing eight or more changes are not uncommon. Actually, only 34 counties (2.9%) never experience any change in their number of newspapers. For 90.9% of the 1,161 remaining counties, $S_g=1$, namely they experience an increase in the number of newspapers at first switch. However, 78.7% of counties experience at least one decrease in their number of newspapers. Still, the no-crossing condition holds for 93.6% of the counties. Finally, the left panel of Figure (ref) displays the distribution of $D_1$, which also shows significant heterogeneity. Assumption (ref) is easily met, with $\mathcal{D}^{\text{r}}_1=\{0,1,...,7\}$.
Estimates of $\text{AVSQ}_\ell$ and pre-trends estimates are shown in Figure (ref) below. Being exposed to a weakly larger number of newspapers for one electoral cycle increases turnout by 1.44 percentage point, and the effect is statistically significant (s.e.=0.43 percentage point). That effect can be estimated for 1,119 out of the 1,195 counties in the data: 34 counties never experience a change in their number of newspapers, and 42 counties that do experience a change cannot be matched with a not-yet-switcher with the same number of newspapers at baseline. Being exposed to a weakly larger number of newspapers for two, three, and four electoral cycles also significantly increases turnout. Effects increase with exposure length, but one cannot reject the null that all effects are equal (p-value=0.40). As $\ell$ increases, effects mechanically apply to fewer and fewer counties, but the effect after four electoral cycles still applies to 917 counties.
Pre-trend estimates are small and individually and jointly insignificant. However, their confidence intervals are quite large. While the first pre-trend estimator applies to 906 of the 1,119 counties for which $\widehat{\text{AVSQ}}_{1}$ is estimated, the fourth pre-trend estimator only applies to 447 of the 917 counties for which $\widehat{\text{AVSQ}}_{4}$ is estimated. The confidence interval of $\widehat{\text{AVSQ}}_{-4}$ is already quite large, but that of $\widehat{\text{AVSQ}}_{-5}$ is substantially larger, so we have very little power to detect differential trends over more than five election cycles. This is why we only report four placebo and four event-study estimators.
To better interpret the $\text{AVSQ}_\ell$, we report the underlying most common treatment paths. For AVSQ$_1$, the three most common effects are effects of having one versus zero newspapers (64% of the cases), two versus zero newspapers (12% of the cases) and two versus one newspapers (5% of the cases). For AVSQ$_2$, the three most common effects are effects of having $D_F=D_{F+1}=1$ instead of $D_F=D_{F+1}=0$ (32% of the cases), $D_F=1, D_{F+1}=0$ instead of $D_F=D_{F+1}=0$ (18% of the cases), and $D_F=1, D_{F+1}=2$ instead of $D_F=D_{F+1}=0$ (12% of the cases). Finally, for AVSQ$_4$, the three most common effects are effects of having $D_F=...=D_{F+3}=1$ instead of $D_F=...=D_{F+3}=0$ (15% of the cases), $D_F=1, D_{F+1}=D_{F+2}=D_{F+3}=0$ instead of $D_F=...=D_{F+3}=0$ (14% of the cases) and $D_F=1, D_{F+1}=D_{F+2}=D_{F+3}=2$ instead of $D_F=...=D_{F+3}=0$ (5% of the cases). As $\ell$ increases, AVSQ$_\ell$ averages the effects of more and more heterogeneous paths across groups, and the three most common paths account for a smaller fraction of all the effects averaged in AVSQ$_{\ell}$. For most paths, too few groups have that path to estimate reasonably precisely a path-specific effect.
Normalized event-study and pre-trends estimates are shown in Figure (ref) below. Normalized event-study estimates are decreasing with $\ell$, but one cannot reject the null that all effects are equal (p-value=0.17). $\widehat{\Omega}^1_0=1$: the first event-study estimate is an effect of contemporaneous newspapers on turnout. $\widehat{\Omega}^2_0=0.48$ and $\widehat{\Omega}^2_1=0.52$: the second normalized event-study estimate is a weighted average of the effects of contemporaneous newspapers and of the first lag of newspapers on turnout, with approximately equal weights. $\widehat{\Omega}^3_0=0.35$, $\widehat{\Omega}^3_1=0.31$, and $\widehat{\Omega}^3_2=0.33$: the third normalized event-study estimate is a weighted average of the effects of contemporaneous newspapers and of the first and second lag of newspapers, with approximately equal weights. Finally, $\widehat{\Omega}^4_0=0.28$, $\widehat{\Omega}^4_1=0.26$, $\widehat{\Omega}^4_2=0.23$, and $\widehat{\Omega}^4_3=0.24$: the fourth normalized event-study estimate is a weighted average of the effects of contemporaneous newspapers and of the first, second, and third lag of newspapers, again with approximately equal weights. Then, the fact that normalized event-study estimates are decreasing with $\ell$ may suggest that lagged newspapers have a smaller effect on turnout than contemporaneous newspapers.
Given that the treatment may revert to its initial value, our first test is applicable here. The results are displayed in Table (ref). We cannot reject the null of static effects, though the tests do not have much power, given the low number of switchers which eventually revert to their initial treatment, for which $\text{AVSQ}_\ell^o$ can be estimated. We also conduct the joint test based on Proposition (ref) above. When considering $L=2$, we can estimate $\text{AVSQ}^{\text{bal}}_\ell$ using 512 counties. In that subsample, the estimates of $\text{AVSQ}^{\text{bal}}_1$ and $\text{AVSQ}^{\text{bal}}_2$ are close and not significantly different (p-value=0.83). Again, this suggests that lagged newspapers do not affect turnout.
Next, we consider a distributed-lag TWFE estimator with just one lag. We obtain $\widehat{\beta}_0=-0.0008$ (s.e.$=0.0014$) and $\widehat{\beta}_1=0.0050$ (s.e.$=0.0015$). Hence, according to this regression, increasing the current number of newspapers insignificantly reduces turnout by 0.08 percentage points. On the other hand, increasing its first lag significantly increases turnout by 0.5 percentage points. These results may look surprising: the normalized event-study effects suggested the opposite, namely that the current number of newspapers would have a larger impact than the number of newspapers four years before. The tests above also do not reject the hypothesis that only the current number of newspapers affects turnout.
However, conclusions drawn from the distributed-lag TWFE estimator are unwarranted if treatment effects are heterogeneous and correlated with the weights. To assess whether this could be the case, we compute the weights in the decomposition (ref). Their distribution is plotted in Figure (ref). Recall that $\widehat{\beta}_0$ estimates the sum of two terms. The first term is a weighted average of 1,195 county-specific effects of current newspapers. Even if most of the weights $(W_g^{0,0})_g$ are non-negative (1,186 of the 1,195 counties), their standard deviation is large (3.19), with a first quartile equal to 0.22, a 99th percentile equal to 8.8 and a maximum as large as 87.8. This implies that the first term upweights very substantially the effects of a small number of counties, whose effects could well differ from the average effect. The “weights” $(W_g^{0,1})_g$, which are centered, also have a large standard deviation (2.09), and vary between -4.65 and 12: contamination of $\widehat{\beta}_0$ by effects of the lagged number of newspapers is substantial.
We obtain similar results for $\widehat{\beta}_1$. The standard deviation of the $(W_g^{1,1})_g$ is also large (2.42), with a first quartile equal to 0.16, a 99th percentile equal to 11.8 and a maximum of 26.1. Finally, the “weights” $(W_g^{1,0})_g$ also have a large standard deviation (1.39), and vary between -7.1 and 15.8.
Next, we turn to the estimation of heterogeneous distributed-lag regressions, following Section (ref) above. We consider both $K=1$, as above, and $K=2$. The results are displayed in Table (ref) below. Note that in both cases, the averages in $\widehat{\overline{\beta}}_k$ apply to a large part of the sample: even $\widehat{\overline{\beta}}_2$ applies to 1,037 counties, namely 90.4% of the sample. The estimators' standard errors are around 50% larger than those of the distributed lag TWFE estimates, but they remain relatively precise. Finally, the results are much more in line with the event-study estimates above, and point towards a positive effect of the current number of newspapers, and little effects of the lagged number of newspapers.
\doublespacing