Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
52,704 characters · 10 sections · 55 citation commands
An Effective Treatment Approach to Difference-in-Differences with General Treatment Patterns
\doublespacing
\pagestyle{plain}
Difference-in-differences (DiD) is one of the most popular empirical strategies using panel data. Under the parallel trends assumption, one can identify meaningful treatment parameters by comparing the time evolution of the outcomes between the treatment and comparison groups. Most of the recent literature has been devoted to the development of methods in the staggered adoption setting, in which each unit continues to receive a binary treatment after the first treatment receipt (cf. de2022two; sun2022linear; callaway2023difference; roth2023s).
Although staggered adoption is a typical setting in applied research, there are many empirical situations in which this is not the case.\footnote{ de2022difference conduct a survey of the 100 most-cited articles in the American Economic Review from 2015 to 2019. They document that 26 papers estimate a two-way fixed effects (TWFE) regression, but only 4 papers fit the standard staggered adoption setting. } For example, vella1998whose estimate the effect of union membership on wages. Approximately 28% of the individuals in their sample change union membership status at least twice, implying non-staggered treatment adoption. In another example, deryugina2017fiscal examines the impact of being hit by a hurricane on the fiscal cost for a US county. In this scenario, the treatment variable of interest indicates whether (or the number of times) a county was hit by a hurricane in a year. This empirical situation also does not fit perfectly into the standard setting of staggered adoption because a county damaged by a hurricane this year is not necessarily hit by a hurricane in the following year.
We consider a general DiD model in which the treatment variable of interest may be non-binary and time-varying, in that the treatment realization of each unit may change in each period, which encompasses the staggered setting as a special case. It is generally difficult to estimate sensible treatment parameters using DiD methods conditional on the entire path of treatment adoption. This is because there may be a large number of treatment paths, so that each treatment path may be experienced by only a small set of observations. As a concrete example, in the dataset of vella1998whose, there are 95 different paths of union membership status, 53 of which are experienced by only one individual. In such a case, it is challenging to develop a DiD method that can accurately estimate meaningful treatment parameters.
We propose a DiD method using the concept of effective treatment, which converts the entire treatment path into an empirically tractable low-dimensional variable. In an influential study, manski2013identification develops the concept of effective treatment to summarize potentially complicated treatment spillovers among individuals in the presence of social interactions.\footnote{ The effective treatment is also referred to as exposure mapping in the literature on causal inference under interference (cf., aronow2021spillover). } Building on this idea, we propose an effective treatment approach to summarize the treatment path over time in the DiD setting, while assuming no interference between units. In particular, we consider the following effective treatment specifications useful for estimating the instantaneous and dynamic treatment effects: whether some treatment realization has been received at least once so far, the earliest period of receiving some treatment realization so far, and the number of treatment adoptions so far.
Given an effective treatment specification, we set our target parameter as the average treatment effect for movers (ATEM), who switch from the comparison group to receiving an effective treatment realization between two periods. We provide a set of sufficient conditions, including a conditional parallel trends assumption given observed covariates, under which the ATEM is identified by each of outcome regression (OR), inverse probability weighting (IPW), and doubly robust (DR) estimands. An important feature of our identification analysis is that it allows the chosen effective treatment specification to be “incorrect” in that the potential outcome is not well defined with respect to this effective treatment. This feature is empirically attractive, because researchers generally do not have prior knowledge of the true form of effective treatment. Furthermore, as the parallel trends assumption is key to our analysis, we also derive its testable implication, which suggests a “pre-trends” test similar to existing ones.
Based on the identification result using the DR estimand, we propose a two-step estimation procedure. First, we estimate the OR function for stayers and a generalized propensity score (GPS) under parametric assumptions. Then, we construct an ATEM estimator using the sample counterpart of the identification result. Desirably, the resulting ATEM estimator has the DR property: it is consistent and asymptotically normally distributed if either the parametric assumptions of the OR function or GPS are correct. We also consider a multiplier bootstrap inference procedure for constructing uniform confidence bands (UCB).
As an empirical illustration, we apply our methods to the dataset of vella1998whose. Using each effective treatment specification discussed above, we find no statistical evidence to support significant effects of union membership on current and future wages. Furthermore, the pre-trends testing result is consistent with the parallel trends assumption.
This study builds on a number of recent contributions, including de2020two, goodman2021difference, imai2021use, sun2021estimating, wooldridge2021two, athey2022design, and borusyak2023revisiting. Among them, our methods extend in particular sant2020doubly and callaway2021difference, who develop the DiD methods using the DR estimands in the canonical two-period model and the staggered setting, to our situation with possibly non-binary time-varying treatment.
Another closely related study is de2022difference, who propose DiD estimation for certain instantaneous and dynamic treatment effects in a setting similar to ours. They begin by showing that commonly used TWFE regressions produce biased estimates for treatment effects. Importantly, this negative result is applicable to our situation, and thus, we cannot recover sensible treatment parameters from such TWFE regressions. See de2020two,de2022several, goodman2021difference, and ishimaru2022what for related results. de2022difference then propose an event study estimation method by defining the “event” as the period in which the unit receives a treatment realization for the first time. In comparison, we consider general forms of effective treatment, and our methods have an empirically attractive DR property in the same spirits as sant2020doubly and callaway2021difference.
This study is also related to several recent works that propose DiD methods when the treatment variable of interest is non-binary. callaway2021continuous and de2022continuous consider settings in which the treatment variable has continuous realizations. de2022several study DiD estimation in the presence of several binary treatment variables with staggered adoption. These studies are tailored to specific forms of treatment realizations, whereas our effective treatment approach handles essentially arbitrary treatment realizations.
The rest of this paper is organized as follows. In Section (ref), we introduce the setup and formalize the concept of effective treatment. Sections (ref) and (ref) develop the identification and estimation methods. Section (ref) presents the testable implication for parallel trends. Section (ref) provides the empirical illustration. The supplementary appendix contains technical proofs and simulation results. A companion R package is available from the author's website.
We have panel data $\{ (Y_{it}, D_{it}, X_i): i \in \mathcal{N}, t \in \mathcal{T} \}$, where $Y_{it} \in \mathcal{Y} \subseteq \mathbb{R}$ and $D_{it} \in \mathcal{D} \subseteq \mathbb{R}^{\dim(D)}$ are the outcome and treatment, respectively, for unit $i$ in period $t$, $X_i \in \mathcal{X} \subseteq \mathbb{R}^{\dim(X)}$ is a vector of unit-$i$-specific covariates, and $\mathcal{N} = \{ 1, \dots, N \}$ and $\mathcal{T} = \{ 1, \dots, T \}$ denote the sets of units and periods, respectively. Some unobserved factors may be correlated with $Y_{it}$ and $D_{it}$, implying the possible treatment endogeneity.
The treatment $D_{it}$ may be binary, discrete, continuous, or multidimensional. When $D_{it} = 0$, unit $i$ receives no treatment in period $t$ (throughout the paper, for notational simplicity, we write a generic vector of zeros by $0$). By contrast, $D_{it} = d$ with $d \neq 0$ indicates that unit $i$ receives treatment intensity $d$ in period $t$. The realization of $D_{it}$ can vary over time in that each unit can move into and out of receiving each treatment intensity in each period. For each $t$, let $\bm{D}_{it} = (D_{i1}, \dots, D_{it}) \in \mathcal{D}^t$.
Let $Y_{it}(\bm{d}_T)$ denote the potential outcome given the entire treatment path $\bm{D}_{iT} = \bm{d}_T$. This notation makes explicit that the current outcome may depend on past, current, and future treatment intensity. For example, $Y_{it}(0)$ denotes the potential outcome in period $t$ if unit $i$ has never been treated in $T$ periods. By construction, $Y_{it} = Y_{it}(\bm{D}_{iT})$.
In general, it is difficult to estimate treatment parameters defined with respect to $Y_{it}(\bm{d}_T)$. This is because the number of units that follow a specific treatment path $\bm{d}_T$ may be small in practice, especially when there is much variation in treatment intensity or when the length of the time series is not very small. For example, when $D_{it}$ has $\bar d$ discrete realizations, there may be $(\bar d)^T$ potential outcomes for each $i$ and $t$. In such a situation, only a limited number of units would experience $\bm{d}_T$, making it difficult to obtain precise causal estimates.
To circumvent this problem, we adopt the concept of effective treatment, which converts the treatment path $\bm{D}_{iT}$ into an empirically tractable low-dimensional variable $E_{it}$ that researchers believe to be relevant to the outcome in period $t$. Specifically, for each $t$, consider a user-specified function $E_t: \mathcal{D}^T \to \mathcal{E}_t$, where $\mathcal{E}_t \subset \mathbb{R}^{\dim(E)}$ denotes the range of $E_t$. Here, the functional form of $E_t$ may change depending on $t$. We call $E_t$ the effective treatment function, the corresponding realization $E_{it} \coloneqq E_t(\bm{D}_{iT})$ the realized effective treatment for unit $i$ in period $t$, and the values in $\mathcal{E}_t$ the effective treatment intensity. When $E_{it} = 0$, we consider unit $i$ in period $t$ as the comparison unit under this $E_t$. We write the set of nonzero realizations of $E_t$ as $\widetilde{\mathcal{E}}_t \coloneqq \mathcal{E}_t \setminus \{ 0 \}$.
The following is a set of examples of the functional form of effective treatment.
Next, we formalize the correctness of the effective treatment specification. The following definition builds on the treatment effects literature on cross-unit interference (cf., manski2013identification; vazquez2022identification; hoshino2023randomization).
There may be several correct specifications. For example, if one of (ref)--(ref) is correct, so is the identity mapping $E_t(\bm{D}_{iT}) = \bm{D}_{iT}$, which is the “finest” correct specification. As the identity mapping is always correct, the true specification always exists, which may or may not equal the identity mapping.
Throughout the paper, we write the true effective treatment function as $E_t^*: \mathcal{D}^T \to \mathcal{E}_t^*$, which is generally unknown to researchers. Let $E_{it}^* = E_t^*(\bm{D}_{iT})$ denote the realized true effective treatment for unit $i$ in period $t$. From the definition of the correct specification, we can define the potential outcome for unit $i$ in period $t$ given $E_{it}^* = e$, that is, $Y_{it}^*(e) = Y_{it}(\bm{d}_T)$ for any $\bm{d}_T \in \mathcal{D}^T$ such that $E_t^*(\bm{d}_T) = e$. For example, $Y_{it}^*(0)$ denotes the potential outcome when unit $i$ in period $t$ is a comparison unit under the true $E_t^*$. The individual treatment effect in period $t$ can be written as $Y_{it}^*(e) - Y_{it}^*(0)$, where $e \in \widetilde{\mathcal{E}}_t^*$ and $\widetilde{\mathcal{E}}_t^* \coloneqq \mathcal{E}_t^* \setminus \{ 0 \}$, which may be heterogeneous across $i$, $t$, and $e$.
Ideally, we would like to adopt the true effective treatment function $E_t^*$ for the DiD analysis to obtain meaningful causal estimates. In reality, however, we generally have no prior knowledge of the true specification, unless the canonical two-period model or the staggered case. Furthermore, even if we know a correct specification, such as identity mapping, it may be too “fine” to obtain accurate DiD estimates. Thus, instead of pursing the true specification, we explicitly allow for potential misspecification and aim to obtain relatively precise DiD estimates equipped with sensible causal interpretation. Similar approaches have been advocated in the literature of causal inference in the presence of treatment spillovers (aronow2017estimating; leung2022causal; vazquez2022identification; hoshino2023causal).
We begin by imposing the following condition:
Assumption (ref)(i) requires the pre-specified $E_t$ to have discrete realizations. Under Assumption (ref)(ii), any comparison unit under the pre-specified $E_t$ is a comparison unit even under the true $E_t^*$. This should be a mild requirement, because each unit $i$ should be considered as a comparison unit when $\bm{D}_{it} = 0$ or when $\bm{D}_{i,t+\delta} = 0$ with some $\delta > 0$.
For periods $t$ and $s$ such that $1 \le s < t$, a user-specified effective treatment function $E_t: \mathcal{D}^T \to \mathcal{E}_t$, and effective treatment intensity $e \in \widetilde{\mathcal{E}}_t$, the ATEM is defined by
where $M_{i,t,s,e} \coloneqq \bm{1}\{ E_{it} = e, E_{is} = 0 \}$ indicates whether unit $i$ is a mover who moves from the comparison group to receiving effective treatment intensity $e$ between periods $s$ and $t$.
The next theorem is our first main result, which shows that the causal interpretation of $\mathrm{ATEM}(t, s, e)$ is determined from the relationship between the pre-specified $E_t$ and true $E_t^*$.
Result (i) considers the case where the pre-specified $E_t$ is actually true. In this case, the ATEM indicates the ATE of receiving the true effective treatment intensity $e$ for movers defined with the true $E_t^*$.
Result (ii) corresponds to the situation where the pre-specified $E_t$ is correct but “finer” than the true $E_t^*$. Then, the ATEM can be interpreted as the ATE of receiving the true effective treatment intensity $c_t(e)$ for movers in terms of the pre-specified $E_t$
Result (iii) allows for the possibility that the pre-specified $E_t$ is neither true nor correct. Even in this case, the ATEM has causal interpretation as a weighted average of the ATE conditional on both the “true mover” ($M_{i,t,s,e^*}^* = 1$) and the “pre-specified mover” ($M_{i,t,s,e} = 1$). The weight function is $\operatorname*{\mathbb{P}}(M_{i,t,s,e^*}^* = 1 \mid M_{i,t,s,e} = 1)$, that is, the probability of being the true mover conditional on being the pre-specified mover. Importantly, this weight function is non-negative and sums up to one so that the ATEM satisfies the so-called no-sign-reversal property. More precisely, $\mathrm{ATEM}(t, s, e)$ is positive (resp. negative) if the individual treatment effect $Y_{it}^*(e^*) - Y_{it}^*(0)$ is almost surely (a.s.) positive (resp. negative) for all $e^* \in \widetilde{\mathcal{E}}_t^*$. Thus, in conjunction with the identification results given in the next subsection, our methods do not suffer from the negative weighting or sign-reversal problem, unlike commonly used TWFE regressions (cf. de2020two,de2022difference,de2022several; goodman2021difference; ishimaru2022what).
The choice of the effective treatment function $E_t$, periods $t > s$, and effective treatment intensity $e \in \widetilde{\mathcal{E}}_t$ determine what $\mathrm{ATEM}(t, s, e)$ measures more specifically.
Let $S_{i,t,s} \coloneqq \bm{1}\{ E_{it} = 0, E_{is} = 0 \}$ indicate a stayer who remains in the comparison group in periods $s$ and $t$. Define the GPS by
which is the probability of being a mover with receiving effective treatment intensity $e \in \widetilde{\mathcal{E}}_t$ between periods $s$ and $t$ conditional on $X_i$ and on being either a mover or a stayer between periods $s$ and $t$. Let $\mathcal{A}$ denote a set of triplets $(t, s, e)$ at which we would like to identify $\mathrm{ATEM}(t, s, e)$.
Assumption (ref) is an overlap condition under which there are non-negligible proportions of stayers and movers. This is a common requirement, but clearly restricts the data generating process and functional form of effective treatment. For instance, it will be violated if there are few treated units in each period or if $E_t$ is time invariant (i.e., $E_{it} = E_{is}$ for all $i, t, s$).
The parallel trends condition in Assumption (ref) is essential for our analysis. This assumption states that the evolution of the untreated potential outcome is mean-independent of the entire treatment path, conditional on the unit-specific covariates. Note that this is a condition on the data generating process and does not restrict the specification of effective treatment. For a better understanding, consider the following specific model:
where $\gamma_t$ is a non-stochastic coefficient vector, $m$ represents a treatment choice equation, $\alpha_i$ is a unit FE, $\eta_t$ and $\lambda_t$ are non-stochastic time FEs, and $v_{it}$ and $u_{it}$ are idiosyncratic error terms that vary across both $i$ and $t$. The essential requirement is that $\alpha_i$ is additively separable from the other components in the untreated potential outcome equation. If the treatment status in $t = 0$ is deterministic, then $\bm{D}_{iT}$ is determined from $(\alpha_i, \lambda_1, \dots, \lambda_T, u_{i1}, \dots, u_{iT})$ given $X_i$. Then, Assumption (ref) is fulfilled if $(\alpha_i, u_{i1}, \dots, u_{iT})$ are independent of $(v_{i1}, \dots, v_{iT})$ conditional on $X_i$. Importantly, this situation allows for the endogenous treatment choice in that $\alpha_i$ can be correlated with both $Y_{it}$ and $D_{it}$.
As the second main result in this paper, the next theorem shows that the ATEM is identifiable from each of OR, IPW, and DR estimands. For a generic $A_{it}$ and periods $t > s$, denote $\Delta A_{i,t,s} \coloneqq A_{it} - A_{is}$. Define
where
In words, $m_{t,s}(X_i)$ is the OR function for stayers, $r_{t,s,e}(X_i)$ is the ratio of GPS, and $w_{i,t,s,e}^{M}$ and $w_{i,t,s,e}^{S}$ are weights for movers and stayers, respectively.
In Theorem (ref), the three estimands serve the same role from the perspective of identification analysis, but we prefer the DR estimand for estimation purposes. This is because only it has the DR property. To see this, we introduce parametric models for the OR function and GPS, say $m_{t,s}(x; \gamma_{t,s})$ and $p_{t,s,e}(x; \pi_{t,s,e})$. Here, $m_{t,s}(x; \gamma_{t,s})$ and $p_{t,s,e}(x; \pi_{t,s,e})$ are functions known up to the finite-dimensional parameters $\gamma_{t,s} \in \Gamma_{t,s}$ and $\pi_{t,s,e} \in \Pi_{t,s,e}$. Let $\gamma_{t,s}^*$ and $\pi_{t,s,e}^*$ denote the pseudo-true values. The DR estimand under these parametric specifications is given by
where
This result, in turn, ensures that the ATEM can be consistently estimated via the DR estimand if either the OR function or GPS is correctly specified.
We consider estimation and inference methods based on the identification result using the DR estimand in (ref). We focus on the following data generating process:
We propose a two-step estimation procedure. First, we estimate the finite-dimensional parameters $\gamma_{t,s}$ and $\pi_{t,s,e}$ using parametric estimation methods. For instance, $\gamma_{t,s}$ may be estimated by running a linear regression with the dependent variable $\Delta Y_{i,t,s}$ and the explanatory vector $X_i$ using the observations satisfying $S_{i,t,s} = 1$. Meanwhile, using observations with $M_{i,t,s,e} + S_{i,t,s} = 1$, we can estimate $\pi_{t,s,e}$ using the parametric maximum likelihood method, such as logit or probit estimation, in which the binary response variable is $M_{i,t,s,e}$ and the explanatory vector is $X_i$. Let $\widehat \theta_{t,s,e} \coloneqq (\widehat \gamma_{t,s}^\top, \widehat \pi_{t,s,e}^\top)^\top$ denote the vector of the first-step estimators.
In the second step, we compute the sample analog of the DR estimand in (ref) as follows:
where ${\operatorname*{\mathbb{E}}}_N[A_{it}] = N^{-1} \sum_{i=1}^N A_{it}$ for generic $A_{it}$ and
Let $\widehat \psi_{i,t,s,e}(\widehat \theta_{t,s,e})$ be an estimator of $\psi_{i,t,s,e}(\theta_{t,s,e}^*)$ obtained by replacing the pseudo-true parameter value $\theta_{t,s,e}^*$ and population mean $\operatorname*{\mathbb{E}}$ with the first-step estimator $\widehat \theta_{t,s,e}$ and sample mean ${\operatorname*{\mathbb{E}}}_N$, respectively. Let $\{ V_i^{\star} \}_{i=1}^N$ be a set of IID bootstrap weights independent of the original sample such that $\operatorname*{\mathbb{E}}[V_i^{\star}] = 0$, $\operatorname*{\mathbb{E}}|V_i^{\star}|^2 = 1$, and $\operatorname*{\mathbb{E}}|V_i^{\star}|^{2 + \varepsilon} < \infty$ for some $\varepsilon > 0$. A common choice is mammen1993bootstrap's (mammen1993bootstrap) weight, such that $\operatorname*{\mathbb{P}}( V_i^{\star} = 1 - \kappa ) = \kappa / \sqrt{5}$ and $\operatorname*{\mathbb{P}}( V_i^{\star} = \kappa ) = 1 - \kappa / \sqrt{5}$ with $\kappa = (\sqrt{5} + 1) / 2$. The bootstrap analog of the ATEM estimator is defined by
We consider the following multiplier bootstrap inference procedure, which is similar to those of belloni2017program and callaway2021difference.
It can be shown that the bootstrap distribution mimics the asymptotic distribution of the ATEM estimator. The proof is analogous to Theorem 3 of callaway2021difference and thus is omitted here. This result, in turn, ensures the asymptotic validity of the multiplier bootstrap inference procedure. More precisely, the coverage error of the $1 - \alpha$ UCB vanishes asymptotically:
To begin with, we add a mild requirement to the pre-specified $E_t$. The following assumption states that a comparison unit in the current period should be a comparison unit even in past periods.
The basic idea is as follows. For a period $r$ such that $2 \le r \le s < t$, we denote $\Delta Y_{ir} \coloneqq Y_{ir} - Y_{i,r-1}$. Under the parallel trends condition in Assumption (ref), we can show that
Thus, it is necessary for parallel trends that, in any “pre-effective treatment” period $r$, the outcome evolution for movers should follow the same path as the outcome evolution for stayers. If we find statistical evidence against (ref), the parallel trends assumption would be implausible.
The next theorem formalizes this idea using the DR estimand. For each $(t, s, e) \in \mathcal{A}$ and $r \ge 2$, define
and
where $m_{t,s,r}(X_i) \coloneqq \operatorname*{\mathbb{E}}[ \Delta Y_{ir} \mid S_{i,t,s} = 1, X_i ]$ is the OR function in period $r$ for stayers between periods $s$ and $t$. Note that $\mathrm{ATEM}(t, s, r, e)$ for $t = r$ reduces to $\mathrm{ATEM}(t, s, e)$ in (ref).
This testable implication suggests the following procedure for assessing parallel trends. First, we construct the $1 - \alpha$ UCB for $\mathrm{ATEM}(t, s, r, e)$ using the multiplier bootstrap. Then, the parallel trends assumption should be refuted if a number of the resulting confidence bands exclude zeros.
It should be cautious that “passing” this type of test does not necessarily ensure the validity of parallel trends. This is because the testable implication in (ref) is merely a necessary condition for parallel trends. Indeed, recent studies point out that this type of test often has low power against the violation of parallel trends (cf., roth2022pretest).
vella1998whose find a positive effect of union membership status on wages by estimating a dynamic selection model. In the DiD literature, de2020two revisit the same dataset and demonstrate that current union membership may not significantly influence current wages. Here, we complement these previous results with the proposed methods, which enable the estimation of both the instantaneous and dynamic treatment effects.
The data record 545 full-time working males taken from the National Longitudinal Survey (Youth Sample) for 1980--1987.\footnote{ The data set is available through the Inter-university Consortium for Political and Social Research (de2020data). } The binary treatment $D_{it}$ indicates the union membership status of individual $i$ in year $t$, reflecting whether $i$ is covered by a collective bargaining agreement. The outcome $Y_{it}$ is the logarithm of the hourly wages for individual $i$ in year $t$. As the set of unit-specific covariates $X_i$, we consider race, years of schooling, and years of labor market experience in the first period.
Figures (ref), (ref), and (ref) present the estimates of the three ATEM parameters in Examples (ref), (ref), and (ref) with the corresponding 95% UCBs obtained from 5,000 bootstrap replications using mammen1993bootstrap's (mammen1993bootstrap) weight. For the first-step estimation, we estimate the OR function for stayers and GPS using the method of least squares and logit estimation, respectively. We can see that there are both positive and negative estimates, but many of them are small in the absolute sense. Indeed, all the resulting UCBs include zeros, that is, we find no statistical evidence for current and future union wage premiums. We also find a statistically insignificant estimate of the aggregation parameter $\mathrm{ATEM}^{\mathrm{A}}$ defined in (ref). Specifically, the estimate is $0.041$ and the 95% confidence interval is $[-0.076, 0.159]$.
The pre-trends testing result for the event specification is shown in Figure (ref). No UCBs for “pre-effective treatment” periods exclude zeros, which is consistent with the parallel trends in Assumption (ref).
Accordingly, our DiD analysis suggests that union membership might not substantially improve current and future wages, which is in line with the finding of de2020two.