EconBase
← Back to paper

An Effective Treatment Approach to Difference-in-Differences with General Treatment Patterns

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

52,704 characters · 10 sections · 55 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

An Effective Treatment Approach to Difference-in-Differences with General Treatment Patterns

\doublespacing

abstractWe consider a general difference-in-differences model in which the treatment variable of interest may be non-binary and its value may change in each period. It is generally difficult to estimate treatment parameters defined with the potential outcome given the entire path of treatment adoption, because each treatment path may be experienced by only a small number of observations. We propose an alternative approach using the concept of effective treatment, which summarizes the treatment path into an empirically tractable low-dimensional variable, and develop doubly robust identification, estimation, and inference methods. We also provide a companion R software package. Keywords: dynamic treatment effects, effective treatment, event study, panel data, parallel trends JEL Classification: C14, C21, C23

\pagestyle{plain}

Introduction

Difference-in-differences (DiD) is one of the most popular empirical strategies using panel data. Under the parallel trends assumption, one can identify meaningful treatment parameters by comparing the time evolution of the outcomes between the treatment and comparison groups. Most of the recent literature has been devoted to the development of methods in the staggered adoption setting, in which each unit continues to receive a binary treatment after the first treatment receipt (cf. de2022two; sun2022linear; callaway2023difference; roth2023s).

Although staggered adoption is a typical setting in applied research, there are many empirical situations in which this is not the case.\footnote{ de2022difference conduct a survey of the 100 most-cited articles in the American Economic Review from 2015 to 2019. They document that 26 papers estimate a two-way fixed effects (TWFE) regression, but only 4 papers fit the standard staggered adoption setting. } For example, vella1998whose estimate the effect of union membership on wages. Approximately 28% of the individuals in their sample change union membership status at least twice, implying non-staggered treatment adoption. In another example, deryugina2017fiscal examines the impact of being hit by a hurricane on the fiscal cost for a US county. In this scenario, the treatment variable of interest indicates whether (or the number of times) a county was hit by a hurricane in a year. This empirical situation also does not fit perfectly into the standard setting of staggered adoption because a county damaged by a hurricane this year is not necessarily hit by a hurricane in the following year.

We consider a general DiD model in which the treatment variable of interest may be non-binary and time-varying, in that the treatment realization of each unit may change in each period, which encompasses the staggered setting as a special case. It is generally difficult to estimate sensible treatment parameters using DiD methods conditional on the entire path of treatment adoption. This is because there may be a large number of treatment paths, so that each treatment path may be experienced by only a small set of observations. As a concrete example, in the dataset of vella1998whose, there are 95 different paths of union membership status, 53 of which are experienced by only one individual. In such a case, it is challenging to develop a DiD method that can accurately estimate meaningful treatment parameters.

We propose a DiD method using the concept of effective treatment, which converts the entire treatment path into an empirically tractable low-dimensional variable. In an influential study, manski2013identification develops the concept of effective treatment to summarize potentially complicated treatment spillovers among individuals in the presence of social interactions.\footnote{ The effective treatment is also referred to as exposure mapping in the literature on causal inference under interference (cf., aronow2021spillover). } Building on this idea, we propose an effective treatment approach to summarize the treatment path over time in the DiD setting, while assuming no interference between units. In particular, we consider the following effective treatment specifications useful for estimating the instantaneous and dynamic treatment effects: whether some treatment realization has been received at least once so far, the earliest period of receiving some treatment realization so far, and the number of treatment adoptions so far.

Given an effective treatment specification, we set our target parameter as the average treatment effect for movers (ATEM), who switch from the comparison group to receiving an effective treatment realization between two periods. We provide a set of sufficient conditions, including a conditional parallel trends assumption given observed covariates, under which the ATEM is identified by each of outcome regression (OR), inverse probability weighting (IPW), and doubly robust (DR) estimands. An important feature of our identification analysis is that it allows the chosen effective treatment specification to be “incorrect” in that the potential outcome is not well defined with respect to this effective treatment. This feature is empirically attractive, because researchers generally do not have prior knowledge of the true form of effective treatment. Furthermore, as the parallel trends assumption is key to our analysis, we also derive its testable implication, which suggests a “pre-trends” test similar to existing ones.

Based on the identification result using the DR estimand, we propose a two-step estimation procedure. First, we estimate the OR function for stayers and a generalized propensity score (GPS) under parametric assumptions. Then, we construct an ATEM estimator using the sample counterpart of the identification result. Desirably, the resulting ATEM estimator has the DR property: it is consistent and asymptotically normally distributed if either the parametric assumptions of the OR function or GPS are correct. We also consider a multiplier bootstrap inference procedure for constructing uniform confidence bands (UCB).

As an empirical illustration, we apply our methods to the dataset of vella1998whose. Using each effective treatment specification discussed above, we find no statistical evidence to support significant effects of union membership on current and future wages. Furthermore, the pre-trends testing result is consistent with the parallel trends assumption.

This study builds on a number of recent contributions, including de2020two, goodman2021difference, imai2021use, sun2021estimating, wooldridge2021two, athey2022design, and borusyak2023revisiting. Among them, our methods extend in particular sant2020doubly and callaway2021difference, who develop the DiD methods using the DR estimands in the canonical two-period model and the staggered setting, to our situation with possibly non-binary time-varying treatment.

Another closely related study is de2022difference, who propose DiD estimation for certain instantaneous and dynamic treatment effects in a setting similar to ours. They begin by showing that commonly used TWFE regressions produce biased estimates for treatment effects. Importantly, this negative result is applicable to our situation, and thus, we cannot recover sensible treatment parameters from such TWFE regressions. See de2020two,de2022several, goodman2021difference, and ishimaru2022what for related results. de2022difference then propose an event study estimation method by defining the “event” as the period in which the unit receives a treatment realization for the first time. In comparison, we consider general forms of effective treatment, and our methods have an empirically attractive DR property in the same spirits as sant2020doubly and callaway2021difference.

This study is also related to several recent works that propose DiD methods when the treatment variable of interest is non-binary. callaway2021continuous and de2022continuous consider settings in which the treatment variable has continuous realizations. de2022several study DiD estimation in the presence of several binary treatment variables with staggered adoption. These studies are tailored to specific forms of treatment realizations, whereas our effective treatment approach handles essentially arbitrary treatment realizations.

The rest of this paper is organized as follows. In Section (ref), we introduce the setup and formalize the concept of effective treatment. Sections (ref) and (ref) develop the identification and estimation methods. Section (ref) presents the testable implication for parallel trends. Section (ref) provides the empirical illustration. The supplementary appendix contains technical proofs and simulation results. A companion R package is available from the author's website.

Setup and Effective Treatment

We have panel data $\{ (Y_{it}, D_{it}, X_i): i \in \mathcal{N}, t \in \mathcal{T} \}$, where $Y_{it} \in \mathcal{Y} \subseteq \mathbb{R}$ and $D_{it} \in \mathcal{D} \subseteq \mathbb{R}^{\dim(D)}$ are the outcome and treatment, respectively, for unit $i$ in period $t$, $X_i \in \mathcal{X} \subseteq \mathbb{R}^{\dim(X)}$ is a vector of unit-$i$-specific covariates, and $\mathcal{N} = \{ 1, \dots, N \}$ and $\mathcal{T} = \{ 1, \dots, T \}$ denote the sets of units and periods, respectively. Some unobserved factors may be correlated with $Y_{it}$ and $D_{it}$, implying the possible treatment endogeneity.

The treatment $D_{it}$ may be binary, discrete, continuous, or multidimensional. When $D_{it} = 0$, unit $i$ receives no treatment in period $t$ (throughout the paper, for notational simplicity, we write a generic vector of zeros by $0$). By contrast, $D_{it} = d$ with $d \neq 0$ indicates that unit $i$ receives treatment intensity $d$ in period $t$. The realization of $D_{it}$ can vary over time in that each unit can move into and out of receiving each treatment intensity in each period. For each $t$, let $\bm{D}_{it} = (D_{i1}, \dots, D_{it}) \in \mathcal{D}^t$.

Let $Y_{it}(\bm{d}_T)$ denote the potential outcome given the entire treatment path $\bm{D}_{iT} = \bm{d}_T$. This notation makes explicit that the current outcome may depend on past, current, and future treatment intensity. For example, $Y_{it}(0)$ denotes the potential outcome in period $t$ if unit $i$ has never been treated in $T$ periods. By construction, $Y_{it} = Y_{it}(\bm{D}_{iT})$.

remark[No anticipation] The no-anticipation condition is commonly assumed in the literature, under which the current outcome does not depend on future treatment adoption. Then, the potential outcome in period $t$ given $\bm{D}_{it} = \bm{d}_t$ is well defined. We will reflect our belief in no anticipation when choosing the functional form of effective treatment.

In general, it is difficult to estimate treatment parameters defined with respect to $Y_{it}(\bm{d}_T)$. This is because the number of units that follow a specific treatment path $\bm{d}_T$ may be small in practice, especially when there is much variation in treatment intensity or when the length of the time series is not very small. For example, when $D_{it}$ has $\bar d$ discrete realizations, there may be $(\bar d)^T$ potential outcomes for each $i$ and $t$. In such a situation, only a limited number of units would experience $\bm{d}_T$, making it difficult to obtain precise causal estimates.

To circumvent this problem, we adopt the concept of effective treatment, which converts the treatment path $\bm{D}_{iT}$ into an empirically tractable low-dimensional variable $E_{it}$ that researchers believe to be relevant to the outcome in period $t$. Specifically, for each $t$, consider a user-specified function $E_t: \mathcal{D}^T \to \mathcal{E}_t$, where $\mathcal{E}_t \subset \mathbb{R}^{\dim(E)}$ denotes the range of $E_t$. Here, the functional form of $E_t$ may change depending on $t$. We call $E_t$ the effective treatment function, the corresponding realization $E_{it} \coloneqq E_t(\bm{D}_{iT})$ the realized effective treatment for unit $i$ in period $t$, and the values in $\mathcal{E}_t$ the effective treatment intensity. When $E_{it} = 0$, we consider unit $i$ in period $t$ as the comparison unit under this $E_t$. We write the set of nonzero realizations of $E_t$ as $\widetilde{\mathcal{E}}_t \coloneqq \mathcal{E}_t \setminus \{ 0 \}$.

The following is a set of examples of the functional form of effective treatment.

example[name = The once specification, label = ex:once] If we are interested in the relationship between the outcome and whether the unit has received some treatment intensity at least once so far, we may consider \begin{align} E_{it}^{\mathrm{O}} = \bm{1}\{ \bm{D}_{it} \neq 0 \}. \end{align} The set of nonzero realizations is $\widetilde{\mathcal{E}}_t^{\mathrm{O}} = \{ 1 \}$.
example[name = The event specification, label = ex:event] When the period so far in which each unit is treated for the first time is relevant to the outcome, we may use the following specification: \begin{align} E_{it}^{\mathrm{E}} = \begin{cases} \min \{ s \le t : D_{is} \neq 0 \} & if $\bm{D}_{it} \neq 0$ \\ 0 & otherwise \end{cases} \end{align} We call $E_{it}^{\mathrm{E}}$ and $t - E_{it}^{\mathrm{E}}$ the event date and event time, respectively (cf. miller2023introductory). The set of nonzero realizations in period $t$ is $\widetilde{\mathcal{E}}_t^{\mathrm{E}} = \{ 1, \dots, t \}$.
example[name = The number specification, label = ex:number] When the number of treatment adoptions so far matters for the outcome, we set \begin{align} E_{it}^{\mathrm{N}} = \sum_{s=1}^t \bm{1}\{ D_{is} \neq 0 \}, \end{align} with the corresponding $\widetilde{\mathcal{E}}_t^{\mathrm{N}} = \{ 1, \dots, t \}$.
remark[Limited anticipation] Under limited anticipation, there is some known $\delta > 0$ such that the outcome in period $t$ is determined from the path of treatment adoptions over $t + \delta$ periods (cf. callaway2021difference). We can easily incorporate our belief in it into the functional form of effective treatment. For example, (ref) can be modified to $E_{it}^{\mathrm{O}} = \bm{1}\{ \bm{D}_{i,t+\delta} \neq 0 \}$.

Next, we formalize the correctness of the effective treatment specification. The following definition builds on the treatment effects literature on cross-unit interference (cf., manski2013identification; vazquez2022identification; hoshino2023randomization).

definitionConsider effective treatment functions $E_t^*: \mathcal{D}^T \to \mathcal{E}_t^*$ and $E_t: \mathcal{D}^T \to \mathcal{E}_t$. \begin{enumerate}[(i)] • $E_t^*$ is correct if for all $i \in \mathcal{N}$ and $\bm{d}_T, \bm{d}_T' \in \mathcal{D}^{T}$, $E_t^*(\bm{d}_T) = E_t^*(\bm{d}_T')$ implies $Y_{it}(\bm{d}_T) = Y_{it}(\bm{d}_T')$. • $E_t^*$ is coarser than $E_t$ if there is a surjective mapping $c_t: \mathcal{E}_t \to \mathcal{E}_t^*$ such that $E_t^*(\bm{d}_T) = c_t(E_t(\bm{d}_T))$ for any $\bm{d}_T \in \mathcal{D}^T$. • A correct $E_t^*$ is true if $E_t^*$ is coarser than any other correct specification. \end{enumerate}

There may be several correct specifications. For example, if one of (ref)--(ref) is correct, so is the identity mapping $E_t(\bm{D}_{iT}) = \bm{D}_{iT}$, which is the “finest” correct specification. As the identity mapping is always correct, the true specification always exists, which may or may not equal the identity mapping.

exampleSince $E_{it}^{\mathrm{O}} = \bm{1}\{ E_{it}^{\mathrm{E}} \neq 0 \} = \bm{1}\{ E_{it}^{\mathrm{N}} \neq 0 \}$, the once specification is coarser than the event and number specifications. In general, there is no relationship between the coarseness of the event and number specifications.

Throughout the paper, we write the true effective treatment function as $E_t^*: \mathcal{D}^T \to \mathcal{E}_t^*$, which is generally unknown to researchers. Let $E_{it}^* = E_t^*(\bm{D}_{iT})$ denote the realized true effective treatment for unit $i$ in period $t$. From the definition of the correct specification, we can define the potential outcome for unit $i$ in period $t$ given $E_{it}^* = e$, that is, $Y_{it}^*(e) = Y_{it}(\bm{d}_T)$ for any $\bm{d}_T \in \mathcal{D}^T$ such that $E_t^*(\bm{d}_T) = e$. For example, $Y_{it}^*(0)$ denotes the potential outcome when unit $i$ in period $t$ is a comparison unit under the true $E_t^*$. The individual treatment effect in period $t$ can be written as $Y_{it}^*(e) - Y_{it}^*(0)$, where $e \in \widetilde{\mathcal{E}}_t^*$ and $\widetilde{\mathcal{E}}_t^* \coloneqq \mathcal{E}_t^* \setminus \{ 0 \}$, which may be heterogeneous across $i$, $t$, and $e$.

Ideally, we would like to adopt the true effective treatment function $E_t^*$ for the DiD analysis to obtain meaningful causal estimates. In reality, however, we generally have no prior knowledge of the true specification, unless the canonical two-period model or the staggered case. Furthermore, even if we know a correct specification, such as identity mapping, it may be too “fine” to obtain accurate DiD estimates. Thus, instead of pursing the true specification, we explicitly allow for potential misspecification and aim to obtain relatively precise DiD estimates equipped with sensible causal interpretation. Similar approaches have been advocated in the literature of causal inference in the presence of treatment spillovers (aronow2017estimating; leung2022causal; vazquez2022identification; hoshino2023causal).

Identification Analysis

We begin by imposing the following condition:

assumption[Effective treatment 1] Let $E_t: \mathcal{D}^T \to \mathcal{E}_t$ be a user-specified effective treatment function. \begin{enumerate}[(i)] • For each $t \in \mathcal{T}$, $\mathcal{E}_t$ is a finite set on $\mathbb{R}^{\dim(E)}$. • For any $t \in \mathcal{T}$ and $\bm{d}_T \in \mathcal{D}^T$, $E_t(\bm{d}_T) = 0$ implies $E_t^*(\bm{d}_T) = 0$. \end{enumerate}

Assumption (ref)(i) requires the pre-specified $E_t$ to have discrete realizations. Under Assumption (ref)(ii), any comparison unit under the pre-specified $E_t$ is a comparison unit even under the true $E_t^*$. This should be a mild requirement, because each unit $i$ should be considered as a comparison unit when $\bm{D}_{it} = 0$ or when $\bm{D}_{i,t+\delta} = 0$ with some $\delta > 0$.

Target parameter and its causal interpretation

For periods $t$ and $s$ such that $1 \le s < t$, a user-specified effective treatment function $E_t: \mathcal{D}^T \to \mathcal{E}_t$, and effective treatment intensity $e \in \widetilde{\mathcal{E}}_t$, the ATEM is defined by

align[align omitted — 135 chars of source]

where $M_{i,t,s,e} \coloneqq \bm{1}\{ E_{it} = e, E_{is} = 0 \}$ indicates whether unit $i$ is a mover who moves from the comparison group to receiving effective treatment intensity $e$ between periods $s$ and $t$.

The next theorem is our first main result, which shows that the causal interpretation of $\mathrm{ATEM}(t, s, e)$ is determined from the relationship between the pre-specified $E_t$ and true $E_t^*$.

theoremSuppose that Assumption (ref) holds. \begin{enumerate}[(i)] • If $E_t = E_t^*$, \begin{align*} \mathrm{ATEM}(t, s, e) = \mathrm{ATEM}^*(t, s, e) \coloneqq \operatorname*{\mathbb{E}}[ Y_{it}^*(e) - Y_{it}^*(0) \mid M_{i,t,s,e}^* = 1 ], \end{align*} where $M_{i,t,s,e}^* \coloneqq \bm{1}\{ E_{it}^* = e, E_{is}^* = 0 \}$. • If $E_t$ is not the true but a correct specification, there is a surjective $c_t: \mathcal{E}_t \to \mathcal{E}_t^*$ such that \begin{align*} \mathrm{ATEM}(t, s, e) = \operatorname*{\mathbb{E}}[ Y_{it}^*(c_t(e)) - Y_{it}^*(0) \mid M_{i,t,s,e} = 1 ]. \end{align*} • Assume that $E_t^*$ has discrete realizations (for exposition purposes). If $E_t$ is incorrect, \begin{align*} \mathrm{ATEM}(t, s, e) = \sum_{e^* \in \mathcal{E}_t^*} \operatorname*{\mathbb{E}}[ Y_{it}^*(e^*) - Y_{it}^*(0) \mid M_{i,t,s,e^*}^* = 1, M_{i,t,s,e} = 1 ] \cdot \operatorname*{\mathbb{P}}(M_{i,t,s,e^*}^* = 1 \mid M_{i,t,s,e} = 1 ). \end{align*} \end{enumerate}

Result (i) considers the case where the pre-specified $E_t$ is actually true. In this case, the ATEM indicates the ATE of receiving the true effective treatment intensity $e$ for movers defined with the true $E_t^*$.

Result (ii) corresponds to the situation where the pre-specified $E_t$ is correct but “finer” than the true $E_t^*$. Then, the ATEM can be interpreted as the ATE of receiving the true effective treatment intensity $c_t(e)$ for movers in terms of the pre-specified $E_t$

Result (iii) allows for the possibility that the pre-specified $E_t$ is neither true nor correct. Even in this case, the ATEM has causal interpretation as a weighted average of the ATE conditional on both the “true mover” ($M_{i,t,s,e^*}^* = 1$) and the “pre-specified mover” ($M_{i,t,s,e} = 1$). The weight function is $\operatorname*{\mathbb{P}}(M_{i,t,s,e^*}^* = 1 \mid M_{i,t,s,e} = 1)$, that is, the probability of being the true mover conditional on being the pre-specified mover. Importantly, this weight function is non-negative and sums up to one so that the ATEM satisfies the so-called no-sign-reversal property. More precisely, $\mathrm{ATEM}(t, s, e)$ is positive (resp. negative) if the individual treatment effect $Y_{it}^*(e^*) - Y_{it}^*(0)$ is almost surely (a.s.) positive (resp. negative) for all $e^* \in \widetilde{\mathcal{E}}_t^*$. Thus, in conjunction with the identification results given in the next subsection, our methods do not suffer from the negative weighting or sign-reversal problem, unlike commonly used TWFE regressions (cf. de2020two,de2022difference,de2022several; goodman2021difference; ishimaru2022what).

remark[Test of significance] If there is no treatment effect heterogeneity across units and true effective treatment intensity such that $Y_{it}^*(e^*) - Y_{it}^*(0) = \tau_t$ with some non-stochastic $\tau_t$, we have $\mathrm{ATEM}(t, s, e) = \tau_t$ for any $E_t$. In particular, this result holds for $\tau_t = 0$. Thus, in practice, we can test for the significance of treatment effects by performing statistical inference for $\mathrm{ATEM}(t, s, e)$ even when the pre-specified $E_t$ is incorrect.
comment\begin{remark}[Other interpretation of Theorem (ref)] The results in Theorem (ref) are still valid even when $E_t$ is an observable proxy measure for the unobserved true treatment $E_t^*$, as long as Assumption (ref) holds. For instance, in the scenario of union membership, union membership status ($E_t$) may be a slightly noisy proxy for collective bargaining coverage ($E_t^*$). Together with the identification analysis in the next subsection, Theorem (ref) provides interpretation of DiD estimation in such a situation. \end{remark}

The choice of the effective treatment function $E_t$, periods $t > s$, and effective treatment intensity $e \in \widetilde{\mathcal{E}}_t$ determine what $\mathrm{ATEM}(t, s, e)$ measures more specifically.

example[continues = ex:once] Let $\mathrm{ATEM}^{\mathrm{O}}(t, s, 1) \coloneqq \operatorname*{\mathbb{E}}[ Y_{it} - Y_{it}^*(0) \mid M_{i,t,s,1}^{\mathrm{O}} = 1 ]$, where $M_{i,t,s,1}^{\mathrm{O}} \coloneqq \bm{1}\{ E_{it}^{\mathrm{O}} = 1, E_{is}^{\mathrm{O}} = 0 \}$. By setting $s = 1$, we have \begin{align*} \mathrm{ATEM}^{\mathrm{O}}(t, 1, 1) = \operatorname*{\mathbb{E}}[ Y_{it} - Y_{it}^*(0) \mid D_{i1} = 0, \bm{D}_{it} \neq 0 ], \end{align*} which captures the composition of the instantaneous and dynamic treatment effects, that is, the effect of receiving some treatment intensity at least once between periods $2$ and $t$ for those who did not receive the treatment in the first period.
example[continues = ex:event] Let $\mathrm{ATEM}^{\mathrm{E}}(t, s, e) \coloneqq \operatorname*{\mathbb{E}}[ Y_{it} - Y_{it}^*(0) \mid M_{i,t,s,e}^{\mathrm{E}} = 1 ]$, where $M_{i,t,s,e}^{\mathrm{E}} \coloneqq \bm{1}\{ E_{it}^{\mathrm{E}} = e, E_{is}^{\mathrm{E}} = 0 \}$ with $e \ge 2$ and $t \ge e$. Setting $s = e - 1$, we have \begin{align*} \mathrm{ATEM}^{\mathrm{E}}(t, e-1, e) = \operatorname*{\mathbb{E}} \left[ Y_{it} - Y_{it}^*(0) \mid \bm{D}_{i,e-1} = 0, D_{ie} \neq 0 \right]. \end{align*} This indicates the ATE in period $t$ for those who received some treatment intensity for the first time in period $e$. In the same manner as in the staggered case, estimating this ATEM over a set of $(t, e)$ allows us to understand the instantaneous and dynamic treatment effects separately. Specifically, $\mathrm{ATEM}^{\mathrm{E}}(t, e-1, e)$ with $t = e$ (resp. with $t > e$) is informative about the instantaneous (resp. dynamic) effect of the first treatment receipt in period $e$.
example[continues = ex:number] Define $\mathrm{ATEM}^{\mathrm{N}}(t, s, e) \coloneqq \operatorname*{\mathbb{E}}[ Y_{it} - Y_{it}^*(0) \mid M_{i,t,s,e}^{\mathrm{N}} = 1 ]$, where $M_{i,t,s,e}^{\mathrm{N}} \coloneqq \bm{1}\{ E_{it}^{\mathrm{N}} = e, E_{is}^{\mathrm{N}} = 0 \}$ with $e \ge 1$ and $t \ge e + 1$. When we set $s = 1$, we have \begin{align*} \mathrm{ATEM}^{\mathrm{N}}(t, 1, e) = \operatorname*{\mathbb{E}} \left[ Y_{it} - Y_{it}^*(0) \; \middle| \; D_{i1} = 0, \sum_{r=2}^t \bm{1}\{ D_{ir} \neq 0 \} = e \right]. \end{align*} This ATEM recovers the ATE in period $t$ for those who were not treated in the first period and received treatment $e$ times in $t$ periods. By estimating this ATEM over a set of $(t, e)$, we can separately assess the instantaneous and dynamic effects of the number of treatment adoptions.
remark[Relationship between different specifications] Interestingly, $\mathrm{ATEM}^{\mathrm{O}}(t, 1, 1)$ can be obtained from aggregating $\mathrm{ATEM}^{\mathrm{E}}(t, e-1, e)$ or $\mathrm{ATEM}^{\mathrm{N}}(t, 1, e)$ in a certain way. See the supplementary appendix for more details.
remark[Practical recommendations] It is common practice to perform event study analyses prior to formal DiD estimation (cf. miller2023introductory). Given this, it would be natural to take the event specification as a good starting point and to plot $\mathrm{ATEM}^{\mathrm{E}}(t, e-1, e)$ over a set of $(t, e)$, including “pre-trends” to assess the validity of parallel trends (see Figure (ref) in the empirical illustration section). It would then be desirable to further examine treatment effect heterogeneity with other effective treatment functions, such as the number specification. Finally, to obtain more precise estimates, it would be recommended to estimate $\mathrm{ATEM}^{\mathrm{O}}(t, 1, 1)$ and its time-series average as useful aggregation parameters, that is, \begin{align} \mathrm{ATEM}^{\mathrm{A}} \coloneqq \frac{1}{(T - 1)} \sum_{t=2}^T \mathrm{ATEM}^{\mathrm{O}}(t, 1, 1). \end{align}

Identification of ATEM

Let $S_{i,t,s} \coloneqq \bm{1}\{ E_{it} = 0, E_{is} = 0 \}$ indicate a stayer who remains in the comparison group in periods $s$ and $t$. Define the GPS by

align*[align* omitted — 122 chars of source]

which is the probability of being a mover with receiving effective treatment intensity $e \in \widetilde{\mathcal{E}}_t$ between periods $s$ and $t$ conditional on $X_i$ and on being either a mover or a stayer between periods $s$ and $t$. Let $\mathcal{A}$ denote a set of triplets $(t, s, e)$ at which we would like to identify $\mathrm{ATEM}(t, s, e)$.

assumption[Overlap] For each $(t, s, e) \in \mathcal{A}$, there exists a constant $\varepsilon > 0$ such that $\varepsilon \le \operatorname*{\mathbb{E}}[ S_{i,t,s} \mid X_i ] \le 1 - \varepsilon$ and $\varepsilon \le p_{t,s,e}(X_i) \le 1 - \varepsilon$ a.s.
assumption[Parallel trends] For any $t \ge 2$, \begin{align*} \operatorname*{\mathbb{E}}[ Y_{it}^*(0) - Y_{i,t-1}^*(0) \mid \bm{D}_{iT}, X_i ] = \operatorname*{\mathbb{E}}[ Y_{it}^*(0) - Y_{i,t-1}^*(0) \mid X_i] \quad a.s. \end{align*}

Assumption (ref) is an overlap condition under which there are non-negligible proportions of stayers and movers. This is a common requirement, but clearly restricts the data generating process and functional form of effective treatment. For instance, it will be violated if there are few treated units in each period or if $E_t$ is time invariant (i.e., $E_{it} = E_{is}$ for all $i, t, s$).

The parallel trends condition in Assumption (ref) is essential for our analysis. This assumption states that the evolution of the untreated potential outcome is mean-independent of the entire treatment path, conditional on the unit-specific covariates. Note that this is a condition on the data generating process and does not restrict the specification of effective treatment. For a better understanding, consider the following specific model:

equation*[equation* omitted — 145 chars of source]

where $\gamma_t$ is a non-stochastic coefficient vector, $m$ represents a treatment choice equation, $\alpha_i$ is a unit FE, $\eta_t$ and $\lambda_t$ are non-stochastic time FEs, and $v_{it}$ and $u_{it}$ are idiosyncratic error terms that vary across both $i$ and $t$. The essential requirement is that $\alpha_i$ is additively separable from the other components in the untreated potential outcome equation. If the treatment status in $t = 0$ is deterministic, then $\bm{D}_{iT}$ is determined from $(\alpha_i, \lambda_1, \dots, \lambda_T, u_{i1}, \dots, u_{iT})$ given $X_i$. Then, Assumption (ref) is fulfilled if $(\alpha_i, u_{i1}, \dots, u_{iT})$ are independent of $(v_{i1}, \dots, v_{iT})$ conditional on $X_i$. Importantly, this situation allows for the endogenous treatment choice in that $\alpha_i$ can be correlated with both $Y_{it}$ and $D_{it}$.

commentStrictly speaking, Assumption (ref) can be replaced with the following weaker condition: for any $t, s \in \mathcal{T}$ and $r \ge 2$, \begin{align*} \operatorname*{\mathbb{E}}[ Y_{ir}^*(0) - Y_{i,r-1}^*(0) \mid E_{it}, E_{is}, X_i] = \operatorname*{\mathbb{E}}[ Y_{ir}^*(0) - Y_{i,r-1}^*(0) \mid X_i] \quad a.s. \end{align*} That said, we consider Assumption (ref) to be our identifying assumption because, unlike the condition stated here, Assumption (ref) does not depend on the functional form of effective treatment.

As the second main result in this paper, the next theorem shows that the ATEM is identifiable from each of OR, IPW, and DR estimands. For a generic $A_{it}$ and periods $t > s$, denote $\Delta A_{i,t,s} \coloneqq A_{it} - A_{is}$. Define

equation[equation omitted — 579 chars of source]

where

align*[align* omitted — 400 chars of source]

In words, $m_{t,s}(X_i)$ is the OR function for stayers, $r_{t,s,e}(X_i)$ is the ratio of GPS, and $w_{i,t,s,e}^{M}$ and $w_{i,t,s,e}^{S}$ are weights for movers and stayers, respectively.

theoremSuppose that Assumptions (ref)--(ref) hold. For each $(t, s, e) \in \mathcal{A}$, \begin{align*} \mathrm{ATEM}(t, s, e) = \mathrm{ATEM}^{\mathrm{OR}}(t, s, e) = \mathrm{ATEM}^{\mathrm{IPW}}(t, s, e) = \mathrm{ATEM}^{\mathrm{DR}}(t, s, e). \end{align*}
remark[Identification in the absence of covariates] If the identification conditions are fulfilled without covariates, the identification result reduces to \begin{align*} \mathrm{ATEM}(t, s, e) = \operatorname*{\mathbb{E}}[\Delta Y_{i,t,s} \mid M_{i,t,s,e} = 1] - \operatorname*{\mathbb{E}}[\Delta Y_{i,t,s} \mid S_{i,t,s} = 1]. \end{align*}

In Theorem (ref), the three estimands serve the same role from the perspective of identification analysis, but we prefer the DR estimand for estimation purposes. This is because only it has the DR property. To see this, we introduce parametric models for the OR function and GPS, say $m_{t,s}(x; \gamma_{t,s})$ and $p_{t,s,e}(x; \pi_{t,s,e})$. Here, $m_{t,s}(x; \gamma_{t,s})$ and $p_{t,s,e}(x; \pi_{t,s,e})$ are functions known up to the finite-dimensional parameters $\gamma_{t,s} \in \Gamma_{t,s}$ and $\pi_{t,s,e} \in \Pi_{t,s,e}$. Let $\gamma_{t,s}^*$ and $\pi_{t,s,e}^*$ denote the pseudo-true values. The DR estimand under these parametric specifications is given by

align[align omitted — 273 chars of source]

where

align[align omitted — 309 chars of source]
assumption[Parametric model] For each $(t, s, e) \in \mathcal{A}$, either condition is satisfied. \begin{enumerate}[(i)] • There exists a unique $\gamma_{t,s}^* \in \Gamma_{t,s}$ such that $m_{t,s}(X_i) = m_{t,s}(X_i; \gamma_{t,s}^*)$ a.s. • There exists a unique $\pi_{t,s,e}^* \in \Pi_{t,s,e}$ such that $p_{t,s,e}(X_i) = p_{t,s,e}(X_i; \pi_{t,s,e}^*)$ a.s. \end{enumerate}
theoremSuppose that Assumptions (ref)--(ref) hold. Fix arbitrary $(t, s, e) \in \mathcal{A}$. \begin{enumerate}[(i)] • Under Assumption (ref)(i), $\mathrm{ATEM}(t, s, e) = \mathrm{ATEM}^{\mathrm{DR}}(t, s, e; \gamma_{t,s}^*, \pi_{t,s,e})$ for all $\pi_{t,s,e} \in \Pi_{t,s,e}$. • Under Assumption (ref)(ii), $\mathrm{ATEM}(t, s, e) = \mathrm{ATEM}^{\mathrm{DR}}(t, s, e; \gamma_{t,s}, \pi_{t,s,e}^*)$ for all $\gamma_{t,s} \in \Gamma_{t,s}$. \end{enumerate}

This result, in turn, ensures that the ATEM can be consistently estimated via the DR estimand if either the OR function or GPS is correctly specified.

Estimation and Statistical Inference

We consider estimation and inference methods based on the identification result using the DR estimand in (ref). We focus on the following data generating process:

assumption[DGP] The panel data $\{ (Y_{it}, D_{it}, X_i): i \in \mathcal{N}, t \in \mathcal{T} \}$ are independent and identically distributed (IID) across $i$.

Estimation procedure

We propose a two-step estimation procedure. First, we estimate the finite-dimensional parameters $\gamma_{t,s}$ and $\pi_{t,s,e}$ using parametric estimation methods. For instance, $\gamma_{t,s}$ may be estimated by running a linear regression with the dependent variable $\Delta Y_{i,t,s}$ and the explanatory vector $X_i$ using the observations satisfying $S_{i,t,s} = 1$. Meanwhile, using observations with $M_{i,t,s,e} + S_{i,t,s} = 1$, we can estimate $\pi_{t,s,e}$ using the parametric maximum likelihood method, such as logit or probit estimation, in which the binary response variable is $M_{i,t,s,e}$ and the explanatory vector is $X_i$. Let $\widehat \theta_{t,s,e} \coloneqq (\widehat \gamma_{t,s}^\top, \widehat \pi_{t,s,e}^\top)^\top$ denote the vector of the first-step estimators.

In the second step, we compute the sample analog of the DR estimand in (ref) as follows:

align[align omitted — 281 chars of source]

where ${\operatorname*{\mathbb{E}}}_N[A_{it}] = N^{-1} \sum_{i=1}^N A_{it}$ for generic $A_{it}$ and

align*[align* omitted — 318 chars of source]
remark[Overview of the asymptotic properties] When $N$ tends to infinity and $T$ is fixed, under a set of mild regularity conditions, we can show that the proposed ATEM estimator is $1/\sqrt{N}$-consistent and asymptotically normally distributed. More precisely, we can prove the following influence function representation: \begin{align*} \sqrt{N} \left( \widehat{\mathrm{ATEM}}(t, s, e) - \mathrm{ATEM}(t, s, e) \right) = \frac{1}{\sqrt{N}} \sum_{i=1}^N \psi_{i,t,s,e}( \theta_{t,s,e}^*) + o_{\operatorname*{\mathbb{P}}}(1) \quad for each $(t, s, e) \in \mathcal{A}$, \end{align*} where $\theta_{t,s,e}^* \coloneqq (\gamma_{t,s}^{*\top}, \pi_{t,s,e}^{*\top})^\top$ denotes the pseudo-true parameter vector in the first-stage parametric estimation and $\psi_{i,t,s,e}( \theta_{t,s,e}^*)$ is an influence function, which has a zero mean and a finite variance. See the supplementary appendix for the formal asymptotic analysis.

Multiplier bootstrap inference

Let $\widehat \psi_{i,t,s,e}(\widehat \theta_{t,s,e})$ be an estimator of $\psi_{i,t,s,e}(\theta_{t,s,e}^*)$ obtained by replacing the pseudo-true parameter value $\theta_{t,s,e}^*$ and population mean $\operatorname*{\mathbb{E}}$ with the first-step estimator $\widehat \theta_{t,s,e}$ and sample mean ${\operatorname*{\mathbb{E}}}_N$, respectively. Let $\{ V_i^{\star} \}_{i=1}^N$ be a set of IID bootstrap weights independent of the original sample such that $\operatorname*{\mathbb{E}}[V_i^{\star}] = 0$, $\operatorname*{\mathbb{E}}|V_i^{\star}|^2 = 1$, and $\operatorname*{\mathbb{E}}|V_i^{\star}|^{2 + \varepsilon} < \infty$ for some $\varepsilon > 0$. A common choice is mammen1993bootstrap's (mammen1993bootstrap) weight, such that $\operatorname*{\mathbb{P}}( V_i^{\star} = 1 - \kappa ) = \kappa / \sqrt{5}$ and $\operatorname*{\mathbb{P}}( V_i^{\star} = \kappa ) = 1 - \kappa / \sqrt{5}$ with $\kappa = (\sqrt{5} + 1) / 2$. The bootstrap analog of the ATEM estimator is defined by

align*[align* omitted — 215 chars of source]

We consider the following multiplier bootstrap inference procedure, which is similar to those of belloni2017program and callaway2021difference.

algorithm[algorithm omitted — 1,568 chars of source]

It can be shown that the bootstrap distribution mimics the asymptotic distribution of the ATEM estimator. The proof is analogous to Theorem 3 of callaway2021difference and thus is omitted here. This result, in turn, ensures the asymptotic validity of the multiplier bootstrap inference procedure. More precisely, the coverage error of the $1 - \alpha$ UCB vanishes asymptotically:

align*[align* omitted — 208 chars of source]

Testable Implication for Parallel Trends: Pre-trends Test

To begin with, we add a mild requirement to the pre-specified $E_t$. The following assumption states that a comparison unit in the current period should be a comparison unit even in past periods.

assumption[Effective treatment 2] For any $t \ge 2$, $E_{it} = 0$ implies $E_{i,t-1} = 0$.

The basic idea is as follows. For a period $r$ such that $2 \le r \le s < t$, we denote $\Delta Y_{ir} \coloneqq Y_{ir} - Y_{i,r-1}$. Under the parallel trends condition in Assumption (ref), we can show that

equation[equation omitted — 229 chars of source]

Thus, it is necessary for parallel trends that, in any “pre-effective treatment” period $r$, the outcome evolution for movers should follow the same path as the outcome evolution for stayers. If we find statistical evidence against (ref), the parallel trends assumption would be implausible.

The next theorem formalizes this idea using the DR estimand. For each $(t, s, e) \in \mathcal{A}$ and $r \ge 2$, define

align*[align* omitted — 124 chars of source]

and

align*[align* omitted — 202 chars of source]

where $m_{t,s,r}(X_i) \coloneqq \operatorname*{\mathbb{E}}[ \Delta Y_{ir} \mid S_{i,t,s} = 1, X_i ]$ is the OR function in period $r$ for stayers between periods $s$ and $t$. Note that $\mathrm{ATEM}(t, s, r, e)$ for $t = r$ reduces to $\mathrm{ATEM}(t, s, e)$ in (ref).

theoremSuppose that Assumptions (ref), (ref), and (ref) hold. Under Assumption (ref), \begin{align} \mathrm{ATEM}(t, s, r, e) = \mathrm{ATEM}^{\mathrm{DR}}(t, s, r, e) = 0 \end{align} for all $(t, s, e) \in \mathcal{A}$ and $r$ such that $2 \le r \le s < t$.

This testable implication suggests the following procedure for assessing parallel trends. First, we construct the $1 - \alpha$ UCB for $\mathrm{ATEM}(t, s, r, e)$ using the multiplier bootstrap. Then, the parallel trends assumption should be refuted if a number of the resulting confidence bands exclude zeros.

comment\begin{algorithm}[Pre-trends test] \hfil \begin{enumerate} • Estimate $\mathrm{ATEM}(t, s, r, e)$ in the same way as in Section (ref). • Construct the $1 - \alpha$ UCB for $\mathrm{ATEM}(t, s, r, e)$ using the multiplier bootstrap inference procedure. The bootstrap critical value $\widehat c_{1-\alpha}$ should be computed by the empirical $(1 - \alpha)$-quantile of the maximum of the bootstrapped $t$ test statistics over $(t, s, e) \in \mathcal{A}$ and $r$. • Check if the resulting confidence bands for the “pre-effective treatment” periods include zeros. If a number of confidence bands exclude zeros, the parallel trends assumption would be implausible. \end{enumerate} \end{algorithm}

It should be cautious that “passing” this type of test does not necessarily ensure the validity of parallel trends. This is because the testable implication in (ref) is merely a necessary condition for parallel trends. Indeed, recent studies point out that this type of test often has low power against the violation of parallel trends (cf., roth2022pretest).

Empirical Illustration: Union Wage Premium

vella1998whose find a positive effect of union membership status on wages by estimating a dynamic selection model. In the DiD literature, de2020two revisit the same dataset and demonstrate that current union membership may not significantly influence current wages. Here, we complement these previous results with the proposed methods, which enable the estimation of both the instantaneous and dynamic treatment effects.

The data record 545 full-time working males taken from the National Longitudinal Survey (Youth Sample) for 1980--1987.\footnote{ The data set is available through the Inter-university Consortium for Political and Social Research (de2020data). } The binary treatment $D_{it}$ indicates the union membership status of individual $i$ in year $t$, reflecting whether $i$ is covered by a collective bargaining agreement. The outcome $Y_{it}$ is the logarithm of the hourly wages for individual $i$ in year $t$. As the set of unit-specific covariates $X_i$, we consider race, years of schooling, and years of labor market experience in the first period.

commentWe focus on the three ATEM parameters considered in Examples (ref), (ref), and (ref). \begin{align} \begin{split} \mathrm{ATEM}^{\mathrm{O}}(t, 1, e) & = \operatorname*{\mathbb{E}}[ Y_{it} - Y_{it}^*(0) \mid M_{i,t,1,e}^{\mathrm{O}} = 1], \\ \mathrm{ATEM}^{\mathrm{E}}(t, e-1, e) & = \operatorname*{\mathbb{E}}[ Y_{it} - Y_{it}^*(0) \mid M_{i,t,e-1,e}^{\mathrm{E}} = 1], \\ \mathrm{ATEM}^{\mathrm{N}}(t, 1, e) & = \operatorname*{\mathbb{E}}[ Y_{it} - Y_{it}^*(0) \mid M_{i,t,1,e}^{\mathrm{N}} = 1]. \end{split} \end{align} The first parameter recovers the composition of the instantaneous and dynamic treatment effects, while the latter two can be used to examine these effects separately.

Figures (ref), (ref), and (ref) present the estimates of the three ATEM parameters in Examples (ref), (ref), and (ref) with the corresponding 95% UCBs obtained from 5,000 bootstrap replications using mammen1993bootstrap's (mammen1993bootstrap) weight. For the first-step estimation, we estimate the OR function for stayers and GPS using the method of least squares and logit estimation, respectively. We can see that there are both positive and negative estimates, but many of them are small in the absolute sense. Indeed, all the resulting UCBs include zeros, that is, we find no statistical evidence for current and future union wage premiums. We also find a statistically insignificant estimate of the aggregation parameter $\mathrm{ATEM}^{\mathrm{A}}$ defined in (ref). Specifically, the estimate is $0.041$ and the 95% confidence interval is $[-0.076, 0.159]$.

The pre-trends testing result for the event specification is shown in Figure (ref). No UCBs for “pre-effective treatment” periods exclude zeros, which is consistent with the parallel trends in Assumption (ref).

Accordingly, our DiD analysis suggests that union membership might not substantially improve current and future wages, which is in line with the finding of de2020two.

figure[figure omitted — 420 chars of source]
figure[figure omitted — 640 chars of source]
figure[figure omitted — 426 chars of source]
center[center omitted — 260 chars of source]