Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
55,239 characters · 10 sections · 39 citation commands
Finitely Heterogeneous Treatment Effect in Event-study
Keywords: event-study, difference-in-differences, panel data, heterogeneity, \\ classification, $K$-means clustering
JEL classification codes: C13, C14, C23
The event-study design, which utilizes the variation in the timing of an event, is an empirical framework whose popularity among empirical researchers has risen tremendously over the time. In applied microeconomics, the event-study design is most often implemented with the difference-in-differences (diff-in-diff) approach. A key identifying assumption of the diff-in-diff style event-study research design is the parallel trend assumption: temporal differences of untreated potential outcomes are mean independent of treatment timing.
The traditional parallel trend assumption that assumes parallel trend with unconditional means is limited in its scope; it is directly violated when there exist heterogeneous patterns of time trends across units. A widely used alternative to this unconditional parallel trend assumption is to assume a conditional parallel trend with observable conditioning covariates (see Abadie2005,CS among others). Following the same spirit, we also assume conditional parallel trend assumption to relax the unconditional parallel trend assumption, but assume that the conditioning variable is latent. Throughout the paper, we call the conditioning variable a latent `type' variable and call the conditional parallel trend a `type-specific parallel trend' assumption. To illustrate the assumption, let us consider a simple two time periods case: with a latent type variable $k_i$,
$Y_{it}(\infty)$ denotes untreated potential outcome for unit $i$ at time $t$ and $D_{i}$ denotes an indicator whether unit $i$ is treated at time $t=2$. When $k_i$ is observable, Equation (ref) is the conditional parallel trend assumption as in Abadie2005,CS.
In this paper, allowing for the conditioning variable $k_i$ to be latent comes at three additional requirements on the model and data. Firstly, we assume that the latent type variable $k_i$ has a finite support. Secondly, we assume that the types are sufficiently separated in terms of pretreatment outcomes. Lastly, we assume that the number of pretreatment periods grows to infinity. In this sense, the type-specific parallel trend assumption can be understood as an alternative to the unconditional parallel trend or the conditional parallel trend with observable conditioning covariates; it allows for heterogeneity patterns that may not be captured by the observable information, at the price of assuming that the heterogeneity is finite and the pretreatment time series is long. Moreover, the finite support assumption motivates the use of a simple within-type diff-in-diff style estimator, which is potentially less susceptible to the overlap problem compared to the conditional diff-in-diff estimators with a continuous conditioning variable.
The finite support assumption gives our framework a unique merit; it helps us analyze the patterns of treatment effect heterogeneity. While the type-specific parallel trend framework of this paper does not put any restriction on the treatment effect heterogeneity, thus allowing for fully flexible treatment effect heterogeneity, the finite support assumption provides a model-based stratifying structure that can be helpful in summarizing the treatment effect heterogeneity. Let $Y_{i2}(2)$ denote the treated potential outcome of unit $i$ at time $t=2$. The existing literature mostly focuses on estimating the conditional expectation of $Y_{i2}(2) - Y_{i2}(\infty)$ given some observable information: e.g., $\mathbf{E} \left[ Y_{i2}(2) - Y_{i2}(\infty) | X_i, D_i=1 \right]$. With the type-specific parallel trend framework, we document treatment effect heterogeneity along the latent type variable $k_i$: $$ \mathbf{E} \left[ Y_{i2}(2) - Y_{i2}(\infty) | k_i, D_i = 1 \right].\footnotemark $$ The stratifying structure induced by the finite type can be particularly helpful for policy recommendation when policymakers have little to zero constraints in policy assignment but do not have perfect information for individual-level responsiveness; it provides a feasible alternative to individual-level treatment assignment, by identifying the treatment effect at the type level. \footnotetext{The treatment effect parameter is comparable to those defined in Abadie2005,SZ,CS when $k_i$ is observable.}
To apply the type-specific parallel trend framework to datasets, we propose a two-step estimation procedure. In the first step, we classify units into the finite number of types, using the $K$-means clustering algorithm on the long pretreatment periods. Given the classification result, the second step of the estimation procedure estimates the treatment effect, using the estimated types as given.\footnotemark \footnotetext{This two-step property of the estimation procedure closely relates to the stratification exercise used in estimating subpopulation treatment effect (see ACW). The goal of the stratification (i.e. classification in this paper's terminology) is to find groups of units whose (estimated) counterfactual untreated outcomes are similar. The type-specific parallel trend assumption directly relates to this since under the type-specific parallel trend assumption, units with the same type share the same time trend.} In the estimation step, a variety of existing estimation strategies can be used by treating the type as a given categorical variable: CD,BJS,CS,SA. As with CD and CS, our `type-specific diff-in-diff' estimator takes an average of the canonical diff-in-diff estimates with two time periods and two treatment timings.
To discuss asymptotic properties of the treatment effect estimators, we first show that the probability of first-step misclassification goes to zero when the number of pretreatment time periods grows at a polynomial rate of the number of units. Given that the number of pretreatment time periods grows sufficiently fast compared to the number of units, the type-specific diff-in-diff estimators are consistent and asymptotically normal under some regularity conditions. These asymptotic results are supported by Monte Carlo simulations.
This paper contributes to the large literature of panel data models where interactive fixed-effects models are used to control for unit heterogeneity across treatment cohorts, relaxing the two-way fixed-effect (TWFE) specification used in the parallel trend approach: see ADH10,Aetal,ABDIK,HCW,FHS,X,CHLZ,CKarami,JS among others. The interactive fixed-effect model often assumes a factor model for the interactive fixed-effects; our framework can be thought of as a special case of the factor model with a finite support on the factor. JS is closest to this paper in that it takes the same approach; however, JS neither explores treatment effect heterogeneity nor develop a full asymptotic theory. As with this paper, most of the papers in the interactive fixed-effect model literature rely on large pretreatment periods. Notable exceptions are CKarami and FHS. In both papers, informative external variables or instruments are used, instead of assuming long pretreatment periods.
This paper also closely relates to the rapidly growing literature on heterogeneous treatment effect: see CD,SA,CS,G,BJS,BLW,GHK, among others. CS is particularly close to this paper in the sense that they also consider a conditional parallel trend assumption. This literature discusses the negative weighting problem that arises in the standard TWFE specification when there is treatment effect heterogeneity across units. We build upon this literature and construct the type-specific diff-in-diff estimator to be robust to the treatment effect heterogeneity.
The rest of the paper is organized as follows. In Section \hyperlink{Sec2}{2}, we formally discuss the type-specific parallel trend assumption. In Section \hyperlink{Sec3}{3}, we propose the two-step estimation procedure for treatment effect estimation. In Section \hyperlink{Sec4}{4}-\hyperlink{Sec5}{5}, we discuss the asymptotic properties of the estimator. In Section \hyperlink{Sec6}{6}, we present Monte Carlo simulation results on the finite-sample performance of the estimator. In Section \hyperlink{Sec7}{7}, we provide an empirical illustration of the type-specific diff-in-diff estimator by revisiting L.
\hypertarget{Sec2}
In the main model of the paper, an econometrician observes a panel data with a binary treatment: $\left\lbrace \left\lbrace Y_{it}, D_{it} \right\rbrace_{t=-T_0-1}^{T_1-1}\right\rbrace_{i=1}^n$. $Y_{it}$ is the outcome variable for unit $i$ at time $t$ and $D_{it} \in \{0,1\}$ is the binary treatment variable for unit $i$ at time $t$. $D_{it}$ follows the staggered adoption scheme; $D_{it} \leq D_{it+1}$. $E_i = \min \{t: D_{it}=1\}$ denotes the treatment timing of unit $i$; this is the same notation as in SA and equivalent with treatment `group' notation $G$ in CS. For never-treated units, let $E_i = \infty$. There are $n$ units and $T+1=T_0+T_1+1$ time periods, with the unit index ranging $i=1, \cdots, n$ and the time index ranging $t=-T_0-1,\cdots,T_1-1$. $T_0+1$ denotes the number of population pretreatment periods and $T_1$ denotes the number of population treatment periods; $\sum_{i=1}^n D_{it} = 0$ for all $t < 0$. $t < 0$ denotes pretreatment periods at the population level and $t \geq 0$ denotes treatment periods at the population level. $T_1$ is fixed. Throughout the paper, we use the potential outcome framework to discuss treatment effect:
$Y_{it}(e)$ is the potential outcome of unit $i$ at time $t$ when their treatment timing is $e$. Thus, for some $Y_{it}(e)$, $t<e$ means untreated potential outcome and $t \geq e$ means treated potential outcome.
The key assumption of this paper is that there exists a unit-level latent type variable. Conditional upon the latent type, the parallel trend assumption and the no anticipation assumption hold.
Assumption \hyperlink{A1}{1} assumes that there is some latent type variable $k_i$, conditioning on which the time trend of the never-treated potential outcome $Y_{it}(\infty) - Y_{is}(\infty)$ is mean independent of the treatment timing $E_i$.\footnotemark \footnotetext{The parallel trend is assumed for every time period, across every treatment timing; when $K=1$, this is equivalent to the conventional parallel trend assumption assumed in CD,SA, but different from CS, which restricts the window of time periods to apply the parallel trend to.} Assumption \hyperlink{A2}{2} assumes that the never-treated potential outcome $Y_{it}(\infty)$ and the pretreatment potential outcome $Y_{it}(e)$ for $t < e$ have the same conditional mean, given the latent type $k_i$ and the treatment timing $E_i$.
In addition to Assumptions \hyperlink{A1}{1}-\hyperlink{A2}{2}, we assume the following two assumptions, to have a type classification results on $k_i$.
Assumption \hyperlink{finite}{3} assumes that the latent type variable $k_i$ has a finite support, as in the group fixed-effect literature or the finite mixture literature. Assumption \hyperlink{separation}{4} assumes that the types are well-separated in terms of $\mathbf{E}[Y_{it}(\infty) - Y_{it-1}(\infty) | k_i=k]$; for any two different types; the $l_2$ norm of the difference between their conditional means is strictly nonzero. Note that the separation assumption is in relation to time trends of the never-treated potential outcomes $Y_{it}(\infty)$. From Assumptions \hyperlink{A1}{1}-\hyperlink{A2}{2}, we have
whenever $t < e$. Thus, Assumption \hyperlink{separation}{4} can be extended to the pretreatment outcomes of all units regardless of their treatment timings.
Given the finite type structure, let us define type-specific treatment effect parameters. Fix two time periods $(s,t)$ and a treatment cohort $\{i:E_i=e\}$ such that $s < e \leq t$. The type-cohort-and-time specific average treatment effect on treated units (ATT) for time $t$, type $k$ and treatment timing $e$ can be written as follows:
The third equality holds from Assumptions \hyperlink{A1}{1}-\hyperlink{A2}{2} and thus $ATT_t(k,e)$ is written as a function of the conditional moments of $\{Y_{it}\}_{t}$ given the latent type variable $k_i$ and the treatment timing variable $E_i$. $ATT_t(k,e)$ is a type-specific version of the treatment effect estimands used in the diff-in-diff literature: see CS,SA for more.
The type-cohort-and-time specific ATT parameter in (ref) takes treatment timing $E_i$ as a conditioning variable and focuses on a specific time period $t$. The full-fledgedness of the type-cohort-and-time specific ATT is useful when the researcher is interested in treatment effect heterogeneity across both time periods and types. Though both dimensions of the treatment effect heterogeneity may be of interest depending on contexts, we focus on an aggregated ATT parameter in this paper, to highlight the treatment effect heterogeneity across types. To aggregate, we take the average of (ref) across $(t,e)$ while maintaining the relative treatment timing $t-e$ fixed: for some $r \geq 0$,
$\beta_r(k)$ is the $r$-times lagged ATT for type $k$. In aggregating $ATT_t(k,e)$, we use the weights discussed in CS,SA. We call $\beta_r(k)$ as `dynamic' type-specific ATT since the parameter compares instantaneous treatment effect and long-run treatment effect for each type $k$, as we compare $\beta_r(k)$ with $r \approx 0$ and $\beta_r(k)$ with $r \gg 0$.
It is straightforward to see that $ATT_t(k,e)$ and $\beta_r(k)$ are identified from the distribution of $\left( \left\lbrace Y_{it} \right\rbrace_{t=-T_0-1}^{T_1-1}, E_i, k_i \right)$ whenever their respective overlap conditions discussed in Section \hyperlink{Aol}{C.1} are satisfied and $\{k_i\}_{i=1}^n$ is observed for a given $T_0$. Though the type assignment $\{k_i\}_{i=1}^n$ is not directly observed from data, the following lemma shows that the treatment effect parameters are identified when $T_0 \to \infty$ and the pretreatment time series shows weak dependence.
Specifics of the overlap condition are discussed in Section \hyperlink{Aol}{C.1} of the Appendix. The proof for Lemma \hyperlink{L1}{1} is found in Section B of the Supplementary Appendix. Lemma \hyperlink{L1}{1} is an identification counterpart of the type classification consistency result in Section \hyperlink{Sec4}{4}: Theorem \hyperlink{T1}{1}.
\hypertarget{Sec3}
The estimation procedure is two-step. The first step is to estimate the type using the $K$-mean minimization problem. The second step is to take the estimated type as given and estimate ATT. To describe the estimation procedure, let us adopt the following notations:
$\gamma$ is a $n \times 1$ vector of a type assignment while $k_i$ denotes the type of unit $i$. $\Gamma$ is a set of all possible type assignments where $n$ units are assigned to $K$ different types. $\delta$ is a collection of $\delta_t(k)$, the type-specific time trend given time $t$ and type $k$: $$ \delta_t(k) = \mathbf{E} \left[ Y_{it}(\infty) - Y_{it-1}(\infty) | k_i=k \right]. $$
In the classification step, we only use a subset of the given data: population pretreatment periods. With the population pretreatment periods, we construct an objective function:
$\mathcal{D}=[-M,M]^{T_0 K}$ with some $M>0$. The minimization problem in (ref) is called $K$-mean minimization problem; the solution to the $K$-means minimization problem is a grouping structure with $K$ groups, defined with $K$ centeroids. In our minimization problem (ref), the centeroids are denoted with $\left\lbrace \delta_t(1) \right\rbrace_{t< 0}, \cdots, \left\lbrace \delta_t(K) \right\rbrace_{t < 0}$ and the grouping structure is denoted with $k_1, \cdots, k_n$.
The algorithm that we use to obtain (ref) is a conventional $K$-means clustering algorithm. Given an initial type assignment ${\gamma}^{(0)} = \left( k_1^{(0)}, \cdots, k_n^{(0)} \right)$,
The iterative algorithm proposed here has two stages. In the first stage, the algorithm estimates $\delta$ by taking sample means. In the second stage, the algorithm reassigns a type for each unit, by finding the type that minimizes the distance between $\{Y_{it}-Y_{it-1}\}_{t<0}$ and $\{\delta_t(k)\}_{t<0}$. The algorithm quickly attains a local minimum of the minimization problem (ref). In the application we used in Section \hyperlink{Sec7}{7}, the algorithm mostly converged within 20 iterations.
Since the iterative algorithm does not conduct an exhaustive search, it may not converge to a global minimum; the computational burden of the exhaustive search is extremely heavy since the space for the type assignment has cardinality of $n^K$. Thus, we recommend that a random initial type assignment be drawn multiple times and the associated local minima be compared.
Another concern is the choice of $K$. So far, the number of types $K$ has been treated as known. When there is no natural choice for $K$, an information criterion can be used to estimate the number of type $K$ (see BN,BM,JS). To construct the information criterion, first we set $K_{\max}$, an upper bound on the number of types. Then, we estimate the error variance using the classification with $K_{\max}$: with $\Gamma(K) = \{1,\cdots, K\}^n$, $$ \hat{\sigma}^2 = \min_{\delta, \gamma \in \Gamma(K_{\max})} \widehat{Q} (\delta, \gamma) $$ Given the estimate, we estimate $K$ to be the minimizer of the Bayesian information criterion: $$ \hat{K} = \arg \min_{K \leq K_{\max}} \left( \min_{\delta, \gamma \in \Gamma(K)} \widehat{Q}(\delta, \gamma) + \hat{\sigma}^2 \frac{K T_0 + n}{nT_0} \log n T_0 \right). $$ See the Supplementary Appendix for more discussion.
Given the first-step result, the type-specific diff-in-diff estimator for $ATT_t(k,e)$ defined in (ref) can be constructed by taking sample means for each type: $$ \widehat{ATT}_t(k,e) = \sum_{i=1}^n \left( Y_{it} -Y_{i,e-1} \right) \left( \frac{ \mathbf{1}\{\hat{k}_i=k,E_i=e\}}{\sum_{i=1}^n \mathbf{1}\{\hat{k}_i=k,E_i=e\}} - \frac{\mathbf{1}\{\hat{k}_i=k,E_i>t\}}{\sum_{i=1}^n \mathbf{1}\{\hat{k}_i=k,E_i>t\}} \right).\footnotemark $$ In the case of the dynamic ATT parameter $\beta_r(k)$ defined in (ref), the estimator is
where $\hat{\mu}(k,e) = \frac{1}{n} \sum_{i=1}^n \mathbf{1}\{\hat{k}_i=k, E_i=e\}$. $\hat{\mu}(k,e)$ is the estimator for the probability $\mu(k,e)= \Pr \left\lbrace k_i=k, E_i=e \right\rbrace$.
\footnotetext{As discussed in CS, there is no straightforward choice in which time differences to use in a diff-in-diff type approach. Though the estimator described above takes one period before the treatment timing to construct a time difference, other choices are equally valid as long as the parallel trend assumption holds for every time period. RSjpem discusses efficiency of these diff-in-diff type estimates when the treatment timing is truly random. More discussion on this is given the \hyperlink{Asec1}{Appendix}. Also, the estimator uses the not-yet-treated units as control units; alternatively, one can use never-treated units or any other subset of $\{i: E_i > t\}$.}
Similarly, we can extend $\hat{\beta}_r(k)$ for $r < -1$ and construct estimators for $$ \beta_r(k) = \sum_{e=0}^{T-1} \frac{\Pr \left\lbrace E_i = e |k_i=k\right\rbrace}{\Pr \left\lbrace E_i \leq T_1 - r |k_i=k\right\rbrace} \cdot \mathbf{E} \left[ Y_{i,e+r}(e) - Y_{i,e+r}(\infty) | k_i=k, E_i=e \right] $$ for some $r < -1$. From Assumptions \hyperlink{A1}{1}-\hyperlink{A2}{2}, $\beta_r(k)= 0$ whenever $r < -1$. Thus, we can use $\hat{\beta}_r(k)$ for $r < -1$ to test Assumptions \hyperlink{A1}{1}-\hyperlink{A2}{2}, equivalent to the widely used `no pretreatment test' in the event-study literature.
\hypertarget{Sec4}
In this section, we discuss the asymptotic properties of the estimator proposed in Section \hyperlink{Sec3}{3}. As we extend the main model and the asymptotic results of this section in Section \hyperlink{Sec5}{5} and Sections \hyperlink{At}{C.2-4} of the Appendix, the proofs are implied by the more general results. The proofs for the general results and how they connect to the asymptotic results of this section are discussed in Sections C-D of the Supplementary Appendix.
Firstly, to derive the classification result for the type estimator defined in (ref), let us adopt following assumptions.
Assumption \hyperlink{A5}{5-c} assumes that the number of population pretreatment periods $T_0$ grows to infinity as $n$ goes to infinity. Assumption \hyperlink{A5}{5-d} assumes that each type realizes with positive probability. Assumption \hyperlink{A5}{5-e} assumes that for $t < 0$, tail probability of $Y_{it}(e) - \mathbf{E} [ Y_{it}(\infty)|k_i]$ goes to zero exponentially and the first difference of $Y_{it}(e) - \mathbf{E} [ Y_{it} (\infty)|k_i]$ is weakly dependent in the sense that it is strongly mixing with mixing coefficient decreasing exponentially in $t$.
Theorem \hyperlink{T1}{1} puts a bound on the misclassification probability; the rate is identical to the rate found in the group fixed-effect literature.
The classification of $n$ units into $K$ types is a crucial part of the estimation procedure that the performance of the treatment effect estimators depends on. Consider a very simple case where $K=2$ and model the untreated potential outcomes as follows: for $t \leq 0$, $$ Y_{it}(\infty) = \delta(k_i) + U_{it}, \hspace{8mm} U_{it} \overset{iid}{\sim} \mathcal{N} (0,1). $$ WLOG let $\delta(1) < \delta(2)$. Find that $\bar{Y}_i(\infty) = \frac{1}{T_0+1} \sum_{t=-T_0-1}^{-1} Y_{it}(\infty) \sim \mathcal{N} \left( \delta(k_i), \frac{1}{T_0+1} \right)$. It is easy to see that for any fixed $T_0$,
is nonzero, with $\Phi$ being the distribution function of the standard normal $\mathcal{N}(0,1)$; the probability of imperfect classification is nonzero. Thus, we need large pretreatment periods $(\Leftrightarrow T_0 \gg 0)$, in addition to the strong separation $(\Leftrightarrow \delta(2) -\delta(1) > 0 )$.\footnotemark \footnotetext{By evaluating the CDF function $\Phi$, we can see that ${T_0}^{\nu} \Phi \left( \sqrt{\frac{T_0+1}{2}}\left( \delta(2) - \delta(1) \right) \right)$ goes to zero for any $\nu > 0$ as $T_0$ grows, as stated in Theorem \hyperlink{T1}{1}.} When we do not have both conditions satisfied and thus units are potentially misclassified, the treatment effect estimator suffers from a non-classical measurement error problem.\footnotemark \footnotetext{The measurement error problem may be severe for a type-specific ATT even when the misclassification rate is small and the tail of the potential outcome distribution is thin, if the within-type overlap is weak. On the other hand, when the misclassification rate is negligible compared to the within-type shares of treated units and control units, the bias will be small.}
Given the long pretreatment, the bound on the misclassification probability from Theorem \hyperlink{T1}{1} can be used to derive asymptotic properties of the type-specific diff-in-diff estimator from (ref).
Remark 1. The asymptotic variance has a consistent estimator, whose expression is given in the Supplementary Appendix, along with the proof of Corollary \hyperlink{C3}{3}.
Remark 2. In formulating the dynamic ATT parameter $\beta_r(k)$, treatment timing distribution is used as weights. Similar asymptotic results as in Corollary \hyperlink{C1}{1} hold for many other choices of weights: e.g. uniform weights across treatment timing.
\hypertarget{Sec5}
In this section, we extend our main model by adding observed covariates $X_{it} \in \mathbb{R}^p$ to the model. The control covariates $X_{it}$ gives us an extra source of heterogeneity in outcomes across different units and different times. For the classification to be successful, we need to decompose the variation in the outcome variable into the variation from the control covariates $X_{it}$ and the variation from the latent type $k_i$. For that end, we assume the following linear model for untreated potential outcome: for $t < 0$,
Note that the interpretation of $\delta_t(k)$ is changed. In the linear model (ref), $\delta_t(k)$ is not the conditional mean of first-differenced potential outcome anymore since there exists ${X_{it}}^\intercal \theta$. Thus, we call $\delta_t(k)$ the type-specific time fixed-effects. The type-specific time fixed-effects explains heterogeneity across units that is not explained by the (linear) observable control covariates $X_{it}$.
Given the model (ref), we can construct a similar objective function from before and solve the $K$-means minimization problem for classification: $$ \left( \theta, \delta, \gamma \right) = \arg \min_{\theta, \delta, \gamma} \frac{1}{nT_0} \sum_{i=1}^n \sum_{t=-T_0}^{-1} \Big( Y_{it} - Y_{it-1} - \delta_t(k_i) - {X_{it}}^\intercal \theta \Big)^2. $$ The objective function includes $X_{it}$. Given an initial type assignment ${\gamma}^{(0)} = \left( k_1^{(0)}, \cdots, k_N^{(0)} \right)$,
In Appendix, we discuss Assumption \hyperlink{A7}{7} which extends Assumptions \hyperlink{separation}{4}-\hyperlink{A5}{5}. Under Assumption \hyperlink{A7}{7}, we have the following classification result.
Remark 3. When $X_{it}$ is time-invariant, i.e. $X_{it} = X_i$, the linear model (ref) and Assumption \hyperlink{A7}{7} can be understood as a parametric special case of the conditional parallel trend assumption: for $t < 0$, $$ \mathbf{E} \left[ Y_{it} (E_i) - Y_{it-1} (E_i) | k_i, X_i \right] = \delta_t(k_i) + {X_{i}}^\intercal \theta. $$
Remark 4. Instead of assuming a linear structure on the first difference as in (ref), we can consider a linear model on the level of the outcome:
The assumptions for the linear model in level is discussed in the Appendix along with Assumption \hyperlink{A7}{7}.
Theorem \hyperlink{T2}{2} finds the same rate on the misclassification probability as Theorem \hyperlink{T1}{1}. The key part of the proof utilizes the linear separability of $k_i$ and $X_{it}$. The proof firstly shows that $\theta$ is consistently estimated. Then, $Y_{it} - Y_{it-1} - {X_{it}}^\intercal \hat{\theta}$ is sufficiently close to $Y_{it} - Y_{it-1} - {X_{it}}^\intercal \theta$ so that the classification using $\hat{\theta}$ and the one using the true $\theta$ are the same.
Theorem \hyperlink{T2}{2} implies that we can take the estimated types as given and apply the available treatment effect estimation methods when the rate given in Corollary \hyperlink{C1}{1} is satisfied. There are largely two ways to incorporate the control covariate $X_{it}$ in the treatment effect estimation. Firstly, we can follow an outcome model approach and impose a parametric model for the post-treatment outcome as we do for the pretreatment outcome in (ref). Given the parametric model, we plug in the estimated types as true types and estimate the model. A large variety of parametric models with a finite grouping structure can be used for the outcome model approach. For example, we can assume a linear model with type-specific coefficients for the treatment effect: for $t \geq 0$,
A more discussion on the outcome model approach is discussed in Section \hyperlink{Aom}{C.3} of the Appendix. The outcome model approach has the merit of developing a parsimonious model for treatment effect; using the outcome model approach, we can impose some structure over how the latent type $k_i$ and the observable information $X_{it}$, $E_i$ interact in terms of the treatment effect heterogeneity.
Alternatively, we can use an assignment model approach. In the assignment model approach, instead of imposing a parametric model for the post-treatment outcomes, we impose a parametric model for the treatment timing. Suppose that we are given a time-invariant control covariate $X_i$ and that the conditional parallel trend assumption holds with $X_i$: for every $t, s \geq -1$, $$ \mathbf{E} \left[ Y_{it}(\infty) - Y_{is}(\infty) | k_i, X_i, E_i \right] = \mathbf{E} \left[ Y_{it}(\infty) - Y_{is} (\infty) | k_i, X_i \right]. $$ Then, we can apply the results of CS by assuming an assignment model and estimating the propensity to be treated given the type $k_i$ and the control covariate $X_i$. For example, when $E_i \in \{0,\infty\}$, the logistic model can be used: $$ \Pr \left\lbrace E_i = 0 | k_i, X_i \right\rbrace = \frac{\exp \left( {X_i}^\intercal \theta + \delta(k_i) \right) }{1+\exp \left( {X_i}^\intercal \theta + \delta(k_i) \right)}. $$ The benefit of the assignment model approach is that we allow for flexible interaction between the observable control covariates $X_i$ and the latent type $k_i$, in terms of the treatment effect heterogeneity. The assignment model approach is an extension of Corollary \hyperlink{C1}{1} since the type-specific diff-in-diff estimator defined in Section \hyperlink{Sec2}{2} is what we get when we assume the propensity score to be a trivial function of $X_i$: $\Pr \left\lbrace E_i = e | k_i, X_i \right\rbrace = \Pr \left\lbrace E_i = e | k_i \right\rbrace$. A more discussion on the assignment model approach is discussed in Section \hyperlink{Aam}{C.4}.
\hypertarget{Sec6}
In this section, we present simulation results to discuss the finite-sample performance of the type-specific diff-in-diff estimator, compared to some existing estimators in the literature. For that, we constructed a random sample using the following data generating process: for $t = -T_0-1, \cdots, 0$,
$D_i, \alpha_i, U_{i,-T_0-1}, \{V_{it}\}_{t \leq 0}$ are mutually independent given $k_i$. $D_i \big| k_i \sim \text{Bernoulli} \big(\pi(k_i) \big)$ and
The values of the DGP parameters that pertain the classification step are taken from the empirical moments of the dataset used in the next section: $\sigma = 1.85$ and $\rho = 0.60$ for the error distribution and $\min_{k \neq k'} | \delta(k) - \delta(k') | = 1.32$ for the type separation.\footnotemark \footnotetext{The rest of the simulation parameters are as follows. For a two-types DGP where $K=2$, we set $\left( \pi(1), \pi(2) \right) = \left(1/3, 2/3 \right)$, $\left( \alpha(1), \alpha(2) \right) = \left(37,39\right)$, $\left( \delta(1), \delta(2) \right) = \left( 1.66, 0 \right)$ and $\left( \beta(1), \beta(2) \right) = \left( 4, 1 \right)$. $\Pr \left\lbrace k_i = 1 \right\rbrace = \Pr \left\lbrace k_i = 2 \right\rbrace=1/2$. For a three-types DGP where $K=3$, we set $\left( \pi(1), \pi(2), \pi(3) \right) = \left(1/3, 1/2, 1/2 \right)$, $\left( \alpha(1), \alpha(2), \alpha(3) \right) = \left(37,39,35 \right)$, $\left( \delta(1), \delta(2), \delta(3) \right) = \left( 2.74, 1.42, 0 \right)$ and $\left( \beta(1), \beta(2), \beta(3) \right) = \left( 5, 1, 0 \right)$. $\Pr \left\lbrace k_i = 1 \right\rbrace = \Pr \left\lbrace k_i = 2 \right\rbrace = 2/5$ and $\Pr \left\lbrace k_i = 3 \right\rbrace = 1/5$. The parameters concerning classification\textemdash $\beta, \sigma, \rho$\textemdash are taken from the dataset used in Section \hyperlink{Sec7}{7}.} Note that a simple mean comparison of the treated units and untreated units is a biased estimator for the treatment effect when $\pi(k)$ is not constant in $k$.
In the classification step, two different specifications for the type-specific time trend $\delta_t(k)$ were used. Firstly, we used the most flexible specification where $\delta_t(k)$ is allowed to vary across every $t$: $\{\delta_t(k)\}_{t\leq 0}$. Secondly, we assumed that the researcher has a priori knowledge on the DGP and imposes a constant slope restriction $\delta_t(k) = \delta_{t'}(k)$ for every $t,t'$, estimating only one parameter for each type: $\delta(k)$. Given the type classifications, we estimated the ATT using the type-specific diff-in-diff estimators; for comparison, we considered the unconditional diff-in-diff, the synthetic control and the synthetic diff-in-diff estimators.
Table (ref) and Table (ref) contain the simulated bias and the simulated MSE from 500 random samples. Also, they contain the type classification success rate: the first and the second columns of the bottom panel denote probability of a random sample having less than 5% misclassification and the third and the fourth columns denote probability of perfect classification. For large $T_0=20,30$, both the type-specific diff-in-diff estimator and the synthetic diff-in-diff estimator perform well since there are many pretreatment outcomes to be used to control for the unit-level heterogeneity. However, in the case of the two-types DGP, the type-specific diff-in-diff estimator outperforms the other estimators for small $T_0=10$ since it best reflects the finite type structure in dataset. As for the classification, long pretreatments ($T_0=20, 30$) always achieves perfect classification, except when no smoothness restriction is imposed to the three-types DGP; even in that case, perfect classification probability is 0.80 for $T_0=20$ and 0.94 for $T_0=30$.
In addition to the simulation specifications discussed above, we consider two additional DGPs to discuss misspecification in classification: firstly, we let $K=5$; secondly, we let the latent type be continuous. For both DGPs, we applied the two-types type-specific diff-in-diff estimator; the first DGP is a case of a misspecified number of types and the second DGP is a case of weakly separated types. The type-specific diff-in-diff estimator is more robust to the first type of misspecification, where the types are still strongly separated: for more discussion, see Section A.2.3. of the Supplementary Appendix.
\hypertarget{Sec7}
To show how the type-specific diff-in-diff estimator applies to a real dataset, we revisit L. Since the Supreme Court ruling on Brown v. Board of Education of Topeka in 1954 that found state laws in the United States enabling racial segregation in public schools unconstitutional, various efforts have been made to desegregate public schools, including court-ordered desegregation plans. After several decades, another important Supreme Court case was made in 1991; the ruling on Board of Education of Oklahoma City v. Dowell in 1991 stated that school districts could terminate the court-ordered plans once it successfully removed the effects of the segregation. Since the second Supreme Court ruling, school districts started to file for dismissal of court-ordered desegregation plans, mostly in southern states.
L used the variation in timing of the district court rulings on the desegregation plan to estimate the effect on racial composition and education outcomes in public schools. The paper uses annual data on mid- and large-sized school districts from 1987 to 2006, obtained from the Common Core of Data (CCD), which contains data on school districts from 1987 to 2006, and the School District Databook (SDDB) of the US census, which contains data on school districts in 1990 and in 2000. To document if a school district was under a court-ordered desegregation plan at the time of the Supreme Court ruling in 1991 and when and if the school district got the desegregation plan dismissed at the district courts, L collected data from various published and unpublished sources, including a survey by Rosell and Armor (1996) and the Harvard Civil Rights Project.
Though L looks at several outcome variable, we focus on one outcome variable, the dissimilarity index: the dissimilarity index for school district $i$ is
The dissimilarity index ranges from 0 to 100, with 100 being perfectly segregated schools and 0 being perfectly representative schools.\footnotemark \footnotetext{In L, the dissimilarity index ranges from zero to one but we rescaled the index for more visibility.}
We followed the data cleaning process in the paper and chose the timespan of 1988-2007 to form a balanced panel of school districts that were under a court-ordered desegregation plan in 1988-1999, which gave us 50 school districts. We use the following linear model for the pretreatment outcomes: for $t=1989,\cdots,1999$,$$ Y_{it} (E_i) - Y_{it-1}(E_i) = \delta_t(k_i) + {X_{it}}^\intercal \theta + U_{it}. $$ The effective number of pretreatment outcomes is 11. The control covariates $X_{it}$ contain a central city indicator variable, percentage of students who are white, percentage of students who are hispanic, percentage of students with free/reduced-priced lunch and number of students. For the purpose of comparison, here we present the main empirical specification of L:
Though two specifications look alike, there are some differences. Firstly, though L and we use the same control covariates, L only used their values from the first year, with time-varying coefficient $\theta_t$: $X_i = X_{i,-T_0-1}$.\footnotemark \footnotetext{Also, L used three additional variables: squared number of students, cubed number of students and squared percentage of students with free/reduced-price lunch.} On the other hand, we use time-varying control covariates $X_{it}$, with time-invariant coefficient $\theta$. Secondly, L uses time fixed-effects $\delta_{jt}$ based on census region, which assigns every school district into one of the four regions. In the terminology of the model used in this paper, L took the census region as the true type assignment whereas we used the data to estimate the type assignment. Lastly, the regression specification in L has a dynamic treatment effect $\beta_r$ whereas we only impose linearity on pretreatment outcomes and therefore do not have any treatment effect term. Since $n=50$ is relatively small, we imposed an additional smoothness restriction on the type-specific time fixed-effects: $\delta_t(k) = \delta(k)$.\footnotemark \footnotetext{The constant slope restriction was chosen out of five specifications\textemdash constant slope, constant slope with a break, constant slope with two breaks, linear, linear with one break\textemdash, based on cross-validated mean-squared forecasting error. For more discussion, see the Supplementary Appendix.} Then, we applied the $K$-means clustering classifier with $K=2$.\footnotemark \footnotetext{As robustness check, we also considered $K=3$ and $K=4$. The Bayesian information criterion selected $K=3$ and the qualitative result remains the same for both $K=2$ and $K=3$. For more discussion, see Supplementary Appendix.} The first-step classification assigns 8 treated units and 14 never-treated units to Type 1 and 13 treated units and 15 never-treated units to Type 2; Type 2 school districts are slightly more likely to be treated.
Given the first-step classification result, we conducted a within-type balancedness test; Table (ref) contains within-type balancedenss test results using control covariates from $t=1988$. Within the two types, the control covariates are well-balanced across treatment status: treated v. never-treated. Thus, we did not use $X_{it}$ in the estimation step and use the unconditional type-specific diff-in-diff estimator as defined in (ref); the target parameter captures the type-specific dynamic ATT that assigns equal weights to each treated unit within the same type.
Figure (ref) contains the type-specific diff-in-diff estimates for Type 1 and Type 2 school districts. From Figure (ref), we see that the treatment effect is bigger for Type 1 and smaller for Type 2; the termination of court-ordered desegregation plans exacerbated racial segregation more severely for Type 1. The type-specific estimation provide a sensible aggregate estimate; averaging the type-specific diff-in-diff estimates across types gives us $(8 \hat{\beta}_r(1) + 13 \hat{\beta}_r(2))/21 = 4.37$ at $r=4$ whereas the pooled regression estimate from L is around 4-5 at $r=4$, depending on specifications, and the conditional diff-in-diff estimate from CS is 4.85 at $r=4$. For reference on the magnitude, the mean of the dissimilarity index was 34 and its standard deviation was around 16 in 1988. Also, Figure (ref) contains estimates for $\beta_r(k)$ such that $r < 0$; none of the pretreatment trend was found to be away from zero at 0.05 significance level.
So, estimates on treatment effect suggest that Type 1 and Type 2 are different; the Type 1 school districts are more responsive to the treatment. How are these types different in other regards? Firstly, Table (ref) shows us some descriptive statistics on the outcome variable and other control covariates for each type, using year 1988 data. The null hypothesis that the entire vector of mean differences between Type 1 and Type 2 is zero is rejected with a $t$-test at size 0.05; the Type 1 school districts are different from the Type 2 school districts in terms of their observable characteristics. For instance, Type 1 school districts have higher proportion of white students and lower proportion of hispanic students. Secondly, in terms of the unobserved heterogeneity captured by the latent type variable, Type 1 has seen a steeper increase in the dissimilarity index while the slope was smaller for Type 2: $\Big( \hat{\delta}(1), \hat{\delta}(2) \Big) = (3.59, 1.93)$. This implies that the dismissal of desegregation plans had a bigger impact on Type 1, where the dissimilarity index was already rising faster. This observation presents future research questions: for example, why do the school districts that were getting more segregated also get affected more from the dismissal of the desegregation plan?
In this paper, we introduce a type-specific parallel trend assumption in a panel data model with a latent type. By assuming the latent type variable has a finite support and is well-separated in the long pretreatment time series, the $K$-means classifier estimates the true types consistently. Also, based on the estimated types, we estimate the type-specific treatment effect. The type-specific diff-in-diff estimator is useful when we suspect heterogeneity in time trends across units and want to explore the associated treatment effect heterogeneity. By applying the estimation method to an empirical application, we find some interesting empirical results where the estimates on the type-specific treatment effects and those on the type-specific time trend tell a story: the effect of terminating court-mandated desegregation plans were bigger for school districts where the dissimilarity index was growing.