EconBase
← Back to paper

The Chained Difference-in-Differences

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

102,205 characters · 17 sections · 80 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The Chained Difference-in-Differences

\newgeometry{top=1cm}

abstractThis paper studies the identification, estimation, and inference of long-term (binary) treatment effect parameters when balanced panel data is not available, or consists of only a subset of the available data. We develop a new estimator: the chained difference-in-differences, which leverages the overlapping structure of many unbalanced panel data sets. This approach consists in aggregating a collection of short-term treatment effects estimated on multiple incomplete panels. Our estimator accommodates (1) multiple time periods, (2) variation in treatment timing, (3) treatment effect heterogeneity, (4) general missing data patterns, and (5) sample selection on observables. We establish the asymptotic properties of the proposed estimator and discuss identification and efficiency gains in comparison to existing methods. Finally, we illustrate its relevance through (i) numerical simulations, and (ii) an application about the effects of an innovation policy in France. \\ Keywords: Event study, Unbalanced panel, Attrition, Treatment effect heterogeneity, GMM. JEL Codes: C14 C2 C23 O3 J38.

\thispagestyle{empty}

\newgeometry{top=4cm}

\parskip=5pt {0pt}

Introduction

Many public policies take years before having effects on their targeted outcomes. For instance, most innovation policies have long run objectives such as making new discoveries, enhancing knowledge, and increasing technology development, but induce only little innovation in the short-run bakker_2013,oconnor_etal_2013,gross_2018. Measuring these long-term effects is however difficult for two principal reasons. First, identification of treatment effects from observational data raises significant challenges, whereas randomized controlled experiments are often too costly, or raise ethical concerns. Second, treatment effects are frequently estimated from panel survey data, where the subjects of interest (e.g. individuals or firms) are not consistently observed over the entire time frame. This problem can be caused by attrition hausman1979attrition, if individuals drop out of the sample, or because of the survey design itself, if individuals are frequently replaced to prevent attrition or over-soliciting respondents nijman1991efficiency.

In this paper, we study the identification, estimation, and inference of long-term treatment effects in settings where balanced panel data is not available, or consists of only a subset of the available data. We develop a new estimator, the chained difference-in-differences (DiD), which leverages the overlapping structure of many unbalanced panel data sets baltagi_song_2006. The intuition behind the chained DiD estimator is simple. Letting $Y_t$ be an outcome of interest observed in three distinct time periods $t_0<t_1<t_2$, the long difference $Y_{t_2}-Y_{t_0}$ can always be decomposed into a sum of short differences $(Y_{t_2}-Y_{t_1})+(Y_{t_1}-Y_{t_0})$. Our estimator generalizes this simple idea by optimally aggregating short-term treatment effect parameters, or “chain links”, obtained from (possibly many) overlapping incomplete panels. For instance, one subsample of individuals might be used to estimate the DiD from $t_0$ to $t_1$ while another would be used to estimate the DiD from $t_1$ to $t_2$. The sum of both effects identifies the long-term DiD from $t_0$ to $t_2$, which may not be feasible or may suffer from efficiency losses if a large part of the sample is discarded by using only a balanced subsample. Our estimator is designed for multiple time periods, it accommodates variations in treatment timing, treatment effect heterogeneity, general missing data patterns, sample selection on observables, and it may deliver substantial efficiency gains compared to estimators using only subsets of the data or treating the unbalanced panel as repeated cross-sectional data.

Building upon the influential work of callaway2021difference, we show that our estimator is consistent, asymptotically normal, and computationally simple. Their multiplier bootstrap and pre-trend tests remain asymptotically valid in our setting. We identify three main advantages of our approach: (1) it does not require having a balanced panel subsample as it is the case with a standard DiD; (2) identifying assumption allows for sample selection on time-persistent unobservable (or latent) factors, unlike the cross-section DiD which treats the sample as repeated cross-sectional data; and (3) it may also deliver efficiency gains compared to existing methods, notably when the outcome variable is highly time-persistent.

The proposed approach is especially relevant with regard to the significant interest for long-term evaluations of interventions, with many applications related to education and labor economics angrist2006long,kahn_2010,oreopoulos_etal_2012,autor_houseman_2010,garcia_Perez_etal_2020,lechner2011long,mroz_savage_2006,stevens_1997. The estimation of treatment effects generally consists in performing a DiD focusing only on units that are observed over the entire time frame: the balanced subsample. This approach, hereafter referred to as the long DiD, allows getting rid of both time- and individual-specific unobservable heterogeneity, but may also involve discarding many individuals with missing observations when estimating long-term effects. The evaluation of long-term effects is sometimes difficult, if not infeasible, because of such missing data problems. Missing observations in panel data sets may exist for two principal reasons: (1) by design or (2) because of attrition.

First, the design of rotating panel surveys attempts to alleviate the burden of administering and responding to statistical surveys, and to prevent attrition, by replacing subjects regularly heshmati1998efficiency. The survey is administered to a cohort of subjects in a limited number of periods. The cohort is then replaced by another cohort randomly drawn from the population of interest. The resulting data set is hence composed of a collection of incomplete panels. The data is said to have an overlapping structure if there is at least one period in which two separate cohorts are administered the survey.

The most famous example of rotating panel survey is the Current Population Survey (CPS), where each cohort is interviewed for a total of 8 months (over a period of 16 months), and part of the sample is replaced each month by a new subsample.\footnote{The design is detailed at \url{https://www.bls.gov/opub/hom/cps/design.htm}.} This survey is one of the most widely used data sources in economic and social research. It has been used or cited in 1000+ articles between 2000 and 2022, including publications in top journals such as the American Economic Review, Journal of Political Economy, Quarterly Journal of Economics, American Sociological Review, and Demography.\footnote{This figure is based on a Google Scholar search conducted on December 5, 2022.} Other rotating panel surveys include, but are not limited to, the Medical Expenditure Panel Survey, the General Social Survey since 2006, the Consumer Expenditure Survey blundell_etal_2008 in the U.S., the Labour Force Survey in the U.K., and the Enqu\^{e}te Emploi (Labour Force Survey) and Enqu\^{e}te sur les Moyens Consacr\'{e}s \`{a} la R&D (R&D Survey) in France.

Despite their prevalence, the econometrics literature on rotating panels is almost non-existent baltagi_song_2006. heshmati1998efficiency uses rotating panels for production function estimation. nijman1991efficiency study the optimal choice of the rotation period for estimating a linear combination of period means. Likewise, our estimator consists in a linear combination of parameters corresponding to the long-term average treatment effect parameter. To the best of our knowledge, our paper is the first to study identification, estimation, and inference of treatment effects in this context.

Second, individuals may also drop out of the surveys and cause attrition, which can be particularly severe for long panels. This attrition raises some concerns because it is associated with selection. Attrition can be due to either “ignorable” or “non-ignorable” selection rules verbeek1996incomplete. Ignorable attrition implies that missing data occurs completely at random, and as such focusing on a balanced panel subsample does not threaten identification. verbeek1992testing, among others, propose a test of the ignorability assumption. Non-ignorable attrition means that missing data is related to either observable or unobservable factors. hirano2001combining provide important identification results for the additive non-ignorable class of attrition models, which nest several well-known methods hausman1979attrition,little1989analysis. These results have been extended to multi-periods panels by hoonhout2019nonignorable and apply to our setting, in particular we consider their Sequential Missing At Random assumption. bhattacharya2008inference studies the properties of a sieves-based semi-parametric estimator of the attrition function initially proposed by hirano2001combining. These methods are often referred to as models of selection on unobservables because the probability of attrition is allowed to depend on variables that are not observed when an individual drops out, but not on unobservable error terms moffit1999sample. Inverse propensity weighting is the most popular method for addressing non-ignorable attrition caused by observable factors, even beyond the estimation of average treatment effects chaudhuri2018indirect. There also exist methods using instrumental variables to address attrition due to latent factors frolich2014treatment. Other approaches focus on particular structures of attrition, like monotonically missing data chaudhuri2020efficiency,barnwell2021note.

Attrition is generally addressed before estimating treatment effects with the long DiD, either by reweighting the observations in the balanced subsample with inverse propensity scores hirano2003efficient, or by imputing the missing observations hirano1998combining. However, discarding a possibly large proportion of the data by focusing on the balanced subsample may lead to significant efficiency losses baltagi_song_2006,chaudhuri2020efficiency,barnwell2021note, or may lead to discarding the complete dataset, as illustrated in our application using the French R&D survey. If there are too few individuals observed over the entire time horizon, the data is typically treated as repeated cross-sections and treatment effects are estimated with the cross-section DiD abadie_2005,callaway2021difference.\footnote{This approach consists in taking the difference of differences of averages, where different sets of units are used to compute each of the four averages.}

The identification of the average treatment effect using the cross-section DiD requires not only a parallel trends assumption on the population, but also that the sampling process be (conditionally) independent of the levels of the outcome variable. This assumption is violated if, for instance, treated units with larger unobserved individual shocks are relatively more likely to be sampled in later periods. In such a case, the treatment groups observed early on will differ in unobservable ways from those observed later on. An identification problem arises as soon as the sampling process does not affect the observed control groups in the exact same way. Instead, our approach requires that the sampling process be (conditionally) independent of the trend of the outcomes, because the chained DiD allows eliminating individual heterogeneity like the long DiD before taking expectations. Therefore, the chained DiD is robust to some forms of attrition caused by unobservable heterogeneity, unlike the cross-section DiD. Remark that these two identifying assumptions are non-nested in general. Note further that the attrition models discussed earlier can also be used as a first-step in our framework. Although we do not correct for the efficiency losses of using such first-step plug-in estimates in our paper, we identify some suitable options to do so frazier2017efficient.

The chained DiD rests on a parallel trends assumption, conditioned on being sampled. The parallel trends assumption posits that in the absence of treatment, the trends in potential outcomes for the treated and non-treated groups in the population would have been similar. It offers a weaker alternative to the assumption that the potential outcomes for both treated and untreated groups in the population would have been identical. Extending this logic to unbalanced panel data, our additional assumption pertains to the composition of the sample. If the population is composed of different unobservable types that correlate with treatment assignment and the likelihood of being sampled, then a selection bias may exist. To address this problem, we assume that all unobservable types exhibit similar trends in outcomes conditional on their treatment status (and possibly other covariates) so that the parallel trend assumption also holds in the sample.

Recent developments in the literature about treatments effects with multiple periods and treatment heterogeneity have revealed that two-way fixed-effects estimators fail to identify the treatment effect parameters of interest in many contexts chaisemartin_haultfoeuille_2020,goodman_bacon_2018,borusiak_jaravel_2017,sun_abraham_2020.\footnote{de2022survey provide an excellent survey of the literature, which we briefly summarize in Section (ref).} Our paper builds upon callaway2021difference which address this issue by generalizing the approach developed in abadie_2005 to multiple periods and varying treatment timing. We contribute further to this literature by extending this framework to settings with incomplete panel data. An alternative approach would have been to adapt the general framework developed by chaisemartin_haultfoeuille_2020 to our setting, or the other related papers focused on staggered adoption designs and event studies with multiple periods. However, the chained DiD fits very well within the framework developed by callaway2021difference which focuses on estimating all group-time average treatment effects before aggregating all those parameters into summary parameters of interest. Likewise, our approach focus on (even smaller) building blocks: one-period-difference group-time average treatment effects, which measure the increase of average treatment effect of group $g$ from period $t-1$ to period $t$. In the general case, these blocks correspond to k-period-difference group-time average treatment effects. We also discuss regression alternatives inspired by borusiak_jaravel_2017 and wooldridge2021two in the paper.

We illustrate the performance of the chained DiD in two ways. First, we use simulations to compare the long DiD, chained DiD and cross-section DiD in terms of bias and variance under several data generating processes (DGP). Our simulations include a stratified panel data set, composed of a balanced panel and a rotating panel. Second, we study the long-term employment effects of a large-scale innovation policy in France giving grants to collaborative R&D projects. Technical progress and innovation, stimulated by R&D activities, are key to economic growth scherer_1982,aghion_howitt_1998,griffith_etal_2004 but firms tend to invest too little because of the public good nature of innovations arrow_1962,nelson_1959. Therefore, measuring the long term effects of subsidized R&D is important to inform policymakers and improve future policies.

This application is especially relevant because the French R&D survey, which contains firm-level data on R&D activities, consists of rotating panels but does not include a balanced subsample except for the largest firms. In addition, it is possible to fully observe a limited number of variables for all firms close to those provided by the R&D survey by using administrative data. We are thus able to compare the results of each of the three estimators by focusing on two of those variables: total employment and highly qualified workforce. The long DiD is applied to the complete data and serves as the benchmark estimates. The chained DiD and cross-section DiD estimators are applied to both data sets, and to an “artificial” unbalanced panel which is generated by discarding all observations from the complete administrative data set which are missing in the R&D survey. This application is somehow comparable to a simulation exercise but uses real data.

Our results show that the policy had a positive effect on employment for firms that received a grant to participate to a collaborative R&D project. We find that the estimates as well as their standard errors obtained with the chained DiD estimator are close to those obtained with the long DiD. In contrast, the cross-section estimator delivers biased estimates which also lack sufficient precision to detect any statistically significant effect associated with the policy.

The remainder of the paper is organized as follows. Section (ref) presents our methodology and asymptotic results. Numerical simulations are in Section (ref). Section (ref) contains the application to R&D policy. Section (ref) concludes the paper.

Identification, Estimation and Inference

Basic Framework

We first present the main insights in a simple framework. The main notation is as follows. There are $\mathcal{T}$ periods and each particular time period is denoted by $t = 1,...,\mathcal{T}$. In a standard DiD setup, $\mathcal{T}=2$, no one is treated in $t=1$, and all treatments take place in $t=2=\mathcal{T}$. To gain intuition, we first focus on the case where all treatments take place in $t=2$ but assume $\mathcal{T} > 2$.

Define $G$ to be a binary variable equal to one if an individual is in the treatment group, and $C=1-G$ as a binary variable equal to one for individuals in the control group. Also, define $D_t$ to be a binary variable equal to one if an individual is treated in $t$, and equal to zero otherwise. Let $Y_t(0)$ denote the potential outcome at time $t$ without the treatment and $Y_t(1)$ denote its counterpart with the treatment. The realized outcome in each period can be written as $Y_t = D_t Y_t(1) + (1-D_t)Y_t(0)$.

We focus on the long-term average treatment effect on the treated which corresponds to the average treatment effect in period $t>2$ on individuals in the treatment group, hence first treated in period 2. It is formally defined by

equation[equation omitted — 62 chars of source]

The identification of $ATT(t)$ with panel data has attracted much research attention. In this paper, we are mainly interested in a case where balanced panel data is not available.

\paragraph{The missing data pattern.}

We assume that a new random sample of $n_t$ individuals is drawn at each period $t$ to replace a subsample of previously observed individuals. This replacement can occur due to attrition or by design. For the moment, we assume that subsamples may differ in size $n_t$ but individuals are only observed for two consecutive periods as in many rotating panel design.\footnote{In practice, rotating panels may involve subgroups sampled over a different number of consecutive periods. The result of this paper continues to apply, so the same estimator and inference procedure can be used even with a more sophisticated sampling process as presented in Section (ref).} Therefore, one cannot observe the entire path $\{Y_1,Y_2,...,Y_{\mathcal{T}}\}$ for any individual. The key feature of our approach is that there is some overlap across subsamples, that is there are at least two different subsamples observed in each period $1<t< \mathcal{T}$.

This structure is in stark contrast with the literature about attrition in panel data, which typically assumes that there always exists a balanced subsample where individuals are observed throughout the entire time frame hirano2001combining,hoonhout2019nonignorable. Although there is no need for such a balanced subsample here, we extend our method to more general settings in Section (ref). This extension is illustrated in Sections (ref) and (ref).

To characterize the sampling process, we define $S_t$ to be a binary variable equal to one for individuals observed at $t$ and zero otherwise, and use $S_{t,t+1}$ to denote $S_tS_{t+1}$ which indicates if an individual is observed at both $t$ and $t+1$. The missing data pattern is summarized in Table (ref) for $\mathcal{T}=3$. In the general framework, we will assume that treatments can also vary with some observable covariates $X$, that individuals can be observed in non-consecutive periods, and that sampling probabilities may also be functions of (possibly different) observable covariates or past outcomes.\footnote{In this paper, we do not consider cases where treatment status (or covariates) may also be missing, but relevant solutions already exist for difference-in-differences methods botosaru2018difference.}

table[table omitted — 554 chars of source]

\paragraph{The long DiD.} The common approach to identify treatment effects is to consider the long difference-in-differences defined by

equation[equation omitted — 124 chars of source]

under the standard parallel trend assumption $E[Y_t(0)-Y_1(0)|G=1]=E[Y_t(0)-Y_1(0)|C=1]$. $ATT(t)$ corresponds to the long-term effect for $t>2$. Unfortunately, individuals are never observed more than two consecutive periods in this framework. Calculating the averages of $Y_{it}-Y_{i1}$ for $t>2$ for the treatment and control groups is hence infeasible.

\paragraph{The cross-section DiD.} If the panel consists of incomplete panel data, the identification of the parameter of interest can be achieved by assuming that the sampling process $S_{t}$ is independent of $(Y_{t},D_{1},D_{2},...,D_{\mathcal{T}})$ in addition to the parallel trend assumption abadie_2005. In this case, $ATT(t)$ is identified by the “cross-section DiD” given by

equation*[equation* omitted — 164 chars of source]

The sampling assumption allows replacing the averages of differences, $E[Y_t(0)-Y_1(0)|a=1]$ for $a \in\{G,C \}$, by the difference of averages, e.g. $E[Y_t(0)|S_ta=1]-E[Y_1(0)|S_1a=1]$ for $a \in\{G,C \}$. This approach does not eliminate the individual-specific unobservable heterogeneity.\footnote{This approach is readily implemented in the R/Stata package did/csdid and corresponds to the repeated cross sections version of callaway2021difference.}

\paragraph{The chained DiD.} Our approach takes advantage of the overlapping panel structure. Remark that each term in (ref) can be decomposed into

equation[equation omitted — 175 chars of source]

for $a \in\{G,C \}$. Thus, identification of $ATT(t)$ is obtained by summing the short-term DiD as in

equation*[equation* omitted — 444 chars of source]

where the second equality holds under the assumption that the sampling process $S_{t,t+1}$ is independent of $(Y_{t+1}-Y_{t},D_{1},D_{2},...,D_{\mathcal{T}})$. This approach not only allows eliminating the individual-specific heterogeneity, it also makes use of a different identifying assumption than the one for the cross-section DiD in this setting. Remark that these assumptions are generally non-nested. However, replacement of individuals in panel survey data is often based on some variables observed before the replacement period but not on their evolution per se, as illustrated in our application.\footnote{Surveys are usually not specifically designed to use panel econometric methods.} Our application provides an illustration of that argument.

\paragraph{Simple model.} In order to gain intuition about the identification, estimation and inference of $ATT(t)$ in this context, we suppose that the potential outcome is generated by a components of variance process

equation[equation omitted — 148 chars of source]

where $D_i = (D_{i1},D_{i2},...,D_{i\mathcal{T}})$ denotes the vector of treatment status, $ATT(t) = \sum_{\tau=2}^t \beta_{\tau}$ is the impact of the treatment evaluated at $t$. $\alpha_i$ is an individual-specific time-persistent unobservable component of variance $\sigma_{\alpha}^2$, $\delta_t$ is a time-specific unobservable component, and $\varepsilon_{it}$ is a mean-zero individual-transitory shock that is an auto-regressive error process of order 1 represented by $\varepsilon_{it+1} = \rho\varepsilon_{it} + \eta_{it+1}$ with $\rho \in [0,1]$ and $\eta_{it+1}$ being a white noise of variance $\sigma_{\eta}^2$. Assume further that $D_{it} \protect\mathpalette{\protect\independenT}{\perp} (Y_{it}(0),Y_{it}(1))$, for all $t>2$. These assumptions imply $V(\varepsilon_{it}) = \sigma_{\varepsilon}^2 = \sigma_{\eta}^2/(1-\rho^2)$ if $0\leq\rho <1$, and $V(\varepsilon_{it}) = \sigma^2_{\varepsilon_1} + (t-1) \sigma^2_{\eta}$ if $\rho =1$.

Identification

If the sampling process is independent of all the components of $Y_{it}(D_i)$, then both $ATT_{CD}$ and $ATT_{CS}$ identify the parameter of interest. However, if the unobservable individual-specific component $\alpha_i$ is correlated with the sampling process, $ATT_{CS}$ admits the bias term

equation[equation omitted — 189 chars of source]

Furthermore, if the individual transitory shock $\varepsilon_{it}$ is correlated with the sampling process, $ATT_{CS}$ may also admit the bias term

equation[equation omitted — 221 chars of source]

These bias terms are non-zero if the compositions of the sampled treatment and control groups evolve differently through time, due to either the time-persistent heterogeneity $\alpha_i$ or the time-varying error $\varepsilon_{it}$. The former situation happens, for instance, if individuals with larger $\alpha_i$ are more likely to be sampled in the control group in period 1 but become relatively more likely to be sampled in the treatment group in period $t$. The latter case occurs, for instance, if $\rho>0$ and untreated individuals with larger $\varepsilon_{i\tau-1}$ are more likely to be sampled in period $\tau$. If $\varepsilon_{i\tau-1}$ is highly persistent, the control group observed in later periods will be increasingly composed of individuals with larger past idiosyncratic shocks. Addressing these biases can be difficult, if not impossible, due to the unobservability of these errors.

In contrast, $ATT_{CD}$ cancels out all $\alpha_i$'s, hence it is immune to biases caused by the time-persistent individual heterogeneity like (ref). However, the time-varying errors may not fully cancel out and a bias term related to (ref) could exist, as shown by

equation[equation omitted — 249 chars of source]

if $\rho<1$, and for instance, individuals with larger $\varepsilon_{i\tau-1}$ are more likely to be sampled in period $\tau$. This bias disappears if $\varepsilon_{it}$ follows a random walk ($\rho=1$).

Therefore, if the sampling process depends on unobservable components of the outcomes $Y_{it}$ but is not correlated with the first-differences in outcomes $Y_{it+1}-Y_{it}$ conditional on treatment status, then $ATT(t)$ is still identified by $ATT_{CD}(t)$ but not necessarily by $ATT_{CS}(t)$. This result implies that identification with $ATT_{CD}(t)$ is more robust to some potential dependence between the sampling process and unobservable time-persistent and time-varying factors. ghanem2023selection show that relying on parallel trends assumptions in difference-in-differences studies, with balanced panel data, imply restrictions on selection into treatment and the time-series properties of the outcomes similar to ours.\footnote{Specifically, they identify two key scenarios: (1) selection into treatment based on fixed effects, leading to a stationarity restriction; and (2) selection on time-varying pre-treatment factors, implying a martingale property.}

\paragraph{The R&D survey example.} To illustrate the practical significance of this discussion about identification, let us consider the example of R&D subsidy programs and their impact on firms' labor forces, as explored in our application section. Consider that the firm population is divided into two distinct, unobservable (time-persistent) types: high-quality and low-quality research teams. The parallel trends assumption would require that, in the absence of subsidies, firms with high-quality and low-quality research teams would have followed similar trends in outcomes.

The risk of selection bias emerges if the allocation of subsidies is not random and correlates with unobservable team quality. The long DiD inherently neutralizes this bias if the selection into treatment is solely linked to individual fixed-effects ghanem2023selection. Nonetheless, this selection bias warrants further scrutiny if the propensity of high and low-quality teams to consistently participate in R&D surveys varies, possibly due to differences in resource availability, past experience with surveys, corporate culture, or differing levels of motivation across types. In this context, the cross-section DiD method becomes susceptible to selection bias. However, the chained DiD approach, which compares the changes in outcomes over multiple time periods, remains unaffected by this bias, as long as the variation in survey response rates does not correlate with how outcomes evolve conditional on treatment assignment.

Estimation

A direct estimation method for both $ATT_{CD}(t)$ and $ATT_{CS}(t)$ is to substitute expectations by their sample counterparts in each expression. These Horvitz-Thompson estimators can be written as the weighted averages

equation[equation omitted — 294 chars of source]
equation[equation omitted — 279 chars of source]

where $y_{it}$ denote the outcome variable for individual $i$ in period $t$, and the weights are defined as

equation*[equation* omitted — 273 chars of source]

with $a_i \in\{ G,C \}$, and for all $\tau = 1,...\mathcal{T}-1$. A formal treatment is presented in the general framework. Remark that these two estimators can be written as elementwise products and divisions of matrices.\footnote{For example, the chained DiD can be written in matrix form using the Hadamard product $\odot$ and division $\oslash$ as $\widehat{ATT_{CD}}(t) = \mathbb{1}_n' \left[ \Delta Y \odot \left( W_1\oslash \mathbb{1}_n\mathbb{1}_n'W_1 - W_0\oslash \mathbb{1}_n\mathbb{1}_n'W_0 \right) \right] \mathbb{1}_{t-1},$ where $\mathbb{1}_l$ denotes a $l$-dimensional vector of $1$, $\Delta Y$ is a $n\times (t-1)$ matrix where each row is the difference $(y_{i2},...,y_{it})-(y_{i1},...,y_{it-1})$, $W_1$ a $n\times (t-1)$ matrix with $i,\tau$ element $S_{i\tau\tau+1}G_i$, and similarly for $W_0$.}

In this simple context, estimators of $ATT_{CD}(t)$ and $ATT_{CS}(t)$ can also be obtained by linear regression using a least-squares dummy variables (LSDV) approach as follows.\footnote{We thank a referee for this valuable insight.} The $ATT_{CD}(t)$'s are obtained by estimating the two-way fixed-effect (TWFE) event-study regression model

equation[equation omitted — 257 chars of source]

by defining $\mathbf{1}_{\{.\}}$ as a dummy variable for event $\{.\}$, and $g_i$ as the first period where individual $i$ received the treatment. In this setting, $g_i=2$ for all $i$ in the treatment group and $g_i>\mathcal{T}$ for all $i$ in the control group. This regression makes it explicit that each $\gamma_{l+1}$ corresponds to the effect $l$-periods after the treatment.

If the treatment is randomly assigned, and in the absence of variations in treatment timing, Theorem 3 of de2023two confirms that the parallel trend assumption yields $ E\left[ \hat{\gamma}_{t} \right] = ATT(t),$ with

equation[equation omitted — 194 chars of source]

corresponding to the long DiD estimator, if feasible. In contrast, if the sample is limited to two consecutive observations for each individual then one can show that $\hat{\gamma}_{t}$ corresponds exactly to the chained DiD estimator in (ref).\footnote{Remark, nevertheless, that there is no need to sum the $\hat{\gamma}_{t}$'s because they already correspond to the $ATT(t)$.} Similarly, the cross-section DiD estimator in (ref) can be obtained by substituting the fixed-effects $\sum_{i'=1}^n\alpha_{i'}\mathbf{1}_{\{i=i' \}}$ by a constant term $\alpha$ in (ref).

However, these TWFE event-study regressions are not heterogeneity-robust de2023two. We identify two key limitations: (1) controlling for covariates,\footnote{Theorem S4 of chaisemartin_haultfoeuille_2020 shows that adding a term $X'\theta$ may not deliver unbiased ATT estimates under a conditional parallel trend assumption with a linear functional form for the covariates $X$.} and (2) settings with staggered adoption designs and heterogenous treatment effects.\footnote{Theorem 4 of de2023two shows that the TWFE event-study may be biased if the treatment effect is heterogenous across groups or over time, essentially because it includes comparisons of groups going from untreated to treated to groups treated at both periods--the so-called “forbidden comparison” borusiak_jaravel_2017--,and can be contaminated by the treatment effects at other periods de2023severaltreat.} We will address these issues in the general case for the Horvitz-Thompson estimator (ref), and discuss hetereogeneity-robust regression alternatives borusiak_jaravel_2017,wooldridge2021two.

Inference

In this simple design, the long DiD estimator $\widehat{ATT_{LD}}(t)$ is an efficient estimator of $ATT(t)$ when feasible. We compare the efficiency of the two estimators $\widehat{ATT_{CD}}(t)$ and $\widehat{ATT_{CS}}(t)$ under standard assumptions in Proposition (ref).\footnote{In this paper, 'efficient estimator' refers to one with the lowest variance among all unbiased estimators.}

propositionAssuming potential outcomes to be specified as in (ref), in addition to the standard assumptions of DiD settings as presented in the proof, and a Missing Completely At Random assumption, then $\widehat{ATT_{CD}}(t)$ is efficient for $\rho=1$. In addition, it is more efficient than $\widehat{ATT_{CS}}(t)$ if \begin{equation*} \begin{aligned} \frac{2(t-1)}{1+\rho}\sigma_\eta^2 & \leq \frac{ 1.5\sigma_\eta^2 }{1-\rho^2} + \sigma_{\alpha}^2, \quad for $\rho \in [0,1)$. \end{aligned} \end{equation*}

On the one hand, the chained DiD introduces additional noise by adding up individual-transitory shocks $\varepsilon_{it}$. On the other hand, the cross-section DiD's precision depends on the variance of unobserved individual heterogeneity $\alpha_i$ and serial correlation of individual shocks $\varepsilon_{it}$. It also makes use of twice as many observations to calculate each empirical expectation since two cohorts are observed at each $1<t<\mathcal{T}$, i.e. $t-1,t$ and $t,t+1$. In comparison, the long DiD would not suffer from any of those two efficiency losses.

According to Proposition (ref), $\widehat{ATT_{CD}}(t)$ is efficient under random-walk idiosyncratic errors. harmon2022difference shows this efficiency result of the chained DiD estimator in the case with variations in treatment timing and heterogeneity, under the random walk assumption. In addition, he shows that the methods proposed by borusiak_jaravel_2017 and wooldridge2021two, although efficient under spherical errors, are not efficient under the random walk assumption if there are multiple pre-treatment periods.

Intuitively, $\widehat{ATT_{CD}}(t)$ delivers a more precise estimate only if the sum of the extra individual-transitory shocks has a smaller variance than that of the individual-specific heterogeneity and individual transitory shock. This condition is plausible as long as $t$ is not too large, that is not too many incremental effects are aggregated together, or if $\sigma_\alpha^2$ is relatively large. As $t$ increases, a new term is added to the sum and results in a marginal increase of the variance proportional to $ \frac{2}{1+\rho}\sigma_\eta^2 $. For example, for $\rho=0$, the condition in Proposition (ref) becomes $\sigma_\eta^2/2 \leq \sigma_\alpha^2$ for $t=2$, and $5\sigma_\eta^2/2 \leq \sigma_\alpha^2$ for $t=6$.

Therefore, $\widehat{ATT_{CD}}(t)$ should largely dominate in terms of precision in settings with autocorrelated idiosyncratic shocks, and where the variance of unobserved individual heterogeneity is large.

General framework

In the general framework, we allow for: 1) heterogeneity in treatment effects; 2) variation in treatment timing; 3) sample selection on exogenous observables and past outcomes; and 4) general missing data patterns.

We first introduce additional notations. Let us substitute the treatment group dummy $G$ by $G_g$, a binary variable that is equal to one if an individual is first treated in period $g$. There are hence several cohorts of treatment groups. The control group binary variable $C$ denotes individuals who are never treated. Notice that for each individual $\sum_{g=2}^T G_g +C=1$. Also, define $D_t$ to be a binary variable equal to one if an individual is treated in $t$, and equal to zero otherwise. This variable will be useful to denote if the individual was first treated for some $g \leq t$. Also, define the generalized propensity score as $p_g(X) = P(G_g =1|X,G_g+C=1)$. This score measures the probability of an individual with covariates $X$ to be treated conditional on being in the treated cohort $g$ or the control group.

There are many parameters of interest in this setting, such as, for example, the average treatment effect $k$ periods after the treatment date. All possible parameters consist of aggregates of the most basic parameter: the average treatment effect in period $t$ for a cohort treated in date $g$, denoted by

equation[equation omitted — 56 chars of source]

This parameter is referred to as the group-time average treatment effect in callaway2021difference, and we will build upon their results to study the collection of such parameters. When only unbalanced panel data is available and $t>g+1$, two methods are possible: (1) the cross-section DiD; and (2) the chained DiD.

In what follows, we assume that individuals are sampled only for two consecutive periods and then drop out forever, so the long DiD is never feasible. We will introduce more general missing data patterns in the next subsection.

Rotating panel data

In order to identify the $ATT(g,t)$ and accommodate varying treatment timing and treatment effect heterogeneity on observable covariates $X$, we impose assumptions as follows.

assumption[Sampling] For all $t=1,...,\mathcal{T}-1$, $\left\{ Y_{it},Y_{it+1},X_i,D_{i1},D_{i2},...,D_{i\mathcal{T}} \right\}_{i=1}^{n_t}$ is independent and identically distributed (iid) conditional on $S_{it,t+1} =1 $.
assumption[Missing Trends At Random] For all $t=1,...,\mathcal{T}-1$, \begin{equation} S_{it,t+1} \perp Y_{it+1}-Y_{it},X_i | D_i. \end{equation}
assumption[Conditional Parallel Trends] For all $t=2,...,\mathcal{T}$, $g=2,...,\mathcal{T}$, such that $g \leq t$, \begin{equation} E \left[ Y_{t+1}(0) - Y_{t}(0) | X,G_g=1\right] = E \left[ Y_{t+1}(0) - Y_{t}(0) | X,C=1\right] a.s.. \end{equation}
assumption[Irreversibility of Treatment] For all $t=2,...,\mathcal{T}$, \begin{equation} D_{t-1} =1 implies that D_t = 1. \end{equation}
assumption[Overlap] For all $g=2,...,\mathcal{T}$, $P(G_g=1) > 0$ and for some $\varepsilon > 0$, $p_g(X) < 1 -\varepsilon$ a.s.

Assumption (ref) means that we are considering a rotating panel structure. Conditional on being sampled in two consecutive periods, individuals are assumed to be iid. However, unlike Assumption 3.3 in abadie_2005, it does not imply that observations are representative of the population of interest because we do not assume that the iid draws are taken from the population distribution. Instead, we focus on the case where the identification with a cross-section DiD can fail by considering Assumption (ref), under which $E\left[Y_{it}|S_{it},X_i,D_i\right] = E\left[Y_{it}|X_i,D_i\right]$ may not hold.

Assumption (ref) constitutes the principal departure from callaway2021difference. It states that the sampling process is independent of the first-differences of individual outcomes, and the covariates $X$, conditionally on treatment assignment, i.e. trends are missing at random. In particular, it implies that $E[Y_{it+1}-Y_{it}|X,a=1,S_{t,t+1}=1]$ for $a \in \lbrace G_2,...,C \rbrace$ and $E[Y_{it+1}-Y_{it}|a=1,S_{t,t+1}=1]$ for $a \in \lbrace G_2,...,C \rbrace$ correspond to their population counterpart. Assumption (ref) is used to simplify exposition by ruling out sample selection on observables, apart from treatment status. In practice, we only require conditional mean-independence between $S_{t,t+1}$ and $Y_{t+1}-Y_{t}$. We will present extensions to sample selection on observables shortly.

Assumption (ref) is a key identifying assumption in DiD settings with treatment heterogeneity. It means that the average outcomes for the treatment and control groups, conditional of observables, would have followed parallel paths in absence of the treatment. It is extensively discussed in abadie_2005, callaway2021difference. ghanem2023selection and de2022not provide insightful results about the underlying treatment selection mechanisms.

Assumption (ref) implies that once an individual is first treated, that individual will continue to be treated in the following periods. In other words, there is no exit from the treatment.\footnote{Departures from this assumption are considered in de2022difference.} Finally, Assumption (ref) ensures that there are positive probabilities to belong to the control and treatment groups for any possible value of $X$. Remark that $X$ and $D_i$ are assumed to be observed for all individuals.

\paragraph{Data generating process.} Under Assumption (ref), the data generating process consists of random draws from the mixture distribution $F_M(\cdot)$ defined as\footnote{In the application, we will discuss more complicated situations in which the data is generated by stratified sampling. The same results apply using a suitably reweighted sample wooldridge2010econometric,davezies2009faut.}

equation*[equation* omitted — 197 chars of source]

where $\lambda_{t,t+1}=P(S_{t,t+1} = 1)$ is the sampling probability, $y_t$ and $y_{t+1}$ denote the outcome, respectively, for an individual sampled at $t$ and $t+1$. Expectations under the mixture distribution does not correspond to population expectations. This difference arises because of different sampling probabilities $\lambda_{t,t+1}=P(S_{t,t+1} = 1)$ across time periods and because Assumption (ref) does not preclude from some forms of dependence between the sampling process and the unobservable heterogeneity in $Y_{it}$. We introduce conditioning of $P(S_{t,t+1})$ on observables in Corollary (ref) and (ref). However, this assumption ensures that expectations of first-differences under the mixture correspond to population expectations once conditioned on the time periods. Thereafter, $E_M[\cdot]$ denotes expectations with respect to the mixture distribution $F_M(\cdot)$, its empirical counterpart being the sample mean.

An important result of the paper is given in Theorem (ref). We define the weights

equation*[equation* omitted — 91 chars of source]

and

equation*[equation* omitted — 128 chars of source]
theoremUnder Assumptions (ref) - (ref), and for $2 \leq g \leq t \leq \mathcal{T}$, the long-term average treatment effect in period $t$ is nonparametrically identified, and given by \begin{equation*} \begin{aligned} ATT_{CD}(g,t) = \sum_{\tau=g}^{t}\Delta ATT(g,\tau) , \end{aligned} \end{equation*} where $\Delta ATT(g,\tau) = E_M\left[ w^G_{\tau\tau-1}(g)\left(Y_{\tau} - Y_{\tau-1}\right) \right] - E_M\left[ w^C_{\tau\tau-1}(g,X) \left(Y_{\tau} - Y_{\tau-1}\right) \right]$ is the 1-period-difference group-time average treatment effects, which measure the increase of average treatment effect of group $g$ from period $t-1$ to period $t$. In addition, the cross-section DiD does not identify ATT(g,t) under Assumption (ref).

Those identification results suggest the two-step estimator

equation*[equation* omitted — 274 chars of source]

where

equation*[equation* omitted — 139 chars of source]

and

equation*[equation* omitted — 202 chars of source]

with $\hat{p}_g(\cdot)$ being an estimated parametric propensity score function, such as logit or probit, obtained in a first step.

Let us denote $\widehat{ATT}_{g \leq t}$ the vector of $ \widehat{ATT}_{CD}(g,t)$'s for $g\leq t$. The next theorem establishes its joint limiting distribution.

theoremUnder Assumptions (ref) - (ref) and a standard assumption on the parametric estimates of the propensity scores (Assumption 5 in callaway2021difference or 4.4 in abadie_2005), for all $2 \leq g \leq t \leq \mathcal{T}$, \begin{equation*} \sqrt{n}\left( \widehat{ATT}_{g \leq t}-ATT_{g \leq t} \right) \overset{d}{\to} N(0,\Sigma), \end{equation*} as $n \to \infty$ and where the covariance $\Sigma$ is detailed in the proof.

In the proof of the above theorem, we also show how the multiplier bootstrap procedure proposed by callaway2021difference adapts to this asymptotic result. The main difference comes from redefining the influence function, but the result about its asymptotic validity applies without other modification so it is not repeated here. We also refer the reader to their paper for a complete discussion of summary parameters which use the $ATT(g,t)$ as building-blocks. We provide some results for these summary parameters and an extension to using the not yet treated as the control group in Online Appendices (ref) and (ref).

\paragraph{Sample selection on observables.} We extend these results to address sample selection on observables by considering two alternative MAR assumptions. Let us first relax Assumption (ref) so we have sample selection on $X$ and $D$.

assumption[Missing Trends At Random 2] For all $t=2,...,\mathcal{T}$, \begin{equation} S_{itt+1} \perp Y_{it+1}-Y_{it} | X_i, D_i. \end{equation}

Assumption (ref) states that trends are now missing at random conditional on $X$ and $D$. In particular, it implies that $E[Y_{it+1}-Y_{it}|X,a=1,S_{t,t+1}=1]$ for $a \in \lbrace G_2,...,C \rbrace$ corresponds to its population counterpart but requires outcome trends to be reweighted using $X$. Notice that it is possible to have different subsets of covariates $X_1,X_2\in X$ for selection into treatment and sample selection with minor modifications.

corollaryUnder the conditions stipulated in Theorem (ref), and by replacing Assumption (ref) with Assumption (ref), the theorem continues to hold. This is achieved through the application of modified weights \begin{equation} \tilde w^G_{\tau\tau-1}(g,X) = \frac{E[S_{\tau} S_{\tau-1}|G_g]}{E[S_{\tau} S_{\tau-1}|X,G_g]} w^G_{\tau\tau-1}(g,X), \end{equation} \begin{equation} \tilde w^C_{\tau\tau-1}(g,X) = \frac{E[S_{\tau} S_{\tau-1}|C]}{E[S_{\tau} S_{\tau-1}|X,C]} w^C_{\tau\tau-1}(g,X). \end{equation}

Corollary (ref) extends the two-step estimator derived from Theorem (ref) to account for sample selection based on observables. This can be achieved by reweighting each difference $Y_t-Y_{t-1}$ with the corresponding factor $\frac{E[S_{\tau} S_{\tau-1}|a]}{E[S_{\tau} S_{\tau-1}|X,a]}$ for $a \in \{G_2,...,C \}$ before applying the same chained DiD estimator. By doing so, we control for potential biases arising from observable characteristics influencing sample selection in addition to unobservable factors, such as unobservable individual heterogeneity.\footnote{Remark that one can also opt for applying the stabilized weights $P(S_{\tau} S_{\tau-1}|X_i,a_i)^{-1}/n^{-1}\sum_{i=1}^nP(S_{\tau} S_{\tau-1}|X_i,a_i)^{-1}$ ex-ante.}

Finally, we extend our framework by relaxing Assumption (ref) to accommodate sample selection on variables $X$, $D$, and past outcomes, aligning with the Sequential Missing at Random (SMAR) assumption described in hoonhout2019nonignorable.\footnote{We thank a referee for suggesting this valuable extension.} For the sake of clarity, we maintain the assumption that an individual is sampled for at most two consecutive periods.

assumption[Sequential Missing At Random] For all $t=1,...,\mathcal{T-1}$, \begin{equation} Y_t\perp S_t|X,D,S_{t-1}=0, \end{equation} \begin{equation} Y_{t+1}\perp S_{t+1}|Y_{t},X,D,S_{t}=1. \end{equation}

Assumption (ref) states that the first time a unit is sampled is a function of $X$ and $D$. However, for subsequent periods, the likelihood of being sampled again also depends on the past observed outcome.

corollaryUnder the conditions stipulated in Theorem (ref), and by replacing Assumption (ref) with Assumption (ref), the theorem continues to hold. This is achieved through the application of modified weights \begin{equation} \tilde{\tilde w}^G_{\tau\tau-1} = \frac{E[S_{\tau} S_{\tau-1}|G_g]}{E[S_{\tau} |X,Y_{\tau-1},S_{\tau-1}=1,G]E[S_{\tau-1} |X,G]} w^G_{\tau\tau-1}(g,X), \end{equation} \begin{equation} \tilde{\tilde w}^C_{\tau\tau-1} = \frac{E[S_{\tau} S_{\tau-1}|C]}{E[S_{\tau} |X,Y_{\tau-1},S_{\tau-1}=1,C]E[S_{\tau-1} |X,C]} w^C_{\tau\tau-1}(g,X). \end{equation}

Corollary (ref) extends the two-step estimator derived from Theorem (ref) to account for sample selection based on observables and past outcomes by using modified weights. The chained DiD estimator is otherwise unchanged.

General missing data patterns

This framework naturally extends to general missing data patterns beyond rotating panel structures. For simplicity of exposition, we remove the dependence of $g$ in our notations. We now consider not only first-differences $\Delta_1 ATT(t)$ but also $k$-period differences $\Delta_k ATT(t)$ from $t-k$ to $t$ for $k=1,...,t-1$. The key requirement is that individuals must be observed at least twice to be used.

Table (ref) provides an example of different missing data patterns with 4 subsamples and 3 periods, and all treatments take place at $t=1$. The first subsample is a balanced panel (BP). It identifies $\Delta_1 ATT(2)$, $\Delta_1 ATT(3)$, and $\Delta_2 ATT(3)$. The second, third and fourth subsamples are incomplete panels (IP1, IP2, IP3) which only identifies a single parameter.\footnote{Note that we do not consider refreshment samples because they do not identify any parameter on their own. They could still be used to address attrition ex-ante hoonhout2019nonignorable.}

table[table omitted — 726 chars of source]

There are multiple ways to identify the ATT from period $1$ to period $3$ in this example. We can identify $ATT(3)$ with (1) BP alone since $ATT(3)=\Delta_2 ATT(3)$ or (2) $ATT(3)=\Delta_1 ATT(2) + \Delta_1 ATT(3)$; or by combining (3) BP and IP3; or (4) BP and IP1; or (5) IP1 and IP2; and with (6) IP3 alone. The optimal combination of the $\Delta_k ATT(t)$ parameters into $ATT(t)$ parameters for general missing data patterns hence involves solving an (overidentified) linear inverse problem. Doing so will make use of all subsamples and deliver efficiency gains compared to focusing on one possible solution, e.g. using only the balanced panel.

The inverse problem arises as follows. Consider the estimates of all possible $\Delta_k ATT(t)$, for all $t\geq 2$ and $t-1\geq k\geq 1$, stacked altogether into a vector $\mathbf{\Delta ATT}$ of length $L_{\Delta}$. $\Delta_k ATT(t)$ is called the k-period-difference group-time average treatment effects, which measure the increase of average treatment effect of group $g$ from period $t-k$ to period $t$. By definition $\Delta_k ATT(t) = ATT(t) - ATT(t-k)$, hence we can write

equation[equation omitted — 84 chars of source]

where $\mathbf{ATT}$ is the vector of $ATT(t)$ for $t\geq2$ of length $L\leq L_{\Delta}$ and $W$ is a matrix where each element takes value in $\lbrace -1,0,1 \rbrace$. Following this reasoning, the example in Table (ref) can be written as

equation[equation omitted — 375 chars of source]

where we pool all subsamples that identify each $\Delta_k ATT(t)$.\footnote{Remark that we could also estimate these parameters separately for each subsample before stacking them into $\mathbf{\Delta ATT}$. That would require to estimate propensity scores conditional on subsample membership.} We propose to solve this problem using a GMM approach. Denoting $\Omega$ the covariance matrix of $\mathbf{\Delta ATT}$, the optimal GMM estimator of $\mathbf{ATT}$ corresponds to $\left(W'\Omega^{-1}W\right)^{-1}W'\Omega^{-1}\mathbf{\Delta ATT}$ since $W\bm{ATT}$ is non-random.\footnote{Remark that if an element of $\mathbf{ATT}$ is not identified, the matrix $W'\Omega^{-1}W$ (or $W'W$) will not be invertible but a (Moore-Penrose) pseudo-inverse can still be used to identify the other elements. In addition, pseudo-inverses, e.g. Tikhonov's, will deliver a stable inverse of $\Omega$ if $\mathbf{\Delta ATT}$ is high-dimensional carrasco2007linear.} This method allows delivering efficiency gains by using all individual time-series with at least two observations, without much additional computational complexity.

\paragraph{Identification, estimation, and inference.}

In order to identify the $ATT(g,t)$ in this general framework, we must modify Assumptions (ref) to (ref) as follows.

assumption[Sampling] For all $t=1,...,\mathcal{T}$, $\left\{ Y_{it-k},Y_{it},X_i,D_{i1},D_{i2},...,D_{i\mathcal{T}} \right\}_{i=1}^{n_{t,k}}$ is independent and identically distributed (iid) conditional on $S_{it-k,t} =1 $, for $k=1,...,t-1$.
assumption[Missing Trends At Random 3] For all $t=1,...,\mathcal{T}$ and $k=1,...,t-1$, \begin{equation} S_{it-k,t} \perp Y_{it}-Y_{it-k}, X_i | D_i. \end{equation}

Assumption (ref) means that we are considering an incomplete panel structure. Conditional on being sampled in the same two periods, individuals are assumed to be iid. However, it does not imply that individuals are sampled in two periods only, nor that observations are representative of the population of interest because we do not assume that the iid draws are taken from the population distribution. Instead, Assumption (ref) states that the sampling process is statistically independent of the joint distribution of k-period-differences of individual outcomes, and observables, conditionally on treatment status in any period. In particular, it implies that $E[Y_{it}-Y_{it-k}|X,a=1,S_{t-k,t}=1]$ for $a \in \lbrace G_1,G_2,...,C \rbrace$ corresponds to its population counterpart. However, this needs not be true for $E[Y_{it}|X,a=1,S_{t}=1]$ for $a \in \lbrace G_1,G_2,...,C \rbrace$.

A critical insight here is that under general missing data patterns, the sampling assumption can lead to testable hypotheses about trends in various data subsets. Indeed, Assumption (ref) implies a certain level of homogeneity across different subpanels. This assumption implies, for the above example, that the trends in outcomes observed in the balanced panel aligns with those in incomplete panel 1, between periods 1 and 2, for both untreated and treated groups. However, this assumption can be relaxed to allow for sample selection based on observables, including subpanel membership, as described in Corollary (ref).

Our estimation procedure is as follows:

enumerate• Compute $\Delta_k ATT(g,t) = 1/n\sum_{i=1}^n[(\hat{w}(g)^{G}_{it,t-k}-\hat{w}(g)^{C}_{it,t-k})(Y_{it} - Y_{it-k})]$, for all $k,g,t$, and stack them into a $L_{\Delta}$-dimensional vector $\bm{\Delta ATT}$; • Define the matrix $W$ appropriately; • Estimate the asymptotic covariance matrix $\hat{\Omega} = n^{-1} \Psi\Psi'$, where $\Psi$ is a $L_{\Delta} \times n$ matrix with elements defined in (ref). • Estimate the optimal GMM estimator $\widehat{\bm{ATT}} = (W'\hat{\Omega}^{-1}W)^{-1}W'\hat{\Omega}^{-1}\bm{\Delta ATT}$;\footnote{Note further that this approach embeds the estimator proposed above in the rotating panel data setting when using only the $\Delta ATT_1(t)$ and replacing $\Omega$ by the identity matrix.}

The next theorem establishes the joint limiting distribution of this estimator.

theoremUnder Assumptions (ref) - (ref), a standard assumption on the parametric estimates of the propensity scores (Assumption 5 in callaway2021difference or 4.4 in abadie_2005), and Assumptions (ref) and (ref), for all $2 \leq g \leq t \leq \mathcal{T}$, \begin{equation*} \sqrt{n}\left( \widehat{\bm{ATT }}-\bm{ATT} \right) \overset{d}{\to} N(0,\bm{\Sigma}), \end{equation*} as $n \to \infty$ and where the covariance $\bm{\Sigma}$ is detailed in the proof.

The theorem shows that this estimator is consistent and asymptotically normal. The bootstrap procedures for $ATT(g,t)$'s (Online Appendix (ref)) and for summary parameters (Online Appendix (ref)) apply with minor modifications, as discussed in the proof. Although our estimator brings efficiency gains by using an optimal weighting matrix, it could be further improved by addressing the efficiency losses from first-step plug-in estimates of propensity scores to adjust for sample selection on observables. This is left for future research.\footnote{We identified two means to address this caveat. frazier2017efficient propose a computationally simple yet general approach that involves targeting and penalization to enforce the asymptotic efficiency for two-step extremum estimators such as ours, whereas chaudhuri2019review and sant2020doubly propose doubly-robust estimators which allow preserving efficiency.}

\paragraph{Regression alternatives.} wooldridge2021two provides a TWFE regression alternative to callaway2021difference which results in estimators of the $ATT$'s that are numerically identical to the imputation approach of borusiak_jaravel_2017. In the absence of covariates $X$, this TWFE regression involves individual, period, and the interaction of cohort with time-since-adoption fixed effects, as shown by

equation[equation omitted — 286 chars of source]

This TWFE event-study regression is equivalent to the chained DiD estimator, provided we use individual fixed-effects (and not cohort fixed-effects), as well as exclude covariates and focus on cases where only two consecutive periods are observed per individual.

A notable advantage of the regression approaches borusiak_jaravel_2017,wooldridge2021two is their flexibility in accommodating unbalanced panel data, unlike the method of callaway2021difference. In these frameworks, the choice between chained DiD and cross-section DiD estimators in these methods hinges on selecting between individual and cohort fixed-effects, each with its own set of identifying assumptions about the sampling process, as discussed earlier.

Despite their adaptability, these alternatives are not without limitations. When incorporating covariates, the model complexity increases significantly due to the need for many interaction terms. They also rely on potentially restrictive linearity assumptions in potential outcomes and treatment effects instead of using propensity scores. Furthermore, aggregating the estimated parameters to derive the target parameter of interest may not be trivial. It is therefore unclear that these methods are more computationally efficient than our estimators.

Numerical Simulations

We propose a simulation design adapted from the first section. Let us specify the potential outcome as a components of variance:

equation[equation omitted — 114 chars of source]

where $D_{i\tau} \in \lbrace 0,1 \rbrace$ denotes whether individual $i$ has been treated in $\tau$ or earlier. Let us assume that $t \in \lbrace 0,...,T+1 \rbrace$ and treatments can only occur in $t\geq 2$ so that $G \in \lbrace 2,...,T+1 \rbrace$. The data generating process is characterized by the following assumptions:

itemize• The individual-specific unobservable heterogeneity is iid gaussian: $\alpha_i \sim N(1,\sigma_{\alpha}^2)$, where $\sigma_{\alpha}^2 = 2$ ; • The time-specific unobservable heterogeneity is iid gaussian: $\delta_t \sim N(1,1)$; • The error term is iid gaussian: $\varepsilon_{it} \sim N(0,\sigma_{\varepsilon}^2)$, where $\sigma_{\varepsilon}^2 = 0.5$; • The probability to receive the treatment at time $g$, conditional on being treated at $g$ or in the control group, is defined as \begin{equation*} Pr(G_{ig}=1|X_i,\alpha_i,G_{ig}+C_i=1) = \frac{1}{1+\exp\left(\theta_0 + \theta_1 X_i + \theta_2 \alpha_i \times g \right)}, \end{equation*} where $X_{i} \sim N(1,1)$ is observable for every $i$, unlike $\alpha_i$, and $\theta_0 = -1$, $\theta_1 = 0.4$ and $\theta_2 = 0$ or $\theta_2 = 0.2$. In the latter case, the treatment probability varies with treatment timing and the unobserved individual heterogeneity; • The sampling probability in the consecutive periods $t,t+1$ conditional on $\alpha_i$ is given by \begin{equation*} Pr(S_{itt+1}=1|\alpha_i) = \frac{1}{1+\exp\left(\lambda_0 + \lambda_1 \alpha_i\times t \right)}, \end{equation*} with $\lambda_0 = -1$, and $\lambda_1 = 0$ or $\lambda_1 = 0.2$, so that the sampling process can also vary with time and the unobserved individual heterogeneity.

We simulate the sampled data in two steps. First, we generate a population sample for each period $t$ to represent individuals that are either treated at $t$ or in the control group. Second, we sample from this population using the specified process. We formalize this procedure as follows:

enumerate• Generate a population of individuals \begin{enumerate} • Draw $N = 2 \times \max\limits_t\{ \frac{n}{E_{\alpha}\left[Pr(S_{itt+1})\right] } \}$ individuals per period in order to have $T+2$ population samples of $N$ individuals, where each individual is characterized by a vector $(\alpha_i,\delta_t,X_i,\varepsilon_{it})$; • Separately for each population sample $g$, draw a uniform random number $\xi_{i}\in[0,1]$ per individual. If $\xi_{i}\leq Pr(G_{it}=1|X_i,\alpha_i,G_{it}+C_i=1)$, then set $(G_{ig} = 1,C_i=0)$, otherwise set $(G_{ig} = 0,C_i=1)$. • Compute $Y_{it}$ from $(\alpha_i,\delta_t,X_i,\varepsilon_{it},G_{i0},...,G_{iT+1},C_i)$; \end{enumerate} • Sample from this population \begin{enumerate} • Draw a uniform random number $\eta_{it}\in[0,1]$ per individual $i$ and period $t$. If $\eta_{it}\leq Pr(S_{itt+1}=1|\alpha_i)$, then set $S_{itt+1} = 1$ and $S_{i\tau \tau+1}=0$ for $\tau \neq t$; • Draw (without replacement) $n$ individuals per period $t$ from the population for which $S_{it t+1}=1$; • Compute the different estimators. • Repeat steps 1(b)-2(c) 1,000 times and report the mean and standard deviation of the estimators. \end{enumerate}

We consider several simulation designs:

itemize• DGP 1: $\theta_2=0$ and $\lambda_1=0$. This is the baseline case where the probability of treatment and the sampling process do not depend on individual heterogeneity so all estimators are unbiased. • DGP 2: $\theta_2=0.2$ and $\lambda_1=0.2$. In this case, both the probability of treatment and the sampling process depend on individual heterogeneity leading to biased estimates for the cross-section DiD. • DGP 3: $\theta_2=0.2$ and $\lambda_1=0.2$. In this case, we simulate a stratified sample where 90% of the individuals are sampled on a rotating basis as before and those with $\alpha_i$ superior to the 90th percentile are always observed (10%). • DGP 4: $\theta_2=0.2$ and $\lambda_1=0.2$. We also simulate a stratified sample where individuals with $\alpha_i$ superior to the 60th percentile are always observed (40%) and the rest is drawn as a rotating subpanel (60%).

For all simulations, we set $T=6$, so there is a total of 8 periods. In each period, the population size is 4800, and we draw 150 individuals such that $S_{itt+1}=1$. Finally, we set $\beta_{\tau}$ to take the values $\{1.75, 1.50, 1.25, 1.00, 0.75, 0.50\}$ for $\tau=\{1,2,3,4,5,6\}$, that is, the treatment effect is positive and decreasing over time, relative to the treatment starting date. For each sample, we estimate the chained DiD (using simulated $S_{itt+1}=1$) and the cross-section DiD (using $S_{it}=1$ that are obtained from $S_{itt+1}=1$).

For DGPs 1 and 2, we also estimate the long DiD assuming that all sampled individuals are observed for the entire time frame. The simulation results are given in Table (ref). This long DiD here is infeasible but serves as a benchmark to illustrate the significant loss of information resulting from having an unbalanced panel. However, the chained DiD delivers unbiased estimates in both cases unlike the cross-section DiD. Remark that we assume iid errors although the chained DiD perform best for random-walk errors.

table[table omitted — 1,802 chars of source]

For DGPs 3 and 4, the long DiD is estimated only for individuals that belong to the balanced subpanel. The simulation results are given in Table (ref). We show estimates from the chained DiD GMM estimator using two weighting matrices: (1) Ch DiD uses the identity matrix, and (2) CD-GMM uses the optimal weighting matrix presented earlier. It appears that when the balanced subpanel consists of only 10% of the data, the chained DiD estimators outperform the long DiD. However, the (asymptotically) optimal weighting matrix does not always deliver more precise estimates than the identity matrix in small samples, at least for this simulation design. In comparison to the identity matrix, It seems that the optimal weights allow mitigating the precision loss for longer-term effects (e.g. $\beta_6$) at the cost of losing some precision for smaller-term effects (e.g. $\beta_1$).

table[table omitted — 1,465 chars of source]

Application: The Employment Effects of an Innovation Policy in France

We now turn to an application of these methods for estimating the causal impact of a French innovation policy supporting collaborative R&D projects over the period 2010-2016.

Background

This innovation policy is made up of different subsidy schemes aimed at developing R&D collaborations between firms and, often, public organizations. These schemes aim at subsidizing collaborative projects oriented towards applied research and experimental development.\footnote{The schemes are FUI, ISI, PSPC, PIAVE, RAPID and ADEME. Although they share the same general objective, they support different forms of R&D projects. For example, PSPC projects are much larger in size than others, FUI projects systematically involve companies and public research organizations, ADEME projects have environmental objectives, and so on. A detailed description of the schemes is available in rapport_dge_2020.} To obtain funding, a firm must set up a research project in partnership with at least one other institution. The project is then submitted to one specific subsidy scheme, often following calls for proposals, much like research grant in academic research. The selection of projects, and the associated funding, is based on a list of criteria including, but not limited to, the innovative nature of the project, its credibility, maturity, or commercial character.

This innovation policy provides an ideal setting to apply our method because R&D projects take several years to complete. Evaluating the effectiveness of the policy hence requires estimating its long-term effects. Unfortunately, one of the main data sources about firm-level R&D activities comes from a survey with a rotating panel design with multiple strata. Firms investing a large amount in R&D expenditure, i.e. large companies, are systematically surveyed, while those spending less, i.e. small and medium-sized enterprises (SMEs) and intermediate-sized enterprises (ISEs), are surveyed only two consecutive years and then dropped out of the sample. Furthermore, large firms are always involved in at least one R&D project, so it is not possible to find a plausible counterfactual for them. Consequently, our policy evaluation focus on SMEs and ISEs, which form an unbalanced panel.

The mechanism for selecting projects submitted by companies should ideally lead to the acceptance of good projects and the rejection of weaker ones. It is conceivable that the accepted projects would have been implemented regardless of subsidies, indicating that what we observe is somehow also reflective of the role played by project quality rather than simply the impact of subsidies. However, the associated risk with these projects is high, and it is unclear that they would have been carried out without subsidies, which account for more than 30% of the total project financing on average. More importantly, the primary goal of our empirical analysis is to assess whether subsidies for collaborative projects are useful in boosting R&D activities within SMEs and ISEs. In this context, the outcome variable is not directly measuring the project's success. Instead, we use the number of researchers and highly qualified workforce, which serves as an indicator of the additional investment in R&D activities. Empirical evidence also indicates significant variability within companies, where one project may be accepted while another is rejected. This variability suggests that it would be too simplistic to label a company as uniformly successful or unsuccessful at attracting subsidies for its projects.\footnote{Note that the treatment is properly defined as receiving a first subsidy for a collaborative project in our application. }

Furthermore, the gradual implementation of the studied mechanisms means that a treated company at a given period often serves as a control for other companies in the period preceding its treatment (the “not yet treated”). This is particularly true given the prolonged maturation of projects, and it is common for the same project to be submitted multiple times in various calls for projects before a definitive acceptance or rejection. Rejected projects may also find better alignment with other funding opportunities and schemes that better suit their objectives. Consequently, many projects are declined not necessarily due to their intrinsic quality but rather because of insufficient maturity or misalignment with the goals of the various subsidy schemes.

The application presented in this paper focuses on the average treatment effect of participating for the first time in any of these schemes without distinguishing their individual effects. It is relevant to analyze this average effect to the extent that all the schemes contribute to the common objective of subsidizing collaborative R&D projects. This aggregation leads to more precise estimates because it maximizes the sample size, but at the same time requires to account for treatment effect heterogeneity and variations in treatment timing.\footnote{It is almost impossible to precisely estimate the individual effect of the smallest schemes as they have subsidized a very small number of projects.}

We estimate the treatment effect of this policy on employment. We focus on employment because it is the main economic variable for which it is also possible to consistently observe an almost identical measure for all firms using administrative data. Therefore, by focusing on such outcome variable, we can compare the results obtained with the chained DiD estimator to the unfeasible estimates obtained using the long DiD. The complete results of this policy evaluation, including the effects on a larger number of economic variables, are available in rapport_dge_2020.

Data

Our data contains information about all R&D projects financed under these schemes over the period 2010-2016. The data includes an unique identifier for each partner participating in a project.\footnote{This identifier corresponds to the SIREN number, a unique identification number for French businesses supervised by the French national institute of statistics.} Using this identifier, we have collected exhaustive firm-level data from administrative sources that provide the main annual indicators on the economic activity of companies over the period 2007-2017. In particular, we collected the following variables: total workforce and the number of engineers in the workforce from administrative records on firms' employment as outcomes of interest, and other variables used in the propensity scores.\footnote{Firms' revenues come from annual tax data restated by INSEE (FICUS/FARE datasets). Employment information comes from administrative records on firms' employment (DADS datasets). Data on cluster policy comes from the French competitiveness cluster (“P\^{o}le de Comp\'etitivit\'e") management database. Finally, data on support for innovation comes from the “Cr\'edit Imp\^{o}t Recherche” (CIR), a research tax credit, and the “Jeunes Entreprises Innovantes” (JEI) scheme, a tax and social exemption aimed at young innovative firms. The CIR is the main tool for supporting innovation in France. Contrary to the devices evaluated in this article, the CIR is an indirect tax aid, in the sense that it is automatically distributed to companies making eligible R&D expenditures and that apply for it.}

Data on the number of researchers is obtained from the R&D survey.\footnote{This survey also provides detailed information on R&D expenditures, the financing of these expenditures, and some outputs.} This annual survey collects information from about 9000 firms companies each year. The survey has a stratified sampling design: firms with intramural R&D expenditures above 750,000 euros are systematically surveyed the following year, while others are surveyed only two years in a row. The vast majority of the SMEs and ISEs in the scope of the study are part of the second stratum. In this context, the application of the standard long DiD method is not possible, which justifies the use of the method developed in this article.

Having knowledge of the amount of CIR (research tax credit) paid to companies and their participation in the French cluster policy is essential. Indeed, the amount of tax credit granted is a good proxy for a company's propensity to be active in R&D. In addition, competitiveness clusters aim to create a network of firms and research organizations to facilitate the formation of collaborative R&D projects. Participation in one of these clusters also reveals a firm's tendency for this form of R&D. These different variables are therefore suitable candidates to explain the probability of receiving the treatment in the propensity score.\footnote{More specifically, the propensity score includes the log of R&D grants, the log of the number of engineers, the log of investment, the log of the variation in R&D grants, the log of the variation in turnover, an indicator for being in a competitiveness cluster, and an indicator for being in the IT sector.} It is important to observe some variables comprehensively across firms so they can be included in the propensity score, as they must be observed in $g-1$, $t$, and $t+1$ to compute the elementary building block constituting each chain link of the chained DiD.

The number of firms that can potentially carry out an innovative R&D activity is very small compared to the total number of firms in France (about 3,000,000). Therefore, in order to avoid comparing firms participating in a collaborative R&D project to other firms which are unlikely to pursue such activity, we restrict the scope of the study to firms active in R&D at least one year over the whole period considered. This activity is measured by merging together all the sources of information available to us for this purpose: the databases of the research tax credit, the JEI scheme, and the R&D survey. Doing so leaves us with about 30,000 firms in the sample.

The schemes covered by the study are the main support mechanisms for collaborative R&D in France and involve the highest amounts of public support. However, there are other alternatives not discussed here. Altogether, the various schemes considered have provided funding for 1697 projects over the period, and have involved 8724 partners.\footnote{A firm can participate in several projects.} These projects received a total aid of 3.6 billion euros and involved total expenditures of 10.4 billion euros from their partners.

\paragraph{Sample selection on observables.} The chained difference-in-differences method relies on the identifying assumption that changes in the potential outcomes are (conditionally) independent of being sampled twice in a row. A key advantage of the method is that it is robust to sample selection on unobservable time-persistent factors. However, the design of the R&D survey implies that the probability of observing the same company twice is positively correlated with observables, in particular past R&D expenditures. Surveyed companies with larger (or increasing) R&D in $t$ are more likely to be surveyed again in $t+1$. Furthermore, companies likely to engage in R&D, identified through auxiliary information, are all surveyed and integrated into the survey data if they conduct R&D. The compilation of various annual surveys therefore results in some sample selection on past R&D expenditures at the company level.

We address this selection issue using inverse propensity weighting of the first-differences $Y_t - Y_{t-1}$ following Corollary (ref) under a variation of Assumption (ref). For tractability, we specify a logit model for the conditional probability $P(S_{tt+1} = 1| Z)$, where $Z$ represents the set of variables determining the survey sampling design, including the level and changes in R&D expenditures and participation in known R&D support mechanisms up to date $t-1$.\footnote{Remark that this simplified approach assumes sampling $S_{t-1}S_t$ to be mean-independent of treatment status $D$ and treatment selection variables $X$, conditional on $Z$.} This approach controls for the effects of selection in the R&D survey sample and stabilizes the scope of the evaluation.

Results

In this application, we focus on the effects on (1) total workforce and (2) the employment of managers and highly qualified workers. Those variables are observed exhaustively from administrative data (DADS). We estimate the dynamic treatment effect on these two observed outcomes by using three estimators: the long DiD, the chained DiD, and the cross-section DiD.

Although these outcomes are consistently observed through time, attrition can still occur. For example, firms may disappear over time because of economic difficulties or because they are acquired by another firm. These companies are not taken into account by the long DiD estimator, unless with an attrition model, whereas they are accounted for by the chained DiD and cross-section DiD estimators.

We balance the data to both facilitate the comparison across estimators and get closer to our theoretical framework. That is, we keep the firms that are consistently observed from 2007 to 2017 in the administrative data. This balanced panel is referred to as the exhaustive panel. We use the long DiD, the chained DiD and the cross-section DiD on this exhaustive panel.

Then, we construct an unbalanced version of this panel by discarding all observations whenever a firm was not sampled in the R&D survey in a given year. Both the chained DiD and cross section DiD estimators are used on this (artificial) unbalanced panel. The objective is to study how each estimator is affected by discarding observations, and compare its performance to the long DiD.

Results are summarized in Figures (ref) and (ref) for the effects on total workforce and Figures (ref) and (ref) for highly qualified workers, along with 95% confidence intervals obtained from the multiplier bootstrap.\footnote{To maximize readability, Figures 2 and 4 are the same as Figures 1 and 3, respectively, to which we have added the cross-section DiD on the unbalanced panel, whose variance is very high. } They show the dynamic treatment effects relative to the beginning of the treatment obtained on a panel balanced on exhaustive variables.\footnote{“Exhaustive" refers to the exhaustively observed outcome. “Unbalanced" refers to the use of an exhaustively observed outcome from which we artificially discard all observations identified as missing in the R&D survey. } Table (ref) and Table (ref) in the Online Appendix provide additional details, including pre-trend tests rejecting the null hypothesis of pre-trends.

figure[figure omitted — 187 chars of source]
figure[figure omitted — 184 chars of source]
figure[figure omitted — 198 chars of source]
figure[figure omitted — 195 chars of source]

The long DiD reveals that total employment has increased by 5.7% the year after the project started ($\beta_2$ in Table (ref)). The long-term increase amounts to 12% five years after the start. For highly qualified workers, the effect is also positive but not statistically significant during the first two years. It becomes significant from the third year onwards.

As expected, the estimates obtained using the complete (exhaustive) panel are very similar between the long DiD, the chained DiD and the cross-section DiD estimators. The standard errors are also similar between the long DiD and the chained DiD estimators but they are much higher for the cross section DiD estimator, which suggests that the variance of the unobserved heterogeneity is large in this data. On the one hand, the estimated coefficients remain quite similar and the standard errors are only slightly higher with the chained DiD estimator on the unbalanced panel. On the other hand, the estimates are considerably worse when obtained with the cross-section DiD estimator on the unbalanced panel (Figures (ref) and (ref)). The point estimates are different and the standard errors become too large to appreciate the effects of the policy.

Finally, we apply the chained DiD and cross-section DiD estimators on outcomes similar to those studied just above but coming from the R&D survey. In this context, it is not possible to apply the standard long DiD estimator because there are too few observations to calculate long differences. For consistency reasons, we present estimates obtained on the same set of firms as those used for the tables (ref) and (ref).\footnote{That is, we use the set of data that is balanced on the exhaustive variables, which is then merged with the R&D survey, and estimate the effects on the variables reported in the R&D survey.}

The outcome variables do not correspond to the exact same definition of employment depending on whether they come from the exhaustive administrative source or from the R&D survey. Total employees headcount from DADS administrative data corresponds to observations at the legal unit level, whereas the R&D survey sometimes provide information on employment at the group level.\footnote{A legal unit is a legal entity of public or private law. A firm, in the sense of a group, is an economic entity that may comprise several legal units thanks to financial links.} The employment of highly qualified workers from the DASD is close to the number of R&D researchers and engineers filled in the R&D survey. Highly qualified workers include engineers but also qualified workers dedicated to other tasks than R&D. Conversely, R&D researchers and engineers include researchers who are specifically assigned to research tasks. Despite their difference, these variables measure similar outcomes and are highly correlated, which justifies their comparison.

Results using outcome variables observed from R&D survey are presented in Figure (ref), and details are provided in Table (ref) in the Online Appendix. The effects obtained with the chained DiD on total employment are somewhat less significant than those presented in Figure (ref), but the coefficients have the same order of magnitude. As might be expected since the policy directly aims at fostering R&D activities, the effects on the number of researchers are stronger and more significant. On the other hand, the effects obtained with the cross section DiD are, once again, much less precise.

figure[figure omitted — 174 chars of source]

We identify two primary reasons for the superior performance of the chained DiD estimator over the cross section DiD estimator in this application. First, employment typically exhibits high temporal persistence, coupled with considerable unobserved heterogeneity. Proposition (ref) demonstrates that the time-dependence of the outcome variable is a crucial determinant in the relative precision of these estimators. Second, there may be sample selection due to unobservable time-persistent factors in the R&D survey. Such selection arises if, for example, the propensity to consistently participate in R&D surveys varies, possibly due to differences in resource availability, past experience with surveys, corporate culture, or differing levels of motivation across research teams in firms.

The time-persistent differences between high-performing and non-performing companies are eliminated by first-differencing. If there are distinct dynamics between treated and untreated firms, then differences could be observed even before treatment initiation. However, placebo tests for pre-trends show no significant employment differences upstream of initial treatment period between treatment and control groups (effects $\beta_{-3}$, $\beta_{-2}$ and $\beta_{-1}$ in Tables (ref) to (ref)). This empirical test supports the absence of a violation of the parallel trend assumption.

Finally, we reproduce the results of Tables (ref), (ref), and (ref) by estimating the treatment effects with the complete original sample from the administrative data, without discarding the individual firms that are not consistently observed throughout the period to create a balanced panel. The results are presented in Online Appendix (ref) and confirm the better performance of the chained DiD estimator.

Conclusion

In this paper, we have developed a new estimator to identify long-term treatment effects in unbalanced panel data sets. This is an important issue, not only because of attrition but also due to how surveys are designed. Common practices are either to use a long DiD estimator by balancing the data, at the cost of losing precision and possibly biasing the results, or to use a cross-section DiD estimator at the cost of not accounting for unobserved heterogeneity. We introduce a new method that simply consists of aggregating short-term DiD estimators obtained from two periods. Our theoretical results show that this estimator identifies the average treatment effects of interest, is consistent and asymptotically normal, accounts for treatment heterogeneity and varying treatment timing, as well as general missing data patterns, and may deliver efficiency gains. An application to an innovation policy implemented in France reveals that, indeed, this estimator allows identifying statistically significant long-term treatment effects where previous methods fail to do so.

Acknowledgments

We are indebted to Serena Ng for her guidance, an associate editor, and two anonymous referees for their valuable comments and suggestions that greatly improved the quality of this work. We would also like to thank Brantly Callaway, Joel Cuerrier, Laurent Davezies, Cl\'{e}ment de Chaisemartin, Xavier D'Haultef{\oe}uille, Yannick Guyonvarch, Xavier Jaravel, and all participants to seminars and conferences for insightful discussions and comments.

The R package cdid is available on CRAN.

Funding

This work was supported by the Social Sciences and Humanities Research Council of Canada (430-2022-00544).