Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
117,838 characters · 13 sections · 84 citation commands
Doubly Robust Difference-in-Differences Estimators
{3pt} {3pt}
Difference-in-differences (DID) methods are among the most popular procedures practitioners adopted to conduct policy evaluation with observational data. In its canonical form, DID identifies the average treatment effect on the treated (ATT) by comparing the difference in pre and post-treatment outcomes of two groups: one that receives and one that does not receive the treatment (the treated and comparison group, respectively). In order to attach a causal interpretation to DID estimators, researchers routinely invoke the (unconditional) parallel trends assumption (PTA): in the absence of the treatment, the average outcome for the treatment and comparison groups would have followed parallel paths over time. Although the PTA is fundamentally untestable, its plausibility is usually questioned if the observed characteristics that are thought to be associated with the evolution of the outcome are not balanced between the treated and comparison group. In such cases, researchers usually deviate from the canonical DID setup and incorporate pre-treatment covariates into the DID analysis and assume that the PTA is satisfied only after conditioning on these covariates.
In this paper, we study the robustness and efficiency properties of DID estimators for the ATT when the PTA holds after conditioning on covariates. We consider both settings where panel data are available and settings where only repeated cross-section data are available. We contribute to the DID literature in different fronts. First, we derive doubly robust (DR) estimands for the ATT under DID settings and propose DR DID estimators for the ATT that are consistent when either a working (parametric) model for the propensity score or a working (parametric) model for the outcome evolution for the comparison group is correctly specified. The setting where only repeated cross-section data are available is particularly interesting. We propose two different DR DID estimators for the ATT that differ from each other depending on whether or not one models the outcome regression for the treated group in both pre and post-treatment periods. Nonetheless, we show DR property does not depend on such a choice.
Second, we derive the semiparametric efficiency bounds for the ATT under DID designs. The semiparametric efficiency bounds we derive are nonparametric in the sense that we do not assume researchers have additional knowledge about outcome regressions or the propensity score functional forms. As so, these bounds provide a standard against which one can compare the efficiency of any (regular) semiparametric DID estimator for the ATT. Here, it is also worth stressing that these semiparametric efficiency bounds explicitly incorporate all the restrictions implied by the invoked identification assumptions. Importantly, these restrictions differ depending on whether panel or repeated cross-section data are available. In both cases they involve the moment restrictions implied by the conditional PTA, though, when repeated cross-section data are available, they also include the restrictions implied by the identifying assumption that the joint distribution of covariates and treatment status is invariant to the sampling period (pre and post-treatment). We emphasize that failing to account for all these implied restrictions can lead to discrepancies on the derived efficiency bound, which, in turn, may suggest that some estimator is semiparametrically efficient when, in fact, it is not.
With the semiparametric efficiency bounds at hand, we can answer several questions that one may have. For instance, one may wonder whether there are efficiency gains associated with having access to panel instead of repeated cross-section data. By directly comparing the efficiency bounds under these two setups, we not only show that the answer to the aforementioned question is yes, but also show that such gains tend to be larger when the sample sizes of the pre and post-treatment repeated cross-section data are more imbalanced.
Another natural question that arises is whether our proposed DR DID estimators can attain the semiparametric efficiency bound. We show that when the working models for the propensity score and for the outcome evolution for the comparison group are correctly specified, our proposed DR DID estimator for the panel data setup is locally efficient, though the DR DID estimators for the cross-section setup are not. In fact, when only repeated cross-section data are available, we show that our proposed DR DID estimator that relies on modelling the propensity score and the outcome evaluation of both the treated and comparison groups attains the semiparametric efficiency bound when all working models are correctly specified. We quantify the loss of efficiency associated with using the inefficient\ DR DID estimator instead of the locally efficient one, and illustrate via Monte Carlo simulations that such a loss can indeed be large.
Our proposed methodology accommodates linear and nonlinear working models for the nuisance functions. We establish $\sqrt{n}$-consistency and asymptotic normality of the proposed DR DID estimators when generic parametric working models are used for the nuisance functions. In doing so we emphasize that, in general, the DR property of our estimators is with respect to consistency and not to inference. In other words, the exact form of the asymptotic variance of our proposed estimators depends on whether the propensity score and/or the outcome regression models are correctly specified. Given that, in practice, one does not know a priori which models are correctly specified, one should consider the estimation effects from all first-step estimators when estimating the asymptotic variance. Failing to do so may lead to invalid inference procedures.
Motivated by this observation, a third contribution of this paper is to show that, by paying particular attention to the estimation method used for estimating the nuisance parameters, it is sometimes possible to construct computationally simple DID estimators for the ATT that are not only DR consistent and locally semiparametric efficient, but are also doubly robust for inference. These further improved DR DID estimators are particularly attractive and easy to implement when researchers are comfortable with a logistic working model for the propensity score and with linear regression working models for the outcome of interest.
Related literature: Our proposal builds on two branches of the causal inference literature. First, our methodological results are intrinsically related to other DID papers; for an overview, see e.g., Section 6.5 of Imbens2009 and references therein. Two leading contributions in this branch of literature that are particularly relevant to this paper are Heckman1997, who propose kernel-based DID regression estimators, and Abadie2005, who proposes (parametric and nonparametric) DID inverse probability weighted (IPW) estimators. We note that when the dimension of available covariates is high or even moderate, fully nonparametric procedures usually do not lead to informative inference because of the \textquotedblleft curse of dimensionality\textquotedblright . In these cases, researchers often adopt parametric methods. Our DR DID estimators fall in this latter category.
Second, our results are also directly related to the literature on doubly robust estimators, see Robins1994, Scharfstein1999, Bang2005, Wooldridge2007a, Chen2008, Cattaneo2010, Graham2012, Graham2016, Vermeulen2015, Lee2017, Sloczynski2018, Rothe2018, Muris2019, among many others; for an overview, see section 2 of Sloczynski2018, and Seaman2018. Recently, DR estimators have also been playing an important role when one uses data-adaptive, \textquotedblleft machine learning\textquotedblright\ estimators for the nuisance functions, see e.g., Belloni2014, Farrell2015, Chernozhukov2017, Belloni2017, and Tan2019. As so, these papers are also broadly related to our proposal, even though we use parametric first-step estimators. On the other hand, we note that the aforementioned papers focus on either the \textquotedblleft selection on observables\textquotedblright\ or \textquotedblleft IV/LATE\textquotedblright\ type assumptions, whereas we pay particular attention to the conditional DID design. Thus, our results complement theirs.
To derive the semiparametric efficiency bounds for the ATT under\ the DID framework, we build on Hahn1998 and Chen2008. Although we follow the structure of semiparametric efficiency bound derivation of the aforementioned papers (which, in turn, follow Newey1990), our derived semiparametric efficiency bounds complement theirs as we focus on DID designs while Hahn1998 and Chen2008 results rely on \textquotedblleft selection on observables\textquotedblright\ type assumptions in cross-section setups.
Our results for the further improved DR DID estimators build on Vermeulen2015, who propose estimators that are DR for inference in cross-section setups under selection on observables type assumptions. We extend Vermeulen2015 proposal to DID settings with both panel and repeated cross-section data. Our further improved DR DID estimators also build on Graham2012, as their proposed propensity score estimator is one important component of our proposal.
Finally, in work related but independent from ours, Zimmert2019 provides high-level conditions under which one can use \textquotedblleft machine-learning\textquotedblright\ first-step estimators when estimating the ATT in DID setups. His results complement ours, though we note that his proposed estimators for the repeated cross-section case do not attain the semiparametric efficiency bound derived in this paper, and the loss of efficiency can be of first-order importance. We also note that Zimmert2019 does not provide a detailed comparison between the panel and repeated cross-section data setups like we do, nor discusses DR inference procedures, which are particularly relevant under model misspecifications.
Organization of the paper: In the next section, we describe this paper's framework, briefly give an overview of the existing DID estimators and describe how we combine the strengths of each method to form our DR DID estimands. We also derive semiparametric efficiency bounds for the ATT in Section (ref). In Section (ref), we propose different DR DID estimators, derive their large sample properties, and show that we can get improved DR DID estimators by paying particular attention to the estimation method used for estimating the nuisance parameters. We examine the finite sample properties of our proposed methodology by means of a Monte Carlo study in Section (ref), and provide an empirical illustration in Section (ref). Section (ref) concludes. Mathematical proofs are gathered in the Supplemental Appendix.\footnote{ The Supplemental Appendix is available at \url{https://pedrohcgs.github.io/files/DR-DID Appendix.pdf}} Finally, all proposed policy evaluation tools discussed in this article can be implemented via the open-source R package DRDID, which is freely available from GitHub (https://github.com/pedrohcgs/DRDID).
We first introduce the notation we use throughout the article. We focus on the case where there are two treatment periods and two treatment groups. Let $Y_{it}$ be the outcome of interest for unit $i$ at time $t$. We assume that researchers have access to outcome data in a pre-treatment period $t=0$ and in a post-treatment period $t=1$. Let $D_{it}=1$ if unit $i$ is treated before time $t$ and $D_{it}=0$ otherwise. Note that $D_{i0}=0$ for every $i$ , allowing us to write $D_{i}=D_{i1}$. Using the potential outcome notation, denote $Y_{it}\left( 0\right) $ the outcome of unit $i$ at time $t$ if it does not receive treatment by time $t$ and $Y_{it}\left( 1\right) $ the outcome for the same unit if it receives treatment. Thus, the realized outcome for unit $i$ at time $t$ is $Y_{it}=D_{i}Y_{it}\left( 1\right) +\left( 1-D_{i}\right) Y_{it}\left( 0\right) $. A vector of pre-treatment covariates $X_{i}$ is also available. Henceforth, we assume that the first element of $X_{i}$ is a constant.
In the rest of the article, we assume that either panel or repeated cross-section data on $\left( Y_{it},D_{i},X_{i}\right) $, $t=0,1$ are available. When repeated cross-section data are available, we follow Abadie2005 and assume that covariates and treatment status are stationary. We formalize these conditions in the following assumption. Let $T_{i}$ be a dummy variable that takes value one if the observation $i$ is only observed in the post-treatment period, and zero if observation $i$ is only observed in the pre-treatment period. Define $Y_{i}=T_{i}Y_{i1}+\left( 1-T_{i}\right) Y_{i0}$, and let $n_{1}$ and $n_{0}$ be the sample sizes of the post-treatment and pre-treatment periods such that $n=n_{1}+n_{0}$. Finally, let $\lambda =\mathbb{P}\left( T=1\right) \in \left( 0,1\right) $.
Assumption (ref)$(a)$ covers the case where panel data are available, whereas Assumption (ref)$(b)$ covers the case where repeated cross-section data are available, and allows for different sampling schemes. For instance, it accommodates the binomial sampling scheme where an observation $i$ is randomly drawn from either $\left( Y_{1},D,X\right) $ or $ \left( Y_{0},D,X\right) $ with fixed probability $\lambda $ (here, $T$ is a non-degenerated random variable). It also accommodates the \textquotedblleft conditional\textquotedblright\ sampling scheme where $n_{1}$ observations are sampled from $\left( Y_{1},D,X\right) $, $n_{0}$ observations are sampled from $\left( Y_{0},D,X\right) $ and $\lambda =n_{1}/n$ (here, $T$ is treated as fixed). On the other hand, Assumption (ref)$\left( b\right) $ rules out settings with compositional changes in $\left( D,X\right) $, see e.g. Hong2013 for a discussion.
The parameter of interest is the average treatment effect on the treated,
As expectations are linear operators and $Y_{i1}\left( 1\right) =Y_{i1}$ if $ D_{i}=1$, we can rewrite the ATT as\footnote{ Throughout the rest of the paper, to ease the notation burden we denote $ \mathbb{E}\left[ \cdot \right] $ as generic expectations. In the case of panel data, such expectations are with respect to the distribution of $ \left( Y_{0},Y_{1},D,X\right) $. In the case of repeated cross-section data, the expectations are with respect to the mixture distribution $\sum_{t=0}^{1} \mathbb{P}\left( T=t\right) \cdot \mathbb{P}\left( Y_{t}\leq y,D=d,X\leq x|T=t\right) $.}
where we drop subscript $i$ to ease notation; we follow this convention throughout the paper. From the above representation, it is clear that the main challenge in identifying the ATT is to compute $\mathbb{E} [Y_{i1}(0)|D_{i}=1]$ from the observed data. To overcome this challenge, we invoke the following assumptions.
Assumption (ref), which we refer to as the conditional PTA throughout the paper, states that in the absence of treatment, the average conditional outcome of the treated and the comparison groups would have evolved in parallel. Note that Assumption (ref) allows for covariate-specific time trends, though it rules out unit specific trends. Assumption (ref) is an overlap condition and states that at least a small fraction of the population is treated and that for every value of the covariates $X$, there is at least a small probability that the unit is not treated. These two assumptions are standard in conditional DID methods, see e.g. Heckman1997, Heckman1998, Blundell2004a, Abadie2005 and Bonhomme2011.
Under Assumptions (ref)-(ref), there are two main flexible estimation procedures to estimate the ATT: the outcome regression (OR) approach, see e.g. Heckman1997, and the IPW approach, see e.g. Abadie2005. The OR approach relies on researchers ability to model the outcome evolution. In such cases, under the aforementioned assumptions one can estimate the ATT using
where $\bar{Y}_{d,t}=\sum_{i\left\vert D_{i}=d,T_{i}=t\right. }Y_{it}/n_{d,t} $ is the sample average outcome among units in treatment group $d$ and time $t$, and $\widehat{\mu }_{d,t}\left( x\right) $ is an estimator of the true, unknown $m_{d,t}\left( x\right) \equiv \mathbb{E} [Y_{t}|D=d,X=x]$,\footnote{ In the repeated cross-section case, $m_{d,t}\left( x\right) =\mathbb{E}\left[ Y|D=d,T=t,X=x\right] $. In the next section, we differentiate the notation for the panel data and repeated cross-section case to avoid potential confusions.} see e.g. Heckman1997.
The IPW approach proposed by Abadie2005 avoids directly modelling the outcome evolution and exploits that, under Assumptions (ref)- (ref), the ATT can be expressed as
when panel data are available, and as
when repeated cross-section data are available, where $p\left( X\right) \equiv \mathbb{P}\left( D=1|X\right) $ is the true, unknown propensity score. Abadie's identification results suggest simple two-step estimators for the ATT that do not involve outcome regressions. For instance, when panel data are available, Abadie2005 proposes the following Horvitz1952 type IPW estimator,
where $\widehat{\pi }\left( x\right) $ is an estimator of the true, unknown $ p\left( x\right) $, and for a generic random variable $Z$, $\mathbb{E}_{n} \left[ Z\right] =n^{-1}\sum_{i=1}^{n}Z_{i};$ the estimator for the repeated cross-section case is formed using the analogous procedure.
It is important to emphasize that the reliability of ATT estimators based on the OR and the IPW approaches depends on different, non-nested conditions. For the OR approach, the consistency of the ATT estimator ((ref)) relies on the estimators of $m_{d,t}\left( \cdot \right) $, $\widehat{\mu } _{d,t}\left( \cdot \right) $, being correctly specified, whereas the IPW estimator ((ref)) relies on the propensity score estimator $\widehat{ \pi }(\cdot )$ of $p\left( \cdot \right) $ being correctly specified. As so, in practice, it may be hard to \textquotedblleft rank\textquotedblright\ these two approaches in terms of their robustness to model misspecification.
In this section, we argue that instead of choosing between the OR and the IPW approaches, one can combine them to form doubly robust (DR) moments/estimands for the ATT. Here, double robustness means that the resulting estimand identifies the ATT even if either (but not both) the propensity score model or the outcome regression models are misspecified. As so, the DR DID estimand for the ATT shares the strengths of each individual DID method and, at the same time, avoids some of their weaknesses.
Before describing how we exactly combine the OR and the IPW approaches to form our DR DID estimand, we need to introduce some additional notation. Let $\pi \left( X\right) $ be an arbitrary model for the true, unknown propensity score. When panel data are available, let $\Delta Y=Y_{1}-Y_{0}$ and define $\mu _{d,\Delta }^{p}\left( X\right) \equiv \mu _{d,1}^{p}\left( X\right) -\mu _{d,0}^{p}\left( X\right) $, $\mu _{d,t}^{p}\left( x\right) $ being a model for the true, unknown outcome regression $m_{d,t}^{p}\left( x\right) \equiv \mathbb{E}[Y_{t}|D=d,X=x]$, $d,t=0,1$. When only repeated cross-section data are available, let $\mu _{d,t}^{rc}\left( x\right) $ be an arbitrary model for the true, unknown regression $m_{d,t}^{rc}\left( x\right) \equiv \mathbb{E}[Y|D=d,T=t,X=x]$, $d,t=0,1$, and for, $d=0,1,$ $ \mu _{d,Y}^{rc}\left( T,X\right) \equiv T\cdot \mu _{d,1}^{rc}\left( X\right) +\left( 1-T\right) \cdot \mu _{d,0}^{rc}\left( X\right) $, and $\mu _{d,\Delta }^{rc}\left( X\right) \equiv \mu _{d,1}^{rc}\left( X\right) -\mu _{d,0}^{rc}\left( X\right) $.
For the case in which panel data are available, we consider the estimand
where, for a generic $g$,
For the repeated cross-section case, we consider two different estimands,
and
where, for a generic $g$,
and, for $t=0,1$,
Theorem (ref) states that provided that at least one of the working nuisance models is correctly specified, we can recover the ATT with either panel or repeated cross-section data. Thus, our proposed DR DID estimands are \textquotedblleft less demanding\textquotedblright\ in terms of the researchers' ability to correctly specify models for the nuisance functions than either the OR or the IPW approach.
Given that we consider two different estimands for the case of repeated cross-section, it is interesting to use Theorem (ref) to compare them. Given that $\tau _{1}^{dr,rc}$ does not rely on OR models for the treated group but $\tau _{2}^{dr,rc}$ does, one could a priori expect that $\tau _{1}^{dr,rc}$ would be more robust\ against model misspecification than $ \tau _{2}^{dr,rc}$. Nonetheless, Theorem (ref) states that this is not the case as they identify the ATT under the same conditions. At this stage, one may wonder how this is possible. To answer such a query, it suffices to remember that, under the stationarity condition in Assumption (ref)$\left( b\right) $, for any generic integrable and measurable function $g$, $\mathbb{E}\left[ \left. g\left( X\right) \right\vert D=1 \right] =\mathbb{E}\left[ \left. g\left( X\right) \right\vert D=1,T=t\right] $, $t=0,1$. Given that this holds for any generic function $g$, it must also hold for $\mu _{1,t}^{rc}\left( \cdot \right) -\mu _{0,t}^{rc}\left( \cdot \right) ,$ $t=0,1$, even when $\mu _{d,t}^{rc}\left( \cdot \right) $ are misspecified models of $m_{d,t}^{rc}\left( \cdot \right) $. Such a result reveals that modeling the OR for the treat group can be \textquotedblleft harmless\textquotedblright\ in terms of identification, provided that these additional models are incorporated into $\tau _{1}^{dr,rc}$ in an appropriate manner.
In the previous subsection, we derived DR moment equations for the ATT under the DID framework and showed that the resulting estimands are more robust against model misspecifications than DID estimands based on either the OR or the IPW approach. In this subsection, we shift our attention from \textquotedblleft robustness\textquotedblright\ to efficiency. More precisely, we calculate the semiparametric efficiency bound for the ATT under Assumptions (ref)-(ref) when either panel or repeated cross-section data are available. These results provide the semiparametric analog of the Cram\'{e}r--Rao lower bound commonly used in fully parametric procedures. As so, they provide a benchmark that researchers can use to assess whether any given (regular) semiparametric DID estimator for the ATT is fully exploiting the empirical content of Assumptions (ref)-(ref).
Let $m_{0,\Delta }^{p}\left( x\right) \equiv m_{0,1}^{p}\left( x\right) -m_{0,0}^{p}\left( x\right) $, and, for $d=0,1,$ $m_{d,\Delta }^{rc}\left( X\right) \equiv m_{d,1}^{rc}\left( X\right) -m_{d,0}^{rc}\left( X\right) $. Recall that $\lambda \equiv \mathbb{P}\left( T=1\right) $. Next proposition displays the semiparametric efficiency bound for the ATT when one has access to panel data and when one has access to repeated cross-section data. To simplify exposition, we abstract from additional technical discussions related to the conditions to guarantee quadratic mean differentiability and their implications for the precise definition of efficient influence function ; see, e.g., Chapter 3 of Bickel1998 for additional details.
It is interesting to compare $\eta ^{e,p}\left( D,X\right) $ with $\eta ^{e,rc}\left( D,T,X\right) $. First, note that the first component of their efficient influence functions are analogous to each other, and depends on the true, unknown conditional ATT, $m_{1,\Delta }\left( X\right) -m_{0,\Delta }\left( X\right) $.\footnote{ To avoid excessive notational burden, we supress the \textquotedblleft $p$ \textquotedblright\ and \textquotedblleft $rc$\textquotedblright\ superscripts unless their omission leads to confusion.} The second and third terms in ((ref)) and ((ref)) are more different from each other. For $\eta ^{e,p}$, the availability of panel data implies that $Y_{1}$ and $Y_{0}$ are observed for all units, and, therefore, we can directly reweight $\Delta Y-m_{1,\Delta }\left( X\right) $ and $ \Delta Y-m_{0,\Delta }\left( X\right) $. In contrast, when only repeated cross-section data are available, one observes $Y_{t}$ only if $T=t$, $t=0,1$ , and, therefore, the efficient influence function ((ref)) depends on different weights for each pair $\left( D,T\right) $ $\in \left\{ 0,1\right\} ^{2}$. In this latter case, we also stress the importance of imposing the stationarity condition in Assumption (ref)(b) when deriving the efficient influence function ((ref)) -- failing to do so will suggest an \textquotedblleft efficiency bound\textquotedblright\ that is wider than ((ref)).
It is also worth mentioning that the efficient influence functions ((ref)) and ((ref)) depend on the true, unknown, outcome regression functions for the treated group, $m_{1,1}\left( \cdot \right) $ and $m_{1,0}\left( \cdot \right) $, in an asymmetric manner. On one hand, when panel data are available, by simple manipulation, we can rewrite $\eta ^{e,p}$ as
emphasizing that the efficient influence function for the ATT when panel data are available does not depend on $m_{1,1}\left( \cdot \right) $ and $m_{1,0}\left( \cdot \right) $. This is in sharp contrast to the case where only repeated cross-section data are available.
Another interesting question raised by Proposition (ref) is whether the semiparametric efficiency bound for the case of repeated cross-section data is larger than the one for the case of panel data. In order to answer this question, we consider the case where $T$ is independent of $\left( Y_{1},Y_{0},D,X\right) $, so that Assumptions (ref)$ \left( a\right) $ and (ref)$\left( b\right) $ are compatible with each other.\footnote{ This \textquotedblleft restriction\textquotedblright\ does not affect the semiparametric efficiency bound for the case where only repeated cross-section data are available, as it does not impose additional restrictions on the observed data.}
In other words, under the DID framework it is possible to form more efficient estimators for the ATT when panel data are available than when only repeated cross-section data are available. In addition, from Corollary (ref), we can also see that the efficiency loss is convex in $ \lambda \,$, implying that the loss of efficiency is bigger when the pre and post-treatment sample sizes are more imbalanced. In fact, when
we can show that $\lambda =0.5$ is optimal . However, when ((ref)) does not hold, the optimal $\lambda $ depends on the data in a more complicated manner, and is given by $\lambda =\left. \tilde{\sigma} _{1}\right/ \left( \tilde{\sigma}_{0}+\tilde{\sigma}_{1}\right) $, where, for $t=0,1$
These results suggest that, in principle, one may benefit from \textquotedblleft oversampling\textquotedblright\ from either the pre or post-treatment period. However, it is, in general, not feasible to know the optimal\ $\lambda $ during the design stage, i.e., at the pre-treatment period, since $\tilde{\sigma}_{1}^{2}$ depends on the outcome data from the post-treatment period. Thus, if one were to design the DID study with repeated cross-section units, it seems that setting $\lambda =0.5$ would be a \textquotedblleft reasonable\textquotedblright\ choice.
In this section, we build on the DR DID estimands in Theorem (ref) and the semiparametric efficiency bounds in Proposition (ref), and discuss estimation and inference procedures for the ATT in DID designs. Indeed, the moment equations ((ref)), ((ref)), and ((ref) ) suggest a simple two-step strategy to estimate the ATT. In the first step, one estimates the true, unknown $p\left( \cdot \right) $ using $\pi (\cdot ), $ and the true, unknown $m_{d,t}^{p}\left( \cdot \right) $ ($ m_{d,t}^{rc}\left( \cdot \right) $) using $\mu _{d,t}^{p}\left( \cdot \right) $($\mu _{d,t}^{rc}\left( \cdot \right) $), $d,t=0,1$, when panel data (repeated cross-section data) are available. In the second step, one plugs the fitted values of the estimated propensity score and regression models into the sample analogue of $\tau ^{dr,p},$ $\tau _{1}^{dr,rc}$, or $ \tau _{2}^{dr,rc}$.
Although, in principle, one can use semi/non-parametric estimators for both the outcome regressions and the propensity score, see e.g. Heckman1997 , Abadie2005, Chen2008 and Rothe2018, in what follows ,we focus our attention on generic parametric first-step estimators. More precisely, we assume that $\pi \left( x;\gamma ^{\ast }\right) $ is a parametric model for $p\left( x\right) ,$ such that $\pi $ is known up to the finite dimensional pseudo-true parameter $\gamma ^{\ast }$. Analogously, for $d,t=0,1$, $\mu _{d,t}^{p}\left( x;\beta _{d,t}^{\ast ,p}\right) $ (and $ \mu _{d,t}^{rc}\left( x;\beta _{d,t}^{\ast ,rc}\right) $) is a parametric model for $m_{d,t}^{p}\left( x\right) $ ($m_{d,t}^{rc}\left( x\right) $), such that $\mu _{d,t}^{p}$ ($\mu _{d,t}^{rc}$) is known up to the finite dimensional pseudo-true parameter $\beta _{d,t}^{\ast ,p}$ ($\beta _{d,t}^{\ast ,rc}$). This is perhaps the most popular approach adopted by practitioners, particularly when the available sample size is moderate and/or the dimension of available covariates is high or even moderate, as the \textquotedblleft curse of dimensionality\textquotedblright\ usually prevents one to adopt fully nonparametric procedures.\footnote{ Let $g\left( x\right) $ be a generic notation for $p\left( x\right) $, $ m_{d,t}^{l}\left( X\right) ,$ $m_{d,t}^{l}\left( X\right) $, $d,t=0,1$, $ l=p,rc.$ From Newey1994, Chen2003, Ai2003, Ai2007, Ai2012, and Chen2008, one can see that the use of nonparametric first-step estimators $\widehat{g}\left( x\right) $ of $g\left( x\right) $ is warranted provided that $\left\Vert \widehat{g}\left( x\right) -g\left( x\right) \right\Vert _{\mathcal{H}}=o_{p}\left( n^{-1/4}\right) $ for a pseudo-metric $\left\Vert \cdot \right\Vert _{\mathcal{H}}$, $\mathcal{H}$ being a vector space of functions. However, when the dimension of $X$ is moderate or large, as is usually the case in many empirical applications, conditons ensuring that $\left\Vert \widehat{g}\left( x\right) -g\left( x\right) \right\Vert _{\mathcal{H}}=o_{p}\left( n^{-1/4}\right) $ can be rather stringent because of the \textquotedblleft curse of dimensionality\textquotedblright .}
In the case when panel data are available, our proposed DR DID estimator for the ATT is based on ((ref)) and is given by
where
$\widehat{\gamma }$ is an estimator for the pseudo-true $\gamma ^{\ast }$, $ \widehat{\beta }_{0,t}^{p}$ is an estimator for pseudo-true $\beta _{0,t}^{\ast ,p}$, $t=0,1$, and for a generic $\beta _{0}$ and $\beta _{1}$, $\mu _{0,\Delta }^{p}(\cdot ;\beta _{0},\beta _{1})=\mu _{0,1}^{p}\left( \cdot ;\beta _{1}\right) -\mu _{0,0}^{p}\left( \cdot ;\beta _{0}\right) $.
When only repeated cross-section data are available, we propose two different DR DID estimators for the ATT. The first one, which is based on ( (ref)) and can be interpreted as the analogue of $\widehat{\tau } ^{dr,p}$, is given by
where $\mu _{0,Y}^{rc}\left( T,\cdot ;\beta _{0,0}^{rc},\beta _{0,1}^{rc}\right) =T\cdot \mu _{0,1}^{rc}\left( \cdot ;\beta _{0,1}^{rc}\right) +\left( 1-T\right) \cdot \mu _{0,0}^{rc}\left( \cdot ;\beta _{0,0}^{rc}\right) $, $\widehat{\beta }_{d,t}^{rc}$ is an estimator for the pseudo-true $\beta _{d,t}^{\ast ,rc}$, $d,t=0,1$, and the weights $ \widehat{w}_{1}^{rc}\left( D,T\right) $ and $\widehat{w}_{0}^{rc}\left( D,T,X;\widehat{\gamma }\right) $ are, respectively, defined as the sample analogues of $w_{1}^{rc}\left( D,T\right) $ and $w_{0}^{rc}\left( D,T,X;g\right) $ defined in ((ref)), but with $\pi \left( x; \widehat{\gamma }\right) $ playing the role of $g$.
The second DR DID estimator for the case of repeated cross-section builds on ((ref)) and is given by
where $\mu _{d,\Delta }^{rc}\left( \cdot ;\beta _{d,1}^{rc},\beta _{d,0}^{rc}\right) =\mu _{d,1}^{rc}\left( \cdot ;\beta _{d,1}^{rc}\right) -\mu _{d,0}^{rc}\left( \cdot ;\beta _{d,0}^{rc}\right) $, and the weights $ \widehat{w}_{1,t}^{rc}\left( D,T\right) $ and $\widehat{w}_{0,t}^{rc}\left( D,T,X;\widehat{\gamma }\right) $ are, respectively, defined as the sample analogues of $w_{1,t}^{rc}\left( D,T\right) $ and $w_{0,t}^{rc}\left( D,T,X;g\right) $, $t=0,1$, defined below ((ref)), but with $\pi \left( x;\widehat{\gamma }\right) $ playing the role of $g$.
As we show in the Appendix (ref), it is relatively straightforward to derive the asymptotic properties of $\widehat{\tau } ^{dr,p}$, $\widehat{\tau }_{1}^{dr,rc}$ and $\widehat{\tau }_{2}^{dr,rc}$ using generic first-step estimators that satisfy some relatively weak, high-level conditions; see Theorems (ref) and (ref) in Appendix (ref). Indeed, Theorem (ref) indicates that $\widehat{\tau }^{dr,p}$ is doubly robust, and also locally semiparametrically efficient, i.e., its asymptotic variance achieves the semiparametric efficiency bound when the working models for the nuisance functions are correctly specified. Theorem (ref) also indicates that both $\widehat{\tau }_{1}^{dr,rc}$ and $\widehat{\tau }_{2}^{dr,rc}$ are doubly robust when repeated cross-section data are available. However, Theorem (ref) also highlights that $\widehat{\tau }_{2}^{dr,rc}$ is locally semiparametrically efficient, whereas $\widehat{\tau } _{1}^{dr,rc} $ is not. In other words, when repeated cross-section data are available, $\widehat{\tau }_{2}^{dr,rc}$ tends to have more attractive properties than $\widehat{\tau }_{1}^{dr,rc}$, regardless of the first-step estimators used.
Although the results in Theorems (ref) and (ref) accommodate a variety of different first-step estimators, in practice, one still needs to choose a particular estimation procedure to be implemented. In what follows, we attempt to provide some guidance on the choice of first-step estimators with the goal of further improving the (generic) DR DID estimators. We are particularly interested in forming DR DID estimators that are not only doubly robust in terms of consistency---like described above---but also doubly robust for inference, i.e., their asymptotic linear representation is also doubly robust. The attractiveness of forming estimators that are DR for inference is that there is no estimation effect from first-step estimators, which, in turn, implies that the asymptotic variance of the results DR DID estimator for the ATT is invariant to which working models for the nuisance functions are correctly specified. In practice, this usually translates to simpler and more stable inference procedures.
To derive these improved DR DID estimators, we focus on the case where a researcher is comfortable with linear regression working models for the outcome of interest, a logistic working model for the propensity score, and with covariates $X$ entering all the nuisance models in a symmetric manner. Although these modelling conditions are more stringent than those allowed by our generic DR DID estimators discussed in Appendix (ref), they are much weaker than those implicitly imposed in the TWFE specification ((ref)), and can be seen as the default choice in many applications. Hence, these extra assumptions can be seen as a reasonable compromise to get further improved DR DID estimators that are also computationally tractable and easy to implement in practice.
As discussed above, we consider the following working models for the nuisance functions:
Our proposed improved DR DID estimator is given by the three-step estimator
where the first two-steps consist of computing
while in the third and last step, one plugs the fitted values of the working models ((ref)) into the sample analogue of $\tau ^{dr,p}$ . {Here, note that }$\widehat{\gamma }^{ipt}$ is the inverse probability tilting estimator proposed by Graham2012 in a different context, while $\widehat{\beta }_{0,\Delta }^{wls,p}$ is simply the weighted least squares estimator for $\beta _{0,\Delta }^{\ast ,p}$.
At this point, one may wonder why we use the estimators $\widehat{\gamma } ^{ipt}$ and $\widehat{\beta }_{0,\Delta }^{wls,p}$ instead of other available alternatives. To answer such a query, recall that the main goal here is to propose DID estimators for the ATT that are not only DR consistent but also DR for inference, i.e., the exact form of their asymptotic variance does not depend on which working models for the nuisance functions are correctly specified. As it turns out, the key to obtain DID estimators for the ATT that are also DR\ for inference is to choose first-step estimators for the nuisance parameters, say $\widehat{\gamma }$ and $\widehat{\beta }$, such that the limiting distribution of the resulting DR DID estimator $\widehat{\tau }^{dr,p}$ is equivalent to that of the infeasible DR DID estimator that uses the pseudo-true values of $\widehat{ \gamma }$ and $\widehat{\beta }$, say $\gamma ^{\ast }$ and $\beta ^{\ast }$ . In a more precise manner, in order to get DID estimators that are DR for inference, we need to guarantee that there will be no estimation effect from the first stage.
In Appendix (ref), we show that the estimation effect associated with using generic first-step estimators $\widehat{\gamma }$ and $ \widehat{\beta }$ is given by $\eta _{est}^{p}\left( W;\gamma ^{\ast },\beta ^{\ast }\right) $ as defined in ((ref)). By paying closer attention to the exact form of $\eta _{est}^{p}\left( W;\gamma ^{\ast },\beta ^{\ast }\right) $, one can see that if
then there will be no estimation effect from the first stage. As the first component of $X$ is assumed to be constant and we adopt the working models ( (ref)), it follows that ((ref)) reduces to
However, as $n\rightarrow \infty $, these two vectors of moment conditions follow from the first-order conditions of the optimization problems associated with $\widehat{\gamma }^{ipt}$ and $\widehat{\beta }_{0,\Delta }^{wls,p}$, respectively, even when these working models are misspecified. Hence, by using $\widehat{\gamma }^{ipt}$ and $\widehat{\beta }_{0,\Delta }^{wls,p}$, we guarantee that $\widehat{\tau }_{imp}^{dr,p}$ is doubly robust for inference as there is no estimation effect from replacing the pseudo-true parameters $\gamma ^{\ast ,ipt}$ and $\beta _{0,\Delta }^{\ast ,wls,p}$ with their estimators $\widehat{\gamma }^{ipt}$ and $\widehat{\beta }_{0,\Delta }^{wls,p}$, respectively.
The next theorem formalizes this discussion. Define
and let
Part $\left( a\right) $ of Theorem (ref) generalizes the cross-section results of Vermeulen2015 to the DID framework. It states that the proposed DR DID estimators for the ATT, $\widehat{\tau } _{imp}^{dr,p}$, is doubly robust, $\sqrt{n}$-consistent and asymptotically normal. It also states that the exact form of $V_{imp}^{p}$ does not depend on which working models are correctly specified, implying that $\widehat{ \tau }_{imp}^{dr,p}$ is doubly robust not only in terms of consistency but also terms of inference. An important consequence of this DR-for-inference property is that it allows one to treat the summands of $\widehat{\tau } _{imp}^{dr,p}$ as if they were independent and identically distributed, and, therefore, estimate $V_{imp}^{p}$ by
Part $\left( b\right) $ of Theorem (ref) indicates that $ \widehat{\tau }_{imp}^{dr,p}$ is semiparametrically efficient when the working model for the propensity score, and the working models for the outcome regression for the comparison units are correctly specified.
In this section, we turn our attention to our proposed improved DR DID estimators for the ATT when only repeated cross-section data are available. Similar to the panel data case, we consider the case where a researcher is comfortable with the following specifications,
We consider two improved DR DID estimators based on ((ref)) and ((ref)), namely
and
where
Here, note that $\widehat{\tau }_{1,imp}^{dr,rc}$ does not rely on OR models for the treated group while $\widehat{\tau }_{2,imp}^{dr,rc}$ does. In addition, when one compares $\widehat{\tau }_{1,imp}^{dr,rc}$ and $\widehat{ \tau }_{2,imp}^{dr,rc}$ with $\widehat{\tau }_{imp}^{dr,p}$, it is evident that the latter relies on a single OR model since we observe $Y_{1}$ and $ Y_{0}$ for all units; when only repeated cross-section data are available, one needs to model the OR in each time period (and each treatment group). Another interesting feature worth mentioning is that we estimate the OR parameters for the treated group via ordinary least squares, whereas we estimate the OR parameters for the control group with weighted least squares. This follows from the fact that estimating the pseudo-true parameters $\beta _{1,t}^{\ast ,rc}$, $t=0,1$, does not lead to any estimation effect, and therefore one can choose her favorite estimation method. Given this observation and the linear specification in ((ref)), we find it natural to estimate $\beta _{1,t}^{\ast ,rc}$, $t=0,1$, via OLS as this is the most widespread estimation procedure adopted by practitioners.
Let
and for $\beta _{imp}^{\ast ,rc}=\left( \beta _{0,1}^{\ast ,wls,rc},\beta _{0,0}^{\ast ,wls,rc},\beta _{1,1}^{\ast ,ols,rc},\beta _{1,0}^{\ast ,ols,rc}\right) ,$ define
where $\eta _{1}^{rc,1}$, $\eta _{0}^{rc,1}$, $\eta _{1}^{rc,2}$, and $\eta _{0}^{rc,2}$ are defined as in ((ref))-((ref)) in the Appendix (ref).
Next theorem states that $\widehat{\tau }_{1,imp}^{dr,rc}$ and $\widehat{ \tau }_{2,imp}^{dr,rc}$ are not only doubly robust consistent but also doubly robust for inference. Furthermore, it states that $\widehat{\tau } _{2,imp}^{dr,rc}$ is locally semiparametrically efficient, whereas $\widehat{ \tau }_{1,imp}^{dr,rc}$ is not.
In other words, Theorem (ref) states that both $\widehat{\tau }_{1,imp}^{dr,rc}$ and $\widehat{\tau }_{2,imp}^{dr,rc}$ are doubly robust for the ATT, $\sqrt{n}$-consistent and asymptotically normal. Similar to the panel data case, the exact form of the $V_{j,imp}^{rc}$, $j=1,2$, does not depend on which working models are correctly specified, implying that both $ \widehat{\tau }_{1,imp}^{dr,rc}$ and $\widehat{\tau }_{2,imp}^{dr,rc}$ are also doubly robust in terms of inference.
Part $\left( b\right) $ of Theorem (ref) indicates that $ \widehat{\tau }_{2,imp}^{dr,rc}$ is semiparametrically efficient when the working model for the propensity score, and all working models for the outcome regressions, for both treated and comparison units, are correctly specified. When compared to Theorem (ref)$\left( b\right) $ , it is evident that such a requirement is much stronger than when panel data are available. Part $\left( b\right) $ of Theorem (ref) also indicates that, in general, $\widehat{\tau }_{1,imp}^{dr,rc}$ is not locally semiparametrically efficient. As so, we argue that, in practice, one should favor $\widehat{\tau }_{2,imp}^{dr,rc}$ with respect to $\widehat{ \tau }_{1,imp}^{dr,rc}$, as both estimators are doubly robust in terms of consistency and inference, but the former is locally semiparametrically efficiency whereas the latter is not.
We conclude this section by providing a precise characterization of the efficiency loss associated with using $\widehat{\tau }_{1,imp}^{dr,rc}$ instead of $\widehat{\tau }_{2,imp}^{dr,rc}$ when all working models are correctly specified. Here, our main goal is to illustrate that by using an estimator that attempts to mimic the panel data setup and that does not explicitly exploit the stationarity condition in Assumption (ref)(b), one may incur in substantial efficiency loss. As so, we argue that, in practice, one should favor estimators based on the DR moment ( (ref))---such as $\widehat{\tau }_{2,imp}^{dr,rc}$---with respect to estimators based on the DR moment ((ref))---such as $\widehat{\tau } _{1,imp}^{dr,rc}$.
In this section, we conduct a series of Monte Carlo experiments in order to study the finite sample properties of our proposed DR DID estimators. When panel data are available, we compare our proposed DR DID estimators $ \widehat{\tau }^{dr,p}$ and $\widehat{\tau }_{imp}^{dr,p}$ given in ((ref)) and ((ref)), respectively, to the OR DID estimator ((ref)), the Horvitz1952 type IPW estimator ( (ref)), and the TWFE regression model ((ref)). Given that the weights of the IPW estimator ((ref)) are not normalized to sum up to one, $\widehat{\tau }^{ipw,p}$ can be unstable particularly when propensity score estimates are relatively close to one. To assess the role played by the weights, we also consider the Hajek1971 type IPW estimator for the ATT
where the weights $\widehat{w}_{1}^{p}\left( D\right) $ and $\widehat{w} _{0}^{p}\left( D,X;\widehat{\gamma }\right) $ are given by ((ref)) and are normalized to sum up to one.
When only repeated cross-section data are available, we compare our proposed DR DID estimators $\widehat{\tau }_{1}^{dr,rc}$ and $\widehat{\tau } _{2}^{dr,rc}$ given in ((ref)) and ((ref) ), and their further improved versions $\widehat{\tau }_{1,imp}^{dr,rc}$ and $\widehat{\tau }_{2,imp}^{dr,rc}$ given in ((ref)) and ( (ref)), to the OR DID estimator ((ref)), the plug-in IPW estimator based on ((ref)), and the TWFE regression model ((ref)). As in the case of panel data, we also consider the Hajek1971 type IPW estimator for the ATT
where the weights are the same as those in $\widehat{\tau }_{1}^{dr,rc}$.
In all simulation exercises, we consider a logistic propensity score working model and a linear regression working model for the outcome evolution. All observed covariates enter the working models linearly. With the exception of $\widehat{\tau }_{imp}^{dr,p}$, $\widehat{\tau }_{j,imp}^{dr,rc}$, $j=1,2$, where we use the estimation methods proposed in Section (ref) and in Section (ref), the OR models are estimated using ordinary least squares, and the propensity score working model is estimated using maximum likelihood estimation. When panel data are available, we consider OR models for $\Delta Y\,$ instead of OR models for $Y_{0}$ and $ Y_{1}$ separately.
We consider sample size $n$ equal to $1000$. For each design, we conduct $ 10,000$ Monte Carlo simulations. We compare the various DID estimators for the ATT in terms of average bias, median bias, root mean square error (RMSE), empirical 95% coverage probability, the average length of a 95% confidence interval, and the average of their plug-in estimator for the asymptotic variance. The confidence intervals are based on the normal approximation, with the asymptotic variances being estimated by their sample analogues. We also compute the semiparametric efficiency bound under each design to allow one to assess the potential loss of efficiency/accuracy associated with using inefficient DID estimators for the ATT.
We first discuss the case where panel data are available. For a generic $ W=\left( W_{1},W_{2},W_{3},W_{4}\right) ^{\prime },$ let
Let $\mathbf{X}=\left( X_{1},X_{2},\right. $ $\left. X_{3},X_{4}\right) ^{\prime }$ be distributed as $N\left( 0,I_{4}\right) $, and $I_{4}$ be the $ 4\times 4$ identity matrix. For $j=1,2,3,4$, let $Z_{j}=\left. \left( \tilde{ Z}-\mathbb{E}\left[ \tilde{Z}\right] \right) \right/ \sqrt{Var\left( \tilde{Z }\right) }$, where $\tilde{Z}_{1}=\exp \left( 0.5X_{1}\right) $, $\tilde{Z} _{2}=10+X_{2}/\left( 1+\exp \left( X_{1}\right) \right) $, $\tilde{Z} _{3}=\left( 0.6+\left. X_{1}X_{3}\right/ 25\right) ^{3}$ and $\tilde{Z} _{4}=\left( 20+X_{2}+X_{4}\right) ^{2}$.
Building on Kang2007, we consider the following data generating processes (DGPs):
where $\varepsilon _{0}$, $\varepsilon _{1}\left( d\right) $, $d=0,1$ are independent standard normal random variables, $U$ is an independent standard uniform random variable, and for a generic $W$, $v\left( W,D\right) $ is an independent normal random variable with mean $D\cdot f_{reg}\left( W\right) $ and variance one. The available data are $\left\{ Y_{0,i},Y_{1,i},D_{i},Z_{i}\right\} _{i=1}^{n}$, where $Y_{0}=Y_{0}\left( 0\right) $, and $Y_{1}=DY_{1}\left( 1\right) +\left( 1-D\right) Y_{1}\left( 0\right) $. In the aforementioned DGPs, the true ATT is zero, and $v$ plays the role of time-invariant unobserved heterogeneity.
Given that we focus on the empirically relevant setting where the observed covariates $Z$ enter all working models linearly, it is clear that in $DPG1$ , both propensity score (PS) and OR working models are correctly specified. In $DGP2$, only the OR working model is correctly specified, whereas in $ DGP3 $ only the PS working model is correctly specified. In $DGP4$, all working models are misspecified. The simulation results are presented in Table (ref).
First, note that the TWFE estimator $\widehat{\tau }^{fe}$ is severely biased and its confidence interval for the ATT has almost zero coverage in all analyzed DGPs. These results should not be unexpected, because, as discussed in Remark (ref), $\widehat{\tau }^{fe}$ implicitly rules out covariate-specific trends, and when these are relevant, like in the considered DGPs, the estimand associated with $\widehat{\tau }^{fe}$ is not the ATT. As so, policy evaluations based on $\widehat{\tau }^{fe}$ can be misleading.
The results in Table (ref) also suggest that, when both the OR and PS working models are correctly specified, all semiparametric estimators for the ATT show little to no Monte Carlo bias, but $\widehat{\tau }^{reg}$, $\widehat{\tau }^{dr,p}$ and $\widehat{\tau }_{imp}^{dr,p}$ dominate the IPW DID estimators $\widehat{\tau }^{ipw,p}$ and $\widehat{\tau }_{std}^{ipw,p}$ on the basis of bias, root mean square error, asymptotic variance, and length of the confidence interval. Indeed, both IPW DID estimator seem to be substantially less efficient than $\widehat{\tau }^{reg}$, $\widehat{\tau } ^{dr,p}$ and $\widehat{\tau }_{imp}^{dr,p}$. The performance of these last three estimators are very close, though $\widehat{\tau }^{reg}$ tends to be more efficient than the other two DR DID estimators. Given that $\widehat{ \tau }^{reg}$ exploits additional assumptions when compared to $\widehat{ \tau }^{dr,p}$ and $\widehat{\tau }_{imp}^{dr,p}$, such a result is not unexpected. Also note that the Hajek1971 type IPW estimator $\widehat{ \tau }_{std}^{ipw,rc}$ is more stable than the Horvitz1952 type IPW estimator $\widehat{\tau }^{ipw,rc}$: the RMSE (the asymptotic variance) of $ \widehat{\tau }^{ipw,rc}$ are more than two (four) times bigger than that of $\widehat{\tau }_{std}^{ipw,rc}$. Such a finding highlights the practical importance of using weights that are normalized to sum up to one.
When only the OR working model is correctly specified, our proposed DR DID estimators $\widehat{\tau }^{dr,p}$ and $\widehat{\tau }_{imp}^{dr,p}$ are competitive with the OR DID estimator $\widehat{\tau }^{reg}$, while the IPW DID estimators are biased, as one should expect. On the other hand, when only the PS working model is correctly specified, the IPW and DR estimators show little to no bias, while $\widehat{\tau }^{reg}$ displays non-negligible bias. Here, it is worth emphasizing that $\widehat{\tau } ^{dr,p}$ and $\widehat{\tau }_{imp}^{dr,p}$ drastically outperform $\widehat{ \tau }^{ipw,p}$ and $\widehat{\tau }_{std}^{ipw,p}$, with $\widehat{\tau } _{imp}^{dr,p}$ also showing substantial improvements with respect to both $ \widehat{\tau }^{dr,p}$ and $\widehat{\tau }_{std}^{ipw,p}$. When one compares the two IPW estimators, the role played by the normalized weights is again clear, as $\widehat{\tau }_{std}^{ipw,p}$ is again much more \textquotedblleft stable\textquotedblright\ than $\widehat{\tau }^{ipw,p}$.
When both OR and PS working models are misspecified, not unexpectedly all estimators have non-negligible biases and inference procedures are, in general, misleading. In this scenario, our DR DID estimators have smaller biases and RMSE than the OR and the normalized IPW DID estimators, with $ \widehat{\tau }_{imp}^{dr,p}$ strictly dominating $\widehat{\tau }^{dr,p}.$ However, the Horvitz1952 IPW DID estimator $\widehat{\tau }^{ipw,p}$ seems to perform best in this DGP.
In terms of efficiency, the results in Table (ref) show that the estimated asymptotic variance of $\widehat{\tau }^{dr,p}$ and $\widehat{\tau }_{imp}^{dr,p}$ are very close to the semiparametric efficiency bound when both the PS and OR regression are correctly specified, which is in agreement with our locally efficiency results in Theorems (ref) and (ref) (in the Appendix). When the PS is misspecified but the OR is not, the estimated asymptotic variances of $\widehat{\tau }^{dr,p}$ and $ \widehat{\tau }_{imp}^{dr,p}$ are still close to the semiparametric efficiency bound in this particular DGP, though we emphasize that this is not predicted by our results and can be a feature of this particular DGP. Finally, we note that when the OR is misspecified but the PS is not, the estimated asymptotic variances of our proposed DR DID estimators $\widehat{ \tau }^{dr,p}$ and $\widehat{\tau }_{imp}^{dr,p}$ are far from the semiparametric efficiency bound, with $\widehat{\tau }_{imp}^{dr,p}$ outperforming $\widehat{\tau }^{dr,p}$ in terms of efficiency in this particular DGP.
We now analyze the performance of the DID estimators for the ATT when one only observes repeated cross-section data. To do so, we consider the same DGPs as in the panel data framework, but instead of observing data on $ \left( Y_{0},Y_{1},D,Z\right) ,$ one observes data on $\left( Y_{0},D,Z\right) $ if $T=0$, or on $\left( Y_{1},D,Z\right) $ if $T=1$, where $T=1\left\{ U_{T}\leq \lambda \right\} $, and $U_{T}$ is a standard uniform random variable, and $\lambda \in \left( 0,1\right) $ a fixed constant.
Table (ref) presents the simulation results with $\lambda =0.5$ and with $n\equiv n_{1}+n_{0}=1,000$.\footnote{ Simulation results with $\lambda =0.25$ and $\lambda =0.75$ reached analogous conclusions to those discussed below and are available upon request.} Overall, the simulation exercise reveals that the efficiency bound, RMSE, asymptotic variance, and confidence interval length of the considered DID estimators are much larger when only repeated cross-section data are available than when panel data are available. In light of Corollary (ref), such a result should be expected, though the magnitude of such loss of efficiency can be striking. In addition, the results in Table (ref) reveal that: $\left( i\right) $ the TWFE estimator $ \widehat{\tau }^{fe}$ is severely biased for the ATT in all DGPs, just like in the panel data case; $\left( ii\right) $ the IPW estimator with standardized weights $\widehat{\tau }_{std}^{ipw,rc}$ is much more stable and efficient than $\widehat{\tau }^{ipw,rc}$ in all DGPs, and, as one should expect, when the PS working model is misspecified, these IPW estimators display non-negligible biases; $\left( iii\right) $ as one should expect, the OR DID estimator displays non-negligible bias when the OR working models are misspecified; $\left( iv\right) $ all four DR DID estimators display little to no bias when one of the working models is correctly specified, but the locally efficient DR DID estimators $\widehat{ \tau }_{2}^{dr,rc}$ and $\widehat{\tau }_{2,imp}^{dr,rc}$ present important efficiency gains when compared to all other DID estimators, including $ \widehat{\tau }_{1}^{dr,rc}$ and $\widehat{\tau }_{1,imp}^{dr,rc}$. These gains in efficiency are more pronounced when the OR models are correctly specified. The simulation results also show that $\left( v\right) $ when one compares the performance of the further improved DR DID estimators $\widehat{ \tau }_{1,imp}^{dr,rc}$ and $\widehat{\tau }_{2,imp}^{dr,rc}$ with the \textquotedblleft traditional\textquotedblright\ DR DID estimators $\widehat{ \tau }_{1}^{dr,rc}$ and $\widehat{\tau }_{2}^{dr,rc}$, it is clear that appropriately choosing the estimation methods for the nuisance parameters can have practical consequences, especially when the outcome regression working models are misspecified.
In terms of efficiency, the results in Table (ref) highlight that, when all working models are correctly specified, the estimated asymptotic variances of $\widehat{\tau }_{2}^{dr,rc}$ and $\widehat{\tau } _{2,imp}^{dr,rc}$ are indeed close to the semiparametric efficiency bound, but the asymptotic variances $\widehat{\tau }_{1}^{dr,rc}$ and $\widehat{ \tau }_{1,imp}^{dr,rc}$ are substantially higher than the semiparametric efficiency bound; these findings are in agreement with our locally efficiency results in Theorems (ref) and (ref) (in the Appendix). Similarly to the panel data case, we find that, in this specific DGP, the estimated asymptotic variances of $\widehat{\tau } _{2}^{dr,rc}$ and $\widehat{\tau }_{2,imp}^{dr,rc}$ are still close to the semiparametric efficiency bound when the outcome regressions are correctly specified but the PS is not, but not when the PS is correctly specified but the outcome regressions are not.
In a very influential study, Lalonde1986 analyzes whether different treatment effect estimators based on observational data are able to replicate the experimental findings of the NSW job training program on post treatment earnings. His negative results led to an increased awareness of the potential pitfalls of observational data and helped spur the use of randomized controlled trials among economists. In addition, alternative policy evaluation tools arose to overcome \textquotedblleft LaLonde's critique\textquotedblright\ of observational estimators. Two prominent examples are the propensity score matching (PSM), see e.g. Dehejia1999, Dehejia2002 (henceforth DW) and the difference-in-differences matching, see e.g. Heckman1997 and Smith2005 (henceforth ST). For instance, DW show that PSM can replicate the experimental benchmark of the NSW for a particular subsample of the original data. ST, on the other hand, cast doubt on the \textquotedblleft generalizability\textquotedblright\ of DW PSM results to a larger population and argue that the conclusions may be sensitive to the propensity score specification. ST also argue that for the NSW data, difference-in-differences matching estimators may be more suitable than cross-section PSM, as they can account for time-invariant unobserved confounding factors.
Motivated by ST findings, in what follows, we focus on DID estimators and evaluate whether our proposed DR DID estimators can better reduce the selection bias when compared to other DID estimation procedures. We analyze three different experimental samples --- the original LaLonde experimental sample, the DW sample, and the \textquotedblleft early random assignment\textquotedblright\ (early RA) subsample of the DW sample considered by ST --- and consider data from the Current Population Survey (CPS) to form a non-experimental comparison group. The pre-treatment covariates in the data include age, years of education, real earnings in 1974, and dummy variables for high school dropout, married, black, and Hispanic. The outcome of interest is real earnings in 1978. We also observe real earnings in 1975, which we use as the pre-treatment outcome $Y_{0}$. The experimental benchmark for the ATT is equal to $\$886$ (s.e. $\$488$), $ \$1794$ (s.e. $\$671$), and $\$2748$ (s.e. $\$1005$) for the LaLonde, DW, and early RA sample, respectively. For additional description and summary statistics for each sample, see Smith2005.
Following ST, we focus on estimating the average \textquotedblleft evaluation bias\textquotedblright\ of different DID estimators. This is only made possible given the availability of experimental data. First, randomization ensures that both \textquotedblleft treatment\textquotedblright\ groups are comparable in terms of self selection. Second, given that randomized-out individuals did not receive training via NSW, the impact of NSW is known to be zero in this group. Thus, applying different DID estimators to data from randomized-out individuals (our pseudo treated group in this exercise) and nonexperimental CPS comparison observations (our comparison group in this exercise) should produce an estimated ATT equal to zero, if these DID estimators are consistent. Deviations from zero are what we call evaluation bias.\footnote{ An alternative way to estimate \textquotedblleft evaluation bias\textquotedblright\ is to compare the ATT using the experimental data with ATT using data from randomized-in and nonexperimental comparison units. This is the approach taken by Lalonde1986 and Dehejia1999, Dehejia2002. A disadvantage of this approach compared to the one we and Smith2005 use is that experimental ATT estimates are also random and may differ from the \textquotedblleft true\textquotedblright\ ATT. Thus, the computation of \textquotedblleft true\textquotedblright\ evaluation biases is much more challenging if not impossible. In any case, results treating the experimental ATT as true effects lead to similar conclusions and are available upon request.}
Like in the Monte Carlo simulation exercises, we compare our proposed DR DID estimators $\widehat{\tau }^{dr,p}$ and $\widehat{\tau }_{imp}^{dr,p}$ with the TWFE estimator $\widehat{\tau }^{fe}$ based on ((ref)), the OR DID estimator $\widehat{\tau }^{reg}$ as defined in ((ref)), and the Horvitz1952 type IPW DID estimator proposed by Abadie2005, $ \widehat{\tau }^{ipw,p}$, as defined in ((ref)). We also consider the Hajek1971 type IPW estimator $\widehat{\tau }_{std}^{ipw,p}$ as defined in ((ref)). We assume that the outcome models are linear in parameters and that the propensity score follows a logistic specification. The unknown parameters are estimated using ordinary least squares (OLS) and maximum likelihood, respectively, except in $\widehat{\tau }_{imp}^{dr,p}$, where we use the estimation methods described in Section (ref).
In order to assess the sensitivity of the findings with respect to the model specifications, we consider three different specifications for how covariates enter into each model: $\left( i\right) $ a linear specification where all covariates enter the models linearly; $\left( ii\right) $ a specification in the spirit of DW, which adds to the linear specification a dummy for zero earnings in 1974, age squared, age cubed divided by $1000$, years of schooling squared, and an interaction term between years of schooling and real earnings in 1974; and $\left( iii\right) $ an \textquotedblleft augmented DW\textquotedblright\ specification, which adds to the \textquotedblleft DW\textquotedblright\ specification the interactions between married and real earnings in 1974, and between married and zero earnings in 1974 --- these two interaction terms were used in Firpo2007.
Table (ref) summarizes the results. Standard errors are reported in parentheses and the estimated evaluation biases relative to the experimental ATT benchmark are reported in brackets. As argued by ST, these \textquotedblleft relative biases\textquotedblright\ are useful for comparing DID estimators within each sample, but as the experimental benchmark estimates for the ATT vary substantially among the three experimental samples, they should not be used for comparing DID estimators across samples. \afterpage{
}
Table (ref) highlights some interesting patterns. First, estimators based on two-way fixed effect regression models tend to be very stable across specifications, but usually display large positive and statistically significant evaluation biases. Second, DID estimators based on the regression approach tend to lead to the most precise estimates. However, for the LaLonde sample, point estimates are severely biased downward, leading to statistically significant evaluation biases. Abadie's IPW estimators $\widehat{\tau }^{ipw,p}$ for the ATT tend to have the largest standard errors across all considered estimators, but their evaluation biases are relatively small. Like in our Monte Carlo simulation results, considering normalized weights as in $\widehat{\tau }_{std}^{ipw,p}$ can improve the stability of the IPW estimators $\widehat{\tau }^{ipw,p}$. Finally, note that our proposed DR DID estimators share the favorable bias properties of Abadie's IPW estimator, but at the same time, have smaller standard errors than IPW estimators. When we compare $\widehat{\tau }^{dr,p}$ with $\widehat{\tau }_{imp}^{dr,p}$, we note that the further improved DR DID estimator $\widehat{\tau }_{imp}^{dr,p}$ tends to have smaller standard errors, particularly when one adopts the \textquotedblleft DW\textquotedblright\ or the \textquotedblleft augmented DW\textquotedblright\ specifications. Taken together, the results using the NSW job training data suggest that our proposed DR DID estimators are an attractive alternative to existing DID procedures.
In this article, we proposed doubly robust estimators for the ATT in difference-in-differences settings where the parallel trends assumption holds only after conditioning on a vector of pre-treatment covariates. Our proposed estimators remain consistent for the ATT when either (but not necessarily both) a propensity score model or outcome regression models are correctly specified, and achieve the semiparametric efficiency bound when the working models for the nuisance functions are correctly specified. We derived the large sample properties of the proposed estimators in situations where either panel data or repeated cross-section data are available, and showed that by paying particular attention to the estimation methods used to estimate the nuisance parameters, one can form DID estimators for the ATT that are not only DR consistent and locally semiparametric efficient, but also DR for inference. We illustrated the attractiveness of our proposed causal inference tools via a simulation exercise and with an empirical application.
Our results can be extended to other situations of practical interest. A leading case is when researchers are interested in understanding treatment effect heterogeneity with respect to continuous covariates $X_{1}$, where $ X_{1}$ is a (strict) subset of available covariates $X$. Here, the parameter of interest is the conditional average treatment effect on the treated CATT$ \left( X_{1}\right) \equiv \mathbb{E}\left[ Y\left( 1\right) -Y\left( 0\right) |X_{1},D=1\right] $ and because of its infinite dimensional nature, the estimation and inference tools proposed in this paper are not directly applicable. However, by combining the DR DID formulation proposed in this paper with the methodology put forward by Chen2018a, one can propose uniformly valid inference procedures not only for the CATT but also for possibly nonlinear functionals of the CATT such as (higher order) partial derivatives, conditional average (higher order) partial derivatives, and partial derivatives of its $log$.
Another interesting extension is when researchers want to adopt data-adaptive, \textquotedblleft machine learning\textquotedblright\ first-step estimators instead of the parametric models discussed in this paper. Here, the main challenge is to derive the influence function of the DR DID estimator for the ATT, as \textquotedblleft machine learning\textquotedblright\ estimators are, in general, in a non-Donsker classes of functions. We envision that one can bypass such technical complications by combining the results derived in this paper with those in Chernozhukov2017, Belloni2017, and Tan2019, for example; see e.g. Zimmert2019 for some recent results in this direction. We leave the detailed analysis of these extensions to future work.