Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
63,103 characters · 14 sections · 122 citation commands
Covariate Balancing and the Equivalence of Weighting and Doubly Robust Estimators of Average Treatment Effects
\thispagestyle{empty}
\onehalfspacing \setcounter{page}{2}
Covariate adjustment is central to causal inference, yet the choice of method remains contested. Much recent research has highlighted the shortcomings of a number of well-established estimation methods in reproducing suitable averages of heterogeneous treatment effects. A key lesson from this literature is that additive linear models may often fail to properly adjust for covariates when those covariates are relevant for identification. This concern arises not only under unconfoundedness Sloczynski2022,GPHK2024,Chen2025, but also in instrumental variables Sloczynski2024,BBMT2025 and difference-in-differences settings CC2024.
When considering alternatives to standard methods, researchers face a wide range of options, including regression adjustment, matching, weighting, and doubly robust estimators, as well as related approaches based on machine learning. Such variety is not necessarily desirable: the existence of a common standard facilitates comparability across studies, while additional researcher degrees of freedom invite specification searching SNS2011,Vivalt2019,FPP2020.
In this paper, we build on the fact that many of these alternative methods require first-step estimation of the propensity score. We assume that a researcher would be willing to commit to estimating the propensity score using a particular method of moments approach with desirable properties, namely the inverse probability tilting (IPT) estimator of EGP2008 and GPE2012,GPE2016. (To be clear, a researcher using IPT still needs to choose a model for the propensity score, perhaps logit or probit, but would then estimate this model using the method of moments instead of maximum likelihood.) Our main contribution is to demonstrate that commitment to using IPT substantially reduces the choice set (i.e., the number of alternative estimators) available to researchers. Specifically, we show that using the IPT moment conditions to estimate the propensity score leads to numerical equivalence between members of several classes of estimators of average treatment effects: inverse probability weighting (IPW), augmented inverse probability weighting (AIPW), and inverse probability weighted regression adjustment (IPWRA), with the latter two classes using linear models for potential outcome means. Our equivalence results are very general, as they are valid for any propensity score model having the index form (e.g., logit or probit). In addition, they apply to both normalized and unnormalized versions of IPW and AIPW, as well as to estimators of both the average treatment effect (ATE) and the average treatment effect on the treated (ATT)\@.
Our equivalence results are meaningful insofar as IPT is appealing in its own right. Indeed, a major reason for this appeal is that the moment conditions underlying IPT impose a desirable property known as “exact balancing.” Specifically, IPT estimates the parameters of the propensity score model such that, after reweighting, the mean covariate values are identical across key groups: between the (weighted) treatment group, the (weighted) comparison group, and the (unweighted) full sample when estimating the ATE, and between the (unweighted) treatment group and the (weighted) comparison group when estimating the ATT\@. In this sense, IPT is a prototypical “covariate balancing” estimator, ensuring that causal comparisons are made only between groups with identical (mean) characteristics. As shown by EGP2008 and GPE2012,GPE2016, IPT estimators of the ATE and ATT also enjoy local efficiency and double robustness properties. That is, if both the propensity score model chosen by the researcher (e.g., logit or probit) and the linear specification of potential outcome means are correct, the estimators are semiparametrically efficient; if only one is correctly specified, the estimators remain consistent. As we review below, the IPT estimator of the ATT is also identical to the subsequent proposals of Hainmueller2012 and IR2014.\footnote{Specifically, IPT coincides with Hainmueller2012's (Hainmueller2012) entropy balancing estimator when the propensity score is estimated with the logit model.}
We also translate our equivalence results to instrumental variables and difference-in-differences settings. In both contexts, weighting and doubly robust estimators have played a prominent role in the recent literature, and our results again simplify the set of alternative estimators available to empirical researchers. In particular, our results imply that the doubly robust estimator proposed by SAZ2020, which uses IPT-estimated propensity scores, is numerically equivalent to the simple IPW estimator of Abadie2005 with IPT weights. It follows that in implementations of SAZ2020 it is redundant to estimate the model for the untreated potential change in outcomes.
We illustrate our findings with two empirical applications. In the first application, we revisit the study of the causal effects of cash transfers on longevity in AEFLM2016. In line with an earlier replication in Sloczynski2022, we conclude that there is insufficient evidence to reject the null hypothesis of zero average effects. Our numerical equivalences simplify analysis and interpretation, and none of the IPT estimates is significantly different from zero despite usually having smaller standard errors than the corresponding estimates based on maximum likelihood. In the second application, we replicate SAZ2020's (SAZ2020) analysis of the NSW--CPS data, which was previously analyzed by LaLonde1986 and many others. We show that many of the estimates reported by SAZ2020 would have been identical to their preferred estimates had they used IPT to estimate the propensity score in all cases.
We also supplement this paper with a companion Stata package, teffects2, available at the Statistical Software Components (SSC) Archive. Our package implements IPW, AIPW, and IPWRA estimators of the ATE and ATT under unconfoundedness, with several approaches to estimate the weights, including IPT\@. The package can also be used to estimate the ATT in difference-in-differences settings after a suitable transformation of the outcome variable. This provides a novel implementation of the doubly robust difference-in-differences (DRDID) estimator proposed by SAZ2020.
This paper builds on a body of research on estimating average treatment effects under the assumption of unconfoundedness. We focus on three classes of estimators: inverse probability weighting (IPW), as in HIR2003, augmented inverse probability weighting (AIPW), as in RRZ1994, and inverse probability weighted regression adjustment (IPWRA), as in Wooldridge2007 and SW2018. These classes of estimators, as well as several others, were surveyed by IW2009, AC2018, and Uysal2024.
Each of the classes of estimators we consider requires a first-step estimation of the propensity score. This is typically done using maximum likelihood estimation (MLE) of a standard binary response model for treatment assignment (e.g., logit or probit). However, this estimation approach does not guarantee that any desirable balancing properties are satisfied in finite samples. In contrast, various “covariate balancing” estimators of the propensity score are explicitly constructed with these properties in mind; they are also tailored to the specific parameter of interest (e.g., ATE or ATT) to improve the statistical properties of the corresponding treatment effect estimator.
Following the early work on inverse probability tilting (IPT) by EGP2008 and GPE2012,GPE2016, many papers have proposed alternative covariate balancing procedures for estimating average treatment effects. Hainmueller2012 suggested estimating the inverse probability weights directly---subject to balancing and normalizing constraints---rather than estimating the propensity score first and then inverting it to obtain the weights. Known as “entropy balancing,” this procedure was designed to estimate the ATT and was later shown to be identical to IPT when the latter uses the logit model ZP2017,Tan2020. IR2014 proposed using different moment conditions than those in EGP2008 and GPE2012 when estimating the ATE; however, the resulting estimator lacks some desirable theoretical properties of IPT, such as double robustness. On the other hand, IR2014's (IR2014) moment conditions for estimating the ATT are the same as in IPT, which implies that the resulting estimator is also equivalent to IPT (as well as to entropy balancing when the logit model is used). Zubizarreta2015 relaxed the exact balancing requirements of earlier methods and proposed estimating weights that minimize variance subject to approximate balancing constraints. Zhao2019 unified and generalized much of the earlier work by introducing a covariate balancing framework based on optimizing loss functions tailored to a given estimand. SASX2022 proposed estimating the propensity score by maximizing balance across the entire covariate distribution rather than in selected functions of the covariates.
Most of the early work on covariate balancing focused on addressing missing data problems and estimating average treatment effects under unconfoundedness. However, recent research has also applied similar ideas to estimating various parameters of interest in difference-in-differences SAZ2020,CSA2021 and instrumental variables settings Heiler2022,SASX2022,SS2024,SUW2025.
This paper is also related to the important work of RSLGR2007, Kline2011, CZ2023, and BSDFO2025 demonstrating numerical equivalences between regression adjustment and weighting estimators of average treatment effects, under the constraint that both the potential outcome means and the weights are linear in covariates. While the weights may generally be approximated as a linear function of a high-dimensional dictionary, as in CNS2022 and BSDFO2025, the corresponding parametric restriction is unlikely to be plausible in low-dimensional settings, especially since it is equivalent to assuming an inverse linear model for the propensity score. In such settings, some of the estimated weights are likely to be negative, which invalidates the sample boundedness property of the resulting estimator RSLGR2007.
In this paper, we extend and generalize this earlier work by demonstrating that the numerical equivalences between the IPW, AIPW, and IPWRA estimators are driven by the IPT moment conditions rather than parametric restrictions on the propensity score or the weights. Unlike our paper, the results in RSLGR2007, Kline2011, CZ2023, and BSDFO2025 are specific to the inverse linear model for the propensity score. (Most of these results are also limited to IPW\@.) Under this strong parametric restriction, which our paper does not make, IPW, AIPW, and IPWRA with IPT moment conditions are also numerically equivalent to (linear) regression adjustment.
We organize the paper as follows. In Section (ref), we review the estimation problems solved by the IPT weights for the ATE and the ATT and discuss some simple implications. In Section (ref), we derive the equivalences among various IPT-based estimators of the ATE and then the IPT-based estimators of the ATT\@. We emphasize that since we are establishing numerical equivalences, we do not need to, and do not, state the assumptions under which the estimators are consistent. These have been covered elsewhere and are well known. In Section (ref), we discuss the consequences that the algebraic equivalence results have for estimating local average treatment effects with instrumental variables and heterogeneous treatment effects in difference-in-differences settings. In Section (ref), we discuss our empirical applications. In Section (ref), we conclude. Our proofs are provided in the Appendix. In the Supplemental Appendix, we briefly discuss implementation of IPT in R and Stata.
In a general missing data setting, EGP2008 and GPE2012,GPE2016 introduced inverse probability tilting (IPT) as a method for estimating the propensity score, along with other parameters of interest. In this section, we review the estimation problems solved by this method in the binary treatment case.
Let $W$ denote the binary treatment indicator, and define the propensity score as $\mathrm{P} ( W=1 | \mathbf{X} = \mathbf{x} )$. Assume an index model, $p ( \mathbf{x\gamma} )$, where $\mathbf{x}$ is $1 \times K$, $\mathbf{\gamma}$ is $K \times 1$, and $x_{1} \equiv 1$ ensures an intercept. Although $p ( \mathbf{x\gamma} )$ is typically taken to be logit, our results apply more generally to any index model, including probit, complementary log-log, linear, and inverse linear models. Let $Y(0)$ and $Y(1)$ denote the potential outcomes. Recall that the ATE and ATT are defined as
and
However, we emphasize that the results in this paper pertain to algebraic equivalences, and therefore, we do not discuss the identification of population parameters.
To fix ideas, consider the case where $p ( \mathbf{x\gamma} )$ is logit, i.e., $p ( \mathbf{x\gamma} ) = \exp ( \mathbf{x\gamma} ) / \left[ 1 + \exp ( \mathbf{x\gamma} ) \right]$. In practice, the logit model is often conflated with its estimation via maximum likelihood, in which case, for a sample of size $N$, the maximum likelihood estimator $\mathbf{\hat{\gamma}}_{mle}$ solves the first-order condition:
IPT replaces this condition with a different set of moment equations for estimating $\mathbf{\gamma}$. While our discussion below is not limited to the logit case, the point is that even when the logit model is used, estimation need not rely on maximum likelihood; it can proceed via the method of moments instead. When $W$ is a treatment indicator, the IPT moment conditions proposed by EGP2008 and GPE2012 for estimating $\mathrm{E} [ Y(1) ]$ are
which follow immediately by iterated expectations when $p ( \mathbf{X\gamma} ) = \mathrm{P} ( W=1 | \mathbf{X} ) = \mathrm{E} ( W | \mathbf{X} )$. (If we were considering identification, we would need to assume, at a minimum, that $p ( \mathbf{X\gamma} ) >0$ with probability one.) The sample analog of (ref) is
and these equations define the IPT estimator of $\mathbf{\gamma}$, $\mathbf{\hat{\gamma}}_{1,ipt}$, regardless of the specific model chosen for $\mathrm{P} ( W=1 | \mathbf{X} = \mathbf{x} )$. Note that we have put a “1” subscript on $\mathbf{\hat{\gamma}}_{1,ipt}$ because, in the treatment effects setting, there is another set of moment conditions for estimating $\mathrm{E} [ Y(0) ]$ that leads to a different IPT estimator of $\mathbf{\gamma}$. Again, by iterated expectations,
and this leads to the sample analog:
In general, $\mathbf{\hat{\gamma}}_{0,ipt} \neq \mathbf{\hat{\gamma}}_{1,ipt}$. However, because $1 \in \mathbf{X}_{i}$, it follows immediately that
and
These two equations are key, as the summands in (ref) are the weights for estimating $\mathrm{E} [ Y(1) ]$ in IPW estimation, and those in (ref) are the weights used in estimating $\mathrm{E} [ Y(0) ]$, as we review in Section (ref). Equations (ref) and (ref) show that the IPT weights are automatically normalized for estimating the ATE\@. That is, the sample mean of the weights is not stochastic but instead equal to one by construction. These equations also indicate that the IPT estimator of the ATE will require estimating the propensity score twice, with one set of predicted probabilities used to estimate $\mathrm{E} [ Y(1) ]$ and another to estimate $\mathrm{E} [ Y(0) ]$.
For estimating the ATT, the moment equations used by EGP2008 and GPE2016 are
where $\rho = \mathrm{P} ( W=1 )$. Using $\hat{\rho} = N_{1}/N$, where $N_{1}$ is the number of treated units, the $K$ sample moment conditions are
or
where $\mathbf{\bar{X}}_{1} = N_{1}^{-1} \sum_{i=1}^{N} W_{i} \mathbf{X}_{i}$\@. Because $1 \in \mathbf{X}_{i}$, (ref) implies
which implies that the weights used in the IPW estimation of the ATT sum to the number of treated units. In other words, like in the case of the ATE, the IPT weights for estimating the ATT are automatically normalized.
It may seem surprising that we use $\mathbf{\hat{\gamma}}_{0,ipt}$ to denote the IPT estimator of $\mathbf{\gamma}$ in the context of estimating the ATT, given that we used the same notation in equations (ref) and (ref) above. This is fully warranted, however, because this estimator is in fact the same as the IPT estimator defined by equation (ref). Indeed, as shown by Tan2020, the moment conditions in (ref), which balance the weighted covariates of the comparison group with those of the overall sample, are algebraically equivalent to the conditions in (ref), which instead balance them with the treated group, but using a different set of weights. To see this equivalence, we can rewrite the sample moment conditions in (ref) as
which, after simple algebra, can be expressed as
and this, in turn, is easily seen as equivalent to equation (ref).
It is also useful to briefly compare the IPT moment conditions for estimating the ATE with a subsequent proposal by IR2014, known as the “covariate balancing propensity score (CBPS),” which uses different moment conditions to obtain a single estimator of $\mathbf{\gamma}$. Again, if the propensity score is correctly specified then, by iterated expectations,
Rather than using the implications of (ref) separately, which is what IPT does, IR2014 use the second equality to obtain the following sample moment conditions:
After simple algebra, the moment conditions can be expressed as
Comparing (ref) with (ref), we can see that the CBPS approach---when applied to the logit model---weights the MLE moment conditions by the estimated inverse conditional variance, $\mathrm{Var} ( W_{i} | \mathbf{X}_{i} )$.
Because the first element of $\mathbf{X}_{i}$ is unity, (ref) also implies that
Equation (ref) shows that the weights appearing in the IPW estimates of $\mathrm{E} [ Y(1) ]$ and $\mathrm{E} [ Y(0) ]$ sum to the same value, but that common value is not necessarily the sample size, $N$. In other words, when these are used as weights in IPW, the CBPS weights are not automatically normalized.
Finally, when estimating the ATT, IR2014 suggest using the moment conditions in equation (ref), following the approach of EGP2008 and GPE2016. This implies that the IPT and CBPS estimators of the ATT are the same. When using the logit model, as shown by Tan2020, both approaches are also numerically identical to the entropy balancing estimator of Hainmueller2012.\footnote{The three estimators of the ATT will no longer coincide if, in the case of CBPS, the IPT moment conditions are combined with the first-order conditions of the maximum likelihood estimator (the so-called “overidentified CBPS”)\@. Likewise, the entropy balancing estimator will differ from IPT and CBPS if its implementation constrains higher moments of the covariates to be balanced, too.}
In this section, we establish numerical equivalences among three different classes of estimators that incorporate inverse probability weighting, starting with estimators of the ATE.
The three estimators we consider are among the most popular alternatives to OLS estimation of an additive linear model when unconfoundedness is assumed to hold: IPW, AIPW, and IPWRA\@. As we establish algebraic equivalence, we do not impose assumptions beyond those necessary for the existence of estimates for a given sample. This simply means that the estimated propensity scores are strictly between zero and one for all $i$.
The IPW estimator of $\tau _{ate}$ using the IPT weights is
where the subscript “ipt” indicates the use of IPT weights. See, e.g., Wooldridge2010 for a variant of this estimator with MLE-based weights and EGP2008 and GPE2012 for IPT\@. We know from (ref) and (ref) that the weights in both weighted averages are automatically normalized.
The AIPW estimator with IPT weights, which we refer to as AIPT, is also the difference in estimates of $\mu_{1} \equiv \mathrm{E} [ Y(1) ]$ and $\mu_{0} \equiv \mathrm{E} [ Y(0) ]$; that is, $\hat{\tau}_{ate,aipt} = \hat{\mu}_{1,aipt} - \hat{\mu}_{0,aipt}$. For $\mu_{1}$,
where remember that $1 \in \mathbf{X}_{i}$. Although it is not important for the equivalence result, the estimates $\mathbf{\hat{\beta}}_{1}$ typically come from an OLS regression of $Y_{i}$ on $\mathbf{X}_{i}$ using $W_{i}=1$ (treated units). The first term in (ref) is a weighted average of the resulting residuals over the treated units. The weights are exactly those appearing in $\hat{\mu}_{1,ipt}$ and are therefore normalized.\footnote{Normalization is less important in AIPW than in IPW\@. Knaus2024 shows that “unnormalized” AIPW, unlike unnormalized IPW, can still be expressed as a weighted average of observed outcomes with weights that sum to one. This normalization is automatic under standard implementations of outcome regressions.}
For $\mu_{0}$, the AIPT estimator is
where $\mathbf{\hat{\beta}}_{0}$ are probably the OLS estimates from a regression of $Y_{i}$ on $\mathbf{X}_{i}$ using $W_{i}=0$.
The third estimator we consider is the IPWRA estimator with IPT weights, which we refer to as IPTRA\@. For $\mu_{1}$, we first solve a weighted least squares (WLS) problem,
where $\hat{p}_{i} = p ( \mathbf{X}_{i} \mathbf{\hat{\gamma}}_{1,ipt} )$ are the IPT propensity score estimates. Given the WLS estimates $\mathbf{\tilde{\beta}}_{1}$ from (ref), $\mu_{1}$ is estimated by averaging the fitted values across all observations, as in the case of linear regression adjustment:
The IPTRA estimator of $\mu_{0}$, $\hat{\mu}_{0,iptra}$, uses the untreated units with weights $\left( 1-p ( \mathbf{X}_{i} \mathbf{\hat{\gamma}}_{0,ipt} ) \right)^{-1}$, and produces $\mathbf{\tilde{\beta}}_{0}$. The final IPTRA estimator of the ATE is given by $\hat{\tau}_{ate,iptra} = \hat{\mu}_{1,iptra} - \hat{\mu}_{0,iptra} = \mathbf{\bar{X}\tilde{\beta}}_{1} - \mathbf{\bar{X}\tilde{\beta}}_{0}$.
When the inverse probability weights are obtained using MLE, CBPS, or some other method of moments procedure, $\hat{\tau}_{ate,ipw}$, $\hat{\tau}_{ate,aipw}$, and $\hat{\tau}_{ate,ipwra}$ are generally different. In fact, $\hat{\tau}_{ate,ipw}$ and $\hat{\tau}_{ate,aipw}$ do not generally use normalized weights, and so one could have five different estimates using the same estimated weights: IPW, normalized IPW (NIPW), AIPW, normalized AIPW (NAIPW), and IPWRA\@. (IPWRA is always normalized.) Strikingly, when IPT weights are used instead, all of these estimates are identical.
The implication of Proposition (ref) is that if one uses the IPT weights in estimating both $\mu_{0}$ and $\mu_{1}$, where conditional means $\mathrm{E} [ Y(0) | \mathbf{X} ]$ and $\mathrm{E} [ Y(1) | \mathbf{X} ]$ are modeled linearly, then three prominent estimators of the ATE are numerically identical; moreover, the IPW and AIPW versions are automatically normalized.
We now establish the equivalence of several prominent estimators of the ATT when the IPT weights from equation (ref) are used. Recall that
and the first term is always consistently estimated using the sample mean of $Y_{i}$ over the treated units, $\bar{Y}_{1}$. The IPW estimator for the second term, using the IPT weights, is
As noted earlier, the weights in (ref) sum to the number of treated units; thus, they are automatically normalized. The same weights also appear in the AIPW estimator. Therefore, the normalized and unnormalized IPW estimators coincide, as do the normalized and unnormalized AIPW estimators. Specifically, the AIPW estimator of $\mu _{0|1}$ is \begingroup \allowdisplaybreaks
\endgroup where $\hat{p}_{i} = p ( \mathbf{X}_{i} \mathbf{\hat{\gamma}}_{0,ipt} )$ are now the IPT propensity score estimates and $\mathbf{\hat{\beta}}_{0}$ is typically the OLS estimator from regressing $Y_{i}$ on $\mathbf{X}_{i}$ using $W_{i}=0$.
Finally, the IPWRA estimator of $\mu _{0|1}$, using the weights from (ref), is
where $\mathbf{\tilde{\beta}}_{0}$ now solves the WLS problem:
where $\hat{p}_{i} = p ( \mathbf{X}_{i} \mathbf{\hat{\gamma}}_{0,ipt} )$. We have the following equivalence result.
Similar to Proposition (ref), the implication of Proposition (ref) is that if one uses the IPT weights in estimating $\mu_{0|1}$ as well as a linear model for $\mathrm{E} [ Y(0) | \mathbf{X} ]$, then the IPW, AIPW, and IPWRA estimators of the ATT are numerically identical; moreover, the IPW and AIPW versions are automatically normalized.
In this section, we briefly discuss the implications of the results in Section (ref) for estimating local average treatment effects with instrumental variables and heterogeneous treatment effects in difference-in-differences settings.
The results in Section (ref) have implications for estimators of the local average treatment effect (LATE) and the local average treatment effect on the treated (LATT) when using control variables $\mathbf{X}$; a recent treatment is SUW2022, which we follow here. As before, $W$ is a treatment variable. We assume it to be binary, although this can be easily relaxed. We also have a binary instrumental variable, $Z$.
It follows from Frolich2007 that many estimators of the LATE are ratios of estimators of the ATE,
where $\hat{\tau}_{ate,Y|Z}$ is an estimator of the ATE where $Y$ is the outcome, $Z$ plays the role of the treatment, and the covariates $\mathbf{X}$ are used to account for confounders of $Z$\@. Again, we are only concerned with equivalences and not statistical properties. The denominator, $\hat{\tau}_{ate,W|Z}$, is an estimated ATE where $W$ is the outcome and $Z$ again is the treatment indicator, with covariates $\mathbf{X}$. It follows from Proposition (ref) that when linear conditional means are used for both $Y$ and $W$, and IPT is used for the weights, estimators of the LATE based on IPW, AIPW, and IPWRA are all identical. The inverse probability weights, in this case, for both the numerator and the denominator, are based on the instrument propensity score:
It should be noted, however, that unlike in the case of the ATE, where the nature of $Y$ is generally unspecified, here it may be impractical to use the linear model in the denominator when $W$ is binary. See SUW2022 for using other doubly robust estimators to exploit the binary nature of $W$ and maybe special features of $Y$. In such cases, however, the numerical equivalence results no longer hold.
Estimators of the LATT that incorporate control variables $\mathbf{X}$ can be written as the ratio of estimators of the ATT, where the instrument plays the role of the treatment variable:
where $\hat{\tau}_{att,Y|Z}$ and $\hat{\tau}_{att,W|Z}$ are both estimators of the ATT with “treatment” variable $Z$ and outcome variables $Y$ and $W$, respectively. If these estimators use the appropriate IPT weights, as in Proposition (ref), then it follows immediately that the estimators of $\tau_{latt}$ based on IPW, AIPW with linear regression functions, and IPWRA with linear regression functions are all numerically the same. Also, recall that the normalized and unnormalized estimators of the ATT are identical when using these weights.
Some popular estimators in difference-in-differences (DID) settings are based on applying standard treatment effect estimators after suitably transforming the outcome variable. For example, Abadie2005, with two time periods, proposes applying IPW to the differences $Y_{i2}-Y_{i1}$, where $Y_{it}$ is the outcome for unit $i$ in period $t$. Abadie2005 uses maximum likelihood estimation of the propensity score to construct the inverse probability weights. SAZ2020 instead develop a doubly robust estimator of the ATT\@. The estimator uses a structure similar to equation (ref), although SAZ2020 also normalize the weights and replace the OLS estimates of the conditional mean of the untreated potential change in outcomes with the WLS estimates similar to equation (ref); they also use the IPT weights from equation (ref). Strikingly, the results in Section (ref) imply that all these additional modifications have no impact on the final estimate of the ATT when the IPT weights are used; in other words, the simple IPW estimator in Abadie2005 is numerically identical to SAZ2020 when both use the IPT moment conditions to estimate the propensity score.
Similar conclusions hold with many periods and staggered interventions. CSA2021 extend SAZ2020 to estimate ATTs by treatment cohort (i.e., the first period of treatment), $g$, and calendar time, $t$. To estimate these ATTs, $\tau_{gt}$, CSA2021 apply different versions of AIPW and IPWRA to differences $Y_{it}-Y_{i,g-1}$, where $Y_{i,g-1}$ is the outcome in the period just before the first treatment period for treatment cohort $g$. CSA2021 emphasize that the comparison group can either consist of the never treated (NT) cohort or the NT cohort plus other cohorts that are first treated in period $t+1$ or later (“not yet treated”). One of the estimators recommended by CSA2021 is AIPW with the IPT weights from equation (ref). It follows immediately that applying IPW, AIPW, or IPWRA to $Y_{it}-Y_{i,g-1}$ with treatment indicator $D_{ig}$ (indicating treatment cohort) and controls $\mathbf{X}_{i}$ delivers identical estimates of the $\tau_{gt}$ when IPT weights are used. This is true even when $t<g-1$, which provides event study graphs for studying the existence of pre-trends.
An alternative transformation in the staggered intervention settings uses data on all pre-treatment outcomes by removing the average of the outcomes over all pre-treatment periods: $\dot{Y}_{itg} \equiv Y_{it} - \left( g-1 \right) ^{-1} \sum_{s=1}^{g-1} Y_{is}$. As shown in LW2024, under standard no anticipation and conditional parallel trends assumptions, one can apply various treatment effect estimators to the cross-sectional data $\left\{ \left( \dot{Y}_{itg},D_{ig},\mathbf{X}_{i}\right) :i=1,...,N\right\} $ to consistently estimate $\tau _{gt}$. Again, the results in Section (ref) immediately imply that the IPW, AIPW, and IPWRA estimators with the IPT weights from equation (ref) are all identical when applied to these data once one chooses a suitable comparison group.
In this section, we illustrate our findings with two empirical applications, beginning with a replication of a prominent study of the causal effects of cash transfers on longevity AEFLM2016 and concluding with a reanalysis of the empirical application in SAZ2020.
AEFLM2016 study the long-run impacts of the Mothers' Pension (MP) program, which was the first government-sponsored welfare program in the prewar U.S\@. The outcome studied by the authors is the log age at death of children of the program participants. A key strength of the original study is in its careful construction of the comparison group, which consists only of mothers who were initially deemed eligible for participation but were later rejected. Still, Sloczynski2022 argues that some of the conclusions of this paper are not robust to treatment effect heterogeneity.
In our application, we use the same data as AEFLM2016 and Sloczynski2022. We consider three covariate specifications and two sources of information on dates of death: program records and death certificates. In our first specification, we control for cohort and state fixed effects. In our second specification, in line with AEFLM2016, we replace state fixed effects with a battery of individual, county, and state characteristics. In our final specification, we reintroduce state fixed effects without dropping any other covariates.\footnote{The final specification in AEFLM2016 uses county rather than state fixed effects but is otherwise identical. In our application, using county fixed effects is not feasible because, in several counties, every eligible applicant was treated, resulting in a failure of overlap.}
Table (ref) reports a number of estimates of the effects of cash transfers on longevity. As in AEFLM2016, the OLS estimates from an additive model, which controls for program participation and covariates but not interactions between the two, strongly suggest that cash transfers positively influenced the longevity of the children of their beneficiaries. However, in line with the replication in Sloczynski2022, the majority of the estimates of the ATE and ATT are smaller or much smaller than the OLS estimates and not statistically significant. At the same time, there is a clear difference between the two panels of Table (ref) that report weighting and doubly robust estimators based on MLE and IPT weights. In the case of MLE, there is a wide variation in estimates, which range from 0.0014 to 0.0597 for the ATE and from --0.0014 to 0.0645 for the ATT\@. Conditional on choosing a specific covariate specification, the choice of an estimator can have a profound impact on the researcher's conclusion. On the other hand, in the case of IPT, this choice is entirely inconsequential, as there are no differences across estimators conditional on a particular specification choice. This illustrates Propositions (ref) and (ref), which demonstrate the underlying numerical equivalences. Moreover, in the case of IPT, the standard errors are usually slightly smaller than in the case of the corresponding estimates based on MLE\@.
A large number of papers, originating with LaLonde1986, combine experimental data from the evaluation of the National Supported Work (NSW) program with a nonexperimental comparison group from the Current Population Survey (CPS) or the Panel Study of Income Dynamics (PSID)\@. The premise of this literature is that a successful nonexperimental estimation method should closely replicate the experimental estimate of the effect of the program when combining the original treatment group with an artificial comparison group LaLonde1986,DW1999 or the “effect” of zero when combining the latter with the original control group ST2005.
In our application, we closely follow a recent reanalysis of these data in SAZ2020, who restrict their attention to samples combining the CPS comparison group and variants of the original control group, previously analyzed by LaLonde1986, DW1999 (the “DW” sample), and ST2005 (the “early RA” sample). As is standard in this literature, the outcome of interest is real earnings in 1978. Because SAZ2020 focus on various difference-in-differences estimators, they often use the transformed outcome, $Y_{i2}-Y_{i1}$; here, this is equal to the difference between real earnings in 1978 and real earnings in 1975. The baseline covariates include age, years of education, real earnings in 1974, and indicator variables for less than 12 years of education, being married, being Black, and being Hispanic. Other covariate specifications also include additional higher-order and interaction terms.
Table (ref) replicates every estimate and standard error in SAZ2020's Table 3, while also reporting a number of additional results.\footnote{“TWFE” corresponds to $\hat{\tau}^{fe}$ in SAZ2020. This is the OLS estimate from a panel data specification with real earnings in 1975 and 1978, a “treatment” indicator, and unit and year fixed effects. “RA” corresponds to $\hat{\tau}^{reg}$ in SAZ2020. IPW with MLE weights corresponds to $\hat{\tau}^{ipw,p}$. NIPW with MLE weights corresponds to $\hat{\tau}^{ipw,p}_{std}$. NAIPW with MLE weights corresponds to $\hat{\tau}^{dr,p}$. Finally, all the estimates in the “IPT weights” panel are identical to SAZ2020's preferred estimator, $\hat{\tau}^{dr,p}_{imp}$.} The bottom line is that weighting and doubly robust estimators perform quite well in replicating the true effect of zero; except for the simplest covariate specification applied to the LaLonde1986 sample, none of these estimates are significantly different from the true effect. In addition, SAZ2020 argue that their preferred estimator based on IPT weights tends to have smaller standard errors than estimators based on MLE weights. While this is true, Table (ref) also illustrates our previous point that SAZ2020's (SAZ2020) estimator might be unnecessarily complex; when using the IPT weights, even the simplest “unnormalized” IPW estimator is numerically equivalent to it. Our standard errors, obtained together with the point estimates using our companion Stata package, teffects2, are also identical to those reported by SAZ2020.
Applied researchers face many ways to adjust for covariates, but in this paper, we show that several popular estimators are in fact identical under a simple condition. Specifically, our results assume that the propensity score is estimated using inverse probability tilting, a method of moments approach developed by EGP2008 and GPE2012,GPE2016. Estimators based on or equivalent to this approach have already become popular in difference-in-differences settings SAZ2020 and outside economics Hainmueller2012,IR2014. Our results, simplifying the set of alternative estimators available to researchers, offer a novel rationale for adopting this approach in various contexts, such as under unconfoundedness and in instrumental variables and difference-in-differences settings.
\onehalfspacing