Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
91,466 characters · 15 sections · 201 citation commands
Abadie's Kappa and Weighting Estimators of the Local Average Treatment Effect
\singlespacing
\setcounter{page}{2} \doublespacing
The validity of many instrumental variables, as applied in economics and related fields, requires conditioning on additional covariates. In such cases empirical researchers often approximate the causal effects of interest using additive linear models and two-stage least squares (2SLS) estimation. However, recent work by Sloczynski2018,Sloczynski2021 and BBMT2022 questions the general validity of this approach and, in particular, the ability of the 2SLS estimand to uncover the local average treatment effect (LATE), that is, the average effect of treatment for “compliers,” as defined by IA1994 and AIR1996. One concern is that covariate specifications used by empirical researchers are insufficiently flexible BBMT2022. Another concern is that even when they are flexible, the 2SLS estimand does not generally correspond to the LATE or any other parameter of interest Sloczynski2018,Sloczynski2021.
In this paper we study a class of simple yet flexible weighting estimators of the LATE, which are robust to the aforementioned limitations of 2SLS\@. The estimators we consider can be motivated by the identification result in Abadie2003, which applies to any parameter defined in terms of moments of the joint distribution of the data for compliers, including the LATE\@. The result in Abadie2003 is based on “kappa weighting,” with weights that depend on the instrument propensity score. Some of the estimators we consider can alternatively be motivated by the identification result in Frolich2007, which suggests a simple approach to estimating the LATE using the ratio of two conventional weighting estimators. Although the recent literature in econometrics and statistics has adopted this approach, it focuses primarily on the ratio of two unnormalized weighting estimators Tan2006,Frolich2007,MCH2011,DHL2014a,DHL2014b,AANP2017, despite the fact that the lack of normalization leads to poor finite sample properties in related contexts Imbens2004,MT2009,BDM2014. Here, normalization means rescaling the weights so that they sum to one in each sample.
In this paper we unify and provide a comprehensive treatment of the two approaches to constructing weighting estimators of the LATE\@. We begin with an observation that the existing identification results enable the construction of multiple consistent estimators of the LATE, only two of which are normalized. One normalized estimator is the sample analogue of a particular expression in AC2018, based on Abadie2003. However, it is also straightforward, as in UysalDiss, to construct a normalized version of Tan2006's Tan2006 and Frolich2007's Frolich2007 estimator and to interpret it through the lens of “kappa weighting.” We argue that these two normalized estimators are likely to dominate the unnormalized weighting estimators of the LATE in many cases. Unlike most other papers that stress the importance of normalization, we also provide an objective and intuitively appealing criterion that differentiates the normalized from the unnormalized estimators; see also Tille1998 and AM2013. Indeed, we demonstrate that the former class of estimators, unlike the latter, satisfies the properties of (i) translation invariance and (ii) scale invariance with respect to the natural logarithm. This ensures that the normalized estimators are not sensitive to the centering of the outcome variable or, when estimating the LATE in logs, to the units of measurement of the untransformed outcome CR2023.
We also identify an important context, namely settings with one-sided noncompliance, in which certain estimators have an additional advantage: they are based on a denominator that is strictly greater than zero by construction. This is the case for (i) Tan2006's Tan2006 and Frolich2007's Frolich2007 unnormalized estimator whenever there are no always-takers, that is, individuals who participate in the treatment regardless of the value of the instrument; (ii) a different unnormalized estimator whenever there are no never-takers, that is, individuals who never participate in the treatment; and (iii) the normalized estimator originally proposed by UysalDiss in both of these cases. We recommend this last estimator for wider use in practice.
Our observations about translation and scale invariance as well as settings with one-sided noncompliance apply equally when the instrument propensity score is known and when it is estimated using standard methods. In practice, the instrument propensity score is rarely known, and its estimation can greatly influence the properties of the final estimator of the LATE\@. We consider maximum likelihood and covariate balancing estimation of the instrument propensity score, where the latter approach follows GPE2012,GPE2016, IR2014, Heiler2022, and SASX2022, among others. Either approach is compatible with the construction of the estimator in UysalDiss, and when appropriate covariate balancing propensity scores are used, this estimator is also equivalent to Heiler2022's Heiler2022.
Aside from the finite sample properties of weighting estimators of the LATE, we also study their asymptotic properties in a unified framework of M-estimation. Under standard regularity conditions, our weighting estimators are asymptotically normal, and we derive their asymptotic variances. To illustrate our findings, we also use three empirical applications and a simulation study. The simulations confirm the very good relative performance of our preferred normalized estimator, especially with covariate balancing propensity scores, which appear to be more robust to misspecification than their maximum likelihood counterparts.
Our empirical applications focus on causal effects of military service Angrist1990, college education Card1995, and childbearing AE1998. In each of these cases, we document what we regard as superiority of normalized weighting. The bottom line is that unnormalized estimators are very sensitive to how the outcome variable is coded. In each application, the estimates are sensitive to the units of measurement (cents, dollars, \$1,000s, \$100,000s) of the income variable prior to the log transformation. In our replication of AE1998, we also consider labor force participation as a binary outcome, and we document that unnormalized estimators are highly sensitive to whether working for pay is coded as, say, 1 or 0.
Our application of weighting to estimate the LATE appears to be somewhat rare in practice, although Abadie2003's Abadie2003 result is more commonly used to estimate mean characteristics of compliers, as also recommended by AP2009. We analyze two samples of applications of instrumental variables to verify this claim. First, our reading of the 30 papers replicated by Young2022, each of which uses 2SLS, suggests that none of these papers uses weighting estimators of the LATE or applies Abadie2003's Abadie2003 result for any other purpose. Second, we have also examined whether any of the papers published in journals of the American Economic Association in 2019 and 2020 consider weighting estimators of the LATE\@. Our best assessment is that the answer is likewise negative. Still, MT2019, GGS2020, LOL2020, and LVRS2020 apply Abadie2003's Abadie2003 result to estimate mean characteristics of compliers, while Cohodes2020 uses this result to estimate the control complier mean (CCM), a parameter introduced by KKL2001. In this paper we argue that “kappa weighting” can also be used more widely as a flexible alternative to 2SLS, and we provide a practical guide to using this method to estimate the LATE\@.
The remainder of the paper is organized as follows. Section (ref) introduces our framework. Section (ref) provides our theoretical results on estimation and inference. Section (ref) illustrates our results with three empirical applications. Section (ref) discusses our simulation study. Section (ref) concludes. Proofs and derivations are collected in the appendix unless noted otherwise. The estimators considered in this paper are also implemented in the companion Stata package kappalate.
Our framework broadly follows Abadie2003. Let $Y$ denote the outcome variable of interest, $D$ the binary treatment, and $Z$ the binary instrument for $D$. We also introduce a vector of observed covariates, $X$, that predict $Z$\@. The instrument propensity score is written as $p(X) = \pr (Z=1 \mid X)$.
There are two potential outcomes, $Y_{1}$ and $Y_{0}$, only one of which is observed for a given individual, $Y = D \cdot Y_{1} + \left( 1-D \right) \cdot Y_{0}$. Similarly, there are two potential treatments, $D_{1}$ and $D_{0}$, and it is $Z$ that determines which of them is observed, $D = Z \cdot D_{1} + \left( 1-Z \right) \cdot D_{0}$. It will also be useful to include $Z$ in the definition of potential outcomes, letting $Y_{zd}$ denote the potential outcome that a given individual would obtain if $Z=z$ and $D=d$.
AIR1996 divide the population into four mutually exclusive subgroups based on the latent values of $D_{1}$ and $D_{0}$. Individuals with $D_{1} = D_{0} = 1$ are referred to as always-takers, as they get treatment regardless of whether they are encouraged to do so or not; similarly, individuals with $D_{1} = D_{0} =0$ are referred to as never-takers. Individuals with $D_{1} = 1$ and $D_{0} = 0$ are referred to as compliers, as they comply with their instrument assignment; they get treatment if they are encouraged to do so but not otherwise. Analogously, individuals with $D_{1} = 0$ and $D_{0} = 1$ are referred to as defiers, as they defy their instrument assignment.
As usual, we define the treatment effect as the difference in the outcomes with and without treatment, $Y_{1} - Y_{0}$. Following IA1994, a large literature has focused on identification and estimation of the local average treatment effect (LATE), defined as
i.e. as the average treatment effect for compliers or, in other words, for those individuals who would be induced to get treatment by the change in $Z$ from zero to one.
Next, we review a general identification result due to Abadie2003, which we will use, in turn, to discuss identification of $\late$. We begin by restating Abadie2003's Abadie2003 assumptions.
These assumptions are standard in the recent literature. Assumption (ref)(i) states that, conditional on covariates, the instrument is “as good as randomly assigned.” Assumption (ref)(ii) implies that the instrument only affects the outcome through its effect on treatment status; it follows that $Y_{0} = Y_{10} = Y_{00}$ and $Y_{1} = Y_{11} = Y_{01}$. Assumption (ref)(iii) combines an overlap condition with a requirement that the instrument affects the conditional probability of treatment. Finally, Assumption (ref)(iv) rules out the existence of defiers, and implies that the population consists of always-takers, never-takers, and compliers. Under Assumption (ref), as demonstrated by Abadie2003, any feature of the joint distribution of $\left( Y,D,X \right)$, $\left( Y_{0},X \right)$, or $\left( Y_{1},X \right)$ is identified for compliers.
Both Abadie2003 and the subsequent applied literature have focused on the implications of Lemma (ref)(a). On the other hand, Lemma (ref)(b) and (c) have been used in the econometrics literature to identify and estimate $\late$ and quantile treatment effects FM2013,AC2018,SASX2022,SS2024.
To see how Lemma (ref)(b) and (c) identifies $\late$, take $g_0(Y_{0},X) = Y_{0}$ and $g_1(Y_{1},X) = Y_{1}$, and write:
We can also rewrite equation ((ref)) to obtain the following expression for $\late$:
As we will see later, it is useful to treat equations ((ref)) and ((ref)) as distinct. In any case, it is clear that $\late$ is identified as long as $\pr(D_{1} > D_{0})$ is identified. As noted by Abadie2003, Lemma (ref)(a) implies that $\pr(D_{1} > D_{0}) = \e(\kappa)$, which follows from taking $g(Y,D,X) = 1$. Similarly, however, we can use Lemma (ref)(b) and (c) to obtain $\pr(D_{1} > D_{0}) = \e(\kappa_{1})$ and $\pr(D_{1} > D_{0}) = \e(\kappa_{0})$. This is not a novel observation but we will provide a more comprehensive discussion of its consequences than has been done in previous work. We conclude this section with the following remark.
The proof of Remark (ref) follows from simple algebra and is omitted. The facts that $\e \left[ \frac{Z - p(X)}{p(X)} \right] = 0$ and $\e \left[ \frac{Z - p(X)}{p(X) \left( 1 - p(X) \right)} \right] = 0$ hold by iterated expectations. It follows that $\e(\kappa) = \e(\kappa_{1}) = \e(\kappa_{0})$. Additionally, Lemma (ref) implies that each of these objects identifies $\pr(D_{1} > D_{0})$.
In this section we study estimation and inference for $\late$. We begin by introducing our preferred weighting estimator of this parameter. Then, we develop the argument in favor of this estimator, beginning with the case where $p(X)$ is known and later explaining how $p(X)$ can be estimated when it is not known. While $p(X)$ is rarely known in practice, our novel insights in Sections (ref) and (ref) apply equally in that case and when $p(X)$ is estimated using standard methods.
Given a random sample $\big\{ (D_{i},Z_{i},X_{i},Y_{i}):i=1,\ldots,N \big\}$, and assuming that the instrument propensity score is known, our recommended weighting estimator of $\late$ can be written as:
This estimator was proposed by UysalDiss, and is easily implementable as a function of six sample means. It is also implementable as the coefficient on $D$ in a weighted IV regression of $Y$ on $D$, with $Z$ as the instrument and weights equal to $\frac{Z}{p(X)} + \frac{1-Z}{1-p(X)}$. When the instrument propensity score is not known, a possibility we consider explicitly in Sections (ref) and (ref), we would adopt a parametric model for $p(X)$, $F(X,\alpha)$, estimate the unknown parameters by an appropriate method, and replace the instrument propensity scores in equation ((ref)) with their estimates, $\hat{p}(X) = F(X,\hat{\alpha})$. The leading model for $p(X)$ is logit, $F(X,\alpha) = \exp(X\alpha)/[1+\exp(X\alpha)]$, and the natural estimation methods are maximum likelihood and covariate balancing. Appropriate covariate balancing approaches include those in GPE2012,GPE2016 and IR2014, both of which would lead to simple method of moments estimators of $\alpha$. We defer further details on estimation of $\alpha$ to Section (ref). Note that $\hat{\tau}_{u}$ with covariate balancing propensity scores is also recommended by Heiler2022 but we are the first to determine its advantages given in the analysis below.
Recent software implements $\hat{\tau}_{u}$ in R and Stata. Specifically, BH2018 implement this estimator in their causalweight package in R, although covariate balancing estimation of $\alpha$ is not currently supported and inference is based on the bootstrap. Our companion Stata package kappalate implements $\hat{\tau}_{u}$ and other weighting estimators, and we allow both maximum likelihood and covariate balancing estimation of $\alpha$, as well as computation of analytical standard errors. The package is downloadable from the Statistical Software Components (SSC) Archive.
Two further comments about $\hat{\tau}_{u}$ are in order. First, this is our preferred member of the class of weighting estimators, but there are other classes of estimators one may be willing to consider. One such class is doubly robust estimators, which combine weighting and models for conditional expectations of $Y$ and $D$\@. Doubly robust estimators of $\late$ have been developed by Tan2006, UysalDiss, ORR2015, BCFVH2017, SUW2022, MSSU2023, and others. In this paper, however, we restrict our attention to the class of weighting estimators.
Second, a prototypical weighting or doubly robust estimator, such as $\hat{\tau}_{u}$, might be poorly behaved when some instrument propensity scores are close to 0 or 1 KT2010, even if Assumption (ref) is not violated. In this scenario, usually referred to as “limited” or “weak” overlap, it might be preferable to use estimators of $\late$ that were designed to alleviate this problem, such as those in Hongetal2020 and MSSU2023. See also CH2016, Rothe2017, MW2020, HK2021, and SU2022 for settings with limited overlap and exogenous $D$, as well as LDDFS2021 and MSW2022 for formal statistical tests of limited overlap.
In this section we introduce several seemingly intuitive weighting estimators of $\late$, which we will later show to have some undesirable finite sample properties. For now, we continue to assume that the instrument propensity score is known. In this case, equation ((ref)) suggests that we can consistently estimate $\late$ as follows:
where $\hat{\pr}(D_{1} > D_{0}) \overset{p}{\rightarrow} \pr(D_{1} > D_{0}) > 0$. Our discussion in Section (ref) also implies that there are at least three candidate estimators for $\pr(D_{1} > D_{0})$, namely $N^{-1}\sum_{i=1}^{N} \kappa_{i}$, $N^{-1}\sum_{i=1}^{N} \kappa_{i1}$, and $N^{-1}\sum_{i=1}^{N} \kappa_{i0}$, where $\kappa_{i} = 1 - \frac{D_{i} \left( 1 - Z_{i} \right)}{1 - p(X_{i})} - \frac{\left( 1 - D_{i} \right) Z_{i}}{p(X_{i})}$, $\kappa_{i1} = D_{i} \frac{Z_{i} - p(X_{i})}{p(X_{i}) \left( 1 - p(X_{i}) \right)}$, and $\kappa_{i0} = \left( 1 - D_{i} \right) \frac{\left( 1 - Z_{i} \right) - \left( 1 - p(X_{i}) \right)}{p(X_{i}) \left( 1 - p(X_{i}) \right)}$. Consequently, we have the following consistent estimators of $\late$: \begingroup \allowdisplaybreaks
\endgroup One might mistakenly expect that the choice of the estimator for $\pr(D_{1} > D_{0})$ is largely inconsequential. We discuss this issue extensively in what follows. For now, it should suffice to note that $N^{-1}\sum_{i=1}^{N} \frac{Z_{i} - p(X_{i})}{p(X_{i})}$ and $N^{-1}\sum_{i=1}^{N} \frac{Z_{i} - p(X_{i})}{p(X_{i}) \left( 1 - p(X_{i}) \right)}$ are not generally equal to zero or to each other, and hence $N^{-1}\sum_{i=1}^{N} \kappa_{i}$, $N^{-1}\sum_{i=1}^{N} \kappa_{i1}$, and $N^{-1}\sum_{i=1}^{N} \kappa_{i0}$ will also generally be different, unlike their population counterparts (cf. Remark (ref)).
Lemma (ref) is not the only identification result that allows us to construct consistent estimators of the LATE\@. An alternative result is provided by Frolich2007. An implication of this result is that the ratio of any consistent estimator of the average treatment effect (ATE) of $Z$ on $Y$ and any consistent estimator of the ATE of $Z$ on $D$ is consistent for the LATE\@. Given our interest in weighting estimators, a natural candidate estimator is
as suggested by Tan2006 and Frolich2007. This estimator is equal to the ratio of two weighting estimators of the ATE of $Z$ (on $Y$ and $D$) under unconfoundedness HIR2003. The following remark, which has not been precisely stated in previous work, clarifies the relationship between $\hat{\tau}_{t}$ and the other estimators introduced above.
Remark (ref) states that $\hat{\tau}_{t}$ and $\hat{\tau}_{a,1}$ are numerically identical, which can be seen by plugging in the expression for $\kappa_{i1}$ into equation ((ref)):
As is easy to see, expressions ((ref)) and ((ref)) are equivalent. It is also important to note that $\hat{\tau}_{t}$ ($= \hat{\tau}_{a,1}$), or at least its variant where $p(X)$ is estimated, is by far the most popular weighting estimator of the LATE in the econometrics literature. It has been considered by Tan2006, Frolich2007, MCH2011, DHL2014a,DHL2014b, and AANP2017, among others. As we will see in the next section, however, this estimator has a major drawback in practice.
Following Imbens2004, MT2009, and BDM2014, it is widely understood that weighting estimators of the ATE under unconfoundedness should be normalized, i.e. their weights should sum to unity, an idea that is often attributed to Hajek1971. More recently, KU2023 provide a general treatment of normalization under unconfoundedness while SAZ2020 and CSA2021 stress the importance of normalization in difference-in-differences methods. It is natural to expect that normalization will also be important when estimating the LATE Heiler2022.
It follows immediately that $\hat{\tau}_{t}$ is likely inferior to the ratio of two normalized, Hajek1971-type estimators of the ATE of $Z$ under unconfoundedness:
This estimator, first proposed by UysalDiss, was introduced in equation ((ref)) as our preferred estimator. It might not be immediately obvious how the importance of normalization affects our understanding of $\hat{\tau}_{a}$, $\hat{\tau}_{a,1}$, and $\hat{\tau}_{a,0}$. To see this, note that these estimators can equivalently be represented as sample analogues of equation ((ref)): \begingroup \allowdisplaybreaks
\endgroup None of these estimators is normalized. First, $\hat{\tau}_{a}$ uses weights of $\left[ \sum_{i=1}^{N} \kappa_{i} \right]^{-1} \kappa_{i1}$ and $\left[ \sum_{i=1}^{N} \kappa_{i} \right]^{-1} \kappa_{i0}$, which do not necessarily sum to unity across $i$. Second, $\hat{\tau}_{a,1}$ is based on weights of $\left[ \sum_{i=1}^{N} \kappa_{i1} \right]^{-1} \kappa_{i1}$, which are properly normalized, and $\left[ \sum_{i=1}^{N} \kappa_{i1} \right]^{-1} \kappa_{i0}$, which are not. Finally, $\hat{\tau}_{a,0}$ uses weights of $\left[ \sum_{i=1}^{N} \kappa_{i0} \right]^{-1} \kappa_{i1}$, which do not necessarily sum to unity across $i$, and $\left[ \sum_{i=1}^{N} \kappa_{i0} \right]^{-1} \kappa_{i0}$, which are properly normalized.
It is straightforward to construct a normalized estimator based on equation ((ref)). To do this, the two denominators need to be estimated separately, using different estimators of $\pr(D_{1} > D_{0})$, $N^{-1}\sum_{i=1}^{N} \kappa_{i1}$ and $N^{-1}\sum_{i=1}^{N} \kappa_{i0}$. The resulting estimator becomes
where both sets of weights, $\left[ \sum_{i=1}^{N} \kappa_{i1} \right]^{-1} \kappa_{i1}$ and $\left[ \sum_{i=1}^{N} \kappa_{i0} \right]^{-1} \kappa_{i0}$, are properly normalized. This estimator has been considered by AC2018 and SASX2022. While the literature on quantile treatment effects studies normalized kappa weighting estimators somewhat more often FM2013, the importance of normalization is not explicitly recognized. Interestingly, if the goal is to estimate $\e \left( X \mid D_{1} > D_{0} \right)$ rather than $\late$ or quantile treatment effects, as in MT2019, GGS2020, LOL2020, and LVRS2020, among others, then three normalized estimators of this object can readily be constructed: $\left[ \sum_{i=1}^{N} \kappa_{i} \right]^{-1} \sum_{i=1}^{N} \kappa_{i} X_i$, $\left[ \sum_{i=1}^{N} \kappa_{i0} \right]^{-1} \sum_{i=1}^{N} \kappa_{i0} X_i$, and $\left[ \sum_{i=1}^{N} \kappa_{i1} \right]^{-1} \sum_{i=1}^{N} \kappa_{i1} X_i$.
It should also be noted that $\hat{\tau}_{u}$ can likewise be interpreted as a normalized “Abadie” or “kappa weighting” estimator. To see this, note that $N^{-1} \sum_{i=1}^{N} \frac{Z_{i}}{p(X_{i})} \overset{p}{\rightarrow} 1$ and $N^{-1} \sum_{i=1}^{N} \frac{1 - Z_{i}}{1 - p(X_{i})} \overset{p}{\rightarrow} 1$. This implies that $\hat{\tau}_{u} \overset{p}{\rightarrow} \frac{\e \left[ \frac{YZ}{p(X)} \right] - \e \left[ \frac{Y\left( 1 - Z \right)}{1 - p(X)} \right]}{\e \left[ \frac{DZ}{p(X)} \right] - \e \left[ \frac{D\left( 1 - Z \right)}{1 - p(X)} \right]} = \frac{\e \left[ Y \frac{Z - p(X)}{p(X) \left( 1 - p(X) \right)} \right]}{\e(\kappa_{1})}$, which is the same as the expression for $\late$ in equation ((ref)), subject to $\pr(D_{1} > D_{0}) = \e(\kappa_{1})$.
So far, we have made it seem obvious that weighting estimators should be normalized. Yet, it is natural to ask: Why is it so important that weights sum to unity? Many of the recommendations to date are based on simulation results MT2009,BDM2014, and it is not clear to what extent such evidence should guide estimator choice AKS2019. In what follows, we provide an objective and intuitively appealing criterion that differentiates the normalized from the unnormalized estimators.
To present our criterion, we need to introduce some additional notation. Let $\mathbf{Y}$ be a column vector of observed data on outcomes and $\mathbf{W} = \left( \mathbf{D} \; \mathbf{Z} \; \mathbf{X} \right)$ be a matrix of observed data on the remaining variables, namely the treatment status, the instrument, and the covariates. We postulate that any reasonable estimator of $\late$ should be translation invariant.
The property of translation invariance is defined as the invariance of an estimator to an additive change of the outcome values for all units by a fixed amount. Put differently, estimators that are not translation invariant will generally depend on how the outcome variable is centered. If this variable is binary, the estimate may change when we relabel the zeros and ones, on top of the obvious sign change that is due to relabeling. If the outcome is a logarithm of some other variable, the estimator is also not invariant to scale transformations of that variable.
The property of scale equivariance, if satisfied by a given estimator, gives a guarantee that a broad class of multiplicative, power, and additive transformations of the outcome data can only lead to specific, intuitively sensible changes in the final estimate. An important special case of scale equivariance is scale invariance with respect to the natural logarithm, which follows from setting $\alpha_1 \to 0$, $\alpha_2 = 1/\alpha_1$, and $\alpha_3 = \alpha_2$ in Definition (ref)\@. To be clear, the idea here is as follows: the researcher transforms the outcome data prior to analysis, perhaps because they want to interpret the estimates as percentages, in which case they would use $g(Y) = \log(Y)$; however, if their estimator is not scale invariant with respect to the natural logarithm, the resulting estimates will depend on the units of $Y$, which directly contradicts the idea of interpreting them as percentages.
The following result demonstrates that the unnormalized weighting estimators discussed so far are not translation invariant and not scale equivariant. Thus, they are also not scale invariant with respect to the natural logarithm. On the other hand, the normalized estimators, $\hat{\tau}_{u}$ and $\hat{\tau}_{a,10}$, satisfy the properties of translation invariance and scale equivariance, which means that they are also scale invariant with respect to $g(Y) = \log(Y)$\@.
The properties of translation invariance and scale equivariance are very appealing, and it makes intuitive sense to only use estimators that satisfy them. To conclude this section, we make three final observations. First, the point of Proposition (ref) is similar but distinct from that of CR2023, who focus on the sensitivity to scaling of $\log(1+Y)$ and similar transformations, and do not restrict their attention to any specific estimators (including weighting). Unlike in CR2023, the problem we describe disappears in large samples. On the other hand, the problem described by CR2023 disappears when the outcome only assumes strictly positive values, which is not the case in Proposition (ref). Second, it is useful to note that doubly robust estimators of $\late$, which we mentioned briefly in Section (ref), are generally translation invariant and scale equivariant, subject to mild conditions on the outcome model. Finally, several previous papers, including Tille1998 and AM2013, note that the usual unnormalized weighting estimator is not translation invariant in settings with exogenous $D$\@. We extend this result to a class of weighting estimators of the LATE and additionally examine the more general property of scale equivariance.
Weighting estimators of $\late$, like two-stage least squares and many other IV methods, are an example of ratio estimators. A common problem with such estimators is that they behave badly if their denominator is close to zero ASS2019. In this section we document that in settings with one-sided noncompliance, i.e. when units with $Z=1$ or units with $Z=0$ fully comply with their instrument assignment, there is a choice of weighting estimators that have an important advantage: they are based on a denominator that is strictly greater than zero by construction.
To see this, note that Table (ref) provides simplified formulas for $\kappa$, $\kappa_{1}$, and $\kappa_{0}$ in each of the four subpopulations defined by their values of $Z$ and $D$. For example, $\kappa = 1$ if $Z=1$ and $D=1$ or $Z=0$ and $D=0$; moreover, $\kappa = - \frac{1 - p(X)}{p(X)}$ if $Z=1$ and $D=0$, and $\kappa = - \frac{p(X)}{1 - p(X)}$ if $Z=0$ and $D=1$. It follows that $N^{-1}\sum_{i=1}^{N} \kappa_{i}$ is the mean of a collection of positive and negative values, and hence it can be positive, negative, or zero. This is despite the fact that $N^{-1}\sum_{i=1}^{N} \kappa_{i}$ is also a consistent estimator of $\pr(D_{1} > D_{0})$, which is strictly positive under Assumption (ref)\@. Similarly, $N^{-1}\sum_{i=1}^{N} \kappa_{i1}$ and $N^{-1}\sum_{i=1}^{N} \kappa_{i0}$ are also not guaranteed to be positive in general.
However, the situation is different in settings with one-sided noncompliance. If all individuals with $Z=1$ get treatment or, equivalently, there are no never-takers, the second row of Table (ref) is empty and $\pr (\kappa_{0} \geq 0) = 1$. This is the case, for example, in studies that use twin births as an instrument for fertility AE1998. Similarly, if there are no always-takers, then $\pr (\kappa_{1} \geq 0) = 1$. This is the case, for example, in randomized trials with noncompliance that make it impossible to access treatment if not offered. An implication of these observations is that in settings with one-sided noncompliance there exist estimators of $\pr(D_{1} > D_{0})$, and perhaps also the LATE, that have some desirable properties in finite samples.
Proposition (ref) demonstrates that settings with one-sided noncompliance offer a choice of estimators of $\pr(D_{1} > D_{0})$, based on $\kappa_{1}$ and $\kappa_{0}$, that are strictly greater than zero by construction. Interestingly, the denominator of $\hat{\tau}_{u}$ is also strictly greater than zero when noncompliance is one sided, and this is true regardless of whether there are no always-takers or no never-takers.
An implication of Propositions (ref) and (ref) is that certain weighting estimators have the advantage of avoiding near-zero denominators when noncompliance is one sided. There are two unnormalized estimators that have this property, $\hat{\tau}_{a,1}$ when there are no always-takers and $\hat{\tau}_{a,0}$ when there are no never-takers, and one normalized estimator, $\hat{\tau}_{u}$, which retains this property in both cases. The other normalized estimator, $\hat{\tau}_{a,10}$, does not generally share this property with $\hat{\tau}_{u}$. Indeed, if $N^{-1}\sum_{i=1}^{N} \kappa_{i1}$ is away from zero but $N^{-1}\sum_{i=1}^{N} \kappa_{i0}$ is not, then this may affect the performance of not only $\hat{\tau}_{a,0}$ but also $\hat{\tau}_{a,10}$. Likewise, if $N^{-1}\sum_{i=1}^{N} \kappa_{i1}$ is close to zero, then both $\hat{\tau}_{a,1}$ and $\hat{\tau}_{a,10}$ are affected.
Our discussion in Sections (ref), (ref), and (ref) assumed that $p(X)$ is known, which is often unrealistic. In practice, researchers typically adopt a parametric model for $p(X)$, say the logit, $F(X,\alpha) = \exp(X\alpha)/[1+\exp(X\alpha)]$, and estimate $\alpha$ by maximum likelihood (cf. Section (ref)). Our observations above apply equally in this case. Indeed, the normalized estimators are translation invariant and scale equivariant while the unnormalized estimators are not. At the same time, two specific unnormalized estimators and one normalized estimator avoid near-zero denominators in settings with one-sided noncompliance. From now on, if we wish to specify that $\alpha$ is estimated using maximum likelihood, we use an “ml” subscript or superscript. Thus, $\hat{\alpha}_{ml}$ is the maximum likelihood estimator of $\alpha$, $\hat{p}_{ml}(X) = F(X,\hat{\alpha}_{ml})$ are the estimated propensity scores, and $\hat{\tau}_{u}^{ml}$, $\hat{\tau}_{a,10}^{ml}$, $\hat{\tau}_{a}^{ml}$, $\hat{\tau}_{t}^{ml}$ ($= \hat{\tau}_{a,1}^{ml}$), and $\hat{\tau}_{a,0}^{ml}$ are the analogues of the previously introduced estimators, with $\hat{p}_{ml}(X)$ replacing $p(X)$.
Alternatively, we can estimate $\alpha$ using covariate balancing methods, such as those studied by GPE2012,GPE2016, IR2014, Heiler2022, and SASX2022. Following Heiler2022, we focus on the approach of IR2014, which amounts to estimating $\alpha$ using a different set of moment conditions than maximum likelihood. Indeed, the population moment conditions in IR2014 are
and the corresponding sample moment conditions can be written as
where $\hat{\alpha}_{cb}$ is the method of moments estimator of $\alpha$. We also use $\hat{p}_{cb}(X) = F(X,\hat{\alpha}_{cb})$ to denote the covariate balancing propensity scores, and $\hat{\tau}_{u}^{cb}$, $\hat{\tau}_{a,10}^{cb}$, $\hat{\tau}_{a}^{cb}$, $\hat{\tau}_{t}^{cb}$ ($= \hat{\tau}_{a,1}^{cb}$), and $\hat{\tau}_{a,0}^{cb}$ to denote the analogues of the previously introduced estimators, with $\hat{p}_{cb}(X)$ replacing $p(X)$.
In a recent paper, $\hat{\tau}_{u}^{cb}$ is also recommended by Heiler2022, who shows that it is numerically identical to $\hat{\tau}_{t}^{cb}$, as long as $X$ includes a constant. We add to Heiler2022's Heiler2022 observation and determine that, when $X$ includes a constant, $\hat{\tau}_{u}^{cb}$ is also identical to $\hat{\tau}_{a,10}^{cb}$ and $\hat{\tau}_{a,0}^{cb}$.
Proposition (ref) demonstrates that using covariate balancing propensity scores solves the problem of choosing an appropriate weighting estimator of $\late$, because all the estimators we previously determined to have some desirable finite sample properties are identical when $\hat{p}_{cb}(X)$ replaces $p(X)$.
So far, we have focused on the finite sample properties of several weighting estimators of $\late$. To determine the asymptotic distribution of each estimator, we apply general results on M-estimation Wooldridge2010,BoosStefanski2013, as all the weighting estimators considered in this paper can be represented as an M-estimator.
Weighting estimators are all functions of the instrument propensity score, $p(X)$\@. As in Section (ref), we assume a parametric model, $F(X,\alpha)$, for $p(X)$\@. Thus, the LATE can be estimated by a two-step procedure where $\alpha$ is estimated in the first step and the unknown $F(X,\alpha)$ is replaced with its estimate in the second step. Alternatively, one could jointly estimate $\alpha$ and $\late$ within an M-estimation framework using moment functions related to both $\alpha$ and $\late$. The moment function related to the estimation of $\alpha$ is either the score from the maximum likelihood estimation or the covariate balancing condition from IR2014. The moment functions related to $\late$ are derived from the identification results in Section (ref). All moment functions are summarized in Table (ref). For different weighting estimators, different combinations of moment functions will be necessary. Provided that the standard regularity conditions NeweyMcFadden1994 are satisfied and the relevant moments exist, all the estimators considered here are asymptotically normal. The derivation of the asymptotic variance for each of the estimators is presented in the appendix. These variances are also estimated in our companion Stata package kappalate.
Although it would be interesting to compare the asymptotic variances of the different weighting estimators considered in this paper, we leave this task to future research. At this time, we instead make three additional points. First, we conjecture, as in KM2016 and KU2023, that normalization may help reduce the asymptotic variance of an estimator, in which case $\hat{\tau}_{u}^{ml}$ would be more efficient than $\hat{\tau}_{t}^{ml}$ ($= \hat{\tau}_{a,1}^{ml}$). Second, we note that $\hat{\tau}_{u}^{cb}$ attains the semiparametric efficiency bound in Frolich2007 and HN2010 as long as the number of balancing constraints grows appropriately with the sample size Heiler2022. Third, we recognize that our asymptotic analysis implicitly requires a restriction stronger than Assumption (ref)(iii), namely the “strong overlap” assumption of KT2010.
In this section we use three empirical applications to illustrate our findings from Section (ref). The bottom line is that the proportion of compliers is sufficiently large in every application (i.e. the instruments are sufficiently strong) so that the phenomenon of dividing by “near zero” never occurs. Ultimately, the three normalized estimators that we consider, $\hat{\tau}_{u}^{cb}$, $\hat{\tau}_{u}^{ml}$, and $\hat{\tau}_{a,10}^{ml}$, are practically indistinguishable from one another in all applications. At the same time, we document the lack of translation invariance and scale equivariance of the unnormalized estimators. We also report the corresponding 2SLS estimates, which are obtained with the covariates appearing additively in the linear equation. Both in this context and in the case of parametric estimation of the instrument propensity score, the relevant model may be misspecified in the absence of sufficiently flexible covariate specifications.
In our first application, we revisit Angrist1990's Angrist1990 study of causal effects of military service using the draft eligibility instrument. In the early 1970s, priority for induction in the U.S. was determined in a sequence of lotteries. The instrument in Angrist1990 takes the value 1 for individuals with dates of birth that were randomly determined as draft eligible and 0 otherwise. Because the fraction of eligible dates of birth was cohort specific, it is essential to control for age in this application.
In what follows, we use a sample of 3,027 individuals from the 1984 Survey of Income and Program Participation (SIPP), which is also considered by MW2017. Our outcome of interest is log wage. To illustrate the invariance properties in Proposition (ref), we consider the natural logarithm of hourly wages as measured in cents or dollars. We also consider three sets of covariates: age, a cubic in age, and a set of indicator variables for each value of age. Summary statistics for these data are reported in Table 6 of MW2017.
Table (ref) reports our estimates of causal effects of military service. Panels A and B, which report 2SLS and normalized weighting estimates, suggest that these effects were positive and economically meaningful in the period under study, with a narrow range of estimates from 20--25 log points. The differences between the 2SLS and weighting estimates (as well as their standard errors) are always very minor. Although the estimated effects are all positive, they are not statistically significant. The estimates do not depend on whether we measure wages in cents or dollars.
Panel C of Table (ref) reports unnormalized weighting estimates. Unlike in panels A and B, these estimates are heavily dependent on the exact specification and, except in the case of the saturated specification, on whether we measure wages in cents or dollars prior to the log transformation. For example, in columns 1 and 2, we only control for age, and yet the estimates are negative and marginally significant when wages are measured in cents prior to the log transformation, while becoming marginally positive when wages are measured in dollars. When the covariate specification is saturated, as in columns 5 and 6, the unnormalized estimates do not depend on the units of measurement of the original outcome variable; they also become identical to each other and to the normalized estimates. This demonstrates the virtue of flexible covariate specifications.
In our second application, we revisit Card1995's Card1995 study of causal effects of education using the college proximity instrument. Card1995 uses data from the National Longitudinal Survey of Young Men (NLSYM) and restricts his attention to a subsample of 3,010 individuals who were interviewed in 1976 and reported valid information on wage and education. His endogenous variable of interest is years of schooling, which is instrumented by an indicator for the presence of a four-year college in the respondent's local labor market in 1966.
This study has been revisited by numerous papers, many of which focus on binarized versions of Card1995's Card1995 education variable. For example, Tan2006 and Sloczynski2021 study the effects of having at least thirteen years of schooling (“some college attendance”) while HM2015, Kitagawa2015, MW2017, and AH2021 focus on having at least sixteen years of schooling (“college completion”). In what follows, we consider both binarizations. Our outcome of interest is log hourly wage, with wages measured either in cents or in dollars. We also consider two sets of covariates: a quadratic in experience, nine regional indicators, and indicators for whether Black, whether lived in an SMSA in 1966 and 1976, and whether lived in the South in 1976, as in Card1995; and indicators for whether Black, whether lived in an SMSA in 1966 and 1976, and whether lived in the South in 1966 and 1976, as in Kitagawa2015. Summary statistics for these data are reported in Table 1 of Card1995.
Table (ref) reports our estimates of causal effects of college education on log wages. Many of these estimates seem implausible, often because they are “too large.” This is unsurprising given the possible failures of the exclusion restriction and monotonicity in this application AH2021, Sloczynski2021. From our perspective, these concerns are less relevant, however, because we use Table (ref) as another illustration of Proposition (ref). The normalized estimates (as well as 2SLS) clearly do not depend on the units of measurement of the outcome variable prior to the log transformation. This is no longer the case for the unnormalized estimates, as reported in Panel C of Table (ref). For example, when focusing on the “some college attendance” treatment and using Card1995's Card1995 specification, we obtain negative estimates when wages are measured in cents but positive when they are measured in dollars. Both sets of estimates are economically meaningful even if insignificant; regardless, the lack of invariance is disconcerting. When we use Kitagawa2015's Kitagawa2015 specification instead, all estimates are positive and statistically different from zero, but more than twice as large when wages are originally measured in cents rather than dollars.
In our third empirical application, we revisit AE1998's AE1998 study of causal effects of childbearing using the sibling sex composition instrument. AE1998 use the incidence of a twin birth and the sex of the first two children as two alternative instruments for having at least three children in a sample of women with two or more children. In what follows, we restrict our attention to the sex composition instrument.
This study has been revisited in many papers, including FGV2018. In what follows, we use FGV2018's FGV2018 subsample of the 1980 US Census that consists of all women aged 21--35 with at least two children. The number of observations is 394,840, which is nearly identical to the sample size in AE1998. Summary statistics for these data are reported in Table 2 of AE1998. Our outcomes of interest are log annual income and an indicator for labor force participation. In the case of log income, we implicitly condition on reported income being greater than zero (as in Sections (ref) and (ref)). The treatment is having more than two children. The set of covariates consists of age, age at first birth, sex of the first and second children, and indicators for whether Black, whether Hispanic, and whether another race. The instrument is an indicator for whether the first two children are of the same sex.
We consider a broader set of transformations of the outcome variables relative to the previous applications. In the case of labor force participation, we originally code working for pay as 1 and not working for pay as 0. Subsequently, however, we also recode working for pay as 2 and not working for pay as 1, as well as not working for pay as 1 and working for pay as 0. In the case of income, we consider four different units of measurement: cents, dollars, thousands of dollars, and hundreds of thousands of dollars. While the first and the last unit of measurement may appear impractical for annual income, our goal is to demonstrate the fragility of the unnormalized estimates with respect to such transformations.
Table (ref) reports our estimates of causal effects of childbearing on labor market outcomes. Panels A and B, which report 2SLS and normalized weighting estimates, respectively, suggest that these effects are negative and economically meaningful, although the effects on log income are not statistically different from zero. As in our replication of Angrist1990, the differences between the 2SLS and weighting estimates (as well as their standard errors) are always very minor. Transformations of the outcome variables do not influence any of the estimates.
Panel C of Table (ref) reports the unnormalized estimates. The fragility of these estimates is immediately evident. In the case of income, the estimated effects of childbearing are positive and highly significant when income is measured in cents, positive and insignificant when in dollars, negative and insignificant when in thousands of dollars, and negative and highly significant when in hundreds of thousands of dollars. This is obviously very disconcerting. Likewise, in the case of labor force participation, the estimates are quite fragile, although less so than in the case of income, perhaps because of the binary nature of the outcome. Still, the estimates in column 3 are nearly twice larger than those in column 2, even though the only difference between these two columns is in a particular recoding of the binary outcome.
In this section we use a simulation study to illustrate our findings on the properties of weighting estimators of the LATE\@. To reduce the number of researcher degrees of freedom, we focus on data-generating processes from Heiler2022, which leads to the following system of equations: \begingroup \allowdisplaybreaks
\endgroup where $u$ and $X$ are i.i.d. standard uniform, $\left(
\right) \sim \mathcal{N} \left( \left[
\right], \left[
\right] \right)$, $\theta_0 = \ln((1-\delta)/\delta)$, and $\delta \in \left\lbrace 0.01,0.02,0.05 \right\rbrace$. What remains to be specified is three functions, namely $\mu_d(x,z)$, $\mu_{y_1}(x)$, and $\mu_z(x)$. Our choices for these functions are listed in Table \ref{tab:DGPs}. It is useful to note that, given these choices and the fact that $X$ has a standard uniform distribution, $\delta$ is equal to the lowest possible value of the instrument propensity score and (symmetrically) one minus the instrument propensity score, that is, $\delta \leq \pr (Z=1 \mid X) \leq 1 - \delta$. Thus, $\delta$ controls the degree of overlap in the data.
Note that Designs A.1, B, C, and D in Table (ref) are identical to Designs A, B, C, and D, respectively, in Heiler2022. It is easy to see that Design A.1 corresponds to a setting with (near) one-sided noncompliance, as $\pr(D=1 \mid Z=1) = \Phi(4) = 0.99997$, where $\Phi(\cdot)$ is the standard normal cdf. It follows that there are essentially no never-takers in Design A.1. To illustrate our findings from Section (ref) on near-zero denominators, we are also interested in a design with (nearly) no always-takers. This is accomplished by Design A.2, which is identical to Design A.1 except for a small change to $\mu_d(x,z)$ that reverses the direction of noncompliance. Indeed, in Design A.2, $\pr(D=1 \mid Z=0) = \Phi(-4) = 0.00003$, which means that there are essentially no always-takers.
It is also useful to note that Designs A.1 and A.2 correspond to the case of a fully independent instrument while in the remaining designs the instrument is conditionally independent. Additionally, in Designs A.1, A.2, and B, treatment effect heterogeneity is only due to the correlation between $\varepsilon_1$ and $v$; in Designs C and D, on the other hand, the dependence of $\mu_{y_1}(X)$ on $X$ constitutes another source of heterogeneity. In the end, the 2SLS estimator that controls for $X$ is expected to perform very well in Designs A.1, A.2, and B but not necessarily elsewhere Heiler2022.
In our simulations, similar to Heiler2022, we thus use the 2SLS estimator as a benchmark that the weighting estimators will not be able to outperform in Designs A.1, A.2, and B while almost certainly being able to do so in Designs C and D\@. We also consider $\hat{\tau}_{u}^{cb}$, $\hat{\tau}_{u}^{ml}$, $\hat{\tau}_{a,10}^{ml}$, $\hat{\tau}_{a}^{ml}$, $\hat{\tau}_{a,1}^{ml}$ ($= \hat{\tau}_{t}^{ml}$), and $\hat{\tau}_{a,0}^{ml}$, also controlling for $X$\@. This leads to a misspecification in Design D, where $\mu_z(X)$ is quadratic in $X$ but we mistakenly omit the quadratic term. We consider three sample sizes, $N=500$, $N=1{,}000$, and $N=5{,}000$, and 10,000 replications for each combination of a design, a value of $\delta$, and a sample size.
Our main results are reported in Tables (ref) to (ref). For each estimator, we report the mean squared error (MSE), normalized by the MSE of the 2SLS estimator, the absolute bias, and the coverage rate for a nominal 95% confidence interval.
In Design A.1, as expected, the 2SLS estimator outperforms all weighting estimators of the LATE, with MSEs of these estimators always at least 31% larger, and sometimes orders of magnitude larger, than that of 2SLS\@. With better overlap and larger sample sizes, all estimators have small biases. When overlap is poor and/or samples small, 2SLS is better than the weighting estimators in terms of bias, too. Coverage rates are close to the nominal coverage rate for all estimators in all cases. At the same time, in a comparison of different weighting estimators, three of them, $\hat{\tau}_{t}^{ml}$, $\hat{\tau}_{a}^{ml}$, and $\hat{\tau}_{a,10}^{ml}$, are very unstable when overlap is sufficiently poor, $\delta \in \left\lbrace 0.01,0.02 \right\rbrace$, and samples are small, $N=500$. This is documented by very large MSEs in these cases. However, as predicted by Section (ref), $\hat{\tau}_{a,0}^{ml}$, $\hat{\tau}_{u}^{ml}$, and $\hat{\tau}_{u}^{cb}$ do not suffer from instability, even in the most challenging case with $\delta = 0.01$ and $N=500$. This is because there are (nearly) no never-takers in Design A.1. More generally, $\hat{\tau}_{u}^{cb}$ and $\hat{\tau}_{u}^{ml}$ perform better than $\hat{\tau}_{a,0}^{ml}$, which is likely due to normalization.
Our results for Design A.2 are generally similar, except for the relative performance of 2SLS in terms of bias and, especially, the exact list of weighting estimators that suffer from instability. Unlike in Design A.1, when overlap is poor and/or samples small, the bias of 2SLS is not clearly smaller than that of (most of) the weighting estimators. Also, it is $\hat{\tau}_{a,0}^{ml}$, $\hat{\tau}_{a,10}^{ml}$, and perhaps $\hat{\tau}_{a}^{ml}$ that suffer from instability in such cases---but clearly not $\hat{\tau}_{t}^{ml}$. As discussed in Section (ref), this is because there are (nearly) no always-takers in Design A.2. As before, $\hat{\tau}_{u}^{cb}$ and $\hat{\tau}_{u}^{ml}$ perform marginally better than the best unnormalized estimator (in this case, $\hat{\tau}_{t}^{ml}$).
In Design B, the instrument is no longer fully independent and noncompliance is no longer one sided. While 2SLS remains dominant in terms of MSE, it is always outperformed by most of the weighting estimators in terms of bias, often substantially and sometimes by all of them. In a comparison of different weighting estimators, $\hat{\tau}_{u}^{cb}$ and $\hat{\tau}_{u}^{ml}$ remain best overall while $\hat{\tau}_{t}^{ml}$, $\hat{\tau}_{a}^{ml}$, and $\hat{\tau}_{a,10}^{ml}$ clearly suffer from instability when overlap is sufficiently poor and samples sufficiently small. The case of $\hat{\tau}_{a,0}^{ml}$ is borderline, which is perhaps due to the fact that there are many more always-takers than never-takers in this design (although both groups clearly exist, unlike before).
Next, in Design C, we introduce another source of treatment effect heterogeneity through the dependence of $\mu_{y_1}(X)$ on $X$\@. The 2SLS estimator is no longer consistent for the LATE, which is illustrated by its large bias in all cases, including the least challenging case with $\delta = 0.05$ and $N=5{,}000$. Given that we define the coverage rate as the fraction of replications in which the LATE is contained in a nominal 95% confidence interval, we also obtain very low coverage rates for 2SLS, never exceeding 66% and approaching 0% when the sample size is sufficiently large. Coverage rates for all the weighting estimators are close to the nominal level when overlap is good and samples large enough. The only weighting estimators that never suffer from instability are $\hat{\tau}_{u}^{cb}$ and $\hat{\tau}_{u}^{ml}$, although $\hat{\tau}_{u}^{cb}$ is now dominant, with substantial improvements in MSE in all cases.
Finally, in Design D, the instrument propensity score is misspecified, as we mistakenly omit the quadratic in $X$\@. The 2SLS estimator remains inconsistent, too, and its coverage rates are close to 0% in all cases. While the weighting estimators clearly differ in performance, sometimes in unexpected ways, the most striking feature of the simulation results for Design D is the dominance of $\hat{\tau}_{u}^{cb}$, in terms of MSE, bias, and coverage. The relative efficiency of $\hat{\tau}_{u}^{cb}$, here and elsewhere, can be understood through the lens of a heuristic argument in Heiler2022, who explained that covariate balancing implicitly regularizes the propensity score estimates away from the boundary and thereby decreases variance. It is also useful to note that, despite misspecification of the instrument propensity score, the coverage rate for $\hat{\tau}_{u}^{cb}$ approaches the nominal level when overlap is sufficiently good and samples sufficiently large, which is not the case for any other estimator.
It seems natural to interpret the instability of different weighting estimators of the LATE as a consequence of near-zero denominators, as we have done so far. To corroborate this interpretation, in Figures (ref) to (ref), we present box plots with simulation evidence on all estimators of the proportion of compliers that we consider: the first-stage coefficient on $Z$ in 2SLS; the denominator of $\hat{\tau}_{u}^{ml}$; $N^{-1}\sum_{i=1}^{N} \hat{\kappa}_{i1}$, $N^{-1}\sum_{i=1}^{N} \hat{\kappa}_{i0}$, and $N^{-1}\sum_{i=1}^{N} \hat{\kappa}_{i}$, with the maximum likelihood propensity scores; the denominator of $\hat{\tau}_{u}^{cb}$; and $N^{-1}\sum_{i=1}^{N} \hat{\kappa}_{i1} = N^{-1}\sum_{i=1}^{N} \hat{\kappa}_{i0}$, with the covariate balancing propensity scores. A straightforward comparison of Tables (ref) to (ref) with Figures (ref) to (ref) reveals that instability of weighting estimators of the LATE is indeed associated with situations in which the supports of their denominators, the estimators of the proportion of compliers, are crossing zero. In fact, it is not negative estimates of this proportion that are particularly problematic, even if they make no logical sense, but rather those estimates that are very close to zero, as this results in dividing by “near zero” to construct an estimate of the LATE, which leads to instability. Additional simulation evidence is also provided in Figures (ref) to (ref), which present histograms for each combination of an estimator, a design, a value of $\delta$, and a sample size. In cases with instability, the normal approximation to the sampling distribution is clearly inappropriate.
In this paper we study the properties of several weighting estimators of the local average treatment effect (LATE), which are based on the identification results of Abadie2003 and Frolich2007. We make several novel observations. First, we show that some of the most popular weighting estimators of the LATE are not translation invariant or scale invariant with respect to the natural logarithm, which translates to their sensitivity to the units of measurement when estimating the LATE in logs and the centering of the outcome variable more generally. In contrast, normalized weighting estimators generally have these important properties. Second, we demonstrate that certain weighting estimators of the LATE have an advantage of being based on a denominator that is strictly greater than zero in settings with one-sided noncompliance. There is only one estimator under consideration in this paper, originally proposed by UysalDiss, that possesses both these advantages. When the instrument propensity score is estimated using an appropriate covariate balancing approach, this estimator is also equivalent to the one in Heiler2022.
We illustrate our findings with three empirical applications and a simulation study. In simulations, our preferred estimator performs relatively well in every setting under consideration. In empirical applications, we clearly document the lack of translation invariance and scale equivariance of the unnormalized estimators. Our preferred estimator is fully robust to the underlying transformations of the outcome data.
\singlespacing
\setlength\bibsep{0pt}