EconBase
← Back to paper

Efficient Covariate Balancing for the Local Average Treatment Effect

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

93,316 characters · 15 sections · 93 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Efficient Covariate Balancing for the Local Average Treatment Effect

titlepage\thispagestyle{empty} \begin{abstract} \singlespacing This paper develops an empirical balancing approach for the estimation of treatment effects under two-sided noncompliance using a binary conditionally independent instrumental variable. The method weighs both treatment and outcome information with inverse probabilities to produce exact finite sample balance across instrument level groups. It is free of functional form assumptions on the outcome or the treatment selection step. By tailoring the loss function for the instrument propensity scores, the resulting treatment effect estimates exhibit both low bias and a reduced variance in finite samples compared to conventional inverse probability weighting methods. The estimator is automatically weight normalized and has similar bias properties compared to conventional two-stage least squares estimation under constant causal effects for the compliers. We provide conditions for asymptotic normality and semiparametric efficiency and demonstrate how to utilize additional information about the treatment selection step for bias reduction in finite samples. The method can be easily combined with regularization or other statistical learning approaches to deal with a high-dimensional number of observed confounding variables. Monte Carlo simulations suggest that the theoretical advantages translate well to finite samples. The method is illustrated in an empirical example. \end{abstract} Keywords: Instrumental variable; Inverse probability weighting; Treatment effect\\ JEL classification: C21, C26

\pagenumbering{arabic}

Introduction

Estimation of causal effects is at the heart of modern empirical research in economics. In particular, evaluating the impact of policies or programs on units with heterogeneous preferences and characteristics such as households, workers, unemployed, firms, or students is necessary to develop a thorough understanding of fundamental economic relationships. This paper deals with the problem of estimating a causal effect of a treatment using variation from a conditionally independent instrumental variable when units have heterogeneous responses to the instrument and to the treatment. We develop a semiparametric estimation method for the local average treatment effect (LATE) based on inverse probability weighting (IPW). The method is designed to increase internal validity by reducing the finite sample bias in estimation of the treatment effect while allowing for dependent observable and unobservable variables to affect both treatment participation and potential outcomes of interest (selection on unobservables). It is more robust than conventional instrumental variable methods as it does not require parametric functional form assumptions about outcome or treatment selection steps or other restrictions such as constant causal effects. Moreover, it has appealing point estimation properties compared to conventional IPW estimators for the LATE.

The key insight required is that IPW estimation of the LATE fundamentally rests on balancing the distribution of observable confounders for two reduced form type components. In particular, for the reduced form type component of the outcome, observations that are instrumented at a given level are weighted by their inverse probability of being in that state, i.e. by their instrument propensity scores. This assures that differences in observed confounders are correctly taken into account when estimating the reduced form type effect for units that choose treatment in accordance with the instrument (compliers). For the overall LATE, the remaining differences due to heterogeneous treatment selection responses to the instrument are then taken into account by another reduced form type component that applies the same inverse probability weights to the corresponding treatment indicators.

On a technical level, balancing means that the empirical distributions of inverse probability weighted observable covariates between the groups that are defined by the instrument are imposed to be identical. However, conventional methods for estimation of the weights such as maximum likelihood or even true weights do not yield perfect balance in finite samples. We exploit recent advances in the literature on the evaluation of causal effects that have been established under the more restrictive conditionally independent treatment and overlap assumptions (selection on observables) to construct empirical balancing conditions for the estimation of inverse probability weights. We achieve exact finite sample balance through tailoring the loss function for the instrument propensity scores that are used for estimation of the LATE. The method compares favorable to conventional inverse probability weighting methods as the tailored loss approach minimizes approximate bias while simultaneously favoring weights that do not exhibit too much variance. In addition, it preserves the design philosophy by rubin2007design as it does not require the use of any outcome or treatment data when selecting a model for the instrument propensities which helps to avoid p-hacking and other post-model-selection problems.

Throughout the paper we show that the proposed balancing estimator has desirable bias properties that compare to the conventional two-stage least squares estimator under homogeneous causal effects for the compliers. Moreover, the bias can be further reduced by incorporating additional information into the balancing constraints, in particular from the treatment selection step. The balancing approach implicitly penalizes deviations from moderate inverse probability weights and thus helps to reduce the variance in estimation of the LATE. Moreover, the balancing estimator is asymptotically normal and reaches the semiparametric efficiency bound if the number of balancing constraints grows appropriately with the sample size. The method can be easily combined with regularization or other statistical learning approaches to deal with a high-dimensional number of observed confounding variables. Monte Carlo simulations suggest that the theoretical advantages over conventional methods translate well to finite samples. The method is applied to a re-evaluation of the causal effect of 401(k) participation on total financial assets.

Identification and non/semiparametric estimation of the LATE under a binary conditionally independent instrument has been first considered by abadie2003semiparametric and frolich2007nonparametric. abadie2003semiparametric relies on implicit identification of complying units by inverse probability weights to estimate or approximate conditional expectation functions for compliers that can consequently be used to construct estimates for local average treatment effects. His method requires an (approximate) model for complier outcomes and is thus more prone to misspecification and post-model-selection problems. In practice, the method often yields estimates close or identical to conventional two-stage least squares if linear models are used angrist2009mostly. frolich2007nonparametric proposes nonparametric imputation/matching estimators for estimation of the LATE and shows that the semiparametric efficiency bound is not affected by knowledge of the instrument propensity scores similar to hahn1998role in the context of selection on observables. frolich2007nonparametric also suggests the use of an IPW estimator but does not provide any theory for estimation. donald2014testing close this gap in the literature by using an IPW estimator for the LATE that relies on nonparametric series estimation for the instrument propensity scores similar to hirano2003efficient in the context of selection on observables and provide conditions for semiparametric efficiency. donald2014inverse suggest to estimate the LATE semiparametrically efficient via IPW with instrument propensity scores estimated by local polynomial regression and provide a higher order mean squared error expansion. The IPW estimators in frolich2007nonparametric, donald2014inverse, and donald2014testing are all of the “IPWI”-type, i.e. do not impose normalization of the inverse probability weights. Our approach is closest to donald2014testing but uses balancing conditions instead of a series logistic estimator. Moreover, our balancing estimator is automatically weight-normalized, favors moderate instrument propensities, and imposes exact mean balance in finite samples. Thus, despite their asymptotic equivalence, we expect major differences in bias and overall point estimation risk in finite samples.

Improving IPW estimation through imposing exact or approximate balancing constraints has been considered in the literature on estimation of treatment effect under the more restrictive selection on observables assumptions. graham2012inverse considers tilted moment conditions for estimation of the conventional propensity scores that bear close resemblance to balancing approaches. hainmueller2012entropy proposes direct optimization of a distance criterion depending on the inverse probability weights (e.g. Kullback entropy divergence) subject to approximate empirical balancing and positivity constraints. zhao2017entropy provide conditions under which the balancing method by hainmueller2012entropy is doubly robust for the average treatment effect on the treated. imai2014covariate consider exact balancing of covariates through propensity scores in a parametric GMM or empirical likelihood framework. zubizarreta2015stable proposes to minimize a quadratic problem in terms the inverse probability weights subject to similar approximate mean balancing constraints. athey2018approximate employ approximate balancing weights for bias-correction of a high-dimensional linear model for estimation of treatment effects. zhao2019covariate develops a unifying framework for empirical balancing approaches and demonstrates how to tailor the (negative) loss function to produce weights that correspond to the treatment effect of interest. Moreover, zhao2019covariate proposes methods for regularization and shows how the bias is affected by relaxing exact balancing conditions to approximate balancing. All of these contributions suggest that, for selection on observables, empirical balancing can substantially outperform conventional weighting estimators by reducing differences between the weighted empirical distributions of treatment and control units. Our balancing approach in its most basic form is a repeated application of the exact balancing method by imai2014covariate and zhao2019covariate using the instrument instead of the treatment indicators. However, it has different bias and weight normalization properties that do not apply in the context of selection on observables. Moreover, for the LATE there is an extended hierarchy in terms of the available information that consists of three different levels: The basic instrument assignment and the higher-order treatment selection and outcome generation steps. We demonstrate that using information from the higher-order steps can help to achieve approximately unbiased estimates for the causal effect. In particular, balancing estimated treatment participation probabilities for a given instrument level can help to reduce point estimation risk even when the treatment selection model is misspecified.

Applications that rely on identification of a conditionally independent instrument are manifold. angrist1990lifetime studies the effect of military service on lifetime earnings using the Vietnam era draft lottery as an instrument. card1995using uses college proximity as an instrument for education, see also card2001estimating. poterba1995401 and abadie2003semiparametric exploit eligibility to 401(k) programs as an instrument to investigate whether contributions to these plans crowd out other personal savings. cawley2007impact and cawley2013impact use minimum requirements for physical education (PE) as an instrument for actual time spend in PE to evaluate the impact on health outcomes for elementary school children and high school students respectively. Based on individual data from different federal states in Germany, knaus2018better exploit between and within variation of compulsory PE lessons as an instrument to estimate the effects of PE on multiple measurements for child development. The working paper by knaus2018better is an application of an idea similar to the one proposed in this paper. They combine two steps of the inverse probability tilting method by graham2012inverse to obtain a balanced estimate for the LATE. However, they do not provide any theory for estimation or statistical inference.

The paper is structured as follows: Section (ref) introduces the identification assumptions and provides the basic arguments behind the balancing approach. Section (ref) presents the estimation strategy. Section (ref) provides the statistical properties of the balancing estimator. Section (ref) demonstrates how to use higher-order information for balancing and outlines extensions to the case of high-dimensional observed confounding variables. Section (ref) contains the Monte Carlo simulations. Section (ref) provides the application and Section (ref) concludes. All major proofs and derivations are collected in the Appendix.

Identification and Balancing

We consider standard identification conditions for treatment effects under unobservable heterogeneity and a binary conditionally independent instrument often referred to as assignment. Assume we observe independent data $(Y_i,D_i,Z_i,X_i')$ for units $i=1,\dots,n$. $X_i$ is a vector of (causally) predetermined\footnote{Predetermined variables are set before the cause can have an effect, e.g. a priori attributes, see imbens2015causal.} random variables supported on $\mathcal{X} \subset \mathbb{R}^{\dim{X_i}}$, $Z_i$ a binary instrument, $D_i$ a binary treatment indicator, and $Y_i$ a real-valued outcome variable. In principle, there are four potential outcome states $Y_i(d,z)$ for $d,z \in \{0,1\}$ but only $Y_i = Y_i(D_i,Z_i)$ is observed. For each instrument level $z\in \{0,1\}$, there is a potential treatment status $D_i(z)$ which yields the observed treatment $D_i = D_i(Z_i)$ according to

align[align omitted — 45 chars of source]

The following identification assumptions along the lines of abadie2003semiparametric and frolich2007nonparametric are imposed, see also donald2014inverse,donald2014testing:

enumerate• (Conditional Independence) $\{Y_i(d,z);\forall d,z\}, D_i(1), D_i(0) \ \rotatebox[origin=c]{90}{$\models$} \ Z_i|X_i$. • (Exclusion) $P(Y_i(d,1) = Y_i(d,0)) = 1 \text{ for } d=0,1$. • (Monotonicity) $ P(D_i(1) - D_i(0) \geq 0) = 1$. • (First Stage) $E[D_i(1) - D_i(0)] \neq 0$. • (Strong Instrument Overlap) Let $\pi(x) = P(Z_i=1|X_i=x)$. There exists a $\delta > 0$ such that $\delta < \pi(x) < 1-\delta$ for all $x \in \mathcal{X}$. • (Stable Unit Treatment Values) $(\{Y_i(d,z);\forall d,z\}, D_i(1), D_i(0))$, $i=1,\dots,n$ are independent for $i\neq j$.

} A.1-A.4 together are a weaker version of the conditions for identification of the LATE in the seminal paper by imbens1994identification that rely on a completely independent instrument. Assumption A.1 is the fundamental identification assumption. It implies that after controlling for a sufficient set of observed confounding variables, all residual variation in potential outcomes and potential treatment statuses is completely independent of the instrument. Thus, conditional on observed confounders, the instrument can be thought of as being allocated like in a completely randomized experiment. Assumption A.2 rules out any direct effects of the instrument on potential outcomes other than through the indirect treatment channel. Thus, there cannot be any unobserved confounders affected by the instrument and no feedback from potential outcomes to the instrument. Under this exclusion, the fundamental problem of causal inference holland1986statistics determines the observation rule for the outcome, i.e.

align[align omitted — 94 chars of source]

Assumption A.3 imposes a monotonous effect of the instrument on the treatment choice, i.e. receiving instrument $Z_i = 1$ makes any unit at least as likely to select itself or to be selected into treatment compared to $Z_i = 0$. This is sometimes referred to as the “no defiers” assumption, see angrist1996identification. Assumption A.4 assures that there is overall variation in potential treatment statuses as a result from variation in the instrument. Together with monotonicity, this requires the data to have a nonzero share of compliers, i.e. there must be units for which $D_i(1) > D_i(0)$. Thus, the population cannot only be comprised of units that are always treated or never treated independently of their instrument level. Assumption A.5 requires that potentially each unit could have been exposed to a different instrument level. In principle, point identification only requires instrument overlap ($\delta = 0$). However, for regular behavior of the point estimators considered throughout the paper, strong instrument overlap is required.\footnote{Irregular estimators for the case of identification under selection on observables are considered in heiler2020inference.} Assumption A.6 rules out any spillover or general equilibrium effects in terms of both potential treatment states and potential outcomes. This, together with the assumption that covariates are predetermined, allows for a coherent definition of individual causal effects as differences in potential outcomes rubin1974estimating,imbens2015causal. The causal effect of the treatment on the outcome for a unit $i$ is then given by

align[align omitted — 40 chars of source]

Note that there are no further restrictions on the functional relationship between a unit's potential outcomes, instrument, or covariates. Thus, the framework allows for almost any type of observable and unobservable heterogeneity in potential outcomes and causal effects, e.g. it allows for nonconstant treatment effects even conditional on the observed confounders. frolich2007nonparametric shows that under similar assumptions as A.1-A.6, the local average treatment effect (LATE)

equation[equation omitted — 64 chars of source]

is nonparametrically identified. $\tau_{LATE}$ is the average treatment effect for the subpopulation of compliers, i.e. the expected causal effect for a unit randomly drawn from the population of units that alter their potential treatment choice in accordance with the instrument. Without further assumptions, identification does not extend to more general causal effects such as the average treatment effect (ATE) or the treatment effect on the treated (TT). Under treatment effect homogeneity, the LATE is equal to the ATE and the TT. Moreover, in the case of one-sided noncompliance, i.e. $P(D_i(0)=1) = 0$, the LATE equals the TT bloom1984accounting,frolich2013identification. One-sided noncompliance often occurs in randomized field experiments with imperfect compliance if the treatment can only be provided by the experimenter.

The identification results in abadie2003semiparametric and frolich2007nonparametric allow for the construction of matching, model-based imputation, and inverse probability weighting estimators for the LATE. In this paper we focus on the latter as they only require estimation of the instrument propensity scores $\pi(X_i)$ and no information from outcome or treatment selection steps. Both treatment selection and in particular potential outcome mechanisms can be more difficult to model as they are often the results of complicated mechanisms such as markets, search and matching processes, or social and biological structures and interactions. Thus, focusing on the instrument or assignment phase avoids misspecification and post-model-selection problems with statistical inference as outcome and treatment data are not used for obtaining the instrument propensities. This is along the lines of the argument for focusing on the design phase of experimental and observational studies made by rubin2007design.

In the following we outline the balancing principle behind inverse probability weighting and show how this allows us to extract population constraints that can be used for estimation of the instrument propensity scores. For exploiting identification via inverse probability weighting note that

align[align omitted — 97 chars of source]

with

align[align omitted — 203 chars of source]

Thus both numerator and denominator of the LATE can be written as a difference of two expectations of inverse probability weighted quantities that can be identified from the joint distribution of $(Y_i,D_i,Z_i,X_i')$. The fundamental mechanism behind identification via inverse probability weighting is the balancing property, i.e. the inverse probability weights will balance the distribution of any function of covariates across instrument levels. Let $f_x(x)$ denote the density of the observed confounders and $f_{x|z=1}(x)$ the density of the observed confounders conditional on $Z_i=1$. By Bayes' Law it follows that for all $x \in \mathcal{X}$

align[align omitted — 90 chars of source]

Thus, under Assumptions A.1 and A.5 we have that

align[align omitted — 304 chars of source]

Note that the conditional mean of the observed treatment status for the units with $Z_i=1$ receives a weight such that it corresponds to the conditional mean of the potential treatment status (normalized) over the full population. Equivalently, it holds that ${(1-\pi(x))}^{-1}{f_{x|z=0}(x)}P(Z_i=0) = f_x(x)$. Thus, similar derivations as in (ref) can be done for all components in (ref). In general, property (ref) of the inverse probability weights implies a mean balancing for any function of the covariates: Let $g:\mathcal{X}\rightarrow\mathbb{R}$ be a measurable function with $E[|g(X_i)|] < \infty$. The mean of $g(\cdot)$ is balanced across inverse probability weighted instrument groups and corresponds to the unweighted population mean, i.e.

align[align omitted — 145 chars of source]

Thus, despite allowing for correlated unobservables driving potential outcome and treatment decision, under Assumptions A.1-A.6 balancing on observed confounders is enough to achieve causal identification for the compliers. This is due to the fact that conditional on observables the instrument is as good as randomly allocated. Thus, any imbalances in unobservables that would introduce a bias when comparing means between units from different treatment levels for always-takers, never-takers, and compliers combined do not matter for the (unidentified) population of compliers as for the latter we have that treatment is chosen according to the instrument, i.e. $D_i(z) = z$. The remaining differences due to varying treatment selection probabilities is then accounted for by the denominator in (ref) that identifies the share of compliers. Thus, balancing for compliers boils down to balancing two instrument assignment groups twice for different outcomes. In its mechanic, each step corresponds to balancing different treatment groups via inverse propensity score weighting under the more restrictive selection on observables assumptions for identification of general average treatment effects imai2014covariate,zhao2019covariate.

While population balance (ref) is certainly present when true instrument propensities are used, note that when using true scores or when choosing a model for $\pi(X_i)$ in finite samples, the ultimate goal is to impose balance on conditional means, not to necessarily choose a model that best predicts $\pi(X_i)$ or $Z_i$ in a given sample according to standard loss functions such as entropy/likelihood, accuracy or mean squared error. From a population perspective or asymptotically these goals are usually aligned. For example, a correctly specified instrument propensity score obtained via maximum likelihood will eventually converge to the maximizer of the population likelihood that is the true instrument propensity score. In finite samples, however, balance is not guaranteed and hence any differences between the weighted distributions can still substantially compromise estimation and causal inference. Thus, to reduce bias choosing an estimator that exploits (ref) directly should be beneficial.

The Balancing Estimator

To impose balancing for the functions of choice we propose to use a tailored loss approach that explicitly exploits sample equivalents of (ref) for estimation of the LATE. For selection on observables, this in spirit of the tailored loss framework by zhao2019covariate and the covariate balancing propensity score method by imai2014covariate. In particular, we would like to estimate instrument propensity scores that yield exact balance in finite samples for a vector-valued function $\phi:\mathcal{X}\rightarrow \mathbb{R}^r$ with $r < n$. This amounts to adapting the loss function and to choosing a logistic link for the instrument propensity score with regressors $\phi(X_i) = (\phi_1(X_i),\dots,\phi_r(X_i))'$. In particular, for some $\theta \in \mathbb{R}^r$ let $L(\phi(X_i)'\theta) = 1/(1+\exp(-\phi(X_i)'\theta))$ denote the standard logistic cumulative distribution function at index $\phi(X_i)'\theta$. The balancing estimator for $\theta$ is obtained via maximizing the following tailored (negative) loss function:

align[align omitted — 354 chars of source]

This function is globally concave and has the same population maximizer as the maximum likelihood estimator for an equivalent logit model. Choosing the logistic link has the advantage of imposing exact balance in finite samples as can be seen from the first order condition of the maximization problem. Setting the first derivative of (ref) equal to zero at $\theta = \hat{\theta}$ yields

align[align omitted — 171 chars of source]

with $\hat{\pi}(X_i) = L(\phi(X_i)'\hat{\theta})$. Thus, the tailored loss with logistic link chooses the propensity score model such that balance holds exactly for the empirical counterparts of (ref). In principle, other link functions or balancing approaches are possible as well, see e.g. imai2014covariate. However, the logistic link yields a transparent connection between the choice of balancing functions and the choice of regressors in the instrument propensity score model. As problem (ref) can be understood as a just-identified moment problem, standard model diagnostics or significance tests can be applied, see also imai2014covariate for a GMM version of this argument for selection on observables. The balanced IPW estimator for the LATE is then given by

align[align omitted — 92 chars of source]

with

align[align omitted — 288 chars of source]

A particular feature that is unique to the balanced LATE approach is that if $\phi(X_i)$ contains an intercept, then the LATE estimator in (ref) is automatically weight-normalized, i.e. numerically identical to its “IPWII” version that uses weights of the form $Z_i/\hat{\pi}(X_i)/[n^{-1}\sum_{i=1}^{n}Z_i/\hat{\pi}(X_i)]$ and $(1-Z_i)/(1-\hat{\pi}(X_i))/[n^{-1}\sum_{i=1}^{n}(1-Z_i)/(1-\hat{\pi}(X_i))]$. This follows from the ratio form of $\hat{\tau}_{LATE}$ together with (ref) for $\phi(X_i) = c$ for any $c\neq 0$. It is not the case that balanced inverse probability weights themselves are normalized to unity, i.e. in general $n^{-1}\sum_{i=1}^{n}Z_i/\hat{\pi}(X_i) \neq 1$ and equivalently for the complementary weights. Under selection on observables, there is clear evidence that weight-normalized versions of inverse probability weighting estimators generally outperform their unweighted counterparts due to a reduction in variance busso2014new,pohlmeier2016simple. We expect these results to translate to the case of selection on unobservables. The Monte Carlo study in Section (ref) suggests that a non-negligible part of the superior finite sample performance of the balanced IPW compared to the maximum likelihood based IPW is due to this default normalization.

Statistical Properties of the Balancing Estimator

Approximate Finite Sample Bias

A central goal of imposing the balancing constraint (ref) is to reduce estimation bias in finite samples. Under the more restrictive selection on observables assumptions and treatment effect homogeneity, zhao2019covariate demonstrates that a sufficiently flexible model for the conventional propensity score is enough to achieve an unbiased estimator of the treatment effect up to weight normalization. As discussed in Section (ref), in contrast to the estimator by zhao2019covariate, the balanced IPW estimator of the LATE achieves weight normalization by construction if an intercept is included into the model. Thus, it would seem straightforward to assume that the balanced LATE estimator should be equal to a ratio of two unbiased estimators if both conditional treatment effects $E[Y_i(1)-Y_i(0)|X_i]$ and conditional first stage $E[D_i(1)-D_i(0)|X_i]$ are constant independently of $X_i$. It turns out, however, that it is sufficient to have treatment effect homogeneity within the groups of compliers only and that there are no restrictions required for the conditional potential treatment to instrument responses. In particular, we obtain the following proposition:

propFor any $z\in\{0,1\}$, let $E[D_i(z)|X_i], E[D_i(z)(Y_i(1)-Y_i(0))|X_i]$ and $E[Y_i(0)|X_i]$ be in the linear span of $\{\phi_1(X_i),\phi_2(X_i),\dots,\phi_r(X_i)\}$. Under constant conditional causal effects for the compliers $E[Y_i(1)-Y_i(0)|X_i,D_i(1)>D_i(0)] = \tau_{LATE}(X_i) = \tau_{LATE}$, the inverse probability weighting estimator (ref) that uses balanced instrument propensity scores (ref) is a ratio of two unbiased estimators, i.e. \begin{align} \frac{E[\hat{\Delta}]}{E[\hat{\Gamma}]} = \tau_{LATE}. \end{align}

The span condition demonstrates the requirement of the instrument propensity score model to be able to contain elements that capture variation in the conditional mean of the control outcome $E[Y_i(0)|X_i]$. Moreover, the balancing components have be flexible enough to cover the relevant building blocks for the conditional mean of a potential treatment level and the conditional causal effect for a combination of different units. These conditions can be reformulated by using the monotonicity assumption. For the case with $z=0$, it reduces to the assumption that $E[D_i(0)|X_i]$ and the product between $E[D_i(0)|X_i]$ and the conditional causal effect for the always-takers $\tau_{AT}(X_i)$ have to be contained in the linear span as

align[align omitted — 312 chars of source]

with the second equality following from monotonicity. For the case of $z=1$, it is required that the linear combination of probability weighted versions of the causal effects of compliers and always-takers are captured since

align[align omitted — 273 chars of source]

with the second equality following from constant complier causal effects and monotonicity. Thus, the conditions in Proposition (ref) effectively demand a special type of flexibility regarding which transformations of regressors $\phi(X_i)$ should be contained in the balancing constraints for the instrument propensity score. Depending on the application at hand, these conditions can be rather restrictive. However, note that the derivations required for Proposition (ref) reveal that it is actually not necessary that the instrument scores used in the denominator have the same degree of flexibility as the ones in the numerator. In particular, for denominator scores only a condition for $E[D_i(z)|X_i]$ for a $z\in\{0,1\}$ is required. This can be exploited by choosing different instrument propensity score models for numerator and denominator. We return to this point in Section (ref).

The robustness property in Proposition (ref) is best understood in comparison with the two-stage least squares (2SLS) estimator. In general, if the first stage for the endogenous treatment variable is fully saturated, the model behind 2SLS identifies a weighted version of conditional LATEs with weights being proportional to the conditional variances of the first stages $V[E[D_i|X_i,Z_i]|X_i]$, see angrist1995two. Under homogeneous treatment effects, however, this corresponds to the LATE. Morever, the parameters from the reduced forms for both outcome and endogenous variable can be estimated without bias. Thus, 2SLS is a ratio of two unbiased estimators with the ratio of the expectations being equal to the LATE. Proposition (ref) states that the balanced IPW estimator has an equivalent property under comparable assumptions. However, instead of having a flexible model for the treatment choice $E[D_i|X_i,Z_i]$, the flexibility is incorporated through the choice of the variables and transformations $\phi(X_i)$ in the model for instrument propensity score $E[Z_i|X_i]$. For IPW estimators of the LATE, the finite sample bias property in Proposition (ref) is unique to the balanced instrument propensities and generally does not apply to any other estimation approach such as maximum likelihood or even true instrument propensity scores. We also expect this property to be beneficial in finite samples if there are moderate deviations from homogeneous causal effects for the compliers. In general, the ratio of the expectations is only an approximation to the expectation of the ratio in finite samples. Therefore, the actual finite sample bias for has to be further investigated. We return to this point in Section (ref).

Duality and Variance Reduction

In this section, we show that in finite samples the balancing approach favors moderate instrument propensity scores in the sense of being close to one half. As instrument propensity scores are inversely related to the (conditional) variance of the IPW estimator for the LATE, there are potential gains in terms of point estimation risk compared to using likelihood scores. For moderate inverse probability weights, the dual problem of maximizing the (negative) tailored loss in (ref) penalizes deviations from the unconditional mean in a manner that is proportional to the conditional variance of the estimator for the LATE. Let the inverse probability weights be defined as

align[align omitted — 80 chars of source]

and denote $W = (W_1,\dots,W_n)$ with $W_i = (X_i',Z_i)$ for $i=1,\dots,n$. Conditional on $W$, the balanced instrument propensity scores are known. Thus, a second order Taylor expansion of the variance of the balanced LATE (ref) conditional on $W$ yields

align[align omitted — 118 chars of source]

with

align[align omitted — 229 chars of source]

which is strictly greater than zero. Hence, approximately the conditional variance is bounded from above by

align[align omitted — 157 chars of source]

which is proportional to the average of the squared inverse probability weights. The conditional variance is directly proportional to the latter, i.e. $nV[{\hat{\Delta}}/{\hat\Gamma}|W] \sim n^{-1}\sum_{i=1}^{n}w_i^2$ if there is homoskedasticity within outcomes and treatment levels and constant correlation across $Y_i$ and $D_i$ conditional on $W_i$ or if the corresponding components in the numerator of (ref) are proportional such that $a(x,z) = a$ for all $x\in\mathcal{X}$ and $z\in\{0,1\}$. To approximate the effect of using a balancing estimator on the squared weights, it is insightful to study the Lagrangian dual problem of the maximization problem (ref) in terms of the weights $w_i$. If $\phi(X_i)$ contains an intercept, the dual problem is given by the following constraint optimization:

align[align omitted — 307 chars of source]

see also chan2016globally and zhao2019covariate. While this is generally not equivalent to minimizing the sum of squared weights, consider the case of an independent instrument such as assignment to treatment in a randomized experiment with imperfect compliance. Under such a circumstance, units receive instrument levels with constant likelihood, e.g. $P(Z_i=1) = 0.5$. The true inverse probability weights are then given by $1/P(Z_i=1) = 2$. Naturally, even under perfect randomization there can be imbalances im terms of the covariates across the two instrument groups in finite samples. Thus, balancing scores $\hat{\pi}(X_i)$ will generally differ from the true inverse probability weights. Consider a deviation from the true probability of one half to a more extreme one, i.e. an inverse probability weight exceeding two. A Taylor series of the unconstrained loss function for a single observation in the dual problem (ref) at $w = 2$ yields

align[align omitted — 102 chars of source]

which is a convergent series for moderate deviations, i.e. $|w -2| < 1$. For more general deviations from moderate instrument propensities, the tailored loss behaves qualitatively similar. Expansion (ref) reveals that the dual problem locally penalizes deviations from the equal weighting in a quadratic manner. As equal weighting is variance minimizing by construction\footnote{Equal weighting corresponds to the simple Wald estimator that only relies on binary reduced forms and does not use any covariates. The conditional variance of the latter is proportional to $1/(P(Z_i=1)(1-P(Z_i=1)))$ and thus always below the conditional variance of the IPW estimator that is equally proportional to $n^{-1}\sum_{i=1}^{n}w_i^2$ using true instrument propensities under homoskedasticity for the different assignment groups.}, this implies that, approximately, the balancing weights seek to minimize an upper bound for the conditional variance of the LATE estimator. They do not succeed in imposing an exactly constant weighting scheme due to the constraints in (ref) that are designed to minimize any bias from imbalances across assignment groups.

Moderate weights and balancing covariates without perfect randomization are generally opposing goals. In principal, one could also use the characterization of the conditional variance in (ref) to choose minimizing weights within a class of balancing weights as proposed in the context of selection on observables by li2016balancing. This strategy, however, changes the definition of the underlying identified causal parameter from the LATE to a ratio of two weighted reduced form estimates which put a higher weight to individuals with instrument propensities close to one half. The resulting identified parameter does not lend itself to an intuitive causal interpretation with similar policy relevance compared to the conventional LATE.

Nonparametric Estimation and Large Sample Properties

In light of Proposition (ref) it seems reasonable to choose a model for the instrument propensity score that eventually incorporates a large set of balancing constraints as the sample size increases. In this section we show that using a nonparametric approach within the tailored loss framework can efficiently incorporate all required information for estimation of the LATE in a semiparametric sense. It does so while still retaining the exact finite sample balancing property in (ref). We rely on a series approach using power series similar to the series logistic approach proposed by hirano2003efficient and donald2014testing that rely on a local maximum likelihood step for estimation of the (instrument) propensity scores. For $K> 0$, let $\phi^K(x) = (x^{\lambda(1)},\dots,x^{\lambda(K)})'$ be a vector of power functions such that $||\lambda(k)||_1 \leq ||\lambda(k+1)||_1$ for $k \in \mathbb{N}_0^r$ with $\lambda = (\lambda_1,\dots,\lambda_r)'$ nonnegative and $||\cdot||_1$ denoting the $\ell_1$-norm, i.e. $||\lambda||_1 = \sum_{j=1}^{r}|\lambda_j|$. We assume that the series is orthogonalized with respect to a weight function such that $E[\phi^K(X_i)\phi^K(X_i)'] = I_K$. This is always possible since a logistic link function is used and thus we approximate the log-odds ratio by a linear single index $\phi^K(x)'\theta_K = \theta_K'A_k^{-1}A_k\phi^K(x)$. Therefore, one can always use basis $A_k\phi^K(x)$ for approximation instead, see also Appendix A in hirano2003efficient. The balanced instrument propensity scores are then given by $\hat{\pi}(x) = L(\phi^K(x)'\hat{\theta}_K)$ with

align[align omitted — 372 chars of source]

Let $C^d$ denote the space of $d$-times continuously differentiable functions. We impose the following regularity and smoothness assumptions:

enumerate$X_i$ are $r$-dimensional random variables compactly supported on $\mathcal{X}$ with absolutely continuous density $f(x)$ in $C^2$ and bounded away from zero. • $E[Y_i|X_i,Z_i=z] = m_z(X_i)$ and $E[D_i|X_i,Z_i=z] = \mu_z(X_i)$ are in $C^1$. • $\pi(X_i)$ are in $C^q$ with $q\geq 7r$. • All second moments of potential outcome levels exist and are finite. • $K = O(n^{v})$ with $1/(4(q/r-1))< v < 1/9$.

} The assumptions are standard in the literature, see hirano2003efficient and donald2014testing or li2009efficient and donald2014inverse,donald2014testing for possible restrictions on the series terms and adaptations to discrete covariates and kernel methods. Assumption B.1 assures a uniform approximation of any continuous function of the covariates, in particular conditional means such as the instrument propensity score. Assumptions B.2 and B.3 are smoothness conditions on the observed outcome and treatment status conditional on covariates and instrument level and on the instrument propensity score. The higher the dimensionality of the covariates, the more smoothness conditions are required. B.4 is a regularity condition that is necessary to assure a finite asymptotic variance for the LATE estimator. Assumption B.5 controls the rate at which the order of basis functions is allowed to grow with increasing sample size depending on the degree of smoothness. We obtain the following theorem:

theorem[Efficient Balancing] Under Assumptions A.1-A.6, B.1-B.5, and instrument propensity scores estimated according to (ref), the inverse probability weighting estimator for the LATE \begin{enumerate} • is asymptotically normal $$\sqrt{n}(\hat{\tau}_{LATE} - \tau_{LATE}) \overset{d}{\rightarrow} \mathcal{N}(0,V)$$ • and reaches the semiparametric efficiency bound \begin{align*} V &= \frac{1}{\Gamma^2}\bigg(E[(m_1(X_i) - m_0(X_i) - \tau_{LATE}\mu_1(X_i) + \tau_{LATE}\mu_0(X_i))^2] \\ &\quad + \sum_{z=0,1} E\bigg[\frac{\sigma_{Y_z}^2(X_i) - 2\tau_{LATE}\sigma^2_{Y_zD_z}(X_i) + \tau_{LATE}^2\sigma_{D_z}^2(X_i)}{P(Z_i=z|X_i)}\bigg] \bigg) \end{align*} with $\sigma_{Y_z}^2(X_i) = V[Y_i|X_i,Z_i=z]$, $\sigma_{D_z}^2(X_i) = V[D_i|X_i,Z_i=z]$, and $\sigma_{Y_zD_z}^2(X_i) = Cov[Y_i,D_i|X_i,Z_i=z]$ for any $z\in\{0,1\}$. \end{enumerate}

Theorem (ref) shows that the inverse probability weighting estimator using a sufficiently flexible nonparametric model for the instrument propensity score that imposes empirical balancing constraints efficiently incorporates all information available for estimation of the LATE and is consistent. This is in line with the insights from the bias characterizations in Section (ref). There, the instrument propensity scores yield first-order unbiased estimates if they incorporate the structure of the potential treatment effects and potential treatment levels sufficiently. If these conditional mean functions are continuous, then a series approximation will eventually contain all the components necessary to uniformly approximate them on compact sets. The proof of Theorem (ref) mainly relies on the global concavity of the tailored loss function together with the strong instrument overlap assumption. Theorem (ref) also implies that in large samples there is no qualitative difference between tailored instrument propensity scores and standard nonparametric approaches. However, the balancing approach additionally guarantees finite sample balance. The efficiency bound for the LATE has originally been derived by frolich2007nonparametric, see also hong2010semiparametric. The asymptotic variance $V$ can be consistently estimated using the series approach as in hirano2003efficient and donald2014testing by replacing the population quantities with sample estimates, see Appendix (ref) for more details.

Extensions

Using Higher-order Information to Improve Balance

The results on nonparametric estimation in Section (ref) make a strong case for eventual flexibility of the balancing constraints used for estimation of instrument propensity scores. In large samples, however, the exact choice or order of inclusion of transformations of regressors is not of primary concern. Thus, the question remains, which empirical means should be prioritized for balancing in finite samples, i.e. how to pick $\phi(X_i)$? From a model selection perspective, choosing informative balancing variables first will be beneficial if they do not render the estimation step infeasible or instable due to e.g. multicollinearity. It might even be useful to exploit a set of generated regressors whose balancing is beneficial for precise estimation of the treatment effect. Looking at the LATE from an applied perspective, it is often reasonable to assume a hierarchy in terms of the knowledge about the different steps of the underlying causal mechanism. This hierarchy usually goes in ascending order from a) assignment over b) treatment choice to c) outcome process. Consider the example of evaluating the causal impact of a job search training program on future earnings for the unemployed offered by an employment agency. In this case, the decision to assign units to a program might be based on a set of observable characteristics such as age, employment history, and qualifications available to the agent responsible for assignment. The instrument assignment on level a) is conditionally independent if all other variables that affect the assignment decision are independent of the potential earnings and potential treatment levels. On hierarchy level b), the unemployed units choose whether to comply with the assignment. If enough information about their trade-offs and restrictions are available, we can model the choice problem and its constraints guided by economic theory. The outcome process determining earnings, however, is likely generated as a consequence of a complicated searching and matching process on the labor market under additional constraints. Thus, relying only on data from the employment agency might not be enough to really inform a model for hierarchy level c).

In light of Proposition (ref), balancing the following quantities would in principle be desirable:

enumerate$E[Y_i(0)|X_i], E[D_i(z)(Y_i(1)-Y_i(0))]$ and • $E[D_i(z)|X_i]$

for any $z \in \{0,1\}$. (i) requires outcome data to generate e.g. model-based quantities. Hence, using them for empirical balancing constraints operates on the highest hierarchy level c) and goes against the design arguments outlined by rubin2007design. (ii) on the other hand is a better candidate as it only concerns the treatment choice step. For simplicity assume that $X_i$ is discrete. From the results of vytlacil2002independence it follows that Assumptions A.1-A.4 (conditional on $X_i$) imply the existence a nonparametric single index model that rationalizes the identical choices as the LATE framework conditional on each $x \in \mathcal{X}$ and vice versa, see also kline2019heckits for estimation and numerical equivalence. Thus, it motivates a fully nonparametric model for the first stage. In particular, one obtains

align[align omitted — 84 chars of source]

with $v_i \ \rotatebox[origin=c]{90}{$\models$} \ Z_i|X_i$, $\mu(x,z)$ non-degenerate conditional on $x$, and $F_{v|x}(v)$ continuous. Without any further restrictions and $X_i$ discrete, this would correspond to the first stage of a fully saturated instrumental variables approach. The estimates for selection probabilities $E[D_i(z)|X_i]$ are obtained as the estimates for $F_{v|x}(\mu(X_i,z))$. They can then be included into the vector of transformations of regressors $\phi(X_i)$ to balance the empirical counterpart of $E[D_i(z)|X_i]$ in finite samples.\footnote{For a regressor with many categories or multiple discrete regressors, using additional smoothing methods as in ouyang2009nonparametric or heiler2018shrinkage for estimation of (ref) might be desirable.} The proof of Proposition (ref) (see Appendix (ref)) reveals that this strategy is enough for the denominator to fulfill its role in having an estimator for the LATE under conditional independence that is given by the ratio of two unbiased estimators. For the numerator, however, information about the outcome process as in (i) would be required. If one is willing to impose a model for the potential outcomes, then from a bias perspective it would be sufficient to include them into the balancing constraints for the instrument propensity scores that enter the numerator only. Thus one can in principal operate with two different instrument propensity scores for numerator and denominator that differ by using balancing constraints for the empirical counterparts of either $E[D_i(z)|X_i]$ (denominator) or $E[Y_i(0)|X_i]$ and $E[D_i(z)(Y_i(1)-Y_i(0))|X_i]$ (numerator). Note that under constant treatment effects for some $z\in\{0,1\}$, i.e. for always-takers and compliers or never-takers and compliers we have that $E[D_i(z)(Y_i(1)-Y_i(0))|X_i] = E[D_i(z)|X_i]\tau$. Thus, the second component in (i) is balanced by simply including $E[D_i(z)|X_i]$ into the instrument propensity score model. Therefore, using a single instrument propensity score for both numerator and denominator that only additionally incorporates $E[D_i(z)|X_i]$ for balancing seems like a reasonable middle-ground between the fully agnostic approach and imposing a lot of structure on the outcome process.

If a parametric model for $\mu(X_i,Z_i)$ is used, there are two additional aspects to consider. First, statistical inference has to be adjusted to the presence of the generated regressors.\footnote{This can be done by a standard asymptotic expansion that captures the additional parametric uncertainty in estimating $\mu(X_i,Z_i)$ or by using a bootstrap approach.} Second, there are two potentially highly correlated choices for inclusion into the balancing constraints, $E[D_i(1)|X_i]$ and $E[D_i(0)|X_i]$. In general, one of these components is enough for approximate unbiasedness. If $\phi(X_i)$ also contains other transformations of regressors, it seems reasonable to include the balancing constraint that adds more information. Including both can quickly lead to multicollinearity problems, in particular if variation in the instrument only has a mild effect on the conditional probability of choosing a certain treatment, i.e. if the expected share of compliers is not very large. This phenomenon and general finite sample performance of the different strategies are further investigated in Section (ref).

High-dimensional Models and Regularization

If the dimensionality of the transformation of regressors is large relative to the sample size, i.e. if $r > n$, solving for the parameters of the balancing constraints in (ref) is generally infeasible. However, with the cost of imposing some bias, balancing can be relaxed to tolerance level which makes estimation feasible. There is a growing literature on approximate balancing approaches in the context of estimation of treatment effects under selection on observables using standard propensity scores. zubizarreta2015stable proposes to minimize the squared $\ell_2$-norm of the balancing weights subject to an empirical balancing constraint. athey2018approximate combine approximate balancing using a supremum-type constraint with a high-dimensional linear model for the outcome under mild sparsity conditions for bias correction. zhao2019covariate suggests the use of regularized generalized linear models, reproducing kernel Hilbert spaces, and boosted trees as penalization strategies for general tailored loss functions. There are no intrinsic differences to the case of the LATE using balanced instrument propensity scores. We briefly outline the general principle of approximate balancing adapted to the tailored loss considered in this paper with a focus on commonly used parameter norm constraints. Let the transformation of regressors $\phi(X_i)$ be standardized. Consider the penalized optimization problem

align[align omitted — 174 chars of source]

with $J:\mathbb{R}^r\rightarrow{\mathbb{R}}$ being a (convex) penalty function. Typical choices are lasso $J(\theta) = ||\theta||_1$, ridge regression $J(\theta) = ||\theta||_2^2/2$, combinations thereof (elastic net), or any $\ell_a$-norm $J(\theta) = ||\theta||_a^a/a$ for $a \geq 1$. Convexity of $J(\cdot)$ assures that (ref) is equivalent to the following dual problem zhao2019covariate:

align[align omitted — 370 chars of source]

with weights $w_i$ depending on $\theta_{\lambda}$ as before. Thus, penalization relaxes the empirical balancing constraint from exact balance to tolerance level depending on the penalty of choice. In the simple case of a lasso type penalty, i.e. $a=1$, the empirical balancing is relaxed to tolerance level $\lambda$. For $\lambda = 0$, (ref) is equivalent to exact balancing while for $\lambda \rightarrow \infty$ all weights are set to a constant by the same argument as for the non-regularized dual problem in Section (ref). In case of the latter, the variance of the estimator is minimized as in Section (ref). However, as balancing constraints are imposed to decrease bias, see Section (ref) and Section (ref), a larger amount of regularization will introduce some bias into the estimation of the LATE. To solve this trade-off, zhao2019covariate proposes to select $\lambda$ via cross-validation using the covariate imbalance in the validation set for tuning and to compare standardized differences along different choices of $\lambda$.

Monte Carlo Study

Design

In this section, we compare the finite sample performance of the balancing estimator and some of the proposed extensions to standard estimation approaches from the literature. In particular, the question arises whether balancing approaches can outperform conventional methods in terms of point estimation risk and if so, what are the driving forces behind it. In Sections (ref) and (ref) we provide some reasoning why reductions in points estimation risk from balancing could be due to both channels, bias and variance. We evaluate the impact of using higher-order information from the treatment selection step as proposed in Section (ref) and disentangle it from the additional estimation noise and the effects of simple weight normalization. We consider multiple designs that highlight different features of the balancing approaches compared to conventional methods. Table (ref) contains the basic Monte Carlo design in spirit of a generalized Roy model:\footnote{See heckman2005structural for a simulation of a simple generalized Roy model using a continuous instrument.}

table[table omitted — 818 chars of source]

with $\rho = 0.5$, $\theta_0 = \ln((1-\delta)/\delta)$ and $\mu_{d}(\cdot)$ and $\mu_{y_1}(\cdot)$ see below. The parameter $\theta_0$ controls the degree of overlap based on the desired bounds $(\delta,1-\delta)$ for the instrument propensity score. $\rho \neq 0$ allows for selection on unobservables. We assume a nonzero correlation between unobservables driving the potential treatment outcome and treatment selection only to simplify analysis. This is without loss of generality as in the generalized Roy model under normality the functional form of the treatment effect is not affected by this choice up to a scaling factor depending on $\rho$.\footnote{The contribution of the unobservables to the conditional LATE is of the form $\rho f(X_i)$. If we assume that $Cov(\varepsilon_i(0),v_i) = \rho_0$, then the control function simply changes to $(\rho - \rho_0) f(X_i)$. As these correlations can have different signs, one could in principle increase the overall contribution of selection from unobservables to the conditional treatment effect for a given share of compliers.} Moreover, the model allows for additive separable contributions of observables and unobservables to the overall treatment effect which simplifies calculations. In particular, the LATE parameter for the model in Table (ref) is given by

align[align omitted — 231 chars of source]

with $F(\cdot)$ and $f(\cdot)$ being the cumulative distribution function and the density function of the univariate standard normal distribution respectively. Table (ref) contains the different design specifications for the treatment selection step and the potential treatment outcome.

table[table omitted — 330 chars of source]

The Roy model together with the functional form assumptions assure that none of the identification conditions for the LATE are violated. In particular, the linear index for the treatment choice with the given parameters imposes monotonicity and a share of compliers $E[D_i(1)-D_i(0)] = 0.5$ for all designs. In design A and design B, heterogeneity in causal effects is achieved through the correlation $\rho$ of the unobservable variables $\varepsilon_i(1)$ and $v_i$. For design C, there is heterogeneity stemming from both the direct contribution of observables to the potential outcome one and from the correlated unobservable variables. Design A represents the special case of a fully independent instrument as the observed confounder does not affect the treatment choice. It also corresponds to a controlled randomized experiment with close to one-sided noncompliance as $P(D_i=1|Z_i=1) > 0.9999$. The choice of the homogeneous mean function for the outcome assures that the unconditional LATE in (ref) is composed of two equally sized components, i.e. the overall contributions of the potential outcome mean and the “control function” part are identical. For this design, a simple Wald estimator would yield precise estimates for the LATE. Design B is a conditionally independent design with a homogeneous potential outcome mean function. The parameters are chosen to be favorable towards linear IV with exogenous covariates $X_i$ and instrument $Z_i$. Design C is a conditionally independent design with a nonlinear potential outcome mean. In this design, the semiparametric approaches that do not require specification of an outcome model still yield asymptotically unbiased results. The precise functional form of the potential outcome mean is not crucial. In general, other nonlinear mean functions with sufficient heterogeneity will produce qualitatively similar results.

table[table omitted — 1,105 chars of source]

Table (ref) contains all estimation approaches used in the Monte Carlo study. IV denotes the standard instrumental variables estimator that differs from the Wald estimator by additionally including $X_i$ as exogenous variable into the first-stage and the outcome equation. MLE and MLE(2) are the IPW estimators for the LATE using likelihood instrument propensities with (MLE(2)) and without (MLE) normalization of the inverse probability weights. B($X$) is the basic balancing estimator using the same regressors as the maximum likelihood approaches. As it contains an intercept, it is automatically weight normalized. B($D$) and B($D,X$) try to exploit the conditions for approximate unbiasedness in Proposition (ref) by using $E[D_i(0)|X_i]$ as additional (or only) variable for balancing. Theoretically, both should be able to reduce the bias of the components of the LATE. Despite their general infeasibility, they are included to see the potential gains from including additional information without the statistical noise from estimation of the generated regressor $E[D_i(0)|X_i]$. B($\hat{D}$) is similar to B($D$) but with estimated quantities, i.e. it uses a correctly specified model for $E[D_i(0)|X_i]$ by extracting the probability predictions from a probit model that uses $Z_i$ and $X_i$ as regressors at $Z_i = 0$. B($\hat{D}_m$) is similar to B($\hat{D}$) but uses a misspecified logit model instead of the true probit model. Note that, when incorrect specifications for the treatment selection are used as additional balancing constraints, this does not affect the asymptotic validity of the balancing IPW estimator for the LATE as even a misspecified quantity as a function of $X_i$ still leads to a valid balancing constraint but not necessarily one that minimizes bias in finite samples in the sense of Proposition (ref).

Results

Tables (ref), (ref), and (ref) contain the mean squared errors and absolute biases for designs A, B, and C and all estimation approaches from Table (ref). Mean squared errors are all normalized by the MSE of the linear instrumental variables estimator. Results are obtained from an experiment using 20000 Monte Carlo replications with overlap parameters $\delta =0.01, 0.02, 0.05$ and sample sizes $n=500, 1000$ yielding an expected number of $250$ and $500$ complier units respectively.

table[table omitted — 2,350 chars of source]
table[table omitted — 2,377 chars of source]
table[table omitted — 2,387 chars of source]

Overall, the results suggest that the unnormalized MLE approach is outclassed by all other methods due to its large finite sample variance. In general, most balancing approaches outperform the conventional MLE based IPW estimators MLE and MLE(2) by a substantial margin depending on the design. The differences are most pronounced for small sample sizes and small $\delta$. Adding higher-order information is generally helpful to reduce points estimation risk if there are no estimation problems due to a strong correlation between the different balancing variables as for method B($D,X$). While some of the differences between MLE(2) and the balancing approaches are due to a slightly reduced bias, the reduction in variance seems to be the main driver, in particular for designs with strong treatment effect heterogeneity. Unsurprisingly, for all approaches point estimation risk drops with an increase in the sample size.

For design A, the IV estimator serves as a benchmark as it is correctly specified. The MLE(2) and all balancing approaches except for $B(D)$ are very close in terms of point estimation risk. This is not surprising as under perfect randomization differences in the distributions of observed covariates across instrument levels can only happen by chance. Note that the infeasible $B(D)$ outclasses even the (correctly specified) parametric IV estimator. This is due to the fact that in design A, both the true instrument propensity scores and treatment selection probability are constant and thus the balanced IPW estimator $B(D)$ collapses to the standard Wald estimator. The latter efficiently incorporates the information on independence of the instrument in this homogeneous design and thus yields lower point estimation risk compared to the overparameterized IV.

For design B, the IV estimator only has a very small bias and thus can serve as a benchmark method. The balancing approaches other than B($D,X$) outperform MLE and, more importantly, MLE(2). The differences are particularly substantial for $n=500$ and $\delta = 0.01$. Here, even the worst balancing approach leads to a reduction in MSE of $92.5\%$ compared to MLE(2). Including higher-order information seems to be beneficial both theoretically (B($D$)) and empirically (B($\hat{D}$) and B($\hat{D}_m$)). In fact, using estimated treatment selection probabilities for balancing seems to be even slightly superior over using true probabilities by a margin of $0$ to $3$ percentage points. Surprisingly, the slightly misspecified model for balancing (B($\hat{D}_m$)) outperforms all other IPW approaches by a margin of at least $1$ to $9$ percentage points. The infeasible B($D,X$) has convergence problems for $n=500$ and $\delta=0.01, 0.02$ due to the strong correlation between $X_i$ and $E[D_i(0)|X_i]$ in the Roy model under normality similar to multicollinearity problems of standard control function approaches in the spirit of heckman1979sample.

For design C, the IV estimator will be severely biased due to the nonlinear potential outcome mean. Here, the differences between balancing and MLE based methods are most pronounced. In fact, all balancing approaches lead to a substantial reduction in MSE compared to IV, MLE and MLE(2). In fact, the consistent MLE(2) needs a much larger sample size to make up for its larger variance compared to the biased IV estimator. All balancing approaches, however, outperform IV by 27 to 69 percentage points in terms of MSE. In general, the gains over IV are more pronounced for larger samples and larger values of $\delta$. Balancing methods also outperform MLE(2) by a factor between $1.7$ ($\delta = 0.05$, $n = 1000$) to $17.4$ ($\delta= 0.01$, $n = 500$) in terms of MSE. There is a clear hierarchy between the different balancing and MLE methods similar but not identical to design B, i.e. in terms of MSE we have that B($\hat{D}_m$) < B($\hat{D}$) < B(D) < B(X) < MLE(2) < MLE. Again, the slightly misspecified model comes out on top with a margin of at least $1$ to $3$ percentage points. As in design B, the infeasible B($D,X$) has convergence problems for $n=500$ and $\delta=0.01, 0.02$. If it is feasible it ranks as the worst balancing approach but still substantially above MLE(2). It is interesting that in this design the superior performance of the balancing approaches compared to MLE(2) is mainly due to a reduction in variance as they sometimes have a larger bias than MLE(2). This is not a contradiction to the bias Proposition (ref), as in this design treatment effects are highly heterogeneous even for the compliers. Therefore, there is no guarantee for a small finite sample bias while the effects on the variance through pushing the inverse probability weights towards more moderate values as outlined in Section (ref) are still in place.

The simulation results are evidence that balancing is a superior strategy compared to using likelihood instrument propensity scores for the inverse probability weighting estimator for the LATE in finite samples. The main channel is the reduction in variance, in particular for heterogeneous designs and small sample sizes. This is in line with the results in the literature on balancing weights for selection on observables, see e.g. imai2014covariate. Moreover, if one is willing to model the treatment selection process, then including higher-order information through estimates of $E[D_i(0)|X_i]$ as balancing constraints seems to be beneficial for estimation of the LATE even under mild misspecification. However, if the information that enters the treatment step is heavily correlated with other regressors used for balancing, then using the generated regressors only is superior to incorporating them jointly in all designs considered. In applied research, the severity can be examined by looking at the Hessian of the empirical (negative) tailored loss function. Alternatively, all information can be used jointly with regularization as suggested in Section (ref).

Empirical Application

This section contains an illustration of the empirical balancing method to the case of one-sided noncompliance. We re-evaluate of the causal effect of 401(k) retirement plans on private net total financial assets. 401(k) plans are employer-provided, tax-deferred saving plans with partially matched monetary contributions of the employee. The question is whether these plans help to increase net private savings or asset holdings. The evaluation of the effect is complicated by both observed and unobserved individual heterogeneity such as income levels or preferences that might be related to both asset level and the propensity to select the plan. This implies that a simple comparison of savings between participating and non-participating units is likely to yield a biased estimate for the causal effect. We use 401(k) eligibility as as a conditionally independent instrument for participation as the former is only provided through the employer. Identification then rests on the assumption that eligibility is independent of possible confounding unobserved heterogeneity after conditioning on a set of observed characteristics such as income and other financial and socioeconomic background variables. This and similar identification strategies are widely used in the literature benjamin2003does,abadie2003semiparametric,chernozhukov2004effects,chernozhukov2017double, see also poterba1995401 for a critical assessment.

The data is an excerpt from the Survey of Income and Program Participation (SIPP) of 1991. We apply the same sample restrictions as in abadie2003semiparametric. The final data set contains 9275 observations. The outcome is measured as total net financial assets in US dollars. The treatment variable is an indicator for participation in a 401(k) plan. The instrument is a binary indicator for eligibility. Observed confounding variables are age, income, family size, education, and other financial and socioeconomic background variables as in chernozhukov2017double. Note that instead of using income measured in dollars we use the natural logarithm of income instead.\footnote{Income has a heavily right-skewed distribution that leads to instability in the nonparametric instrument propensity score estimation step. The log transformation stabilizes the distribution and avoids extreme propensity scores as a result of heavy extrapolation to extremely high income levels.}

We estimate the causal effect of 401(k) participation on net total financial assets using two inverse probability weighting approaches. We also provide standard Wald estimates and replicate the IV result with covariates as in abadie2003semiparametric for comparison. For the estimators using inverse probability weights we consider both a simple logistic model and the semiparametric empirical balancing method for estimation of the instrument propensity scores. For both methods we include all possible interactions of the binary regressors. For the empirical balancing method we employ an additive spline basis with degree and number of nodes selected via leave-one-out cross-validation. The loss function used in the cross-validation is the tailored loss function in (ref), i.e. the model is chosen to be optimal in terms of its out-of-sample balance. We also experimented with a non-additive tensor basis but it performed worse in the out-of-sample evaluation.

figure[figure omitted — 728 chars of source]

Figures (ref) and (ref) contain the kernel density estimates of the instrument propensity scores for eligible and non-eligible units for both the logistic and the empirical balancing scores. Contrary to the more restrictive logistic model, the semiparametric balancing estimator seems to suggest an almost bimodal distribution for the eligible units with more probability mass around probabilities of 0.65 to 0.75, less between 0.25 and 0.5, and more extreme propensities in the left tail. One can see that for both models there is sufficient overlap in the distribution and minimum scores are sufficiently far away from the boundary. Thus we do not expect any irregularity problems regarding the identification of the treatment effect parameters khan2010irregular,heiler2020inference.

table[table omitted — 782 chars of source]

Table (ref) contains the point estimates of the causal effect and the corresponding standard errors for all methods. The estimates are all significant on a 1% significance level and broadly consistent with the results previously reported in the literature. The Wald estimate of 26771\$ differs severely from the other estimates as it is based on the overly restrictive identification assumption of unconditionally independent eligibility. The IV estimates using the model by abadie2003semiparametric and the parametric IPW estimator yield a very similar result of an around 9500\$ increase in total net financial assets. The more robust semiparametric balancing estimator suggest about a 30% larger effect. Note that, despite its higher flexibility, empirical balancing has a lower standard error than the more restricted logistic model. This is a consequence of the variance stabilizing effect outlined in Section (ref). Empirical balancing reduces the presence of extreme propensity scores that can lead to extreme weights which heavily affect the point estimates in finite samples.

Concluding Remarks

In this paper we develop an estimation method for the local average treatment effect that relies on empirical balancing constraints to improve the internal validity of the causal estimand. It has favorable bias and variance properties compared to conventional approaches, is asymptotically normal, and reaches the semiparametric efficiency bound if a sufficiently flexible model is used. Moreover, it does not rely on the use of outcome or treatment selection information and is easy to implement. Monte Carlo simulations suggest that the theoretical advantages translate well to finite samples.

As the estimator for the local average treatment effect has the typical ratio form, future work should be done to see whether finite sample bias can be further reduced by imposing alternative balancing constraints that do not have to go through two separate reduced form estimates but instead minimize the bias of the ratio directly. In addition, the empirical performance of different modifications and extensions require further attention. For example, using different instrument propensities for the two reduced form type components that also require different normalization schemes could be beneficial in designs with strong heterogeneities in the potential outcomes. Moreover, the simple balancing approach that uses higher-order information and basic covariates together used in the Monte Carlo study could be combined with e.g. $\ell_2$-regularization to overcome multicollinearity problems. An additional contribution would be to derive conditions required for the higher-order approaches to conduct semiparametric inference in the presence of generated regressors from nonparametric models.