Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
37,260 characters · 11 sections · 47 citation commands
Double Robustness for Complier Parameters and a Semiparametric Test for Complier Characteristics
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 \fi
\if01 {
} \fi
keywords: Instrumental Variable; Kappa Weight; Semiparametric Efficiency.
Average complier characteristics help to assess the external validity of any study that uses instrumental variable identification angrist1998children,angrist2013extrapolate,swanson2013commentary,baiocchi2014instrumental,marbach2020profiling; whose treatment effects are we estimating when we use a particular instrument? We propose a semiparametric hypothesis test, free of functional form restrictions, to evaluate (i) whether two different instruments induce subpopulations of compliers with the same observable characteristics on average, and (ii) whether compliers have observable characteristics that are the same as the full population on average. It appears that no semiparametric test previously exists for this important question about the external validity of instruments, despite the popularity of reporting average complier characteristics in empirical research, e.g. abdulkadirouglu2014elite. By developing this hypothesis test, we equip empirical researchers with a new robustness check.
Equipped with this new test, we replicate, extend, and test previous findings about the impact of childbearing on female labor supply. In a seminal paper, angrist1998children use two different instrumental variables: twin births and same-sex siblings. The two instruments give rise to two substantially different local average treatment effect (LATE) estimates for the reduction in weeks worked due to a third child: -3.28 (0.63) and -6.36 (1.18), respectively, where the standard errors are in parentheses. angrist2013extrapolate attribute the difference in LATE estimates to a difference in average complier characteristics, i.e. a difference in average covariates for instrument specific complier subpopulations, writing that “twins compliers therefore are relatively more likely to have a young second-born and to be highly educated.” We find weak evidence in favor of the explanation that twins compliers are more likely to have a young second-born. We do not find evidence that twins compliers have a significantly different education level than same-sex compliers.
Our test is based on a new doubly robust estimator, which we call the automatic $\kappa$ weight (Auto-$\kappa$). To prove the validity of the test, we characterize the doubly robust moment function for average complier characteristics, which appears to have been previously unknown. More generally, we study low dimensional complier parameters that are identified using a binary instrumental variable $Z$, which is valid conditional on a possibly high dimensional vector of covariates $X$. angrist1996identification prove that identification of LATE based on the instrumental variable does not require any functional form restrictions. Using $\kappa$ weighting, abadie2003semiparametric extends identification for a broad class of complier parameters. As our main theoretical result, we characterize the doubly robust moment function for this class of complier parameters by augmenting $\kappa$ weighting with the classic Wald formula. Our main result answers the open question posed by sloczynski2018general of how to characterize the doubly robust moment function for the full class, and it generalizes the well known result of tan2006regression, who characterizes the doubly robust moment function for LATE. By characterizing the doubly robust moment function for abadie2003semiparametric's class of complier parameters, we handle the new and economically important case of average complier characteristics.
The doubly robust moment function confers many favorable properties for estimation. As its name suggests, it provides double robustness to misspecification robins1995semiparametric as well as the mixed bias property chernozhukov2018original,rotnitzky2021characterization. As such, it allows for estimation of models in which the treatment effect for different individuals may vary flexibly according to their covariates frolich2007nonparametric,ogburn_doubly_2015. It also allows for nonlinear models abadie2003semiparametric,cheng2009efficient, which are often appropriate when outcome $Y$ and treatment $D$ are binary, and therefore avoids the issue of negative weights in misspecified linear models blandhol2022tsls. Moreover, it allows for model selection of covariates and their transformations using machine learning, as emphasized in the targeted machine learning van2006targeted,zheng2011cross,luedtke2016statistical,van2018targeted and debiased machine learning belloni2017program,chernozhukov2016locally,chernozhukov2018original,chernozhukov2021simple literatures. A doubly robust estimator that combines both the $\kappa$ weight and Wald formulations not only guards against misspecification but also debiases machine learning. Finally, it is semiparametrically efficient in many cases hasminskii1979nonparametric,robinson1988root,bickel1993efficient,newey1994asymptotic,robins1995semiparametric,hong2010semiparametric.
The structure of the paper is as follows. Section (ref) defines the class of complier parameters from abadie2003semiparametric. Section (ref) summarizes our main insight: the doubly robust moment for a complier parameter combines the familiar Wald and $\kappa$ weight formulations. Section (ref) formalizes this insight for the full class of complier parameters. Section (ref) develops the practical implication of our main insight: a semiparametric test to evaluate differences in observable complier characteristics, which we use to revisit angrist1998children. Section (ref) concludes. Appendix (ref) proposes a machine learning estimator that we call the automatic $\kappa$ weight (Auto-$\kappa$), which we use to implement our proposed test.
This paper was previously circulated under a different title singh2019biased.
Suppose we are interested in the effect of a binary treatment $D$ on a continuous outcome $Y$ in $\mathcal{Y}$, a subset of $\mathbb{R}$. There is a binary instrumental variable $Z$ available, as well as a potentially high dimensional covariate $X$ in $\mathcal{X}$, a subset of $\mathbb{R}^{dim(X)}$. We observe $n$ independent and identically distributed observations $(W_i)$, $(i=1,...,n)$, where $W=(Y,D,Z,X^{\top})^{\top}$ concatenates the random variables. Following the notation of angrist1996identification, we denote by $Y^{(z,d)}$ the potential outcome under the intervention $Z=z$ and $D=d$. We denote by $D^{(z)}$ the potential treatment under the intervention $Z=z$. Compliers are the subpopulation for whom $D^{(1)}>D^{(0)}$. We place standard assumptions for identification.
Independence states that the instrument $Z$ is as good as randomly assigned conditional on covariates $X$. Exclusion imposes that the instrument $Z$ only affects the outcome $Y$ via the treatment $D$. We can therefore simplify notation: $Y^{(d)}=Y^{(1,d)}=Y^{(0,d)}$. Overlap ensures that there are no covariate values for which the instrument assignment is deterministic. Monotonicity rules out the possibility of defiers: individuals who will always pursue an opposite treatment status from their instrument assignment.
angrist1996identification prove identification of the local average treatment effect (LATE) using Assumption (ref). abadie2003semiparametric extends identification for a broad class of complier parameters.
For a given instrumental variable $Z$, one may define the average complier characteristics as a special case of Definition (ref). This causal parameter summarizes the observable characteristics of the subpopulation of compliers who are induced to take up or refuse treatment $D$ based on the instrument assignment $Z$. It is an important parameter to estimate because it aids the interpretation of LATE. As we will see in Section (ref), this causal parameter can help to reconcile different LATE estimates obtained with different instruments.
We provide intuition for our key insight that a doubly robust moment for a complier parameter has two components: the Wald formula and the $\kappa$ weight. For clarity, we focus on the familiar example of local average treatment effect (LATE) in this initial discussion: $\theta_0=E\{Y^{(1)}-Y^{(0)}\mid D^{(1)}>D^{(0)}\}$. In subsequent sections, we study the entire class of complier parameters in Definition (ref), including the new case of average complier characteristics.
Under Assumption (ref), LATE can be identified as $$ \theta_0 =\frac{E\left\{E(Y\mid Z=1,X)-E(Y\mid Z=0,X)\right\}}{E\left\{E(D\mid Z=1,X)-E(D\mid Z=0,X)\right\}} $$ following frolich2007nonparametric. We call this expression the expanded Wald formula.
The direct Wald approach involves estimating the reduced form regression $E(Y\mid Z,X)$ and first stage regression $E(D\mid Z,X)$, then plugging these estimates into the expanded Wald formula. Such an approach is called the plug-in, and it is valid only when both regressions are estimated with correctly specified and unregularized models. It is not a valid approach when either regression is incorrectly specified, leading to the name “forbidden regression” angrist2008mostly. It is also invalid when the covariates are high dimensional and a regularized machine learning estimator is used to estimate either regression. The matching procedure of frolich2007nonparametric faces similar limitations.
In seminal work, abadie2003semiparametric proposes an alternative formulation in terms of the $\kappa$ weights $$ \kappa^{(0)}(W)= (1-D)\frac{(1-Z)-\{1-\pi_0(X)\}}{\{1-\pi_0(X)\}\pi_0(X)},\quad \kappa^{(1)}(W)= D\frac{Z - \pi_0(X)}{\{1-\pi_0(X)\}\pi_0(X)} $$ where $\pi_0(X)=\text{\normalfont pr}(Z=1\mid X)$ is the instrument propensity score. The $\kappa$ weights have the property that $$ \theta_0=\omega^{-1}E\{\kappa^{(1)}(W) Y-\kappa^{(0)}(W) Y\},\quad \omega= E\left\{1-\frac{D(1-Z)}{1-\pi_0(X)}-\frac{(1-D)Z}{\pi_0(X)}\right\}. $$ In words, the mean of the product of $Y$ and $\kappa^{(d)}(W)$ gives, up to a scaling, the expected potential outcome $Y^{(d)}$ of compliers when treatment is $D=d$. As an aside, abadie2003semiparametric also introduces a third weight $\kappa(W)$ for parameters that belong to the third case in Definition (ref).
The $\kappa$ weight approach would involve estimating the propensity score $\hat{\pi}$ and plugging this estimate into the $\kappa$ weight formula. Intuitively, the $\kappa$ weight approach is like a multistage inverse propensity weighting. Impressively, it remains agnostic about the functional form of the reduced form regression $E(Y\mid Z,X)$ and first stage regression $E(D\mid Z,X)$. It is valid only when $\hat{\pi}$ is estimated with a correctly specified and unregularized model. It is invalid if $\hat{\pi}$ is incorrectly specified or if covariates are high dimensional and a regularized machine learning estimator is used to estimate $\hat{\pi}$. Moreover, the inversion of $\hat{\pi}$ can lead to numerical instability in high dimensional settings.
Next, we introduce the moment function and doubly robust moment function formulations of LATE. For the special case of LATE, these formulations were first derived by tan2006regression with the goal of addressing misspecification of the regressions and the propensity score. Consider the expanded Wald formula. Rearranging and using the notation $V=(Y,D)^{\top}$ as a column vector, $ \gamma_0(Z,X)=E(V\mid Z,X)$ as a vector valued regression, and $
$ as a row vector, we arrive at the moment function formulation of LATE: $$ E\left[
\{\gamma_0(1,X)-\gamma_0(0,X)\} \right] =0 if and only if \theta=\theta_0. $$ Denote the the Horvitz-Thompson balancing weight as $$ \alpha_0(Z,X)=\frac{Z}{\pi_0(X)}-\frac{1-Z}{1-\pi_0(X)},\quad \pi_0(X)=\normalfont pr(Z=1\mid X). $$ \cite{tan2006regression} shows that for LATE, the doubly robust moment function is $$ E\left[
\{\gamma_0(1,X)-\gamma_0(0,X)\} +\alpha_0(Z,X)
\{V-\gamma_0(Z,X)\}\right] =0 if and only if \theta=\theta_0. $$ The doubly robust formulation remains valid if either the vector valued regression $\gamma_0$ or propensity score $\pi_0$ is incorrectly specified.
Our key observation is the connection between the $\kappa$ weight and the balancing weight $\alpha_0$. This simple observation will allow us to characterize the doubly robust moment function for a broad class of complier parameters, generalizing tan2006regression to the full class defined by abadie2003semiparametric.
Next, we formalize the sense in which the balancing weight $\alpha_0$ represents the functional $\gamma\mapsto E\left\{
\gamma(1,X)-\gamma(0,X)\right\}$ that appears in the moment formulation of LATE and the extended Wald formula.
An immediate consequence of Proposition (ref) is that $$ E\left\{
\gamma(1,X)-\gamma(0,X)\right\}=E\left\{\alpha_0(Z,X)
\gamma(Z,X)\right\} for any \gamma. $$
In summary, Proposition (ref) shows that the $\kappa$ weight is a reparametrization of the balancing weight $\alpha_0$. Meanwhile, Proposition (ref) shows that the balancing weight appears in the Riesz representer to the moment formulation of LATE, i.e. the expanded Wald formula. We conclude that the $\kappa$ weight is essentially the Riesz representer to the Wald formula. In seminal work, newey1994asymptotic demonstrates that a doubly robust moment is constructed from a moment formulation and its Riesz representer. Therefore the doubly robust moment for complier parameters must combine the Wald formula and the $\kappa$ weight.
With the general doubly robust moment function, one can propose flexible, semiparametric tests for complier parameters. In particular, the semiparametric tests may involve regularized machine learning for flexible estimation and model selection of (i) the regression $\hat{\gamma}$ in a way that approximates nonlinearity and heterogeneity, and (ii) the balancing weight $\hat{\alpha}$ in a way that guarantees balance. In Section (ref), we instantiate such a test to compare observable characteristics of compliers.
As explained in Appendix (ref), we avoid the numerically unstable step of estimating and inverting $\hat{\pi}$ that appears in tan2006regression,belloni2017program,chernozhukov2018original. We replace it with the numerically stable step of estimating $\hat{\alpha}$ directly, extending techniques of chernozhukov2018learning to the instrumental variable setting. We call this extension automatic $\kappa$ weighting (Auto-$\kappa$), and demonstrate how it applies to the new and economically important case of average complier characteristics.
In summary, our main theoretical result allows us to combine the classic Wald and $\kappa$ weight formulations for the entire class of complier parameters in Definition (ref), including average complier characteristics, while also updating them to incorporate machine learning.
We now state our main theoretical result, which is the doubly robust moment for the class of complier parameters in Definition (ref). This result formalizes the intuition of Section (ref), and it justifies the hypothesis test in Section (ref). It is convenient to divide the main result into two statements for clarity. Theorem (ref) handles the first and second cases in Definition (ref), while Theorem (ref) handles the third case in Definition (ref).
In the doubly robust moment function $\psi(w,\gamma,\alpha,\theta)=m(w,\gamma,\theta)+\phi(w,\gamma,\alpha,\theta)$, we generalize our insight from Section (ref). The first term $m(w,\gamma,\theta)$ is essentially a generalized Wald formula. The second term $\phi(w,\gamma,\alpha,\theta)$ is essentially a product between the $\kappa$ weight and a generalized regression residual. In the language of semiparametrics, we augment the $\kappa$ weight with the Wald formula. Equivalently, we debias the Wald formula with the $\kappa$ weight.
The doubly robust moment function $\psi$ remains valid if either $\gamma_0$ or $\alpha_0$ is misspecified, i.e. $$ 0=E\{\psi(W,\gamma,\alpha_0,\theta_0)=E[\psi(W,\gamma_0,\alpha,\theta_0)\} \text{ for any } \gamma,\alpha. $$ In the former expression, $\gamma_0$ may be misspecified yet $\psi$ remains valid as an estimating equation. In the latter, $\alpha_0$ may be misspecified yet $\psi$ remains valid as an estimating equation. Theorem (ref) demonstrates that all complier parameters in cases 1 and 2 of Definition (ref) have a doubly robust moment function $\psi$ with a common structure. As such, we are able to analyze all of these causal parameters with the same argument. Case 3 of Definition (ref) is more involved, but we show that it shares the common structure as well.
This time, the doubly robust moment function $\psi$ remains valid if either $\tilde{\gamma}_0$ or $\tilde{\alpha}_0$ is misspecified, i.e. $$ 0=E\{\psi(W,\tilde{\gamma},\tilde{\alpha}_0,\theta_0)=E[\psi(W,\tilde{\gamma}_0,\tilde{\alpha},\theta_0)\} \text{ for any } \tilde{\gamma},\tilde{\alpha}. $$ In the former expression, $\tilde{\gamma}_0$ may be misspecified yet $\psi$ remains valid as an estimating equation. In the latter, $\tilde{\alpha}_0$ may be misspecified yet $\psi$ remains valid as an estimating equation.
In Section (ref), we translate this general characterization of the doubly robust moment into a practical hypothesis test to evaluate the external validity of instruments. In Appendix (ref), we translate this general characterization into general machine learning estimators for complier parameters, which we use to implement the hypothesis test. In particular, we consider direct estimation of the balancing weight, a procedure that we call automatic $\kappa$ weighting (Auto-$\kappa$).
As a corollary, we characterize the doubly robust moment for average complier characteristics, which appears to have been previously unknown. Using the new doubly robust moment, we propose a hypothesis test, free of functional form restrictions, to evaluate (i) whether two different instruments induce subpopulations of compliers with the same observable characteristics on average, and (ii) whether compliers have observable characteristics that are the same as the full population on average.
Suppose we wish to test the null hypothesis that two different instruments $Z_1$ and $Z_2$ induce complier subpopulations with the same observable characteristics on average. Denote by $\hat{\theta}_1$ and $\hat{\theta}_2$ the estimators for average complier characteristics using the different instruments $Z_1$ and $Z_2$, respectively. One may construct machine learning estimators $\hat{\theta}_1$ and $\hat{\theta}_2$ based on the doubly robust moment function in Corollary (ref). In Appendix (ref), we instantiate automatic $\kappa$ weight (Auto-$\kappa$) estimators of this type. The following procedure allows us to test the null hypothesis from some estimator $\hat{C}$ for the asymptotic variance $C$ of $\hat{\theta}=(\hat{\theta}^{\top}_1,\hat{\theta}^{\top}_2)^{\top}$. In Appendix (ref), we provide an explicit variance estimator $\hat{C}$ based on Auto-$\kappa$ as well.
Algorithm (ref) can also test the null hypothesis that compliers have observable characteristics that are the same as the full population on average. $\hat{\theta}_1$ is as before, $\hat{\theta}_2=n^{-1}\sum_{i=1}^nf(X_i)$, and $\hat{C}$ updates accordingly.
Corollary (ref) is our main practical result: justification of a flexible hypothesis test to evaluate a difference in average complier characteristics. It appears that no semiparametric test previously exists for this important question about the external validity of instruments. By developing this hypothesis test, we equip empirical researchers with a new robustness check. This practical result follows as a consequence of our main insight in Section (ref) and our main theoretical result in Section (ref). In Appendix (ref), we verify the conditions of Corollary (ref) for Auto-$\kappa$ under weak regularity assumptions.
With this practical result, we revisit a classic empirical paper in labor economics to test whether two different instruments induce different average complier characteristics. angrist1998children estimate the impact of childbearing $D$ on female labor supply $Y$ in a sample of 394,840 mothers, aged 21--35 with at least two children, from the 1980 Census. The first instrument $Z_1$ is twin births: $Z_1$ indicates whether the mother's second and third children were twins. The second instrument $Z_2$ is same-sex siblings: $Z_2$ indicates whether the mother's initial two children were siblings with the same sex. The authors reason that both $(Z_1,Z_2)$ are quasi random events that induce having a third child.
The two instruments give rise to two LATE estimates for the reduction in weeks worked due to a third child: -3.28 (0.63) for $Z_1$ and -6.36 (1.18) for $Z_2$, where the standard errors are in parentheses. angrist2013extrapolate attribute the difference in LATE estimates to a difference in average complier characteristics, i.e. a difference in average covariates for instrument specific complier subpopulations. The authors use parametric $\kappa$ weights, report point estimates without standard errors, and conclude that “twins compliers therefore are relatively more likely to have a young second-born and to be highly educated.”
We replicate, extend, and test these previous findings. In their parametric $\kappa$ weight approach, angrist2013extrapolate estimate $\pi_{0}(X)$ using a logistic model with polynomials of continuous covariates. In our semiparametric Auto-$\kappa$ approach, we expand the dictionary to higher order polynomials, include interactions between the instrument and covariates, and directly estimate and regularize the balancing weights. Crucially, our main result allows us to conduct inference, and to test whether the instruments $Z_1$ and $Z_2$ induce differences in the observable complier characteristics suggested by previous work.
Table (ref) summarizes results. In Columns 1, 2, 5, and 6, we find similar point estimates to angrist2013extrapolate, given in Row 1. Columns 3, 4, 7, and 8 report $p$ values for tests of the null hypothesis that average complier characteristics are equal for the twins and same-sex instruments. We find weak evidence in favor of the explanation that twins compliers are more likely to have a young second-born. We do not find evidence that twins compliers have a significantly different education level than same-sex compliers.
We propose a semiparametric test to evaluate (i) whether two different instruments induce subpopulations of compliers with the same observable characteristics on average, and (ii) whether compliers have observable characteristics that are the same as the full population on average. This hypothesis test is a flexible and practical robustness check for the external validity of instrumental variables. We use the test to reinterpret the difference in LATE estimates that angrist1998children obtain when using two different instrumental variables. Specifically, we implement a machine learning update to $\kappa$ weighting that we call the automatic $\kappa$ weight (Auto-$\kappa$). To justify the test, we develop new econometric theory. Most notably, we characterize the doubly robust moment function for the entire class of complier parameters from abadie2003semiparametric, answering an open question in the semiparametric literature in order to handle the new and economically important case of average complier characteristics.