The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
37,260 characters
Double Robustness for Complier Parameters and a Semiparametric Test for Complier Characteristics
\def\spacingset#1{\renewcommand{\baselinestretch}
{#1}\small\normalsize} \spacingset{1}
\if11
{
\title{\bf Double Robustness for Complier Parameters and a Semiparametric Test for Complier Characteristics}
\author{Rahul Singh \hspace{.2cm}\\
Department of Economics, Massachusetts Institute of Technology,\\ Cambridge, Massachusetts 02142, U.S.A.\\
and \\
Liyang Sun \\
Center for Monetary and Financial Studies,\\ Madrid, 28014, Spain }
\maketitle
} \fi
\if01
{
\bigskip
\bigskip
\bigskip
\begin{center}
{\LARGE\bf Double Robustness for Complier Parameters and a Semiparametric Test for Complier Characteristics}
\end{center}
\medskip
} \fi
\maketitle
\begin{abstract}
We propose a semiparametric test to evaluate (i) whether different instruments induce subpopulations of compliers with the same observable characteristics on average, and (ii) whether compliers have observable characteristics that are the same as the full population on average. The test is a flexible robustness check for the external validity of instruments. We use it to reinterpret the difference in LATE estimates that \cite{angrist1998children} obtain when using different instrumental variables. To justify the test, we characterize the doubly robust moment for \cite{abadie2003semiparametric}’s class of complier parameters, and we analyze a machine learning update to $\kappa$ weighting.
\end{abstract}
\textit{keywords:} Instrumental Variable; Kappa Weight; Semiparametric Efficiency.
\section{Introduction and related work}\label{sec:intro}
Average complier characteristics help to assess the external validity of any study that uses instrumental variable identification \citep{angrist1998children,angrist2013extrapolate,swanson2013commentary,baiocchi2014instrumental,marbach2020profiling}; whose treatment effects are we estimating when we use a particular instrument? We propose a semiparametric hypothesis test, free of functional form restrictions, to evaluate (i) whether two different instruments induce subpopulations of compliers with the same observable characteristics on average, and (ii) whether compliers have observable characteristics that are the same as the full population on average. It appears that no semiparametric test previously exists for this important question about the external validity of instruments, despite the popularity of reporting average complier characteristics in empirical research, e.g. \citet[Table 2]{abdulkadirouglu2014elite}. By developing this hypothesis test, we equip empirical researchers with a new robustness check.
Equipped with this new test, we replicate, extend, and test previous findings about the impact of childbearing on female labor supply. In a seminal paper, \cite{angrist1998children} use two different instrumental variables: twin births and same-sex siblings. The two instruments give rise to two substantially different local average treatment effect (LATE) estimates for the reduction in weeks worked due to a third child: -3.28 (0.63) and -6.36 (1.18), respectively, where the standard errors are in parentheses. \cite{angrist2013extrapolate} attribute the difference in LATE estimates to a difference in average complier characteristics, i.e. a difference in average covariates for instrument specific complier
subpopulations, writing that ``twins compliers therefore are relatively more
likely to have a young second-born and to be highly educated.'' We find weak evidence in favor of the explanation that twins compliers are more
likely to have a young second-born. We do not find evidence that twins compliers have a significantly different education level than same-sex compliers.
Our test is based on a new doubly robust estimator, which we call the automatic $\kappa$ weight (Auto-$\kappa$). To prove the validity of the test, we characterize the doubly robust moment function for average complier characteristics, which appears to have been previously unknown. More generally, we study low dimensional complier parameters that are identified using a binary instrumental variable $Z$, which is valid conditional on a possibly high dimensional vector of covariates $X$. \cite{angrist1996identification} prove that identification of LATE based on the instrumental variable does not require any functional form restrictions. Using $\kappa$ weighting, \cite{abadie2003semiparametric} extends identification for a broad class of complier parameters. As our main theoretical result, we characterize the doubly robust moment function for this class of complier parameters by augmenting $\kappa$ weighting with the classic Wald formula. Our main result answers the open question posed by \cite{sloczynski2018general} of how to characterize the doubly robust moment function for the full class, and it generalizes the well known result of \cite{tan2006regression}, who characterizes the doubly robust moment function for LATE. By characterizing the doubly robust moment function for \cite{abadie2003semiparametric}'s class of complier parameters, we handle the new and economically important case of average complier characteristics.
The doubly robust moment function confers many favorable properties for estimation. As its name suggests, it provides double robustness to misspecification \citep{robins1995semiparametric} as well as the mixed bias property \citep{chernozhukov2018original,rotnitzky2021characterization}. As such, it allows for estimation of models in which the treatment effect for different individuals may vary flexibly according to their covariates \citep{frolich2007nonparametric,ogburn_doubly_2015}. It also allows for nonlinear models \citep{abadie2003semiparametric,cheng2009efficient}, which are often appropriate when outcome $Y$ and treatment $D$ are binary, and therefore avoids the issue of negative weights in misspecified linear models \citep{blandhol2022tsls}. Moreover, it allows for model selection of covariates and their transformations using machine learning, as emphasized in the targeted machine learning \citep{van2006targeted,zheng2011cross,luedtke2016statistical,van2018targeted} and debiased machine learning \citep{belloni2017program,chernozhukov2016locally,chernozhukov2018original,chernozhukov2021simple} literatures. A doubly robust estimator that combines both the $\kappa$ weight and Wald formulations not only guards against misspecification but also debiases machine learning. Finally, it is semiparametrically efficient in many cases \citep{hasminskii1979nonparametric,robinson1988root,bickel1993efficient,newey1994asymptotic,robins1995semiparametric,hong2010semiparametric}.
The structure of the paper is as follows. Section~\ref{sec:framework} defines the class of complier parameters from \cite{abadie2003semiparametric}. Section~\ref{sec:insight} summarizes our main insight: the doubly robust moment for a complier parameter combines the familiar Wald and $\kappa$ weight formulations. Section~\ref{sec:double} formalizes this insight for the full class of complier parameters. Section~\ref{sec:test} develops the practical implication of our main insight: a semiparametric test to evaluate differences in observable complier characteristics, which we use to revisit \cite{angrist1998children}. Section~\ref{sec:conc} concludes. Appendix~\ref{sec:estimation} proposes a machine learning estimator that we call the automatic $\kappa$ weight (Auto-$\kappa$), which we use to implement our proposed test.
This paper was previously circulated under a different title \citep{singh2019biased}.
\section{Framework}\label{sec:framework}
Suppose we are interested in the effect of a binary treatment $D$ on a continuous outcome $Y$ in $\mathcal{Y}$, a subset of $\mathbb{R}$. There is a binary instrumental variable $Z$ available, as well as a potentially high dimensional covariate $X$ in $\mathcal{X}$, a subset of $\mathbb{R}^{dim(X)}$. We observe $n$ independent and identically distributed observations $(W_i)$, $(i=1,...,n)$, where $W=(Y,D,Z,X^{\top})^{\top}$ concatenates the random variables. Following the notation of \cite{angrist1996identification}, we denote by $Y^{(z,d)}$ the potential outcome under the intervention $Z=z$ and $D=d$. We denote by $D^{(z)}$ the potential treatment under the intervention $Z=z$. Compliers are the subpopulation for whom $D^{(1)}>D^{(0)}$. We place standard assumptions for identification.
\begin{assumption}[Instrumental variable identification] \label{assumption:id}
Assume
\begin{enumerate}
\item Independence: $\{Y^{(z,d)}\},\{D^{(z)}\} \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} Z\mid X$ for $d=0,1$ and $z=0,1$.
\item Exclusion: $\text{\normalfont pr}\{Y^{(1,d)}=Y^{(0,d)}\mid X\}=1$ for $d=0,1$.
\item Overlap: $\pi_0(X)=\text{\normalfont pr}(Z=1\mid X)$ is in $(0,1)$.
\item Monotonicity: $\text{\normalfont pr}\{D^{(1)} \geq D^{(0)}\mid X\}=1$ and $\text{\normalfont pr}\{D^{(1)} > D^{(0)}\mid X\}>0$.
\end{enumerate}
\end{assumption}
Independence states that the instrument $Z$ is as good as randomly assigned conditional on covariates $X$. Exclusion imposes that the instrument $Z$ only affects the outcome $Y$ via the treatment $D$. We can therefore simplify notation: $Y^{(d)}=Y^{(1,d)}=Y^{(0,d)}$. Overlap ensures that there are no covariate values for which the instrument assignment is deterministic. Monotonicity rules out the possibility of defiers: individuals who will always pursue an opposite treatment status from their instrument assignment.
\cite{angrist1996identification} prove identification of the local average treatment effect (LATE) using Assumption~\ref{assumption:id}. \cite{abadie2003semiparametric} extends identification for a broad class of complier parameters.
\begin{definition}[General class of complier parameters \citep{abadie2003semiparametric}]\label{def:class}
Let $g(y,d,x,\theta)$ be a measurable, real valued function such that $E\{g(Y,D,X ,\theta)^2\}<\infty$ for all $\theta$ in $\Theta$. Consider complier parameters $\theta_0$ implicitly defined by any of the following expressions:
\begin{enumerate}
\item $E\{ g(Y^{(0)},X, \theta)\mid D^{(1)}>D^{(0)}\} =0$ if and only if $\theta=\theta_0$;
\item $E\{ g(Y^{(1)},X ,\theta)\mid D^{(1)}>D^{(0)}\} =0$ if and only if $\theta=\theta_0$;
\item $E\{g(Y,D,X,\theta)\mid D^{(1)}>D^{(0)}\}=0$ if and only if $\theta=\theta_0$.
\end{enumerate}
We subsequently refer to these expressions as the three possible cases for complier parameters.
\end{definition}
For a given instrumental variable $Z$, one may define the average complier characteristics as a special case of Definition~\ref{def:class}. This causal parameter summarizes the observable characteristics of the subpopulation of compliers who are induced to take up or refuse treatment $D$ based on the instrument assignment $Z$. It is an important parameter to estimate because it aids the interpretation of LATE. As we will see in Section~\ref{sec:test}, this causal parameter can help to reconcile different LATE estimates obtained with different instruments.
\begin{definition}[Average complier characteristics]\label{def:characteristics}
Average complier characteristics are $\theta_0=E\{f(X)\mid D^{(1)}>D^{(0)}\}$ for any measurable function $f$ of covariate $X$ that may have a finite dimensional, real vector value such that $E\{f_j(X)^2\}<\infty$.
\end{definition}
\section{Key insight}\label{sec:insight}
\subsection{Classic approaches: Wald formula and $\kappa$ weight}
We provide intuition for our key insight that a doubly robust moment for a complier parameter has two components: the Wald formula and the $\kappa$ weight. For clarity, we focus on the familiar example of local average treatment effect (LATE) in this initial discussion: $\theta_0=E\{Y^{(1)}-Y^{(0)}\mid D^{(1)}>D^{(0)}\}$. In subsequent sections, we study the entire class of complier parameters in Definition~\ref{def:class}, including the new case of average complier characteristics.
Under Assumption~\ref{assumption:id}, LATE can be identified as
$$
\theta_0
=\frac{E\left\{E(Y\mid Z=1,X)-E(Y\mid Z=0,X)\right\}}{E\left\{E(D\mid Z=1,X)-E(D\mid Z=0,X)\right\}}
$$
following \citet[Theorem 1]{frolich2007nonparametric}. We call this expression the expanded Wald formula.
The direct Wald approach involves estimating the reduced form regression $E(Y\mid Z,X)$ and first stage regression $E(D\mid Z,X)$, then plugging these estimates into the expanded Wald formula. Such an approach is called the plug-in, and it is valid only when both regressions are estimated with correctly specified and unregularized models. It is not a valid approach when either regression is incorrectly specified, leading to the name ``forbidden regression'' \citep{angrist2008mostly}. It is also invalid when the covariates are high dimensional and a regularized machine learning estimator is used to estimate either regression. The matching procedure of \cite{frolich2007nonparametric} faces similar limitations.
In seminal work, \cite{abadie2003semiparametric} proposes an alternative formulation in terms of the $\kappa$ weights
$$
\kappa^{(0)}(W)= (1-D)\frac{(1-Z)-\{1-\pi_0(X)\}}{\{1-\pi_0(X)\}\pi_0(X)},\quad
\kappa^{(1)}(W)= D\frac{Z - \pi_0(X)}{\{1-\pi_0(X)\}\pi_0(X)}
$$
where $\pi_0(X)=\text{\normalfont pr}(Z=1\mid X)$ is the instrument propensity score.
The $\kappa$ weights have the property that
$$
\theta_0=\omega^{-1}E\{\kappa^{(1)}(W) Y-\kappa^{(0)}(W) Y\},\quad \omega= E\left\{1-\frac{D(1-Z)}{1-\pi_0(X)}-\frac{(1-D)Z}{\pi_0(X)}\right\}.
$$
In words, the mean of the product of $Y$ and $\kappa^{(d)}(W)$ gives, up to a scaling, the expected potential outcome $Y^{(d)}$ of compliers when treatment is $D=d$. As an aside, \cite{abadie2003semiparametric} also introduces a third weight $\kappa(W)$ for parameters that belong to the third case in Definition~\ref{def:class}.
The $\kappa$ weight approach would involve estimating the propensity score $\hat{\pi}$ and plugging this estimate into the $\kappa$ weight formula. Intuitively, the $\kappa$ weight approach is like a multistage inverse propensity weighting. Impressively, it remains agnostic about the functional form of the reduced form regression $E(Y\mid Z,X)$ and first stage regression $E(D\mid Z,X)$. It is valid only when $\hat{\pi}$ is estimated with a correctly specified and unregularized model. It is invalid if $\hat{\pi}$ is incorrectly specified or if covariates are high dimensional and a regularized machine learning estimator is used to estimate $\hat{\pi}$. Moreover, the inversion of $\hat{\pi}$ can lead to numerical instability in high dimensional settings.
\subsection{Doubly robust moment for a special case}
Next, we introduce the moment function and doubly robust moment function formulations of LATE. For the special case of LATE, these formulations were first derived by \cite{tan2006regression} with the goal of addressing misspecification of the regressions and the propensity score. Consider the expanded Wald formula. Rearranging and using the notation $V=(Y,D)^{\top}$ as a column vector,
$
\gamma_0(Z,X)=E(V\mid Z,X)$ as a vector valued regression, and $\begin{pmatrix}1, & -\theta \end{pmatrix}$ as a row vector,
we arrive at the moment function formulation of LATE:
$$
E\left[
\begin{pmatrix}1, & -\theta \end{pmatrix}
\{\gamma_0(1,X)-\gamma_0(0,X)\}
\right]
=0\text{ if and only if }\theta=\theta_0.
$$
Denote the the Horvitz-Thompson balancing weight as
$$
\alpha_0(Z,X)=\frac{Z}{\pi_0(X)}-\frac{1-Z}{1-\pi_0(X)},\quad \pi_0(X)=\text{\normalfont pr}(Z=1\mid X).
$$
\cite{tan2006regression} shows that for LATE, the doubly robust moment function is
$$
E\left[
\begin{pmatrix}1, & -\theta \end{pmatrix}
\{\gamma_0(1,X)-\gamma_0(0,X)\}
+\alpha_0(Z,X)
\begin{pmatrix}1, & -\theta \end{pmatrix}
\{V-\gamma_0(Z,X)\}\right]
=0\text{ if and only if }\theta=\theta_0.
$$
The doubly robust formulation remains valid if either the vector valued regression $\gamma_0$ or propensity score $\pi_0$ is incorrectly specified.
\subsection{A new synthesis that allows for machine learning}
Our key observation is the connection between the $\kappa$ weight and the balancing weight $\alpha_0$. This simple observation will allow us to characterize the doubly robust moment function for a broad class of complier parameters, generalizing \cite{tan2006regression} to the full class defined by \cite{abadie2003semiparametric}.
\begin{proposition}[$\kappa$ weight as balancing weight]\label{rewrite_kappa}
The $\kappa$ weights can be rewritten as
\begin{align*}
&\kappa^{(0)}(W)=\alpha_0(Z,X)(D-1),\quad
\kappa^{(1)}(W)=\alpha_0(Z,X) D,\quad
\kappa(W)=1-\frac{D(1-Z)}{1-\pi_{0}(X)}-\frac{(1-D)Z}{\pi_{0}(X)}.
\end{align*}
\end{proposition}
\begin{proof}
Observe that
$$\alpha_0(z,x)=\frac{z}{\pi_0(x)}-\frac{1-z}{1-\pi_0(x)}=\frac{z-\pi_0(x)}{\pi_0(x)\{1-\pi_0(x)\}}$$
which proves the expression for $\kappa^{(0)}$ and $\kappa^{(1)}$. Using these expressions, we have
$$\kappa(w)= \{1-\pi_0(x)\}\alpha_0(z,x)(d-1)+\pi_0(x)\alpha_0(z,x) d=1-\frac{d(1-z)}{1-\pi_{0}(x)}-\frac{(1-d)z}{\pi_{0}(x)}.$$\end{proof}
Next, we formalize the sense in which the balancing weight $\alpha_0$ represents the functional $\gamma\mapsto E\left\{\begin{pmatrix}1, & -\theta \end{pmatrix}\gamma(1,X)-\gamma(0,X)\right\}$ that appears in the moment formulation of LATE and the extended Wald formula.
\begin{proposition}[Balancing weight as Riesz representer]\label{prop:rr}
$\alpha_0(z,x)$ is the Riesz representer to the continuous linear functional $\gamma\mapsto E\{\gamma(1,X)-\gamma(0,X)\}$, i.e. for all $\gamma$ such that $E\{\gamma(Z,X)^2\}<\infty$,
$$
E\{\gamma(1,X)-\gamma(0,X)\}=E\{\alpha_0(Z,X)\gamma(Z,X)\}.
$$
Similarly, $Z/\pi_0(X)$ is the Riesz representer to the continuous linear functional $\gamma\mapsto E\{\gamma(1,X)\}$, and $(1-Z)/\{1-\pi_0(X)\}$ is the Riesz representer to the continuous linear functional $\gamma\mapsto E\{\gamma(0,X)\}$.
\end{proposition}
\begin{proof}
This result is well known in semiparametrics. We provide the proof for completeness. Observe that
\begin{align*}
E\left\{\gamma(Z,X) \frac{Z}{\pi_0(X)}\mid X\right\}
&= E\left\{\gamma(Z,X) \frac{1}{\pi_0(X)}\mid Z=1,X\right\}\text{\normalfont pr}(Z=1\mid X) \\
&= E\left\{\gamma(Z,X) \frac{1}{\pi_0(X)}\mid Z=1,X\right\}\pi_0(X)
=\gamma(1,X)
\end{align*}
and likewise
$$
E\left\{\gamma(Z,X) \frac{1-Z}{1-\pi_0(X)}\mid X\right\}=\gamma(0,X).
$$
Combining these two terms, we have by the law of iterated expectations
\begin{align*}
&E\{\gamma(1,X)-\gamma(0,X)\}
= \int \{\gamma(1,x)-\gamma(0,x)\} \mathrm{d}\text{\normalfont pr}(x) \\
&=\int \left[E\left\{\gamma(Z,X) \frac{Z}{\pi_0(X)}\mid X=x\right\}-E\left\{\gamma(Z,X) \frac{1-Z}{1-\pi_0(X)}\mid X=x\right\}\right] \mathrm{d}\text{\normalfont pr}(x) \\
&=E\left\{\gamma(Z,X)\frac{Z}{\pi_0(X)}\right\}-E\left\{\gamma(Z,X)\frac{1-Z}{1-\pi_0(X)}\right\}.
\end{align*}
\end{proof}
An immediate consequence of Proposition~\ref{prop:rr} is that
$$
E\left\{\begin{pmatrix}1, & -\theta \end{pmatrix}\gamma(1,X)-\gamma(0,X)\right\}=E\left\{\alpha_0(Z,X)\begin{pmatrix}1, & -\theta \end{pmatrix}\gamma(Z,X)\right\} \text{ for any } \gamma.
$$
In summary, Proposition~\ref{rewrite_kappa} shows that the $\kappa$ weight is a reparametrization of the balancing weight $\alpha_0$. Meanwhile, Proposition~\ref{prop:rr} shows that the balancing weight appears in the Riesz representer to the moment formulation of LATE, i.e. the expanded Wald formula. We conclude that the $\kappa$ weight is essentially the Riesz representer to the Wald formula. In seminal work, \cite{newey1994asymptotic} demonstrates that a doubly robust moment is constructed from a moment formulation and its Riesz representer. Therefore the doubly robust moment for complier parameters must combine the Wald formula and the $\kappa$ weight.
With the general doubly robust moment function, one can propose flexible, semiparametric tests for complier parameters. In particular, the semiparametric tests may involve regularized machine learning for flexible estimation and model selection of (i) the regression $\hat{\gamma}$ in a way that approximates nonlinearity and heterogeneity, and (ii) the balancing weight $\hat{\alpha}$ in a way that guarantees balance. In Section~\ref{sec:test}, we instantiate such a test to compare observable characteristics of compliers.
As explained in Appendix~\ref{sec:estimation}, we avoid the numerically unstable step of estimating and inverting $\hat{\pi}$ that appears in \cite{tan2006regression,belloni2017program,chernozhukov2018original}. We replace it with the numerically stable step of estimating $\hat{\alpha}$ directly, extending techniques of \cite{chernozhukov2018learning} to the instrumental variable setting. We call this extension automatic $\kappa$ weighting (Auto-$\kappa$), and demonstrate how it applies to the new and economically important case of average complier characteristics.
In summary, our main theoretical result allows us to combine the classic Wald and $\kappa$ weight formulations for the entire class of complier parameters in Definition~\ref{def:class}, including average complier characteristics, while also updating them to incorporate machine learning.
\section{The doubly robust moment}\label{sec:double}
We now state our main theoretical result, which is the doubly robust moment for the class of complier parameters in Definition~\ref{def:class}. This result formalizes the intuition of Section~\ref{sec:insight}, and it justifies the hypothesis test in Section~\ref{sec:test}. It is convenient to divide the main result into two statements for clarity. Theorem~\ref{thm:general_kappa} handles the first and second cases in Definition~\ref{def:class}, while Theorem~\ref{thm:general_kappa3} handles the third case in Definition~\ref{def:class}.
\begin{theorem}[Cases 1 and 2]\label{thm:general_kappa}
Suppose Assumption~\ref{assumption:id} holds. Let $g(y,d,x,\theta)$ be a measurable, real valued function such that $E\{g(Y,D,X ,\theta)^2\}<\infty$ for all $\theta$ in $\Theta$.
\begin{enumerate}
\item If $\theta_0$ is defined by $E[ g\{Y^{(0)},X ,\theta_0\}\mid D^{(1)}>D^{(0)}] =0$, let
$v(w,\theta)=(d-1) g(y,x,\theta).$
\item If $\theta_0$ is defined by $E[ g\{Y^{(1)},X ,\theta_0\}\mid D^{(1)}>D^{(0)}] =0$, let
$v(w,\theta)=d g(y,x,\theta).$
\end{enumerate}
Then the doubly robust moment function $\psi$ for $\theta_0$ is of the form
\begin{align*}
&\psi(w,\gamma,\alpha,\theta)=m(w,\gamma,\theta)+\phi(w,\gamma,\alpha,\theta),\quad
m(w,\gamma,\theta)=\gamma(1,x,\theta)-\gamma(0,x,\theta),\\
&\phi(w,\gamma,\alpha,\theta)=\alpha(z,x)\{v(w,\theta)-\gamma(z,x, \theta)\}
\end{align*}
where $\gamma_0(z,x,\theta)=E\{v(W,\theta)\mid z,x\}$ is a vector valued regression and $\alpha_0(z,x)=z/\pi_0(x)-(1-z)/\{1-\pi_0(x)\}$ is the Riesz representer of the functional $\gamma\mapsto E\{\gamma(1,X,\theta)-\gamma(0,X,\theta)\}$.
\end{theorem}
\begin{proof}
Consider the first case. Under Assumption~\ref{assumption:id}, we can appeal to \citet[Theorem 3.1]{abadie2003semiparametric}:
$$
0=E[ g\{Y^{(0)},X ,\theta_0\}\mid D^{(1)}>D^{(0)}]=\dfrac{E\{\kappa ^{(0)}(W)g(Y,X ,\theta_0)\}}{\text{\normalfont pr}\{D^{(1)}>D^{(0)}\}}.
$$
Hence
\begin{align*}
0&=E\{\kappa ^{(0)}(W)g(Y,X ,\theta_0)\}
=E\{\alpha_0(Z,X)(D-1)g(Y,X ,\theta_0)\}
=E\{\alpha_0(Z,X)v(W,\theta_0)\} \\
&=E\{\alpha_0(Z,X)\gamma_0(Z,X,\theta_0)\}
=E\{\gamma_0(1,X,\theta_0)-\gamma_0(0,X,\theta_0)\}
\end{align*}
appealing to the previous statement, Proposition~\ref{rewrite_kappa}, the definition of $v(W,\theta_0)$, the law of iterated expectations, and Proposition~\ref{prop:rr}. Likewise for the second case.
\end{proof}
In the doubly robust moment function $\psi(w,\gamma,\alpha,\theta)=m(w,\gamma,\theta)+\phi(w,\gamma,\alpha,\theta)$, we generalize our insight from Section~\ref{sec:insight}. The first term $m(w,\gamma,\theta)$ is essentially a generalized Wald formula. The second term $\phi(w,\gamma,\alpha,\theta)$ is essentially a product between the $\kappa$ weight and a generalized regression residual. In the language of semiparametrics, we \textit{augment} the $\kappa$ weight with the Wald formula. Equivalently, we \textit{debias} the Wald formula with the $\kappa$ weight.
The doubly robust moment function $\psi$ remains valid if either $\gamma_0$ or $\alpha_0$ is misspecified, i.e.
$$
0=E\{\psi(W,\gamma,\alpha_0,\theta_0)=E[\psi(W,\gamma_0,\alpha,\theta_0)\} \text{ for any } \gamma,\alpha.
$$
In the former expression, $\gamma_0$ may be misspecified yet $\psi$ remains valid as an estimating equation. In the latter, $\alpha_0$ may be misspecified yet $\psi$ remains valid as an estimating equation. Theorem~\ref{thm:general_kappa} demonstrates that all complier parameters in cases 1 and 2 of Definition~\ref{def:class} have a doubly robust moment function $\psi$ with a common structure. As such, we are able to analyze all of these causal parameters with the same argument. Case 3 of Definition~\ref{def:class} is more involved, but we show that it shares the common structure as well.
\begin{theorem}[Case 3]\label{thm:general_kappa3}
Suppose Assumption~\ref{assumption:id} holds. Let $g(y,d,x,\theta)$ be a measurable, real valued function such that $E\{g(Y,D,X ,\theta)^2\}<\infty$ for all $\theta$ in $\Theta$.
If $\theta_{0}$ is defined by the moment condition $E\{g(Y,D,X,\theta_{0})\mid D^{(1)}>D^{(0)}\}=0$,
then the doubly robust moment function for $\theta_{0}$ is
of the form
\begin{align*}
&\psi(w,\tilde{\gamma},\tilde{\alpha},\theta) =m(w,\tilde{\gamma},\theta)+\phi(w,\tilde{\gamma},\tilde{\alpha},\theta),\quad
m(w,\tilde{\gamma},\theta) =\gamma(z,x,\theta)-\gamma^{0}(1,x,\theta)-\gamma^{1}(0,x,\theta)\\
&\phi(w,\tilde{\gamma},\tilde{\alpha},\theta) =\{g(y,d,x,\theta)-\gamma(z,x,\theta)\} -\alpha^0(z,x)\{(1-d) g(y,d,x,\theta)-\gamma^{0}(z,x,\theta)\}\\
&\quad\quad\quad\quad\quad\quad -\alpha^1(z,x)\{d g(y,d,x,\theta)-\gamma^{1}(z,x,\theta)\}
\end{align*}
where $\tilde{\gamma}$ concatenates $(\gamma,\gamma^0,\gamma^1)$ and $\tilde{\alpha}$ concatenates $(\alpha^0,\alpha^1)$. These functions are defined by
\begin{align*}
&\gamma_{0}(z,x,\theta)=E\{g(Y,D,X,\theta)\mid z,x\},\quad \gamma_{0}^{0}(z,x,\theta)=E\{(1-D) g(Y,D,X,\theta)\mid z,x\},\\ &\gamma_{0}^{1}(z,x,\theta)=E\{D g(Y,D,X,\theta)\mid z,x\},\quad \alpha_0^0(z,x)=z/\pi_0(x),\quad \alpha_0^1(z,x)=(1-z)/\{1-\pi_0(x)\}.
\end{align*}
\end{theorem}
\begin{proof}
A similar argument extends to the third case. Under Assumption~\ref{assumption:id},
we can appeal to \citet[Theorem 3.1]{abadie2003semiparametric}:
$$
0=E\{g(Y,D,X,\theta_{0})\mid D^{(1)}>D^{(0)}\}=\dfrac{E\{\kappa(W)g(Y,D,X,\theta_{0})\}}{\text{\normalfont pr}\{D^{(1)}>D^{(0)}\}}.
$$
Hence
\begin{align*}
0 & =E\{\kappa(W)g(Y,D,X,\theta_{0})\}\\
& =
E\left\{g(Y,D,X,\theta_{0})
-\frac{Z}{\pi_{0}(X)}(1-D) g(Y,D,X,\theta_{0})
-\frac{1-Z}{1-\pi_{0}(X)}D g(Y,D,X,\theta_{0})\right\}\\
& =E\left\{\gamma_{0}(Z,X,\theta_0)-\frac{Z}{\pi_{0}(X)}\gamma_{0}^{0}(Z,X,\theta_{0})-\frac{1-Z}{1-\pi_{0}(X)}\gamma_{0}^{1}(Z,X,\theta_{0})\right\}\\
& =E\{\gamma_{0}(Z,X,\theta_0)-\gamma_{0}^{0}(1,X,\theta_{0})-\gamma_{0}^{1}(0,X,\theta_{0})\}
\end{align*}
appealing to the previous statement, Proposition~\ref{rewrite_kappa}, the definitions of $(\gamma_0,\gamma_0^0,\gamma_0^1)$ together with the law of iterated expectations, and Proposition~\ref{prop:rr}.
\end{proof}
This time, the doubly robust moment function $\psi$ remains valid if either $\tilde{\gamma}_0$ or $\tilde{\alpha}_0$ is misspecified, i.e.
$$
0=E\{\psi(W,\tilde{\gamma},\tilde{\alpha}_0,\theta_0)=E[\psi(W,\tilde{\gamma}_0,\tilde{\alpha},\theta_0)\} \text{ for any } \tilde{\gamma},\tilde{\alpha}.
$$
In the former expression, $\tilde{\gamma}_0$ may be misspecified yet $\psi$ remains valid as an estimating equation. In the latter, $\tilde{\alpha}_0$ may be misspecified yet $\psi$ remains valid as an estimating equation.
In Section~\ref{sec:test}, we translate this general characterization of the doubly robust moment into a practical hypothesis test to evaluate the external validity of instruments. In Appendix~\ref{sec:estimation}, we translate this general characterization into general machine learning estimators for complier parameters, which we use to implement the hypothesis test. In particular, we consider direct estimation of the balancing weight, a procedure that we call automatic $\kappa$ weighting (Auto-$\kappa$).
\section{A hypothesis test to compare observable characteristics}\label{sec:test}
\subsection{Corollaries for average complier characteristics}
As a corollary, we characterize the doubly robust moment for average complier characteristics, which appears to have been previously unknown. Using the new doubly robust moment, we propose a hypothesis test, free of functional form restrictions, to evaluate (i) whether two different instruments induce subpopulations of compliers with the same observable characteristics on average, and (ii) whether compliers have observable characteristics that are the same as the full population on average.
\begin{corollary}[Average complier characteristics]\label{cor:characteristics}
The doubly robust moment for average complier characteristics is
$$
\psi(w,\gamma,\alpha,\theta)=A(\theta)\{\gamma(1,x)-\gamma(0,x)\}+\alpha(z,x)A(\theta)\{v-\gamma(z,x)\},\quad A(\theta)=\begin{pmatrix}I, & -\theta \end{pmatrix}
$$
where
$
v=\{df(x)^{\top}, d\}^{\top}$, $\gamma_0(z,x)=E(V\mid z,x)$, and $\alpha_0(z,x)=z/\pi_0(x)-(1-z)/\{1-\pi_0(x)\}$.
\end{corollary}
\begin{proof}
The result is a special case of Corollary~\ref{cor:LATE} in Appendix~\ref{sec:estimation}.
\end{proof}
Suppose we wish to test the null hypothesis that two different instruments $Z_1$ and $Z_2$ induce complier subpopulations with the same observable characteristics on average. Denote by $\hat{\theta}_1$ and $\hat{\theta}_2$ the estimators for average complier characteristics using the different instruments $Z_1$ and $Z_2$, respectively. One may construct machine learning estimators $\hat{\theta}_1$ and $\hat{\theta}_2$ based on the doubly robust moment function in Corollary~\ref{cor:characteristics}. In Appendix~\ref{sec:estimation}, we instantiate automatic $\kappa$ weight (Auto-$\kappa$) estimators of this type. The following procedure allows us to test the null hypothesis from some estimator $\hat{C}$ for the asymptotic variance $C$ of $\hat{\theta}=(\hat{\theta}^{\top}_1,\hat{\theta}^{\top}_2)^{\top}$. In Appendix~\ref{sec:estimation}, we provide an explicit variance estimator $\hat{C}$ based on Auto-$\kappa$ as well.
\begin{algorithm}[Hypothesis test for difference of average complier characteristics]\label{alg:hyp}
Given $\hat{\theta}$ and $\hat{C}$, which may be based on Auto-$\kappa$ as in Appendix~\ref{sec:estimation},
\begin{enumerate}
\item Calculate the statistic $T=n(\hat{\theta}_1-\hat{\theta}_2)^{\top}(R
\hat{C}
R^{\top})^{-1}(\hat{\theta}_1-\hat{\theta}_2)$ where $ R=\begin{pmatrix}I, & -I\end{pmatrix}$.
\item Compute the value $c_{a}$ as the $(1-a)$ quantile of $\chi^2\{dim(\theta_1)\}$.
\item Reject the null hypothesis if $T>c_{a}$.
\end{enumerate}
\end{algorithm}
Algorithm~\ref{alg:hyp} can also test the null hypothesis that compliers have observable characteristics that are the same as the full population on average. $\hat{\theta}_1$ is as before, $\hat{\theta}_2=n^{-1}\sum_{i=1}^nf(X_i)$, and $\hat{C}$ updates accordingly.
\begin{corollary}[Hypothesis test for difference of average complier characteristics]\label{cor:hyp}
If $\hat{\theta}=\theta_0+o_p(1)$, $n^{1/2}(\hat{\theta}-\theta_0)\rightsquigarrow \mathcal{N}(0,C)$, and $\hat{C}=C+o_p(1)$, then the hypothesis test in Algorithm~\ref{alg:hyp}
falsely rejects the null hypothesis $H_0$ with
probability approaching the nominal level, i.e.
$
\text{\normalfont pr}(T>c_{a}\mid H_0)\rightarrow a.
$
\end{corollary}
\begin{proof}
The result is immediate from \citet[Section 9]{newey1994large}.
\end{proof}
Corollary~\ref{cor:hyp} is our main practical result: justification of a flexible hypothesis test to evaluate a difference in average complier characteristics. It appears that no semiparametric test previously exists for this important question about the external validity of instruments. By developing this hypothesis test, we equip empirical researchers with a new robustness check. This practical result follows as a consequence of our main insight in Section~\ref{sec:insight} and our main theoretical result in Section~\ref{sec:double}. In Appendix~\ref{sec:estimation}, we verify the conditions of Corollary~\ref{cor:hyp} for Auto-$\kappa$ under weak regularity assumptions.
\subsection{Empirical application}
With this practical result, we revisit a classic empirical paper in labor economics to test whether two different instruments induce different average complier characteristics. \cite{angrist1998children} estimate the impact of childbearing $D$ on female labor supply $Y$ in a sample of 394,840 mothers, aged 21--35 with at least two children, from the 1980 Census. The first instrument $Z_1$ is twin births: $Z_1$ indicates whether the mother's second and third children were twins. The second instrument $Z_2$ is same-sex siblings: $Z_2$ indicates whether the mother's initial two children were siblings with the same sex. The authors reason that both $(Z_1,Z_2)$ are quasi random events that induce having a third child.
\begin{table}[h]
\caption{Comparison of average complier characteristics}{
\begin{tabular}{lcccccccc}
\\
& \multicolumn{4}{c}{Average age of second child} & \multicolumn{4}{c}{Average schooling of mother} \\
& Twins & Same-sex & 2 sided & 1 sided & Twins & Same-sex & 2 sided & 1 sided\\[5pt]
$\kappa$ weight & 5.51 & 7.14 & - & - & 12.43 & 12.09 & - & - \\
Auto-$\kappa$ & 4.52 & 6.92 & 0.13 & 0.07 & 9.84 & 12.10 & 0.54 & 0.27 \\
Auto-$\kappa $ (S.E.) & (0.70) & (1.43) & - & - & (2.47) & (2.78) & - & -
\end{tabular}}
\label{tab:1}
\textit{Notes:}
S.E., standard error; Auto-$\kappa$, automatic $\kappa$ weighting. See Supplement~\ref{sec:application} for estimation details.
\end{table}
The two instruments give rise to two LATE estimates for the reduction in weeks worked due to a third child: -3.28 (0.63) for $Z_1$ and -6.36 (1.18) for $Z_2$, where the standard errors are in parentheses. \cite{angrist2013extrapolate} attribute the difference in LATE estimates to a difference in average complier characteristics, i.e. a difference in average covariates for instrument specific complier
subpopulations. The authors use parametric $\kappa$ weights, report point estimates without standard errors, and conclude that
``twins compliers therefore are relatively more
likely to have a young second-born and to be highly educated.''
We replicate, extend, and test these previous findings. In their parametric $\kappa$ weight approach, \cite{angrist2013extrapolate} estimate $\pi_{0}(X)$ using a logistic model with polynomials of continuous covariates. In our semiparametric Auto-$\kappa$ approach, we expand the dictionary to higher order polynomials, include interactions between the instrument and covariates, and directly estimate and regularize the balancing weights. Crucially, our main result allows us to conduct inference, and to test whether the instruments $Z_1$ and $Z_2$ induce differences in the observable complier characteristics suggested by previous work.
Table~\ref{tab:1} summarizes results. In Columns 1, 2, 5, and 6, we find similar point estimates to \cite{angrist2013extrapolate}, given in Row 1. Columns 3, 4, 7, and 8 report $p$ values for tests of the null hypothesis that average complier characteristics are equal for the twins and same-sex instruments. We find weak evidence in favor of the explanation that twins compliers are more
likely to have a young second-born. We do not find evidence that twins compliers have a significantly different education level than same-sex compliers.
\section{Conclusion}\label{sec:conc}
We propose a semiparametric test to evaluate (i) whether two different instruments induce subpopulations of compliers with the same observable characteristics on average, and (ii) whether compliers have observable characteristics that are the same as the full population on average. This hypothesis test is a flexible and practical robustness check for the external validity of instrumental variables. We use the test to reinterpret the difference in LATE estimates that \cite{angrist1998children} obtain when using two different instrumental variables. Specifically, we implement a machine learning update to $\kappa$ weighting that we call the automatic $\kappa$ weight (Auto-$\kappa$). To justify the test, we develop new econometric theory. Most notably, we characterize the doubly robust moment function for the entire class of complier parameters from \cite{abadie2003semiparametric}, answering an open question in the semiparametric literature in order to handle the new and economically important case of average complier characteristics.