EconBase
← Back to paper

Efficient and Robust Estimation of the Generalized LATE Model

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

79,626 characters · 13 sections · 78 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Efficient and Robust Estimation of the Generalized LATE Model

\def\spacingset#1{ {#1}} \spacingset{1}

\if00 \fi

\if10 {

center[center omitted — 35 chars of source]

} \fi

abstractThis paper studies the estimation of causal parameters in the generalized local average treatment effect (GLATE) model, a generalization of the classical LATE model encompassing multi-valued treatment and instrument. We derive the efficient influence function (EIF) and the semiparametric efficiency bound (SPEB) for two types of parameters: local average structural function (LASF) and local average structural function for the treated (LASF-T). The moment condition generated by the EIF satisfies two robustness properties: double robustness and Neyman orthogonality. Based on the robust moment condition, we propose the double/debiased machine learning (DML) estimators for LASF and LASF-T. The DML estimator is semiparametric efficient and suitable for high dimensional settings. We also propose null-restricted inference methods that are robust against weak identification issues. As an empirical application, we study the effects across different sources of health insurance by applying the developed methods to the Oregon Health Insurance Experiment.

{\it Keywords:} Causal Inference, Double Robustness, Efficient Influence Function, Multi-valued Treatment, Neyman Orthogonality, Oregon Health Insurance Experiment, Unordered Monotonicity, Weak Identification.

\spacingset{1.45}

Introduction

Since the seminal works of imbens1994identification and angrist1996identification, the local average treatment effect (LATE) model has become popular for causal inference in economics. Instead of imposing homogeneity of the treatment effects as in the classical instrumental variable (IV) regression model, the LATE framework allows the treatment effect to vary across individuals. Under the monotonicity condition, the average treatment effect can be identified for a subgroup of individuals whose treatment choice complies with the change in instrument levels.

The current form of the LATE model only accepts binary treatment variables. This restriction is inconvenient in many economic settings where the treatment is multi-leveled in nature. For example, parents select different preschool programs for their kids, schools assign students to different classroom sizes, families relocate to various neighborhoods in housing experiments, and people choose different sources of health insurance. To apply the LATE model to these settings, researchers often need to redefine the treatment so that there are only two treatment levels. However, merging the treatment levels can complicate the task of program evaluation and dampen the causal interpretation of the estimates. As pointed out by kline2016evaluating, if the original treatment levels are substitutes, then there is ambiguity regarding which causal parameters are of interest. After merging the treatment levels, the heterogeneity in the treatment effect across different treatment levels is lost.

This paper addresses the above issues by generalizing the LATE framework to incorporate the potential multiplicity in treatment levels directly. We call the new framework the generalized LATE (GLATE) model. The main assumption of the GLATE model is the unordered monotonicity assumption proposed by heckman2018unordered, which is a generalization of the monotonicity assumption in the binary LATE model.\footnote{To distinguish with the GLATE model, we sometimes use the terminology “binary LATE model" to refer to the LATE model studied by imbens1994identification and abadie2003semiparametric.}

We generalize the identification results in heckman2018unordered to explicitly account for the presence of conditioning covariates, which is often important in practical settings. Recently, blandhol2022tsls point out that linear TSLS, the common way to control for covariates in empirical studies, does not bear the LATE interpretation. The only specifications that have LATE interpretations are the ones that control for covariates nonparametrically. Therefore, it is essential from the causal analysis perspective to incorporate the covariates into the GLATE framework in a nonparametric way.

The causal parameters identifiable in the GLATE model include local average structural function (LASF) and local average structural function for the treated (LASF-T). LASF is the mean potential outcome for specific subpopulations. These subpopulations are defined by their treatment choice behaviors and are generalizations of the concepts always takers, compliers, and never takers in the binary LATE model. The parameter LASF-T further restricts the subpopulation to exclude individuals who do not take up the treatment.

The paper is concerned with the econometric aspects of the GLATE model. The analysis begins by deriving efficient influence function (EIF) and semiparametric efficiency bound (SPEB) for the identified parameters. The calculation is based on the method outlined in Chapter 3 of bickel1993efficient and newey1990semiparametric. We then verify that the conditional expectation projection (CEP) estimator chen2008semiparametric, constructed directly from the identification result, achieves the SPEB and hence is semiparametric efficient. Using these results, we may efficiently estimate other important parameters of interest by the plug-in method since a standard delta-method argument preserves semiparametric efficiency.

The EIF not only facilitates the efficiency calculation but can also serve as the moment condition for estimation. This is because the EIF is mean zero by construction and is equal to the original identification result plus an adjustment term due to the presence of infinite-dimensional parameters. We show that the moment condition constructed from the EIF satisfies two related robustness properties: double robustness and Neyman orthogonality. Double robustness guarantees that the moment condition is correctly specified in a parametric setting even when some nuisance parameters are not.

The Neyman orthogonality condition means that the moment condition is insensitive to the nuisance parameters. This condition is particularly useful when the conditioning covariates are of high dimension. To further utilize this condition, we study the double/debiased machine learning (DML) estimator chernozhukov2018double in the GLATE setting. Under certain conditions regarding the convergence rate of the first-step nonparametric estimators, the DML estimator is asymptotically normally uniformly over a large class of data generating processes (DGPs).

The weak identification issue is a practical concern of the GLATE model. This is because both the treatment and instrument are multi-valued, and hence the subpopulation on which LASF and LASF-T are defined can be small in size. To deal with this issue, we propose null-restricted test statistics in one-sided and two-sided testing problems. This procedure is the generalization of the well-known Anderson-Rubin (AR) test. We show that the proposed tests are consistent and uniformly control size across a large class of DGPs, in which the size of the subpopulation mentioned above can be arbitrarily close to zero.

The paper is organized as follows. The remaining part of this section discusses the literature. Section (ref) introduces the GLATE model and the nonparametric identification results. Section (ref) calculates the EIF and SPEB. Section (ref) discusses the robustness properties of the moment condition generated by the EIF. Section (ref) proposes inference procedures under weak identification issues. Section (ref) presents the empirical application. Section (ref) concludes. The proofs for theoretical results in the main text are collected in Appendix (ref).

Literature Review

The GLATE model provides a way to conduct causal inference under endogeneity when the treatment is multi-valued and unordered. As mentioned above, the identification result (conditional on the covariates) is first established in heckman2018unordered by using the unordered monotonicity condition. lee2018identifying proposes another method of identification in a similar model of multi-valued treatment. Their method is concerned with continuous instruments, while the GLATE is framed in terms of discrete-valued instruments. When the treatment levels are ordered, angrist1995two derives the identification and estimation results for the causal parameter, which is a weighted average of LATEs across different treatment levels.

The literature on semiparametric efficiency in program evaluation starts with the seminal work of hahn1998role, which studies the benchmark case of estimating the average treatment effect (ATE) under unconfoundedness. For multi-level treatment, cattaneo2010efficient studies the efficient estimation of causal parameters implicitly defined through over-identified non-smooth moment conditions. In the case where unconfoundedness fails and instruments are present, frolich2007nonparametric calculates the SPEB for the LATE parameter, and hong2010semiparametric extend to the estimation of parameters implicitly defined by moment restrictions. In a more general framework encompassing missing data, chen2004semiparametric and chen2008semiparametric studies semiparametric efficiency bounds and efficient estimation of parameters defined through overidentifying moment restrictions. However, there is currently no theoretical research on semiparametric efficient estimation in models that encompasses endogeneity and unordered multiple treatment levels.

Several ways are available for calculating the EIF for semiparametric estimators, as illustrated by newey1990semiparametric and ichimura2022influence. Semiparametric efficiency calculations can be used to construct robust (Neyman orthogonal) moment conditions. This method is illustrated in newey1994asymptotic and chernozhukov2016locally. Based on the Neyman orthogonality condition, chernozhukov2018double introduces the DML method that suits high dimensional settings. This is because Donsker properties and stochastic equicontinuity conditions are no longer required in deriving the asymptotic distribution of the semiparametric estimator.

For testing the GLATE model, sun2020instrument proposes a bootstrap test which is the generalization and improvement of the test studied by kitagawa2015test in the binary LATE model.

The GLATE model has received attention in the recent empirical literature due to its ability to model multi-valued treatment. kline2016evaluating evaluate the cost-effectiveness of Head Start, classifying Head Start and other preschool programs as different treatment levels against the control group of no preschool. galindo2020empirical assesses the impact of different childcare choice in Colombia on children’s development. pinto2014noncompliance studies the neighborhood effects and voucher effects in housing allocations using data from the Moving to Opportunity experiment. Our theoretical analysis of the GLATE model presents important tools for estimation and inference that can be applied to those empirical settings.

Identification in the GLATE Model

This section describes the generalized local average treatment effect (GLATE) model, discusses identification of the local average structural function (LASF) and other parameters, and introduces the notation.

The model

We assume a finite collection of instrument values $ \mathcal{Z} = \{z_1, \cdots, z_{N_Z}\}$ and a finite collection of treatment values $ \mathcal{T} = \{t_1, \cdots, t_{N_T}\}$, where $N_Z$ and $N_T$ are respectively the total number of instrument and treatment levels. The sets $\mathcal{T}$ and $\mathcal{Z}$ are categorical and unordered. The instrumental variable $Z$ denotes which of the $N_Z$ instrument levels is realized. The random variables $T_{z_1}, \cdots, T_{z_{N_Z}}$, each taking values in $\mathcal{T}$, denote the collection of potential treatments under each instrument status. Thus, the observed treatment level is the random variable $T = T_Z = \sum_{z \in \mathcal{Z}} \mathbf{1}\{Z=z\} T_z$. For each given treatment level $t \in \mathcal{T}$, there is a potential outcome $Y_t \in \mathcal{Y} \subset \mathbb{R}$. The observed outcome is denoted by $Y = Y_T = \sum_{t \in \mathcal{T}} \mathbf{1}\{T=t\} Y_t$. The random vector $X \in \mathcal{X} \subset \mathbb{R}^{d_X}$ contains the set of covariates. The observed data is a random sample $(Y_i,T_i,Z_i,X_i), 1 \leq i \leq n$.

The description above establishes a random sampling model where the researcher only observes one potential outcome, the one associated with the observed treatment. This implies that the sample of $Y$, observed from an individual with treatment $T=t$, comes from the conditional distribution of $Y_t$ given $T=t$ rather than from the marginal distribution of $Y_t$. In general, this fact leads to identifications issues and presents challenges for causal inference. To overcome these problems, we impose further structures on the model.

assumption[Conditional Independence] $( \{Y_{t}: t \in \mathcal{T} \}, \{T_{z}: z \in \mathcal{Z} \} ) \perp Z \mid X$.
assumption[Unordered Monotonicity] For any $t \in \mathcal{T}, z,z' \in \mathcal{Z}$, either \begin{align*} \mathbb{P}(\mathbf{1}\{T_{z} = t\} \geq \mathbf{1}\{T_{z'} = t\} \mid X) = 1 \end{align*} or \begin{align*} \mathbb{P}(\mathbf{1}\{T_{z} = t\} \leq \mathbf{1}\{T_{z'} = t\} \mid X) = 1. \end{align*}

Assumption (ref) and (ref) provide the multi-valued analog of Assumption 2.1 in abadie2003semiparametric. Assumption (ref) restricts that the instrument $Z$ is independent with the potential treatments and outcomes once we condition on $X$. Assumption (ref) is the conditional version of the unordered monotonicity condition proposed by heckman2018unordered. It means that when we focus on a particular treatment level $t$ and a pair $(z,z')$ of instrument values, the binary environment should satisfy the usual monotonicity constraint in the LATE model. Specifically, the unordered monotonicity condition requires that a shift in the instrument moves all agents uniformly toward or against each possible treatment value.\footnote{As pointed out by vytlacil2002independence, the LATE monotonicity condition is a restriction across individuals on the relationship between different hypothetical treatment choices defined in terms of an instrument.}

We define the type $S$ of an individual as the vector of the potential treatments, that is,

align*[align* omitted — 50 chars of source]

By construction, $S$ is not observed. Assumption (ref), the unordered monotonicity condition, is essentially a restriction on $\mathcal{S} \equiv \textit{supp}(S)$, the support of $S$. Denote the elements in $\mathcal{S}$ by $s_1,\cdots,s_{N_S}$, where $N_S$ is the cardinality of $\mathcal{S}$. A convenient way to characterize $\mathcal{S}$ is by using the $N_Z \times N_S$ matrix $R \equiv (s_1,\cdots,s_{N_S})$. The matrix $R$ is referred to as the response matrix since it describes how each type of individuals' treatment choice responds to the instrument.

The role of $S$ is to assist the identification of the counterfactual outcomes by dividing the population into a finite number of groups, where identification can be achieved within specific groups. Those groups are defined as follows. For $k = 0,\cdots,N_Z$, let $\Sigma_{t,k}$ be the set of types in which the treatment level $t$ appears exactly $k$ times. That is,

align*[align* omitted — 119 chars of source]

where $s[i]$ denotes the $i$th element of the vector $s$. In particular, the collection $\Sigma_{t,k}, k = 0,\cdots,N_Z$ forms a partition of $\mathcal{S}$.

For individuals with type $S$ in the same type set $\Sigma_{t,k}$, their treatment response in terms of $T=t$ is in a way homogeneous. Thus, it is easier intuitively to identify the marginal distribution of the potential outcome $Y_t$ within each $\Sigma_{t,k}$. More specifically, we define the local average structural functions (LASF) and the local average structural functions for the treated (LASF-T) as follows.

align*[align* omitted — 176 chars of source]

Before presenting the identification results for the above two classes of parameters, we illustrate the GLATE model in the following two examples.

example[Binary LATE model] In the binary LATE model of imbens1994identification, there are two treatment levels $\mathcal{T} = \{0,1\}$ and two instrument levels $\mathcal{Z} = \{0,1\}$. There are three types: $\mathcal{S} = \{ s_1 = (0,0)', s_2 = (0,1)', s_3 = (1,1)' \}$, which are referred to in the literature as defiers, compliers, and always-takers, respectively. The type set $\Sigma_{1,0} = \{s_1\}$ contains the defiers, $\Sigma_{1,1} = \{s_2\}$ the compliers, and $\Sigma_{1,2} = \{s_3\}$ the always-takers. The response matrix is the following binary matrix \begin{align*} R = (s_1,s_2,s_3) = \begin{pmatrix} 0 & 0 & 1 \\ 0 & 1 & 1 \end{pmatrix}. \end{align*} The local average treatment effect is the treatment effect for the compliers, which can be written as the difference between two LASFs: \begin{align*} \mathbb{E}[Y_1 - Y_0 \mid S = compliers] = \mathbb{E}[Y_1 - Y_0 \mid T_1 > T_0] = \beta_{1,1} - \beta_{0,1}. \end{align*}
example[Three treatment levels and two instrument levels] The simplest GLATE model (excluding the binary case in Example (ref)) has three treatment levels $\mathcal{T}=\{t_1,t_2,t_3\}$ and two instrument levels $\mathcal{Z}=\{z_1,z_2\}$. There are five types specified as the columns in the following response matrix \begin{align*} R = (s_1,s_2,s_3,s_4,s_5) = \begin{pmatrix} t_1 & t_2 & t_3 & t_1 & t_2 \\ t_1 & t_2 & t_3 & t_3 & t_3 \end{pmatrix}. \end{align*} In this example, a shift from $z_1$ to $z_2$ moves all agents uniformly toward the treatment level $t_3$. The type set $\Sigma_{t_1,2} = \{s_1\}$ contains the type that always choose the treatment $t_1$ and thus can be referred to as $t_1$-always taker. The same applies to $\Sigma_{t_2,2} = \{s_2\}$ and $\Sigma_{t_3,2} = \{s_3\}$. The type set $\Sigma_{t1,1} = \{s_4\}$ switches from $t_1$ to $t_3$ and hence can be considered as $t_1$-swticher (or $t_1$-compliter). Similarly, we can refer to $\Sigma_{t_2,1} = \{s_5\}$ as $t_2$-switcher and $\Sigma_{t_2,1} = \{s_5\}$ as $t_3$-switcher. This model is used in kline2016evaluating to study the causal effect of the Head Start preschool program. The instrument indicates whether the household receives a Head Start offer, and the treatment levels are $t_1=$ Head Start, $t_2 = $ other preschool programs, and $t_3 = $ no preschool. The unordered monotonicity condition means that anyone who changes behavior as a result of the Head Start offer does so to attend Head Start.

Identification Results

We introduce some matrix notations related to the type $S$. For each treatment level $t \in \mathcal{T}$, let $B_t$ be a binary matrix of the same dimension as the response matrix $R$ with each element of $B_t$ signifying whether the corresponding element in the response matrix is $t$. That is, $B_t[i,j]$, the $(i,j)$th element of $B_t$, is whether $T_{z_i}$ equals $t$ for the subpopulation $S=s_j$. Define $b_{t,k} \equiv \left(\mathbf{1}\{s_1 \in \Sigma_{t,k}\}, \cdots, \mathbf{1}\{s_{N_S} \in \Sigma_{t,k}\} \right) B_t^+,$ where $B_t^+$ is the Moore-Penrose inverse of $B_t$.

For convenience, we also need some notations regarding conditional expectations. Let

align*[align* omitted — 130 chars of source]

be the vector of functions that describes the conditional distribution of the instrument $Z$. For each treatment level $t \in \mathcal{T}$, let

align*[align* omitted — 122 chars of source]

be the vector that describes the conditional treatment probabilities given each level of the instrument. Denote

align*[align* omitted — 149 chars of source]

as the vector that contains the conditional outcomes for each treatment level $t$. Notice that the functions $\pi$, $P_{t}$, and $Q_t$ are all identified.

theorem[Identification of LASF] Let Assumptions (ref) - (ref) hold. Let $t \in \mathcal{T}$ and $k \in \{1,\cdots,N_Z\}$. \begin{enumerate}[label = (\roman*)] • The type set probability is identified by \begin{equation*} p_{t,k} \equiv \mathbb{P}(S \in \Sigma_{t,k}) = b_{t,k} \mathbb{E} \left[ P_{t}(X)\right]. \end{equation*} • If $p_{t,k} > 0$, the LASF is identified by: \begin{equation*} \beta_{t,k} = b_{t,k} \mathbb{E} \left[ Q_{t}(X) \right] / p_{t,k}. \end{equation*} \end{enumerate}

Theorem (ref) identifies $p_{t,k}$, the size of the subpopulation $\Sigma_{t,k}$, and the local structural function for that subpopulation. The only exception when the identification fails is when the type set $\Sigma_{t,0}$, in which case the individual never chooses the treatment $t$. This identification result is a modification of Theorem T-6 in heckman2018unordered that explicitly accounts for the presence of covariates $X$. Bayes rule is applied to convert the conditional result into the unconditional one. The following theorem presents the identification result for the LASF-T.

Let $\mathcal{Z}_{t,k} \subset \mathcal{Z}$ be the set of instrument values that induces the treatment level $t$ in the type set $\Sigma_{t,k}$. That is, $\mathcal{Z}_{t,k} \equiv \left\{ z_i \in \mathcal{Z} : s[i] = t \text{, for all } s \in \Sigma_{t,k} \right\}$, where $s[i]$ denotes the $i$th element of the vector $s$. Then define $\pi_{t,k} \equiv \sum_{z \in \mathcal{Z}_{t,k}} \pi_z$ as the total probability of those instrument values.

theorem[Identification of LASF-T] Let Assumptions (ref) - (ref) hold. Let $t \in \mathcal{T}$ and $k \in \{1,\cdots,N_Z\}$. Then $\mathcal{Z}_{t,k}$ is nonempty. \begin{enumerate}[label = (\roman*)] • The treatment probability within the type set is identified by \begin{equation*} q_{t,k} \equiv P\left( T = t, S \in \Sigma_{t,k} \right) = b_{t,k} \mathbb{E} \left[ P_{t}(X) \pi_{t,k}(X)\right]. \end{equation*} • If $q_{t,k} >0 $, then the LASF-T is identified by \begin{equation} \gamma_{t,k} = b_{t,k} \mathbb{E} \left[ Q_t(X) \pi_{t,k}(X) \right] / q_{t,k}. \end{equation} \end{enumerate}

The identification results are illustrated using the two examples.

example[continues = eg:binary] Since the treatment is binary, the matrix $B_1$ is equal to the response matrix $R$. The matrix $B_1$ and its generalized inverse $B_1^+$ are respectively \begin{align*} B_1 = \begin{pmatrix} 0 & 0 & 1 \\ 0 & 1 & 1 \end{pmatrix}, and (B_1^+)' = \begin{pmatrix} 0 & -1 & 1 \\ 0 & 1 & 0 \end{pmatrix}. \end{align*} The matrix $B_0$ and its generalized inverse $B_0^+$ are respectively \begin{align*} B_0 = \begin{pmatrix} 1 & 1 & 0 \\ 1 & 0 & 0 \end{pmatrix}, and (B_0^+)' = \begin{pmatrix} 0 & 1 & 0 \\ 1 & -1 & 0 \end{pmatrix}. \end{align*} The vectors $b_{1,1}$ and $b_{0,1}$ are respectively \begin{align*} b_{1,1} = (-1,1), and b_{0,1} = (1,-1). \end{align*} Theorem (ref) implies that \begin{align*} \beta_{1,1} = \frac{\mathbb{E}[Q_{1,1}(X)] - \mathbb{E}[Q_{1,0}(X)]}{\mathbb{E}[P_{1,1}(X)] - \mathbb{E}[P_{1,0}(X)]}, and \beta_{0,1} = \frac{\mathbb{E}[Q_{0,0}(X)] - \mathbb{E}[Q_{0,1}(X)]}{\mathbb{E}[P_{0,0}(X)] - \mathbb{E}[P_{0,1}(X)]}. \end{align*} The two denominators in the above expressions are both equal to the type probability of compliers. Then the usual identification of the LATE parameter frolich2007nonparametric follows: \begin{align*} \mathbb{E}[Y_1 - Y_0 \mid T_1 > T_0] = \frac{\int (\mathbb{E}[Y \mid Z=1,X=x] - \mathbb{E}[Y \mid Z=0,X=x]) f_X(x)dx}{\int (\mathbb{E}[T \mid Z=1,X=x] - \mathbb{E}[T \mid Z=0,X=x]) f_X(x)dx}, \end{align*} where $f_X$ denotes the marginal density function of $X$.
example[continues = eg:3t2z] Recall that $\Sigma_{t_1,1} = \{s_4\}$ contains the $t_1$-switcher. By Theorem (ref), the LASF for the treatment level $t$ and the subpopulation $S = s_4$ is identified by\footnote{The calculation of $b_{t,k}$ is omitted for brevity, but it can be done in the same way as Example (ref).} \begin{align*} p_{t_1,1} &= \mathbb{E}[P_{t_1,z_1}(X)] - \mathbb{E}[P_{t_1,z_2}(X)], \\ \beta_{t_1,1} &= \frac{\mathbb{E}[Q_{t_1,z_1}(X)] - \mathbb{E}[Q_{t_1,z_2}(X)]}{\mathbb{E}[P_{t_1,z_1}(X)] - \mathbb{E}[P_{t_1,z_2}(X)]}. \end{align*} Notice that $\mathcal{Z}_{t_1,1} = \{z_1\}$. Then by Theorem (ref) we have \begin{align*} q_{t_1,1} &= \mathbb{E}[(P_{t_1,z_1}(X) - P_{t_1,z_2}(X))\pi_{z_1}(X)], \\ \gamma_{t_1,1} &= \frac{\mathbb{E}[(Q_{t_1,z_1}(X) - Q_{t_1,z_2}(X))\pi_{z_1}(X)]}{\mathbb{E}[(P_{t_1,z_1}(X) - P_{t_1,z_2}(X))\pi_{z_1}(X)]}. \end{align*}

Semiparametric Efficiency

In this section, we calculate the semiparametric efficiency bound (SPEB) and propose estimators that achieve such bounds. We focus on the parameters LASF and LASF-T. In Appendix (ref), we study general parameters implicitly defined through moment restrictions.

LASF and LASF-T

For the rest of the paper, we assume that $Y_t, t \in \mathcal{T}$ have finite second moments. This is necessary since we are studying efficiency. Let $\iota$ denote the column vector of ones and $\zeta(Z,X,\pi)$ the diagonal matrix with the diagonal elements being $\mathbf{1}\{Z=z\}/\pi_{z}(X), z \in \mathcal{Z}$. The following theorem gives the efficient influence function (EIF) and the SPEB for the parameters identified in the preceding section.

theorem[SPEB for LASF and LASF-T] Let Assumptions (ref) - (ref) hold. Let $t \in \mathcal{T}$ and $k \in \{1,\cdots,N_Z\}$. Assume that $p_{t,k}, q_{t,k} > 0$. \begin{enumerate}[label = (\roman*)] • The semiparametric efficiency bound for $\beta_{t,k}$ is given by the variance of the efficient influence function \begin{align} \begin{split} &\psi^{\beta_{t,k}}(Y,T,Z,X,\beta_{t,k},p_{t,k},Q_t,P_{t},\pi)\\ = & \frac{1}{p_{t,k}} b_{t,k} \left( \zeta(Z,X, \pi) \left( \iota (Y \mathbf{1}\{T=t\}) - Q_t(X) \right) + Q_t(X) \right) \\ -& \frac{\beta_{t,k}}{p_{t,k}} b_{t,k} \left( \zeta(Z,X, \pi) \left( \iota \mathbf{1}\{T=t\} - P_t(X) \right) + P_t(X) \right). \end{split} \end{align} • The semiparametric efficiency bound for $\gamma_{t,k}$ is given by the variance of the efficient influence function \begin{align*} \begin{split} &\psi^{\gamma_{t,k}}(Y,T,Z,X,\gamma_{t,k},q_{t,k},Q_t,P_t,\pi) \\ =& \frac{1}{q_{t,k}} b_{t,k} \left( \zeta(Z,X, \pi) \left( \iota (Y \mathbf{1}\{T=t\}) - Q_t(X) \right)\pi_{{t,k}}(X) + Q_t(X)\mathbf{1}\{Z \in \mathcal{Z}_{t,k}\} \right) \\ -& \frac{\gamma_{t,k}}{q_{t,k}} b_{t,k} \left( \zeta(Z,X, \pi) \left( \iota \mathbf{1}\{T=t\} - P_t(X) \right)\pi_{{t,k}}(X) + P_t(X)\mathbf{1}\{Z \in \mathcal{Z}_{t,k}\} \right). \end{split} \end{align*} • The semiparametric efficiency bound for $p_{t,k}$ is given by the variance of the efficient influence function \begin{align*} \begin{split} \psi^{p_{t,k}}(T,Z,X,p_{t,k},P_t,\pi) = b_{t,k} \left( \zeta(Z,X, \pi) \left( \iota \mathbf{1}\{T=t\} - P_t(X) \right) + P_t(X) \right)- p_{t,k}. \end{split} \end{align*} • The semiparametric efficiency bound for $q_{t,k}$ is given by the variance of the efficient influence function \begin{align*} \begin{split} &\psi^{q_{t,k}}(T,Z,X,q_{t,k},P_t,\pi)\\ =& b_{t,k} \left( \zeta(Z,X, \pi) \left( \iota \mathbf{1}\{T=t\} - P_t(X) \right)\pi_{{t,k}}(X) + P_t(X)\mathbf{1}\{Z \in \mathcal{Z}_{t,k}\} \right) - q_{t,k}. \end{split} \end{align*} \end{enumerate}

The EIF in Theorem (ref) can be interpreted as the moment condition from the identification results modified by an adjustment term due to the presence of unknown infinite-dimensional parameters. Take $\psi^{\beta_{t,k}}$ as an example, the terms

align*[align* omitted — 115 chars of source]

and

align*[align* omitted — 123 chars of source]

are respectively the adjustment terms due to the presence of $Q_t$ and $P_t$.

From the expression of $\psi^{\beta_{t,k}}$, we can see that the SPEB would be large when $p_{t,k}$ is small. This is because $p_{t,k}$ measures the size of the subpopulation $S \in \Sigma_{t,k}$ on which the LASF is estimated. When $p_{t,k}$ is small, we run into the weak identification issue. In Section (ref), we study inference procedures that are robust against weak identification issues.

One benefit of the EIFs is that we can easily calculate the covariance matrix of different estimators. Consider an example where we are interested in two LASFs $\beta_1$ and $\beta_2$, whose EIF is given by $\psi_1$ and $\psi_2$, respectively. If the two estimators $\hat{\beta}_1$ and $\hat{\beta}_2$ are both semiparametric efficient, then their covariance matrix equals $\mathbb{E}[\psi_1\psi_2']$.

example[continues = eg:binary] In the binary LATE model, the first two parts of Theorem (ref) reduce to Theorem 2 of hong2010semiparametric. If we assume unconfoundedness by having $T = Z$, then the result further reduces to Theorem 1 of hahn1998role.

The derived SPEB helps determine whether an estimation procedure is efficient. In this section, we focus on the condition expectation projection (CEP) estimator.\footnote{The terminology “condition expectation projection” is adopted from the papers chen2008semiparametric and hong2010semiparametric, whereas hahn1998role refers to these estimators as “nonparametric imputation based estimators.”} Define

align*[align* omitted — 202 chars of source]

The CEP procedure first estimates $\pi_z$, $h_{Y,t,z}$, and $h_{t,z}$ by using nonparametric estimators $\hat{\pi}_z$, $\hat{h}_{Y,t,z}$, and $\hat{h}_{t,z}$ respectively. These estimators can be constructed based on series or local polynomial estimation. Then $Q_{t,z}$ and $P_{t,z}$ are estimated using $\hat{Q}_{t,z} = \hat{h}_{Y,t,z} / \hat{\pi}_z$ and $\hat{P}_{t,z} = \hat{h}_{t,z} / \hat{\pi}_z$. The vectors of estimators $\hat{Q}_{t}$ and $\hat{P}_{t}$, $\hat{\pi}$ are stacked in an obvious way. Let $\hat{\pi}_{{t,k}} = \sum_{z \in \mathcal{Z}_{t,k}}\hat{\pi}_{z}$. The CEP estimators for the structural parameters are defined by

align*[align* omitted — 399 chars of source]

The next proposition shows that the CEP estimators are semiparametrically efficient. The result is similar in style to hahn1998role's (hahn1998role) Proposition 4 that the low-level regularity conditions are omitted. Instead, the proposition assumes the high-level condition that the CEP estimators are asymptotically linear, which means they are asymptotically equivalent to sample averages. More formally, an estimator $\hat{\beta}$ of $\beta$ is asymptotically linear if it admits an influence function. That is, there exists an iid sequence $\psi_i$ with zero mean and finite variance such that

align*[align* omitted — 94 chars of source]

Since each element of the conditional expectations $h_{Y,t,z}$, $h_{t,z}$, and $\pi_z$ can be considered as coming from a binary LATE model, the regularity conditions in hong2010supplement should work with little modification.

propositionSuppose the CEP estimators are asymptotically linear, then they achieve the semiparametric efficiency bound.
commentThe efficient influence functions also provide a way for conducting optimal joint inferences. The efficient influence function for a vector of parameters is the collection of the efficient influence functions corresponding to the parameters. The variance-covariance matrix for efficient estimators can be calculated accordingly. More concretely, let $ \kappa = (\kappa_1, \cdots, \kappa_{K})$ be a vector of parameters selected from the identified ones $\{\beta_{t,k}, \gamma_{t,k}, p_{t,k}, q_{t,k}\}$. Then the efficient influence function of $\kappa$ is $\Psi_\kappa = \left(\Psi_{\kappa_1}, \cdots, \Psi_{\kappa_K} \right)' $, and the efficiency bound is $\mathbb{E} \left[ \Psi_\kappa \Psi_\kappa' \right] $. A natural plug-in estimator $\hat{V}_\kappa$ defined below can be employed to consistently estimate this bound under mild conditions, \begin{equation} \hat{V}_\kappa = \frac{1}{n} \sum_{i=1}^n \hat{\Psi}_{\kappa,i} \hat{\Psi}_{\kappa,i}' \end{equation} where $\hat{\Psi}_{\kappa,i}$ is defined by plugging-in the data observation of index $i$, the CEP estimates $\hat{\kappa}$, and the nonparametric estimators $\hat{\mathcal{I}}_{t,z}$, $\hat{P}_{t,z}$, and $\hat{\pi}_z$. For example, when $\kappa = \beta_{t,k}$, then $$ \hat{\Psi}_{\beta_{t,k},i} = \Psi_{\beta_{t,k}}(Y_i,T_i,Z_i,X_i,\hat{\beta}_{t,k},\hat{p}_{t,k},\hat{\mathcal{I}}_{t,Z},\hat{P}_{t,Z},\hat{\pi}). $$ Since the CEP estimators are efficient, their asymptotic covariance matrix can be estimated by the plug-in estimators for efficiency bounds. This leads to optimal joint inferences, where the optimality is by Section 25.6 of van1998asymptotic, where it is stated that the semiparametric efficiency of the estimators leads to (locally) asymptotically uniformly most powerful tests. The following theorem summarizes the properties of the CEP estimation procedure defined above. For each $t$ and $z$, let $\mathscr{H}_{Y,t,z}$, $\mathscr{H}_{t,z}$, and $\Pi_{z}$ be the space of functions containing the true nuisance parameters $h_{Y,t,z}^o$, $h^o_{t,z}$, and $\pi_{z}^o$ respectively. For any small enough $\delta > 0$, denote the nonparametric parameter space $\mathscr{H}_{Y,t,z}^\delta= \left\{ h_{Y,t,z} \in \mathscr{H}_{Y,t,z} : \norm{h_{Y,t,z} - h^o_{Y,t,z}} \leq \delta \right\} $, and let $\mathscr{H}_{t,z}^\delta$ and $\Pi_{z}^\delta$ be defined analogously. \begin{theorem} Consider $t \in \mathcal{T}$. For any $z \in \mathcal{Z}$, assume the following conditions hold. \begin{enumerate}[label = (\roman*)] • The convergence rate of nonparametric estimators satisfy $\sqrt{n} \norm{\hat{h}_{Y,t,z} - h_{Y,t,z}^o}^2 = o_p(1)$,\\ $\sqrt{n} \norm{\hat{h}_{t,z} - h_{t,z}^o}^2 = o_p(1)$, and $\sqrt{n} \norm{\hat{\pi}_{z} - \pi_{z}^o}^2 = o_p(1)$. • There exists some $\delta>0$ such that the classes $\mathscr{H}_{Y,t,z}^\delta$, $\mathscr{H}_{t,z}^\delta$, and $\Pi_{z}^\delta$ are Donsker, and satisfying the condition $ \mathbb{E} \left[ \sup_{h_{Y,t,z} \in \mathscr{H}_{Y,t,z}^\delta} \abs{h_{Y,t,z}(X)} \right] < \infty $. \end{enumerate} Then, for $k = 1, \cdots, N_Z$, estimators $\hat{\beta}_{t,k}$, $\hat{\gamma}_{t,k}$, $\hat{p}_{t,k}$, and $\hat{q}_{t't,k}$ are semiparametric efficient for $\beta_{t,k}$, $\gamma_{t,k}$, $p_{t,k}$, and $q_{t,k}$, respectively. Moreover, the plug-in estimator $\hat{V}_\kappa$ for efficiency bound, defined in equation((ref)), is consistent. \end{theorem} Condition (ref)(ref) is a standard requirement on the convergence rate of nonparametric estimators in the semiparametric two-step estimation literature newey1994asymptotic,newey1994large. Condition (ref)(ref) is also standard that requires the functional spaces containing the infinite-dimensional nuisance parameters are not too complex, for the stochastic equicontinuity condition to hold.

The reason that this type of estimator is efficient is well explained in ackerberg2014asymptotic. The estimation problem here falls into their general semiparametric model, where the finite-dimensional parameter of interest is defined by unconditional moment restrictions. They show that the semiparametric two-step optimally weighted GMM estimators, the CEP estimators in this case, achieve the efficiency bound since the parameters of interest are exactly identified. Discussions related to this phenomenon can also be found in chen2018overidentification.

We next examine the efficient estimation of other policy-relevant parameters that can be derived from the parameters $\left( \beta_{t,k},\gamma_{t,k},p_{t,k},q_{t,k} \right) $. As an example, consider the type set $\Sigma_t \equiv \cup_{k=1}^{N_{Z-1}} \Sigma_{t,k}$, which is referred to as $t$-switchers. This subpopulation contains individuals who switch between $t$ and other treatments when given different levels of instruments. It is a generalization of the concept of compliers in the binary LATE framework.\footnote{Recall that switchers are also illustrated in Example (ref).} The LASF for the subpopulation $\Sigma_t$ is given by

align*[align* omitted — 159 chars of source]

Similarly, one can also define

equation[equation omitted — 163 chars of source]

which represents the LASF-T for the subpopulation of $t$-treated $t$-switchers.

For some subpopulations, a treatment effect can be identified. This point is already illustrated with Example (ref) in the discussion of the identification of the usual LATE parameter. We further illustrate this point with Example (ref).

example[continues = eg:3t2z] The quantity \begin{align*} \beta_{t_3,1} - \frac{\beta_{t_1,1} p_{t_1,1} + \beta_{t_2,1} p_{t_2,1}}{p_{t_1,1} + p_{t_2,1}} \end{align*} represents the local average treatment effect of $t_3$ against other treatments within the subpopulation of $t_3$-switchers. Analogously, the parameter \begin{align*} \gamma_{t_3,1} - \frac{\gamma_{t_3,t_1,1} q_{t_3,t_1,1} + \gamma_{t_3,t_2,1} q_{t_3,t_2,1}}{q_{t_3,t_1,1} + q_{t_3,t_2,1}} \end{align*} is the local average treatment effect of $t_3$ against other treatments within the subpopulation of $t_3$-treated $t_3$-switchers.

To summarize the above examples using a general expression, let $\phi = \phi(\underline{p},\underline{q},\underline{\beta},\underline{\gamma})$ be a finite-dimensional parameter, where $\phi(\cdot)$ is a known continuously differentiable function, and $\underline{p}$ is the vector containing all identifiable $p_{t,k}$'s, that is, $\underline{p} \equiv \{p_{t,k} : t \in \mathcal{T}, 1 \leq k \leq N_Z\}$. Let $\underline{q},\underline{\beta}$, and $\underline{\gamma}$ be defined analogously. A natural estimator can be defined through the CEP estimates, $\phi(\hat{\underline{p}},\hat{\underline{q}},\hat{\underline{\beta}},\hat{\underline{\gamma}})$. The delta method can help calculate the efficiency bound of $\phi$ and show the efficiency of $\phi(\hat{\underline{p}},\hat{\underline{q}},\hat{\underline{\beta}},\hat{\underline{\gamma}})$. In fact, by Theorem 25.47 of van1998asymptotic, we immediately have the following corollary, which shows that plug-in estimators are efficient.

corollaryThe semiparametric efficiency bound of $\phi$ is given by the variance of efficient influence function \begin{equation} \psi^\phi = \sum_{p \in p} \frac{\partial \phi}{\partial p} \psi^p + \sum_{q \in q} \frac{\partial \phi}{\partial q} \psi^q + \sum_{\beta \in \beta} \frac{\partial \phi}{\partial \beta} \psi^\beta + \sum_{\gamma \in \gamma} \frac{\partial \phi}{\partial \gamma} \psi^\gamma \end{equation} where the partial derivatives are evaluated at the true parameter value. Moreover, the plug-in estimator $\phi(\hat{\underline{p}},\hat{\underline{q}},\hat{\underline{\beta}},\hat{\underline{\gamma}})$, based on the CEP estimators $\hat{\underline{p}},\hat{\underline{q}},\hat{\underline{\beta}},\hat{\underline{\gamma}}$, achieves the efficiency bound.

Robustness

In the previous section, the EIF is used as a tool for computing the SPEB. In this section, we directly use the EIF as the moment condition for estimation. These moment conditions are appealing because they satisfy double robustness and local robustness --- the two topics of this section.

A word on notation: in the rest of the paper, we use a superscript $o$ to signify the true value whenever necessary. For example, when both $\pi^o$ and $\pi$ appear, the former means the true probability while the latter denotes a generic function.

Double Robustness

We focus on the LASF $\beta_{t,k}$. The same analysis can be applied to the other parameters. To avoid notational burden in the main text, we drop the subscript $(t,k)$ in $\beta_{t,k}$, $p_{t,k}$, and $b_{t,k}$, and the subscript $t$ in $P_t$ and $Q_t$.\footnote{The full subscripts are kept in the Appendices.} It is straightforward to verify that the EIF $\psi^{\beta}$ has zero mean. However, we do not want to use $\psi^{\beta}$ itself as the estimating equation since it contains $1/p$ as a factor. To deal with this problem, we simply multiply $\psi^{\beta}$ by $p$ and define

align*[align* omitted — 290 chars of source]

The corresponding moment condition is

align[align omitted — 104 chars of source]

This moment condition is doubly robust, as demonstrated in the following proposition.

proposition[Double Robustness] Let $\left( Q,P,\pi \right)$ be an arbitrary vector of functions and $( Q^o,P^o,\pi^o ) $ the true vector of conditional expectations. Then \begin{align*} \mathbb{E}\left[ \psi(Y,T,Z,X,\beta^o,Q^o,P^o,\pi) \right]=0 \end{align*} and \begin{align*} \mathbb{E}\left[ \psi(Y,T,Z,X,\beta^o,Q,P,\pi^o) \right]=0. \end{align*}

The above proposition divides the nonparametric nuisance parameters into two groups, $\pi$ and $(Q,P)$. The doubly robust moment condition is valid if either of these two groups of nuisance parameters is true. On the other hand, if the researcher uses parametric models for these nuisance parameters, then the structural parameter $\beta$ can be recovered provided that at least one of the working nuisance models is correctly specified. Therefore, the doubly robust moment condition is “less demanding" on the researcher's ability to devise a correctly specified model for the nuisance parameters. The double robustness result in Proposition (ref) can be seen as the GLATE extension of the existing results in the binary LATE literature tan2006regression,okui2012doubly.

Neyman Orthogonality

The second robustness property is Neyman orthogonality. Moment conditions with this property have reduced sensitivity with respect to the nuisance parameters. Formally, Neyman orthogonality means that the moment condition has zero Gateaux derivative with respect to the nuisance parameters. The result is presented in the following proposition.

proposition[Neyman Orthogonality] Let $\left( Q,P,\pi \right) $ be an arbitrary set of functions. For $r \in [0,1)$, define $Q^r = Q^o + r(Q - Q^o),$ $P^r = P^o + r(P - P^o),$ and $\pi^r = \pi^o + r(\pi - \pi^o)$. Suppose that $\sup_{r \in [0,1]}\big|\frac{\partial}{\partial r} \psi(Y,T,Z,X,\beta,Q^r,P^r,\pi^r) \big|$ is integrable, then \begin{align*} \frac{\partial}{\partial r} \mathbb{E} \left[ \psi(Y,T,Z,X,\beta,Q^r,P^r,\pi^r) \right] \Big|_{r=0} = 0, \end{align*} where $\beta$ does not need to be the true parameter value.

In many econometrics models, double robustness and Neyman orthogonality come in pairs. Discussions about their general relationships can be found in chernozhukov2016locally. In practice, double robustness is often used for parametric estimation, as previously explained, whereas Neyman orthogonality is used in estimation with the presence of possibly high-dimensional nuisance parameters.

Next, we apply the double/debiased machine learning (DML) method developed by chernozhukov2018double to the moment condition ((ref)). This estimation method works even when the nuisance parameter space is complex enough that the traditional assumptions, e.g., Donsker properties, are no longer valid.\footnote{In two-step semiparametric estimations, Donsker properties are usually required so that a suitable stochastic equicontinuity condition is satisfied. See, for example, Assumption 2.5 in chen2003estimation.} The implementation details are explained below.

The nuisance parameters $Q$, $P$, and $\pi$ are estimated using a cross-fitting method: Take an $L$-fold random partition of the data such that the size of each fold is $n/L$. For $l = 1, \cdots, L$, let $I_l$ denote the set of observation indices in the $l$th fold and $I^c_l = \bigcup_{l' \ne l} I_{l'}$ the set of observation indices not in the $l$th fold. Define $\check{Q}^l$, $\check{P}^l$, and $\check{\pi}^l$ to be the estimates constructed by using data from $I_l^c$. The DML estimator of $\beta$ is constructed following the moment condition ((ref)):\footnote{This is the DML2 estimator defined in chernozhukov2018double. Another estimator, the DML1 estimator, is proposed in the same paper. We do not study the DML1 estimator since it is asymptotically equivalent to DML2, and the authors generally recommend DML2.}

align[align omitted — 364 chars of source]

To conduct inference, we also need an estimate for the asymptotic variance of $\check{\beta}$, which we denote by $\sigma^2$. The asymptotic variance equals to the expectation of the squared efficient influence function: $\sigma^2 = \mathbb{E}\left[ \psi^{\beta} \right]^2 = \mathbb{E}[\psi^2]/p^2$. We first estimate $p$ by using the cross-fitting method, which is essentially given by the denominator of ((ref)):

align[align omitted — 208 chars of source]

Then the asymptotic variance can be estimated by

align*[align* omitted — 363 chars of source]

We want to establish the convergence results for the DML estimator uniformly over a class of data generating processes (DGPs) defined as follows. For any two constants $c_1 > c_0 >0$, let $\mathcal{P}(c_1,c_0)$ be the set of joint distributions of $(Y,T,Z,X)$ such that

enumerate[label = (\roman*)] • $p \in [c_0,1]$, • $\mathbb{E}[\psi^2],\pi_z^o(X) \geq c_0,z \in \mathcal{Z}$, and $|Y \mathbf{1}\{T=t\}|,|Y \mathbf{1}\{T=t\} - Q_t^o(X)| \leq c_1$.

The first condition excludes the case where $\beta$ is weakly identified (when $p$ can be arbitrarily close to zero). Inference under weak identification is studied in the next section. The following theorem establishes the asymptotic properties of the DML estimation procedure. In particular, the estimator achieves the SPEB.

theoremLet Assumptions (ref) and (ref) hold. Assume the following conditions on the nuisance parameter estimators $(\check{Q}^l,\check{P}^l,\check{\pi}^l)$: \begin{enumerate}[label = (\roman*)] • For $z \in \mathcal{Z}$, $|\check{Q}^l|$ is bounded, $\check{P}^l_{z}$ and $\check{\pi}^l_z \in [0,1]$, and $\check{\pi}^l_z$ is bounded away from zero. • $\max_{z \in \mathcal{Z}} \big( \Vert \hat{Q} - Q^o \rVert_2 \vee \Vert \hat{P} - P^o \rVert_2 \vee \lVert \hat{\pi} - \pi^o \rVert_2 \big) = o_p\big( n^{-1/4} \big)$. \end{enumerate} Then the estimator $\check{\beta}$ obeys that \begin{equation*} \sigma^{-1} \sqrt{n} \big( \check{\beta} - \beta \big) \Rightarrow N(0,1), \end{equation*} uniformly over the DGPs in $\mathcal{P}(c_0,c_1)$. Moreover, the above convergence result continues to hold when $\sigma$ is replaced by the estimator $\check{\sigma}$.
comment\begin{align} \begin{split} & \mathbb{E} \Big[ b_{t,k} \left( \zeta(Z,X, \pi) \left( \iota (Y \mathbf{1}\{T=t\}) - Q_t(X) \right) + Q_t(X) \right) \\ - & \beta_{t,k} b_{t,k} \left( \zeta(Z,X, \pi) \left( \iota \mathbf{1}\{T=t\} - P_t(X) \right) + P_t(X) \right) \Big] = 0. \end{split} \end{align} Using this moment condition, the estimator $\check{\beta}_{t,k}$ is defined by The variance estimator is given by \begin{equation} \check{V}_{\beta_{t,k}} = \frac{1}{nL} \sum_{l=1}^L \sum_{i=1}^n \left( \Psi_{\beta_{t,k}}\left( Y_i,T_i,Z_i,X_i, \check{\beta}_{t,k}, \check{p}_{t,k}, \check{\mathcal{I}}_{t,Z}, \check{P}_{t,Z}, \check{\pi} \right) \right)^2 \end{equation} where \begin{equation} \check{p}_{t,k} = \frac{1}{nL} \sum_{l=1}^L \sum_{i=1}^n b_{t,k} \left( \zeta(Z_i,X_i, \check{\pi}) \left( \iota \mathbf{1}\{T_i=t\} - \check{P}_{t,Z}(X_i) \right) + \check{P}_{t,Z}(X_i) \right) \end{equation} \begin{theorem} Let $\delta_n \geq n^{-1/2}$ and $\Delta_n$ be some sequences of positive constants approaching zero. Also, let $C>0$ and $q>2$ be fixed constants, and $L \geq 2$ be fixed integer. Assume the following conditions hold for any joint distribution $P \in \mathcal{P}$ for the quadruple $(Y,T,Z,X)$. \begin{enumerate} [label = (\roman*)] • The variance bound for $\beta_{t,k}$ calculated in Theorem (ref) is strictly positive. • $\max \left\{ \norm{Q_t^o}_q ,\norm{ \iota Y \mathbf{1}\{T = t\} - Q_t^o}_q \right\} \leq C$. • With probability no less than $ 1 - \Delta_n$, $\max \left\{ \norm{\check{\mathcal{I}}_{t,Z} - Q_t^o}_q, \norm{\check{P}_{t,Z} - P_t^o}_q, \norm{ \check{\pi} - \pi^o}_q \right\} \leq C$, $\max \left\{ \norm{\check{\mathcal{I}}_{t,Z} - Q_t^o}_2, \norm{\check{P}_{t,Z} - P_t^o}_2, \norm{ \check{\pi} - \pi^o}_2 \right\} \leq \delta_n$, and for any $z \in \mathcal{Z}$, $\check{\pi}_z \geq \underline{\pi}$ and $\norm{ \check{\pi}_z - \pi_z^o}_2 \times \left( \norm{\check{\mathcal{I}}_{t,z} - Q_t^o}_2 + \norm{\check{P}_{t,z} - P_t^o}_2 \right) \leq n^{-1/2} \delta_n $. \end{enumerate} Then the estimator $\check{\beta}_{t,k}$ obey \begin{equation} V^{-1/2}_{\beta_{t,k}} \sqrt{n} \left( \check{\beta}_{t,k} - \beta_{t,k} \right) \Rightarrow N(0,1), \end{equation} uniformly over $\mathcal{P}$, where $V_{\beta_{t,k}} = \mathbb{E}\left[ \Psi_{\beta_{t,k}}(Y,T,Z,X,\beta_{t,k}^o,p_{t,k}^o, Q^o_t,P^o_t,\pi^o_{t,Z})^2 \right] $. Moreover, the results continue to hold when $V_{\beta_{t,k}}$ is replaced by $\check{V}_{\beta_{t,k}}$. \end{theorem}

The proof verifies the conditions of Theorem 3.1 in chernozhukov2018double. The essential restriction is on the uniform convergence rate for the estimators of the nuisance parameters. In low-dimensional settings, one can consider the local polynomial regression for estimation of the conditional expectations. Under suitable conditions hansen2008uniform,masry1996multivariate, the uniform convergence rate of the local polynomial estimators is $(\log n / n)^{2/(d_X+4)}$, which is $o(n^{-1/4})$ if $d_X \leq 3$. In high-dimensional settings, as pointed out by chernozhukov2018double, the rate $o(n^{-1/4})$ is often available for common machine learning methods under structured assumptions on the nuisance parameters.\footnote{This includes the LASSO method under sparsity of the nuisance space. See, for example, buhlmann2011statistics, belloni2011l1, and belloni2013least. However, chernozhukov2018double also indicate that to prove that machine learning methods achieve the $o(n^{-1/4})$ rate, one will eventually have to use related entropy conditions.} This means that the asymptotic normality of the DML estimator continues to hold.

Theorem (ref) can be directly used to conduct inference on $\beta$. Confidence regions can be constructed by inverting the usual $t$-tests. These confidence regions are uniformly valid since the convergence results in the above theorem hold uniformly over $\mathcal{P}$. In the next section, we explain why uniform validity is crucial when dealing with weak identification issues.

Weak Identification

The convergence result established in Theorem (ref) is uniform over the set of DGPs with type probability $p$ bounded away from zero. However, the identification of $\beta$ would be weak in the case where $p$ can be arbitrarily close to zero. This leads to distortion of the uniform size of the test and poor asymptotic approximation in finite-sample settings. This section studies this weak identification issue and proposes an inference procedure that is robust against such a problem.

We begin with a heuristic illustration of the weak identification problem. To ease notation, define $\upsilon = \beta p$ and

align*[align* omitted — 215 chars of source]

After a simple calculation, we can write

align*[align* omitted — 148 chars of source]

In the above expression, we can interpret the estimation errors $\sqrt{n}(\check{\upsilon}-\upsilon)$ and $\sqrt{n}(\check{p}-p)$ as the noises, while the signal is the term $\sqrt{n}p$. Under the usual asymptotics where $p > 0$ is fixed, the noise terms are bounded in probability, whereas the signal term $\sqrt{n}p \rightarrow \infty$. Hence, the signal dominates the noise, and the estimator $\check{\beta}$ is consistent. However, under asymptotics with a drifting sequence $p = p_n \rightarrow 0$ and $\sqrt{n}p$ converging to a finite constant, the signal and the noise are of the same magnitude, which results in the inconsistency of $\check{\beta}$. This problem is the weak identification issue. In the weak IV literature, a common measure of identification strength is the so-called concentration parameter. In our case, the concentration parameter is given by $\sqrt{n}p$ where $\sqrt{n}p \rightarrow \infty$ corresponds to strong identification, and identification is weak when the limit of $\sqrt{n}p$ is finite.

While weak identification is a finite-sample issue, it is formalized using the asymptotic framework. However, the illustration above using asymptotics under drifting sequences is not meant to model DGPs that vary with the sample size $n$. Instead, it is a tool used to detect the lack of uniform convergence. In fact, controlling the uniform size of the test is the key to solving weak identification problems.\footnote{See, for example, imbens2004confidence, mikusheva2007uniform, and andrews2020generic.} Formally, the uniform size of a test is the large sample limit of the supremum of the rejection probability under the null hypothesis, where the supremum is taken over the nuisance parameter space. When testing a null hypothesis on $\beta$ in the GLATE model, the supremum mentioned above is taken over all values of $p>0$. That is, a desirable test should have rejection probability under the null converge to the nominal size uniformly over $p \in (0,1]$. From the previous discussion, we can see that the uniform size can not be controlled using the usual $t$-statistic $\sqrt{n}(\check{\beta} - \beta)/\check{\sigma}$. This failure of uniform convergence, however, does not conflict with Theorem (ref), where the uniform convergence of $\check{\beta}$ is established only after restricting $p$ to be bounded away from zero.

Inference procedures that are robust against weak identification can be obtained by directly imposing the null hypothesis in the construction of the test statistic. One such example is the well-known Anderson-Rubin (AR) statistic in the weak IV literature. Its idea can be generalized to the GLATE model. We first consider testing the two-sided hypothesis $H_0:\beta = \beta_0$ versus $H_1:\beta \ne \beta_0$. To control the uniform size of the test, we need the test statistic to converge uniformly on the parameter space where (1) $\beta = \beta_0$, and (2) $p$ is allowed to be arbitrarily close to zero. A null-restricted $t$-statistic can be obtained as follows. Notice that when $p>0$, $\beta = \beta_0$ is equivalent to

align[align omitted — 129 chars of source]

Its estimate can be written as

align[align omitted — 151 chars of source]

Under the null hypothesis $\beta = \beta_0$, the above estimate does not depend on the concentration parameter $\sqrt{n}p$ and consists only of the noise terms $\check{\upsilon} - \upsilon$ and $\check{p} - p$, whose uniform convergence can be established directly.

For implementation, this test statistic can be obtained as a straightforward application of the DML procedure described in the previous section to the moment condition ((ref)). As a consequence of Proposition (ref), the above moment condition satisfies the Neyman orthogonality condition regardless of the true value of $\beta$. More specifically, the null-restricted $t$-statistic is defined to be

align*[align* omitted — 99 chars of source]

where

align*[align* omitted — 153 chars of source]

The corresponding test of $H_0:\beta = \beta_0$ against $H_1:\beta \ne \beta_0$ rejects for large values of $|\check{\rho}|$.

The same methodology can be applied to testing one-sided hypothesis $H_0:\beta \leq \beta_0$ versus $H_1:\beta > \beta_0$. Under the null hypothesis, $(\beta - \beta_0)p$ is non-positive, suggesting that the test should reject for large values of $\check{\rho}$. Notice that this relies on knowing the sign of $p$ due to the GLATE model structure. This restriction on the sign of $p$ is similar to knowing the first-stage sign in the linear IV model, which is studied by andrews2017unbiased in the context of unbiased estimation.

We now define the set of DGPs that allows $p$ to be arbitrarily close to zero. For any two constants $c_1>c_0>0$, let $\mathcal{P}^{\text{WI}}(c_0,c_1)$ be the set of joint distributions of $(Y,T,Z,X)$ such that

enumerate[label = (\roman*)] • $p \in (0,1]$, • $\mathbb{E}[\psi^2],\pi_z^o(X) \geq c_0,z \in \mathcal{Z}$, and $\abs{Y \mathbf{1}\{T=t\}},|Y \mathbf{1}\{T=t\} - Q_t^o(X)| \leq c_1$.

For any $\beta' \in \mathbb{R}$, let $\mathcal{P}^{\text{WI}}_{\beta'}(c_0,c_1)$ be the subset of $\mathcal{P}^{\text{WI}}(c_0,c_1)$ in which the true value of the parameter $\beta$ is $\beta'$. In particular, $\mathcal{P}^{\text{WI}}_{\beta_0}(c_0,c_1)$ denotes the subset where the null hypothesis is true. The superscript “WI” denotes weak identification. The difference between $\mathcal{P}(c_0,c_1)$ and $\mathcal{P}^{\text{WI}}(c_0,c_1)$ is that $\mathcal{P}^{\text{WI}}(c_0,c_1)$ allows the type probability $p$ to be arbitrarily small, whereas the type probabilities in $\mathcal{P}(c_0,c_1)$ are uniformly bounded away from zero. Denote $\mathcal{N}_{\nu}$ as the $\nu$th quantile of the standard normal distribution. The following theorem establishes that the above testing procedures have uniformly correct sizes and are consistent.

theoremSuppose the conditions on the nuisance parameter estimates in Theorem (ref) hold. Let $\alpha \in (0,1)$ be the nominal size of the tests. \begin{enumerate} [label = (\roman*)] • The test that rejects $H_0:\beta = \beta_0$ in favor of $H_0:\beta \ne \beta_0$ when $|\check{\rho}| > \mathcal{N}_{1-\frac{\alpha}{2}}$ has (asymptotically) uniformly correct size and is consistent. That is, \begin{align*} \sup \big\{ \mathbb{P}_P\big(|\check{\rho}| > \mathcal{N}_{1-\frac{\alpha}{2}}\big): P \in \mathcal{P}^{WI}_{\beta_0}(c_0,c_1) \big\} \rightarrow \alpha \end{align*} and \begin{align*} \mathbb{P}_P \big(|\check{\rho}| > \mathcal{N}_{1-\frac{\alpha}{2}} \big) \rightarrow 1, P \in \mathcal{P}^{WI}_{\beta}(c_0,c_1), \beta \ne \beta_0. \end{align*} • The test that rejects $H_0:\beta \leq \beta_0$ in favor of $H_0:\beta > \beta_0$ when $\check{\rho} > \mathcal{N}_{1-\alpha}$ has (asymptotically) uniformly correct size and is consistent. That is, \begin{align*} \sup \big\{ \mathbb{P}_P\big(\check{\rho} > \mathcal{N}_{1-\alpha} \big): P \in \mathcal{P}^{WI}_{\beta}(c_0,c_1),\beta \leq \beta_0 \big\} \rightarrow \alpha \end{align*} and \begin{align*} \mathbb{P}_P \big(\check{\rho} > \mathcal{N}_{1-\alpha} \big) \rightarrow 1, P \in \mathcal{P}^{WI}_{\beta}(c_0,c_1), \beta > \beta_0. \end{align*} \end{enumerate}
comment\section{Numerical Results} \subsection{Empirical Application} In this section, the estimation methods discussed in the paper are applied to study the return to schooling using proximity to college as an instrument card1993using. The data comes from the National Longitudinal Survey on the original cohort of young men at the time of 1966 and 1976. It is shown in kitagawa2015test that the instrument introduced is only valid after conditioning on covariates such as race and region of residence, making this example appropriate for illustrating the usefulness of including conditioning covariates in the framework. The model is the same as the main example, and the variables are explained as follows. The outcome $Y$ is the log of weekly earning in 1976. The instrument is the binary $Z$ indicating whether a four-year local college was present for the individual in 1966. It is relevant because the presence of a college nearby gets the individual to know about college in early life and reduces the cost of receiving some college education, and valid since the individual's ability is presumed to be independent of the place of residence, given race and the extent to which this region is developed. The treatment $T$ describes the education received by the individual up to the time of 1976. Instead of being binary, $T$ takes on three unordered values $\{t_1,t_2,t_3\}$, where $t_1$ means the student receives college education in fields including engineering, mathematics, law, social sciences, and others, $t_2$ means the student receives college education in other fields including business, education, and public services among others, and $t_3$ means the student does not receive college education.\footnote{This classification of the fields of study is for balancing the sample size in the dataset.} In the binary LATE case, $t_1$ and $t_2$ would be collapsed into one treatment indicating college-level education for the study of return to college schooling. The unobserved type $S$ is defined by the way education decisions vary with the proximity to the college. The conditioning covariates $X$ includes race, whether residence in southern states, and whether residence in the standard metropolitan area. The available sample size is 2930, 381 of which are treated by $t_1$, and 487 by $t_2$. For estimation, no advanced nonparametric method is needed since $X$ is discrete. The $P_t$'s are estimated using the linear probability model with five dummies as in kitagawa2015test. The $\pi_z$'s are estimated with the sample means. Then we use the CEP estimators to evaluate the parameters of interest. The estimation results for the LASFs and LASF-Ts are displayed in Table (ref). The asymptotic covariance matrix can be easily estimated using the efficient influence functions, as in Table (ref), leading to joint inferences of the parameters, which is another benefit of the methodology developed in this paper. Table (ref) shows the results of some selected statistical tests comparing both between and within LASFs and LASF-Ts. The differences between the pairs $\beta_{t_1,1}, \beta_{t_2,1}$ and $\gamma_{t_1,1}, \gamma_{t_2,1}$ are both insignificant, indicating that the income after receiving college education on the two categories of fields are similar for their own switchers respectively. As mentioned before, these facts can be turned into overidentifying restrictions ($\beta_{t_1,1}= \beta_{t_2,1}$, $\gamma_{t_1,1}= \gamma_{t_2,1}$) to improve efficiency, if supported by underlying economic theory. A college education is generally perceived as a causal factor for increasing income. Thus it should be the case that the LASF-T is higher than the LASF when the treatment belongs to $\{t_1,t_2\}$, and lower when treatment is $t_3$, which is consistent with the testing results. Also, both $\beta_{t_1,1}$ and $\beta_{t_1,1}$ are higher than $\beta_{t_3,1}$, where the insignificance is due to the fact that these three parameters are averages over different subpopulations. Comparisons among $\beta_{t_1,2}$, $\beta_{t_2,2}$, and $\beta_{t_3,2}$ shows similar intuition on the effect of schooling with a tendency that the outcomes for always-takers are less spread across treatments. The ratios $P(T=t_1 \mid S = s_4)$ and $P(T=t_2 \mid S = s_5)$ are estimated to be 0.68 and 0.92, revealing a significant amount of individuals receiving college education in the subpopulations of switchers. At this point, one might question whether the parameter estimates are of policy interest. The argument here is that clearly by the identification results, the parameters discussed here are at least as informative as the LATE (and LATT) parameters in the binary LATE case. And one of the themes of LATE and other causal inference models is to tradeoff informativeness with removing incredible assumptions. \begin{table} \caption{Empirically Estimated LASFs and LASF-Ts} \begin{tabular}{c c c c} \toprule Parameter & Estimate & Std Error & 95% CI \\ \midrule $\beta_{t1,1}$ & 5.318 & 0.689 & $[3.968,6.668]$ \\ $\beta_{t1,2}$ & 5.539 & 0.347 & $[4.859,6.219]$ \\ $\beta_{t2,1}$ & 5.785 & 0.698 &$[4.417,7.153]$ \\ $\beta_{t2,2}$ & 5.525 & 0.034 &$[5.458,5.592]$ \\ $\beta_{t3,1}$ & 4.965 & 0.624 &$[3.742,6.188]$ \\ $\beta_{t3,2}$ & 5.383 & 0.015 &$[5.354,5.412]$ \\ $\gamma_{t1,1}$ & 6.299 & 1.168 &$[4.010,8.588]$ \\ $\gamma_{t2,1}$ & 5.682 & 0.498 &$[4.706,6.658]$ \\ $\gamma_{t3,1}$ & 2.909 & 1.028 &$[ 0.894,4.924]$ \\ \bottomrule \end{tabular} \end{table} \begin{table} \caption{Empirically Estimated Asymptotic Covariance Matrix} \begin{tabular}{c c c c c c c c c c} \toprule & $\beta_{t1,2}$ & $\beta_{t2,1}$ & $\beta_{t2,2}$ & $\beta_{t3,1}$ & $\beta_{t3,2}$ & $\gamma_{t1,1}$ & $\gamma_{t2,1}$ & $\gamma_{t3,1}$ \\ \midrule $\beta_{t1,1}$ & -0.2339 &0.0050 & $<0.0001$ & 0.3425 & -0.0008&0.7555 & 0.0013 & -0.0090\\ $\beta_{t1,2}$ & & 0.0005 & $<0.0001$& -0.1703& 0.0001 &-0.3927 &0.0010 & 0.0274\\ $\beta_{t2,1}$ & & &-0.0145 & -0.0419& -0.0001 & -0.0246& 0.3252& -0.2186 \\ $\beta_{t2,2}$ & & & & -0.0004& $<0.0001$& 0.0002& -0.0116& 0.0001 \\ $\beta_{t3,1}$ & & & & & -0.0026&0.5415 &-0.0194 &0.1688 \\ $\beta_{t3,2}$ & & & & & & -0.0005& $<0.0001$&-0.0038 \\ $\gamma_{t1,1}$ & & & & & & &-0.0164 &-0.1673 \\ $\gamma_{t2,1}$ & & & & & & & &-0.0813 \\ \bottomrule \end{tabular} \end{table} \begin{table} \begin{threeparttable} \caption{Differences between Parameter Values} \begin{tabular}{c c c c } \toprule \multicolumn{2}{c}{Two-sided } & \multicolumn{2}{c}{One-sided} \\ \midrule $\beta_{t1,1} - \beta_{t2,1}$ & -0.467 (0.976) & $\beta_{t1,1} - \beta_{t3,1}$ & 0.353 (0.424) \\ $\gamma_{t1,1} - \gamma_{t2,1}$ &0.617 (1.282) & $\beta_{t2,1} - \beta_{t3,1}$ & 0.821 (0.980)\\ & & $\gamma_{t1,1} - \beta_{t1,1}$ & 0.981 (0.573)$*$\\ & & $\gamma_{t2,1} - \beta_{t2,1}$ & -0.104 (0.291) \\ & & $\beta_{t3,1} - \gamma_{t3,1} $ & 2.056 (1.053)$*$ \\ \bottomrule \end{tabular} \begin{tablenotes} • Standard deviation of the estimate is presented in the parenthesis, and “$*$” indicates significance at the 5% level. \end{tablenotes} \end{threeparttable} \end{table} \fi

Empirical Application

In this section, we apply the theoretical results to data from the Oregon Health Insurance Experiment finkelstein2012oregon and examine the effects on the health of different sources of health insurance. The experiment is conducted by the state of Oregon between March and September 2008. A series of lottery draws were administered to award the participants the option of enrolling in the Oregon Health Plan Standard, which is a Medicaid expansion program available for Oregon adult residents that have limited income. Follow-up surveys were sent out in several waves to record, among many variables, the participants' insurance plan and health status. finkelstein2012oregon obtain the effects of insurance coverage by using a LATE model. We apply the GLATE model can study the effect heterogeneity across different sources of insurance.

According to the data, many lottery winners did not choose to participate in the Medicaid program. Instead, they went with other insurance plans or chose not to have any health insurance. Based on this observation, we can set up the GLATE model. The instrument $Z$ is the binary lottery that determines whether an individual is selected. The covariates $X$ include the number of household members and survey waves. Given $X$, $Z$ is randomly assigned finkelstein2012oregon.\footnote{Though the covariates are discrete, the methods developed in this paper are still different from linear regressions in finkelstein2012oregon.} The treatment $T$ is the insurance plan, which contains three categories: Medicaid ($m$), non-Medicaid insurance plans ($nm$), and no health insurance ($no$). The second category includes Medicare, private plans, employer plans, and other plans. The counterfactual health plan choices under different lottery results are the variables $T_0$ and $T_1$. The unordered monotonicity condition requires that any participant who changes insurance plan due to winning the lottery does so to enroll in the Medicaid program.

The above setup is the same as Example (ref), with types. We follow the terminologies in kline2016evaluating and define the following six type sets by their counterfactual insurance plan choices:

enumerate$no$-never takers: $S \in \Sigma_{no,2} = \{ s_1 \}$, $T_{0} = T_1 = no$; • $nm$-never takers: $S \in \Sigma_{nm,2} = \{ s_2 \}$, $T_{0} = T_1 = nm$; • always takers: $S \in \Sigma_{m,2} = \{ s_3 \}$, $T_{0} = T_1 = m$; • $no$-compliers: $S \in \Sigma_{no,1} = \{ s_4 \}$, $T_0 = no$, $T_1 = m$; • $nm$-compliers: $S \in \Sigma_{nm,1} = \{ s_5 \}$, $T_0 = nm$, $T_1 = m$; • compliers: $S \in \Sigma_{m,1} = \{ s_4, s_5 \}$, $T_{0} \ne m$, $T_1 = m$.

The two groups of never takers choose not to join Medicaid regardless of the offer. Always takers manage to enroll in Medicaid even without an offer. The $no$- and $nm$- compliers switch to Medicaid from no insurance plan and other plans, respectively, upon winning the lottery. Combining these two groups gives the larger set of compliers.

Table (ref) shows the estimated probabilities of the six types.\footnote{We use the data from the 12-month survey. After taking care of the missing values, we are left with $23290$ observations. For cross-fitting, we choose $L=10$.} We can see that half of the population are $no$-never takers, who are never covered by any insurance plan. The compliers make up around one-fifth of the population. There are effectively no $nm$-compliers, meaning that the experiment does not crowd out other insurance plan choices. These findings are consistent with finkelstein2012oregon.

table[table omitted — 499 chars of source]

The outcome of interest $Y$ is health status, which is (inversely) measured by the number of days (out of past 30) when poor health impaired regular activities.\footnote{Other types of outcomes are also studied by finkelstein2012oregon, including health care utilization and financial strain. Here we only focus on health status for simplicity.} The potential outcomes are denoted by $Y_{no}$, $Y_{nm}$, and $Y_{m}$. By Theorem (ref), we can identify the distribution of $Y_{no}$ for $no$-never takers and $no$-compliers, the distribution of $Y_{nm}$ for $nm$-never takers and $nm$-compliers, and the distribution of $Y_{nm}$ for always takers and compliers. Table (ref) reports the estimated LASFs.\footnote{The LASF $\beta_{nm,1}$ is excluded because there are few $nm$-compliers as reported in Table (ref).} We can clearly see a pattern of self-selection into the treatment. For example, when there is no insurance coverage, the potential health status of $no$-compliers is worse than $no$-never takers and therefore choose to enroll in Medicaid.

table[table omitted — 487 chars of source]

Concluding Remarks

In this paper, we considered the estimation of the causal parameters, LASF and LASF-T, in the GLATE model by using the EIF. The proposed DML estimator satisfies the SPEB and can be applied in situations, such as high-dimensional settings, where Donsker properties fail. For inference, we proposed generalized AR tests robust against weak identification issues. Currently, empirical researchers use the TSLS and control the covariates linearly in models with multi-valued treatments and instruments. This linear specification does not have LATE interpretation, as pointed out by blandhol2022tsls. Therefore, we advocate using the semiparametric methods studied by this paper in those cases.

center[center omitted — 48 chars of source]