Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
79,626 characters · 13 sections · 78 citation commands
Efficient and Robust Estimation of the Generalized LATE Model
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} Causal Inference, Double Robustness, Efficient Influence Function, Multi-valued Treatment, Neyman Orthogonality, Oregon Health Insurance Experiment, Unordered Monotonicity, Weak Identification.
\spacingset{1.45}
Since the seminal works of imbens1994identification and angrist1996identification, the local average treatment effect (LATE) model has become popular for causal inference in economics. Instead of imposing homogeneity of the treatment effects as in the classical instrumental variable (IV) regression model, the LATE framework allows the treatment effect to vary across individuals. Under the monotonicity condition, the average treatment effect can be identified for a subgroup of individuals whose treatment choice complies with the change in instrument levels.
The current form of the LATE model only accepts binary treatment variables. This restriction is inconvenient in many economic settings where the treatment is multi-leveled in nature. For example, parents select different preschool programs for their kids, schools assign students to different classroom sizes, families relocate to various neighborhoods in housing experiments, and people choose different sources of health insurance. To apply the LATE model to these settings, researchers often need to redefine the treatment so that there are only two treatment levels. However, merging the treatment levels can complicate the task of program evaluation and dampen the causal interpretation of the estimates. As pointed out by kline2016evaluating, if the original treatment levels are substitutes, then there is ambiguity regarding which causal parameters are of interest. After merging the treatment levels, the heterogeneity in the treatment effect across different treatment levels is lost.
This paper addresses the above issues by generalizing the LATE framework to incorporate the potential multiplicity in treatment levels directly. We call the new framework the generalized LATE (GLATE) model. The main assumption of the GLATE model is the unordered monotonicity assumption proposed by heckman2018unordered, which is a generalization of the monotonicity assumption in the binary LATE model.\footnote{To distinguish with the GLATE model, we sometimes use the terminology “binary LATE model" to refer to the LATE model studied by imbens1994identification and abadie2003semiparametric.}
We generalize the identification results in heckman2018unordered to explicitly account for the presence of conditioning covariates, which is often important in practical settings. Recently, blandhol2022tsls point out that linear TSLS, the common way to control for covariates in empirical studies, does not bear the LATE interpretation. The only specifications that have LATE interpretations are the ones that control for covariates nonparametrically. Therefore, it is essential from the causal analysis perspective to incorporate the covariates into the GLATE framework in a nonparametric way.
The causal parameters identifiable in the GLATE model include local average structural function (LASF) and local average structural function for the treated (LASF-T). LASF is the mean potential outcome for specific subpopulations. These subpopulations are defined by their treatment choice behaviors and are generalizations of the concepts always takers, compliers, and never takers in the binary LATE model. The parameter LASF-T further restricts the subpopulation to exclude individuals who do not take up the treatment.
The paper is concerned with the econometric aspects of the GLATE model. The analysis begins by deriving efficient influence function (EIF) and semiparametric efficiency bound (SPEB) for the identified parameters. The calculation is based on the method outlined in Chapter 3 of bickel1993efficient and newey1990semiparametric. We then verify that the conditional expectation projection (CEP) estimator chen2008semiparametric, constructed directly from the identification result, achieves the SPEB and hence is semiparametric efficient. Using these results, we may efficiently estimate other important parameters of interest by the plug-in method since a standard delta-method argument preserves semiparametric efficiency.
The EIF not only facilitates the efficiency calculation but can also serve as the moment condition for estimation. This is because the EIF is mean zero by construction and is equal to the original identification result plus an adjustment term due to the presence of infinite-dimensional parameters. We show that the moment condition constructed from the EIF satisfies two related robustness properties: double robustness and Neyman orthogonality. Double robustness guarantees that the moment condition is correctly specified in a parametric setting even when some nuisance parameters are not.
The Neyman orthogonality condition means that the moment condition is insensitive to the nuisance parameters. This condition is particularly useful when the conditioning covariates are of high dimension. To further utilize this condition, we study the double/debiased machine learning (DML) estimator chernozhukov2018double in the GLATE setting. Under certain conditions regarding the convergence rate of the first-step nonparametric estimators, the DML estimator is asymptotically normally uniformly over a large class of data generating processes (DGPs).
The weak identification issue is a practical concern of the GLATE model. This is because both the treatment and instrument are multi-valued, and hence the subpopulation on which LASF and LASF-T are defined can be small in size. To deal with this issue, we propose null-restricted test statistics in one-sided and two-sided testing problems. This procedure is the generalization of the well-known Anderson-Rubin (AR) test. We show that the proposed tests are consistent and uniformly control size across a large class of DGPs, in which the size of the subpopulation mentioned above can be arbitrarily close to zero.
The paper is organized as follows. The remaining part of this section discusses the literature. Section (ref) introduces the GLATE model and the nonparametric identification results. Section (ref) calculates the EIF and SPEB. Section (ref) discusses the robustness properties of the moment condition generated by the EIF. Section (ref) proposes inference procedures under weak identification issues. Section (ref) presents the empirical application. Section (ref) concludes. The proofs for theoretical results in the main text are collected in Appendix (ref).
The GLATE model provides a way to conduct causal inference under endogeneity when the treatment is multi-valued and unordered. As mentioned above, the identification result (conditional on the covariates) is first established in heckman2018unordered by using the unordered monotonicity condition. lee2018identifying proposes another method of identification in a similar model of multi-valued treatment. Their method is concerned with continuous instruments, while the GLATE is framed in terms of discrete-valued instruments. When the treatment levels are ordered, angrist1995two derives the identification and estimation results for the causal parameter, which is a weighted average of LATEs across different treatment levels.
The literature on semiparametric efficiency in program evaluation starts with the seminal work of hahn1998role, which studies the benchmark case of estimating the average treatment effect (ATE) under unconfoundedness. For multi-level treatment, cattaneo2010efficient studies the efficient estimation of causal parameters implicitly defined through over-identified non-smooth moment conditions. In the case where unconfoundedness fails and instruments are present, frolich2007nonparametric calculates the SPEB for the LATE parameter, and hong2010semiparametric extend to the estimation of parameters implicitly defined by moment restrictions. In a more general framework encompassing missing data, chen2004semiparametric and chen2008semiparametric studies semiparametric efficiency bounds and efficient estimation of parameters defined through overidentifying moment restrictions. However, there is currently no theoretical research on semiparametric efficient estimation in models that encompasses endogeneity and unordered multiple treatment levels.
Several ways are available for calculating the EIF for semiparametric estimators, as illustrated by newey1990semiparametric and ichimura2022influence. Semiparametric efficiency calculations can be used to construct robust (Neyman orthogonal) moment conditions. This method is illustrated in newey1994asymptotic and chernozhukov2016locally. Based on the Neyman orthogonality condition, chernozhukov2018double introduces the DML method that suits high dimensional settings. This is because Donsker properties and stochastic equicontinuity conditions are no longer required in deriving the asymptotic distribution of the semiparametric estimator.
For testing the GLATE model, sun2020instrument proposes a bootstrap test which is the generalization and improvement of the test studied by kitagawa2015test in the binary LATE model.
The GLATE model has received attention in the recent empirical literature due to its ability to model multi-valued treatment. kline2016evaluating evaluate the cost-effectiveness of Head Start, classifying Head Start and other preschool programs as different treatment levels against the control group of no preschool. galindo2020empirical assesses the impact of different childcare choice in Colombia on children’s development. pinto2014noncompliance studies the neighborhood effects and voucher effects in housing allocations using data from the Moving to Opportunity experiment. Our theoretical analysis of the GLATE model presents important tools for estimation and inference that can be applied to those empirical settings.
This section describes the generalized local average treatment effect (GLATE) model, discusses identification of the local average structural function (LASF) and other parameters, and introduces the notation.
We assume a finite collection of instrument values $ \mathcal{Z} = \{z_1, \cdots, z_{N_Z}\}$ and a finite collection of treatment values $ \mathcal{T} = \{t_1, \cdots, t_{N_T}\}$, where $N_Z$ and $N_T$ are respectively the total number of instrument and treatment levels. The sets $\mathcal{T}$ and $\mathcal{Z}$ are categorical and unordered. The instrumental variable $Z$ denotes which of the $N_Z$ instrument levels is realized. The random variables $T_{z_1}, \cdots, T_{z_{N_Z}}$, each taking values in $\mathcal{T}$, denote the collection of potential treatments under each instrument status. Thus, the observed treatment level is the random variable $T = T_Z = \sum_{z \in \mathcal{Z}} \mathbf{1}\{Z=z\} T_z$. For each given treatment level $t \in \mathcal{T}$, there is a potential outcome $Y_t \in \mathcal{Y} \subset \mathbb{R}$. The observed outcome is denoted by $Y = Y_T = \sum_{t \in \mathcal{T}} \mathbf{1}\{T=t\} Y_t$. The random vector $X \in \mathcal{X} \subset \mathbb{R}^{d_X}$ contains the set of covariates. The observed data is a random sample $(Y_i,T_i,Z_i,X_i), 1 \leq i \leq n$.
The description above establishes a random sampling model where the researcher only observes one potential outcome, the one associated with the observed treatment. This implies that the sample of $Y$, observed from an individual with treatment $T=t$, comes from the conditional distribution of $Y_t$ given $T=t$ rather than from the marginal distribution of $Y_t$. In general, this fact leads to identifications issues and presents challenges for causal inference. To overcome these problems, we impose further structures on the model.
Assumption (ref) and (ref) provide the multi-valued analog of Assumption 2.1 in abadie2003semiparametric. Assumption (ref) restricts that the instrument $Z$ is independent with the potential treatments and outcomes once we condition on $X$. Assumption (ref) is the conditional version of the unordered monotonicity condition proposed by heckman2018unordered. It means that when we focus on a particular treatment level $t$ and a pair $(z,z')$ of instrument values, the binary environment should satisfy the usual monotonicity constraint in the LATE model. Specifically, the unordered monotonicity condition requires that a shift in the instrument moves all agents uniformly toward or against each possible treatment value.\footnote{As pointed out by vytlacil2002independence, the LATE monotonicity condition is a restriction across individuals on the relationship between different hypothetical treatment choices defined in terms of an instrument.}
We define the type $S$ of an individual as the vector of the potential treatments, that is,
By construction, $S$ is not observed. Assumption (ref), the unordered monotonicity condition, is essentially a restriction on $\mathcal{S} \equiv \textit{supp}(S)$, the support of $S$. Denote the elements in $\mathcal{S}$ by $s_1,\cdots,s_{N_S}$, where $N_S$ is the cardinality of $\mathcal{S}$. A convenient way to characterize $\mathcal{S}$ is by using the $N_Z \times N_S$ matrix $R \equiv (s_1,\cdots,s_{N_S})$. The matrix $R$ is referred to as the response matrix since it describes how each type of individuals' treatment choice responds to the instrument.
The role of $S$ is to assist the identification of the counterfactual outcomes by dividing the population into a finite number of groups, where identification can be achieved within specific groups. Those groups are defined as follows. For $k = 0,\cdots,N_Z$, let $\Sigma_{t,k}$ be the set of types in which the treatment level $t$ appears exactly $k$ times. That is,
where $s[i]$ denotes the $i$th element of the vector $s$. In particular, the collection $\Sigma_{t,k}, k = 0,\cdots,N_Z$ forms a partition of $\mathcal{S}$.
For individuals with type $S$ in the same type set $\Sigma_{t,k}$, their treatment response in terms of $T=t$ is in a way homogeneous. Thus, it is easier intuitively to identify the marginal distribution of the potential outcome $Y_t$ within each $\Sigma_{t,k}$. More specifically, we define the local average structural functions (LASF) and the local average structural functions for the treated (LASF-T) as follows.
Before presenting the identification results for the above two classes of parameters, we illustrate the GLATE model in the following two examples.
We introduce some matrix notations related to the type $S$. For each treatment level $t \in \mathcal{T}$, let $B_t$ be a binary matrix of the same dimension as the response matrix $R$ with each element of $B_t$ signifying whether the corresponding element in the response matrix is $t$. That is, $B_t[i,j]$, the $(i,j)$th element of $B_t$, is whether $T_{z_i}$ equals $t$ for the subpopulation $S=s_j$. Define $b_{t,k} \equiv \left(\mathbf{1}\{s_1 \in \Sigma_{t,k}\}, \cdots, \mathbf{1}\{s_{N_S} \in \Sigma_{t,k}\} \right) B_t^+,$ where $B_t^+$ is the Moore-Penrose inverse of $B_t$.
For convenience, we also need some notations regarding conditional expectations. Let
be the vector of functions that describes the conditional distribution of the instrument $Z$. For each treatment level $t \in \mathcal{T}$, let
be the vector that describes the conditional treatment probabilities given each level of the instrument. Denote
as the vector that contains the conditional outcomes for each treatment level $t$. Notice that the functions $\pi$, $P_{t}$, and $Q_t$ are all identified.
Theorem (ref) identifies $p_{t,k}$, the size of the subpopulation $\Sigma_{t,k}$, and the local structural function for that subpopulation. The only exception when the identification fails is when the type set $\Sigma_{t,0}$, in which case the individual never chooses the treatment $t$. This identification result is a modification of Theorem T-6 in heckman2018unordered that explicitly accounts for the presence of covariates $X$. Bayes rule is applied to convert the conditional result into the unconditional one. The following theorem presents the identification result for the LASF-T.
Let $\mathcal{Z}_{t,k} \subset \mathcal{Z}$ be the set of instrument values that induces the treatment level $t$ in the type set $\Sigma_{t,k}$. That is, $\mathcal{Z}_{t,k} \equiv \left\{ z_i \in \mathcal{Z} : s[i] = t \text{, for all } s \in \Sigma_{t,k} \right\}$, where $s[i]$ denotes the $i$th element of the vector $s$. Then define $\pi_{t,k} \equiv \sum_{z \in \mathcal{Z}_{t,k}} \pi_z$ as the total probability of those instrument values.
The identification results are illustrated using the two examples.
In this section, we calculate the semiparametric efficiency bound (SPEB) and propose estimators that achieve such bounds. We focus on the parameters LASF and LASF-T. In Appendix (ref), we study general parameters implicitly defined through moment restrictions.
For the rest of the paper, we assume that $Y_t, t \in \mathcal{T}$ have finite second moments. This is necessary since we are studying efficiency. Let $\iota$ denote the column vector of ones and $\zeta(Z,X,\pi)$ the diagonal matrix with the diagonal elements being $\mathbf{1}\{Z=z\}/\pi_{z}(X), z \in \mathcal{Z}$. The following theorem gives the efficient influence function (EIF) and the SPEB for the parameters identified in the preceding section.
The EIF in Theorem (ref) can be interpreted as the moment condition from the identification results modified by an adjustment term due to the presence of unknown infinite-dimensional parameters. Take $\psi^{\beta_{t,k}}$ as an example, the terms
and
are respectively the adjustment terms due to the presence of $Q_t$ and $P_t$.
From the expression of $\psi^{\beta_{t,k}}$, we can see that the SPEB would be large when $p_{t,k}$ is small. This is because $p_{t,k}$ measures the size of the subpopulation $S \in \Sigma_{t,k}$ on which the LASF is estimated. When $p_{t,k}$ is small, we run into the weak identification issue. In Section (ref), we study inference procedures that are robust against weak identification issues.
One benefit of the EIFs is that we can easily calculate the covariance matrix of different estimators. Consider an example where we are interested in two LASFs $\beta_1$ and $\beta_2$, whose EIF is given by $\psi_1$ and $\psi_2$, respectively. If the two estimators $\hat{\beta}_1$ and $\hat{\beta}_2$ are both semiparametric efficient, then their covariance matrix equals $\mathbb{E}[\psi_1\psi_2']$.
The derived SPEB helps determine whether an estimation procedure is efficient. In this section, we focus on the condition expectation projection (CEP) estimator.\footnote{The terminology “condition expectation projection” is adopted from the papers chen2008semiparametric and hong2010semiparametric, whereas hahn1998role refers to these estimators as “nonparametric imputation based estimators.”} Define
The CEP procedure first estimates $\pi_z$, $h_{Y,t,z}$, and $h_{t,z}$ by using nonparametric estimators $\hat{\pi}_z$, $\hat{h}_{Y,t,z}$, and $\hat{h}_{t,z}$ respectively. These estimators can be constructed based on series or local polynomial estimation. Then $Q_{t,z}$ and $P_{t,z}$ are estimated using $\hat{Q}_{t,z} = \hat{h}_{Y,t,z} / \hat{\pi}_z$ and $\hat{P}_{t,z} = \hat{h}_{t,z} / \hat{\pi}_z$. The vectors of estimators $\hat{Q}_{t}$ and $\hat{P}_{t}$, $\hat{\pi}$ are stacked in an obvious way. Let $\hat{\pi}_{{t,k}} = \sum_{z \in \mathcal{Z}_{t,k}}\hat{\pi}_{z}$. The CEP estimators for the structural parameters are defined by
The next proposition shows that the CEP estimators are semiparametrically efficient. The result is similar in style to hahn1998role's (hahn1998role) Proposition 4 that the low-level regularity conditions are omitted. Instead, the proposition assumes the high-level condition that the CEP estimators are asymptotically linear, which means they are asymptotically equivalent to sample averages. More formally, an estimator $\hat{\beta}$ of $\beta$ is asymptotically linear if it admits an influence function. That is, there exists an iid sequence $\psi_i$ with zero mean and finite variance such that
Since each element of the conditional expectations $h_{Y,t,z}$, $h_{t,z}$, and $\pi_z$ can be considered as coming from a binary LATE model, the regularity conditions in hong2010supplement should work with little modification.
The reason that this type of estimator is efficient is well explained in ackerberg2014asymptotic. The estimation problem here falls into their general semiparametric model, where the finite-dimensional parameter of interest is defined by unconditional moment restrictions. They show that the semiparametric two-step optimally weighted GMM estimators, the CEP estimators in this case, achieve the efficiency bound since the parameters of interest are exactly identified. Discussions related to this phenomenon can also be found in chen2018overidentification.
We next examine the efficient estimation of other policy-relevant parameters that can be derived from the parameters $\left( \beta_{t,k},\gamma_{t,k},p_{t,k},q_{t,k} \right) $. As an example, consider the type set $\Sigma_t \equiv \cup_{k=1}^{N_{Z-1}} \Sigma_{t,k}$, which is referred to as $t$-switchers. This subpopulation contains individuals who switch between $t$ and other treatments when given different levels of instruments. It is a generalization of the concept of compliers in the binary LATE framework.\footnote{Recall that switchers are also illustrated in Example (ref).} The LASF for the subpopulation $\Sigma_t$ is given by
Similarly, one can also define
which represents the LASF-T for the subpopulation of $t$-treated $t$-switchers.
For some subpopulations, a treatment effect can be identified. This point is already illustrated with Example (ref) in the discussion of the identification of the usual LATE parameter. We further illustrate this point with Example (ref).
To summarize the above examples using a general expression, let $\phi = \phi(\underline{p},\underline{q},\underline{\beta},\underline{\gamma})$ be a finite-dimensional parameter, where $\phi(\cdot)$ is a known continuously differentiable function, and $\underline{p}$ is the vector containing all identifiable $p_{t,k}$'s, that is, $\underline{p} \equiv \{p_{t,k} : t \in \mathcal{T}, 1 \leq k \leq N_Z\}$. Let $\underline{q},\underline{\beta}$, and $\underline{\gamma}$ be defined analogously. A natural estimator can be defined through the CEP estimates, $\phi(\hat{\underline{p}},\hat{\underline{q}},\hat{\underline{\beta}},\hat{\underline{\gamma}})$. The delta method can help calculate the efficiency bound of $\phi$ and show the efficiency of $\phi(\hat{\underline{p}},\hat{\underline{q}},\hat{\underline{\beta}},\hat{\underline{\gamma}})$. In fact, by Theorem 25.47 of van1998asymptotic, we immediately have the following corollary, which shows that plug-in estimators are efficient.
In the previous section, the EIF is used as a tool for computing the SPEB. In this section, we directly use the EIF as the moment condition for estimation. These moment conditions are appealing because they satisfy double robustness and local robustness --- the two topics of this section.
A word on notation: in the rest of the paper, we use a superscript $o$ to signify the true value whenever necessary. For example, when both $\pi^o$ and $\pi$ appear, the former means the true probability while the latter denotes a generic function.
We focus on the LASF $\beta_{t,k}$. The same analysis can be applied to the other parameters. To avoid notational burden in the main text, we drop the subscript $(t,k)$ in $\beta_{t,k}$, $p_{t,k}$, and $b_{t,k}$, and the subscript $t$ in $P_t$ and $Q_t$.\footnote{The full subscripts are kept in the Appendices.} It is straightforward to verify that the EIF $\psi^{\beta}$ has zero mean. However, we do not want to use $\psi^{\beta}$ itself as the estimating equation since it contains $1/p$ as a factor. To deal with this problem, we simply multiply $\psi^{\beta}$ by $p$ and define
The corresponding moment condition is
This moment condition is doubly robust, as demonstrated in the following proposition.
The above proposition divides the nonparametric nuisance parameters into two groups, $\pi$ and $(Q,P)$. The doubly robust moment condition is valid if either of these two groups of nuisance parameters is true. On the other hand, if the researcher uses parametric models for these nuisance parameters, then the structural parameter $\beta$ can be recovered provided that at least one of the working nuisance models is correctly specified. Therefore, the doubly robust moment condition is “less demanding" on the researcher's ability to devise a correctly specified model for the nuisance parameters. The double robustness result in Proposition (ref) can be seen as the GLATE extension of the existing results in the binary LATE literature tan2006regression,okui2012doubly.
The second robustness property is Neyman orthogonality. Moment conditions with this property have reduced sensitivity with respect to the nuisance parameters. Formally, Neyman orthogonality means that the moment condition has zero Gateaux derivative with respect to the nuisance parameters. The result is presented in the following proposition.
In many econometrics models, double robustness and Neyman orthogonality come in pairs. Discussions about their general relationships can be found in chernozhukov2016locally. In practice, double robustness is often used for parametric estimation, as previously explained, whereas Neyman orthogonality is used in estimation with the presence of possibly high-dimensional nuisance parameters.
Next, we apply the double/debiased machine learning (DML) method developed by chernozhukov2018double to the moment condition ((ref)). This estimation method works even when the nuisance parameter space is complex enough that the traditional assumptions, e.g., Donsker properties, are no longer valid.\footnote{In two-step semiparametric estimations, Donsker properties are usually required so that a suitable stochastic equicontinuity condition is satisfied. See, for example, Assumption 2.5 in chen2003estimation.} The implementation details are explained below.
The nuisance parameters $Q$, $P$, and $\pi$ are estimated using a cross-fitting method: Take an $L$-fold random partition of the data such that the size of each fold is $n/L$. For $l = 1, \cdots, L$, let $I_l$ denote the set of observation indices in the $l$th fold and $I^c_l = \bigcup_{l' \ne l} I_{l'}$ the set of observation indices not in the $l$th fold. Define $\check{Q}^l$, $\check{P}^l$, and $\check{\pi}^l$ to be the estimates constructed by using data from $I_l^c$. The DML estimator of $\beta$ is constructed following the moment condition ((ref)):\footnote{This is the DML2 estimator defined in chernozhukov2018double. Another estimator, the DML1 estimator, is proposed in the same paper. We do not study the DML1 estimator since it is asymptotically equivalent to DML2, and the authors generally recommend DML2.}
To conduct inference, we also need an estimate for the asymptotic variance of $\check{\beta}$, which we denote by $\sigma^2$. The asymptotic variance equals to the expectation of the squared efficient influence function: $\sigma^2 = \mathbb{E}\left[ \psi^{\beta} \right]^2 = \mathbb{E}[\psi^2]/p^2$. We first estimate $p$ by using the cross-fitting method, which is essentially given by the denominator of ((ref)):
Then the asymptotic variance can be estimated by
We want to establish the convergence results for the DML estimator uniformly over a class of data generating processes (DGPs) defined as follows. For any two constants $c_1 > c_0 >0$, let $\mathcal{P}(c_1,c_0)$ be the set of joint distributions of $(Y,T,Z,X)$ such that
The first condition excludes the case where $\beta$ is weakly identified (when $p$ can be arbitrarily close to zero). Inference under weak identification is studied in the next section. The following theorem establishes the asymptotic properties of the DML estimation procedure. In particular, the estimator achieves the SPEB.
The proof verifies the conditions of Theorem 3.1 in chernozhukov2018double. The essential restriction is on the uniform convergence rate for the estimators of the nuisance parameters. In low-dimensional settings, one can consider the local polynomial regression for estimation of the conditional expectations. Under suitable conditions hansen2008uniform,masry1996multivariate, the uniform convergence rate of the local polynomial estimators is $(\log n / n)^{2/(d_X+4)}$, which is $o(n^{-1/4})$ if $d_X \leq 3$. In high-dimensional settings, as pointed out by chernozhukov2018double, the rate $o(n^{-1/4})$ is often available for common machine learning methods under structured assumptions on the nuisance parameters.\footnote{This includes the LASSO method under sparsity of the nuisance space. See, for example, buhlmann2011statistics, belloni2011l1, and belloni2013least. However, chernozhukov2018double also indicate that to prove that machine learning methods achieve the $o(n^{-1/4})$ rate, one will eventually have to use related entropy conditions.} This means that the asymptotic normality of the DML estimator continues to hold.
Theorem (ref) can be directly used to conduct inference on $\beta$. Confidence regions can be constructed by inverting the usual $t$-tests. These confidence regions are uniformly valid since the convergence results in the above theorem hold uniformly over $\mathcal{P}$. In the next section, we explain why uniform validity is crucial when dealing with weak identification issues.
The convergence result established in Theorem (ref) is uniform over the set of DGPs with type probability $p$ bounded away from zero. However, the identification of $\beta$ would be weak in the case where $p$ can be arbitrarily close to zero. This leads to distortion of the uniform size of the test and poor asymptotic approximation in finite-sample settings. This section studies this weak identification issue and proposes an inference procedure that is robust against such a problem.
We begin with a heuristic illustration of the weak identification problem. To ease notation, define $\upsilon = \beta p$ and
After a simple calculation, we can write
In the above expression, we can interpret the estimation errors $\sqrt{n}(\check{\upsilon}-\upsilon)$ and $\sqrt{n}(\check{p}-p)$ as the noises, while the signal is the term $\sqrt{n}p$. Under the usual asymptotics where $p > 0$ is fixed, the noise terms are bounded in probability, whereas the signal term $\sqrt{n}p \rightarrow \infty$. Hence, the signal dominates the noise, and the estimator $\check{\beta}$ is consistent. However, under asymptotics with a drifting sequence $p = p_n \rightarrow 0$ and $\sqrt{n}p$ converging to a finite constant, the signal and the noise are of the same magnitude, which results in the inconsistency of $\check{\beta}$. This problem is the weak identification issue. In the weak IV literature, a common measure of identification strength is the so-called concentration parameter. In our case, the concentration parameter is given by $\sqrt{n}p$ where $\sqrt{n}p \rightarrow \infty$ corresponds to strong identification, and identification is weak when the limit of $\sqrt{n}p$ is finite.
While weak identification is a finite-sample issue, it is formalized using the asymptotic framework. However, the illustration above using asymptotics under drifting sequences is not meant to model DGPs that vary with the sample size $n$. Instead, it is a tool used to detect the lack of uniform convergence. In fact, controlling the uniform size of the test is the key to solving weak identification problems.\footnote{See, for example, imbens2004confidence, mikusheva2007uniform, and andrews2020generic.} Formally, the uniform size of a test is the large sample limit of the supremum of the rejection probability under the null hypothesis, where the supremum is taken over the nuisance parameter space. When testing a null hypothesis on $\beta$ in the GLATE model, the supremum mentioned above is taken over all values of $p>0$. That is, a desirable test should have rejection probability under the null converge to the nominal size uniformly over $p \in (0,1]$. From the previous discussion, we can see that the uniform size can not be controlled using the usual $t$-statistic $\sqrt{n}(\check{\beta} - \beta)/\check{\sigma}$. This failure of uniform convergence, however, does not conflict with Theorem (ref), where the uniform convergence of $\check{\beta}$ is established only after restricting $p$ to be bounded away from zero.
Inference procedures that are robust against weak identification can be obtained by directly imposing the null hypothesis in the construction of the test statistic. One such example is the well-known Anderson-Rubin (AR) statistic in the weak IV literature. Its idea can be generalized to the GLATE model. We first consider testing the two-sided hypothesis $H_0:\beta = \beta_0$ versus $H_1:\beta \ne \beta_0$. To control the uniform size of the test, we need the test statistic to converge uniformly on the parameter space where (1) $\beta = \beta_0$, and (2) $p$ is allowed to be arbitrarily close to zero. A null-restricted $t$-statistic can be obtained as follows. Notice that when $p>0$, $\beta = \beta_0$ is equivalent to
Its estimate can be written as
Under the null hypothesis $\beta = \beta_0$, the above estimate does not depend on the concentration parameter $\sqrt{n}p$ and consists only of the noise terms $\check{\upsilon} - \upsilon$ and $\check{p} - p$, whose uniform convergence can be established directly.
For implementation, this test statistic can be obtained as a straightforward application of the DML procedure described in the previous section to the moment condition ((ref)). As a consequence of Proposition (ref), the above moment condition satisfies the Neyman orthogonality condition regardless of the true value of $\beta$. More specifically, the null-restricted $t$-statistic is defined to be
where
The corresponding test of $H_0:\beta = \beta_0$ against $H_1:\beta \ne \beta_0$ rejects for large values of $|\check{\rho}|$.
The same methodology can be applied to testing one-sided hypothesis $H_0:\beta \leq \beta_0$ versus $H_1:\beta > \beta_0$. Under the null hypothesis, $(\beta - \beta_0)p$ is non-positive, suggesting that the test should reject for large values of $\check{\rho}$. Notice that this relies on knowing the sign of $p$ due to the GLATE model structure. This restriction on the sign of $p$ is similar to knowing the first-stage sign in the linear IV model, which is studied by andrews2017unbiased in the context of unbiased estimation.
We now define the set of DGPs that allows $p$ to be arbitrarily close to zero. For any two constants $c_1>c_0>0$, let $\mathcal{P}^{\text{WI}}(c_0,c_1)$ be the set of joint distributions of $(Y,T,Z,X)$ such that
For any $\beta' \in \mathbb{R}$, let $\mathcal{P}^{\text{WI}}_{\beta'}(c_0,c_1)$ be the subset of $\mathcal{P}^{\text{WI}}(c_0,c_1)$ in which the true value of the parameter $\beta$ is $\beta'$. In particular, $\mathcal{P}^{\text{WI}}_{\beta_0}(c_0,c_1)$ denotes the subset where the null hypothesis is true. The superscript “WI” denotes weak identification. The difference between $\mathcal{P}(c_0,c_1)$ and $\mathcal{P}^{\text{WI}}(c_0,c_1)$ is that $\mathcal{P}^{\text{WI}}(c_0,c_1)$ allows the type probability $p$ to be arbitrarily small, whereas the type probabilities in $\mathcal{P}(c_0,c_1)$ are uniformly bounded away from zero. Denote $\mathcal{N}_{\nu}$ as the $\nu$th quantile of the standard normal distribution. The following theorem establishes that the above testing procedures have uniformly correct sizes and are consistent.
In this section, we apply the theoretical results to data from the Oregon Health Insurance Experiment finkelstein2012oregon and examine the effects on the health of different sources of health insurance. The experiment is conducted by the state of Oregon between March and September 2008. A series of lottery draws were administered to award the participants the option of enrolling in the Oregon Health Plan Standard, which is a Medicaid expansion program available for Oregon adult residents that have limited income. Follow-up surveys were sent out in several waves to record, among many variables, the participants' insurance plan and health status. finkelstein2012oregon obtain the effects of insurance coverage by using a LATE model. We apply the GLATE model can study the effect heterogeneity across different sources of insurance.
According to the data, many lottery winners did not choose to participate in the Medicaid program. Instead, they went with other insurance plans or chose not to have any health insurance. Based on this observation, we can set up the GLATE model. The instrument $Z$ is the binary lottery that determines whether an individual is selected. The covariates $X$ include the number of household members and survey waves. Given $X$, $Z$ is randomly assigned finkelstein2012oregon.\footnote{Though the covariates are discrete, the methods developed in this paper are still different from linear regressions in finkelstein2012oregon.} The treatment $T$ is the insurance plan, which contains three categories: Medicaid ($m$), non-Medicaid insurance plans ($nm$), and no health insurance ($no$). The second category includes Medicare, private plans, employer plans, and other plans. The counterfactual health plan choices under different lottery results are the variables $T_0$ and $T_1$. The unordered monotonicity condition requires that any participant who changes insurance plan due to winning the lottery does so to enroll in the Medicaid program.
The above setup is the same as Example (ref), with types. We follow the terminologies in kline2016evaluating and define the following six type sets by their counterfactual insurance plan choices:
The two groups of never takers choose not to join Medicaid regardless of the offer. Always takers manage to enroll in Medicaid even without an offer. The $no$- and $nm$- compliers switch to Medicaid from no insurance plan and other plans, respectively, upon winning the lottery. Combining these two groups gives the larger set of compliers.
Table (ref) shows the estimated probabilities of the six types.\footnote{We use the data from the 12-month survey. After taking care of the missing values, we are left with $23290$ observations. For cross-fitting, we choose $L=10$.} We can see that half of the population are $no$-never takers, who are never covered by any insurance plan. The compliers make up around one-fifth of the population. There are effectively no $nm$-compliers, meaning that the experiment does not crowd out other insurance plan choices. These findings are consistent with finkelstein2012oregon.
The outcome of interest $Y$ is health status, which is (inversely) measured by the number of days (out of past 30) when poor health impaired regular activities.\footnote{Other types of outcomes are also studied by finkelstein2012oregon, including health care utilization and financial strain. Here we only focus on health status for simplicity.} The potential outcomes are denoted by $Y_{no}$, $Y_{nm}$, and $Y_{m}$. By Theorem (ref), we can identify the distribution of $Y_{no}$ for $no$-never takers and $no$-compliers, the distribution of $Y_{nm}$ for $nm$-never takers and $nm$-compliers, and the distribution of $Y_{nm}$ for always takers and compliers. Table (ref) reports the estimated LASFs.\footnote{The LASF $\beta_{nm,1}$ is excluded because there are few $nm$-compliers as reported in Table (ref).} We can clearly see a pattern of self-selection into the treatment. For example, when there is no insurance coverage, the potential health status of $no$-compliers is worse than $no$-never takers and therefore choose to enroll in Medicaid.
In this paper, we considered the estimation of the causal parameters, LASF and LASF-T, in the GLATE model by using the EIF. The proposed DML estimator satisfies the SPEB and can be applied in situations, such as high-dimensional settings, where Donsker properties fail. For inference, we proposed generalized AR tests robust against weak identification issues. Currently, empirical researchers use the TSLS and control the covariates linearly in models with multi-valued treatments and instruments. This linear specification does not have LATE interpretation, as pointed out by blandhol2022tsls. Therefore, we advocate using the semiparametric methods studied by this paper in those cases.