Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
96,902 characters · 12 sections · 56 citation commands
Treatment Effect Estimation with Noisy Conditioning Variables
In observational studies, controlling for confounding factors is crucial to identify causal effects of interest. The main challenge is that measuring these confounding elements are often difficult, if not impossible. A widely used approach to address this issue is to use proxy variables in place of the unobserved characteristics. For instance, to account for differences in unobserved worker's ability, researchers routinely include test scores in their regression controls. However, these measurements are generally contaminated with noise. In simple linear regression settings, it is well-known that using proxies as regressors induces attenuation bias. In nonparametric settings, Battistin-Chesher_2014_JoE showed that average treatment effects are not identified under covariate measurement errors and that the direction of the bias cannot be determined a priori. Therefore, simply controlling for proxy variables does not lead to reliable inference on causal effects. In this paper, I provide a novel identification strategy for causal effects when important control variables are imperfectly measured. This new approach has a wide range of applications since coarse measurements are prevalent in practice Gillen-Snowberg-Yariv_2019_JPE.
In its simplest form, the identification problem of interest is captured in the model
where $D$ is the treatment of interest and $X$ is a proxy variable for the unobserved confounding factor $X^*$. If $X^*$ were observed, one would estimate the equation $Y=\beta_0 + \beta_1 D+\beta_2X^*+\epsilon$. In practice, however, one regresses $Y$ on $D$ and $X$, using the error-ridden variable in place of $X^*$. The regression estimate using $X$ suffers from measurement error bias, and the treatment effect $\beta_1$ is not consistently estimated. A textbook solution to this problem is to use a repeated measurement as an instrumental variable (IV) for $X$. Denoting the second measurement of $X^*$ by $Z$, the two-stage least squares (2SLS) method is equivalent to estimating
where $\mathbb{E}[X|D,Z]$ is used as it is the fitted value in the first-stage equation. Here $\mathbb{E}[X|D,Z]$ plays the role of a control function HeckmanRobb1985; see MATZKIN2007 for a definition of control functions.\footnote{Here I take the perspective of control function rather than IV because the control function approach extends to non-linear settings naturally whereas IV counterparts face some difficulties Blundell-Powell_2003. For instance, nonparametric IV methods do not identify causal effects unless the unobserved heterogeneity is additively separable.}
The most novel contribution of this paper is to extend the above control function method to non-linear settings, including general nonparametric models. Specifically, I show that under appropriate assumptions, the random vector $V=(\mathbb{E}[X|D,Z],\dots,\mathbb{E}[X^k|D,Z])'$ for some $k\in\mathbb{N}$ is a valid control function in the sense that the treatment variable becomes conditionally independent of unobserved confounding factors given $V$. This is a natural extension of (ref) since the 2SLS method controls for only the first conditional moment of proxy variables linearly whereas my identification approach also controls for the higher conditional moments in a nonparametric manner. Accommodating non-linear models is especially important for treatment effect estimation because the linear model (ref) imposes constant treatment effects, and as well-documented in the literature, ignoring treatment effect heterogeneity may produce misleading estimates.
The control function approach may be particularly appealing to applied researchers for its simplicity: the estimation procedure boils down to running regressions and computing sample averages. To facilitate implementation of the new control function method, I discuss semiparametric/flexible parametric modelling of the outcome equation, leveraging existing results in the control function literature. In addition, I develop a Lasso-based estimation procedure to flexibly choose regression specifications, building on Chernozhukov-Newey-Singh_2022_ecma. With cross-validation, it is easy to choose Lasso tuning parameters, and I provide a closed-form variance estimator. I characterize a set of sufficient conditions for $\sqrt{n}$-consistency and asymptotic normality of the proposed estimator as well as consistency of the variance estimator.
To demonstrate the empirical relevance of the new results, I apply the control function method to estimating causal effects of attending selective college on post-college earnings, building on the empirical framework of Dale-Krueger_2002. Their identification strategy controls for the set of colleges to which students applied and got admitted, and it is considered one of most credible identification strategies in non-experimental contexts Chettyetal2020. Yet, the method requires detailed information about college application processes, and it is not applicable in many settings due to data limitations. In my empirical application, I follow the main idea of Dale and Krueger but use the control function approach to overcome the data issue. Namely, I observe only partial information about college applications and use the noisy measures to construct control functions. My empirical results indicate that ignoring measurement errors in control variables lead to non-trivial differences in causal effect estimates.
\paragraph{Relation to existing literature} The main alternatives to my control function approach are the integral equation approach of Deaner2021,Miao-Geng-TchetgenTchetgen and the operator diagonalization method of HuSchennach2008.
My identifying assumptions greatly overlap with those of the integral equation approach, yet there is one important difference: the integral equation approach depends on a difficult-to-verify high-level condition, which I do not impose. My method may be more appealing because of the difficulty in verifying this high-level condition in general. As discussed in the appendix, with a similar high-level condition, my method yields the same identification result as the integral equation approach, and thus, this high-level condition is the key difference between the two approaches. Also, I exhibit a simple example where the key high-level condition of the integral equation approach fails whereas my identifying assumptions hold.
HuSchennach2008 pioneered the use of completeness conditions in nonparametric measurement error models and their methods have been successfully applied beyond measurement error contexts Arellano-Blundell-Bobhomme_2017,Bohhomme-Lamadon-Manresa_2018,Sasaki2015. Building on their idea, I also use a completeness condition to formalize the notion of a proxy variable. HuSchennach2008 established powerful nonparametric identification results, and although the identifying assumptions are non-nested, their approach identifies a larger class of parameters than my approach e.g.,\ the distribution of unobserved confounding factors. On the other hand, my method has practical advantages because it can relax some of the identifying assumptions by exploiting additional structures on the outcome equation. Such relaxation of identifying assumptions does not seem to hold for the operator diagonalization technique (i.e.,\ identifying assumptions for a specific model are not weaker than for a general model). In addition, my control function method is easy to implement as estimation boils down to running regressions and computing sample averages.
This paper also contributes to the extensive literature on control functions Blundell-Powell_2003,MATZKIN2007,Wooldridge2015JHR. Many existing results construct a control variable from IV, whereas I use proxy variables. Since the identifying assumptions are quite different, the new result in this paper complements existing results by expanding the scope of applicability of control function methods. Many existing approaches use the invertibility of the first-stage equation to construct a control function, and this feature excludes the case of multi-dimensional unobserved heterogeneity with a scalar treatment variable. Provided that appropriate proxies are available, my method allows for multiple unobserved heterogeneity and does not require the invertibility of the first-stage equation. This aspect is practically relevant as imposing the scalar restriction on unobserved confounding factors is not appealing in some settings. A common drawback of the control function method is that it has difficulty handling discrete endogenous variables. Although my method shares this feature to an extent, I discuss how additional assumptions help achieve the identification of causal parameters with a discrete endogenous variable.
As discussed below, my identification strategy is closely related to the method of AltonjiMatzkin2005. In settings with group structure, ArkhangelskyImbens2019 provided a framework where the key condition of Altonji and Matzkin (the exchangeability condition) follows from model primitives. Similarly, I provide a framework where the exchangeability-type condition holds, but my model differs in not having an explicit group structure.
I motivated my identification strategy as a generalization of 2SLS estimation using repeated measurements. While existing literature on non-linear models with measurement errors is extensive SCHENNACH2020, my approach is unique in the use of what I call excluded variable, which includes repeated measurements and IVs but also other types of variables. In addition, my approach is distinct from deconvolution methods.
\paragraph{Roadmap} In the next section, I describe the econometric model and discuss the identification results. Semiparametric and flexible parametric estimation methods are developed in Section (ref), and I apply the results of this paper to estimating causal effects of attending selective college on post-college earnings in Section (ref). Section (ref) concludes.
I employ the potential outcome notation. $\{Y(d):d\in\mathcal{D}\}$ denotes the set of potential outcomes, $\mathcal{D}$ is the set of possible treatment levels, $D$ is the realized treatment level, and $X^*$ represents unobserved confounding factors. The key identification challenge I focus on is endogeneity due to statistical dependence between $D$ and $X^*$. If $D$ and $X^*$ are independent (e.g.,\ in a randomized control trial), simple comparison of means for different treatment levels suffices for identification of treatment effects. For nonparametric outcome models, I focus on continuous treatment, but discrete cases are also discussed. $X$ denotes a proxy variable for $X^*$, and $Z$ is what I call excluded variable. A leading example of $Z$ is a repeated measurement as in the linear 2SLS case of (ref). Yet, there are other types of variables that can be used in the place of $Z$. Since $Z$ does not need to be a repeated measurement, I refer to $Z$ as excluded variable to emphasize this aspect. I will discuss examples of excluded variables in Section (ref).
I note that the dimension of proxy variable $X$ should be at least as large as that of $X^*$ in general (precise assumptions below). An important consequence of this feature is that to allow for rich unobserved heterogeneity (e.g.,\ multi-dimensional, continuous $X^*$), the proxy variable has to exhibit sufficient variation (multi-dimensional, continuous $X$). Conversely, if one is willing to restrict variation in $X^*$, a discrete proxy $X$ suffices for my identification results below. Also, covariates without measurement errors can be introduced. For brevity, I do not consider covariates explicitly. The analysis below goes through by conditioning on the covariates throughout.
To ground discussion on concrete terms, I use the empirical setting of Dale-Krueger_2002 (DK henceforth) as a running example.
The causal parameter I focus on is the average structural function (ASF) Blundell-Powell_2003:
In DK's setting, $\vartheta(d)$ is the earnings function with respect to college selectivity measure $d$. As one varies school-level mean SAT score $d$, the distribution of unobserved ability is held constant, and thus, the change in $\vartheta(d)$ represents the ceteris-paribus effect of college selectivity on future earnings. In different empirical settings with a binary treatment $\mathcal{D}=\{0,1\}$, $\vartheta(1)-\vartheta(0)$ is the average treatment effect. As studied in the literature, other distributional features of the potential outcomes may be of interest e.g.,\ the quantile and distributional structural functions. Using my approach, these objects are identifiable under the same assumptions that identify the ASF. As the identification arguments are similar, I focus on the ASF.
I denote statistical independence by $\protect\mathpalette{\protect\independenT}{\perp}$. For a random vector $A$, I write $F_A,f_A$ for the distribution function and density (with respect to some measure) of $A$, respectively. Let $\operatorname*{supp}(A)$ be the support of $A$. Following are the identifying assumptions I impose on econometric models.
To interpret Assumption (ref), note that it is implied by $Y(d)\protect\mathpalette{\protect\independenT}{\perp} D|X^*$ and $Y(d)\protect\mathpalette{\protect\independenT}{\perp} Z|D,X^*$. The first part states that once conditional on $X^*$, the treatment variable becomes exogenous i.e.,\ $X^*$ represents the confounding factors. The other condition is that $Z$ satisfies an exclusion restriction in the sense that $Z$ does not affect the outcome given the treatment and confounding factors. This exclusion restriction on $Z$ is distinct from the standard IV exclusion restriction because the restriction needs to hold only conditional on confounding factors $X^*$. In my setting, excluded variables $Z$ are allowed to be correlated with confounding factors, which is not the case for IVs.
Assumption (ref) states that given the “correctly measured” variable, its noisy measurement is independent of other variables. Intuitively, it requires that measurement errors in $X$ be orthogonal to treatment and excluded variables conditional on confounding factors. This type of restriction is common in the measurement error literature HuSchennach2008. As a visual aid to understand Assumptions (ref)-(ref), Figure (ref) provides directed acyclical graphs that are compatible with the imposed conditional independence restrictions.
To assess the plausibility of these assumptions in empirical settings, it is important to specify some elements of the treatment assignment mechanism as well as the outcome determination process. I illustrate this point using DK's empirical setting.
In this model of college admission, Assumption (ref) is plausible. Holding $X^*$ constant, the remaining variation in the college admission process is due to $\epsilon$'s, which represents a sum of idiosyncratic shocks. Therefore, given $X^*$, $D$ becomes orthogonal to the unobserved term $\varepsilon$ in the outcome equation (ref), impliying $Y(d)\protect\mathpalette{\protect\independenT}{\perp} D|X^*$. To assess the other condition $Y(d)\protect\mathpalette{\protect\independenT}{\perp} Z|D,X^*$, note $Z$ corresponds to an SAT score. Since SAT is a measurement of academic ability, once conditioning on the college selectivity and academic ability, SAT scores in high school are unlikely to have predictive power on post-college earnings, satisfying the exclusion restriction.
To evaluate Assumption (ref), we think of test scores as a noisy measurement of academic ability. One way to express this is that $X=X^*+U_x,Z=X^*+U_z$ where $U_x,U_z$ denote measurement errors. Then, Assumption (ref) requires that $U_x$, residual variation in test scores, is independent of $(D,U_z)$ conditional on $X^*$. The requirement that $U_x\protect\mathpalette{\protect\independenT}{\perp} U_z|X^*$ is plausible as test scores may fluctuate around its mean (ability) for idiosyncratic reasons e.g.,\ whether students were feeling sick on the day of test taking. Regarding $U_x\protect\mathpalette{\protect\independenT}{\perp} D|X^*$, in my empirical application, academic tests were administered by an agency of the U.S.\ Department of Education solely for the survey purpose, and they were not used for college admission. Thus, once conditional on the ability $X^*$, these test scores are likely to be independent of college admission processes. Note that when evaluating Assumptions (ref)-(ref), the functional forms in DK's model did not play any role. The above reasoning is equally valid for nonparametric models.
Assumption (ref) formalizes the idea that the proxy variable $X$ has a strong relationship with $X^*$. It is the $L^2$-completeness of the conditional distributions of $X^*$ given $X$. Completeness conditions can be thought of as a generalization of the IV rank condition in linear models to nonparametric settings NeweyPowell2003ECMA. Since this is a rank condition, the dimension of $X$ should be at least as large as that of $X^*$ in general. In the literature, there are several known sufficient conditions for completeness Andrews2017JoE,DHaultfoeuille2011ET,Hu-Schennach-Shiu2017. For instance, if researchers are willing to impose the measurement error structure such as $X=\chi(X^*+\eta)$ where $\chi$ is invertible and $X^*\protect\mathpalette{\protect\independenT}{\perp} \eta$, then primitive sufficient conditions for completeness exist (see Lemma 4 in the appendix). In a panel data setting, Wilhelm2015 discussed justifications for completeness assumptions using past observations as proxies. In my empirical application, test scores are a strong indicator of academic ability, and the rank condition seems reasonable.
A key aspect of Assumptions (ref)-(ref) is that $X^*$ includes important unobserved confounding factors and that we observe a proxy variable for each element of $X^*$. In DK's setting, one might argue that colleges look at not only academic ability but also some personality trait (e.g.,\ “grit”) that is reflected in essays and recommendation letters. If the non-cognitive skill also affects future earnings, then we need to modify the identifying assumptions by including the non-cognitive ability as part of $X^*$. What this means in practice is that researchers need to observe a proxy variable for the non-cognitive skill in addition to a proxy for academic ability. In my empirical application, I use survey questions intended to measure psychological attributes as a noisy measurement of non-cognitive skills. Since I use two-dimensional proxy variables for the two-dimensional ability, Assumptions (ref)-(ref) continute to hold even if the non-cognitive skill is important part of the selection mechanism.
Assumption (ref) is a collection of regularity conditions. It imposes that outcome has a finite second moment, some ratios of densities are bounded, and some functions of the data generating process are continuous. These requirements are standard and mild. The formulation handles discrete and continuous treatment variables in a unified framework.
Assumption (ref) is a common support condition for the control function. Theorem (ref) below states that under Assumptions (ref)-(ref), the stochastic process $\mathfrak{V}=\{f_{X|DZ}(x|D,Z):x\in\mathcal{X}\}$ is a valid control function in that the potential outcome is conditionally independent of the treatment variable given $\mathfrak{V}$. As well-known in the control function literature, a common support condition is crucial to identify the ASF with a general nonparametric outcome equation Blundell-Powell_2003,ImbensNewey2009. See Assumption 2 of ImbensNewey2009 and their discussion. Informally speaking, to identify causal effects, one needs to be able to vary $D$ while holding constant $\mathfrak{V}$. Since $\mathfrak{V}$ is a function of $D$, this is possible only if the excluded variable $Z$ can “undo” the effect of $D$ on $\mathfrak{V}$, which is the content of Assumption (ref).
Generally, the common support condition requires the large support of $Z$ and some structure on the treatment equation. To illustrate this point, consider the following example:
Then, Assumption (ref) holds with $\phi_{\tilde{d}}(d,z)= \tilde{d}-d+z$ if the support of $Z$ is the entire real line. This claim holds because there exists a bijection between $f_{X|DZ}(\cdot|d,z)$ and $f_{X^*|DZ}(\cdot|d,z)$ by Assumptions (ref)-(ref). The conditional distribution of $X^*$ is determined by the value $D-Z$, and $Z$ having a large support ensures that the effect of $D$ on the conditional distribution on $X^*$ can be undone by a change in $Z$. This example resembles the type of restrictions imposed by ImbensNewey2009. Although my setting shares common elements with the existing control function literature, a key distinction is that my approach accommodates multi-dimensional unobserved heterogeneity in the treatment equation provided that appropriate proxy variables exist. I elaborate on this point in Section (ref).
The common support condition may be more stringent for a discrete treatment variable. With a binary treatment, Assumption (ref) translates to $0<\mathrm{Pr}[D=1|\mathfrak{V}]<1$ with probability one, the well-known overlap condition. This overlap assumption may fail to hold if the treatment assignment follows a threshold-crossing model $D=\mathbbm{1}\{Z'\delta \geq X^*\}$. An intuitive explanation for this failure is that the conditional distributions of $X^*$ differ considerably across the treated and control groups; the conditional support of $X^*$ for $D=1$ is $(-\infty,Z'\delta]$ while the conditional support for $D=0$ is $(Z'\delta,\infty)$. The difficulty in handling discrete endogenous variables is not specific to my approach but generally shared by control function methods Wooldridge2015JHR. Fortunately, the conditional independence result below holds without the overlap assumption, and one can replace Assumption (ref) with alternative restrictions to achieve the identification of causal effects. One possibility is to rely on identification-at-infinity type arguments. Another approach is to impose additional structures on the outcome equation, which I discuss in Section (ref).
The following is the main identification result of this paper. A formal proof is presented in the appendix, and the main ideas are sketched in Section (ref).
The statement above focuses on the ASF, but as standard in the literature, other parameters of interest such as the quantile structural function can be identified using the same argument. Note that the conditional independence result $Y(d)\protect\mathpalette{\protect\independenT}{\perp} D|\mathfrak{V}$ holds without Assumption (ref), and it is possible to replace the common support condition with restrictions on the outcome equation as discussed in the next section.
In devising an estimation method based on Theorem (ref), the potentially high dimension of $\mathfrak{V}$ appears to pose a challenge. However, the sigma-field generated by $\mathfrak{V}$ is coarser than the sigma-filed generated by $(D,Z)$. Thus, intuitively, the effective dimension of $\mathfrak{V}$ is at most the dimension of $(D,Z)$ and potentially smaller. To make a heuristic argument, suppose that a single-index restriction holds for the conditional distribution of $X$ i.e.,\ $f_{X|DZ}(x|d,z)= f(x,d'\gamma+z'\delta)$ for some fixed function $f$ and vectors $\gamma,\delta$. If $f$ is continuously differentiable in the second argument and $\partial f(x_1,u)/\partial u\neq 0$ for some $x_1$, then there is a bijection $T:f_{X|DZ}(x_1|d,z)\rightarrow d'\gamma +z'\delta$. In turn, we have a representation $f_{X|DZ}(x|D,Z) = f(x,T(f_{X|DZ}(x_1|D,Z)))$ for all $x\in \mathcal{X}$. Thus, conditioning on $f_{X|DZ}(x_1|D,Z)$ is equivalent to conditioning on $\mathfrak{V}$. Generalizing this argument to multi-index cases can be done (locally) via the inverse function theorem.
To formalize the above idea, I impose the following conditions. For simplicity, I assume that all elements of $(D,Z)$ are continuous, but when there are discrete elements, I fix values of discrete elements and consider each probability event separately.
For the last part of the assumption, the maximal set of points is finite as the derivative is a $l\times (\mathrm{dim}(D)+\mathrm{dim}(Z))$ matrix and its rank cannot be larger than $\mathrm{dim}(D)+\mathrm{dim}(Z)$. As the single-index example above indicates, the maximal set may be a singleton, which formalizes the idea that $\mathfrak{V}$ has a dimension smaller than $(D,Z)$. The following theorem shows that the sigma-field generated by $\mathfrak{V}$ is equivalent to the sigma-field of a finite-dimensional random vector.
This theorem operationalizes the control function result of Theorem (ref) by providing a method to compute conditional expectations given $\mathfrak{V}$. Elements of $V$ have the form $\int g(x) f_{X|DZ}(x|D,Z)d\lambda(x)$ where $g$ is some $X$-measurable function. In Theorem (ref), I take $g(\cdot) = \mathbbm{1}\{\cdot\leq x\}$ for some $x$. This specific choice is not essential, and there are many other possibilities. For example, one can use monomial transformation i.e.,\ $g(x) = x^j$ and have $V = (\mathbb{E}[X|D,Z],\dots, \mathbb{E}[X^k|D,Z])'$. In the introduction, I used this choice to preview the identification result. In theory, any choice of $g$'s is valid provided that continuous differentiability and the maximal rank condition as in Assumption (ref) hold. In the sequel, let
be a choice of transformations of $x$ and set $V=\mathbb{E}[B(X)|D,Z]$. In the empirical application, I use $g_{\ell}(x)=x^{\ell}$. Ex ante, the dimension $k$ is not known to researchers, and a practical approach is to include a sufficient number of elements in the control functions and use a model selection method to choose relevant regressors. I develop this method formally in Section (ref) where I use Lasso to select relevant control functions.
Another consequence of Theorem (ref) is that the estimation problem is well-posed, as opposed to ill-posed, under regularity conditions. Well-posedness refers to the situation where estimand of interest can be represented as a unique solution to some estimation equation and the solution varies continuously as some feature of the underlying data generating process changes CarrascoFlorensRenault2007. If an estimation problem is not well-posed, it is ill-posed. A well-known example of ill-posedness is nonparametric IV estimation: the parameter of interest (a structural function) can be recovered by computing an inverse of an integral operator, but this inverse is known to be discontinuous, leading to ill-posedeness. Ill-posedness of estimation problems may lead to slower convergence rates, and obtaining reliable estimates might be difficult unless the sample size is large. In my identification argument, I use a completeness condition, which guarantees a unique solution of an integral operator, but I do not need to compute the inverse of such integral operator. Instead, I construct a control function $V$. As Theorem (ref) shows that this control function is finite-dimensional, it fits into the framework of HahnRidder2013 and they provided sufficient conditions under which causal parameters are continuous (even differentiable) in the control function.
The assumption of compact support of $(D,Z)$ may be at odds with Assumption (ref) as the common support condition often requires a large support of $Z$. In the empirical application below, I replace Assumption (ref) with restrictions on the outcome equation to maintain the compact support assumption. Yet, it is possible to relax the compact support condition while retaining the conclusion of Theorem (ref). An alternative assumption to the compact support of $(D,Z)$ is that the conditional proxy distribution satisfies index restrictions
for some fixed (but unknown) functions $F,\iota_1,\dots,\iota_k$ where $\iota_l$ is real-valued for $l=1,\dots,k$, $\iota_l(D,Z)$ is continuously distributed and has a compact support for each $l=1,\dots, k$, and $F$ is continuously differentiable with respect to values of $(\iota_1,\dots,\iota_k)$. Here the dimension $k$ is unknown. Then, with the rank condition as in Assumption (ref), a finite-dimensional representation of the control function holds. Note that compact support of $\iota_l(D,Z)$ does not contradict with Assumption (ref).
In the next section, I discuss how additional structures on the outcome equation can replace the common support condition. As demonstrated below, this approach maintains quite flexible specifications of econometric models while relaxing some of the identifying assumptions.
The existing work on control function methods has recognized that the common support condition may fail in some empirical applications Chernozhukov-etal2019,Florens-Heckman-Meghir-Vytlacil_2008. Building on the available results, I consider additional structures on the outcome equation, which can substantially relax the support condition.
\paragraph{Random coefficient model} I consider a version of the random coefficient model
with $D\in\mathbb{R}$. As discussed above, this model allows for heterogeneous treatment effects. The causal parameter of interest I consider is the average of the random slope on the treatment variable:
This parameter represents the average marginal effect of the treatment $D$. Under Assumptions (ref)-(ref) and (ref), the model (ref) yields
where $\mu_1,\mu_2$ are unknown functions. NeweyStouli2020 showed that if the conditional variance of $D$ given $V=v$ is positive for almost all $v$, then $\theta_0$ is identified under regularity conditions. This new identifying condition is substantially weaker than the common support condition: while the common support condition demands that $\operatorname*{supp}(D|V)$ equals $\operatorname*{supp}(D)$ almost surely, the new condition only requires $\operatorname*{supp}(D|V)$ has more than one point almost surely. Therefore, with the random coefficient specification (ref), an excluded variable $Z$ may have discrete variation and one can still achieve the identification of the causal parameter $\theta_0$.\footnote{Let $\nu(D,Z)=\mathbb{E}[B(X)|D,Z]=V$. For $v\in\operatorname*{supp}(V)$, if there exist two distinct points $(d_1,z_1),(d_2,z_2)\in\operatorname*{supp}(D,Z)$ such that $v=\nu(d_1,z_1)=\nu(d_2,z_2)$, then the conditional variance of $D$ given $V=v$ is positive.}
With a binary treatment $D$, the model (ref) places few restrictions on the outcome determination process, and one needs to impose additional structures to relax the overlap condition. One possibility is to impose that $\mu_1(v),\mu_2(v)$ are real analytic and the conditional support of $V$ given $D$ contains an open set. This identification argument was used by Arellano-Bonhomme_2017. A nice feature of this approach is that the assumptions encompass the cases where $\mu_1(v),\mu_2(v)$ are polynomial functions of $v$, a widely used specification in empirical studies.
\paragraph{Binary outcome} The random coefficient model above is natural for a continuous outcome variable but may not be appropriate for limited dependent variables. For binary outcomes, Blundell-Powell_2004 considered flexible yet tractable outcome models. Suppose that $D\in\mathbb{R}^{\dim(D)}$, $\dim(D)\geq 2$, has at least one element that is continuously distributed and that
where $\gamma_0$ is a vector of coefficient parameters. Here, one can think of $D$ containing additional covaraites as well as treatment variables of interest. Under Assumptions (ref)-(ref) and (ref), the above model implies
where $V=\mathbb{E}[B(X)|D,Z]$ and $F_{\varepsilon|V}(\cdot|v)$ is the conditional CDF of $\varepsilon$ given $V=v$. Existing studies have established sufficient conditions for the identification of coefficient parameters and the ASF. For instance, Proposition 4.1 of Jochmans2013 states that if the first element of $\gamma_0$ is non-zero and the first element of $D$ has everywhere positive Lebesgue density conditional on $(D_2,V)$, where $D_2\in\mathbb{R}^{\dim(D)-1}$ is the remaining elements of $D$, and the conditional support of $D$ given $V$ is not contained in a proper linear subspace of $\mathbb{R}^{\dim(D)}$, then the coefficient vector $\gamma_0$ is identified up to scale. With $\gamma_0$ identified, the ASF can be computed by averaging $\mathbb{E}[Y|D,V] = F_{\varepsilon|V}(u|V)$ over $V$ fixing $u=D'\gamma_0$. Here, the index structure mitigates the possible curese of dimensionality.
I should note that with a binary endogenous variable following a threshold crossing model, the full-rank condition on $D$ given $V$ may fail and additional structures on $F_{\varepsilon|V}(e|v)$ may be necessary to achieve the identification of the coefficient vector. This aspect is related to the results of Blundell-Powell_2004, where they assumed that endogenous variables have a continuous distribution.
\paragraph{Flexible parametric model} Let $q(d,v)$ be a vector of some transformations of $(d,v)$ e.g.,\ $q(d,v) = (1,d,v,d*v)'$. If desired, higher-order polynomial transformations may be included in $q(d,v)$. I postulate that the outcome equation satisfy
where $\Lambda$ is a known, strictly monotonic link function and $\gamma_0$ is the parameter to be estimated. This model encompasses a parametric version of the models considered above. For the random coefficient model, suppose
where $p(v)$ is a vector of transformations of $v$ and the conditional independence $\varepsilon_l\protect\mathpalette{\protect\independenT}{\perp} D|V$, $l=1,2$ follows under Assumptions (ref)-(ref) and (ref). Relative to the above semiparametric model, the additional restriction here is $\mathbb{E}[\varepsilon_l|V]=p(V)'\gamma_l$ for $l=1,2$. This random coefficient model fits into (ref), where $\Lambda$ equals the identity function, $q(D,V) = (p(V)',Dp(V)')'$, and $\gamma_0=(\gamma_1',\gamma_2')'$.
For the case of a binary outcome, one may consider the outcome equation
and using Theorem (ref), $\mathbb{E}[Y|D,V] = F(\gamma_1'D |V)$ where $F$ is the conditional distribution of $\varepsilon$ given $V$. For tractability, we may assume that the conditional distribution of $\varepsilon$ given $V$ is a location family i.e.,\ $F(\,\cdot\,|V) = \Lambda(\,\cdot\,-p(V)'\gamma_2)$ and specify $\Lambda$ as the normal/logit CDF, which gives rise to the specification (ref) with $q(D,V)=(D',p(V)')'$ and $\gamma_0=(\gamma_1,\gamma_2')'$.
Under (ref), identification of the ASF holds if the parameter $\gamma_0$ is identified. In turn, the coefficients $\gamma_0$ is identified if the matrix $\mathbb{E}[q(D,V)q(D,V)']$ is non-singular. This full-column rank condition is weaker than the support invariance condition, and Chernozhukov-etal2019 and NeweyStouli2020 provided sufficient conditions for the non-singularity of $\mathbb{E}[q(D,V)q(D,V)']$.
To see what types of restrictions imply (ref), consider the binary outcome model
Suppose that $F_{\varepsilon|X^*}(e|x^*)=F(e+\delta'x^*)$ for some function $F$ and coefficient vector $\delta\in\mathbb{R}^{\dim(X^*)}$ and that $f_{X^*|DZ}(x^*|d,z) = f(x^*-\iota(d,z))$ for some fixed functions $f,\iota$. Also, $X=X^*+U_x$ with $U_x\protect\mathpalette{\protect\independenT}{\perp} (X^*,D,Z)$ and $\mathbb{E}[U]=0$. Then, $V=\mathbb{E}[X|D,Z]=\iota(D,Z)$ and (ref) holds with
With $\dim(X^*)=1$, if $F$ and $f$ are standard normal CDF and density respectively, then $\Lambda(y) = F(y/\sqrt{1+\delta^2})$, implying a probit model. The non-singularity of $\mathbb{E}[q(D,V)q(D,V)']$ holds in this case if for each $d\in\mathcal{D}$, there exist $z_1,z_2$ in $\mathcal{Z}_d$ such that $\iota(d,z_1)\neq\iota(d,z_2)$.
As discussed above, a repeated measurement is a leading example for excluded variables $Z$. Yet, there exist other types of variables that qualify as excluded variables. For $Z$ to be an excluded variable, it needs to satisfy the exclusion restriction Assumption (ref). In addition, excluded variables should be correlated with $X^*$ and/or $D$ to help satisfy Assumption (ref) in the nonparametric model or appropriate full-rank conditions for semiparametric/parametric models.
One example of excluded variables is an IV for the treatment variable. Suppose $D=h(Z,\eta)$ where $h$ is a non-stochastic function and $\eta$ is some unobserved heterogeneity. By the standard IV exclusion restriction, $(Y(d),X^*,\eta)\protect\mathpalette{\protect\independenT}{\perp} Z$ is plausible, and we may additionally impose $Y(d)\protect\mathpalette{\protect\independenT}{\perp} \eta|X^*$. Then, $Y(d)\protect\mathpalette{\protect\independenT}{\perp} (D,Z)|X^*$ (Assumption (ref)) follows. If $h(z,\eta)$ is non-trivial function in $z$, $Z$ is correlated with $D$, which helps to satisfy the support condition.
In DK's setting, one may use educational attainment of worker's parents as excluded variables. The exclusion restriction (i.e.,\ $Y(d)\protect\mathpalette{\protect\independenT}{\perp} Z|D,X^*$) is plausible as potential wage levels are unlikely to be affected by parents' education level once you control for ability and own education as well as other observed characteristics. The relevance condition (i.e.,\ $Z$ correlated with $X^*$ and/or $D$) is also plausible because parents' education is likely to influence the probability of college attendance (e.g.,\ having college-graduate parents makes going to college a norm). Note that educational attainment of worker's parents may not satisfy the IV exclusion restriction (i.e.,\ $(Y(d),X^*)\protect\mathpalette{\protect\independenT}{\perp} Z$) as it is potentially correlated with worker's unobserved ability through human capital formation. This example demonstrates that there exist empirically relevant excluded variables other than repeated measurements and IVs.
\paragraph{Integral equation approach} My identifying assumptions are closely related to those used by Deaner2021,Miao-Geng-TchetgenTchetgen. Both my method and the integral equation approach impose Assumptions (ref), (ref), and (ref). The main difference is Assumption (ref). In the integral equation approach, instead of Assumption (ref), the identifying assumptions are
See the appendix for details. The first condition turns out to be implied by Assumption (ref).\footnote{This implication was pointed out by a referee. I thank them for very helpful feedback.} Then, the key difference lies in the second condition, which imposes some high-level restrictions on the outcome equation.
An important advantage of the control function approach is that it avoids the high-level assumption (ref), which may be difficult to verify in practice. For example, consider the following simple model:
$(X^*,D)\protect\mathpalette{\protect\independenT}{\perp} (U_x,U_z)$, $U_x\protect\mathpalette{\protect\independenT}{\perp} U_z$, and $U_x,U_z$ both follows the standard normal distribution. In this model, for any value of $\rho \in (-1,1)$, the high-level condition (ref) fails while Assumptions (ref)-(ref) hold. See the appendix for derivations. In this specific example, explicit calculation of singular values and associated orthonormal functions is possible due to the normality assumption. In more general settings, it seems difficult to provide easy-to-verify sufficient conditions for the finiteness of the infinite sum unless one imposes specific structures on the outcome equation.
If one is willing to impose the high-level condition (ref), then the integral equation approach produces an estimation equation that is linear in nuisance parameters, which may facilitate estimation. Also, it has no difficulties with discrete endogenous variables. Yet, as the above example suggests, verifying (ref) needs to be done on a case-by-case basis, and it is not easy to do so. The control function method applies to a general nonparametric outcome equation if Assumption (ref) holds, and in Section (ref), I demonstrate that even when the common support condition does not hold, the control function method can effectively deal with non-linear outcome models. In addition, my method can leverage available results from the control function literature to provide a variety of modeling devices that facilitate implementations of my identification result.
\paragraph{Operator diagonalization method} The identifying assumptions of my approach and those of the operator diagonalization method HuSchennach2008 are non-nested. In particular, Assumption (ref) does not imply and is not implied by a completeness condition on $(X^*,Z)$ with respect to $Z$ (conditional on $D$), although Assumption (ref) implies a weaker version of completeness (see the appendix Section C).
The nonparametric identification results of HuSchennach2008 are more general than Theorem (ref) in the sense that they identify a larger class of parameters, including the distribution of $X^*$. Also, I assume that the treatment variable $D$ does not have a measurement error issue, whereas the operator diagonalization method can handle measurement errors in the treatment variable if there exists an appropriate proxy for mismeasured treatment.
The main advantage of my control function approach lies in practical implementation. It can effectively use additional structures on the outcome equation to relax some of the identifying assumptions. This point is practically relevant because in the majority of empirical applications, sample sizes are not large enough to warrant fully nonparametric estimation. In addition, control function methods are easy to implement with the help of cross validation as shown in Section (ref), whereas the operator diganalization method relies on sieve estimation methods with multiple tuning parameters.
To be concrete, compare a general nonparametric model with the random coefficient model discussed above. Whereas identification in a general nonparametric model requires $X$ and $Z$ to satisfy completeness-like conditions for $X^*$, the additional structure afforded by random coefficient models enables identification with only completeness condition on $X$ and very mild relevance condition for $Z$. In particular, discrete $Z$ is accommodated, even when the unobserved heterogeneity $X^*$ is continuously distributed. The operator diagnozliation method seems to require the same strong identifying assumptions in models with additional restrictions as in a general nonparametric model, because it does not exploit specific features of those models.
\paragraph{Control function method in triangular models} Most constructions of control functions rely on strict monotonicity of the first-stage equation. To be specific, consider the model
where $h$ is a fixed function and $\eta$ is an unobserved variable. ImbensNewey2009 (IN henceforth) showed that if $Z\protect\mathpalette{\protect\independenT}{\perp} (Y(d),\eta)$ and $h(z,\cdot)$ is invertible in the second argument for almost all $z$, then $F_{D|Z}(D|Z)$ is a valid control function.
For the sake of comparison, I consider how IN's approach might be used to construct a valid control function (i.e.,\ $Y(d)\protect\mathpalette{\protect\independenT}{\perp} D|F_{D|Z}(D|Z)$) in the measurement error problem of this paper. The main difference in the identifying assumptions lies in (i) the scalar restriction on the unobserved heterogeneity and (ii) the exclusion restriction on $Z$.
Suppose $Z$ satisfies the IV exclusion restriction in the sense that $(Y(d),X^*)\protect\mathpalette{\protect\independenT}{\perp} Z$. Then, IN's approach produces a valid control function if $X^*$ enters into the first-stage equation as a scalar e.g.,\ there exists some real-valued function $\iota$ such that
where $U$ is another unobserved term independent of other variables, and $h(z,\cdot)$ is strictly monotonic in its second argument. In my empirical setting, $X^*$ is two-dimensional (cognitive and non-cognitive ability), so the scalar restriction is indeed a substantive restriction. In the simple model of college admission considered above, the scalar restriction and strict monotonicity may be plausible as admission decision is based on the scalar admission score. In other settings, the scalar restriction might be violated Florens-Heckman-Meghir-Vytlacil_2008. My approach can relax this scalar restriction, provided that appropriate proxy variables are available, because it does not rely on the invertibility of the first-stage equation. On the other hand, if the invertibility holds, IN does not require a proxy variable as it exploits the available structure.
Another important distinction is the exclusion restriction on $Z$. Whereas my approach allows for statistical dependence between $X^*$ and $Z$, IN requires $Z\protect\mathpalette{\protect\independenT}{\perp} X^*$. In fact, correlation between $X^*$ and $Z$ helps with the common support condition for my approach. This distinction is empirically relevant. In DK's setting, a valid IV is not readily available because in observational settings, it is difficult to find an instrument that is independent of unobserved ability $X^*$ and yet affects college attendance. Thus, IN's approach may not be applicable in DK's setting. My approach remains valid if measurement errors in $Z$ (SAT score) are orthogonal to the outcome conditional on the unobserved ability, which seems plausible as I argued in Section (ref).
\paragraph{Group-level correlated random effect models} Theorem (ref) has a close connection with the identification results of AltonjiMatzkin2005. To explain, it is helpful to use the notation
where $\mathscr{Y}$ is an unknown non-stochastic function and $\varepsilon$ is an unobserved heterogeneity. Altonji and Matzkin based their identification results on finding pairs $(d_1,z_1),(d_2,z_2)$ such that
where $f_{\varepsilon|DZ}$ is the conditional density of $\varepsilon$ given $(D,Z)$ (see their equation (1.4)). In proving Theorem (ref), I show that (ref) is implied by
Thus, my identification strategy uses the proxy distribution to find pairs $(d_1,z_1),(d_2,z_2)$ such that Altonji and Matzkin's exchangeability condition holds. In this sense, my paper provides a framework in which the exchangeability condition (ref) follows from model primitives.
ArkhangelskyImbens2019 also provided a framework that implies a version of exhcnageability condition, and they presented additional identification results. They focused on settings where observational units belong to groups and there exists a group-level unobserved heterogeneity. My model does not have an explicit group structure, and what my identification strategy does is to form groups based on the value of $V$. That is, two observations belong to the same group if they have the same value of $V$. Similar to the setup in Arkhangelsky and Imbens, the treatment assignment becomes exogenous within groups, and treatment effects can be identified using this induced group structure.
I sketch the identification argument using a simple setting. I focus on the case where $D,X,X^*$ are all discrete. Specifically, the supports of $X$ and $X^*$ are $\mathcal{X}=\{x_1,\dots,x_L\}$ and $\mathcal{X}^*=\{x_1^*,\dots,x_L^*\}$ for some $L$. Note that in this special case, Assumption (ref) reduces to the full-column rank of the matrix
Define
which is the conditional distribution of the proxy variable $X$. Now I show that $V$ is a valid control function in the sense that $Y(d)\protect\mathpalette{\protect\independenT}{\perp} D|V$. To verify this claim, it suffices to show $X^*\protect\mathpalette{\protect\independenT}{\perp} D|V$ since Assumption (ref) imply $\mathrm{Pr}[Y(d)\leq y|D,V] = \mathbb{E}[\mathrm{Pr}[Y(d)\leq y|X^*]|D,V]$ and $X^*\protect\mathpalette{\protect\independenT}{\perp} D|V$ implies the desired result.
The law of total probabilities and Assumption (ref) imply
and stacking this equation for different values of $x\in\{x_1,\dots,x_L\}$,
By the full-rank condition,
This equality indicates that if two groups of workers have the same conditional distribution of test scores, then they also have the same conditional distribution of unobserved ability since $\bm{\Pi}$ is non-stochastic. This in turn implies that the conditional proxy distribution has the balancing property. To substantiate this last claim, for any $l\in\{1,\dots,L\}$,
where $e_l\in\mathbb{R}^L$ is the unit vector whose $l$th element is unity, the first equality holds as $V$ is a function of $(D,Z)$, the third equality follows from (ref), the fourth equality is by $\bm{\Pi}$ being non-random, and the fifth equality applies (ref) again. The conclusion $\mathrm{Pr}[X^*=x_l|D,V] = \mathrm{Pr}[X^*=x_l|V]$ establishes the desired result $X^*\protect\mathpalette{\protect\independenT}{\perp} D|V$.
In applications, it is important to allow for covariates. Denoting additional covariates by $S$, Assumptions (ref) and (ref) may be modified to $Y(d)\protect\mathpalette{\protect\independenT}{\perp} (D,Z)|X^*,S$ and $X\protect\mathpalette{\protect\independenT}{\perp} (D,Z,S)|X^*$, respectively. Then, I include additional covariates in the control function regression and the outcome regression i.e,\ $V=\mathbb{E}[B(X)|D,Z,S]$ and $\mathbb{E}[Y|D,V,S]$. To lighten the notation, I again make covariates implicit, but including them does not change estimation results below.
Theorem (ref) indicates that instead of conditioning on $\mathfrak{V}$, it suffices to condition on a finite-dimensional vector $V=\mathbb{E}[B(X)|D,Z]$ with some choice of $B(x) = (g_1(x),\dots,g_k(x))'$. While $k$ is guaranteed to be finite, it is not known a priori. Therefore, I develop practical estimation procedures that automatically select relevant control functions using Lasso. To this goal, I build on the recent developments in estimation methods using Neyman-orthogonal moments. The treatment here mostly follows Chernozhukov-Newey-Singh_2022_ecma (henceforth CNS) and I refer interested readers to the paper for details. For concreteness, I focus on the random coefficient model $Y=\varepsilon_1 + \varepsilon_2 D$ as considered in Section (ref) with $D\in\mathbb{R}$, where the parameter of interest is $\theta_0=\mathbb{E}[\varepsilon_2]$. Although the discussion in this section focuses on this specific model, the estimation method below can be easily extended to causal parameters in other models.
Let $\mu_0(d,v) = \mathbb{E}[Y|D=d,V=v]=\mu_{0,1}(v) + \mu_{0,2}(v)d$ and $V=\nu_0(D,Z)$ with $\nu_0(D,Z)=\mathbb{E}[B(X)|D,Z]$ where $B(x)$ potentially contains redundant elements, which may be discarded via a model selection procedure. A moment condition that identifies $\theta_0$ is
where $W=(Y,D,Z',B(X)')'$ is a data observation and $m(w,\mu,\nu)=\mu(d+1,\nu(d,z))-\mu(d,\nu(d,z))$ is the moment function. Having characterized the moment function $m$, a simple approach to estimate $\theta_0$ is to first estimate $\mu_0,\nu_0$ and form a sample analogue of $\mathbb{E}[m(W,\mu_0,\nu_0)]$. Yet, such plug-in approach is known to suffer from bias arising from noises in the first-stage estimates of $(\mu_0,\nu_0)$, especially when one uses Lasso or other machine learning methods to select regressors. Neyman-orthogonal estimation equations can be used to address this issue.
To describe Neyman-orthogonal moments, note that for any square-integrable function $\mu$,
holds with some regularity conditions. The function $f_{DV}(D-1,V)/f_{DV}(D,V)-1$ plays the role of a Riesz representer as discussed in CNS. Let $\Gamma$ be a function class that we use to estimate $\mu_0$, and denote by $\alpha_0$ the $L^2$ projection of a Riesz representer onto $\Gamma$. The exact form of $\Gamma$ will be specified below. Using the pathwise derivative calculation of HahnRidder2013 and Newey1994a with the assumptions below, one can show that
is an orthogonal score function for the parameter $\theta_0$. $\psi$ being an orthogonal score means that (i) $\mathbb{E}[\psi(W,\mu_0,\nu_0,\alpha_0)]=\theta_0$ and (ii) for a path $\{\mu_t,\nu_t,\alpha_t:t\in [0,\epsilon),\epsilon>0\}$,
holds. This second property suggests that the score function is insensitive to the first-stage estimation errors in $(\mu_0,\nu_0,\alpha_0)$. Then, one can form an estimator of $\theta_0$ by first estimating $(\mu_0,\nu_0,\alpha_0)$ and plugging them into $\psi$ to form a sample analogue of $\mathbb{E}[\psi(W,\mu_0,\nu_0,\alpha_0)]$.
Following CNS, I use cross-fitting. Let $\{I_l\}_{l=1}^L$ be a $L$-fold random partition of the sample. The number of partition $L$ is fixed (a common choice is $L=5,$ or $10$), and the size of each partition should be similar. Estimation involves multiple steps. In the first step, I estimate the control function $V=\nu_0(D,Z)$. Any nonparametric/machine learning method may be used to construct estimates of $V$, and I write $\widehat{V}_{ll'}=\widehat{\nu}_{ll'}(D,Z)$ for the estimate of $V$ using observations not in $I_l,I_{l'}$. Here, all the elements in $\mathbb{E}[B(X)|D,Z]$ are estimated, some of which may be redundant and discarded later. In the second step, I estimate $(\mu_0,\alpha_0)$ using Lasso. Let $p(v)=(p_1(v),\dots,p_K(v))'$ be a vector of approximating functions (e.g.,\ polynomial splines) whose dimension $K=K_n$ depends on the sample size. Define
where $n_l$ is the size of $I_l$. Then, the Lasso estimators of $(\mu_0,\alpha_0)$ are given by $\widehat{\mu}_l(d,v) = (p(v)',p(v)'d) \widehat{\rho}_l$, $\widehat{\alpha}_l(d,v) = (p(v)',p(v)'d) \widehat{\pi}_l$ where
$\kappa_n$ is a sequence of penalty terms that shrinks to zero, and $\|\cdot\|_1$ is the $L^1$ norm on $\mathbb{R}^{2K}$. Finally, an estimator of the causal parameter is formed by
where $\widehat{\nu}_l$ is an estimate of $\nu_0$ using observations not in $I_l$. Also, one can estimate the asymptotic variance of $\widehat{\theta}_n$ by
The proposed procedure allows researchers to be agnostic about which elements of the control function $V$ to be included in the regression. This aspect can be practically important as there is potential bias-variance trade-off. When the effective dimension of $V$ is low, including redundant elements may lead to noisier estimates in finite samples, whereas omitting relevant controls can lead to biased estimates. The automatic selection using Lasso may offer efficiency gains while ensuring that the bias does not arise from omitting important controls.
To present a formal result on the asymptotic distribution of $\widehat{\theta}_n$, I define additional notations. Write $\|\cdot\|_{\ell}$ for the $\ell$-norm $\ell=1,2$ on the Euclidean space. Let $q(d,v) = (p(v)'\ dp(v)')'$ be the vector of approximating functions and define $\Gamma$ to be the $L^2$ closure of linear combinations of $\{q_l(d,v): l\in\mathbb{N}\}$. For $u=(u_1,\dots,u_l)\in\mathbb{Z}_{\geq0}^l$, write $|u|=\sum_{\ell=1}^ku_{\ell}$ and $\partial^u f(x)=\partial^{|u|}f(x)/\partial^{u_1}x_1\dots\partial^{u_l}x_l$. Below $C>1$ denotes a positive constant independent of the sample size and represents a different number at different places.
This assumption imposes regularity conditions on the underlying data generating process, which are mild and standard in the literature. One non-standard aspect is that it strengthens Assumption (ref) to $\{Y(d):d\in \mathcal{D}\}\protect\mathpalette{\protect\independenT}{\perp} (D,Z)|X^*$. Although it is a strictly stronger condition mathematically, in many econometric models used in practice, it holds when Assumption (ref) does. This conditional independence restriction implies $\mathbb{E}[\varepsilon|D,Z]=0$ (Lemma A.3 in the appendix), which is crucial for the form of the asymptotic variance HahnLiaoRidderShi.
The next assumption imposes that the first-stage estimate of the control function is consistent and converges at a certain rate in $L^2$ norm.
Define
where $\partial^0 p(u) = p(u)$. Then, let
This rate plays an important role in the rest of assumptions. With specific $p(\cdot)$, bounds on $\Xi_{a,n}$ are available in the literature. For instance, with polynomial splines, $\Xi_{a,n}=K_n^{1/2+a}$ NEWEY1997.
The next assumption needs some additional notation: given a set of indices $J\subset\{1,\dots,L\}$ and $v\in\mathbb{R}^L$, let $\#|J|$ be the cardinality of $J$, $v_J$ be the $\#|J|\times 1$ subvector of $v$ whose elements are $\{v_l:l\in J\}$, and $v_{J^c}$ be the subvector of $v$ that consists of elements not in $v_J$.
This condition is a sparse eigenvalue condition commonly used in the Lasso literature (see Assumption 3 of CNS).
Assumption (ref) imposes that $\mu_0,\alpha_0$ admit sparse approximations. I follow Bradicetal in the formulation of this condition (their Assumptions 3 and 8). One aspect that differs from CNS is that I need to estimate the derivative of $\mu_0$. To accommodate this change, I impose that $p(\cdot)$ can approximate smooth functions and their derivatives in the $L^2$ norm. This approximation property holds for many of standard choices of approximating functions e.g.,\ polynomial splines.
The theorem states that the estimator $\widehat{\theta}_n$ is $\sqrt{n}$-consistent and asymptotically normal under the rate restrictions on the first-stage estimator, the Lasso penalty term, and the sup norm on the $\partial^u p(v)$ with $|u|\leq 2$. To illustrate, suppose we use polynomial splines as approximating functions and $\tilde{\omega}_n=n^{-\kappa_1}$, $K_n = n^{\kappa_2}$ for some $\kappa_1,\kappa_2>0$. Then, the hypothesis of Theorem (ref) implies $2\kappa_1-3\kappa_2 > \frac{1}{2} + \frac{1}{4\xi}$. The restriction is somewhat stringent on the rate at which the number of approximation terms can grow (i.e.,\ $K_n$) because both the first and second stages are non-parametric in $V$. It might be possible to relax rate restrictions and to obtain a better finite-sample distribution approximation by carefully analyzing the higher-order terms of the estimator CattaneoJanssonMa2019. I leave such analysis for future work.
In this section, I consider a flexible parametric estimation procedure based on the model (ref) in Section (ref). Recall that (ref) posits $\mathbb{E}[Y|D,V]=\Lambda(q(D,V)'\gamma_0)$ where $\Lambda(\cdot),q(\cdot)$ are specified by researchers and $\gamma_0$ is to be estimated. Let $d_v$ be the dimension of $B(x)$ and $d_q$ be some positive integer. For the control function $V$, I specify the model
where $Q:\mathcal{Z}\mathcal{D}\to\mathbb{R}^{ d_v\times d_q}$ is a matrix-valued transformation of $(D,Z)$ and $\delta_{0}\in\mathbb{R}^{d_q}$ is the parameter to be estimated. As a baseline, one may use $Q(D,Z)=I\otimes q_2(D,Z)$ with $I$ being the identity matrix, $q_2(D,Z)=(1,D',Z')$, and $\otimes$ denoting the Kronecker product. Researchers can include higher-order polynomial terms to enhance flexibility. Note that this is a reduced form equation, and as long as the model has good predictive power, the procedure is expected to work reasonably well.
For implementation, first estimate $\delta_{0}$ by least squares and form $\widehat{V}_i = Q(D_i,Z_i)\widehat{\delta}_{n}$. Then, estimate $\gamma_0$ in $\mathbb{E}[Y|D,V]=\Lambda(q(D,V)'\gamma_0)$ by (non-linear) regression of $Y$ on $q(D,\widehat{V})$. Finally, the estimator for the ASF is formed by
For inference, it is useful to have a closed-form variance estimator. Let $\dot{\Lambda}$ be the first derivative of $\Lambda$,
and $\partial q$ be the derivative of $q$ with respect to $V$. Then,
is an estimator for the asymptotic variance of $\sqrt{n}(\widehat{\vartheta}(d)-\vartheta(d))$. Note that the variance estimator does not require additional nuisance parameter estimation.
Since the asymptotic distributional theory for $\widehat{\vartheta}_n(d)$ is well-established NeweyMcfadden1994, I relegate the discussion of the asymptotic theory to the appendix. Under the assumptions stated there, the ASF estimator $\widehat{\vartheta}_n(d)$ is asymptotically normal and the variance estimator is consistent.
I apply the new results of this paper to studying causal effects of attending selective college on earnings. I build on the empirical framework of DK, but my approach has an advantage that data requirements are less demanding. Specifically, DK's main identification strategy relies on detailed information about college admission processes, which is not available in my setting. My control function method overcomes the data limitations, and my empirical result suggests that measurement errors indeed have non-trivial impact on causal effect estimates.
As discussed in Section (ref), DK's empirical framework posits that a college computes an admission score for each student and admits students if and only if their admission score is above the college's cutoff. That is,
where $C$ denotes the cutoff level, $X^*_1 +X^*_2 + \epsilon$ denotes the admission score, $X_1^*$ represents academic ability, $X_2^*$ is non-cognitive skill, and $\epsilon$ is an idiosyncratic term.
DK pointed out that students who applied to and got admitted to the same set of colleges must have similar admission scores e.g.,\ if students got admitted to college 1 but rejected from college 2 with cutoffs $C_1<C_2$, their admission scores would lie in the interval $(C_1,C_2]$. Using this insight, DK controlled for average SAT scores of colleges to which students applied and got admitted in order to account for the non-random selection. Therefore, for DK's empirical strategy, it is crucial to observe college application choices and admission results.
However, such detailed information is not often available. In my dataset, the National Longitudinal Study of 1972 (NLS72), some information about college admission process is observed, but there are two issues regarding the measurement. First, many missing observations on admission results make it difficult to control for the mean SAT score of colleges to which students got admitted. In fact, DK also analyzed the NLS72 dataset as a robustness check to their main analysis, but they only controlled for the mean SAT score of colleges to which students applied, what they call self-revelation models. In their empirical result, self-revelation models yielded similar estimates to their main specifications, yet theoretically, not conditioning on admission results might induce selection bias. Another issue is that the survey elicited only up to three schools to which students applied, but a non-negligible fraction of respondents applied to four or more colleges. In the final sample I use for analysis, $18.6\%$ ($210$ observations) said they applied to more than three colleges. Thus, the variable of mean SAT score of colleges to which students applied may have non-negligible measurement errors.
The control function method developed in this paper can address these data limitations by using partial information as proxy variables. In particular, I use math and reading test scores at 12th grade as proxies for academic ability $X_1^*$, a survey question that intended to measure psychological traits as a proxy for non-cognitive skill $X_2^*$, and the error-ridden measurement of the average SAT score of colleges to which students applied as another proxy variable. For the excluded variable $Z$, I use an SAT test score, another survey question as a measure of psychological traits, and whether worker's parents attended college. I impose Assumptions (ref)-(ref), (ref) as well as the outcome equation
The outcome model strikes the balance between flexibility and practicality. Importantly, the random coefficient model allows for heterogeneous treatment effects that can vary across the unobserved ability levels. The plausibility of the substantive assumptions (ref)-(ref) have been discussed in Sections (ref) and (ref). One caveat I should note is that the proxy for non-cognitive ability is a discrete variable, and thus, I impose that the unobserved non-cognitive skill has only moderate variation.
Table (ref) presents summary statistics of variables used in the analysis. The sample size is $1130$. Earnings, the outcome of interest, were measured when survey respondents were around 32 years old. The average of school mean of SAT (the treatment variable) is $1029$ with standard deviation $137$. In the College and Beyond survey analyzed by DK, the mean of the same variable was $1194$ with standard deviation $92.8$. Thus, workers in my sample attended less selective colleges on average relative to those in the College and Beyond survey. This feature of the sample could be important as DK found null effects of college selectivity on future earnings among workers who attended highly selective colleges while other studies found positive, sometime large, effects for workers with relatively low ability Hoekstra2009,Zimmerman2014.
Table (ref) presents the estimation results for the average effect on log earnings of attending college with $100$ higher mean SAT score. To better understand the performance of my proposed method, I also estimated the log earnings regression with the proxy and excluded variables as control regressors using the Lasso-based method of CNS. This second approach is prone to measurement errors as test scores are noisy measurements of ability. The control function estimate indicates that earnings increases by $7.9\% \approx 100*(\exp(0.076)-1)$ when a worker attended college with $100$ higher mean SAT score, whereas the “na\"{i}ve” estimate indicates the increase is $3.7\%$. Thus, the control function estimate of the causal effect is more than twice the measurement-error prone estimate. As expected, the control function estimate has a larger standard error with a wider confidence interval. Using the influence function representations, I tested the null hypothesis of the two estimators converging to the same value, and the $p$-value for the two-sided alternative was $0.083$. Thus, my empirical result indicates that measurement errors in control variables have non-trivial impact on causal effect estimates.
I developed a new identification strategy for causal effects such as the average structural functions by exploiting proxy variables for unobserved confounding factors. This new approach identifies causal parameters with relatively weak assumptions by exploiting additional structures on the outcome equation. I developed a Lasso-based method to flexibly choose regression specifications for my control function method. As illustrated through an empirical application, measurement errors in covariates can have non-negligible impact on causal estimates. My method can be used to apply the main idea of Dale-Krueger_2002 under less demanding data requirements.