EconBase
← Back to paper

IV Regressions without Exclusion Restrictions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

121,290 characters · 22 sections · 47 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

IV Regressions without Exclusion Restrictions

\sloppy

\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long\global\long\global\long\global\long \global\long\global\long \global\long\global\long \global\long \global\long \global\long\global\long \global\long\global\long\global\long\global\long\global\long \global\long \global\long\def\mc#1{\mathscr{#1}}

\global\long\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long\def\abs#1{\left|#1\right|}

\global\long\def\norm#1{\left\Vert #1\right\Vert }

\global\long\def\rest#1{\left.#1\right|}

\global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle }

\global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle }

\global\long\def\turd#1{\frac{#1}{3}}

\global\long \global\long\def\sand#1{\left\lceil #1\right\vert }

\global\long\def\wich#1{\left\vert #1\right\rfloor }

\global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor }

\global\long\def\abs#1{\left|#1\right|}

\global\long\def\norm#1{\left\Vert #1\right\Vert }

\global\long\def\rest#1{\left.#1\right|}

\global\long\def\inprod#1{\left\langle #1\right\rangle }

\global\long\def\ol#1{\overline{#1}}

\global\long\def\ul#1{#1}

\global\long\def\td#1{\tilde{#1}} \global\long\def\bs#1{\boldsymbol{#1}}

\global\long \global\long \global\long \global\long \global\long {6pt} {6pt}

abstractWe study identification and estimation of endogenous linear and nonlinear regression models without excluded instrumental variables, based on the standard mean independence condition and a nonlinear relevance condition. Based on the identification results, we propose two semiparametric estimators as well as a discretization-based estimator that does not require any nonparametric regressions. We establish their asymptotic normality and demonstrate via simulations their robust finite-sample performances with respect to exclusion restrictions violations and endogeneity. Our approach is applied to study the returns to education, and to test the direct effects of college proximity indicators as well as family background variables on the outcome. \\ Keywords: linear regression, quantile regression, endogeneity, instrumental variable, exclusion restriction, semiparametric two-stage estimation

Introduction

The method of instrumental variables (IV) has been a central approach to identify and estimate linear regression models with endogeneity. The conventional IV regression exploits excluded instrumental variables that have no direct effects on the outcome variable. However, finding valid instruments that satisfy the exclusion restriction can be challenging in many applications.

In this paper, we show that even in the absence of excluded instruments, the endogenous linear regression model can still be identified by leveraging the nonlinear relevance between the included exogenous regressor and the endogenous variable. In contrast to the traditional IV regression that uses a linear first-stage projection, our approach applies a mean projection of the endogenous variable on only the included exogenous regressor in the first stage. More generally, this approach can also be applied to nonlinear regressions with known functional form, which may naturally arise from structure models. In such cases, we provide local identification of the model parameter under a full-rank condition.

To illustrate, let us consider the following simple linear regression model:

equation[equation omitted — 98 chars of source]

with a scalar endogenous variable $X_{i}$ and a scalar exogenous variable $Z_{i}$ such that \[ \mathbb{E}\left[\rest{\epsilon_{i}}Z_{i}\right]=0. \]

To identify the coefficient $\theta_0:=\left(\alpha_{0},\beta_{0},\gamma_{0}\right)$, the standard IV regression, or the two-stage least square (2SLS) regression, relies on the availability of an additional variable $Z_{i, exc}$ and applies a linear first-stage projection as follows: \[ X_{i}=\lambda_{0}+\lambda_{1}Z_{i}+\lambda_{2}Z_{i, exc}+U_{i}, \] where the instrumental variable $Z_{i, exc}$ is required to be exogenous with respect to $\epsilon_{i}$, relevant for $X_{i}$, and excluded from the regression model (ref).

By contrast, this paper investigates the identification and estimation of $\theta_0$ without excluded instrumental variables. To see the idea, take conditional expectations of both sides of (ref) given $Z_{i}$, we have

equation[equation omitted — 123 chars of source]

because the term $\mathbb{E}\left[\rest{\epsilon_i}Z_i \right]=0$ by the exogeneity of $Z_i$.

Instead of linearly projecting endogenous variable $X_i$ on $Z_i$, we adopt the mean projection of $X_i$ on $Z_i$, i.e., $\pi_{0}\left(Z_{i}\right):=\mathbb{E}\left[\rest{X_{i}}Z_{i}\right]$. Then the moment restriction in (ref) can be written as

equation[equation omitted — 144 chars of source]

Our key idea is based on the simple observation that condition (ref) can be viewed as a linear regression of $Y_i$ on 1, $Z_{i}$, and $\pi_{0}\left(Z_{i}\right)$ with no endogeneity issue since $Z_i$ satisfies the exogeneity condition. It is thus clear that the parameter $\theta_0$ is identified in (ref) as long as $\left(1,Z_{i},\pi_{0}\left(Z_{i}\right)\right)$ are not (perfectly) multicollinear, which is equivalent to the following requirement: \[ \pi_{0}\left(z\right)\text{ is nonlinear in }z, \] i.e, there exist no constants $a,b\in\mathbb{R}$ such that $\pi_{0}\left(z\right)=a+bz$ for any $z$ in the support of $Z_i$. This nonlinearity condition is testable, as $\pi_0$ only involves observed variables.

One natural setting this nonlinear relationship arises is when the endogenous regressor $X_{i}\in \{0, 1\}$ is a binary variable. Then the propensity score function, $\pi_0(z)$, is naturally nonlinear in $z$. For example, consider the following binary choice model for $X_i$, \[ X_{i}=\mathbf{\mathbbm1}\left\{ \eta_0+\eta_1 Z_i \geq u_{i}\right\}, \] where $u_i \perp Z_i$ and $ u_i$ follows some distribution $F_u$ (e.g., a normal distribution for the Probit model). So \[ \pi_{0}\left(z\right)=\mathbb{E}\left[\rest{X_{i} }Z_{i}=z\right]=F_u(\eta_0+\eta_1z). \] Then, the function $\pi_{0}\left(z\right)$ is nonlinear as long as $Z_i$ takes at least three values and $F_u$ is not a uniform distribution. Note that in our identification approach, the distribution $F_u$ does not need to be known, which is distinct from the Heckman correction approach.\footnote{Section (ref) provides a more detailed discussion of the differences between our approach and the Heckman correction approach.}

As a more concrete example, suppose we are interested in the effect of a college degree $X_i$ on (log) wage $Y_i$. Then the included instrument $Z_i$ could be years of parents' education, which takes more than three values. Alternatively, $Z_i$ can include two binary variables $Z_{i1}, Z_{i2}$, with $Z_{i1}$ representing gender and $Z_{i2}$ representing whether one's mother has a college degree.\footnote{The nonlinearity of $\pi_0$ can be satisfied under a mild condition on their coefficients: $\Phi(\eta_0+\eta_1+\eta_2)+\Phi(\eta_0)-\Phi(\eta_0+\eta_1)-\Phi(\eta_0+\eta_2)\neq 0$, where $\eta_1, \eta_2$ are the coefficients of $Z_{i1}, Z_{i2}$, respectively.}

More generally, the analysis extends to endogenous nonlinear and quantile regression. By adopting a mean projection of the nonlinear function onto the included exogenous regressor, we show that local identification is achieved under a full-rank condition. In particular, for quantile regression, we show that the full-rank condition is equivalent to a different nonlinear relevance condition.

commentThe identification strategy can be generalized to endogenous nonlinear regressions: \[Y_i=f(Z_i, X_i, \theta_0)+\epsilon_i,\] where $f$ is known up to the parameter $\theta_0$. We can adopt the same approach by applying the mean projection of the entire term $f(Z_i, X_i, \theta_0)$ on the included regressor $Z_i$. Let $m_0(Z_i, \theta_0):=\mathbb{E}\left[ \rest{f\left(Z_i, X_i, \theta_0 \right)}Z_i\right]$, the exogeneity condition $\mathbb{E}\left[ \rest{\epsilon_i} Z_i \right]=0$ yields \[ \mathbb{E}\left[ \rest{Y_i-m_0(Z_i, \theta_0)}Z_i \right]=0. \] Similarly, the above condition can be viewed as a standard nonlinear regression of $Y_i$ on $m_0(Z_i, \theta_0)$ with no endogeneity. Then $\theta_0$ is locally identified if the matrix $\mathbb{E}\left[\nabla_\theta m_0(Z_i, \theta_0)\nabla_{\theta^{'}} m_0(Z_i, \theta_0)\right]$ has full rank. In the example of the endogenous quantile regression, the moment restriction is given as: \begin{equation} \mathbb{E}\left[\rest{\mathbf{\mathbbm1}\{Y_i\leq \alpha_{0}+\beta_0 Z_{i}+\gamma_0X_{i} \}-\tau} Z_i \right]=0. \end{equation} In this example, the indicator term $\mathbf{\mathbbm1}\{Y_i\leq \alpha_{0}+\beta_0 Z_{i}+\gamma_0X_{i} \}$ is nonlinear and nonseparable in all variables $(Y_i, Z_i, X_i)$. After projecting the indicator term $\mathbf{\mathbbm1}\{Y_i\leq \alpha_{0}+\beta_0 Z_{i}+\gamma_0X_{i} \}$ on $Z_i$ with $m_0(Z_i, \theta_0):=\mathbb{E}\left[\rest{\mathbf{\mathbbm1}\{Y_i\leq \alpha_{0}+\beta_0 Z_{i}+\gamma_0X_{i} \}} Z_i \right]$, the moment restriction in (ref) can be expressed as \[ m_0(Z_i, \theta_0)-\tau=0. \] The full rank condition of $\mathbb{E}\left[\nabla_\theta m_0(Z_i, \theta_0)\nabla_{\theta^{'}} m_0(Z_i, \theta_0)\right]$ is shown to be equivalent to the no multicollinearity of $1$, $Z_i$, and $\tilde{\pi}_0\left(Z_i\right):=\mathbb{E}\left[\rest{X_i} Z_i, \epsilon_i=0 \right]$, which is also equivalent to the following nonlinear relevance condition: \[\tilde{\pi}_0\left(z\right) \text{ is nonlinear in }z\text{ on the support of }Z_{i}. \] The above nonlinearity condition is similar to the nonlinearity requirement of $\pi_0(z)$ in the linear regression model, except that $\tilde{\pi}_0$ is defined conditional on the error term $\epsilon_i=0$. The nonlinearity of $\tilde{\pi}_0$ also naturally arises when the endogenous regressor $X_i$ is binary or discrete.

The identification result suggests a natural semiparametric two-step estimator. We describe the estimator for the linear regression model, and the estimator for the nonlinear and quantile regression is provided in Section (ref). Specifically, given $\hat{\pi}$ obtained via first-stage nonparametric regression of $X_{i}$ on $Z_{i}$, we construct our first estimator $\hat{\theta}$ by \[ \text{regressing \ensuremath{Y_{i}\text{ on }1,\ }}Z_{i}\ \text{and}\ \hat{\pi}_{0}\left(Z_{i}\right)\ \ \text{via OLS}. \]

We also propose an estimator that uses the mean projection $h_0\left( Z_i\right):=\mathbb{E}\left[\rest{Y_{i}}Z_{i}\right]$ of $Y_i$ on $Z_i$ as the dependent variable. Let $\hat{h}$ denote the estimator of nonparametric regression of $Y_i$ on $Z_i$, the second estimator $\hat{\theta}^{*}$ is constructed by \[ \text{regressing \ensuremath{\hat{h}\left(Z_{i}\right)\text{ on }1,\ }}Z_{i}\ \text{and}\ \hat{\pi}_{0}\left(Z_{i}\right)\ \text{via OLS}. \]

The only difference between $\hat{\theta}$ and $\hat{\theta}^{*}$ lies in the dependent variable used in the second step: $\hat{\theta}$ uses the raw observed variable $Y_{i}$, while $\hat{\theta}^{*}$ uses the fitted value $\hat{h}\left(Z_{i}\right)$ obtained through nonparametric regression of $Y_{i}$ on $Z_{i}$. We propose the second estimator $\hat{\theta}^{*}$, as it can perform slightly better than $\hat{\theta}$ under some specifications in simulations.

We further propose a third estimator, $\hat{\theta}_{disc}$, which does not require any nonparametric estimation, based on a discretization of the support of $Z_{i}$ into $K$ (finite and fixed) partitions. Under this discretization, the first-stage estimation simplifies to sample averages in each partition. Furthermore, the estimator can be computed as a standard 2SLS estimator with partition dummies as IVs. While the discretization results in some loss of information and asymptotic efficiency, there is no “discretization bias” in our setting and the number of partition cells $K$ is not required to grow large with the sample size.

We establish the $\sqrt{n}$-consistency of our three proposed estimators $\hat{\theta}$, $\hat{\theta}^{*}$, and $\hat{\theta}_{disc}$ for $\theta_{0}$, along with their asymptotic normality. We show that $\hat{\theta}$ and $\hat{\theta}^{*}$ share exactly the same asymptotic variance, while that of $\hat{\theta}_{disc}$ is in general different and, when error are homoskedastic, larger under the partial order of positive semi-definiteness.

Monte Carlo simulations support our theoretical results, and demonstrate the good finite-sample performance of the three estimators $\hat{\theta}$, $\hat{\theta}^{*}$, and $\hat{\theta}_{disc}$ with the presence of violation of the exclusion restriction and endogeneity. For comparison, we also implement the standard 2SLS estimator which treats the included regressor as excluded instrument, as well as the OLS estimator which does not account for endogeneity. The root mean squared error (RMSE) of the three estimators $\hat{\theta}$, $\hat{\theta}^{*}$, and $\hat{\theta}_{disc}$ are reasonably small, and the coverage probabilities of the 95% confidence intervals are very close to their nominal level, even with a relatively modest sample size of $n=250$. In contrast, the standard 2SLS estimator has much larger bias and standard errors when the exclusion restriction is violated. As expected, the OLS estimator perform poorly in the presence of endogeneity.

Our approach is applied to study the returns to education and to test the direct effects of different instruments. Our first application, in line with card1993using, studies two indicators of college proximity: the presence of a nearby 2-year college and a nearby 4-year college. Our findings show that after controlling for regional characteristics, the two college proximity indicators have no significant effects on the outcome. However, the 2SLS estimator varies substantially when using different instruments, while our estimators remain robust under various specifications. In the second application, we investigate two family background variables as potential instruments: parents' education and number of siblings. The results indicate that the number of siblings exerts no significant effect on wages, while parents' education significantly increases income. The estimated returns to education based on our three estimators appear to be smaller than those of 2SLS estimators, as our methods account for the direct effects of the two instruments.

The main contribution of our paper is to provide identification and estimation of endogenous linear and nonlinear regression models using only included exogenous regressors.\footnote{Our proposed method can also be applied to test exclusion restrictions. One could estimate the regression model using our method, treating all exogenous variables as included IVs. Then testing the exclusion restrictionna of a specific exogenous variable corresponds to testing whether its coefficient is zero. Compared to the classic overidentification testing, our method only requires one instrument to conduct the test.} Our approach offers an alternative solution for endogeneity when it is challenging to find excluded IVs. In such scenarios, empirical researchers may consider using the included exogenous variables $Z_{i}$ as the “included IVs” for $X_i$ (or the entire term containing $X_i$), and adopt the mean projection of $X_i$ on $Z_i$. We provide a corresponding set of easy-to-use semiparametric estimation and inference procedures, along with their theoretical properties. Hence, we believe our results could have broad applicability, given the general relevance of endogenous linear and nonlinear regression models in applied work.

Our paper is most closely related to the line of econometric literature on the identification of endogenous regression models without exclusion restrictions. See lewbel2019identification for a comprehensive survey of related work on this topic. In the standard linear regression setting, rigobon2003identification, klein2010estimating, lewbel2012using, and lewbel2018identification utilize heteroskedasticity of error terms, while \citet*{lewbel2023identification} works with a specific decomposition of error and imposes independence between them. Beyond the standard linear regression setting, \citet*{dong2010endogenous} considers a binary response model with imposed independence assumption among error terms, kolesar2015identification studies a linear regression with “many IVs” under an orthogonality condition between the IVs' direct effects on the outcome variable and the effects on the endogenous covariates, and \citet*{d2021testing} considers a linear random coefficient model with exogeneity and independence assumptions on the random coefficients. Another relevant paper is \citet*{escanciano2016identification}, who studies a more general framework of semiparametric conditional moment models and provides high-level conditions for identification without exclusions. They adopt the control function approach and impose the (full) conditional independence assumption of errors. Furthermore, they work with the moment equation conditional on both the endogenous regressor and the included exogenous regressor. In contrast, our approach is based on the moment equation given only the included exogenous regressor under the mean independence assumption of this regressor. More differently, \citet*{honore2020selection,honore2022sample} investigate partial identification of sample selection models without exclusion.

Another relevant line of literature is on optimal IV and asymptotic efficiency in the estimation of conditional moment restriction models: see, for example, \citet*{amemiya1974nonlinear,amemiya1977maximum}, \citet*{chamberlin1987asymptotic} and \citet*{newey1990efficient,NEWEY1993419}, \citet*{ai2003efficient}, and \citet*{newey2004efficient}. The main focus of this line of literature is on asymptotic efficiency and typically assumes identification as a starting point. Consequently, this literature does not explicitly distinguish between included and excluded IVs or between linear and nonlinear revelance of IVs. In addition, many papers in this literature, such as \citet*{donald2001choosing}, \citet*{hahn2002optimal} and \citet*{stock2005asymptotic}, are more concerned with the scenario where there are many IVs (which are often implicitly excluded IVs), while we focus on exactly the opposite scenario, where researchers do not have any excluded IVs. In addition, \citet*{escanciano2018simple} considers endogenous linear regressions and proposes the “integrated IV estimator” as a simple and robust alternative to the optimal IV approach. However, the focus of \citet*{escanciano2018simple} is on robustness (especially with weak instruments), and similarly it does not distinguish between included/excluded IVs or linear/nonlinear relevance of IVs.

In the special case where $X_{i}$ is binary, there is also a connection between our paper and the literature on heterogeneous treatment effects. This literature, as exemplified by imbens1994identification, \citet*{angrist1996identification}, and heckman2005structural, studies endogenous selection and instrumental variables within the potential outcome framework. See, e.g., imbens2014instrumental, \citet*{imbens2015causal}, mogstad2018identification, and abadie2018econometric for more comprehensive reviews. This framework allows for nonparametrically heterogeneous treatment effects, but usually imposes full conditional independence assumptions along with monotone relevance conditions on the IVs. Under this framework, the most closely related line of work is on the identification of treatment effects without exclusion restrictions: manski2000monotone, flores2013partial, and mealli2013using establish partial identification without exclusion. Moreover, hirano2000assessing relaxes the exclusion condition by applying the Bayesian approach, while wang2022 employs an additional instrument for identification.

The paper also relates to work on endogenous nonlinear and quantile regression models, such as newey2003instrumental, chernozhukov2005iv, chernozhukov2007instrumental, and chernozhukov2008instrumental. The existing studies explore nonparametric identification with excluded instruments. In contrast, our paper focuses on parametric models and investigates identification using only included exogenous regressors.

Our semiparametric two-stage estimation procedure with nonparametric regression of the endogenous/outcome variables on the included exogenous variables are also reminiscent of robinson1988root, who considers a partially linear regression model without endogeneity. However, one of the key steps in robinson1988root is to transform the regression equation into a “differenced form” that is free of the unknown nonparametric function in the original equation. In contrast, the identification arguments in our linear regression setup does not involve the “differenced form” equation. A recent paper by \citet*{antoine2022partially} studies the partially linear model with endogenous covariates. They again work with the differenced form in the style of robinson1988root, and then rely on excluded IVs for identification.

Lastly, our discretization-based estimator bears some resemblance to the inferential methods for conditional moment inequalities, as studied in khan2009inference and andrews2013, for example.

The rest of the paper is organized as follows. Section (ref) introduces the identification of linear regression models, along with further discussions about the identification condition. Section (ref) discusses the comparison of our approach with existing methods in the literature. Section (ref) derives asymptotic distributions of our three proposed estimators and provides corresponding variance estimators. Section (ref) explores the identification of nonlinear and quantile regressions. Section (ref) presents simulation results about the finite-sample performances of our estimators. Section (ref) studies the returns to education and examines the direct effects of various instruments. We conclude with Section (ref).

Endogenous Linear Regression without Exclusion

Model and Identification

Consider the following linear regression model with endogeneity:

equation[equation omitted — 101 chars of source]

where $X_{i}$ is a $d_{x}$-dimensional endogenous regressor that can be dependent with $\epsilon_{i}$, while $Z_{i}$ is a $d_{z}$-dimensional included exogenous regressor satisfying the following mean independence, or strict exogeneity, assumption:

assumption[Mean Independence] $\mathbb{E}\left[\rest{\epsilon_{i}}Z_{i}=z\right]=0\label{eq:ExoZ} $ for any $z\in{\cal Z}:=\text{Supp}\left(Z_{i}\right)$.

Writing $\theta_{0}:=\left(\alpha_{0},\beta_{0}^{'},\gamma_{0}^{'}\right)^{'}\in\mathbb{R}^{d:=1+d_{x}+d_{z}}$, we are interested in identifying and estimating $\theta_{0}$. Section (ref) explores the extension of endogenous nonlinear and quantile regressions.

Assumption (ref) on $Z_i$ leads to the following conditional moment restriction:

equation[equation omitted — 128 chars of source]

which characterizes the identified set for $\theta_0$. We show that the above restriction can point identify $\theta_0$ under the no multicollinearity condition.

Define $\pi_{0}\left(z\right):=\mathbb{E}\left[\rest{X_{i}}Z_{i}=z\right]$. By employing the mean projection of $X_i$ on $Z_i$, we can rewrite (ref) by replacing the endogenous regressor $X_i$ with $\pi_0\left(Z_i\right) $: \[ \mathbb{E}\left[\rest{Y_{i}-\alpha_{0}-Z_{i}^{'}\beta_{0}-\pi_0\left(Z_i\right)^{'}\gamma_{0} }Z_{i}=z\right]=0.\label{eq:RegZwNoError} \] When treating $\pi_0\left(Z_i\right)$ as a regressor, the above condition transforms into the moment restriction of a standard linear regression, which regresses $Y_i$ on 1, $Z_i$, $\pi_{0}\left(Z_{i}\right)$. After applying the mean projection, there is no endogeneity since $Z_i$ satisfies the strict exogeneity condition.

Let $W_{i}:=\left(1,Z_{i}^{'},\pi_{0}\left(Z_{i}\right)^{'}\right)'$. We then apply the usual identification strategy by premultiplying both sides of the above equation by $W_{i}$ and then taking unconditional expectations: \[ \mathbb{E}\left[W_{i}Y_i\right]=\mathbb{E}\left[W_{i}W_{i}^{'}\right]\theta_{0}.\label{eq:IDeq_theta0} \] Since $\pi_{0}\left(z\right)$ is nonparametrically identified from data, the terms $W_{i}, \mathbb{E}\left[W_{i}W_{i}^{'}\right]$, and $\mathbb{E}\left[W_{i}Y_i\right]$ are also identified. It is then clear that $\theta_{0}$ is identified whenever $\mathbb{E}\left[W_{i}W_{i}^{'}\right]$ is invertible, which boils down to the familiar requirement of no multicollinearity condition:

assumption[No Multicollinearity] $\left(1,Z_{i}^{'},\pi_{0}\left(Z_{i}\right)^{'}\right)$ are not (perfectly) multicollinear. Or equivalently, $\mathbb{E}\left[W_{i}W_{i}^{'}\right]$ has full rank.

The discussion regarding Assumption (ref) is presented in Section (ref). Under this assumption, $\theta_0$ is identified as the standard OLS formula with $W_i$ as the regressor: \[ \theta_{0} =\left(\mathbb{E}\left[W_{i}W_{i}^{'}\right]\right)^{-1}\mathbb{E}\left[W_{i}Y_{i}\right]. \]

Since $W_{i}$ is a deterministic function of $Z_{i}$, we can also project $Y_i$ on $Z_i$ and obtain an alternative expression for $\theta_0$. Defining $ h_0\left(Z_i \right):=\mathbb{E}\left[\rest{Y_i} Z_i\right]$, then $\theta_0$ can be expressed as \[\theta_{0}=\left(\mathbb{E}\left[W_{i}W_{i}^{'}\right]\right)^{-1}\mathbb{E}\left[W_{i}h_{0}\left(Z_{i}\right)\right], \] which follows from the Law of Iterated Expectations. We conduct this additional projection because, through simulation, we find that the estimator based on this formula can exhibit slightly better performance under some specifications.

thm[Identification with Included IV] Under Assumptions (ref) and (ref), \begin{align} \theta_{0} &=\left(\mathbb{E}\left[W_{i}W_{i}^{'}\right]\right)^{-1}\mathbb{E}\left[W_{i}Y_{i}\right] =\left(\mathbb{E}\left[W_{i}W_{i}^{'}\right]\right)^{-1}\mathbb{E}\left[W_{i}h_{0}\left(Z_{i}\right)\right]. \end{align}

Theorem (ref) suggests two natural semiparametric two-step estimators for $\theta_{0}$. Specifically, given first-stage nonparametric estimators $\hat{\pi}$ for $\pi_{0}$ and $\hat{h}$ for $h_{0}$, the second-stage plug-in estimators for $\theta_{0}$ is given by, with $\hat{W}_{i}:=\left(1,Z_{i}^{'},\hat{\pi}\left(Z_{i}\right)^{'}\right)'$, \[

aligned\hat{\theta} & :=\left(\frac{1}{n}\sum_{i=1}^{n}\hat{W}_{i}\hat{W}_{i}^{'}\right)^{-1}\frac{1}{n}\sum_{i=1}^{n}\hat{W}_{i}Y_{i},\\ \hat{\theta}^{*} & :=\left(\frac{1}{n}\sum_{i=1}^{n}\hat{W}_{i}\hat{W}_{i}^{'}\right)^{-1}\frac{1}{n}\sum_{i=1}^{n}\hat{W}_{i}\hat{h}\left(Z_{i}\right).

\]

As shown in Section (ref), both $\hat{\theta}$ and $\hat{\theta}^*$ are $\sqrt{n}$-consistent and asymptotically normal. Furthermore, they share the same asymptotic variance, and are thus asymptotically equally efficient. In the meanwhile, $\hat{\theta}$ does not need nonparametric estimation of $h_{0}$, and is thus simpler and faster to compute than $\hat{\theta}^*$. However, we do find that $\hat{\theta}^*$ can have better finite-sample performance under certain simulation setups. Hence, we keep the estimator $\hat{\theta}^*$ in our paper and provide results for it along with $\hat{\theta}$.

In Section (ref), we propose a third estimator $\hat{\theta}_{disc}$ based on a discretization of the support of $Z_i$, which does not require any nonparametric regressions in the first stage. However, $\hat{\theta}_{disc}$ is not directly based on the sample analog of (ref). Hence, we defer $\hat{\theta}_{disc}$ to Section (ref).

An Alternative Perspective

In Section (ref), we establish point identification of $\theta_0$ from the perspective of the standard linear regression models, which naturally leads to the familiar “no multicollinearity" or “full rank" condition in Assumption (ref). A slightly different perspective is to exploit the fact that the conditional moment equation (ref) is a system of deterministic linear equations in $\theta$ across all $z\in{\cal Z}$. Therefore $\theta_0$ is uniquely determined if the following condition holds:

condition[Full-Dimensional Support] There exist $d=1+d_{x}+d_{z}$ distinct points $z_{1},...,z_{d}\in{\cal Z}$ such that \[ \text{rank}\left(\begin{array}{ccc} 1 & z_{1}^{'} & \pi_{0}\left(z_{1}\right)^{'}\\ 1 & z_{2}^{'} & \pi_{0}\left(z_{2}\right)^{'}\\ \vdots & \vdots & \vdots\\ 1 & z_{d}^{'} & \pi_{0}\left(z_{d}\right)^{'} \end{array}\right)=d. \]

It turns out that Condition (ref) is equivalent to Assumption (ref), which is also intuitively so under linearity. Hence, the two perspectives for identification are equivalent.

lemAssumption (ref) $\Leftrightarrow$ Condition (ref).

Condition (ref) provides an alternative perspective for identification from the support of the included instrument $Z_i$, under the feature that $\left(1, Z_{i}^{'}, \pi_{0}\left(Z_{i}\right)^{'}\right)$ is a deterministic function of $Z_{i}$. We see that even though the dimension $d_{z}$ of the included instrument $Z_{i}$ is by construction smaller than the number of parameters $d$ (e.g., a scalar $Z_i$), it is still possible for us to find $d$ linearly independent realizations of $\left(1, Z_{i}^{'}, \pi_{0}\left(Z_{i}\right)^{'}\right)$ on the support of $Z_{i}$, which will guarantee the required “no multicollinearity” assumption.

commentWhile the sufficiency of Condition (ref) for Assumption (ref) is relatively trivial, the necessity of it in our current setup is more tightly based on the fact that $\left(1,z^{'},\pi_{0}\left(z\right)^{'}\right)$ is a deterministic function of $z$. Hence, even though the dimension $d_{z}$ of included IVs $Z_{i}$ is by construction smaller than the number of parameters $d$, it is still possible for us to find $d$ linearly independent realizations of $\left(1,Z_{i}^{'},\pi_{0}\left(Z_{i}\right)^{'}\right)$ on the support of $Z_{i}$.

This perspective also motivates our third estimator $\hat{\theta}_{disc}$ by transforming the conditional moment equation into the following unconditional moment equation: \[\mathbb{E}\left[\left(Y_i-\alpha_{0}-Z_i^{'}\beta_{0}-X_i'\gamma_{0} \right)\mathbf{\mathbbm1}\{Z_i\in {\cal Z}_k\} \right]=0, \] where $({\cal Z}_k)_{k=1}^K$ is a finite partition of the support of $Z_i$ with $K\geq d$. The idea of transforming conditional moments into unconditional ones using instrumental functions such as indicator functions, has been well studied and applied in the literature: e.g., khan2009, andrews2013, and shi2018estimating. See Section (ref) for more details about the discretization-based estimator $\hat{\theta}_{disc}$.

Discussion about Assumption (ref)

Since Assumption (ref) is the foundation for the identification of $\theta_{0}$, we now provide some necessary and/or sufficient conditions for it, along with some more detailed discussions on its relationship to nonlinearity, relevance, and order condition:

commentWe first provide the following equivalent condition for Assumption (ref). \begin{condition}[Full-Dimensional Support] There exist $d=1+d_{x}+d_{z}$ distinct points $z_{1},...,z_{d}\in{\cal Z}$ such that \[ \text{rank}\left(\begin{array}{ccc} 1 & z_{1}^{'} & \pi_{0}\left(z_{1}\right)^{'}\\ 1 & z_{2}^{'} & \pi_{0}\left(z_{2}\right)^{'}\\ \vdots & \vdots & \vdots\\ 1 & z_{d}^{'} & \pi_{0}\left(z_{d}\right)^{'} \end{array}\right)=d. \] \end{condition} \begin{lem} Assumption (ref) $\Leftrightarrow$ Condition (ref). \end{lem} While the sufficiency of Condition (ref) for Assumption (ref) is relatively trivial, the necessity of it in our current setup is more tightly based on the special feature that $\left(1,Z_{i}^{'},\pi_{0}\left(Z_{i}\right)^{'}\right)$ is a deterministic function of $Z_{i}$. Hence, even though the dimension $d_{z}$ of included IVs $Z_{i}$ is by construction smaller than the number of parameters $d=1$, it is still possible for us to find $d$ linearly independent realizations of $\left(1,Z_{i}^{'},\pi_{0}\left(Z_{i}\right)^{'}\right)$ on the support of $Z_{i}$, which will guarantee the required “no multicollinearity” assumption.
condition[No Multicollinearity in $Z_{i}$] $\left(1,Z_{i}^{'}\right)$ are not multicollinear.
condition[Nonlinearity] $\pi_{0,k}\left(z\right):=\mathbb{E}\left[\rest{X_{i,k}}Z=z\right]$ is nonlinear in $z$ on ${\cal Z}$, for each component $k=1,...,d_{x}$.
condition[Relevance] $\pi_{0,k}\left(z\right):=\mathbb{E}\left[\rest{X_{i,k}}Z=z\right]$ is not constant in $z$ on ${\cal Z}$, for each component $k=1,...,d_{x}$.
condition[Order Condition on ${\cal Z}$] The support of $Z_{i}$ must contain $d$ distinct points, i.e., $\#\left({\cal Z}\right)\geq d = 1+d_{x}+d_{z}$.

Clearly, all of the above are necessary conditions for Assumption (ref):

lem(a) Assumption (ref) implies Conditions (ref) and (ref); (b) Condition (ref) implies Conditions (ref) and (ref).

The no multicollinearity condition and the relevance condition are standard for linear regression models. Condition (ref) requires $Z_{i}$ to be relevant for $X_{i}$ in a nonlinear manner. This requirement of nonlinearity marks the departure of our approach from the standard IV approach which utilizes a linear projection of $X_i$ on $Z_i$.

The requirement of nonlinearity also imposes a restriction on the cardinality of the support of $Z_{i}$ as in Condition (ref). This is because it is always possible to fit a straight line between any two distinct points, and more generally, to fit a linear $d$-dimensional hyperplane across any $d$ distinct points in $\mathbb{R}^d$. Hence, our order condition is on the cardinality of the support of $Z_i$, rather than the number of variables. Of course, if ${\cal Z}$ is a continuum, then the order condition is automatically satisfied.

When there is only one endogenous variable, then the converse of Lemma (ref)(a) is also true, effectively establishing the sufficiency of nonlinearity for point identification.

lem[Sufficient Condition with Scalar $X_{i}$] Suppose that $X_{i}$ is scalar-valued, i.e. $d_{x}=1$. Then, Conditions (ref) and (ref) $\Rightarrow$ Assumption (ref).

Lemma (ref) is particularly relevant when we are primarily worried about the endogeneity of a single treatment status variable $X_{i}$, which is often a discrete random variable. Then, if there exists some exogenous shifter $Z_{i}$ that is relevant for $X_{i}$, $\pi_{0}\left(z\right)$ is naturally nonlinear given the discreteness of $X_{i}$.

example[Linear Treatment Effect Model with Selection] Consider \begin{align*} Y_{i} & =\alpha_{0}+Z_{i}^{'}\beta_{0}+X_{i}\gamma_{0}+\epsilon_{i},\\ X_{i} & =\mathbf{\mathbbm1}\left\{ \varphi_{0}\left(Z_{i}\right)\geq u_{i}\right\}, \end{align*} with $\mathbb{E}\left[\rest{\epsilon_{i}}Z_{i}\right]=0$, $u_{i}\perp Z_{i}$, and $u_{i}\sim F_{u}$. Then, the propensity score function $\pi_{0}\left(z\right):=\mathbb{E}\left[\rest{X_{i}}Z_{i}=z\right] $ is naturally nonlinear in $z$ when $\#(Z_i) \geq 3$, i.e., the support of $Z_i$ contains at least three points. As discussed in the introduction, the order condition $\#(Z_i) \geq 3$ can be satisfied even if $Z_i$ just consists of two dummy variables. Hence, Condition (ref) can be thought as a mild condition in this setting.

Lastly, we note that, when $d_{x}>1$, we not only need each $\pi_{0,k}$ to be nonlinear in $z$, but also need each $\pi_{0,k}$ to be linearly independent (as a function) from 1, $z$, and all other $\left(\pi_{0,j}\right)_{j\neq k}$ as well. We consider this condition relatively mild and easy to verify. Heuristically, whenever ${\cal Z}$ is a continuum, the space of functions on ${\cal Z}$ (under some regularity conditions) can be often viewed as an infinite-dimensional Hilbert space that admits a linear series representation under a certain orthonormal basis of functions $\left(b_{k}\left(\cdot\right)\right)_{k=1}^{\infty}$ on ${\cal Z}$: \[ {\cal F}=\left\{ \sum_{k=1}^{\infty}c_{k}b_{k}\left(\cdot\right):\sum_{k=1}^{\infty}c_{k}^{2}<\infty\right\} . \] Hence, linear independence among a finite number $\left(d=1+d_{x}+d_{z}\right)$ of “generic" functions from $\cal{F}$ seems heuristically as a “generic property".

comment\subsection{Panel Data Models} The identification strategy also applies to the panel data model with fixed effects: \[ Y_{it}=\alpha_0+Z_{it}'\beta_0+X_{it}'\gamma_0+a_i+\epsilon_{it}, \] where $a_i\in \mathbb{R}$ denotes an unobserved fixed effect which can be arbitrarily correlated with the covariates $(Z_{it}, X_{it})$. By taking the difference over two periods, we can cancel out the fixed effects $a_i$, , which yields \[\Delta Y_{it}=\Delta Z_{it}'\beta_0+\Delta X_{it}'\gamma_0+\Delta \epsilon_{it}, \] where $\Delta Y_{it}=Y_{it}-Y_{it-1}$. To identify the coefficient $(\beta_0, \gamma_0)$, traditional methods typically relies on the exogeneity assumption of all covariates $(\Delta Z_{it}, \Delta X_{it})$ or exploit excluded instruments for identification. Our method allows for potential endogeneity of the covariate $\Delta X_{it}$ and requires the following exogeneity assumption of $\Delta Z_{it}$ for identification: \[ \mathbb{E}\left[ \rest{\Delta \epsilon_{it} } {\Delta Z_{it} }\right]=0 \] Similar to Section (ref), we adopt the mean projection of the endogenous covariate $\Delta X_{it}$ on the exogenous regressor $\Delta Z_{it}$, denoted as $\pi_0\left(\Delta Z_{it} \right):=\mathbb{E}\left[ \rest{\Delta X_{it} } {\Delta Z_{it} }\right]$. Then we can run the standard OLS regression of $Y_i$ on $1, \Delta Z_{it}, \pi_0\left(\Delta Z_{it} \right)$ and identify $(\beta_0, \gamma_0)$ as long as all regressors are not multicollinear. One example for the endogenous covariate $X_{it}$ is the lagged dependent variable $Y_{it-1}$ and the differenced model is given as \[ \Delta Y_{it}=\Delta Z_{it}'\beta_0+\Delta Y_{it-1}'\gamma_0+\Delta \epsilon_{it}. \] In this example, the lagged dependent variable $\Delta Y_{it-1}$ is naturally correlated with $\Delta \epsilon_{it}$ since $Y_{it-1}$ depends on the error term $\epsilon_{it-1}$. Our method enables to identify $\gamma_0$ without excluded instruments, by using the regressor $\Delta Z_{it}$ as the included instrument for $\Delta Y_{it-1}$. \subsection{Interpretation under Misspecification} When the true model is nonlinear and unknown, our estimators can be interpreted as the optimal linear approximation within a subspace of the true regression model. Let $\theta_{lim}$ denote the limit of the two proposed estimators $\hat{\theta}, \hat{\theta}^{*}$, given as \[\theta_{lim}=\left(\mathbb{E}\left[W_{i}W_{i}^{'}\right]\right)^{-1}\mathbb{E}\left[W_{i}Y_{i}\right]. \] Consider that the true regression model is presented as follows: \[Y_i=g\left(\tilde{W}_i\right)+\epsilon_i,\] where $\tilde{W}_i:=(1, Z_i', X_i')'$ denote the vector of all regressors, and $g$ is the unknown function that we are interested in identifying and estimating. Given the computational difficulties of nonparametric estimation, we may want to estimate the optimal linear approximation to the true regression model: \[\begin{aligned} \theta_{opt}&:= \arg\min_{\theta}\mathbb{E}\left[ \left(g\left({\tilde{W}_i}\right)-\tilde{W}^{'}\theta \right)^2 \right] \\ &=\left( \mathbb{E}\left[ \tilde{W}_i \tilde{W}_i^{'} \right] \right)^{-1}\mathbb{E}\left[\tilde{W}_i g\left({\tilde{W}_i}\right)\right]. \end{aligned} \] However, this parameter $\theta_{opt}$ is not feasible since $g\left({\tilde{W}_i}\right)$ is not identified without further assumptions. \footnote{ escanciano2021optimal provides identification of $\theta_{opt}$ with an excluded instrument $Z_{i, exc}$ and the assumption that there exists a function $h$ satisfying $ \mathbb{E}\left[ \rest{h(Z_i, Z_{i, exc})} \tilde{W}_i \right]=\tilde{W}_i. $ However, this assumption cannot hold without exclusion instruments, since $ \mathbb{E}\left[ \rest{h(Z_i)} \tilde{W}_i \right]=h(Z_i)\neq \tilde{W}_i. $ } The parameter $\theta_{lim}$ of our approach can be interpreted as the optimal linear approximation within the mean projection space $W_i=\mathbb{E}\left[ \rest{ \tilde{W}_i}{Z_i}\right]$ of the original regressor $\tilde{W}_i$ to the true regression model: \[ \theta_{lim}= \arg\min_{\theta}\mathbb{E}\left[ \left(g\left({\tilde{W}_i}\right)-W_i^{'}\theta \right)^2 \right], \] which uses the fact that $\mathbb{E}\left[ W_i g\left({\tilde{W}_i}\right) \right]=\mathbb{E}\left[W_iY_i \right]-\mathbb{E}\left[ W_i \epsilon_i \right]=0$. The parameter $\theta_{lim}$ is the feasible sub-optimal linear approximation to the true regression $g\left(\tilde{W}_i \right)$, and it is identified without the need for exclusion restrictions or additional assumptions. When the true model is approximately linear: $g\left(\tilde{W}_i \right)\approx \tilde{W}_i'\theta_0$, then the two parameters $\theta_{lim}$ and $\theta_{opt}$ coincide and provide an approximation of the true parameter $\theta_{lim}\approx \theta_{opt}\approx \theta_0$. The standard 2SLS estimator $\theta_{2sls}$ with excluded instruments $Z_{i, exc}$ has a similar interpretation in terms of approximation of nonparametric regression,\footnote{imbens1994identification provides a causal interpretation for the 2SLS estimator with binary treatment and instruments.} when the instrument $Z_{i, all}:=(1, Z_i', Z_{i, exc}')$ is linearly relevant (i.e., $\mathbb{E}\left[\rest{\tilde{W}_i} Z_{i, all} \right]=\Pi_0Z_{i, all}$) and $Z_{i, all}$ has the same dimension with $\tilde{W}_i$. Then the 2SLS can also be interpreted as the optimal linear approximation within the mean/linear space to the true regression model: suppose $\Pi_0$ is invertible, \[\begin{aligned} \theta_{2sls}&:=\left(\mathbb{E}\left[ Z_{i, all} \tilde{W}_{i}' \right] \right)^{-1}\mathbb{E}\left[ Z_{i, all} Y_i \right] \\ &=\arg\min_{\theta} \mathbb{E}\left[ \left(g\left({\tilde{W}_i}\right)-\left(\Pi_0Z_{i, all}\right)' \theta\right)^2 \right]. \end{aligned} \] When the linear model is misspecified, it is unclear whether $\theta_{lim}$ or $\theta_{2sls}$ is closer to the true regression model $g\left({\tilde{W}_i}\right)$, which depends on the joint distribution of $(X_i, Z_i, Z_{i, exc}, Y_i)$ and the true functional form of $g$.

Discussion

Comparison: IV regression with Excluded Instrument

The canonical IV approach utilizes a linear projection of the endogenous regressor on the exogenous regressors. This approach requires the presence of an excluded instrument, as otherwise all regressors will exhibit perfect multicollinearity. In principle, this method can achieve nonparametric identification under the completeness condition and is robust to the misspecification of function forms.\footnote{In practice, however, nonparametric IV estimation is not commonly used, partly due to the computational difficulties and inference complexities.}

In contrast, our approach exploits a mean projection of the endogenous regressor on the included exogenous regressor, which allows us to extract more information for identification through the nonlinear dependence between the exogenous variable and the endogenous variable. Our approach relies on a parametric (e.g., linear) assumption on the functional form, but it enables identification without exclusion restrictions. We believe our method could be a viable alternative to the standard IV approach in situations where there is a natural parametric specification, and where finding excluded IVs is challenging.

Comparison: Heckman Correction Approach

The conventional Heckman correction approach can also achieve identification without exclusion restrictions, under distributional assumptions or parametric functions. This approach typically focuses on a binary endogenous regressor and examines the following specification: \[

alignedY_{i} & =\alpha_{0}+Z_{i}^{'}\beta_{0}+X_{i}\gamma_{0}+\epsilon_{i},\\ X_{i} & =\mathbf{\mathbbm1}\left\{ Z_{i}'\eta_0 \geq u_{i}\right\}, \\ (\epsilon_i, u_i)' & \sim \mathcal{N}\left((0, 0)', (1, \rho_0; \rho_0, 1) \right).

\]

Under the joint distribution of the two error terms $(\epsilon_i, u_i)$, it yields the following conditional moment restriction: \[ \mathbb{E}[\rest{Y_i}{X_i, Z_i}]=\alpha_0+Z_i'\beta_0+X_i'\gamma_0+\rho_0 \frac{\phi(Z_{i}'\eta_0)}{\Phi(Z_{i}'\eta_0)}, \] which can identify $(\theta_0, \rho_0)$ without exclusion restrictions. The Heckman correction approach can be extended to nonbinary and multi-dimensional endogenous regressor $X_i$. We can still look at the conditional expectation of $Y_i$ given all regressors $(X_i, Z_i)$: \[ \mathbb{E}[\rest{Y_i}{X_i, Z_i}]=\alpha_0+Z_i'\beta_0+X_i'\gamma_0+\mathbb{E}[\rest{\epsilon_i}{X_i, Z_i}]. \] If a parametric form on the selection bias term $\mathbb{E}[\rest{\epsilon_i}{X_i, Z_i}]$ is imposed as follows: \[\mathbb{E}[\rest{\epsilon_i}{X_i, Z_i}]=s(X_i, Z_i, \eta_0), \] and this function $s$ is nonlinear in $(X_i, Z_i)$, then the coefficient $\theta_0$ is identified.

The Heckman correction approach exploits the mean projection of the error $\epsilon_i$ on all regressors $(X_i, Z_i)$. To achieve identification, this approach requires a parametric form (or parametric distributions of errors) for $s$ as well as the nonlinearity of $s$. However, since the function $s$ involves the unobserved error term $\epsilon_i$, its nonlinearity cannot be directly tested.

In contrast, our approach applies the mean projection of the endogenous $X_i$ on the exogenous $Z_i$. The identification relies on the nonlinearity of the function $\pi_0(Z_i)=\mathbb{E}[\rest{X_i}{Z_i}]$, but does not require further functional form assumption on $\pi_0$. Moreover, the function $\pi_0$ only depends on observed variables $(X_i, Z_i)$, making its nonlinearity a testable condition.

Relationship to the Instrumental Function Approach

We establish identification of $\theta_0$ from the viewpoint of a standard linear regression model, while treating $\pi_0(Z_i)$ as an exogenous regressor. Point identification is then obtained under the no-multicollinearity condition, which translates into a nonlinearity requirement on $\pi_0$. Another approach to address endogeneity is to use a (known) nonlinear function $g(Z_i)$ as an instrument for the endogenous regressor $X_i$. In this section, we will discuss the connections between our identification approach and this alternative approach.

To illustrate, consider the case where both $Z_i$ and $X_i$ are scalar variables. Using $(1, Z_i, g(Z_i))$ as IVs, we can obtain the following moment restrictions for $\theta_0$: \[ \mathbb{E}\left[(Y_i-\alpha_0-\beta_0Z_i-\gamma_0X_i)\left(

array[array omitted — 33 chars of source]

\right) \right]=0.\] The parameter $\theta_0$ is identified from the above equation if \[H_g:= \left[

array[array omitted — 181 chars of source]

\right] =\left[

array[array omitted — 201 chars of source]

\right] \] has full rank, which depends on the functional form of $\pi_0$ and the choice of $g$.

Clearly, a necessary condition for the full rank requirement is the nonlinearity of $\pi_0$. Otherwise, the third column of $H_g$ will be a linear combination of the first two columns,\footnote{Writing $\pi_0(Z_i) = a + bZ_i$, we have $\mathbb{E}[\pi_0(Z_i)] = a + b\mathbb{E}[Z_i]$, $\mathbb{E}[\pi_0(Z_i)Z_i] = a\mathbb{E}[Z_i] + b\mathbb{E}[Z_i^2]$, and $\mathbb{E}[\pi_0(Z_i)g(Z_i)] = a\mathbb{E}[g(Z_i)] + b\mathbb{E}[Z_i g(Z_i)]$.} and hence $H_g$ cannot have full rank, irrespective of the choice of $g$. Our identification results thus make explicit the dependence of the identifiability on the nonlinearity of the $\pi_0$ function.

Moreover, it is worth noting that not all nonlinear functions can serve as valid IVs for $X_i$ in the sense of satisfying the full rank condition on $H_g$. For example, if $Z_i \sim \mathcal{N}(0,1)$, $\pi_0(z)=z^3$, and $g(z)=z^2$, then $H_g$ has deficient rank, and thus $Z_i^2$ is not a valid IV.

Our identification results in Theorem (ref) can be interpreted as using $\pi_0(Z_i)$ as an instrument for the endogenous regressor $X_i$, which is an unknown function that can be identified from data. As shown in chamberlin1987asymptotic and newey1990efficient, $\pi_0$ is in fact the optimal instrument under homoskedasticity. The identification results in our paper, Lemma (ref) in particular, further imply that, with $\pi_0(Z_i)$ used as the IV, the nonlinearity of $\pi_0$ becomes sufficient for the full rank condition. Therefore, $\pi_0(Z_i)$ is not only the instrumental function that minimizes the asymptotic variance under homoskedasticity, but also the instrumental function that requires the minimum assumption for identification.

Estimation and Inference

Semiparametric Estimators $\hat{\theta}$ and $\hat{\theta}^*$

Based on our identification result, we propose the following two semiparametric estimators: \[

aligned\hat{\theta} & =\left(\frac{1}{n}\sum_{i=1}^{n}\hat{W}_{i}\hat{W}_{i}^{'}\right)^{-1}\frac{1}{n}\sum_{i=1}^{n}\hat{W}_{i}Y_{i},\\ \hat{\theta}^{*} & =\left(\frac{1}{n}\sum_{i=1}^{n}\hat{W}_{i}\hat{W}_{i}^{'}\right)^{-1}\frac{1}{n}\sum_{i=1}^{n}\hat{W}_{i}\hat{h}\left(Z_{i}\right).

\] We now lay out the regularity conditions for the $\sqrt{n}$-consistency and asymptotic normality of $\hat{\theta}$ and $\hat{\theta}^{*}$. The first one is a standard one on the existence of moments.\footnote{We impose this assumption on the fourth moment for subsequent variance estimation.}

assumption[Finite Fourth Moments] $\mathbb{E}\left|\epsilon_{i}\right|^{4},\mathbb{E}\norm{X_{i}}^{4}$, and $\mathbb{E}\norm{Z_{i}}^{4}$ are finite.

Below we give some high-level conditions about the first-stage nonparametric regressions, which can be satisfied with a wide variety of lower-level conditions and many types of nonparametric estimators. See, for example, \citet*{newey1994large} and \citet*{chen2007large} for more information.

assumption[Smoothness and Nonparametric Convergence] Suppose that: \begin{itemize} • $h_{0},\pi_{0}\in{\cal H}$, where ${\cal H}$ is a Sobolev function space of order $s>\frac{d_{z}}{2}$ on ${\cal Z}$. • The nonparametric estimators $\hat{h}$ and $\hat{\pi}$ belong to ${\cal H}$ (with probability approaching 1) and are asymptotically linear. • The nonparametric estimators $\hat{h}$ for $h_{0}$ and $\text{\ensuremath{\hat{\pi}}}$ for $\pi_{0}$ converge in $L_{2}\left(Z\right)$-norm faster than the $n^{-1/4}$ rate: $\norm{\hat{h}-h_{0}}_{L_{2}\left(Z\right)}=o_{p}\left(n^{-\frac{1}{4}}\right)$, $\norm{\hat{\pi}-\pi_{0}}_{L_{2}\left(Z\right)}=o_{p}\left(n^{-\frac{1}{4}}\right)$. \end{itemize}
thm[Asymptotic Normality] Under Assumptions (ref) - (ref), we have: \begin{align*} \sqrt{n}\left(\hat{\theta}-\theta_{0}\right)\overset{d}{\longrightarrow}\mathcal{N}\left({\bf 0},V_{0}\right) & ,\quad\sqrt{n}\left(\hat{\theta}^{*}-\theta_{0}\right)\overset{d}{\longrightarrow}\mathcal{N}\left({\bf 0},V_{0}\right), \end{align*} with $$ V_{0}:= \mathbb{E}\left[W_{i}W_{i}^{'}\right]^{-1} \mathbb{E}\left[\epsilon_{i}^{2}W_{i}W_{i}^{'}\right] \mathbb{E}\left[W_{i}W_{i}^{'}\right]^{-1}. $$ The assumptions about $h_{0}$ and $\hat{h}$ can be dropped for $\hat{\theta}$ since it does not involve $\hat{h}$. Also, if $\mathbb{E}[\rest{\epsilon_i^2} Z_i]\equiv\sigma^2_\epsilon$ (errors are homoskedastic), $V_0$ simplifies to $ \sigma^2_\epsilon \mathbb{E}\left[W_{i}W_{i}^{'}\right]^{-1}$.

The asymptotic variance can then be easily estimated via standard plug-in methods as in the following theorem. Based on the standard error estimates, confidence intervals and various test statistics can be computed in the standard manner.

thm[Variance Estimation] Let $\hat{\epsilon}_{i}:=Y_{i}-\hat{\alpha}-Z_{i}^{'}\hat{\beta}-X_{i}^{'}\hat{\gamma},$ and \begin{align*} \hat{\Sigma}:=\frac{1}{n}\sum_{i=1}^{n}\hat{W}_{i}\hat{W}_{i}^{'},\quad & \hat{\Omega}:=\frac{1}{n}\sum_{i=1}^{n}\hat{\epsilon}_{i}^{2}\hat{W}_{i}\hat{W}_{i}^{'},\quad \hat{V}:=\hat{\Sigma}^{-1}\hat{\Omega}\hat{\Sigma}^{-1}. \end{align*} With homoskedasticity, $\hat{V}:= \left(\frac{1}{n}\sum_{i=1}^{n}\hat{\epsilon}_{i}^{2}\right)\hat{\Sigma}^{-1}$. Under Assumptions (ref) - (ref), $\hat{V}\overset{p}{\longrightarrow} V_{0}.$

Discretization-Based Estimator $\hat{\theta}_{disc}$

The two estimators $\hat{\theta}$ and $\hat{\theta}^{*}$ we proposed before both involve nonparametric regressions in the first stage. Alternatively, we propose a third estimator $\hat{\theta}_{disc}$ that does not require any nonparametric regression at all.

Specifically, let $\left({\cal Z}_{k}\right)_{k=1}^{K}$ be a partition of ${\cal Z}$ with $K$ being a finite and fixed number such that $K\geq d=1+d_{z}+d_{x}$. To rule out redundant cells, we require each cell to have a positive probability.

assumption[Positive Probabilities] $p_k := \mathbb{P}\left(Z_{i}\in{\cal Z}_{k}\right)>0$ for $k=1,...,K$.

Define the dummy variable for each of the $K$ partition cells as $D_{i,k}:=\mathbf{\mathbbm1}\left\{ Z_{i}\in{\cal Z}_{k}\right\}$, and write $D_{i}:=\left(D_{i1},...,D_{iK}\right)^{'}$. We can then use $D_{i}$ as IVs to identify and estimate $\theta_0$ based on the following transformation of equation (ref):

equation*[equation* omitted — 185 chars of source]

Since $Z_{i}$ is averaged out within each partition cell, there is some information loss, and the no-multicollinearity condition for identification of $\theta_{0}$ in Assumption (ref) needs to be strengthened to a partitional version. To state the condition in a “lower-level” form, write $\ol Z_{k}:=\mathbb{E}\left[\rest{Z_{i}}Z_{i}\in{\cal Z}_{k}\right]$, $\ol X_{k}:=\mathbb{E}\left[\rest{X_{i}}Z_{i}\in{\cal Z}_{k}\right]$, and $\ol W_{k}:=\left(1,\ol Z_{k}^{'},\ol X_{k}^{'}\right)^{'}$.

assumption[No Partitional Multicollinearity] Suppose that $\left(1,\ol Z_{k}^{'},\ol X_{k}^{'}\right)$ are not multicollinear across $k=1,...,K$, or equivalently, \[ \text{rank}\left(\begin{array}{ccc} 1 & \ol Z_{1}^{'} & \ol X_{1}^{'}\\ \vdots & \vdots & \vdots\\ 1 & \ol Z_{K}^{'} & \ol X_{K}^{'} \end{array}\right)=d. \]

We note that Assumption (ref) translates into the following standard full-rank condition written in terms of expectations (i.e., probability-weighted sums under discreteness), provided that each cell has a strictly positive probability.

lemSuppose that Assumption (ref) holds. Then Assumption (ref) holds if and only if $\sum_{k=1}^{K}p_k\ol W_{k}\ol W_{k}^{'}$ is invertible.

Note that a necessary condition for Assumption (ref) is the order condition $K\geq d$ already mentioned above. It is also easy to verify that Assumption (ref) implies Assumption (ref), but the converse is not generally true. However, Assumption (ref) still remains as a condition nonparametrically identified from the observable distribution of data.

We can then construct $\hat{\theta}_{disc}$ as the standard two-stage least square (2SLS) estimator with the $K$-dimensional vector $D_{i}$ as instruments. Formally, write $\tilde{W}_{i}:=\left(1,Z_{i}^{'},X_{i}^{'}\right)^{'}$, and let $Y,D,\tilde{W}$ denote the vector/matrix concatenation of the variables across all $i=1,...,n$, and each row of which contains $Y_{i},D_{i}^{'},\tilde{W}_{i}^{'}$, respectively. Then

align[align omitted — 126 chars of source]

where $P_{D}:=D\left(D^{'}D\right)^{-1}D^{'}$. Since $D$ consists of partition cell dummies, the projection matrix $P_{D}$ is essentially computing cell-wise averages, and thus $\hat{\theta}_{disc}$ can be equivalently written as

equation[equation omitted — 191 chars of source]

where $\hat{p}_k := \frac{n_k}{n}$, $n_{k}:=\sum_{i=1}^{n}D_{ik}$, $\hat{\mu}_{y,k}:=\frac{1}{n_{k}}\sum_{i=1}^{n}D_{ik}Y_{i},$ $\hat{\mu}_{z,k}:=\frac{1}{n_{k}}\sum_{i=1}^{n}D_{ik}Z_{i}$, $\hat{\mu}_{x,k}:=\frac{1}{n_{k}}\sum_{i=1}^{n}D_{ik}X_{i}$, and $\hat{\ol W}_{k}:=\left(1,\hat{\mu}_{z,k}^{'},\hat{\mu}_{x,k}^{'}\right)^{'}$.

Clearly, $\hat{\theta}_{disc}$ is very easy to compute. Researchers may use any standard 2SLS command with $D_{i}$ as IVs (or with one of $D_{ik}$'s dropped if the constant is included), which yields equivalent results as (ref) and (ref). The asymptotic distribution of $\hat{\theta}_{disc}$ is derived as follows:

thm[Asymptotic Normality of $\hat{\theta}_{disc}$] Under Assumptions (ref), (ref), (ref), and (ref), $\sqrt{n}\left(\hat{\theta}_{disc}-\theta_{0}\right)\overset{d}{\longrightarrow}\mathcal{N}\left({\bf 0},V_{0,disc}\right)$ with \begin{equation} V_{0,disc}:=\left(\sum_{k=1}^{K}p_{k}\ol W_{k}\ol W_{k}^{'}\right)^{-1}\sum_{k=1}^{K}p_{k}\ol{\sigma}_{\epsilon,k}^{2}\ol W_{k}\ol W_{k}^{'}\left(\sum_{k=1}^{K}p_{k}\ol W_{k}\ol W_{k}^{'}\right)^{-1} \end{equation} where $\ol{\sigma}_{\epsilon,k}^{2}:=\mathbb{E}\left[\rest{\epsilon_{i}^{2}}Z_{i}\in{\cal Z}_{k}\right].$ Furthermore, a consistent estimator $\hat{V}_{disc}$ for $V_{0,disc}$ can be constructed by plugging $\hat{p}_{k}:=\frac{n_{k}}{n}$ in place of $p_{k}$, $\hat{\ol W_{k}}$ in place of $\ol W_{k}$, and $\hat{\ol{\sigma}}_{\epsilon,k}^{2}:=\frac{1}{n_{k}}\sum_{i:Z_{i}\in{\cal Z}_{k}}(Y_{i}-\tilde{W}_{i}^{'}\hat{\theta}_{disc})^{2}$ in place of $\ol{\sigma}_{\epsilon,k}^{2}$ in the formula (ref) above. Under homoskedasticity $\mathbb{E}[\rest{\epsilon_i^2} Z_{i}\in{\cal Z}_{k}]\equiv\sigma^2_\epsilon$, the asymptotic variance simplifies to $V_{0,disc} = \sigma^2_\epsilon \left(\sum_{k=1}^{K}p_{k}\ol W_{k}\ol W_{k}^{'}\right)^{-1}$ with $\hat{\sigma}^2_\epsilon := \frac{1}{n}\sum_{i=1}^n(Y_{i}-\tilde{W}_{i}^{'}\hat{\theta}_{disc})^{2}$ consistent for $\sigma^2_\epsilon$.

Since $\hat{\theta}_{disc}$ is constructed based on averages over each partition cell ${\cal Z}_{k}$, there is in general some information loss, and thus $\hat{\theta}_{disc}$ tends to be less efficient than $\hat{\theta}$ and $\hat{\theta}^*$. Our next result formalizes this efficiency loss in the setting where $\epsilon_i$ is homoskedastic.

thmSuppose that $\mathbb{E}[\rest{\epsilon_i^2} Z_i]\equiv\sigma^2_\epsilon$. Then $V_{0, disc} - V_0$ is positive semi-definite.

Despite the efficiency loss, the discretization-based estimator $\hat{\theta}_{disc}$ is very simple and user-friendly. Applied researchers just need to create dummy variables for a chosen partition of ${\cal Z}$ and run a standard 2SLS command. Furthermore, in our simulations, we find that $\hat{\theta}_{disc}$ performs surprisingly well in finite sample: the efficiency loss of $\hat{\theta}_{disc}$ tends to be quite small and more than compensated by its smaller finite-sample bias as a 2SLS estimator that does not require nonparametric regressions.

Furthermore, in the special case where $Z_{i}$ are discrete variables with finite support, there is clearly no information loss from discretization. In this case, expectations simplify to weighted sums over the $K$ realizations of $Z_i$ (weighted by the probability mass $p_k$), and it can be easily shown that the asymptotic variance of $\hat{\theta}_{disc}$ coincides with the one for $\hat{\theta}$ and $\hat{\theta}^{*}$ in Theorem (ref) (without the homoskedasticity assumption).

corSuppose that ${\cal Z}=\left\{ z_{1},...,z_{K}\right\} $ for some finite $K$. Then, under the element-by-element partition, i.e., ${\cal Z}_{k}:=\left\{ z_{k}\right\}$, we have $V_{0,disc}=V_{0}.$

While the established results for $\hat{\theta}_{disc}$ hold for any choice of partition $({\cal Z}_k)_{k=1}^K$ that satisfy Assumptions (ref) and (ref), it is recommended in practice to choose ${\cal Z}_k$ in such a way that the cell probabilities $p_k$ are comparable in magnitude. To illustrate, consider a simple example where we partition ${\cal Z}$ into $K = 10$ cells with $p_1 = 0.91$ but $p_2=...=p_{10} = 0.01.$ Despite Assumption (ref) holds (so that $\theta_0$ is identified), the estimator $\hat{\theta}_{disc}$ is likely to perform badly, since there are only a few observations in cell $2,...,K$ and thus the sample average estimation in those cells could be highly imprecise. Moreover, since $p_2,...,p_{10}$ are close to zero, the smallest eigenvalue of $\sum_{k=1}^{K}p_{k}\ol W_{k}\ol W_{k}^{'}$ may be close to zero (unless the corresponding $\ol{W}_k$'s are very large in magnitude, so that the product terms $p_k\ol W_{k}\ol W_{k}^{'}$ stay comparable across $k$). Since the estimator $\hat{\theta}_{disc}$ is based on the inverse of $\sum_{k=1}^{K}p_{k}\ol W_{k}\ol W_{k}^{'}$, its variance can be large because of the imbalance between $p_1$ and $p_2,...,p_{10}$.

If $Z_i$ is a scalar, a natural strategy would be to choose the partition $({\cal Z}_k)$ to be the $K$ equally-sized quantile ranges, which would ensure that $p_k \equiv 1/K$ (or at least asymptotically so when sample quantiles are used in finite samples). If $Z_i$ is vector, one could work with (empirical) vector quantiles as developed relatively recently in the literature based on the theory of optimal transport: see galichon2018optimal for an introduction, and, e.g., chernozhukov2017monge, hallin2021distribution, and ghosal2022multivariate for detailed discussions. Alternatively, one could start with a partition of the support of $Z_i$ obtained as products of partitions in each dimension of $Z_i$, and adjust and/or merge certain cells (if necessary) to ensure that the sample proportions of observations in each cell are comparable across $k = 1,...,K$.

We emphasize again that our results above apply for any choice of the partition $({\cal Z}_k)$ as long as Assumptions (ref) and (ref) are satisfied. Hence, while we provide some suggestions for the choice of partitions above, there may be more appropriate partition choices depending on the specific applications and contexts.

comment\begin{proof} When $Z_i$ is discrete and we use element-wise partition ${\cal Z}_{k}=z_k$, then we have $\bar{W}_k=W_k:=(1, z_k', \mathbb{E}[\rest{X_i} Z_i=z_k]')'$. The variance $V_{0, disc}$ becomes \[ V_{0,disc}:=\left(\sum_{k=1}^{K}p_{k} W_{k} W_{k}^{'}\right)^{-1}\sum_{k=1}^{K}p_{k}\ol{\sigma}_{\epsilon,k}^{2} W_{k} W_{k}^{'}\left(\sum_{k=1}^{K}p_{k} W_{k} W_{k}^{'}\right)^{-1}, \] where $\ol{\sigma}_{\epsilon,k}^{2}=\mathbb{E}[\rest{\epsilon_i^2} Z_i=z_k]$. Recall that the variance $V_0$ is given as \[ V_{0}=\left(\mathbb{E}\left[W_{i}W_{i}^{'}\right] \right)^{-1} \mathbb{E}\left[\epsilon_{i}^{2}W_{i}W_{i}^{'}\right] \left(\mathbb{E}\left[W_{i}W_{i}^{'}\right] \right)^{-1}. \] The first term $\mathbb{E}\left[W_{i}W_{i}^{'}\right]$ can be expressed as \[\begin{aligned} \mathbb{E}\left[W_{i}W_{i}^{'}\right]&=\sum_{k=1}^K\mathbb{E}\left[ \rest{W_{i}W_{i}^{'}} Z_{i}=z_{k} \right] p_k \\ &=\sum_{k=1}^K p_k W_k W_k'. \end{aligned} \] Similarly, the middle term $ \mathbb{E}\left[\epsilon_{i}^{2}W_{i}W_{i}^{'}\right]$ is given as \[\begin{aligned} \mathbb{E}\left[\epsilon_{i}^{2}W_{i}W_{i}^{'}\right]&=\sum_{k=1}^K\mathbb{E}\left[\rest{ \epsilon_{i}^{2}W_{i}W_{i}^{'}} Z_{i}=z_{k} \right] p_k \\ &=\sum_{k=1}^K p_kW_kW_k' \mathbb{E}\left[\rest{ \epsilon_{i}^{2} } Z_{i}=z_{k} \right] \\ &=\sum_{k=1}^K p_k\ol{\sigma}_{\epsilon,k}^{2}W_kW_k'. \end{aligned} \] Therefore, $V_0=V_{0, disc}$. \end{proof}

Extension: Nonlinear and Quantile Regressions

The identification strategy can be extended to analyze the following endogenous nonlinear regression models without exclusion restrictions: \[ Y_{i}=f(Z_{i}, X_i, \theta_{0})+\epsilon_{i}, \] where $X_{i}$ is a $d_{x}$-dimensional endogenous regressor, $Z_{i}$ is a $d_{z}$-dimensional exogenous regressor satisfying Assumption (ref), and the function $f$ is known up to the $d$-dimensional parameter $\theta_0$. The function $f$ can be nonlinear and nonseparable in the covariates $Z_i$ and $X_i$.

To identify $\theta_0$, we adopt a similar strategy by projecting the entire functional term $f(Z_{i}, X_i, \theta_{0})$ onto the included instrument $Z_i$. Let the function $m_0$ be defined as \[m_0(Z_i, \theta_0):=\mathbb{E}[\rest{f(Z_{i}, X_i, \theta_{0})}Z_i], \] which is identified up to the parameter $\theta_0$. Then under Assumption (ref) (exogeneity) of the included instrument $Z_i$, we have the following conditional moment condition: \[\mathbb{E}[\rest{Y_i-m_0(Z_i, \theta_0)}Z_i]=0.\] The above moment condition can be viewed as the moment restriction of the standard nonlinear regression model without endogeneity, while treating $m_0(Z_i, \theta_0)$ as the nonlinear regressor. The key distinction is that the function $m_0$ needs to be estimated. For standard nonlinear regression, local identification of $\theta_0$ can be attained under the following condition.

thmSuppose that Assumption (ref) holds, the function $m_0(z, \cdot)$ is continuously differentiable for any $z$, and $\mathbb{E}\left[\nabla_\theta m_0(Z_i, \theta_0)\nabla_{\theta^{'}} m_0(Z_i, \theta_0)\right]$ has full rank, then $\theta_0$ is locally identified.

Local identification of $\theta_0$ ensures that there exists a neighborhood $\Theta_0$ of $\theta_0$ on which $\theta_0$ is identified. This result can be expanded to achieve global identification under additional assumptions, by invoking the global inversion theorem in ambrosetti1995primer (Chapter 3, Theorem 1.8). In the case of general nonlinear regressions, the interpretation of the full-rank condition depends on the specific functional form of $f(Z_{i}, X_i, \theta_{0})$. This feature also applies to the standard IV regression with excluded instruments, where the identification conditions are contingent upon the specification of $f$ as well.

remarkWe focus on the nonlinear regression model, while the analysis also applies to a more general parametric model. Consider that we have the following moment condition: \[ \tilde{m}_0(Z_i, \theta_0):= \mathbb{E}[ \rest{g\left(Y_i, Z_i, X_i, \theta_0\right) }Z_i ]=0, \] where the function $g$ is known up to the parameter $\theta_0$ and $g$ can be nonlinear and nonseparable in all variables $(Y_i, Z_i, X_i)$. The moment function $g$ may be derived from structural models, which is naturally nonlinear in all variables. In terms of the nonlinear regression model, the function $g$ is given as $g\left(Y_i, Z_i, X_i, \theta_0\right)=Y_i-f(Z_{i}, X_i, \theta_{0})$. Following Theorem (ref), the parameter $\theta_0$ is locally identified if $\mathbb{E}\left[\nabla_\theta \tilde{m}_0(Z_i, \theta_0)\nabla_{\theta^{'}} \tilde{m}_0(Z_i, \theta_0)\right]$ has full rank. Section (ref) explores the endogenous quantile regression model, where the moment condition is nonseparable in all observed variables $(Y_i, Z_i, X_i)$.

Similar to linear regression models, we propose two semiparametric two-step nonlinear regression estimators. In the first step, nonparametrically regress $f(X_i, Z_i, \theta)$ on $Z_i$ for each $\theta$ and get the predicted value $\hat{m}(Z_i, \theta)$; nonparametrically regress $Y_i$ on $Z_i$ and get the predicted value $\hat{h}(Z_i)$. In the second step, run the standard nonlinear regression using $Y_i$ and $\hat{h}(Z_i)$ as the dependent variable, respectively: \[

aligned\hat{\theta}_{nl}&=\arg\min_{\theta \in \Theta_0} \frac{1}{n}\sum_i \left(Y_i- \hat{m}(Z_i, \theta) \right)^2, \\ \hat{\theta}^{*}_{nl}&=\arg\min_{\theta \in \Theta_0 } \frac{1}{n}\sum_i \left(\hat{h}(Z_i)- \hat{m}(Z_i, \theta) \right)^2.

\] Similar to Section (ref), the two estimators $\hat{\theta}_{nl}, \hat{\theta}^{*}_{nl}$ are $\sqrt{n}$-consistent and have the same asymptotic variance: \[\sqrt{n} \left( \hat{\theta}_{nl}-\theta_0 \right) \overset{d}{\longrightarrow} \mathcal{N}\left({\bf 0},V_{0}\right), \quad \sqrt{n} \left( \hat{\theta}_{nl}^{*}-\theta_0 \right) \overset{d}{\longrightarrow} \mathcal{N}\left({\bf 0},V_{0}\right), \] with \[ V_0:=M_0^{-1} \mathbb{E}\left[ \epsilon_i^2 \nabla_{\theta} m_0(Z_i, \theta_0) \nabla_{\theta'} m_0(Z_i, \theta_0) \right] M_0^{-1}, \] where $M_0=\mathbb{E}\left[\nabla_\theta m_0(Z_i, \theta_0)\nabla_{\theta^{'}} m_0(Z_i, \theta_0)\right]$. A consistent estimator $\hat{V}$ for the variance matrix $V_0$ can be developed by replacing $\epsilon_i$ with $Y_i-f(Z_i, X_i, \hat{\theta}_{nl})$, $\nabla_{\theta} m_0(Z_i, \theta_0)$ with its estimator $\nabla_{\theta} \hat{m}(Z_i, \hat{\theta}_{nl})$, and expectation with the sample mean.

Next, we investigates endogenous quantile regression as an illustration of Theorem (ref).

Endogenous Quantile Regressions

We study the following endogenous quantile regression model: \[

aligned&Y_{i}=\alpha_{0}+Z_{i}^{'}\beta_{0}+X_{i}^{'}\gamma_{0}+\epsilon_{i}, &\ Quan_{\tau}(\rest{\epsilon_i} Z_i)=0,

\] where $\text{Quan}_{\tau}(\rest{\epsilon_i} Z_i)$ denotes the $\tau$-th quantile of the conditional distribution of $\epsilon_i$ given $Z_i$. In this example, $X_i$ is the potentially endogenous regressor and $Z_i$ is the exogenous regressor that satisfies the conditional quantile restriction. We still study the identification and estimation of the coefficient $\theta_0$ using only the included instrument $Z_i$.

By the quantile exogeneity of $Z_i$, it yields the conditional moment restriction as follows: \[ \mathbb{E}\left[\rest{\mathbf{\mathbbm1}\{Y_i\leq \alpha_{0}+Z_{i}^{'}\beta_{0}+X_{i}^{'}\gamma_{0}\}-\tau} Z_i \right]=0. \] The moment condition above is naturally nonlinear and nonseparable in all variables $(Y_i, Z_i, X_i)$. We project the whole indicator term on $Z_i$ and define the function $m_0$ as \[m_0(Z_i, \theta_0):=\mathbb{E}\left[\rest{\mathbf{\mathbbm1}\{Y_i\leq \alpha_{0}+Z_{i}^{'}\beta_{0}+X_{i}^{'}\gamma_{0}\}} Z_i \right]. \] According to Theorem (ref), the coefficient $\theta_0$ is locally identified if $\mathbb{E}\left[\nabla_\theta m_0(Z_i, \theta_0)\nabla_{\theta^{'}} m_0(Z_i, \theta_0)\right]$ has full rank. Lemma (ref) presents an alternative condition that is equivalent to this full rank condition, making it easier to interpret.

assumption[Continuous Errors] The error term $\epsilon_{i}$ conditional on $(x, z)$ is continuously distributed with the density function $f_{\epsilon\mid X, Z}\left(\rest{\epsilon}x,z\right)$.

Assumption (ref) is a standard assumption that simplifies the calculation of $\nabla_{\theta}m_0 \left(Z_i,\theta_0\right)$.

lemSuppose that Assumption (ref) holds and $f_{\rest{\epsilon}{Z}}(\rest{0}z)>0$ for any $z\in \cal Z$, then $\mathbb{E}\left[\nabla_\theta m_0(Z_i, \theta_0)\nabla_{\theta^{'}} m_0(Z_i, \theta_0)\right]$ has full rank if and only if $1, Z_i$, and $\tilde{\pi}_0\left(Z_i \right):=\mathbb{E}\left[\rest{X_i} Z_i, \epsilon_i=0\right]$ are not multicollinear.

In the case where the endogenous regressor $X_i$ is a scalar, the full-rank condition is equivalent to the nonlinearity of $\tilde{\pi}_0\left(Z_i \right)$. This nonlinearity condition is analogous to the nonlinearity requirement of $\pi_0\left(Z_i \right)$ in the linear regression model, except it is also conditional on $\epsilon_i=0$. Similarly, this nonlinearity relationship naturally arises when the endogenous regressor $X_i$ is binary or discrete.

Based on the identification results, a natural two-step quantile regression estimator $\hat{\theta}_{q}$ can be obtained: in the first step, nonparametrically regress $\mathbf{\mathbbm1}\{Y_i\leq \alpha+Z_{i}^{'}\beta +X_{i}^{'}\gamma\}$ on $Z_i$ for each $\theta$ and compute $\hat{m}(Z_i, \theta)$; in the second step, obtain the quantile estimator $\hat{\theta}_q$ as \[ \hat{\theta}_{q}:=\arg\min_{\theta\in \Theta_0} \frac{1}{n}\sum_i \left(\hat{m}(Z_i, \theta)-\tau \right)^2. \]

Writing $S_i:=f_{\epsilon\mid Z}\left(\rest 0 Z_i\right)(1, Z_i', \tilde{\pi}_0\left(Z_i \right)' )'$, the asymptotic distribution of the two-step quantile estimator $\hat{\theta}_q$ is given by \[\sqrt{n} \left( \hat{\theta}_{q}-\theta_0 \right) \overset{d}{\longrightarrow} \mathcal{N}\left({\bf 0},V_{0, q}\right), \] with $ V_{0, q}:=\left( \mathbb{E}\left[S_i S_i' \right] \right)^{-1} \mathbb{E}\left[(\mathbf{\mathbbm1}\{\epsilon_i\leq 0 \}-\tau )^2 S_i S_i' \right] \left( \mathbb{E}\left[S_i S_i' \right] \right)^{-1} =\tau (1-\tau) \left( \mathbb{E}\left[ S_i S_i' \right] \right)^{-1}.$

In contrast to the standard quantile regression without endogeneity, our approach allows for potential endogeneity in covariate $X_i$. While chernozhukov2005iv examines endogenous quantile models with excluded instruments, the key distinction is that our method establishes identification of $\theta_0$ using solely included regressors. Our approach can be viewed as leveraging the derivative term $\nabla_{\theta} m_0(Z_i, \theta_0)$ as an instrumental function, which is more informative than using $Z_i$ as an instrument since it exploits the dependence between the indicator term $\mathbf{\mathbbm1}\{Y_i\leq \alpha_{0}+Z_{i}^{'}\beta_{0}+X_{i}^{'}\gamma_{0}\}$ and the included regressor $Z_i$. Thus, our approach enables identification without exclusion restrictions. On the other hand, our method does require a parametric specification for $Y_i$, whereas chernozhukov2005iv allows for nonparametric identification with excluded instruments.

Simulation

This section examines the finite sample performances of $\hat{\theta}, \hat{\theta}^*$, and $\hat{\theta}_{disc}$, the three estimators proposed in Section (ref). We compare their performance with both the standard 2SLS estimator $\hat{\theta}_{2sls}$, which treats the included regressor $Z_i$ as an excluded instrument, and the OLS estimator $\hat{\theta}_{ols}$ obtained by regressing $Y_i$ on $(1, Z_i', X_i')$. We report four finite-sample performance measures for every estimator: “Bias”, “SD" (standard deviation), “RMSE" (root mean squared error), and “CP" (coverage probability of 95% confidence interval). The confidence intervals are constructed using the standard $\pm 1.96\times\text{SE}$ formula, where the standard error estimates SE are obtained based on the asymptotic variance estimators proposed in Section (ref), all of which allow for heteroskedasticity.\footnote{The standard errors of the OLS estimator are also calculated under heteroskedasticity.} The four performance measures are computed based on $B=2000$ simulations, and we table the performance measures under three sample sizes: $n = 250, 500, 1000$.

Binary $X_i$ with Two Binary $Z_{i1}, Z_{i2}$

Our first simulation setup is as follows. In this setup, there are two binary included IVs $Z_{i1}, Z_{i2}$, randomly generated from $Bernoulli(0.5)$ independently. The endogenous regressor $X_i$ and the outcome variable $Y_i$ are generated by \[

alignedX_i&=\mathbbm{1}\{2Z_{i1}Z_{i2}+2(1-Z_{i1})(1-Z_{i2})-1 \geq u_i\}, \\ Y_i&=\alpha_0+\beta_{01}Z_{i1}+\beta_{02}Z_{i2}+\gamma_0 X_i+\epsilon_i,

\] where $\alpha_0=\gamma_0=1$. To compare with the 2SLS estimator, which treats $(Z_{i1}, Z_{i2})$ as excluded instruments, we examine different values of the coefficients of the included regressors $(Z_{i1}, Z_{i2})$: $\beta_{01}=\beta_{02}= \{1, 0.5, 0\}$. The values of the coefficients $(\beta_{01}, \beta_{02})$ represent the degree of violation of the exclusion restriction, and the 2SLS estimator is only consistent when $\beta_{01}=\beta_{02}=0$.

The two error terms $(\epsilon_i, u_i)$ are drawn, independently from $(Z_{i1}, Z_{i2})$, from the joint normal distribution with mean $(0, 0)$, variance $(1, 1)$, and correlation parameter $\rho$, which captures the extent of endogeneity between $X_i$ and $\epsilon_i$. We also consider different levels of endogeneity $\rho\in \{0.5, 0, -0.5\}$, with $\rho=0$ corresponding to the case with no endogeneity issue (where OLS becomes unbiased and consistent).

Since the instrument $Z_i=(Z_{i1}, Z_{i2})$ is discrete, we use the sample averages to estimate the conditional expectations: \[ \hat{\pi}(z)=\frac{\sum_{i=1}^n X_i \mathbf{\mathbbm1}\{Z_i=z\} } {\sum_{i=1}^n\mathbf{\mathbbm1}\{Z_i=z\} }, \quad \hat{h}(z)=\frac{\sum_{i=1}^n Y_i \mathbf{\mathbbm1}\{Z_i=z\} } {\sum_{i=1}^n\mathbf{\mathbbm1}\{Z_i=z\} }. \] In this case, the three estimators $\hat{\theta}, \hat{\theta}^{*}, \hat{\theta}_{disc}$ are numerically equivalent.

Tables (ref) and (ref) report the performance of the five different estimators for $\gamma_0$ under various degrees of exclusion violations (with endogeneity $\rho=0.5$) and different levels of endogeneity (with $\beta_{01}=\beta_{02}=1$), respectively.\footnote{Tables (ref)-(ref) in the Online Appendix report the performances of the estimators for the remaining coefficients $\alpha_0, \beta_{01}$, and $\beta_{02}$.} The results demonstrate the robust performance of our estimators in the presence of violations of the exclusion restriction and endogeneity. The root mean squared error (RMSE) of the three estimators are reasonably small, and the coverage probabilities of the 95% confidence intervals are close to the nominal level.

In contrast, the 2SLS estimator has a very large standard deviation and bias in this simulation setup, even when the exclusion is satisfied $\beta_{01}=\beta_{02}=0$, due to the small determinant of the matrix $X'P_ZX$. Additionally, the OLS estimator has a very small (close to zero) coverage probability with the presence of endogeneity. Our estimators' advantages become more significant as the sample size $n$ increases due to the fast reduction in both bias and standard deviation, but the 2SLS and OLS estimators remain biased regardless of sample size. Also, the $\sqrt{n}$ convergence rate of our three estimators, are strongly demonstrated by the almost exact 50% reduction in SD and RMSE from $n=250$ to $n=1000$.

sidewaystable[!htbp] \caption{Bin $X$ with Bin $Z_1, Z_2$: Performance of $\hat{\gamma}$ (Coef. of $X$) \\ Different Degrees of Exclusion Violations} \begin{tabular}{c|cccc|cccc|cccc} \hline \hline & \multicolumn{4}{c|}{$n=250$ } & \multicolumn{4}{c|}{$n=500$ }& \multicolumn{4}{c}{$n=1000$ } \\ \hline Est& Bias & SD & RMSE & CP &Bias & SD & RMSE & CP & Bias & SD & RMSE & CP \\ \hline &\multicolumn{12}{c}{ $\beta_{01}=\beta_{02}=1$} \\ \hline $\hat{\theta}$ & -0.003 & 0.182 & 0.182 & 0.956& 0.002& 0.137& 0.137& 0.939 & 0.005& 0.094 & 0.094 & 0.952 \\[1.5ex] $\hat{\theta}^{*}$ & -0.003 & 0.182 & 0.182 & 0.956& 0.002& 0.137& 0.137& 0.939 & 0.005& 0.094 & 0.094 & 0.952 \\[1.5ex] $\hat{\theta}_{disc}$ & -0.003 & 0.182 & 0.182 & 0.956& 0.002& 0.137& 0.137& 0.939 & 0.005& 0.094 & 0.094 & 0.952 \\[1.5ex] $\hat{\theta}_{2sls}$ & -0.826 & 34.302 & 34.312 & 0.889 & -3.105 & 116.242 & 116.284 & 0.893 & 2.563 & 65.802 & 65.852 & 0.878 \\[1.5ex] $\hat{\theta}_{ols}$ & -0.485& 0.121 & 0.500 & 0.024 & -0.485 & 0.088& 0.493 & 0.000 & -0.485 & 0.062 & 0.489 & 0.000 \\[1.5ex] \hline &\multicolumn{12}{c}{ $\beta_{01}=\beta_{02}=0.5$} \\ \hline $\hat{\theta}$ & -0.003 & 0.182 & 0.182 & 0.956& 0.002 & 0.137 & 0.137 &0.939& 0.005 & 0.094 & 0.094& 0.952 \\[1.5ex] $\hat{\theta}^{*}$ & -0.003 & 0.182 & 0.182 & 0.956& 0.002 & 0.137 & 0.137 &0.939&0.005 & 0.094 & 0.094& 0.952 \\[1.5ex] $\hat{\theta}_{disc}$ & -0.003 & 0.182 & 0.182 & 0.956& 0.002 & 0.137 & 0.137 &0.939&0.005 & 0.094 & 0.094& 0.952 \\[1.5ex] $\hat{\theta}_{2sls}$ & -0.678 & 17.872 & 17.885 & 0.932&-1.759& 59.329 & 59.355 & 0.917& 0.984 & 32.927 & 32.942 & 0.896 \\[1.5ex] $\hat{\theta}_{ols}$ & -0.485 & 0.121 & 0.500 & 0.024& -0.485 & 0.088 &0.493 & 0.000 & -0.485 & 0.062 & 0.489 &0.000 \\[1.5ex] \hline &\multicolumn{12}{c}{ $\beta_{01}=\beta_{02}=0$} \\ \hline $\hat{\theta}$ & -0.003& 0.182 & 0.182& 0.956& 0.002& 0.137& 0.137 & 0.939 & 0.005 & 0.094 &0.094 &0.952 \\[1.5ex] $\hat{\theta}^{*}$ & -0.003& 0.182 & 0.182& 0.956& 0.002& 0.137& 0.137 & 0.939& 0.005 & 0.094 &0.094 &0.952 \\[1.5ex] $\hat{\theta}_{disc}$ & -0.003& 0.182 & 0.182& 0.956& 0.002& 0.137& 0.137 & 0.939& 0.005 & 0.094 &0.094 &0.952 \\[1.5ex] $\hat{\theta}_{2sls}$ & -0.530 & 3.734& 3.771 & 0.991&-0.413 & 4.745& 4.763 & 0.996& -0.594 & 3.571 & 3.620 & 0.993 \\[1.5ex] $\hat{\theta}_{ols}$ & -0.485 & 0.121& 0.500 & 0.024& -0.485& 0.088 & 0.493& 0.000 & -0.485 & 0.062 & 0.489 & 0.000 \\[1.5ex] \hline \end{tabular}
sidewaystable[!htbp] \caption{Bin $X$ with Bin $Z_1, Z_2$: Performance of $\hat{\gamma}$ (Coef. of $X$)\\ Different Degrees of Endogeneity} \begin{tabular}{c|cccc|cccc|cccc} \hline \hline & \multicolumn{4}{c|}{$n=250$ } & \multicolumn{4}{c|}{$n=500$ }& \multicolumn{4}{c}{$n=1000$ } \\ \hline Est& Bias & SD & RMSE & CP &Bias & SD & RMSE & CP & Bias & SD & RMSE & CP \\ \hline &\multicolumn{12}{c}{$\rho=0.5$} \\ \hline $\hat{\theta}$ & -0.003 & 0.182 & 0.182 & 0.956& 0.002& 0.137& 0.137& 0.939 & 0.005& 0.094 & 0.094 & 0.952 \\[1.5ex] $\hat{\theta}^{*}$ & -0.003 & 0.182 & 0.182 & 0.956& 0.002& 0.137& 0.137& 0.939 & 0.005& 0.094 & 0.094 & 0.952 \\[1.5ex] $\hat{\theta}_{disc}$ & -0.003 & 0.182 & 0.182 & 0.956& 0.002& 0.137& 0.137& 0.939 & 0.005& 0.094 & 0.094 & 0.952 \\[1.5ex] $\hat{\theta}_{2sls}$ & -0.826 & 34.302 & 34.312 & 0.889 & -3.105 & 116.242 & 116.284 & 0.893 & 2.563 & 65.802 & 65.852 & 0.878 \\[1.5ex] $\hat{\theta}_{ols}$ & -0.485& 0.121 & 0.500 & 0.024 & -0.485 & 0.088& 0.493 & 0.000 & -0.485 & 0.062 & 0.489 & 0.000 \\[1.5ex] \hline &\multicolumn{12}{c}{$\rho=0$} \\ \hline $\hat{\theta}$ & -0.007 & 0.181& 0.181& 0.956& -0.000 & 0.137& 0.137 & 0.938& 0.004& 0.093 & 0.093 & 0.952 \\[1.5ex] $\hat{\theta}^{*}$ & -0.007 & 0.181& 0.181& 0.956& -0.000 & 0.137& 0.137 & 0.938& 0.004& 0.093 & 0.093 & 0.952 \\[1.5ex] $\hat{\theta}_{disc}$ & -0.007 & 0.181& 0.181& 0.956& -0.000 & 0.137& 0.137 & 0.938& 0.004& 0.093 & 0.093 & 0.952 \\[1.5ex] $\hat{\theta}_{2sls}$ & -1.315& 32.234 & 32.261 & 0.904 & -0.301 & 38.531 &38.532 & 0.894 & 0.362 & 74.036 & 74.036 & 0.877 \\[1.5ex] $\hat{\theta}_{ols}$ & -0.002 & 0.127 &0.127 &0.948 & -0.001& 0.091 &0.091 & 0.946 & 0.002 &0.064 &0.064 &0.942 \\[1.5ex] \hline &\multicolumn{12}{c}{$\rho=-0.5$} \\ \hline $\hat{\theta}$ & -0.011 & 0.182 &0.183& 0.954& -0.003& 0.138& 0.138 & 0.937 &0.003 & 0.093 & 0.093 &0.950 \\[1.5ex] $\hat{\theta}^{*}$ & -0.011 & 0.182 &0.183& 0.954& -0.003& 0.138& 0.138 & 0.937 &0.003 & 0.093 & 0.093 &0.950 \\[1.5ex] $\hat{\theta}_{disc}$ & -0.011 & 0.182 &0.183& 0.954& -0.003& 0.138& 0.138 & 0.937 &0.003 & 0.093 & 0.093 &0.950 \\[1.5ex] $\hat{\theta}_{2sls}$ & 0.913 & 59.738 & 59.745 & 0.900 & 1.477& 58.951 & 58.969 & 0.882 & 0.576 & 74.167 & 74.169 & 0.872 \\[1.5ex] $\hat{\theta}_{ols}$ & 0.483 & 0.124& 0.499 & 0.023 & 0.485 & 0.087 & 0.492& 0.002 & 0.485 & 0.062 & 0.489 & 0.000 \\[1.5ex] \hline \end{tabular}

Binary $X_i$ with Continuous $Z_i$

In this subsection, we consider a different DGP in which there is a continuous IV $Z_i$ drawn from $\mathcal{N}(0, 2)$. The variables $X_i$ and $Y_i$ are generated by \[

alignedX_i =\mathbbm{1}\{2Z_i \geq u_i\}, \quad Y_i =\alpha_0+\beta_0Z_i+\gamma_0 X_i+\epsilon_i,

\] where $\alpha_0=\gamma_0=1$, and we consider three values for $\beta_0=\{1, 0.5, 0\}$. The error terms $(u_i,\epsilon_i)$ are again drawn from the joint normal distribution as in Section (ref), independently from $Z_i$, with $\rho=\{0.5, 0, -0.5\}$.

For the two semiparametric estimators $\hat{\theta}$ and $\hat{\theta}^*$, we use the Nadaraya-Watson kernel estimator to nonparametrically estimate $\pi_0$ and $h_0$ in the first stage. We use the standard Gaussian kernel and set the bandwidth based on least square cross validation. We also find that the performances of the final estimators do not change much with other choices of kernels (e.g., an Epanechnikov kernel). For the discretization-based estimator $\hat{\theta}_{disc}$, we partition the support of $Z_i$ into $K=10$ cells defined by the (empirical) decile ranges. Our results stay similar if $K$ is set to be larger, say, $30$.

As shown in Tables (ref) and (ref) (and Tables (ref)-(ref) in the Online Appendix), all our three estimators perform uniformly well across different values of $\beta_0$ and $\rho$. Although the two semiparametric estimators $\hat{\theta}, \hat{\theta}^{*}$ involve nonparametric regressions in the first stage, they have reasonably good performances even with a small sample size ($n=250$), with the corresponding CI coverage probabilities for $\gamma_0$ close to their nominal level 95%. In contrast, the 2SLS and OLS estimators are significantly biased under exclusion restriction violation and endogeneity, and their CI coverage probabilities are close to zero for all the sample sizes.

We also find that, the discretization-based estimator $\hat{\theta}_{disc}$ performs (surprisingly) well in finite samples. While the two semiparametric estimators $\hat{\theta}$ and $\hat{\theta}^*$ perform well overall, their small-sample biases induced by the first-stage nonparametric regressions are fairly noticeable when compared to that of $\hat{\theta}_{disc}$, especially in the estimation of $\gamma_0$ under $n = 250$. In contrast, the loss of asymptotic efficiency in $\hat{\theta}_{disc}$ seems to be fairly small and more than compensated by its smaller finite-sample bias.

sidewaystable[!htbp] \caption{Bin $X$ with Cts $Z$: Performance of $\hat{\gamma}$ (Coef. of $X$) \\ Different Degrees of Exclusion Violations} \begin{tabular}{c|cccc|cccc|cccc} \hline \hline & \multicolumn{4}{c|}{$n=250$ } & \multicolumn{4}{c|}{$n=500$ }& \multicolumn{4}{c}{$n=1000$ } \\ \hline Est& Bias & SD & RMSE & CP &Bias & SD & RMSE & CP & Bias & SD & RMSE & CP \\ \hline &\multicolumn{12}{c}{$\beta_0=1$} \\ \hline $\hat{\theta}$ & 0.044 & 0.326 & 0.329 & 0.942 & 0.036 &0.220 & 0.223 & 0.954 & 0.024 & 0.155 & 0.156 & 0.948 \\[1.5ex] $\hat{\theta}^{*}$ & -0.099 & 0.306 & 0.321 & 0.942 & -0.072 & 0.209 & 0.221 & 0.948 & -0.059 & 0.148 & 0.159 & 0.942 \\[1.5ex] $\hat{\theta}_{disc}$ & -0.031 & 0.321 & 0.323 & 0.950 & -0.016 & 0.223 & 0.223 & 0.955 & -0.014 & 0.161 & 0.161 & 0.951 \\[1.5ex] $\hat{\theta}_{2sls}$ & 5.165 &0.298 & 5.174 & 0.000 & 5.159 &0.200 & 5.163& 0.000 & 5.165 & 0.147& 5.167& 0.000 \\[1.5ex] $\hat{\theta}_{ols}$ & -0.477 & 0.198 & 0.516 & 0.322 & -0.485 & 0.140 & 0.505 & 0.060 & -0.486 & 0.097 & 0.495 & 0.000 \\[1.5ex] \hline &\multicolumn{12}{c}{$\beta_0=0.5$} \\ \hline $\hat{\theta}$ & 0.044 & 0.326 &0.329 & 0.942& 0.036 &0.220 & 0.223 &0.954 & 0.024& 0.155 & 0.156 & 0.948 \\[1.5ex] $\hat{\theta}^{*}$ & -0.150 & 0.293 & 0.329 & 0.930& -0.114 & 0.204 & 0.234 &0.936& -0.092& 0.146 & 0.173 & 0.906 \\[1.5ex] $\hat{\theta}_{disc}$ & -0.031 & 0.321& 0.323& 0.950& -0.016 & 0.223 & 0.223 & 0.955& -0.014 & 0.161 & 0.161 & 0.951 \\[1.5ex] $\hat{\theta}_{2sls}$ & 2.584 & 0.206 & 2.592 & 0.000 & 2.578& 0.140 &2.582 & 0.000 & 2.582 & 0.102 & 2.584 & 0.000 \\[1.5ex] $\hat{\theta}_{ols}$ & -0.477 & 0.198& 0.516& 0.322& -0.485& 0.140 & 0.505 &0.060 & -0.486& 0.097& 0.495 & 0.000 \\[1.5ex] \hline &\multicolumn{12}{c}{$\beta_0=0$} \\ \hline $\hat{\theta}$ & 0.044 &0.326 & 0.329 & 0.942 & 0.036 &0.220& 0.223 & 0.954& 0.024 & 0.155 & 0.156 & 0.948 \\[1.5ex] $\hat{\theta}^{*}$ & -0.287 & 0.310 & 0.422 & 0.828 &-0.220 & 0.223 &0.313& 0.802& -0.168& 0.160 & 0.232& 0.782 \\[1.5ex] $\hat{\theta}_{disc}$ & -0.031 & 0.321 & 0.323 & 0.950 & -0.016 & 0.223 & 0.223 & 0.955& -0.014& 0.161 & 0.161& 0.951 \\[1.5ex] $\hat{\theta}_{2sls}$ & 0.003 & 0.164 & 0.164 &0.948 & -0.002 & 0.113 &0.113 &0.954& -0.001 & 0.081& 0.081 & 0.950 \\[1.5ex] $\hat{\theta}_{ols}$ & -0.477 &0.198 & 0.516 & 0.322 &-0.485 & 0.140 & 0.505 &0.060 & -0.486 & 0.097 & 0.495 & 0.000 \\[1.5ex] \hline \end{tabular}
sidewaystable[!htbp] \caption{Bin $X$ with Cts $Z$: Performance of $\hat{\gamma}$ (Coef. of $X$) \\ Different Degrees of Endogeneity} \begin{tabular}{c|cccc|cccc|cccc} \hline \hline & \multicolumn{4}{c|}{$n=250$ } & \multicolumn{4}{c|}{$n=500$ }& \multicolumn{4}{c}{$n=1000$ } \\ \hline Est& Bias & SD & RMSE & CP &Bias & SD & RMSE & CP & Bias & SD & RMSE & CP \\ \hline &\multicolumn{12}{c}{$\rho=0.5$} \\ \hline $\hat{\theta}$ & 0.044 & 0.326 & 0.329 & 0.942 & 0.036 &0.220 & 0.223 & 0.954 & 0.024 & 0.155 & 0.156 & 0.948 \\[1.5ex] $\hat{\theta}^{*}$ & -0.099 & 0.306 & 0.321 & 0.942 & -0.072 & 0.209 & 0.221 & 0.948 & -0.059 & 0.148 & 0.159 & 0.942 \\[1.5ex] $\hat{\theta}_{disc}$ & -0.031 & 0.321 & 0.323 & 0.950 & -0.016 & 0.223 & 0.223 & 0.955 & -0.014 & 0.161 & 0.161 & 0.951 \\[1.5ex] $\hat{\theta}_{2sls}$ & 5.165 &0.298 & 5.174 & 0.000 & 5.159 &0.200 & 5.163& 0.000 & 5.165 & 0.147& 5.167& 0.000 \\[1.5ex] $\hat{\theta}_{ols}$ & -0.477 & 0.198 & 0.516 & 0.322 & -0.485 & 0.140 & 0.505 & 0.060 & -0.486 & 0.097 & 0.495 & 0.000 \\[1.5ex] \hline &\multicolumn{12}{c}{$\rho=0$} \\ \hline $\hat{\theta}$ & 0.074 &0.325 & 0.334 & 0.933 & 0.056 & 0.218 & 0.225 &0.948& 0.035 & 0.153 & 0.157 & 0.943 \\[1.5ex] $\hat{\theta}^{*}$ & -0.093 & 0.305 & 0.319 & 0.946 & -0.065 &0.207 & 0.218 & 0.953 & -0.056 & 0.147 & 0.158 &0.943 \\[1.5ex] $\hat{\theta}_{disc}$ & 0.001 & 0.325& 0.325 & 0.953 & 0.002 & 0.222 & 0.222 & 0.958& -0.005 & 0.161& 0.161 & 0.954 \\[1.5ex] $\hat{\theta}_{2sls}$ & 5.165 & 0.298 & 5.173 & 0.000 & 5.158 & 0.198& 5.161 & 0.000 & 5.165 & 0.147 & 5.167 & 0.000 \\[1.5ex] $\hat{\theta}_{ols}$ & -0.004 & 0.204 &0.204 & 0.938 & -0.003 & 0.146 & 0.146 & 0.940 &-0.003 & 0.101 & 0.101& 0.948 \\[1.5ex] \hline &\multicolumn{12}{c}{$\rho=-0.5$} \\ \hline $\hat{\theta}$ & 0.108& 0.323 & 0.340 &0.917 & 0.074 & 0.217 & 0.230& 0.937& 0.046 &0.153 & 0.160 & 0.936 \\[1.5ex] $\hat{\theta}^{*}$ & -0.081& 0.302 & 0.312 & 0.955 & -0.060 & 0.206 & 0.214 & 0.959 & -0.054 & 0.147 & 0.156 & 0.947 \\[1.5ex] $\hat{\theta}_{disc}$ & 0.038 & 0.321& 0.324 & 0.952 & 0.019 & 0.222 & 0.223 & 0.960 & 0.004 & 0.161 & 0.161 & 0.952 \\[1.5ex] $\hat{\theta}_{2sls}$ & 5.163 & 0.294 & 5.172 &0.000 & 5.157 & 0.196 & 5.161& 0.000 &5.165 & 0.147 & 5.167 & 0.000 \\[1.5ex] $\hat{\theta}_{ols}$ & 0.481 & 0.197 & 0.519 & 0.298 & 0.477 & 0.139 & 0.497 & 0.068 & 0.479 & 0.098 & 0.489 & 0.002 \\[1.5ex] \hline \end{tabular}

Continuous $X_i$ with Continuous IV $Z_i$

In this subsection, we consider another DGP setup, where the endogenous covariate $X_i$ is a continuous random variable generated as $$X_i = \cos(Z_i) + \sqrt{0.5\left|Z_i\right| + 0.5} \cdot u_i,\quad \text{ with } Z_i \sim U\left[-\pi,\pi\right],$$ where $u_i$ and $\epsilon_i$ are jointly normal as before. As before, $(\alpha_0, \gamma_0) = (1,1)$ and we study three values of $\beta_0$: $\beta_0=\{1, 0.5, 0\}$. To also illustrate the point that our method works well under different choices of the first-stage nonparametric estimation methods, here we run cubic spline regressions with cross-validated choices of degrees of freedom.\footnote{Our results do not change substantially if the Nadaraya-Watson estimator is used instead.}

Tables (ref) and (ref) (along with Tables (ref)-(ref) in the Online Appendix) show that our estimators continue to perform well under different nonparametric regression methods and DGP designs. It is worth noting that the 2SLS estimator yields very large standard errors (even when the exclusion restriction is satisfied). This is because, even though $Z_i$ is by construction relevant for $X_i$ (in a nonlinear manner), $Z_i$ is only “weak IV" for $X_i$ via linear projection. In contrast, our three proposed estimators are able to capture the nonlinear relevance of $Z_i$, and deliver small standard errors across all the simulation configurations.

sidewaystable[!htbp] \caption{Cts $X$ with Cts $Z$: Performance of $\hat{\gamma}$ (Coef. of $X$) \\ Different Degrees of Exclusion Violations} \begin{tabular}{c|cccc|cccc|cccc} \hline \hline & \multicolumn{4}{c|}{$n=250$ } & \multicolumn{4}{c|}{$n=500$ }& \multicolumn{4}{c}{$n=1000$ } \\ \hline Est& Bias & SD & RMSE & CP &Bias & SD & RMSE & CP & Bias & SD & RMSE & CP \\ \hline &\multicolumn{12}{c}{$\beta_0=1$} \\ \hline $\hat{\theta}$ & 0.061 & 0.094 & 0.112 & 0.871 & 0.035 & 0.064 & 0.072 & 0.901 & 0.020 & 0.045 & 0.049 & 0.911 \\[1.5ex] $\hat{\theta}^{*}$ & -0.044 & 0.112 & 0.120 & 0.909 & -0.029 & 0.073 & 0.079 & 0.922 & -0.018 & 0.049 & 0.052 & 0.925\\[1.5ex] $\hat{\theta}_{disc}$ & 0.024 & 0.089 & 0.092 & 0.932 & 0.013 & 0.062 & 0.064 & 0.947 & 0.005 & 0.045 & 0.045 & 0.936 \\[1.5ex] $\hat{\theta}_{2sls}$ & 39.48 & 1367 & 1368 & 0.946 & 35.02 & 1609 & 1609 & 0.943 & -21.83 & 2032 & 2032 & 0.956\\[1.5ex] $\hat{\theta}_{ols}$ & 0.311 & 0.042 & 0.314 & 0.000 & 0.312 & 0.030 & 0.313 & 0.000 & 0.312 & 0.022 & 0.313 & 0.000\\[1.5ex] \hline &\multicolumn{12}{c}{$\beta_0=0.5$} \\ \hline $\hat{\theta}$ & 0.061 & 0.094 & 0.112 & 0.871 & 0.035 & 0.064 & 0.072 & 0.901 & 0.020 & 0.045 & 0.049 & 0.911 \\[1.5ex] $\hat{\theta}^{*}$ & -0.044 & 0.111 & 0.120 & 0.909 & -0.029 & 0.073 & 0.079 & 0.922 & -0.018 & 0.049 & 0.052 & 0.925 \\[1.5ex] $\hat{\theta}_{disc}$ & 0.024 & 0.089 & 0.092 & 0.932 & 0.013 & 0.062 & 0.064 & 0.947 & 0.005 & 0.045 & 0.045 & 0.936 \\[1.5ex] $\hat{\theta}_{2sls}$ & 19.50 & 670.8 & 670.9 & 0.945 & 17.06 & 785.3 & 785.3 & 0.943 & -10.21 & 982.7 & 982.5 & 0.957 \\[1.5ex] $\hat{\theta}_{ols}$ & 0.311 & 0.042 & 0.314 & 0.000 & 0.312 & 0.030 & 0.313 & 0.000 & 0.312 & 0.022 & 0.313 & 0.000 \\[1.5ex] \hline &\multicolumn{12}{c}{$\beta_0=0$} \\ \hline $\hat{\theta}$ & 0.061 & 0.094 & 0.112 & 0.871 & 0.035 & 0.064 & 0.072 & 0.901 & 0.020 & 0.045 & 0.049 & 0.911 \\[1.5ex] $\hat{\theta}^{*}$ & -0.044 & 0.112 & 0.120 & 0.909 & -0.029 & 0.073 & 0.079 & 0.922 & -0.018 & 0.049 & 0.052 & 0.925 \\[1.5ex] $\hat{\theta}_{disc}$ & 0.024 & 0.089 & 0.092 & 0.932 & 0.013 & 0.062 & 0.064 & 0.947 & 0.005 & 0.045 & 0.045 & 0.936 \\[1.5ex] $\hat{\theta}_{2sls}$ & -0.487 & 49.20 & 49.19 & 0.990 & -0.899 & 54.68 & 54.67 & 0.987 & 1.403 & 72.53 & 72.52 & 0.996 \\[1.5ex] $\hat{\theta}_{ols}$ & 0.311 & 0.042 & 0.314 & 0.000 & 0.312 & 0.030 & 0.313 & 0.000 & 0.312 & 0.022 & 0.313 & 0.000 \\[1.5ex] \hline \end{tabular}
sidewaystable[!htbp] \caption{Cts $X$ with Cts $Z$: Performance of $\hat{\gamma}$ (Coef. of $X$) \\ Different Degrees of Endogeneity} \begin{tabular}{c|cccc|cccc|cccc} \hline \hline & \multicolumn{4}{c|}{$n=250$ } & \multicolumn{4}{c|}{$n=500$ }& \multicolumn{4}{c}{$n=1000$ } \\ \hline Est& Bias & SD & RMSE & CP &Bias & SD & RMSE & CP & Bias & SD & RMSE & CP \\ \hline &\multicolumn{12}{c}{$\rho=0.5$} \\ \hline $\hat{\theta}$ & 0.061 & 0.094 & 0.112 & 0.871 & 0.035 & 0.064 & 0.072 & 0.901 & 0.020 & 0.045 & 0.049 & 0.911 \\[1.5ex] $\hat{\theta}^{*}$ & -0.044 & 0.112 & 0.120 & 0.909 & -0.029 & 0.073 & 0.079 & 0.922 & -0.018 & 0.049 & 0.052 & 0.925\\[1.5ex] $\hat{\theta}_{disc}$ & 0.024 & 0.089 & 0.092 & 0.932 & 0.013 & 0.062 & 0.064 & 0.947 & 0.005 & 0.045 & 0.045 & 0.936 \\[1.5ex] $\hat{\theta}_{2sls}$ & 39.48 & 1367 & 1368 & 0.946 & 35.02 & 1609 & 1609 & 0.943 & -21.83 & 2032 & 2032 & 0.956\\[1.5ex] $\hat{\theta}_{ols}$ & 0.311 & 0.042 & 0.314 & 0.000 & 0.312 & 0.030 & 0.313 & 0.000 & 0.312 & 0.022 & 0.313 & 0.000\\[1.5ex] \hline &\multicolumn{12}{c}{$\rho=0$} \\ \hline $\hat{\theta}$ & 0.052 & 0.099 & 0.112 & 0.913 &0.030 & 0.068 & 0.074 & 0.918 & 0.016 & 0.047 & 0.050 & 0.928 \\[1.5ex] $\hat{\theta}^{*}$ & -0.025 & 0.105 & 0.108 & 0.916 & -0.017 & 0.072 & 0.074 & 0.916 & -0.011 & 0.048 & 0.049 & 0.929 \\[1.5ex] $\hat{\theta}_{disc}$ & -0.002 & 0.090 & 0.090 & 0.950 & -0.000 & 0.065 & 0.065 & 0.942 & -0.002 & 0.046 & 0.046 & 0.945 \\[1.5ex] $\hat{\theta}_{2sls}$ & -17.63 & 1306 & 1306 & 0.935 & -74.79 & 2165 & 2166 & 0.946 & -37.09 & 882.2 & 882.8 & 0.951\\[1.5ex] $\hat{\theta}_{ols}$ & -0.001 & 0.047 & 0.047 & 0.942 & 0.001 & 0.034 & 0.034 & 0.940 & -0.001 & 0.024 & 0.024 & 0.947\\[1.5ex] \hline &\multicolumn{12}{c}{$\rho=-0.5$} \\ \hline $\hat{\theta}$ & 0.044 & 0.104 & 0.113 & 0.936 &0.026 & 0.070 & 0.074 & 0.934 & 0.014 & 0.047 & 0.049 & 0.941 \\[1.5ex] $\hat{\theta}^{*}$ & -0.003 & 0.106 & 0.106 & 0.918 &-0.001 & 0.070 & 0.070 & 0.918 &-0.002 & 0.048 & 0.048 & 0.937 \\[1.5ex] $\hat{\theta}_{disc}$ & -0.026 & 0.090 & 0.094 & 0.919 &-0.013 & 0.065 & 0.066 & 0.926 &-0.007 & 0.046 & 0.046 & 0.942 \\[1.5ex] $\hat{\theta}_{2sls}$ & -14.81 & 597.6 & 597.6 & 0.946 &35.31 & 1988 & 1988 & 0.946 & -68.94 & 1710 & 1711 & 0.952 \\[1.5ex] $\hat{\theta}_{ols}$ & -0.313 & 0.043 & 0.315 & 0.000 &-0.311 & 0.031 & 0.313 & 0.000 & -0.313 & 0.021 & 0.314 & 0.000 \\[1.5ex] \hline \end{tabular}

Empirical Applications

We apply our methodology to examine the returns to education, a topic of substantial attention in the literature (see, e.g., card2001estimating for a review of various studies on this topic). A key concern in investigating the causal impact of education is its potential endogeneity, and it is challenging to find valid instruments that are excluded from the model. Our approach allows us to include all potential instruments in the regression model and test their validity.

We conduct two applications and test the direct effects of different instruments for education. The first application explores the college proximity indicators, which are proposed in card1993using. We find that after controlling for regional characteristics, the presence of a nearby college does not significantly affect income. However, the 2SLS estimator varies substantially with different instruments and can become insignificant, while our estimators remain more robust regardless of the choice of the instruments. In the second application, we examine the validity of family background variables as instruments. Our findings show that the number of siblings has no significant effect on wages rates, while parents' education significantly increases wages.

Application I: College Proximity Indicators

We use the same dataset as in card1993using, drawn from the National Longitudinal Survey of Young Men (NLSYM). This data contains information of $n=3010$ male observations in 1976, documenting their educational attainment, wage, race, age, and assorted demographic characteristics. card1993using proposes to use the presence of a 4-year college as an instrument for education, which is likely to affect an individual's educational attainment but may not have direct effects on their earnings. However, card1993using also raises a potential concern with this instrument, as the presence of a college might be correlated with superior school quality and, consequently, could lead to higher earnings. We study two specifications that investigate one of the two indicators of college proximity respectively: the presence of a nearby 2-year college (nearc2) and the presence of a nearby 4-year college (nearc4).

Following card1993using, the dependent variable is the log of hourly wage in 1976, the endogenous variable is education, and the control variables include experience, experience squared, a black indicator, indicators for southern residence and residence in an SMSA in 1976, and indicators for region in 1966 and living in an SMSA in 1966. Distinct from card1993using, our approach also includes the college proximity instrument in the model and allows for testing its direct effect on the outcome by examining the significance of the coefficient. Table (ref) presents the summary statistics of the primary variables.

table[table omitted — 981 chars of source]

We present the results of six different estimators. The first two estimators $\hat{\theta}, \hat{\theta}^{*}$ are introduced in Section (ref). For the estimation of $\hat{\pi}(Z_i), \hat{h}(Z_i)$, we employ the Support Vector Machine (SVM) method, a broadly applied machine learning technique for high-dimensional regressors. The neural network approach is also implemented, yielding similar results and the same significance of all coefficients.\footnote{We adopt the function `svm' from the e1071 package and the function `neuralnet' from the neuralnet package in R to implement the SVM approach and the neural network method.} For the discretization estimator $\hat{\theta}_{disc}$, we divide the experience variable into three partitions using empirical quantiles and generate dummy variables for each partition. Then we construct the instrument for education using the product of any two indicator variables. The estimator $\hat{\theta}_{ols}$ is the OLS estimator that includes the instrument in the regression, while $\tilde{\theta}_{ols}$ represents the OLS estimator that does not include the instrument. The last one $\hat{\theta}_{2sls}$ is the 2SLS estimator using the college proximity indicator as the excluded instrument for education.

Table (ref) and Table (ref) display the outcomes of the coefficients for education and the college proximity instruments using SVM and neural network methods. The results from our three estimators $\hat{\theta}, \hat{\theta}^{*}, \hat{\theta}_{disc}$ demonstrate that, after controlling for all the regional factors in 1966 and 1976, having a nearby 2-year college or 4-year college has no significant effects on wages. This finding supports the validity of using college proximity indicators as excluded instruments. The standard deviations of the instrument coefficients with $\hat{\theta},\hat{\theta}^{*}, \hat{\theta}_{disc}$ are very close to that of $\hat{\theta}_{ols}$, reinforcing the good performance of these estimators.

The estimated returns to education from $\hat{\theta}, \hat{\theta}^{*}, \hat{\theta}_{disc}$ are uniformly positive and significant under various specifications and nonparametric estimation methods. Moreover, the standard deviations of the education coefficients, derived from the three estimators, are also reasonably small across different specifications and are smaller than the one obtained from the 2SLS estimator. The coefficients on education from $\hat{\theta}, \hat{\theta}_{disc}$ are all higher than those from the two OLS estimators, suggesting that the OLS estimators may underestimate education's impact. The coefficient from $\hat{\theta}^{*}$ can be lower than $\hat{\theta}_{ols}$, as it uses a different dependent variable. Overall, the three estimators $\hat{\theta}, \hat{\theta}^{*}, \hat{\theta}_{disc}$ yield very similar results, which corroborates our theoretical results on their asymptotic variances in Section (ref).

For the 2SLS estimator $\hat{\theta}_{2sls}$, the coefficients on education vary substantially when using the two different instruments. It becomes insignificant when using the presence of a nearby 2-year college as an instrument, due to the large standard deviation. In addition, the estimated coefficients on education from $\hat{\theta}, \hat{\theta}^{*}, \hat{\theta}_{disc}$ are uniformly smaller than the one obtained from the 2SLS estimator, since our three estimators allow for the direct effect of the instrument.

table[table omitted — 1,259 chars of source]
table[table omitted — 1,270 chars of source]

Application II: Family Background Variables

As data about nearby colleges might not always be available, family background variables are often used as instruments for education. In this application, we explore two family background variables as instruments: parents' average education and number of siblings. To conduct this study, we utilize the dataset `NLSY79', which conducts interviews of $n = 10800$ young individuals, both male and female, ranging in age from 14 to 21 in 1979. This survey records various characteristics of the individuals, such as gender, marriage status, work-related factors, region indicators, as well as family background variables. Table (ref) displays the summary statistics of the key variables.

table[table omitted — 1,005 chars of source]

We compare our three estimators $\hat{\theta}, \hat{\theta}^{*}, \hat{\theta}_{disc}$ with two OLS estimators and three 2SLS estimators. We still apply both the SVM and the neural network methods to estimate $\hat{\pi}(Z_i)$ and $ \hat{h}(Z_i)$. For the discretization estimator, we construct dummy variables for experience, hour, parents' average education, and number of siblings, based on whether each variable is above the median. The included instruments for education are then constructed using the product of any two variables. For the two OLS estimators, $\hat{\theta}_{ols}$ includes the two family background variables, while $\tilde{\theta}_{ols}$ does not. Additionally, we evaluate three 2SLS estimators. The first one $\hat{\theta}_{2sls}^{both}$ uses both parents' education and number of siblings as excluded instruments for education. The second estimator $\hat{\theta}_{2sls}^{edu}$ employs only parents' education as an instrument, while the third one $\hat{\theta}_{2sls}^{sib}$ utilizes solely the number of siblings.

Table (ref) and Table (ref) display the results of the eight estimators. The findings from the three estimators $\hat{\theta}, \hat{\theta}^{*}, \hat{\theta}_{disc}$ show that the number of siblings does not have significant effects on wages, whereas parents' education significantly increases income. This result is consistent with our intuitive reasoning, as parents' education could influence an individual's wage by creating a more favorable educational environment. The three 2SLS estimators appear to overestimate the returns to education, especially the two estimators $\hat{\theta}_{2sls}^{both}, \hat{\theta}_{2sls}^{edu}$ which involve using parents' education as instruments. The three estimators $\hat{\theta}, \hat{\theta}^{*}, \hat{\theta}_{disc}$ all have significantly positive coefficients on education, and their results are smaller than those of the three 2SLS estimators, given that they control for direct effects of parents' education.

table[table omitted — 1,375 chars of source]
table[table omitted — 1,395 chars of source]

Conclusion

This paper offers an alternative approach to achieve point identification of endogenous regression models in the absence of excluded instruments. The key idea of this approach is to leverage the nonlinear dependence between the included exogenous regressor and the endogenous variable. For estimation, we introduce two semiparametric estimators and a easy-to-compute discretization-based estimator. The asymptotic properties of all three estimators are derived and their robust finite sample performances are demonstrated through Monte Carlo simulations. We apply the approach to study returns to education, and to test the direct effects of college proximity indicators as well as family background variables.