EconBase
← Back to paper

Asymptotic Properties of Endogeneity Corrections Using Nonlinear Transformations

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

58,711 characters · 7 sections · 70 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Asymptotic Properties of Endogeneity Corrections Using Nonlinear Transformations

\newtheorem{theorem}{Theorem} \newtheorem{definition}{Definition} \newtheorem{assumption}{Assumption} \newtheorem{example}{Example} \newtheorem{remark}{Remark}

{

abstractThis paper studies the asymptotic properties of endogeneity corrections based on nonlinear transformations without external instruments, which were originally proposed by Park and Gupta (2012) and have become popular in applied research. In contrast to the original copula-based estimator, our approach is based on a nonparametric control function and does not require a conformably specified copula. Moreover, we allow for exogenous regressors, which may be (linearly) correlated with the endogenous regressor(s). We establish consistency, asymptotic normality and validity of the bootstrap for the unknown model parameters. An empirical application on wage data of the US current population survey demonstrates the usefulness of the method.

Keywords: IV regression, generated regressors, wage equations }

Introduction

Employing instrumental variables is a classical econometric tool for identifying causal effects if no randomized controlled trial or suitable observed control variables are available (see e.g. angristpischke:2008). Instruments allow for constructing quasi-experiments and yield consistent parameter estimators for endogenous regressors as long as they are uncorrelated with the error term. In empirical practice, however, a common problem is the choice of suitable instruments. As reviewed in ebbes:2009, there are different approaches to construct consistent estimators in a regression with endogenous regressors without using (external) instruments, the `higher moments' approach (see lewbel:1997), the `identification through heteroscedasticity' approach (see rigobon:2003) and the `latent instrumental variable' approach (see ebbes:2005). Alternative approaches to instrument-free identification have been recently proposed by lewetal:22 or kiviet:23.

The present paper follows up on guptapark:2012 who obtain identification via nonlinearity. The copula based endogeneity correction suggested by guptapark:2012 has recently become very popular in marketing research (see e.g. burmeister:15, datta:2015, vomberg:20, or butt:21) but also in economics (see e.g. aloui:16, blauw:16, or bhatt:21). In contrast to the copula-based approach in earlier work, our endogeneity correction relies on a nonparametric control function approach (see e.g. navarro:2010 and wooldridge:2015 for an overview on control functions). We just assume that the endogeneity is due to an error component that enters the endogenous regressor in a nonlinear way. This approach does not need to assume that the nonlinear relationship between the endogenous regressor and the error term is known. Rather we just assume that it is monotonously increasing or decreasing. Hence our endogeneity adjustment is basically nonparametric, although the underlying regression model is fully parametric.\newline The idea goes back at least to guptapark:2012 who propose a related copula-based approach. In their estimation method, the joint distribution of the endogenous regressor and the error is represented by a copula. The parameters are estimated by maximum likelihood, where the likelihood function is based on the estimated joint distribution. There are two potential drawbacks in their approach. On the one hand, they do not study the asymptotic properties of their estimator and hence it is not clear under which conditions their approach is consistent and how to conduct statistical inference. On the other hand, they (implicitly) assume that any additional exogenous regressor is independent of the endogenous regressor. These drawbacks are partly addressed by haschka:2021 and yang:2022 who allow for a particular form of dependence between the exogenous and endogenous regressors. However, the proof for consistency in yang:2022 assumes that the distribution of the endogenous regressor is known which is typically not the case in empirical practice. Furthermore they do not study other asymptotic properties such as the convergence rate and asymptotic normality of their estimator. haschka:2022 provides an extension of the idea to linear panel models based on maximum likelihood, but does not study the asymptotic properties.\newline We study the asymptotic properties of a popular (control function based) variant of the estimator which does not require to specify a suitable copula. Furthermore, it allows for a correlation between the endogenous regressor and other exogenous regressors, and we rigorously study the asymptotic properties of our estimator. The estimator is calculated in two steps: First, we remove the (linear) dependence among the exogenous and endogenous regressors by computing the residuals from an auxiliary regression. Second, the endogenous part of the error is estimated by applying the quantile function of the normal distribution on the ranks of the residuals from the first step. In contrast, yang:2022 (2sCOPE) calculate residuals from a regression of the transformed regressors, thereby assuming a particular nonlinear dependence among the exogenous and endogenous regressors.\newline Due to the two-step estimation, it is not straightforward to derive asymptotic results for our estimator. A particular problem is that both the residuals and the ranks are dependent so that no simple law of large number or central limit theorem can be applied. We rely on results for stochastic integrals in order to analyze the asymptotic properties. In particular, we derive consistency and asymptotic normality of the structural parameters based on recent results on (residual) empirical process theory. The calculations reveal that the asymptotic distribution of the estimator depends on nuisance parameters arising from estimating a latent error term. This is similar in nature to the so-called `{\it generated regressor issue}' originating from the work of pagan:1984 and which has since then been discussed elsewhere in the literature, see, among others, chen2003estimation, hahnridder:13, mammenetal:12, or mammenetal:16.\newline The plan of our paper is as follows. In Section (ref), we develop our empirical endogeneity correction, discuss connections to former copula-based approaches as well as to instrumental variable estimation. Asymptotic properties are derived in Section (ref). Next, we study the small sample properties by means of Monte Carlo simulation in Section (ref) and provide an empirical application to wage data (Section (ref)). This application is well suited for our approach because the potentially endogenous regressor (years of education) is apparently not normal and because there is an ongoing debate in the empirical literature about the choice of appropriate instruments.

Model, Identification, and Estimator

Consider a structural linear model

align[align omitted — 53 chars of source]

where \(x\) is a \(k \times 1\) vector of exogenous regressors (including a constant) and \(z\) is a scalar endogenous regressor correlated with the error term \(u\). Let us further assume that the endogenous variate and the error can be decomposed as

align[align omitted — 98 chars of source]

where \(e\) and \(\varepsilon\) are mean-zero error terms satisfying $\mathbb{E}[\varepsilon\mid x,z] = 0$ and $e \perp x$. If the function $f$ is linear, the parameter $\rho$ is not identified in general without using external instrumental variables. lewetal:22 consider a somewhat different triangular system in which $\rho$ is identified without external instrumental variables if the latent error terms are not normally distributed and the sign of $\rho$ is fixed. We pursue a different path, i.e. we can allow for normally distributed \(f(e)\) and \(\varepsilon\), but require that the function \(f(\cdot)\) is nonlinear so that \(e\) must not be Gaussian. If the function $f(\cdot)$ is nonlinear and known, then it is possible to identify and correct for endogeneity by just augmenting the regression by $f(e)$ and in this case we refer to $f(e)$ as the `control function'. This result is summarized by the following theorem.

theoremAssume that Eqs.\ (ref) and (ref) hold, with $\mathbb{E}[\varepsilon\mid x,z] = 0$, $e \perp x$, $\mathbb{V}[e]>0$, $f(\cdot)$ striclty monotone, and $f(e) \sim \mathcal{N}(0, 1)$.\footnote{Since we can absorb mean and variance of \(f(e)\) into the intercept and \(\rho\) in Eqs. (ref) and (ref), we shall, without loss of generality, assume in the following that \(\mathbb{E}[f(e)] = 0\) and \(\mathbb{V}[f(e)] = 1.\)} Then, $(\beta',\gamma,\rho)'$ are identified if and only if the distribution of $e$ is not normal.

The setting is mainly inspired by guptapark:2012 who consider the special case of independence between endogenous and exogenous regressor, i.e. \(\delta = 0\) in Eq. (ref). Under the assumption that the marginal cumulative distribution function (c.d.f.) \(F_u(t) \coloneqq \mathbb{P}(u \leq t)\) as well as the copula governing the joint distribution of \((u,e)\) are normal, they propose an estimation method using a Gaussian copula. Since the resulting likelihood depends on the unknown c.d.f. \(F_e(t) \coloneqq \mathbb{P}(e \leq t)\), the latter is estimated by integrating a kernel density estimator. To illustrate, assume for simplicity that \(\mathbb{V}[e] = \mathbb{V}[u] = 1\) and let \(\{y_i,x_i,e_i\}_{i = 1}^n\) be a sample of length \(n\) of independent and identically distributed (i.i.d.) copies of \((y,x,e)\). The resulting estimator of the true parameter \(\alpha_0 \coloneqq (\beta_0',\gamma_0)'\), say, maximizes the approximate log-likelihood

align*[align* omitted — 323 chars of source]

where \(u_i(\alpha) \coloneqq y_i - \beta'x_i - \gamma z_i\) is a generalized residual, \(\Phi(\cdot)\) and \(\phi(\cdot)\) are the standard normal c.d.f. and density, respectively, \(H^{-1}\) is the left-continuous generalized inverse of any c.d.f. \(H\), and we use the notation \(\ell_n(\alpha; \hat{F}_e)\) to indicate that the log-likelihood depends on the nonparametric estimator \(\hat{F}_e\) of the unknown distribution \(F_e\). Since the (approximate) likelihood function is quite complicated, it needs to be maximized by numerical techniques.

Notwithstanding its great conceptual value, this procedure also has at least two major drawbacks. First, the approach does not account for dependence between endogenous and exogenous regressors. Second, and perhaps the most important practical concern, it is a priori unclear under which assumptions the usual properties of ML estimation carry over (if at all) to the case with nonparameterically estimated infinite dimensional nuisance parameter \(F_e\); see e.g. genest:1995. As we demonstrate below, deriving precise statements about limiting properties in the presence of a nonparameterically generated regressor is a highly non-trivial undertaking. Thus, there is no asymptotic theory in guptapark:2012 to enable and/or justify statistical inference.

In order to overcome these issues, we propose an easy-to-implement alternative based on a simple augmented regression that does not require the specification of the joint distribution \((u,e)\). Akin to guptapark:2012, we assume that there exists a strictly monotonic and nonlinear function \(f(\cdot)\) such that $f(e)$ is distributed according to some c.d.f. $F_e$. In the following, we assume that $f(e)$ is Gaussian, which is a plausible assumption in many applications. For example, in our empirical application below, $f(e)$ represents latent information about personal information such as intelligence, other skills or diligence, compare the discussion in Section (ref). It then follows by standard arguments that \(\Phi^{-1}(F_e(e)) = f(e)\). Hence, if \(e\) and its distribution \(F_e\) were known, then one can identify \(\beta\) and \(\gamma\) by augmenting Eq. (ref) with \(\eta \coloneqq \Phi^{-1}(F_e(e))\); see Theorem (ref).

Note that these assumptions are similar but not identical to the assumptions of guptapark:2012. For example, whereas in Eq. (ref) we assume that $f(e)$ is normal (or has some other known distribution), guptapark:2012 assume that $u$ is normally distributed. For ensuring the latter, one needs to assume that also $\varepsilon$ is normally distributed, which we do not. Another difference is that we allow for \(\delta \neq 0\), while this sort of dependence between \(z\) and \(x\) has been ruled out by guptapark:2012. Moreover, and in contrast to the aforementioned study, we allow for conditionally heteroscedasticity of $\varepsilon$.

In practice, \(F_e\) is rarely known and has to be estimated. Here, we approach this problem nonparametrically using the empirical distribution. That is, we use the ordinary least-squares (OLS) estimator for Eq. (ref) augmented with a feasible sample counterpart of \(\eta\). More specifically, in a preliminary step the (linear) dependence between \(z\) and \(x\) is eliminated using the OLS estimator \(\hat\delta\) in the regression of \(z\) on \(x\). From the first-stage residuals \(\hat e_i \coloneqq z_i - \hat\delta'x_i\), \(i \in \{1,\dots,n\},\) we then compute the empirical distribution and obtain the following nonparametrically generated regressor

align[align omitted — 105 chars of source]

where $\hat{F}_{\hat{e}}(\hat{e}_i)$ denotes the (rescaled) empirical distribution function of $e$, i.e.

align*[align* omitted — 149 chars of source]

We adopt common practice and rescale by \(n+1\) in order to escape the possibility of $\hat{F}_{\hat{e}}(\hat{e}_i)=1$ (note that $\Phi^{-1}(a)\to \infty$ as $a\to 1$).\footnote{Note that dividing by $n+1$ instead of $n$ is innocuous as \[\operatorname*{\textsf{sup}}\limits_{t \in \mathbb{R}}|\hat{F}_{\hat e}(t)-\tilde{F}_{\hat e}(t)| \leq 1/(n+1),\] where \(\tilde{F}_{\hat e}(\cdot) \coloneqq (n+1)\hat{F}_{\hat e}(\cdot)/n \). Moreover, note that the fact that $\Phi^{-1}(a)\to -\infty$ as $a\to 0$ does not yield any difficulties in the definition of the estimator as $n \geq 1$.} Finally, we can estimate \(\theta \coloneqq (\beta',\gamma,\rho)'\) using the OLS estimator of the model in Eq. (ref) augmented by \(\widehat \eta_i\).

remarkAssuming that the endogeneity enters the error term in an additive manner, it is relatively easy to extend our approach in order to accommodate further endogenous regressors. To fix ideas, consider the special case of two endogenous regressors $z_1$ and $z_2$ so that \[ y = \beta'x + \gamma_1z_1+\gamma_2z_2+u, \quad u = \rho_1f_1(e_1)+\rho_2f_2(e_2) + \varepsilon, \] where $z_j = \delta_j'x+e_j$, $j \in \{1,2\}.$ Let $F_j(\cdot)$ be the c.d.f. of $e_j$, $j \in \{1,2\}$. Assuming that $f_j(e_j)$ are normally distributed and $f_j(\cdot)$ are strictly monotone, one obtains \[ f_j(e_j) = \Phi^{-1}(F_j(e_j)) \eqqcolon \eta_j, \] say. For this additive model specification we can easily replace $f_j(e_j)$ by $\hat\eta_{j}$, where $\hat\eta_{j}$ is constructed from the emprical c.d.f. of residuals from individual first stage regressions of $z_j$ on $x$. Thus, we could estimate the $(k+4) \times 1$ parameter vector $$\theta \coloneqq (\beta',\gamma',\rho')', \quad \gamma \coloneqq (\gamma_1,\gamma_2)',\; \rho \coloneqq (\rho_1,\rho_2)'$$ from the augmented regression $y = \beta'x + \gamma'z + \rho'\hat\eta + \textnormal{\sf error}$, $z \coloneqq (z_1,z_2)'$, $\hat\eta \coloneqq (\hat\eta_1,\hat\eta_2)'$. If the vector $(x,z,e,\varepsilon)$ satisfies correspondingly Assumption (ref) specified below, then our subsequent results carry over.
remarkIt is also possible to relate our approach to the standard instrumental variable estimation. To simplify the exposition we focus on the model with a single endogenous regressor which we write in vector format as \(y = z\gamma + \widehat \eta \rho + \tilde \varepsilon,\) with \(\varepsilon \coloneqq \varepsilon + \rho(\eta - \hat\eta),\) where $y\coloneqq(y_1,\ldots,y_n)'$, $z\coloneqq(z_1,\ldots,z_n)'$, $\widehat \eta \coloneqq (\widehat \eta_1,\ldots,\widehat \eta_n)'$. The estimator can be represented as \begin{align*} \widehat \gamma & \coloneqq \frac{ z'M_{\hat \eta} y }{ z'M_{\hat \eta} z } = \frac{ \widehat v'y }{ \widehat v'z } \end{align*} where $\widehat v \coloneqq M_{\hat \eta} z$ and $M_{\hat \eta} \coloneqq I_n - \widehat \eta \widehat \eta'/\widehat \eta'\widehat \eta$. This representation of the estimator gives rise to the interpretation of a just-identified IV estimator using the residuals from a regression of $z$ on $\widehat \eta$ as instrumental variable vector $\widehat v$. Accordingly, for the full model we may alternatively estimate the coefficients $\beta$ and $\gamma$ from an IV regression using $(x_i, \widehat v_i)$ as instruments with $\hat v_i$ as the $i$'th element of the residual vector $\widehat v$. This allows us to combine the `internal instrument' $\widehat v_i$ with possible external instruments resulting in an over-identified IV (or GMM) estimator. Furthermore the approach of stockyogo:2005 may be adapted to test against `weak instruments'. It should be noted, however, that $\widehat v_i$ is an estimated instrumental variable and its estimation error affects the asymptotic distribution, as will be shown in more detail below using recent findings on residual empirical processes.

Asymptotic Properties

Consider the joint OLS estimator of \(\theta \coloneqq (\beta',\gamma,\rho)'\) given by \(\hat\theta \coloneqq (W'W)^{-1}W'y\), where \(y \coloneqq (y_1,\dots,y_n)'\) and \(W \coloneqq (w_1',\dots,w_n')'\) is a \((k+2)\times n\) matrix for \(w_i \coloneqq (x_i',z_i,\hat{\eta}_i)'\), \(i \in \{1,\dots,n\}\). We aim at deriving the asymptotic distribution of

align[align omitted — 133 chars of source]

where \(\eta \coloneqq (\eta_1,\ldots,\eta_n)'\) and \(\hat\eta \coloneqq (\hat\eta_1,\ldots,\hat\eta_n)'\). Technically, this represents a nontrivial task because of the term \(W'(\eta-\hat\eta)/\sqrt{n}\) that arises due to a nonparametrically generated regressor. The presence of these normal quantile transforms of estimated residuals considerably complicates the derivation of an appropriate asymptotic theory. We approach this problem by combining seminal results on empirical quantile process theory (see e.g. koenkerxiao:2002) and stochastic calculus with recent findings of zhao:2020 on residual rank approximations. In order to do so, we have to impose some assumptions.

assumption\textcolor[rgb]{1,1,1}{.} \begin{enumerate}[label=A\arabic*,ref=A\arabic*] • \(\{y_i,x_i,z_i\}_{i=1}^n\) is a sample of i.i.d. copies of \((y,x,z)\). • $\sum_{i=1}^nx_ix_i'/n \stackrel{p}{\rightarrow} \mathbb{E}[xx'] \eqqcolon \Sigma_x$, and $\sum_{i=1}^nx_ix_i'\varepsilon_i^2/n \stackrel{p}{\rightarrow} \mathbb{E}[xx'\varepsilon^2]$ where $\Sigma_x$ and $\mathbb{E}[xx'\varepsilon^2]$ are \(k \times k\) positive definite matrices and $\mathbb{E}[x] \eqqcolon\mu_x$. • There exists a strictly monotonous function $f(\cdot)$ such that $f(e) \sim {\cal N}(0,1)$. • There exists a decomposition of the endogenous regressor such that $z = \delta'x + e$, where $e$ is independent of $x$, with \(\mathbb{E}[e] = 0\), \(\mathbb{V}[e] \eqqcolon \sigma_e^2 > 0\), and $\mathbb{E}[e^4]<\infty$. • (i) \(e\) has differentiable c.d.f. \(F_e\) that does not coincide with the normal distribution. (ii) For the density $f_e$ of $e$, the following must hold: There is a constant $\gamma \in (1/2,1)$ such that, for $a \rightarrow 0$ \begin{equation*} \sup_{u \in (a,1-a)} \frac{f_e(F^{-1}_e(u))}{\mathsf{min}(u,1-u)} = o\left(a^{-1/(2\gamma)} \right). \end{equation*} • $\mathbb{E}[\varepsilon \mid x,z]=0$, $\mathbb{E}[\varepsilon^2 \mid x,z] \in (0,\infty)$ a.s., and $\mathbb{E}[\varepsilon^4] <\infty$. \end{enumerate}

The first part of Assumption (ref) is similar to the first-stage monotonicity condition employed frequently in the literature on causal inference (see e.g. imbang:1994 or imbnew:2009) and justifies the representation \(\Phi^{-1}(F_e(e)) = f(e)\). Assumption (ref) implies that $x$ affects $z$ only linearly. In applications, this limitation can be addressed by adding nonlinear transformations of the exogenous regressors to $x$. We could, in principle, relax this assumption and model the relation between exogeneous and endogenous regressor nonparametrically using recent results of zhao:2022. Since deriving an asymptotic theory is already fairly complex, we leave this issue to future research. The first part of Assumption (ref) is essential for identification as mentioned above. Without any further specification of the distribution of \(\varepsilon\) the normality of \(f(e)\) is not directly testable; see also the discussion in becker:22. Still, in practice one could use the first-stage residuals \(\hat e\) to test whether \(e\) is Gaussian, which would lead to non-identification. The second part of Assumption (ref) requires that the density of $e$ decays sufficiently fast to $0$ at the boundary. The condition is satisfied by many parametric distributions like the Gamma distribution with shape parameter of at least two or the Chi-square distribution with degrees of freedom parameter of at least three. This is a technical requirement needed to deal with the estimation error of the normal scores introduced by the first-stage residuals; see zhao:2020 for details. Finally, the mean-independence in Assumption (ref) ensures that $\varepsilon$ is orthogonal to all regressors in the augmented regression, while conditional heteroscedasticity is allowed for.

propositionIf Assumption (ref) is satisfied, then \( \sqrt{n}(\hat\theta-\theta) \stackrel{d}{\rightarrow} \mathcal{N}(0_{k+2},\Sigma), \quad \Sigma \coloneqq M^{-1}\Omega M^{-1}, \) where \[ M \coloneqq \begin{bmatrix} \Sigma_x & \Sigma_x'\delta & 0_{k} \\ \delta'\Sigma_x & \delta'\Sigma_x\delta+\sigma_e^2 & c_2 \\ 0_{k}' & c_2 & 1 \end{bmatrix}\quad \text{ and } \quad \Omega \coloneqq \begin{bmatrix} \Omega_1 & \Omega_1\delta & 0_{k} \\ \delta'\Omega_1 & \delta'\Omega_1\delta+\omega_1 & \omega_{12} \\ 0_{k}' &\omega_{12} & \omega_2\end{bmatrix}, \] with \(\Omega_1 \coloneqq \mathbb{E}[xx'\varepsilon^2]+\Sigma_x\rho^2c_1^2\sigma_e^2\), \(\omega_1 \coloneqq \mathbb{E}[e^2\varepsilon^2]+\rho^2c_3\), \(\omega_2 \coloneqq \mathbb{E}[\eta^2\varepsilon^2]+\rho^2/2\), $\omega_{12} \coloneqq \mathbb{E}[e\eta\varepsilon^2] + \rho^2c_2/2$, and \[ c_1 \coloneqq \int_0^1\frac{f_e(F_e^{-1}(u))}{\phi(\Phi^{-1}(u))} du, \quad c_2 \coloneqq \int_0^1F_e^{-1}(u)\Phi^{-1}(u) du, \] and \[\normalfont c_3 \coloneqq \int_0^1\int_0^1 \frac{F_e^{-1}(u)}{\phi(\Phi^{-1}(u))}\frac{F_e^{-1}(v)}{\phi(\Phi^{-1}(v))} (\textsf{min}(s,u)-su) ds du. \]

Proposition 3,1 reveals that the limiting distribution is affected by finite (\(\delta\)) and infinite dimensional (\(F_e(\cdot)\)) nuisance parameters that arise due to not knowing the (linear) relationship between exogenous and endogenous regressors and the unknown functional form of the endogeneity, respectively. One way in which this so-called `generated regressor' issue (see e.g. pagan:1984) could be mitigated consists in expressing our two-step estimator in terms of a (joint) moment estimator whose moment conditions are orthogonalized using suitable influence functions as proposed recently by cheetal:23. Finally, we note that the identification failure of \(e\) being Gaussian leads to \(c_j = 1\), \(j \in \{1,2,3\}\), so that \(M\) is not invertible.

The limiting variance-covariance matrix could be estimated by a sample analogy principle. The only parameters difficult to estimate are \(c_j\), \(j \in \{1,2,3\}\); but they could, in general, be replaced by sample counterparts using \(\hat{F}_{\hat{e}}(\cdot)\) and a suitable density estimator \[ \hat{f}_{\hat{e}}(y) \coloneqq \frac{1}{nh_n}\sum_{i = 1}^nK\left(\frac{\hat{e}_i-y}{h_n}\right), \] say, where \(K(\cdot)\) is a Kernel density function and \(h_n\) is the bandwidth such that \(h_n \rightarrow 0\) as \(n \rightarrow \infty\), see e.g. hardlint:1994. Given all the difficulties involved in density estimation and its oftentimes poor performance in finite samples, we refrain from this idea and resort to bootstrap standard errors instead. We propose a resampling scheme that is able to account for the sampling uncertainty of the nonparametrically generated regressor thereby reproducing the correct distribution of Proposition 1. Similar to the control function literature (see e.g. wooldridge:2015), we simply generate bootstrap samples by drawing with replacement from the original data; see also chenetal:03 for a more general treatment of bootstrap inference with nonparametrically estimated nuisance parameters. More specifically, let \({\cal X}^*_1,\dots,{\cal X}^*_n\) be drawn with replacement from the empirical distribution of the vectors $\{ {\cal X}_1,\ldots,{\cal X}_n\},$ ${\cal X}_i \coloneqq (y_i,x_i',z_i)'$, and define, analogously to \(\hat\theta\), \(\hat\theta^*\) based on the bootstrap data. Corollary 1 below then justifies the use of bootstrap (percentile) confidence intervals as well as standard errors constructed from a large number of the bootstrap estimators \(\hat\theta^*\).

corollary(i) If Assumption (ref) is satisfied, then \(\sqrt{n}(\hat\theta^*-\hat\theta) \stackrel{d}{\rightarrow} \mathcal{N}(0_{k+2},\Sigma)\) under the probability measure \(\mathbb{P}_{\cal X}\) implied by the bootstrap. (ii) If, in addition, there exists some \(\nu > 0\) such that \(\mathbb{E}^*[\lVert\sqrt{n}(\hat\theta^*-\hat\theta)\rVert^{2+\nu}] < \infty\), then \(n\mathbb{E}^*[(\hat\theta^*-\hat\theta)(\hat\theta^*-\hat\theta)'] \rightarrow_{\mathbb{P}_{\cal X}} \Sigma\), where \(\mathbb{E}^*\) is the expectation induced by \(\mathbb{P}_{\cal X}\).

The moment condition in part ($ii$) of Corollary 3.1 could, in principle, be derived from more primitive assumptions. If this condition fails, test statistics equipped with bootstrap standard errors tend to bee too conservative, as shown recently by hahnliao:2021.

Although, it is generally true that hypothesis tests equipped with naive OLS standard errors yield asymptotically inaccurate results, an exception is given for a Durbin-Hausman-Wu type-test of the null hypothesis \(H_0\): \(\rho = 0\) of no endogeneity. More specifically, as summarized by Corollary 3.2 below, the usual approach based on a textbook \(t\)-statistic remains valid; a situation reminiscent of the `generated regressors' literature, see e.g. pagan:1984.

corollaryLet Assumptions (ref) be fulfilled. Under $H_0:$ $\rho=0$ the regressor $z$ is exogenous and the respective $t$-statistic has a standard normal limiting distribution as $n\to \infty$.

Small Sample Properties

In order to investigate the small sample properties of our estimator, we consider a simple linear model inspired by yang:2022:

equation[equation omitted — 173 chars of source]

with $(\beta_0,\beta_1,\gamma)'=(1,-1,1)'.$ To this end, we follows yang:2022 by setting $x \sim \Gamma(1,1)$. Moreover, we distinguish between two different (gamma) distributions for the component $e$, i.e. ($i$) $\Gamma(1,1)$ or ($ii$) $\Gamma(3,2)$. Although ($i$) violates Assumption (ref), we report results to alleviate comparison with yang:2022. As can be deduced from Figure (ref), passing from ($ii$) to ($i$), the distribution of $e$ (and therefore $z$) becomes highly skewed. A more challenging situation results if $e \sim \Gamma(3,2)$ as this distribution is more similar to that of a Gaussian random variable. For the remaining distributional specifications, two different data generating processes are considered ({\sf DGP1} and {\sf DGP2}).

figure[figure omitted — 755 chars of source]

{\sf DGP1.} According to Eq. (ref), the error term is generated as \(u = \rho \eta + \varepsilon,\) $\varepsilon \sim {\cal N}(0,1)$, and the standard normal distributed random variable $\eta$ is given by $\eta = \Phi^{-1}(F_{e}(e))$, where $F_{e}$ is the c.d.f. of the respective $\Gamma$-distribution specified above. We distinguish the cases of uncorrelated ($\delta=0$) and correlated ($\delta=1$) regressors as well as of no ($\rho=0$), moderate ($\rho=0.5$) and strong ($\rho=0.9$) endogeneity.\newline

{\sf DGP2.} The second design obtains with $\delta = 0$ and is taken from yang:2022 who assume that \[ e = F_{e}^{-1}(\Phi(e^*)),\; x = F_{x}^{-1}(\Phi(x^*)),\quad (e^*,x^*,u)' \sim \mathcal{N}(0_3,\Xi), \; \Xi \coloneqq

bmatrix[bmatrix omitted — 65 chars of source]

, \] where, as pointed out above, $F_{e}$ varies across the different gamma distributions and $F_x = \Gamma(1,1)$. The endogeneity and the dependence between $x$ and $e$ is captured by the correlation coefficients $\rho$ and $\alpha$, respectively. Again, we distinguish the cases of uncorrelated ($\alpha=0$) and correlated ($\alpha=1/2$) regressors as well as of no ($\rho=0$) or strong ($\rho=1/2$) endogeneity.

table[table omitted — 4,462 chars of source]
table[table omitted — 3,778 chars of source]

For all Monte Carlo experiments, we draw $n=250$ IID random draws $\{x_i,z_i,u_i\}_{i=1}^n$ from $(x,z,u)$ according to {\sf DGP1} and {\sf DGP2}. The proposed nonparametric control function estimator ({\sf npCF}) is compared with OLS (i.e. ignoring the endogeneity of $z$), with the estimator by guptapark:2012 ({\sf GP}) and with the {\sf 2sCOPE} estimator by yang:2022.\footnote{We employ the {\sf R}-package `REndo' for computing {\sf GP}. All computations were parallelized and performed using CHEOPS, the DFG-funded (Funding number: INST 216/512/1FUGG) High Performance Computing (HPC) system of the Regional Computing Center at the University of Cologne (RRZK) using iterations.} The main difference between {\sf 2sCOPE} and our method is that the rank transformations as in (ref) of the present paper are applied already in the first step to both $x$ and $z$. Accordingly, in the first step the transformed endogenous regressor ($\widehat \eta$) is regressed on the corresponding transformations of the exogenous regressors. In the second step, the OLS residuals of this regression are added to the original equation. This procedure is designed to capture the dependence among the regressors as in {\sf DGP2}, but the procedure is expected to be less suitable for {\sf DGP1}.

In order to investigate the validity of statistical inference we consider $t$-type test statistics for the hypotheses $\beta_1 = -1$ and $\gamma=1$ respectively at a nominal significance level of 5%. For the OLS estimator the usual $t$-statistics are used, whereas the other tests are performed using \(t\)-statistics equipped with bootstrap standard errors. Tests based on bootstrap percentile intervals did perform very similar so that the corresponding results are omitted to save space. Our experience suggests that using 99 bootstrap replications is typically enough to obtain reliable estimates for the unknown standard deviations of the coefficients. Finally, all results are based on 1,000 Monte Carlo repetitions.

{\bf Discussion of \sf DGP1.} We first discuss the results in Table (ref) for the case $e \sim \Gamma(1,1)$ if $\delta=0$ . That is, the setup with uncorrelated regressors that is assumed by guptapark:2012. Not surprisingly, OLS performs best in case of no endogeneity with a RMSE less than the half of the other three approaches. This highlights the efficiency loss resulting from the endogeneity correction if it is in fact not necessary. A similar efficiency loss occurs when employing external instruments.\footnote{It should be noted, however, that it is easy to verify that the endogeneity correction is unnecessary by observing an insignificant $t$-statistic for the correction term $\widehat \eta_i$.} Note that in this case the asymptotic distribution of the latter IV estimator with $v$ as some external instrument is given by

align*[align* omitted — 121 chars of source]

where $V_{\textsf{OLS}}$ is the asymptotic variance of the OLS estimator and $r_{zv}^2$ is the $R^2$ from the (first stage) regression of $z$ on $v$. Accordingly, a comparable loss of efficiency results from an IV estimator if the first stage $R^2$ corresponds to $(0.062/0.147)^2 = 0.178$.

In the presence of endogeneity (i.e. $\rho\ne 0$) the OLS estimator of the parameter $\gamma$ is highly biased. Note that for the case of uncorrelated regressors, the coefficient $\beta$ can be estimated unbiasedly even if the regressor $z$ is endogenous. Accordingly, all estimators (including OLS) for $\beta$ are unbiased in this case. On the other hand, the OLS estimator for $\gamma$ is severely biased as the corresponding regressor is endogenous. The copula based endogeneity corrections {\sf GP} and {\sf 2sCOPE} and our {\sf npCF} estimator effectively remove the bias and perform more or less similarly. Furthermore, the empirical sizes of the (bootstrap) $t$-tests are close to the nominal size. Although not reported to save space, using classical OLS standard errors resulted in case of {\sf npCF} in severe size distortions.

The case of correlated regressors with $\delta=1$ is considered in the lower panel of Table (ref). As expected, in this case the OLS estimator is biased for both coefficients $\beta$ and $\gamma$. The bias of the {\sf GP} estimator is similar to the OLS bias for $\beta$ and only slightly smaller for $\gamma$. A substantial bias reduction is obtained by applying the {\sf 2sCOPE} estimator but some bias remains in particular if $\rho=0.9$. This is due to the fact that the {\sf 2sCOPE} estimator assumes a particular nonlinear dependence between $x$ and $z$, whereas in our data generating process the dependence between $z$ and $x$ is linear. The {\sf npCF} estimator is the only estimator that efficiently removes the bias from both parameters.

Turning to the case where $e \sim \Gamma(3,2)$, we expect that the endogeneity corrections have problems to empirically identify the endogenous component of the error term. Indeed, the standard deviations of the endogeneity corrected standard deviations for the estimators of $\gamma$ increase about 50 percent. Apart from the higher standard deviations the findings are qualitatively similar to the results presented in Table (ref). It is interesting to note, however, that the magnitude of the bias of the {\sf 2sCOPE} estimator is similar to the OLS estimator suggesting that this estimator is not able to cope with the (linear) correlation between the regressors $x$ and $z$.

{\bf Discussion of \sf DGP2.} Similar to {\sf DGP1}, the results for {\sf DGP2} confirm the shortcomings of {\sf GP} in case of correlated regressors ($\alpha \neq 0$ and $\rho=0.5$). Given that {\sf DGP2} mimics the model framework in yang:2022 that corresponds to the model assumptions of {\sf 2sCOPE}, it is not surprising that {\sf 2sCOPE} performs best in this setting. Interestingly, our {\sf npCF} estimator effectively removes the bias even in the setting of {\sf DGP2}, whereas the {\sf 2sCOPE} estimator fails to remove the bias resulting from {\sf DGP1}. The standard deviations for {\sf DGP2} are however somewhat smaller for the {\sf 2sCOPE} estimator compared to the {\sf npCF} estimator due to the fact that the former estimator takes care of the particular nonlinear dependence in the {\sf DGP2}.

Empirical Application

In this section, we apply our new estimator to empirical wage data, thereby revisiting the classical economic problem of estimating the returns to schooling (see e.g. harmon:2000). Similarly as in chernozhukov2013inference and rothewied:2013, we analyze micro-level data from the US Current Population Survey in 1988. The sample size is $n=144,750$ and the model is given by

equation[equation omitted — 390 chars of source]

Accordingly, we consider a linear regression in which the logarithm of hourly wages for individual $i$ is explained by the years of education and different control variables. The latter consist of the years of working experience (linearly and quadratically) and several binary variables (married, working in part time, member in a union, living in a standard metropolitan statistical area, being non-white). The control variables are assumed to be exogenous. The years of education is assumed to be an endogenous variable as it might be correlated with unobservable worker's characteristics in the error term $u_i$ such as ability or motivation. The key goal is to find a reliable estimate for $\gamma$.

The literature in labor economics discusses several approaches for estimating $\gamma$ and provides different empirical results for different data sets. For example, if educ$_i$ is instrumented with family characteristics such as the years of education of the parents, the IV estimates are often smaller than the OLS estimates (see wooldridge:textbook). This fits to the economic intuition that years of education might be positively correlated with ability/motivation. If institutional characteristics are used as instruments, the IV estimates are often larger (see card:2000). lemke:2003 try to merge these results with the conclusion that the OLS estimators might be biased upwards.

This analysis contributes to this discussion by providing results of a consistent estimator for $\gamma$ that do not depend on external instruments. The starting point is the observation that the distribution of the regressor educ is obviously not normal, as the histogram in Figure (ref) (left) shows. On the one hand, the distribution is discrete and on the other hand, it is highly skewed. This is a simple argument for justifying the augmentation of the regression (ref) by the normalizing transformation of educ as discussed earlier in the paper. In the first step, residuals $\hat e_i$ from an OLS regression of educ on all exogenous control variables are obtained. Also the distribution of these residuals is highly skewed, as Figure (ref) (right) shows. In the second step, the augmented regressor $\widehat \eta_i = \Phi^{-1}(\widehat F_{\hat e}(\widehat e_i) )$ is calculated. The distribution of $\widehat \eta_i$ is approximately normal by construction. It is the difference between the two distributions which enables identification of $\gamma$.

Note that at this point, we assume that also the latent counterpart $\eta$, which corresponds to personal characteristics such as intelligence, other skills and diligence, is normally distributed. In our point of view, this makes sense for several reasons. On the one hand, the single characteristics can be assumed as being normally distributed, e.g. intelligence: While standardized intelligence quotient (IQ) ratios are by construction normally distributed, this is less clear for the intelligence itself. However, early research in psychology about intelligence in the beginning of the 20th century (see e.g. burt:17), provided evidence that also intelligence is normally distributed. In later research, psychologists found some more mixed evidence, in particular the question arose whether the actual tails of the intelligence distribution might be heavier than expected under the normal distribution assumption. Nevertheless, there are also more recent studies which still support the normality assumption, see e.g. warne:13.

Moreover, one can rely on an argument based on the central limit theorem: If several sources have an additive influence on $\eta$, the sum of these sources can be expected to be approximately normal. On the other hand, this latent variable is not a linear transformation of the years of education due to rules concerning compulsory school attendance, for example.

figure[figure omitted — 342 chars of source]

Table (ref) presents the results (estimates, standard errors and $t$-statistics) of the standard OLS estimation of model (ref) and the results of the robustified estimation, where $\widehat \eta_i$ is included. In the latter case, we report both the bootstrap standard errors and $t$-statistics as well as the OLS $t$-statistics. In the latter case, the estimation error for $\widehat \eta_i$ is ignored, but the differences are small and not systematic. Due to the large sample size, the $p$-values for all coefficients are smaller than $2 \cdot 10^{-16}$. One exception is the coefficient $\rho_{\eta}$ for $\widehat \eta_i$ which is somewhat larger with $2.5 \cdot 10^{-5}$. This reflects that, on the one hand, all control variables were used to calculate $\widehat \eta_i$, but, on the other hand, $\widehat \eta_i$ still provides considerable explanatory power. The latter observation implies that it is reasonable to assume that the regressor educ is endogenous.

In the case of OLS, we have $\hat \gamma = 0.083$, whereas we have $\hat \gamma = 0.073$ when correcting for endogeneity. So, we observe a slight reduction of the estimated return for education which is what lemke:2003 also suggest. The other coefficients remain very similar, which could have been expected as the residuals $\hat e_i$ are orthogonal to the exogenous regressors and the adjusted $R^2$ of the first stage regression is rather low ($0.1132$). For all coefficients, the $t$-statistics in absolute values are lower and the standard errors are higher in the augmented regression compared to OLS. This is in particular the case for the intercept and $\gamma$ and can be explained with the high correlation ($0.91$) of educ and $\widehat \eta_i$. A similar phenomenon is observed in other studies of endogeneity-corrected estimators such as guptapark:2012. With 2sCOPE, the estimate for $\gamma$ is less than half as large compared to npCF and the estimate for $\rho$ is more than six times larger. A direct comparison is difficult, however, as $\rho$ has a different interpretation for both estimation approaches.

table[table omitted — 2,579 chars of source]

Concluding Remarks

This paper proposes a method for solving the endogeneity problem without external instruments or specifying a copula. The endogeneity correction is obtained by simply augmenting the regression with a conformably transformed regressor. We study the asymptotic properties and analyse the small sample properties of the resulting estimator. Our results suggest that the proposed estimation framework may provide a useful tool for estimating regression models with endogenous regressors if no suitable instruments are available. Even if instrumental variables are at hand, additional nonlinear instruments can be constructed for improve the efficiency of the IV estimator. Furthermore, the nonparametric control function may be employed for testing the exogeneity of the regressors in empirical situations where valid instruments are missing.

A drawback of our approach is, however, that the distribution of the endogenous regressor needs to be different from that of the latent variable. If the distributions are close, then the estimator suffers from multicollinearity among the endogenous regressor and the correction term resulting in potentially large standard errors. It would be interesting to develop more diagnostic tests that allow for checking how appropriate the new method is for a particular data set. An alternative strategy would relax the normality condition of \(f(e)\) and specify instead a class of permissible distributions indexed by some finite dimensional parameter \(\tau\), say. Using nonlinear estimation techniques, one could then estimate \(\tau\) alongside the remaining model parameters \(\theta\).

ACKNOWLEDGEMENTS

We are grateful to the co-editor, Petra Todd, and two referees, as well as to Julien Bergeot and Rouven E. Haschka for helpful comments.