Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
94,002 characters · 11 sections · 67 citation commands
Iterative Estimation of Nonparametric Regressions with Continuous Endogenous Variables and Discrete Instruments
\linespread{1.25}
\address[S.\ Centorrino]{$(^\ast)$ Corresponding Author. International Monetary Fund and Economics Department, Stony Brook University, USA. \\Email address: [email removed].} \address[F. F\`eve]{Toulouse School of Economics, University of Toulouse Capitole, Toulouse, France. \\Email address: [email removed].}
\address[J.\ P.\ Florens]{Toulouse School of Economics, University of Toulouse Capitole, Toulouse, France. \\Email address: [email removed].}
\setstretch{1.5}
Instrumental variables are a workhorse of applied research in economics. The instrumental variable approach is based on two main assumptions that instruments need to satisfy: exogeneity, i.e., instruments are uncorrelated with or otherwise independent of the structural error term, and relevance, i.e., instruments are correlated with or otherwise dependent on the included endogenous regressor. The ability to identify the main causal effect of interest depends on the complexity of the model we wish to estimate and on the strength of these two requirements.
In this work, we consider the nonparametric regression model
where the explanatory variable $X$ is endogenous. This model is usually identified and estimated by taking the instrumental variable, $W$, as mean-independent of the error term. That is, $E(U\vert W)=0$, an exogeneity restriction on $W$.
Under this relatively weak hypothesis, the regression function $\varphi_{\dagger}$ is the solution of the integral equation
An estimator of $\varphi_{\dagger}$ is obtained by solving a regularized version of this functional equation chen2012,darolles2011,florens2003,hall2005,horowitz2011,newey2003.
Uniqueness of the solution to equation (ref) is usually obtained under the so-called completeness condition, or strong identification, which is the relevance assumption in this model (see \citet*{florens1990} and more recently \citet*{andrews2011}, \citet*{canay2013}, \citet*{xavier2011} and \citet*{freyberger2015}, among others). This condition dictates that
which implies a strong dependence between the endogenous variable, $X$, and the instrument, $W$. In particular, whenever $X$ is a continuous variable, completeness requires the instrument $W$ to be itself continuous.\footnote{If $W\in \{1,...,L\}$ the set of equations $\int \varphi(x) f(x \vert W=l) dx=0 \; \forall l$ cannot characterize a function $\varphi$. This is true also when $L = \infty$, as we show in Example (ref) in Appendix (ref).}
This condition is strong and may pose some challenges to the empirical application of fully nonparametric estimators of model (ref). Many instruments employed in empirical work are only binary or discrete. This is true, for instance, for randomized experiments with partial compliance, where the intent-to-treat provides an obvious source of exogenous variation krueger1999,torgovitsky2015.
In many of these cases, the endogenous variable is continuous, while the instrument is binary and, therefore, cannot satisfy the completeness condition stated above. In general, under standard assumptions of uncorrelated or mean-independent instruments, it is not possible to identify any, albeit known, function of $X$ with more than two parameters when only a binary instrument is available lochner2015,loken2012.
Discrete instruments can be accommodated in this setting either by considering a pseudo-solution to equation (ref) babii2017a; or by strengthening the exogeneity condition and, as a consequence, weakening the relevance condition. The latter route is the one we undertake in this paper.
We consider the estimation of the model in equation (ref) where the instrument, $W$, is taken to be fully independent of the error term $U$. That is, $U\upmodels W$. This strong exogeneity condition characterizes $\varphi_{\dagger}$ as a solution to a nonlinear integral equation of the first kind dunker2014,dunker2018,kalten2008.
We detail the implementation of an iterative Landweber-Fridman (LF) algorithm that can be used to estimate the function $\varphi_{\dagger}$ under the restriction of independence, and we derive some of its asymptotic properties.
In a previous paper, dunker2014 have considered the estimation and asymptotic properties of nonparametric estimators for nonlinear integral equations. The model analyzed in our paper is presented there as an example dunker2018. The approach of dunker2014 is extremely general and oriented to a class of regularization schemes less familiar to econometricians. We believe, however, that the case of discrete instruments is sufficiently important for applied research in econometrics to be separately and carefully developed. loh2022 takes a first step in this direction by considering nonparametric identification and estimation in a model with discrete instruments and endogenous variables. Our paper focuses on the case where $X$ is continuous loh2023.
We contribute to the existing literature in several directions. First, the LF technique that we propose is both computationally efficient and easy to implement. Its properties in this class of statistical problems are new to the best of our knowledge. Finally, while our approach is not as general as those provided by dunker2014 and dunker2018 and is limited to the statistical model at hand, our convergence rates clearly spell out the sources of the estimation error and allow us to provide better guidance for implementation in this particular example. Ultimately, our paper has both theoretical and empirical value when compared to existing literature.
The independence assumption is often used in the following nonseparable triangular model
with $(\varepsilon,V) \upmodels W$, $\varepsilon$ and $V$ are standard uniform random variables, and $g_1(x,\cdot)$ and $g_2(w,\cdot)$ are monotone in their second argument fevrier2015,torgovitsky2015. Identification and estimation of this model is usually based on a control function assumption \`a la imbens2009 (see also chernozhukov2007 and horowitzlee2007). The relative advantage of instrumental variables over control functions is that we do not need to specify a triangular structure, i.e., we do not need to specify a first-stage equation that relates the endogenous variable to the instrument. More generally, we allow for a simultaneous dependence between the endogenous regressor, $X$, and the dependent variable, $Y$, that is excluded by a control function approach. However, our model assumes that the function $g (X, \varepsilon)$ has a separable form, with $U = F^{-1}_U(\varepsilon)$, and $F_U$ the cumulative distribution function of $U$. In this respect, our approach is less general than control functions that can account for unobserved multivariate heterogeneity.
The paper is organized as follows. In Section (ref), we briefly discuss some local identification conditions in the independence case in relation to the usual completeness condition. In Section (ref), we present the practical implementation of our estimator, whose properties are detailed in Section (ref). We discuss the specific case of a binary instrument in Section (ref). We conclude our work with an empirical application estimating the returns to education, using data from card1995.
In the following, we let $L^2$ be the space of square-integrable functions with respect to the Lebesgue measure. For a real-valued function $\varphi \in L^2$, we let $\Vert \varphi \Vert$ be the $L^2$-norm. For an operator $T$, we let $\Vert T \Vert$ be the operator norm induced by the $L^2$ norm. For two random variables $X_1$ and $X_2$, we denote $X_1 \upmodels X_2$, if $X_1$ and $X_2$ are independent. If $a$ and $b$ are scalar, $a \vee b = \max\{ a,b\}$, and $a \wedge b = \min\{a,b\}$. For two sequences $a_n$ and $b_n$, we use the notation $a_n \asymp b_n$ to signify that the ratio $a_n/b_n$ is bounded away from zero and infinity. Finally, for an even integer $a \geq 2$, we let $E \Vert \cdot \Vert^a = E \left( \Vert \cdot \Vert^a \right)$.
We consider a random element $(Y,X,W) \in \mathbb{R} \times \mathbb{R}^p\times\mathbb{R}^q$. We analyze the model
where the random vector $X$ only contains endogenous regressors, i.e., $X$ and $W$ do not have elements in common.
This model has been considerably studied under the following mean independence condition
This condition characterizes $\varphi_{\dagger}$ as the solution of the linear equation
We replace Assumption (ref) with a stronger Assumption of full independence.
The condition $E(U) = 0$, which implies $E(Y) = E (\varphi_{\dagger}(X))$, is a normalization condition. It may be thought to be implicit in the definition of the error term $U$, and if it is not satisfied, $\varphi$ may be identified up to location only. In the following, and without loss of generality, we assume that $E(Y) = E (\varphi_{\dagger}(X)) = 0$, as the mean of $Y$ can always be identified (provided it exists) and estimated at a parametric rate. Therefore, we assume that $\varphi_{\dagger} \in \mathcal{E}$, where $\mathcal{E} = \lbrace \varphi \in L^2_X : E(\varphi) = 0 \rbrace$
The (ref) condition implies (ref). Thus, if the completeness condition between $X$ and $W$ is verified, $\varphi_{\dagger}$ is identified under (ref), while the reverse cannot be true.
We, therefore, explore weaker local identification conditions under (ref), which may be satisfied in particular in the case where $X$ is continuous and $W$ discrete.\footnote{Global identification conditions could also be analyzed, but they are less interpretable and difficult to verify in practice chernozhukov2005,feve2015,beyhum2024. loh2023 studies global identification conditions in a similar setting, but he focuses on the space of essentially bounded functions.}
We recast the independence condition in (ref) in the following way. We use the notations
and
Roughly speaking $F$ is a cdf as a function of $y$, and a density as a function of $x$.
The condition in (ref) is equivalent to
or
Equation (ref) may be written as
and is a nonlinear integral equation, where $\varphi_{\dagger}$ denotes its solution, and $T:\mathcal{E} \rightarrow L^2_{U \times W}$ is a nonlinear operator.
We consider the following definition of local identification.
Let \[ \mathcal{B}_R =\lbrace \varphi \in \mathcal{E}: \Vert \varphi - \varphi_{\dagger} \Vert < R \rbrace, \] an open ball of radius $R$ around $\varphi_{\dagger}$. Local identification of $\varphi_{\dagger}$ is given in the following Proposition which recasts a more general result given by chenlee2013.
Condition (i) is about the Fr\'echet differentiability of the operator $T$ chenlee2013, which helps control the behavior of our nonlinear problem in the vicinity of the solution. Moreover, the Fr\'echet derivative is taken to be injective, which is a rank condition on $T^\prime_{\varphi_{\dagger}}$ chenlee2013. Condition $(iii)$ restricts the amount of nonlinearity that is allowed for the ill-posed inverse problem at hand by requiring the continuity of the Fr\'echet derivative chenlee2013. Finally, the restriction on the set $\mathcal{B}_M$ is a source condition, as explained by chenlee2013. In Hilbert spaces, a sufficient condition for $(iv)$ to hold is to restrict the identification set to an ellipsoid or a hyper-rectangle whose elements are sufficiently smooth in the sense that we specify below.
We now provide more explicit sufficient conditions on the primitives of our model to obtain local identification according to Proposition (ref).
Under Assumption (ref), condition (i) of Proposition (ref) is satisfied and
where $T^\prime_\varphi: \mathcal{E} \rightarrow L^2_{U \times W}$ denotes the Fr\'echet derivative of $T$ computed in $\varphi$ as a linear function of $\tilde{\varphi}$.
Assumption (ref) states that the projection of $\varphi$ under $L^2_{U\times W}$ differs from the projection of $\varphi$ under $L^2_U$ except if $\varphi$ is constant. Moreover, this constant equals $0$ under the additional condition $E(U) = 0$. In other words, this condition requires that the dependence structure between $X$ and $U$ changes with the value of the instrument $W$. This Assumption implies condition $(ii)$ in Proposition (ref).
Assumption (ref) implies condition $(iii)$ in Proposition (ref) for some radius $R>0$ chernozhukov2007, and, more generally, it implies that $T^\prime_{\varphi}$ is a Hilbert-Schmidt operator, for all $\varphi \in \mathcal{E}$, and thus compact carrasco2007h. The last condition also implies the local ill-posedness of the inverse problem as the eigenvalues of $T^\prime_{\varphi_{\dagger}}$ have zero as an accumulation point. Proofs of these statements are given in Section (ref) of the Appendix.
Finally, we restrict the set $\mathcal{B}_M$ in a way that is amenable to imposing a source condition. That is, we impose a restriction on the smoothness properties of local deviations from $\varphi_{\dagger}$ as a function of the local ill-posedness of the inverse problem, as determined by the speed of convergence of the eigenvalues of $T^\prime_{\varphi_{\dagger}}$ to $0$.
Definition (ref) imposes that the set of functions in the ellipsoid $\mathcal{B}^\ast_{M}$ around $\varphi_{\dagger}$ be sufficiently smooth to guarantee that its Fourier coefficients go to zero faster than the eigenvalues of $T^\prime_{\varphi_{\dagger}}$. We have that $\mathcal{B}^\ast_{M} \subseteq \mathcal{B}_{M}$, so that this condition is sufficient to satisfy the requirement of Proposition 2.1(iv).
Let $s$ and $a$ be positive constants. When $b_j \asymp \exp( - s j)$, then this condition is satisfied whenever the eigenvalues of $T^\prime_{\varphi_{\dagger}}$ decay polynomially ($t_j \asymp j^{-a}$, mildly ill-posed inverse problem) or exponentially ($t_j \asymp \exp( - a j)$, severely ill-posed inverse problem) as long as $s > a$. If the $b_j$'s have polynomial decay, that is $b_j \asymp j^{-s}$, then Definition (ref) cannot be satisfied if the inverse problem is severely ill-posed. If the inverse problem is mildly ill-posed with $t_j\asymp j^{-a}$, then a sufficient condition to characterize the ellipsoid $\mathcal{B}^\ast_M$ is to impose that $s > a + 0.5$. In particular, if we interpret $s$ as the smoothness of $\varphi - \varphi_{\dagger}$ with respect to the Hilbert scale generated by the operator $\left(T^{\prime \ast}_{\varphi_{\dagger}} T^\prime_{\varphi_{\dagger}} \right)^{-1/2}$, then $s$ is determined by the least smooth element of this difference.\footnote{Results exist about the relationship between the number of continuous derivative of a function and the decay of its Fourier coefficients grafakos2014,katznelson2004. For instance, a function whose Fourier coefficients decay as $j^{-s}$ has $s-1$ continuous derivatives, and its $s$-th derivative is integrable (that is, having a continuous $s$-th derivative is sufficient but not necessary for the function's integrability). Similarly, analytic functions have exponentially decaying Fourier coefficients. For a general Fourier decomposition wrt the eigenvectors of $T^{\prime \ast}_{\varphi_{\dagger}} T^\prime_{\varphi_{\dagger}}$, one can define a differential operator $D =\left(T^{\prime \ast}_{\varphi_{\dagger}} T^\prime_{\varphi_{\dagger}} \right)^{-1/2}$, i.e., the inverse of our integral operator, so that one could interpret smoothness in a standard sense.}
We thus revisit the conditions in Proposition (ref) to obtain the following.
Recall that the model is written as $Y=\varphi_{\dagger}(X)+U$, where $X$ potentially depends on $U$, and $\varphi_{\dagger}$ is the true value of the function of interest. We assume there exists an instrumental variable, $W$, such that $U\upmodels W$ with $E(U) =0$.
We observe an IID sample $\lbrace (Y_i, X_i, W_i), i=1,\dots,n \rbrace$ from the joint distribution of the random vector $(Y,X,W)$, and we take the support of $X$ to be compact and restricted to be the interval $[0,1]^p$ without loss of generality.
The estimation procedure is based on equation (ref), where $F$ is replaced by a nonparametric estimator. Equation (ref) is a nonlinear integral equation that defines an ill-posed inverse problem.
In the case of nonlinear equations, we may consider both concepts of global and local ill-posedness. Global ill-posedness implies that the equation does not have a unique solution that depends continuously on $F$. This property has been studied by gagliardini2012 in a more general case. Our approach is based on local ill-posedness. This property is characterized by the linear approximation at the true value, $T^\prime_{\varphi_{\dagger}}$, which exists, and it is a compact operator under Assumptions (ref) and (ref). Then $T^\prime_{\varphi_{\dagger}}$ does not have a continuous inverse, and the problem is locally ill-posed.
The local ill-posedness of equation (ref) requires using a regularization technique to obtain consistent estimators. Many solutions exist, but it is commonly accepted in the mathematical literature that iterative methods are the easiest to implement. We limit ourselves to the LF method, which is the simplest iterative method. More sophisticated methods can be used that do not require some of the smoothness restrictions imposed in our work. However they seem less suited for econometric problems, in which both the operator $T$ and the function of interest $\varphi_{\dagger}$ can be taken to be very regular. The main issue with LF regularization for nonlinear inverse problems is that it cannot take full advantage of the smoothness of the regression function (i.e., it saturates at a certain level of regularity). In particular, Definition (ref) requires that $s$, the decay of the generalized Fourier coefficients of $\varphi_{\dagger}$, is larger than $a$, the decay of the eigenvalues of $T^\prime_{\varphi_{\dagger}}$ to zero. However in this context and under the additional conditions imposed below, LF regularization saturates at $s = a$ for both the mildly and the severely ill-posed inverse problem and prevents achieving a faster convergence rate. Therefore, to better accommodate the smoothness imposed by our identification conditions, we consider the estimation of the first partial derivatives of $\varphi_{\dagger}$, rather than of the function itself. As we explain in more detail below, this approach allows us to exploit the additional regularity of the class of continuously differentiable functions. Once the first derivative is estimated, we integrate with respect to the corresponding endogenous variable to recover the function of interest. Moreover, the first derivative has an important meaning in empirical analysis, as it measures marginal effects.
Let us denote by $\varphi_{\dagger}^\prime = D\varphi_{\dagger}$ the first partial derivative of $\varphi_{\dagger}$ with respect to any of its arguments. $D$ is a differential operator, with $D^{-1}$ being an integral operator. Without loss of generality, we assume that the partial derivative is always taken with respect to the first component of $X$, $X_1$, and we denote as $X_{-1}$ the remaining components. That is, \[ (D^{-1} \varphi_{\dagger}^\prime)(x) = \int^{x_1}_{-\infty} \varphi_{\dagger}^\prime(\xi,x_{-1})d\xi. \] We let $\mathcal{E}^\prime \subseteq \mathcal{E}$ be the space of centered, one-time differentiable functions of $X$, whose derivative is square-integrable. $\mathcal{E}^\prime$ is dense in $\mathcal{E}$. $D^{-1}$ is a linear operator, and the Fr\'echet derivative of the operator $T$ exists under the same conditions as above. $D$ is playing the role of Hilbert scale, and the ill-posedness of the inverse problem and the smoothness of $\varphi_{\dagger}$ are going to be linked to the properties of $D$. Finally, we define $A_{\varphi_{\dagger}} = T^\prime_{\varphi_{\dagger}} D^{-1}$, as the Fr\'echet derivative of $T$ with respect to $\varphi_{\dagger}^\prime$.
We first describe the numerical algorithm and its implementation. We defer the study of the properties of this estimator to Section (ref). The LF algorithm is based on the following recursive definition
where $j = 0,1,2,\dots,N(c)-1$ is an integer and $N(c) >0$ is the total number of iterations, our regularization parameter, and $c > 0$ determines the size of the step between consecutive iterations.
The algorithm in (ref) requires a closed-form expression for $A^{\ast}_{\varphi}$, the adjoint of $A_\varphi$. $A_\varphi$ is a linear operator from $L^2_{X}$ into $L^2_{U\times W}$. Therefore, $A^{\ast}_{\varphi}$ is a linear operator from $L^2_{U\times W}$ into $L^2_{X}$ which ought to satisfy the following relation
with
Some algebra leads to
which reduces to
where $E_W$ denotes the expectation taken with respect to the marginal distribution of $W$, and we inverted the role of $\xi$ and $x_1$ to simplify notations. The LF estimator of $\varphi_{\dagger}^\prime$ is therefore given by
with $\hat{T}(\hat\varphi_j)$ an estimator of $T(\varphi)$ computed at the point $\hat\varphi_j$, and $\hat{A}^\ast_{\hat{\varphi}^\prime_j}$ an estimator of the adjoint of $A_{\hat\varphi_j}$. These objects are constructed as follows.
For nonlinear inverse problems, iteration methods like the one used here would generally not converge globally. We can prove local convergence by appropriately restricting the initial condition and controlling the behavior of the Fr\'echet derivative of the operator $\hat{T}$. In the following, the number of iterations, $N$, should be taken as a function of $c$ and the sample size, $n$, although we shall not make this dependence explicit. We also let $\varphi_0$ and $\varphi^\prime_0$ denote the probability limit of the (random) initial conditions $\hat\varphi_0$, and $\hat\varphi^\prime_0$. We assume the following.
Assumption (ref) imposes smoothness restrictions and local identification of the parameter of interest (see Section (ref)). Assumption (ref) is a sufficient condition to guarantee local convergence of the iterative procedure. Under these conditions, we prove the following result.
This Proposition states that, at each iteration, the LF algorithm approaches $\varphi_{\dagger}$. Under Assumption (ref), this also implies that the LF algorithm stays in $\mathcal{B}_\dagger$. The proof is provided in Section (ref) of the Appendix. The condition in (ref) implies that the eigenvalues of $T^{\prime\ast}_{\varphi_j} T^{\prime}_{\varphi_j}$ go to zero at a speed that is at most the one of the eigenvalues of $ T^{\prime\ast}_{\varphi_{\dagger}} T^{\prime}_{\varphi_{\dagger}}$. As the decay of the singular values of $T^{\prime}_{\varphi_{\dagger}}$ determines the ill-posedness of the inverse problem, we essentially require that the degree of ill-posedness is no larger than the true one at each step of the LF algorithm. We further impose the following.
Assumption (ref) is a H\"older source condition engl2000. This condition is widely used in linear inverse problems to relate the properties of the function of interest with the properties of the conditional expectation operator carrasco2007h. In the nonlinear inverse problem literature, a similar condition is imposed dunker2014,dunker2018,kalten2008. However this condition does not apply directly to the smoothness properties of $\varphi_{\dagger}^\prime$, but it is rather a condition on the local properties of our initial guess. With $a$ measuring the rate of decay of the eigenvalues of $T^{\prime}_{\varphi_{\dagger}}$, Assumption (ref) links the smoothness of the conditional expectation operator to the one of the integral operator, $D^{-1}$ (engl2000, Section 8.1, and chen2011). The notation $\sim$ is used here to indicate that the eigenvalues of $D^{-a}$ decay as fast as the singular values of $T^{\prime}_{\varphi_{\dagger}}$, i.e. the operator $T^{\prime}_{\varphi_{\dagger}}$ smooths a function at least as much as $D^{-a}$. If $s$ is the speed of decay of the generalized Fourier coefficients of $\varphi_0 - \varphi_{\dagger}$, $\beta = \frac{s-1}{a+1}$, then Assumption (ref) is equivalent to $\varphi^\prime_0 - \varphi_{\dagger}^\prime = \left( T^{\prime \ast}_{\varphi_{\dagger}} T^{\prime}_{\varphi_{\dagger}} \right)^{\frac{(s-1)}{2a}}v$ and $\varphi_0 - \varphi_{\dagger} = \left( T^{\prime \ast}_{\varphi_{\dagger}} T^{\prime}_{\varphi_{\dagger}} \right)^{\frac{s}{2a}}v$. That is, if $\varphi^\prime_0 - \varphi_{\dagger}^\prime$ has regularity $\beta$, then $\varphi_0 - \varphi_{\dagger}$ has regularity $\beta + (a+1)^{-1}$ with respect to the operator $A_{\varphi_{\dagger}}$. Moreover, Assumption (ref) also implies $\varphi_0 - \varphi_{\dagger} = \left( T^{\prime \ast}_{\varphi_{\dagger}} T^{\prime}_{\varphi_{\dagger}} \right)^{\frac{s}{2a}}v = D^{-s}v$, which implies differentiability of $\varphi_0 - \varphi_{\dagger}$ up to the $s^{th}$ order. Assumption (ref) could be extended to encompass a logarithmic source condition as in dunker2018. This is a generalization that we do not pursue here.
The conditions imposed so far are not sufficient to derive an upper bound in the $L^2$ norm for our estimator. We also need to restrict further the local behavior of the Fr\'echet derivative and its adjoint. For any pair of functions $\varphi,\tilde{\varphi} \in \mathcal{E}$, and $\psi \in L^2_{U\times W}$, we can write
where we take
to be the kernels of $\hat{T}^\prime_\varphi$ and $\hat{A}^{\ast}_\varphi$, respectively. Since $W$ is a discrete variable, integrals can be replaced by sums, where appropriate, and we use the integral notation without loss of generality.
The marginal density of $W$ is estimated using a discrete kernel $L(\cdot)$ with bandwidth $h_w$ such that \[ \hat{f}_W(w) = \frac{1}{n} \sum_{i = 1}^n L_{h_w}(w - W_i). \] The properties of this function are given in Assumption (ref) aitch1976,liracine2007. We only note here that when the bandwidth $h_w = 0$, $L(\cdot)$ reduces to the product of indicator functions, and $\hat{f}_W(w)$ is the usual frequency estimator. Then we let
We choose the points $u$ in expanding sets of the form $\{ u: \vert u \vert \leq \ell_n\}$, where $\ell_n$ is a sequence that is either bounded or diverging slowly to infinity so that we can allow for error distributions with unbounded support hansen2008. Finally, we make additional assumptions about the estimated operators and their convergence, and we appropriately restrict the behavior of the tuning parameters.
Assumption (ref) guarantees that we can locally control the behavior of the estimated operators. Assumption (ref) gives the rate of convergence of the operators in Mean Integrated Squared Error (MISE) li2008,darolles2011,florens2012. As all conditioning IVs are discrete, and following li2008, the bias term only depends on $h_w$, and it can be made asymptotically negligible by choosing $h_w = O(n^{-1})$. This result does not depend on the dimension of $W$. Similarly, one can choose $h_u = O(n^{-1/2\rho})$ so that $\hat{T}(\varphi)$ has parametric rates of convergence for all $\varphi \in \mathcal{E}^\prime$. $\hat{A}^{\ast}_\varphi$ is the partial integral of a conditional expectation operator, and its rate depends on the dimension of the conditioning variable. Its estimation also depends on $\hat{f}_{U}$, the density of the error term at $\varphi$. The latter is negligible as long as $h_u = o(1)$, and $nh^p_x \rightarrow \infty$. This result can be proven using a standard argument for U-statistics, and it is omitted here for brevity. In darolles2011, with $W$ continuous, the variance term depends on $h_w^q$. However, with $W$ discrete, the variance term does not depend on $h_w$, and the bias generated by smoothing with respect to the discrete instruments is negligible when $h_w = O(n^{-1})$. We specialize this example in Section (ref) to the case of a binary instrument $W$. Assumption (ref) imposes restrictions on the tuning parameters. Proposing a data-driven procedure for the choice of these parameters is an essential step to be pursued in future research.
Proof of this Proposition is in Section (ref) of the Appendix. This Proposition restricts the local behavior of the Fr\'echet derivatives to ensure convergence of the LF algorithm. It also shows that we have to pay a price to approximate the density of the unknown error term by the density of the regression residuals. This is tantamount to the usual estimation error with nonparametric generated regressors mammen2012. When the density of the error term is uniformly bounded away from $0$ over its support, $\kappa_n$ is a constant. However, as we want to allow for distributions with unbounded support, we let $\kappa_n$ converge to $0$ as $n$ approaches $\infty$.
The following Theorem contains the main result of this Section.
The result of this Theorem gives an upper bound on the Mean Integrated Squared Error (MISE) of our estimator. Loosely speaking, the MISE depends on a regularization bias, $N^{-\beta_\epsilon/2}$, which decreases with the number of iterations, $N$, and varies with the regularity of the function, $\beta$; and a variance term, which is a function of the estimation error, and increases as $N$ diverges. The restriction $\beta \leq 1$ for LF regularization in nonlinear ill-posed inverse problems is not original to us. It is well-established in the mathematical literature that the regularization bias at iteration $j+1$ is of order $(j+1)^{-\beta/2}$. This regularization bias accumulates across $N$ iterations, and, for $\beta$ sufficiently large, it would converge to a non-zero constant as $N\rightarrow \infty$. This could be considered a saturation effect for LF regularization in the context of nonlinear ill-posed inverse problems carrasco2007h. Moreover, there is an additional saturation effect on $\beta$ in our case, as we need the regularization bias to go to zero fast enough so that the additional error from the nonparametric estimation of the unknown density of the residuals is negligible. Therefore, we need $\beta \leq 1-\epsilon$, where $\epsilon > 0$ satisfies Assumption (ref)(iii). For $\beta > 1 - \epsilon$, we simply cannot take full advantage of the source condition in Assumption (ref), and the bias converges at a lower speed. Under this restriction on $\beta$, the upper bound is the same one we would get in the context of a linear ill-posed inverse problem with H\"older source condition florens2012. Finally, the estimation of the first derivative allows us to extend our result to more regular functions, as the order of convergence of $\hat{\varphi}_N$ to $\varphi_{\dagger}$ is improved by integration. This result could be further extended by considering a general approach to LF regularization in Hilbert scales kalten2008.
Finally, the rate of convergence depends on the approximation properties of the initial condition. One could choose $\hat\varphi_0$ in a way that it converges to $\varphi_0$ at a rate faster than $\delta_n \sqrt{N} \vee \gamma_n N^{\frac{1-\beta}{2}}$. This is satisfied by parametric regression models and nonparametric regressions.
There are several important differences between our result and the one in dunker2018. First of all, the result in dunker2018 is of more general scope than the one provided here and obtained under high-level regularity conditions. The Newton method used by dunker2018 does not suffer from the saturation effect of LF regularization, although its implementation is more cumbersome as it involves outer and inner iterations of the regularization scheme. Moreover, to the best of our understanding, dunker2018 directly treats $\hat{T}^{\prime}_{\hat\varphi_j}$ and $\hat{T}^{\prime}_{\varphi_{\dagger}}$ as operators with the same range and does not consider the estimation error that arises from the use of the regression residuals instead of the true error. In his running example of quantile IV regression, this approach seems appropriate. However, in our case, it is important to explicitly consider the nonlinearity in the approximation to provide guidance for the choice of tuning parameters.
Finally, dunker2018 provides conditions under which its rate is optimal, i.e., \[ \left( E\Vert \hat\varphi^\prime_N - \varphi_{\dagger}^\prime \Vert^2\right)^{1/2} = O\left( \delta^{\frac{\beta}{\beta+1}}_n\right). \] It is unclear whether our rates are optimal. When $\beta < 1-\epsilon$, the dominating variance term depends on $\beta$, and the relationship between $\delta_n$ and $\gamma_n$. When $\beta \geq 1-\epsilon$, there is a saturation effect that does not allow us to reach the optimal rate. We have the following.
Provided the variance term $\delta_n \sqrt{N}$ dominates, the result of this Corollary is obtained by balancing the variance and the squared bias terms in Theorem (ref). E.g., if $\beta= 1$, condition (ref) is satisfied if $\gamma_n =o \left( \delta_n \sqrt{N} \right)$. When $\beta \leq 1 - \epsilon$, i.e. $\beta_\epsilon = \beta$, we can reach the optimal rate. In particular, with $\beta = (s-1)/(a+1)$, we have
However we cannot reach the optimal rate for any $\beta > 1 - \epsilon$. Whether LF regularization in Hilbert scales can reach the optimal rate for nonlinear ill-posed problems is an interesting question to be pursued in future research.
We consider in this Section a particular example where $X\in \mathbbm{R}$ is a continuous variable and $W$ is a binary instrument. That is, $W\in \{0,1\}$.
Recall that, for simplicity, we take $E(Y) = E(\varphi_{\dagger}) = 0$. Under mean independence, the model is not identified. The restriction $E(U\vert W) =0$ reduces to two conditions \[ \int \varphi_{\dagger} (x) f_{X\vert W} (x\vert w=0) dx = \int \varphi_{\dagger} (x) f_{X\vert W} (x\vert w=1) dx =0 \] which cannot imply $\varphi_{\dagger}=0$, except when $X$ is also binary or when $\varphi_{\dagger}$ is a two-parameter function in $X$.
When the completeness condition fails, a smooth regularized estimator of $\varphi$ obtained from the sample counterpart of $E(Y\vert W) = E(\varphi_{\dagger}(X)\vert W)$ has a well-defined probability limit, which is the pseudo-true value \[ \varphi_{\dagger}^{P}(x) = \lambda_0 \frac{f_{X\vert W} (x\vert w=0)}{f_X(x)} + \lambda_1 \frac{f_{X\vert W} (x\vert w=1)}{f_X(x)}, \] defined as the projection of $\varphi_{\dagger}$ on the orthogonal complement of the null space of the conditional expectation operator. $\lambda_0$ and $\lambda_1$ are the unique solutions of the system \[ \int \varphi_{\dagger}(x) f_{X\vert W} (x\vert w=l)dx = \sum_{j=0,1} \lambda_j\int \frac{f_{X\vert W} (x\vert w=l) f_{X\vert W} (x\vert w=j)}{f_X (x)} dx, \text{ for } l = \{0,1 \}. \]
We provide proof of this result in Section (ref) of the Appendix and refer to babii2017a for a more general framework. Then to estimate $\varphi_{\dagger}$ using $W$ as an instrument, we require additional assumptions, and the assumption $U\upmodels W$ with $E(U) = 0$ is convenient for this purpose. Nonetheless, we need some auxiliary completeness conditions to obtain identification.
Let us consider first the local condition in Assumption (ref). In this case, it is equivalent to \[ E(\varphi_{\dagger} (X)\vert U=u, W=1) \overset{a.s.}{=} E(\varphi_{\dagger}(X) \vert U=u, W=0) \Rightarrow \varphi_{\dagger} \overset{a.s.}{=} 0. \] We can, therefore, write the following decomposition \[ \frac{f_{U,X\vert W} (u,x\vert j)}{f_{X\vert W} (x\vert j) f_U (u)} = \sum_{k=1}^{\infty}\lambda^{(j)}_{k} \chi_{k,j} (x) \psi_{k,j} (u) = \sum_{k=1}^{\infty}\lambda^{(j)}_{k} \chi_{k} (x) \psi_{k} (u), \] where the functions $\chi$ and $\psi$ form an orthonormal basis in $L^2_{U\times X}$, and they do not depend on $j$ (this is for the sake of exposition and slightly generalizes our Example (ref), where the densities of $(X,U) \vert W = 0$ and $(X,U) \vert W = 1$ have the same eigenfunctions).
Assumption (ref) in this context can be interpreted as stating that if the singular value decomposition of the conditional expectation operators is the same when $W$ is equal to $0$ or to $1$, then the instrument does not have any identifying power. This is true, whenever $\vert \lambda^{(0)}_{k}\vert=\vert\lambda^{(1)}_{k}\vert$, for all $k \geq 0$. On the contrary, whenever $\vert \lambda^{(0)}_{k} \vert \neq \vert \lambda^{(1)}_{k}\vert$, for all $k \geq 1$, then the condition in Assumption (ref) is true only if $\varphi_{\dagger}\overset{a.s.}{=}0$. Hence, the model is locally identified if the dependence structure between $X$ and $U$, as captured by the sequence $\lambda^{(j)}_{k}$, is different for $W=1$ and $W=0$.
We now consider estimation in the particular case of binary instruments. First, consider a parametric approach with \[ \varphi(x) = b(x)^\prime \theta, \] where $\theta \in \mathbb{R}^k$ and $b$ is a given vector of $k$ basis functions (polynomials, splines, etc.). This model is parametric if $k$ is fixed but may be viewed as a sieve estimation if $k$ is allowed to increase with $n$, although we do not pursue this interpretation further. The parametric estimation is based on the exogeneity restriction \[ P(U\leq u\vert W=1) = P(U\leq u\vert W=0). \] Let us define
and we estimate $P(U\leq u\vert W=j)$ by
where $\bar{C}$ is the cdf of a kernel function of order $\rho \geq 2$, $\hbox{\it 1\hskip-3pt I}$ is the indicator function, $h_u$ is a bandwidth, and $n_j = \sum^n_{i=1} \hbox{\it 1\hskip-3pt I} (W_i =j)$.
We minimize the Cram\'er-von Mises distance between the two conditional cdfs with respect to $\theta$
for a suitable grid of points $u_l$. An example of the application of this method is provided below.
A more flexible method (not dependent on a parametric functional form) follows from the application of the LF algorithm described in Section (ref). One detail specific to the implementation with a scalar binary instrument is that we do not smooth the conditional cdf with respect to $W$, but we sort the data for $W=0$ and $W=1$. The Fr\'echet differentiability of the operator $T$ holds as long as the joint pdf of $(Y,X)$ given $W = w$ satisfies Assumption (ref) dunker2014,dunker2018.
Upon an appropriate choice of the bandwidth parameter, i.e., $h_u \asymp n^{-1/2\rho}$, we can take $\delta_n \asymp n^{-1/2}$ in Assumption (ref). Similarly, for $A^\ast_\varphi$, we write \[ \left( \hat{A}^\ast_{\varphi_{\dagger}} \psi \right)(x) = \int_{-\infty}^x \int \frac{\left[\psi(u,1) \hat{f}_{Y,X,W}(\varphi_{\dagger}(x) + u,\xi,1) - \psi(u,0) \hat{f}_{Y,X,W}(\varphi_{\dagger}(x) + u,\xi,0) \right]\hat{f}_{U}(u)}{\hat{f}_{X}(x)} dud \xi, \] with $\psi \in L^2_{U\times W}$, $\hat{f}_{Y,X,\cdot}(\varphi_{\dagger}(x) + u,x,\cdot)$, $\hat{f}_{X}(x)$ nonparametric estimators of the densities of observable components of the model, as described in Section (ref), with bandwidth parameters $(h_x,h_u)$; and $\hat{f}_{U}(u)$, nonparametric estimator of the density of the error term, with bandwidth $h_u$. Using a second order kernel for the estimation of $\hat{f}_{Y,X,\cdot}(\varphi_{\dagger}(x) + u,x,\cdot)$ and $\hat{f}_{X}(x)$, and using Assumption (ref) \[ E \Vert \hat{A}^\ast_{\varphi_{\dagger}} - A^\ast_{\varphi_{\dagger}} \Vert^2 = O\left( \frac{1}{nh_x} + h_x^{4}\right), \] so that, upon an appropriate choice of the bandwidth parameter, i.e., $h_x \asymp n^{-1/5}$, we can take $\gamma_n = n^{-2/5}$ in Assumption (ref). Both Silverman's rule-of-thumb and leave-one-out cross-validation satisfy these requirements. We use the former in our simulation study and empirical application to speed up computations.
To satisfy the conditions of Assumption (ref)(ii), we then require \[ \frac{n^{-\frac{2}{5} + \frac{1}{\rho}}\sqrt{N}}{\kappa_n} = o(1), \] because $\delta_n \vee \gamma_n = \gamma_n$. Let $\kappa_n \asymp n^{-\nu}$, with $\nu > 0$, we can pick $N$ such that
A sufficient restriction for this condition to hold is \[ \rho > \frac{5}{2 - 5\nu}, \] with $\nu \in (0,2/5)$, which implies it may be advisable to take higher-order kernels for the estimation of the cdf and pdf of $\hat{U}$.
The condition on the bias term given in Assumption (ref)(iii) has a more cumbersome interpretation, as it depends on the unknown parameters $\epsilon$ and $\kappa_n$. However we would like to choose $h_u$ that goes to $0$ as slowly as possible. In this respect, using higher-order kernel should also help reduce the bias of our estimator. Moreover, the order of the bias depends nonlinearly on $\beta$, and it is, therefore, more challenging to obtain an optimal order for $N$ based on the usual squared bias-variance trade-off. We thus propose to select the maximum number of iterations, $N_{max}$, according to equation (ref). This rule applies up to a constant, which is inversely proportional to the value of $c$ chosen (the smaller $c$, the higher the number of iterations we need for convergence), and we increase it with the variation in the dependent variable $Y$. In practice, we take this constant to be equal to $C_N = \alpha \left( Y_{max} - Y_{min} \right)/c$, with $\alpha>0$. We further take $\rho = 8$. As $\nu$, the decay at the tails of the distribution of $U$, is unknown in practice, we take it to be equal to $0$ (which implicitly assumes compact support), and then floor the maximum number of iterations $N_{max}$. That is, we take \[ N_{max} \ln(N_{max}) = \lfloor C_N n^{\frac{4 \rho - 10}{5\rho}} \rfloor = \lfloor C_N n^{\frac{11}{20}} \rfloor, \] where, for a scalar $a$, $\lfloor a \rfloor$ denotes the integer smaller than or equal to $a$. This choice of $N_{max}$ avoids overfitting whenever the error term does not have bounded support. For every $N \leq N_{max}$, we proceed as described in Section (ref): we compute the Euclidean norm of $\hat{T}(\hat\varphi_{j})$ at every iteration $j$ and take $N_0$ as the argmin of $\Vert \hat{T}(\hat\varphi_{j})\Vert^2$, for $j \in [1, N_{max}]$.
The finite sample properties are illustrated via the following simulations. The instrument $W_i\in \{0,1\}$ with probabilities $\left(\frac{1}{3}, \frac{2}{3}\right)$ and $U_i \sim N(0,1)$. The endogenous variable is given by \[ X_i = 1 + 0.5 U_i - 0.1 U_i^2 + (2 + 0.5 U_i - 0.1 U_i^2) W_i + \varepsilon_i, \] where $\varepsilon_i \sim N(0,1)$. The dependent variable is defined as follows
In both DGPs, the noise-to-signal ratio is about equal to one. We run $1000$ simulations for each DGP. The initial condition, $\hat\varphi^\prime_0$, is selected in three different ways.
All starting values are inconsistent estimators of $\varphi_{\dagger}$. The control function assumption is not satisfied in our simulations, as there is no monotone relation between $X$ and $U$. TSLS is the furthest from the true regression function, and it serves as a benchmark to assess the performance of our estimator with a poorly chosen starting value.
We start by presenting the results for $DGP_1$. The left panel of Figure (ref) shows the sample, the true curve, and the parametric and nonparametric estimators for one sample of simulated data. The dotted lines represent the starting values of our iterative procedure. The solid lines are the LF estimators obtained using each of the initial conditions, and the gray line is the correctly specified parametric estimator. Our nonparametric estimator behaves reasonably well compared to the fully parametric estimator. On the right panel, we show the decline of the Euclidean norm of $\hat{T}(\varphi)$ normalized in the interval $[0</->,1]$ for the last $500$ iterations, using different starting values, with $N_{max} = 4655$. For starting value 1, the stopping rule reaches its minimum at $N_0=1465$. Meanwhile, for starting values 2 and 3, the stopping rule is binding. Increasing the number of iterations would have further increased the variance of our estimator without necessarily improving the bias, as discussed above.
For $DGP_2$, we have similar results for starting values 1 and 2, but they substantially differ when starting from 3. The left panel of Figure (ref) shows the sample, the true curve, and the nonparametric estimator for one sample of simulated data (where the dotted lines are the starting values of our iterative procedure). Because of the approximation error, the parametric estimator is farther from the true value and has a substantial bias. The nonparametric procedure appears to behave better in this case. On the right panel, we show the behavior of the empirical square norm of $\hat{T}(\varphi)$, which reaches its minimum at $N_0 = 944$ for starting value 1, and at $N_0= 443$ for starting value $2$. For starting value $3$, we do not obtain a consistent estimator of $\varphi_{\dagger}$. $N_0 = 263$, and the final estimator is a noisy version of the starting value, and the norm of $\hat{T}(\hat\varphi_j)$ diverges as $j$ grows. This last case exemplifies the scenario in which a bad initial condition is chosen, and the LF algorithm does not converge.
The sampling distributions of both our nonparametric estimates are presented in Figure (ref), where the dotted lines are the pointwise $2.5\%$ and $97.5\%$ quantiles of the values of the estimates over $1000$ simulation draws.
We also provide evidence of our estimator performance as the sample size increases. We take the same setting and number of simulations as above, with $n = \lbrace 500,1000,2000 \rbrace$. For each DGP and initial condition, we compute the Mean Integrated Squared Error (MISE). Results are reported in Table (ref).
The estimator behaves as expected with the MISE decreasing and the median optimal number of iterations increasing with the sample size. The only exception is when we initialize $DGP_2$ using the TSLS estimator. In this case, the estimator is not consistent, the MISE increases, and $N$ decreases with $n$.
In Table (ref), we also report the median value and $95\%$ range of the chosen bandwidth parameters. We do not smooth with respect to $W$, and we use indicator functions instead. That is, $h_w = 0$ for all DGPs and sample sizes. We use a kernel of order $8$ for the residuals, which, in line with our theoretical results, leads to a rate for $h_u$ proportional to $n^{-1/16}$. The bandwidth $h_x$ is instead chosen to be proportional to $n^{-1/5}$. As $h_u$ is chosen according to the distribution of the residuals, its value changes depending on the initial condition taken. Moreover, for $DGP_2$, when starting from $\hat{\varphi}^\prime_{0,3}$, the LF algorithm does not converge. In this case, the median value of $h_u$ diverges instead of converging to $0$ at the chosen rate.
We illustrate our methodology for estimating returns to education using data from the 1979 National Longitudinal Survey of Youth card1995,torgovitsky2017. The model is the one used in the main body of the paper \[ Y_{i} = \varphi_{\dagger}(X_{i}) + U_{i}, \text{ with } i = 1,\dots,n, \] where $Y$ denotes the logarithm of the wage; $X$ represents years of education, and $E\left[ U_{i} \vert X_{i}\right] \neq 0$. To address the concern about years of schooling being discrete, we add an idiosyncratic $Unif[-1,1]$ noise to each observation. As an instrument, $W$, we use a dummy variable equal to $1$ if the individual grew up near an accredited four-year college and $0$ otherwise. The sample is restricted to the $n = 939$ individuals older than 29 and who have completed at least $8$ years of education. As the instrument is binary, we follow the nonparametric approach outlined in Section (ref). Summary statistics for these variables are provided in Table (ref).
The initial condition is taken as the first derivative of a local linear nonparametric regression ($\hat\varphi^\prime_{0,1}$), a control function approach, where the control function is the conditional cdf of $X \mid W$ ($\hat\varphi^\prime_{0,2}$), and a flexible parametric model (third order polynomial in $X$), which uses the assumption of independence between $U$ and $W$ (see Section (ref), $\hat\varphi^\prime_{0,3}$). Tuning parameters (bandwidths and regularization parameters) are selected as explained in Section (ref).
We report in Figure (ref) the results of our empirical exercise. In the figure, the solid blue, red, and green lines are our LF estimators, $\hat\varphi_{N,1}$, $\hat\varphi_{N,2}$, and $\hat\varphi_{N,3}$, respectively. The dashed blue, red, and green lines are the initial conditions, while the black solid line is the TSLS estimator. The dotted lines show the $95\%$ pointwise confidence intervals for $\hat\varphi_{N,1}$, $\hat\varphi_{N,2}$, $\hat\varphi_{N,3}$ obtained by nonparametric bootstrap with $99$ replications. The confidence intervals for the nonparametric estimator are presented only to illustrate its variability but should not be considered valid.
When starting from $\hat\varphi^\prime_{0,1}$ (blue line), we stop the LF algorithm at $N_0 = 40$ iterations. However, when starting from $\hat\varphi^\prime_{0,2}$ and $\hat\varphi^\prime_{0,3}$ (red and green lines, respectively), the LF algorithm remains at the initial condition. All nonparametric estimators confirm the log-linear relationship between education and wages, though some variation exists depending on the initial condition used. Marginal effects remain positive and are comparable to the constant effect estimated by TSLS. The wide confidence intervals (dashed-dotted lines) indicate higher uncertainty in the estimates. Assuming the instrumental variable satisfies our relevance assumptions, the exogeneity of the instrument can be violated, for instance, because of omitted variables.
In this paper, we propose a nonparametric IV estimator with continuous endogenous variables and discrete instruments. We take the IVs and the structural error term to be independent, and we discuss conditions for identification and present an estimator of the regression function and its first derivative based on a smooth iterative regularization procedure. We derive an upper bound for the Mean Integrated Squared Error and present the finite sample properties in a simulation with one continuous endogenous regressor and a binary instrument. An empirical application to estimating the returns to schooling demonstrates its features compared with a linear instrumental variable estimator. Future work should focus on providing valid inference procedures and carefully studying the case with additional exogenous controls.
\setcounter{section}{0}