EconBase
← Back to paper

Identification and Estimation of Nonseparable Triangular Equations with Mismeasured Instruments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

85,270 characters · 10 sections · 52 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Estimation of Nonseparable Triangular Equations with Mismeasured Instruments

\abovedisplayskip=15pt \belowdisplayskip=15pt \centerline{Abstract} In this paper, I study the nonparametric identification and estimation of the marginal effect of an endogenous variable $X$ on the outcome variable $Y$, given a potentially mismeasured instrument variable $W^*$, without assuming linearity or separability of the functions governing the relationship between observables and unobservables. To address the challenges arising from the co-existence of measurement error and nonseparability, I first employ the deconvolution technique from the measurement error literature to identify the joint distribution of $Y, X, W^*$ using two error-laden measurements of $W^*$. I then recover the structural derivative of the function of interest and the “Local Average Response” (LAR) from the joint distribution via the “unobserved instrument” approach in matzkin2016independence. I also propose nonparametric estimators for these parameters and derive their uniform rates of convergence. Monte Carlo exercises show evidence that the estimators I propose have good finite sample performance.

Introduction

This paper studies the nonparametric identification and estimation of the marginal effect of an endogenous variable $X$ on the outcome variable $Y$, given a potentially mismeasured instrument variable $W^*$, without assuming linearity or separability of the functions governing the relationship between observables and unobservables. Without measurement error, this type of model is referred to as “nonseparable triangular equations” and its identification was studied by chesher2003identification, imbens2009identification, and shaikh2011partial. Measurement error on the instrument variable poses additional challenges to the identification and estimation of the model primitives. Because of nonseparability, simply using an error-laden measurement as the instrument leads to inconsistent results. Because of measurement error, the true value of $W^*$ is unobserved, making it impossible to proceed using existing methods in imbens2009identification. In this paper, I propose a way to deal with the two difficulties mentioned above and show the identification of average and individual-level marginal effects. I also propose estimators for these parameters. To illustrate ideas, I study the following model:

align*[align* omitted — 61 chars of source]

where $X$ is endogenous in the sense that it's correlated with $\epsilon$, and $W^*$ is the instrument variable independent with both of the error terms $\epsilon, \eta$ but cannot be measured accurately. Following the measurement error literature, I assume that there are two error-laden measurements of $W^*$, denoted as $W_1, W_2$. Denote the corresponding measurement errors as $\Delta W_1, \Delta W_2$, i.e. $W^*=W_1+\Delta W_1$, $W^*=W_2+\Delta W_2$. I could allow a vector of observed exogenous control variables $Z$ to enter both equations and all the arguments will carry over by conditioning on $Z$, so they are omitted here for simplicity. \\ To give an example of where this model can be used, consider the Engel curve estimation studied in blundell2007semi. They used Sieve Minimum Distance to estimate a semi-nonparametric model, but the curve can be estimated under less restrictive modeling assumptions as done in imbens2009identification. $Y$ here is the share of expenditure on a commodity of a household. $X$ is the log of the household’s total expenditure. $Z$ is a vector capturing the household's demographic composition and $\epsilon$ captures the household's unobserved heterogeneity. Researchers might be interested in the average response of households' expenditure on food $Y$, to changes in household's total expenditure $X$, holding the distribution of unobserved household heterogeneity $F_{\epsilon\mid X=x}$ fixed. This parameter is called “Local Average Response” (LAR) by altonji2005cross and if the function $m$ is differentiable, it can be written as

align[align omitted — 118 chars of source]

In addition, researchers might also be interested in the structural derivative of the function $m$. This is a more disaggregate level parameter. It stands for the marginal response of $Y$ to changes in $X$ for a specific household, say the household with log total expenditure equal $\bar x$, and share of food expenditure equal $\bar y$. Denote the structural derivative of this household as $\rho(\bar y, \bar x)$. If $m$ is differentiable and strictly monotone in its second argument for all values of $X$, $\rho(\bar y, \bar x)$ can be written as

align[align omitted — 160 chars of source]

where $m^{-1}$ is the inverse of $m$ with respect to its second argument, and $m^{-1}(\bar y, \bar x)$ is the value of $\epsilon$ of the specific household one is interested in. To estimate these parameters, under the assumption that heterogeneity in earnings is not correlated with households’ preferences over consumption, one can use the income of the head of the household as the IV $W^*$. The income variable is likely to suffer from measurement errors. Researchers could obtain multiple measurements of it from panel data, for example.\\ To see how nonseparability makes it harder to identify parameters like the LAR and the structural derivative when the IV $W^*$ is mismeasured, consider the case when both equations are linear:

align*[align* omitted — 88 chars of source]

where $E[\epsilon W^*]=0$ (exclusion restriction) and $E[XW^*]\neq E[X]E[W^*]$ (relevance condition). Suppose I have $W_2=W^*+\Delta W_2$ as an error-laden measurement of $W^*$, it's not hard to verify that $E[\epsilon W_2]=0$ and $X\not\perp W_2 $ still hold under mild assumptions on $\Delta W_2$ (e.g. $E\left[\Delta W_2|Y,X\right]=0$). This means in a linear model, even if the instrument variable suffers from measurement errors, one can still use the error-laden measurement $W_2$ as an IV and proceed as usual. However, this is not the case in nonseparable models. Plugging $W_2$ into the second equation yields $X=h(W_2-\Delta W_2, \eta)$, where both $\Delta W_2$ and $\eta$ are unobervable and importantly, $W_2\not\perp\Delta W_2$. This means if I were to use $W_2$ as an IV, both of the two equations in the triangular system would contain endogenous variables and nothing can be done without additional IVs outside of this system. This also means that if one simply uses $W_2$ as the instrument variable and proceeds with the standard techniques in imbens2009identification, they would get inconsistent results.\\ Given that nonseparability makes the problem much harder to solve, a natural question is why one wants to deal with nonseparable models instead of an additive separable or linear model. There are multiple reasons why researchers might prefer a nonseparable model. First, nonseparable models allow the observed variable $X$ and the unobserved variable $\epsilon$ to interact in a flexible way. For example, a recent paper by brancaccio2020geography estimated the matching function $m(s,e)$ between ships (of number $s$) and exporters (of number $e$ which is unobserved) at a seaport. Not imposing functional form assumptions (including separability) is important. It allows the authors to remain agnostic about the nature of the meeting process. In addition, the flexibility of functional forms can be key when deriving welfare and policy implications (see brancaccio2020geography). Second, nonseparability is also important when $X$ and $\epsilon$ are correlated, since in many cases the source of endogeneity is the nonseparable nature of the model. For example, $X$ could come from the optimization problem of maximizing the expected value of $Y$ minus the production cost, given an exogenous variable $W$, and some noisy information about $\epsilon$. Think of $X$ as an individual's education level, or a firm's input level, and $Y$ as the individual's lifetime earnings, or the firm's output. Then $X$ could be the solution of $\max_{x}\left\{E[m(x,\epsilon)|\eta,W]-c(x,W)\right\}$, leading to $X=h(W,\eta)$, where $\eta$ is some noisy proxy of $\epsilon$, and $c(x,W)$ is the cost function. If the function $m$ were additively separable in $\epsilon$, the optimal choice of $x$ would not even depend on $\eta$.\\ When there is no measurement error, imbens2009identification proposed a way to identify the model primitives in a nonseparable triangular system, making use of the fact that $X\perp\epsilon|\eta$. They first estimate the control variable $\eta$ (or a strictly monotone function of $\eta$, denoted as $V$ in their paper) from the second equation, and then estimate the model primitives in the first equation by first conditioning on the estimated $\eta$, and then integrate it out. Their method cannot directly apply when the instrument variable $W^*$ in the second equation is mismeasured, because the control variable (or control function) cannot be observed from error-laden measurements of $W^*$. A recent paper by aradillas2022inference studies inference in models where control functions are unobserved. The setup of his paper is different from this paper in many aspects. He requires the availability of observable or estimable bounds for the unobserved control functions, which is not required by the model in this paper. Instead, this paper requires the availability of error-laden measurements and builds upon the measurement error literature. Also, his focus is on constructing confidence sets for finite-dimensional parameters, while my focus is on point identification and estimation of infinite-dimensional parameters.\\ In this paper, I propose a method that makes use of the same intuition as in imbens2009identification, but can deal with mismeasured $W^*$. Same as imbens2009identification, I utilize the fact that the correlation between $X$ and $\epsilon$ is merely coming from $\eta$, but instead of estimating and conditioning on $\eta$, I separate out this correlation by writing $\epsilon$ as a function of $\eta$ and a uniformly distributed random variable which is independent with $X$ and $W^*$. In this way, I am able to write the model primitives like the structural derivative and LAR as functionals of the joint distribution of $Y,X,W^*$, which can be recovered using the two error-laden measurements and the deconvolution technique developed in the measurement error literature (e.g. fan1991asymptotic, fan1993nonparametric, schennach2004estimation, schennach2004nonparametric). I can thus identify the model primitives in a constructive way and estimate them using plug-in estimators. I also derive uniform rates of convergence of the estimators.\\ This paper is most related to schennach2012local (SWC, hereafter), where the authors also consider a triangular simultaneous equations model with a mismeasured exogenous instrument. This paper differs from their paper in the modeling assumptions, parameters that can be identified, and also theoretical methods. SWC shows that under separability assumption on the second equation, the instrument conditioned marginal response\footnote{$m_x$ denotes the partial derivative of the $m$ function with respect to its first argument.} $E\left[m_x(X, \epsilon)|W^*=w^*\right]$ can be written as a ratio of the derivative of the conditional mean of $Y$ given $W^*$ over the derivative of the conditional mean of $X$ given $W^*$, both of which can be recovered from the data given two error-laden measurements $W_1, W_2$. Integrating out $W^*$, they can also recover the average response $E\left[m_x(X, \epsilon)\right]$. However, under the nonseparability of both of the equations, their method cannot recover either of the two parameters. Different from SWC, this paper shows that it's possible to identify not only the instrument-conditioned marginal response and the average marginal response but also the structural derivative and LAR, even when both of the equations are nonseparable. This conclusion, however, comes at the expense of more assumptions on the function $m$ and unobservables compared with SWC. In particular, this paper assumes strict monotonicity of the function $m$ on its second argument, and scalar unobservables, which are not required in SWC. Regarding estimation, both SWC and this paper employ plug-in estimators, but the asymptotic analysis in this paper is a bit more complex than in SWC, because of the observed variable components $Y,X$ of the joint density $f_{Y,X,W^*}$. They have non-trivial implications on the asymptotic treatment (including the convergence rates) so that SWC's asymptotic analysis cannot directly apply here. \\ This paper is also closely related to song2015estimating. In their paper, the coauthors study a nonseparable model with mismeasured endogenous variable, assuming a correctly measured control variable is available. The setup of this paper is different from theirs. This paper also studies nonseparable models with endogeneity, but instead of mismeasured endogenous variable, this paper studies the case when the endogenous variable is correctly measured, and instead of assuming the existence of a correctly measured control variable, this paper assumes that a potentially mismeasured instrument variable is available. The two papers are complementary depending on the availability of data. Regarding the parameters of interest, in addition to the various parameters studied in song2015estimating, including the average marginal response conditional on the control variable, the average marginal response, the LAR, and the weighted average version of these variables, this paper additionally show identification of the structural derivative, which is more disaggregate and could be helpful when researchers are interested heterogeneous effects.\\ The rest of the paper consists of six parts. Section 2 introduces the model and assumptions. Section 3 talks about the identification of the model primitives. Section 4 proposes plug-in estimators of the model primitives identified in Section 3, and talks about their asymptotic properties. Section 5 conducts limited Monte Carlo studies to show the finite sample performance of the estimator. Section 6 concludes.

The Model

I consider the following triangular model in this paper:

align*[align* omitted — 63 chars of source]

where $X$ is an observed endogenous variable that is correlated with $\epsilon$. In the returns to education example, $X$ stands for years of education, which is correlated with $\epsilon$ since it's chosen by the agent as an equilibrium outcome. Z is a vector of observed exogenous variables which are independent with $\epsilon$ and $\eta$. $W^*$ is an instrument for $X$ and satisfies $W^*\perp (\epsilon, \eta)$. Researchers cannot measure $W^*$ exactly but have two error-laden measurements of it:

align*[align* omitted — 76 chars of source]

The measurement errors satisfy $E[\Delta W_1\mid W^*, \Delta W_2]=0$ and $\Delta W_2\perp W^*, Y, X, Z$. As in the literature studying the identification of nonseparable models (chesher2003identification, matzkin2003nonparametric,altonji2005cross,matzkin2015estimation), I impose monotonicity assumptions on the structural functions. I assume that $m$ is strictly increasing in $\epsilon$ and that $h$ is strictly increasing in $\eta$. Since the vector of covariates $Z$ is correctly measured and exogenous, all the analysis can be done conditional on $Z$. For brevity, from now on I omit the vector $Z$ in my notations and work on the simplified model below, while other assumptions on the structural functions and distributions of variables remain unchanged.

align[align omitted — 58 chars of source]

The assumptions mentioned above are stated formally below:

assumptionFunction $m$ and $h$ are continuously differentiable with respect to both of their arguments and are strictly increasing in their respective second argument.
assumption$W^*\perp (\epsilon,\eta)$.
assumption$W_1 = W^*+\Delta W_1$ and $W_2 = W^*+\Delta W_2$.
assumption$E[\Delta W_1\mid W^*, \Delta W_2]=0$, $\Delta W_2\perp W^*, Y, X$.

Following the measurement error literature, I also impose:

assumptionFor any finite $t\in\mathbb{R}$, $\left|E[\exp\left(\mathbf{i}tW_2\right)]\right|>0$.

Note that for the measurement error $\Delta W_1$, only the mean independence assumption is imposed, which is weaker than the assumption on $\Delta W_2$. This weaker assumption is sufficient to identify the distribution of $W^*$ (schennach2004nonparametric). On the other hand, although $\Delta W_2$ satisfies strong independence assumptions, making $W_2$ a measurement with classical measurement error, one still cannot use $W_2$ as the instrument because of the nonseparability of the model. More specifically, plugging $W_2$ into the second equation in the triangular system yields

align*[align* omitted — 44 chars of source]

This is a nonseparable model with two unobservables $\Delta W_2$ and $\eta$, and $W_2$ is not independent with $\Delta W_2$. This means both of the two equations in the system contain endogenous variables and thus cannot be identified without further assumptions.\\ In addition, I also impose the following assumptions on the distribution of $Y,X,W^*$:

assumptionThe distribution of $(Y,X,W^*)$ has compact support, denoted as $\mathbb{S}_{(Y,X,W^*)}$ and has a joint density $f_{Y,X,W^*}$ which is twice continuously differentiable on $R^3$.
assumptionFor any $x,w^*$ belonging to the support of $X,W^*$, $\left|\frac{\partial F_{X\mid W^*=w^*}(x)}{\partial w^*}\right|>0$.

Assumption (ref) is a sufficient condition to ensure $\left.\frac{\partial h(w^*,\eta)}{\partial w^*}\right|_{\eta = r(x,w^*)}\neq 0$ for all $x,w^*$ on the support of $X,W^*$. This assumption is like the rank condition imposed in the linear instrument variable models. To illustrate this idea, suppose $h$ is a linear function such that $X= \pi W^*+\eta$, then $F_{X\mid W^*=w^*}(x) = F_{\eta}(x-\pi w^*)$, and $\pi\neq 0$ is a neccessary condition for Assumption (ref).

Identification

In this section, I show the identification of the model parameters in two steps. First, assuming that the joint distribution of $Y, X, W^*$ is known, I show identification of the derivative of the structural function $m$ with respect to $x$, when the value of $\epsilon$ is fixed, and the identification of other model parameters including the LAR and AR. Then I show how to identify the joint distribution of $Y, X, W^*$ from the two measurements $W_1$ and $W_2$.

Identification of parameters assuming the joint distribution of $Y, X, W^*$ is known

The independence between $W^*$ and $\epsilon,\eta$ implies that conditional on $\eta$, $W^*$ and $\epsilon$ are independent. The next lemma uses this fact to show that one can write $\epsilon$ as a function of $\eta$ and another random variable which is independent with $X, W^*$. This lemma is similar to Proposition 5.1 in matzkin2016independence. Before stating the lemma, I first impose the following assumption on the conditional distribution of $\epsilon$ given $\eta$:

assumption$F_{\epsilon\mid \eta=\bar\eta}(\bar\epsilon)$ is strictly increasing in $\bar\epsilon$, given any values of $\bar\eta$.
lemmaUnder the model setup, suppose Assumption (ref) is satisfied. There exists a function $s:\mathbb{R}^2\rightarrow\mathbb{R}$ strictly increasing in its second argument and an unobservable random term $\delta$ such that \begin{align} \epsilon = s(\eta, \delta), \end{align} and $\delta$ is independent of $(X,W^*)$ and is $U(0,1)$.
proofSee Appendix B.

Next, I plug ((ref)) into the first structural equation. The assumption that $h$ is strictly increasing in $\eta$ implies that one can write the inverse of $h(W^*,\eta)$ w.r.t. $\eta$ as $r(X, W^*)$. Then I have

align[align omitted — 119 chars of source]

Equation ((ref)) builds a bridge between the structural function $m$ and the reduced form function $v$. Note that $v$ is a function of observable variables $X,W^*$ and an unobservable variable $\delta$ which is independent of the observables. The derivatives of $v$ can be identified by applying identification techniques in standard nonseparable models (matzkin2003nonparametric). To identify the main parameter of interest, the derivative of the structural function $m$, one can utilize the last equality in (ref): $m(X,s(r(X,W^*),\delta))\equiv v(X,W^*,\delta)$. Taking derivative w.r.t $X$ and $W^*$ yields:

footnotesize\begin{align} &\left.\frac{\partial m(x,\epsilon)}{\partial x}\right|_{\epsilon =s(r(x, w^*),\delta)} + \left.\frac{\partial m(x,\epsilon)}{\partial \epsilon}\right|_{\epsilon =s(r(x, w^*),\delta)}\cdot\left.\frac{\partial s(\eta, \delta)}{\partial \eta}\right|_{\eta = r(x,w^*)}\cdot \frac{\partial r(x,w^*)}{\partial x} = \frac{\partial v(x,w^*,\delta)}{\partial x}\\ &\left.\frac{\partial m(x,\epsilon)}{\partial \epsilon}\right|_{\epsilon =s(r(x, w^*),\delta)}\cdot\left.\frac{\partial s(\eta, \delta)}{\partial \eta}\right|_{\eta = r(x,w^*)}\cdot \frac{\partial r(x,w^*)}{\partial w^*} = \frac{\partial v(x,w^*,\delta)}{\partial w^*} \end{align}

Plug ((ref)) into ((ref)) to cancel $\left.\frac{\partial m(x,\epsilon)}{\partial \epsilon}\right|_{\epsilon =s(r(x, w^*),\delta)}\cdot\left.\frac{\partial s(\eta, \delta)}{\partial \eta}\right|_{\eta = r(x,w^*)}$ yields

align*[align* omitted — 282 chars of source]

Taking derivative w.r.t. $W^*$ on both sides of $X\equiv h(W^*,r(X,W^*))$ and cancelling out the unobserved $\left.\frac{\partial h(w^*,\eta)}{\partial \eta}\right|_{\eta=r(x,w^*)}$, one can get $\left.\frac{\partial h(w^*,\eta)}{\partial w^*}\right|_{\eta=r(x,w^*)} = -\frac{\frac{\partial r(x,w^*)}{\partial w^*}}{\frac{\partial r(x,w^*)}{\partial x}}$. Then one can write

align[align omitted — 290 chars of source]

The right-hand side of the equation ((ref)) can be identified from the data. To show this, note that

align[align omitted — 472 chars of source]

Similarly

align*[align* omitted — 355 chars of source]

where $\delta$ is the unique value such that $y=v(x,w^*,\delta)$ and $\eta$ is the unique value such that $ x=h(w^*,\eta) $. Plug into ((ref)), one can get

align[align omitted — 573 chars of source]

for all $\bar y, \bar x$ belongs to support, and values of $\bar\epsilon$ such that $\bar y = m(\bar x, \bar\epsilon)$. Note that as long as $\bar y$ and $\bar x$ are fixed, $\bar\epsilon$ is fixed, no matter which value is picked for $w^*$. This means there's actually overidentification for the structural derivative. To identify the LAR and AR, one needs independent variations of $\epsilon$ given $X$. Rewrite ((ref)) with slightly different notations yields:

align*[align* omitted — 157 chars of source]

so that one can write

align[align omitted — 359 chars of source]

To show the identification of the LAR and AR, one needs to show that the support of $s(r(X, W^*), \delta)$ given $X=x$ is the same as the support of $\epsilon$ given $X=x$. This is stated formally in the lemma below.

lemmaDefine random variable $\mathbf{s}\equiv s(r(X, W^*), \delta)$. Denote the support of $\mathbf{s}$ and $\epsilon$ conditional on $X=x$ as $\mathbb{S}_{\mathbf{s}\mid X=x}$ and $\mathbb{S}_{\epsilon\mid X=x}$, respectively. If Assumption (ref) holds, then $\mathbb{S}_{\mathbf{s}\mid X=x} = \mathbb{S}_{\epsilon\mid X=x}$, for all $x$ belonging to its support.
proofSee Appendix B.

Lemma (ref) and (ref) ensures that $F_{\mathbf{s}\mid X}$ is the same as $F_{\epsilon\mid X}$, so that integrating $\frac{\partial m(x,\epsilon)}{\partial x}$ over $\epsilon$ given $X=x$ will be equivalent to integrating $\frac{\partial m(x,\mathbf{s})}{\partial x}$ over $\mathbf{s}$ given $X=x$. Then I have the following identification results:

align[align omitted — 466 chars of source]

Identification of the joint density of $Y, X, W^*$

First note that by Theorem 1 in schennach2004estimation

align*[align* omitted — 131 chars of source]

Then I can identify the density $f_{W^*}(w^*)$ by taking the Inverse Fourier Transform:

align*[align* omitted — 77 chars of source]

For the joint density $f_{Y,X,W^*}(y, x, w^*)$, by Assumption (ref), I have the following convolution:

align*[align* omitted — 90 chars of source]

Applying Fourier transformation on both sides (holding $x$ constant) yields

align*[align* omitted — 105 chars of source]

which, by Assumption (ref) implies

align[align omitted — 201 chars of source]

Then applying the inverse Fourier transform, one can get

align[align omitted — 151 chars of source]

and similarly,

align[align omitted — 161 chars of source]

Then $F_{X\mid W^*=w^*}(x)$ follows from

align[align omitted — 141 chars of source]

and $F_{Y\mid X=x, W^*=w^*}(y)$ follows from

align*[align* omitted — 142 chars of source]

From the equations above, one can write the derivatives of the conditional CDFs as functionals of the joint densities:

footnotesize\begin{align*} &\frac{\partial F_{Y\mid X=x, W^*=w^*}(y)}{\partial w^*} \\ &= \frac{\left(\int_{-\infty}^y \frac{\partial f_{Y,X,W^*}(s,x,w^*)}{\partial w^*}ds\right)f_{X,W^*}(x,w^*)-\left(\int_{-\infty}^y f_{Y,X,W^*}(s,x,w^*)ds\right)\frac{\partial f_{X,W^*}(x,w^*)}{\partial w^*}}{f^2_{X,W^*}(x,w^*)}\\ &\frac{\partial F_{Y\mid X=x, W^*=w^*}(y)}{\partial x}\\ &=\frac{\left(\int_{-\infty}^y \frac{\partial f_{Y,X,W^*}(s,x,w^*)}{\partial x}ds\right)f_{X,W^*}(x,w^*)-\left(\int_{-\infty}^y f_{Y,X,W^*}(s,x,w^*)ds\right)\frac{\partial f_{X,W^*}(x,w^*)}{\partial x}}{f^2_{X,W^*}(x,w^*)}\\ &\frac{\partial F_{Y\mid X=x, W^*=w^*}(y)}{\partial y}=\frac{f_{Y,X,W^*}(y,x,w^*)}{f_{X,W^*}(x,w^*)}\\ &\frac{\partial F_{X\mid W^*=w^*}(x)}{\partial w^*}=\frac{\left(\int_{-\infty}^{x} \frac{\partial f_{X,W^*}(s,w^*)}{\partial w^*}ds\right) f_{W^*}(w^*)-\left(\int_{-\infty}^{x} f_{X,W^*}(s,w^*)ds\right)\frac{\partial f_{W^*}(w^*)}{\partial w^*}}{f^2_{W^*}(w^*)}\\ &\frac{\partial F_{X\mid W^*=w^*}(x)}{\partial x}=\frac{f_{X,W^*}(x,w^*)}{ f_{W^*}(w^*)}. \end{align*}

This means that the right-hand side of the equation ((ref)) can be written as a functional of the joint density $f_{Y,X, W^*}$, which can be identified from the data. To show identification of the left-hand side of the equation ((ref)), one needs to build the relationship between the derivatives of conditional quantiles and the joint densities, which is shown in the following lemma:

lemmaLet $\delta\in[0,1]$ be a constant, then \begin{align*} &\frac{\partial F^{-1}_{Y|X=x, W^*=w^*}(\delta)}{\partial w^*} = \frac{\delta \frac{\partial f_{X,W^*}(x,w^*)}{\partial w^*}-\int_{-\infty}^{F^{-1}_{Y|X=x, W^*=w^*}(\delta)}\frac{\partial f_{Y,X,W^*}(y,x,w^*)}{\partial w^*}dy}{f_{Y,X,W^*}\left(F^{-1}_{Y|X=x, W^*=w^*}(\delta),x,w^*\right)}\\ &\frac{\partial F^{-1}_{Y|X=x, W^*=w^*}(\delta)}{\partial x} = \frac{\delta \frac{\partial f_{X,W^*}(x,w^*)}{\partial x}-\int_{-\infty}^{F^{-1}_{Y|X=x, W^*=w^*}(\delta)}\frac{\partial f_{Y,X,W^*}(y,x,w^*)}{\partial x}dy}{f_{Y,X,W^*}\left(F^{-1}_{Y|X=x, W^*=w^*}(\delta),x,w^*\right)}. \end{align*}
proofsee Appendix B.

The identification of the left-hand side of ((ref)) and ((ref)) is a straightforward implication of Lemma (ref) and equation ((ref)).

Estimation

The Estimator

From now on, I use the following function to denote the joint density of $Y, X,W^*$, or its derivative with respect to $Y, X,W^*$:

align*[align* omitted — 193 chars of source]

where $\lambda_1, \lambda_{2,1}, \lambda_{2,2}\in\{0,1\}$ and $\lambda\equiv\max\{\lambda_1, \lambda_{2,1},\lambda_{2,2}\}\leq 1$. The $0-th$ order derivative of a function is defined as the function itself. For convenience of notation, I also define $\lambda_2\equiv\max\{ \lambda_{2,1},\lambda_{2,2}\}$. In this section, I first propose an estimator for $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$, and then propose plug-in estimators for the structural derivative and a weighted average version of the LAR since they can be written as known functionals of $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$. \\ To make the expressions more transparent and clear, I show the explicit forms of the functionals by taking the structural derivative as an example. I write each component on the right-hand-side of equation (ref) as

align*[align* omitted — 1,678 chars of source]

For the LAR, I will introduce a weighted average version of it, and propose an estimator for this weighted average version of the LAR. The aim of introducing the weight is to ensure that the integration is taken on a set where the joint density $f_{Y,X,W^*}$ is bounded away from zero. This has the benefit of guaranteeing that the relevant functional is Fr\'echet differentiable when $f_{Y,X,W^*}$ appears in the denominator of it. The weighted average version of the LAR is defined as follows:

align[align omitted — 312 chars of source]

where $\omega(x,w^*,\epsilon)$ is a known or estimable weighting function taking value 0 outside a compact set $\mathbb M$. As a weighting function, $\omega(x,w^*,\epsilon)$ also satisfies that given any values of $x$ such that $\omega(x,\cdot,\cdot)$ could obtain non-zero values, $\int_{w^*}\int_\epsilon\omega(x,w^*,\epsilon)f_{\epsilon, W^*|X=x}(\epsilon, w^*) d\epsilon dw^*=1$. The specific form of $\omega$ is up to the researcher's choice, however, to ensure easy calculation of $\omega(x,w^*,s(r(x,w^*),\delta))$ on the right-hand side of ((ref)), I require that function $\omega(x,w^*,\epsilon)$ satisfy the following restriction: $\omega(x,w^*,\epsilon)=\tilde \omega(x,w^*)$ if $\epsilon \in \left[q_{x,w^*}(\tau_l), q_{x,w^*}(\tau_u)\right]$ and $\omega(x,w^*,\epsilon)=0$ otherwise, where $\tilde\omega(x,w^*)$ is a known function with compact support, and $q_{x,w^*}(\tau)$ denotes the conditional-$\tau$ quantile of $\epsilon$ given $X=x, W^*=w^*$. The specific form of function $\tilde w$ and the specific values of $\tau_l$ and $\tau_u$ are up to the researcher's choice. Under this restriction, $\omega(x,w^*,s(r(x,w^*),\delta))$ is equal to $\tilde w(x,w^*)$ if $\delta$ is between $\tau_l$ and $\tau_u$, and is equal to 0 otherwise. An example of function $\omega(x,w^*,\epsilon)$ is when $\tilde \omega(x,w^*)$ is a constant equal to $\left[(\tau_u-\tau_l)\int_{\underline{w}^*}^{\bar w^*} f_{W^*|X=x}(w^*)dw^*\right]^{-1}$. The WLAR in this case can be interpreted as the Local Average Response of a subgroup of individuals whose unobserved $\epsilon$ is between its $\tau_l$ and $\tau_u$ conditional quantile given $X, W^*$ and whose $W^*$ is between $\underline w^*$ and $\bar w^*$. Note that even though $X$ is correlated with $\epsilon$, neither the distribution of $\epsilon|X$ nor the subgroup of individuals changes as I take the derivative with respect to $x$. This means that my objective of interest is the same group of people when studying the effect of the counterfactual changes of the endogenous $X$. The WLAR can also be written as a functional of $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$. The explicit form can be derived from Lemma (ref) and equation ((ref)).\\ So far I have written the parameters of interest as functionals of $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$ in explicit forms. Next I focus on the estimation of $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$. To deal with the well-known ill-posed inverse problem when inverting a convolution operator, I base my estimator on a smoothed version of $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$:

align*[align* omitted — 208 chars of source]

and use kernel functions satisfying the following assumption:

assumptionThe kernel functions $K: \mathbb{R}\rightarrow \mathbb{R}$, $G_Y: \mathbb{R}\rightarrow \mathbb{R}$ and $G_X: \mathbb{R}\rightarrow \mathbb{R}$ are measurable, symmetric. $\int K(x)dx=1$, $\int G_Y(y)dy=1$, $\int G_X(x)dx=1$. $G_Y$, $G_X$ are differentiable. $G_Y$, $G_X$, and their derivatives are bounded, and denote the maximum of these bounds as $\bar G$. Their Fourier transforms $\xi\rightarrow\phi_{K}(\xi)$, $\xi\rightarrow\phi_{G_Y}(\xi)$ and $\xi\rightarrow\phi_{G_X}(\xi)$ obey: (i) $\phi_{F}$ is compactly supported (without loss of generality, the support is $[-1,1]$) for $F\in \{K, G_X, G_Y\}$; and (ii) there exists $\bar \xi_F$ such that $\phi_{F}(\xi)=1$ for $|\xi|\leq\bar\xi_F$, where $F\in \{K,G_X,G_Y\}$.

By assuming (ii), I use flat-top kernels proposed by politis1999multivariate. The benefit of using flat-top kernels is that the rate of decrease of the bias term will not be affected by the order of the kernel, and will only be affected by the smoothness of the function to be estimated. The restriction of compact support of the Fourier transform is without loss of generality. As discussed by schennach2004nonparametric, given any kernel $K$, one can always create a modified kernel $\tilde K$ that satisfies the assumption by using a “windowing” function. For example,

align*[align* omitted — 53 chars of source]

where

align*[align* omitted — 239 chars of source]

for some $\bar t\in(0,1)$.

assumption$E[|\Delta W_1|]<\infty$.
comment\textcolor{red}{see song2015estimating P758 for arguments}

The following lemma shows that the smoothed version $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_1)$ can be written as an expression involving the Fourier transform of the kernel function and the characteristic functions of different variables. This expression makes it more straightforward to introduce the estimator that will appear soon below.

lemmaFor $(y,x,w^*)\in\mathbb{S}_{(Y,X,W^*)}$ and $h>0$, let \begin{align*} &g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_1)\equiv\int \frac{1}{h_1}K\left(\frac{\tilde w^*-w^*}{h_1}\right)g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,\tilde w^*)d\tilde w^*. \end{align*} where $K$ satisfies Assumption (ref). Then under Assumptions (ref), (ref), and (ref), \begin{align*} &g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_1) = \frac{1}{2\pi}\int \left(-\mathbf{i}t\right)^{\lambda_1}e^{-\mathbf{i}tw^*} \phi_K(h_1t)\frac{\phi_{f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W_2}(y,x, \cdot)}(t)\phi_{W^*}(t)}{\phi_{W_2}(t)} dt \end{align*} where $f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W_2}(y,x,w_2)\equiv \frac{\partial^{\lambda_2}f_{Y,X,W_2}(y,x,w_2)}{\partial y^{\lambda_{2,1}}\partial x^{\lambda_{2,2}}}$.
proofSee Appendix B.

I now define my estimator for $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_1)$ by replacing $\phi_{f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W_2}(y,x,\cdot)}(t)$, $\phi_{W^*}(t)$, and $\phi_{W_2}(t)$ in Lemma (ref) with their sample analogues. Formally,

definitionLet $h_n\equiv(h_{1n}, h_{2n,1}, h_{2n,2})\rightarrow 0 $ as $n\rightarrow 0$. The estimator for $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$ is defined as \begin{align*} &\hat g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_n)\equiv \frac{1}{2\pi}\int\left(-\mathbf{i}t\right)^{\lambda_1} e^{-\mathbf{i}tw^*} \phi_K(h_{1n}t)\frac{\hat\phi_{f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W_2}(y,x,\cdot)}(t)\hat\phi_{W^*}(t)}{\hat\phi_{W_2}(t)} dt, \end{align*} where \begin{align*} &\hat\phi_{f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W_2}(y,x,\cdot)}(t) \equiv \hat E\left[e^{\mathbf{i}tW_{2}}\frac{1}{h_{2n,1}^{1+\lambda_{2,1}}h_{2n,2}^{1+\lambda_{2,2}}}G_Y^{(\lambda_{2,1})}\left(\frac{y-Y}{h_{2n,1}}\right)G_X^{(\lambda_{2,2})}\left(\frac{x-X}{h_{2n,2}}\right)\right],\\ &\hat\phi_{W_2}(t)\equiv\hat E\left[e^{\mathbf{i}tW_2}\right],\\ &\hat\phi_{W^*}(t) \equiv \exp\left(\int_0^{t}\frac{\hat E\left[iW_1 e^{i\xi W_2}\right]}{\hat E\left[e^{i\xi W_2}\right]}d\xi\right), \end{align*} where $\hat E$ denotes the sample average, $G_Y, G_X$ denotes the kernel functions for $Y,X$, respectively, and $G_X^{(\lambda_{2,2})}(x)\equiv\frac{\partial^{\lambda_{2,2}} G_X(x)}{\partial x^{\lambda_{2,2}}}$. $G_Y, G_X$ are not necessarily the same as $K$.

Having defined $\hat g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x ,w^*,h_n)$, I can now define the estimator for $\rho(\bar y, \bar x)$ by replacing $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$ with its estimator on the right-hand-side of Equation ((ref)). Note that by Fubini's Theorem

footnotesize\begin{align*} &\quad\int_{-\infty}^y \hat g_{\lambda_1,0,0, \lambda_{2,2}}(s,x,w^*,h_n)ds=\frac{1}{2\pi}\int\left(-\mathbf{i}t\right)^{\lambda_1} e^{-\mathbf{i}tw^*} \phi_K(h_{1n}t)\frac{\int_{-\infty}^y\hat\phi_{f^{(0,\lambda_{2,2})}_{Y, X,W_2}(s,x, \cdot)}(t)ds\hat\phi_{W^*}(t)}{\hat\phi_{W_2}(t)} dt\\ &\quad\int_{-\infty}^\infty \hat g_{\lambda_1,0,0, \lambda_{2,2}}(y,x,w^*,h_n)dy=\frac{1}{2\pi}\int\left(-\mathbf{i}t\right)^{\lambda_1} e^{-\mathbf{i}tw^*} \phi_K(h_{1n}t)\frac{\hat\phi_{f^{(0,\lambda_{2,2})}_{X,W_2}(x, \cdot)}(t)\hat\phi_{W^*}(t)}{\hat\phi_{W_2}(t)} dt\\ &\quad\int_{-\infty}^x\int_{-\infty}^\infty\hat g_{\lambda_1,0,0}(y,s,w^*,h_n)dyds=\frac{1}{2\pi}\int\left(-\mathbf{i}t\right)^{\lambda_1} e^{-\mathbf{i}tw^*} \phi_K(h_{1n}t)\frac{\int_{-\infty}^x\hat\phi_{f_{X,W_2}(s, \cdot)}(t)ds\hat\phi_{W^*}(t)}{\hat\phi_{W_2}(t)} dt\\ &\quad\int_{-\infty}^\infty\int_{-\infty}^\infty\hat g_{\lambda_1,0,0}(y,x,w^*,h_n)dydx=\frac{1}{2\pi}\int\left(-\mathbf{i}t\right)^{\lambda_1} e^{-\mathbf{i}tw^*} \phi_K(h_{1n}t)\hat\phi_{W^*}(t) dt, \end{align*}

where I've let $\tilde G_Y(y)\equiv\int_{-\infty}^y G_Y\left(u\right)du$ and $\tilde G_X(x)\equiv\int_{-\infty}^x G_X\left(u\right)du$ denote the kernel CDFs, and I've defined:

align*[align* omitted — 606 chars of source]

The Asymptotic Properties of the Estimator

In this subsection, I analyze the asymptotic properties for $\hat g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_n)$. I begin by decomposing the difference between the estimator and the true value of the parameter as the sum of a bias term, a variance term, and a remainder term.

lemmaSuppose $\left\{Y_i, X_i, W^*_i, \Delta W_{1,i}, W_{2,i}\right\}$ is an IID sequence satisfying Assumptions (ref), (ref), (ref), (ref), and (ref), and that Assumption (ref) holds. Then for $(y,x,w^*)\in\mathbb{S}_{(Y,X,W^*)}$, and $h>0$, \begin{align*} &\quad\hat g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*,h)-g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*) \\ &= B_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*, h)+L_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*, h)+R_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*, h) \end{align*} where $B_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*, h)$ is the bias term admitting the linear representation: \begin{footnotesize} \begin{align*} &\quad B_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*, h) \\ &= E\left[\bar g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*,h)\right] - g_{\lambda_1,\lambda_{2,1},\lambda_{2,2}}(y,x,w^*)\\ &= g_{\lambda_1,\lambda_{2,1},\lambda_{2,2}}(y,x,w^*,h_1)-g_{\lambda_1,\lambda_{2,1},\lambda_{2,2}}(y,x,w^*)\\ &\quad+\frac{1}{2\pi}\int\left(-\mathbf{i}t\right)^{\lambda_1} e^{-\mathbf{i}tw^*} \phi_K(h_{1n}t)\frac{E\left[\hat\phi_{f^{(\lambda_{2,1}, \lambda_{2,2})}_{Y,X,W_2}(y,x,\cdot)}(t)-\phi_{f^{(\lambda_{2,1}, \lambda_{2,2})}_{Y,X,W_2}(y,x,\cdot)}(t)\right]\phi_{W^*}(t)}{\phi_{W_2}(t)} dt; \end{align*} \end{footnotesize} $L_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*, h)$ is the variance term admitting the linear representation: \begin{footnotesize} \begin{align*} &\quad L_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*, h) \\ &= \bar g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*,h)-E\left[\bar g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*,h)\right] \\ &=\hat E\Bigg[\int \Psi_{1, \lambda_1,\lambda_{2,1},\lambda_{2,2}}(\xi, y, x, w^*, h_1)\left(W_1e^{\mathbf{i}\xi W_2}-E[W_1e^{\mathbf{i}\xi W_2}]\right)d\xi\notag\\ &\quad+\int \Psi_{2,\lambda_1,\lambda_{2,1},\lambda_{2,2}}(\xi, y, x, w^*, h_1)\left(e^{\mathbf{i}\xi W_2}-E[e^{\mathbf{i}\xi W_2}]\right)d\xi\notag\\ &\quad+\int \Psi_{3,\lambda_1,\lambda_{2,1},\lambda_{2,2}}(\xi, y, x, w^*, h_1)\times\\ &\quad\quad\quad\quad\left(e^{\mathbf{i}\xi W_2}\frac{1}{h_{2,1}^{1+\lambda_{2,1}}h_{2,2}^{1+\lambda_{2,2}}}G_Y^{(\lambda_{2,1})}\left( \frac{y-Y}{h_{2,1}}\right)G_X^{(\lambda_{2,2})}\left(\frac{x-X}{h_{2,2}}\right)\right.\\ &\quad\quad\quad\quad\quad\quad\left.-E\left[e^{\mathbf{i}\xi W_2}\frac{1}{h_{2,1}^{1+\lambda_{2,1}}h_{2,2}^{1+\lambda_{2,2}}}G_Y^{(\lambda_{2,1})}\left( \frac{y-Y}{h_{2,1}}\right)G_X^{(\lambda_{2,2})}\left(\frac{x-X}{h_{2,2}}\right)\right]\right) d\xi\Bigg]\notag\\ &=\hat E[l_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*,h;Y, X, W_1, W_2)], \end{align*} \end{footnotesize} where I've let $\theta(\xi) \equiv E\left[W_1 e^{\mathrm{i} \xi W_{2}}\right]$ and defined \begin{align*} &\Psi_{1, \lambda_1,\lambda_{2,1},\lambda_{2,2}}(\xi, y, x, w^*, h_1) = \frac{1}{2 \pi} \frac{\mathbf{i}}{\phi_{W_2}(\xi)}\int_{\xi}^{\pm\infty} \left(-\mathbf{i}t\right)^{\lambda_1} e^{-\mathbf itw^*} \phi_{K}(h_1t)\phi_{f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W^*}(y,x,\cdot)}(t)dt \\ &\Psi_{2,\lambda_1,\lambda_{2,1},\lambda_{2,2}}(\xi, y, x, w^*, h_1) \\ &= -\frac{1}{2 \pi} \frac{\mathbf{i}\theta(\xi)}{(\phi_{W_2}(\xi))^2}\int_{\xi}^{\pm\infty} \left(-\mathbf{i}t\right)^{\lambda_1}e^{-\mathbf itw^*} \phi_{K}(h_1 t)\phi_{f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W^*}(y,x,\cdot)}(t)dt\\ &\quad -\frac{1}{2 \pi} \left(-\mathbf{i}t\right)^{\lambda_1}e^{-\mathbf i\xi w^*} \phi_{K}(h _1\xi)\frac{\phi_{f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W^*}(y,x,\cdot)}(\xi)}{\phi_{W_2}(\xi)}\\ &\Psi_{3,\lambda_1,\lambda_{2,1},\lambda_{2,2}}(\xi, y, x, w^*, h_1) = \frac{1}{2 \pi} \left(-\mathbf{i}t\right)^{\lambda_1}e^{-\mathbf i\xi w^*} \phi_{K}(h_1\xi)\frac{\phi_{W^*}(\xi)}{\phi_{W_2}(\xi)}. \end{align*} where for a given function $\zeta\rightarrow f(\zeta)$, I've written $\int_{\xi}^{\pm \infty}f(\zeta)d\zeta\equiv\lim_{c\rightarrow+\infty}\int_{\xi}^{c\xi}f(\zeta)d\zeta$; and $R_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*, h)$ is the remainder term: \begin{align*} R_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*, h) = \hat g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*,h)-\bar g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x, w^*,h). \end{align*}
proofSee Appendix B.

To derive the uniform rate of convergence, I impose some assumptions on the smoothness of $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$, stated in terms of the rate of decay of the tail of its Fourier transform:

assumption(i) There exist constants $C_{\phi}>0$, $\alpha_{\phi}\leq0$, $\beta_{\phi}\geq0$ and $\gamma_{\phi}\in\mathbb{R}$ such that $\beta_{\phi}\gamma_{\phi}\geq 0$ and for $j=1,2$ \begin{align*} &\max_{\lambda_2\in\{0,1\}}\sup_{(y,x)\in\mathbb{S}_{(Y,X)}}\left|\phi_{f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W^*}(y,x, \cdot)}(t)\right| \leq C_{\phi}(1+|t|)^{\gamma_{\phi}} \exp \left(\alpha_{\phi}|t|^{\beta_{\phi}}\right)\\ &\left|\phi_{W^*}(t)\right| \leq C_{\phi}(1+|t|)^{\gamma_{\phi}} \exp \left(\alpha_{\phi}|t|^{\beta_{\phi}}\right), \end{align*} where $f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W_2}(y,x,w_2)\equiv \frac{\partial^{\lambda_2}f_{Y,X,W_2}(y,x,w_2)}{\partial y^{\lambda_{2,1}}\partial x^{\lambda_{2,2}}}$. Moreover, if $\beta_{\phi}=0$, then for given $\lambda \in\{0,1\}, \gamma_{\phi}<-\lambda_1-1$.\\ (ii) Denote the Fourier transform of $f^{(\lambda_{2,1}, \lambda_{2,2})}_{Y,X,W^*}(y,x,w^*)$ as $\phi_{f^{(\lambda_{2,1}, \lambda_{2,2})}}(t_1,t_2,\xi)$. There exists constants $C_{f_1}>0$, $C_{f_2}>0$, $\alpha_{f_1}\leq 0$, $\alpha_{f_2}\leq 0$, $\beta_{f_1}\geq 0$, $\beta_{f_2}\geq 0$ and $\gamma_{f_1}\in\mathbb{R}$, $\gamma_{f_2}\in\mathbb{R}$ such that $\beta_{f_1}\gamma_{f_1}\geq 0$, $\beta_{f_2}\gamma_{f_2}\geq 0$, and \begin{small} \begin{align*} \left|\phi_{f^{(\lambda_{2,1}, \lambda_{2,2})}}(t_1,t_2,\xi)\right| \leq C_{f_1}C_{f_2}(1+|t_1|)^{\gamma_{f_1}}(1+|t_2|)^{\gamma_{f_2}}\exp\left(\alpha_{f_1}|t_1|^{\beta_{f_1}}+\alpha_{f_2}|t_2|^{\beta_{f_2}}\right)\left|\phi_{W^*}(t)\right|. \end{align*} \end{small} Moreover, if $\beta_{f_1}=0$, $\gamma_{f_1}<-1$, and if $\beta_{f_2}=0$, $\gamma_{f_2}<-1$.

In Assumption (ref)(i) I impose the same bound for the tail behavior of $\phi_{f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W^*}(y,x,\cdot)}$ and $\phi_{W^*}(t)$. This is without loss of generality since they have the same effect on the convergence rate. Assumption (ref)(ii) is similar to Assumption A7 in li2002robust, and can be viewed as a generalization of smoothness condition from the univariate case to the multivariate case.\\ I next state the first main result of this paper, the uniform asymptotic rate of the bias term:

theoremLet the conditions of Lemma (ref) hold with $\{Y_i,X_i,W_{i}^*, \Delta W_{1,i}, \Delta W_{2,i}\}$ IID, and suppose in addition that Assumption (ref) holds. Then for $h>0$, \begin{align*} \sup_{(y,x,w^*)\in\mathbb{S}_{(Y,X,W^*)}}\left|B_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h)\right| = O\left(\left(h_1^{-1}\right)^{1+\gamma_{\phi}+\lambda_1}\exp\left(\alpha_{\phi}\bar\zeta_K^{\beta_{\phi}}\left(h_1^{-1}\right)^{\beta_{\phi}}\right)\right). \end{align*}
proofSee Appendix B.

I impose the following assumption to ensure finite variance:

assumptionFor some $\delta>0$, $E[\left|W_1\right|^{2+\delta}]<\infty$, $\sup_{w_2\in\mathbb{S}_{W_2}}E\left[\left|W_1\right|^{2+\delta}\mid W_2=w_2\right]<\infty$ .

To derive the rate for $L_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h)$, I impose the following assumption on the tail behavior of the Fourier transforms involved. These are common in the deconvolution literature (e.g. fan1991optimal, fan1993nonparametric, and schennach2004nonparametric).

assumption\begin{enumerate}[(i)] • There exist constants $C_{2}>0$, $\alpha_{2}\leq 0$, $\beta_{2}\geq\beta_{\phi}\geq 0$ and $\gamma_{2}\in\mathbb{R}$ such that $\beta_{2}\gamma_{2}\geq 0$ and \begin{align*} \left|\phi_{W_2}(t)\right| \geq C_{2}(1+|t|)^{\gamma_{2}} \exp \left(\alpha_{2}|t|^{\beta_{2}}\right). \end{align*} Moreover, if $\beta_2=0$, $\lambda_1-\gamma_2+\gamma_{\phi}>0$. • There exist constants $C_*>0$ and $\gamma_*\geq 0$ such that \begin{align*} \left|\frac{\phi'_{W^*}(t)}{\phi_{W^*}(t)}\right|\leq C_*\left(1+|t|\right)^{\gamma_*}. \end{align*} \end{enumerate}

I explicitly impose $\beta_2\geq\beta_{\phi}$ because

align*[align* omitted — 409 chars of source]

The following theorem states the asymptotic properties of the linear term \\ $L_{\lambda_1, \lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h)$ and facilitates the analysis of various quantities of interest later.

theoremSuppose the conditions of Lemma (ref) hold. (i) Then for each $(y, x, w^*)\in\mathbb{S}_{(y, x, w^*)}$, $E[L_{\lambda_1, \lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h)]=0$, and if Assumption (ref) also holds, then \begin{footnotesize} $E[L_{\lambda_1, \lambda_{2,1}, \lambda_{2,2}}^2(y,x,w^*,h)]$$=n^{-1}\Omega_{\lambda_1, \lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h)$ \end{footnotesize}, where \begin{footnotesize} $\Omega_{\lambda_1, \lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h)\equiv E[\left(l_{\lambda_1, \lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h;Y, X, W_1, W_2)\right)^2]$ \end{footnotesize}.\\ Further, if Assumption (ref) and (ref) also holds then \begin{footnotesize} \begin{align} &\quad\sqrt{\sup_{(y, x, w^*)\in\mathbb{S}_{(Y,X,W^*)}}\Omega_{\lambda_1, \lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h)}\notag\&= O\bigg(\max\left\{(h_1^{-1})^{1+\gamma_*},(h_{2,1}^{-1})^{1+\lambda_{2,1}}(h_{2,2}^{-1})^{1+\lambda_{2,2}}\right\}\notag\\ &\quad\quad\quad\quad\left.(h_1^{-1})^{1-\gamma_2+\gamma_{\phi}+\lambda_1}\exp\left(\left(\alpha_{\phi}\mathbf{1}\{\beta_{\phi}=\beta_2\}-\alpha_2\right)\left(h_1^{-1}\right)^{\beta_2}\right)\right). \end{align} \end{footnotesize} I also have \begin{footnotesize} \begin{align} &\quad\sup_{(y, x, w^*)\in\mathbb{S}_{(Y,X,W^*)}} \left|L_{\lambda_1, \lambda_{2,1}, \lambda_{2,2}}(y, x, w^*,h)\right|\notag\\ &=O\bigg(n^{-\frac{1}{2}}\max\left\{(h_1^{-1})^{1+\gamma_*},(h_{2,1}^{-1})^{1+\lambda_{2,1}}(h_{2,2}^{-1})^{1+\lambda_{2,2}}\right\}\notag\\ &\quad\quad\quad\quad\left.(h_1^{-1})^{1-\gamma_2+\gamma_{\phi}+\lambda_1}\exp\left(\left(\alpha_{\phi}\mathbf{1}\{\beta_{\phi}=\beta_2\}-\alpha_2\right)\left(h_1^{-1}\right)^{\beta_2}\right)\right). \end{align} \end{footnotesize} (ii) If Assumption (ref) also holds, and if for each $(y,x,w^*)\in\mathbb{S}_{(y, x, w^*)}$, $\Omega_{\lambda_1, \lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_n)>0$ for all $n$ sufficiently large, then for each $(y, x, w^*)\in\mathbb{S}_{(y,x,w^*)}$ \begin{align*} n^{1/2}\left(\Omega_{\lambda_1, \lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_n)\right)^{-1/2}L_{\lambda_1, \lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_n)\xrightarrow{d}N(0,1). \end{align*}
proofSee Appendix B.
comment\begin{assumption}\textcolor{red}{for asymptotic normality} If $\beta_2=0$ in Assumption (ref), then $h_n^{-1}=O\left(n^{-\eta} n^{(3/2)/(4-\gamma_2+\gamma_{W^*}+\gamma_*)}\right)$ for some $\eta>0$; otherwise $h^{-1}_n=O\left(\left(\ln n\right)^{\beta_2^{-1}-\eta}\right)$ for some $\eta>0$. \end{assumption}

Finally, I bound the remainder term. To do that, I first impose some restrictions on the moments of $W_2$:

assumption$E[|W_2|]< \infty, E[|W_1W_2|]<\infty$.

I then impose the following bounds for the bandwidths:

assumptionIf $\beta_2=0$ in Assumption (ref), then $h_{1n}^{-1}=O\left(n^{(4+4\gamma_*-4\gamma_{2})^{-1}-\eta}\right)$ and $h_{2n,j}^{-1}=O\left(n^{(16+8\lambda_2)^{-1}-\eta}\right)$, for some $\eta>0$, $j=1,2$; otherwise $h^{-1}_{1n}=O\left(\left(\ln n\right)^{\beta_2^{-1}-\eta}\right)$ and $h_{2n,j}^{-1}=O\left(n^{(8+4\lambda_2)^{-1}-\eta}\right)$, for some $\eta>0$, $j=1,2$.
theorem(i) Suppose the conditions of Lemma (ref) and Assumption (ref), (ref), (ref), and (ref) hold. Then \begin{footnotesize} \begin{align*} &\quad\sup_{(y,x,w^*)\in\mathbb{S}_{(Y,X,W^*)}}\left|R_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_n)\right| \\ &=o_p\left(h_{2,1}^{-2}(h_{2,2}^{-1})^{2+2\lambda_2}n^{-1+2\epsilon}(h_{1n}^{-1})^{1+\gamma_{*}-\gamma_2}\exp\left(-\alpha_2\left(h_{1n}^{-1}\right)^{\beta_2}\right)\right)\\ &\quad\times O\bigg(\max\left\{(h_1^{-1})^{1+\gamma_*},(h_{2,1}^{-1})^{1+\lambda_{2,1}}(h_{2,2}^{-1})^{1+\lambda_{2,2}}\right\}\\ &\quad\quad\quad\quad\left.(h_1^{-1})^{1-\gamma_2+\gamma_{\phi}+\lambda_1}\exp\left(\left(\alpha_{\phi}\mathbf{1}\{\beta_{\phi}=\beta_2\}-\alpha_2\right)\left(h_1^{-1}\right)^{\beta_2}\right)\right) \end{align*} \end{footnotesize} for arbitrarily small $\epsilon>0$. (ii) If Assumption (ref) also holds, then \begin{footnotesize} \begin{align*} &\quad\sup_{(y,x,w^*)\in\mathbb{S}_{(Y,X,W^*)}}\left|R_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_n)\right| =\\ &o_p\bigg(n^{-\frac{1}{2}}\max\left\{(h_1^{-1})^{1+\gamma_*},(h_{2,1}^{-1})^{1+\lambda_{2,1}}(h_{2,2}^{-1})^{1+\lambda_{2,2}}\right\}\\ &\quad\quad\quad\quad\left.(h_1^{-1})^{1-\gamma_2+\gamma_{\phi}+\lambda_1}\exp\left(\left(\alpha_{\phi}\mathbf{1}\{\beta_{\phi}=\beta_2\}-\alpha_2\right)\left(h_1^{-1}\right)^{\beta_2}\right)\right). \end{align*} \end{footnotesize}
proofSee Appendix B.

Collecting the results from Theorem (ref),(ref) and (ref) yields a straightforward corollary of the uniform rate of convergence of $\hat g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_n)$:

corollaryIf the conditions of Theorem (ref) (ii) hold, then \begin{align*} &\quad\sup_{(y,x,w^*)\in \mathbb{S}_{(Y,X,W^*)}}\left|\hat g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_n)-g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)\right|\\ &=O_p\left(\epsilon_{n,\lambda_1}\right)+O_p\left(\tilde\epsilon_{n,\lambda_1,\lambda_{2,1},\lambda_{2,2}}\right). \end{align*} where \begin{footnotesize} \begin{align*} &\epsilon_{n,\lambda_1}\equiv\left(h_1^{-1}\right)^{1+\gamma_{\phi}+\lambda_1}\exp\left(\alpha_{\phi}\bar\zeta_K^{\beta_{\phi}}\left(h_1^{-1}\right)^{\beta_{\phi}}\right)\\ &\tilde\epsilon_{n,\lambda_1,\lambda_{2,1},\lambda_{2,2}}\\ &\equiv n^{-1/2}\max\left\{(h_1^{-1})^{1+\gamma_*},(h_{2,1}^{-1})^{1+\lambda_{2,1}}(h_{2,2}^{-1})^{1+\lambda_{2,2}}\right\}(h_1^{-1})^{1-\gamma_2+\gamma_{\phi}+\lambda_1}\\ &\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\quad\times\exp\left(\left(\alpha_{\phi}\mathbf{1}\{\beta_{\phi}=\beta_2\}-\alpha_2\right)\left(h_1^{-1}\right)^{\beta_2}\right). \end{align*} \end{footnotesize}

The uniform (over the whole support) rate of convergence of the kernel estimators of the density or its derivatives has been considered in andrews1995nonparametric, hansen2008uniform and schennach2012local\footnote{schennach2012local studies the case both with and without measurement errors. Their Theorem 3.2 corresponds to the no-measurement error case, and Corollary 4.7 corresponds to the case with measurement errors.}. hansen2008uniform obtained a faster rate than andrews1995nonparametric and schennach2012local (Theorem 3.2, for the no-measurement error case), but requires more assumptions on the kernel functions (their Assumption 1 and 3), including bounded support or an integrable tail, which, unfortunately, are not satisfied by infinite order kernels. andrews1995nonparametric's conclusion is stated assuming a kernel with finite order. Infinite order kernels are not essential when there is no measurement error (like in andrews1995nonparametric and hansen2008uniform), but are especially advantageous when there are measurement errors. If infinite order kernels are used, only the smoothness of the functions, but not the order of the kernels will affect the rate of convergence. Corollary (ref) is similar to Corollary 4.7 in schennach2012local.\\ With the uniform rate of convergence of $\hat g$, I can then derive the uniform rate of convergence of the plug-in estimators of the structural derivative $\rho(y,x)$ and the WLAR, denoted as $\widehat{\rho(y,x)}$ and $\widehat{\mathbf{WLAR}(x)}$, respectively.

theoremSuppose that $\left\{Y_j, X_j, W^*_j, \Delta W_{1,j}, W_{2,j}\right\}$ is an IID sequence satisfying the conditions of Corollary (ref) with $\lambda_1,\lambda_{2,1},\lambda_{2,2}\in\{0,1\}$ and $\max\{\lambda_1,\lambda_{2,1},\lambda_{2,2}\}\leq 1$. Suppose in addition, Assumption (ref) and (ref) hold. \\ (i) Define $\mathbb S_{\tau_n}\equiv \left\{(y,x,w^*)\in\mathbb{S}_{(Y,X,W^*)}: \left|\frac{\partial F_{X\mid W^*=w^*}(x)}{\partial w^*}\right|>\tau_n \text{ and } f_{Y,X,W^*}(y,x, w^*)>\tau_n\right\}$. Then \begin{align*} &\sup_{(y,x,w^*)\in \mathbb S_{\tau_n}}\left|\widehat{\rho(y,x)}-\rho(y,x)\right|\leq O_p\left(\tilde\epsilon_{n,0,0,1}\right)+\frac{O_p\left(\epsilon_{n,1}\right)+O_p\left(\tilde\epsilon_{n,1,0,0}\right)}{\tau_n^2}. \end{align*} and there exists $\{\tau_n\}$ such that $\tau_n>0, \tau_n\rightarrow 0$ as $n\rightarrow\infty$, and \begin{align*} &\sup_{(y,x,w^*)\in \mathbb S_{\tau_n}}\left|\widehat{\rho(y,x)}-\rho(y,x)\right|= o_p(1). \end{align*} (ii) \begin{align*} \sup_{x\in\mathbb{S}_X}\left|\widehat{\mathbf{WLAR}(x)}-\mathbf{WLAR}(x)\right|\leq O_p\left(\epsilon_{n,1}\right)+O_p\left(\tilde\epsilon_{n,0,0,1}\right)+O_p\left(\tilde\epsilon_{n,1,0,0}\right). \end{align*}
proofSee Appendix B.

In Appendix C, I state an asymptotic normality result for $\hat g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}} (y,x,w^*,h_n)$ and $\widehat{\rho(y,x)}$. The proof of them requires a lower bound on $\Omega_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*.h_n)$ relative to $B_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*.h_n)$ and $R_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*.h_n)$. The assumptions to ensure the lower bound are stated at a high level, and more primitive sufficient conditions need to be derived.

Monte Carlo Simulations

In this section, I conduct some Monte Carlo simulation exercises to study the finite sample performance of my proposed estimators. I consider two simulation designs. Design 1 is when both of the equations are linear, and Design 2 is when the equations are nonlinear.\\ Design 1

align*[align* omitted — 61 chars of source]

Design 2

align*[align* omitted — 108 chars of source]

In both cases, $\epsilon$ and $\eta$ are correlated through a common component $\theta$: $\epsilon = \theta + \epsilon_1$, and $\eta = \theta + \eta_1$, where $\theta$, $\epsilon_1$ and $\eta_1$ are mutually independent, and are independent with $W^*$. In both designs, I have two error-laden measurements $W_1 = W^*+\Delta W_1, W_2 = W^*+\Delta W_2$, where $\Delta W_1$ and $\Delta W_2$ are independent and they are independent with all other variables. The distributions of variables in both designs are listed in Table (ref) below:

table[table omitted — 662 chars of source]

In the estimation, I use the following flat-top kernel\footnote{I may use different flat-top kernels for $G_X$ and $G_Y$. For simplicity I use the same flat-top kernel.} proposed by politis1999multivariate:

align*[align* omitted — 69 chars of source]

I use the sample size of 500 and replicate the estimation 500 times. \\ First I show the performance of my proposed method versus two other methods for estimating the structural derivative: (1) 2SLS using the error-laden measurement $W_2$ as IV; (2) a plug-in estimator replacing $W^*$ with $W_2$ in my identification equation ((ref)). Method (1) will be valid under the linear design (Design 1), since $W_2$, although error-laden, still satisfies the exclusion restriction and the rank condition for linear IV estimation. However, because of its misspecification of the model, it won't be valid under the nonlinear design (Design 2). Method (2) won't be valid under either the linear or nonlinear design, as the estimator will converge to a population value that is not equal to the right-hand side of equation ((ref)) in general. I estimate $\rho(y,x)$ evaluated at $y=0, x=0, w^*=0.7$ for Design 1, and $y=0.6, x=0.6, w^*=7$ for Design 2, yielding a true value of $0.25$ and $0.4512$, respectively. There are three bandwidths used in the estimation: $h_1$, $h_{21}$, $h_{22}$. The latter two correspond to estimating the unknown distribution of observed variables $X$ and $Y$. I thus choose the values of $h_{21}$ and $h_{22}$ by cross-validation (i.e. minimizing the estimated MISE of $\hat f_{Y,X}(y,x)$). For the bandwidth $h_1$, I scan a set of values ranging from 0.5 to 3 for my proposed estimator and a set of values ranging from 0.5 to 6 for the method that uses the error-laden measurement. Table (ref), (ref), and (ref) show the MSE, VAR, and the absolute value of BIAS of my proposed estimator, Method (2), and Method (1), respectively. \\ From Table (ref), one can see that the bias from the error-contaminated estimator does not shrink toward zero as bandwidth decreases. Comparing results from Table (ref) and Table (ref), one can see that my estimator also gives smaller variances than the error-contaminated method. As a result, my estimator performs better than the error-contaminated method in terms of MSE under both the linear and the nonlinear designs. Comparing Table (ref) with Table (ref), one can see that under the linear design, both my estimator and 2SLS work well except that my estimator has a slightly larger variance. Under the nonlinear design, the bias and MSE of 2SLS are much larger than my estimator, which aligns with the theory.\\ Table (ref) below shows the Monte Carlo simulation results of my proposed estimator as a function of the sample size. The value of $h_{21}$ and $h_{22}$ are selected by cross-validation for the respective sample sizes and the optimal values of $h_1$ are selected by minimizing the MSE in the corresponding set of values the same as when $N=500$. The MSE, VAR, and Bias decrease as the sample size increases, which is in accordance with the theory.

landscape\begin{table}[p] \caption{Monte Carlo simulation results of my proposed estimator for $\rho(y,x)$} \begin{tabular}{cccccc|cccccc} \hline \hline \multicolumn{6}{c|}{Design 1 (linear)} & \multicolumn{6}{c}{Design 2 (nonlinear)} \\ $h_1$ & $h_{21}$ & $h_{22}$ & MSE & VAR & abs(BIAS) & $h_1$ & $h_{21}$ & $h_{22}$ & MSE & VAR & abs(BIAS) \\ \hline \multicolumn{1}{l}{0.5} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 0.01072 & 0.01072 & 0.00009 & \multicolumn{1}{l}{0.5} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 0.31292 & 0.30489 & 0.08961 \\ \multicolumn{1}{l}{0.75} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 0.00766 & 0.00766 & 0.00050 & \multicolumn{1}{l}{0.75} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 0.17724 & 0.16339 & 0.11769 \\ \multicolumn{1}{l}{1} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 0.00728 & 0.00728 & 0.00138 & \multicolumn{1}{l}{1} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 0.15425 & 0.13973 & 0.12050 \\ \multicolumn{1}{l}{1.25} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 0.00738 & 0.00738 & 0.00253 & \multicolumn{1}{l}{1.25} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 0.14824 & 0.13360 & 0.12098 \\ \multicolumn{1}{l}{1.5} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 0.00755 & 0.00754 & 0.00328 & \multicolumn{1}{l}{1.5} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 0.14611 & 0.13129 & 0.12175 \\ \multicolumn{1}{l}{1.75} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 0.00772 & 0.00771 & 0.00376 & \multicolumn{1}{l}{1.75} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 0.14540 & 0.13034 & 0.12273 \\ \multicolumn{1}{l}{2} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 0.00789 & 0.00787 & 0.00408 & \multicolumn{1}{l}{2} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 0.14540 & 0.13008 & 0.12380 \\ \multicolumn{1}{l}{2.25} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 0.00806 & 0.00804 & 0.00431 & \multicolumn{1}{l}{2.25} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 0.14581 & 0.13023 & 0.12482 \\ \multicolumn{1}{l}{2.5} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 0.00822 & 0.00820 & 0.00447 & \multicolumn{1}{l}{2.5} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 0.14640 & 0.13060 & 0.12569 \\ \multicolumn{1}{l}{2.75} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 0.00836 & 0.00834 & 0.00461 & \multicolumn{1}{l}{2.75} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 0.14707 & 0.13109 & 0.12640 \\ \multicolumn{1}{l}{3} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 0.00851 & 0.00849 & 0.00474 & \multicolumn{1}{l}{3} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 0.14777 & 0.13163 & 0.12703 \\ \multicolumn{6}{c|}{Optimal} & \multicolumn{6}{c}{Optimal} \\ $h_1$ & $h_{21}$ & $h_{22}$ & MSE & VAR & BIAS & $h_1$ & $h_{21}$ & $h_{22}$ & MSE & VAR & BIAS \\ 1 & 1.05 & 2.92 & 0.00728 & 0.00728 & 0.00138 & 1.75 & 2.09 & 0.88 & 0.14540 & 0.13034 & 0.12273 \\ \hline \hline \end{tabular} \end{table}
landscape\begin{table}[p] \caption{Monte Carlo simulation results of plugging in error-laden $W_2$ to ((ref)) to estimate $\rho(y,x)$ (Method (2) above)} \begin{tabular}{cccccc|cccccc} \hline \hline \multicolumn{6}{c|}{Design 1 (linear)} & \multicolumn{6}{c}{Design 2 (nonlinear)} \\ $h_1$ & $h_{21}$ & $h_{22}$ & MSE & VAR & abs(BIAS) & $h_1$ & $h_{21}$ & $h_{22}$ & MSE & VAR & abs(BIAS) \\ \hline \multicolumn{1}{l}{0.5} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 831.49239 & 827.49615 & 1.99906 & \multicolumn{1}{l}{0.5} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 342390.30868 & 342380.03811 & 3.20477 \\ \multicolumn{1}{l}{1} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 15.36480 & 15.18324 & 0.42611 & \multicolumn{1}{l}{1} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 3268.04002 & 3264.78862 & 1.80316 \\ \multicolumn{1}{l}{1.5} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 1609.24991 & 1605.98439 & 1.80708 & \multicolumn{1}{l}{1.5} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 795.12130 & 795.09105 & 0.17391 \\ \multicolumn{1}{l}{2} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 981.89810 & 980.63434 & 1.12417 & \multicolumn{1}{l}{2} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 1113.49474 & 1113.43014 & 0.25417 \\ \multicolumn{1}{l}{2.5} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 2739.03247 & 2729.13653 & 3.14578 & \multicolumn{1}{l}{2.5} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 7439.31131 & 7399.31171 & 6.32452 \\ \multicolumn{1}{l}{3} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 48.28533 & 48.27034 & 0.12243 & \multicolumn{1}{l}{3} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 2589.00861 & 2583.05235 & 2.44054 \\ \multicolumn{1}{l}{3.5} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 227.69069 & 227.32253 & 0.60677 & \multicolumn{1}{l}{3.5} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 7367.37049 & 7363.00243 & 2.08999 \\ \multicolumn{1}{l}{4} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 326.60891 & 325.63719 & 0.98576 & \multicolumn{1}{l}{4} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 4810.95117 & 4805.88266 & 2.25133 \\ \multicolumn{1}{l}{4.5} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 29.37044 & 29.30256 & 0.26055 & \multicolumn{1}{l}{4.5} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 12005.60726 & 11951.43663 & 7.36007 \\ \multicolumn{1}{l}{5} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 7904.23502 & 7887.80380 & 4.05354 & \multicolumn{1}{l}{5} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 863.16882 & 862.91399 & 0.50480 \\ \multicolumn{1}{l}{5.5} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 16.49962 & 16.48066 & 0.13769 & \multicolumn{1}{l}{5.5} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 6884.05497 & 6883.55213 & 0.70911 \\ \multicolumn{1}{l}{6} & \multicolumn{1}{l}{1.05} & \multicolumn{1}{l}{2.92} & 54.17185 & 54.16807 & 0.06147 & \multicolumn{1}{l}{6} & \multicolumn{1}{l}{2.09} & \multicolumn{1}{l}{0.88} & 844.96901 & 843.85308 & 1.05637 \\ \multicolumn{6}{c|}{Optimal} & \multicolumn{6}{c}{Optimal} \\ $h_1$ & $h_{21}$ & $h_{22}$ & MSE & VAR & BIAS & $h_1$ & $h_{21}$ & $h_{22}$ & MSE & VAR & BIAS \\ 1 & 1.05 & 2.92 & 15.36480 & 15.18324 & 0.42611 & 1.5 & 2.09 & 0.88 & 795.12130 & 795.09105 & 0.17391 \\ \hline \hline \end{tabular} \end{table}
landscape\begin{table}[p] \caption{Monte Carlo simulation results of using 2SLS and error-laden $W_2$ to estimate $\rho(y,x)$ (Method (1) above)} \begin{tabular}{ccc|ccc} \hline \hline \multicolumn{3}{c|}{Design 1 (linear)} & \multicolumn{3}{c}{Design 2 (nonlinear)} \\ MSE & VAR & abs(BIAS) & MSE & VAR & abs(BIAS) \\ \hline 0.00267 & 0.00267 & 0.00190 & 1.11718 & 0.00147 & 1.05627 \\ \hline \hline \end{tabular} \end{table} \begin{table}[p] \caption{Monte Carlo simulation results of my $\hat\rho(y,x)$ as a function of the sample size} \begin{tabular}{ccccccc|ccccccc} \hline \hline \multicolumn{7}{c|}{Design 1 (linear)} & \multicolumn{7}{c}{Design 2 (nonlinear)} \\ $N$ & $h_1$ & $h_{21}$ & $h_{22}$ & MSE & VAR & abs(BIAS) & $N$ & $h_1$ & $h_{21}$ & $h_{22}$ & MSE & VAR & abs(BIAS) \\ \hline 500 & 1 & 1.05 & 2.92 & 0.00728 & 0.00728 & 0.00138 & 500 & 1.75 & 2.09 & 0.88 & 0.14540 & 0.13034 & 0.12273 \\ 1000 & 1 & 0.97 & 2.61 & 0.00354 & 0.00353 & 0.00359 & 1000 & 1.5 & 1.58 & 0.87 & 0.08341 & 0.06822 & 0.12323 \\ 5000 & 1.25 & 0.82 & 2.05 & 0.00102 & 0.00102 & 0.00026 & 5000 & 1.75 & 0.93 & 0.65 & 0.03053 & 0.03048 & 0.00708 \\ \hline \hline \end{tabular} \end{table}

Next, I study the performance of my estimator of the weighted LAR. For Design 1 (linear), I evaluate the WLAR at $x=0$ and set the weighting function to be $\omega(0,w^*,\epsilon)=\omega_1$ if $\epsilon\in\left[q_{0,w^*}(0.25),q_{0,w^*}(0.35)\right]$ and $w^*\in\left[0.70, 0.90\right]$. The constant $\omega_1$ is selected such that $\int_{0.7}^{0.9}\int_{q_{0,w^*}(0.25)}^{q_{0,w^*}(0.35)}\omega_1f_{\epsilon,W^*|X=0}(\epsilon,w^*)d\epsilon d w^*=1$. For Design 2 (nonlinear), I evaluate the WLAR at $x=0.6$ and set the weighting function to be $\omega(0.6,w^*,\epsilon)=\omega_2$ if $\epsilon\in\left[q_{0.6,w^*}(0.25),q_{0.6,w^*}(0.35)\right]$ and $w^*\in\left[6, 6.23\right]$. The constant $\omega_2$ is selected to ensure that $\int_{6}^{6.23}\int_{q_{0.6,w^*}(0.25)}^{q_{0.6,w^*}(0.35)}\omega_2f_{\epsilon,W^*|X=0.6}(\epsilon,w^*)d\epsilon d w^*=1$. These choices of the weighting functions yield a true value of the WLAR of 0.25 and 0.5032 for Design 1 and 2, respectively. In the estimation, I use the optimal values of $h_1$ coming from Table (ref), and the same values of $h_{21}$ and $h_{22}$ as in Table (ref). Table (ref) below shows the results of my estimators of the WLAR under both designs. The performance of my estimator of the WLAR is comparable to the performance of my estimator for the point-wise structural derivative $\rho(y,x)$ (shown in Table (ref)). This shows evidence that my estimator could not only perform well at single points but also could maintain good performance in a global sense. The supports of the weighting functions I used here are not large. This is mainly due to computation constraints. One direction of future work is to look at the performance of my estimator for WLAR when the support of the weighting function is larger. This will give us a better understanding of the global performance of my estimator.

table[table omitted — 573 chars of source]

The Monte Carlo simulation exercise shown in this section is very limited. To have a comprehensive examination of the performance of my estimators, there are several future directions that I can work on. First, in my current exercise, the distribution of the measurement error $\Delta W_2$ is set to be $\chi^2$, which is ordinarily smooth. I can further explore the performance of my estimators under different distributions of the measurement error $\Delta W_2$ and other variables. Second, when estimating the WLAR, I set the support for the weighting function to be small because of the computation constraint. With more time and more computation power, I can look at the performance of my estimator when the support of the weighting function in the WLAR is larger, i.e. the WLAR is taking the average marginal effect over a larger population. Third, the way I select $h_1$ is by minimizing the actual MSE, i.e., the mean squared difference between my estimator and the true value of the parameter. This is infeasible in practice when one deals with real-world data. It will be helpful if I could propose a feasible way to select the bandwidth.

Conclusion

I study the nonparametric identification and estimation of the nonseparable triangular equations model when the instrument variable $W^*$ is mismeasured. I don't assume linearity or separability of the functions governing the relationship between observables and unobservables. To deal with the challenges caused by the co-existence of the measurement error and nonseparability, I first employ the deconvolution technique developed in the measurement error literature (schennach2004estimation,schennach2004nonparametric) to identify the joint distribution of $Y, X, W^*$ using two error-laden measurements of $W^*$. I then recover the structural derivative of the function of interest and the “Local Average Response” (LAR) via the “unobserved instrument” approach in matzkin2016independence. Based on the constructive identification results, I propose plug-in nonparametric estimators for these parameters and derive their uniform rates of convergence. I also conducted limited Monte Carlo exercises to show the finite sample performance of my estimators. \\ I recognize that there are some important future directions for this paper. First, in the main text, I only demonstrated the uniform rate of convergence of the estimators. The appendix contains proofs of their asymptotic normality, which rely on high-level assumptions. To further enhance my results, it would be beneficial to derive more primitive sufficient conditions for these assumptions. Second, as discussed in Section (ref), to conduct a comprehensive examination of the performance of my estimators, additional Monte Carlo studies are necessary. It will also be extremely useful if I could propose a feasible way to select the bandwidth $h_1$. Last, it will be an interesting exercise to apply my proposed method to real-world data.