Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
85,270 characters · 10 sections · 52 citation commands
Identification and Estimation of Nonseparable Triangular Equations with Mismeasured Instruments
\abovedisplayskip=15pt \belowdisplayskip=15pt \centerline{Abstract} In this paper, I study the nonparametric identification and estimation of the marginal effect of an endogenous variable $X$ on the outcome variable $Y$, given a potentially mismeasured instrument variable $W^*$, without assuming linearity or separability of the functions governing the relationship between observables and unobservables. To address the challenges arising from the co-existence of measurement error and nonseparability, I first employ the deconvolution technique from the measurement error literature to identify the joint distribution of $Y, X, W^*$ using two error-laden measurements of $W^*$. I then recover the structural derivative of the function of interest and the “Local Average Response” (LAR) from the joint distribution via the “unobserved instrument” approach in matzkin2016independence. I also propose nonparametric estimators for these parameters and derive their uniform rates of convergence. Monte Carlo exercises show evidence that the estimators I propose have good finite sample performance.
This paper studies the nonparametric identification and estimation of the marginal effect of an endogenous variable $X$ on the outcome variable $Y$, given a potentially mismeasured instrument variable $W^*$, without assuming linearity or separability of the functions governing the relationship between observables and unobservables. Without measurement error, this type of model is referred to as “nonseparable triangular equations” and its identification was studied by chesher2003identification, imbens2009identification, and shaikh2011partial. Measurement error on the instrument variable poses additional challenges to the identification and estimation of the model primitives. Because of nonseparability, simply using an error-laden measurement as the instrument leads to inconsistent results. Because of measurement error, the true value of $W^*$ is unobserved, making it impossible to proceed using existing methods in imbens2009identification. In this paper, I propose a way to deal with the two difficulties mentioned above and show the identification of average and individual-level marginal effects. I also propose estimators for these parameters. To illustrate ideas, I study the following model:
where $X$ is endogenous in the sense that it's correlated with $\epsilon$, and $W^*$ is the instrument variable independent with both of the error terms $\epsilon, \eta$ but cannot be measured accurately. Following the measurement error literature, I assume that there are two error-laden measurements of $W^*$, denoted as $W_1, W_2$. Denote the corresponding measurement errors as $\Delta W_1, \Delta W_2$, i.e. $W^*=W_1+\Delta W_1$, $W^*=W_2+\Delta W_2$. I could allow a vector of observed exogenous control variables $Z$ to enter both equations and all the arguments will carry over by conditioning on $Z$, so they are omitted here for simplicity. \\ To give an example of where this model can be used, consider the Engel curve estimation studied in blundell2007semi. They used Sieve Minimum Distance to estimate a semi-nonparametric model, but the curve can be estimated under less restrictive modeling assumptions as done in imbens2009identification. $Y$ here is the share of expenditure on a commodity of a household. $X$ is the log of the household’s total expenditure. $Z$ is a vector capturing the household's demographic composition and $\epsilon$ captures the household's unobserved heterogeneity. Researchers might be interested in the average response of households' expenditure on food $Y$, to changes in household's total expenditure $X$, holding the distribution of unobserved household heterogeneity $F_{\epsilon\mid X=x}$ fixed. This parameter is called “Local Average Response” (LAR) by altonji2005cross and if the function $m$ is differentiable, it can be written as
In addition, researchers might also be interested in the structural derivative of the function $m$. This is a more disaggregate level parameter. It stands for the marginal response of $Y$ to changes in $X$ for a specific household, say the household with log total expenditure equal $\bar x$, and share of food expenditure equal $\bar y$. Denote the structural derivative of this household as $\rho(\bar y, \bar x)$. If $m$ is differentiable and strictly monotone in its second argument for all values of $X$, $\rho(\bar y, \bar x)$ can be written as
where $m^{-1}$ is the inverse of $m$ with respect to its second argument, and $m^{-1}(\bar y, \bar x)$ is the value of $\epsilon$ of the specific household one is interested in. To estimate these parameters, under the assumption that heterogeneity in earnings is not correlated with households’ preferences over consumption, one can use the income of the head of the household as the IV $W^*$. The income variable is likely to suffer from measurement errors. Researchers could obtain multiple measurements of it from panel data, for example.\\ To see how nonseparability makes it harder to identify parameters like the LAR and the structural derivative when the IV $W^*$ is mismeasured, consider the case when both equations are linear:
where $E[\epsilon W^*]=0$ (exclusion restriction) and $E[XW^*]\neq E[X]E[W^*]$ (relevance condition). Suppose I have $W_2=W^*+\Delta W_2$ as an error-laden measurement of $W^*$, it's not hard to verify that $E[\epsilon W_2]=0$ and $X\not\perp W_2 $ still hold under mild assumptions on $\Delta W_2$ (e.g. $E\left[\Delta W_2|Y,X\right]=0$). This means in a linear model, even if the instrument variable suffers from measurement errors, one can still use the error-laden measurement $W_2$ as an IV and proceed as usual. However, this is not the case in nonseparable models. Plugging $W_2$ into the second equation yields $X=h(W_2-\Delta W_2, \eta)$, where both $\Delta W_2$ and $\eta$ are unobervable and importantly, $W_2\not\perp\Delta W_2$. This means if I were to use $W_2$ as an IV, both of the two equations in the triangular system would contain endogenous variables and nothing can be done without additional IVs outside of this system. This also means that if one simply uses $W_2$ as the instrument variable and proceeds with the standard techniques in imbens2009identification, they would get inconsistent results.\\ Given that nonseparability makes the problem much harder to solve, a natural question is why one wants to deal with nonseparable models instead of an additive separable or linear model. There are multiple reasons why researchers might prefer a nonseparable model. First, nonseparable models allow the observed variable $X$ and the unobserved variable $\epsilon$ to interact in a flexible way. For example, a recent paper by brancaccio2020geography estimated the matching function $m(s,e)$ between ships (of number $s$) and exporters (of number $e$ which is unobserved) at a seaport. Not imposing functional form assumptions (including separability) is important. It allows the authors to remain agnostic about the nature of the meeting process. In addition, the flexibility of functional forms can be key when deriving welfare and policy implications (see brancaccio2020geography). Second, nonseparability is also important when $X$ and $\epsilon$ are correlated, since in many cases the source of endogeneity is the nonseparable nature of the model. For example, $X$ could come from the optimization problem of maximizing the expected value of $Y$ minus the production cost, given an exogenous variable $W$, and some noisy information about $\epsilon$. Think of $X$ as an individual's education level, or a firm's input level, and $Y$ as the individual's lifetime earnings, or the firm's output. Then $X$ could be the solution of $\max_{x}\left\{E[m(x,\epsilon)|\eta,W]-c(x,W)\right\}$, leading to $X=h(W,\eta)$, where $\eta$ is some noisy proxy of $\epsilon$, and $c(x,W)$ is the cost function. If the function $m$ were additively separable in $\epsilon$, the optimal choice of $x$ would not even depend on $\eta$.\\ When there is no measurement error, imbens2009identification proposed a way to identify the model primitives in a nonseparable triangular system, making use of the fact that $X\perp\epsilon|\eta$. They first estimate the control variable $\eta$ (or a strictly monotone function of $\eta$, denoted as $V$ in their paper) from the second equation, and then estimate the model primitives in the first equation by first conditioning on the estimated $\eta$, and then integrate it out. Their method cannot directly apply when the instrument variable $W^*$ in the second equation is mismeasured, because the control variable (or control function) cannot be observed from error-laden measurements of $W^*$. A recent paper by aradillas2022inference studies inference in models where control functions are unobserved. The setup of his paper is different from this paper in many aspects. He requires the availability of observable or estimable bounds for the unobserved control functions, which is not required by the model in this paper. Instead, this paper requires the availability of error-laden measurements and builds upon the measurement error literature. Also, his focus is on constructing confidence sets for finite-dimensional parameters, while my focus is on point identification and estimation of infinite-dimensional parameters.\\ In this paper, I propose a method that makes use of the same intuition as in imbens2009identification, but can deal with mismeasured $W^*$. Same as imbens2009identification, I utilize the fact that the correlation between $X$ and $\epsilon$ is merely coming from $\eta$, but instead of estimating and conditioning on $\eta$, I separate out this correlation by writing $\epsilon$ as a function of $\eta$ and a uniformly distributed random variable which is independent with $X$ and $W^*$. In this way, I am able to write the model primitives like the structural derivative and LAR as functionals of the joint distribution of $Y,X,W^*$, which can be recovered using the two error-laden measurements and the deconvolution technique developed in the measurement error literature (e.g. fan1991asymptotic, fan1993nonparametric, schennach2004estimation, schennach2004nonparametric). I can thus identify the model primitives in a constructive way and estimate them using plug-in estimators. I also derive uniform rates of convergence of the estimators.\\ This paper is most related to schennach2012local (SWC, hereafter), where the authors also consider a triangular simultaneous equations model with a mismeasured exogenous instrument. This paper differs from their paper in the modeling assumptions, parameters that can be identified, and also theoretical methods. SWC shows that under separability assumption on the second equation, the instrument conditioned marginal response\footnote{$m_x$ denotes the partial derivative of the $m$ function with respect to its first argument.} $E\left[m_x(X, \epsilon)|W^*=w^*\right]$ can be written as a ratio of the derivative of the conditional mean of $Y$ given $W^*$ over the derivative of the conditional mean of $X$ given $W^*$, both of which can be recovered from the data given two error-laden measurements $W_1, W_2$. Integrating out $W^*$, they can also recover the average response $E\left[m_x(X, \epsilon)\right]$. However, under the nonseparability of both of the equations, their method cannot recover either of the two parameters. Different from SWC, this paper shows that it's possible to identify not only the instrument-conditioned marginal response and the average marginal response but also the structural derivative and LAR, even when both of the equations are nonseparable. This conclusion, however, comes at the expense of more assumptions on the function $m$ and unobservables compared with SWC. In particular, this paper assumes strict monotonicity of the function $m$ on its second argument, and scalar unobservables, which are not required in SWC. Regarding estimation, both SWC and this paper employ plug-in estimators, but the asymptotic analysis in this paper is a bit more complex than in SWC, because of the observed variable components $Y,X$ of the joint density $f_{Y,X,W^*}$. They have non-trivial implications on the asymptotic treatment (including the convergence rates) so that SWC's asymptotic analysis cannot directly apply here. \\ This paper is also closely related to song2015estimating. In their paper, the coauthors study a nonseparable model with mismeasured endogenous variable, assuming a correctly measured control variable is available. The setup of this paper is different from theirs. This paper also studies nonseparable models with endogeneity, but instead of mismeasured endogenous variable, this paper studies the case when the endogenous variable is correctly measured, and instead of assuming the existence of a correctly measured control variable, this paper assumes that a potentially mismeasured instrument variable is available. The two papers are complementary depending on the availability of data. Regarding the parameters of interest, in addition to the various parameters studied in song2015estimating, including the average marginal response conditional on the control variable, the average marginal response, the LAR, and the weighted average version of these variables, this paper additionally show identification of the structural derivative, which is more disaggregate and could be helpful when researchers are interested heterogeneous effects.\\ The rest of the paper consists of six parts. Section 2 introduces the model and assumptions. Section 3 talks about the identification of the model primitives. Section 4 proposes plug-in estimators of the model primitives identified in Section 3, and talks about their asymptotic properties. Section 5 conducts limited Monte Carlo studies to show the finite sample performance of the estimator. Section 6 concludes.
I consider the following triangular model in this paper:
where $X$ is an observed endogenous variable that is correlated with $\epsilon$. In the returns to education example, $X$ stands for years of education, which is correlated with $\epsilon$ since it's chosen by the agent as an equilibrium outcome. Z is a vector of observed exogenous variables which are independent with $\epsilon$ and $\eta$. $W^*$ is an instrument for $X$ and satisfies $W^*\perp (\epsilon, \eta)$. Researchers cannot measure $W^*$ exactly but have two error-laden measurements of it:
The measurement errors satisfy $E[\Delta W_1\mid W^*, \Delta W_2]=0$ and $\Delta W_2\perp W^*, Y, X, Z$. As in the literature studying the identification of nonseparable models (chesher2003identification, matzkin2003nonparametric,altonji2005cross,matzkin2015estimation), I impose monotonicity assumptions on the structural functions. I assume that $m$ is strictly increasing in $\epsilon$ and that $h$ is strictly increasing in $\eta$. Since the vector of covariates $Z$ is correctly measured and exogenous, all the analysis can be done conditional on $Z$. For brevity, from now on I omit the vector $Z$ in my notations and work on the simplified model below, while other assumptions on the structural functions and distributions of variables remain unchanged.
The assumptions mentioned above are stated formally below:
Following the measurement error literature, I also impose:
Note that for the measurement error $\Delta W_1$, only the mean independence assumption is imposed, which is weaker than the assumption on $\Delta W_2$. This weaker assumption is sufficient to identify the distribution of $W^*$ (schennach2004nonparametric). On the other hand, although $\Delta W_2$ satisfies strong independence assumptions, making $W_2$ a measurement with classical measurement error, one still cannot use $W_2$ as the instrument because of the nonseparability of the model. More specifically, plugging $W_2$ into the second equation in the triangular system yields
This is a nonseparable model with two unobservables $\Delta W_2$ and $\eta$, and $W_2$ is not independent with $\Delta W_2$. This means both of the two equations in the system contain endogenous variables and thus cannot be identified without further assumptions.\\ In addition, I also impose the following assumptions on the distribution of $Y,X,W^*$:
Assumption (ref) is a sufficient condition to ensure $\left.\frac{\partial h(w^*,\eta)}{\partial w^*}\right|_{\eta = r(x,w^*)}\neq 0$ for all $x,w^*$ on the support of $X,W^*$. This assumption is like the rank condition imposed in the linear instrument variable models. To illustrate this idea, suppose $h$ is a linear function such that $X= \pi W^*+\eta$, then $F_{X\mid W^*=w^*}(x) = F_{\eta}(x-\pi w^*)$, and $\pi\neq 0$ is a neccessary condition for Assumption (ref).
In this section, I show the identification of the model parameters in two steps. First, assuming that the joint distribution of $Y, X, W^*$ is known, I show identification of the derivative of the structural function $m$ with respect to $x$, when the value of $\epsilon$ is fixed, and the identification of other model parameters including the LAR and AR. Then I show how to identify the joint distribution of $Y, X, W^*$ from the two measurements $W_1$ and $W_2$.
The independence between $W^*$ and $\epsilon,\eta$ implies that conditional on $\eta$, $W^*$ and $\epsilon$ are independent. The next lemma uses this fact to show that one can write $\epsilon$ as a function of $\eta$ and another random variable which is independent with $X, W^*$. This lemma is similar to Proposition 5.1 in matzkin2016independence. Before stating the lemma, I first impose the following assumption on the conditional distribution of $\epsilon$ given $\eta$:
Next, I plug ((ref)) into the first structural equation. The assumption that $h$ is strictly increasing in $\eta$ implies that one can write the inverse of $h(W^*,\eta)$ w.r.t. $\eta$ as $r(X, W^*)$. Then I have
Equation ((ref)) builds a bridge between the structural function $m$ and the reduced form function $v$. Note that $v$ is a function of observable variables $X,W^*$ and an unobservable variable $\delta$ which is independent of the observables. The derivatives of $v$ can be identified by applying identification techniques in standard nonseparable models (matzkin2003nonparametric). To identify the main parameter of interest, the derivative of the structural function $m$, one can utilize the last equality in (ref): $m(X,s(r(X,W^*),\delta))\equiv v(X,W^*,\delta)$. Taking derivative w.r.t $X$ and $W^*$ yields:
Plug ((ref)) into ((ref)) to cancel $\left.\frac{\partial m(x,\epsilon)}{\partial \epsilon}\right|_{\epsilon =s(r(x, w^*),\delta)}\cdot\left.\frac{\partial s(\eta, \delta)}{\partial \eta}\right|_{\eta = r(x,w^*)}$ yields
Taking derivative w.r.t. $W^*$ on both sides of $X\equiv h(W^*,r(X,W^*))$ and cancelling out the unobserved $\left.\frac{\partial h(w^*,\eta)}{\partial \eta}\right|_{\eta=r(x,w^*)}$, one can get $\left.\frac{\partial h(w^*,\eta)}{\partial w^*}\right|_{\eta=r(x,w^*)} = -\frac{\frac{\partial r(x,w^*)}{\partial w^*}}{\frac{\partial r(x,w^*)}{\partial x}}$. Then one can write
The right-hand side of the equation ((ref)) can be identified from the data. To show this, note that
Similarly
where $\delta$ is the unique value such that $y=v(x,w^*,\delta)$ and $\eta$ is the unique value such that $ x=h(w^*,\eta) $. Plug into ((ref)), one can get
for all $\bar y, \bar x$ belongs to support, and values of $\bar\epsilon$ such that $\bar y = m(\bar x, \bar\epsilon)$. Note that as long as $\bar y$ and $\bar x$ are fixed, $\bar\epsilon$ is fixed, no matter which value is picked for $w^*$. This means there's actually overidentification for the structural derivative. To identify the LAR and AR, one needs independent variations of $\epsilon$ given $X$. Rewrite ((ref)) with slightly different notations yields:
so that one can write
To show the identification of the LAR and AR, one needs to show that the support of $s(r(X, W^*), \delta)$ given $X=x$ is the same as the support of $\epsilon$ given $X=x$. This is stated formally in the lemma below.
Lemma (ref) and (ref) ensures that $F_{\mathbf{s}\mid X}$ is the same as $F_{\epsilon\mid X}$, so that integrating $\frac{\partial m(x,\epsilon)}{\partial x}$ over $\epsilon$ given $X=x$ will be equivalent to integrating $\frac{\partial m(x,\mathbf{s})}{\partial x}$ over $\mathbf{s}$ given $X=x$. Then I have the following identification results:
First note that by Theorem 1 in schennach2004estimation
Then I can identify the density $f_{W^*}(w^*)$ by taking the Inverse Fourier Transform:
For the joint density $f_{Y,X,W^*}(y, x, w^*)$, by Assumption (ref), I have the following convolution:
Applying Fourier transformation on both sides (holding $x$ constant) yields
which, by Assumption (ref) implies
Then applying the inverse Fourier transform, one can get
and similarly,
Then $F_{X\mid W^*=w^*}(x)$ follows from
and $F_{Y\mid X=x, W^*=w^*}(y)$ follows from
From the equations above, one can write the derivatives of the conditional CDFs as functionals of the joint densities:
This means that the right-hand side of the equation ((ref)) can be written as a functional of the joint density $f_{Y,X, W^*}$, which can be identified from the data. To show identification of the left-hand side of the equation ((ref)), one needs to build the relationship between the derivatives of conditional quantiles and the joint densities, which is shown in the following lemma:
The identification of the left-hand side of ((ref)) and ((ref)) is a straightforward implication of Lemma (ref) and equation ((ref)).
From now on, I use the following function to denote the joint density of $Y, X,W^*$, or its derivative with respect to $Y, X,W^*$:
where $\lambda_1, \lambda_{2,1}, \lambda_{2,2}\in\{0,1\}$ and $\lambda\equiv\max\{\lambda_1, \lambda_{2,1},\lambda_{2,2}\}\leq 1$. The $0-th$ order derivative of a function is defined as the function itself. For convenience of notation, I also define $\lambda_2\equiv\max\{ \lambda_{2,1},\lambda_{2,2}\}$. In this section, I first propose an estimator for $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$, and then propose plug-in estimators for the structural derivative and a weighted average version of the LAR since they can be written as known functionals of $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$. \\ To make the expressions more transparent and clear, I show the explicit forms of the functionals by taking the structural derivative as an example. I write each component on the right-hand-side of equation (ref) as
For the LAR, I will introduce a weighted average version of it, and propose an estimator for this weighted average version of the LAR. The aim of introducing the weight is to ensure that the integration is taken on a set where the joint density $f_{Y,X,W^*}$ is bounded away from zero. This has the benefit of guaranteeing that the relevant functional is Fr\'echet differentiable when $f_{Y,X,W^*}$ appears in the denominator of it. The weighted average version of the LAR is defined as follows:
where $\omega(x,w^*,\epsilon)$ is a known or estimable weighting function taking value 0 outside a compact set $\mathbb M$. As a weighting function, $\omega(x,w^*,\epsilon)$ also satisfies that given any values of $x$ such that $\omega(x,\cdot,\cdot)$ could obtain non-zero values, $\int_{w^*}\int_\epsilon\omega(x,w^*,\epsilon)f_{\epsilon, W^*|X=x}(\epsilon, w^*) d\epsilon dw^*=1$. The specific form of $\omega$ is up to the researcher's choice, however, to ensure easy calculation of $\omega(x,w^*,s(r(x,w^*),\delta))$ on the right-hand side of ((ref)), I require that function $\omega(x,w^*,\epsilon)$ satisfy the following restriction: $\omega(x,w^*,\epsilon)=\tilde \omega(x,w^*)$ if $\epsilon \in \left[q_{x,w^*}(\tau_l), q_{x,w^*}(\tau_u)\right]$ and $\omega(x,w^*,\epsilon)=0$ otherwise, where $\tilde\omega(x,w^*)$ is a known function with compact support, and $q_{x,w^*}(\tau)$ denotes the conditional-$\tau$ quantile of $\epsilon$ given $X=x, W^*=w^*$. The specific form of function $\tilde w$ and the specific values of $\tau_l$ and $\tau_u$ are up to the researcher's choice. Under this restriction, $\omega(x,w^*,s(r(x,w^*),\delta))$ is equal to $\tilde w(x,w^*)$ if $\delta$ is between $\tau_l$ and $\tau_u$, and is equal to 0 otherwise. An example of function $\omega(x,w^*,\epsilon)$ is when $\tilde \omega(x,w^*)$ is a constant equal to $\left[(\tau_u-\tau_l)\int_{\underline{w}^*}^{\bar w^*} f_{W^*|X=x}(w^*)dw^*\right]^{-1}$. The WLAR in this case can be interpreted as the Local Average Response of a subgroup of individuals whose unobserved $\epsilon$ is between its $\tau_l$ and $\tau_u$ conditional quantile given $X, W^*$ and whose $W^*$ is between $\underline w^*$ and $\bar w^*$. Note that even though $X$ is correlated with $\epsilon$, neither the distribution of $\epsilon|X$ nor the subgroup of individuals changes as I take the derivative with respect to $x$. This means that my objective of interest is the same group of people when studying the effect of the counterfactual changes of the endogenous $X$. The WLAR can also be written as a functional of $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$. The explicit form can be derived from Lemma (ref) and equation ((ref)).\\ So far I have written the parameters of interest as functionals of $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$ in explicit forms. Next I focus on the estimation of $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$. To deal with the well-known ill-posed inverse problem when inverting a convolution operator, I base my estimator on a smoothed version of $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$:
and use kernel functions satisfying the following assumption:
By assuming (ii), I use flat-top kernels proposed by politis1999multivariate. The benefit of using flat-top kernels is that the rate of decrease of the bias term will not be affected by the order of the kernel, and will only be affected by the smoothness of the function to be estimated. The restriction of compact support of the Fourier transform is without loss of generality. As discussed by schennach2004nonparametric, given any kernel $K$, one can always create a modified kernel $\tilde K$ that satisfies the assumption by using a “windowing” function. For example,
where
for some $\bar t\in(0,1)$.
The following lemma shows that the smoothed version $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_1)$ can be written as an expression involving the Fourier transform of the kernel function and the characteristic functions of different variables. This expression makes it more straightforward to introduce the estimator that will appear soon below.
I now define my estimator for $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_1)$ by replacing $\phi_{f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W_2}(y,x,\cdot)}(t)$, $\phi_{W^*}(t)$, and $\phi_{W_2}(t)$ in Lemma (ref) with their sample analogues. Formally,
Having defined $\hat g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y, x ,w^*,h_n)$, I can now define the estimator for $\rho(\bar y, \bar x)$ by replacing $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$ with its estimator on the right-hand-side of Equation ((ref)). Note that by Fubini's Theorem
where I've let $\tilde G_Y(y)\equiv\int_{-\infty}^y G_Y\left(u\right)du$ and $\tilde G_X(x)\equiv\int_{-\infty}^x G_X\left(u\right)du$ denote the kernel CDFs, and I've defined:
In this subsection, I analyze the asymptotic properties for $\hat g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_n)$. I begin by decomposing the difference between the estimator and the true value of the parameter as the sum of a bias term, a variance term, and a remainder term.
To derive the uniform rate of convergence, I impose some assumptions on the smoothness of $g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*)$, stated in terms of the rate of decay of the tail of its Fourier transform:
In Assumption (ref)(i) I impose the same bound for the tail behavior of $\phi_{f^{(\lambda_{2,1},\lambda_{2,2})}_{Y,X,W^*}(y,x,\cdot)}$ and $\phi_{W^*}(t)$. This is without loss of generality since they have the same effect on the convergence rate. Assumption (ref)(ii) is similar to Assumption A7 in li2002robust, and can be viewed as a generalization of smoothness condition from the univariate case to the multivariate case.\\ I next state the first main result of this paper, the uniform asymptotic rate of the bias term:
I impose the following assumption to ensure finite variance:
To derive the rate for $L_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h)$, I impose the following assumption on the tail behavior of the Fourier transforms involved. These are common in the deconvolution literature (e.g. fan1991optimal, fan1993nonparametric, and schennach2004nonparametric).
I explicitly impose $\beta_2\geq\beta_{\phi}$ because
The following theorem states the asymptotic properties of the linear term \\ $L_{\lambda_1, \lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h)$ and facilitates the analysis of various quantities of interest later.
Finally, I bound the remainder term. To do that, I first impose some restrictions on the moments of $W_2$:
I then impose the following bounds for the bandwidths:
Collecting the results from Theorem (ref),(ref) and (ref) yields a straightforward corollary of the uniform rate of convergence of $\hat g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*,h_n)$:
The uniform (over the whole support) rate of convergence of the kernel estimators of the density or its derivatives has been considered in andrews1995nonparametric, hansen2008uniform and schennach2012local\footnote{schennach2012local studies the case both with and without measurement errors. Their Theorem 3.2 corresponds to the no-measurement error case, and Corollary 4.7 corresponds to the case with measurement errors.}. hansen2008uniform obtained a faster rate than andrews1995nonparametric and schennach2012local (Theorem 3.2, for the no-measurement error case), but requires more assumptions on the kernel functions (their Assumption 1 and 3), including bounded support or an integrable tail, which, unfortunately, are not satisfied by infinite order kernels. andrews1995nonparametric's conclusion is stated assuming a kernel with finite order. Infinite order kernels are not essential when there is no measurement error (like in andrews1995nonparametric and hansen2008uniform), but are especially advantageous when there are measurement errors. If infinite order kernels are used, only the smoothness of the functions, but not the order of the kernels will affect the rate of convergence. Corollary (ref) is similar to Corollary 4.7 in schennach2012local.\\ With the uniform rate of convergence of $\hat g$, I can then derive the uniform rate of convergence of the plug-in estimators of the structural derivative $\rho(y,x)$ and the WLAR, denoted as $\widehat{\rho(y,x)}$ and $\widehat{\mathbf{WLAR}(x)}$, respectively.
In Appendix C, I state an asymptotic normality result for $\hat g_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}} (y,x,w^*,h_n)$ and $\widehat{\rho(y,x)}$. The proof of them requires a lower bound on $\Omega_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*.h_n)$ relative to $B_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*.h_n)$ and $R_{\lambda_1,\lambda_{2,1}, \lambda_{2,2}}(y,x,w^*.h_n)$. The assumptions to ensure the lower bound are stated at a high level, and more primitive sufficient conditions need to be derived.
In this section, I conduct some Monte Carlo simulation exercises to study the finite sample performance of my proposed estimators. I consider two simulation designs. Design 1 is when both of the equations are linear, and Design 2 is when the equations are nonlinear.\\ Design 1
Design 2
In both cases, $\epsilon$ and $\eta$ are correlated through a common component $\theta$: $\epsilon = \theta + \epsilon_1$, and $\eta = \theta + \eta_1$, where $\theta$, $\epsilon_1$ and $\eta_1$ are mutually independent, and are independent with $W^*$. In both designs, I have two error-laden measurements $W_1 = W^*+\Delta W_1, W_2 = W^*+\Delta W_2$, where $\Delta W_1$ and $\Delta W_2$ are independent and they are independent with all other variables. The distributions of variables in both designs are listed in Table (ref) below:
In the estimation, I use the following flat-top kernel\footnote{I may use different flat-top kernels for $G_X$ and $G_Y$. For simplicity I use the same flat-top kernel.} proposed by politis1999multivariate:
I use the sample size of 500 and replicate the estimation 500 times. \\ First I show the performance of my proposed method versus two other methods for estimating the structural derivative: (1) 2SLS using the error-laden measurement $W_2$ as IV; (2) a plug-in estimator replacing $W^*$ with $W_2$ in my identification equation ((ref)). Method (1) will be valid under the linear design (Design 1), since $W_2$, although error-laden, still satisfies the exclusion restriction and the rank condition for linear IV estimation. However, because of its misspecification of the model, it won't be valid under the nonlinear design (Design 2). Method (2) won't be valid under either the linear or nonlinear design, as the estimator will converge to a population value that is not equal to the right-hand side of equation ((ref)) in general. I estimate $\rho(y,x)$ evaluated at $y=0, x=0, w^*=0.7$ for Design 1, and $y=0.6, x=0.6, w^*=7$ for Design 2, yielding a true value of $0.25$ and $0.4512$, respectively. There are three bandwidths used in the estimation: $h_1$, $h_{21}$, $h_{22}$. The latter two correspond to estimating the unknown distribution of observed variables $X$ and $Y$. I thus choose the values of $h_{21}$ and $h_{22}$ by cross-validation (i.e. minimizing the estimated MISE of $\hat f_{Y,X}(y,x)$). For the bandwidth $h_1$, I scan a set of values ranging from 0.5 to 3 for my proposed estimator and a set of values ranging from 0.5 to 6 for the method that uses the error-laden measurement. Table (ref), (ref), and (ref) show the MSE, VAR, and the absolute value of BIAS of my proposed estimator, Method (2), and Method (1), respectively. \\ From Table (ref), one can see that the bias from the error-contaminated estimator does not shrink toward zero as bandwidth decreases. Comparing results from Table (ref) and Table (ref), one can see that my estimator also gives smaller variances than the error-contaminated method. As a result, my estimator performs better than the error-contaminated method in terms of MSE under both the linear and the nonlinear designs. Comparing Table (ref) with Table (ref), one can see that under the linear design, both my estimator and 2SLS work well except that my estimator has a slightly larger variance. Under the nonlinear design, the bias and MSE of 2SLS are much larger than my estimator, which aligns with the theory.\\ Table (ref) below shows the Monte Carlo simulation results of my proposed estimator as a function of the sample size. The value of $h_{21}$ and $h_{22}$ are selected by cross-validation for the respective sample sizes and the optimal values of $h_1$ are selected by minimizing the MSE in the corresponding set of values the same as when $N=500$. The MSE, VAR, and Bias decrease as the sample size increases, which is in accordance with the theory.
Next, I study the performance of my estimator of the weighted LAR. For Design 1 (linear), I evaluate the WLAR at $x=0$ and set the weighting function to be $\omega(0,w^*,\epsilon)=\omega_1$ if $\epsilon\in\left[q_{0,w^*}(0.25),q_{0,w^*}(0.35)\right]$ and $w^*\in\left[0.70, 0.90\right]$. The constant $\omega_1$ is selected such that $\int_{0.7}^{0.9}\int_{q_{0,w^*}(0.25)}^{q_{0,w^*}(0.35)}\omega_1f_{\epsilon,W^*|X=0}(\epsilon,w^*)d\epsilon d w^*=1$. For Design 2 (nonlinear), I evaluate the WLAR at $x=0.6$ and set the weighting function to be $\omega(0.6,w^*,\epsilon)=\omega_2$ if $\epsilon\in\left[q_{0.6,w^*}(0.25),q_{0.6,w^*}(0.35)\right]$ and $w^*\in\left[6, 6.23\right]$. The constant $\omega_2$ is selected to ensure that $\int_{6}^{6.23}\int_{q_{0.6,w^*}(0.25)}^{q_{0.6,w^*}(0.35)}\omega_2f_{\epsilon,W^*|X=0.6}(\epsilon,w^*)d\epsilon d w^*=1$. These choices of the weighting functions yield a true value of the WLAR of 0.25 and 0.5032 for Design 1 and 2, respectively. In the estimation, I use the optimal values of $h_1$ coming from Table (ref), and the same values of $h_{21}$ and $h_{22}$ as in Table (ref). Table (ref) below shows the results of my estimators of the WLAR under both designs. The performance of my estimator of the WLAR is comparable to the performance of my estimator for the point-wise structural derivative $\rho(y,x)$ (shown in Table (ref)). This shows evidence that my estimator could not only perform well at single points but also could maintain good performance in a global sense. The supports of the weighting functions I used here are not large. This is mainly due to computation constraints. One direction of future work is to look at the performance of my estimator for WLAR when the support of the weighting function is larger. This will give us a better understanding of the global performance of my estimator.
The Monte Carlo simulation exercise shown in this section is very limited. To have a comprehensive examination of the performance of my estimators, there are several future directions that I can work on. First, in my current exercise, the distribution of the measurement error $\Delta W_2$ is set to be $\chi^2$, which is ordinarily smooth. I can further explore the performance of my estimators under different distributions of the measurement error $\Delta W_2$ and other variables. Second, when estimating the WLAR, I set the support for the weighting function to be small because of the computation constraint. With more time and more computation power, I can look at the performance of my estimator when the support of the weighting function in the WLAR is larger, i.e. the WLAR is taking the average marginal effect over a larger population. Third, the way I select $h_1$ is by minimizing the actual MSE, i.e., the mean squared difference between my estimator and the true value of the parameter. This is infeasible in practice when one deals with real-world data. It will be helpful if I could propose a feasible way to select the bandwidth.
I study the nonparametric identification and estimation of the nonseparable triangular equations model when the instrument variable $W^*$ is mismeasured. I don't assume linearity or separability of the functions governing the relationship between observables and unobservables. To deal with the challenges caused by the co-existence of the measurement error and nonseparability, I first employ the deconvolution technique developed in the measurement error literature (schennach2004estimation,schennach2004nonparametric) to identify the joint distribution of $Y, X, W^*$ using two error-laden measurements of $W^*$. I then recover the structural derivative of the function of interest and the “Local Average Response” (LAR) via the “unobserved instrument” approach in matzkin2016independence. Based on the constructive identification results, I propose plug-in nonparametric estimators for these parameters and derive their uniform rates of convergence. I also conducted limited Monte Carlo exercises to show the finite sample performance of my estimators. \\ I recognize that there are some important future directions for this paper. First, in the main text, I only demonstrated the uniform rate of convergence of the estimators. The appendix contains proofs of their asymptotic normality, which rely on high-level assumptions. To further enhance my results, it would be beneficial to derive more primitive sufficient conditions for these assumptions. Second, as discussed in Section (ref), to conduct a comprehensive examination of the performance of my estimators, additional Monte Carlo studies are necessary. It will also be extremely useful if I could propose a feasible way to select the bandwidth $h_1$. Last, it will be an interesting exercise to apply my proposed method to real-world data.