Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
37,680 characters · 3 sections · 10 citation commands
Nonparametric Identification and Estimation with Non-Classical Errors-in-Variables
Regression is a fundamental tool for empirical analysis. Errors-in-Variables (EIV) are a widespread problem in empirical applications. Mismeasurement of a covariate, when not accounted for, may lead to biased estimates and invalid inferences.
The goal of this paper is to study the nonparametric identification and estimation of the regression function when a covariate is mismeasured. Importantly, the measurement error need not be classical and can be correlated with the mismeasured covariate. In this paper, we adopt the small measurement error approximation, which allows us to provide a simple nonparametric characterization of the problem. Then we provide transparent and constructive identification analysis under weak and easy-to-interpret conditions on the instrumental variable.
First, we focus on the Weakly Classical Measurement Error (WCME) model, where the measurement error is uncorrelated with the true covariate but generally is not independent from it. We show that the skedastic function of the measurement errors, together with its derivative, plays a key role in determining the bias of the naive regression estimator. We also show how the EIV skedastic function can be recovered from the distribution of the observables using a (possibly discrete) instrument, and how one can construct a bias-corrected estimator of the regression function. We derive its rate of convergence and provide conditions under which the approximation error becomes negligible compared to the errors arising from the nonparametric estimation of unknown functions in large samples.
Next, we consider the general Non-Classical Measurement Error (NCME) model, which allows for a very broad form of EIV. In particular, the measurement error can be correlated with the true covariate. Even though the NCME model is much more general than the WCME model, we demonstrate how the results from our analysis of the WCME model can be utilized to establish identification of the general NCME model.
Importantly, our approach only requires an instrumental variable that can be discrete. This allows for a broader range of applications compared to the methods that require a continuously distributed instrument or the availability of repeated (multiple) measurements of the true covariate. In Section (ref), we discuss in detail the exclusion and relevance conditions that the instrumental variable needs to satisfy, and consider some examples.
Our paper contributes to the large literature studying models with mismeasured data. CarrollEtAl2006Book-ME,ChenHongNekipelov2011JEL; and Schennach2020HB-ME,Schennach2022JEcLit provide excellent literature overviews.
Our main focus is on the settings where the distribution of measurement error is unknown. Nonparametric analysis of the EIV problem in such settings requires additional information to separate the true covariate from the measurement error. In Economics, most commonly instrumental variables are used for this purpose ( HINP1991JoE,HausmanNeweyPowell1995JoE,Newey2001REStat,Schennach2007Ecta,Hu2008JoE,HuSchennach2008Ecta,Wilhelm2019WP-TestingForME , among others). Repeated measurements can also be utilized ( HINP1991JoE,LiVuong1998JoE,Schennach2004Ecta, among others) but are less frequently available. Note that repeated measurements can serve as valid instruments in our analysis.
Small measurement error (SME) approximation has been widely employed in Statistics and Econometrics to study the effect of EIV on various estimators and to bias-correct them (e.g., WolterFuller1982AS,CarrollStefanski1990JASA,Chesher1991Biomet,Chesher2000WP,CarrollEtAl2006Book-ME,ChesherSchluter2002ReStud , among others). BoundBrownMathiowetz2001HBoE document that the EIV in economic applications are typically relatively small although are often non-classical, which suggests that our analysis should be useful in many applied settings. This paper differs from the previous literature in two ways. First, it appears to be the first paper to study the nonparametric SME approximation with non-classical EIV. Second, previously developed SME bias reduction techniques usually assume that the EIV variance is either known or can be directly estimated from an available dataset, e.g., using repeated measurements. In contrast, this paper demonstrates how the whole EIV skedastic function can be identified and estimated using only a (possibly discrete) instrumental variable.
The analysis of this paper complements the existing \textquotedblleft large\textquotedblright\ measurement error literature. By focusing on a narrower range of settings, the paper provides simpler characterizations of the problem and estimators, which are valid under very weak and easy-to-interpret conditions on the instrumental variable. In particular, our identification results do not rely on the completeness conditions, and estimation does not involve solving ill-posed inverse problems or deconvolution.
The rest of this paper is organized as follows. Section (ref) studies the Weakly Classical Measurement Error (WCME) model. Section (ref) considers the general Non-Classical Measurement Error (NCME) model. The proofs are collected in the Appendix.
We consider the regression model
where $Y_{i}\in \mathbb{R}$ is the outcome variable, and $X_{i}^{\ast }\in \mathbb{R}$ is the true value of the covariate for individual $i$. The researcher observed a mismeasured version of $X_{i}^{\ast }$:
where ${\Greekmath 0122} _{i}$ is the measurement error. The researcher has a random sample of $\left( Y_{i},X_{i},Z_{i}\right) $, where $Z_{i}$ are instrumental variables that are used to identify the model and will be discussed later. It is straightforward to also include correctly measured covariates into the model, see Remark (ref) for details.
In this section we consider the Weakly Classical Measurement Error (WCME)\ model:
The measurement error ${\Greekmath 0122} _{i}$ is uncorrelated with the true covariate $X_{i}^{\ast }$. Assumption (ref) is significantly weaker than the (Strongly) Classical Measurement Error (CME) assumption, since $ {\Greekmath 0122} _{i}$ need not be independent from $X_{i}^{\ast }$. For example, the measurement error can be conditionally heteroskedastic, i.e., its conditional variance
may depend on $x$. Function $v\left( x^{\ast }\right) $ is usually unknown.
In this paper we use the Small Measurement Error (SME)\ approximation (e.g., WolterFuller1982AS) for the analysis, i.e., we will consider the approximations of the model when $v(x^*)$ and the higher conditional moments of ${\Greekmath 0122}_i$ are small.
Specifically, we model the measurement error as ${\Greekmath 0122} _{i}={\Greekmath 011C} {\Greekmath 0118} _{i}$ where the distribution of ${\Greekmath 0118} _{i}$ is fixed, and ${\Greekmath 011C} $ is a non-stochastic parameter. Assumption (ref) requires $E[{\Greekmath 0118} _{i}|X_{i}^{\ast }]=0$. The conditional variance of ${\Greekmath 0122} _{i}$ is given by $v(x)={\Greekmath 011C} ^{2}V[{\Greekmath 0118} _{i}|X_{i}^{\ast }=x]=O({\Greekmath 011C} ^{2})$.
We study the properties of the model when ${\Greekmath 011C} \rightarrow 0$. Under some smoothness conditions,
Thus, a naive regression estimator of ${\Greekmath 011A} $ that ignores the presence of the measurement errors in $X_{i}$ has a bias of order $O\left( {\Greekmath 011C} ^{2}\right) $, e.g., see Chesher1991Biomet.
The goal of the small measurement error analysis is to provide a function $ \widetilde{{\Greekmath 011A} }\left( x\right) $ that has a smaller bias, i.e., satisfies
for some $p\geq 3$.
To identify the model we will rely on an observed instrumental variable (instrument) $Z_{i}$ that satisfies the following exogeneity assumption.
\setcounter{assumption}{0}
This assumption states that $Z_{i}$ is an \textquotedblleft excluded\textquotedblright\ variable: given $X_{i}^{\ast }$, instrument $ Z_{i}$ has no effect on the conditional mean of $Y_{i}$. Without loss of generality, we can assume that $Z_{i}$ is discrete. (The instrument also needs to satisfy a \textquotedblleft relevance\textquotedblright condition: it needs to affect the conditional distribution $f_{X^{\ast }|Z}\left( x|z\right) $. This condition will appear in Theorem (ref) .)
Assumption (ref) says that the measurement error $ {\Greekmath 0122} _{i}$ is nondifferential: conditional on $\left( X_{i}^{\ast },Z_{i}\right) $, $X_{i}$ provides no additional information about (the conditional mean of) $Y_{i}$. This assumption can be equivalently stated as $ E\left[ Y_{i}|X_{i}^{\ast },Z_{i},{\Greekmath 0122} _{i}\right] =E\left[ Y_{i}|X_{i}^{\ast },Z_{i}\right] =E\left[ Y_{i}|X_{i}^{\ast }\right] $.
Assumption (ref) and the smoothness conditions below are stated using the auxiliary variable ${\Greekmath 0118} _{i}$, whose variance does not shrink. This allows formulating the smoothness conditions in the conventional form. For example, we will assume that the density $f_{{\Greekmath 0118} |X^{\ast }}\left( u|x\right) $ is bounded. In contrast, the density of ${\Greekmath 0122} _{i}={\Greekmath 011C} {\Greekmath 0118} _{i}$ is $f_{{\Greekmath 0122} |X^{\ast }}\left( e|x\right) =\frac{1}{{\Greekmath 011C} } f_{{\Greekmath 0118} |X^{\ast }}\left( \left. \frac{e}{{\Greekmath 011C} }\right\vert x\right) $ and is not bounded as ${\Greekmath 011C} \rightarrow 0$. Assumption (ref) also implies that ${\Greekmath 0122} _{i}\perp Z_{i}|X_{i}^{\ast }$.
Finally, the following two assumptions are smoothness conditions.
Assumption (ref) is a weak restriction imposed on the conditional moments of ${\Greekmath 0118} _{i}$. Appendix (ref) provides a set of primitive conditions that guarantee that Assumption (ref) holds. Also notice that Assumption (ref) would automatically hold if the support of ${\Greekmath 0118} _{i}$ is bounded, since $f_{{\Greekmath 0118} |X^{\ast }}(u |x )$ and its derivatives are uniformly bounded under Assumption (ref).
We can now state the first main result of the paper. Let
Let $\mathcal{S}_{X^{\ast }}(z)$ denote the conditional support of $ X_{i}^{\ast }|Z_{i}=z$. Consider any two values $z_{1}$ and $z_{2}$ the instrument can take.
Theorem (ref) demonstrates that $\widetilde {\Greekmath 011A} (x,z_1)$ identifies ${\Greekmath 011A}(x)$ up to an error of order $O({\Greekmath 011C}^p)$ when ${\Greekmath 011C} \rightarrow 0$. This is a substantial improvement over naive regression $ q(x) $ which has a bias of order $O({\Greekmath 011C}^2)$. The improvement in the magnitude of the approximation error (from $O({\Greekmath 011C}^2)$ to $O({\Greekmath 011C}^4)$) is especially noticeable when $E[{\Greekmath 0118}_i^3|X_i^*] = 0$, e.g., when the measurement error is symmetric.
To establish the desired result, we first characterize the bias of $q(x,z)$ up to an error of order $O({\Greekmath 011C}^p)$. The bias of $q(x,z)$ is of order $ O({\Greekmath 011C}^2)$ and determined by the conditional variance of the measurement error $v(x)$ and its derivative $v^{\prime }(x)$, which are unknown. Then, we show that $\widetilde v (x)$ identifies $v(x)$ up to an error of order $ O({\Greekmath 011C}^p)$.\footnote{ We also demonstrate that $\widetilde v^{\prime }(x) = v^{\prime }(x) + O({\Greekmath 011C}^p)$. This is an important step of the proof.} This allows us to approximate the bias of $q(x,z)$ with a sufficient precising using $ \widetilde v(x)$ in place of $v(x)$. Finally, we construct $\widetilde {\Greekmath 011A} (x,z_1)$ by bias correcting $q(x,z_1)$ and demonstrate that it approximates the true regression function ${\Greekmath 011A}(x)$ up to an error of order $O({\Greekmath 011C}^p)$.
The idea behind nonparametric identification is that although function $E \left[ Y_{i}|X_{i}^{\ast }=x,Z_{i}=z\right] $ does not depend on $z$, function $q\left( x,z\right) \equiv E\left[ Y_{i}|X_{i}=x,Z_{i}=z\right] $ does vary with $z$. The theorem shows how this variation allows recovering $ v\left( x\right) $. Specifically, the proof of the Theorem shows that
Then it is shown that replacing the derivatives of ${\Greekmath 011A} $ with those of $q$ and $s_{X^{\ast }|Z}$ with $s_{X|Z}$ on the right-hand side in the above equation does not increase the magnitude of the approximation error, i.e., that
Note that only the second term on the right-hand side depends on $z$. Since $ q$, $q^{\prime }$, and $s_{X|Z}$ are directly identified from the joint distribution of the observables, considering the differences $q\left( x,z_{1}\right) -q\left( x,z_{2}\right) $ then allows identification of $ v\left( x\right) $ by $\widetilde{v}\left( x\right) $ up to an error of order $\ O\left( {\Greekmath 011C} ^{p}\right) $. This identification approach requires the rank condition that $s_{X|Z}\left( x|z\right) $ depends on $z$, which is ensured by equation ((ref)). In addition, it is necessary that $q^{\prime }\left( x\right) \neq 0$, which is also ensured by equation ( (ref)). The latter condition is weak: ${\Greekmath 011A} ^{\prime }\left( x\right) =0$ for all $x$ only if ${\Greekmath 011A} \left( x\right) $ is a constant.\footnote{ For an analysis of the role of the conditions such as ${\Greekmath 011A} ^{\prime }\left( x\right) \neq 0$ in the measurement error literature see EvdokimovZeleneev2018WP-Inference.}
When the measurement error is classical, $v\left( x\right) $ is constant, and hence the term containing $v^{\prime }\left( x\right) $ is absent from equation ((ref)). In addition, Corollary (ref) requires only a single point $\dot{x}$ satisfying the rank condition ((ref)) for identification of ${\Greekmath 011A} \left( x\right) $ for all $x$, since $v\left( x\right) =v\left( \dot{x}\right) $ for all $x$.\footnote{ For the result of Corollary (ref) to hold, it is sufficient to require $v(x)$ to be constant, i.e., to assume that ${\Greekmath 0122}_i$ is homoskedastic, instead of requiring ${\Greekmath 0122} _{i}\perp \left( X_{i}^{\ast },Z_{i},Y_{i}\right)$.} In contrast, in the general case of Theorem (ref), nonparametric identification of $v\left( x\right) $ and ${\Greekmath 011A} \left( x\right) $ for a given $x$ requires condition ( (ref)) to hold at that point $x$. Identification of the Classical Measurement Error model has been previously established in EvdokimovZeleneev-Estimation-WP-2022.
\paragraph{Nonparametric Estimation}
Theorem (ref) suggests an analogue estimator of $ \widetilde{{\Greekmath 011A} }$ by replacing functions that appear in equations ((ref))-((ref)) with their standard nonparametric estimators (e.g., kernel or sieve), with optimally chosen tuning parameters. Let $ \widehat{{\Greekmath 011A} }\left( x\right) $ denote this estimator. As an alternative to $\widehat{{\Greekmath 011A} }\left( x\right)$, we also consider $\widehat{{\Greekmath 011A} }^{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ Naive}}\left( x\right)$, a naive nonparametric estimator of ${\Greekmath 011A}(x)$ ignoring the presence of the measurement error.
To approximate the finite sample properties of the studied estimators when $ {\Greekmath 011C}$ is small, we consider a triangular asymptotic framework with drifting $ {\Greekmath 011C} = {\Greekmath 011C}_n$ converging to zero as the sample size $n \rightarrow \infty$.
Lemma (ref) establishes the rates of convergence for the proposed and naive estimators. For each estimator, the rate of convergence is determined by two components: the standard nonparametric learning rate and the EIV (errors-in-variables) bias due to the presence of the measurement error.
The nonparametric learning rate for $\widehat {\Greekmath 011A} (x)$ is slower than for the naive estimator because it involves nonparametric estimation of derivatives such as $q^{\prime }(x)$ and $s_{X|Z}(x|z)$. However, as Theorem (ref) suggests, the EIV bias of the proposed estimator $ \widehat {\Greekmath 011A} (x)$ is of order $O({\Greekmath 011C}_n^p)$, whereas the naive estimator has a much larger bias of order $O({\Greekmath 011C}_n^2)$. Thus, despite the slower nonparametric learning rate, $\widehat {\Greekmath 011A} (x)$ has a faster rate of convergence than $\widehat{{\Greekmath 011A} }^{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Naive}}\left( x\right)$ unless $ {\Greekmath 011C}_n$ is very small, i.e., the measurement error is negligible.
To illustrate this result, suppose the conditions of Theorem (ref)(ii) hold, $m=p=4$, and ${\Greekmath 011C} _{n}={}O\left( n^{-\frac{1 }{12}}\right) $. Then $\widehat{{\Greekmath 011A} }\left( x\right) -{\Greekmath 011A} \left( x\right) =O_{p}\left( n^{-\frac{1}{3}}\right) $, but the naive estimator has a much slower rate of convergence: $\widehat{{\Greekmath 011A} }^{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Naive}}\left( x\right) -{\Greekmath 011A} \left( x\right) =O_{p}\left( n^{-\frac{1}{6}}\right) $, because of the EIV bias.
In this section, we will use notation $\mathcal{X}_{i}^{\ast }$ for the true mismeasured covariate and consider the general measurement model
where $\mathscr{m}$ is an unknown function and ${\Greekmath 0120} _{i}$ is a random vector independent from $\left( Y_{i},\mathcal{X}_{i}^{\ast },Z_{i}\right) $. Function $\mathscr{m}$ need not be monotone in any of the arguments. The measurement error is non-classical: $X_{i}- \mathcal{X}_{i}^{\ast }$ and $ \mathcal{X}_{i}^{\ast }$ are generally correlated.
As before, we want to identify and estimate the regression function
We will assume that the measurement is sufficiently informative about the true $\mathcal{X}_{i}^{\ast }$. Define
Note that since functions $\mathscr{m}$ and ${\Greekmath 011A} _{\mathcal{X}^{\ast }}$ are unrestricted, the model (ref)-(ref) cannot be identified without some normalization or additional information. Specifically, for any strictly increasing function ${\Greekmath 0115} $, we can define an observationally equivalent model with $\mathcal{\tilde{X}}_{i}^{\ast }\equiv {\Greekmath 0115} \left( \mathcal{X}_{i}^{\ast }\right) $, ${\Greekmath 011A} _{\mathcal{ \tilde{X}}^{\ast }}\left( \tilde{\varkappa}^{\ast }\right) \equiv {\Greekmath 011A} _{ \mathcal{X}^{\ast }}\left( {\Greekmath 0115} ^{-1}\left( \tilde{\varkappa}^{\ast }\right) \right) $ and $\mathscr{m}_{\mathcal{\tilde{X}}_{i}^{\ast }}\left( \tilde{ \varkappa}^{\ast },{\Greekmath 0120} \right) \equiv \mathscr{m}\left( {\Greekmath 0115} ^{-1}\left( \tilde{ \varkappa}^{\ast }\right) ,{\Greekmath 0120} \right) $.
Let us define random variable
Then,
where the first equality follows by ((ref)), the second equality follows from the strict monotonicity of ${\Greekmath 0116} (\cdot)$, the third equality is the definition of ${\Greekmath 0116} \left( \cdot \right) $, and the last equality follows from ((ref)).
Thus, for the general measurement error model we can consider an observationally equivalent model that defines $X_{i}^{\ast }$ as in equation ((ref)):
Since in this model Assumption (ref) holds, we can apply the result of Theorem (ref) to identify ${\Greekmath 011A} _{X^{\ast }}\left( x\right) $ and $v\left( x\right) \equiv E\left[ {\Greekmath 0122} _{i}^{2}|X_{i}^{\ast }=x\right] $ (up to an error of order $O({\Greekmath 011C} ^{p})$) using $\widetilde{v}(x)$ and $\widetilde{{\Greekmath 011A} }(x)$ defined in equations (ref) and (ref), respectively. Notice that $ {\Greekmath 011A} _{\mathcal{X}^{\ast }}\left( \varkappa \right) ={\Greekmath 011A} _{X^{\ast }}\left( {\Greekmath 0116} \left( \varkappa \right) \right) $. However, since $\mathcal{X} _{i}^{\ast }$ is not observed, one cannot identify ${\Greekmath 0116} \left( \varkappa \right) $ and ${\Greekmath 011A} _{\mathcal{X}^{\ast }}(\varkappa )$ without some sadditional information.
Suppose for a moment that the marginal distribution $F_{\mathcal{X}^{\ast }}$ of $\mathcal{X}_{i}^{\ast }$ is known (for example, from a separate dataset, e.g., administrative records). In this case, we can identify ${\Greekmath 011A} _{ \mathcal{X}^{\ast }}(\varkappa )$ up to an error of order $O({\Greekmath 011C} ^{p})$ using
where
Here $F_{\mathcal{X}^{\ast }}(\cdot )$ denotes the CDF of $\mathcal{X} _{i}^{\ast }$, and $Q_{X}(\cdot )$ denotes the quantile function (QF) of $ X_{i}$. Note that all functions on the right-hand side of equation ((ref)) are identified directly from the observed data.
First, we demonstrate that $\widetilde Q_{X^*} (s) = Q_{X^*}(s) + O({\Greekmath 011C}^p)$ , where $Q_{X^*}(\cdot)$ denotes the quantile function of $X_i^*$. Combining this with the result of Theorem (ref) allows us to establish the desired result formalized by the theorem below.
\setcounter{assumption}{0}
Theorem (ref) demonstrates that, if the marginal distribution of $\mathcal{X}_{i}^{\ast }$ is given, it is possible to identify ${\Greekmath 011A} (\varkappa )$ up to an error of order $O({\Greekmath 011C} ^{p})$ in the general NCME model (ref) building on the identification results for the WCME model.
Note that obtaining (an estimate of) the marginal distribution $F_{\mathcal{X }^{\ast }}$ is a much simpler task than obtaining a validation sample, i.e., the data on $\left( X_{i},\mathcal{X}_{i}^{\ast }\right) $ jointly. For example, suppose $\mathcal{X}_{i}^{\ast }$ are individual wages, and $X_{i}$ are self-reported wages in a survey. The marginal distribution $F_{\mathcal{X }^{\ast }}$ can be provided by the Social Security Administration or similar tax authorities in other countries. Providing such marginal distribution does not pose any privacy risks. In contrast, obtaining a validation sample that links individual's responses $X_{i}$ to the individual's social security records $\mathcal{X}_{i}^{\ast }$ is a difficult task that in particular faces major challenges concerning privacy.
If the distribution of $\mathcal{X}_{i}^{\ast }$ is unknown we can still apply Theorem (ref) at $\varkappa =Q_{\mathcal{X}^{\ast }}(q)$ for any quantile $q\in \left( 0,1\right) $ to identify $E[Y_{i}| \mathcal{X}_{i}^{\ast }=Q_{\mathcal{X}^{\ast }}(q)]={\Greekmath 011A} _{\mathcal{X}^{\ast }}(Q_{\mathcal{X}^{\ast }}(q))$, where $Q_{\mathcal{X}^{\ast }}(\cdot )$ is the (unknown) quantile function of $\mathcal{X}_{i}^{\ast }$:
Corollary (ref) demonstrates that even if $F_{\mathcal{X} ^*}$ is unknown, we can still identify the conditional expectation of $Y_i^*$ given the $q$'th quantile of $\mathcal{X}_{i}^{\ast }$. Notice that in some applications, the unobserved variable $\mathcal{X}_i^*$, for example an individual's ability, might not even have well-defined economic units. In such settings, identification of $E[Y_i | \mathcal{X}_i^* = Q_{\mathcal{X} ^*} (q)]$ is fully exhaustive.