EconBase
← Back to paper

Simple Estimation of Semiparametric Models with Measurement Errors

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

111,930 characters · 13 sections · 36 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

-1inSimple Estimation of Semiparametric Models with Measurement Errors

abstractWe develop a practical way of addressing the Errors-In-Variables (EIV) problem in the Generalized Method of Moments (GMM) framework. We focus on the settings in which the variability of the EIV is a fraction of that of the mismeasured variables, which is typical for empirical applications. For any initial set of moment conditions our approach provides a “corrected” set of moment conditions that are robust to the EIV. We show that the GMM estimator based on these moments is $\sqrt{n}$-consistent, with the standard tests and confidence intervals providing valid inference. This is true even when the EIV are so large that naive estimators (that ignore the EIV problem) are heavily biased with their confidence intervals having 0% coverage. Our approach involves no nonparametric estimation, which is especially important for applications with many covariates and settings with multivariate EIV. In particular, the approach makes it easy to use instrumental variables to address EIV in nonlinear models. \noindentKeywords: errors-in-variables, nonstandard asymptotic approximation, nonparametric identification, instrumental variables

Introduction

Measurement errors are a common problem for empirical studies. Addressing the Errors-In-Variables (EIV) bias in nonlinear models requires elaborate strategies.\footnote{ See HINP1991JoE,HausmanNeweyPowell1995JoE,Newey2001REStat,Schennach2007Ecta,Li2002JoE,Schennach2004Ecta,ChenHongTamer2005ReStud,HuSchennach2008Ecta,Schennach2014Ecta-ELVIS,Wilhelm2019WP-TestingForME , among others.} Despite the fundamental theoretical progress in identification and estimation of nonlinear models with EIV, the problem of EIV is still rarely addressed in empirical work outside of linear specifications.

{8.0pt plus 2.0pt minus 7.0pt} {6.0pt plus 2.0pt minus 5.0pt}

The goal of this paper is to develop a simple and practical approach to estimation of nonlinear semiparametric models that can be expressed in the form of general moment conditions

equation[equation omitted — 280 chars of source]

where $g\left( \cdot \right) $ is a vector of moment functions and ${\Greekmath 0112} _{0}$ is the parameter vector of interest. The researcher has a random sample of $\left\{ X_{i},S_{i}\right\} _{i=1}^{n}$, where scalar or vector $ X_{i}$ is a mismeasured version of unobserved $X_{i}^{\ast }$ with measurement error ${\Greekmath 0122} _{i}$:

equation*[equation* omitted — 60 chars of source]

We will refer to $ g\left( \cdot \right) $ as the original moment function, since it would have been valid had the researcher observed $X_{i}^{\ast }$. A naive GMM estimator (that ignores the EIV and uses $X_{i}$ in place of $ X_{i}^{\ast }$) based on $g\left( \cdot \right) $ is biased because $\mathbb{ E}[g(X_{i},S_{i},{\Greekmath 0112} _{0})]\neq 0$, in contrast to equation ((ref)).

example*[Nonlinear Regression, NLR] Let $Y_{i}$ denote a scalar outcome, and let $X_{i}^{\ast }$ and $W_{i}$ be the covariates. Suppose \begin{equation} E\left[ Y_{i}|X_{i}^{\ast },W_{i}\right] ={\Greekmath 011A} \left( X_{i}^{\ast },W_{i},{\Greekmath 0112} _{0}\right) \end{equation} for some function ${\Greekmath 011A} $ known up to the parameter ${\Greekmath 0112} $. For example, in the Logit model, $Y_{i}$ is binary, ${\Greekmath 011A} \left( x,w,{\Greekmath 0112} \right) \equiv 1\left/ \left( 1+\exp \left( -\left( {\Greekmath 0112} _{x}^{\prime }x+{\Greekmath 0112} _{w}^{\prime }w\right) \right) \right) \right. $, and ${\Greekmath 0112} \equiv \left( {\Greekmath 0112} _{x}^{\prime },{\Greekmath 0112} _{w}^{\prime }\right) ^{\prime }$. Suppose the researcher has an instrumental variable $Z_{i}$. Then, they can use \begin{equation*} g\left( y,x,w,z;{\Greekmath 0112} \right) \equiv \left( y-{\Greekmath 011A} \left( x,w,{\Greekmath 0112} \right) \right) {\Greekmath 0127} \left( x,w,z\right) \end{equation*} as the original moment function, where ${\Greekmath 0127} \left( x,w,z\right) $ is a vector that, for example, can include $x$, $z$, $w$, their powers and/or interactions.\footnote{ Note that the moment condition ((ref)) is stated in terms of the true (correctly measured) $X_{i}^{\ast }$. Determining what functions $g\left( \cdot \right) $ (or $h\left( \cdot \right) $ in the NLR model) satisfy this moment condition does not involve any consideration of the measurement errors and hence is straightforward.} Here $S_{i}=\left( Y_{i},W_{i},Z_{i}\right) $. \ensuremath{\blacksquare}

Even in this well-studied example of nonlinear regression, estimation in the presence of the EIV is a difficult problem. Importantly, nonlinear instrumental variable regression estimator cannot be used, since it is inconsistent in the presence of EIV Amemiya1985JoE. The existing approaches typically require nonparametric estimation that can be impractical in many empirical applications. In contrast, in this paper, we develop an alternative class of estimators, that are essentially GMM estimators that modify the original moment functions $g\left( \cdot \right) $ in a way that makes the moment conditions robust to the EIV. In particular, our approach makes it easy to use instrumental variables to address EIV\ in nonlinear models.

To provide a practical estimation approach\ for the general class of models ( (ref)), we focus on empirical settings in which the researcher believes the variability of the measurement error to be at most a fraction of the variability of the mismeasured variable, i.e., the noise-to-signal ratio ${\Greekmath 011C} \equiv {\Greekmath 011B} _{{\Greekmath 0122} }/{\Greekmath 011B} _{X^{\ast }} $ to be moderate. The absolute magnitude of the measurement error ${\Greekmath 011B} _{{\Greekmath 0122} }$ does not need to be small. Existing validation studies provide insights into the magnitude of ${\Greekmath 011C} $ for some key economic variables and datasets. BoundKrueger1991JoLaborEcon consider log-earnings in the Current Population Survey (CPS) data matched to the Social Security payroll records. Their estimates of the variance of the measurement errors correspond to $ {\Greekmath 011C} $ of approximately $0.47$ and $0.30$ for the subsamples of men and women, respectively. BoundBrownDuncanRodgers1994JoLaborEcon consider a validation study of Panel Study of Income Dynamics (PSID). Their estimates imply ${\Greekmath 011C} =0.39-0.66$ for log-earnings and ${\Greekmath 011C} =0.63-0.76$ for hours worked. Pischke1995JBES estimates correspond to ${\Greekmath 011C} =0.39-0.50$ for the log-earnings in PSID. AshenfelterKrueger1994AER assess the mismeasurement in the years of education; their estimates correspond to $ {\Greekmath 011C} =0.30-0.37$.

Focusing on these settings allows us to isolate the most important aspects of the problem and, as result, to develop a simple estimator, which does not require any nonparametric estimation or simulation. Such simple estimation becomes possible because in these settings we can obtain a simple approximation of the EIV\ bias of the moment conditions as a function of $ {\Greekmath 0112} $.

We propose to bias correct the original moments $g\left( \cdot \right) $, which in turn removes the bias of the corresponding estimator of ${\Greekmath 0112} _{0} $. This bias correction depends on some moments of the distribution of the measurement errors that are unknown. Another difficulty is that the estimators of some components of the bias correction themselves may need to be bias corrected. To address these issues, we develop the corrected moment conditions, which depend on ${\Greekmath 0112} $ and additional parameters ${\Greekmath 010D} $ that govern the bias correction. The true parameter value ${\Greekmath 010D} _{0}$ is associated with (possibly conditional) low-order moments of ${\Greekmath 0122} _{i}$. Despite some theoretical subtleties with the construction of the corrected moment conditions, their practical implementation is straightforward and they can be automatically computed for any original moment function $g\left( \cdot \right) $.

We introduce the Measurement Error Robust Moments (MERM) estimator, which is a GMM estimator that uses the corrected moment conditions to jointly estimate parameters ${\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{0}$. The estimator can be computed using any standard software for GMM estimation. Joint estimation of parameters ${\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{0}$ using the corrected moment conditions effectively robustifies moment conditions $g\left( \cdot \right) $ against the impact of the measurement errors.

To make these ideas precise and to study the properties of the proposed estimators, we develop an asymptotic theory using a nonstandard asymptotic approximation that models ${\Greekmath 011C} $ as slowly shrinking with the sample size. Standard asymptotics considers ${\Greekmath 011C} $ to be constant, which implies that as $n\rightarrow \infty $ the bias of a naive estimator dwarfs its sampling variability: the bias is constant while the standard errors shrink proportionally to $1/\sqrt{n}$. As a result, under the standard asymptotics, the problem of removing the EIV bias becomes central in the analysis, with relatively little attention paid to the sampling variability of estimators. However, this focus does not seem to be appropriate in many empirical applications, in which the researcher does not expect the potential EIV bias to be several orders of magnitude larger than the standard errors.\footnote{ Such empirical settings appear to be widespread. Although the concerns about measurement errors are often raised, the majority of applied work does not explicitly correct the EIV bias in nonlinear models, and instead implicitly or explicitly argues or conjectures that the EIV bias is likely not to be too large. } By considering ${\Greekmath 011C} $ as drifting towards zero with the sample size, our approach provides a better guidance on construction of EIV robust estimators with good finite sample properties when ${\Greekmath 011C} $ is small or moderate. \footnote{ Nonstandard asymptotic approximations with drifting parameters are often used to obtain better approximations of the finite sample behavior of estimators and tests. For example, in the instrumental variable regression settings, to consider the settings with relatively small first stage coefficients, StaigerStock1997 model them as shrinking with $n$. It is important to keep in mind that such nonstandard asymptotic approximations are merely mathematical tools. One should not take them literally and think of parameters somehow changing if more data is collected. }

Using this approximation, we show that the proposed estimation approach indeed addresses the EIV problem. The MERM estimator is shown to be $\sqrt{n} $-consistent and asymptotically normal and unbiased. The standard confidence intervals and tests for GMM estimators are also valid for the MERM\ estimator. Additionally, the standard GMM arsenal of assessment tools can be applied to the MERM estimator, allowing one to test model identification, conduct valid inference, and perform model specification diagnostics.

The usefulness of a large sample theory is measured by its ability to approximate the finite sample properties of the estimators and inference procedures. Thus, we study the MERM estimators in a variety of simulation experiments. The results confirm that the nonstandard asymptotic theory indeed provides a good approximation of the finite sample properties of the estimators even in the settings with relatively large EIV. In some of the simulation experiments, the EIV are so large that for the naive estimators' standard $95\%$ confidence intervals have actual coverages of $0\%$ in finite samples, due to the magnitude of the EIV\ bias. At the same time, even in these settings the MERM estimators perform well, removing the EIV bias and providing confidence intervals with the correct coverage. In particular, the simulation results show that despite the simplicity of implementation, the MERM estimators can compete with and outperform semi-nonparametric estimators.

The MERM estimator is structurally different from the existing approaches that require nonparametric estimation of some nuisance parameters, for example, of the density $f_{X^{\ast }|Z,W}$. Avoiding nonparametric estimation has at least two advantages. First, since the majority of empirical applications include additional covariates $ W_{i}$, nonparametric estimation is often infeasible due to the curse of dimensionality. Because the MERM estimator does not involve any nonparametric estimation, it can be used in applications with a relatively large number of additional covariates $W_{i}$, and remains feasible even in the more complicated settings, including multi-equation and structural models, and applications with multiple mismeasured variables $X_{i}$. Second, estimation of infinite-dimensional nuisance parameters is typically more demanding towards the sources of identification available in the data, for example, requiring an instrumental variable with a large support (continuously distributed). In contrast, having a discrete instrument is sufficient for the MERM approach because the nuisance parameter ${\Greekmath 010D} _{0}$ is finite-dimensional.

For example, in Section (ref) we consider estimation of the model of multinomial choice among three modes of transportation. A leading alternative approach to the errors-in-variables problem in this model is the semi-nonparametric sieve-MLE estimator advocated by HuSchennach2008Ecta,CarrollChenHu2010JoNS, among others. This approach requires, among other things, estimating the conditional density $ f_{X^{\ast }|Z,W}$ of $X_{i}^{\ast }$ given the instrument $Z_{i}$ and covariates $W_{i}$. In this empirical example, $W_{i}$ includes four continuously distributed covariates (two continuously distributed characteristics per choice) and a discrete one, while scalar $X_{i}^{\ast }$ and $Z_{i}$ are also continuous. Thus, $f_{X^{\ast }|Z,W}$ is a function of six continuous and one discrete variable. Hence, for typical sample sizes, estimating $f_{X^{\ast }|Z,W}$ in this example is infeasible due to the curse of dimensionality. In contrast, as the results of Section (ref) demonstrate, the MERM approach is practical and effective in this application, in part because it avoids estimation of the high-dimensional nuisance functions like $f_{X^{\ast }|Z,W}$ altogether.

The simplicity and practicality of the MERM\ approach do come at a cost: there is a limit on the magnitude of the measurement errors it can handle. For example, one generally should not expect the MERM\ approach to work well when ${\Greekmath 011C} >1$, i.e., when the noise dominates the signal; in this case the researcher should seek an alternative estimation method.

We view the MERM\ approach as providing a bridge between the settings in which the measurement errors are guaranteed to be absent or negligible, and the settings where the measurement errors are so large that one has to use the relatively more complicated estimators from the earlier literature (if they exist at all for the model of interest).

{ Related Literature} ChenHongNekipelov2011JEL, Schennach2016AnnRev, and Schennach2020HB-ME provide excellent overviews of the measurement error literature.

The existing semiparametric approaches to estimation and inference in models with EIV\ involve nonparametric estimation of infinite-dimensional nuisance parameters (e.g., \QTR{citealp}{ Chesher2000WP,Li2002JoE,Schennach2004Ecta,Schennach2007Ecta,HuSchennach2008Ecta,SchennachHu2013JASA,Song2015JoE }), simulation (e.g., \QTR{citealp}{Schennach2014Ecta-ELVIS}), or both (e.g., \QTR{citealp}{Newey2001REStat,WangHsiao2011JoE}$)$. The exceptions include models with linear and polynomial regression functions (see \QTR{citealp}{HINP1991JoE,HausmanNeweyPowell1995JoE}$)$, and Gaussian control variable models such as Probit and Tobit with endogeneity (see \QTR{citealp}{SmithBlundell1986Ecta,RiversVuong1988JoE}).

To the best of our knowledge, this paper is the first to provide an approach for $\sqrt{n}$-consistent and asymptotically normal and unbiased estimation of general GMM models with EIV that does not require any nonparametric estimation (or simulation).

We are able to provide such an estimator because we focus on the models with moderate measurement errors. Modeling the variance of the measurement error as shrinking to zero with the sample size is a popular approach in Statistics. The method has been proposed by WolterFuller1982AS, who used it to construct an approximate MLE\ estimator of a nonlinear regression model with Gaussian errors. Following their approach, the Statistics literature has mainly focused on the settings where the moments of the EIV needed to bias correct the estimators are either known or can be directly estimated from the available data such as repeated measurements (e.g., \QTR{citealp}{CarrollStefanski1990JASA,CarrollEtAl2006Book-ME}). In Economics, such data are relatively rare. The use of approximations with shrinking variance of measurement errors in Econometrics literature has been pioneered by Kadane1971Ecta, Amemiya1985JoE, and Chesher1991Biomet. Such approximations have been used to check the sensitivity of naive estimators to the EIV by considering how the estimates change as the unknown moments of the measurement errors vary within some set of plausible values, e.g., see ChesherSchluter2002ReStud, ChesherDumanganeSmith2002JoE, BattistinChesher2014JoE, Chesher2017JoE, and HongTamer2003JoE. BoundBrownMathiowetz2001HBoE review a broad list of validation studies matching standard economic dataset to administrative records. The estimates they report suggest that the measurement errors of moderate magnitude are typical for empirical applications. This suggests that the approach developed in this paper could prove valuable for a wide range of applied work.

This paper differs from the earlier literature in several ways. First, it presents a way to estimate the unknown nuisance parameters (moments of the measurement errors) jointly with the parameters of interest. As a result, the approach can, for example, use instrumental variables as a source of identification. Second, the method applies to a very general class of semiparametric models specified by moment conditions. Third, the MERM approach allows the measurement errors to have larger magnitudes than most of the papers in the earlier literature; this is achieved by the MERM approach recursively bias correcting the bias correction terms.

The most widespread approach to identification of the EIV\ models in economic applications is to use instrumental variables, e.g., see HINP1991JoE,Newey2001REStat,Schennach2007Ecta,WangHsiao2011JoE. In a recent paper, HahnHausmanKim2021EL reconsider the regression model in Amemiya1990JoE using a bias correction similar to ours. When proper excluded variables are not available, researchers have considered using higher moments of $X_{i}$ as instruments, e.g., see Reiersol1950Ecta,Lewbel1997Ecta,EricksonWhited2002ET,SchennachHu2013JASA,BenMosheDHaultfeuilleLewbel2017JoE . When available, repeated measurements can also be used to identify the model, e.g., see HINP1991JoE,LiVuong1998JoE,Li2002JoE,Schennach2004Ecta. The MERM estimator accommodates these identification approaches within a unified estimation framework.

The power of the general MERM\ approach can be illustrated in the NLR model. For example, when a candidate instrumental variable is available, the conditions it needs to satisfy are much weaker than what is required by many existing approaches. Availability of a discrete instrument is sufficient for identification; and the instrument is allowed to have heterogeneous impact on covariates $X_{i}^{\ast }$. \ One can also take a nonclassical, nonlinear (e.g., discretized or censored), or biased measurement of $X_{i}^{\ast }$ as an instrument in the MERM approach. We discuss identification in Section (ref). In addition, in a related paper EvdokimovZeleneev2022WP-NPID study nonparametric regression with EIV using the ${\Greekmath 011C} \rightarrow 0$ approximation, and demonstrate that the MERM approach can also be motivated from a nonparametric perspective.

KitamuraOtsuEvdokimov2013Ecta,AndrewsGentzkowShapiro2017QJE,ArmstrongKolesar2021QE,BonhommeWeidner2021QE , among others, develop tools for estimation and inference in GMM, which are robust to general perturbation or misspecification of the true data generating process. They focus on the settings in which these perturbations are sufficiently small, so that naive estimators remain $\sqrt{n}$ -consistent, and their biases are of the same order of magnitude as their standard errors. In contrast, we focus on more specific forms of data contamination due to the EIV. This allows the MERM approach to remain valid even in the settings with larger measurement errors, in which naive estimators may have slower than $\sqrt{n}$ rates of convergence.

The MERM approach also provides a useful foundation for dealing with EIV in more complicated settings. EvdokimovZeleneev2018WP-Inference utilize the MERM framework to address an issue of nonstandard inference, which turns out to arise generally when EIV models are identified using instrumental variables. EvdokimovZeleneev2019WP-Panel extend the analysis of this paper to long panel and network settings.

{ Organization of the paper} Section (ref) introduces the Moderate Measurement Error framework and the proposed MERM\ estimator. Section (ref) presents several Monte Carlo experiments that illustrate finite sample properties of the MERM\ estimators. Section (ref) considers several extensions of the framework. A supplementary appendix contains all proofs and additional results for the numerical and empirical illustrations.

Moderate Measurement Errors Framework

To present the main ideas we first consider the case of univariate $ X_{i}^{\ast }$. We will consider multivariate $X_{i}^{\ast }$ later. We assume that the measurement error is classical, i.e., that ${\Greekmath 0122}_i$ is independent of $X_i^*$ and $S_i$; later we will discuss how this assumption can be relaxed. Following the rest of the literature, we assume that $\mathbb{E}\left[ {\Greekmath 0122} _{i}\right] =0$.\footnote{ A location normalization such as $\mathbb{E}\left[ {\Greekmath 0122} _{i}\right] =0 $ is usually necessary because it is not possible to separately identify the means $\mathbb{E}\left[ X_{i}^{\ast }\right] $ and $\mathbb{E}\left[ {\Greekmath 0122} _{i}\right] $.}

To develop a practical estimation approach for general moment condition models we focus on the settings in which ${\Greekmath 011C} \equiv {\Greekmath 011B} _{{\Greekmath 0122} }/{\Greekmath 011B} _{X^{\ast }}$ is small or moderate. We consider an asymptotic approximation with ${\Greekmath 011C} _{n}\equiv {\Greekmath 011C} \rightarrow 0$ as $n\rightarrow \infty $.

Note that economically meaningful parameters are usually invariant to rescaling of $X_{i}^{\ast }$. Likewise, the extent of the EIV problem does not change with such rescaling. For simplicity of exposition, it is convenient to assume that $X_{i}^{\ast }$ is scaled so that ${\Greekmath 011B} _{X^{\ast }}$ is of order one and, correspondingly, the moments $\mathbb{E}[\left\vert {\Greekmath 0122} _{i}\right\vert ^{k}]\propto {\Greekmath 011C} _{n}^{k}$ decrease with $k$ when ${\Greekmath 011C} _{n}<1$. For example, this could be ensured by normalizing observed $X_{i}$ to have ${\Greekmath 011B} _{X}=1$. Let us stress that this normalization is used only to simplify the exposition; as we show in Appendix (ref), the proposed MERM\ estimator does not require any normalizations in practice.

Special Case: Quadratic Expansion

For clarity, we first consider a simple special case of the general approach. Let us denote $g_{x}^{(k)}\left( x,s,{\Greekmath 0112} \right) \equiv \partial ^{k}g\left( x,s,{\Greekmath 0112} \right) /\partial x^{k}$. Since $\mathbb{E} [\left\vert {\Greekmath 0122} _{i}\right\vert ^{k}]\propto {\Greekmath 011C} _{n}^{k}\rightarrow 0$ as $n\rightarrow \infty $, under some regularity conditions, we can write the quadratic Taylor expansion of function $ g(X_{i},S_{i},{\Greekmath 0112} )=g(X_{i}^{\ast }+{\Greekmath 0122} _{i},S_{i},{\Greekmath 0112} )$ around ${\Greekmath 0122} _{i}=0$ as

align[align omitted — 609 chars of source]

where the second equality holds because ${\Greekmath 0122} _{i}$ and $ \left( X_{i}^{\ast },S_{i}\right) $ are independent, and ${\mathbb{E}\left[ {\Greekmath 0122} _{i}\right] =0}$.

Evaluating the expansion above at ${\Greekmath 0112} = {\Greekmath 0112}_0$ gives $\mathbb{E}[g(X_{i},S_{i},{\Greekmath 0112} _{0})]=O\left( {\Greekmath 011B} _{{\Greekmath 0122} }^{2}\right) =O({\Greekmath 011C} _{n}^{2})$, because $\mathbb{E}[g(X_{i}^*,S_{i},{\Greekmath 0112} _{0})] = 0$. As a result, a naive estimator that ignores the EIV and uses $X_{i}$ in place of $X_{i}^{\ast }$ has EIV\ bias of order ${\Greekmath 011C} _{n}^{2}$.\footnote{ For example, consider a linear regression with a scalar mismeasured regressor. The bias of the naive OLS\ estimator of the slope parameter $ {\Greekmath 0112} _{01}$ is $-{\Greekmath 0112} _{01}\frac{{\Greekmath 011C} _{n}^{2}}{1+{\Greekmath 011C} _{n}^{2}}=-{\Greekmath 0112} _{01}{\Greekmath 011C} _{n}^{2}+O\left( {\Greekmath 011C} _{n}^{4}\right) $.} The bias of the naive estimator should be compared with its standard error, which is of order $ n^{-1/2}$. Thus, the bias of the naive estimator is not negligible, unless the measurement error is rather small (theoretically, unless ${\Greekmath 011C} _{n}^{2}=o\left( n^{-1/2}\right) $). In particular, tests and confidence intervals based on the naive estimator are invalid and can provide highly misleading results. Moreover, if ${\Greekmath 011C} _{n}^{2}$ shrinks at a rate slower than $O\left( n^{-1/2}\right) $, the rate of convergence of the naive estimator is slower than $\sqrt{n}$.

Suppose ${\Greekmath 011C} _{n}=o\left( n^{-1/6}\right) $. Then, $O({\Greekmath 011C} _{n}^{3})=o\left( n^{-1/2}\right) $ and we can rearrange equation ((ref)) as

equation[equation omitted — 281 chars of source]

The left-hand side of this equation is exactly the moment condition ((ref)) that we would like to use for estimation of ${\Greekmath 0112} _{0} $. The first term on the right-hand side involves only observed variables, and can be estimated by the sample average $\overline{g}({\Greekmath 0112} )\equiv n^{-1}\sum_{i=1}^{n}g(X_{i},S_{i},{\Greekmath 0112} ).$ The second term on the right-hand side can be thought of as a bias correction that removes the EIV-bias from the expected moment function $\mathbb{E}[g(X_{i},S_{i},{\Greekmath 0112} )]$.

The idea of the MERM\ estimator we propose is to make use of expansions such as ((ref)) to bias correct the moment condition $\mathbb{ E}[g(X_{i},S_{i},{\Greekmath 0112} )]$, which in turn removes the bias of the estimator of the parameters of interest ${\Greekmath 0112} _{0}$. To perform the bias correction we need to estimate two quantities: $\mathbb{E}[{\Greekmath 0122} _{i}^{2}]$ and $ \mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$.

First, we show that in equation ((ref)) we can substitute $\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$ with $ \mathbb{E}[g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]$, which in turn can be estimated by $\overline{g}_{x}^{(2)}({\Greekmath 0112} )\equiv n^{-1}\sum_{i=1}^{n}g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )$. By the Taylor expansion around ${\Greekmath 0122} _{i}=0$ similar to equation ((ref)), we can show that $\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]=\mathbb{E}[g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]+O({\Greekmath 011C} _{n}^{2})$ and hence

equation[equation omitted — 394 chars of source]

Here $O\left( {\Greekmath 011C} _{n}^{4}\right) =o\left( n^{-1/2}\right) $ because we assume that ${\Greekmath 011C} _{n}=o\left( n^{-1/6}\right) $. The idea behind this substitution is that the bias of order $O({\Greekmath 011C} _{n}^{2})$ in $\mathbb{E} [g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]$ can be ignored because it is multiplied by $E\left[ {\Greekmath 0122} _{i}^{2}\right] =O\left( {\Greekmath 011C} _{n}^{2}\right) $. \footnote{ Such substitutions of $X^{\ast }$ with $X$ have been used in other contexts, e.g., ChesherSchluter2002ReStud.} With the substitution, we can rearrange equation ((ref)) and write it as

equation[equation omitted — 276 chars of source]

Second, we propose estimating the unknown $\mathbb{E}[{\Greekmath 0122} _{i}^{2}]$ together with the parameter of interest ${\Greekmath 0112} $. Specifically, let ${\Greekmath 010D} _{02}\equiv \mathbb{E}[{\Greekmath 0122} _{i}^{2}]/2$ denote the true value of parameter ${\Greekmath 010D} _{2}$, and consider the following corrected moment function:

equation[equation omitted — 232 chars of source]

Function ${\Greekmath 0120} $ is a moment function parameterized by ${\Greekmath 0112} $ and ${\Greekmath 010D} $, and

equation[equation omitted — 252 chars of source]

where the first equality follows from equation ((ref)) and the definition of ${\Greekmath 010D} _{02}$, and the second equality follows from equation ((ref)). Hence, the corrected moment conditions ${\Greekmath 0120} $ can be used to jointly estimate the true parameters $ {\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{02}$ by a GMM\ estimator.\footnote{ In the moment condition settings, having $o\left( n^{-1/2}\right) $ is equivalent to having $0$ on the right-hand side\ of equation ((ref)).}

remarkIf $\mathbb{E}[{\Greekmath 0122} _{i}^{3}]=0$ (e.g., if the distribution of ${\Greekmath 0122} _{i}$ is symmetric), the remainder in equation ((ref)) is of a smaller order $O({\Greekmath 011C} _{n}^{4})$. Hence, the corrected moments ((ref)) remain valid for larger values of ${\Greekmath 011C} _{n}$, requiring only the weaker condition ${\Greekmath 011C} _{n}=o(n^{-1/8})$. The bias of the naive estimators in this case can be as large as $o(n^{-1/4})$.

General Case: Expansion of order $K$

The quadratic expansion of equation ((ref)) can be extended to general order $K\geq 2$. Considering larger $K$ theoretically allows ${\Greekmath 011C} _{n}$ converging to zero at a slower rate. In finite samples this corresponds to the asymptotics providing good approximations for larger values of ${\Greekmath 011C} _{n}$, i.e., large measurement errors. Expanding $g(X_{i}^{\ast }+{\Greekmath 0122} _{i},S_{i},{\Greekmath 0112} )$ around ${\Greekmath 0122} _{i}=0$ we have,

equation[equation omitted — 365 chars of source]

The above special case of quadratic expansion corresponds to $K=2$.

The approximation we consider is formalized by the following assumption.

assumption[MME] (Moderate Measurement Errors) \namedlabel{ass:MME}{MME} (i) ${\Greekmath 011C} _{n}=o(n^{-1/\left( 2K+2\right) })$ for some integer $K\geq 2$; and (ii) $\mathbb{E}[\left\vert {\Greekmath 0122} _{i}\right\vert ^{L}] \leq C {\Greekmath 011B}_{\Greekmath 0122}^L $ for some $L\geq K+1$ and $C > 0$.

Assumption (ref)(i) limits the magnitude of the measurement errors and implies that ${\Greekmath 011C} _{n}^{K+1}=o\left( n^{-1/2}\right) $. Assumption (ref)(ii) implies that $\mathbb{E}[\left\vert {\Greekmath 0122} _{i}\right\vert ^{k}]=O\left( {\Greekmath 011B} _{{\Greekmath 0122} }^{k}\right) $, and requires the tails of ${\Greekmath 0122} _{i}/{\Greekmath 011B} _{{\Greekmath 0122} }$ to be sufficiently thin. Together, parts (i) and (ii) imply that $\mathbb{E} [\left\vert {\Greekmath 0122} _{i}\right\vert ^{K+1}]=O\left( {\Greekmath 011C} _{n}^{K+1}\right) =o\left( n^{-1/2}\right) $, and hence ensure that the remainder in equation ((ref))$\ $is negligible. Using $\mathbb{E}\left[ {\Greekmath 0122} _{i}|X_{i}^{\ast },S_{i} \right] =0$ to further simplify this expansion and rearranging the terms we obtain

equation[equation omitted — 288 chars of source]

This equation is the general expansion analog of equation ((ref)). The summation on the right hand side is the bias correction term, which we use to construct the MERM\ estimator.\footnote{It is useful to get a sense of the magnitudes of the coefficients $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{k}\right] /k!$ in equation ((ref)). Suppose ${\Greekmath 0122} _{i}\sim N\left( 0,{\Greekmath 011B} _{{\Greekmath 0122} }^{2}\right) $, ${\Greekmath 011B} _{{\Greekmath 0122} }=0.5$, and $ {\Greekmath 011B} _{X^{\ast }}=1$, so$\ {\Greekmath 011C} ={\Greekmath 011B} _{{\Greekmath 0122} }=0.5$. Then the coefficients in front of $g_{x}^{\left( 2\right) }$, $g_{x}^{\left( 4\right) }$, and $g_{x}^{\left( 6\right) }$ are $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{2} \right] /2!=0.125$, $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{4}\right] /4!\approx 0.008$, and $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{6}\right] /6!\approx 0.0003$ . }

It turns out that for $K\geq 4$, estimation of $\mathbb{E} [g_{x}^{(k)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$ is more intricate than in the case of $K=2$, and the substitution we made in equation ((ref)) no longer works. Larger values of $K$ allow for larger values of ${\Greekmath 011C} _{n}$ and hence larger EIV\ biases of naive estimators $ n^{-1}\sum_{i=1}^{n}g_{x}^{(k)}(X_{i},S_{i},{\Greekmath 0112} )$. The expansion of order $K$ includes terms up to the order ${\Greekmath 011C} _{n}^{K}$, with the asymptotically negligible remainder of order $O\left( {\Greekmath 011C} _{n}^{K+1}\right) $. For $K\geq 4$, terms of order ${\Greekmath 011C} _{n}^{4}$ are not negligible. This implies that we cannot ignore the EIV bias that would arise from substituting $\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$ with $ \mathbb{E}[g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]$ in equation ((ref) ), because this bias is of order $O\left( {\Greekmath 011C} _{n}^{4}\right) $ according to equation ((ref)). To address this problem, we instead replace $\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$ with the bias corrected expression $\mathbb{E} [g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]-\left( \mathbb{E}[{\Greekmath 0122} _{i}^{2}]/2\right) \mathbb{E}[g_{x}^{(4)}(X_{i},S_{i},{\Greekmath 0112} )]$. Thus, for $ K\geq 4$, one needs to bias correct the estimator of the bias correction term. Moreover, for larger $K$ one needs to bias correct the bias correction of the bias correction term and so on.

Fortunately, we show that these bias corrections can be constructed as linear combinations of the expectations of the higher order derivatives of $ g_{x}^{\left( k\right) }(X_{i},S_{i},{\Greekmath 0112} )$. Let us define the following corrected moment function:

equation[equation omitted — 238 chars of source]

where ${\Greekmath 010D} =({\Greekmath 010D} _{2},\dots ,{\Greekmath 010D} _{K})^{\prime }$ is a $K-1$ dimensional vector of parameters. Let ${\Greekmath 010D} _{0}\equiv ({\Greekmath 010D} _{02},\dots ,{\Greekmath 010D} _{0K})^{\prime }$ denote the vector of true parameters ${\Greekmath 010D} _{0k}$ , defined as

equation[equation omitted — 683 chars of source]

We formalize this discussion below.

assumption[CME] \namedlabel{ass: CME}{CME} (Classical Measurement Error) ${\Greekmath 0122} _{i}$ is independent from $ (X_{i}^{\ast },S_{i})$ and $\mathbb{E}[{\Greekmath 0122} _{i}]=0$.

The following lemma establishes validity of the corrected moment conditions under Assumptions (ref), (ref), and some mild regularity conditions provided in Appendix (ref).

lemmaUnder Assumptions (ref), (ref) and (ref) in Appendix (ref), \begin{equation*} \mathbb{E}[{\Greekmath 0120} (X_{i},S_{i},{\Greekmath 0112} _{0},{\Greekmath 010D} _{0})]=\mathbb{E} [g(X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0})]+o\left( n^{-1/2}\right) =o\left( n^{-1/2}\right). \end{equation*}

Lemma (ref) implies that the corrected moment conditions ${\Greekmath 0120} $ are valid and can potentially be used to jointly estimate parameters ${\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{0}$. The total number of parameters to be estimated is now $\dim \left( {\Greekmath 0112} \right) +K-1$. Thus, joint estimation of ${\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{0}$ requires that $\dim \left( {\Greekmath 0120} \right) =\dim \left( g\right) \geq \dim \left( {\Greekmath 0112} \right) +K-1$, i.e., that the original moment conditions $g$ include sufficiently many overidentifying restrictions. For example, the overidentifying restrictions can be constructed by using an instrumental variable; we discuss this in more detail below.

remarkConstruction of the corrected moment conditions ${\Greekmath 0120}$ requires the original moment function $g(x,s,{\Greekmath 0112})$ to have a sufficient number of derivatives with respect to $x$. Thus, the proposed correction method does not apply to settings with non-differentiable moment functions, for example, those arising in the instrumental variable quantile regression (IVQR).

Measurement Error Robust Moments (MERM) estimator

The MERM estimator jointly estimates the parameters ${\Greekmath 0112} _{0}$ and $ {\Greekmath 010D} _{0}$ using moment conditions ${\Greekmath 0120} $. It is convenient to define the joint vector of parameters

equation*[equation* omitted — 400 chars of source]

and the parameter space $\mathcal{B}\equiv \Theta \times \Gamma $, where $ \Theta $ and $\Gamma $ are the parameter spaces for ${\Greekmath 0112} $ and ${\Greekmath 010D} $. Then, MERM estimator is the GMM\ estimator (\QTR{citealp}{Hansen1982Ecta}):

equation[equation omitted — 339 chars of source]

where $\overline{{\Greekmath 0120} }({\Greekmath 010C} )\equiv n^{-1}\sum_{i=1}^{n}{\Greekmath 0120} _{i}({\Greekmath 010C} )$ , ${\Greekmath 0120} _{i}({\Greekmath 010C} )\equiv {\Greekmath 0120} \left( X_{i},S_{i},{\Greekmath 010C} \right) $, $\hat{\Xi }$ is a weighting matrix, and $\hat{Q}({\Greekmath 010C} )$ is the standard GMM objective function.

While Lemma (ref) establishes validity of the corrected moment restrictions ${\Greekmath 0120} $, the MERM estimator also relies on $ {\Greekmath 010C} _{0}$ being identified from ${\Greekmath 0120} $. This requirement is formalized by the following assumption.

assumption[ID] (Identification) \namedlabel{ass: ID}{ID} \begin{enumerate}[(i)] • the Jacobian $\Psi ^{\ast }$ has full column rank, where \begin{align*} \Psi ^{\ast }& \equiv \mathbb{E}\left[ {\Greekmath 0272} _{{\Greekmath 0112} }{\Greekmath 0120} (X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0},0),{\Greekmath 0272} _{{\Greekmath 010D} }{\Greekmath 0120} (X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0},0)\right] \\ & =\mathbb{E}\left[ {\Greekmath 0272} _{{\Greekmath 0112} }g(X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}),-g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}),\dots ,-g_{x}^{(K)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0})\right] ; \end{align*} • $\mathbb{E}\left[ {\Greekmath 0120} (X_{i}^{\ast },S_{i},{\Greekmath 0112} ,{\Greekmath 010D} )\right] =0$ iff ${\Greekmath 0112} ={\Greekmath 0112} _{0}$ and ${\Greekmath 010D} =0$. \end{enumerate}

Assumptions (ref)((ref)) and ((ref)) are the standard GMM local and global identification conditions applied to the moment function ${\Greekmath 0120} (X_{i}^{\ast },S_{i},{\Greekmath 0112} ,{\Greekmath 010D} )$. These are high-level conditions, which we will return to later in the paper. The moment conditions formulation is sufficiently general to encompass a wide variety of sources of identification. In Section (ref), we discuss identification in detail and illustrate the construction of the moment function using an instrumental variable or a second measurement.

Under some additional regularity conditions, estimator $\hat{{\Greekmath 010C}}$ behaves as a standard GMM-type estimator: it is $\sqrt{n}$-consistent and asymptotically normal and unbiased. This result is formalized by the following theorem.

theorem[Asymptotic Normality] Suppose that $\{(X_{i}^{\ast },S_{i}^{\prime },{\Greekmath 0122} _{i})\}_{i=1}^{n}$ are i.i.d.. Then, under Assumptions (ref), (ref), (ref), and (ref)-(ref) in Appendix (ref), \begin{equation} n^{1/2}\Sigma ^{-1/2}(\hat{{\Greekmath 010C}}-{\Greekmath 010C} _{0})\overset{d}{\rightarrow } N(0,I_{\dim ({\Greekmath 010C} )}),\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ where} \end{equation} \begin{equation*} \Sigma \equiv (\Psi ^{\prime }\Xi \Psi )^{-1}\Psi ^{\prime }\Xi \Omega _{{\Greekmath 0120} {\Greekmath 0120} }\Xi \Psi (\Psi ^{\prime }\Xi \Psi )^{-1}. \end{equation*}

Theorem (ref) shows that the MERM\ approach addresses the EIV bias problem, and in particular provides a $\sqrt{n}$-consistent asymptotically normal and unbiased estimator $\hat{{\Greekmath 0112}}$, which can be used to conduct inference about the true parameters ${\Greekmath 0112} _{0}$. The asymptotic variance $\Sigma $ takes the standard sandwich\ form, with $\Psi \equiv \mathbb{E}\left[ {\Greekmath 0272} _{{\Greekmath 010C} }{\Greekmath 0120} _{i}({\Greekmath 010C} _{0})\right] $, $ \Omega _{{\Greekmath 0120} {\Greekmath 0120} }\equiv \mathbb{E}\left[ {\Greekmath 0120} _{i}\left( {\Greekmath 010C} _{0}\right) {\Greekmath 0120} _{i}^{\prime }\left( {\Greekmath 010C} _{0}\right) \right] $, and $\hat{ \Xi}\rightarrow _{p}\Xi $.

remarkNotice that the bias of naive estimators (such as a GMM estimator based on the original moment conditions) is $O({\Greekmath 011C} _{n}^{2})$, so their rate of convergence is $O_{p}({\Greekmath 011C} _{n}^{2}+n^{-1/2})$. The bias dominates sampling variability and naive estimators are not $\sqrt{n}$-consistent unless ${\Greekmath 011C} _{n}=O(n^{-1/4})$, i.e., unless the magnitude of the measurement error is rather small. At the same time, the MERM estimator remains $\sqrt{n}$ -consistent for much larger values of ${\Greekmath 011C} _{n}$, up to ${\Greekmath 011C} _{n}=O(n^{-1/(2K+2)})$, whereas the rate of convergence of naive estimators is only $O_{p}(n^{-1/(K+1)})$ in this case.

Once the corrected moment condition ${\Greekmath 0120} $ is constructed, estimation of and inference about parameters ${\Greekmath 010C} _{0}$ can be performed using any standard software package for GMM\ estimation. In other words, the proposed estimator can be simply treated as a standard GMM estimator based on the corrected moment conditions ${\Greekmath 0120} $, and the conventional standard errors, tests, and confidence intervals are valid.

In addition to the estimation of the parameters ${\Greekmath 0112}_{0}$, researchers are often interested in average effects of the form ${\Greekmath 0115}_{0}\equiv \mathbb{E}\left[ {\Greekmath 0115}\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) \right] $. For instance, in the NLR model, one may be interested in the average partial effect of $x$ (i.e., ${\Greekmath 0115} _{0}\equiv \mathbb{E}\left[ {\Greekmath 0272} _{x}{\Greekmath 011A} \left( X_{i}^{\ast },S_{i},{\Greekmath 0112}_{0}\right) \right] $) or another covariate. The naive average partial effect estimator $\hat{{\Greekmath 0115}}_{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Naive}}\equiv \frac{1}{n} \sum_{i=1}^{n}{\Greekmath 0115}(X_{i},S_{i},\hat{{\Greekmath 0112}})$ suffers from the EIV bias, unless function ${\Greekmath 0115}$ is linear in $X_{i}^{\ast }$. Instead, one should use estimates $\hat{{\Greekmath 010D}}$ to construct the bias-corrected estimator $\hat{ {\Greekmath 0115}}_{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{MERM}}\equiv \frac{1}{n}\sum_{i=1}^{n}\left\{ {\Greekmath 0115}(X_{i},S_{i},\hat{{\Greekmath 0112}})-\sum_{k=2}^{K}\hat{{\Greekmath 010D}} _{k}{\Greekmath 0115}_{x}^{(k)}(X_{i},S_{i},\hat{{\Greekmath 0112}})\right\} $.

remarkThe standard $J$-test of overidentifying restrictions remains valid in the MERM settings, and can be used to check the model specification. The $J$ -test jointly tests the following hypotheses: (i) $K$ is sufficiently large to correct the EIV\ bias; (ii) assumptions on the EIV are valid; and (iii) the original moment conditions $g$ are correctly specified so equation ((ref)) holds, i.e., that the original economic model is correctly specified aside from the presence of the EIV in $X_{i}$. Thus, if the $J$-test rejects the validity of the corrected moment conditions ${\Greekmath 0120}$, the researcher might want to (i) consider taking a larger $K$; (ii) employ a different correction method; or (iii) consider an alternative specification of the original moments $g$.
remarkConsidering larger $K$ allows for ${\Greekmath 011C} _{n}\ $converging to zero at a slower rate, which in finite samples corresponds to the asymptotics providing better approximations for larger magnitudes of measurement errors. On the other hand, taking a larger $K$ increases the dimension of the nuisance parameter ${\Greekmath 010D} _{0}$ and thus typically increases the variance of $\hat{ {\Greekmath 0112}}$. We consider this issue in more detail and provide a data-driven method for choosing $K$ in Section (ref).
remarkThe MERM framework can be extended to the case of non-classical measurement errors; see EvdokimovZeleneev2022WP-NPID for details and a fully nonparametric analysis. In this paper, we focus on the classical measurement errors, developing a practical bias correction approach, which can be easily implemented in a wide range of economic applications. Even when Assumption (ref) is violated, the deviations from it are often limited in magnitude, so the corrections based on the (ref) assumption remove most of the EIV bias. Thus, in practice, using the estimator designed for classical measurement errors is typically preferable to ignoring mismeasurement altogether.
remarkIt is important to note that ${\Greekmath 010D} _{0k}\neq \mathbb{E}\left[ {\Greekmath 0122} _{i}^{k}\right] \left/ k!\right. $ for $k\geq 4$,\ contrary to what equation ((ref)) might suggest. For example, ${\Greekmath 010D} _{04}=\left( \mathbb{E}\left[ {\Greekmath 0122} _{i}^{4}\right] -6{\Greekmath 011B} _{{\Greekmath 0122} }^{4}\right) \left/ 24\right. $ is negative for many distributions, including normal. The reason that generally ${\Greekmath 010D} _{0k}\neq \mathbb{E}\left[ {\Greekmath 0122} _{i}^{k}\right] \left/ k!\right. $ is that the estimators of the correction terms themselves need a correction, which is accounted for by the form of ${\Greekmath 010D} _{0k}$. Since there is a one-to-one relationship between ${\Greekmath 010D} _{0}$ and the moments $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{\ell }\right] $, parameter space $\Gamma $ for ${\Greekmath 010D} _{0}$ can incorporate restrictions that the moments must satisfy (e.g., ${\Greekmath 011B} _{{\Greekmath 0122} }^{2}\geq 0$ and $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{4}\right] \geq {\Greekmath 011B} _{{\Greekmath 0122} }^{4}$). Such restrictions can increase the efficiency of the estimator and the power of tests.
remarkNo parametric assumptions are imposed on the distribution of ${\Greekmath 0122} _{i}$, i.e. the distribution of ${\Greekmath 0122} _{i}$ is treated nonparametrically. The regularity conditions restrict only the magnitude of the moments of ${\Greekmath 0122} _{i}$. The approach imposes no restrictions on the smoothness of the distributions of $X_{i}^{\ast }$ and ${\Greekmath 0122} _{i}$ , which are not even required to be continuous. Examples in which this can be useful include individual wages (whose distributions may have point masses at round numbers), and allowing the measurement error ${\Greekmath 0122} _{i}$ to have a point mass at zero (a fraction of the population may have a zero measurement or recall error).
remarkThe formulas of the derivatives $g_{x}^{(k)}(\cdot )$ are typically easy to compute analytically or using symbolic algebra software. Alternatively, these derivatives can be computed using numerical differentiation. Thus, the corrected set of moments can be automatically produced for a generic moment function $g(\cdot )$ provided by the user.

Model Identification:\ Jacobian $\Psi $

Theorem (ref) requires ${\Greekmath 010C} _{0}$ to be identified and the Jacobian matrix $\Psi $ to be full rank. Notably, the MERM framework encompasses many possible sources of identification at once, including instrumental variables, additional measurements, or nonlinearities of the functional form. The identifying information is incorporated in the moment functions. Essentially, our approach first characterizes in what directions the measurement errors can bias the moment conditions $\mathbb{E}\left[ g\left( X_{i},S_{i},{\Greekmath 0112} \right) \right] $, and then uses the moments orthogonal to those directions for identification of ${\Greekmath 0112} _{0}$. To be more specific, we will now consider identification when in addition to the error-laden $X_{i}$ we have either (i) a general instrument $Z_{i}$ or (ii) a second measurement $Q_{i}$.

\paragraph{Identification Using A General Instrument $Z_{i}$}

Many applications can be formulated as the following conditional moment restriction:

equation[equation omitted — 251 chars of source]

for some moment function $u$. For example, consider the nonlinear regression model $\mathbb{E}\left[ Y_{i}|X_{i}^{\ast }=x\right] ={\Greekmath 011A} \left( x,{\Greekmath 0112} _{0}\right) $, then $u\left( x,y,{\Greekmath 0112} \right) ={\Greekmath 011A} \left( x,{\Greekmath 0112} \right) -y$.\footnote{ For simplicity of the exposition, in the expectation in equation ((ref)) we only condition on $X_{i}^{\ast }$. The discussion applies in a straightforward way to the settings with additional correctly measured variables $W_{i}$ in the conditioning set, i.e., the model $\mathbb{E}\left[ u\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) |X_{i}^{\ast },W_{i}\right] =0 $. For example, in the nonlinear regression example with additional covariates $W_{i}$ we have $\mathbb{E}\left[ Y_{i}|X_{i}^{\ast }=x,W_{i}=w \right] ={\Greekmath 011A} \left( x,w,{\Greekmath 0112} _{0}\right) $, so $u\left( x,y,w,{\Greekmath 0112} \right) ={\Greekmath 011A} \left( x,w,{\Greekmath 0112} \right) -y$.}

In applications, identification of the models with EIV would typically rely on an instrumental variable $Z_{i}$. Suppose the instrument satisfies the exclusion restriction $\mathbb{E}\left[ u\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} \right) |X_{i}^{\ast },Z_{i}\right] =\mathbb{E}\left[ u\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} \right) |X_{i}^{\ast }\right] $, i.e., conditional on the true $X_{i}^{\ast }$ the instrument has no further effect on the moment conditions $u$. Consider the moment functions $h\left( x,s,{\Greekmath 0112} \right) \equiv u\left( x,s,{\Greekmath 0112} \right) \otimes {\Greekmath 0127} _{X}\left( x\right) $, where ${\Greekmath 0127} _{X}\left( x\right) $ is a vector of functions of $x$, e.g., $ {\Greekmath 0127} _{X}\left( x\right) \equiv \left( 1,x,\ldots ,x^{J}\right) ^{\prime } $. By the Law of Iterated Expectations, $\mathbb{E}\left[ h\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) |Z_{i}=z\right] =0$ for all $z$. However, the same expectation with $X_{i}^{\ast }$ replaced by the observed $ X_{i}$, $\mathbb{E}\left[ h\left( X_{i},S_{i},{\Greekmath 0112} _{0}\right) |Z_{i}=z \right] $, will depend on $z$. One way to see this is to consider a Taylor expansion similar to equation ((ref)):{

equation*[equation* omitted — 470 chars of source]

}which shows that $\mathbb{E}\left[ h\left( X_{i},S_{i},{\Greekmath 0112} _{0}\right) |Z_{i}=z\right] $ is zero for all $z$ (up to a negligible remainder) unless $ {\Greekmath 010D} _{0}\neq 0$, where ${\Greekmath 010D} _{0}={\Greekmath 011B} ^{2}/2$. Thus, $\mathbb{E} \left[ h\left( X_{i},S_{i},{\Greekmath 0112} _{0}\right) |Z_{i}=z\right] $ varies with $ z$ only because of the presence of the measurement error. Intuitively, the magnitude of this variation then identifies the nuisance parameters ${\Greekmath 010D} _{0}$. Thus, one can rely on the original moment functions of the typical form $g\left( x,s,{\Greekmath 0112} \right) =h\left( x,s,{\Greekmath 0112} \right) \otimes {\Greekmath 0127} _{Z}\left( z\right) $, where ${\Greekmath 0127} _{Z}\left( z\right) $ is a vector of functions of $z$.

The above discussion provides the intuition for identification of the nonlinear moment condition models with EIV. It is important to note that identification of general nonlinear moment condition models is a complicated problem. Even in the settings without measurement errors, it is generally not possible to give low-level conditions guaranteeing that a specific set of nonlinear moment conditions identifies the parameter vector. The presence of EIV\ makes the question of identification even harder.

We attempt to address this concern and make the above intuitions more precise in two ways. First, in the following subsection we consider a specific (but frequently employed) kind of an instrument: a second measurement (possibly non-classical). The specific form of the excluded variable allows us to provide more transparent identification conditions. Second, in EvdokimovZeleneev2022WP-NPID we study nonparametric regression model with EIV using the ${\Greekmath 011C} _{n}\rightarrow 0$ approximation. We show that the model is identified using an instrument (even a discrete one), and motivate the MERM approach from a nonparametric perspective. Finally, since MERM estimator is a standard GMM\ estimator, one can test the strength of identification of the model parameters, or conduct identification-robust inference using the standard methods (e.g., \QTR{citealp}{ StockWright2000Ecta,Kleibergen2005Ecta,GuggenbergerSmith2005ET,GuggenbergerRamalhoSmith2012JoE,AndrewsMikusheva2016Ecta-ConditionalFunctionalNuisance,AndrewsI-2016-Ecta-CLC,AndrewsGuggenberger2019QE }).

\paragraph{Identification Using A Second Measurement}

Suppose we observe a second measurement

equation*[equation* omitted — 84 chars of source]

where ${\Greekmath 010B} _{1}$ may not be known. Assume that ${\Greekmath 010B} _{1}\neq 0$ and $ \mathbb{E}\left[ {\Greekmath 0122} _{Q,i}|X_{i}^{\ast },S_{i},{\Greekmath 0122} _{i} \right] =0$. The variance of ${\Greekmath 0122} _{Q,i}$ does not need to be small. Note that the measurement error in $Q_{i}$ can be non-classical: $ Q_{i}-X_{i}^{\ast }$ and $X_{i}^{\ast }$ are correlated unless ${\Greekmath 010B} _{1}=1 $.

Consider the conditional moment restrictions ((ref)). If $ X_{i}^{\ast }$ were observed, we could have constructed the unconditional moments

equation*[equation* omitted — 311 chars of source]

for some $J\geq \dim \left( {\Greekmath 0112} \right) -1$. Suppose that the model is identified if $X_{i}^{\ast }$ observed, which means that the Jacobian of these moment conditions has full rank:

equation*[equation* omitted — 297 chars of source]

To deal with the error-laden $X_{i}$, consider the MERM estimator with $K=2$ based on the following moment function

equation[equation omitted — 267 chars of source]

Here the total number of moments is $m=2J+1$. The additional $J$ moments added in equation ((ref)) use $Q_{i}$, which will allow identifying ${\Greekmath 010D} _{0}=\mathbb{E}[{\Greekmath 0122} _{i}^{2}]/2$.

It turns out that in these settings there is a simple sufficient condition for Assumption (ref)((ref)) to hold. Appendix (ref) demonstrates that $\Psi ^{\ast }$ will have full rank if

equation[equation omitted — 250 chars of source]

Condition ((ref)) has a very simple interpretation: it essentially it means that $\mathbb{E}\left[ \left. u_{x}^{\left( 1\right) }\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) \right\vert X_{i}^{\ast } \right] $ should not be identically zero. For example, in the nonlinear regression model $u\left( x,y,{\Greekmath 0112} \right) ={\Greekmath 011A} \left( x,{\Greekmath 0112} \right) -y $, and condition ((ref)) is satisfied as long as $ \mathbb{E}\left[ {\Greekmath 011A} _{x}^{(1)}\left( X_{i}^{\ast },{\Greekmath 0112} _{0}\right) \left( X_{i}^{\ast }\right) ^{j}\right] \neq 0$ for some $j\in \left\{ 0,\ldots ,J-1\right\} $.

Numerical Evidence

Comparison with a Semi-Nonparametric Estimation Approach

We compare MERM estimator with the state-of-the-art semiparametric estimator of Schennach2007Ecta for nonlinear regression models. The Monte Carlo designs are taken from S07, and include a polynomial, rational fraction, and Probit nonlinear regression models. Identification of the model is ensured by the availability of an instrument.

equation[equation omitted — 182 chars of source]

$(Z_i,V_i,{\Greekmath 0122}_i)' \sim N \left((0, 0, 0)', \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Diag}(1, 1/4, 1/4)\right)$, ${\Greekmath 0119}_1=1$, and $n = 1000$. The conditional expectation function ${\Greekmath 011A}$, the true value of the parameter of interest ${\Greekmath 0112}_0$, and the conditional distribution of the regression error $U_i$ are design-specific and reported in Tables (ref)-(ref) below. In all designs, ${\Greekmath 011C} = {\Greekmath 011B}_{{\Greekmath 0122}}/{\Greekmath 011B}_{X}^* \approx 0.45$, so the measurement error is “fairly large” Schennach2007Ecta.

We report simulation results for the MERM estimator considering correction schemes with $K=2$ and $K=4$. The original moment function is

equation*[equation* omitted — 118 chars of source]

where we use ${\Greekmath 0127}(x,z) = \left(1, x, z, x^2, z^2, x^3, z^3\right)'$ for $K=2$ and ${\Greekmath 0127}(x,z) = \left(1, x, z, x^2, xz, z^2, x^3, x^2 z, x z^2, z^3\right)'$ for $K=4$.

The finite sample properties of the MERM estimators (evaluated based on 5,000 replications) are reported in Tables (ref)-(ref) below. For comparison, we also provide the same statistics for naive estimators (OLS/NLLS) and for the benchmark estimator of S07 (as reported in the original paper). For the polynomial model (Table (ref)), both $K=2$ and $K=4$ MERM estimators effectively remove the EIV bias. Component-wise, the MERM estimators perform similarly (for ${\Greekmath 0112}_2$ and ${\Greekmath 0112}_4$) or better (for ${\Greekmath 0112}_1$ and ${\Greekmath 0112}_3$) compared to the benchmark estimator of S07. For the rational fraction model (Table (ref)), both the MERM estimators are vastly superior to the benchmark estimator both in terms of the bias and the standard deviation. For the probit model (Table (ref)), the MERM estimator with $K=2$ removes a large fraction of the EIV bias compared to the NLLS estimator. However, the EIV bias remains non-negligible when this simplest correction scheme is used. Employing a higher order correction scheme with $K=4$ completely eliminates the remaining EIV bias, while at the same time having smaller standard deviations (than the benchmark estimator of S07) . Overall, in the considered designs, the MERM estimator with $K=4$ consistently outperforms the benchmark estimator. It also proves to be more effective in removing the EIV bias compared to the $K=2$ estimator, especially in the highly nonlinear settings of the considered probit design.

table[table omitted — 1,583 chars of source]
table[table omitted — 1,464 chars of source]
table[table omitted — 1,445 chars of source]

Estimation and Inference in a Multinomial Choice Model

Consider the standard multinomial logit model, in which an agent chooses between 3 available options. For an agent $i$ with characteristics $(X_i^*,W_i)$, the utility of option $j$ is given by

equation*[equation* omitted — 234 chars of source]

and $U_{i0} = {\Greekmath 010F}_{i0}$ for the outside option $j = 0$, where ${\Greekmath 010F}_{ij}$ are i.i.d. (across $i$ and $j$) draws from a standard type-1 extreme value distribution. The researcher observes $\left\{(X_i,W_i,Y_{i1},Y_{i2},Y_{i0})\right\}_{i=1}^n$, where $Y_{ij}$ is a binary variable indicating whether agent $i$ chooses option $j$, i.e. $Y_{ij} = 1$ if and only if $j = \operatorname*{\mathrm{arg}\!\max\limits}_{j' \in \{0,1,2\}} U_{i j'}$. In addition,

equation*[equation* omitted — 209 chars of source]

and $\left(V_{i1}, V_{i0}, Z_i,{\Greekmath 0122}_i,{\Greekmath 0117}_{i1},{\Greekmath 0117}_{i2}\right)' \sim N\left((1,0,0,0,0,0)', \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Diag}({\Greekmath 011B}_{V1}^2,{\Greekmath 011B}_{V0}^2, {\Greekmath 011B}_Z^2, {\Greekmath 011B}_{\Greekmath 0122}^2, {\Greekmath 011B}_{\Greekmath 0117}^2, {\Greekmath 011B}_{\Greekmath 0117}^2)\right)$. In all of the designs, we fix $({\Greekmath 0112}_{011},{\Greekmath 0112}_{012},{\Greekmath 0112}_{013},{\Greekmath 0112}_{021},{\Greekmath 0112}_{022},{\Greekmath 0112}_{023},{\Greekmath 011A},{\Greekmath 011B}_{V1}^2,{\Greekmath 011B}_{V0}^2,{\Greekmath 011B}_Z^2,{\Greekmath 011B}_{\Greekmath 0117}^2) = (1,0,0,0,0,0,0.7,1/2,1/2,1,1)$ and $n=2000$. We consider ${\Greekmath 011C} = {\Greekmath 011B}_{{\Greekmath 0122}}/{\Greekmath 011B}_{X^*} \in \{1/4, 1/2, 3/4\}$. Setting ${\Greekmath 011B}_{V1}=0$ would correspond to the additive control variable model. We omit such simulation results for brevity.

Similarly to Section (ref), we report results for the MERM estimators with $K=2$ and $K=4$ based on the following original moment function

align*[align* omitted — 534 chars of source]

where ${\Greekmath 0127}_j(x,z,w) = \left(1, x, z, x^2, z^2, x^3, z^3, w_j\right)'$ for $K=2$ and ${\Greekmath 0127}_j(x,z,w) = \left(1, x, z, x^2, xz, z^2, x^3, x^2 z, x z^2, z^3, w_j\right)'$ for $K=4$.

We report the results on estimation and inference on the partial derivatives of the conditional choice probabilities $p_j(x,w_1,w_2)$ with respect to $x$, $w_1$, and $w_2$, evaluated at the population means.

Table (ref) reports the finite sample biases, standard deviations, and RMSE of the MERM estimators, as well as the sizes of the corresponding t-tests with nominal size of 5%. To illustrate the importance of dealing with EIV, we also report the same statistics for the standard (naive) MLE estimator that ignores the presence of the measurement errors.

In all designs, the MLE estimator is biased, and the corresponding t-tests over-reject. Note that failing to account for the EIV in the mismeasured variable $X_i^*$ generally biases estimators of all of the parameter, including those corresponding to the correctly measured variables $W_{i1}$ and $W_{i2}$. In particular, the t-tests may falsely reject true null hypotheses $\partial p_j/\partial w_\ell=0$ up to nearly $100\%$ of the time.

The MERM estimator with $K=2$ removes a large fraction of the EIV bias in all of the designs. While this proves to be enough to achieve accurate size control when the magnitude of the measurement error is moderate (${\Greekmath 011C} = 1/4$), the remaining EIV bias may still result in size distortions of the t-tests with larger measurement errors, especially ${\Greekmath 011C} = 3/4$. Using the higher order correction scheme with $K=4$ effectively removes the EIV bias in all of the simulation designs for all of the parameters. Remarkably, the corresponding finite sample null rejection probabilities remain close to the nominal $5\%$ rate even when the standard deviation of the measurement error is as large as $75\%$ of the standard deviation of the mismeasured $X^*$.

To further check the limits of applicability of our method, we also consider larger values of ${\Greekmath 011C} \in \{1,3/2,2\}$. The numerical results analogous to the ones reported in Table (ref) are provided in Table (ref) in Appendix (ref). Specifically, we find that inference results based on the correction scheme with $K=4$ remain accurate even for ${\Greekmath 011C} = 1$. Unsurprisingly, inference becomes less reliable for bigger ${\Greekmath 011C} = 3/2$ and ${\Greekmath 011C} = 2$. At the same time, while the $K=4$ correction scheme fails to entirely eliminate the EIV bias in these designs, it still removes a big fraction of the bias and greatly improves on MLE in terms of the RMSE. Thus, while inference based on our estimator might be less reliable in extreme settings when the measurement error overwhelms the signal, the MERM estimator using $K=4$ appears to be sufficiently accurate over a wide range of ${\Greekmath 011C}$.

\afterpage{

landscape\begin{table}[h!] \begin{footnotesize} \begin{threeparttable} \caption{Simulation results for the multinomial logit model} \begin{tabular}{ l c c c c| c c c c | c c c c} \toprule \toprule & \multicolumn{4}{c}{MLE} & \multicolumn{4}{c}{$K=2$} & \multicolumn{4}{c}{$K=4$}\\ \cmidrule(lr){2-5} \cmidrule(lr){6-9} \cmidrule(lr){10-13} &{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}, $10^{-2}$}&{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{std}, $10^{-2}$}&{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}, $10^{-2}$}&{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{size}}&{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}, $10^{-2}$}&{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{std}, $10^{-2}$}&{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}, $10^{-2}$}&{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{size}}&{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}, $10^{-2}$}&{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{std}, $10^{-2}$}&{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}, $10^{-2}$}&{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{size}}\\ \midrule \multicolumn{13}{c}{${\Greekmath 011C} = 1/4$}\\ \midrule $\partial p_1/\partial x$&-3.24&1.36&3.51&66.98&0.74&2.63&2.74&4.30&1.13&2.73&2.95&7.86\\ $\partial p_1/\partial w_1$&2.32&1.64&2.84&30.74&-0.11&2.30&2.30&4.82&-0.31&2.29&2.31&6.54\\ $\partial p_1/\partial w_2$&0.48&0.75&0.90&9.40&-0.04&0.87&0.87&4.82&-0.08&0.87&0.87&5.36\\ $\partial p_2/\partial x$&1.96&1.17&2.28&39.44&-0.40&1.88&1.92&4.72&-0.63&1.93&2.03&6.74\\ $\partial p_2/\partial w_1$&-1.16&0.82&1.42&30.66&0.06&1.15&1.15&4.84&0.15&1.15&1.16&6.48\\ $\partial p_2/\partial w_2$&-0.96&1.50&1.78&9.48&0.09&1.74&1.74&4.98&0.17&1.74&1.75&5.44\\ $\partial p_0/\partial x$&1.28&1.01&1.63&25.28&-0.34&1.43&1.47&5.08&-0.50&1.48&1.56&7.36\\ $\partial p_0/\partial w_1$&-1.16&0.82&1.42&30.60&0.06&1.15&1.15&4.82&0.15&1.15&1.16&6.46\\ $\partial p_0/\partial w_2$&0.48&0.74&0.88&9.46&-0.05&0.87&0.87&4.90&-0.09&0.87&0.88&5.32\\ \midrule \multicolumn{13}{c}{${\Greekmath 011C} = 1/2$}\\ \midrule $\partial p_1/\partial x$&-8.97&1.09&9.04&100.00&-1.69&2.60&3.10&9.84&0.97&2.89&3.05&6.04\\ $\partial p_1/\partial w_1$&6.44&1.53&6.62&98.96&1.39&2.44&2.81&12.14&-0.21&2.41&2.42&5.54\\ $\partial p_1/\partial w_2$&1.28&0.72&1.47&42.54&0.29&0.90&0.94&7.00&-0.06&0.92&0.92&5.00\\ $\partial p_2/\partial x$&5.22&0.97&5.31&99.98&1.05&1.89&2.16&10.16&-0.53&2.08&2.15&6.14\\ $\partial p_2/\partial w_1$&-3.21&0.77&3.30&98.96&-0.69&1.22&1.40&12.10&0.10&1.21&1.21&5.52\\ $\partial p_2/\partial w_2$&-2.52&1.41&2.88&42.82&-0.56&1.79&1.87&7.20&0.13&1.84&1.84&4.98\\ $\partial p_0/\partial x$&3.75&0.86&3.85&98.78&0.64&1.45&1.59&8.18&-0.44&1.59&1.65&6.50\\ $\partial p_0/\partial w_1$&-3.23&0.78&3.32&98.96&-0.70&1.22&1.41&12.06&0.10&1.21&1.21&5.48\\ $\partial p_0/\partial w_2$&1.23&0.69&1.41&42.90&0.28&0.89&0.93&7.20&-0.07&0.92&0.92&4.94\\ \midrule \multicolumn{13}{c}{${\Greekmath 011C} = 3/4$}\\ \midrule $\partial p_1/\partial x$&-13.35&0.86&13.38&100.00&-6.83&2.64&7.32&80.32&0.71&3.22&3.29&4.74\\ $\partial p_1/\partial w_1$&9.69&1.45&9.80&100.00&4.95&2.65&5.61&65.52&0.01&2.62&2.62&5.34\\ $\partial p_1/\partial w_2$&1.81&0.69&1.94&75.30&1.01&0.89&1.35&26.08&-0.01&0.98&0.98&5.24\\ $\partial p_2/\partial x$&7.48&0.79&7.52&100.00&4.06&1.83&4.45&68.82&-0.37&2.32&2.35&5.74\\ $\partial p_2/\partial w_1$&-4.83&0.73&4.88&100.00&-2.47&1.32&2.81&65.46&-0.01&1.31&1.31&5.28\\ $\partial p_2/\partial w_2$&-3.51&1.33&3.76&75.60&-1.99&1.76&2.66&26.38&0.03&1.97&1.97&5.32\\ $\partial p_0/\partial x$&5.87&0.73&5.92&100.00&2.77&1.47&3.14&56.28&-0.34&1.77&1.80&5.82\\ $\partial p_0/\partial w_1$&-4.87&0.75&4.93&100.00&-2.48&1.33&2.81&65.40&-0.01&1.31&1.31&5.32\\ $\partial p_0/\partial w_2$&1.70&0.64&1.82&75.76&0.98&0.87&1.31&26.50&-0.02&0.99&0.99&5.30\\ \bottomrule \end{tabular} \begin{tablenotes} • This table reports the simulated finite sample bias, standard deviation, RMSE, and size of the MLE and the MERM estimators and the corresponding t-tests for the partial derivatives $\partial p_j(x,w,{\Greekmath 0112}_0)/\partial x$, $\partial p_j(x,w,{\Greekmath 0112}_0)/\partial w_1$, $\partial p_j(x,w,{\Greekmath 0112}_0)/\partial w_2$ for $j \in \{1,2,0\}$ evaluated at the population mean. The true values of the marginal effects are $(\partial p_1/\partial x, \partial p_2/\partial x, \partial p_0/\partial x) = ( 0.222, -0.111, -0.111)$ and zeros for the rest. The results are based on 5,000 replications. \end{tablenotes} \end{threeparttable} \end{footnotesize} \end{table}

}

Empirical Illustration: Choice of Transportation Mode

In this section, we illustrate the finite sample properties of the MERM estimator in the context of a classical multinomial choice application: choice of transportation mode (e.g., McFadden1974JPubE).

To calibrate the numerical experiment, we use the ModeCanada dataset, a survey of business travelers for the Montreal-Toronto corridor. We focus on the subset of travelers choosing between train, air, and car ($n = 2769$), and estimate the conditional logit model with traveler $i$'s utilities given in the table below.

center[center omitted — 764 chars of source]

To generate the simulated samples, we randomly draw covariates from their joint empirical distribution. To generate the simulated outcomes, we draw ${\Greekmath 010F}_{ij}$ from the standard type-I extreme value distribution. The true value of ${\Greekmath 0112}_0$ is set to be the MLE estimate based on the original dataset. More details about this numerical experiment are given in Appendix (ref).

To evaluate the performance of the MERM estimator in these settings, we generate mismeasured $Income_i = Income_i^* + {\Greekmath 0122}_i$. We focus on the individual income because it is often mismeasured. We report the results for ${\Greekmath 011C} = {\Greekmath 011B}_{\Greekmath 0122}/{\Greekmath 011B}_{Income^*} \in \{1/4, 1/2, 3/4\}$.

Table (ref) reports the simulation results for the (naive) MLE estimator and for the MERM estimators with $K=2$ and $K=4$. We focus on estimation of and inference on the income elasticities (evaluated at the population mean of the covariates). The MLE estimator is considerably biased for ${\Greekmath 011C} \in \{1/2, 3/4\}$, which results in substantial size distortions of the MLE based t-tests. The MERM estimator with $K=4$ effectively eliminates the EIV bias and the corresponding t-tests provide accurate size control in all of the considered designs. The estimator with $K=2$ is more precise, while successfully removing the EIV bias for ${\Greekmath 011C}\le1/2$.

Overall, the MERM estimators perform well in the considered empirical context, providing a basis for estimation and inference even for quite large values of ${\Greekmath 011C}$.

table[table omitted — 3,120 chars of source]

Extensions

Data-driven choice of $K$

Making an appropriate choice of the expansion order $K$ is important for the estimation procedure. One has to be cautious not to take $K$ too small, as this may result in an estimator that only partially removes the EIV bias. On the other hand, picking a larger $K$ than needed might inflate standard errors and result in less powerful inference.

In this section, we address this issue by providing a data-dependent procedure for selecting $K$. We demonstrate that our procedure has desirable theoretical properties. We also find that the procedure has good finite sample properties in a set of Monte Carlo simulation experiments across different values of ${\Greekmath 011C}$ and sample sizes.

Consider two alternative values of the expansion order: $L$ and $K$, where $2 \leq L<K$. In practice, even-order biases tend to dominate, so to reduce the set of choices, it is useful to focus on even values of the expansion orders $L$ and $K$. For example, one would typically be interested in choosing between $L=2$ and $K=4$.

Let $\hat {\Greekmath 010C}_L$ and $\hat {\Greekmath 010C}_K$ denote the corresponding MERM estimators. The estimator $\hat {\Greekmath 010C}_L$ should be preferred as having smaller asymptotic variance provided that its remaining EIV bias is negligible relative to its standard error. Otherwise, the more conservative $\hat {\Greekmath 010C}_K$ should be used instead.

Note that Lemma (ref) suggests that the remaining asymptotic bias of $\hat {\Greekmath 010C}_L$ (due to the additional terms accounted for when the expansion of higher order $K$ is used) is given by (up to an $o(n^{-1/2})$ remainder)

equation*[equation* omitted — 306 chars of source]

where matrix $B$ is based on the moments used for estimation of $\hat {\Greekmath 010C}_L$.

Importantly, ${\Greekmath 010D} _{0k}=O\left( {\Greekmath 011B}_{{\Greekmath 0122}}^{k}\right) $, and hence for $k>2$ we can estimate a bound on $\sqrt{n}{\Greekmath 010D} _{0k}$ sufficiently quickly to provide a valid procedure for choosing $K$. To this end, we first estimate the model using the larger $K$. Let $\hat {\Greekmath 010C}_K = (\hat {\Greekmath 0112}', \hat {\Greekmath 010D}')'$ and $\hat {\Greekmath 011B}_{\Greekmath 0122}^2 \equiv 2 \hat {\Greekmath 010D}_{2}$, where we dropped the additional subscripts $K$ for notation simplicity. Then, we can estimate ${\Greekmath 011B}_{\Greekmath 0122}^K$ by $\hat {\Greekmath 011B}_{\Greekmath 0122}^K \equiv (\hat {\Greekmath 011B}_{\Greekmath 0122}^2)^{K/2}$. Using $\hat {\Greekmath 011B}_{\Greekmath 0122}^2 = {\Greekmath 011B}_{\Greekmath 0122}^2 + O_p (n^{-1/2})$, in the appendix we show that

equation[equation omitted — 265 chars of source]

Next, consider a sequence $\varkappa _{n}\rightarrow 0$, which we will specify precisely later, and let

equation[equation omitted — 352 chars of source]

where $\hat \Sigma$ and $\hat B$ are consistent estimators of the asymptotic variance of $\hat {\Greekmath 010C}_L$ and of $B$, with $\hat \Sigma_{\ell \ell}$ denoting its $\ell$-th diagonal element, $\hat B_{\ell \cdot}$ denoting its $\ell$-th row, ${\overline g_x^{(K)} (\hat {\Greekmath 0112}) \equiv n^{-1} \sum_{i=1}^n g^{(K)}(X_i,S_i,\hat {\Greekmath 0112})}$, and $c_K > 0$ is a constant that we will calibrate below after we state the main theoretical result of this section.

If ${\Greekmath 010E}_n = 1$, the researchers should select $\hat {\Greekmath 010C}_L$. Otherwise, $\hat {\Greekmath 010C}_K$ should be used. The following lemma demonstrates that the proposed selection procedure has desirable theoretical properties.

lemmaSuppose the hypotheses of Theorem 2 hold for some even $ K\geq 4$, and consider $L$ satisfying $2 \leq L<K$. Suppose $\varkappa _{n}n^{\left( K-L-1\right) /\left( 2L+2\right) }\rightarrow 0$. Then $\sqrt{n}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{AsyB}(\hat {\Greekmath 010C}_L) {\Greekmath 010E} _{n}=o_{p}\left(1\right)$. Moreover, consider $\varkappa _{n} = n^{-\left( K-L-1\right) /\left( 2L+2\right) }\left( \ln n\right) ^{-a}$ for any $a>0$. Then the criterion is consistent, in the sense that if ${\Greekmath 011C}_n =o( n^{-\frac{1}{2L+2} -{\Greekmath 010F} }) $ for any ${\Greekmath 010F} >0$, the criterion will choose $\hat {\Greekmath 010C}_L$ with probability approaching one.

The first part of Lemma (ref) shows that, provided that $\varkappa_n$ goes to zero sufficiently fast, $\hat {\Greekmath 010C}_L$ is selected (i.e., ${\Greekmath 010E}_n = 1$) only when its asymptotic bias is negligible. Next, note that $\hat {\Greekmath 010C}_L$ is asymptotically unbiased as long as ${\Greekmath 011C}_n = ( n^{-\frac{1}{2L+2}})$. The second part of the lemma shows that, for the suggested choices of $\varkappa_n$, the criterion is non-vacuous, i.e., that it does select the MERM estimator with a smaller expansion order $L$ when this is appropriate. Typically, one would pick $L=K-2$, so the lemma suggests taking $\varkappa _{n} = n^{-1/\left( 2K-2\right) }\left( \ln n\right) ^{-a}$ in this case.

In practice, it is important to pick an appropriate constant $c_K$ used in the construction of ${\Greekmath 010E}_n$ in equation (ref). When constructing ${\Greekmath 010E}_n$, we used $\left\vert {\Greekmath 010D}_{0K}\right\vert \propto {\Greekmath 011B}_{\Greekmath 0122}^K$, which motivates choosing $c_K = {\Greekmath 011B}_{\Greekmath 0122}^K / \left\vert {\Greekmath 010D}_{0K}\right\vert$. It is convenient to use a rule-of-thumb approach, using a reference distribution to determine $c_K$. It turns out that the normal distribution is not only convenient, but also sufficiently conservative (note that the bigger ${\Greekmath 010D}_{0K}$ is, the smaller $c_K$ is, resulting in a more conservative selection procedure). For example, suppose $K=4$, and consider using Student's $t({\Greekmath 0117} )$ distribution as the reference distribution. Then, for ${\Greekmath 0117} \geq 5$, the biggest $\left\vert {\Greekmath 010D}_{04}\right\vert/{\Greekmath 011B}_{\Greekmath 0122}^4$ and the smallest $c_4$ correspond to ${\Greekmath 0117} = \infty$ matching the normal distribution.\footnote{Note that Student's $t({\Greekmath 0117} )$ distribution does not have a finite $5$-th moment ${\Greekmath 0117} \leq 5$.}

Since for the normal distribution we have ${\Greekmath 010D}_{0K} = {\Greekmath 011B}_{\Greekmath 0122}^K / K!!$ for even $K$, we recommend using $c_K = K!!$ in equation (ref). Finally, while we recommend using normal distribution as the reference distribution, we also stress that the results of Lemma (ref) hold even if the measurement error is not normal or if it is skewed.

We summarize the proposed procedure in the following suggested algorithm.

algorithm[algorithm omitted — 635 chars of source]

To illustrate the performance of the algorithm provided above, we revisit the numerical experiment considered in Section (ref). We focus on choosing between $K=4$ and $L=2$, and, as in Section (ref), we report results for $n=2000$ in Table (ref) below. To ensure that the proposed algorithm performs well in a variety of sample sizes, we also consider $n=1000$ and $n=4000$ and report the corresponding results in Tables (ref) and (ref) in Appendix (ref).

We report the finite sample bias and RMSE for the naive MLE estimator, as well as for the MERM estimators using $K=2$ and $K=4$, and for the adaptive MERM estimator using data-driven $K$ following Algorithm (ref). Notice that for the considered sample sizes, $K=2$ is preferred when ${\Greekmath 011C} = 1/4$, and $K=4$ is preferred when ${\Greekmath 011C} = 3/4$. We find that in both of these regimes and for all sample sizes, the adaptive estimator using data-driven $K$ is essentially equivalent to the preferred estimators, i.e., our procedure selects the appropriate $K$. Interestingly, in the intermediate regime with ${\Greekmath 011C} = 1/2$, the adaptive estimator also has the smallest bias among the considered estimators, coming at the cost of a slightly bigger RMSE compared to the MERM estimator using $K=4$. Thus, the considered numerical experiment suggests that our algorithm has good finite sample properties supporting the findings of Lemma (ref).

table[table omitted — 4,184 chars of source]

Multiple Mismeasured Variables

It is easy to use the MERM framework to deal with multiple mismeasured variables. This is useful in many applications, including not only settings with multiple mismeasured covariates, but also settings with serially correlated measurement errors, settings where repeated measurements are available, and panel data models. Using the MERM approach is particularly advantageous in such applications, since it avoids nonparametric estimation of multivariate unobserved distributions.

Suppose $X_i^*$, ${\Greekmath 0122}_{i}$, and $X_i$ are $d \times 1$ vectors. Let ${\Greekmath 011C}_n \equiv \max_{j\le d} {\Greekmath 011B}_{{\Greekmath 0122}_j}/{\Greekmath 011B}_{X^*_j}$, where ${\Greekmath 011B}_{{\Greekmath 0122}_j}$ and ${\Greekmath 011B}_{X^*_j}$ denote the standard deviations of the $j$-th components of ${\Greekmath 0122}_i$ and $X_i^*$, so $\mathbb{E}\left[\left\vert {\Greekmath 0122}_{ij}\right\vert^k\right] = O({\Greekmath 011C}_n^k)$ for $k\in\{1,\ldots,K\}$.

For a $d \times 1$ vector of non-negative integers ${\Greekmath 0114} = ({\Greekmath 0114}_1,\ldots,{\Greekmath 0114}_d) \in \mathbb Z_+^{d}$, let

equation*[equation* omitted — 357 chars of source]

Also, for a positive integer $k$, let $\mathcal K_k = \{{\Greekmath 0114} \in \mathbb Z_+^d: \left\vert {\Greekmath 0114}\right\vert = k\}$. Then, we consider the following corrected moment function

equation*[equation* omitted — 259 chars of source]

where, with some abuse of notation, ${\Greekmath 010D}$ is a collection of all ${\Greekmath 010D}_{\Greekmath 0114}$ with ${\Greekmath 0114} \in \mathcal K_k$ and $k \in \{2, \dots, K\}$.

Under mild smoothness conditions

equation*[equation* omitted — 216 chars of source]

where the second equality holds provided that $O({\Greekmath 011C}_n^{K+1}) = o(n^{-1/2})$. Similarly to the scalar case, components of ${\Greekmath 010D}_{0}$ are determined by the moments of ${\Greekmath 0122}_i$. Specifically, let ${\Greekmath 0116}_{{\Greekmath 0114}} \equiv \mathbb{E}\left[{\Greekmath 0122}_{i 1}^{{\Greekmath 0114}_1} \dots {\Greekmath 0122}_{i d}^{{\Greekmath 0114}_d}\right]$, then

equation[equation omitted — 289 chars of source]

where ${\Greekmath 0114}! \equiv {\Greekmath 0114}_1! \ldots {\Greekmath 0114}_d!$. For $\left\vert {\Greekmath 0114}\right\vert \geq 4$, the coefficients can be computed by the following formulas. For example, for ${\Greekmath 0114} \in \mathcal K_4$, let $\mathcal K_{2, {\Greekmath 0114}} = \{\tilde {\Greekmath 0114} \in \mathcal K_2: {\Greekmath 0114} - \tilde {\Greekmath 0114} \in \mathcal K_2 \}$. Then,

equation*[equation* omitted — 483 chars of source]

More generally, for ${\Greekmath 0114} \in \mathcal K_{k}$ with $k \geq 4$, let $\mathcal K_{\ell, {\Greekmath 0114}} = \{\tilde {\Greekmath 0114} \in \mathcal K_\ell, {\Greekmath 0114} - \tilde {\Greekmath 0114} \in \mathcal K_{\left\vert {\Greekmath 0114}\right\vert - \ell} \}$ for $\ell \leq \left\vert {\Greekmath 0114}\right\vert - 2$. Then,

equation*[equation* omitted — 390 chars of source]
example*[Bivariate $X$, $K = 4$] \\ Suppose $X$ is bivariate (i.e., $d = 2$) and $K = 4$. For ${\Greekmath 0114} \in \mathcal K_2 = \{(2,0),(1,1),(0,2)\}$ and ${\Greekmath 0114} \in \mathcal K_3 = \{(3,0), (2,1), (1,2), (0,3)\}$, ${\Greekmath 010D}_{0 {\Greekmath 0114}}$ is given by (ref). For ${\Greekmath 0114} \in \mathcal K_4$, ${\Greekmath 010D}_{0 {\Greekmath 0114}}$ is given by \begin{center} \begin{tabular}{c|c} ${\Greekmath 0114}$ & ${\Greekmath 010D}_{0 {\Greekmath 0114}}$ \\ \midrule (4,0) & $\left(\mathbb{E}[{{\Greekmath 0122}_{i1}^4}] - 6 \mathbb{E}[{{\Greekmath 0122}_{i1}^2}]^2\right)/24$ \\ (3,1) & $\left(\mathbb{E}[{{\Greekmath 0122}_{i1}^3 {\Greekmath 0122}_{i2}}] - 6 \mathbb{E}[{{\Greekmath 0122}_{i1}^2}] \mathbb{E}[{{\Greekmath 0122}_{i1} {\Greekmath 0122}_{i2}}]\right)/6$ \\ (2,2) & $\left(\mathbb{E}[{{\Greekmath 0122}_{i1}^2 {\Greekmath 0122}_{i2}^2}] - 2 \mathbb{E}[{{\Greekmath 0122}_{i1}^2}] \mathbb{E}[{{\Greekmath 0122}_{i2}^2}] - 4 \mathbb{E}[{{\Greekmath 0122}_{i1} {\Greekmath 0122}_{i2}}]^2\right)/4$ \\ (1,3) & $\left(\mathbb{E}[{{\Greekmath 0122}_{i1} {\Greekmath 0122}_{i2}^3}] - 6 \mathbb{E}[{{\Greekmath 0122}_{i2}^2}] \mathbb{E}[{{\Greekmath 0122}_{i1} {\Greekmath 0122}_{i2}}]\right)/6$ \\ (0,4) & $\left(\mathbb{E}[{{\Greekmath 0122}_{i2}^4}] - 6 \mathbb{E}[{{\Greekmath 0122}_{i2}^2}]^2\right)/24$ \end{tabular} \end{center} If in addition measurement errors ${\Greekmath 0122}_{i1}$ and ${\Greekmath 0122}_{i2}$ are independent, ${\Greekmath 010D}_{0 {\Greekmath 0114}} = 0$ for ${\Greekmath 0114} \in \{(1,1),(2,1),(1,2),(3,1),(1,3)\}$. In this case, the total number of the nuisance parameters to be estimated is 6. \ensuremath{\blacksquare}
singlespacing