Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,930 characters · 13 sections · 36 citation commands
-1inSimple Estimation of Semiparametric Models with Measurement Errors
Measurement errors are a common problem for empirical studies. Addressing the Errors-In-Variables (EIV) bias in nonlinear models requires elaborate strategies.\footnote{ See HINP1991JoE,HausmanNeweyPowell1995JoE,Newey2001REStat,Schennach2007Ecta,Li2002JoE,Schennach2004Ecta,ChenHongTamer2005ReStud,HuSchennach2008Ecta,Schennach2014Ecta-ELVIS,Wilhelm2019WP-TestingForME , among others.} Despite the fundamental theoretical progress in identification and estimation of nonlinear models with EIV, the problem of EIV is still rarely addressed in empirical work outside of linear specifications.
{8.0pt plus 2.0pt minus 7.0pt} {6.0pt plus 2.0pt minus 5.0pt}
The goal of this paper is to develop a simple and practical approach to estimation of nonlinear semiparametric models that can be expressed in the form of general moment conditions
where $g\left( \cdot \right) $ is a vector of moment functions and ${\Greekmath 0112} _{0}$ is the parameter vector of interest. The researcher has a random sample of $\left\{ X_{i},S_{i}\right\} _{i=1}^{n}$, where scalar or vector $ X_{i}$ is a mismeasured version of unobserved $X_{i}^{\ast }$ with measurement error ${\Greekmath 0122} _{i}$:
We will refer to $ g\left( \cdot \right) $ as the original moment function, since it would have been valid had the researcher observed $X_{i}^{\ast }$. A naive GMM estimator (that ignores the EIV and uses $X_{i}$ in place of $ X_{i}^{\ast }$) based on $g\left( \cdot \right) $ is biased because $\mathbb{ E}[g(X_{i},S_{i},{\Greekmath 0112} _{0})]\neq 0$, in contrast to equation ((ref)).
Even in this well-studied example of nonlinear regression, estimation in the presence of the EIV is a difficult problem. Importantly, nonlinear instrumental variable regression estimator cannot be used, since it is inconsistent in the presence of EIV Amemiya1985JoE. The existing approaches typically require nonparametric estimation that can be impractical in many empirical applications. In contrast, in this paper, we develop an alternative class of estimators, that are essentially GMM estimators that modify the original moment functions $g\left( \cdot \right) $ in a way that makes the moment conditions robust to the EIV. In particular, our approach makes it easy to use instrumental variables to address EIV\ in nonlinear models.
To provide a practical estimation approach\ for the general class of models ( (ref)), we focus on empirical settings in which the researcher believes the variability of the measurement error to be at most a fraction of the variability of the mismeasured variable, i.e., the noise-to-signal ratio ${\Greekmath 011C} \equiv {\Greekmath 011B} _{{\Greekmath 0122} }/{\Greekmath 011B} _{X^{\ast }} $ to be moderate. The absolute magnitude of the measurement error ${\Greekmath 011B} _{{\Greekmath 0122} }$ does not need to be small. Existing validation studies provide insights into the magnitude of ${\Greekmath 011C} $ for some key economic variables and datasets. BoundKrueger1991JoLaborEcon consider log-earnings in the Current Population Survey (CPS) data matched to the Social Security payroll records. Their estimates of the variance of the measurement errors correspond to $ {\Greekmath 011C} $ of approximately $0.47$ and $0.30$ for the subsamples of men and women, respectively. BoundBrownDuncanRodgers1994JoLaborEcon consider a validation study of Panel Study of Income Dynamics (PSID). Their estimates imply ${\Greekmath 011C} =0.39-0.66$ for log-earnings and ${\Greekmath 011C} =0.63-0.76$ for hours worked. Pischke1995JBES estimates correspond to ${\Greekmath 011C} =0.39-0.50$ for the log-earnings in PSID. AshenfelterKrueger1994AER assess the mismeasurement in the years of education; their estimates correspond to $ {\Greekmath 011C} =0.30-0.37$.
Focusing on these settings allows us to isolate the most important aspects of the problem and, as result, to develop a simple estimator, which does not require any nonparametric estimation or simulation. Such simple estimation becomes possible because in these settings we can obtain a simple approximation of the EIV\ bias of the moment conditions as a function of $ {\Greekmath 0112} $.
We propose to bias correct the original moments $g\left( \cdot \right) $, which in turn removes the bias of the corresponding estimator of ${\Greekmath 0112} _{0} $. This bias correction depends on some moments of the distribution of the measurement errors that are unknown. Another difficulty is that the estimators of some components of the bias correction themselves may need to be bias corrected. To address these issues, we develop the corrected moment conditions, which depend on ${\Greekmath 0112} $ and additional parameters ${\Greekmath 010D} $ that govern the bias correction. The true parameter value ${\Greekmath 010D} _{0}$ is associated with (possibly conditional) low-order moments of ${\Greekmath 0122} _{i}$. Despite some theoretical subtleties with the construction of the corrected moment conditions, their practical implementation is straightforward and they can be automatically computed for any original moment function $g\left( \cdot \right) $.
We introduce the Measurement Error Robust Moments (MERM) estimator, which is a GMM estimator that uses the corrected moment conditions to jointly estimate parameters ${\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{0}$. The estimator can be computed using any standard software for GMM estimation. Joint estimation of parameters ${\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{0}$ using the corrected moment conditions effectively robustifies moment conditions $g\left( \cdot \right) $ against the impact of the measurement errors.
To make these ideas precise and to study the properties of the proposed estimators, we develop an asymptotic theory using a nonstandard asymptotic approximation that models ${\Greekmath 011C} $ as slowly shrinking with the sample size. Standard asymptotics considers ${\Greekmath 011C} $ to be constant, which implies that as $n\rightarrow \infty $ the bias of a naive estimator dwarfs its sampling variability: the bias is constant while the standard errors shrink proportionally to $1/\sqrt{n}$. As a result, under the standard asymptotics, the problem of removing the EIV bias becomes central in the analysis, with relatively little attention paid to the sampling variability of estimators. However, this focus does not seem to be appropriate in many empirical applications, in which the researcher does not expect the potential EIV bias to be several orders of magnitude larger than the standard errors.\footnote{ Such empirical settings appear to be widespread. Although the concerns about measurement errors are often raised, the majority of applied work does not explicitly correct the EIV bias in nonlinear models, and instead implicitly or explicitly argues or conjectures that the EIV bias is likely not to be too large. } By considering ${\Greekmath 011C} $ as drifting towards zero with the sample size, our approach provides a better guidance on construction of EIV robust estimators with good finite sample properties when ${\Greekmath 011C} $ is small or moderate. \footnote{ Nonstandard asymptotic approximations with drifting parameters are often used to obtain better approximations of the finite sample behavior of estimators and tests. For example, in the instrumental variable regression settings, to consider the settings with relatively small first stage coefficients, StaigerStock1997 model them as shrinking with $n$. It is important to keep in mind that such nonstandard asymptotic approximations are merely mathematical tools. One should not take them literally and think of parameters somehow changing if more data is collected. }
Using this approximation, we show that the proposed estimation approach indeed addresses the EIV problem. The MERM estimator is shown to be $\sqrt{n} $-consistent and asymptotically normal and unbiased. The standard confidence intervals and tests for GMM estimators are also valid for the MERM\ estimator. Additionally, the standard GMM arsenal of assessment tools can be applied to the MERM estimator, allowing one to test model identification, conduct valid inference, and perform model specification diagnostics.
The usefulness of a large sample theory is measured by its ability to approximate the finite sample properties of the estimators and inference procedures. Thus, we study the MERM estimators in a variety of simulation experiments. The results confirm that the nonstandard asymptotic theory indeed provides a good approximation of the finite sample properties of the estimators even in the settings with relatively large EIV. In some of the simulation experiments, the EIV are so large that for the naive estimators' standard $95\%$ confidence intervals have actual coverages of $0\%$ in finite samples, due to the magnitude of the EIV\ bias. At the same time, even in these settings the MERM estimators perform well, removing the EIV bias and providing confidence intervals with the correct coverage. In particular, the simulation results show that despite the simplicity of implementation, the MERM estimators can compete with and outperform semi-nonparametric estimators.
The MERM estimator is structurally different from the existing approaches that require nonparametric estimation of some nuisance parameters, for example, of the density $f_{X^{\ast }|Z,W}$. Avoiding nonparametric estimation has at least two advantages. First, since the majority of empirical applications include additional covariates $ W_{i}$, nonparametric estimation is often infeasible due to the curse of dimensionality. Because the MERM estimator does not involve any nonparametric estimation, it can be used in applications with a relatively large number of additional covariates $W_{i}$, and remains feasible even in the more complicated settings, including multi-equation and structural models, and applications with multiple mismeasured variables $X_{i}$. Second, estimation of infinite-dimensional nuisance parameters is typically more demanding towards the sources of identification available in the data, for example, requiring an instrumental variable with a large support (continuously distributed). In contrast, having a discrete instrument is sufficient for the MERM approach because the nuisance parameter ${\Greekmath 010D} _{0}$ is finite-dimensional.
For example, in Section (ref) we consider estimation of the model of multinomial choice among three modes of transportation. A leading alternative approach to the errors-in-variables problem in this model is the semi-nonparametric sieve-MLE estimator advocated by HuSchennach2008Ecta,CarrollChenHu2010JoNS, among others. This approach requires, among other things, estimating the conditional density $ f_{X^{\ast }|Z,W}$ of $X_{i}^{\ast }$ given the instrument $Z_{i}$ and covariates $W_{i}$. In this empirical example, $W_{i}$ includes four continuously distributed covariates (two continuously distributed characteristics per choice) and a discrete one, while scalar $X_{i}^{\ast }$ and $Z_{i}$ are also continuous. Thus, $f_{X^{\ast }|Z,W}$ is a function of six continuous and one discrete variable. Hence, for typical sample sizes, estimating $f_{X^{\ast }|Z,W}$ in this example is infeasible due to the curse of dimensionality. In contrast, as the results of Section (ref) demonstrate, the MERM approach is practical and effective in this application, in part because it avoids estimation of the high-dimensional nuisance functions like $f_{X^{\ast }|Z,W}$ altogether.
The simplicity and practicality of the MERM\ approach do come at a cost: there is a limit on the magnitude of the measurement errors it can handle. For example, one generally should not expect the MERM\ approach to work well when ${\Greekmath 011C} >1$, i.e., when the noise dominates the signal; in this case the researcher should seek an alternative estimation method.
We view the MERM\ approach as providing a bridge between the settings in which the measurement errors are guaranteed to be absent or negligible, and the settings where the measurement errors are so large that one has to use the relatively more complicated estimators from the earlier literature (if they exist at all for the model of interest).
{ Related Literature} ChenHongNekipelov2011JEL, Schennach2016AnnRev, and Schennach2020HB-ME provide excellent overviews of the measurement error literature.
The existing semiparametric approaches to estimation and inference in models with EIV\ involve nonparametric estimation of infinite-dimensional nuisance parameters (e.g., \QTR{citealp}{ Chesher2000WP,Li2002JoE,Schennach2004Ecta,Schennach2007Ecta,HuSchennach2008Ecta,SchennachHu2013JASA,Song2015JoE }), simulation (e.g., \QTR{citealp}{Schennach2014Ecta-ELVIS}), or both (e.g., \QTR{citealp}{Newey2001REStat,WangHsiao2011JoE}$)$. The exceptions include models with linear and polynomial regression functions (see \QTR{citealp}{HINP1991JoE,HausmanNeweyPowell1995JoE}$)$, and Gaussian control variable models such as Probit and Tobit with endogeneity (see \QTR{citealp}{SmithBlundell1986Ecta,RiversVuong1988JoE}).
To the best of our knowledge, this paper is the first to provide an approach for $\sqrt{n}$-consistent and asymptotically normal and unbiased estimation of general GMM models with EIV that does not require any nonparametric estimation (or simulation).
We are able to provide such an estimator because we focus on the models with moderate measurement errors. Modeling the variance of the measurement error as shrinking to zero with the sample size is a popular approach in Statistics. The method has been proposed by WolterFuller1982AS, who used it to construct an approximate MLE\ estimator of a nonlinear regression model with Gaussian errors. Following their approach, the Statistics literature has mainly focused on the settings where the moments of the EIV needed to bias correct the estimators are either known or can be directly estimated from the available data such as repeated measurements (e.g., \QTR{citealp}{CarrollStefanski1990JASA,CarrollEtAl2006Book-ME}). In Economics, such data are relatively rare. The use of approximations with shrinking variance of measurement errors in Econometrics literature has been pioneered by Kadane1971Ecta, Amemiya1985JoE, and Chesher1991Biomet. Such approximations have been used to check the sensitivity of naive estimators to the EIV by considering how the estimates change as the unknown moments of the measurement errors vary within some set of plausible values, e.g., see ChesherSchluter2002ReStud, ChesherDumanganeSmith2002JoE, BattistinChesher2014JoE, Chesher2017JoE, and HongTamer2003JoE. BoundBrownMathiowetz2001HBoE review a broad list of validation studies matching standard economic dataset to administrative records. The estimates they report suggest that the measurement errors of moderate magnitude are typical for empirical applications. This suggests that the approach developed in this paper could prove valuable for a wide range of applied work.
This paper differs from the earlier literature in several ways. First, it presents a way to estimate the unknown nuisance parameters (moments of the measurement errors) jointly with the parameters of interest. As a result, the approach can, for example, use instrumental variables as a source of identification. Second, the method applies to a very general class of semiparametric models specified by moment conditions. Third, the MERM approach allows the measurement errors to have larger magnitudes than most of the papers in the earlier literature; this is achieved by the MERM approach recursively bias correcting the bias correction terms.
The most widespread approach to identification of the EIV\ models in economic applications is to use instrumental variables, e.g., see HINP1991JoE,Newey2001REStat,Schennach2007Ecta,WangHsiao2011JoE. In a recent paper, HahnHausmanKim2021EL reconsider the regression model in Amemiya1990JoE using a bias correction similar to ours. When proper excluded variables are not available, researchers have considered using higher moments of $X_{i}$ as instruments, e.g., see Reiersol1950Ecta,Lewbel1997Ecta,EricksonWhited2002ET,SchennachHu2013JASA,BenMosheDHaultfeuilleLewbel2017JoE . When available, repeated measurements can also be used to identify the model, e.g., see HINP1991JoE,LiVuong1998JoE,Li2002JoE,Schennach2004Ecta. The MERM estimator accommodates these identification approaches within a unified estimation framework.
The power of the general MERM\ approach can be illustrated in the NLR model. For example, when a candidate instrumental variable is available, the conditions it needs to satisfy are much weaker than what is required by many existing approaches. Availability of a discrete instrument is sufficient for identification; and the instrument is allowed to have heterogeneous impact on covariates $X_{i}^{\ast }$. \ One can also take a nonclassical, nonlinear (e.g., discretized or censored), or biased measurement of $X_{i}^{\ast }$ as an instrument in the MERM approach. We discuss identification in Section (ref). In addition, in a related paper EvdokimovZeleneev2022WP-NPID study nonparametric regression with EIV using the ${\Greekmath 011C} \rightarrow 0$ approximation, and demonstrate that the MERM approach can also be motivated from a nonparametric perspective.
KitamuraOtsuEvdokimov2013Ecta,AndrewsGentzkowShapiro2017QJE,ArmstrongKolesar2021QE,BonhommeWeidner2021QE , among others, develop tools for estimation and inference in GMM, which are robust to general perturbation or misspecification of the true data generating process. They focus on the settings in which these perturbations are sufficiently small, so that naive estimators remain $\sqrt{n}$ -consistent, and their biases are of the same order of magnitude as their standard errors. In contrast, we focus on more specific forms of data contamination due to the EIV. This allows the MERM approach to remain valid even in the settings with larger measurement errors, in which naive estimators may have slower than $\sqrt{n}$ rates of convergence.
The MERM approach also provides a useful foundation for dealing with EIV in more complicated settings. EvdokimovZeleneev2018WP-Inference utilize the MERM framework to address an issue of nonstandard inference, which turns out to arise generally when EIV models are identified using instrumental variables. EvdokimovZeleneev2019WP-Panel extend the analysis of this paper to long panel and network settings.
{ Organization of the paper} Section (ref) introduces the Moderate Measurement Error framework and the proposed MERM\ estimator. Section (ref) presents several Monte Carlo experiments that illustrate finite sample properties of the MERM\ estimators. Section (ref) considers several extensions of the framework. A supplementary appendix contains all proofs and additional results for the numerical and empirical illustrations.
To present the main ideas we first consider the case of univariate $ X_{i}^{\ast }$. We will consider multivariate $X_{i}^{\ast }$ later. We assume that the measurement error is classical, i.e., that ${\Greekmath 0122}_i$ is independent of $X_i^*$ and $S_i$; later we will discuss how this assumption can be relaxed. Following the rest of the literature, we assume that $\mathbb{E}\left[ {\Greekmath 0122} _{i}\right] =0$.\footnote{ A location normalization such as $\mathbb{E}\left[ {\Greekmath 0122} _{i}\right] =0 $ is usually necessary because it is not possible to separately identify the means $\mathbb{E}\left[ X_{i}^{\ast }\right] $ and $\mathbb{E}\left[ {\Greekmath 0122} _{i}\right] $.}
To develop a practical estimation approach for general moment condition models we focus on the settings in which ${\Greekmath 011C} \equiv {\Greekmath 011B} _{{\Greekmath 0122} }/{\Greekmath 011B} _{X^{\ast }}$ is small or moderate. We consider an asymptotic approximation with ${\Greekmath 011C} _{n}\equiv {\Greekmath 011C} \rightarrow 0$ as $n\rightarrow \infty $.
Note that economically meaningful parameters are usually invariant to rescaling of $X_{i}^{\ast }$. Likewise, the extent of the EIV problem does not change with such rescaling. For simplicity of exposition, it is convenient to assume that $X_{i}^{\ast }$ is scaled so that ${\Greekmath 011B} _{X^{\ast }}$ is of order one and, correspondingly, the moments $\mathbb{E}[\left\vert {\Greekmath 0122} _{i}\right\vert ^{k}]\propto {\Greekmath 011C} _{n}^{k}$ decrease with $k$ when ${\Greekmath 011C} _{n}<1$. For example, this could be ensured by normalizing observed $X_{i}$ to have ${\Greekmath 011B} _{X}=1$. Let us stress that this normalization is used only to simplify the exposition; as we show in Appendix (ref), the proposed MERM\ estimator does not require any normalizations in practice.
For clarity, we first consider a simple special case of the general approach. Let us denote $g_{x}^{(k)}\left( x,s,{\Greekmath 0112} \right) \equiv \partial ^{k}g\left( x,s,{\Greekmath 0112} \right) /\partial x^{k}$. Since $\mathbb{E} [\left\vert {\Greekmath 0122} _{i}\right\vert ^{k}]\propto {\Greekmath 011C} _{n}^{k}\rightarrow 0$ as $n\rightarrow \infty $, under some regularity conditions, we can write the quadratic Taylor expansion of function $ g(X_{i},S_{i},{\Greekmath 0112} )=g(X_{i}^{\ast }+{\Greekmath 0122} _{i},S_{i},{\Greekmath 0112} )$ around ${\Greekmath 0122} _{i}=0$ as
where the second equality holds because ${\Greekmath 0122} _{i}$ and $ \left( X_{i}^{\ast },S_{i}\right) $ are independent, and ${\mathbb{E}\left[ {\Greekmath 0122} _{i}\right] =0}$.
Evaluating the expansion above at ${\Greekmath 0112} = {\Greekmath 0112}_0$ gives $\mathbb{E}[g(X_{i},S_{i},{\Greekmath 0112} _{0})]=O\left( {\Greekmath 011B} _{{\Greekmath 0122} }^{2}\right) =O({\Greekmath 011C} _{n}^{2})$, because $\mathbb{E}[g(X_{i}^*,S_{i},{\Greekmath 0112} _{0})] = 0$. As a result, a naive estimator that ignores the EIV and uses $X_{i}$ in place of $X_{i}^{\ast }$ has EIV\ bias of order ${\Greekmath 011C} _{n}^{2}$.\footnote{ For example, consider a linear regression with a scalar mismeasured regressor. The bias of the naive OLS\ estimator of the slope parameter $ {\Greekmath 0112} _{01}$ is $-{\Greekmath 0112} _{01}\frac{{\Greekmath 011C} _{n}^{2}}{1+{\Greekmath 011C} _{n}^{2}}=-{\Greekmath 0112} _{01}{\Greekmath 011C} _{n}^{2}+O\left( {\Greekmath 011C} _{n}^{4}\right) $.} The bias of the naive estimator should be compared with its standard error, which is of order $ n^{-1/2}$. Thus, the bias of the naive estimator is not negligible, unless the measurement error is rather small (theoretically, unless ${\Greekmath 011C} _{n}^{2}=o\left( n^{-1/2}\right) $). In particular, tests and confidence intervals based on the naive estimator are invalid and can provide highly misleading results. Moreover, if ${\Greekmath 011C} _{n}^{2}$ shrinks at a rate slower than $O\left( n^{-1/2}\right) $, the rate of convergence of the naive estimator is slower than $\sqrt{n}$.
Suppose ${\Greekmath 011C} _{n}=o\left( n^{-1/6}\right) $. Then, $O({\Greekmath 011C} _{n}^{3})=o\left( n^{-1/2}\right) $ and we can rearrange equation ((ref)) as
The left-hand side of this equation is exactly the moment condition ((ref)) that we would like to use for estimation of ${\Greekmath 0112} _{0} $. The first term on the right-hand side involves only observed variables, and can be estimated by the sample average $\overline{g}({\Greekmath 0112} )\equiv n^{-1}\sum_{i=1}^{n}g(X_{i},S_{i},{\Greekmath 0112} ).$ The second term on the right-hand side can be thought of as a bias correction that removes the EIV-bias from the expected moment function $\mathbb{E}[g(X_{i},S_{i},{\Greekmath 0112} )]$.
The idea of the MERM\ estimator we propose is to make use of expansions such as ((ref)) to bias correct the moment condition $\mathbb{ E}[g(X_{i},S_{i},{\Greekmath 0112} )]$, which in turn removes the bias of the estimator of the parameters of interest ${\Greekmath 0112} _{0}$. To perform the bias correction we need to estimate two quantities: $\mathbb{E}[{\Greekmath 0122} _{i}^{2}]$ and $ \mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$.
First, we show that in equation ((ref)) we can substitute $\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$ with $ \mathbb{E}[g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]$, which in turn can be estimated by $\overline{g}_{x}^{(2)}({\Greekmath 0112} )\equiv n^{-1}\sum_{i=1}^{n}g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )$. By the Taylor expansion around ${\Greekmath 0122} _{i}=0$ similar to equation ((ref)), we can show that $\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]=\mathbb{E}[g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]+O({\Greekmath 011C} _{n}^{2})$ and hence
Here $O\left( {\Greekmath 011C} _{n}^{4}\right) =o\left( n^{-1/2}\right) $ because we assume that ${\Greekmath 011C} _{n}=o\left( n^{-1/6}\right) $. The idea behind this substitution is that the bias of order $O({\Greekmath 011C} _{n}^{2})$ in $\mathbb{E} [g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]$ can be ignored because it is multiplied by $E\left[ {\Greekmath 0122} _{i}^{2}\right] =O\left( {\Greekmath 011C} _{n}^{2}\right) $. \footnote{ Such substitutions of $X^{\ast }$ with $X$ have been used in other contexts, e.g., ChesherSchluter2002ReStud.} With the substitution, we can rearrange equation ((ref)) and write it as
Second, we propose estimating the unknown $\mathbb{E}[{\Greekmath 0122} _{i}^{2}]$ together with the parameter of interest ${\Greekmath 0112} $. Specifically, let ${\Greekmath 010D} _{02}\equiv \mathbb{E}[{\Greekmath 0122} _{i}^{2}]/2$ denote the true value of parameter ${\Greekmath 010D} _{2}$, and consider the following corrected moment function:
Function ${\Greekmath 0120} $ is a moment function parameterized by ${\Greekmath 0112} $ and ${\Greekmath 010D} $, and
where the first equality follows from equation ((ref)) and the definition of ${\Greekmath 010D} _{02}$, and the second equality follows from equation ((ref)). Hence, the corrected moment conditions ${\Greekmath 0120} $ can be used to jointly estimate the true parameters $ {\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{02}$ by a GMM\ estimator.\footnote{ In the moment condition settings, having $o\left( n^{-1/2}\right) $ is equivalent to having $0$ on the right-hand side\ of equation ((ref)).}
The quadratic expansion of equation ((ref)) can be extended to general order $K\geq 2$. Considering larger $K$ theoretically allows ${\Greekmath 011C} _{n}$ converging to zero at a slower rate. In finite samples this corresponds to the asymptotics providing good approximations for larger values of ${\Greekmath 011C} _{n}$, i.e., large measurement errors. Expanding $g(X_{i}^{\ast }+{\Greekmath 0122} _{i},S_{i},{\Greekmath 0112} )$ around ${\Greekmath 0122} _{i}=0$ we have,
The above special case of quadratic expansion corresponds to $K=2$.
The approximation we consider is formalized by the following assumption.
Assumption (ref)(i) limits the magnitude of the measurement errors and implies that ${\Greekmath 011C} _{n}^{K+1}=o\left( n^{-1/2}\right) $. Assumption (ref)(ii) implies that $\mathbb{E}[\left\vert {\Greekmath 0122} _{i}\right\vert ^{k}]=O\left( {\Greekmath 011B} _{{\Greekmath 0122} }^{k}\right) $, and requires the tails of ${\Greekmath 0122} _{i}/{\Greekmath 011B} _{{\Greekmath 0122} }$ to be sufficiently thin. Together, parts (i) and (ii) imply that $\mathbb{E} [\left\vert {\Greekmath 0122} _{i}\right\vert ^{K+1}]=O\left( {\Greekmath 011C} _{n}^{K+1}\right) =o\left( n^{-1/2}\right) $, and hence ensure that the remainder in equation ((ref))$\ $is negligible. Using $\mathbb{E}\left[ {\Greekmath 0122} _{i}|X_{i}^{\ast },S_{i} \right] =0$ to further simplify this expansion and rearranging the terms we obtain
This equation is the general expansion analog of equation ((ref)). The summation on the right hand side is the bias correction term, which we use to construct the MERM\ estimator.\footnote{It is useful to get a sense of the magnitudes of the coefficients $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{k}\right] /k!$ in equation ((ref)). Suppose ${\Greekmath 0122} _{i}\sim N\left( 0,{\Greekmath 011B} _{{\Greekmath 0122} }^{2}\right) $, ${\Greekmath 011B} _{{\Greekmath 0122} }=0.5$, and $ {\Greekmath 011B} _{X^{\ast }}=1$, so$\ {\Greekmath 011C} ={\Greekmath 011B} _{{\Greekmath 0122} }=0.5$. Then the coefficients in front of $g_{x}^{\left( 2\right) }$, $g_{x}^{\left( 4\right) }$, and $g_{x}^{\left( 6\right) }$ are $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{2} \right] /2!=0.125$, $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{4}\right] /4!\approx 0.008$, and $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{6}\right] /6!\approx 0.0003$ . }
It turns out that for $K\geq 4$, estimation of $\mathbb{E} [g_{x}^{(k)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$ is more intricate than in the case of $K=2$, and the substitution we made in equation ((ref)) no longer works. Larger values of $K$ allow for larger values of ${\Greekmath 011C} _{n}$ and hence larger EIV\ biases of naive estimators $ n^{-1}\sum_{i=1}^{n}g_{x}^{(k)}(X_{i},S_{i},{\Greekmath 0112} )$. The expansion of order $K$ includes terms up to the order ${\Greekmath 011C} _{n}^{K}$, with the asymptotically negligible remainder of order $O\left( {\Greekmath 011C} _{n}^{K+1}\right) $. For $K\geq 4$, terms of order ${\Greekmath 011C} _{n}^{4}$ are not negligible. This implies that we cannot ignore the EIV bias that would arise from substituting $\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$ with $ \mathbb{E}[g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]$ in equation ((ref) ), because this bias is of order $O\left( {\Greekmath 011C} _{n}^{4}\right) $ according to equation ((ref)). To address this problem, we instead replace $\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$ with the bias corrected expression $\mathbb{E} [g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]-\left( \mathbb{E}[{\Greekmath 0122} _{i}^{2}]/2\right) \mathbb{E}[g_{x}^{(4)}(X_{i},S_{i},{\Greekmath 0112} )]$. Thus, for $ K\geq 4$, one needs to bias correct the estimator of the bias correction term. Moreover, for larger $K$ one needs to bias correct the bias correction of the bias correction term and so on.
Fortunately, we show that these bias corrections can be constructed as linear combinations of the expectations of the higher order derivatives of $ g_{x}^{\left( k\right) }(X_{i},S_{i},{\Greekmath 0112} )$. Let us define the following corrected moment function:
where ${\Greekmath 010D} =({\Greekmath 010D} _{2},\dots ,{\Greekmath 010D} _{K})^{\prime }$ is a $K-1$ dimensional vector of parameters. Let ${\Greekmath 010D} _{0}\equiv ({\Greekmath 010D} _{02},\dots ,{\Greekmath 010D} _{0K})^{\prime }$ denote the vector of true parameters ${\Greekmath 010D} _{0k}$ , defined as
We formalize this discussion below.
The following lemma establishes validity of the corrected moment conditions under Assumptions (ref), (ref), and some mild regularity conditions provided in Appendix (ref).
Lemma (ref) implies that the corrected moment conditions ${\Greekmath 0120} $ are valid and can potentially be used to jointly estimate parameters ${\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{0}$. The total number of parameters to be estimated is now $\dim \left( {\Greekmath 0112} \right) +K-1$. Thus, joint estimation of ${\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{0}$ requires that $\dim \left( {\Greekmath 0120} \right) =\dim \left( g\right) \geq \dim \left( {\Greekmath 0112} \right) +K-1$, i.e., that the original moment conditions $g$ include sufficiently many overidentifying restrictions. For example, the overidentifying restrictions can be constructed by using an instrumental variable; we discuss this in more detail below.
The MERM estimator jointly estimates the parameters ${\Greekmath 0112} _{0}$ and $ {\Greekmath 010D} _{0}$ using moment conditions ${\Greekmath 0120} $. It is convenient to define the joint vector of parameters
and the parameter space $\mathcal{B}\equiv \Theta \times \Gamma $, where $ \Theta $ and $\Gamma $ are the parameter spaces for ${\Greekmath 0112} $ and ${\Greekmath 010D} $. Then, MERM estimator is the GMM\ estimator (\QTR{citealp}{Hansen1982Ecta}):
where $\overline{{\Greekmath 0120} }({\Greekmath 010C} )\equiv n^{-1}\sum_{i=1}^{n}{\Greekmath 0120} _{i}({\Greekmath 010C} )$ , ${\Greekmath 0120} _{i}({\Greekmath 010C} )\equiv {\Greekmath 0120} \left( X_{i},S_{i},{\Greekmath 010C} \right) $, $\hat{\Xi }$ is a weighting matrix, and $\hat{Q}({\Greekmath 010C} )$ is the standard GMM objective function.
While Lemma (ref) establishes validity of the corrected moment restrictions ${\Greekmath 0120} $, the MERM estimator also relies on $ {\Greekmath 010C} _{0}$ being identified from ${\Greekmath 0120} $. This requirement is formalized by the following assumption.
Assumptions (ref)((ref)) and ((ref)) are the standard GMM local and global identification conditions applied to the moment function ${\Greekmath 0120} (X_{i}^{\ast },S_{i},{\Greekmath 0112} ,{\Greekmath 010D} )$. These are high-level conditions, which we will return to later in the paper. The moment conditions formulation is sufficiently general to encompass a wide variety of sources of identification. In Section (ref), we discuss identification in detail and illustrate the construction of the moment function using an instrumental variable or a second measurement.
Under some additional regularity conditions, estimator $\hat{{\Greekmath 010C}}$ behaves as a standard GMM-type estimator: it is $\sqrt{n}$-consistent and asymptotically normal and unbiased. This result is formalized by the following theorem.
Theorem (ref) shows that the MERM\ approach addresses the EIV bias problem, and in particular provides a $\sqrt{n}$-consistent asymptotically normal and unbiased estimator $\hat{{\Greekmath 0112}}$, which can be used to conduct inference about the true parameters ${\Greekmath 0112} _{0}$. The asymptotic variance $\Sigma $ takes the standard sandwich\ form, with $\Psi \equiv \mathbb{E}\left[ {\Greekmath 0272} _{{\Greekmath 010C} }{\Greekmath 0120} _{i}({\Greekmath 010C} _{0})\right] $, $ \Omega _{{\Greekmath 0120} {\Greekmath 0120} }\equiv \mathbb{E}\left[ {\Greekmath 0120} _{i}\left( {\Greekmath 010C} _{0}\right) {\Greekmath 0120} _{i}^{\prime }\left( {\Greekmath 010C} _{0}\right) \right] $, and $\hat{ \Xi}\rightarrow _{p}\Xi $.
Once the corrected moment condition ${\Greekmath 0120} $ is constructed, estimation of and inference about parameters ${\Greekmath 010C} _{0}$ can be performed using any standard software package for GMM\ estimation. In other words, the proposed estimator can be simply treated as a standard GMM estimator based on the corrected moment conditions ${\Greekmath 0120} $, and the conventional standard errors, tests, and confidence intervals are valid.
In addition to the estimation of the parameters ${\Greekmath 0112}_{0}$, researchers are often interested in average effects of the form ${\Greekmath 0115}_{0}\equiv \mathbb{E}\left[ {\Greekmath 0115}\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) \right] $. For instance, in the NLR model, one may be interested in the average partial effect of $x$ (i.e., ${\Greekmath 0115} _{0}\equiv \mathbb{E}\left[ {\Greekmath 0272} _{x}{\Greekmath 011A} \left( X_{i}^{\ast },S_{i},{\Greekmath 0112}_{0}\right) \right] $) or another covariate. The naive average partial effect estimator $\hat{{\Greekmath 0115}}_{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Naive}}\equiv \frac{1}{n} \sum_{i=1}^{n}{\Greekmath 0115}(X_{i},S_{i},\hat{{\Greekmath 0112}})$ suffers from the EIV bias, unless function ${\Greekmath 0115}$ is linear in $X_{i}^{\ast }$. Instead, one should use estimates $\hat{{\Greekmath 010D}}$ to construct the bias-corrected estimator $\hat{ {\Greekmath 0115}}_{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{MERM}}\equiv \frac{1}{n}\sum_{i=1}^{n}\left\{ {\Greekmath 0115}(X_{i},S_{i},\hat{{\Greekmath 0112}})-\sum_{k=2}^{K}\hat{{\Greekmath 010D}} _{k}{\Greekmath 0115}_{x}^{(k)}(X_{i},S_{i},\hat{{\Greekmath 0112}})\right\} $.
Theorem (ref) requires ${\Greekmath 010C} _{0}$ to be identified and the Jacobian matrix $\Psi $ to be full rank. Notably, the MERM framework encompasses many possible sources of identification at once, including instrumental variables, additional measurements, or nonlinearities of the functional form. The identifying information is incorporated in the moment functions. Essentially, our approach first characterizes in what directions the measurement errors can bias the moment conditions $\mathbb{E}\left[ g\left( X_{i},S_{i},{\Greekmath 0112} \right) \right] $, and then uses the moments orthogonal to those directions for identification of ${\Greekmath 0112} _{0}$. To be more specific, we will now consider identification when in addition to the error-laden $X_{i}$ we have either (i) a general instrument $Z_{i}$ or (ii) a second measurement $Q_{i}$.
\paragraph{Identification Using A General Instrument $Z_{i}$}
Many applications can be formulated as the following conditional moment restriction:
for some moment function $u$. For example, consider the nonlinear regression model $\mathbb{E}\left[ Y_{i}|X_{i}^{\ast }=x\right] ={\Greekmath 011A} \left( x,{\Greekmath 0112} _{0}\right) $, then $u\left( x,y,{\Greekmath 0112} \right) ={\Greekmath 011A} \left( x,{\Greekmath 0112} \right) -y$.\footnote{ For simplicity of the exposition, in the expectation in equation ((ref)) we only condition on $X_{i}^{\ast }$. The discussion applies in a straightforward way to the settings with additional correctly measured variables $W_{i}$ in the conditioning set, i.e., the model $\mathbb{E}\left[ u\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) |X_{i}^{\ast },W_{i}\right] =0 $. For example, in the nonlinear regression example with additional covariates $W_{i}$ we have $\mathbb{E}\left[ Y_{i}|X_{i}^{\ast }=x,W_{i}=w \right] ={\Greekmath 011A} \left( x,w,{\Greekmath 0112} _{0}\right) $, so $u\left( x,y,w,{\Greekmath 0112} \right) ={\Greekmath 011A} \left( x,w,{\Greekmath 0112} \right) -y$.}
In applications, identification of the models with EIV would typically rely on an instrumental variable $Z_{i}$. Suppose the instrument satisfies the exclusion restriction $\mathbb{E}\left[ u\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} \right) |X_{i}^{\ast },Z_{i}\right] =\mathbb{E}\left[ u\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} \right) |X_{i}^{\ast }\right] $, i.e., conditional on the true $X_{i}^{\ast }$ the instrument has no further effect on the moment conditions $u$. Consider the moment functions $h\left( x,s,{\Greekmath 0112} \right) \equiv u\left( x,s,{\Greekmath 0112} \right) \otimes {\Greekmath 0127} _{X}\left( x\right) $, where ${\Greekmath 0127} _{X}\left( x\right) $ is a vector of functions of $x$, e.g., $ {\Greekmath 0127} _{X}\left( x\right) \equiv \left( 1,x,\ldots ,x^{J}\right) ^{\prime } $. By the Law of Iterated Expectations, $\mathbb{E}\left[ h\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) |Z_{i}=z\right] =0$ for all $z$. However, the same expectation with $X_{i}^{\ast }$ replaced by the observed $ X_{i}$, $\mathbb{E}\left[ h\left( X_{i},S_{i},{\Greekmath 0112} _{0}\right) |Z_{i}=z \right] $, will depend on $z$. One way to see this is to consider a Taylor expansion similar to equation ((ref)):{
}which shows that $\mathbb{E}\left[ h\left( X_{i},S_{i},{\Greekmath 0112} _{0}\right) |Z_{i}=z\right] $ is zero for all $z$ (up to a negligible remainder) unless $ {\Greekmath 010D} _{0}\neq 0$, where ${\Greekmath 010D} _{0}={\Greekmath 011B} ^{2}/2$. Thus, $\mathbb{E} \left[ h\left( X_{i},S_{i},{\Greekmath 0112} _{0}\right) |Z_{i}=z\right] $ varies with $ z$ only because of the presence of the measurement error. Intuitively, the magnitude of this variation then identifies the nuisance parameters ${\Greekmath 010D} _{0}$. Thus, one can rely on the original moment functions of the typical form $g\left( x,s,{\Greekmath 0112} \right) =h\left( x,s,{\Greekmath 0112} \right) \otimes {\Greekmath 0127} _{Z}\left( z\right) $, where ${\Greekmath 0127} _{Z}\left( z\right) $ is a vector of functions of $z$.
The above discussion provides the intuition for identification of the nonlinear moment condition models with EIV. It is important to note that identification of general nonlinear moment condition models is a complicated problem. Even in the settings without measurement errors, it is generally not possible to give low-level conditions guaranteeing that a specific set of nonlinear moment conditions identifies the parameter vector. The presence of EIV\ makes the question of identification even harder.
We attempt to address this concern and make the above intuitions more precise in two ways. First, in the following subsection we consider a specific (but frequently employed) kind of an instrument: a second measurement (possibly non-classical). The specific form of the excluded variable allows us to provide more transparent identification conditions. Second, in EvdokimovZeleneev2022WP-NPID we study nonparametric regression model with EIV using the ${\Greekmath 011C} _{n}\rightarrow 0$ approximation. We show that the model is identified using an instrument (even a discrete one), and motivate the MERM approach from a nonparametric perspective. Finally, since MERM estimator is a standard GMM\ estimator, one can test the strength of identification of the model parameters, or conduct identification-robust inference using the standard methods (e.g., \QTR{citealp}{ StockWright2000Ecta,Kleibergen2005Ecta,GuggenbergerSmith2005ET,GuggenbergerRamalhoSmith2012JoE,AndrewsMikusheva2016Ecta-ConditionalFunctionalNuisance,AndrewsI-2016-Ecta-CLC,AndrewsGuggenberger2019QE }).
\paragraph{Identification Using A Second Measurement}
Suppose we observe a second measurement
where ${\Greekmath 010B} _{1}$ may not be known. Assume that ${\Greekmath 010B} _{1}\neq 0$ and $ \mathbb{E}\left[ {\Greekmath 0122} _{Q,i}|X_{i}^{\ast },S_{i},{\Greekmath 0122} _{i} \right] =0$. The variance of ${\Greekmath 0122} _{Q,i}$ does not need to be small. Note that the measurement error in $Q_{i}$ can be non-classical: $ Q_{i}-X_{i}^{\ast }$ and $X_{i}^{\ast }$ are correlated unless ${\Greekmath 010B} _{1}=1 $.
Consider the conditional moment restrictions ((ref)). If $ X_{i}^{\ast }$ were observed, we could have constructed the unconditional moments
for some $J\geq \dim \left( {\Greekmath 0112} \right) -1$. Suppose that the model is identified if $X_{i}^{\ast }$ observed, which means that the Jacobian of these moment conditions has full rank:
To deal with the error-laden $X_{i}$, consider the MERM estimator with $K=2$ based on the following moment function
Here the total number of moments is $m=2J+1$. The additional $J$ moments added in equation ((ref)) use $Q_{i}$, which will allow identifying ${\Greekmath 010D} _{0}=\mathbb{E}[{\Greekmath 0122} _{i}^{2}]/2$.
It turns out that in these settings there is a simple sufficient condition for Assumption (ref)((ref)) to hold. Appendix (ref) demonstrates that $\Psi ^{\ast }$ will have full rank if
Condition ((ref)) has a very simple interpretation: it essentially it means that $\mathbb{E}\left[ \left. u_{x}^{\left( 1\right) }\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) \right\vert X_{i}^{\ast } \right] $ should not be identically zero. For example, in the nonlinear regression model $u\left( x,y,{\Greekmath 0112} \right) ={\Greekmath 011A} \left( x,{\Greekmath 0112} \right) -y $, and condition ((ref)) is satisfied as long as $ \mathbb{E}\left[ {\Greekmath 011A} _{x}^{(1)}\left( X_{i}^{\ast },{\Greekmath 0112} _{0}\right) \left( X_{i}^{\ast }\right) ^{j}\right] \neq 0$ for some $j\in \left\{ 0,\ldots ,J-1\right\} $.
We compare MERM estimator with the state-of-the-art semiparametric estimator of Schennach2007Ecta for nonlinear regression models. The Monte Carlo designs are taken from S07, and include a polynomial, rational fraction, and Probit nonlinear regression models. Identification of the model is ensured by the availability of an instrument.
$(Z_i,V_i,{\Greekmath 0122}_i)' \sim N \left((0, 0, 0)', \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Diag}(1, 1/4, 1/4)\right)$, ${\Greekmath 0119}_1=1$, and $n = 1000$. The conditional expectation function ${\Greekmath 011A}$, the true value of the parameter of interest ${\Greekmath 0112}_0$, and the conditional distribution of the regression error $U_i$ are design-specific and reported in Tables (ref)-(ref) below. In all designs, ${\Greekmath 011C} = {\Greekmath 011B}_{{\Greekmath 0122}}/{\Greekmath 011B}_{X}^* \approx 0.45$, so the measurement error is “fairly large” Schennach2007Ecta.
We report simulation results for the MERM estimator considering correction schemes with $K=2$ and $K=4$. The original moment function is
where we use ${\Greekmath 0127}(x,z) = \left(1, x, z, x^2, z^2, x^3, z^3\right)'$ for $K=2$ and ${\Greekmath 0127}(x,z) = \left(1, x, z, x^2, xz, z^2, x^3, x^2 z, x z^2, z^3\right)'$ for $K=4$.
The finite sample properties of the MERM estimators (evaluated based on 5,000 replications) are reported in Tables (ref)-(ref) below. For comparison, we also provide the same statistics for naive estimators (OLS/NLLS) and for the benchmark estimator of S07 (as reported in the original paper). For the polynomial model (Table (ref)), both $K=2$ and $K=4$ MERM estimators effectively remove the EIV bias. Component-wise, the MERM estimators perform similarly (for ${\Greekmath 0112}_2$ and ${\Greekmath 0112}_4$) or better (for ${\Greekmath 0112}_1$ and ${\Greekmath 0112}_3$) compared to the benchmark estimator of S07. For the rational fraction model (Table (ref)), both the MERM estimators are vastly superior to the benchmark estimator both in terms of the bias and the standard deviation. For the probit model (Table (ref)), the MERM estimator with $K=2$ removes a large fraction of the EIV bias compared to the NLLS estimator. However, the EIV bias remains non-negligible when this simplest correction scheme is used. Employing a higher order correction scheme with $K=4$ completely eliminates the remaining EIV bias, while at the same time having smaller standard deviations (than the benchmark estimator of S07) . Overall, in the considered designs, the MERM estimator with $K=4$ consistently outperforms the benchmark estimator. It also proves to be more effective in removing the EIV bias compared to the $K=2$ estimator, especially in the highly nonlinear settings of the considered probit design.
Consider the standard multinomial logit model, in which an agent chooses between 3 available options. For an agent $i$ with characteristics $(X_i^*,W_i)$, the utility of option $j$ is given by
and $U_{i0} = {\Greekmath 010F}_{i0}$ for the outside option $j = 0$, where ${\Greekmath 010F}_{ij}$ are i.i.d. (across $i$ and $j$) draws from a standard type-1 extreme value distribution. The researcher observes $\left\{(X_i,W_i,Y_{i1},Y_{i2},Y_{i0})\right\}_{i=1}^n$, where $Y_{ij}$ is a binary variable indicating whether agent $i$ chooses option $j$, i.e. $Y_{ij} = 1$ if and only if $j = \operatorname*{\mathrm{arg}\!\max\limits}_{j' \in \{0,1,2\}} U_{i j'}$. In addition,
and $\left(V_{i1}, V_{i0}, Z_i,{\Greekmath 0122}_i,{\Greekmath 0117}_{i1},{\Greekmath 0117}_{i2}\right)' \sim N\left((1,0,0,0,0,0)', \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Diag}({\Greekmath 011B}_{V1}^2,{\Greekmath 011B}_{V0}^2, {\Greekmath 011B}_Z^2, {\Greekmath 011B}_{\Greekmath 0122}^2, {\Greekmath 011B}_{\Greekmath 0117}^2, {\Greekmath 011B}_{\Greekmath 0117}^2)\right)$. In all of the designs, we fix $({\Greekmath 0112}_{011},{\Greekmath 0112}_{012},{\Greekmath 0112}_{013},{\Greekmath 0112}_{021},{\Greekmath 0112}_{022},{\Greekmath 0112}_{023},{\Greekmath 011A},{\Greekmath 011B}_{V1}^2,{\Greekmath 011B}_{V0}^2,{\Greekmath 011B}_Z^2,{\Greekmath 011B}_{\Greekmath 0117}^2) = (1,0,0,0,0,0,0.7,1/2,1/2,1,1)$ and $n=2000$. We consider ${\Greekmath 011C} = {\Greekmath 011B}_{{\Greekmath 0122}}/{\Greekmath 011B}_{X^*} \in \{1/4, 1/2, 3/4\}$. Setting ${\Greekmath 011B}_{V1}=0$ would correspond to the additive control variable model. We omit such simulation results for brevity.
Similarly to Section (ref), we report results for the MERM estimators with $K=2$ and $K=4$ based on the following original moment function
where ${\Greekmath 0127}_j(x,z,w) = \left(1, x, z, x^2, z^2, x^3, z^3, w_j\right)'$ for $K=2$ and ${\Greekmath 0127}_j(x,z,w) = \left(1, x, z, x^2, xz, z^2, x^3, x^2 z, x z^2, z^3, w_j\right)'$ for $K=4$.
We report the results on estimation and inference on the partial derivatives of the conditional choice probabilities $p_j(x,w_1,w_2)$ with respect to $x$, $w_1$, and $w_2$, evaluated at the population means.
Table (ref) reports the finite sample biases, standard deviations, and RMSE of the MERM estimators, as well as the sizes of the corresponding t-tests with nominal size of 5%. To illustrate the importance of dealing with EIV, we also report the same statistics for the standard (naive) MLE estimator that ignores the presence of the measurement errors.
In all designs, the MLE estimator is biased, and the corresponding t-tests over-reject. Note that failing to account for the EIV in the mismeasured variable $X_i^*$ generally biases estimators of all of the parameter, including those corresponding to the correctly measured variables $W_{i1}$ and $W_{i2}$. In particular, the t-tests may falsely reject true null hypotheses $\partial p_j/\partial w_\ell=0$ up to nearly $100\%$ of the time.
The MERM estimator with $K=2$ removes a large fraction of the EIV bias in all of the designs. While this proves to be enough to achieve accurate size control when the magnitude of the measurement error is moderate (${\Greekmath 011C} = 1/4$), the remaining EIV bias may still result in size distortions of the t-tests with larger measurement errors, especially ${\Greekmath 011C} = 3/4$. Using the higher order correction scheme with $K=4$ effectively removes the EIV bias in all of the simulation designs for all of the parameters. Remarkably, the corresponding finite sample null rejection probabilities remain close to the nominal $5\%$ rate even when the standard deviation of the measurement error is as large as $75\%$ of the standard deviation of the mismeasured $X^*$.
To further check the limits of applicability of our method, we also consider larger values of ${\Greekmath 011C} \in \{1,3/2,2\}$. The numerical results analogous to the ones reported in Table (ref) are provided in Table (ref) in Appendix (ref). Specifically, we find that inference results based on the correction scheme with $K=4$ remain accurate even for ${\Greekmath 011C} = 1$. Unsurprisingly, inference becomes less reliable for bigger ${\Greekmath 011C} = 3/2$ and ${\Greekmath 011C} = 2$. At the same time, while the $K=4$ correction scheme fails to entirely eliminate the EIV bias in these designs, it still removes a big fraction of the bias and greatly improves on MLE in terms of the RMSE. Thus, while inference based on our estimator might be less reliable in extreme settings when the measurement error overwhelms the signal, the MERM estimator using $K=4$ appears to be sufficiently accurate over a wide range of ${\Greekmath 011C}$.
\afterpage{
}
In this section, we illustrate the finite sample properties of the MERM estimator in the context of a classical multinomial choice application: choice of transportation mode (e.g., McFadden1974JPubE).
To calibrate the numerical experiment, we use the ModeCanada dataset, a survey of business travelers for the Montreal-Toronto corridor. We focus on the subset of travelers choosing between train, air, and car ($n = 2769$), and estimate the conditional logit model with traveler $i$'s utilities given in the table below.
To generate the simulated samples, we randomly draw covariates from their joint empirical distribution. To generate the simulated outcomes, we draw ${\Greekmath 010F}_{ij}$ from the standard type-I extreme value distribution. The true value of ${\Greekmath 0112}_0$ is set to be the MLE estimate based on the original dataset. More details about this numerical experiment are given in Appendix (ref).
To evaluate the performance of the MERM estimator in these settings, we generate mismeasured $Income_i = Income_i^* + {\Greekmath 0122}_i$. We focus on the individual income because it is often mismeasured. We report the results for ${\Greekmath 011C} = {\Greekmath 011B}_{\Greekmath 0122}/{\Greekmath 011B}_{Income^*} \in \{1/4, 1/2, 3/4\}$.
Table (ref) reports the simulation results for the (naive) MLE estimator and for the MERM estimators with $K=2$ and $K=4$. We focus on estimation of and inference on the income elasticities (evaluated at the population mean of the covariates). The MLE estimator is considerably biased for ${\Greekmath 011C} \in \{1/2, 3/4\}$, which results in substantial size distortions of the MLE based t-tests. The MERM estimator with $K=4$ effectively eliminates the EIV bias and the corresponding t-tests provide accurate size control in all of the considered designs. The estimator with $K=2$ is more precise, while successfully removing the EIV bias for ${\Greekmath 011C}\le1/2$.
Overall, the MERM estimators perform well in the considered empirical context, providing a basis for estimation and inference even for quite large values of ${\Greekmath 011C}$.
Making an appropriate choice of the expansion order $K$ is important for the estimation procedure. One has to be cautious not to take $K$ too small, as this may result in an estimator that only partially removes the EIV bias. On the other hand, picking a larger $K$ than needed might inflate standard errors and result in less powerful inference.
In this section, we address this issue by providing a data-dependent procedure for selecting $K$. We demonstrate that our procedure has desirable theoretical properties. We also find that the procedure has good finite sample properties in a set of Monte Carlo simulation experiments across different values of ${\Greekmath 011C}$ and sample sizes.
Consider two alternative values of the expansion order: $L$ and $K$, where $2 \leq L<K$. In practice, even-order biases tend to dominate, so to reduce the set of choices, it is useful to focus on even values of the expansion orders $L$ and $K$. For example, one would typically be interested in choosing between $L=2$ and $K=4$.
Let $\hat {\Greekmath 010C}_L$ and $\hat {\Greekmath 010C}_K$ denote the corresponding MERM estimators. The estimator $\hat {\Greekmath 010C}_L$ should be preferred as having smaller asymptotic variance provided that its remaining EIV bias is negligible relative to its standard error. Otherwise, the more conservative $\hat {\Greekmath 010C}_K$ should be used instead.
Note that Lemma (ref) suggests that the remaining asymptotic bias of $\hat {\Greekmath 010C}_L$ (due to the additional terms accounted for when the expansion of higher order $K$ is used) is given by (up to an $o(n^{-1/2})$ remainder)
where matrix $B$ is based on the moments used for estimation of $\hat {\Greekmath 010C}_L$.
Importantly, ${\Greekmath 010D} _{0k}=O\left( {\Greekmath 011B}_{{\Greekmath 0122}}^{k}\right) $, and hence for $k>2$ we can estimate a bound on $\sqrt{n}{\Greekmath 010D} _{0k}$ sufficiently quickly to provide a valid procedure for choosing $K$. To this end, we first estimate the model using the larger $K$. Let $\hat {\Greekmath 010C}_K = (\hat {\Greekmath 0112}', \hat {\Greekmath 010D}')'$ and $\hat {\Greekmath 011B}_{\Greekmath 0122}^2 \equiv 2 \hat {\Greekmath 010D}_{2}$, where we dropped the additional subscripts $K$ for notation simplicity. Then, we can estimate ${\Greekmath 011B}_{\Greekmath 0122}^K$ by $\hat {\Greekmath 011B}_{\Greekmath 0122}^K \equiv (\hat {\Greekmath 011B}_{\Greekmath 0122}^2)^{K/2}$. Using $\hat {\Greekmath 011B}_{\Greekmath 0122}^2 = {\Greekmath 011B}_{\Greekmath 0122}^2 + O_p (n^{-1/2})$, in the appendix we show that
Next, consider a sequence $\varkappa _{n}\rightarrow 0$, which we will specify precisely later, and let
where $\hat \Sigma$ and $\hat B$ are consistent estimators of the asymptotic variance of $\hat {\Greekmath 010C}_L$ and of $B$, with $\hat \Sigma_{\ell \ell}$ denoting its $\ell$-th diagonal element, $\hat B_{\ell \cdot}$ denoting its $\ell$-th row, ${\overline g_x^{(K)} (\hat {\Greekmath 0112}) \equiv n^{-1} \sum_{i=1}^n g^{(K)}(X_i,S_i,\hat {\Greekmath 0112})}$, and $c_K > 0$ is a constant that we will calibrate below after we state the main theoretical result of this section.
If ${\Greekmath 010E}_n = 1$, the researchers should select $\hat {\Greekmath 010C}_L$. Otherwise, $\hat {\Greekmath 010C}_K$ should be used. The following lemma demonstrates that the proposed selection procedure has desirable theoretical properties.
The first part of Lemma (ref) shows that, provided that $\varkappa_n$ goes to zero sufficiently fast, $\hat {\Greekmath 010C}_L$ is selected (i.e., ${\Greekmath 010E}_n = 1$) only when its asymptotic bias is negligible. Next, note that $\hat {\Greekmath 010C}_L$ is asymptotically unbiased as long as ${\Greekmath 011C}_n = ( n^{-\frac{1}{2L+2}})$. The second part of the lemma shows that, for the suggested choices of $\varkappa_n$, the criterion is non-vacuous, i.e., that it does select the MERM estimator with a smaller expansion order $L$ when this is appropriate. Typically, one would pick $L=K-2$, so the lemma suggests taking $\varkappa _{n} = n^{-1/\left( 2K-2\right) }\left( \ln n\right) ^{-a}$ in this case.
In practice, it is important to pick an appropriate constant $c_K$ used in the construction of ${\Greekmath 010E}_n$ in equation (ref). When constructing ${\Greekmath 010E}_n$, we used $\left\vert {\Greekmath 010D}_{0K}\right\vert \propto {\Greekmath 011B}_{\Greekmath 0122}^K$, which motivates choosing $c_K = {\Greekmath 011B}_{\Greekmath 0122}^K / \left\vert {\Greekmath 010D}_{0K}\right\vert$. It is convenient to use a rule-of-thumb approach, using a reference distribution to determine $c_K$. It turns out that the normal distribution is not only convenient, but also sufficiently conservative (note that the bigger ${\Greekmath 010D}_{0K}$ is, the smaller $c_K$ is, resulting in a more conservative selection procedure). For example, suppose $K=4$, and consider using Student's $t({\Greekmath 0117} )$ distribution as the reference distribution. Then, for ${\Greekmath 0117} \geq 5$, the biggest $\left\vert {\Greekmath 010D}_{04}\right\vert/{\Greekmath 011B}_{\Greekmath 0122}^4$ and the smallest $c_4$ correspond to ${\Greekmath 0117} = \infty$ matching the normal distribution.\footnote{Note that Student's $t({\Greekmath 0117} )$ distribution does not have a finite $5$-th moment ${\Greekmath 0117} \leq 5$.}
Since for the normal distribution we have ${\Greekmath 010D}_{0K} = {\Greekmath 011B}_{\Greekmath 0122}^K / K!!$ for even $K$, we recommend using $c_K = K!!$ in equation (ref). Finally, while we recommend using normal distribution as the reference distribution, we also stress that the results of Lemma (ref) hold even if the measurement error is not normal or if it is skewed.
We summarize the proposed procedure in the following suggested algorithm.
To illustrate the performance of the algorithm provided above, we revisit the numerical experiment considered in Section (ref). We focus on choosing between $K=4$ and $L=2$, and, as in Section (ref), we report results for $n=2000$ in Table (ref) below. To ensure that the proposed algorithm performs well in a variety of sample sizes, we also consider $n=1000$ and $n=4000$ and report the corresponding results in Tables (ref) and (ref) in Appendix (ref).
We report the finite sample bias and RMSE for the naive MLE estimator, as well as for the MERM estimators using $K=2$ and $K=4$, and for the adaptive MERM estimator using data-driven $K$ following Algorithm (ref). Notice that for the considered sample sizes, $K=2$ is preferred when ${\Greekmath 011C} = 1/4$, and $K=4$ is preferred when ${\Greekmath 011C} = 3/4$. We find that in both of these regimes and for all sample sizes, the adaptive estimator using data-driven $K$ is essentially equivalent to the preferred estimators, i.e., our procedure selects the appropriate $K$. Interestingly, in the intermediate regime with ${\Greekmath 011C} = 1/2$, the adaptive estimator also has the smallest bias among the considered estimators, coming at the cost of a slightly bigger RMSE compared to the MERM estimator using $K=4$. Thus, the considered numerical experiment suggests that our algorithm has good finite sample properties supporting the findings of Lemma (ref).
It is easy to use the MERM framework to deal with multiple mismeasured variables. This is useful in many applications, including not only settings with multiple mismeasured covariates, but also settings with serially correlated measurement errors, settings where repeated measurements are available, and panel data models. Using the MERM approach is particularly advantageous in such applications, since it avoids nonparametric estimation of multivariate unobserved distributions.
Suppose $X_i^*$, ${\Greekmath 0122}_{i}$, and $X_i$ are $d \times 1$ vectors. Let ${\Greekmath 011C}_n \equiv \max_{j\le d} {\Greekmath 011B}_{{\Greekmath 0122}_j}/{\Greekmath 011B}_{X^*_j}$, where ${\Greekmath 011B}_{{\Greekmath 0122}_j}$ and ${\Greekmath 011B}_{X^*_j}$ denote the standard deviations of the $j$-th components of ${\Greekmath 0122}_i$ and $X_i^*$, so $\mathbb{E}\left[\left\vert {\Greekmath 0122}_{ij}\right\vert^k\right] = O({\Greekmath 011C}_n^k)$ for $k\in\{1,\ldots,K\}$.
For a $d \times 1$ vector of non-negative integers ${\Greekmath 0114} = ({\Greekmath 0114}_1,\ldots,{\Greekmath 0114}_d) \in \mathbb Z_+^{d}$, let
Also, for a positive integer $k$, let $\mathcal K_k = \{{\Greekmath 0114} \in \mathbb Z_+^d: \left\vert {\Greekmath 0114}\right\vert = k\}$. Then, we consider the following corrected moment function
where, with some abuse of notation, ${\Greekmath 010D}$ is a collection of all ${\Greekmath 010D}_{\Greekmath 0114}$ with ${\Greekmath 0114} \in \mathcal K_k$ and $k \in \{2, \dots, K\}$.
Under mild smoothness conditions
where the second equality holds provided that $O({\Greekmath 011C}_n^{K+1}) = o(n^{-1/2})$. Similarly to the scalar case, components of ${\Greekmath 010D}_{0}$ are determined by the moments of ${\Greekmath 0122}_i$. Specifically, let ${\Greekmath 0116}_{{\Greekmath 0114}} \equiv \mathbb{E}\left[{\Greekmath 0122}_{i 1}^{{\Greekmath 0114}_1} \dots {\Greekmath 0122}_{i d}^{{\Greekmath 0114}_d}\right]$, then
where ${\Greekmath 0114}! \equiv {\Greekmath 0114}_1! \ldots {\Greekmath 0114}_d!$. For $\left\vert {\Greekmath 0114}\right\vert \geq 4$, the coefficients can be computed by the following formulas. For example, for ${\Greekmath 0114} \in \mathcal K_4$, let $\mathcal K_{2, {\Greekmath 0114}} = \{\tilde {\Greekmath 0114} \in \mathcal K_2: {\Greekmath 0114} - \tilde {\Greekmath 0114} \in \mathcal K_2 \}$. Then,
More generally, for ${\Greekmath 0114} \in \mathcal K_{k}$ with $k \geq 4$, let $\mathcal K_{\ell, {\Greekmath 0114}} = \{\tilde {\Greekmath 0114} \in \mathcal K_\ell, {\Greekmath 0114} - \tilde {\Greekmath 0114} \in \mathcal K_{\left\vert {\Greekmath 0114}\right\vert - \ell} \}$ for $\ell \leq \left\vert {\Greekmath 0114}\right\vert - 2$. Then,