The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
111,916 characters
-1inSimple Estimation of Semiparametric Models with Measurement Errors
\title{\vspace*{-1in}\textsc{Simple Estimation of Semiparametric Models with
Measurement Errors}}
\author{Kirill S. \textsc{Evdokimov}\thanks{
Universitat Pompeu Fabra and Barcelona School of Economics: \textsf{
[email removed]}.} \and Andrei \textsc{Zeleneev}\thanks{
University College London: \textsf{[email removed]}.} \thanks{
First version: November 7, 2016. A part of the material of this paper was
previously circulated as a part of Evdokimov and Zeleneev (2018).} \thanks{
We thank the participants of the numerous seminars and conferences for
helpful comments and suggestions. We are also grateful to the Gregory C.
Chow Econometrics Research Program at Princeton University and the
Department of Economics at the Massachusetts Institute of Technology for
their hospitality and support. Evdokimov also gratefully acknowledges the
support from the National Science Foundation via grant SES-1459993, and from
the Spanish MCIN/AEI via grants RYC2020-030623-I, PID2019-107352GB-I00, PID2022-140825NB-I00, and Severo Ochoa Programme CEX2024-001476-S, funded by MICIU/AEI/10.13039/501100011033.}}
\date{This version: \today}
\maketitle
\begin{abstract}
We develop a practical way of addressing the Errors-In-Variables (EIV)
problem in the Generalized Method of Moments (GMM) framework. We focus on
the settings in which the variability of the EIV is a fraction of that of
the mismeasured variables, which is typical for empirical applications. For
any initial set of moment conditions our approach provides a ``corrected''
set of moment conditions that are robust to the EIV. We show that the GMM
estimator based on these moments is $\sqrt{n}$-consistent, with the standard
tests and confidence intervals providing valid inference. This is true even
when the EIV are so large that naive estimators (that ignore the EIV
problem) are heavily biased with their confidence intervals having 0\%
coverage. Our approach involves no nonparametric estimation, which is
especially important for applications with many covariates and settings
with multivariate EIV. In particular, the approach makes it
easy to use instrumental variables to address EIV in nonlinear models.
\bigskip\noindent\textbf{Keywords:} errors-in-variables, nonstandard
asymptotic approximation, nonparametric
identification, instrumental variables
\end{abstract}
\newpage
\section{Introduction}
Measurement errors are a common problem for empirical studies. Addressing
the Errors-In-Variables (EIV) bias in nonlinear models requires elaborate
strategies.\footnote{
See \cite
{HINP1991JoE,HausmanNeweyPowell1995JoE,Newey2001REStat,Schennach2007Ecta,Li2002JoE,Schennach2004Ecta,ChenHongTamer2005ReStud,HuSchennach2008Ecta,Schennach2014Ecta-ELVIS,Wilhelm2019WP-TestingForME}
, among others.} Despite the fundamental theoretical progress in
identification and estimation of nonlinear models with EIV, the problem of
EIV is still rarely addressed in empirical work outside of linear
specifications.
\setlength{\belowdisplayskip}{8.0pt plus 2.0pt minus 7.0pt}
\setlength{\abovedisplayskip}{6.0pt plus 2.0pt minus 5.0pt}
The goal of this paper is to develop a simple and practical approach to
estimation of nonlinear semiparametric models that can be expressed in the
form of general moment conditions
\begin{equation}
\mathbb{E}[g(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]=0\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ iff }{\Greekmath 0112} ={\Greekmath 0112} _{0}
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,} \label{eq:g descr moment}
\end{equation}
where $g\left( \cdot \right) $ is a vector of moment functions and ${\Greekmath 0112}
_{0}$ is the parameter vector of interest. The researcher has a random
sample of $\left\{ X_{i},S_{i}\right\} _{i=1}^{n}$, where scalar or vector $
X_{i}$ is a mismeasured version of unobserved $X_{i}^{\ast }$ with
measurement error ${\Greekmath 0122} _{i}$:
\begin{equation*}
X_{i}=X_{i}^{\ast }+{\Greekmath 0122} _{i}.
\end{equation*}
We will refer to $
g\left( \cdot \right) $ as the \emph{original} moment function, since it
would have been valid had the researcher observed $X_{i}^{\ast }$. A naive
GMM estimator (that ignores the EIV and uses $X_{i}$ in place of $
X_{i}^{\ast }$) based on $g\left( \cdot \right) $ is biased because $\mathbb{
E}[g(X_{i},S_{i},{\Greekmath 0112} _{0})]\neq 0$, in contrast to equation~(\ref{eq:g
descr moment}).
\begin{example*}[Nonlinear Regression, NLR]
\label{eg:NLR}Let $Y_{i}$ denote a scalar outcome, and let $X_{i}^{\ast }$
and $W_{i}$ be the covariates. Suppose
\begin{equation}
E\left[ Y_{i}|X_{i}^{\ast },W_{i}\right] ={\Greekmath 011A} \left( X_{i}^{\ast
},W_{i},{\Greekmath 0112} _{0}\right) \label{eq:Intro:eg NLR}
\end{equation}
for some function ${\Greekmath 011A} $ known up to the parameter ${\Greekmath 0112} $. For example,
in the Logit model, $Y_{i}$ is binary, ${\Greekmath 011A} \left( x,w,{\Greekmath 0112} \right)
\equiv 1\left/ \left( 1+\exp \left( -\left( {\Greekmath 0112} _{x}^{\prime }x+{\Greekmath 0112}
_{w}^{\prime }w\right) \right) \right) \right. $, and ${\Greekmath 0112} \equiv \left(
{\Greekmath 0112} _{x}^{\prime },{\Greekmath 0112} _{w}^{\prime }\right) ^{\prime }$.
Suppose the researcher has an instrumental variable $Z_{i}$. Then, they can
use
\begin{equation*}
g\left( y,x,w,z;{\Greekmath 0112} \right) \equiv \left( y-{\Greekmath 011A} \left( x,w,{\Greekmath 0112}
\right) \right) {\Greekmath 0127} \left( x,w,z\right)
\end{equation*}
as the original moment function, where ${\Greekmath 0127} \left( x,w,z\right) $ is a
vector that, for example, can include $x$, $z$, $w$, their powers and/or
interactions.\footnote{
Note that the moment condition~(\ref{eq:g descr moment}) is stated in terms
of the true (correctly measured) $X_{i}^{\ast }$. Determining what functions
$g\left( \cdot \right) $ (or $h\left( \cdot \right) $ in the NLR model)
satisfy this moment condition does not involve any consideration of the
measurement errors and hence is straightforward.} Here $S_{i}=\left(
Y_{i},W_{i},Z_{i}\right) $. \hfill\ensuremath{\blacksquare}
\end{example*}
Even in this well-studied example of nonlinear regression, estimation in the
presence of the EIV is a difficult problem. Importantly, nonlinear
instrumental variable regression estimator cannot be used, since it is
inconsistent in the presence of EIV \citep{Amemiya1985JoE}. The existing
approaches typically require nonparametric estimation that can be
impractical in many empirical applications. In contrast, in this paper, we
develop an alternative class of estimators, that are essentially GMM
estimators that modify the original moment functions $g\left( \cdot \right) $
in a way that makes the moment conditions robust to the EIV. In particular,
our approach makes it easy to use instrumental variables to address EIV\ in
nonlinear models.
To provide a practical estimation approach\ for the general class of models~(
\ref{eq:g descr moment}), we focus on empirical settings in which the
researcher believes the variability of the measurement error to be at most a
fraction of the variability of the mismeasured variable, i.e., the
noise-to-signal ratio ${\Greekmath 011C} \equiv {\Greekmath 011B} _{{\Greekmath 0122} }/{\Greekmath 011B} _{X^{\ast
}} $ to be moderate. {}
The absolute magnitude of the measurement error ${\Greekmath 011B} _{{\Greekmath 0122} }$
does not need to be small. Existing validation studies provide insights into
the magnitude of ${\Greekmath 011C} $ for some key economic variables and datasets.
\citet{BoundKrueger1991JoLaborEcon} consider log-earnings in the Current
Population Survey (CPS) data matched to the Social Security payroll records.
Their estimates of the variance of the measurement errors correspond to $
{\Greekmath 011C} $ of approximately $0.47$ and $0.30$ for the subsamples of men and
women, respectively. \citet{BoundBrownDuncanRodgers1994JoLaborEcon} consider
a validation study of Panel Study of Income Dynamics (PSID). Their estimates
imply ${\Greekmath 011C} =0.39-0.66$ for log-earnings and ${\Greekmath 011C} =0.63-0.76$ for hours
worked. \citet{Pischke1995JBES} estimates correspond to ${\Greekmath 011C} =0.39-0.50$
for the log-earnings in PSID. \citet{AshenfelterKrueger1994AER} assess the
mismeasurement in the years of education; their estimates correspond to $
{\Greekmath 011C} =0.30-0.37$.
Focusing on these settings allows us to isolate the most important aspects
of the problem and, as result, to develop a simple estimator, which does not
require any nonparametric estimation or simulation. Such simple estimation
becomes possible because in these settings we can obtain a simple
approximation of the EIV\ bias of the moment conditions as a function of $
{\Greekmath 0112} $.
We propose to bias correct the original moments $g\left( \cdot \right) $,
which in turn removes the bias of the corresponding estimator of ${\Greekmath 0112}
_{0} $. This bias correction depends on some moments of the distribution of
the measurement errors that are unknown. Another difficulty is that the
estimators of some components of the bias correction themselves may need to
be bias corrected. To address these issues, we develop the \emph{corrected
moment conditions}, which depend on ${\Greekmath 0112} $ and additional {}parameters ${\Greekmath 010D} $ that govern the bias correction. The true
parameter value ${\Greekmath 010D} _{0}$ is associated with (possibly conditional)
low-order moments of ${\Greekmath 0122} _{i}$. Despite some theoretical subtleties
with the construction of the corrected moment conditions, their practical
implementation is straightforward and they can be automatically computed for
any original moment function $g\left( \cdot \right) $.
We introduce the Measurement Error Robust Moments (MERM) estimator, which is
a GMM estimator that uses the corrected moment conditions to jointly
estimate parameters ${\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{0}$. The estimator can be
computed using any standard software for GMM estimation. Joint estimation of
parameters ${\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{0}$ using the corrected moment
conditions effectively robustifies moment conditions $g\left( \cdot \right) $
against the impact of the measurement errors.
To make these ideas precise and to study the properties of the proposed
estimators, we develop an asymptotic theory using a nonstandard asymptotic
approximation that models ${\Greekmath 011C} $ as slowly shrinking with the sample size.
Standard asymptotics considers ${\Greekmath 011C} $ to be constant, which implies that as
$n\rightarrow \infty $ the bias of a naive estimator dwarfs its sampling
variability: the bias is constant while the standard errors shrink
proportionally to $1/\sqrt{n}$. As a result, under the standard asymptotics,
the problem of removing the EIV bias becomes central in the analysis, with
relatively little attention paid to the sampling variability of estimators.
However, this focus does not seem to be appropriate in many empirical
applications, in which the researcher does not expect the potential EIV bias
to be several orders of magnitude larger than the standard errors.\footnote{
Such empirical settings appear to be widespread. Although the concerns about
measurement errors are often raised, the majority of applied work does not
explicitly correct the EIV bias in nonlinear models, and instead implicitly
or explicitly argues or conjectures that the EIV bias is likely not to be
too large.
} By considering ${\Greekmath 011C} $ as drifting towards zero with the sample size, our
approach provides a better guidance on construction of EIV robust estimators
with good finite sample properties when ${\Greekmath 011C} $ is small or moderate.
\footnote{
Nonstandard asymptotic approximations with drifting parameters are often
used to obtain better approximations of the finite sample behavior of
estimators and tests. For example, in the instrumental variable regression
settings, to consider the settings with relatively small first stage
coefficients, \cite{StaigerStock1997} model them as shrinking with $n$. It
is important to keep in mind that such nonstandard asymptotic approximations
are merely mathematical tools. One should not take them literally and think
of parameters somehow changing if more data is collected. {}}{}
Using this approximation, we show that the proposed estimation approach
indeed addresses the EIV problem. The MERM estimator is shown to be $\sqrt{n}
$-consistent and asymptotically normal and unbiased. The standard confidence
intervals and tests for GMM estimators are also valid for the MERM\
estimator. Additionally, the standard GMM arsenal of assessment tools can be
applied to the MERM estimator, allowing one to test model identification,
conduct valid inference, and perform model specification diagnostics.
The usefulness of a large sample theory is measured by its ability to
approximate the finite sample properties of the estimators and inference
procedures. Thus, we study the MERM estimators in a variety of simulation
experiments. The results confirm that the nonstandard asymptotic theory
indeed provides a good approximation of the finite sample properties of the
estimators even in the settings with relatively large EIV. In some of the
simulation experiments, the EIV are so large that for the naive estimators'
standard $95\%$ confidence intervals have actual coverages of $0\%$ in
finite samples, due to the magnitude of the EIV\ bias. At the same time,
even in these settings the MERM estimators perform well, removing the EIV
bias and providing confidence intervals with the correct coverage. In
particular, the simulation results show that despite the simplicity of
implementation, the MERM estimators can compete with and outperform
semi-nonparametric estimators.{}
The MERM estimator is structurally different from the existing approaches
that require nonparametric estimation of some nuisance parameters, for
example, of the density $f_{X^{\ast }|Z,W}$. Avoiding nonparametric
estimation has at least two advantages. First, since the majority of
empirical applications include additional covariates $
W_{i}$, nonparametric estimation is often infeasible due to the curse of
dimensionality. Because the MERM estimator does not involve any
nonparametric estimation, it can be used in applications with a relatively
large number of additional covariates $W_{i}$, and remains feasible even in
the more complicated settings, including multi-equation and structural
models, and applications with multiple mismeasured variables $X_{i}$.
Second, estimation of infinite-dimensional nuisance parameters is typically
more demanding towards the sources of identification available in the data,
for example, requiring an instrumental variable with a large support
(continuously distributed). In contrast, having a discrete instrument is
sufficient for the MERM approach because the nuisance parameter ${\Greekmath 010D} _{0}$
is finite-dimensional.
For example, in Section~\ref{sec:empirical} we consider estimation of the
model of multinomial choice among three modes of transportation. A leading
alternative approach to the errors-in-variables problem in this model is the
semi-nonparametric sieve-MLE estimator advocated by
\citet{HuSchennach2008Ecta,CarrollChenHu2010JoNS}, among others. This
approach requires, among other things, estimating the conditional density $
f_{X^{\ast }|Z,W}$ of $X_{i}^{\ast }$ given the instrument $Z_{i}$ and
covariates $W_{i}$. In this empirical example, $W_{i}$ includes four
continuously distributed covariates (two continuously distributed
characteristics per choice) and a discrete one, while scalar $X_{i}^{\ast }$
and $Z_{i}$ are also continuous. Thus, $f_{X^{\ast }|Z,W}$ is a function of
six continuous and one discrete variable. Hence, for typical sample sizes,
estimating $f_{X^{\ast }|Z,W}$ in this example is infeasible due to the
curse of dimensionality. In contrast, as the results of Section~\ref
{sec:empirical} demonstrate, the MERM approach is practical and effective in
this application, in part because
it avoids estimation of the high-dimensional nuisance functions like $f_{X^{\ast }|Z,W}$ altogether.
The simplicity and practicality of the MERM\ approach do come at a cost:
there is a limit on the magnitude of the measurement errors it can handle.
For example, one generally should not expect the MERM\ approach to work well
when ${\Greekmath 011C} >1$, i.e., when the noise dominates the signal; in this case the
researcher should seek an alternative estimation method.
We view the MERM\ approach as providing a bridge between the settings in
which the measurement errors are guaranteed to be absent or negligible, and
the settings where the measurement errors are so large that one has to use
the relatively more complicated estimators from the earlier literature (if
they exist at all for the model of interest).
\smallskip
{\noindent \textbf{Related Literature}} \cite{ChenHongNekipelov2011JEL},
\cite{Schennach2016AnnRev}, and \cite{Schennach2020HB-ME} provide excellent
overviews of the measurement error literature.{}
The existing semiparametric approaches to estimation and inference in models
with EIV\ involve nonparametric estimation of infinite-dimensional nuisance
parameters (e.g., \QTR{citealp}{
Chesher2000WP,Li2002JoE,Schennach2004Ecta,Schennach2007Ecta,HuSchennach2008Ecta,SchennachHu2013JASA,Song2015JoE
}), simulation (e.g., \QTR{citealp}{Schennach2014Ecta-ELVIS}), or both
(e.g., \QTR{citealp}{Newey2001REStat,WangHsiao2011JoE}$)$. The exceptions
include models with linear and polynomial regression functions (see
\QTR{citealp}{HINP1991JoE,HausmanNeweyPowell1995JoE}$)$, and Gaussian
control variable models such as Probit and Tobit with endogeneity (see
\QTR{citealp}{SmithBlundell1986Ecta,RiversVuong1988JoE}).
To the best of our knowledge, this paper is the first to provide an approach
for $\sqrt{n}$-consistent and asymptotically normal and unbiased estimation
of general GMM models with EIV that does not require any nonparametric
estimation (or simulation).
We are able to provide such an estimator because we focus on the models with
moderate measurement errors. Modeling the variance of the measurement error
as shrinking to zero with the sample size is a popular approach in
Statistics. The method has been proposed by \cite{WolterFuller1982AS}, who
used it to construct an approximate MLE\ estimator of a nonlinear regression
model with Gaussian errors. Following their approach, the Statistics
literature has mainly focused on the settings where the moments of the EIV
needed to bias correct the estimators are either known or can be directly
estimated from the available data such as repeated measurements (e.g.,
\QTR{citealp}{CarrollStefanski1990JASA,CarrollEtAl2006Book-ME}). In
Economics, such data are relatively rare. The use of approximations with
shrinking variance of measurement errors in Econometrics literature has been
pioneered by \cite{Kadane1971Ecta}, \cite{Amemiya1985JoE}, and \cite
{Chesher1991Biomet}. Such approximations have been used to check the
sensitivity of naive estimators to the EIV by considering how the estimates
change as the unknown moments of the measurement errors vary within some set
of plausible values, e.g., see \cite{ChesherSchluter2002ReStud}, \cite
{ChesherDumanganeSmith2002JoE}, \cite{BattistinChesher2014JoE}, \cite
{Chesher2017JoE}, and \cite{HongTamer2003JoE}.
\citet{BoundBrownMathiowetz2001HBoE} review a broad list of validation
studies matching standard economic dataset to administrative records. The
estimates they report suggest that the measurement errors of moderate
magnitude are typical for empirical applications. This suggests that the
approach developed in this paper could prove valuable for a wide range of
applied work.
This paper differs from the earlier literature in several ways. First, it
presents a way to estimate the unknown nuisance parameters (moments of the
measurement errors) jointly with the parameters of interest. As a result,
the approach can, for example, use instrumental variables as a source of
identification. {} Second, the method applies to
a very general class of semiparametric models specified by moment
conditions. Third, the MERM approach allows the measurement errors to have
larger magnitudes than most of the papers in the earlier literature; this is
achieved by the MERM approach recursively bias correcting the bias
correction terms.
The most widespread approach to identification of the EIV\ models in
economic applications is to use instrumental variables, e.g., see \cite
{HINP1991JoE,Newey2001REStat,Schennach2007Ecta,WangHsiao2011JoE}. In a
recent paper, \cite{HahnHausmanKim2021EL} reconsider the regression model in
\cite{Amemiya1990JoE} using a bias correction similar to ours. When proper
excluded variables are not available, researchers have considered using
higher moments of $X_{i}$ as instruments, e.g., see \cite
{Reiersol1950Ecta,Lewbel1997Ecta,EricksonWhited2002ET,SchennachHu2013JASA,BenMosheDHaultfeuilleLewbel2017JoE}
. When available, repeated measurements can also be used to identify the
model, e.g., see \cite
{HINP1991JoE,LiVuong1998JoE,Li2002JoE,Schennach2004Ecta}. The MERM estimator
accommodates these identification approaches within a unified estimation
framework.
The power of the general MERM\ approach can be illustrated in the NLR model.
For example, when a candidate instrumental variable is available, the
conditions it needs to satisfy are much weaker than what is required by many
existing approaches. Availability of a discrete instrument is sufficient for
identification; and the instrument is allowed to have heterogeneous impact
on covariates $X_{i}^{\ast }$.
\ One can also take a nonclassical, nonlinear (e.g., discretized or
censored), or biased measurement of $X_{i}^{\ast }$ as an instrument in the
MERM approach. We discuss identification in Section~\ref{ssec:MERM-ID}. In
addition, in a related paper \cite{EvdokimovZeleneev2022WP-NPID} study
nonparametric regression with EIV using the ${\Greekmath 011C} \rightarrow 0$
approximation, and demonstrate that the MERM approach can also be motivated
from a nonparametric perspective.
\cite
{KitamuraOtsuEvdokimov2013Ecta,AndrewsGentzkowShapiro2017QJE,ArmstrongKolesar2021QE,BonhommeWeidner2021QE}
, among others, develop tools for estimation and inference in GMM, which are
robust to general perturbation or misspecification of the true data
generating process. They focus on the settings in which these perturbations
are sufficiently small, so that naive estimators remain $\sqrt{n}$
-consistent, and their biases are of the same order of magnitude as their
standard errors. In contrast, we focus on more specific forms of data
contamination due to the EIV. This allows the MERM approach to remain valid
even in the settings with larger measurement errors, in which naive
estimators may have slower than $\sqrt{n}$ rates of convergence.
The MERM approach also provides a useful foundation for dealing with EIV in
more complicated settings. \cite{EvdokimovZeleneev2018WP-Inference} utilize
the MERM framework to address an issue of nonstandard inference, which turns
out to arise generally when EIV models are identified using instrumental
variables. \cite{EvdokimovZeleneev2019WP-Panel} extend the analysis of this
paper to long panel and network settings.
\smallskip
{\noindent \textbf{Organization of the paper}} Section~\ref{sec:framework}
introduces the Moderate Measurement Error framework and the proposed MERM\
estimator. Section \ref{sec:MC} presents several Monte Carlo experiments
that illustrate finite sample properties of the MERM\ estimators. Section
\ref{sec:Extensions} considers several extensions of the framework. A supplementary appendix contains all proofs and
additional results for the numerical and empirical illustrations.
\section{Moderate Measurement Errors Framework\label{sec:framework}}
To present the main ideas we first consider the case of univariate $
X_{i}^{\ast }$. We will consider multivariate $X_{i}^{\ast }$ later. We assume that the measurement error is classical, i.e., that ${\Greekmath 0122}_i$ is independent of $X_i^*$ and $S_i$; later we will discuss how this assumption can be relaxed. Following the rest of the literature, we assume
that $\mathbb{E}\left[ {\Greekmath 0122} _{i}\right] =0$.\footnote{
A location normalization such as $\mathbb{E}\left[ {\Greekmath 0122} _{i}\right]
=0 $ is usually necessary because it is not possible to separately identify
the means $\mathbb{E}\left[ X_{i}^{\ast }\right] $ and $\mathbb{E}\left[
{\Greekmath 0122} _{i}\right] $.}
To develop a practical estimation approach for general moment condition
models we focus on the settings in which ${\Greekmath 011C} \equiv {\Greekmath 011B} _{{\Greekmath 0122}
}/{\Greekmath 011B} _{X^{\ast }}$ is small or moderate. We consider an asymptotic
approximation with ${\Greekmath 011C} _{n}\equiv {\Greekmath 011C} \rightarrow 0$ as $n\rightarrow
\infty $.
Note that economically meaningful parameters are usually invariant
to rescaling of~$X_{i}^{\ast }$. Likewise, the extent of the EIV problem
does not change with such rescaling. For simplicity of exposition, it is convenient to
assume that $X_{i}^{\ast }$ is scaled so that ${\Greekmath 011B} _{X^{\ast }}$ is of
order one and, correspondingly, the moments $\mathbb{E}[\left\vert {\Greekmath 0122}
_{i}\right\vert ^{k}]\propto {\Greekmath 011C} _{n}^{k}$ decrease with $k$ when ${\Greekmath 011C}
_{n}<1$. For example, this could be ensured by normalizing observed $X_{i}$
to have ${\Greekmath 011B} _{X}=1$. Let us stress that this normalization is used only
to simplify the exposition; as we show in Appendix~\ref{sec:MME expansion
with large ME}, the proposed MERM\ estimator does not require any
normalizations in practice.
\subsection{Special Case: Quadratic Expansion}
For clarity, we first consider a simple special case of the general
approach. Let us denote $g_{x}^{(k)}\left( x,s,{\Greekmath 0112} \right) \equiv
\partial ^{k}g\left( x,s,{\Greekmath 0112} \right) /\partial x^{k}$. Since $\mathbb{E}
[\left\vert {\Greekmath 0122} _{i}\right\vert ^{k}]\propto {\Greekmath 011C}
_{n}^{k}\rightarrow 0$ as $n\rightarrow \infty $, under some regularity
conditions, we can write the quadratic Taylor expansion of function $
g(X_{i},S_{i},{\Greekmath 0112} )=g(X_{i}^{\ast }+{\Greekmath 0122} _{i},S_{i},{\Greekmath 0112} )$
around ${\Greekmath 0122} _{i}=0$ as
\begin{align}
\mathbb{E}[g(X_{i},S_{i},{\Greekmath 0112} )]& =\mathbb{E}\left[ g(X_{i}^{\ast
},S_{i},{\Greekmath 0112} )+g_{x}^{(1)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} ){\Greekmath 0122} _{i}+
\dfrac{1}{2}g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} ){\Greekmath 0122} _{i}^{2}
\right] +O(\mathbb{E}[\left\vert {\Greekmath 0122} _{i}\right\vert ^{3}]) \notag
\\
& =\mathbb{E}[g(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]+\dfrac{\mathbb{E}[{\Greekmath 0122}
_{i}^{2}]}{2}\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]+O({\Greekmath 011C}
_{n}^{3}), \label{eq:moment expectation for K=2}
\end{align}
where the second equality holds because ${\Greekmath 0122} _{i}$ and $
\left( X_{i}^{\ast },S_{i}\right) $ are independent, and ${\mathbb{E}\left[
{\Greekmath 0122} _{i}\right] =0}$.
Evaluating the expansion above at ${\Greekmath 0112} = {\Greekmath 0112}_0$ gives $\mathbb{E}[g(X_{i},S_{i},{\Greekmath 0112} _{0})]=O\left(
{\Greekmath 011B} _{{\Greekmath 0122} }^{2}\right) =O({\Greekmath 011C} _{n}^{2})$, because $\mathbb{E}[g(X_{i}^*,S_{i},{\Greekmath 0112} _{0})] = 0$. As a result, a naive
estimator that ignores the EIV and uses $X_{i}$ in place of $X_{i}^{\ast }$
has EIV\ bias of order ${\Greekmath 011C} _{n}^{2}$.\footnote{
For example, consider a linear regression with a scalar mismeasured
regressor. The bias of the naive OLS\ estimator of the slope parameter $
{\Greekmath 0112} _{01}$ is $-{\Greekmath 0112} _{01}\frac{{\Greekmath 011C} _{n}^{2}}{1+{\Greekmath 011C} _{n}^{2}}=-{\Greekmath 0112}
_{01}{\Greekmath 011C} _{n}^{2}+O\left( {\Greekmath 011C} _{n}^{4}\right) $.} The bias of the naive
estimator should be compared with its standard error, which is of order $
n^{-1/2}$. Thus, the bias of the naive estimator is not negligible, unless the
measurement error is rather small (theoretically, unless ${\Greekmath 011C}
_{n}^{2}=o\left( n^{-1/2}\right) $). In particular, tests and confidence
intervals based on the naive estimator are invalid and can provide highly
misleading results. Moreover, if ${\Greekmath 011C} _{n}^{2}$ shrinks at a rate slower
than $O\left( n^{-1/2}\right) $, {} the rate of convergence of the naive estimator is slower
than $\sqrt{n}$.
Suppose ${\Greekmath 011C} _{n}=o\left( n^{-1/6}\right) $. Then, $O({\Greekmath 011C}
_{n}^{3})=o\left( n^{-1/2}\right) $ and we can rearrange equation~(\ref
{eq:moment expectation for K=2}) as
\begin{equation}
\mathbb{E}[g(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]=\mathbb{E}[g(X_{i},S_{i},{\Greekmath 0112}
)]-\frac{\mathbb{E}[{\Greekmath 0122} _{i}^{2}]}{2}\mathbb{E}\left[
g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )\right] +o(n^{-1/2}).
\label{eq:g MME mu for K=2}
\end{equation}
The left-hand side of this equation is exactly the moment condition~(\ref
{eq:g descr moment}) that we would like to use for estimation of ${\Greekmath 0112}
_{0} $. The first term on the right-hand side involves only observed
variables, and can be estimated by the sample average $\overline{g}({\Greekmath 0112}
)\equiv n^{-1}\sum_{i=1}^{n}g(X_{i},S_{i},{\Greekmath 0112} ).$ The second term on the
right-hand side can be thought of as a bias correction that removes the
EIV-bias from the expected moment function $\mathbb{E}[g(X_{i},S_{i},{\Greekmath 0112}
)]$.
The idea of the MERM\ estimator we propose is to make use of expansions such
as~(\ref{eq:g MME mu for K=2}) to bias correct the moment condition $\mathbb{
E}[g(X_{i},S_{i},{\Greekmath 0112} )]$, which in turn removes the bias of the estimator
of the parameters of interest ${\Greekmath 0112} _{0}$. To perform the bias correction
we need to estimate two quantities: $\mathbb{E}[{\Greekmath 0122} _{i}^{2}]$ and $
\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$. {}
First, we show that in equation~(\ref{eq:g MME mu for K=2}) we can
substitute $\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$ with $
\mathbb{E}[g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]$, which in turn can be
estimated by $\overline{g}_{x}^{(2)}({\Greekmath 0112} )\equiv
n^{-1}\sum_{i=1}^{n}g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )$. By the Taylor
expansion around ${\Greekmath 0122} _{i}=0$ similar to equation~(\ref{eq:moment
expectation for K=2}), we can show that $\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast
},S_{i},{\Greekmath 0112} )]=\mathbb{E}[g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]+O({\Greekmath 011C}
_{n}^{2})$ and hence
\begin{equation}
\frac{1}{2}\mathbb{E}[{\Greekmath 0122} _{i}^{2}]\left( \mathbb{E}\left[
g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )\right] -\mathbb{E}\left[
g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )\right] \right) =\mathbb{E}[{\Greekmath 0122}
_{i}^{2}]O\left( {\Greekmath 011C} _{n}^{2}\right) =O\left( {\Greekmath 011C} _{n}^{4}\right) .
\label{eq:bias replacing Xs with X for K=2}
\end{equation}
Here $O\left( {\Greekmath 011C} _{n}^{4}\right) =o\left( n^{-1/2}\right) $ because we
assume that ${\Greekmath 011C} _{n}=o\left( n^{-1/6}\right) $. The idea behind this
substitution is that the bias of order $O({\Greekmath 011C} _{n}^{2})$ in $\mathbb{E}
[g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]$ can be ignored because it is multiplied
by $E\left[ {\Greekmath 0122} _{i}^{2}\right] =O\left( {\Greekmath 011C} _{n}^{2}\right) $.
\footnote{
Such substitutions of $X^{\ast }$ with $X$ have been used in other contexts,
e.g., \cite{ChesherSchluter2002ReStud}.} With the substitution, we can
rearrange equation~(\ref{eq:g MME mu for K=2}) and write it as
\begin{equation}
\mathbb{E}[g(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]=\mathbb{E}\left[
g(X_{i},S_{i},{\Greekmath 0112} )-\frac{\mathbb{E}[{\Greekmath 0122} _{i}^{2}]}{2}
g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )\right] +o(n^{-1/2}).
\label{eq:expansion - almost psi for K=2}
\end{equation}
Second, we propose estimating the unknown $\mathbb{E}[{\Greekmath 0122} _{i}^{2}]$
together with the parameter of interest ${\Greekmath 0112} $. Specifically, let ${\Greekmath 010D}
_{02}\equiv \mathbb{E}[{\Greekmath 0122} _{i}^{2}]/2$ denote the true value of
parameter ${\Greekmath 010D} _{2}$, and consider the following \emph{corrected moment
function}:
\begin{equation}
{\Greekmath 0120} (X_{i},S_{i},{\Greekmath 0112} ,{\Greekmath 010D} )\equiv g(X_{i},S_{i},{\Greekmath 0112} )-{\Greekmath 010D}
_{2}g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} ).
\label{eq: psi moms definition for K=2}
\end{equation}
Function ${\Greekmath 0120} $ is a moment function parameterized by ${\Greekmath 0112} $ and ${\Greekmath 010D}
$, and
\begin{equation}
\mathbb{E}[{\Greekmath 0120} (X_{i},S_{i},{\Greekmath 0112} _{0},{\Greekmath 010D} _{02})]=\mathbb{E}
[g(X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0})]+o\left( n^{-1/2}\right) =o\left(
n^{-1/2}\right) , \label{eq:E psi = o(n-12) for K=2}
\end{equation}
where the first equality follows from equation~(\ref{eq:expansion - almost
psi for K=2}) and the definition of ${\Greekmath 010D} _{02}$, and the second equality
follows from equation~(\ref{eq:g descr moment}). Hence, the corrected moment
conditions ${\Greekmath 0120} $ can be used to jointly estimate the true parameters $
{\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{02}$ by a GMM\ estimator.\footnote{
In the moment condition settings, having $o\left( n^{-1/2}\right) $ is
equivalent to having $0$ on the right-hand side\ of equation~(\ref{eq:E psi
= o(n-12) for K=2}).}
\begin{remark}
\label{remark: symmetric ME} If $\mathbb{E}[{\Greekmath 0122} _{i}^{3}]=0$ (e.g.,
if the distribution of ${\Greekmath 0122} _{i}$ is symmetric), the remainder in
equation~(\ref{eq:moment expectation for K=2}) is of a smaller order $O({\Greekmath 011C}
_{n}^{4})$. Hence, the corrected moments~(\ref{eq:E psi = o(n-12) for K=2})
remain valid for larger values of ${\Greekmath 011C} _{n}$, requiring only the weaker
condition ${\Greekmath 011C} _{n}=o(n^{-1/8})$. The bias of the naive estimators in this
case can be as large as $o(n^{-1/4})$.
\end{remark}
\subsection{General Case: Expansion of order $K$}
The quadratic expansion of equation~(\ref{eq:moment expectation for K=2})
can be extended to general order $K\geq 2$. Considering larger $K$
theoretically allows ${\Greekmath 011C} _{n}$ converging to zero at a slower rate. In
finite samples this corresponds to the asymptotics providing good
approximations for larger values of ${\Greekmath 011C} _{n}$, i.e., large measurement
errors. {} Expanding $g(X_{i}^{\ast
}+{\Greekmath 0122} _{i},S_{i},{\Greekmath 0112} )$ around ${\Greekmath 0122} _{i}=0$ we have,
\begin{equation}
\mathbb{E}[g(X_{i},S_{i},{\Greekmath 0112} )]=\mathbb{E}\left[ g(X_{i}^{\ast
},S_{i},{\Greekmath 0112} )+\sum_{k=1}^{K}\frac{{\Greekmath 0122} _{i}^{k}}{k!}
g_{x}^{(k)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )\right] +O\left( \mathbb{E}\left[
\left\vert {\Greekmath 0122} _{i}\right\vert ^{K+1}\right] \right) .
\label{eq: moment exp approx - step 1}
\end{equation}
The above special case of quadratic expansion corresponds to $K=2$.
The approximation we consider is formalized by the following assumption.
\begin{assumption}[MME]
\emph{(Moderate Measurement Errors)} \namedlabel{ass:MME}{MME}
(i) ${\Greekmath 011C} _{n}=o(n^{-1/\left( 2K+2\right) })$ for some integer $K\geq 2$;
and (ii) $\mathbb{E}[\left\vert {\Greekmath 0122} _{i}\right\vert ^{L}] \leq C
{\Greekmath 011B}_{\Greekmath 0122}^L $ for some $L\geq K+1$ and $C > 0$.
\end{assumption}
Assumption~\ref{ass:MME}(i) limits the magnitude of the measurement errors
and implies that ${\Greekmath 011C} _{n}^{K+1}=o\left( n^{-1/2}\right) $. Assumption~\ref
{ass:MME}(ii) implies that $\mathbb{E}[\left\vert {\Greekmath 0122}
_{i}\right\vert ^{k}]=O\left( {\Greekmath 011B} _{{\Greekmath 0122} }^{k}\right) $, and
requires the tails of ${\Greekmath 0122} _{i}/{\Greekmath 011B} _{{\Greekmath 0122} }$ to be
sufficiently thin. Together, parts (i) and (ii) imply that $\mathbb{E}
[\left\vert {\Greekmath 0122} _{i}\right\vert ^{K+1}]=O\left( {\Greekmath 011C}
_{n}^{K+1}\right) =o\left( n^{-1/2}\right) $, and hence ensure that the
remainder in equation~(\ref{eq: moment exp approx - step 1})$\ $is
negligible. Using $\mathbb{E}\left[ {\Greekmath 0122} _{i}|X_{i}^{\ast },S_{i}
\right] =0$ to further simplify this expansion and rearranging the terms we
obtain
\begin{equation}
\mathbb{E}[g(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]=\mathbb{E}[g(X_{i},S_{i},{\Greekmath 0112}
)]-\sum_{k=2}^{K}\frac{\mathbb{E}[{\Greekmath 0122} _{i}^{k}]}{k!}\mathbb{E}\left[
g_{x}^{(k)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )\right] +o(n^{-1/2}).
\label{eq:g MME mu}
\end{equation}
This equation is the general expansion analog of equation~(\ref{eq:g MME mu
for K=2}). The summation on the right hand side is the bias correction term,
{}
which we use to construct the MERM\ estimator.\footnote{It is useful to get a sense of the
magnitudes of the coefficients $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{k}\right]
/k!$ in equation~(\ref{eq:g MME mu}). Suppose ${\Greekmath 0122} _{i}\sim N\left(
0,{\Greekmath 011B} _{{\Greekmath 0122} }^{2}\right) $, ${\Greekmath 011B} _{{\Greekmath 0122} }=0.5$, and $
{\Greekmath 011B} _{X^{\ast }}=1$, so$\ {\Greekmath 011C} ={\Greekmath 011B} _{{\Greekmath 0122} }=0.5$. Then the
coefficients in front of $g_{x}^{\left( 2\right) }$, $g_{x}^{\left( 4\right)
}$, and $g_{x}^{\left( 6\right) }$ are $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{2}
\right] /2!=0.125$, $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{4}\right] /4!\approx
0.008$, and $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{6}\right] /6!\approx 0.0003$
. {}} {}
It turns out that for $K\geq 4$, estimation of $\mathbb{E}
[g_{x}^{(k)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$ is more intricate than in the
case of $K=2$, and the substitution we made in equation~(\ref{eq:expansion -
almost psi for K=2}) no longer works. Larger values of $K$ allow for larger
values of ${\Greekmath 011C} _{n}$ and hence larger EIV\ biases of naive estimators $
n^{-1}\sum_{i=1}^{n}g_{x}^{(k)}(X_{i},S_{i},{\Greekmath 0112} )$. The expansion of
order $K$ includes terms up to the order ${\Greekmath 011C} _{n}^{K}$, with the
asymptotically negligible remainder of order $O\left( {\Greekmath 011C} _{n}^{K+1}\right)
$. For $K\geq 4$, terms of order ${\Greekmath 011C} _{n}^{4}$ are not negligible. This
implies that we cannot ignore the EIV bias that would arise from
substituting $\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} )]$ with $
\mathbb{E}[g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]$ in equation~(\ref{eq:g MME mu}
), because this bias is of order $O\left( {\Greekmath 011C} _{n}^{4}\right) $ according
to equation~(\ref{eq:bias replacing Xs with X for K=2}). To address this
problem, we instead replace $\mathbb{E}[g_{x}^{(2)}(X_{i}^{\ast
},S_{i},{\Greekmath 0112} )]$ with the bias corrected expression $\mathbb{E}
[g_{x}^{(2)}(X_{i},S_{i},{\Greekmath 0112} )]-\left( \mathbb{E}[{\Greekmath 0122}
_{i}^{2}]/2\right) \mathbb{E}[g_{x}^{(4)}(X_{i},S_{i},{\Greekmath 0112} )]$. Thus, for $
K\geq 4$, one needs to bias correct the estimator of the bias correction
term. Moreover, for larger $K$ one needs to bias correct the bias correction
of the bias correction term and so on.
Fortunately, we show that these bias corrections can be constructed as
linear combinations of the expectations of the higher order derivatives of $
g_{x}^{\left( k\right) }(X_{i},S_{i},{\Greekmath 0112} )$. Let us define the following
\emph{corrected moment function}:
\begin{equation}
{\Greekmath 0120} (X_{i},S_{i},{\Greekmath 0112} ,{\Greekmath 010D} )\equiv g(X_{i},S_{i},{\Greekmath 0112}
)-\sum_{k=2}^{K}{\Greekmath 010D} _{k}g_{x}^{(k)}(X_{i},S_{i},{\Greekmath 0112} ),
\label{eq: psi moms definition}
\end{equation}
where ${\Greekmath 010D} =({\Greekmath 010D} _{2},\dots ,{\Greekmath 010D} _{K})^{\prime }$ is a $K-1$
dimensional vector of parameters. Let ${\Greekmath 010D} _{0}\equiv ({\Greekmath 010D} _{02},\dots
,{\Greekmath 010D} _{0K})^{\prime }$ denote the vector of true parameters ${\Greekmath 010D} _{0k}$
, defined as
\begin{equation}
{\Greekmath 010D} _{02}\equiv \frac{\mathbb{E}\left[ {\Greekmath 0122} _{i}^{2}\right] }{2}
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,} \;\; {\Greekmath 010D} _{03}\equiv \frac{\mathbb{E}\left[ {\Greekmath 0122} _{i}^{3}
\right] }{6}, \;\; \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{and} \;\; {\Greekmath 010D} _{0k}\equiv
\frac{\mathbb{E}\left[ {\Greekmath 0122} _{i}^{k}\right] }{k!}-\sum_{\ell =2}^{k-2}
\frac{\mathbb{E}\left[ {\Greekmath 0122} _{i}^{k-\ell }\right] }{(k-\ell )!}{\Greekmath 010D}
_{0\ell }\;\; \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for} \;\; k\geq 4. \label{eq: gammas}
\end{equation}
We formalize this discussion below.
\begin{assumption}[CME]
\namedlabel{ass: CME}{CME}
\emph{(Classical Measurement Error)} ${\Greekmath 0122} _{i}$ is independent from $
(X_{i}^{\ast },S_{i})$ and $\mathbb{E}[{\Greekmath 0122} _{i}]=0$.
\end{assumption}
The following lemma establishes validity of the corrected moment conditions
under Assumptions \ref{ass:MME}, \ref{ass: CME}, and some mild regularity
conditions provided in Appendix~\ref{sec:LargeSample}.
\begin{lemma}
\label{lem: corrected moment} Under Assumptions \ref{ass:MME}, \ref{ass: CME}
and \ref{ass: Moment function} in Appendix~\ref{sec:LargeSample},
\begin{equation*}
\mathbb{E}[{\Greekmath 0120} (X_{i},S_{i},{\Greekmath 0112} _{0},{\Greekmath 010D} _{0})]=\mathbb{E}
[g(X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0})]+o\left( n^{-1/2}\right) =o\left(
n^{-1/2}\right).
\end{equation*}
\end{lemma}
Lemma~\ref{lem: corrected moment} implies that the corrected moment
conditions ${\Greekmath 0120} $ are valid and can potentially be used to jointly estimate
parameters ${\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{0}$. The total number of parameters
to be estimated is now $\dim \left( {\Greekmath 0112} \right) +K-1$. Thus, joint
estimation of ${\Greekmath 0112} _{0}$ and ${\Greekmath 010D} _{0}$ requires that $\dim \left(
{\Greekmath 0120} \right) =\dim \left( g\right) \geq \dim \left( {\Greekmath 0112} \right) +K-1$,
i.e., that the original moment conditions $g$ include sufficiently many
overidentifying restrictions. For example, the overidentifying restrictions can be
constructed by using an instrumental variable; we discuss this in more
detail below.
\begin{remark}
Construction of the corrected moment conditions ${\Greekmath 0120}$ requires the original moment function $g(x,s,{\Greekmath 0112})$ to have a sufficient number of derivatives with respect to~$x$. Thus, the proposed correction method does not apply to settings with non-differentiable moment functions, for example, those arising in the instrumental variable quantile regression (IVQR).
\end{remark}
\subsection{Measurement Error Robust Moments (MERM) estimator}
The MERM estimator jointly estimates the parameters ${\Greekmath 0112} _{0}$ and $
{\Greekmath 010D} _{0}$ using moment conditions ${\Greekmath 0120} $. It is convenient to define the
joint vector of parameters
\begin{equation*}
{\Greekmath 010C} \equiv ({\Greekmath 0112} ^{\prime },{\Greekmath 010D} ^{\prime })^{\prime },\ {\Greekmath 010C}
_{0}\equiv ({\Greekmath 0112} _{0}^{\prime },{\Greekmath 010D} _{0}^{\prime })^{\prime },\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }
\hat{{\Greekmath 010C}}\equiv (\hat{{\Greekmath 0112}}^{\prime },\hat{{\Greekmath 010D}}^{\prime })^{\prime },
\end{equation*}
and the parameter space $\mathcal{B}\equiv \Theta \times \Gamma $, where $
\Theta $ and $\Gamma $ are the parameter spaces for ${\Greekmath 0112} $ and ${\Greekmath 010D} $.
Then, MERM estimator is the GMM\ estimator (\QTR{citealp}{Hansen1982Ecta}):
\begin{equation}
\hat{{\Greekmath 010C}}\equiv \operatorname*{\mathrm{arg}\!\min\limits}_{{\Greekmath 010C} \in \mathcal{B}}\hat{Q}({\Greekmath 010C} ),\qquad \hat{
Q}({\Greekmath 010C} )\equiv \overline{{\Greekmath 0120} }({\Greekmath 010C} )^{\prime }\hat{\Xi}\overline{{\Greekmath 0120} }
({\Greekmath 010C} ), \label{eq:MME definition}
\end{equation}
where $\overline{{\Greekmath 0120} }({\Greekmath 010C} )\equiv n^{-1}\sum_{i=1}^{n}{\Greekmath 0120} _{i}({\Greekmath 010C} )$
, ${\Greekmath 0120} _{i}({\Greekmath 010C} )\equiv {\Greekmath 0120} \left( X_{i},S_{i},{\Greekmath 010C} \right) $, $\hat{\Xi
}$ is a weighting matrix, and $\hat{Q}({\Greekmath 010C} )$ is the standard GMM
objective function.
While Lemma~\ref{lem: corrected moment} establishes validity of the
corrected moment restrictions ${\Greekmath 0120} $, the MERM estimator also relies on $
{\Greekmath 010C} _{0}$ being identified from ${\Greekmath 0120} $. This requirement is formalized by
the following assumption.
\begin{assumption}[ID]
\emph{(Identification)}
\namedlabel{ass: ID}{ID}
\begin{enumerate}[(i)]
\item
\label{item: local ID} the Jacobian $\Psi ^{\ast }$ has full column rank,
where
\begin{align*}
\Psi ^{\ast }& \equiv \mathbb{E}\left[ {\Greekmath 0272} _{{\Greekmath 0112} }{\Greekmath 0120} (X_{i}^{\ast
},S_{i},{\Greekmath 0112} _{0},0),{\Greekmath 0272} _{{\Greekmath 010D} }{\Greekmath 0120} (X_{i}^{\ast },S_{i},{\Greekmath 0112}
_{0},0)\right] \\
& =\mathbb{E}\left[ {\Greekmath 0272} _{{\Greekmath 0112} }g(X_{i}^{\ast },S_{i},{\Greekmath 0112}
_{0}),-g_{x}^{(2)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}),\dots
,-g_{x}^{(K)}(X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0})\right] ;
\end{align*}
\item
\label{item: global ID} $\mathbb{E}\left[ {\Greekmath 0120} (X_{i}^{\ast },S_{i},{\Greekmath 0112}
,{\Greekmath 010D} )\right] =0$ iff ${\Greekmath 0112} ={\Greekmath 0112} _{0}$ and ${\Greekmath 010D} =0$.
\end{enumerate}
\end{assumption}
Assumptions \ref{ass: ID}(\ref{item: local ID}) and (\ref{item: global ID})
are the standard GMM local and global identification conditions applied to
the moment function ${\Greekmath 0120} (X_{i}^{\ast },S_{i},{\Greekmath 0112} ,{\Greekmath 010D} )$. These are
high-level conditions, which we will return to later in the paper. The moment conditions formulation is
sufficiently general to encompass a wide variety of sources of identification. In Section~\ref{ssec:MERM-ID}, we discuss identification in detail and illustrate the construction of the moment function using an instrumental variable or a second measurement.
\bigskip
Under some additional regularity conditions, estimator $\hat{{\Greekmath 010C}}$ behaves
as a standard GMM-type estimator: it is $\sqrt{n}$-consistent and
asymptotically normal and unbiased. This result is formalized by the
following theorem.
\begin{theorem}[Asymptotic Normality]
\label{the: asy normality} Suppose that $\{(X_{i}^{\ast },S_{i}^{\prime
},{\Greekmath 0122} _{i})\}_{i=1}^{n}$ are i.i.d.. Then, under Assumptions \ref
{ass:MME}, \ref{ass: CME}, \ref{ass: ID}, and \ref{ass: Moment function}-\ref
{ass: basic conditions} in Appendix~\ref{sec:LargeSample},
\begin{equation}
n^{1/2}\Sigma ^{-1/2}(\hat{{\Greekmath 010C}}-{\Greekmath 010C} _{0})\overset{d}{\rightarrow }
N(0,I_{\dim ({\Greekmath 010C} )}),\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ where} \label{eq:betah asy-N}
\end{equation}
\vspace{-4ex}
\begin{equation*}
\Sigma \equiv (\Psi ^{\prime }\Xi \Psi )^{-1}\Psi ^{\prime }\Xi \Omega
_{{\Greekmath 0120} {\Greekmath 0120} }\Xi \Psi (\Psi ^{\prime }\Xi \Psi )^{-1}.
\end{equation*}
\end{theorem}
Theorem \ref{the: asy normality} shows that the MERM\ approach addresses the
EIV bias problem, and in particular provides a $\sqrt{n}$-consistent
asymptotically normal and unbiased estimator $\hat{{\Greekmath 0112}}$, which can be
used to conduct inference about the true parameters ${\Greekmath 0112} _{0}$. The
asymptotic variance $\Sigma $ takes the standard sandwich\ form, with $\Psi
\equiv \mathbb{E}\left[ {\Greekmath 0272} _{{\Greekmath 010C} }{\Greekmath 0120} _{i}({\Greekmath 010C} _{0})\right] $, $
\Omega _{{\Greekmath 0120} {\Greekmath 0120} }\equiv \mathbb{E}\left[ {\Greekmath 0120} _{i}\left( {\Greekmath 010C}
_{0}\right) {\Greekmath 0120} _{i}^{\prime }\left( {\Greekmath 010C} _{0}\right) \right] $, and $\hat{
\Xi}\rightarrow _{p}\Xi $.
\begin{remark}
Notice that the bias of naive estimators (such as a GMM estimator based on
the original moment conditions) is $O({\Greekmath 011C} _{n}^{2})$, so their rate of
convergence is $O_{p}({\Greekmath 011C} _{n}^{2}+n^{-1/2})$. The bias dominates sampling
variability and naive estimators are not $\sqrt{n}$-consistent unless ${\Greekmath 011C}
_{n}=O(n^{-1/4})$, i.e., unless the magnitude of the measurement error is
rather small. At the same time, the MERM estimator remains $\sqrt{n}$
-consistent for much larger values of ${\Greekmath 011C} _{n}$, up to ${\Greekmath 011C}
_{n}=O(n^{-1/(2K+2)})$, whereas the rate of convergence of naive estimators
is only $O_{p}(n^{-1/(K+1)})$ in this case.
\end{remark}
Once the corrected moment condition ${\Greekmath 0120} $ is constructed, estimation of
and inference about parameters ${\Greekmath 010C} _{0}$ can be performed using any
standard software package for GMM\ estimation. In other words, the proposed
estimator can be simply treated as a standard GMM estimator based on the
corrected moment conditions ${\Greekmath 0120} $, and the conventional standard errors,
tests, and confidence intervals are valid.
In addition to the estimation of the parameters ${\Greekmath 0112}_{0}$, researchers are often interested in average
effects of the form ${\Greekmath 0115}_{0}\equiv \mathbb{E}\left[ {\Greekmath 0115}\left(
X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) \right] $. For instance, in the NLR model, one may be
interested in the average partial effect of $x$ (i.e., ${\Greekmath 0115} _{0}\equiv \mathbb{E}\left[ {\Greekmath 0272} _{x}{\Greekmath 011A} \left( X_{i}^{\ast },S_{i},{\Greekmath 0112}_{0}\right) \right] $) or another covariate.
The naive average partial effect estimator $\hat{{\Greekmath 0115}}_{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Naive}}\equiv \frac{1}{n} \sum_{i=1}^{n}{\Greekmath 0115}(X_{i},S_{i},\hat{{\Greekmath 0112}})$ suffers from the EIV bias,
unless function ${\Greekmath 0115}$ is linear in $X_{i}^{\ast }$. Instead, one should use
estimates $\hat{{\Greekmath 010D}}$ to construct the bias-corrected estimator $\hat{
{\Greekmath 0115}}_{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{MERM}}\equiv \frac{1}{n}\sum_{i=1}^{n}\left\{
{\Greekmath 0115}(X_{i},S_{i},\hat{{\Greekmath 0112}})-\sum_{k=2}^{K}\hat{{\Greekmath 010D}}
_{k}{\Greekmath 0115}_{x}^{(k)}(X_{i},S_{i},\hat{{\Greekmath 0112}})\right\} $.
\begin{remark}
The standard $J$-test of overidentifying restrictions remains valid in the
MERM settings, and can be used to check the model specification. The $J$
-test jointly tests the following hypotheses: \emph{(i)} $K$ is sufficiently
large to correct the EIV\ bias; \emph{(ii)} assumptions on the EIV are
valid; and \emph{(iii)} the original moment conditions $g$ are correctly
specified so equation~(\ref{eq:g descr moment}) holds, i.e., that the
original economic model is correctly specified aside from the presence of
the EIV in $X_{i}$. Thus, if the $J$-test rejects the validity of the corrected moment conditions ${\Greekmath 0120}$, the researcher might want to \emph{(i)}~consider taking a larger $K$; \emph{(ii)} employ a different correction method; or \emph{(iii)}~consider an alternative specification of the original moments $g$.
\end{remark}
\begin{remark}
Considering larger $K$ allows for ${\Greekmath 011C} _{n}\ $converging to zero at a slower
rate, which in finite samples corresponds to the asymptotics providing
better approximations for larger magnitudes of measurement errors. On the
other hand, taking a larger $K$ increases the dimension of the nuisance
parameter ${\Greekmath 010D} _{0}$ and thus typically increases the variance of $\hat{
{\Greekmath 0112}}$. We consider this issue in more detail and provide a data-driven method for choosing $K$ in Section~\ref{ssec:adaptive K}.
{}
\end{remark}
\begin{remark}
The MERM framework can be extended to the case of non-classical measurement errors; see \citet{EvdokimovZeleneev2022WP-NPID} for details and a fully nonparametric analysis. In this paper, we focus on the classical measurement errors, developing a practical bias correction approach, which can be easily implemented in a wide range of economic applications. Even when Assumption~\ref{ass: CME} is violated, the deviations from it are often limited in magnitude, so the corrections based on the \ref{ass: CME} assumption remove most of the EIV bias. Thus, in practice, using the estimator designed for classical measurement errors is typically preferable to ignoring mismeasurement altogether.
\end{remark}
\begin{remark}
It is important to note that ${\Greekmath 010D} _{0k}\neq \mathbb{E}\left[ {\Greekmath 0122}
_{i}^{k}\right] \left/ k!\right. $ for $k\geq 4$,\ contrary to what
equation~(\ref{eq:g MME mu}) might suggest. For example, ${\Greekmath 010D}
_{04}=\left( \mathbb{E}\left[ {\Greekmath 0122} _{i}^{4}\right] -6{\Greekmath 011B}
_{{\Greekmath 0122} }^{4}\right) \left/ 24\right. $ is negative for many
distributions, including normal. The reason that
generally ${\Greekmath 010D} _{0k}\neq \mathbb{E}\left[ {\Greekmath 0122} _{i}^{k}\right]
\left/ k!\right. $ is that the estimators of the correction terms themselves
need a correction, which is accounted for by the form of ${\Greekmath 010D} _{0k}$.
Since there is a one-to-one relationship between ${\Greekmath 010D} _{0}$ and the
moments $\mathbb{E}\left[ {\Greekmath 0122} _{i}^{\ell }\right] $, parameter space
$\Gamma $ for ${\Greekmath 010D} _{0}$ can incorporate restrictions that the moments
must satisfy (e.g., ${\Greekmath 011B} _{{\Greekmath 0122} }^{2}\geq 0$ and $\mathbb{E}\left[
{\Greekmath 0122} _{i}^{4}\right] \geq {\Greekmath 011B} _{{\Greekmath 0122} }^{4}$). Such
restrictions can increase the efficiency of the estimator and the power of
tests.
\end{remark}
\begin{remark}
No parametric assumptions are imposed on the distribution of ${\Greekmath 0122}
_{i}$, i.e. the distribution of ${\Greekmath 0122} _{i}$ is treated
nonparametrically. The regularity conditions restrict only the magnitude of
the moments of ${\Greekmath 0122} _{i}$. The approach imposes no restrictions on
the smoothness of the distributions of $X_{i}^{\ast }$ and ${\Greekmath 0122} _{i}$
, which are not even required to be continuous. Examples in which this can
be useful include individual wages (whose distributions may have point
masses at round numbers), and allowing the measurement error ${\Greekmath 0122}
_{i}$ to have a point mass at zero (a fraction of the population may have a
zero measurement or recall error).
\end{remark}
\begin{remark}
The formulas of the derivatives $g_{x}^{(k)}(\cdot )$ are typically easy to
compute analytically or using symbolic algebra software. Alternatively,
these derivatives can be computed using numerical differentiation. Thus, the corrected set of moments can be automatically produced for a
generic moment function $g(\cdot )$ provided by the user.
\end{remark}
\subsection{\label{ssec:MERM-ID}Model Identification:\ Jacobian $\Psi $}
Theorem~\ref{the: asy normality} requires ${\Greekmath 010C} _{0}$ to be identified and
the Jacobian matrix $\Psi $ to be full rank. Notably, the MERM framework
encompasses many possible sources of identification at once, including
instrumental variables, additional measurements, or nonlinearities of the
functional form. The identifying information is incorporated in the moment
functions. Essentially, our approach first characterizes in what directions
the measurement errors can bias the moment conditions $\mathbb{E}\left[
g\left( X_{i},S_{i},{\Greekmath 0112} \right) \right] $, and then uses the moments
orthogonal to those directions for identification of ${\Greekmath 0112} _{0}$. To be
more specific, we will now consider identification when in addition to the
error-laden $X_{i}$ we have either (i) a general instrument $Z_{i}$ or (ii)
a second measurement $Q_{i}$.
\paragraph{Identification Using A General Instrument $Z_{i}$}
Many applications can be formulated as the following conditional moment
restriction:
\begin{equation}
\mathbb{E}\left[ u\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} \right) |X_{i}^{\ast }
\right] =0\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ iff }{\Greekmath 0112} ={\Greekmath 0112} _{0}, \label{eq: eg-ID mom u}
\end{equation}
for some moment function $u$. For example, consider the nonlinear regression
model $\mathbb{E}\left[ Y_{i}|X_{i}^{\ast }=x\right] ={\Greekmath 011A} \left( x,{\Greekmath 0112}
_{0}\right) $, then $u\left( x,y,{\Greekmath 0112} \right) ={\Greekmath 011A} \left( x,{\Greekmath 0112}
\right) -y$.\footnote{
For simplicity of the exposition, in the expectation in equation~(\ref{eq:
eg-ID mom u}) we only condition on $X_{i}^{\ast }$. The discussion applies
in a straightforward way to the settings with additional correctly measured
variables $W_{i}$ in the conditioning set, i.e., the model $\mathbb{E}\left[
u\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) |X_{i}^{\ast },W_{i}\right]
=0 $. For example, in the nonlinear regression example with additional
covariates $W_{i}$ we have $\mathbb{E}\left[ Y_{i}|X_{i}^{\ast }=x,W_{i}=w
\right] ={\Greekmath 011A} \left( x,w,{\Greekmath 0112} _{0}\right) $, so $u\left( x,y,w,{\Greekmath 0112}
\right) ={\Greekmath 011A} \left( x,w,{\Greekmath 0112} \right) -y$.}
In applications, identification of the models with EIV would typically rely
on an instrumental variable $Z_{i}$. Suppose the instrument satisfies the
exclusion restriction $\mathbb{E}\left[ u\left( X_{i}^{\ast },S_{i},{\Greekmath 0112}
\right) |X_{i}^{\ast },Z_{i}\right] =\mathbb{E}\left[ u\left( X_{i}^{\ast
},S_{i},{\Greekmath 0112} \right) |X_{i}^{\ast }\right] $, i.e., conditional on the
true $X_{i}^{\ast }$ the instrument has no further effect on the moment
conditions $u$. Consider the moment functions $h\left( x,s,{\Greekmath 0112} \right)
\equiv u\left( x,s,{\Greekmath 0112} \right) \otimes {\Greekmath 0127} _{X}\left( x\right) $,
where ${\Greekmath 0127} _{X}\left( x\right) $ is a vector of functions of $x$, e.g., $
{\Greekmath 0127} _{X}\left( x\right) \equiv \left( 1,x,\ldots ,x^{J}\right) ^{\prime
} $. By the Law of Iterated Expectations, $\mathbb{E}\left[ h\left(
X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) |Z_{i}=z\right] =0$ for all $z$.
However, the same expectation with $X_{i}^{\ast }$ replaced by the observed $
X_{i}$, $\mathbb{E}\left[ h\left( X_{i},S_{i},{\Greekmath 0112} _{0}\right) |Z_{i}=z
\right] ${}, will depend on $z$. One way to see this is to consider a Taylor
expansion similar to equation~(\ref{eq:moment expectation for K=2}):{\small
\begin{equation*}
\mathbb{E}\left[ h\left( X_{i},S_{i},{\Greekmath 0112} _{0}\right) |Z_{i}=z\right] =
\underbrace{\mathbb{E}\left[ h\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right)
|Z_{i}=z\right] }_{=0\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ by the LIE}}+{\Greekmath 010D} _{0}\mathbb{E}\left[
h_{x}^{\left( 2\right) }\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right)
|Z_{i}=z\right] +O\left( {\Greekmath 011C} _{n}^{3}\right) ,
\end{equation*}
}which shows that $\mathbb{E}\left[ h\left( X_{i},S_{i},{\Greekmath 0112} _{0}\right)
|Z_{i}=z\right] $ is zero for all $z$ (up to a negligible remainder) unless $
{\Greekmath 010D} _{0}\neq 0$, where ${\Greekmath 010D} _{0}={\Greekmath 011B} ^{2}/2$. Thus, $\mathbb{E}
\left[ h\left( X_{i},S_{i},{\Greekmath 0112} _{0}\right) |Z_{i}=z\right] $ varies with $
z$ only because of the presence of the measurement error. Intuitively, the
magnitude of this variation then identifies the nuisance parameters ${\Greekmath 010D}
_{0}$. Thus, one can rely on the original moment functions of the typical
form $g\left( x,s,{\Greekmath 0112} \right) =h\left( x,s,{\Greekmath 0112} \right) \otimes {\Greekmath 0127}
_{Z}\left( z\right) $, where ${\Greekmath 0127} _{Z}\left( z\right) $ is a vector of
functions of $z$.
The above discussion provides the intuition for identification of the
nonlinear moment condition models with EIV. It is important to note that
identification of general nonlinear moment condition models is a complicated
problem. Even in the settings without measurement errors, it is generally
not possible to give low-level conditions guaranteeing that a specific set
of nonlinear moment conditions identifies the parameter vector. The presence
of EIV\ makes the question of identification even harder.
We attempt to address this concern and make the above intuitions more
precise in two ways. First, in the following subsection we consider a
specific (but frequently employed) kind of an instrument: a second
measurement (possibly non-classical). The specific form of the excluded
variable allows us to provide more transparent identification conditions.
Second, in \cite{EvdokimovZeleneev2022WP-NPID} we study nonparametric
regression model with EIV using the ${\Greekmath 011C} _{n}\rightarrow 0$ approximation.
We show that the model is identified using an instrument (even a discrete
one), and motivate the MERM approach from a nonparametric perspective.
Finally, since MERM estimator is a standard GMM\ estimator, one can test the
strength of identification of the model parameters, or conduct
identification-robust inference using the standard methods (e.g.,
\QTR{citealp}{
StockWright2000Ecta,Kleibergen2005Ecta,GuggenbergerSmith2005ET,GuggenbergerRamalhoSmith2012JoE,AndrewsMikusheva2016Ecta-ConditionalFunctionalNuisance,AndrewsI-2016-Ecta-CLC,AndrewsGuggenberger2019QE
}).
\paragraph{Identification Using A Second Measurement}
Suppose we observe a second measurement
\begin{equation*}
Q_{i}={\Greekmath 010B} _{1}X_{i}^{\ast }+{\Greekmath 0122} _{Q,i},
\end{equation*}
where ${\Greekmath 010B} _{1}$ may not be known. Assume that ${\Greekmath 010B} _{1}\neq 0$ and $
\mathbb{E}\left[ {\Greekmath 0122} _{Q,i}|X_{i}^{\ast },S_{i},{\Greekmath 0122} _{i}
\right] =0$. The variance of ${\Greekmath 0122} _{Q,i}$ does not need to be small.
Note that the measurement error in $Q_{i}$ can be non-classical: $
Q_{i}-X_{i}^{\ast }$ and $X_{i}^{\ast }$ are correlated unless ${\Greekmath 010B}
_{1}=1 $.
Consider the conditional moment restrictions~(\ref{eq: eg-ID mom u}). If $
X_{i}^{\ast }$ were observed, we could have constructed the unconditional
moments
\begin{equation*}
\mathbb{E}\left[ h\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) \right] =0,
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{\qquad }h\left( x,s,{\Greekmath 0112} \right) \equiv u\left( x,s,{\Greekmath 0112} \right)
\times \left( 1,x,\ldots ,x^{J}\right) ^{\prime },
\end{equation*}
for some $J\geq \dim \left( {\Greekmath 0112} \right) -1$. Suppose that the model is
identified if $X_{i}^{\ast }$ observed, which means that the Jacobian of
these moment conditions has full rank:
\begin{equation*}
\limfunc{Rk}\left( H^{\ast }\right) =\dim \left( {\Greekmath 0112} \right) \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,
where }H^{\ast }\equiv E\left[ {\Greekmath 0272} _{{\Greekmath 0112} }h\left( X_{i}^{\ast
},S_{i},{\Greekmath 0112} _{0}\right) \right] .
\end{equation*}
To deal with the error-laden $X_{i}$, consider the MERM estimator with $K=2$
based on the following moment function
\begin{equation}
g\left( x,s,q,{\Greekmath 0112} \right) \equiv \left(
\begin{array}{l}
h\left( x,s,{\Greekmath 0112} \right) \\
u\left( x,s,{\Greekmath 0112} \right) q\times \left( 1,x,\ldots ,x^{J-1}\right)
^{\prime }
\end{array}
\right) . \label{eq:rank-cond:eg g}
\end{equation}
Here the total number of moments is $m=2J+1$. The additional $J$ moments
added in equation~(\ref{eq:rank-cond:eg g}) use $Q_{i}$, which will allow
identifying ${\Greekmath 010D} _{0}=\mathbb{E}[{\Greekmath 0122} _{i}^{2}]/2$.
It turns out that in these settings there is a simple sufficient condition
for Assumption~\ref{ass: ID}(\ref{item: local ID}) to hold. Appendix~\ref{sec:full rank Psi}
demonstrates that $\Psi ^{\ast }$ will have full rank if
\begin{equation}
\mathbb{E}\left[ u_{x}^{\left( 1\right) }\left( X_{i}^{\ast },S_{i},{\Greekmath 0112}
_{0}\right) \times \left( 1,X_{i}^{\ast },\ldots ,\left( X_{i}^{\ast
}\right) ^{J-1}\right) ^{\prime }\right] \neq 0.
\label{eq:eg rk Psi:suff cond}
\end{equation}
Condition~(\ref{eq:eg rk Psi:suff cond}) has a very simple interpretation:
it essentially it means that $\mathbb{E}\left[ \left. u_{x}^{\left( 1\right)
}\left( X_{i}^{\ast },S_{i},{\Greekmath 0112} _{0}\right) \right\vert X_{i}^{\ast }
\right] $ should not be identically zero. For example, in the nonlinear
regression model $u\left( x,y,{\Greekmath 0112} \right) ={\Greekmath 011A} \left( x,{\Greekmath 0112} \right)
-y $, and condition~(\ref{eq:eg rk Psi:suff cond}) is satisfied as long as $
\mathbb{E}\left[ {\Greekmath 011A} _{x}^{(1)}\left( X_{i}^{\ast },{\Greekmath 0112} _{0}\right)
\left( X_{i}^{\ast }\right) ^{j}\right] \neq 0$ for some $j\in \left\{
0,\ldots ,J-1\right\} $.
\bigskip
\section{\label{sec:MC}Numerical Evidence}
\subsection{Comparison with a Semi-Nonparametric Estimation Approach}
\label{ssec: MC S07 comparison}
We compare MERM estimator with the state-of-the-art semiparametric estimator of \citet[henceforth S07]{Schennach2007Ecta} for nonlinear regression models. The Monte Carlo designs are taken from S07, and include a polynomial, rational fraction, and Probit nonlinear regression models. Identification of the model is ensured by the availability of an instrument.
\begin{equation}
Y_i = {\Greekmath 011A}(X_i^*,{\Greekmath 0112}_0) + U_i, \quad X_i^* = {\Greekmath 0119}_1 Z_i + V_i, \quad X_i = X_i^* + {\Greekmath 0122}_i, \label{eq: MC DGP}
\end{equation}
$(Z_i,V_i,{\Greekmath 0122}_i)' \sim N \left((0, 0, 0)', \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Diag}(1, 1/4, 1/4)\right)$, ${\Greekmath 0119}_1=1$, and $n = 1000$. The conditional expectation function ${\Greekmath 011A}$, the true value of the parameter of interest ${\Greekmath 0112}_0$, and the conditional distribution of the regression error $U_i$ are design-specific and reported in Tables \ref{tab: MC poly}-\ref{tab: MC probit} below. In all designs, ${\Greekmath 011C} = {\Greekmath 011B}_{{\Greekmath 0122}}/{\Greekmath 011B}_{X}^* \approx 0.45$, so the measurement error is ``fairly large'' \citep{Schennach2007Ecta}.
We report simulation results for the MERM estimator considering correction schemes with $K=2$ and $K=4$. The original moment function is
\begin{equation*}
g(x,y,z,{\Greekmath 0112}) = (y - {\Greekmath 011A}(x,{\Greekmath 0112})) {\Greekmath 0127}(x,z),
\end{equation*}
where we use ${\Greekmath 0127}(x,z) = \left(1, x, z, x^2, z^2, x^3, z^3\right)'$ for $K=2$ and ${\Greekmath 0127}(x,z) = \left(1, x, z, x^2, xz, z^2, x^3, x^2 z, x z^2, z^3\right)'$ for $K=4$.
The finite sample properties of the MERM estimators (evaluated based on 5,000 replications) are reported in Tables \ref{tab: MC poly}-\ref{tab: MC probit} below. For comparison, we also provide the same statistics for naive estimators (OLS/NLLS) and for the benchmark estimator of S07 (as reported in the original paper). For the polynomial model (Table \ref{tab: MC poly}), both $K=2$ and $K=4$ MERM estimators effectively remove the EIV bias. Component-wise, the MERM estimators perform similarly (for ${\Greekmath 0112}_2$ and ${\Greekmath 0112}_4$) or better (for ${\Greekmath 0112}_1$ and ${\Greekmath 0112}_3$) compared to the benchmark estimator of S07. For the rational fraction model (Table \ref{tab: MC frac}), both the MERM estimators are vastly superior to the benchmark estimator both in terms of the bias and the standard deviation. For the probit model (Table \ref{tab: MC probit}), the MERM estimator with $K=2$ removes a large fraction of the EIV bias compared to the NLLS estimator. However, the EIV bias remains non-negligible when this simplest correction scheme is used. Employing a higher order correction scheme with $K=4$ completely eliminates the remaining EIV bias, while at the same time having smaller standard deviations (than the benchmark estimator of S07) . Overall, in the considered designs, the MERM estimator with $K=4$ consistently outperforms the benchmark estimator. It also proves to be more effective in removing the EIV bias compared to the $K=2$ estimator, especially in the highly nonlinear settings of the considered probit design.
\begin{table}[h!]
\begin{center}
\begin{footnotesize}
\begin{threeparttable}
\caption{Simulation results for the polynomial model of S07}
\label{tab: MC poly}
\begin{tabular*}{\textwidth}{@{\extracolsep{\fill}} l c c c c c c c c c c c c c }
\toprule
\toprule
& \multicolumn{4}{c}{Bias} & \multicolumn{4}{c}{Std. Dev.} & \multicolumn{5}{c}{RMSE} \\
\cmidrule(lr){2-5} \cmidrule(lr){6-9} \cmidrule(lr){10-14}
&\textbf{${\Greekmath 0112}_1$}&\textbf{${\Greekmath 0112}_2$}&\textbf{${\Greekmath 0112}_3$}&\textbf{${\Greekmath 0112}_4$}&\textbf{${\Greekmath 0112}_1$}&\textbf{${\Greekmath 0112}_2$}&\textbf{${\Greekmath 0112}_3$}&\textbf{${\Greekmath 0112}_4$}&\textbf{${\Greekmath 0112}_1$}&\textbf{${\Greekmath 0112}_2$}&\textbf{${\Greekmath 0112}_3$}&\textbf{${\Greekmath 0112}_4$}& All \\ \midrule
OLS &-0.00&-0.43&0.00&0.21&0.07&0.13&0.06&0.04&0.07&0.45&0.06&0.22&0.51\\
S07 &-0.05&-0.07&-0.02&0.05&0.17&0.19&0.24&0.05&0.17&0.20&0.24&0.07&0.36\\
$K=2$ &-0.00& 0.10& 0.00& 0.00& 0.10& 0.23& 0.10& 0.08& 0.10& 0.25& 0.10& 0.08& 0.29 \\
$K=4$ &-0.00& 0.00& 0.00& 0.02& 0.09& 0.21& 0.10& 0.08& 0.09& 0.21& 0.10& 0.08& 0.27 \\
\bottomrule
\end{tabular*}
\begin{tablenotes}
\item \footnotesize The DGP is as in \eqref{eq: MC DGP} with ${\Greekmath 011A}(x,{\Greekmath 0112}) = {\Greekmath 0112}_{1} + {\Greekmath 0112}_{2} x + {\Greekmath 0112}_{3} x^2 + {\Greekmath 0112}_{4} x^3$, ${\Greekmath 0112}_0 = (1,1,0,-0.5)'$, and $U_i \sim N(0,1/4)$.
\end{tablenotes}
\end{threeparttable}
\end{footnotesize}
\end{center}
\end{table}
\begin{table}[h!]
\begin{center}
\begin{footnotesize}
\begin{threeparttable}
\caption{Simulation results for the rational fraction model of S07}
\label{tab: MC frac}
\begin{tabular*}{\textwidth}{@{\extracolsep{\fill}} l c c c c c c c c c c }
\toprule
\toprule
& \multicolumn{3}{c}{{Bias}} & \multicolumn{3}{c}{Std. Dev.} & \multicolumn{4}{c}{RMSE} \\
\cmidrule(lr){2-4} \cmidrule(lr){5-7} \cmidrule(lr){8-11} &\textbf{${\Greekmath 0112}_1$}&\textbf{${\Greekmath 0112}_2$}&\textbf{${\Greekmath 0112}_3$}&\textbf{${\Greekmath 0112}_1$}&\textbf{${\Greekmath 0112}_2$}&\textbf{${\Greekmath 0112}_3$}&\textbf{${\Greekmath 0112}_1$}&\textbf{${\Greekmath 0112}_2$}&\textbf{${\Greekmath 0112}_3$}&{All}\\ \midrule
OLS &0.339&-0.167&-0.644&0.040&0.020&0.076&0.341&0.168&0.648&0.752\\
S07 &0.107&0.117&-0.150&0.146&0.139&0.328&0.181&0.182&0.361&0.443\\
$K=2$ &-0.004&-0.018& 0.014& 0.062& 0.026& 0.139& 0.062& 0.032& 0.139& 0.156\\
$K=4$ &0.014&-0.002&-0.024& 0.062& 0.031& 0.154& 0.063& 0.031& 0.156& 0.171 \\
\bottomrule
\end{tabular*}
\begin{tablenotes}
\item \footnotesize The DGP is as in \eqref{eq: MC DGP} with ${\Greekmath 011A}(x,{\Greekmath 0112}) = {\Greekmath 0112}_{1} + {\Greekmath 0112}_{2} x + \frac{{\Greekmath 0112}_3}{(1 + x^2)^{2}}$, ${\Greekmath 0112}_0 = (1,1,2)'$, and $U_i \sim N(0,1/4)$.
\end{tablenotes}
\end{threeparttable}
\end{footnotesize}
\end{center}
\end{table}
\begin{table}[h!]
\begin{center}
\begin{footnotesize}
\begin{threeparttable}
\caption{Simulation results for the Probit model of S07}
\label{tab: MC probit}
\begin{tabular*}{\textwidth}{@{\extracolsep{\fill}} l c c c c c c c}
\toprule
\toprule
& \multicolumn{2}{c}{Bias} & \multicolumn{2}{c}{Std. Dev.} & \multicolumn{3}{c}{RMSE} \\
\cmidrule(lr){2-3} \cmidrule(lr){4-5} \cmidrule(lr){6-8}
&\textbf{${\Greekmath 0112}_1$}&\textbf{${\Greekmath 0112}_2$}&\textbf{${\Greekmath 0112}_1$}&\textbf{${\Greekmath 0112}_2$}&\textbf{${\Greekmath 0112}_1$}&\textbf{${\Greekmath 0112}_2$}&{All}\\
\midrule
NLLS& 0.38&-0.97&0.06&0.08&0.39&0.98&1.05\\
S07 & 0.05&-0.06&0.39&0.53&0.39&0.53&0.69\\
$K=2$ & 0.11&-0.31& 0.18& 0.34& 0.21& 0.46& 0.51 \\
$K=4$ &-0.01&-0.01& 0.23& 0.42& 0.23& 0.42& 0.48 \\
\bottomrule
\end{tabular*}
\begin{tablenotes}
\item \footnotesize The DGP is as in \eqref{eq: MC DGP} with ${\Greekmath 011A}(x,{\Greekmath 0112}) = \frac{1}{2} (1 + \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{erf}({\Greekmath 0112}_{1} + {\Greekmath 0112}_{2} x))$, ${\Greekmath 0112}_0 = (-1,2)'$, and $U_i = 1 - {\Greekmath 011A}(X_i^*,{\Greekmath 0112}_0)$ with probability ${\Greekmath 011A}(X_i^*,{\Greekmath 0112}_0)$ and $- {\Greekmath 011A}(X_i^*,{\Greekmath 0112}_0)$ otherwise.
\end{tablenotes}
\end{threeparttable}
\end{footnotesize}
\end{center}
\end{table}
\subsection{\label{ssec: MC Mult Logit} Estimation and Inference in a Multinomial Choice Model}
Consider the standard multinomial logit model, in which an agent chooses between 3 available options.
For an agent $i$ with characteristics $(X_i^*,W_i)$, the utility of option $j$ is given by
\begin{equation*}
U_{ij} = {\Greekmath 0112}_{0j1} X_i^* + {\Greekmath 0112}_{0j2} W_{ij} + {\Greekmath 0112}_{0j3} + {\Greekmath 010F}_{ij} \quad \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for } j \in \{1,2\},
\end{equation*}
and $U_{i0} = {\Greekmath 010F}_{i0}$ for the outside option $j = 0$, where ${\Greekmath 010F}_{ij}$ are i.i.d. (across $i$ and $j$) draws from a standard type-1 extreme value distribution. The researcher observes $\left\{(X_i,W_i,Y_{i1},Y_{i2},Y_{i0})\right\}_{i=1}^n$, where $Y_{ij}$ is a binary variable indicating whether agent $i$ chooses option $j$, i.e. $Y_{ij} = 1$ if and only if $j = \operatorname*{\mathrm{arg}\!\max\limits}_{j' \in \{0,1,2\}} U_{i j'}$. In addition,
\begin{equation*}
X_i^* = V_{i1} Z_i + V_{i0}, \quad X_i = X_i^* + {\Greekmath 0122}_i, \quad W_{ij} = {{\Greekmath 011A}} X_i^*/{\Greekmath 011B}_X^* + \sqrt{1 - {\Greekmath 011A}^2} {\Greekmath 0117}_{ij},
\end{equation*}
and $\left(V_{i1}, V_{i0}, Z_i,{\Greekmath 0122}_i,{\Greekmath 0117}_{i1},{\Greekmath 0117}_{i2}\right)' \sim N\left((1,0,0,0,0,0)', \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Diag}({\Greekmath 011B}_{V1}^2,{\Greekmath 011B}_{V0}^2, {\Greekmath 011B}_Z^2, {\Greekmath 011B}_{\Greekmath 0122}^2, {\Greekmath 011B}_{\Greekmath 0117}^2, {\Greekmath 011B}_{\Greekmath 0117}^2)\right)$. In all of the designs, we fix $({\Greekmath 0112}_{011},{\Greekmath 0112}_{012},{\Greekmath 0112}_{013},{\Greekmath 0112}_{021},{\Greekmath 0112}_{022},{\Greekmath 0112}_{023},{\Greekmath 011A},{\Greekmath 011B}_{V1}^2,{\Greekmath 011B}_{V0}^2,{\Greekmath 011B}_Z^2,{\Greekmath 011B}_{\Greekmath 0117}^2) = (1,0,0,0,0,0,0.7,1/2,1/2,1,1)$ and $n=2000$. We consider ${\Greekmath 011C} = {\Greekmath 011B}_{{\Greekmath 0122}}/{\Greekmath 011B}_{X^*} \in \{1/4, 1/2, 3/4\}$. Setting ${\Greekmath 011B}_{V1}=0$ would correspond to the additive control variable model. We omit such simulation results for brevity.
Similarly to Section \ref{ssec: MC S07 comparison}, we report results for the MERM estimators with $K=2$ and $K=4$ based on the following original moment function
\begin{align*}
g (x,w,y,z,{\Greekmath 0112}) &= \left(\left(y_1 - p_1(x,w,{\Greekmath 0112}) \right) {\Greekmath 0127}_1(x,z,w)', \left(y_2 - p_2(x,w,{\Greekmath 0112})\right) {\Greekmath 0127}_2(x,z,w)' \right)',\\
p_j(x,w,{\Greekmath 0112})
&= \frac{\exp({\Greekmath 0112}_{j1} x + {\Greekmath 0112}_{j2} w_{j} + {\Greekmath 0112}_{j3})}{1 + \exp({\Greekmath 0112}_{11} x + {\Greekmath 0112}_{12} w_{1} + {\Greekmath 0112}_{13}) + \exp({\Greekmath 0112}_{21} x + {\Greekmath 0112}_{22} w_{2} + {\Greekmath 0112}_{23})},
\end{align*}
where ${\Greekmath 0127}_j(x,z,w) = \left(1, x, z, x^2, z^2, x^3, z^3, w_j\right)'$ for $K=2$ and ${\Greekmath 0127}_j(x,z,w) = \left(1, x, z, x^2, xz, z^2, x^3, x^2 z, x z^2, z^3, w_j\right)'$ for $K=4$.
We report the results on estimation and inference on the partial derivatives of the conditional choice probabilities $p_j(x,w_1,w_2)$ with respect to $x$, $w_1$, and $w_2$, evaluated at the population means.
Table \ref{tab: MC mult logit 2} reports the finite sample biases, standard deviations, and RMSE of the MERM estimators, as well as the sizes of the corresponding t-tests with nominal size of 5\%. To illustrate the importance of dealing with EIV, we also report the same statistics for the standard (naive) MLE estimator that ignores the presence of the measurement errors.
In all designs, the MLE estimator is biased, and the corresponding t-tests over-reject.
Note that failing to account for the EIV in the mismeasured variable $X_i^*$ generally biases estimators of all of the parameter, including those corresponding to the correctly measured variables $W_{i1}$ and $W_{i2}$. In particular, the t-tests may falsely reject true null hypotheses $\partial p_j/\partial w_\ell=0$ up to nearly $100\%$ of the time.
The MERM estimator with $K=2$ removes a large fraction of the EIV bias in all of the designs. While this proves to be enough to achieve accurate size control when the magnitude of the measurement error is moderate (${\Greekmath 011C} = 1/4$), the remaining EIV bias may still result in size distortions of the t-tests with larger measurement errors, especially ${\Greekmath 011C} = 3/4$. Using the higher order correction scheme with $K=4$ effectively removes the EIV bias in all of the simulation designs for all of the parameters.
Remarkably, the corresponding finite sample null rejection probabilities remain close to the nominal $5\%$ rate even when the standard deviation of the measurement error is as large as $75\%$ of the standard deviation of the mismeasured $X^*$.
To further check the limits of applicability of our method, we also consider larger values of ${\Greekmath 011C} \in \{1,3/2,2\}$. The numerical results analogous to the ones reported in Table~\ref{tab: MC mult logit 2} are provided in Table~\ref{tab: MC mult logit 2, large tau} in Appendix~\ref{ssec: large tau MC}. Specifically, we find that inference results based on the correction scheme with $K=4$ remain accurate even for ${\Greekmath 011C} = 1$. Unsurprisingly, inference becomes less reliable for bigger ${\Greekmath 011C} = 3/2$ and ${\Greekmath 011C} = 2$. At the same time, while the $K=4$ correction scheme fails to entirely eliminate the EIV bias in these designs, it still removes a big fraction of the bias and greatly improves on MLE in terms of the RMSE. Thus, while inference based on our estimator might be less reliable in extreme settings when the measurement error overwhelms the signal, the MERM estimator using $K=4$ appears to be sufficiently accurate over a wide range of ${\Greekmath 011C}$.
\afterpage{
\clearpage
\begin{landscape}
\begin{table}[h!]
\begin{footnotesize}
\begin{threeparttable}
\caption{Simulation results for the multinomial logit model}
\label{tab: MC mult logit 2}
\begin{tabular}{ l c c c c| c c c c | c c c c}
\toprule
\toprule
& \multicolumn{4}{c}{MLE} & \multicolumn{4}{c}{$K=2$} & \multicolumn{4}{c}{$K=4$}\\
\cmidrule(lr){2-5} \cmidrule(lr){6-9} \cmidrule(lr){10-13}
&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{std}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{size}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{std}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{size}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{std}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{size}}\\
\midrule
\multicolumn{13}{c}{${\Greekmath 011C} = 1/4$}\\
\midrule
$\partial p_1/\partial x$&-3.24&1.36&3.51&66.98&0.74&2.63&2.74&4.30&1.13&2.73&2.95&7.86\\
$\partial p_1/\partial w_1$&2.32&1.64&2.84&30.74&-0.11&2.30&2.30&4.82&-0.31&2.29&2.31&6.54\\
$\partial p_1/\partial w_2$&0.48&0.75&0.90&9.40&-0.04&0.87&0.87&4.82&-0.08&0.87&0.87&5.36\\
$\partial p_2/\partial x$&1.96&1.17&2.28&39.44&-0.40&1.88&1.92&4.72&-0.63&1.93&2.03&6.74\\
$\partial p_2/\partial w_1$&-1.16&0.82&1.42&30.66&0.06&1.15&1.15&4.84&0.15&1.15&1.16&6.48\\
$\partial p_2/\partial w_2$&-0.96&1.50&1.78&9.48&0.09&1.74&1.74&4.98&0.17&1.74&1.75&5.44\\
$\partial p_0/\partial x$&1.28&1.01&1.63&25.28&-0.34&1.43&1.47&5.08&-0.50&1.48&1.56&7.36\\
$\partial p_0/\partial w_1$&-1.16&0.82&1.42&30.60&0.06&1.15&1.15&4.82&0.15&1.15&1.16&6.46\\
$\partial p_0/\partial w_2$&0.48&0.74&0.88&9.46&-0.05&0.87&0.87&4.90&-0.09&0.87&0.88&5.32\\
\midrule
\multicolumn{13}{c}{${\Greekmath 011C} = 1/2$}\\
\midrule
$\partial p_1/\partial x$&-8.97&1.09&9.04&100.00&-1.69&2.60&3.10&9.84&0.97&2.89&3.05&6.04\\
$\partial p_1/\partial w_1$&6.44&1.53&6.62&98.96&1.39&2.44&2.81&12.14&-0.21&2.41&2.42&5.54\\
$\partial p_1/\partial w_2$&1.28&0.72&1.47&42.54&0.29&0.90&0.94&7.00&-0.06&0.92&0.92&5.00\\
$\partial p_2/\partial x$&5.22&0.97&5.31&99.98&1.05&1.89&2.16&10.16&-0.53&2.08&2.15&6.14\\
$\partial p_2/\partial w_1$&-3.21&0.77&3.30&98.96&-0.69&1.22&1.40&12.10&0.10&1.21&1.21&5.52\\
$\partial p_2/\partial w_2$&-2.52&1.41&2.88&42.82&-0.56&1.79&1.87&7.20&0.13&1.84&1.84&4.98\\
$\partial p_0/\partial x$&3.75&0.86&3.85&98.78&0.64&1.45&1.59&8.18&-0.44&1.59&1.65&6.50\\
$\partial p_0/\partial w_1$&-3.23&0.78&3.32&98.96&-0.70&1.22&1.41&12.06&0.10&1.21&1.21&5.48\\
$\partial p_0/\partial w_2$&1.23&0.69&1.41&42.90&0.28&0.89&0.93&7.20&-0.07&0.92&0.92&4.94\\
\midrule
\multicolumn{13}{c}{${\Greekmath 011C} = 3/4$}\\
\midrule
$\partial p_1/\partial x$&-13.35&0.86&13.38&100.00&-6.83&2.64&7.32&80.32&0.71&3.22&3.29&4.74\\
$\partial p_1/\partial w_1$&9.69&1.45&9.80&100.00&4.95&2.65&5.61&65.52&0.01&2.62&2.62&5.34\\
$\partial p_1/\partial w_2$&1.81&0.69&1.94&75.30&1.01&0.89&1.35&26.08&-0.01&0.98&0.98&5.24\\
$\partial p_2/\partial x$&7.48&0.79&7.52&100.00&4.06&1.83&4.45&68.82&-0.37&2.32&2.35&5.74\\
$\partial p_2/\partial w_1$&-4.83&0.73&4.88&100.00&-2.47&1.32&2.81&65.46&-0.01&1.31&1.31&5.28\\
$\partial p_2/\partial w_2$&-3.51&1.33&3.76&75.60&-1.99&1.76&2.66&26.38&0.03&1.97&1.97&5.32\\
$\partial p_0/\partial x$&5.87&0.73&5.92&100.00&2.77&1.47&3.14&56.28&-0.34&1.77&1.80&5.82\\
$\partial p_0/\partial w_1$&-4.87&0.75&4.93&100.00&-2.48&1.33&2.81&65.40&-0.01&1.31&1.31&5.32\\
$\partial p_0/\partial w_2$&1.70&0.64&1.82&75.76&0.98&0.87&1.31&26.50&-0.02&0.99&0.99&5.30\\
\bottomrule
\end{tabular}
\begin{tablenotes}
\item \footnotesize This table reports the simulated finite sample bias, standard deviation, RMSE, and size of the MLE and the MERM estimators and the corresponding t-tests for the partial derivatives $\partial p_j(x,w,{\Greekmath 0112}_0)/\partial x$, $\partial p_j(x,w,{\Greekmath 0112}_0)/\partial w_1$, $\partial p_j(x,w,{\Greekmath 0112}_0)/\partial w_2$ for $j \in \{1,2,0\}$ evaluated at the population mean. The true values of the marginal effects are $(\partial p_1/\partial x, \partial p_2/\partial x, \partial p_0/\partial x) = ( 0.222, -0.111, -0.111)$ and zeros for the rest. The results are based on 5,000 replications.
\end{tablenotes}
\end{threeparttable}
\end{footnotesize}
\end{table}
\end{landscape}
}
\subsection{\label{sec:empirical}Empirical Illustration: Choice of Transportation Mode}
In this section, we illustrate the finite sample properties of the MERM estimator in the context of a classical multinomial choice application: choice of transportation mode (e.g., \citealp{McFadden1974JPubE}).
To calibrate the numerical experiment, we use the ModeCanada dataset, a survey of business travelers for the Montreal-Toronto corridor. We focus on the subset of travelers choosing between train, air, and car ($n = 2769$), and estimate the conditional logit model with traveler $i$'s utilities given in the table below.
\begin{center}
\begin{tabular}{l|c}
Mode & Utility \\
\midrule
Air & $U_{i1} = {\Greekmath 0112}_{01} \thinspace Income_{i}^* + {\Greekmath 0112}_{02} \thinspace Urban_{i} + {\Greekmath 0112}_{03} + {\Greekmath 0112}_{07} \thinspace Price_{i1} + {\Greekmath 0112}_{08} \thinspace InTime_{i1} + {\Greekmath 010F}_{i1}$\\
Car & $U_{i2} = {\Greekmath 0112}_{04} \thinspace Income_{i}^* + {\Greekmath 0112}_{05} \thinspace Urban_{i} + {\Greekmath 0112}_{06} + {\Greekmath 0112}_{07} \thinspace Price_{i2} + {\Greekmath 0112}_{08} \thinspace InTime_{i2} + {\Greekmath 010F}_{i2}$\\
Train & $U_{i0} = {\Greekmath 0112}_{07} \thinspace Price_{i0} + {\Greekmath 0112}_{08} \thinspace InTime_{i0} + {\Greekmath 010F}_{i0}$
\end{tabular}
\end{center}
To generate the simulated samples, we randomly draw covariates from their joint empirical distribution. To generate the simulated outcomes, we draw ${\Greekmath 010F}_{ij}$ from the standard type-I extreme value distribution.
The true value of ${\Greekmath 0112}_0$ is set to be the MLE estimate based on the original dataset. More details about this numerical experiment are given in Appendix \ref{sec:Empirical Details}.
To evaluate the performance of the MERM estimator in these settings,
we generate mismeasured $Income_i = Income_i^* + {\Greekmath 0122}_i$. We focus on the individual income because it is often mismeasured.
We report the results for ${\Greekmath 011C} = {\Greekmath 011B}_{\Greekmath 0122}/{\Greekmath 011B}_{Income^*} \in \{1/4, 1/2, 3/4\}$.
Table \ref{tab: MC emprical} reports the simulation results for the (naive) MLE estimator and for the MERM estimators with $K=2$ and $K=4$. We focus on estimation of and inference on the income elasticities (evaluated at the population mean of the covariates).
The MLE estimator is considerably biased for ${\Greekmath 011C} \in \{1/2, 3/4\}$, which results in substantial size distortions of the MLE based t-tests. The MERM estimator with $K=4$ effectively eliminates the EIV bias and the corresponding t-tests provide accurate size control in all of the considered designs.
The estimator with $K=2$ is more precise, while successfully removing the EIV bias for ${\Greekmath 011C}\le1/2$.
Overall, the MERM estimators perform well in the considered empirical context, providing a basis for estimation and inference even for quite large values of ${\Greekmath 011C}$.
\begin{table}[h!]
\begin{footnotesize}
\begin{threeparttable}
\caption{Simulation results for the empirically calibrated conditional logit model}
\label{tab: MC emprical}
\begin{tabular}{ l c c c c| c c c c | c c c c}
\toprule
\toprule
& \multicolumn{4}{c}{MLE} & \multicolumn{4}{c}{$K=2$} & \multicolumn{4}{c}{$K=4$}\\
\cmidrule(lr){2-5} \cmidrule(lr){6-9} \cmidrule(lr){10-13}
&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{std}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{size}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{std}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{size}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{std}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{size}}\\
\midrule
\multicolumn{13}{c}{${\Greekmath 011C} = 1/4$}\\
\midrule
$\partial\ln p_1/\partial\ln I$ &-0.07&0.12&0.14&9.00&0.01&0.14&0.14&5.68&0.02&0.19&0.19&7.02\\
$\partial\ln p_2/\partial\ln I$ &0.03&0.07&0.08&5.84&-0.00&0.08&0.08&5.68&-0.01&0.10&0.10&6.40\\
$\partial\ln p_0/\partial\ln I$ &0.05&0.13&0.13&6.10&0.00&0.14&0.14&5.42&-0.00&0.17&0.17&7.66\\
\midrule
\multicolumn{13}{c}{${\Greekmath 011C} = 1/2$}\\
\midrule
$\partial\ln p_1/\partial\ln I$ &-0.24&0.11&0.27&61.84&-0.05&0.14&0.15&6.96&0.02&0.21&0.21&6.16\\
$\partial\ln p_2/\partial\ln I$ &0.09&0.07&0.11&24.76&0.02&0.09&0.09&5.96&-0.01&0.10&0.10&6.16\\
$\partial\ln p_0/\partial\ln I$ &0.16&0.12&0.20&25.36&0.04&0.15&0.15&6.06&-0.00&0.18&0.18&6.86\\
\midrule
\multicolumn{13}{c}{${\Greekmath 011C} = 3/4$}\\
\midrule
$\partial\ln p_1/\partial\ln I$ &-0.43&0.09&0.44&99.50&-0.19&0.14&0.24&27.46&0.02&0.22&0.22&5.84\\
$\partial\ln p_2/\partial\ln I$ &0.16&0.06&0.17&71.88&0.07&0.08&0.11&13.78&-0.01&0.11&0.11&6.32\\
$\partial\ln p_0/\partial\ln I$ &0.29&0.11&0.31&73.20&0.12&0.15&0.19&13.40&0.00&0.19&0.19&6.32\\
\bottomrule
\end{tabular}
\begin{tablenotes}
\item \footnotesize This table reports the simulated finite sample bias, standard deviation, RMSE, and size of the MLE and the MERM estimators and the corresponding t-tests for the income elasticities $\partial \ln p_j(I,w,{\Greekmath 0112}_0)/\partial \ln I$, $j \in \{1,2,0\}$, evaluated at the population mean. The true values of the income elasticities are $(\partial \ln p_1/\partial \ln I, \partial \ln p_2/\partial \ln I, \partial \ln p_0/\partial \ln I) = (1.11, -0.39, -0.82)$. The results are based on 5,000 replications.
\end{tablenotes}
\end{threeparttable}
\end{footnotesize}
\end{table}
\section{\label{sec:Extensions}Extensions}
\subsection{\label{ssec:adaptive K}Data-driven choice of $K$}
Making an appropriate choice of the expansion order $K$ is important for the estimation procedure.
One has to be cautious not to take $K$ too small, as this may result in an estimator that only partially removes the EIV bias. On the other hand, picking a larger $K$ than needed might inflate standard errors and result in less powerful inference.
In this section, we address this issue by providing a
data-dependent procedure for selecting~$K$.
We demonstrate that our procedure has desirable theoretical properties. We also find that the procedure has good finite sample properties in a set of Monte Carlo simulation experiments across different values of ${\Greekmath 011C}$ and sample sizes.
\bigskip
Consider two alternative values of the expansion order: $L$ and $K$, where $2 \leq L<K$. In practice, even-order biases tend to dominate, so to reduce the set of choices, it is useful to focus on even values of the expansion orders $L$ and $K$. For example, one would typically be interested in choosing between $L=2$ and $K=4$.
Let $\hat {\Greekmath 010C}_L$ and $\hat {\Greekmath 010C}_K$ denote the corresponding MERM estimators. The estimator $\hat {\Greekmath 010C}_L$ should be preferred as having smaller asymptotic variance
provided that its remaining EIV bias is negligible relative to its standard error. Otherwise, the more conservative $\hat {\Greekmath 010C}_K$ should be used instead.
Note that Lemma~\ref{lem: corrected moment} suggests that the remaining asymptotic bias of $\hat {\Greekmath 010C}_L$ (due to the additional terms accounted for when the expansion of higher order $K$ is used) is given by (up to an $o(n^{-1/2})$ remainder)
\begin{equation*}
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{AsyB}(\hat {\Greekmath 010C}_L) \equiv B \sum_{k=L+1}^{K}{\Greekmath 010D} _{0k}\mathbb{E}
[g_{x}^{(k)}(X_{i},S_{i},{\Greekmath 0112} _{0})] ,\quad
B \equiv - (\Psi ^{\prime }\Xi \Psi )^{-1}\Psi ^{\prime }\Xi,
\end{equation*}
where matrix $B$ is based on the moments used for estimation of $\hat {\Greekmath 010C}_L$.
Importantly, ${\Greekmath 010D} _{0k}=O\left( {\Greekmath 011B}_{{\Greekmath 0122}}^{k}\right) $, and hence for $k>2$ we can estimate a bound on $\sqrt{n}{\Greekmath 010D} _{0k}$ sufficiently quickly to provide a valid procedure for choosing $K$. To this end, we first estimate the model using the larger $K$. Let $\hat {\Greekmath 010C}_K = (\hat {\Greekmath 0112}', \hat {\Greekmath 010D}')'$ and $\hat {\Greekmath 011B}_{\Greekmath 0122}^2 \equiv 2 \hat {\Greekmath 010D}_{2}$, where we dropped the additional subscripts $K$ for notation simplicity. Then, we can estimate ${\Greekmath 011B}_{\Greekmath 0122}^K$ by $\hat {\Greekmath 011B}_{\Greekmath 0122}^K \equiv (\hat {\Greekmath 011B}_{\Greekmath 0122}^2)^{K/2}$. Using $\hat {\Greekmath 011B}_{\Greekmath 0122}^2 = {\Greekmath 011B}_{\Greekmath 0122}^2 + O_p (n^{-1/2})$, in the appendix we show that
\begin{equation}
\sqrt{n}\left( \hat{{\Greekmath 011B}}_{\Greekmath 0122}^{K}-{\Greekmath 011B}_{\Greekmath 0122} ^{K}\right) =O_{p}\left( {\Greekmath 011B}_{\Greekmath 0122}
^{K-2}+n^{-\left( K-2\right) /4}\right) =o_{p}\left( 1\right). \label{eq:estn sigK}
\end{equation}
Next, consider a sequence $\varkappa _{n}\rightarrow 0$, which we will specify
precisely later, and let
\begin{equation}
{\Greekmath 010E} _{n}=1\left\{ \max_{1\leq \ell \leq \dim ({\Greekmath 010C} )}\left\vert \hat{
\Sigma}_{\ell \ell }^{-1/2}\sqrt{n}\hat{{\Greekmath 011B}}_{\Greekmath 0122}^{K} \hat B_{\ell \cdot}
\overline{g}_{x}^{(K)}(\hat{{\Greekmath 0112}}) \right\vert \leq c_K \varkappa _{n}\right\}
, \label{eq:crit choice of K}
\end{equation}
where $\hat \Sigma$ and $\hat B$ are consistent estimators of the asymptotic variance of $\hat {\Greekmath 010C}_L$ and of $B$, with $\hat \Sigma_{\ell \ell}$ denoting its $\ell$-th diagonal element, $\hat B_{\ell \cdot}$ denoting its $\ell$-th row, ${\overline g_x^{(K)} (\hat {\Greekmath 0112}) \equiv n^{-1} \sum_{i=1}^n g^{(K)}(X_i,S_i,\hat {\Greekmath 0112})}$, and $c_K > 0$ is a constant that we will calibrate below after we state the main theoretical result of this section.
If ${\Greekmath 010E}_n = 1$, the researchers should select $\hat {\Greekmath 010C}_L$. Otherwise, $\hat {\Greekmath 010C}_K$ should be used. The following lemma demonstrates that the proposed selection procedure has desirable theoretical properties.
\begin{lemma}
\label{lem:choice-of-K}Suppose the hypotheses of Theorem 2 hold for some even $
K\geq 4$, and consider $L$ satisfying $2 \leq L<K$. Suppose $\varkappa
_{n}n^{\left( K-L-1\right) /\left( 2L+2\right) }\rightarrow 0$. Then
$\sqrt{n}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{AsyB}(\hat {\Greekmath 010C}_L) {\Greekmath 010E} _{n}=o_{p}\left(1\right)$.
Moreover, consider $\varkappa _{n} = n^{-\left( K-L-1\right) /\left(
2L+2\right) }\left( \ln n\right) ^{-a}$ for any $a>0$. Then the criterion is
consistent, in the sense that if ${\Greekmath 011C}_n =o( n^{-\frac{1}{2L+2}
-{\Greekmath 010F} }) $ for any ${\Greekmath 010F} >0$, the criterion will choose $\hat {\Greekmath 010C}_L$ with probability approaching one.
\end{lemma}
The first part of Lemma~\ref{lem:choice-of-K} shows that, provided that $\varkappa_n$ goes to zero sufficiently fast, $\hat {\Greekmath 010C}_L$ is selected (i.e., ${\Greekmath 010E}_n = 1$) only when its asymptotic bias is negligible. Next, note that $\hat {\Greekmath 010C}_L$ is asymptotically unbiased as long as ${\Greekmath 011C}_n = ( n^{-\frac{1}{2L+2}})$. The second part of the lemma shows that, for the suggested choices of $\varkappa_n$, the criterion is non-vacuous, i.e., that it does select the MERM estimator with a smaller expansion order $L$ when this is appropriate. Typically, one would pick $L=K-2$, so the lemma suggests taking $\varkappa
_{n} = n^{-1/\left( 2K-2\right) }\left( \ln n\right) ^{-a}$ in this case.
In practice, it is important to pick an appropriate constant $c_K$ used in the construction of ${\Greekmath 010E}_n$ in equation~\eqref{eq:crit choice of K}. When constructing ${\Greekmath 010E}_n$, we used $\left\vert {\Greekmath 010D}_{0K}\right\vert \propto {\Greekmath 011B}_{\Greekmath 0122}^K$, which motivates choosing $c_K = {\Greekmath 011B}_{\Greekmath 0122}^K / \left\vert {\Greekmath 010D}_{0K}\right\vert$. It is convenient to use a rule-of-thumb
approach, using a reference distribution to determine $c_K$. It turns
out that the normal distribution is not only convenient, but also sufficiently
conservative (note that the bigger ${\Greekmath 010D}_{0K}$ is, the smaller $c_K$ is, resulting in a more conservative selection procedure). For example, suppose $K=4$, and consider using
Student's $t({\Greekmath 0117} )$ distribution as the reference distribution. Then, for ${\Greekmath 0117} \geq 5$, the biggest $\left\vert {\Greekmath 010D}_{04}\right\vert/{\Greekmath 011B}_{\Greekmath 0122}^4$ and the smallest $c_4$ correspond to ${\Greekmath 0117} = \infty$ matching the normal distribution.\footnote{Note that Student's $t({\Greekmath 0117} )$ distribution does not have a finite $5$-th moment ${\Greekmath 0117} \leq 5$.}
Since for the normal distribution we have ${\Greekmath 010D}_{0K} = {\Greekmath 011B}_{\Greekmath 0122}^K / K!!$ for even $K$, we recommend using $c_K = K!!$ in equation~\eqref{eq:crit choice of K}. Finally, while we recommend using normal distribution as the reference distribution, we also stress that the results of Lemma~\ref{lem:choice-of-K} hold even if the measurement error is not normal or if it is skewed.
We summarize the proposed procedure in the following suggested algorithm.
\begin{algorithm}
\label{alg: choose K}
\leavevmode
\begin{enumerate}
\item Pick a large enough even $K\geq 4$ and compute $\hat {\Greekmath 010C}_K = (\hat {\Greekmath 0112}', \hat {\Greekmath 010D}')'$ and $\hat{{\Greekmath 011B}}_{\Greekmath 0122}^{2} = 2 \hat {\Greekmath 010D}_{2}$.
\item Compute ${\Greekmath 010E} _{n}$ in \eqref{eq:crit choice of K} with $L=K-2$, $c_K = K !!$, and $\varkappa
_{n}=\left( n\ln n\right) ^{-1/\left( 2K-2\right) }$.
\item Pick the length of expansion $L = K-2$ if ${\Greekmath 010E} _{n}=1$, otherwise keep the initial $K$.
\end{enumerate}
\smallskip
This procedure can be iterated if desired.
\end{algorithm}
\smallskip
To illustrate the performance of the algorithm provided above, we revisit the numerical experiment considered in Section~\ref{ssec: MC Mult Logit}. We focus on choosing between $K=4$ and $L=2$, and, as in Section~\ref{ssec: MC Mult Logit}, we report results for $n=2000$ in Table~\ref{tab: adaptive K mult logit 2000} below. To ensure that the proposed algorithm performs well in a variety of sample sizes, we also consider $n=1000$ and $n=4000$ and report the corresponding results in Tables~\ref{tab: adaptive K mult logit 1000}~and~\ref{tab: adaptive K mult logit 4000} in Appendix \ref{ssec: additional adaptive MCs}.
We report the finite sample bias and RMSE for the naive MLE estimator, as well as for the MERM estimators using $K=2$ and $K=4$, and for the adaptive MERM estimator using data-driven $K$ following Algorithm \ref{alg: choose K}. Notice that for the considered sample sizes, $K=2$ is preferred when ${\Greekmath 011C} = 1/4$, and $K=4$ is preferred when ${\Greekmath 011C} = 3/4$. We find that in both of these regimes and for all sample sizes, the adaptive estimator using data-driven $K$ is essentially equivalent to the preferred estimators, i.e., our procedure selects the appropriate $K$. Interestingly, in the intermediate regime with ${\Greekmath 011C} = 1/2$, the adaptive estimator also has the smallest bias among the considered estimators, coming at the cost of a slightly bigger RMSE compared to the MERM estimator using $K=4$. Thus, the considered numerical experiment suggests that our algorithm has good finite sample properties supporting the findings of Lemma~\ref{lem:choice-of-K}.
\begin{table}[h!]
\begin{scriptsize}
\begin{threeparttable}
\caption{Choice of $K$ simulation results for the multinomial logit model, $n=2000$}
\label{tab: adaptive K mult logit 2000}
\begin{tabular}{ l c c | c c | c c | c c}
\toprule
\toprule
& \multicolumn{2}{c}{MLE} & \multicolumn{2}{c}{$K=2$} & \multicolumn{2}{c}{$K=4$} & \multicolumn{2}{c}{data-driven $K$}\\
\cmidrule(lr){2-3} \cmidrule(lr){4-5} \cmidrule(lr){6-7} \cmidrule(lr){8-9}
&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{bias}, $10^{-2}$}&{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{rmse}, $10^{-2}$}\\
\midrule
\multicolumn{9}{c}{${\Greekmath 011C} = 1/4$}\\
\midrule
$\partial p_1/\partial x$&-3.24&3.51&0.74&2.74&1.13&2.95&0.75&2.75\\
$\partial p_1/\partial w_1$&2.32&2.84&-0.11&2.30&-0.31&2.31&-0.11&2.30\\
$\partial p_1/\partial w_2$&0.48&0.90&-0.04&0.87&-0.08&0.87&-0.04&0.87\\
$\partial p_2/\partial x$&1.96&2.28&-0.40&1.92&-0.63&2.03&-0.41&1.93\\
$\partial p_2/\partial w_1$&-1.16&1.42&0.06&1.15&0.15&1.16&0.06&1.15\\
$\partial p_2/\partial w_2$&-0.96&1.78&0.09&1.74&0.17&1.75&0.09&1.74\\
$\partial p_0/\partial x$&1.28&1.63&-0.34&1.47&-0.50&1.56&-0.34&1.48\\
$\partial p_0/\partial w_1$&-1.16&1.42&0.06&1.15&0.15&1.16&0.06&1.15\\
$\partial p_0/\partial w_2$&0.48&0.88&-0.05&0.87&-0.09&0.88&-0.05&0.87\\
\midrule
\multicolumn{9}{c}{${\Greekmath 011C} = 1/2$}\\
\midrule
$\partial p_1/\partial x$&-8.97&9.04&-1.69&3.10&0.97&3.05&0.88&3.15\\
$\partial p_1/\partial w_1$&6.44&6.62&1.39&2.81&-0.21&2.42&-0.15&2.48\\
$\partial p_1/\partial w_2$&1.28&1.47&0.29&0.94&-0.06&0.92&-0.05&0.93\\
$\partial p_2/\partial x$&5.22&5.31&1.05&2.16&-0.53&2.15&-0.48&2.20\\
$\partial p_2/\partial w_1$&-3.21&3.30&-0.69&1.40&0.10&1.21&0.08&1.24\\
$\partial p_2/\partial w_2$&-2.52&2.88&-0.56&1.87&0.13&1.84&0.11&1.85\\
$\partial p_0/\partial x$&3.75&3.85&0.64&1.59&-0.44&1.65&-0.40&1.68\\
$\partial p_0/\partial w_1$&-3.23&3.32&-0.70&1.41&0.10&1.21&0.08&1.24\\
$\partial p_0/\partial w_2$&1.23&1.41&0.28&0.93&-0.07&0.92&-0.06&0.93\\
\midrule
\multicolumn{9}{c}{${\Greekmath 011C} = 3/4$}\\
\midrule
$\partial p_1/\partial x$&-13.35&13.38&-6.83&7.32&0.71&3.29&0.71&3.29\\
$\partial p_1/\partial w_1$&9.69&9.80&4.95&5.61&0.01&2.62&0.01&2.62\\
$\partial p_1/\partial w_2$&1.81&1.94&1.01&1.35&-0.01&0.98&-0.01&0.98\\
$\partial p_2/\partial x$&7.48&7.52&4.06&4.45&-0.37&2.35&-0.37&2.35\\
$\partial p_2/\partial w_1$&-4.83&4.88&-2.47&2.81&-0.01&1.31&-0.01&1.31\\
$\partial p_2/\partial w_2$&-3.51&3.76&-1.99&2.66&0.03&1.97&0.03&1.97\\
$\partial p_0/\partial x$&5.87&5.92&2.77&3.14&-0.34&1.80&-0.34&1.80\\
$\partial p_0/\partial w_1$&-4.87&4.93&-2.48&2.81&-0.01&1.31&-0.01&1.31\\
$\partial p_0/\partial w_2$&1.70&1.82&0.98&1.31&-0.02&0.99&-0.02&0.99\\
\bottomrule
\end{tabular}
\begin{tablenotes}
\item \scriptsize This table reports the simulated finite sample bias and RMSE of the MLE and the MERM estimators for the partial derivatives $\partial p_j(x,w,{\Greekmath 0112}_0)/\partial x$, $\partial p_j(x,w,{\Greekmath 0112}_0)/\partial w_1$, $\partial p_j(x,w,{\Greekmath 0112}_0)/\partial w_2$ for $j \in \{1,2,0\}$ evaluated at the population mean. The true values of the marginal effects are $(\partial p_1/\partial x, \partial p_2/\partial x, \partial p_0/\partial x) = ( 0.222, -0.111, -0.111)$ and zeros for the rest. The results are based on 5,000 replications.
\end{tablenotes}
\end{threeparttable}
\end{scriptsize}
\end{table}
\subsection{\label{sec:Multivariate}Multiple Mismeasured Variables}
It is easy to use the MERM framework to deal with multiple mismeasured variables. This is useful in many applications, including not only settings with multiple mismeasured covariates, but also settings with serially correlated measurement errors, settings where repeated measurements are available, and panel data models. Using the MERM approach is particularly advantageous in such applications, since it avoids nonparametric estimation of multivariate unobserved distributions.
Suppose $X_i^*$, ${\Greekmath 0122}_{i}$, and $X_i$ are $d \times 1$ vectors. Let ${\Greekmath 011C}_n \equiv \max_{j\le d} {\Greekmath 011B}_{{\Greekmath 0122}_j}/{\Greekmath 011B}_{X^*_j}$, where ${\Greekmath 011B}_{{\Greekmath 0122}_j}$ and ${\Greekmath 011B}_{X^*_j}$ denote the standard deviations of the $j$-th components of ${\Greekmath 0122}_i$ and $X_i^*$, so $\mathbb{E}\left[\left\vert {\Greekmath 0122}_{ij}\right\vert^k\right] = O({\Greekmath 011C}_n^k)$ for $k\in\{1,\ldots,K\}$.
For a $d \times 1$ vector of non-negative integers ${\Greekmath 0114} = ({\Greekmath 0114}_1,\ldots,{\Greekmath 0114}_d) \in \mathbb Z_+^{d}$, let
\begin{equation*}
\partial_{\Greekmath 0114} \equiv \frac{\partial^{\left\vert {\Greekmath 0114}\right\vert}}{\partial {x_{1}}^{{\Greekmath 0114}_1} \ldots \partial {x_d}^{{\Greekmath 0114}_d}},\quad \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{where } \left\vert {\Greekmath 0114}\right\vert \equiv \sum_{j=1}^d {\Greekmath 0114}_j.
\end{equation*}
Also, for a positive integer $k$, let $\mathcal K_k = \{{\Greekmath 0114} \in \mathbb Z_+^d: \left\vert {\Greekmath 0114}\right\vert = k\}$. Then, we consider the following corrected moment function
\begin{equation*}
{\Greekmath 0120}(x,s,{\Greekmath 0112},{\Greekmath 010D}) = g(x,s,{\Greekmath 0112}) - \sum_{k = 2}^{K} \sum_{{\Greekmath 0114} \in \mathcal K_k } {\Greekmath 010D}_{\Greekmath 0114} \partial_{{\Greekmath 0114}} g(x,s,{\Greekmath 0112}),
\end{equation*}
where, with some abuse of notation, ${\Greekmath 010D}$ is a collection of all ${\Greekmath 010D}_{\Greekmath 0114}$ with ${\Greekmath 0114} \in \mathcal K_k$ and $k \in \{2, \dots, K\}$.
Under mild smoothness conditions
\begin{equation*}
\mathbb{E}\left[{\Greekmath 0120}(X_i,S_i,{\Greekmath 0112}_0,{\Greekmath 010D}_{0})\right] = \mathbb{E}\left[g(X_i^*,S_i,{\Greekmath 0112}_0)\right] + O({\Greekmath 011C}_n^{K+1}) = o(n^{-1/2}),
\end{equation*}
where the second equality holds provided that $O({\Greekmath 011C}_n^{K+1}) = o(n^{-1/2})$. Similarly to the scalar case, components of ${\Greekmath 010D}_{0}$ are determined by the moments of ${\Greekmath 0122}_i$. Specifically, let ${\Greekmath 0116}_{{\Greekmath 0114}} \equiv \mathbb{E}\left[{\Greekmath 0122}_{i 1}^{{\Greekmath 0114}_1} \dots {\Greekmath 0122}_{i d}^{{\Greekmath 0114}_d}\right]$, then
\begin{equation}
\label{eq: kappa 2 and 3}
{\Greekmath 010D}_{0 {\Greekmath 0114}} = \frac{{\Greekmath 0116}_{{\Greekmath 0114}}}{{\Greekmath 0114}!}, \quad \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for } {\Greekmath 0114} \in \{\mathcal K_2, \mathcal K_3\},
\end{equation}
where ${\Greekmath 0114}! \equiv {\Greekmath 0114}_1! \ldots {\Greekmath 0114}_d!$.
For $\left\vert {\Greekmath 0114}\right\vert \geq 4$, the coefficients can be computed by the following formulas. For example, for ${\Greekmath 0114} \in \mathcal K_4$, let $\mathcal K_{2, {\Greekmath 0114}} = \{\tilde {\Greekmath 0114} \in \mathcal K_2: {\Greekmath 0114} - \tilde {\Greekmath 0114} \in \mathcal K_2 \}$. Then,
\begin{equation*}
{\Greekmath 010D}_{0 {\Greekmath 0114}} = \frac{{\Greekmath 0116}_{\Greekmath 0114}}{{\Greekmath 0114}!} - \sum_{\tilde {\Greekmath 0114} \in \mathcal K_{2, {\Greekmath 0114}}} \frac{{\Greekmath 0116}_{{\Greekmath 0114} - \tilde {\Greekmath 0114}}}{({\Greekmath 0114} - \tilde {\Greekmath 0114})!} {\Greekmath 010D}_{0 \tilde {\Greekmath 0114}} , \quad \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for } {\Greekmath 0114} \in \mathcal K_4.
\end{equation*}
More generally, for ${\Greekmath 0114} \in \mathcal K_{k}$ with $k \geq 4$, let $\mathcal K_{\ell, {\Greekmath 0114}} = \{\tilde {\Greekmath 0114} \in \mathcal K_\ell, {\Greekmath 0114} - \tilde {\Greekmath 0114} \in \mathcal K_{\left\vert {\Greekmath 0114}\right\vert - \ell} \}$ for $\ell \leq \left\vert {\Greekmath 0114}\right\vert - 2$. Then,
\begin{equation*}
{\Greekmath 010D}_{0 {\Greekmath 0114}} = \frac{{\Greekmath 0116}_{\Greekmath 0114}}{{\Greekmath 0114}!} - \sum_{\ell=2}^{k-2} \sum_{\tilde {\Greekmath 0114} \in \mathcal K_{\ell, {\Greekmath 0114}}} \frac{{\Greekmath 0116}_{{\Greekmath 0114} - \tilde {\Greekmath 0114}}}{({\Greekmath 0114} - \tilde {\Greekmath 0114})!} {\Greekmath 010D}_{0 \tilde {\Greekmath 0114}}.
\end{equation*}
\begin{example*}[Bivariate $X$, $K = 4$]\hfill \\
Suppose $X$ is bivariate (i.e., $d = 2$) and $K = 4$. For ${\Greekmath 0114} \in \mathcal K_2 = \{(2,0),(1,1),(0,2)\}$ and ${\Greekmath 0114} \in \mathcal K_3 = \{(3,0), (2,1), (1,2), (0,3)\}$, ${\Greekmath 010D}_{0 {\Greekmath 0114}}$ is given by \eqref{eq: kappa 2 and 3}. For ${\Greekmath 0114} \in \mathcal K_4$, ${\Greekmath 010D}_{0 {\Greekmath 0114}}$ is given by
\begin{center}
\begin{tabular}{c|c}
${\Greekmath 0114}$ & ${\Greekmath 010D}_{0 {\Greekmath 0114}}$ \\
\midrule
(4,0) & $\left(\mathbb{E}[{{\Greekmath 0122}_{i1}^4}] - 6 \mathbb{E}[{{\Greekmath 0122}_{i1}^2}]^2\right)/24$ \\
(3,1) & $\left(\mathbb{E}[{{\Greekmath 0122}_{i1}^3 {\Greekmath 0122}_{i2}}] - 6 \mathbb{E}[{{\Greekmath 0122}_{i1}^2}] \mathbb{E}[{{\Greekmath 0122}_{i1} {\Greekmath 0122}_{i2}}]\right)/6$ \\
(2,2) & $\left(\mathbb{E}[{{\Greekmath 0122}_{i1}^2 {\Greekmath 0122}_{i2}^2}] - 2 \mathbb{E}[{{\Greekmath 0122}_{i1}^2}] \mathbb{E}[{{\Greekmath 0122}_{i2}^2}] - 4 \mathbb{E}[{{\Greekmath 0122}_{i1} {\Greekmath 0122}_{i2}}]^2\right)/4$ \\
(1,3) & $\left(\mathbb{E}[{{\Greekmath 0122}_{i1} {\Greekmath 0122}_{i2}^3}] - 6 \mathbb{E}[{{\Greekmath 0122}_{i2}^2}] \mathbb{E}[{{\Greekmath 0122}_{i1} {\Greekmath 0122}_{i2}}]\right)/6$ \\
(0,4) & $\left(\mathbb{E}[{{\Greekmath 0122}_{i2}^4}] - 6 \mathbb{E}[{{\Greekmath 0122}_{i2}^2}]^2\right)/24$
\end{tabular}
\end{center}
If in addition measurement errors ${\Greekmath 0122}_{i1}$ and ${\Greekmath 0122}_{i2}$ are independent, ${\Greekmath 010D}_{0 {\Greekmath 0114}} = 0$ for ${\Greekmath 0114} \in \{(1,1),(2,1),(1,2),(3,1),(1,3)\}$. In this case, the total number of the nuisance parameters to be estimated is 6. \hfill\ensuremath{\blacksquare}
\end{example*}
\begin{singlespacing}
\bibliographystyle{ecta}
\bibliography{library_out}
\end{singlespacing}