EconBase
← Back to paper

Optimally-Transported Generalized Method of Moments

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

60,989 characters

Optimally-Transported Generalized Method of Moments



\title{Optimally-Transported\\Generalized Method of Moments}
\author{Susanne Schennach\thanks{
Support from NSF grants SES-1950969 and SES-2150003 is gratefully acknowledged. The authors would like to thank Alfred Galichon, Florian Gunsilius, Heejun Lee and seminar participants at NYU's Advanced mathematical modeling in economics seminar, at the Econometrics and Optimal Transport Workshop at the University of Washington, at the Women in Econometrics Conference at the University of Toronto, at the Workshop on Optimal Transport in Econometrics at Collegio Carlo Alberto, and at the Aarhus Workshop in Econometrics VII for useful comments.}~\thanks{
[email removed]}~ and Vincent Starck\thanks{Financial support from the European Research Council (Starting Grant No. 852332) is gratefully acknowledged.}~\thanks{
[email removed]} \\
Brown University, LMU München}
\maketitle

\begin{abstract}

We propose a novel optimal transport-based version of the Generalized
Method of Moment (GMM).
Instead of handling overidentification by reweighting
the data to satisfy the moment conditions (as in Generalized
Empirical Likelihood methods), this method proceeds by allowing for
errors in the variables of the least mean-square magnitude necessary to
simultaneously satisfy all moment conditions. This approach, based on the
notions of optimal transport and Wasserstein metric, aims to address the problem of assigning a
logical interpretation to GMM results even when overidentification tests
reject the null, a situation that cannot always be avoided in applications.
We illustrate the method by revisiting Duranton, Morrow and Turner's (2014) study of the relationship between a city's exports and the extent of its transportation infrastructure.
Our results corroborate theirs under weaker assumptions and provide
insight into the error structure of the variables.

\medskip

\noindent {\bf Keywords:} Wasserstein metric, GMM, overidentification, misspecification.

\end{abstract}

\section{Introduction}

The Generalized Method of Moment (GMM) (\citet{hansen:gmm}) has long been
the workhorse of statistical modeling in economics and the social sciences.
Its key distinguishing feature, relative to the basic method of moments, is
the presence of overidentifying restrictions that enable the model's
validity to be tested (\citet{Newey:HB}). With this ability to test comes
the obvious practical question of what one should do if an overidentified
GMM model fails overidentification tests, a situation that is not uncommon
(as noted in \citet{hall:missgmm}, \citet{hansen:miss},
\citet{poirier:salvaging}, \citet{conley:plausexo}, \citet{kwon:ineqmiss}),
even for perfectly reasonable, economically grounded, models.

A popular approach has been to find the \textquotedblleft
pseudo-true\textquotedblright\ value of the model parameter (
\citet{sawa:pseudo}, \citet{white:miss}) that minimizes the ``distance''
or \emph{discrepancy} between the data and the moment constraints implied by
the model. This approach has gained further support since the introduction
of Generalized Empirical Likelihoods (GEL) and Minimum Discrepancy
estimators (\citet{owen:el1}, \citet{qinlawless}, \citet{newey:gel}), all of
which provide more readily interpretable pseudo-true values (
\citet{imbens:res}, \citet{kitamura:itgmm}, \citet{schennach:etel}).

GEL implicitly
attributes the mismatch in the moment conditions solely to a biased sampling of the population. While this
is a possible explanation, it is not the only reason a valid model would
fail overidentification tests, when taken to the data. Another natural
possibility is the presence of errors in the variables (\citet{aguiar:revpref},
\citet{doraszelski:megmm}, \citet{schennach:HB}). In this work, we develop
an alternative to GMM that ensures, by construction, that overidentifying
restrictions are satisfied by allowing for possible errors in the
variables instead of sampling bias. We employ the generic term {\em error} to include, not only measurement error, but anything that could cause the recorded data to differ from the value they should have if the model were fully correct, i.e., this could include some model errors.
More generally, we allow distortions in the data generating process whose magnitude is quantified by a Wasserstein-type metric (\citet{villani:new}),
in the spirit of distributionally robust methods (\citet{connault:sens}, \citet{blanchet:distrob}).
In analogy with GEL, which does
not require the form of the sampling bias to be explicitly specified, the
error process does not need to be explicitly specified in our
approach, but is instead inferred from the requirement of satisfying the
overidentifying constraints imposed by the GMM model.
Of course, the accuracy of the
resulting estimated parameters will typically improve with the degree of overidentification.
\footnote{In the case where the statistical properties of the errors are in fact
known a priori, other methods may be more appropriate (e.g. \citet{schennach:nlme}, \citet{schennach:elvis}),
\citet{schennach:annrev}, \citet{schennach:HB} and references therein.}

A fruitful way to accomplish this is to employ concepts from the general
area of optimal transport problems (e.g., \citet{galichon:optran},
\citet{villani:new}, \citet{carlier:vectq}, \citet{galichon:vectquan}, \citet{cherno:breniercons},
\citet{schennach:nlfact}). The idea is to find the parameter value that
minimizes cost of \textquotedblleft transporting\textquotedblright\ the observed
distribution of the data $\mu_x$ onto another distribution $\mu_z$ that satisfies the moment
conditions exactly.
Formally, the true iid data $z_{i}$ is assumed to satisfy $\mathbb{E}\left[ g\left(
z_{i},\theta \right) \right] =0$, where $\mathbb{E}$ is the expectation
operator, for a parameter value $\theta$ in some set $\Theta$ and some given $d_{g}$-dimensional vector $
g\left( z_{i},\theta \right) $ of moment functions. However, we instead
observe an error-contaminated counterpart $x_{i}$ of the true vector $z_{i}$ (both taking value in $\mathcal{X} \subseteq \mathbb{R}^{d_{x}}$). We seek to exploit the model's over-identification to gain information regarding the error in $x_{i}$.
The Euclidean norm $\Vert (z-x) \Vert$ is chosen here for computational convenience, although one could imagine a whole class of related estimators obtained with different choices of metric. Our focus on Euclidean norms parallels the
choice made in common estimators (e.g. least squares regressions,
classical minimum distance and even GMM). Considering a {\it weighted} Euclidean norm can be useful to indicate the relative expected error magnitudes along different dimensions of $x$.

Given a probability measure $\mu_x$ for the random variable $x$, this setup suggests solving the following population optimization problem, for a given $\theta$:
\begin{equation}
\min_{\mu _{zx}}\mathbb{E}_{\mu _{zx}}\left[ \left\Vert z-x\right\Vert ^{2} \label{eqobjgen}
\right]
\end{equation}
subject to $\mu _{zx}$, supported on $\mathcal{X}\times \mathcal{X}$, having marginal $\mu _{x}$ and $\mathbb{E}_{\mu _{zx}} \left[ g\left( z,\theta \right) \right] =0$, where $\mathbb{E}_{\mu}$ denotes an expectation under the measure $\mu$. (This problem is guaranteed to have a solution if there exists at least one measure $\mu_z^*$ such that $\mathbb{E}_{\mu_z^*} \left[ g\left( z,\theta \right) \right] =0$.)
This setup covers the most general case, including both discrete and continuous variables, and can be handled using linear programming techniques
(e.g., \citet{santambrogio:optran}, Section 6.4.1), after observing that the moment constraint is easy to incorporate since it is linear in the probability measure. However, we shall focus on the purely continuous case in the remainder of this paper, because it enables us to express the main ideas more transparently. The fully continuous case indeed admits a convenient treatment, under the following regularity condition:
\begin{condition}
The marginals $\mu_z$ (arising from the solution $\mu_{zx}$ at each $\theta\in\Theta$) and $\mu_x$ have finite variance and $\mu_x$ is absolutely continuous with respect to the Lebesgue measure.
\end{condition}
Under this condition, by Theorem 1.22 in \citet{santambrogio:optran}, there exists a unique $\mu_{zx}$ implied by a deterministic transport map $z=q(x)$ that solves the constrained optimization problem (\ref{eqobjgen}) and yielding a transport cost $\mathbb{E}_{\mu_x} \left[ \left\Vert q(x)-x\right\Vert ^{2}\right]$.
Since determining the function $q$ amounts to finding which value $z$ each point $x$ should be mapped to, the sample version of this problem can be stated as
\begin{equation}
\min_{\left\{ z_{i}\right\} }\frac{1}{2}\hat{\mathbb{E}}\left[ \left\Vert
z-x\right\Vert ^{2}\right]  \label{eqobjf}
\end{equation}
subject to:
\begin{equation}
\hat{\mathbb{E}}\left[ g\left( z,\theta \right) \right] =0,  \label{eqcons}
\end{equation}
where $\hat{\mathbb{E}}$ denotes sample averages (i.e. $\hat{
\mathbb{E}}[a\left( x\right) ]\equiv \frac{1}{n}\sum_{i=1}^{n}a\left(
x_{i}\right) $, where $n$ is sample size and $a(\cdot)$ a given function).
This optimization problem is then nested into an optimization over $\theta $,
which delivers the estimated parameter value $\hat{\theta}$. We call this
estimator an \emph{Optimally-Transported}\ GMM (OTGMM) estimator.

Our approach is conceptually similar to GEL, in that it
minimizes some concept of distributional distance under moment constraints.
Yet, the notion of distance used differs significantly. As shown
in Figure \ref{figotgmm}, the distance here is measured along the
\textquotedblleft observation values\textquotedblright\ axis rather than the
\textquotedblleft observation weights\textquotedblright\ axis (as it would
be in GEL). This feature arguably makes the method a hybrid between GEL and
optimal transport, since GEL's goal of satisfying all the moment
conditions is achieved through optimal transport instead of optimal reweighting.
(The distinction from GEL applies to the discrete case as well,
since the OTGMM objective function depends on both the amount and location of probability mass transfers,
while the GEL objective function is only sensitive on the amount, but not on the specific locations, of the probability mass transfers.)
Most of our regularity conditions will parallel those of optimal GMM and GEL, but
some will not (they are tied to the optimal transport nature of the problem and involve assumptions regarding higher order derivatives).
In analogy with the behavior of GEL estimators, the OTGMM estimator
will be shown to be root-$n$ consistent and asymptotically normal, despite involving an optimization problem having
an infinite-dimensional nuisance parameter (the $z_i$).
In general, however, OTGMM's asymptotic variance does not coincide with that of GEL or efficiently weighted GMM under correct specification.
Hence, OTGMM is most useful when errors in the variables or Wasserstein-type deviations in the data distribution constitute a primary concern.

\begin{figure}
\centerline{\includegraphics[width=3.5in]{otgmm.eps}}
\caption{Comparison between the sample points adjustments for
(a) Generalized Empirical Likelihood (GEL), where observation weights (shown by point size) are adjusted,
and (b) Optimal Transport GMM, where point positions are adjusted.
Simple case of overidentified (parameter-free) model imposing no correlation shown, with original sample in gray
and adjusted sample in black.}
\label{figotgmm}
\end{figure}

The remainder of the paper is organized as follows. We first formally define
and solve the optimization problem defining our estimator, before
considering the limit of small errors (in the spirit of
\citet{chesher:1991})\ to gain some intuition. We then derive the resulting
estimator's asymptotic properties for the general, large error, case.
We show asymptotic normality and root $n$ consistency in both cases.
We then discuss related approaches and some extensions.
The method's practical usefulness is illustrated by revisiting
an influential study of the relationship between a city's exports and the extent of its transportation infrastructure (\citet{duranton2014roads}).
Our results corroborate that study under weaker assumptions and provide insight into the error structure of the variables.

\section{The estimator}

\subsection{Definition}

The Lagrangian associated with the constrained optimization problem defined
in Equations (\ref{eqobjf}) and (\ref{eqcons}) is
\begin{equation*}
\frac{1}{2}\hat{\mathbb{E}}\left[ \left\Vert z-x\right\Vert ^{2}\right]
-\lambda ^{\prime }\hat{\mathbb{E}}\left[ g\left( z,\theta \right) \right],
\end{equation*}
where $\hat{\mathbb{E}}[\ldots ]$ denotes a sample average and where $
\lambda $ is a Lagrange multiplier. The dual problem's first-order conditions
with respect to $\theta $, $\lambda $ and $z_{j}$, respectively, are
then
\begin{eqnarray}
\hat{\mathbb{E}}\left[ \partial_{\theta}g^{\prime }(z,\theta )\right] \lambda
&=&0  \label{eqfoctheta} \\
\hat{\mathbb{E}}\left[ g(z,\theta )\right] &=&0  \label{eqfoclambda} \\
\left( z_{j}-x_{j}\right) -\partial_{z}g^{\prime }(z_{j},\theta )\lambda
&=&0\text{ for }j=1,\ldots ,n  \label{eqfoczj}
\end{eqnarray}
where we let $\partial_{v}$ denote a partial derivative with respect to argument
$v$. We shall use $\partial_{v^{\prime }}$ to
denote a matrix of partial derivatives with respect to a transposed variable
(e.g., $\partial_{\theta^{\prime }}g(z,\theta )\equiv \partial g(z,\theta
)/\partial \theta ^{\prime }$). This formulation of the problem assumes
differentiability of $g\left( z,\theta \right) $ to a sufficiently high
order, as shall be formalized in our asymptotic analysis.

\subsection{Implementation}

The nonlinear system (\ref{eqfoctheta})-(\ref{eqfoczj}) of equations can be
solved numerically. To this effect, we propose an iterative procedure to
determine the $z_{j}$, $\lambda $ for a given $\theta $. This yields an\
objective function $\hat{Q}(\theta )$ that can be minimized to estimate $
\theta $. Let $z_{j}^{t}$ and $\lambda ^{t}$ denote the approximations
obtained after $t$ steps. As shown in Supplement \ref{appiter}, given
tolerances $\epsilon ,\epsilon ^{\prime }$ and a given $\theta $, the
objective function $\hat{Q}(\theta )$ can be determined as follows:

\begin{algorithm}
\label{algoiter}

\begin{enumerate}
\item Start the iterations with $z_{j}^{0}=x_{j}$ and $t=0$.

\item Let $
\lambda ^{t+1}=\left( \hat{\mathbb{E}}\left[ H\left( z^{t},\theta \right)
H^{\prime }\left( z^{t},\theta \right) \right] \right) ^{-1}\left( -\hat{
\mathbb{E}}\left[ g\left( z^{t},\theta \right) \right] +\hat{\mathbb{E}}
\left[ H\left( z^{t},\theta \right) \left( z^{t}-x\right) \right] \right)$ and
$z_{j}^{t+1}=x_{j}+H^{\prime }\left( z_{j}^{t},\theta \right) \lambda ^{t+1}$,
where $H\left( z,\theta \right) =\partial_{z^{\prime }}g\left( z,\theta
\right) =\left( \partial_{z}g^{\prime }\left( z,\theta \right) \right)
^{\prime }$.

\item Increment $t$ by $1$; repeat from step 2 until $\left\Vert
z_{j}^{t+1}-z_{j}^{t}\right\Vert \leq \epsilon $ and $\left\Vert \lambda
^{t+1}-\lambda ^{t}\right\Vert \leq \epsilon ^{\prime }$.

\item The objective function is then:
$\hat{Q}(\theta )=\frac{1}{2}(\lambda ^{t})^{\prime }\hat{\mathbb{E}}\left[
H\left( z^{t},\theta \right) H^{\prime }\left( z^{t},\theta \right) \right]
\lambda ^{t}\text{.}$

\end{enumerate}
\end{algorithm}

This algorithm is obtained by substituting $z_{j}=x_{j}+H^{\prime }\left(
z_{j},\theta \right) \lambda $ obtained from Equation (\ref{eqfoczj}) into
Equation (\ref{eqfoclambda}) and expanding the resulting expression to
linear order in $\lambda $. This linearized expression provides an improved
approximation $\lambda ^{t}$ to the Lagrange multiplier which can, in turn,
yield an improved approximation $z_{j}^{t}$. The process is then iterated to
convergence. The expression for $\hat{Q}(\theta )$ is obtained by
re-expressing $\hat{\mathbb{E}}\left[ \left\Vert z-x\right\Vert ^{2}\right] $
using Equation (\ref{eqfoczj}). Formal sufficient conditions for the
convergence of this iterative procedure can be found in Supplement \ref
{appiterconv}. In cases where this simple approach fails to converge, one can employ
robust schemes based on a combination of discretization and linear programming
(see Section 6.4.1 in \citet{santambrogio:optran}).

To gain some intuition regarding the estimator, it is useful to consider the
limit of small errors when solving Equations (\ref{eqobjf})-(\ref
{eqcons}), in the spirit of \citet{chesher:1991}).
This limit corresponds to assuming that higher-order powers of $\Vert z_i-x_i \Vert$
are negligible relative to $\Vert z_i-x_i \Vert$ itself.
In this limit, the estimator admits a closed form with an
intuitive interpretation, as shown by the following result, shown in
Appendix \ref{applin}.

\begin{proposition}
\label{propsml}To the first order in $z_{i}-x_{i}$ ($i=1,\ldots ,n$) the
estimator is equivalent to minimizing a GMM-like objective function with
respect to $\theta $ with a specific choice of weighting matrix:
\begin{equation}
\hat{\theta}=\arg \min_{\theta }\hat{\mathbb{E}}\left[ g^{\prime }\left(
x,\theta \right) \right] \left( \hat{\mathbb{E}}\left[ H\left( x,\theta
\right) H^{\prime }\left( x,\theta \right) \right] \right) ^{-1}\hat{\mathbb{
E}}\left[ g\left( x,\theta \right) \right] .  \label{eqsmlerr}
\end{equation}
\end{proposition}

From this expression, it is clear that the estimator downweights the
moments that are the most sensitive to errors in $x$, as measured by $
H\left( x,\theta \right) \equiv \partial_{z^{\prime }}g\left( x,\theta
\right) $. This accomplishes the desired goal of minimizing the effect of
the errors when the properties of the error process are unknown.

Although this weighting matrix appears suboptimal (relative to a correctly specified optimally weighted
GMM estimator), one should realize that the notion of optimality depends on what class of data generating processes the ``true model'' encompasses.
Optimal GMM and GEL can be seen as Maximum Likelihood estimators (\citet{chamberlain:gmmeff}, \citet{newey:gel}) under moment conditions expressed in terms of the observed $x$.
In contrast, OTGMM can be interpreted as a Maximum Likelihood estimator for homoskedastic and normally distributed errors ($x-z$) under moment constraints on the unobserved $z$.
There is therefore a clear efficiency-robustness trade-off: OTGMM is less efficient if the observed $x$ satisfy the moment constraints, but allows for additional error terms that maintain the model's correct specification even if the observed $x$ do not satisfy the overidentified moment constraints, a more general setting where optimally weighted GMM or GEL offers no efficiency guarantees.

\section{Asymptotics}

In this section, we show that, despite the estimator's roots in the theory
of optimal transport, its large sample behavior remains amenable to standard
asymptotic tools since our focus is on an estimator of the parameter $\theta $
rather than on an estimator of a distribution. We first consider the case
of small errors, a limiting case that may be especially important in the
relatively common case of applications where overidentifying restrictions
tests are near the rejection region boundary. This limit also parallels the
approach taken in the GEL literature, where asymptotic properties are often
derived in the case where the overidentifying restrictions hold (e.g.,
\citet{newey:gel}).

\subsection{Small errors limit}

Our small error results enable us to illustrate that there is little risk
in using our estimator instead of efficient GMM when one is concerned about
overidentification test failure. If the data were to, in fact,
satisfy the moment conditions, using our approach does not sacrifice
consistency, root $n$ convergence or asymptotic normality. The only possible
drawback would be a suboptimal weighting of overidentifying moment
conditions, potentially leading to an increase in variance if the model happened to be correctly specified.
Conversely, if the model is misspecified, e.g., because the data is error-contaminated, the optimal weighting of efficient GMM is no longer the optimal weighting (since random deviations due to sampling variability are not the main reason for the failure to simultaneously satisfy all moment conditions).
For instance, if the errors are such that there is an unknown bias in the moment conditions that decays to zero asymptotically but at a rate possibly slower than $n^{-1/2}$, then the model is still correctly specified asymptotically but the bias dominates the random sampling error. Then, the optimal weighting should seek to minimize the effect of error-induced bias, which our approach seeks to accomplish by weighting based on the effect of errors in the variables on the moment conditions.
Hence, in that sense, the method provides a complementary alternative to standard GMM estimation offering a different trade-off between efficiency and robustness to misspecification.

Our consistency result requires a number of fairly standard primitive
assumptions.
\begin{condition}
\label{assiid}The random variables $x_{i}$ are iid and take
value in $\mathcal{X}\subset \mathbb{R}^{d_{x}}$.
\end{condition}

\begin{condition}
\label{ass1sml}$\mathbb{E}[g(x_{i};\theta _{0})]=0$, and $\mathbb{E}[g(x_{i};\theta )]\neq 0$
for other $\theta \in \Theta $, a compact set.
\end{condition}

In other words, Assumption \ref{ass1sml} indicates that we consider here the
case where GMM would be consistent, in analogy with the setup traditionally considered in the GEL literature (e.g. \citet{newey:gel}).

\begin{condition}
\label{ass2sml}$\mathbb{V}[g(x_{i};\theta _{0})]<\infty $, where $\mathbb{V}$
denotes the variance operator.
\end{condition}

\begin{condition}
\label{ass5sml}$g(x;\cdot )$ is almost surely continuous and $\Vert
g(x;\theta )\Vert\leq h(x)$ for any $\theta \in \Theta $ and for some function $
h$ satisfying $\mathbb{E}\left[ h\left( x_{i}\right) \right] <\infty $.
\end{condition}

While Assumptions \ref{assiid}, \ref{ass1sml}, \ref{ass2sml} and \ref{ass5sml}
directly parallel those needed to establish the asymptotic properties of a
standard GMM estimator (e.g. Theorems 2.6 and 3.2 in \citet{Newey:HB}),
our estimator requires a few more low-level regularity conditions.
Given that our estimator, in the small error limit (Equation (\ref
{eqsmlerr})), involves a sample average involving derivative $H\left(
x,\theta \right) \equiv \partial_{z^{\prime }}g\left( x,\theta \right) $,
we need to place some constraints on the behavior of that quantity as well.
Below, we let $\left\Vert a\right\Vert =\left(\sum_{i,j}a_{i,j}^2\right) ^{1/2}$ for a matrix $a$.

\begin{condition}
\label{ass3sml}$g$ is differentiable in its first argument and the
derivative satisfies $\mathbb{E}[\Vert \partial_{z^{\prime }}g(x_{i};\theta
_{0})\Vert ^{2}]\allowbreak <\infty $. Moreover, $\Vert \partial_{z^{\prime
}}g(x_{i};\theta )\Vert \neq 0$ almost surely for all $\theta \in \Theta $.
\end{condition}

\begin{condition}
\label{ass4sml}$\partial_{z^{\prime }}g(x;\theta _{0})$ is H\"{o}lder
continuous in $x$.
\end{condition}


\begin{condition}
\label{ass7sml}$\mathbb{E}[\partial_{z^{\prime }}g(x_{i};\theta
_{0})\partial_{z}g^{\prime }(x_{i};\theta _{0})]$ exists and is of full rank.
\end{condition}

These assumptions ensure that the minimization problem defined by (\ref
{eqobjf}) and (\ref{eqcons}) is well-behaved, i.e., small changes in the
values of $x_{i}$ do not lead to jumps in the
solution $z_{i}$ to the optimization problem (aside from zero-probability events). It is likely that these
assumptions can be relaxed using empirical processes techniques. However,
here we favor simply imposing more smoothness (compared to the standard GMM
assumptions), because this leads to more transparent assumptions. They can
all be stated in terms of the basic function $g(x;\theta )$ that defines the
moment condition model, making them fairly primitive. We can then state our
first consistency result.

\begin{theorem}
\label{thconsissml}Under assumptions \ref{assiid}-\ref{ass7sml}, the OTGMM
estimator is consistent for $\theta _{0}$ and $\lambda =O_{p}(n^{-1/2})$.
\end{theorem}

As a by-product, this theorem also secures a convergence rate on the
Lagrange multiplier $\lambda $ which proves useful for establishing our
distributional results. The conditions needed to show asymptotic normality
also closely mimic those of standard GMM estimators (e.g. Theorem 3.2 in
\citet{Newey:HB}):

\begin{condition}
\label{ass8sml}$\theta _{0}\in \Theta ^{\circ }$, the interior of $\Theta $.
\end{condition}

\begin{condition}
\label{ass9sml}$\mathbb{E}[\sup_{\theta \in \eta }\Vert \partial_{\theta^{\prime
}}g(x_{i};\theta )\Vert ]<\infty $ where $\eta \subset \Theta $ is a
neighborhood of $\theta _{0}$.
\end{condition}

\begin{condition}
\label{ass10sml} $\left( \mathbb{E}[\partial_{\theta'}g(x_{i};\theta _{0})^{\prime
}]\left( \mathbb{E}[\partial_{z'}g(x_{i};\theta _{0})\partial
_{z'}g(x_{i},\theta _{0})^{\prime }]\right) ^{-1}\mathbb{E}[\partial
_{\theta'}g(x_{i};\theta _{0})]\right) $ is invertible.
\end{condition}

We can then provide an explicit expression of the estimator's asymptotic variance.

\begin{theorem}
\label{thnormsml}Under Assumptions \ref{assiid}-\ref{ass10sml}, the OTGMM
estimator is asymptotically normal with $\sqrt{n}(\hat{\theta}
_{OTGMM}-\theta _{0})\rightarrow ^{d}\mathcal{N}(0;V)$, where
\begin{eqnarray*}
V &=&\left( \mathbb{E}[G_i^{\prime }]\left( \mathbb{E}[H_i H_i^{\prime
}]\right) ^{-1}\mathbb{E}[G_i]\right) ^{-1}\times
(\mathbb{E}[G_i^{\prime }]\left( \mathbb{E}[H_iH_i^{\prime }]\right)
^{-1}\mathbb{E}[g_ig_i^{\prime }]\left( \mathbb{E}[H_i H_i ^{\prime }]\right)
^{-1}\mathbb{E}[G_i])\times \\
&&\left( \mathbb{E}[G_i^{\prime }]\left( \mathbb{E}[H_i H_i^{\prime
}]\right) ^{-1}\mathbb{E}[G_i]\right) ^{-1},
\end{eqnarray*}
where\ $H_i \equiv \partial_{z^{\prime }}g(x_{i};\theta _{0})$,
       $G_i \equiv \partial_{\theta^{\prime }}g(x_{i};\theta _{0})$ and
       $g_i \equiv g(x_{i};\theta _{0})$.
\end{theorem}

This simple normal limiting distribution with root $n$ convergent behavior is somewhat unexpected from an estimator that involves a high-dimensional optimization over $O(n)$ latent variables. As in GEL, this is made possible thanks to the existence of an equivalent low-dimensional dual optimization problem, which is, in turn, equivalent to a simple GMM estimator, albeit with a nonstandard weighting matrix. Thus, the variance has the expected \textquotedblleft sandwich\textquotedblright\ form, since the reciprocal weights $\mathbb{E}[H_i H_i^{\prime }]$
differs from the moment variance $\mathbb{E}[g_i g_i^{\prime }]$. For comparison, an optimally weighted GMM estimator would have an asymptotic variance of $(\mathbb{E}[G_i]'(\mathbb{E}[g_ig_i'])^{-1}\mathbb{E}[G_i])^{-1}$.

It is difficult to formulate a simple expression for the difference $V_{OTGMM}-V_{GMM}$ between the asymptotic variance of OTGMM and that of optimal GMM, but a simple example suffices to illustrate that this difference could be any non-negative-definite matrix.

Let $x_{i}$ take value in $\mathbb{R}^{d_{x}}$ ($d_{x}\geq 2$) with $\mathbb{E}\left[ x_{i\ell }\right] =0$ and $\mathbb{E}\left[ x_{i\ell }x_{i\ell ^{\prime }}\right] \equiv \omega _{\ell } $ when $\ell=\ell'$ and zero otherwise, for $\ell ,\ell ^{\prime }=1,\ldots ,d_{x}$. For $\theta \in \mathbb{R}$, let the moment functions be given by $g_{\ell }\left( x_{i},\theta \right) =x_{i\ell }-\theta $. We then have
\[
V_{OTGMM}=\frac{1}{d_{x}^{2}}\sum_{\ell =1}^{d_{x}}\omega _{\ell }\text{ and }V_{GMM}=\left( \sum_{\ell =1}^{d_{x}}\omega _{\ell }^{-1}\right) ^{-1}.
\]
Then, by the harmonic-arithmetic mean inequality, $V_{OTGMM}\geq V_{GMM}$ with equality when $\omega _{\ell }=c$ for all $\ell $. If $\omega _{1}\longrightarrow \infty $ leaving all the other $\omega _{\ell }$ finite, $V_{OTGMM}\longrightarrow \infty $ while $V_{GMM}$ remains finite. (General non-negative-definite differences can be obtained by stacking such moment conditions for different parameters $\theta_k$ ($k=1,\ldots,d_\theta$) and possibly linearly transforming the resulting parameter vector $\theta$.) Generally, we expect the difference to be large when some moments have a disproportionally large variance while not being disproportionally sensitive to changes in the underlying variables.

\subsection{Asymptotics under large errors}

In some applications, there may be considerable misspecification or its magnitude may be a priori unknown. It thus proves useful to relax the assumption of small errors in deriving the estimator's asymptotic properties. To handle this more general setup, we employ the following equivalence, demonstrated in Appendix \ref{appproof}.

\begin{theorem}
\label{thjidgmm}If $g\left( z,\theta \right) $ is differentiable in its
arguments, the OTGMM estimator is equivalent to a just-identified GMM
estimator expressed in terms of the modified moment function
\begin{equation}
\tilde{g}\left( x,\theta ,\lambda \right) =\left[
\begin{array}{c}
\partial_{\theta}g^{\prime }\left( q\left( x,\theta ,\lambda \right) ,\theta
\right) ~\lambda \\
g\left( q\left( x,\theta ,\lambda \right) ,\theta \right)
\end{array}
\right]  \label{qauggmm}
\end{equation}
that is a function of the observed data $x$ and the augmented parameter
vector $\tilde{\theta}\equiv \left( \theta ^{\prime },\lambda ^{\prime
}\right) ^{\prime }$ and where
\begin{equation}
q\left( x,\theta ,\lambda \right) \equiv
\arg \min_{z:z-\partial_{z}g^{\prime }\left( z,\theta \right) \lambda =x}\left\Vert z-x\right\Vert
^{2}.  \label{eqinvfoclam}
\end{equation}
\end{theorem}

Note that $q\left( x,\theta ,\lambda \right) $ is essentially the inverse of
the mapping $z-\partial_{z}g^{\prime }\left( z,\theta \right) \lambda =x$
(from Equation (\ref{eqfoczj})), augmented with a rule to select the
appropriate branch in case the inverse is multivalued.

The equivalence result of Theorem \ref{thjidgmm} implies that many of the
asymptotic technical tools used in GMM-type estimators can be adapted to our
setup, with the distinction that the function $q\left( x,\theta ,\lambda
\right) $ is defined only implicitly. Hence, many of our efforts below seek
to recast necessary conditions on $q\left( x,\theta ,\lambda \right) $ in
terms of more primitive conditions on the moment function $g\left( z,\theta
\right) $ whenever possible or in terms of regularity conditions drawn from optimal transport theory.

We start with a standard GMM-like identification condition:

\begin{condition}
\label{assidgmm}For some compact sets $\Theta $ and $\Lambda $, there exists
a unique $\left( \theta _{0},\lambda _{0}\right) \in \Theta \times \Lambda $
solving $\mathbb{E}\left[ \tilde{g}\left( x,\theta ,\lambda \right) \right]
=0$ for $\tilde{g}\left( x,\theta ,\lambda \right) $ defined in Theorem \ref
{thjidgmm}.
\end{condition}

This condition is implied by a natural uniqueness and regularity condition on the solution to the primal optimal transportation problem (Equations (\ref{eqobjgen})):

\begin{condition}\label{asscaff}
Let $\mu_{zx;\theta}$ denote the solution to Problem (\ref{eqobjgen}) for a given $\theta\in\Theta$.
(i) $\mathbb{E}_{\mu_{zx;\theta}}[\Vert z-x \Vert^2]$ is uniquely minimized at $\theta=\theta_0$
(ii) The corresponding marginals $\mu_{z;\theta}$ and $\mu _{x}$ are absolutely continuous with respect to the Lebesgue measure with a density that is
finite, nonvanishing and H\"{o}lder continuous on their convex support.
\end{condition}

Indeed, by Theorem [C3], part b) and d), in \citet{caffarelli:regconv}, Assumption \ref{asscaff}(ii) implies that, at each $\theta$, there exists a unique invertible transport map $z=q\left( x\right) $ from $\mu _{x}$ to $\mu _{z;\theta}$ and both $q$ and its inverse are equal to the gradient of a twice differentiable strictly convex function. The fact that $q^{-1}$ is the gradient of a twice differentiable strictly convex function ensures that the first-order condition in Assumption \ref{assidgmm} has a unique solution (making a rule to handle multivalued inverses unnecessary).


Next, we consider standard continuity and dominance conditions that are used
to establish uniform convergence of the GMM objective function.
These assumptions constitute a superset of those needed for standard GMM because the
modified moment conditions include the additional parameter $\lambda$ and higher-order derivatives of the original moment conditions.
In a high-level form, these conditions read:

\begin{condition}
\label{asssimplegtilda}(i) $\tilde{g}\left( x,\theta ,\lambda \right) $ is
continuous in $\theta $ and $\lambda $ for $\left( \theta ,\lambda \right)
\in \Theta \times \Lambda $ with probability one and (ii) $\mathbb{E}\left[
\sup_{\left( \theta ,\lambda \right) \in \Theta \times \Lambda }\left\Vert
\tilde{g}\left( x,\theta ,\lambda \right) \right\Vert \right] <\infty $.
\end{condition}

Alternatively, Assumption \ref{asssimplegtilda} can be replaced by more
primitive conditions on $g\left( z,\theta \right) $ instead, as given below
in Assumptions \ref{assgcontdiff}, \ref{asseigval} and \ref{assgdom}.

\begin{condition}
\label{assgcontdiff}(i) $g\left( z,\theta \right) $ and
$\partial_{z^{\prime }}g\left( z,\theta \right) $ are differentiable in $\theta $ and
(ii) $\partial_{\theta^{\prime }}g\left( z,\theta \right) $ is continuous in
both arguments.
\end{condition}

This assumption parallels continuity assumptions typically made for GMM, but
higher order derivatives of $g\left( z,\theta \right) $ are needed, because
they enter the moment condition either directly or indirectly via the
function $q\left( x,\theta ,\lambda \right) $.
The next condition ensures that the function $q\left( x,\theta ,\lambda
\right) $ is well behaved.

\begin{condition}
\label{asseigval}$\bar{\nu}\bar{\lambda}<1$ where $\bar{\lambda}
=\max_{\lambda \in \Lambda }\left\Vert \lambda \right\Vert $ and $\bar{\nu}
=\sup_{\theta \in \Theta }\sup_{z\in \mathcal{X}}\max_{k\in \left\{ 1,\ldots
,d_{g}\right\} }\allowbreak \max \operatorname{eigval}\left( \partial_{zz^{\prime
}}g_{k}\left( z,\theta \right) \right) $, in which $\partial_{zz^{\prime
}}g_{k}\left( z,\theta \right) $ exists for $k=1,\ldots ,d_{g}$ and where $
\operatorname{eigval}\left( M\right) $ for some matrix $M$ denotes the set of its
eigenvalues.
\end{condition}

Once again, this condition can be alternatively phrased in terms of optimal transport concepts.
The first order condition which implicitly defines $z=q\left( x,\theta
,\lambda \right)$ can be written in terms of the derivative of a \emph{potential function}
$\psi(z,\theta,\lambda)=z^{\prime }z/2-g^{\prime }\left(z,\theta \right) \lambda$: $\nabla_z \psi(z,\theta,\lambda)=x$.
With the help of Theorem [C3], part b) and d), in \citet{caffarelli:regconv}, Assumption \ref{asscaff}(ii) implies that, at each $\theta$, the above
potential $\psi(z,\theta,\lambda)$ has a positive-definite Hessian (with respect to $z$), which implies Assumption \ref{asseigval}.

In order to state our remaining regularity conditions, it is useful to
introduce a notion of (nonuniform) Lipschitz continuity, combined with dominance conditions.

\begin{definition}
\label{defholderLD}Let $\mathcal{L}$ be the set of functions $h\left(
z,\theta \right) $ such that (i) $\mathbb{E}\left[ \sup_{\theta \in \Theta
}\left\Vert h\left( x,\theta \right) \right\Vert \right]\linebreak <\infty $ and (ii)
there exists a function $\bar{h}\left( x,\theta \right) $ satisfying
\begin{eqnarray}
\mathbb{E}\left[ \sup_{\theta \in \Theta }\bar{h}\left( x,\theta \right)
\left\Vert \partial_{z^{\prime }}g\left( x,\theta \right) \right\Vert
\right] &<&\infty .  \label{eqdomhg} \\
\left\Vert h\left( z,\theta \right) -h\left( x,\theta \right) \right\Vert
&\leq &\bar{h}\left( x,\theta \right) \left\Vert z-x\right\Vert \label{eqnonulip}
\end{eqnarray}
for all $x,z\in \mathcal{X}$ and $\theta\in\Theta$, and where $g\left( x,\theta \right) $ is as in the moment conditions.
\end{definition}

This Lipschitz continuity-type assumption has no parallel in conventional GMM.
It is made here because it ensures that
the behavior of the observed $x$ and the underlying unobserved $z$ will not
differ to such an extent that moments of unobserved variables would be
infinite, while the corresponding observed moments are finite. Clearly,
without such an assumption, observable moments would be essentially
uninformative. The idea underlying Definition \ref{defholderLD} is that we
want to define a property that is akin to Lipschitz continuity but that
allows for some heterogeneity (through the function $\bar{h}\left( x,\theta
\right) $ in Equation (\ref{eqnonulip})). This heterogeneity proves
particularly useful in the case where $\mathcal{X}$ is not compact (for
compact $\mathcal{X}$, one can take $\bar{h}\left( x,\theta \right) $ to be
constant in $x$ with little loss of generality). For a given function $
h\left( x,\theta \right) $ that is finite for finite $x$, membership in $
\mathcal{L}$ is easy to check by inspecting the tail behavior (in $x$) of
the given function $h\left( x,\theta \right) $. Polynomial tails will
suggest a polynomial form for $\bar{h}\left( x,\theta \right) $, for
instance. Equation (\ref{eqdomhg}) strengthens the dominance condition \ref
{defholderLD}(i) to ensure that functions $h\left( x,\theta \right) $ in $
\mathcal{L}$ also satisfy a dominance condition when interacted with other
quantities entering the optimization problem,
i.e. $\partial_{z^{\prime}}g\left( x,\theta \right) $).

With this definition in hand, we can succinctly state a sufficient condition
for $\tilde{g}\left( x,\theta ,\lambda \right) $ to satisfy a dominance
condition:

\begin{condition}
$\,$\label{assgdom}$g\left( \cdot ,\cdot \right) $ and each element of $
\partial_{\theta}g^{\prime }\left( \cdot ,\cdot \right) $ belong to $\mathcal{L}$
.
\end{condition}

We are now ready to state our general consistency result.

\begin{theorem}
\label{thconsislarge}Under Assumptions \ref{assiid}, \ref{assidgmm} and
either Assumption \ref{asssimplegtilda} or Assumptions \ref{assgcontdiff},
\ref{asseigval}, \ref{assgdom}, the OTGMM estimator is consistent
($(\hat{\theta},\hat{\lambda})\overset{p}{\longrightarrow }(\theta _{0},\lambda _{0}) $).
\end{theorem}

We now turn to asymptotic normality. We first need a conventional
\textquotedblleft interior solution\textquotedblright\ assumption.

\begin{condition}
\label{assinterlarge}$\left( \theta _{0},\lambda _{0}\right) $ from
Assumption \ref{assidgmm} lies in the interior of $\Theta \times \Lambda $.
\end{condition}

Next, as in any GMM estimator, we need finite variance of the moment
functions and their differentiability:

\begin{condition}
\label{asshesjaclarge}(i) $\mathbb{V}\left[ \tilde{g}\left( x,\theta
_{0},\lambda _{0}\right) \right] \equiv \Omega $ exists and (ii) $\mathbb{E}
\left[ \partial \tilde{g}\left( x,\theta ,\lambda \right) /\partial \left(
\theta ^{\prime },\lambda ^{\prime }\right) \right] \equiv \tilde{G}$ exists
and is nonsingular.
\end{condition}

Assumption \ref{asshesjaclarge}(ii) can be expressed in a more primitive
fashion using the explicit form for $\tilde{G}$ provided in Theorem \ref
{thanlarge} below.

Next, we first state a high-level dominance condition that ensures uniform
convergence of the Jacobian term $\partial \tilde{g}\left( x,\theta ,\lambda
\right) /\partial \left( \theta ^{\prime },\lambda ^{\prime }\right) $.

\begin{condition}
\label{assdgdom}(i) $\tilde{g}\left( x,\theta ,\lambda \right) $ is
continuously differentiable in $\left( \theta ,\lambda \right) $;\linebreak (ii) $
\mathbb{E[}\sup_{\left( \theta ,\lambda \right) \in \Theta \times \Lambda
}\allowbreak \left\Vert \partial \tilde{g}\left( x,\theta ,\lambda \right)
/\partial \left( \theta ^{\prime },\lambda ^{\prime }\right) \right\Vert
]<\infty $.
\end{condition}

This assumption is implied by the following, more primitive, condition:

\begin{condition}
\label{assdgdomprim}(i) $g\left( z,\theta \right) $ and
$\partial_{\theta}g\left( z,\theta \right) $ are continuously differentiable in $\theta $,
(ii) all elements of $\partial_{\theta}g_{k}\left( z,\theta \right) $ and $
\partial_{\theta\theta^{\prime }}g_{k}\left( z,\theta \right) $ for $k=1,\ldots
,d_{g} $ belong to $\mathcal{L}$ and (iii) Assumptions \ref{assgcontdiff}(i)
and \ref{asseigval} hold.
\end{condition}

We can now state our general asymptotic normality and root-$n$ consistency result, shown in Appendix \ref{appproof}.

\begin{theorem}
\label{thanlarge}Let the assumptions of Theorem \ref{thconsislarge} hold as
well as Assumptions \ref{assinterlarge}, \ref{asshesjaclarge} and either
Assumption \ref{assdgdom} or \ref{assdgdomprim}. Then,
\begin{equation*}
\sqrt{n}\left( \left[
\begin{array}{c}
\hat{\theta} \\
\hat{\lambda}
\end{array}
\right] -\left[
\begin{array}{c}
\theta _{0} \\
\lambda _{0}
\end{array}
\right] \right) \overset{d}{\longrightarrow }\mathcal{N}\left(
0,W^{-1}\right)
\end{equation*}
where $W=\tilde{G}^{\prime }\Omega ^{-1}\tilde{G}$, $\Omega =\mathbb{E}\left[
\tilde{g}\tilde{g}^{\prime }\right],$
\begin{equation*}
\tilde{g} \equiv \tilde{g}(z,(\theta_0,\lambda_0)) = \left[
\begin{array}{c}
\partial_{\theta}g^{\prime }\left( z,\theta_0 \right) \lambda_0 \\
g\left( z,\theta_0 \right)
\end{array}
\right] \text{ and }\tilde{G} \equiv \mathbb{E}[\partial_{\theta} \tilde{g}'] = \left[
\begin{array}{cc}
\tilde{G}_{\theta \theta } & \tilde{G}_{\theta \lambda } \\
\tilde{G}_{\lambda \theta } & \tilde{G}_{\lambda \lambda }
\end{array}
\right]
\end{equation*}
in which
\begin{eqnarray*}
\tilde{G}_{\theta \theta } &\equiv &\mathbb{E}\left[ \partial_{\theta \theta^{\prime
}}\left( \lambda _{0}^{\prime }g\left( z,\theta _{0}\right) \right)
+\partial_{\theta z^{\prime }}\left( \lambda _{0}^{\prime }g\left( z,\theta
_{0}\right) \right) \partial_{\theta^{\prime }}q\left( x,\theta _{0},\lambda
_{0}\right) \right] \\
\tilde{G}_{\lambda \theta } &\equiv &\mathbb{E}\left[ \partial_{\theta^{\prime
}}g\left( z,\theta _{0}\right) +\partial_{z^{\prime }}g\left(
z,\theta _{0}\right) \partial_{\theta^{\prime }}q\left( x,\theta
_{0},\lambda _{0}\right) \right] \\
\tilde{G}_{\theta \lambda } &\equiv &\mathbb{E}\left[ \partial_{\theta}\left(
g^{\prime }\left( z,\theta _{0}\right) \right) +\partial_{\theta z^{\prime
}}\left( \lambda _{0}^{\prime }g\left( z,\theta _{0}\right) \right)
\partial_{\lambda^{\prime }}q\left( x,\theta _{0},\lambda _{0}\right) \right]
\\
\tilde{G}_{\lambda \lambda } &\equiv &\mathbb{E}\left[ \partial_{z^{\prime
}}g\left( z,\theta _{0}\right) \partial_{\lambda^{\prime }}q\left(
x,\theta _{0},\lambda _{0}\right) \right]
\end{eqnarray*}
where $z$ solves $x=z-\partial_{z}g^{\prime }\left( z,\theta \right) \lambda $ for given $x,\theta
,\lambda $ and where
\begin{eqnarray}
\partial_{\theta^{\prime }}q\left( x,\theta ,\lambda \right) &=&\left[ \left(
I-\partial_{zz^{\prime }}\left( \lambda ^{\prime }g\left( z,\theta \right)
\right) \right) ^{-1}\partial_{z\theta^{\prime }}\left( \lambda ^{\prime
}g\left( z,\theta \right) \right) \right] _{z=q\left( x,\theta ,\lambda
\right) } \\
\partial_{\lambda^{\prime }}q\left( x,\theta ,\lambda \right) &=&\left[ \left(
I-\partial_{zz^{\prime }}\left( \lambda ^{\prime }g\left( z,\theta \right)
\right) \right) ^{-1}\partial_{z}g^{\prime }\left( z,\theta \right) \right]
_{z=q\left( x,\theta ,\lambda \right) }.
\end{eqnarray}
In particular, for $\theta $, the partitioned inverse formula gives
\begin{equation*}
\sqrt{n}\left( \hat{\theta}-\theta _{0}\right) \overset{d}{\longrightarrow }
\mathcal{N}\left( 0,\left( W_{\theta \theta }-W_{\theta \lambda }W_{\lambda
\lambda }^{-1}W_{\lambda \theta }\right) ^{-1}\right)
\end{equation*}
where $W$ is similarly partitioned as:
\begin{equation*}
W=\left[
\begin{array}{cc}
W_{\theta \theta } & W_{\theta \lambda } \\
W_{\lambda \theta } & W_{\lambda \lambda }
\end{array}
\right]
\end{equation*}
\end{theorem}

The asymptotic variance stated in Theorem \ref{thanlarge} takes the familiar
form expected from a just-identified GMM estimator: $(\tilde{G}^{\prime
}\Omega ^{-1}\tilde{G})^{-1}$. The relatively lengthy expressions merely come
from explicitly computing the first derivative matrix $\tilde{G}$ in terms
of its constituents. This is accomplished by differentiating $\tilde{g}$
with respect to all parameters using the chain rule and calculating the
derivative of $q\left( x,\theta ,\lambda \right) $ using the implicit
function theorem.

We thus have now completely characterized the first-order asymptotic properties of our estimator in the most general settings of large (i.e., non-local) misspecification. This result thus allows researcher to directly replace their GMM estimator which may happen to fail overidentification tests by another, logically consistent and easy-to-interpret, estimator where the overidentification failure is naturally accounted for by errors in the variables. In addition, researchers can further document the presence of errors via Theorem \ref{thanlarge}, as it enables, as a by-product, a formal test of the absence of error. Under this null hypothesis, which can be stated as $\lambda=0$, we have
\begin{equation}
    n\hat{\lambda}' \left( W_{\lambda\lambda} -
W_{\lambda\theta}W_{\theta\theta}^{-1}W_{\theta\lambda} \right)^{-1} \hat{\lambda}
\overset{d}{\longrightarrow }
\chi^2_{d_g}, \label{eqchi}
\end{equation}
where the above expression can be straightforwardly derived from the partitioned inverse formula applied to the $\lambda$ sub-block of the asymptotic variance.

\section{Discussion and extensions}

On a conceptual level, our use of a so-called Wasserstein metric to measure distance between distributions does provide some desirable theoretical properties.
For instance, the Wasserstein metric metrizes convergence in distribution (see Theorem 6.9 in \citet{villani:new}) under some simple bounded moment assumptions. In contrast, the discrepancies which generate GEL estimators do not admit such an interpretation. In fact, most discrepancies are not metrics, as they lack symmetry. The Kullback-Leibler discrepancy, which is perhaps the best known among them, does not allow comparison between distributions that are not absolutely continuous with respect to one another, whereas the Wassertein metric does. (Of course, leveraging this advantage requires considering the most general transport problem of Equation (\ref{eqobjgen}).) Finally, it is arguably logical to penalize probability transfer over larger distances more than the same probability transfer over smaller distances, as the Wasserstein metric does, while none of GEL-related discrepancies do.


An interesting extension of our approach would be a hybrid method in
which (i) the possibility of general forms of errors is
accounted for with the current method by constructing the equivalent GMM
formulation of the model via Theorem \ref{thjidgmm} and (ii) additional
restrictions on the form of the errors are imposed via additional
moment conditions involving some elements of $z$ and $x$. This could prove a
useful middle ground when a priori information regarding the errors
is available for some, but not all, variables.

As shown in Supplement \ref{secconstr}, it is straightforward to extend our approach to allow error in some, but not all variables.
When our method is used while allowing for errors in only a few variables,
it may not be possible to simultaneously satisfy all moment conditions for any $\theta$. In such cases, it could make sense to consider a hybrid method
where both errors in the variables, handled via our approach, and re-weighting of the
sample, handled via GEL, are simultaneously allowed.

Finally, we should mention other approaches aimed at handling violations of overidentifying restrictions, which
include the use of set identification combined with relaxed moment constraints (\citet{poirier:salvaging}),
placing a Bayesian prior on the magnitude of the deviations from correct specification of the moments (\citet{conley:plausexo}),
distributionally robust approaches that allow for deviations from the data generating process up to a given bound (\citet{connault:sens}, \citet{blanchet:distrob}), sensitivity analysis (\citet{shapiro:sens}, \citet{bonhomme:sens}) and misspecified moment inequality models (e.g. \citet{kwon:ineqmiss}).

\section{Application}

We revisit the study of \citet{duranton2014roads}, who documented evidence that the quantity of goods a city exports is strongly related to the extent of interstate highways present in that city.
Due to simultaneity concerns, the authors adopt an instrumental variable approach to recover the causal effect of building highways. Although the authors mitigate instrument validity concerns with controls, instrument exogeneity or exclusion could remain a concern (as noted by \citet{poirier:salvaging}) and this problem would manifest itself by specification test failure.

\begin{table}[!ht]
\caption{Main results}
\label{main_results}
\begin{equation*}
\begin{array}{|l|l|l|}
\hline
& \mbox{GMM} & \mbox{OTGMM} \\ \hline
\mbox{log highway km} & 0.39 & 0.40 \\
\mbox{se} & (0.12) & (0.11) \\ \hline
\mbox{log employment} & 0.47 & 1.24  \\
\mbox{se} & (0.32) & (0.31) \\ \hline
\mbox{market access (export)} & -0.63 & -0.66  \\
\mbox{se} & (0.11) & (0.10) \\ \hline
\mbox{log 1920 population} & -0.29 & -0.57  \\
\mbox{se} & (0.23) & (0.23) \\ \hline
\mbox{log 1950 population} & 0.65 & 1.14  \\
\mbox{se} & (0.37) & (0.37) \\ \hline
\mbox{log 2000 population} & -0.20 & -1.25  \\
\mbox{se} & (0.44) & (0.35) \\ \hline
\mbox{log \% manuf empl} & 0.64 & 0.58  \\
\mbox{se} & (0.12) & (0.12) \\ \hline
\mbox{Overidentification P-value} & 0.30 & \\ \hline
\end{array}
\end{equation*}
{Main results from Table 5 in \citet{duranton2014roads}. Original GMM estimates and OTGMM estimates. Heteroskedasticity-robust standard errors (GMM) and small-error asymptotic standard error (OTGMM) in parentheses.}
\end{table}

We apply our method to further assess the robustness of the results to potential model misspecification. We seek to recover point estimates that remain interpretable under potential misspecification and account for misspecification by viewing the model's variables as potentially measured with error.
Not only do we consider errors in the regressors, but we also think of potentially invalid instruments as simply mismeasured versions of an underlying valid instrument that is unfortunately not available. Alternatively, one can think of the underlying valid instrument as the counterfactual value of the instrument in a world where the mechanism causing this instrument to be invalid would be absent. This broader interpretation of what could constitute an ``error'' under our framework considerably expands the scope of models that are conceptually consistent with our approach.

\citet{duranton2014roads} consider three instruments: (log) kilometers of railroads in 1898, quantity of historical exploration routes, and planned (log) highway kilometers according to a 1947 construction map.
The validity of these instruments could be criticized, for instance, in a situation where some cities are in proximity to key natural resources, which could cause higher exports and, at the same time, more transportation routes (or plans to build them). If this mechanism is active both in the present and in the past, the causal effect of highways on exports would be overestimated.
In our framework, the true but unavailable instrument could represent a measure of past transportation routes in a counterfactual world where natural resources would be evenly distributed among cities. The actual available instruments represent an approximation to this ideal instrument, a situation which we represent as a potentially non-classical errors-in-variables model.

The model's moment conditions are written in terms of
$g(z_i,\theta)=w_i(y_i-r_i'\theta)$, where the vector of observables
$z_i=(y_i,r_i',w_i')'$ contains the dependent variable $y_i$, the vector of regressors $r_i$ and the vector of instruments $w_i$.
In \citet{duranton2014roads}, the dependent variable $y_i$ measures ``propensity to export'' and is constructed from an auxiliary panel data model, which regresses volume of exports between given cities on distance and trading partner characteristics, modeled as fixed effects. It is the value of these fixed effects that is used as $y_i$. Following \citet{poirier:salvaging}, we take this construction as given and focus on export volume measured by weight.
Explicitly accounting for errors in $y_i$ is superfluous since, in a regression setting, errors in the dependent variable are already allowed for.

In the analysis below, $r_i$ always includes the regressor of main interest: the (logarithm of) the number of kilometers of highway. It also contains a number of controls, which may differ in the different models considered. These controls include: log employment, log market access, log population in 1920, 1950, and 2000, and log manufacturing share in 2003, average distance to the nearest body of water, land gradient, dummy variables for census regions, log share of the fraction of adult population with a college degree or more, log average income per capita, log share of employment in wholesale trade, and log average daily traffic on the interstate highways in 2005.
We allow for errors in all of these variables, except for the constant term and dummies.





We consider two of the specifications employed by \citet{duranton2014roads}: The specification with many covariates that most obviously passes the overidentifying test and the specification with few covariates that most clearly fails this test.
(The other specifications fall in between these extreme cases and are thus not reported here, for conciseness.)
The results from GMM, replicated from the original study, and from OTGMM (allowing for large errors) are reported in Tables \ref{main_results} and \ref{failed_overID}.

\begin{table}[!ht]
\caption{Specification that fails the test for over-identifying restriction}
\label{failed_overID}
\begin{equation*}
\begin{array}{|l|l|l|}
\hline
& \mbox{GMM} & \mbox{OTGMM} \\ \hline
\mbox{log highway km} & 0.57 & 0.65 \\
\mbox{se} & (0.16) & (0.19) \\ \hline
\mbox{log employment} & 0.52 & 0.49  \\
\mbox{se} & (0.11) & (0.09) \\ \hline
\mbox{market access (export)} & -0.45 & -0.46  \\
\mbox{se} & (0.13) & (0.14) \\ \hline
\mbox{Overidentification P-value} & 0.043 & \\ \hline
\end{array}
\end{equation*}
{Specification from column 2, Table 5, in \citet{duranton2014roads}. Replicated estimates from the original paper and OTGMM estimates. Heteroskedasticity-robust standard errors (GMM) and small-error asymptotic standard error (OTGMM) in parentheses.}
\end{table}

The OTGMM estimates are similar to those of \citet{duranton2014roads}. Although some coefficients of the controls are larger in magnitude, the main coefficient --- the elasticity of export weight relative to kilometers of highway --- is almost unchanged and remains statistically significant.
As in the original study, the 95\% confidence intervals of the main elasticity of interest obtained with different sets of controls overlap significantly.

It is also instructive to look at the correction of the underlying variables implied by OTGMM. Table \ref{MEs}, columns 1 and 2, documents this by looking at the standard deviation of different elements of the correction $z_{i} - x_{i}$.
The corresponding quantities for the models of Tables \ref{main_results} and \ref{failed_overID} are reported in columns 1 and 2, respectively.
To get an idea of the scale, the last column reports the standard deviations of the observed variables.
As expected, the errors in column 1 are exceedingly small, reflecting the fact that the model passes the overidentification test.
In contrast, column 2 is particularly interesting for our purposes because it corresponds to a specification that fails the test of over-identifying restrictions. While this may, at first, cast doubt regarding the GMM estimates, OTGMM shows that errors of the order of only 10\% of the observed regressors' magnitude are sufficient to eliminate the mispecification. Such small error magnitudes are highly plausible empirically, thus supporting the plausibility of the OTGMM estimate. As the GMM and OTGMM estimates of the main elasticity of interest are very close in Table \ref{failed_overID}, this also corroborates the authors' GMM estimates. Overall, our approach strongly supports the conclusions of the original study and thus provides an effective robustness test.

The fact that the model with more controls leads to smaller error magnitudes is very consistent with our interpretation: to the extent that including more controls reduces the magnitude of the potentially endogenous residuals, one would expect that our estimator has to perform smaller alterations to the variables to arrive at valid instruments and/or non-endogeneous regressors. More quantitatively, suppose that the instrument $w$ and the error term $\epsilon$ can be decomposed into separate components, say $w=w_1+w_2$ and $\epsilon = \epsilon_1 + \epsilon_2$, where $w_1$ and $\epsilon_1$ are correlated, but the latter is well explained by additional controls. In a model with fewer controls, our method has to identify $w_1$ and/or $\epsilon_1$ as error components to obtain a correctly specified model. However, if including controls already accounts for the $\epsilon_1$ term, then our estimator no longer needs to correct the corresponding components and thus achieves orthogonality with smaller errors.

In contrast, using GEL to address misspecification would effectively assume that the misspecification originates from some form of selection bias. The fact that adding more control eliminates the misspecification indicates that the controls incorporate the variables that explain selection bias. Yet, these controls are not included in the model with few controls, precisely the model where the largest amount of sample reweighting would take place when using GEL. This paradox makes the use of GEL as a remedy difficult to rationalize in a setting where misspecification arises primarily from the unavailability of adequate control variables.



\begin{table}[!ht]
\caption{{Standard deviations of errors and corresponding regressors}}
\label{MEs}
\begin{equation*}
\begin{array}{|l|l|l|l|}
\hline
\mbox{Variables} & \sqrt{\hat{\mathbb{V}}[z-x]} & \hfill \sqrt{\hat{\mathbb{V}}[z-x]} \hfill & \sqrt{\hat{\mathbb{V}}[x]}\\
& \mbox{(main)} &  \mbox{(failed over-ID)} & \\
\hline \hline
\mbox{log highway km} & 0.0034 & 0.0634 & 0.5884 \\ \hline
\mbox{log railroads km 1898} & 0.0048 & 0.0267 & 0.6031 \\ \hline
\mbox{exploration routes} & 0.0005 & 0.0418 & 0.8339 \\ \hline
\mbox{plan 1947} & 0.0075 & 0.0598 & 0.7049 \\ \hline
\mbox{log employment} & 0.0144 &  0.0473 & 0.8573 \\ \hline
\mbox{market access (export)} & 0.0056 & 0.0444 & 0.4864 \\ \hline
\mbox{log 1920 population} & 0.0052 & & 1.0417 \\ \hline
\mbox{log 1950 population} & 0.0097 & & 0.9253 \\ \hline
\mbox{log 2000 population} & 0.0163 & & 0.8083 \\ \hline
\mbox{log \% manuf empl} & 0.0050 & & 0.3707 \\ \hline
\end{array}
\end{equation*}
\end{table}


In Supplement \ref{appallcon}, we perform a number of robustness tests along the lines suggested by \citet{poirier:salvaging}.
We further replicate specifications with many more controls and find that the GMM and OTGMM are still in close agreement.
In Supplement \ref{apprelaxex}, we also consider alternative specifications that allow for the possibility that some instrumental variables should be included as regressors. The performance of OTGMM in other simple models is also investigated via simulations in Supplement \ref{secsimul}.

\section{Conclusion}

We have proposed a novel optimal transport-based version of the Generalized
Method of Moment (GMM) that fulfills, by construction, the overidentifying
moment restrictions by introducing the smallest possible amount of error in the variables or, equivalently, by allowing for the smallest possible Wasserstein metric distortions in the data distribution. This approach
conceptually merges the Generalized Empirical Likelihood (GEL) and optimal
transport methodologies. It provides a theoretically motivated
interpretation to GMM results when standard overidentification tests reject
the null. GEL approaches handle model misspecification by re-weighting the
data, which would be appropriate when misspecification arise from improper
sampling of the population. In contrast, our optimal transport approach is
appropriate when the sources of misspecification are errors or, more generally, Wasserstein-type distortions in the data distribution, which
is arguably a common situation in applications.