EconBase
← Back to paper

Optimally-Transported Generalized Method of Moments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

60,992 characters · 10 sections · 63 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Optimally-Transported Generalized Method of Moments

abstractWe propose a novel optimal transport-based version of the Generalized Method of Moment (GMM). Instead of handling overidentification by reweighting the data to satisfy the moment conditions (as in Generalized Empirical Likelihood methods), this method proceeds by allowing for errors in the variables of the least mean-square magnitude necessary to simultaneously satisfy all moment conditions. This approach, based on the notions of optimal transport and Wasserstein metric, aims to address the problem of assigning a logical interpretation to GMM results even when overidentification tests reject the null, a situation that cannot always be avoided in applications. We illustrate the method by revisiting Duranton, Morrow and Turner's (2014) study of the relationship between a city's exports and the extent of its transportation infrastructure. Our results corroborate theirs under weaker assumptions and provide insight into the error structure of the variables. {\bf Keywords:} Wasserstein metric, GMM, overidentification, misspecification.

Introduction

The Generalized Method of Moment (GMM) (hansen:gmm) has long been the workhorse of statistical modeling in economics and the social sciences. Its key distinguishing feature, relative to the basic method of moments, is the presence of overidentifying restrictions that enable the model's validity to be tested (Newey:HB). With this ability to test comes the obvious practical question of what one should do if an overidentified GMM model fails overidentification tests, a situation that is not uncommon (as noted in hall:missgmm, hansen:miss, poirier:salvaging, conley:plausexo, kwon:ineqmiss), even for perfectly reasonable, economically grounded, models.

A popular approach has been to find the \textquotedblleft pseudo-true\textquotedblright\ value of the model parameter ( sawa:pseudo, white:miss) that minimizes the “distance” or discrepancy between the data and the moment constraints implied by the model. This approach has gained further support since the introduction of Generalized Empirical Likelihoods (GEL) and Minimum Discrepancy estimators (owen:el1, qinlawless, newey:gel), all of which provide more readily interpretable pseudo-true values ( imbens:res, kitamura:itgmm, schennach:etel).

GEL implicitly attributes the mismatch in the moment conditions solely to a biased sampling of the population. While this is a possible explanation, it is not the only reason a valid model would fail overidentification tests, when taken to the data. Another natural possibility is the presence of errors in the variables (aguiar:revpref, doraszelski:megmm, schennach:HB). In this work, we develop an alternative to GMM that ensures, by construction, that overidentifying restrictions are satisfied by allowing for possible errors in the variables instead of sampling bias. We employ the generic term {\em error} to include, not only measurement error, but anything that could cause the recorded data to differ from the value they should have if the model were fully correct, i.e., this could include some model errors. More generally, we allow distortions in the data generating process whose magnitude is quantified by a Wasserstein-type metric (villani:new), in the spirit of distributionally robust methods (connault:sens, blanchet:distrob). In analogy with GEL, which does not require the form of the sampling bias to be explicitly specified, the error process does not need to be explicitly specified in our approach, but is instead inferred from the requirement of satisfying the overidentifying constraints imposed by the GMM model. Of course, the accuracy of the resulting estimated parameters will typically improve with the degree of overidentification. \footnote{In the case where the statistical properties of the errors are in fact known a priori, other methods may be more appropriate (e.g. schennach:nlme, schennach:elvis), schennach:annrev, schennach:HB and references therein.}

A fruitful way to accomplish this is to employ concepts from the general area of optimal transport problems (e.g., galichon:optran, villani:new, carlier:vectq, galichon:vectquan, cherno:breniercons, schennach:nlfact). The idea is to find the parameter value that minimizes cost of \textquotedblleft transporting\textquotedblright\ the observed distribution of the data $\mu_x$ onto another distribution $\mu_z$ that satisfies the moment conditions exactly. Formally, the true iid data $z_{i}$ is assumed to satisfy $\mathbb{E}\left[ g\left( z_{i},\theta \right) \right] =0$, where $\mathbb{E}$ is the expectation operator, for a parameter value $\theta$ in some set $\Theta$ and some given $d_{g}$-dimensional vector $ g\left( z_{i},\theta \right) $ of moment functions. However, we instead observe an error-contaminated counterpart $x_{i}$ of the true vector $z_{i}$ (both taking value in $\mathcal{X} \subseteq \mathbb{R}^{d_{x}}$). We seek to exploit the model's over-identification to gain information regarding the error in $x_{i}$. The Euclidean norm $\Vert (z-x) \Vert$ is chosen here for computational convenience, although one could imagine a whole class of related estimators obtained with different choices of metric. Our focus on Euclidean norms parallels the choice made in common estimators (e.g. least squares regressions, classical minimum distance and even GMM). Considering a {\it weighted} Euclidean norm can be useful to indicate the relative expected error magnitudes along different dimensions of $x$.

Given a probability measure $\mu_x$ for the random variable $x$, this setup suggests solving the following population optimization problem, for a given $\theta$:

equation[equation omitted — 116 chars of source]

subject to $\mu _{zx}$, supported on $\mathcal{X}\times \mathcal{X}$, having marginal $\mu _{x}$ and $\mathbb{E}_{\mu _{zx}} \left[ g\left( z,\theta \right) \right] =0$, where $\mathbb{E}_{\mu}$ denotes an expectation under the measure $\mu$. (This problem is guaranteed to have a solution if there exists at least one measure $\mu_z^*$ such that $\mathbb{E}_{\mu_z^*} \left[ g\left( z,\theta \right) \right] =0$.) This setup covers the most general case, including both discrete and continuous variables, and can be handled using linear programming techniques (e.g., santambrogio:optran, Section 6.4.1), after observing that the moment constraint is easy to incorporate since it is linear in the probability measure. However, we shall focus on the purely continuous case in the remainder of this paper, because it enables us to express the main ideas more transparently. The fully continuous case indeed admits a convenient treatment, under the following regularity condition:

conditionThe marginals $\mu_z$ (arising from the solution $\mu_{zx}$ at each $\theta\in\Theta$) and $\mu_x$ have finite variance and $\mu_x$ is absolutely continuous with respect to the Lebesgue measure.

Under this condition, by Theorem 1.22 in santambrogio:optran, there exists a unique $\mu_{zx}$ implied by a deterministic transport map $z=q(x)$ that solves the constrained optimization problem ((ref)) and yielding a transport cost $\mathbb{E}_{\mu_x} \left[ \left\Vert q(x)-x\right\Vert ^{2}\right]$. Since determining the function $q$ amounts to finding which value $z$ each point $x$ should be mapped to, the sample version of this problem can be stated as

equation[equation omitted — 132 chars of source]

subject to:

equation[equation omitted — 91 chars of source]

where $\hat{\mathbb{E}}$ denotes sample averages (i.e. $\hat{ \mathbb{E}}[a\left( x\right) ]\equiv \frac{1}{n}\sum_{i=1}^{n}a\left( x_{i}\right) $, where $n$ is sample size and $a(\cdot)$ a given function). This optimization problem is then nested into an optimization over $\theta $, which delivers the estimated parameter value $\hat{\theta}$. We call this estimator an Optimally-Transported\ GMM (OTGMM) estimator.

Our approach is conceptually similar to GEL, in that it minimizes some concept of distributional distance under moment constraints. Yet, the notion of distance used differs significantly. As shown in Figure (ref), the distance here is measured along the \textquotedblleft observation values\textquotedblright\ axis rather than the \textquotedblleft observation weights\textquotedblright\ axis (as it would be in GEL). This feature arguably makes the method a hybrid between GEL and optimal transport, since GEL's goal of satisfying all the moment conditions is achieved through optimal transport instead of optimal reweighting. (The distinction from GEL applies to the discrete case as well, since the OTGMM objective function depends on both the amount and location of probability mass transfers, while the GEL objective function is only sensitive on the amount, but not on the specific locations, of the probability mass transfers.) Most of our regularity conditions will parallel those of optimal GMM and GEL, but some will not (they are tied to the optimal transport nature of the problem and involve assumptions regarding higher order derivatives). In analogy with the behavior of GEL estimators, the OTGMM estimator will be shown to be root-$n$ consistent and asymptotically normal, despite involving an optimization problem having an infinite-dimensional nuisance parameter (the $z_i$). In general, however, OTGMM's asymptotic variance does not coincide with that of GEL or efficiently weighted GMM under correct specification. Hence, OTGMM is most useful when errors in the variables or Wasserstein-type deviations in the data distribution constitute a primary concern.

figure[figure omitted — 463 chars of source]

The remainder of the paper is organized as follows. We first formally define and solve the optimization problem defining our estimator, before considering the limit of small errors (in the spirit of chesher:1991)\ to gain some intuition. We then derive the resulting estimator's asymptotic properties for the general, large error, case. We show asymptotic normality and root $n$ consistency in both cases. We then discuss related approaches and some extensions. The method's practical usefulness is illustrated by revisiting an influential study of the relationship between a city's exports and the extent of its transportation infrastructure (duranton2014roads). Our results corroborate that study under weaker assumptions and provide insight into the error structure of the variables.

The estimator

Definition

The Lagrangian associated with the constrained optimization problem defined in Equations ((ref)) and ((ref)) is

equation*[equation* omitted — 164 chars of source]

where $\hat{\mathbb{E}}[\ldots ]$ denotes a sample average and where $ \lambda $ is a Lagrange multiplier. The dual problem's first-order conditions with respect to $\theta $, $\lambda $ and $z_{j}$, respectively, are then

eqnarray[eqnarray omitted — 313 chars of source]

where we let $\partial_{v}$ denote a partial derivative with respect to argument $v$. We shall use $\partial_{v^{\prime }}$ to denote a matrix of partial derivatives with respect to a transposed variable (e.g., $\partial_{\theta^{\prime }}g(z,\theta )\equiv \partial g(z,\theta )/\partial \theta ^{\prime }$). This formulation of the problem assumes differentiability of $g\left( z,\theta \right) $ to a sufficiently high order, as shall be formalized in our asymptotic analysis.

Implementation

The nonlinear system ((ref))-((ref)) of equations can be solved numerically. To this effect, we propose an iterative procedure to determine the $z_{j}$, $\lambda $ for a given $\theta $. This yields an\ objective function $\hat{Q}(\theta )$ that can be minimized to estimate $ \theta $. Let $z_{j}^{t}$ and $\lambda ^{t}$ denote the approximations obtained after $t$ steps. As shown in Supplement (ref), given tolerances $\epsilon ,\epsilon ^{\prime }$ and a given $\theta $, the objective function $\hat{Q}(\theta )$ can be determined as follows:

algorithm[algorithm omitted — 1,088 chars of source]

This algorithm is obtained by substituting $z_{j}=x_{j}+H^{\prime }\left( z_{j},\theta \right) \lambda $ obtained from Equation ((ref)) into Equation ((ref)) and expanding the resulting expression to linear order in $\lambda $. This linearized expression provides an improved approximation $\lambda ^{t}$ to the Lagrange multiplier which can, in turn, yield an improved approximation $z_{j}^{t}$. The process is then iterated to convergence. The expression for $\hat{Q}(\theta )$ is obtained by re-expressing $\hat{\mathbb{E}}\left[ \left\Vert z-x\right\Vert ^{2}\right] $ using Equation ((ref)). Formal sufficient conditions for the convergence of this iterative procedure can be found in Supplement (ref). In cases where this simple approach fails to converge, one can employ robust schemes based on a combination of discretization and linear programming (see Section 6.4.1 in santambrogio:optran).

To gain some intuition regarding the estimator, it is useful to consider the limit of small errors when solving Equations ((ref))-((ref)), in the spirit of chesher:1991). This limit corresponds to assuming that higher-order powers of $\Vert z_i-x_i \Vert$ are negligible relative to $\Vert z_i-x_i \Vert$ itself. In this limit, the estimator admits a closed form with an intuitive interpretation, as shown by the following result, shown in Appendix (ref).

propositionTo the first order in $z_{i}-x_{i}$ ($i=1,\ldots ,n$) the estimator is equivalent to minimizing a GMM-like objective function with respect to $\theta $ with a specific choice of weighting matrix: \begin{equation} \hat{\theta}=\arg \min_{\theta }\hat{\mathbb{E}}\left[ g^{\prime }\left( x,\theta \right) \right] \left( \hat{\mathbb{E}}\left[ H\left( x,\theta \right) H^{\prime }\left( x,\theta \right) \right] \right) ^{-1}\hat{\mathbb{ E}}\left[ g\left( x,\theta \right) \right] . \end{equation}

From this expression, it is clear that the estimator downweights the moments that are the most sensitive to errors in $x$, as measured by $ H\left( x,\theta \right) \equiv \partial_{z^{\prime }}g\left( x,\theta \right) $. This accomplishes the desired goal of minimizing the effect of the errors when the properties of the error process are unknown.

Although this weighting matrix appears suboptimal (relative to a correctly specified optimally weighted GMM estimator), one should realize that the notion of optimality depends on what class of data generating processes the “true model” encompasses. Optimal GMM and GEL can be seen as Maximum Likelihood estimators (chamberlain:gmmeff, newey:gel) under moment conditions expressed in terms of the observed $x$. In contrast, OTGMM can be interpreted as a Maximum Likelihood estimator for homoskedastic and normally distributed errors ($x-z$) under moment constraints on the unobserved $z$. There is therefore a clear efficiency-robustness trade-off: OTGMM is less efficient if the observed $x$ satisfy the moment constraints, but allows for additional error terms that maintain the model's correct specification even if the observed $x$ do not satisfy the overidentified moment constraints, a more general setting where optimally weighted GMM or GEL offers no efficiency guarantees.

Asymptotics

In this section, we show that, despite the estimator's roots in the theory of optimal transport, its large sample behavior remains amenable to standard asymptotic tools since our focus is on an estimator of the parameter $\theta $ rather than on an estimator of a distribution. We first consider the case of small errors, a limiting case that may be especially important in the relatively common case of applications where overidentifying restrictions tests are near the rejection region boundary. This limit also parallels the approach taken in the GEL literature, where asymptotic properties are often derived in the case where the overidentifying restrictions hold (e.g., newey:gel).

Small errors limit

Our small error results enable us to illustrate that there is little risk in using our estimator instead of efficient GMM when one is concerned about overidentification test failure. If the data were to, in fact, satisfy the moment conditions, using our approach does not sacrifice consistency, root $n$ convergence or asymptotic normality. The only possible drawback would be a suboptimal weighting of overidentifying moment conditions, potentially leading to an increase in variance if the model happened to be correctly specified. Conversely, if the model is misspecified, e.g., because the data is error-contaminated, the optimal weighting of efficient GMM is no longer the optimal weighting (since random deviations due to sampling variability are not the main reason for the failure to simultaneously satisfy all moment conditions). For instance, if the errors are such that there is an unknown bias in the moment conditions that decays to zero asymptotically but at a rate possibly slower than $n^{-1/2}$, then the model is still correctly specified asymptotically but the bias dominates the random sampling error. Then, the optimal weighting should seek to minimize the effect of error-induced bias, which our approach seeks to accomplish by weighting based on the effect of errors in the variables on the moment conditions. Hence, in that sense, the method provides a complementary alternative to standard GMM estimation offering a different trade-off between efficiency and robustness to misspecification.

Our consistency result requires a number of fairly standard primitive assumptions.

conditionThe random variables $x_{i}$ are iid and take value in $\mathcal{X}\subset \mathbb{R}^{d_{x}}$.
condition$\mathbb{E}[g(x_{i};\theta _{0})]=0$, and $\mathbb{E}[g(x_{i};\theta )]\neq 0$ for other $\theta \in \Theta $, a compact set.

In other words, Assumption (ref) indicates that we consider here the case where GMM would be consistent, in analogy with the setup traditionally considered in the GEL literature (e.g. newey:gel).

condition$\mathbb{V}[g(x_{i};\theta _{0})]<\infty $, where $\mathbb{V}$ denotes the variance operator.
condition$g(x;\cdot )$ is almost surely continuous and $\Vert g(x;\theta )\Vert\leq h(x)$ for any $\theta \in \Theta $ and for some function $ h$ satisfying $\mathbb{E}\left[ h\left( x_{i}\right) \right] <\infty $.

While Assumptions (ref), (ref), (ref) and (ref) directly parallel those needed to establish the asymptotic properties of a standard GMM estimator (e.g. Theorems 2.6 and 3.2 in Newey:HB), our estimator requires a few more low-level regularity conditions. Given that our estimator, in the small error limit (Equation ((ref))), involves a sample average involving derivative $H\left( x,\theta \right) \equiv \partial_{z^{\prime }}g\left( x,\theta \right) $, we need to place some constraints on the behavior of that quantity as well. Below, we let $\left\Vert a\right\Vert =\left(\sum_{i,j}a_{i,j}^2\right) ^{1/2}$ for a matrix $a$.

condition$g$ is differentiable in its first argument and the derivative satisfies $\mathbb{E}[\Vert \partial_{z^{\prime }}g(x_{i};\theta _{0})\Vert ^{2}]\allowbreak <\infty $. Moreover, $\Vert \partial_{z^{\prime }}g(x_{i};\theta )\Vert \neq 0$ almost surely for all $\theta \in \Theta $.
condition$\partial_{z^{\prime }}g(x;\theta _{0})$ is H\"{o}lder continuous in $x$.
condition$\mathbb{E}[\partial_{z^{\prime }}g(x_{i};\theta _{0})\partial_{z}g^{\prime }(x_{i};\theta _{0})]$ exists and is of full rank.

These assumptions ensure that the minimization problem defined by ((ref)) and ((ref)) is well-behaved, i.e., small changes in the values of $x_{i}$ do not lead to jumps in the solution $z_{i}$ to the optimization problem (aside from zero-probability events). It is likely that these assumptions can be relaxed using empirical processes techniques. However, here we favor simply imposing more smoothness (compared to the standard GMM assumptions), because this leads to more transparent assumptions. They can all be stated in terms of the basic function $g(x;\theta )$ that defines the moment condition model, making them fairly primitive. We can then state our first consistency result.

theoremUnder assumptions (ref)-(ref), the OTGMM estimator is consistent for $\theta _{0}$ and $\lambda =O_{p}(n^{-1/2})$.

As a by-product, this theorem also secures a convergence rate on the Lagrange multiplier $\lambda $ which proves useful for establishing our distributional results. The conditions needed to show asymptotic normality also closely mimic those of standard GMM estimators (e.g. Theorem 3.2 in Newey:HB):

condition$\theta _{0}\in \Theta ^{\circ }$, the interior of $\Theta $.
condition$\mathbb{E}[\sup_{\theta \in \eta }\Vert \partial_{\theta^{\prime }}g(x_{i};\theta )\Vert ]<\infty $ where $\eta \subset \Theta $ is a neighborhood of $\theta _{0}$.
condition$\left( \mathbb{E}[\partial_{\theta'}g(x_{i};\theta _{0})^{\prime }]\left( \mathbb{E}[\partial_{z'}g(x_{i};\theta _{0})\partial _{z'}g(x_{i},\theta _{0})^{\prime }]\right) ^{-1}\mathbb{E}[\partial _{\theta'}g(x_{i};\theta _{0})]\right) $ is invertible.

We can then provide an explicit expression of the estimator's asymptotic variance.

theoremUnder Assumptions (ref)-(ref), the OTGMM estimator is asymptotically normal with $\sqrt{n}(\hat{\theta} _{OTGMM}-\theta _{0})\rightarrow ^{d}\mathcal{N}(0;V)$, where \begin{eqnarray*} V &=&\left( \mathbb{E}[G_i^{\prime }]\left( \mathbb{E}[H_i H_i^{\prime }]\right) ^{-1}\mathbb{E}[G_i]\right) ^{-1}\times (\mathbb{E}[G_i^{\prime }]\left( \mathbb{E}[H_iH_i^{\prime }]\right) ^{-1}\mathbb{E}[g_ig_i^{\prime }]\left( \mathbb{E}[H_i H_i ^{\prime }]\right) ^{-1}\mathbb{E}[G_i])\times \\ &&\left( \mathbb{E}[G_i^{\prime }]\left( \mathbb{E}[H_i H_i^{\prime }]\right) ^{-1}\mathbb{E}[G_i]\right) ^{-1}, \end{eqnarray*} where\ $H_i \equiv \partial_{z^{\prime }}g(x_{i};\theta _{0})$, $G_i \equiv \partial_{\theta^{\prime }}g(x_{i};\theta _{0})$ and $g_i \equiv g(x_{i};\theta _{0})$.

This simple normal limiting distribution with root $n$ convergent behavior is somewhat unexpected from an estimator that involves a high-dimensional optimization over $O(n)$ latent variables. As in GEL, this is made possible thanks to the existence of an equivalent low-dimensional dual optimization problem, which is, in turn, equivalent to a simple GMM estimator, albeit with a nonstandard weighting matrix. Thus, the variance has the expected \textquotedblleft sandwich\textquotedblright\ form, since the reciprocal weights $\mathbb{E}[H_i H_i^{\prime }]$ differs from the moment variance $\mathbb{E}[g_i g_i^{\prime }]$. For comparison, an optimally weighted GMM estimator would have an asymptotic variance of $(\mathbb{E}[G_i]'(\mathbb{E}[g_ig_i'])^{-1}\mathbb{E}[G_i])^{-1}$.

It is difficult to formulate a simple expression for the difference $V_{OTGMM}-V_{GMM}$ between the asymptotic variance of OTGMM and that of optimal GMM, but a simple example suffices to illustrate that this difference could be any non-negative-definite matrix.

Let $x_{i}$ take value in $\mathbb{R}^{d_{x}}$ ($d_{x}\geq 2$) with $\mathbb{E}\left[ x_{i\ell }\right] =0$ and $\mathbb{E}\left[ x_{i\ell }x_{i\ell ^{\prime }}\right] \equiv \omega _{\ell } $ when $\ell=\ell'$ and zero otherwise, for $\ell ,\ell ^{\prime }=1,\ldots ,d_{x}$. For $\theta \in \mathbb{R}$, let the moment functions be given by $g_{\ell }\left( x_{i},\theta \right) =x_{i\ell }-\theta $. We then have \[ V_{OTGMM}=\frac{1}{d_{x}^{2}}\sum_{\ell =1}^{d_{x}}\omega _{\ell }\text{ and }V_{GMM}=\left( \sum_{\ell =1}^{d_{x}}\omega _{\ell }^{-1}\right) ^{-1}. \] Then, by the harmonic-arithmetic mean inequality, $V_{OTGMM}\geq V_{GMM}$ with equality when $\omega _{\ell }=c$ for all $\ell $. If $\omega _{1}\longrightarrow \infty $ leaving all the other $\omega _{\ell }$ finite, $V_{OTGMM}\longrightarrow \infty $ while $V_{GMM}$ remains finite. (General non-negative-definite differences can be obtained by stacking such moment conditions for different parameters $\theta_k$ ($k=1,\ldots,d_\theta$) and possibly linearly transforming the resulting parameter vector $\theta$.) Generally, we expect the difference to be large when some moments have a disproportionally large variance while not being disproportionally sensitive to changes in the underlying variables.

Asymptotics under large errors

In some applications, there may be considerable misspecification or its magnitude may be a priori unknown. It thus proves useful to relax the assumption of small errors in deriving the estimator's asymptotic properties. To handle this more general setup, we employ the following equivalence, demonstrated in Appendix (ref).

theoremIf $g\left( z,\theta \right) $ is differentiable in its arguments, the OTGMM estimator is equivalent to a just-identified GMM estimator expressed in terms of the modified moment function \begin{equation} \tilde{g}\left( x,\theta ,\lambda \right) =\left[ \begin{array}{c} \partial_{\theta}g^{\prime }\left( q\left( x,\theta ,\lambda \right) ,\theta \right) \lambda \\ g\left( q\left( x,\theta ,\lambda \right) ,\theta \right) \end{array} \right] \end{equation} that is a function of the observed data $x$ and the augmented parameter vector $\tilde{\theta}\equiv \left( \theta ^{\prime },\lambda ^{\prime }\right) ^{\prime }$ and where \begin{equation} q\left( x,\theta ,\lambda \right) \equiv \arg \min_{z:z-\partial_{z}g^{\prime }\left( z,\theta \right) \lambda =x}\left\Vert z-x\right\Vert ^{2}. \end{equation}

Note that $q\left( x,\theta ,\lambda \right) $ is essentially the inverse of the mapping $z-\partial_{z}g^{\prime }\left( z,\theta \right) \lambda =x$ (from Equation ((ref))), augmented with a rule to select the appropriate branch in case the inverse is multivalued.

The equivalence result of Theorem (ref) implies that many of the asymptotic technical tools used in GMM-type estimators can be adapted to our setup, with the distinction that the function $q\left( x,\theta ,\lambda \right) $ is defined only implicitly. Hence, many of our efforts below seek to recast necessary conditions on $q\left( x,\theta ,\lambda \right) $ in terms of more primitive conditions on the moment function $g\left( z,\theta \right) $ whenever possible or in terms of regularity conditions drawn from optimal transport theory.

We start with a standard GMM-like identification condition:

conditionFor some compact sets $\Theta $ and $\Lambda $, there exists a unique $\left( \theta _{0},\lambda _{0}\right) \in \Theta \times \Lambda $ solving $\mathbb{E}\left[ \tilde{g}\left( x,\theta ,\lambda \right) \right] =0$ for $\tilde{g}\left( x,\theta ,\lambda \right) $ defined in Theorem (ref).

This condition is implied by a natural uniqueness and regularity condition on the solution to the primal optimal transportation problem (Equations ((ref))):

conditionLet $\mu_{zx;\theta}$ denote the solution to Problem ((ref)) for a given $\theta\in\Theta$. (i) $\mathbb{E}_{\mu_{zx;\theta}}[\Vert z-x \Vert^2]$ is uniquely minimized at $\theta=\theta_0$ (ii) The corresponding marginals $\mu_{z;\theta}$ and $\mu _{x}$ are absolutely continuous with respect to the Lebesgue measure with a density that is finite, nonvanishing and H\"{o}lder continuous on their convex support.

Indeed, by Theorem [C3], part b) and d), in caffarelli:regconv, Assumption (ref)(ii) implies that, at each $\theta$, there exists a unique invertible transport map $z=q\left( x\right) $ from $\mu _{x}$ to $\mu _{z;\theta}$ and both $q$ and its inverse are equal to the gradient of a twice differentiable strictly convex function. The fact that $q^{-1}$ is the gradient of a twice differentiable strictly convex function ensures that the first-order condition in Assumption (ref) has a unique solution (making a rule to handle multivalued inverses unnecessary).

Next, we consider standard continuity and dominance conditions that are used to establish uniform convergence of the GMM objective function. These assumptions constitute a superset of those needed for standard GMM because the modified moment conditions include the additional parameter $\lambda$ and higher-order derivatives of the original moment conditions. In a high-level form, these conditions read:

condition(i) $\tilde{g}\left( x,\theta ,\lambda \right) $ is continuous in $\theta $ and $\lambda $ for $\left( \theta ,\lambda \right) \in \Theta \times \Lambda $ with probability one and (ii) $\mathbb{E}\left[ \sup_{\left( \theta ,\lambda \right) \in \Theta \times \Lambda }\left\Vert \tilde{g}\left( x,\theta ,\lambda \right) \right\Vert \right] <\infty $.

Alternatively, Assumption (ref) can be replaced by more primitive conditions on $g\left( z,\theta \right) $ instead, as given below in Assumptions (ref), (ref) and (ref).

condition(i) $g\left( z,\theta \right) $ and $\partial_{z^{\prime }}g\left( z,\theta \right) $ are differentiable in $\theta $ and (ii) $\partial_{\theta^{\prime }}g\left( z,\theta \right) $ is continuous in both arguments.

This assumption parallels continuity assumptions typically made for GMM, but higher order derivatives of $g\left( z,\theta \right) $ are needed, because they enter the moment condition either directly or indirectly via the function $q\left( x,\theta ,\lambda \right) $. The next condition ensures that the function $q\left( x,\theta ,\lambda \right) $ is well behaved.

condition$\bar{\nu}\bar{\lambda}<1$ where $\bar{\lambda} =\max_{\lambda \in \Lambda }\left\Vert \lambda \right\Vert $ and $\bar{\nu} =\sup_{\theta \in \Theta }\sup_{z\in \mathcal{X}}\max_{k\in \left\{ 1,\ldots ,d_{g}\right\} }\allowbreak \max \operatorname{eigval}\left( \partial_{zz^{\prime }}g_{k}\left( z,\theta \right) \right) $, in which $\partial_{zz^{\prime }}g_{k}\left( z,\theta \right) $ exists for $k=1,\ldots ,d_{g}$ and where $ \operatorname{eigval}\left( M\right) $ for some matrix $M$ denotes the set of its eigenvalues.

Once again, this condition can be alternatively phrased in terms of optimal transport concepts. The first order condition which implicitly defines $z=q\left( x,\theta ,\lambda \right)$ can be written in terms of the derivative of a potential function $\psi(z,\theta,\lambda)=z^{\prime }z/2-g^{\prime }\left(z,\theta \right) \lambda$: $\nabla_z \psi(z,\theta,\lambda)=x$. With the help of Theorem [C3], part b) and d), in caffarelli:regconv, Assumption (ref)(ii) implies that, at each $\theta$, the above potential $\psi(z,\theta,\lambda)$ has a positive-definite Hessian (with respect to $z$), which implies Assumption (ref).

In order to state our remaining regularity conditions, it is useful to introduce a notion of (nonuniform) Lipschitz continuity, combined with dominance conditions.

definitionLet $\mathcal{L}$ be the set of functions $h\left( z,\theta \right) $ such that (i) $\mathbb{E}\left[ \sup_{\theta \in \Theta }\left\Vert h\left( x,\theta \right) \right\Vert \right]\linebreak <\infty $ and (ii) there exists a function $\bar{h}\left( x,\theta \right) $ satisfying \begin{eqnarray} \mathbb{E}\left[ \sup_{\theta \in \Theta }\bar{h}\left( x,\theta \right) \left\Vert \partial_{z^{\prime }}g\left( x,\theta \right) \right\Vert \right] &<&\infty . \\ \left\Vert h\left( z,\theta \right) -h\left( x,\theta \right) \right\Vert &\leq &\bar{h}\left( x,\theta \right) \left\Vert z-x\right\Vert \end{eqnarray} for all $x,z\in \mathcal{X}$ and $\theta\in\Theta$, and where $g\left( x,\theta \right) $ is as in the moment conditions.

This Lipschitz continuity-type assumption has no parallel in conventional GMM. It is made here because it ensures that the behavior of the observed $x$ and the underlying unobserved $z$ will not differ to such an extent that moments of unobserved variables would be infinite, while the corresponding observed moments are finite. Clearly, without such an assumption, observable moments would be essentially uninformative. The idea underlying Definition (ref) is that we want to define a property that is akin to Lipschitz continuity but that allows for some heterogeneity (through the function $\bar{h}\left( x,\theta \right) $ in Equation ((ref))). This heterogeneity proves particularly useful in the case where $\mathcal{X}$ is not compact (for compact $\mathcal{X}$, one can take $\bar{h}\left( x,\theta \right) $ to be constant in $x$ with little loss of generality). For a given function $ h\left( x,\theta \right) $ that is finite for finite $x$, membership in $ \mathcal{L}$ is easy to check by inspecting the tail behavior (in $x$) of the given function $h\left( x,\theta \right) $. Polynomial tails will suggest a polynomial form for $\bar{h}\left( x,\theta \right) $, for instance. Equation ((ref)) strengthens the dominance condition (ref)(i) to ensure that functions $h\left( x,\theta \right) $ in $ \mathcal{L}$ also satisfy a dominance condition when interacted with other quantities entering the optimization problem, i.e. $\partial_{z^{\prime}}g\left( x,\theta \right) $).

With this definition in hand, we can succinctly state a sufficient condition for $\tilde{g}\left( x,\theta ,\lambda \right) $ to satisfy a dominance condition:

condition$\,$$g\left( \cdot ,\cdot \right) $ and each element of $ \partial_{\theta}g^{\prime }\left( \cdot ,\cdot \right) $ belong to $\mathcal{L}$ .

We are now ready to state our general consistency result.

theoremUnder Assumptions (ref), (ref) and either Assumption (ref) or Assumptions (ref), (ref), (ref), the OTGMM estimator is consistent ($(\hat{\theta},\hat{\lambda})\overset{p}{\longrightarrow }(\theta _{0},\lambda _{0}) $).

We now turn to asymptotic normality. We first need a conventional \textquotedblleft interior solution\textquotedblright\ assumption.

condition$\left( \theta _{0},\lambda _{0}\right) $ from Assumption (ref) lies in the interior of $\Theta \times \Lambda $.

Next, as in any GMM estimator, we need finite variance of the moment functions and their differentiability:

condition(i) $\mathbb{V}\left[ \tilde{g}\left( x,\theta _{0},\lambda _{0}\right) \right] \equiv \Omega $ exists and (ii) $\mathbb{E} \left[ \partial \tilde{g}\left( x,\theta ,\lambda \right) /\partial \left( \theta ^{\prime },\lambda ^{\prime }\right) \right] \equiv \tilde{G}$ exists and is nonsingular.

Assumption (ref)(ii) can be expressed in a more primitive fashion using the explicit form for $\tilde{G}$ provided in Theorem (ref) below.

Next, we first state a high-level dominance condition that ensures uniform convergence of the Jacobian term $\partial \tilde{g}\left( x,\theta ,\lambda \right) /\partial \left( \theta ^{\prime },\lambda ^{\prime }\right) $.

condition(i) $\tilde{g}\left( x,\theta ,\lambda \right) $ is continuously differentiable in $\left( \theta ,\lambda \right) $; (ii) $ \mathbb{E[}\sup_{\left( \theta ,\lambda \right) \in \Theta \times \Lambda }\allowbreak \left\Vert \partial \tilde{g}\left( x,\theta ,\lambda \right) /\partial \left( \theta ^{\prime },\lambda ^{\prime }\right) \right\Vert ]<\infty $.

This assumption is implied by the following, more primitive, condition:

condition(i) $g\left( z,\theta \right) $ and $\partial_{\theta}g\left( z,\theta \right) $ are continuously differentiable in $\theta $, (ii) all elements of $\partial_{\theta}g_{k}\left( z,\theta \right) $ and $ \partial_{\theta\theta^{\prime }}g_{k}\left( z,\theta \right) $ for $k=1,\ldots ,d_{g} $ belong to $\mathcal{L}$ and (iii) Assumptions (ref)(i) and (ref) hold.

We can now state our general asymptotic normality and root-$n$ consistency result, shown in Appendix (ref).

theoremLet the assumptions of Theorem (ref) hold as well as Assumptions (ref), (ref) and either Assumption (ref) or (ref). Then, \begin{equation*} \sqrt{n}\left( \left[ \begin{array}{c} \hat{\theta} \\ \hat{\lambda} \end{array} \right] -\left[ \begin{array}{c} \theta _{0} \\ \lambda _{0} \end{array} \right] \right) \overset{d}{\longrightarrow }\mathcal{N}\left( 0,W^{-1}\right) \end{equation*} where $W=\tilde{G}^{\prime }\Omega ^{-1}\tilde{G}$, $\Omega =\mathbb{E}\left[ \tilde{g}\tilde{g}^{\prime }\right],$ \begin{equation*} \tilde{g} \equiv \tilde{g}(z,(\theta_0,\lambda_0)) = \left[ \begin{array}{c} \partial_{\theta}g^{\prime }\left( z,\theta_0 \right) \lambda_0 \\ g\left( z,\theta_0 \right) \end{array} \right] and \tilde{G} \equiv \mathbb{E}[\partial_{\theta} \tilde{g}'] = \left[ \begin{array}{cc} \tilde{G}_{\theta \theta } & \tilde{G}_{\theta \lambda } \\ \tilde{G}_{\lambda \theta } & \tilde{G}_{\lambda \lambda } \end{array} \right] \end{equation*} in which \begin{eqnarray*} \tilde{G}_{\theta \theta } &\equiv &\mathbb{E}\left[ \partial_{\theta \theta^{\prime }}\left( \lambda _{0}^{\prime }g\left( z,\theta _{0}\right) \right) +\partial_{\theta z^{\prime }}\left( \lambda _{0}^{\prime }g\left( z,\theta _{0}\right) \right) \partial_{\theta^{\prime }}q\left( x,\theta _{0},\lambda _{0}\right) \right] \\ \tilde{G}_{\lambda \theta } &\equiv &\mathbb{E}\left[ \partial_{\theta^{\prime }}g\left( z,\theta _{0}\right) +\partial_{z^{\prime }}g\left( z,\theta _{0}\right) \partial_{\theta^{\prime }}q\left( x,\theta _{0},\lambda _{0}\right) \right] \\ \tilde{G}_{\theta \lambda } &\equiv &\mathbb{E}\left[ \partial_{\theta}\left( g^{\prime }\left( z,\theta _{0}\right) \right) +\partial_{\theta z^{\prime }}\left( \lambda _{0}^{\prime }g\left( z,\theta _{0}\right) \right) \partial_{\lambda^{\prime }}q\left( x,\theta _{0},\lambda _{0}\right) \right] \\ \tilde{G}_{\lambda \lambda } &\equiv &\mathbb{E}\left[ \partial_{z^{\prime }}g\left( z,\theta _{0}\right) \partial_{\lambda^{\prime }}q\left( x,\theta _{0},\lambda _{0}\right) \right] \end{eqnarray*} where $z$ solves $x=z-\partial_{z}g^{\prime }\left( z,\theta \right) \lambda $ for given $x,\theta ,\lambda $ and where \begin{eqnarray} \partial_{\theta^{\prime }}q\left( x,\theta ,\lambda \right) &=&\left[ \left( I-\partial_{zz^{\prime }}\left( \lambda ^{\prime }g\left( z,\theta \right) \right) \right) ^{-1}\partial_{z\theta^{\prime }}\left( \lambda ^{\prime }g\left( z,\theta \right) \right) \right] _{z=q\left( x,\theta ,\lambda \right) } \\ \partial_{\lambda^{\prime }}q\left( x,\theta ,\lambda \right) &=&\left[ \left( I-\partial_{zz^{\prime }}\left( \lambda ^{\prime }g\left( z,\theta \right) \right) \right) ^{-1}\partial_{z}g^{\prime }\left( z,\theta \right) \right] _{z=q\left( x,\theta ,\lambda \right) }. \end{eqnarray} In particular, for $\theta $, the partitioned inverse formula gives \begin{equation*} \sqrt{n}\left( \hat{\theta}-\theta _{0}\right) \overset{d}{\longrightarrow } \mathcal{N}\left( 0,\left( W_{\theta \theta }-W_{\theta \lambda }W_{\lambda \lambda }^{-1}W_{\lambda \theta }\right) ^{-1}\right) \end{equation*} where $W$ is similarly partitioned as: \begin{equation*} W=\left[ \begin{array}{cc} W_{\theta \theta } & W_{\theta \lambda } \\ W_{\lambda \theta } & W_{\lambda \lambda } \end{array} \right] \end{equation*}

The asymptotic variance stated in Theorem (ref) takes the familiar form expected from a just-identified GMM estimator: $(\tilde{G}^{\prime }\Omega ^{-1}\tilde{G})^{-1}$. The relatively lengthy expressions merely come from explicitly computing the first derivative matrix $\tilde{G}$ in terms of its constituents. This is accomplished by differentiating $\tilde{g}$ with respect to all parameters using the chain rule and calculating the derivative of $q\left( x,\theta ,\lambda \right) $ using the implicit function theorem.

We thus have now completely characterized the first-order asymptotic properties of our estimator in the most general settings of large (i.e., non-local) misspecification. This result thus allows researcher to directly replace their GMM estimator which may happen to fail overidentification tests by another, logically consistent and easy-to-interpret, estimator where the overidentification failure is naturally accounted for by errors in the variables. In addition, researchers can further document the presence of errors via Theorem (ref), as it enables, as a by-product, a formal test of the absence of error. Under this null hypothesis, which can be stated as $\lambda=0$, we have

equation[equation omitted — 204 chars of source]

where the above expression can be straightforwardly derived from the partitioned inverse formula applied to the $\lambda$ sub-block of the asymptotic variance.

Discussion and extensions

On a conceptual level, our use of a so-called Wasserstein metric to measure distance between distributions does provide some desirable theoretical properties. For instance, the Wasserstein metric metrizes convergence in distribution (see Theorem 6.9 in villani:new) under some simple bounded moment assumptions. In contrast, the discrepancies which generate GEL estimators do not admit such an interpretation. In fact, most discrepancies are not metrics, as they lack symmetry. The Kullback-Leibler discrepancy, which is perhaps the best known among them, does not allow comparison between distributions that are not absolutely continuous with respect to one another, whereas the Wassertein metric does. (Of course, leveraging this advantage requires considering the most general transport problem of Equation ((ref)).) Finally, it is arguably logical to penalize probability transfer over larger distances more than the same probability transfer over smaller distances, as the Wasserstein metric does, while none of GEL-related discrepancies do.

An interesting extension of our approach would be a hybrid method in which (i) the possibility of general forms of errors is accounted for with the current method by constructing the equivalent GMM formulation of the model via Theorem (ref) and (ii) additional restrictions on the form of the errors are imposed via additional moment conditions involving some elements of $z$ and $x$. This could prove a useful middle ground when a priori information regarding the errors is available for some, but not all, variables.

As shown in Supplement (ref), it is straightforward to extend our approach to allow error in some, but not all variables. When our method is used while allowing for errors in only a few variables, it may not be possible to simultaneously satisfy all moment conditions for any $\theta$. In such cases, it could make sense to consider a hybrid method where both errors in the variables, handled via our approach, and re-weighting of the sample, handled via GEL, are simultaneously allowed.

Finally, we should mention other approaches aimed at handling violations of overidentifying restrictions, which include the use of set identification combined with relaxed moment constraints (poirier:salvaging), placing a Bayesian prior on the magnitude of the deviations from correct specification of the moments (conley:plausexo), distributionally robust approaches that allow for deviations from the data generating process up to a given bound (connault:sens, blanchet:distrob), sensitivity analysis (shapiro:sens, bonhomme:sens) and misspecified moment inequality models (e.g. kwon:ineqmiss).

Application

We revisit the study of duranton2014roads, who documented evidence that the quantity of goods a city exports is strongly related to the extent of interstate highways present in that city. Due to simultaneity concerns, the authors adopt an instrumental variable approach to recover the causal effect of building highways. Although the authors mitigate instrument validity concerns with controls, instrument exogeneity or exclusion could remain a concern (as noted by poirier:salvaging) and this problem would manifest itself by specification test failure.

table[table omitted — 1,021 chars of source]

We apply our method to further assess the robustness of the results to potential model misspecification. We seek to recover point estimates that remain interpretable under potential misspecification and account for misspecification by viewing the model's variables as potentially measured with error. Not only do we consider errors in the regressors, but we also think of potentially invalid instruments as simply mismeasured versions of an underlying valid instrument that is unfortunately not available. Alternatively, one can think of the underlying valid instrument as the counterfactual value of the instrument in a world where the mechanism causing this instrument to be invalid would be absent. This broader interpretation of what could constitute an “error” under our framework considerably expands the scope of models that are conceptually consistent with our approach.

duranton2014roads consider three instruments: (log) kilometers of railroads in 1898, quantity of historical exploration routes, and planned (log) highway kilometers according to a 1947 construction map. The validity of these instruments could be criticized, for instance, in a situation where some cities are in proximity to key natural resources, which could cause higher exports and, at the same time, more transportation routes (or plans to build them). If this mechanism is active both in the present and in the past, the causal effect of highways on exports would be overestimated. In our framework, the true but unavailable instrument could represent a measure of past transportation routes in a counterfactual world where natural resources would be evenly distributed among cities. The actual available instruments represent an approximation to this ideal instrument, a situation which we represent as a potentially non-classical errors-in-variables model.

The model's moment conditions are written in terms of $g(z_i,\theta)=w_i(y_i-r_i'\theta)$, where the vector of observables $z_i=(y_i,r_i',w_i')'$ contains the dependent variable $y_i$, the vector of regressors $r_i$ and the vector of instruments $w_i$. In duranton2014roads, the dependent variable $y_i$ measures “propensity to export” and is constructed from an auxiliary panel data model, which regresses volume of exports between given cities on distance and trading partner characteristics, modeled as fixed effects. It is the value of these fixed effects that is used as $y_i$. Following poirier:salvaging, we take this construction as given and focus on export volume measured by weight. Explicitly accounting for errors in $y_i$ is superfluous since, in a regression setting, errors in the dependent variable are already allowed for.

In the analysis below, $r_i$ always includes the regressor of main interest: the (logarithm of) the number of kilometers of highway. It also contains a number of controls, which may differ in the different models considered. These controls include: log employment, log market access, log population in 1920, 1950, and 2000, and log manufacturing share in 2003, average distance to the nearest body of water, land gradient, dummy variables for census regions, log share of the fraction of adult population with a college degree or more, log average income per capita, log share of employment in wholesale trade, and log average daily traffic on the interstate highways in 2005. We allow for errors in all of these variables, except for the constant term and dummies.

We consider two of the specifications employed by duranton2014roads: The specification with many covariates that most obviously passes the overidentifying test and the specification with few covariates that most clearly fails this test. (The other specifications fall in between these extreme cases and are thus not reported here, for conciseness.) The results from GMM, replicated from the original study, and from OTGMM (allowing for large errors) are reported in Tables (ref) and (ref).

table[table omitted — 777 chars of source]

The OTGMM estimates are similar to those of duranton2014roads. Although some coefficients of the controls are larger in magnitude, the main coefficient --- the elasticity of export weight relative to kilometers of highway --- is almost unchanged and remains statistically significant. As in the original study, the 95% confidence intervals of the main elasticity of interest obtained with different sets of controls overlap significantly.

It is also instructive to look at the correction of the underlying variables implied by OTGMM. Table (ref), columns 1 and 2, documents this by looking at the standard deviation of different elements of the correction $z_{i} - x_{i}$. The corresponding quantities for the models of Tables (ref) and (ref) are reported in columns 1 and 2, respectively. To get an idea of the scale, the last column reports the standard deviations of the observed variables. As expected, the errors in column 1 are exceedingly small, reflecting the fact that the model passes the overidentification test. In contrast, column 2 is particularly interesting for our purposes because it corresponds to a specification that fails the test of over-identifying restrictions. While this may, at first, cast doubt regarding the GMM estimates, OTGMM shows that errors of the order of only 10% of the observed regressors' magnitude are sufficient to eliminate the mispecification. Such small error magnitudes are highly plausible empirically, thus supporting the plausibility of the OTGMM estimate. As the GMM and OTGMM estimates of the main elasticity of interest are very close in Table (ref), this also corroborates the authors' GMM estimates. Overall, our approach strongly supports the conclusions of the original study and thus provides an effective robustness test.

The fact that the model with more controls leads to smaller error magnitudes is very consistent with our interpretation: to the extent that including more controls reduces the magnitude of the potentially endogenous residuals, one would expect that our estimator has to perform smaller alterations to the variables to arrive at valid instruments and/or non-endogeneous regressors. More quantitatively, suppose that the instrument $w$ and the error term $\epsilon$ can be decomposed into separate components, say $w=w_1+w_2$ and $\epsilon = \epsilon_1 + \epsilon_2$, where $w_1$ and $\epsilon_1$ are correlated, but the latter is well explained by additional controls. In a model with fewer controls, our method has to identify $w_1$ and/or $\epsilon_1$ as error components to obtain a correctly specified model. However, if including controls already accounts for the $\epsilon_1$ term, then our estimator no longer needs to correct the corresponding components and thus achieves orthogonality with smaller errors.

In contrast, using GEL to address misspecification would effectively assume that the misspecification originates from some form of selection bias. The fact that adding more control eliminates the misspecification indicates that the controls incorporate the variables that explain selection bias. Yet, these controls are not included in the model with few controls, precisely the model where the largest amount of sample reweighting would take place when using GEL. This paradox makes the use of GEL as a remedy difficult to rationalize in a setting where misspecification arises primarily from the unavailability of adequate control variables.

table[table omitted — 959 chars of source]

In Supplement (ref), we perform a number of robustness tests along the lines suggested by poirier:salvaging. We further replicate specifications with many more controls and find that the GMM and OTGMM are still in close agreement. In Supplement (ref), we also consider alternative specifications that allow for the possibility that some instrumental variables should be included as regressors. The performance of OTGMM in other simple models is also investigated via simulations in Supplement (ref).

Conclusion

We have proposed a novel optimal transport-based version of the Generalized Method of Moment (GMM) that fulfills, by construction, the overidentifying moment restrictions by introducing the smallest possible amount of error in the variables or, equivalently, by allowing for the smallest possible Wasserstein metric distortions in the data distribution. This approach conceptually merges the Generalized Empirical Likelihood (GEL) and optimal transport methodologies. It provides a theoretically motivated interpretation to GMM results when standard overidentification tests reject the null. GEL approaches handle model misspecification by re-weighting the data, which would be appropriate when misspecification arise from improper sampling of the population. In contrast, our optimal transport approach is appropriate when the sources of misspecification are errors or, more generally, Wasserstein-type distortions in the data distribution, which is arguably a common situation in applications.