Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
49,081 characters · 11 sections · 40 citation commands
Optimal transportation and the falsifiability of incompletely specified economic models
\titlerunning{Falsifiability of incomplete models} \authorrunning{Ekeland et al.}
In many contexts, the ability to identify econometric models often rests on strong prior assumptions that are difficult to substantiate and even to analyze within the economic decision problem. A recent approach has been to forego such prior assumptions, thus giving up the ability to identify a single value of the parameter governing the model, and allow instead for a set of parameter values compatible with the empirical setup. A variety of models have been analyzed in this way, whether partial identification stems from incompletely specified models (typically models with multiple equilibria) or from structural data insufficiencies (typically cases of data censoring). See Manski:2005 for a recent survey on the topic.
All these incompletely specified models share the basic fundamental structure that a set of unobserved economic variables and a set of observed ones are linked by restrictions that stem from the theoretical economic model. In this paper, we propose a general framework for conducting inference in such contexts. This approach is articulated around the formulation of a hypothesis of compatibility of the true distribution of observable variables with the restrictions implied by the model as an optimal transportation problem. Given a hypothesized distribution for latent variables, compatibility of the true distribution of observed variables with the model is shown to be equivalent to the existence of a zero cost transportation plan from the hypothesized distribution of latent variables to the true distribution of observable variables, where the zero-one cost function is equal to one in cases of violations of the restrictions embodied in the model.
Two distinct types of economic restrictions are considered here. On the one hand, the case where the distribution of unobserved variables is parameterized yields a traditional optimal transportation formulation. On the other hand, the case where the distribution of unobserved economic variables are only restricted by a finite set of moment equalities yields an optimization formulation which is not a classical optimal transportation problem, but shares similar variational properties. In both cases the inspection of the dual of the specification problem's optimization formulation has three major benefits.
First, the optimization formulation relates the problem of falsifying incompletely specified economic models to the growing literature on optimal transportation (see RR:98a and Villani:2003), in particular with relation to the literature on probability metrics (see Zolotarev:97 chapter 1). Second, the dual formulation of the optimization problem provides significant dimension reduction, thereby allowing the construction of computable test statistics for the hypothesis of compatibility of true observable data distribution with the economic model given. Thirdly, and perhaps most importantly, in the case of models with discrete outcomes, the optimal transportation formulation allows to tap into a very rich combinatorial optimization literature relative to the discrete transport problem (see for instance PS:98) thereby allowing inference in realistic models of industrial organization and other areas of economics where sophisticated empirical research is being carried out.
The paper is organized as follows. The next section sets out the framework, notations and defines the problem considered. Section (ref) considers the case of parametric restrictions on the distribution of unobserved variables, gives the optimal transportation formulation of the compatibility of the distribution of observable variables with the economic model at hand, and discusses strategies to falsify the model based on a sample of realizations of the observable variables. Section (ref) similarly considers the case of semiparametric restrictions on the distribution of unobservable variables and the last section concludes.
We consider, as in Jovanovic:89, an economic model that governs the behaviour of a collection of economic variables $(Y,U)$, where $Y$ is a random element taking values in the Polish space ${\cal Y}$ (endowed with its Borel $\sigma$-algebra ${\cal B}_{\cal Y}$) and $U$ is a random element taking values in the Polish space ${\cal U}$ (endowed with its Borel $\sigma$-algebra ${\cal B}_{\cal U}$). $Y$ represents the subcollection of observable economic variables generated by the unknown distribution $P$, and $U$ represents the subcollection of unobservable economic variables generated by a distribution $\nu$. The economic model provides a set of restrictions on the joint behaviour of observable and latent variables, i.e. a subset of ${\cal Y}\times{\cal U}$, which can be represented without loss of generality by a correspondence $G: {\cal U}\rightrightarrows{\cal Y}$.
In all that follows, the correspondence will be assumed non-empty closed-valued and measurable, i.e. $G^{-1}({\cal O}):=\{u\in{\cal U}: G(u)\cap{\cal O}\neq\varnothing\}\in{\cal B}_{\cal U}$ for all open subset ${\cal O}$ of ${\cal Y}$. A measurable selection of a measurable correspondence $G$ is a measurable function $g$ such that $g\in G$ almost surely, and Sel$(G)$ denotes the collection of measurable selections of $G$ (non-empty by the Kuratowski--Ryll-Nardzewski selection theorem). We shall denote by $c(y,u)$ a cost of transportation, i.e. a real valued function on $\mathcal{Y}\times\mathcal{U}$. For any set $A$, we denote by $1_A$ its indicator function, i.e. the function taking value $1$ on $A$ and $0$ outside of $A$. $\mathcal{M}(\mathcal{Y})$ (resp. $\mathcal{M}(\mathcal{U})$) will denote the set of Borel probability measures on $\mathcal{Y}$ (resp. $\mathcal{U}$) and $\mathcal{M}(P,\nu)$ will denote the collection of Borel probability measures on $\mathcal{Y}\times\mathcal{U}$ with marginal distributions $P$ and $\nu$ on $\mathcal{Y}$ and $\mathcal{U}$ respectively. We shall generally denote by $\pi$ a typical element of $\mathcal{M}(P,\nu)$. For a Borel probability measure $\nu$ on $\mathcal{U}$ and a measurable correspondence $G: \mathcal{U}\rightrightarrows\mathcal{Y}$, we denote by $\nu G^{-1}$ the set function that to a set $A$ in $\mathcal{B}_\mathcal{Y}$ associates $\nu(G^{-1}(A))=\nu \left( \left\{ u\in \mathcal{U}:G\left( u\right) \cap A\neq \varnothing \right\} \right)$. Note that the set function $\nu G^{-1}$ is a {\em Choquet capacity functional} (see for instance Choquet:53). The {\em Core} of a Choquet capacity functional $\nu G^{-1}$, denoted $\mathrm{Core}(\nu G^{-1})$ is defined as the collection of Borel probability measures set-wise dominated by $\nu G^{-1}$, i.e. $\mathrm{Core}(\nu G^{-1})=\{Q\in\mathcal{M}(\mathcal{Y}): \forall A\in\mathcal{B}_\mathcal{Y}, Q(A)\leq\nu G^{-1}(A)\}$. In the terminology of cooperative games, if $\nu G^{-1}$ defines a transferable utility game, $\nu G^{-1}(A)$ is the utility value or worth of coalition $A$ and the Core of the game $\nu G^{-1}$ is the collection of undominated allocations (see Moulin:95).
We are interested in characterizing restrictions on the distribution of observables induced by the model, in order to devise methods to falsify the model based on a sample of repeated observations of $Y$. We shall successively consider two leading cases of this framework. First the case where the distribution $\nu$ of unobservable variables is given by the economic model, and second, the case where a finite collection of moments of the distribution $\nu$ of unobservable variables are given by the economic model.
The general principle we shall develop here in both parts is therefore the following. We want to test the compatibility of a reduced-form model , summarized by the distribution $P$ of an observed variable $Y$, with a structural model, summarized by a set $\mathcal{V}$ of distributions $\nu $ for the latent variable $U$. Two leading cases will be considered for the set $\mathcal{V}$: the parametric case, where $\mathcal{V} $ contains one element $\mathcal{V=}\left\{ \nu \right\} $, and the semiparametric case, where the distributions $\nu $ in $\mathcal{V}$ are specified by a finite number of moment restrictions $\mathbb{E}_{\nu }\left[ m_{i}\left( U\right) \right] =0$.
The restriction of the model defines compatibility between outcomes of the reduced-form and the structural models: such outcomes $u$ and $y$ are compatible if and only if the binary relation $y\in G\left( u\right) $ holds (this relation defines $G$).
Now we turn to the compatibility of the probabilistic models, namely of the specification of distributions for $U$ and $Y$. The models $Y\sim P$ and $ U\sim \nu \in \mathcal{V}$ are compatible if there is a joint distribution $ \pi $ for the pair $\left( Y,U\right) $ with respective marginals $P$ and some $\nu \in \mathcal{V}$ such that $Y\in G\left( U\right) $ holds $\pi $ almost surely. In other words, $P$ and $\mathcal{V}$\ are compatible if and only if \[ \exists \nu \in \mathcal{V},\exists \pi \in \mathcal{M}\left( P,\nu \right) :\Pr\nolimits_{\pi }\left\{ Y\notin G\left( U\right) \right\} =0. \] In the sequel we shall examine equivalent formulations of this compatibility principle, first in the parametric case and then in the semiparametric case.
Consider first the case where the economic model consists in the correspondence $G:\mathcal{U}\rightrightarrows \mathcal{Y}$ and the distribution $\nu$ of unobservables. The observables are fully characterized by their distribution $P$, which is unknown, but can be estimated from data. The question of compatibility of the model with the data can be formalized as follows: Consider the restrictions imposed by the model on the joint distribution $\pi$ of the pair $(Y,U)$:
A probability distribution $\pi$ that satisfies the restrictions above may or may not exist. If and only if it does, we say that the distribution $P$ of observable variables is compatible with the economic model $(G,\nu)$.
This hypothesis of compatibility has the following optimization interpretation. The distribution $P$ is compatible with the model $(G,\nu)$ if and only if $$\exists\pi\in{\cal M}(P,\nu): \int_{\mathcal{Y}\times\mathcal{U}}1_{\{y\notin G(u)\}}d\pi(y,u)=0,$$ and thus we see that it is equivalent to the existence of a zero cost transportation plan for the problem of transporting mass $\nu$ into mass $P$ with zero-one cost function $c(y,u)=1_{\{y\notin G(u)\}}$ associated with violations of the restrictions implied by the model.
The two dual formulations of this optimal transportation problem are the following:
Through applications of optimal transportation duality theory, it can be shown that the two programs are equal and that the infimum in $(\mbox{P})$ is attained, so that the compatibility hypothesis of definition (ref) is equivalent to $(\mbox{D})=0$, which in turn can be shown to be equivalent to
using the zero-one nature of the cost function to specialize the test functions $f$ and $h$ to indicator functions of Borel sets. Note that it is relatively easy to show necessity, since the definition of compatibility implies that ${Y\in A}\Rightarrow U\in G^{-1}(A)$, so that $1_{\{Y\in A\}}\leq1_{\{U\in G^{-1}(A)\}}$, $\pi$-almost surely. Taking expectation, we have $\mathbb{E}_\pi(1_{\{Y\in A\}})\leq \mathbb{E}_\pi(1_{\{U\in G^{-1}(A)\}})$, which yields $P(A)\leq\nu(G^{-1}(A))$. The converse relies on the duality of optimal transportation (see theorem 1.27 page 44 of Villani:2003 and GH:2006d for details). Note also that in the particular case where the spaces of the observed and latent variables are the same $ \mathcal{Y}=\mathcal{U}$ and $G$ is the identity function $G\left( u\right) =\left\{ u\right\} $, then ((ref)) defines the Total Variation metric between $P $ and $\nu $. When $\mathcal{Y}=\mathcal{U}$ and $G\left( u\right) =\left\{ y\in \mathcal{Y}:d\left( y,u\right) \leq \varepsilon \right\} $, the above duality boils down to a celebrated theorem due to Strassen (see section 11.6 of Dudley:2002). A closely related result was proven by Artstein in Artstein:83, Theorem 3.1, using an extension of the marriage lemma.
The optimal transportation of the specification problem at hand leads to an interpretation of the latter as a game between the Analyst and a malevolent Nature. This highlights connections between partial identification and robust decision making (in HS:2001) and ambiguity (in MMR:2006). As above, $P$ and $\nu$ are given. In the special case where we want to test whether the true functional relation between observable and unobservable variables is $\gamma_0$ (i.e. the complete specification problem), and where $P$ and $\nu$ are absolutely continuous with respect to Lebesgue measure, the optimal transportation formulation of the specification problem involves the minimization over the set of joint probability measures with marginals $P$ and $\nu$ of the integral $\int 1_\{y\ne\gamma_0(u)\}d\pi(y,u)$. The latter can be written as the minimax problem \[\min_\tau\max_V\int [1_{\{\tau(u)\ne\gamma_0(u)\}}-V(\tau(u))]d\nu(u)+\int V(y)dP(y).\]
This yields the interpretation as a zero-sum game between the Analyst and Nature, where the Analyst pays Nature the amount
$P$ and $\nu$ are fixed. The Analyst is asked to propose a plausible functional relation $y=\tau(u)$ between observed and latent variables, and Nature chooses $V$ in order to maximize transfer ((ref)) from the Analyst. This transfer can be decomposed into two terms. The first term $\int V(y)dP(y)-\int V(\tau(u))d\nu(u)$ is a punishment for guessing the wrong distribution: this term can be arbitrarily large unless $P=\nu\tau^{-1}$. The second term, $\int 1_{\{ \tau(u)\ne\gamma_0(u)\}}d\nu(u)$ is an incentive to guess $\tau$ close to the true functional relation $\gamma_0$ between $u$ and $y$.
The value of this game for Nature is equal to $T(P)=\inf \{\mathbb{P}(\tau(U)\ne\gamma_0(U)):\;U\sim\nu,\;\tau(U)\sim P\}$ and is independent of who moves first. This follows from the Monge-Kantorovitch duality. Indeed, if Nature moves first and plays $V$, the Analyst will choose $\tau$ to minimize $\int \left(1_{\{\tau(u)\ne\gamma_0(u)\}}-V(\tau(u))\right)d\nu(u)$. Denoting $V^\ast(u)=\inf_y\{1_{\{y\ne\gamma_0(u)\}}-V(y)\}$, the value of this game for Nature is \[ \sup_{V^\ast(u)+V(y)\leq1_{\{y\ne\gamma_0(u)\}}}\int V^\ast(u)d\nu(u)+\int V(y)dP(y).\] If, on the other hand, the Analyst moves first and plays $\tau$, then Nature will receive an arbitrarily large transfer if $P\ne\nu\tau^{-1}$, and a transfer of $\int 1_{\{\tau(u)\ne\gamma_0(u)\}}d\nu(u)$ independent of $V$ otherwise. The value of the game for Nature is therefore $\inf \{\mathbb{P}(\tau(U)\ne\gamma_0(U)):\;U\sim\nu,\;\tau(U)\sim P\}$. The Monge-Kantorovitch duality states precisely that the value when Nature plays first is equal to the value when Analyst plays first.
Finally, we have an interpretation of the set of observable distributions $P$ that are compatible with the model $(G,\nu)$ as the set of distributions $P$ such that the Analyst is willing to play the game, i.e. such that the value of the game is zero for some functional relationship $\gamma_0$ among the selections of $G$.
We now consider falsifiability of the incompletely specified model through a test of the null hypothesis that $P$ is compatible with $(G,\nu)$. Falsifying the model in this framework corresponds to the finding that a sample $(Y_1,\ldots,Y_n)$ of $n$ copies of $Y$ distributed according to the unknown true distribution $P$ was not generated as part of a sample $((Y_1,U_1),\ldots,(Y_n,U_n))$ distributed according to a fixed $\pi$ with marginal $\nu$ on ${\cal U}$ and satisfying the restrictions $Y\in G(U)$ almost surely. Using the results of the previous section, this can be expressed in the following equivalent ways.
Call $P_n$ the empirical distribution, defined by $P_n(A) = \sum_{i=1}^{n} 1_{\{Y_i\in A\}}/n$ for all $A$ measurable, and form the empirical analogues of the conditions above as
Note first that by the duality of optimal transportation, the empirical primal (EP) and the empirical dual (ED) are equal. In the case $\mathcal{Y}\subseteq\mathbb{R}^{d_y}$, GH:2006d propose a testing procedure based on the asymptotic treatment of the feasible statistic \[T_n=\sqrt{n}\sup_{A\in\mathcal{C}_n} [P_n(A)-\nu G^{-1}(A)],\hskip20pt\mbox{with }\;\mathcal{C}_n=\{(-\infty,Y_i],(Y_i,\infty):\;i=1,\ldots,n\}.\] More general families of test statistic for this problem can be derived from the following observation: consider the total variation metric defined by
for any two probability measures $\mu_1$ and $\mu_2$ on $(\mathcal{Y}, \cal{B}_\mathcal{Y})$, and \[d_{TV}\left( P,\mathcal{Q}\right) =\inf_{Q\in \mathcal{Q} }d_{TV}\left( P,Q\right) \] for a probability measure $P$ and a set of probability measures $\mathcal{Q}$. GH:2008 derive conditions under which the equalities
hold, so that the empirical dual is equal to the total variation distance between the empirical distribution $P_n$ and Core$(\nu G^{-1})$. Hence, (ED) yields a family of test statistics $d(P_n,\mbox{Core}(\nu G^{-1}))$, for the falsification of the model $(G,\nu)$, where $d$ satisfies $d(x,A)=0$ if $x\in A$ and $1$ otherwise.
Alternatively, a family of statistics can be derived from the empirical primal (EP) if the 0-1 cost is replaced by $d$ as above, yielding the statistics \[\inf_{\pi\in{\cal M}(P_n,\nu)} \int\!\!\!\int_{{\cal Y}\times{\cal U}} d(y,G(u)) d\pi(y,u)\] generalizing goodness-of-fit statistics based on the Wasserstein distance (see for instance BCMR:99).
In addition to producing families of test statistics, hence inference strategies, for partially identified structures, the optimal transportation formulation has clear computational advantages. First of all, efficient algorithms for the computation of the optimal transport map rely on both primal and dual formulations of the optimization problem. More specifically, in cases with discrete observable outcomes, the Monge-Kantorovitch optimal transportation problem reduces to its discrete counterpart, sometimes called the Hitchcock problem (see Hitchcock:41, Kantorovich:42 and Koopmans:49). This problem has a long history of applications in a vast array of fields, and hence spurred the development of many families of algorithms and implementations since FF:57. The optimal transportation formulation therefore allows the development of procedures for testing incomplete structures and estimating partially identified parameters that are vastly more efficient than existing ones (see for instance GH:2008a for the efficient computation of the the identified set in discrete games).
As before, we consider an economic model that governs the behaviour of a collection of economic variables $(Y,U)$. Here, $Y$ is a random element taking values in the Polish space ${\cal Y}$ (endowed with its Borel $\sigma$-algebra ${\cal B}_{\cal Y}$) and $U$ is a random vector taking values in ${\cal U}\subseteq\mathbb{R}^{d_u}$. $Y$ represents the subcollection of observable economic variables generated by the unknown distribution $P$, and $U$ represents the subcollection of unobservable economic variables generated by a distribution $\nu$. As before, the economic model provides a set of restrictions on the joint behaviour of observable and latent variables, i.e. a subset of ${\cal Y}\times{\cal U}\,$ represented by the measurable correspondence $G: {\cal U}\rightrightarrows{\cal Y}$. The distribution $\nu$ of the unobservable variables $U$ is now assumed to satisfy a set of moment conditions, namely
and we denote by $\mathcal{V}$ the set of distributions that satisfy ((ref)), and by $\mathcal{M}(P,\mathcal{V})$ the collection of Borel probability measures with one marginal fixed equal to $P$ and the other marginal belonging to the set $\mathcal{V}$. Note that a limit case of this framework, where an infinite collection of moment conditions uniquely determines the distribution of unobservable variables, i.e. when $\mathcal{V}$ is a singleton, we recover the parametric setup, with a classical optimal transportation formulation as in section (ref).
Finally we turn to an example of binary response, which we shall use as pilot examples for illustrative purposes.
We are now in the case where the economic model consists in the correspondence $G:\mathcal{U}\rightrightarrows \mathcal{Y}$ and a finite set of moment restrictions on the distribution $\nu$ of unobservables. Denote the model $(G,\mathcal{V})$. Again, the observables are fully characterized by their distribution $P$, which is unknown, but can be estimated from data. Consider now the restrictions imposed by the model on the joint distribution $\pi$ of the pair $(Y,U)$:
Again, a probability distribution $\pi$ that satisfies the restrictions above may or may not exist. If and only if it does, we say that the distribution $P$ of observable variables is compatible with the economic model $(G,\mathcal{V})$.
This hypothesis of compatibility has a similar optimization interpretation as in the case of parametric restrictions on unobservables. The distribution $P$ is compatible with the model $(G,\mathcal{V})$ if and only if $$\exists\pi\in{\cal M}(P,\mathcal{V}): \int_{\mathcal{Y}\times\mathcal{U}}1_{\{y\notin G(u)\}}d\pi(y,u)=0,$$ or equivalently
Although this optimization problem differs from the optimal transportation problem considered above, we shall see that inspection of the dual nevertheless provides a dimension reduction which will allow to devise strategies to falsify the model based on a sample of realizations of $Y$. However, before inspecting the dual, we need to show that the minimum in ((ref)) is actually attained, so that compatibility of observable distribution $P$ with the model $(G,\mathcal{V})$ is equivalent to
The following example shows that the infimum is not always attained.
It is clear from example (ref) that we need to make some form of assumption to avoid letting masses drift off to infinity. The theorem below gives formal conditions under which quasi-consistent alternatives are ruled out. It says essentially that the moment functions $m(u)$ need to be bounded.
Assumption (ref) is an assumption of uniform integrability. It is immediate to note that assumptions (ref) and (ref) are satisfied when the moment functions $m(u)$ are bounded and $\mathcal{U}$ is compact.
In example (ref), by Theorem 1.6 page 9 of RW:98, we know that assumption (ref) is satisfied when the moment functions $\varphi_j$, $j=1,\ldots,d_{\varphi}$ are lower semi-continuous.
We can now state the result:
The two dual formulations of this optimization problem are the following:
Since $u$ does not enter in the dual functional, the dual constraint can be rewritten as $f(y)=\inf_{u}\{1_{\{y\notin G(u)\}}-\lambda'm(u)\}$, so that the dual program can be rewritten \[T(P,\mathcal{V}):=\sup_{\lambda\in\mathbb{R}^{d_m}}\int_\mathcal{Y} \left(\inf_{u\in\mathcal{U}}[1_{\{y\notin G(u)\}}-\lambda'm(u)]\right)dP(y),\] which does not involve optimizing over an infinite dimensional space as the primal program did.
However, the dual formulation is useless if primal and dual are not equal. Note first that taking expectation in the dual constraint immediately yields (D)$ \leq $(P), which is the weak duality inequality. The converse inequality is shown below.
The Slater condition is an interior condition, i.e. it ensures there exists a feasible solution to the optimization problem in the interior of the constraints. Notice that when the $m_i$ are bounded, the Slater condition is always satisfied.
As described in the appendix, this result is ensured by the fact that there is no duality gap, i.e. that the statistic obtained by duality is indeed positive when the primal is.
We now consider falsifiability of the model with semiparametric constraints on unobservables through a test of the null hypothesis that $P$ is compatible with $(G,\mathcal{V})$. Falsifying the model in this framework corresponds to the finding that a sample $(Y_1,\ldots,Y_n)$ of $n$ copies of $Y$ distributed according to the unknown true distribution $P$ was not been generated as part of an sample $((Y_1,U_1),\ldots,(Y_n,U_n))$ distributed according to a fixed $\pi$ with $U$-marginal $\nu$ in $\mathcal{V}$ and satisfying the restrictions $Y\in G(U)$ almost surely. Using the results of the previous section, this can be expressed in the following equivalent ways.
Call $P_n$ the empirical distribution, defined by $P_n(A) = \sum_{i=1}^{n} 1_{Y_i\in A}/n$ for all $A$ measurable, and form the empirical analogues of the conditions above as
Note first that by the duality result of theorem (ref), the empirical primal (EP) and the empirical dual (ED) are equal. As in the parametric case, the cost function $c(y,u)=1_{\{y\notin G(u)\}}$ can be replaced by $c(y,u)=d(y,G(u))>0$ if $y\notin G(u)$ and equal to $0$ if $y\in G(u)$, to yield a family of numerically equivalent test statistics. Quantiles of their limiting distribution, or obtained from a bootstrap procedure can be used to form a test of compatibility, however, since (ED) involves two consecutive optimizations, a computationally more appealing procedure called dilation is proposed in GH:2006c. The idea is to control the size of the test nonparametrically so as to compute (ED) only once. For a test with level $1-\alpha$, compute a correspondence $J_n:\mathcal{Y}\rightrightarrows\mathcal{Y}$ such that there exist a pair of random vectors $Y$ and $Y^\ast$ with marginal distributions $P$ and $P_n$ respectively and satisfying $Y^\ast\in J_n(Y)$ with probability $1-\alpha$. The test then consists in rejecting compatibility of the unknown distribution $P$ of the observables with the model $(G,\mathcal{V})$ if and only if the known empirical distribution $P_n$ is not compatible with the model $(J_n\circ G,\mathcal{V})$, i.e. if \[\sup_{\lambda\in\mathbb{R}^{d_m}}\frac{1}{n}\sum_{i=1}^{n}\left(\inf_{u\in\mathcal{U}}[1_{\{Y_i\notin J_n\circ G(u)\}}-\lambda'm(u)]\right)\ne0.\]
We have proposed an optimal transportation formulation of the problem of testing compatibility of an incompletely specified economic model with the distribution of its observable components. In addition to relating this problem to a rich optimization literature, it allows the construction of computable test statistics and the application of efficient combinatorial optimization algorithms to the problem of inference in discrete games with multiple equilibria. A major application of tests of incomplete specifications is the construction of confidence regions for partially identified parameters. In this respect, the optimal transportation formulation proposed here allows the direct application of the methodology proposed in the seminal paper of CHT:2007 to general models with multiple equilibria.