Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
121,652 characters · 21 sections · 107 citation commands
Set identification in models with multiple equilibria
{ Keywords: multiple equilibria, optimal transportation, identification regions, core determining classes.
JEL subject classification: C13, C72}
The empirical study of game theoretic models is complicated by the presence of multiple equilibria. As noted in Jovanovic:89, the existence of multiple equilibria generally leads to a failure of identification of the structural parameters governing the model. BT:2006 and ABBP:2007 give an account of the various ways this identification issue was approached in the literature, where identification of structural parameters is achieved through equilibrium refinements, shape restrictions, informational assumptions or the specification of equilibrium selection mechanisms. An alternative approach is to eschew identification strategies and base inference purely on the identified features of the models with multiple equilibria, which are sets of values rather than a single value of the structural parameter vector. This approach is taken in the context of imperfectly competitive markets by ABJ:2003 and CT:2006. The inferential method they use, however, relies on a set of restrictions which is not guaranteed to exhaust all the restrictions embodied in the model, and hence leads to more conservative inference than could be achieved. This paper proposes a computationally feasible way of recovering the identified feature of a model with multiple equilibria, with specific applications to inference in participation games such as oligopoly entry models and group bargaining models.
We first note that the likelihood implied by a model with multiple equilibria can be represented by a non additive set function called a Choquet capacity. Seminal work on coherency conditions and nonadditive likelihoods can be found in Heckman:78, GLM:80, Manski:90 and HSC:97 among others. In cases where identification holds, DJ:94 and BMT:2002 propose likelihood-based estimation methods that are robust to the lack of coherency. Otherwise, the nonadditive likelihood can be refined to a likelihood represented by a probability measure if there exists a mechanism that picks outcomes among the admissible equilibria in the region of multiplicity. We give a formal definition of an equilibrium selection mechanism, and call such a mechanism compatible with the data if the likelihood of the model augmented with such a mechanism is equal to the probabilities observed in the data. The identified feature of the model therefore is the set of parameter values such that there exists an equilibrium selection mechanism compatible with the data. The main result of this paper is the equivalence between the latter condition and the actual probability of observed outcomes belonging to the core of the likelihood predicted by the model, where the core is a well known and well studied notion in economics since the word was coined in Gillies:53. This results allows the computation of the identified feature of models with multiple equilibria and a finite number of observable outcomes, as it reduces the problem to that of checking a finite number of moment inequalities.
The computational burden remains high in situations with a large number of observable outcomes, since the number of inequalities to be checked is equal to the number of subsets of the set of observable outcomes. When the set of observable outcomes is infinite, the problem remains infinite dimensional. GH:2006a and EGH:2008 include results pertaining to that case. To lift the remaining computational burden, we propose several alternative strategies, and discuss their relative merits.
First, in cases with only pure strategy equilibria, we show that the model likelihood is a submodular function (the set function analogue of a convex function), and that checking that the data distribution belongs to the core of the likelihood is equivalent to minimizing a submodular function, a well studied problem (analogue of minimizing a convex function) with efficient algorithms and easily available off the shelf implementations. Second, we show that if only pure strategy equilibria are considered, a special case of submodular function minimization applies, which relies on optimal transportation methods, and provides more efficient algorithms for the problem of constructing the identified set. Finally, we introduce the notion of core determining classes, which are suitably low cardinality classes of sets that are sufficient to characterize the core, and we provide some results to exhibit such core determining classes when the model satisfies some monotonicity properties. The method is illustrated on the family bargaining game of ES:2002 and the oligopoly entry game of BT:2006.
In cases where mixed strategies are included in the equilibrium concept, the fundamental work by BMM:2008 was the first to address the problem of constructing the identified set (whereas, in the case of pure strategies, to the best our knowledge, our work had been the first to address this problem, in GH:2006a). Our incremental contribution in the case of mixed strategy equilibria is two-fold. When the game satisfies a regularity condition defined in Shapley:71, we show that the identified set is still characterized by the core of the model likelihood, hence by a finite number of inequalities and a submodular optimization problem as before. In all other cases, we give a simple and efficient algorithm to construct the identified set based on convex optimization. This goes beyond the results in BMM that characterize the identified set only by a continuum of inequalities. Thus our results are complementary to the results of BMM.
The methodology proposed here applies to the empirical study of any normal form game, with both pure and mixed strategy equilibria. In economics, coordination games and participation games are the most natural area of application, including models of labour force participation (including BV:85), models of discrete choice with social interactions (including BD:2001), oligopoly entry models (including ABJ:2003, CT:2006), auction models (including BHR:2005), bargaining models (including ES:2002), network effects (including Sweeting:2004).
Beyond the identification issue of computing the identified set given the knowledge of the true distribution of observable variables, the inference issue of constructing confidence regions for structural parameters in models with multiple equilibria is taken up in GH:2006d and GH:2006c to complement the seminal contribution of CHT:2007. Related work on inference in partially identified models include MT:2002, IM:2004, BM:2007, RS:2006a, Rosen:2006, AS:2007, Canay:2007 among many others.
The remainder of the paper is organized as follows. Section (ref) describes the framework and general results in the case where only equilibria in pure strategies are considered, while section (ref) specializes and illustrates them on leading examples of participation games. Section (ref) describes three related methods to efficiently compute the identified feature of the model based on the characterization from section (ref) and discusses their relative merits. Section (ref) illustrates the three methods on an oligopoly entry game with two types of players, section (ref) shows how the results extend to the case where mixed strategy equilibria are also considered. Section (ref) also proposes extensive simulation results and section (ref) applies the methodology to the study of the determinants of long term care option choices for elderly parents in American families. The last section concludes, and proofs of the results are collected in an appendix.
The general framework is that of KR:50 generalized by Jovanovic:89. It applies primarily to the empirical analysis of normal form games, where only equilibria in pure strategies are considered. We defer the extension of our results to mixed strategies to section (ref).
We consider three types of economic variables. Outcome variables $Y$, exogenous explanatory variables $X$, and latent variables, or random shocks, $\epsilon$. Outcome variables and latent variables are assumed to belong to complete and separable metric spaces, so that both outcomes and latent variables could be discrete, continuous, they could be probability distributions or stochastic processes. The economic model consists in a set of restrictions on the joint behaviour of the variables listed above. These restrictions may be induced by assumptions of rationality of agents, and they generally depend on a set of unknown structural parameters $\theta$. Without loss of generality, the model may be formalized as a measurable correspondence (defined in assumption (ref) below) between the latent variables $\epsilon$ and the outcome variables $Y$ indexed by the exogenous variables $X$ and the vector of parameters $\theta$. We call this correspondence $G$, and write $Y\in G(\epsilon\vert X;\theta)$ to indicate admissible values of $Y$ given values of $\epsilon$, $X$ and $\theta$. The econometrician will be assumed to have access to a sample of independent and identically distributed vectors $(Y,X)$, and the problem considered is that of estimating the vector of parameters $\theta$. The latent variables $\epsilon$ is supposed to be distributed according to a parametric distribution $\nu(.\vert X;\theta)$, where the indexing is meant to indicate that the unknown parameters that enter in the distribution of latent variables are contained in the vector $\theta$ of parameters to the estimated. We collect these assumptions next.
To conduct inference on the parameter vector $\theta$, one first needs to determine the identified features of the model. Because the correspondence $G$ may be multi-valued due to the presence of multiple equilibria, the outcomes may not be uniquely determined by the latent variable. In such cases, the likelihood of an outcome falling in the subset $A$ of $\mathcal{Y}$ predicted by the model is $\mathcal{L}(A\vert X;\theta)=\nu(G^{-1}(A\vert X;\theta)\vert X;\theta)$. Because of multiple equilibria, this likelihood may sum to more than one, as we may have $A\cap B=\varnothing$, and yet $G^{-1}(A\vert X;\theta)\cap G^{-1}(B\vert X;\theta)\ne\varnothing$, so that we may have $\mathcal{L}(A\cup B\vert X;\theta)<\mathcal{L}(A\vert X;\theta)+\mathcal{L} (B\vert X;\theta)$. The set function $A\mapsto\mathcal{L}(A\vert X;\theta)=\nu(G^{-1}(A\vert X;\theta)\vert X;\theta)$ is generally not additive, and is called a Choquet capacity (see Choquet:53). This non additivity of the model likelihood is referred to as “lack of coherency” in GLM:80 and Tamer:2003.
As discussed in Jovanovic:89, BT:2006 and CT:2006, the model with multiple equilibria can be completed with an equilibrium selection mechanism. Following Jovanovic:89 and BT:2006 (See for instance the formulation (2.20) page 66 of BT:2006), we define an equilibrium selection mechanism as a conditional distribution $\pi_{Y\vert\epsilon,X}$ over equilibrium outcomes $Y$ in the regions of multiplicity. By construction, an equilibrium selection is allowed to depend on the latent variables $\epsilon$ even after conditioning on $X$. This is summarized in the following definition.
The identified feature of the model is the smallest set of parameters that cannot be rejected by the data. Hence, it is the set of parameters for which one can find an equilibrium selection mechanism which completes the model and equates probabilities of outcomes predicted by the model with the probabilities obtained from the data.
Hence the identified set is the set of parameters $\theta$ such that there exists an equilibrium selection mechanism compatible with the data.
The definition above is not an operational definition, in the sense that it does not alow the computation of the identified set based on the knowledge of the probabilities in the data because the conditional distribution $\pi$ is an infinite dimensional nuisance parameter. We now set out to show how to reduce the dimensionality of the problem. Our equivalent formulation of the identified set is based on an appeal to the notion of core of the Choquet capacity introduced in definition (ref).
In cooperative game theory, a Choquet capacity on a set $\mathcal{Y}$ is interpreted as a game, where $\mathcal{Y}$ is the set of players and $\mathcal{L}$ is the utility value or worth of coalition $A\subseteq\mathcal{Y}$ and the core of the game $\mathcal{L}$ is the collection of allocations that cannot be improved upon by any coalition of players (see Moulin:95).
The result we propose next\footnote{This result appeared as equivalence between (ii') and (iv') in theorem 1' of the first version circulated GH:2006a.} shows the equivalence between the existence of a compatible equilibrium selection mechanism and the fact that the true distribution of the data belongs to the core of the Choquet capacity that characterizes the likelihood predicted by the model (which we shall call core of the likelihood predicted by the model).
The first thing to note from this theorem is that the problem of computing the identified set has been transformed into a finite dimensional problem in the special case where ${\mathcal{Y}}$ is a finite set (or equivalently, the support of the distribution $P$ of observable outcomes has finite cardinality). Indeed, in the latter case, the problem of computing the identified set is reduced to the problem of computing a finite number of inequalities, i.e. $P(A\vert X)\leq\mathcal{L}(A\vert X;\theta)=\nu(\epsilon:\;G(\epsilon\vert X;\theta)\cap A\ne\varnothing\vert X;\theta)$ for each subset $A$ of $\mathcal{Y}$. However, in cases where the cardinality of $\mathcal{Y}$ is large, then the number of inequalities to be checked is $2^{\mathrm{Card}(\mathcal{Y})}-2$, and the computational burden is only partially lifted, and the second section of the paper is devoted to efficient methods of computation of the identified set based on the characterization of theorem (ref). First we turn to the specialization of our results and concepts to some leading examples, and illustrate them with a simplified version of a family bargaining game studied in the literature.
A leading special example of the framework above is that of empirical models of oligopoly entry, proposed in BR:90 and Berry:92, and considered in the framework of partial identification by Tamer:2003, ABJ:2003, BT:2006, CT:2006 and PPHI:2004 among others. In this setup, economic agents are firms who decide whether of not to enter a market. Markets are indexed by $m$, $m=1,\ldots,M$ and firms that could potentially enter the market are indexed by $i$, $i=1,\ldots,J$. $Y_{im}$ is firm $i$'s strategy in market $m$, and it is equal to $1$ if firm $i$ enters market $m$, and zero otherwise. $Y_m$ denotes the vector $(Y_{1m},\ldots,Y_{Jm})^t$ of strategies of all the firms. In standard notation, $Y_{-im}$ denotes the vector of strategies of all firms except firm $i$. In models of oligopoly entry, the profit $\pi_{im}$ of firm $i$ in market $m$ is allowed to depend on strategies $Y_{-im}$ of other firms, as well as on a set of profit shifters $X_{im}$ that are observed by all firms and the econometrician, a profit shifter $\epsilon_{im}$ that is observed by all the firms but not by the econometrician, and a vector of unknown structural parameters, so that it can be written $\pi_{im}(Y_m,X_{im},\epsilon_{im};\theta)=\pi_{im}(Y_{im},Y_{-im},X_{im},\epsilon_{im};\theta)$. If, for instance, firms are assumed to play Nash equilibria in pure strategies in market $m$, their strategies $Y_{im}$ are such that they yield higher profits than $1-Y_{im}$ given other firms' strategies $Y_{-im}$. So the restrictions induced on the strategies and latent profit shifters are $\pi_{im}(Y_{im},Y_{-im},X_{im},\epsilon_{im};\theta)\geq\pi_{im}(1-Y_{im},Y_{-im},X_{im},\epsilon_{im};\theta)$ for all $i=1,\ldots,J$. Hence the model can be written $Y_m\in G(\epsilon_m\vert X_m;\theta)$, where $X_m$ denotes the matrix of observed profit shifters for firms $i=1,\ldots,J$, $\epsilon_m$ denotes the vector of latent profit shifters for firms $i=1,\ldots,J$, and the correspondence $G$ is defined by $G(\epsilon\vert X;\theta)=\{Y:\;\pi_{i}(Y_{i},Y_{-i},X_{i},\epsilon_{i};\theta) \geq\pi_{i}(1-Y_{i},Y_{-i},X_{i},\epsilon_{i};\theta);\mathrm{ all } \;i=1,\ldots,J\}$, where the index $m$ is dropped when considering a generic market.
For illustration purposes, we consider a simplified version of the bargaining model of decision regarding the long term care of an elderly parent in ES:2002.
Consider a family with two children. The issue is which of the children will become the primary care giver of an elderly and disabled parent or whether the parent moves to a nursing home.
The payoff to family member $i$, $i=1,2$ is represented as the sum of three terms. The first term $V_{ij}$ represents the value to child $i$ of care option $j$, where $j>0$ means child $j$ becomes the primary care giver and $j=0$ means the parent is moved to a nursing home. The matrix $V=(V_{ij})_{ij}$ is known to both children. We suppose it takes the form \[V=\left(
\right)\] where $\theta$ is a nonnegative parameter, unknown to the analyst.
Both children simultaneously decide whether or not to take part in the long term care decision. Suppose $M$ is the set of children who participate. The option chosen is option $j$ which maximizes the sum $\sum_{i\in M}V_{ij}$ among the available options (only participating children can become primary care givers). It is assumed that participants abide with the decision and that benefits are then shared equally among children participating in the decision through a monetary transfer $s_i$, which is the second term in the children's payoff.
The third term $\epsilon_i$ in the payoff is a random benefit from participation, which is $0$ for children who decide not to participate and distributed according to $\nu(.\vert \theta)$ for children who participate. All children observe the realizations of $\epsilon$, whereas the analyst only knows its distribution.
The matrix of payoffs of the participation game is given in table (ref).
We derive the equilibrium correspondence, restricting the analysis to pure strategy Nash equilibria for now. We shall later add equilibria in mixed strategies to illustrate results of section (ref).
The equilibrium correspondence $G(\epsilon\vert \theta)$ is represented in figure (ref)(a) as a function of $\epsilon$. The dotted lines represents the axes in the $\epsilon$ space, and the full lines represent the frontiers of the regions defining the correspondence $G$. The shaded area is the area of multiplicity, where $G(\epsilon\vert X;\theta)$ contains two values $(0,1)$ and $(1,0)$.
The model thus described is incomplete in the sense that more information is required in the regions of multiplicity to determine which equilibrium will be selected. Without knowledge of such an equilibrium selection mechanism, the likelihood predicted by the model can be written as follows. As before $\mathcal{Y}$ denotes the set of possible outcomes. The likelihood of observation $y$ is $\mathcal{L}(y\vert \theta)=\nu(\epsilon:\;y\in G(\epsilon\vert \theta)\vert \theta) =\nu(G^{-1}(y\vert \theta)\vert \theta)$, for all $y\in\mathcal{Y}$ and $\sum_{y\in\mathcal{Y}}\mathcal{L}(y\vert \theta)\geq1$, where the inequality may be strict if there are regions of multiplicity.
The model can be completed by adding an equilibrium selection mechanism which will pick out a single equilibrium for each value of the latent variable $\epsilon$ in the region of multiplicity. As formally defined in the previous section, an equilibrium selection mechanism is a conditional probability $\pi(.\vert\epsilon,X)$ with support included in $G(\epsilon\vert X;\theta)$. It is compatible with the data if the probabilities it predicts are equal to the true probabilities of the observable variables.
As noted above, since the model contains no prior information about which outcome is selected in the regions of multiplicity, the identified set $\Theta_I$ for the parameter vector $\theta$ is the set of parameters for which one can find an equilibrium selection mechanism which completes the model and equates probabilities of outcomes predicted by the model with true outcome frequencies. The definition of the identification region using a semiparametric likelihood representation, where the equilibrium selection mechanism is included as the infinite dimensional nuisance parameter $\pi$ is impractical, so we use theorem (ref) to provide an operational method to compute $\Theta_I$. The existence of a compatible selection mechanism is equivalent to the fact that the true distribution $P$ of observed outcomes lies in the core of the Choquet capacity $\mathcal{L}=\nu G^{-1}$ defined by the model. Hence, we have $\Theta_I=\{\theta\in\Theta:\;\left( \forall A\in2^\mathcal{Y};\;P(A\vert X)\leq\mathcal{L}(A\vert X;\theta);\right) \;X\,\mathrm{ a.s.}\}$ where $2^S$ denotes the set of all subsets of a set $S$.
We now describe three approaches to the effective computation of the identified set based on our characterization of theorem (ref).
The first approach, described in section (ref) is based on the observation that the likelihood is a submodular set function (the set function analogue of a convex function), and that checking the condition of theorem (ref) is equivalent to minimizing a submodular function, which is a well studied problem (analogous to minimizing a convex function) and easily and efficiently implementable. It extends readily to the more general case when equilibria in mixed strategies are also allowed, as will be discussed in the computational part of section (ref).
The second approach, described in section (ref), is a special case of the first, which relies on the highly efficient algorithms (and easily available packaged implementations) for optimal transportation problems. As a combinatorial optimization method, it is a well known special case of the second approach, and it is computationally more efficient, and therefore recommended to practitioners who restrict attention to equilibria in pure strategies.
The third approach, based on the notion of core determining classes of sets and providing a dramatic reduction in the computational complexity under specific assumptions on the game under study, is described in section (ref), where it is shown how the exponential problem of checking $2^{|\mathcal{Y}|}-2$ inequalities in theorem (ref) can be replaced by the problem polynomial problem of checking $2|\mathcal{Y}|-2$ inequalities instead.
The first proposal to deal with the complexity of the problem of checking inequalities in theorem (ref) is a method of general validity based on the minimization of a submodular function, the discrete equivalent of a convex function, which is a well known problem in combinatorial optimization, and for which efficient algorithms and easily available off the shelf implementations exist.
Submodularity for set functions is the analogue of convexity, and the problem of minimizing a submodular function is well studied (see for instance Topkis:98 chapter 2). A complete account of the theory can be found in Fujishige:2005 and off the shelf matlab implementation of the most efficient known algorithms can be found in Andreas Krause's SFO toolbox.
We now show that checking inequalities involved in the characterization of the identified set in theorem (ref) is indeed equivalent to the minimization of a submodular function. Theorem (ref) shows that the identified set is the set of values of $\theta$ such that $X$-almost surely, we have the domination $\forall A\subseteq\mathcal{Y}$, $P(A\vert X)\leq\mathcal{L}(A\vert X;\theta)$, or equivalently $\min_{A\subseteq\mathcal{Y}}\left(\mathcal{L}(A\vert X;\theta)-P(A\vert X)\right)\geq0$. We first note that the function in the minimization above is indeed submodular.
The most efficient generic way to check that a convex function is everywhere non negative is to minimize it and to verify that the minimum is non negative. In the same way, we propose to check inequalities of theorem (ref) by minimizing the submodular function defined for all $A\subseteq\mathcal{Y}$ by $A\mapsto\mathcal{L}(A\vert X;\theta)-P(A\vert X)$, and checking that the minimum is indeed non negative. Of course, we can speed up the algorithm by stopping short of the minimum as soon as a negative value is found.
More details of the procedure are given in section (ref), as this method is one of the recommended methods of construction of the identified set in the case where equilibria in mixed strategies are considered. Below is a description of a special case of submodular optimization, which is more efficient, and applies to the case where only equilibria in pure strategies are considered.
In the special case with only pure strategy equilibria that we are considering until section (ref), the model likelihood $\mathcal{L}$ is a very special case of submodular function, since it is derived as the distribution function of a random set, or random correspondence $\mathcal{L}(A\vert X;\theta)=\nu(\epsilon:\;G(\epsilon\vert X;\theta)\cap A\ne\varnothing\vert X;\theta)$. It is that property that allows the following refinement of the method, and the simpler and more efficient technique proposed below. When mixed equilibria are considered, this improvement in efficiency is no longer available, precisely because in general the model likelihood is no longer the distribution of a random set.
To describe the method, we need the following notations and definitions. The terminology used is intended to be self-explanatory. Otherwise, the reader is referred to PS:98 for standard definitions in graph theory. Call $\mathcal{U}^\ast$ the set of predicted combinations of equilibria, formally $\mathcal{U}^\ast=\{G(\epsilon\vert X;\theta);\epsilon\in\mathcal{U}\}$ (we suppress reference to the dependence of $\mathcal{U}^\ast$ on $\theta$ for notational convenience). Hence $\mathcal{U}^\ast$ contains subsets of $\mathcal{Y}$, but is typically of much lower cardinality than the power set $2^\mathcal{Y}$. Further consider the bi-partite graph $\mathcal{G}(\theta,X)$ in $\mathcal{Y}\times\mathcal{U}^\ast$. The latter is defined as the set of pairs $(y,u)\in\mathcal{Y}\times\mathcal{U}^\ast$ such that $y\in u$. Each vertex $y$ in $\mathcal{Y}$ has weight $\mathbb{P}(Y=y\vert X)$ and each vertex $u\in\mathcal{U}^\ast$ has weight $\mathbb{P}(G(\epsilon\vert X;\theta)=u\vert X)$. The graph contains edges $(y,u)$ linking an element $y\in\mathcal{Y}$ to an element $u\in\mathcal{U}^\ast$ if the former is an element of the latter (i.e. $y\in u$). Finally, call $P(y\vert X)=\mathbb{P}(Y=y\vert X)$ the actual probabilities of observable variables $y\in\mathcal{Y}$, and call $Q(.\vert X;\theta)$ the probabilities $Q(u\vert X;\theta)=\mathbb{P}(G(\epsilon\vert X;\theta)=u\vert X)$. If we consider $G$ (keeping the same notation for simplicity) as a correspondence from $\mathcal{U}^\ast$ to $\mathcal{Y}$, then, formally $G(u)=u$, and we have shown in theorem (ref) that $\theta$ belongs to the identified set if and only if for any subset $A$ of $\mathcal{Y}$, $P(A\vert X)\leq Q(G^{-1}(A)\vert X;\theta)$. GH:2006d show that it is equivalent to the existence of a joint probability $\pi$ on $\mathcal{G}(\theta,X)$ with marginal distributions $P(.\vert X)$ and $Q(.\vert X;\theta)$. This is summarized in the following proposition\footnote{ This result is a special case of equivalence between (ii') and (iii') in theorem 1' of the previous version circulated GH:2006a.}.
Note that one implication in theorem (ref) is very easy to prove. Call $U$ the random element with distribution $Q$. If a joint probability exists with the required properties, then ${Y\in A}\Rightarrow U\in G^{-1}(A)$, so that $1_{\{Y\in A\}}\leq1_{\{U\in G^{-1}(A)\}}$, $\pi$-almost surely. Taking expectation, we have $\mathbb{E}_\pi(1_{\{Y\in A\}})\leq \mathbb{E}_\pi(1_{\{U\in G^{-1}(A)\}})$, or equivalently $P(A\vert X)\leq Q(G^{-1}(A))\vert X;\theta)$. The converse is much more involved and relies on optimal transportation theory. A similar result is proved in theorem 3 of Artstein:83, based on an extension of the marriage lemma.
We illustrate this requirement on our example (ref).
Since we have now formulated the problem of computing the identified set as a problem involving the existence of a probability measure with given marginal distributions, we can appeal to efficient computational methods in the optimal transportation literature. The problem of sending $p_y$, $y\in\mathcal{Y}$ units of a good from sources $y\in\mathcal{Y}$ to $q_u$, $u\in\mathcal{U}^\ast$ units in terminals $u\in\mathcal{U}^\ast$ at minimum cost of transportation, where costs are attached to each pair $(y,u)\in\mathcal{Y}\times\mathcal{U}^\ast$ is called an optimal transportation problem, and first appears in the economics literature with Koopmans:49. Our problem can be reduced to an optimal transportation problem with 0-1 cost of transportation, where a pair $(y,u)$ is assigned cost zero if it belongs to $\mathcal{G}(X;\theta)$, and 1 otherwise, and there exists a joint law on $\mathcal{G}(X;\theta)$ with marginals $P$ and $Q$ (in other words, $\theta$ is in the identified set) if and only if there is a zero cost solution to the optimization problem thus defined.
As explained in FF:57 (see also PS:98 section 7.4 page 143), there is an equivalent dual formulation of this minimum cost of transportation problem as a maximum flow problem described in figure (ref)(b). The edges in the graph with zero cost in the minimum cost of transportation problem have infinite carrying capacity (not to be confused with Choquet capacity) in the dual maximal flow problem. Hence efficient maximum flow programs (such as maxflow.m in the Matlab BGL library) can be applied directly to the network described in figure (ref)(b): Mass flows from the source to the sink through the network. The number on each edge is the maximum mass that can flow through that edge. The source sends $p_y$ mass exactly to each node corresponding to elements of $\mathcal{Y}$, and the sink receives $q_u$ mass from each node corresponding to an element of $\mathcal{U}^\ast$. Between edges in $\mathcal{Y}$ and $\mathcal{U}^\ast$, mass can flow freely through pairs $(y,u)$ such that $y\in u$ (full lines with infinite carrying capacity), and not at all through pairs $(y,u)$ such that $y\notin u$. The parameter value $\theta$ is in the identified set if and only if the maximum flow program returns a maximum flow of exactly $1$ (note that the network carrying capacities depend on $\theta$ through the probabilities $q_u$). The full procedure is described in detail on the oligopoly entry example of section (ref).
As we have seen in the first section, theorem (ref) allows to reduce the problem of computing the identified set to that of checking a set of inequalities. However, the computational burden is only partially lifted, as the number of inequalities to check can be very large if the cardinality of the outcome space is large. In this section, we shall analyze ways of reducing this remaining computational burden, by eliminating redundant inequalities in the computation of the identified set. This is formalized with the concept of core determining classes, which was first introduced in section 3.2.2 page 27 of GH:2006a.
A core determining class $\mathcal{A}$ allows the elimination of all the inequalities $Q(A)\leq\mathcal{L}(A)$ for $A\notin\mathcal{A}$ when checking whether a probability $Q$ belongs to the core of a Choquet capacity $\mathcal{L}$. Since the likelihood predicted by the model $Y\in G(\epsilon\vert X;\theta)$ was characterized by the Choquet capacity $A\mapsto\mathcal{L}(A\vert X;\theta)=\nu(\epsilon:G(\epsilon\vert X;\theta)\cap A\ne\varnothing\vert X;\theta)$, a core determining class of sets is sufficient to characterize the identified region $\Theta_I$ as summarized in the following proposition.
The challenge therefore becomes that of finding a core determining class $\mathcal{A}$ in order to reduce the number of inequalities to be checked to the cardinality of $\mathcal{A}$. We first consider the case of our example (ref), before turning to a criterion that will prove useful in exhibiting core determining classes in many important cases.
We now show how to identify core determining classes more generally to avoid painstaking case-by-case elimination of redundant inequalities. To that end, we give general conditions under which one can find a core determining class of low cardinality. Recall that a subset $A$ of an ordered set (with ordering $\preceq$) is said to be connected if any $a$ such that $\inf A\preceq a\preceq \sup A$ belongs to $A$.
We illustrate this assumption on our family bargaining game before stating the theorem and applying it to the more sophisticated case of an oligopoly entry game with two types of players presented in BT:2006.
We are now in a position to state the theorem\footnote{ This result is a reformulation of theorem 3d of the previous version circulated GH:2006a.}, which is the main tool in the construction of core determining classes, and hence in the computation of the identified set.
Theorem (ref) allows to reduce the cardinality of the power set $2^\mathcal{Y}$ to twice the cardinality of $\mathcal{Y}$ minus $2$ (since the inequality needn't be checked on the whole set $\mathcal{Y}$), as we illustrate in our example (ref).
We now turn to a more substantive illustration of our methods to compute the identified set, first, to show the operational usefulness of corollary (ref), and second, to illustrate the power of the combinatorial approach. To do so, we consider the oligopoly entry game with two types of players presented in appendix A of BT:2006. The profit function of type 1 firms depends on the total number of firms in the market, but not on the type of those firms, whereas profits of type 2 firms depend both on the number and on the type of firms present in the market. The latent variable is the fixed cost $f_1$ for firms of type 1 and $f_2$ for firms of type 2. $(f_1,f_2)$ is uniformly distributed over $[0,1]^2$. The model is simplified by assuming linearity of profits in firm number as follows.
with $\alpha_1,\beta_1,\beta_2$ strictly negative and $\beta_2>\beta_1$ to fix ideas (profit of type 2 firms will decrease by a larger amount if a type 1 firm enters the market than if a type 2 firm does). The set of observable outcomes is $\mathcal{Y}=\{(i,j):\;i,j=0,1,2\}$, where $i$ denotes the number of type 1 firms and $j$ the number of type 2 firms present in the market. $\mathcal{Y}$ can be ordered lexicographically, where the number of firms present in the market is considered first, and then the identity of firms (type 1 dominating type 2)\footnote{The order could be rationalized by total profit in the industry, but it is not necessary for the construction of a core determining class nor the computation of the identified set.}. Hence $(0,0)\precsim_\mathcal{Y}$ $(0,1)\precsim_\mathcal{Y}$ $(1,0)\precsim_\mathcal{Y}$ $(0,2)\precsim_\mathcal{Y}$ $(1,1)\precsim_\mathcal{Y}$ $(2,0)\precsim_\mathcal{Y}$ $(1,2)\precsim_\mathcal{Y}$ $(2,1)\precsim_\mathcal{Y}$ $(2,2)$. The model correspondence is represented in figure (ref), which is taken from BT:2006.
The implementation of the submodular optimization method is extremely simple. Once the values of the likelihood have been derived (analytically in the present case, but more often by simulation following the method proposed in BHR:2005), all that remains to do to check whether a value $\theta$ belongs to the identified set, is to minimize the function defined for all $A\subseteq\mathcal{Y}$ by $A\mapsto\mathcal{L}(A\vert X;\theta)-P(A\vert X)$ using for instance the Matlab SFO toolbox routine $SFO-min-norm-point.m$ implementing Fujishige's algorithm (see page 293 of Fujishige:2005).
Consider now the linear programming strategy for computing the identified set. The bipartite graph corresponding to this example is represented in figure (ref).
As shown in theorem (ref), a value of the parameter vector is in the identified set if and only if there exists a zero cost transportation plan for the transfer of masses $p_y$ on the elements of $\mathcal{Y}$ to masses $q_u$ on the elements of $\mathcal{U}^\ast$. A transportation plan is a set of nonnegative numbers attached to all pairs $(y,u)\in\mathcal{Y}\times\mathcal{U}^\ast$ (which represents the amount of mass from $y$ that is transferred to $u$ via the edge $(y,u)$). In our application, the transportation cost from $y$ to $u$ is zero if $y$ and $u$ are connected by an edge in the graph of figure (ref), and 1 otherwise. If the algorithm returns a zero cost transportation plan, it means that mass is transferred through edges of the graph only, and for instance the pair $\left((1,1),\{(0,2),(1,1),(2,0)\}\right)$ is assigned a non negative number (i.e. some mass is transported there), but the pair $((1,1),\{(2,0),(0,2)\})$ is assigned zero (i.e. no mass is transported there). The existence of a zero cost transportation plan is equivalent to the existence of a joint distribution on $\mathcal{Y}\times\mathcal{U}^\ast$ which is supported on the graph of figure (ref) and has the correct marginal distributions, hence, it is equivalent to the fact that $\theta$ is in the identified set, as we showed in theorem (ref).
The minimum cost transportation problem is equivalent to the dual maximum flow problem, as described in the previous section. Mass flows through the network with $25$ nodes, which include the source, the $9$ elements of $\mathcal{Y}$, the $14$ elements of $\mathcal{U}^\ast$ and the sink (mass flows in the direction $\mathrm{Source}\rightarrow\mathcal{Y}\rightarrow\mathcal{U}^\ast\rightarrow\mathrm{Sink}$). A network is characterized by its adjacency matrix, which gives all the links between nodes with their capacity. In the case of interest here, the adjacency matrix is given in table (ref). Maximum flow programs take this adjacency matrix as an input, and return the maximum flow through the network it characterizes. This maximum flow cannot be larger than $\sum_{y\in\mathcal{Y}}p_y=\sum_{u\in\mathcal{U}^\ast}q_u=1$, and it is equal to $1$ if and only if $\theta$ is in the identified set.
We now illustrate the usefulness of corollary (ref) for the determination of a core determining class in this example. Figure (ref) graphs the orderings that satisfy assumption (ref) up to the fact that the set of equilibria is not always connected. Indeed, in the ordering of outcomes, $(1,1)$ comes between $(0,2)$ and $(2,0)$, or more precisely, $(0,2)\precsim_\mathcal{Y}(1,1) \precsim_\mathcal{Y}(2,0)$. Now $(1,1)$ is not an equilibrium when $\alpha_0+3\alpha_1< f_1\leq\alpha_0+2\alpha_1$ and $\beta_0+\beta_1+\beta_2<f_2\leq\beta_0+2\beta_2$, so the set of equilibria $\{(0,2),(2,0)\}$ is disconnected in that case.
However, since $(1,1)$ is observed only when $\epsilon\in\{(0,2),(1,1),(2,0)\}$, the mass $p_{11}$ can be removed from $q_{02,11,20}$. Indeed, $(1,1)$ is isolated in the sense that all its mass is necessarily transferred to $\{(0,2),(1,1),(2,0)\}$, which is the only predicted combination of equilibria in which it appears. Hence, after checking that $p_{11}\leq q_{02,11,20}$, we can remove $p_{11}$ and $q_{02,11,20}$ and replace $q_{02,20}$ by $q_{02,20}+q_{02,11,20}-p_{11}$. Then, the monotonicity and connectedness conditions hold, and theorem (ref) can be applied directly to $\mathcal{Y}\backslash\{(1,1)\}$ and $\mathcal{U}^\ast$, yielding the class $\mathcal{A}$ $=( \{(0,0)\},$ $\{(0,0),$ $(0,1)\},$ $\{(0,0),$ $(0,1),$ $(1,0)\},$ $\{(0,0),$ $(0,1),$ $(1,0),$ $(0,2)\},$ $\{(0,0),$ $(0,1),$ $(1,0),$ $(0,2),$ $(2,0)\},$ $\{(0,0),$ $(0,1),$ $(1,0),$ $(0,2),$ $(2,0),$ $(1,2)\},$ $\{(0,0),$ $(0,1),$ $(1,0),$ $(0,2),$ $(2,0),$ $(1,2),$ $(2,1)\},$ $\{(0,1),$ $(1,0),$ $(0,2),$ $(2,0),$ $(1,2),$ $(2,1),$ $(2,2)\},$ $\{(1,0),$ $(0,2),$ $(2,0),$ $(1,2),$ $(2,1),$ $(2,2)\},$ $\{(0,2),$ $(2,0),$ $(1,2),$ $(2,1),$ $(2,2)\},$ $\{(2,0),$ $(1,2),$ $(2,1),$ $(2,2)\},$ $\{(1,2),$ $(2,1),$ $(2,2)\},$ $\{(2,1),$ $(2,2)\},$ $\{(2,2)\})$. Note that its cardinality is $2\times7=14$, as opposed to the cardinality of the power set of $\mathcal{Y}$ which is $2^9=512$.
As an illustration of the procedure, we compute the identified set for the two-type oligopoly model with the following distributional hypotheses and normalization restrictions. The fixed cost vector $(f_1,f_2)$ is assumed to be uniformly distributed on $[0,1]^2$. $\alpha_0$ and $\beta_0$ are set equal to $1$. As previously noted, we assume that monopoly profits are larger than oligopoly profits, and that a type two firm's profit decreases more if a type one firm enters than a type two firm, hence $0>\alpha_0$ and $0>\beta_2>\beta_1$. We can therefore calculate the probabilities of each combination of equilibria $u\in\mathcal{U}^\ast$. These probabilities are computed in a Matlab program file available on request.
The values of the model likelihood are derived from these probabilities, and the function $\mathcal{L}(.\vert X;\theta)-P(.\vert X)$ is then minimized over all subsets of $\mathcal{Y}$ using the Matlab routine SFO-min-norm-point.m from the SFO toolbox. An idea of the computational efficiency of this procedure is given by the fact that an order of $10^3$ values of the parameter can be tested in one second on a standard laptop computer.
To implement the minimum cost of transportation/maximum flow method, the probabilities of each value of the equilibrium correspondence are entered together with the true frequencies of observable outcomes into the adjacency matrix of table (ref) and the Matlab routine maxflow.m of the BGL library returns a flow of $1$ if the value of $\theta=(\alpha_1,\beta_1,\beta_2)'$ (used to compute the predicted probabilities) belongs to the identified set, and a flow strictly smaller than $1$ if it doesn't. To give an idea of the efficiency of the method, we can test $10^5$ values of $\theta=(\alpha_1,\beta_1,\beta_2)'$ in less than a second on a standard portable computer. Hence this method is faster than the general submodular minimization method, but only applies to pure strategy equilibria.
Finally, the additional information yielded by the existence of a core determining class for this problem allows us to test a value of the parameter by a constrained submodular optimization, where the search is limited to the sets in the core determining class.
For illustration purposes, we compute the identified set for a given choice of the observable frequencies, namely $(p_{00},p_{01},p_{10},p_{02},p_{11},p_{20},p_{12},p_{21},p_{22})= (0.1,0.15,0.15,0.1,0,0.5,0,0,0)$, and compare it to the set obtained by imposing the inequality restrictions on the Singleton class only. The latter corresponds to the set of values of the parameters such that $p_{00}\leq q_{00}$, $p_{01}\leq q_{01}+q_{01,10}$, $p_{10}\leq q_{01,10}+q_{10}+q_{02,10}$, $p_{02}\leq q_{02,10}+q_{02}+q_{02,20}+q_{02,11,20}$, $p_{11}\leq q_{02,11,20}$, $p_{20}\leq q_{02,20}+q_{02,11,20}+q_{20}+q_{20,12}$, $p_{12}\leq q_{20,12}+q_{12}+q_{12,21}$, $p_{21}\leq q_{12,21}+q_{21}$ and $p_{22}\leq q_{22}$. It turns out the values of $\alpha_1$ and $\beta_2$ are identified (under these specific values for the true probabilities of observable variables, which were chosen for the simulation purposes from the parameter values), and all values of $\beta_1<\beta_2$ are compatible with the given frequencies. The set defined by the Singleton class restrictions, however, is much larger, as shown by its projection on the $(\beta_2,\alpha_1)$ space in figure (ref). There are also many values of the observed frequencies, for which the identified set is empty, so that the model is rejected, but the set defined by the Singleton class restrictions is non-empty, so that it fails to reject the model.
We now show how our results extend to the case where mixed strategy equilibria are allowed. As before, we consider a game parameterized by the deterministic parameter vector $\theta$, a vector of covariates $X$ and unobserved heterogeneity parameter $\epsilon$ with distribution $\nu(.\vert X;\theta)$. The observable outcomes of the game are equilibrium actions profiles $Y$ whose realizations belong to the finite set $\mathcal{Y}$. Call $P(.\vert X)$ the true distribution of observable outcomes $Y$ conditionally on $X$. The equilibrium correspondence $G(\varepsilon\vert X;\theta)$ is now the set of Nash equilibria of the game in mixed strategies (including pure strategies as degenerate mixed strategies). Call $\sigma$ the equilibrium strategy profile that is actually selected within the set $G(\epsilon\vert X;\theta)$, i.e. the model predicts $\sigma\in G(\epsilon\vert X;\theta)$. The outcome $Y$ is a random variable with distribution $\sigma$, so it can be written $Y=Q_\sigma(V)$, where $Q_\sigma$ is the quantile transform associated with $\sigma$, and $V$ is a random variable with uniform distribution, independent of $\epsilon$ and $\sigma$, conditionally on $X$. The variable $V$ is traditionally interpreted as an independent signal that agents receive and on whose realization they base their action. The identified set is therefore defined as the set $\Theta_I$ of values of the parameter such that $Y=Q_\sigma(V)$ and $\sigma\in G(\epsilon\vert X;\theta)$ as specified above.
BMM:2008 (hereafter BMM) were the first to extend the characterization of the identified set to the case of equilibria in mixed strategies, which is both a very important and nontrivial problem. It is particularly useful, as it is well known that in normal form games, existence of equilibria is guaranteed in mixed strategies, but not necessarily in pure strategies, except in special classes of games such as supermodular games. It is a nontrivial problem, since the likelihood of the model is no longer the distribution function of a random set. Nonetheless, we go beyond the results of BMM and show that in certain classes of games the identified set can be characterized by a finite set of inequalities. We also show that our efficient submodular optimization method is still applicable to this case, as the identified set continues to be characterized by the fact that true data distribution is contained in the core of the model likelihood as in theorem (ref). In the general case, where the likelihood is not necessarily submodular, we show that the identified set can be computed efficiently through a convex optimization problem, which we describe.
We first show now that the identified set $\Theta_I$ defined in definition (ref) is indeed equal to $\Theta_I^\ast$ defined in BMM. Indeed: suppose $\theta\in\Theta_I$. Call $\mu(.\vert\epsilon, X;\theta)$ the distribution of the random element $\sigma$ (hence a distribution on the simplex). The latter has support $G(\epsilon\vert X;\theta)$. Then, for any value $y\in\mathcal{Y}$, we have
by independence. Conversely, if ((ref)) holds, we can construct the random quadruplet $(Y,V,\epsilon,\sigma)$ with the required properties.
The fundamental work by BMM was the first to derive a characterization of the identified set in the case of multiple mixed strategy equilibria, which we restate here as the following proposition:
For completeness, in the appendix we give a different proof of this result, as it fits more naturally with our presentation than the original random set theoretic proof given in BMM:2008. We now show that this characterization allows efficient computation of the identified set as the solution of a convex optimization problem. In addition, under regularity conditions on the equilibrium correspondence of the game, we show that we can derive an exact characterization of the identified set with a finite collection of inequalities, as was the case when considering equilibria in pure strategies only.
In this section, we give a simple procedure, which is valid for any normal form game, and is recommended to practitioners who do not wish to restrict attention to pure strategy Nash equilibria. In proposition (ref), the functional $\tilde{\mathcal{L}}$ is convex, as a maximum of linear functionals. As a result, we show that the identified set can be computed with a standard convex optimization program. Indeed, $\Theta_I$ is the set of $\theta$'s such that the convex functional $\phi(f)=\tilde{\mathcal{L}}(f\vert X;\theta)-E_P(f(Y)\vert X)$ remains non negative for all function $f$. Suppose there is some $v$ (a function on $\mathcal{Y}$, hence a vector in $\mathbb{R}^{d_y}$) such that $\phi(v)<0$. Since $\phi$ is positively homogeneous (i.e. $\phi(\lambda v)=\lambda\phi(v)$ for all $\lambda\geq0$), then for each $z\in\mathbb{R}^{d_y}$, such that $z\ne0$, one of the following three cases applies:
Hence, if we take any $z$, for instance $z=(1,0,\ldots,0)$, and define $H=\{w: \langle z,w\rangle=0\}$, $\phi$ takes negative values if and only if the following statement is true \[\min\left(\min_{w\in H}\phi(w+z),\hskip5pt\min_{w\in H}\phi(w-z)\right)<0\] Since $H$ is a convex set, the program above is a completely standard convex optimization program (we implement the procedure with the matlab optimization toolbox).
We now turn to the efficient computation of the identified set in the case where the likelihood is submodular. This occurs when the game satisfies the following assumption, taken from Shapley:71. Recall that the upper envelope $\phi$ of a family of probability functions $A\mapsto\mu_i(A)$, $i=1,\ldots,I$ is defined as $A\mapsto\phi(A)=\max_{i\in I}\mu_i(A)$.
We now show a very simple and easy to check sufficient condition for a regular core.
In this case, we give an exact characterization of the identified set with only a finite number of inequalities, unlike BMM's characterization. One implication is trivial: it is obvious to see that $E_{P}{ (f(Y)|X)\leq \tilde{\mathcal{L}} }\left( f\vert X;\theta \right)$ implies $ P(B|X)\leq \mathcal{L}(B|X;\theta )$ (it suffices to take $f\left( {y}\right) =1_{B}\left(y\right) $). The converse is far from trivial, and follows from the characterization of Choquet integrals (see Schmeidler:86).
This implies that the problem of checking whether a value of the parameter is in the identified set is equivalent to the problem of checking whether the distribution of $Y$ is in the core of a Choquet capacity. This was already the case in GH:2006a when we restricted ourselves to pure strategy equilibria. However, in the latter case, the Choquet capacity was infinitely alternating, as it was defined as the distribution of a random set. With mixed strategy equilibria, it no longer has this property, hence it is no longer the distribution of a random set.
As explained in section (ref), the strategy to compute the identified set efficiently in the case, where the likelihood is submodular is based on the following observation.
As stated in definition (ref), a set function $\varphi$ is called submodular if $\varphi(A\cup B)+\varphi(A\cap B)\leq \varphi(A)+\varphi(B)$, and $\varphi(B) = \mathcal{L}(B\vert X;\theta)-P(B\vert X)$ is indeed submodular, as shown in the proof of theorem (ref). The notion of submodularity is the analog of convexity for functions defined on lattices, such as set functions. Hence, minimizing a submodular set function is akin to minimizing a convex function. As noted in section (ref), it is a classical problem in combinatorial optimization (see for instance Topkis:98 chapter 2), and several polynomial-time algorithms exist (see for instance Fujishige:2005). We implement this using Andreas Krause's Matlab SFO toolbox. Note that, as explained in section (ref), the more efficient optimal transportation method does not apply when mixed strategy equilibria are considered.
In order to illustrate our fundamental characterization of the identified set, we propose a series of Monte Carlo simulations based on our example (ref). For three different values of the interaction parameter $\theta=0.25,0.5,0.75$, we represent the core of the likelihood in the four dimensional simplex. We simulate observed strategy profiles according to probability distributions $P$ in the core of the model. We consider three cases. First, the case where the probability distribution $P$ of the simulated data is the barycenter of the core (this case is called “central DGP”). Second a case where $P$ is an extreme point of the core (this case is called “extreme DGP”), and finally an intermediate case (called “intermediate DGP”). In each case, and for each value of $\theta$, we simulated $10000$ samples of size $n=100,1000$, excluded the $5\%$ most distant (case with “confidence=0.95”) or the $10\%$ most distant (case with “confidence=0.9”) and showed graphically how the set of remaining empirical distributions $P_n$ intersects with the core. The $36$ graphics pertaining to all cases described are given in the appendix. In all graphs on figures (ref) to (ref), the red tetrahedron is the four dimensional simplex, whose extreme points correspond to the dirac masses on each of the equilibrium profiles $00,01,10,11$. All points inside the tetrahedron have barycentric coordinates corresponding to the vector of probabilities attached to each profile. The blue diamond-shaped polyhedron is the core computed for each of the values of the parameter $\theta=0.25,0.5,0.75$, i.e. the set of distributions of equilibrium profiles that are compatible with the model for that particular value of $\theta$. The true distribution is a point in the simplex. If the true distribution is a point in the core, then the value of $\theta$ is in the identified set.
In figure (ref), the core is plotted for the case $\theta=0.25$ and $10,000$ data series of size $n=100$ are drawn from the analytical center of the core. For each simulated sample, the empirical distribution is computed, and the green ball-shaped polyhedron is the set of $9,000$ out of the $10,000$ computed empirical distributions. The $1,000$ that are excluded are the most distant (in Euclidian distance) from the data generating process, and they are excluded for visual convenience. The sensitivity to the proportion of plotted empirical distributions (called “confidence” in the captions) can be seen by comparing figures (ref) and (ref), in which only $500$ outliers are excluded. We see that in both figures (ref) and (ref), a sizeable proportion of draws falls outside the core, which does not mean, however, that a confidence ball around such empirical distributions would not intersect with the core, hence it does not necessarily lead to the exclusion of $\theta$ from a confidence region for the identified set. In the following graphics, figures (ref) and (ref), the setup is identical, except for the fact that the samples drawn have size $1,000$. In that case, it is clear that all drawn empirical distributions are well within the core, which would lead to include $\theta$ in any reasonably constructed confidence region for the identified set. The observations are similar for figures (ref) to (ref), where the data are generated from the mid-point between the analytical center and an extreme point of the core. As a result, a larger proportion of simulated draws falls outside the core when $n=100$, but again, all draws a well within the core for $n=1,000$, so that $\theta$ would always be contained in a reasonably constructed confidence region for the identified set. In figures (ref) to (ref), the data is generated from an extreme point of the core. In such a case, a point in the neighbourhood of the data generating process is more likely to fall outside the core than inside it. Hence, for all sample sizes, the proportion of empirical distributions that fall outside the core is relatively stable, and larger than the proportion that falls inside the core. However, for larger sample sizes, we see that simulated empirical distributions are more likely to be within a predetermined neighbourhood of the core, so that the value $\theta$ is less likely to be excluded from a confidence region for the identified set. The cases where $\theta= 0.5$ and $\theta= 0.75$ are very similar and we do not report them here.
The applicability of the methodology proposed is best illustrated on data pertaining to our example (ref). We estimate the determinants of long term care option choices for elderly parents in American families. The data consists of a sample of 948 families with two children drawn from the National Long Term Care Survey, sponsored by the National Institute of Aging and conducted by the Duke University Center for Demographic Studies under Grant number U01-AG007198, LTCS:82. Elderly people were interviewed in 1984 about their living and care arrangements. The survey questions include gender and age of the children, the distance between homes of parents and each of the children, the disability status of the elderly parent (where disability is referred to as problems with “Activities of Daily Living or Instrumental Activities of Daily Living (ADL)”) and the number of hours per week each of the children devote to the care of the elderly parent. The dependent variable is the care provision for the parent. The parent is asked to list children (either at home or away from home) and how much each provides help. If only one child is listed as providing significant help, that child is designated the primary care giver. If more than one child is listed, the one providing the most hours is designated the primary care giver. If no child is listed or if the parent lives in a nursing home, then the parent is designated as “living alone.”
The observable choice of care option is modeled as in ES:2002 as the outcome of a family bargaining game. The payoff to family member $i$, $i=1,2$ is represented as the sum of three terms. The first term $V_{ij}$ represents the value to child $i$ of care option $j$, where $j>0$ means child $j$ becomes the primary care giver and $j=0$ means the parent remains alone or is moved to a nursing home. The matrix $V=(V_{ij})_{ij}$ is known to both children. We suppose it takes the form \[V=\left(
\right)\] where $ADL$ is a dummy variable that takes the value 1 if the parent has problems with activities of daily living or instrumental activities of daily living, $D_i$ is the distance between parent and child $i$, $F=1$ if child $1$ (who is first-born) is female, and zero otherwise and $\theta=(\beta_0,\beta,\psi,\alpha)'$ is a four dimensional parameter vector, unknown to the analyst. The value a family attaches to the fact that a non disabled parent is taken care of by one of the children is measured by $\beta_0$, whereas for a disabled parent, it is $\beta_0+\beta$. $\psi$ measures the disutility to the caregiver of living more than an hour away from the parent, and $\alpha$ measures the incremental utility of taking care of the parent for the firstborn daughter, as compared with a firstborn son.
In the data, $85\%$ of interviewed parents have problems with activities of daily living. $70\%$ of parents live alone. $46\%$ of first-born children are female, whereas $56\%$ of first-born who are primary care-takers are female. The female first-born effect was identified in ES:2002, and we wish to see whether it is robust to our analysis that takes multiple equilibria in the game seriously.
Both children simultaneously decide whether or not to take part in the long term care decision. Suppose $M$ is the set of children who participate. The option chosen is option $j$ which maximizes the sum $\sum_{i\in M}V_{ij}$ among the available options (only participating children can become primary care givers). It is assumed that participants abide with the decision and that benefits are then shared equally among children participating in the decision through a monetary transfer $s_i$, which is the second term in the children's payoff. The third term $\epsilon_i$ in the payoff is a random benefit from participation, which is $0$ for children who decide not to participate and distributed according to $\nu(.\vert \theta)$ for children who participate. All children observe the realizations of $\epsilon$, whereas the analyst only knows its distribution (the negative shock on participation is calibrated with a normal distribution $N(-3,1)$ for $\epsilon$. An alternative would have been a negative exponential distribution for $\epsilon$).
The payoff matrix of the participation game can be determined in the following way.
We do not observe participation, but only the chosen care option. To each equilibrium strategy profile corresponds a (almost surely) unique care option choice, hence for each participation shock $\epsilon$, we can derive the $G(\epsilon\vert F,D,ADL;\theta)$ as the set of probability measures on the set of care options $\{0,1,2\}$ induced by mixed strategy profiles, which are probabilities on the set of participation profiles $\{(N,N), (N,P), (P,N), (P,P)\}$. The Gambit software is a good option for the computation of $G(\epsilon\vert F,D,ADL;\theta)$.
The methodology proposed in the paper allows the construction of the identified set based on the hypothetical knowledge of the true distribution of the data. In order to account for sampling uncertainty, we appeal to the bootstrap procedure proposed in GHQ:2008 to construct confidence regions for the identified set in such situations. The method relies on a bootstrap determination of a set function $A\mapsto\underline{P}(A\vert F,D,ADL)$ which is dominated by $P(A\vert F,D,ADL)$ (uniformly over $A\subseteq\{0,1,2\}$, $X$, $D$ and $ADL$) with probability $1-\alpha$ (the chosen level of significance, here $0.95$). Once $\underline{P}$ is determined, one proceeds as recommended in section (ref). That is to say, we keep in the identified set only values of $\theta$ such that for all observed values of the explanatory variables, the minimum over $A\subseteq\{0,1,2\}$ of the function $\mathcal{L}(A\vert F,D,ADL;\theta)-\underline{P}(A\vert F,D,ADL;\theta)$ is non negative. The search over the 4-dimensional parameter space is conducted in the following way. First one finds a value of the parameter which lies within the confidence region (for instance the corresponding estimates in the analysis of ES:2002 which achieves identification by removing multiplicity of equilibria), then one chooses a coarse discretization of $[-\pi,\pi]^3$ to search in all directions for the frontier point of the confidence region, which is assumed to be arc-connected (on each arc, we use a dichotomic search). The resulting frontier points are the extreme points of the confidence polytope. Each value of the parameter can be tested in a fraction of a second on a standard laptop, and the region can be constructed in a few hours, again on a standard laptop without parallel processing. The confidence region is a four dimensional polytope and 3, 2 or 1-dimensional confidence regions for subsets or linear combinations of parameters can be easily visualized as cuts from the confidence region using the matlab multiparametric MPT toolbox.
The variables chosen were those that were significant in ES:2002. We test the significance of each of the individual parameters by checking whether the hyperplanes defined by $\beta_0=0$, $\beta=0$, $\psi=0$ and $\alpha=0$ intersect with our $95\%$ confidence region. The range of values of $\beta_0$ is $\beta_0\in[-4.297,1.426]$. Hence, the hyperplane $\beta_0=0$ has non empty intersection with the confidence region, which means we fail to reject (at the $5\%$ level) the null hypothesis that the family values identically the fact that a non disabled parent lives alone or is taken care of by a family member. On the other hand, the range of values for $\beta$ is $\beta\in[1.976,7.297]$, so the half space $\beta\leq0$ has empty intersection with the confidence region, and we reject the null hypothesis that a parent's disability does not increase the value of care provided by the family. We also reject the hypothesis that distance between parent and caretaker is insignificant, as the range for $\psi$ is $\psi\in[-2.981,-1.306]$. Finally, we reject the hypothesis that gender of the firstborn child does not affect the chosen care option, as the range for $\alpha$ is $\alpha\in[0.409,2.405]$. Indeed, our results are still consistent with ES:2002 in the finding that families value care provided by a firstborn daughter more than care provided by a firstborn son. Since the hypothesis that $\beta_0=0$ is not rejected, we slice the confidence region at $\beta_0=0$ to obtain the constrained confidence region, which is three dimensional and is represented (from four different angles) in figure (ref). In the constrained region, the ranges for the three remaining parameters are considerably tighter, as $\beta\in[2.717,4.062]$, $\psi\in[-2.867,-1.365]$ and $\alpha\in[0.532,2.290]$. Two dimensional regions are also given for $(\alpha,\psi)$, when $\beta=3$, for $(\beta,\psi)$ when $\alpha=2$ and $(\beta,\alpha)$ when $\psi=-2$. These regions are plotted in figure (ref). Notice the thin diagonal shapes of the regions, which makes them rather informative, as for instance simultaneous large values of $\psi$ and $\alpha$ are rejected, which roughly means that firstborn daughters living close by are not the only possible caregivers. Similarly, $\beta$ and $\psi$ are not simultaneously very large, so that either the disability effect is very large or the distance effect is very large, but not both.
In the context of models with multiple equilibria in pure and mixed strategies, we have proposed an equivalence result between the existence of an equilibrium selection mechanism compatible with the data and a finite set of inequalities characterizing the core of the model likelihood, and provided methods to reduce this number of inequalities to be checked with an appeal to the notion of core determining families and to efficient easily implementable combinatorial methods. The issue of statistical inference on the identified feature thus characterized is taken up in GH:2006d and GH:2006c, which complement the seminal work of CHT:2007.
\markboth{References}{References} \printbibliography