Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
51,880 characters · 12 sections · 35 citation commands
Dilation Bootstrap
\vskip6pt {\scriptsize JEL Classification: C15, C31. \\Keywords: Partial identification, dilation bootstrap, quantile process, optimal matching.}
In several rapidly expanding areas of economic research, the identification problem is steadily becoming more acute. In policy and program evaluation Manski:90 and more general contexts with censored or missing data SV:2005, MM:2005 and measurement error CHT:2005, ad hoc imputation rules lead to fragile inference. In demand estimation based on revealed preference BBC:2005 the data is generically insufficient for identification. In the analysis of social interactions BD:2001, Manski:2004, complex strategies to reduce the large dimensionality of the correlation structure are needed. In the estimation of models with complex strategic interactions and multiple equilibria BV:85, Tamer:2003, assumptions on equilibrium selection mechanisms may not be available or acceptable.
More generally, in all areas of investigation with structural data insufficiencies or incompletely specified economic mechanisms, the hypothesized structure fails to identify a unique possible data generating mechanism for the data that is actually observed. In such cases, many traditional estimation and testing techniques become inapplicable and a framework for inference in incomplete models is developing, with an initial focus on estimation of the set of structural parameters compatible with true data distribution (hereafter {\em identified set}). A question of particular relevance in applied work is how to construct valid confidence regions for the identified set. Formal methodological proposals abound since the seminal work of CHT:2007, but computational efficiency is still a major concern.
In the present work, we propose a methodology that clearly distinguishes how to deal with sampling uncertainty on the one hand, and model uncertainty on the other, so that unlike previous methodological proposals, search in the parameter space is conducted only once, thereby greatly reducing the computational burden. The key to this separation is to deal with sampling variability without any reference to the hypothesized structure, using a methodology we call the {\em dilation method}. This consists in dilating each point in the space of observable variables in such a way that the empirical probability (which is known) of a dilated set dominates the true probability (which is unknown) of the original set (before dilation). The unknown true probability (i.e. the true data generating mechanism) is then removed from the analysis, and we can proceed as if the problem were purely deterministic, hence apply the methods proposed in GH:2008a and BMM:2008.
To construct confidence regions of level $1-\alpha$ for the identified set, such a dilation $y\rightrightarrows J(y)$ (where $\rightrightarrows$ denotes a one-to-many map) must satisfy $\tilde{Y}^\ast\in J(\tilde{Y})$ a.s. for some pair of random vectors $(\tilde{Y}^\ast,\tilde{Y})$, with probability $1-\alpha$, where $\tilde{Y}$ is drawn from the true distribution of observable variables and $\tilde{Y}^\ast$ is drawn from the empirical distribution relative to the observed sample. We propose a dilation bootstrap procedure to construct $J$, in which bootstrap realizations $Y_j^b$, $j=1,\ldots,n$ are matched one-to-one with the original sample points $Y_j$, $j=1,\ldots,n$ so as to minimize $\eta_n^b=\max_{j=1,\ldots,n}\|Y_j^b-Y_{\sigma(j)}\|$, where the permutation $\sigma$ defines the matching. The $\alpha$ quantile of the distribution of $\eta_n^b$ then defines the radius of the dilation.
When the observable $Y$ is a random variable, the dilation bootstrap relies on bootstrapping the quantile process, as proposed by DG:92. However, bootstrapping the quantile process relies on order statistics and had no higher dimensional generalization to date. This is now provided by the dilation bootstrap, which removes the constraint on dimension through the appeal to optimal matching. Although the problem of finding minimum cost matchings (called the {\em assignment} or {\em marriage problem}) is very familiar to economists, as far as we know, its application within an inference procedure is unprecedented.
The rest of the paper is organized as follows. The next section describes the econometric framework and introduces the {\em Composition Theorem} and the dilation method the latter justifies. Section (ref) discusses the application of the Composition Theorem to constructing confidence regions for partially identified parameters. Section (ref) presents the bootstrap feasible dilation and its theoretical underpinnings. Section (ref) presents simulation evidence on the performance of the dilation bootstrap in comparison with alternative methods. Section (ref) explains how the method extends to higher dimensions and discrete choice and the last section concludes.
We consider the problem of inference on the structural parameters of an economic model, when the latter are (possibly) only partially identified. The economic structure is defined as in Jovanovic:89, which generalizes KR:50. Variables under consideration are divided into two groups. Latent variables $U$ capture unobserved heterogeneity in the model. They are typically not observed by the analyst, but some of their components may be observed by the economic actors. Observable variables $Y$ include outcome variables and other observable heterogeneity. They are observed by the analyst and the economic actors. We call observable distribution $P$ the true probability distribution generating the observable variables, and denote by $\nu$ the probability distribution that generated the latent variables $U$. The econometric structure under consideration is given by a binary relation between observable and latent variables, i.e. a subset of $\mathcal{Y}\times\mathcal{U}$, which can be written without loss of generality as a correspondence from $\mathcal{U}$ to $\mathcal{Y}$.
We assume a parametric structure for the unobserved heterogeneity and the model linking unobserved heterogeneity variables to observable ones.
Note that the measurability and closed values assumptions are very mild conditions. The assumption that the correspondence is non-empty, however, may be restrictive. In the revealed preferences example, we require that the demand correspondence be non empty. In the games example, we require existence of equilibrium.
The pair of random vectors $(Y,U)$ involved in the model is generated by a probability distribution, that we denote $\pi$. Since the vector $U$ is unobservable, the probability distribution $\pi$ is not directly identifiable from the data. However, the econometric model imposes restrictions on $\pi$. The distribution of its component $Y$ is the observable distribution $P$. The distribution of its component $U$ is the hypothesized probability distribution $\nu(\cdot\vert\theta)$. Finally, the joint distribution is further restricted by the fact that it gives mass 0 to the event that the relation $Y\in G(U\vert\theta)$ is violated. For any given value of the structural parameter vector $\theta$, a joint distribution satisfying all these restrictions may or may not exist. If it does, it is generally non unique. The identified set $\Theta_I$ is the collection of values of the structural parameter vector $\theta$ for which such a joint probability distribution does indeed exist.
The set $\Theta_I$, first formalized in this way in GH:2006a is sometimes called “sharp identification region” to emphasize the fact that it exhausts all the information on the parameter available in the model. No value $\theta\in\Theta_I$ could be rejected on the basis of the knowledge of the model and the observable distribution $P$ only. Take a parameter value $\theta\in\Theta$. It belongs to the identified set $\Theta_I$ if and only if there exists a joint distribution satisfying the required restrictions, in other words, if and only if there exists a “version” of $U$, i.e. a random vector $\tilde{U}$ with the same distribution as $U$, namely $\nu(\cdot\vert\theta)$, such that $Y\in G(\tilde{U}\vert\theta)$ with probability 1. Hence, denoting by $X\sim\mu$ the statement “the random vector $X$ has probability distribution $\mu$,” we can characterize the identified set in the following way, which we take as our formal definition.
Our inference method on the identified set will be based on a general way of combining sources of uncertainty (sampling uncertainty or data incompleteness) by composition of correspondences. Suppose the probability measure $Q$ on $\mathcal{Y}$ is the known distribution of a random vector $Z$ and that it is related to the true unknown distribution $P$ of the observed variables $Y$ by the following relation:
Assumption (ref) characterizes the additional level of indeterminacy the analyst faces. The structural model is incomplete in the sense that the relation between unobserved heterogeneity $U$ and outcomes $Y$ is a many-to-many mapping. In addition, due to observability issues or sampling uncertainty, the distribution of outcomes $P$ is unknown and the relation between true outcome $Y$ and a variable $Z$ that we can simulate is also many-to-many.
The following theorem shows how the two levels of uncertainty can be combined without loss of information.\footnote{The current proof of Theorem (ref), suggested by Alexei Onatski, is shorter and simpler than our original proof in previous versions of the paper. We are responsible for any remaining errors.}
Theorem (ref) implies that when the distribution $P$ of outcomes is unknown, the infeasible identified set $\Theta_I$ can be replaced by a feasible identified set \[\tilde\Theta_I=\left\{\theta\in\Theta\;\vert\;\;\exists\tilde{Z}\sim Q,\;\tilde{U}\sim\nu(\cdot\vert\theta):\;\mathbb{P}(\tilde{Z}\notin J\circ G(\tilde{U}\vert\theta))\leq\beta \right\}.\]
To illustrate the composition theorem, consider a special case of the revealed preference example (ref) combined with measurement error, as in example (ref). Suppose we observe the share $Y$ of risky assets in the portfolio of investors, who are assumed to maximize the expectation of a CARA utility function $u(Y,A;U)=\exp(-U[(1-Y)+YA])$, hence they are assumed to maximize $Y\mathbb{E}(A)-U Y^2\mbox{var(A)}/2$, where $\mathbb{E}(A)$ is the perceived mean of the risky asset $A$ and $\mbox{var}(A)$ its perceived variance. We further suppose investors differ by their risk aversion $U$, for which the analyst hypothesizes an exponential distribution ($F_U(u;\theta)=\mathbb{P}(U\leq u;\theta)=1-e^{-\theta u}$) and by their perception of the riskiness of the asset, and all the analyst knows is a pair of bounds $(\underline\lambda,\overline\lambda)$ such that $\mathbb{E}(A)/\mbox{var}(A)\in[\underline\lambda,\overline\lambda]$. The investor's maximization yields $Y=\mathbb{E}(A)/U\mbox{var}(A)$, so that the model can be summarized by $Y\in G(U)=[\underline\lambda/U,\overline\lambda/U]$. Values $\underline\lambda=50\%$ and $\overline\lambda=200\%$ can be calibrated according to Weitzman:2007. The true distribution of income $Y$ is unknown, but the true cumulative distribution of a mismeasured version $Z=Y+\epsilon$, with $\|\epsilon\|\leq\eta$ a.s., is $F_Z(y)=\mathbb{P}(Z\leq z)=\exp(-1/z)$. By Theorem (ref), the identified set $\tilde\Theta_I$ can be derived from the composed correspondence $J\circ G: u\rightrightarrows J\circ G(u)=[\underline\lambda/u-\eta,\overline\lambda/u+\eta]$, where $J: y\rightrightarrows J(y)=B(y,\eta)$ is a dilation satisfying Assumption (ref). The cumulative distribution of risk aversion satisfies $1-e^{-\theta u}=\mathbb{P}(U\leq u)\in[\mathbb{P}(\overline\lambda/u+\eta\leq Z),\mathbb{P}(\underline\lambda/u-\eta\leq Z)]=[1-e^{-(\overline\lambda/u+\eta)^{-1}},1-e^{-(\underline\lambda/u-\eta)^{-1}}]$. Hence, for all $u>0$, $(\overline\lambda+\eta u)^{-1}\leq\theta\leq(\underline\lambda-\eta u)^{-1}$. Therefore, the identified set can be derived as $\tilde \Theta_I=[1/\overline\lambda,1/\underline\lambda]$.
The main application of the Composition Theorem (ref) that we consider here is the construction of valid confidence regions for partially identified models, based on a sample of realizations of the observable variables.
We propose a new method to construct a confidence region for the identified set $\Theta_I$ of definition (ref).
As noted in IM:2004, this is not the only way to define confidence regions in a partially identified setting, as one might also consider coverage (point wise or uniform) of each value within the identified set. Here we concentrate on a situation where one cannot assume that any value within the identified set can be construed as the true value, so that the whole set is the object of interest. Moreover, a confidence region for the identified set is also a uniform confidence region for each of its elements. The construction of the confidence region is based on a new nonparametric way of controlling sampling uncertainty and its validity relies on a corollary to the Composition Theorem (Theorem (ref) of Section (ref)). We construct sample based sets $J_n^\alpha$, where $\alpha\in(0,1)$ is the desired confidence level, to account for the discrepancy between the empirical distribution $P_n$ associated with the sample and the true observable distribution $P$. We thereby obtain an analogue of Assumption (ref):
Heuristically, the region $J_n^\alpha$ satisfying Assumption (ref) ensures that with suitable confidence, the realizations of the empirical distribution are caught by the enlarged realizations of the true distribution $J_n^\alpha(\tilde{Y})$. Once the dilation $J_n^\alpha$ is obtained, the Composition Theorem can be applied to prove the following:
The dilation $J_n^\alpha$ is chosen to control the confidence level: indeed, by Proposition 1 of GH:2006d (called P1 in the proof of Theorem (ref)), the statement $\exists\tilde{Y}^\ast\sim P_n,\;\tilde{U}\sim\nu(\cdot\vert\theta):\;\mathbb{P}(\tilde{Y}^\ast\notin J_n^\alpha\circ G(\tilde{U}\vert\theta))=0$ is equivalent to $P(A)\leq P_n(J_n^\alpha(A)),$ for all Borel subset $A$ of $\mathcal{Y}$. Hence, the unknown distribution $P$ of an event $A$ is dominated by the empirical distribution of the dilation $J_n^\alpha(A)$ of the event $A$. As both $P_n$ and $\nu(\cdot\vert\theta)$ are known, the construction of $\Theta_n^\alpha$ is feasible and efficient methods to compute it were proposed in GH:2008a and BMM:2008.
We now turn to the question of how to construct the dilation $J_n^\alpha$ that satisfies Assumption (ref). When $Y$ is a random variable, such dilation will be obtained from uniform confidence bands for the quantile process.
The idea of the construction of dilations satisfying Assumption (ref) is based on the quantile transformation. Indeed, letting $Z$ be a uniform random variable on $[0,1]$ and defining $\tilde{Y}=Q(Z)$ and $\tilde{Y}^\ast=Q_n(Z)$, we have a pair of random variables $\tilde{Y}$ and $\tilde{Y}^\ast$ with respective probability distributions $P$ and $P_n$. Suppose a uniform confidence band is available for the quantile function of the form $\mathbb{P}(\eta_n:=\sup_{0\leq t\leq 1}|q_n(t)|\leq \tilde c_n(\alpha))=1-\alpha_n$. Then, with probability $1-\alpha_n$, we have $|\tilde Y^\ast-\tilde Y|=|Q_n(Z)-Q(Z)|\leq \tilde c_n(\alpha)/\sqrt{n}$ almost surely. Hence, the dilation $J_n^\alpha$ defined for all $y$ by $J_n^\alpha(y)=B(y,\tilde c_n(\alpha)/\sqrt{n})$ satisfies Assumption (ref). Moreover, the choice of dilation $J_n^\alpha(y)=B(y,\tilde c_n(\alpha)/\sqrt{n})$ is optimal in the sense that, under the regularity conditions of Assumption (ref), $|Q_n(Z)-Q(Z)|$ achieves the minimum of $|\tilde Y^\ast-\tilde Y|$ when $\tilde Y^\ast$ (respectively $\tilde Y$) ranges over the set of random variables with distribution $P_n$ (respectively $P$). Note that smaller dilations are desirable, as they maximize informativeness of the resulting confidence region.
The following conditions guarantee the existence of such uniform confidence bands for the quantile process.
A distribution function $F$ satisfying Assumption (ref) is called {\em tail monotonic with index $\gamma$} by Parzen Parzen:79. To indicate the mildness of Assumption (ref), Parzen:79 gives the following example where it fails: $1-F(y)=\exp(-y-C\sin y)$ with $0.5<C<1$. As shown below, under Assumption (ref), asymptotic results on the empirical quantile process allow us to derive a dilation $J_n^\alpha$ that satisfies Assumption (ref) for all $\alpha\in(0,1)$. Define $c(\alpha)$ implicitly by $\mathbb{P}(\sup_{0\leq t\leq 1}|B(t)|\leq c(\alpha))=1-\alpha$, where $B(t)$ is a Gaussian process called a {\em Brownian bridge}. For any $\alpha\in (0,1)$, we then have the following result.
The dilation in proposition (ref) is infeasible, as it depends on the unknown $f$ and it relies on quantiles $c(\alpha)$ that are difficult to compute. We develop a feasible alternative in our dilation bootstrap procedure in section (ref). We resort to a bootstrap matching algorithm to construct feasible versions of the dilation above.
To introduce the simple idea underlying the method, consider the sample $(Y_1,\ldots,Y_n)$ and a given bootstrap realization $(Y_1^b,\ldots,Y_n^b)$ as in figure (ref). As before, $(Y_{(1)},\ldots,Y_{(n)})$ are the order statistics associated with the sample and $(Y_{(1)}^b,\ldots,Y_{(n)}^b)$ are the order statistics associated with the bootstrap realization (with arbitrary ranking of the ties). In the illustrative example of figure (ref), the smallest observation of the initial sample $Y_{(1)}$ was drawn once in the bootstrap sample, the second smallest was not drawn, the third smallest was drawn once, the fourth smallest twice, and the largest $Y_{(n)}$ was drawn twice. The arrows in the figure represent the bijection that matches the $j$'th order statistic of the initial sample $Y_{(j)}$ with the $j$'th order statistic of the bootstrap sample $Y_{(j)}^b$ for each $j=1,\ldots,n$.
To achieve a bootstrap analog of Assumption (ref), we need a dilation $J_n^b$ and a permutation $\sigma$ of $\{1,\ldots,n\}$ such that $Y_{(j)}\in J_n^b(Y_{\sigma(j)}^b)$ for all $j=1,\ldots,n$. One such permutation matches the order statistics of the initial sample with the order statistics of the bootstrap sample. In this matching in the example of figure (ref), $Y_{(1)}$ is matched with $Y_{(1)}^b$, namely with itself. $Y_{(2)}$ was not drawn in the bootstrap sample, so it is matched with $Y_{(2)}^b$, which is equal to $Y_{(3)}$, for whom $Y_{(2)}$ is the second closest neighbor in Euclidian distance. $Y_{(3)}$ is the nearest neighbor of its match $Y_{(3)}^b=Y_{(4)}$, $Y_{(4)}$ is matched with itself, $Y_{(n-1)}$ is the nearest neighbor of its match $Y_{(n-1)}^b=Y_{(n)}$ and finally $Y_{(n)}$ is matched with itself. The longest distance between two matches is $\eta_n^b=|Y^{b}_{(n-1)}-Y_{(n-1)}|=|Y_{(n)}-Y_{(n-1)}|$. Hence, if $J_n^b(y)=B(y,\eta_n^b)$, $Y^b_{(j)}\in J_n^b(Y_{(j)})$ will be satisfied for all $j=1,\ldots,n$ in this particular bootstrap sample realization. The chosen matching in figure (ref) characterizes the bootstrap quantile process (see Definition (ref)) and it minimizes the largest deviation $\eta_n^b$, and hence produces the smallest dilation.
In the illustrative example of figure (ref), the bootstrap quantile process attains its maximum over $t\in[0,1]$ at $t$ such that $n-2<nt\leq n-1$ and $\eta_n^b=Y^b_{(n-1)}-Y_{(n-1)}=Y_{(n)}-Y_{(n-1)}$. In the population of bootstrap realizations, $\eta_n^b$ has distribution with $1-\alpha$ quantile $c_n^\ast(\alpha)$. The latter can be approximated by simulation with a large number $B$ of bootstrap replications. We obtain $\eta_n^b$ for each $b=1,\ldots,B$. Call $\hat{c}_n^\ast(\alpha)$ the $[B\alpha]$-th largest among the $\eta_n^b$'s (where $[.]$ denoted integer part) and $\hat J_n^{\alpha,\ast}(y)=B(y,\hat{c}_n^\ast(\alpha))$, then by construction, a proportion $1-\alpha$ of the bootstrap samples indexed by $b=1,\ldots,n$ will satisfy $Y^b_{(j)}\in J_n^{\alpha,\ast}(Y_{(j)})$ for all $j=1,\ldots,n$. By Theorem 2 of Singh:81 (see also Theorem 5.1 of BF:81), the bootstrap quantile process $(q_n^b(t))_{t\in[0,1]}$ has almost surely the same uniform weak limit as the empirical quantile process $(q_n(t))_{t\in[0,1]}$ and we therefore have the following result on the validity of the bootstrap dilation.
Note that in the univariate case, the simulation approximation $\hat{c}_n^\ast(\alpha)$ to the quantile $c_n^\ast(\alpha)$ is very simple to derive. The simplest algorithm requires ordering the initial sample and each of the bootstrap samples and computing the maximum of $|Y^b_{(j)}-Y_{(j)}|$ over $j=1,\ldots,n$. However, we have introduced, with figure 1 and the discussion above, an equivalent algorithm, which runs as follows: for each $b=1,\ldots,B$, find the permutation $\sigma$ over $\{1,\ldots,n\}$, which minimizes the quantity $\max_j|Y^b_{j}-Y_{\sigma(j)}|$. Unlike the algorithm based on the order statistics, such an {\em optimal matching} or {\em optimal assignment} procedure can be performed regardless of dimension and efficient algorithms and implementations are available.
We assess the small sample performance of the dilation bootstrap on the following simulation design. Observable variables $Y$ have a standard normal distribution, while unobserved heterogeneity variable $U$ is assumed to follow a normal distribution with mean $\theta$ and variance $1$. The cumulative distribution of $U$ is denoted $F_U$. The model correspondence $G$ is defined for each $u$ by $G(u)=[u-1,u+1]$, so that the model is characterized by the relation $Y\in G(U)=[U-1,U+1]$. Therefore, the identified set can be immediately derived as $\Theta_I=[-1,1]$. For $5,000$ initial samples of size $n=50,100,500$, with empirical distributions $P_n$, we compute $\hat{c}_n^\ast(\alpha)$ with $5,000$ bootstrap replications, and use the dilation $\hat J_n^{\alpha,\ast}(y)=B(y,\hat{c}_n^\ast(\alpha))$, so that a parameter value $\theta$ belongs to the $(1-\alpha)$-confidence region $\Theta_{\mathrm{CR}}$ for $\Theta_I$ if and only if there exist $\tilde{Y}^\ast\sim P_n$ and $\tilde{U}\sim F_U(.;\theta)$ such that $\mathbb{P}^\ast(\tilde{Y}^\ast\in \hat J_n^{\alpha,\ast}\circ G(\tilde{U}))=1$. Since $P_n$ and $F_U$ are known, the latter condition can be checked efficiently with the core determining class method of GH:2008a, section 2.3. We report Monte Carlo coverage probabilities in case of significance level $\alpha=0.01$, $0.05$ and $0.1$ in Table (ref).
The most notable feature to note is the tendency to under reject in small samples, especially for true size $\alpha=0.10$ but also for true size $\alpha=0.05$. For true size $\alpha=0.01$ on the other hand, the procedure displays slight over rejection in small samples. For comparison purposes, we also report coverage probabilities from the generic subsampling procedure for set coverage in CHT:2007 based on the criterion function $\sqrt{n}\max(\max_{j=1,\ldots,n}[F_n(Y_j)+F_U(Y_j+1)], \max_{j=1,\ldots,n}[-F_n(Y_j)+F_U(Y_j-1)])$. In order to avoid artificially favoring our results, we report coverage probabilities under several subsample sizes and when the true identified set is known and no initial estimate is needed in the CHT:2007 procedure. The Monte Carlo coverage probabilities for $500$ initial samples and $500$ subsamples of sizes $40,45,48$ when $n=50$, $85,92,95$ for $n=100$ and $425,450,475$ for $n=500$ are reported in Table (ref). We find the procedure over rejects in all but one case, and there is moderate dependence in the choice of subsample size.
The dilation method and dilation bootstrap have natural extensions to the cases, where observable variables $Y$ are multivariate and to the case, where $Y$ is discrete. We consider both extensions in the following subsections.
Consider first the case, where the random vector of observable variables $Y$ has dimension $d\geq2$. This extension allows the consideration of multiple equations models. Moreover, it is particularly relevant in this partially identified framework, as it also allows the consideration of single equations models with endogenous regressors.
In case $Y$ is multivariate, although Theorem (ref) holds irrespective of dimension, the construction of a dilation satisfying Assumption (ref) can no longer rely on the traditional quantile process as in Propositions (ref) and (ref). However, the quantity $\eta_n=\inf\{\|\tilde{Y}^\ast-\tilde{Y}\|_\infty:\;\tilde{Y}^\ast\sim P_n,\;\tilde{Y}\sim P\}$ is still well defined. When attained, it is achieved by a pair of random vectors $(\tilde{Y}^\ast,\tilde{Y})$ with marginal distributions $P_n$ and $P$, which minimizes the largest deviation. Equivalently, there exist $\tilde{Y}^\ast\sim P_n$ and $\tilde{Y}\sim P$ such that $\tilde{Y}^\ast$ belongs to a closed ball $B(\tilde{Y},\eta_n)$ centered on $\tilde{Y}$ and with radius $\eta_n$, i.e. such that $\mathbb{E}[1\{\tilde{Y}^\ast\notin B(\tilde{Y},\eta_n)\}]=0$.
When $Y$ is uniformly distributed on the unit cube $[0,1]^d$, the quantity $\eta_n$ is well studied in the probability literature. Hence, using asymptotic results on the quantity $\eta_n$ in the literature, specifically LS:89 for the case $d=2$ and SY:91 for the case $d\geq3$, we can derive analytical formulae for the dilation $J_n$:
However, the results of Proposition (ref) only pertain to the uniform case and produce conservative confidence regions. More generally, we propose constructing suitable dilations based on the distribution of $\eta_n$.
By construction, we then see that the ball $B(y,c_n(\alpha))$ is a suitable dilation, in the sense that it satisfies Assumption (ref).
As for the approximation of $c_n(\alpha)$ to obtain a feasible dilation, once again, although the quantile process is no longer defined, the matching algorithm described in Section (ref) is easily generalizable and delivers a bootstrap dilation approximation of $J_n^\alpha$. The general procedure is described as follows. \vskip10pt
The problem of finding the permutation that achieves $\eta_n^b$ is called {\em bottleneck bipartite matching} in the combinatorial optimization and operations research literature.
We now turn to the case of aggregate data from discrete choice. To fix ideas, consider a voting model, where $K$ parties are represented in $n$ electoral districts and observations $\hat p_{i,k}$, $i=1,\ldots,n$ and $k=1,\ldots,K$, are reported shares of votes for party $k$ in district $i$. Voter $l$ chooses the party that maximizes their utility $u^l_{i,k}(\theta)+\rho_{i,k}+\epsilon^l_{i,k}$, where $u_{i,k}(\theta)$ is a deterministic function of (observed covariates and) the unknown parameter $\theta$, $\rho_{i,k}$ are random district-party effects (independent of voters) and the $\epsilon^l_{i,k}$'s are i.i.d. type I extreme value random utilities. True vote shares for party $k$ in district $i$ satisfy $\ln p_{i,k}^\ast(\rho_{i,k})=u_{i,k}(\theta)+\rho_{i,k}+ \ln\sum_k\exp(u_{i,k}+\rho_{i,k})$. True shares $p_{i,k}^\ast$ are unobserved, however, due to the possibility of electoral fraud. Reported shares $p_{i,k}$ are assumed to satisfy $p_{i,k}\geq p_{i,k}^\ast$ when a representative of party $k$ is present during the vote count in district $i$. In districts, where no party representative is present, the situation is equivalent to missing data on vote shares. Let $X_{i,k}$ be equal to $1$ if a representative of party $k$ is present in district $i$ during vote count, and zero otherwise. We assume $X=(X_{i,k})_{i=1,\ldots,n;\;k=1,\ldots,K}$ is exogenous. The correspondence characterizing the model is
District $i$ has $n_i$ voters. Call $\hat p_{i,k}$ the proportion of votes in district $i$ reported as going to party $k$ and write $\hat p_i=(\hat p_{i,k})_{k=1,\ldots,K}$. By the central limit theorem, $\sqrt{n_i}(\hat p_i-p_i)$ has Gaussian limiting distribution with zero mean and covariance matrix $V_i$, with diagonal elements $p_{i,k}(1-p_{i,k})$ and off-diagonal elements $-p_{i,k}p_{i,k'}$. Call $Z_i$ a random vector with distribution $N(0,V_i/n_i)$ and let $\eta_i$ be such that $\mathbb{P}(Z_i\notin B(0,\eta_i))=\alpha_i$, where $B(0,\eta_i)$ is the open ball centered at zero with radius $\eta_i$. Define the dilation $J_{n_i}^{\alpha_i}$ defined for each $p$ by $J_{n_i}^{\alpha_i}(p)=B(p,\eta_i)$. Then $J_n^\alpha(p)=\bigcup_iJ_{n_i}^{\alpha_i}(p)=B(p,\max_i\eta_i)$ satisfies Assumption (ref) for $\alpha=\lim\sup_n\Pi_{i=1}^n\alpha_i$.
In the two-party case, call $\hat p_i$ the reported share of votes for party 1, $p_i$ the {\em true} or {\em population} reported share and $p_i^\ast$ the true share (absent reporting fraud). The true share satisfies $\ln p_i^\ast(\rho_i)=u_{i,1}(\theta)+\rho_{i,1}+\ln(\exp(u_{i,1}(\theta)+ \rho_{i,1})+ \exp(u_{i,2}(\theta)+\rho_{i,2}))$. Because of fraud issues, all we know about the relation between $p_i$ and $p_i^\ast$ is the following:
Note that reported vote shares are equal to true vote shares in case both parties have observers present for vote count. Letting $X_{i,k}$ take value $1$ if party $k$ places an observer in district $i$ and zero otherwise, the correspondence characterizing the model is
By the central limit theorem, $\sqrt{n_i}(\hat p_i-p_i)$ has Gaussian limiting distribution with zero mean and variance $p_i(1-p_i)$. Call $c_{\alpha_i/2}$ the quantile of level $1-\alpha_i/2$ of the standard normal distribution. Call $\eta=\max_i\eta_i$ with $\eta_i=c_{\alpha_i/2}\sqrt{p_i(1-p_i)/n_i}$. Then the dilation defined for each $p$ by $J_n^\alpha(p)=[p-\eta,p+\eta]$ satisfies Assumption (ref) with $\alpha=\lim\sup_n\Pi_i\alpha_i$. The composition of the dilation $J_n^\alpha$ and the correspondence $G$ yields \[J_n^\alpha\circ G(\rho\vert X;\theta)=[X_1p^\ast(\rho)-\eta,X_2(1-p^\ast(\rho))+\eta].\] The region $\tilde \Theta_I$ containing all $\theta$ such that $\hat p\in J_n^\alpha\circ G(\rho\vert X,\theta)$ a.s. is therefore a valid confidence region for the identified set and can be computed efficiently using methods proposed in GH:2008a.
We have proposed a method to combine several sources of uncertainty, such as missing or corrupted data and structural incompleteness in the model through a composition of correspondences. We show that our composition theorem applies in particular to the construction of confidence regions in partially identified models of general form. In that case, the composition theorem is applied to the composition of the correspondence that defines the econometric structure and a dilation of the sample space that controls the significance level and allows to replace the unknown distribution of observable data by the empirical distribution of the sample in the characterization of compatibility between model and data. An important computational advantage of this method over previous proposed confidence regions for partially identified parameters is that the dilation is performed independently of the structural parameter, hence needs to be performed only once. The remaining search over the parameter space is purely deterministic. The dilation is obtained through a minimax matching procedure. It is equivalent to a uniform confidence band for the quantile process when the dimension of the endogenous variable is one, however, it has no parallel in higher dimensions. The method is shown to perform well in simulation experiments.
We thank Christian Bontemps, Gary Chamberlain, Victor Chernozhukov, Pierre-Andr\'e Chiappori, Ivar Ekeland, Rustam Ibragimov, Guido Imbens, Thierry Magnac, Francesca Molinari, Alexei Onatski, Geert Ridder, Bernard Salani\'e, participants at the “Semiparametric and Nonparametric Methods in Econometrics” conference in Oberwolfach and seminar participants at BU, Brown, CalTech, Chicago, \'Ecole polytechnique, Harvard, MIT Sloan, Northwestern, Toulouse, UCLA, UCSD and Yale for helpful comments (with the usual disclaimer). Both authors gratefully acknowledge financial support from NSF grant SES 0532398 and from Chaire AXA “Assurance des Risques Majeurs” and Chaire Soci\'et\'e G\'en\'erale “Risques Financiers”. Galichon's research is partly supported by Chaire EDF-Calyon “Finance and D\'{e}veloppement Durable” and FiME, Laboratoire de Finance des March\'es de l'Energie. Henry's research is also partly supported by SSHRC Grant 410-2010-242.
\printbibliography