Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
427,694 characters · 38 sections · 733 citation commands
Microeconometrics with Partial Identification
}
\thispagestyle{empty} \onehalfspacing
\pagenumbering{arabic}
Knowing the population distribution that data are drawn from, what can one learn about a parameter of interest? It has long been understood that assumptions about the data generating process (DGP) play a crucial role in answering this identification question at the core of all empirical research. Inevitably, assumptions brought to bear enjoy a varying degree of credibility. Some are rooted in economic theory (e.g., optimizing behavior) or in information available to the researcher on the DGP (e.g., randomization mechanisms). These assumptions can be argued to be highly credible. Others are driven by concerns for tractability and the desire to answer the identification question with a certain level of precision (e.g., functional form and distributional assumptions). These are arguably less credible.
Early on, koo:rei50 highlighted the importance of imposing restrictions based on prior knowledge of the phenomenon under analysis and some criteria of simplicity, but not for the purpose of identifiability of a parameter that the researcher happens to be interested in, stating (p. 169): “One might regard problems of identifiability as a necessary part of the specification problem. We would consider such a classification acceptable, provided the temptation to specify models in such a way as to produce identifiability of relevant characteristics is resisted."
Much work, spanning multiple fields, has been devoted to putting forward strategies to carry out empirical research while relaxing distributional, functional form, or behavioral assumptions. One example, embodied in the research program on semiparameteric and nonparametric methods, is to characterize sufficient sets of assumptions, that exclude many suspect ones --sometimes as many as possible-- to guarantee that point identification of specific economically interesting parameters attains. This literature is reviewed in, e.g., mat07,mat13, and is not discussed here.
Another example, embodied in the research program on Bayesian model uncertainty, is to specify multiple models (i.e., multiple sets of assumptions), put a prior on the parameters of each model and on each model, embed the various separate models within one large hierarchical mixture model, and obtain model posterior probabilities which can be used for a variety of inferences and decisions. This literature is reviewed in, e.g., was00 and cly:geo04, and is not discussed here.
The approach considered here fixes a set of assumptions and a parameter of interest a priori, in the spirit of koo:rei50, and asks what can be learned about that parameter given the available data, recognizing that even partial information can be illuminating for empirical research, while enjoying wider credibility thanks to the weaker assumptions imposed. The bounding methods at the core of this approach appeared in the literature nearly a century ago. Arguably, the first exemplar that leverages economic reasoning is given by the work of mar:and44. They provided bounds on Cobb-Douglas production functions in models of supply and demand, building on optimization principles and restrictions from microeconomic theory. lea81 revisited their analysis to obtain bounds on the elasticities of demand and supply in a linear simultaneous equations system with uncorrelated errors. The first exemplars that do not rely on specific economic models appear in gin21, fri34, and rei41, who bounded the coefficient of a simple linear regression in the presence of measurement error. These results were extended to the general linear regression model with errors in all variables by kle:lea84 and lea87.
This chapter surveys some of the methods proposed over the last thirty years in the microeconometrics literature to further this approach. These methods belong to the systematic program on partial identification analysis started with man89,man90,man95,man03,man07a,man13book and developed by several authors since the early 1990s. Within this program, the focus shifts from points to sets: the researcher aims to learn what is the set of values for the parameters of interest that can generate the same distribution of observables as the one in the data, for some DGP consistent with the maintained assumptions. In other words, the focus is on the set of observationally equivalent values, which henceforth I refer to as the parameters' sharp identification region. In the partial identification paradigm, empirical analysis begins with characterizing this set using the data alone. This is a nonparametric approach that dispenses with all assumptions, except basic restrictions on the sampling process such that the distribution of the observable variables can be learned as data accumulate. In subsequent steps, one incorporates additional assumptions into the analysis, reporting how each assumption (or set of assumptions) affects what one can learn about the parameters of interest, i.e., how it modifies and possibly shrinks the sharp identification region. Point identification may result from the process of increasingly strengthening the maintained assumptions, but it is not the goal in itself. Rather, the objective is to make transparent the relative role played by the data and the assumptions in shaping the inference that one draws.
There are several strands of independent, but thematically related literatures that are not discussed in this chapter. As a consequence, many relevant contributions are left out of the presentation and the references. One example is the literature in finance. han:jag91 developed nonparametric bounds for the admissible set for means and standard deviations of intertemporal marginal rates of substitution (IMRS) of consumers. The bounds were developed exploiting the condition, satisfied in many finance models, that the equilibrium price of any traded security equals the expectation (conditioned on current information) of the product's future payoff and the IMRS of any consumer.\footnote{ han:jag91 deduce a duality relation with the mean variance theory of mar52 and fam96, but the relation does not apply to the sharp bounds they derive. In the Arbitrage Pricing Model ros76, bounds on extensions of existing pricing functions, consistent with the absence of arbitrage opportunities, were considered by har:kre79 and kre81. } lut96 extended the analysis to economies with frictions. han:hea:lut95 developed econometric tools to estimate the regions, to assess asset pricing models, and to provide nonparametric characterizations of asset pricing anomalies. Earlier on, the existence of volatility bounds on IMRSs were noted by shi82 and han82comment. The bounding arguments that build on the minimum-volatility frontier for stochastic discount factors proposed by han:jag91 have become a litmus test to detect anomalies in asset pricing models shi03. I refer to the textbook presentations in lju:sar04 and coc05, and the review articles by fer03 and cam14, for a careful presentation of this literature.
In macroeconomics, fau98, can:den02, and uhl05 proposed bounds for impulse response functions in sign-restricted structural vector autoregression models, and carried out Bayesian inference with a non-informative prior for the non-identified parameters. I refer to kil:lut17 for a careful presentation of this literature.
In microeconomic theory, bounds were derived from inequalities resulting as necessary and sufficient conditions that data on an individual's choice need to satisfy in order to be consistent with optimizing behavior, as in the research pioneered by sam38 and advanced early on by hou50 and ric66. afr67 and var82 extended this research program to revealed preference extrapolation. Notably, in this work no stochastic terms enter the analysis. blo:mar60, mar60, hal73, mcf75, fal78, and mcf:ric91, extended revealed preference arguments to random utility models, and obtained bounds on the distributions of preferences. I refer to the survey articles by cra:der14 and blu19 for a careful presentation of this literature.
A complementary approach to partial identification is given by sensitivity analysis, advocated for in different ways by, e.g., gil:lea83, ros:rub83, lea85, ros95, imb03, and others. Within this approach, the analysis begins with a fully parametric model that point identifies the parameter of interest. One then reports the set of values for this parameter that result when the more suspicious assumptions are relaxed.
Related literatures, not discussed in this chapter, abound also outside Economics. For example, in probability theory, hoe40 and fre51 put forward bounds on the joint distributions of random variables, and mak81, rus82, and fra:nel:sch87 on the sum of random variables, when only marginal distributions are observed. The literature on probability bounds is discussed in the textbook by sho:wel09. Addressing problems faced in economics, sociology, epidemiology, geography, history, political science, and more, dun:dav53 derived bounds on correlations among variables measured at the individual level based on observable correlations among variables measured at the aggregate level. The so called ecological inference problem they studied, and the associated literature, is discussed in the survey article by cho:man09 and references therein.
To carry out econometric analysis with partial identification, one needs: (1) computationally feasible characterizations of the parameters' sharp identification region; (2) methods to estimate this region; and (3) methods to test hypotheses and construct confidence sets. The goal of this chapter is to provide insights into the challenges posed by each of these desiderata, and into some of their solutions. In order to discuss the partial identification literature in microeconometrics with some level of detail while keeping this chapter to a manageable length, I focus on a selection of papers and not on a complete survey of the literature. As a consequence, many relevant contributions are left out of the presentation and the references. I also do not discuss the important but separate topic of statistical decisions in the presence of partial identification, for which I refer to the textbook treatments in man05,man07a and to the review by hir:por19.
The presumption in identification analysis that the distribution from which the data are drawn is known allows one to keep separate the identification question from the distinct question of statistical inference from a finite sample. I use the same separation in this chapter. I assume solid knowledge of the topics covered in first year Economics PhD courses in econometrics and microeconomic theory.
I begin in Section (ref) with the analysis of what can be learned about features of probability distributions that are well defined in the absence of an economic model, such as moments, quantiles, cumulative distribution functions, etc., when one faces measurement problems. Specifically, I focus on cases where the data is incomplete, either due to selectively observed data or to interval measurements. I refer to man95,man03,man07a for textbook treatments of many other cases. I lay out formally the maintained assumptions for several examples, and then discuss in detail what is the source of the identification problem. I conclude with providing tractable characterizations of what can be learned about the parameters of interest, with formal proofs. I show that even in simple problems, great care may be needed to obtain the sharp identification region. It is often easier to characterize an outer region, i.e., a collection of values for the parameter of interest that contains the sharp one but may contain also additional values. Outer regions are useful because of their simplicity and because in certain applications they may suffice to answer questions of great interest, e.g., whether a policy intervention has a nonnegative effect. However, compared to the sharp identification region they may afford the researcher less useful predictions, and a lower ability to test for misspecification, because they do not harness all the information in the observed data and maintained assumptions.
In Section (ref) I use the same approach to study what can be learned about features of parameters of structural econometric models when the model is incomplete tam03,hai:tam03,cil:tam09. Specifically, I discuss single agent discrete choice models under a variety of challenging situations (interval measured as well as endogenous explanatory variables; unobserved as well as counterfactual choice sets); finite discrete games with multiple equilibria; auction models under weak assumptions on bidding behavior; and network formation models. Again I formally derive sharp identification regions for several examples.
I conclude each of these sections with a brief discussion of further theoretical advances and empirical applications that is meant to give a sense of the breadth of the approach, but not to be exhaustive. I refer to the recent survey by ho:ros17 for a thorough discussion of empirical applications of partial identification methods.
In Section (ref) I discuss finite sample inference. I limit myself to highlighting the challenges that one faces for consistent estimation when the identified object is a set, and several coverage notions and requirements that have been proposed over the last 20 years. I refer to the recent survey by can:sha17 for a thorough discussion of methods to tests hypotheses and build confidence sets in moment inequality models.
In Section (ref) I discuss the distinction between refutable and non-refutable assumptions, and how model misspecification may be detectable in the presence of the former, even within the partial identification paradigm. I then highlight certain challenges that model misspecification presents for the interpretation of sharp identification (as well as outer) regions, and for the construction of confidence sets.
In Section (ref) I highlight that while most of the sharp identification regions characterized in Section (ref) can be easily computed, many of the ones in Section (ref) are more challenging. This is because the latter are obtained as level sets of criterion functions in moderately dimensional spaces, and tracing out these level sets or their boundaries is a non-trivial computational problem. In Section (ref) I conclude providing some considerations on what I view as open questions for future research.
I refer to tam10 for an earlier review of this literature, and to lew18 for a careful presentation of the many notions of identification that are used across the econometrics literature, including an important historical account of how these notions developed over time.
Throughout Sections (ref) and (ref), a simple organizing principle for much of partial identification analysis emerges. The cause of the identification problems discussed can be traced back to a collection of random variables that are consistent with the available data and maintained assumptions. For the problems studied in Section (ref), this set is often a simple function of the observed variables. The incompleteness of the data stems from the fact that instead of observing the singleton variables of interest, one observes set-valued variables to which these belong, but one has no information on their exact value within the sets. For the problems studied in Section (ref), the collection of random variables consistent with the maintained assumptions comprises what the model predicts for the endogenous variable(s). The incompleteness of the model stems from the fact that instead of making a singleton prediction for the variable(s) of interest, the model makes multiple predictions but does not specify how one is chosen.
The central role of set-valued objects, both stochastic and nonstochastic, in partial identification renders random set theory a natural toolkit to aid the analysis.\footnote{The first idea of a general random set in the form of a region that depends on chance appears in kol50, originally published in 1933. For another early example where confidence regions are explicitly described as random sets, see haa44. The role of random sets in this chapter is different.} This theory originates in the seminal contributions of cho53, aum65, and deb67, with the first self contained treatment of the theory given by mat75. I refer to mo1 for a textbook presentation, and to mol:mol14,mol:mol18 for a treatment focusing on its applications in econometrics.
ber:mol08 introduce the use of random set theory in econometrics to carry out identification analysis and statistical inference with incomplete data. ber:mol:mol11,ber:mol:mol12 propose it to characterize sharp identification regions both with incomplete data and with incomplete models. gal:hen11 propose the use of optimal transportation methods that in some applications deliver the same characterizations as the random set methods. I do not discuss optimal transportation methods in this chapter, but refer to gal16 for a thorough treatment.
Over the last ten years, random set methods have been used to unify a number of specific results in partial identification, and to produce a general methodology for identification analysis that dispenses completely with case-by-case distinctions. In particular, as I show throughout the chapter, the methods allow for simple and tractable characterizations of sharp identification regions. The collection of these results establishes that indeed this is a useful tool to carry out econometrics with partial identification, as exemplified by its prominent role both in this chapter and in Chapter XXX in this Volume by che:ros19, which focuses on general classes of instrumental variable models. The random sets approach complements the more traditional one, based on mathematical tools for (single valued) random vectors, that proved extremely productive since the beginning of the research program in partial identification.
This chapter shows that to fruitfully apply random set theory for identification and inference, the econometrician needs to carry out three fundamental steps. First, she needs to define the random closed set that is relevant for the problem under consideration using all information given by the available data and maintained assumptions. This is a delicate task, but one that is typically carried out in identification analysis regardless of whether random set theory is applied. Indeed, throughout the chapter I highlight how relevant random closed sets were characterized in partial identification analysis since the early 1990s, albeit the connection to the theory of random sets was not made. As a second step, the econometrician needs to determine how the observable random variables relate to the random closed set. Often, one of two cases occurs: either the observable variables determine a random set to which the unobservable variable of interest belongs with probability one, as in incomplete data scenarios; or the (expectation of the) (un)observable variable belongs to (the expectation of) a random set determined by the model, as in incomplete model scenarios. Finally, the econometrician needs to determine which tool from random set theory should be utilized. To date, new applications of random set theory to econometrics have fruitfully exploited (Aumann) expectations and their support functions, (Choquet) capacity functionals, and laws of large numbers and central limit theorems for random sets. Appendix (ref) reports basic definitions from random set theory of these concepts, as well as some useful theorems. The chapter explains in detail through applications to important identification problems how these steps can be carried out.
This chapter employs consistent notation that is summarized in Table (ref). Some important conventions are as follows: ${\boldsymbol{y}}$ denotes outcome variables, $({\boldsymbol{x}},{\boldsymbol{w}})$ denote explanatory variables, and ${\boldsymbol{z}}$ denotes instrumental variables (i.e., variables that satisfy some form of independence with the outcome or with the unobservable variables, possibly conditional on ${\boldsymbol{x}},{\boldsymbol{w}}$).
I denote by $\mathsf{P}$ the joint distribution of all observable variables. Identification analysis is carried out using the information contained in this distribution, and finite sample inference is carried out under the presumption that one draws a random sample of size $n$ from $\mathsf{P}$. I denote by $\mathsf{Q}$ the joint distribution whose features the researcher wants to learn. If $\mathsf{Q}$ were identified given the observed data (e.g., if it were a marginal of $\mathsf{P}$), point identification of the parameter or functional of interest would attain. I denote by $\mathsf{R}$ the joint distribution of all variables, observable and unobservable ones; both $\mathsf{P}$ and $\mathsf{Q}$ can be obtained from it. In the context of structural models, I denote by $\mathsf{M}$ a distribution for the observable variables that is consistent with the model. I note that model incompleteness typically implies that $\mathsf{M}$ is not unique. I let $\mathcal{H}_\mathsf{P}[\cdot]$ denote the sharp identification region of the functional in square brackets, and $\mathcal{O}_\mathsf{P}[\cdot]$ an outer region. In both cases, the regions are indexed by $\mathsf{P}$, because they depend on the distribution of the observed data.
The literature reviewed in this chapter starts with the analysis of what can be learned about functionals of probability distributions that are well-defined in the absence of a model. The approach is nonparametric, and it is typically constructive, in the sense that it leads to “plug-in" formulae for the bounds on the functionals of interest.
As in man89, suppose that a researcher is interested in learning the probability that an individual who is homeless at a given date has a home six months later. Here the population of interest is the people who are homeless at the initial date, and the outcome of interest ${\boldsymbol{y}}$ is an indicator of whether the individual has a home six months later (so that ${\boldsymbol{y}}=1$) or remains homeless (so that ${\boldsymbol{y}}=0$). A random sample of homeless individuals is interviewed at the initial date, so that individual background attributes ${\boldsymbol{x}}$ are observed, but six months later only a subset of the individuals originally sampled can be located. In other words, attrition from the sample creates a selection problem whereby ${\boldsymbol{y}}$ is observed only for a subset of the population. Let ${\boldsymbol{d}}$ be an indicator of whether the individual can be located (hence ${\boldsymbol{d}}=1$) or not (hence ${\boldsymbol{d}}=0$). The question is what can the researcher learn about $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}}=x)$, with $\mathsf{Q}$ the distribution of $({\boldsymbol{y}},{\boldsymbol{x}})$? man89 showed that $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}}=x)$ is not point identified in the absence of additional assumptions, but informative nonparametric bounds on this quantity can be obtained. In this section I review his approach, and discuss several important extensions of his original idea.
Throughout the chapter, I formally state the structure of the problem under study as an “Identification Problem", and then provide a solution, either in the form of a sharp identification region, or of an outer region. To set the stage, and at the cost of some repetition, I do the same here, slightly generalizing the question stated in the previous paragraph.
man89's analysis of this problem begins with a simple application of the law of total probability, that yields
Equation (ref) lends a simple but powerful anatomy of the selection problem. While $\mathsf{P}({\boldsymbol{y}}|{\boldsymbol{x}}=x,{\boldsymbol{d}}=1)$ and $\mathsf{P}({\boldsymbol{d}}|{\boldsymbol{x}}=x)$ can be learned from the observable distribution $\mathsf{P}({\boldsymbol{y}}{\boldsymbol{d}},{\boldsymbol{d}},{\boldsymbol{x}})$, under the maintained assumptions the sampling process reveals nothing about $\mathsf{R}({\boldsymbol{y}}|{\boldsymbol{x}}=x,{\boldsymbol{d}}=0)$. Hence, $\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}}=x)$ is not point identified.
If one were to assume exogenous selection (or data missing at random conditional on ${\boldsymbol{x}}$), i.e., $\mathsf{R}({\boldsymbol{y}}|{\boldsymbol{x}},{\boldsymbol{d}}=0)=\mathsf{P}({\boldsymbol{y}}|{\boldsymbol{x}},{\boldsymbol{d}}=1)$, point identification would obtain. However, that assumption is non-refutable and it is well known that it may fail in applications.\footnote{Section (ref) discusses the consequences of model misspecification (with respect to refutable assumptions).} Let $\mathcal{T}$ denote the space of all probability measures with support in $\mathcal{Y}$. The unknown functional vector is $\{\tau(x),\upsilon(x)\}\equiv \{\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}}=x),\mathsf{R}({\boldsymbol{y}}|{\boldsymbol{x}}=x,{\boldsymbol{d}}=0)\}$. What the researcher can learn, in the absence of additional restrictions on $\mathsf{R}({\boldsymbol{y}}|{\boldsymbol{x}}=x,{\boldsymbol{d}}=0)$, is the region of observationally equivalent distributions for ${\boldsymbol{y}}|{\boldsymbol{x}}=x$, and the associated set of expectations taken with respect to these distributions.
These are the worst case bounds, so called because assumptions free and therefore representing the widest possible range of values for the parameter of interest that are consistent with the observed data. A simple “plug-in" estimator for $\mathcal{H}_\mathsf{P}[\mathbb{E}_\mathsf{Q}(g({\boldsymbol{y}})|{\boldsymbol{x}}=x)]$ replaces all unknown quantities in (ref) with consistent estimators, obtained, e.g., by kernel or sieve regression. I return to consistent estimation of partially identified parameters in Section (ref). Here I emphasize that identification problems are fundamentally distinct from finite sample inference problems. The latter are typically reduced as sample size increase (because, e.g., the variance of the estimator becomes smaller). The former do not improve, unless a different and better type of data is collected, e.g. with a smaller prevalence of missing data dom:man17.
man03 shows that the proof of Theorem SIR-(ref) can be extended to obtain the smallest and largest points in the sharp identification region of any parameter that respects stochastic dominance.\footnote{ Recall that a probability distribution $\mathsf{F}\in\mathcal{T}$ stochastically dominates $\mathsf{F}^\prime\in\mathcal{T}$ if $\mathsf{F}(-\infty,t]\le \mathsf{F}^\prime(-\infty,t]$ for all $t\in\mathbb{R}$. A real-valued functional $\mathsf{d}:\mathcal{T}\to\mathbb{R}$ respects stochastic dominance if $\mathsf{d}(\mathsf{F})\ge \mathsf{d}(\mathsf{F}^\prime)$ whenever $\mathsf{F}$ stochastically dominates $\mathsf{F}^\prime$.} This is especially useful to bound the quantiles of ${\boldsymbol{y}}|{\boldsymbol{x}}=x$. For any given $\alpha \in (0,1)$, let $\mathsf{q}_{\mathsf{P}}^{g({\boldsymbol{y}})}(\alpha,1,x)\equiv \left\{\min t:\mathsf{P}(g({\boldsymbol{y}})\le t|{\boldsymbol{d}}=1,{\boldsymbol{x}}=x)\ge \alpha\right\}$. Then the smallest and largest admissible values for the $\alpha$-quantile of $g({\boldsymbol{y}})|{\boldsymbol{x}}=x$ are, respectively,
The lower bound on $\mathbb{E}_\mathsf{Q}(g({\boldsymbol{y}})|{\boldsymbol{x}}=x)$ is informative only if $g_0>-\infty$, and the upper bound is informative only if $g_1<\infty$. By comparison, for any value of $\alpha$, $r(\alpha,x)$ and $s(\alpha,x)$ are generically informative if, respectively, $\mathsf{P}({\boldsymbol{d}}=1|{\boldsymbol{x}}=x) > 1-\alpha$ and $\mathsf{P}({\boldsymbol{d}}=1|{\boldsymbol{x}}=x) \ge \alpha$, regardless of the range of $g$.
sto10 further extends partial identification analysis to the study of spread parameters in the presence of missing data (as well as interval data, data combinations, and other applications). These parameters include ones that respect second order stochastic dominance, such as the variance, the Gini coefficient, and other inequality measures, as well as other measures of dispersion which do not respect second order stochastic dominance, such as interquartile range and ratio.\footnote{ Earlier related work includes, e.g., gas72 and cow91, who obtain worst case bounds on the sample Gini coefficient under the assumption that one knows the income bracket but not the exact income of every household.} sto10 shows that the sharp identification region for these parameters can be obtained by fixing the mean or quantile of the variable of interest at a specific value within its sharp identification region, and deriving a distribution consistent with this value which is “compressed" with respect to the ones which bound the cumulative distribution function (CDF) of the variable of interest, and one which is “dispersed" with respect to them. Heuristically, the compressed distribution minimizes spread, while the dispersed one maximizes it (the sense in which this optimization occurs is formally defined in the paper). The intuition for this is that a compressed CDF is first below and then above any non-compressed one; a dispersed CDF is first above and then below any non-dispersed one. Second-stage optimization over the possible values of the mean or the quantile delivers unconstrained bounds. The main results of the paper are sharp identification regions for the expectation and variance, for the median and interquartile ratio, and for many other combinations of parameters.
Despite how transparent the framework in Identification Problem (ref) is, important subtleties arise even in this seemingly simple context. For a given $t\in\mathbb{R}$, consider the function $g({\boldsymbol{y}})=\mathbf{1}({\boldsymbol{y}}\le t)$, with $\mathbf{1}(A)$ the indicator function taking the value one if the logical condition in parentheses holds and zero otherwise. Then equation (ref) yields pointwise-sharp bounds on the CDF of ${\boldsymbol{y}}$ at any fixed $t\in\mathbb{R}$:
Yet, the collection of CDFs that belong to the band defined by (ref) is not the sharp identification region for the CDF of ${\boldsymbol{y}}|{\boldsymbol{x}}=x$. Rather, it constitutes an outer region, as originally pointed out by man94.
How can one characterize the sharp identification region for the CDF of ${\boldsymbol{y}}|{\boldsymbol{x}}=x$ under the assumptions in Identification Problem (ref)? In general, there is not a single answer to this question: different methodologies can be used. Here I use results in man03 and mol:mol18, which yield an alternative characterization of $\mathcal{H}_\mathsf{P}[\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}}=x)]$ that translates directly into a characterization of $\mathcal{H}_\mathsf{P}[\mathsf{F}({\boldsymbol{y}}|{\boldsymbol{x}}=x)]$.\footnote{Whereas man94 is very clear that the collection of CDFs in (ref) is an outer region for the CDF of ${\boldsymbol{y}}|{\boldsymbol{x}}=x$, and man03 provides the sharp characterization in (ref), man07a does not state all the requirements that characterize $\mathcal{H}_\mathsf{P}[\mathsf{F}({\boldsymbol{y}}|{\boldsymbol{x}}=x)]$.}
This section provides sharp identification regions and outer regions for a variety of functionals of interest. The computational complexity of these characterizations varies widely. Sharp bounds on parameters that respect stochastic dominance only require computing the parameters with respect to two probability distributions. An outer region on the CDF can be obtained by evaluating all tail probabilities of a certain distribution. A sharp identification region on the CDF requires evaluating the probability that a certain distribution assigns to all intervals. I return to computational challenges in partial identification in Section (ref).
The discussion of partial identification of probability distributions of selectively observed data naturally leads to the question of its implications for program evaluation. The literature on program evaluation is vast. The purpose of this section is exclusively to show how the ideas presented in Section (ref) can be applied to learn features of treatment effects of interest, when no assumptions are imposed on treatment selection and outcomes. I also provide examples of assumptions that can be used to tighten the bounds. To keep this chapter to a manageable length, I discuss only partial identification of the average response to a treatment and of the average treatment effect (ATE). There are many different parameters that received much interest in the literature. Examples include the local average treatment effect of imb:ang94 and the marginal treatment effect of hec:vyt99,hec:vyt01,hec:vyt05. For thorough discussions of the literature on program evaluation, I refer to the textbook treatments in man95,man03,man07a and imb:rub15, to the Handbook chapters by hec:vyt07I,hec:vyt07II and abb:hec07, and to the review articles by imb:woo09 and mog:tor18.
Using standard notation ney23, let ${\boldsymbol{y}}:\mathbb{T} \mapsto \mathcal{Y}$ be an individual-specific response function, with $\mathbb{T}=\{0,1,\dots,T\}$ a finite set of mutually exclusive and exhaustive treatments, and let ${\boldsymbol{s}}$ denote the individual's received treatment (taking its realizations in $\mathbb{T}$).\footnote{Here the treatment response is a function only of the (scalar) treatment received by the given individual, an assumption known as stable unit treatment value assumption rub78.} The researcher observes data $({\boldsymbol{y}},{\boldsymbol{s}},{\boldsymbol{x}})\sim\mathsf{P}$, with ${\boldsymbol{y}}\equiv{\boldsymbol{y}}({\boldsymbol{s}})$ the outcome corresponding to the received treatment ${\boldsymbol{s}}$, and ${\boldsymbol{x}}$ a vector of covariates. The outcome ${\boldsymbol{y}}(t)$ for ${\boldsymbol{s}}\neq t$ is counterfactual, and hence can be conceptualized as missing. Therefore, we are in the framework of Identification Problem (ref) and all the results from Section (ref) apply in this context too, subject to adjustments in notation.\footnote{ber:mol:mol12 and mol:mol18 provide a characterization of the sharp identification region for the joint distribution of $[{\boldsymbol{y}}(t),t\in\mathbb{T}]$.} For example, using Theorem SIR-(ref),
where $y_0\equiv\inf_{y\in\mathcal{Y}}y$, $y_1\equiv\sup_{y\in\mathcal{Y}}y$. If $y_0<\infty$ and/or $y_1<\infty$, these worst case bounds are informative. When both are infinite, the data is uninformative in the absence of additional restrictions.
If the researcher is interested in an Average Treatment Effect (ATE), e.g.
with $t_0,t_1\in\mathbb{T}$, sharp worst case bounds on this quantity can be obtained as follows. First, observe that the empirical evidence reveals $\mathbb{E}_\mathsf{P}({\boldsymbol{y}}|{\boldsymbol{x}}=x,{\boldsymbol{s}}=t_j)$ and $\mathsf{P}({\boldsymbol{s}}|{\boldsymbol{x}}=x)$, but is uninformative about $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t_j)|{\boldsymbol{x}}=x,{\boldsymbol{s}}\neq t_j)$, $j=0,1$. Each of the latter quantities (the expectations of ${\boldsymbol{y}}(t_0)$ and ${\boldsymbol{y}}(t_1)$ conditional on different realizations of ${\boldsymbol{s}}$ and ${\boldsymbol{x}}=x$) can take any value in $[y_0,y_1]$. Hence, the sharp lower bound on the ATE is obtained by subtracting the upper bound on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t_0)|{\boldsymbol{x}}=x)$ from the lower bound on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t_1)|{\boldsymbol{x}}=x)$. The sharp upper bound on the ATE is obtained by subtracting the lower bound on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t_0)|{\boldsymbol{x}}=x)$ from the upper bound on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t_1)|{\boldsymbol{x}}=x)$. The resulting bounds have width equal to $(y_1-y_0)[2-\mathsf{P}({\boldsymbol{s}}=t_1|{\boldsymbol{x}}=x)-\mathsf{P}({\boldsymbol{s}}=t_0|{\boldsymbol{x}}=x)]\in[(y_1-y_0),2(y_1-y_0)]$, and hence are informative only if both $y_0>-\infty$ and $y_1<\infty$. As the largest logically possible value for the ATE (in the absence of information from data) cannot be larger than $(y_1-y_0)$, and the smallest cannot be smaller than $-(y_1-y_0)$, the sharp bounds on the ATE always cover zero.
What assumptions may researchers bring to bear to learn more about treatment effects of interest? The literature has provided a wide array of well motivated and useful restrictions. Here I consider two examples. The first one entails shape restrictions on the treatment response function, leaving selection unrestricted. man97:monotone obtains bounds on treatment effects under the assumption that the response functions are monotone, semi-monotone, or concave-monotone. These restrictions are motivated by economic theory, where it is commonly presumed, e.g., that demand functions are downward sloping and supply functions are upward sloping. Let the set $\mathbb{T}$ be ordered in terms of degree of intensity. Then man97:monotone's monotone treatment response assumption requires that
Under this assumption, one has a sharp characterization of what can be learned about ${\boldsymbol{y}}(t)$:
Hence, the sharp bounds on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t)|{\boldsymbol{x}}=x)$ are man97:monotone
This finding highlights some important facts. Under the monotone treatment response assumption, the bounds on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t)|{\boldsymbol{x}}=x)$ are obtained using information from all $({\boldsymbol{y}},{\boldsymbol{s}})$ pairs (given ${\boldsymbol{x}}=x$), while the bounds in (ref) only use the information provided by $({\boldsymbol{y}},{\boldsymbol{s}})$ pairs for which ${\boldsymbol{s}}=t$ (given ${\boldsymbol{x}}=x$). As a consequence, the bounds in (ref) are informative even if $\mathsf{P}({\boldsymbol{s}}= t|{\boldsymbol{x}}=x)=0$, whereas the worst case bounds are not.
Concerning the ATE with $t_1>t_0$, under monotone treatment response its lower bound is zero, and its upper bound is obtained by subtracting the lower bound on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t_0)|{\boldsymbol{x}}=x)$ from the upper bound on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t_1)|{\boldsymbol{x}}=x)$, where both bounds are obtained as in (ref) man97:monotone.
The second example of assumptions used to tighten worst case bounds is that of exclusion restrictions, as in, e.g., man90. Suppose the researcher observes a random variable ${\boldsymbol{z}}$, taking its realizations in $\mathcal{Z}$, such that\footnote{Stronger exclusion restrictions include statistical independence of the response function at each $t$ with ${\boldsymbol{z}}$: $\mathsf{Q}({\boldsymbol{y}}(t)|{\boldsymbol{z}},{\boldsymbol{x}})=\mathsf{Q}({\boldsymbol{y}}(t)|{\boldsymbol{x}})~\forall t \in\mathbb{T},~{\boldsymbol{x}}$-a.s.; and statistical independence of the entire response function with ${\boldsymbol{z}}$: $\mathsf{Q}([{\boldsymbol{y}}(t),t \in\mathbb{T}]|{\boldsymbol{z}},{\boldsymbol{x}})=\mathsf{Q}([{\boldsymbol{y}}(t),t \in\mathbb{T}]|{\boldsymbol{x}}),~{\boldsymbol{x}}$-a.s. Examples of partial identification analysis under these conditions can be found in bal:pea97, man03, kit09, ber:mol:mol12, mac:sha:vyt18, and many others.}
This assumption is treatment-specific, and requires that the treatment response to $t$ is mean independent with ${\boldsymbol{z}}$. It is easy to show that under the assumption in (ref), the bounds on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t)|{\boldsymbol{x}}=x)$ become
These are called intersection bounds because they are obtained as follows. Given ${\boldsymbol{x}}$ and ${\boldsymbol{z}}$, one uses (ref) to obtain sharp bounds on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t)|{\boldsymbol{z}}=z,{\boldsymbol{x}}=x)$. Due to the mean independence assumption in (ref), $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t)|{\boldsymbol{x}}=x)$ must belong to each of these bounds ${\boldsymbol{z}}$-a.s., hence to their intersection. The expression in (ref) follows. If the instrument affects the probability of being selected into treatment, or the average outcome for the subpopulation receiving treatment $t$, the bounds on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t)|{\boldsymbol{x}}=x)$ shrink. If the bounds are empty, the mean independence assumption can be refuted (see Section (ref) for a discussion of misspecification in partial identification). man:pep00,man:pep09 generalize the notion of instrumental variable to monotone instrumental variable, and show how these can be used to obtain tighter bounds on treatment effect parameters.\footnote{See che:ros19 for further discussion.} They also show how shape restrictions and exclusion restrictions can jointly further tighten the bounds. man13social generalizes these findings to the case where treatment response may have social interactions -- that is, each individual's outcome depends on the treatment received by all other individuals.
Identification Problem (ref), as well as the treatment evaluation problem in Section (ref), is an instance of the more general question of what can be learned about (functionals of) probability distributions of interest, in the presence of interval valued outcome and/or covariate data. Such data have become commonplace in Economics. For example, since the early 1990s the Health and Retirement Study collects income data from survey respondents in the form of brackets, with degenerate (singleton) intervals for individuals who opt to fully reveal their income jus:suz95. Due to concerns for privacy, public use tax data are recorded as the number of tax payers which belong to each of a finite number of cells pic05. The Occupational Employment Statistics (OES) program at the Bureau of Labor Statistics BLS collects wage data from employers as intervals, and uses these data to construct estimates for wage and salary workers in more than 800 detailed occupations. man:mol10 and giu:man:mol19round document the extensive prevalence of rounding in survey responses to probabilistic expectation questions, and propose to use a person's response pattern across different questions to infer his rounding practice, the result being interpretation of reported numerical values as interval data. Other instances abound. Here I focus first on the case of interval outcome data.
It is immediate to obtain the sharp identification region
As in the previous section, it is also easy to obtain sharp bounds on parameters that respect stochastic dominance, and pointwise-sharp bounds on the CDF of ${\boldsymbol{y}}$ at any fixed $t\in\mathbb{R}$:
In this case too, however, as in Theorem OR-(ref), the tube of CDFs satisfying equation (ref) for all $t\in\mathbb{R}$ is an outer region for the CDF of ${\boldsymbol{y}}|{\boldsymbol{x}}=x$, rather than its sharp identification region. Indeed, also in this context it is easy to construct examples similar to Example (ref).
How can one characterize the sharp identification region for the probability distribution of ${\boldsymbol{y}}|{\boldsymbol{x}}$ when one observes $({\boldsymbol{y}}_{\mathrm{L}},{\boldsymbol{y}}_{\mathrm{U}},{\boldsymbol{x}})$ and assumes $\mathsf{R}({\boldsymbol{y}}_{\mathrm{L}}\le{\boldsymbol{y}}\le{\boldsymbol{y}}_{\mathrm{U}})=1$? Again, there is not a single answer to this question. Depending on the specific problem at hand, e.g., the specifics of the interval data and whether ${\boldsymbol{y}}$ is assumed discrete or continuous, different methods can be applied. I use random set theory to provide a characterization of $\mathcal{H}_\mathsf{P}[\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}}=x)]$. Let
Then ${\boldsymbol{Y}}$ is a random closed set according to Definition (ref).\footnote{For a proof of this statement, see mol:mol18.} The requirement $\mathsf{R}({\boldsymbol{y}}_{\mathrm{L}}\le{\boldsymbol{y}}\le{\boldsymbol{y}}_{\mathrm{U}})=1$ can be equivalently expressed as
Equation (ref), together with knowledge of $\mathsf{P}$, exhausts all the information in the data and maintained assumptions. In order to harness such information to characterize the set of observationally equivalent probability distributions for ${\boldsymbol{y}}$, one can leverage a result due to art83 nor92, reported in Theorem (ref) in Appendix (ref), which allows one to translate (ref) into a collection of conditional moment inequalities. Specifically, let $\mathcal{T}$ denote the space of all probability measures with support in $\mathcal{Y}$.
Compare equation (ref) with equation (ref). Under the set-up of Identification Problem (ref), when ${\boldsymbol{d}}=1$ we have ${\boldsymbol{Y}}=\{{\boldsymbol{y}}\}$ and when ${\boldsymbol{d}}=0$ we have ${\boldsymbol{Y}}=\mathcal{Y}$. Hence, for any $K \subsetneq \mathcal{Y}$, $\mathsf{P}({\boldsymbol{Y}} \subset K|{\boldsymbol{x}}=x)=\mathsf{P}({\boldsymbol{y}}\in K|{\boldsymbol{x}}=x,{\boldsymbol{d}}=1)\mathsf{P}({\boldsymbol{d}}=1)$.\footnote{For $K = \mathcal{Y}$, both (ref) and (ref) hold trivially.} It follows that the characterizations in (ref) and (ref) are equivalent. If $\mathcal{Y}$ is countable, it is easy to show that (ref) simplifies to (ref) ber:mol:mol12.
An attractive feature of the characterization in (ref) is that it holds regardless of the specific assumptions on ${\boldsymbol{y}}_{\mathrm{L}},\,{\boldsymbol{y}}_{\mathrm{U}}$, and $\mathcal{Y}$. Later sections in this chapter illustrate how Theorem (ref) delivers the sharp identification region in other more complex instances of partial identification of probability distributions, as well as in structural models. In Chapter XXX in this Volume, che:ros19 apply Theorem (ref) to obtain sharp identification regions for functionals of interest in the important class of generalized instrumental variable models. To avoid repetitions, I do not systematically discuss that class of models in this chapter.
When addressing questions about features of $\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}}=x)$ in the presence of interval outcome data, an alternative approach tam10,pon:tam11 looks at all (random) mixtures of ${\boldsymbol{y}}_{\mathrm{L}},{\boldsymbol{y}}_{\mathrm{U}}$. The approach is based on a random variable ${\boldsymbol{u}}$ (a selection mechanism that picks an element of ${\boldsymbol{Y}}$) with values in $[0,1]$, whose distribution conditional on ${\boldsymbol{y}}_{\mathrm{L}},{\boldsymbol{y}}_{\mathrm{U}}$ is left completely unspecified. Using this random variable, one defines
The sharp identification region in Theorem SIR-(ref) can be characterized as the collection of conditional distributions of all possible random variables ${\boldsymbol{y}}_{\boldsymbol{u}}$ as defined in (ref), given ${\boldsymbol{x}}=x$. This is because each ${\boldsymbol{y}}_{\boldsymbol{u}}$ is a (stochastic) convex combination of ${\boldsymbol{y}}_{\mathrm{L}},{\boldsymbol{y}}_{\mathrm{U}}$, hence each of these random variables satisfies $\mathsf{R}({\boldsymbol{y}}_{\mathrm{L}}\le{\boldsymbol{y}}_{\boldsymbol{u}}\le{\boldsymbol{y}}_{\mathrm{U}})=1$. While such characterization is sharp, it can be of difficult implementation in practice, because it requires working with all possible random variables ${\boldsymbol{y}}_{\boldsymbol{u}}$ built using all possible random variables ${\boldsymbol{u}}$ with support in $[0,1]$. Theorem (ref) allows one to bypass the use of ${\boldsymbol{u}}$, and obtain directly a characterization of the sharp identification region for $\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}}=x)$ based on conditional moment inequalities.\footnote{It can be shown that the collection of random variables ${\boldsymbol{y}}_{\boldsymbol{u}}$ equals the collection of measurable selections of the random closed set ${\boldsymbol{Y}}\equiv [{\boldsymbol{y}}_{\mathrm{L}},{\boldsymbol{y}}_{\mathrm{U}}]$ (see Definition (ref)); see ber:mol:mol11. Theorem (ref) provides a characterization of the distribution of any ${\boldsymbol{y}}_{\boldsymbol{u}}$ that satisfies ${\boldsymbol{y}}_{\boldsymbol{u}} \in {\boldsymbol{Y}}$ a.s., based on a dominance condition that relates the distribution of ${\boldsymbol{y}}_{\boldsymbol{u}}$ to the distribution of the random set ${\boldsymbol{Y}}$. Such dominance condition is given by the inequalities in (ref). }
hor:man98,hor:man00 study nonparametric conditional prediction problems with missing outcome and/or missing covariate data. Their analysis shows that this problem is considerably more pernicious than the case where only outcome data are missing. For the case of interval covariate data, man:tam02 provide a set of sufficient conditions under which simple and elegant sharp bounds on functionals of $\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}})$ can be obtained, even in this substantially harder identification problem. Their assumptions are listed in Identification Problem (ref), and their result (with proof) in Theorem SIR-(ref).
Compared to the earlier discussion for the interval outcome case, here there are two additional assumptions. The monotonicity condition (M) is a simple shape restrictions, which however requires some prior knowledge about the joint distribution of $({\boldsymbol{y}},{\boldsymbol{x}})$. The mean independence restriction (MI) requires that if ${\boldsymbol{x}}$ were observed, knowledge of $({\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})$ would not affect the conditional expectation of ${\boldsymbol{y}}|{\boldsymbol{x}}$. The assumption is not innocuous, as pointed out by the authors. For example, it may fail if censoring is endogenous.\footnote{For the case of missing covariate data, which is a special case of interval covariate data similarly to arguments in footnote (ref), auc:bug:hot17 show that the MI restriction implies the assumption that data is missing at random.}
Learning about functionals of $\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}}=x)$ naturally implies learning about predictors of ${\boldsymbol{y}}|{\boldsymbol{x}}=x$. For example, $\mathcal{H}_\mathsf{P}[\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}}=x)]$ yields the collection of values for the best predictor under square loss; $\mathcal{H}_\mathsf{P}[\mathbb{M}_\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}}=x)]$, with $\mathbb{M}_\mathsf{Q}$ the median with respect to distribution $\mathsf{Q}$, yields the collection of values for the best predictor under absolute loss. And so on. A related but distinct problem is that of parametric conditional prediction. Often researchers specify not only a loss function for the prediction problem, but also a parametric family of predictor functions, and wish to learn the member of this family that minimizes expected loss. To avoid confusion, let me clarify that here I am not referring to a parametric assumption on the best predictor, e.g., that $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}})$ is a linear function of ${\boldsymbol{x}}$. I return to such assumptions at the end of this section. For now, in the example of linearity and square loss, I am referring to best linear prediction, i.e., best linear approximation to $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}})$. man03 discusses what can be learned about the best linear predictor of ${\boldsymbol{y}}$ conditional on ${\boldsymbol{x}}$, when only interval data on $({\boldsymbol{y}},{\boldsymbol{x}})$ is available.
I treat first the case of interval outcome and perfectly observed covariates.
For simplicity suppose that ${\boldsymbol{x}}$ is a scalar, and let $\theta=[\theta_0~\theta_1]^\top\in\Theta\subset\mathbb{R}^2$ denote the parameter vector of the best linear predictor of ${\boldsymbol{y}}|{\boldsymbol{x}}$. Assume that $Var({\boldsymbol{x}})>0$. Combining the definition of best linear predictor with a characterization of the sharp identification region for the joint distribution of $({\boldsymbol{y}},{\boldsymbol{x}})$, we have that
where, using an argument similar to the one in Theorem SIR-(ref),
ber:mol08 show that (ref) can be re-written in an intuitive way that generalizes the well-known formula for the best linear predictor that arises when ${\boldsymbol{y}}$ is perfectly observed. Define the random segment ${\boldsymbol{G}}$ and the matrix $\Sigma_\mathsf{P}$ as
where $\operatorname{Sel}({\boldsymbol{Y}})$ is the set of all measurable selections from ${\boldsymbol{Y}}$, see Definition (ref). Then,
In either representation (ref) or (ref), $\mathcal{H}_\mathsf{P}[\theta]$ is the collection of best linear predictors for each selection of ${\boldsymbol{Y}}$.\footnote{Under our assumption that $\mathcal{Y}$ is a bounded interval, all the selections of ${\boldsymbol{Y}}$ are integrable. ber:mol08 consider the more general case where $\mathcal{Y}$ is not required to be bounded.} Why should one bother with the representation in (ref)? The reason is that $\mathcal{H}_\mathsf{P}[\theta]$ is a convex set, as it can be evinced from representation (ref): ${\boldsymbol{G}}$ has almost surely convex realizations that are segments and the Aumann expectation of a convex set is convex.\footnote{In $\mathbb{R}^2$ in our example, in $\mathbb{R}^d$ if ${\boldsymbol{x}}$ is a $d-1$ vector and the predictor includes an intercept.} Hence, it can be equivalently represented through its support function $h_{\mathcal{H}_\mathsf{P}[\theta]}$, see Definition (ref) and equation (ref). In particular, in this example,
where $f({\boldsymbol{x}},u)\equiv [1~{\boldsymbol{x}}]\Sigma_\mathsf{P}^{-1}u$.\footnote{See ber:mol08 and bon:mag:mau12.} The characterization in (ref) results from Theorem (ref), which yields $h_{\mathcal{H}_\mathsf{P}[\theta]}(u)=h_{\Sigma_\mathsf{P}^{-1} \mathbb{E}_\mathsf{P}{\boldsymbol{G}}}(u)=\mathbb{E}_\mathsf{P} h_{\Sigma_\mathsf{P}^{-1} {\boldsymbol{G}}}(u)$, and the fact that $\mathbb{E}_\mathsf{P} h_{\Sigma_\mathsf{P}^{-1} {\boldsymbol{G}}}(u)$ equals the expression in (ref). As I discuss in Section (ref) below, because the support function fully characterizes the boundary of $\mathcal{H}_\mathsf{P}[\theta]$, (ref) allows for a simple sample analog estimator, and for inference procedures with desirable properties. It also immediately yields sharp bounds on linear combinations of $\theta$ by judicious choice of $u$.\footnote{For example, in the case that ${\boldsymbol{x}}$ is a scalar, sharp bounds on $\theta_1$ can be obtained by choosing $u=[0~1]^\top$ and $u=[0~-1]^\top$, which yield $\theta_1\in[\theta_{1L},\theta_{1U}]$ with $\theta_{1L}=\min_{{\boldsymbol{y}}\in[{\boldsymbol{y}}_{\mathrm{L}},{\boldsymbol{y}}_{\mathrm{U}}]}\frac{Cov({\boldsymbol{x}},{\boldsymbol{y}})}{Var({\boldsymbol{x}})}=\frac{\mathbb{E}_\mathsf{P}[({\boldsymbol{x}}-\mathbb{E}_\mathsf{P}{\boldsymbol{x}})({\boldsymbol{y}}_{\mathrm{L}}\mathbf{1}({\boldsymbol{x}} >\mathbb{E}_\mathsf{P}{\boldsymbol{x}})+{\boldsymbol{y}}_{\mathrm{U}}\mathbf{1}({\boldsymbol{x}}\le\mathbb{E}{\boldsymbol{x}}))]}{\mathbb{E}_\mathsf{P}{\boldsymbol{x}}^2-(\mathbb{E}_\mathsf{P}{\boldsymbol{x}})^2}$ and $\theta_{1U}=\max_{{\boldsymbol{y}}\in[{\boldsymbol{y}}_{\mathrm{L}},{\boldsymbol{y}}_{\mathrm{U}}]}\frac{Cov({\boldsymbol{x}},{\boldsymbol{y}})}{Var({\boldsymbol{x}})}=\frac{\mathbb{E}_\mathsf{P}[({\boldsymbol{x}}-\mathbb{E}_\mathsf{P}{\boldsymbol{x}})({\boldsymbol{y}}_{\mathrm{L}}\mathbf{1}({\boldsymbol{x}} <\mathbb{E}_\mathsf{P}{\boldsymbol{x}})+{\boldsymbol{y}}_{\mathrm{U}}\mathbf{1}({\boldsymbol{x}}\ge\mathbb{E}{\boldsymbol{x}}))]}{\mathbb{E}_\mathsf{P}{\boldsymbol{x}}^2-(\mathbb{E}_\mathsf{P}{\boldsymbol{x}})^2}$.} sto07 and mag:mau08 provide the same characterization as in (ref) using, respectively, direct optimization and the Frisch-Waugh-Lovell theorem.
A natural generalization of Identification Problem (ref) allows for both outcome and covariate data to be interval valued.
Abstractly, $\mathcal{H}_\mathsf{P}[\theta]$ is as given in (ref), with
replacing (ref) by an application of Theorem (ref). While this characterization is sharp, it is cumbersome to apply in practice, see hor:man:pon:sto03.
On the other hand, when both ${\boldsymbol{y}}$ and ${\boldsymbol{x}}$ are perfectly observed, the best linear predictor is simply equal to the parameter vector that yields a mean zero prediction error that is uncorrelated with ${\boldsymbol{x}}$. How can this basic observation help in the case of interval data? The idea is that one can use the same insight applied to the set-valued data, and obtain $\mathcal{H}_\mathsf{P}[\theta]$ as the collection of $\theta$'s for which there exists a selection $(\tilde{{\boldsymbol{y}}},\tilde{{\boldsymbol{x}}}) \in \operatorname{Sel}({\boldsymbol{Y}} \times {\boldsymbol{X}})$, and associated prediction error $\varepsilon_\theta=\tilde{{\boldsymbol{y}}}-\theta_0-\theta_1 \tilde{{\boldsymbol{x}}}$, satisfying $\mathbb{E}_\mathsf{P} \varepsilon_\theta=0$ and $\mathbb{E}_\mathsf{P} (\varepsilon_\theta \tilde{{\boldsymbol{x}}})=0$ ber:mol:mol11.\footnote{Here for simplicity I suppose that both ${\boldsymbol{x}}_{\mathrm{L}}$ and ${\boldsymbol{x}}_{\mathrm{U}}$ have bounded support. ber:mol:mol11 do not make this simplifying assumption.} To obtain the formal result, define the $\theta$-dependent set\footnote{Note that while ${\boldsymbol{G}}$ is a convex set, $\mathcal{E}_\theta$ is not.} \[\mathcal{E}_\theta = \left\lbrace
\: : \, (\tilde{{\boldsymbol{y}}},\tilde{{\boldsymbol{x}}}) \in \operatorname{Sel}({\boldsymbol{Y}} \times{\boldsymbol{X}}) \right\rbrace. \]
The support function $h_{\mathcal{E}_\theta}(u)$ is an easy to calculate convex sublinear function of $u$, regardless of whether the variables involved are continuous or discrete. The optimization problem in ((ref)), determining whether $\theta \in \mathcal{H}_\mathsf{P}[\theta]$, is a convex program, hence easy to solve. See for example the CVX software by gra:boy10. It should be noted, however, that the set $\mathcal{H}_\mathsf{P}[\theta]$ itself is not necessarily convex. Hence, tracing out its boundary is non-trivial. I discuss computational challenges in partial identification in Section (ref).
I conclude this section by discussing parametric regression. man:tam02 study identification of parametric regression models under the assumptions in Identification Problem (ref); Theorem SIR-(ref) below reports the result. The proof is omitted because it follows immediately from the proof of Theorem SIR-(ref).
auc:bug:hot17 study Identification Problem (ref) for the case of missing covariate data without imposing the mean independence restriction of man:tam02 (Assumption MI in Identification Problem (ref)). As discussed in footnote (ref), restriction MI is undesirable in this context because it implies the assumption that data are missing at random. auc:bug:hot17 characterize $\mathcal{H}_\mathsf{P}[\theta]$ under the weaker assumptions, but face the problem that this characterization is usually too complex to compute or to use for inference. They therefore provide outer regions that are easier to compute, and they show that these regions are informative and relatively easy to use.
One of the first examples of bounding analysis appears in fri34, to assess the impact in linear regression of covariate measurement error. This analysis was substantially extended in gil:lea83, kle:lea84, and lea87. The more recent literature in partial identification has provided important advances to learn features of probability distributions when the observed variables are error-ridden measures of the variables of interest. Here I briefly mention some of the papers in this literature, and refer to Chapter XXX in this Volume by sch19 for a thorough treatment of identification and inference with mismeasured and unobserved variables. In an influential paper, hor:man95 study what can be learned about features of the distribution of ${\boldsymbol{y}}|{\boldsymbol{x}}$ in the presence of contaminated or corrupted outcome data. Whereas a contaminated sampling model assumes that data errors are statistically independent of sample realizations from the population of interest, the corrupted sampling model does not. These models are regularly used in the important literature on robust estimation hub64,hub04,ham:ron:rou:sta11. However, the goal of that literature is to characterize how point estimators of population parameters behave when data errors are generated in specified ways. As such, the inference problem is approached ex-ante: before collecting the data, one looks for point estimators that are not greatly affected by error. The question addressed by hor:man95 is conceptually distinct. It asks what can be learned about specific population parameters ex-post, that is, after the data has been collected. For example, whereas the mean is well known not to be a robust estimator in the presence of contaminated data, hor:man95 show that it can be (non-trivially) bounded provided the probability of contamination is strictly less than one. dom:she04,dom:she05 and kre:pep07,kre:pep08 extend the results of hor:man95 to allow for (partial) verification of the distribution from which the data are drawn. They apply the resulting sharp bounds to learn about school performance when the observed test scores may not be valid for all students. mol08 provides sharp bounds on the distribution of a misclassified outcome variable under an array of different assumptions on the extent and type of misclassification.
A completely different problem is that of data combination. Applied economists often face the problem that no single data set contains all the variables that are necessary to conduct inference on a population of interest. When this is the case, they need to integrate the information contained in different samples; for example, they might need to combine survey data with administrative data rid:mof07. From a methodological perspective, the problem is that while the samples being combined might contain some common variables, other variables belong only to one of the samples. When the data is collected at the same aggregation level (e.g., individual level, household level, etc.), if the common variables include a unique and correctly recorded identifier of the units constituting each sample, and there is a substantial overlap of units across all samples, then exact matching of the data sets is relatively straightforward, and the combined data set provides all the relevant information to identify features of the population of interest. However, it is rather common that there is a limited overlap in the units constituting each sample, or that variables that allow identification of units are not available in one or more of the input files, or that one sample provides information at the individual or household level (e.g., survey data) while the second sample provides information at a more aggregate level (e.g., administrative data providing information at the precinct or district level). Formally, the problem is that one observes data that identify the joint distributions $\mathsf{P}({\boldsymbol{y}},{\boldsymbol{x}})$ and $\mathsf{P}({\boldsymbol{x}},{\boldsymbol{w}})$, but not data that identifies the joint distribution $\mathsf{Q}({\boldsymbol{y}},{\boldsymbol{x}},{\boldsymbol{w}})$ whose features one wants to learn. The literature on statistical matching has aimed at using the common variable(s) ${\boldsymbol{x}}$ as a bridge to create synthetic records containing $({\boldsymbol{y}},{\boldsymbol{x}},{\boldsymbol{w}})$ okn72. As sim72 points out, the inherent assumption at the base of statistical matching is that conditional on ${\boldsymbol{x}}$, ${\boldsymbol{y}}$ and ${\boldsymbol{w}}$ are independent. This conditional independence assumption is strong and untestable. While it does guarantee point identification of features of the conditional distributions $\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}},{\boldsymbol{w}})$, it often finds very little justification in practice. Early on, dun:dav53 provided numerical illustrations on how one can bound the object of interest, when both ${\boldsymbol{y}}$ and ${\boldsymbol{w}}$ are binary variables. cro:man02 provide a general analysis of the problem. They obtain bounds on the long regression $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}},{\boldsymbol{w}})$, under the assumption that ${\boldsymbol{w}}$ has finite support. They show that sharp bounds on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}},{\boldsymbol{w}}=w)$ can be obtained using the results in hor:man95, thereby establishing a connection with the analysis of contaminated data. They then derive sharp identification regions for $[\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}}=x,{\boldsymbol{w}}=w),x\in\mathcal{X},w\in\mathcal{W}]$. They show that these bounds are sharp when ${\boldsymbol{y}}$ has finite support, and mol:pes06 establish sharpness without this restriction. fan:she:shu14 address the question of what can be learned about counterfactual distributions and treatment effects under the data scenario just described, but with ${\boldsymbol{x}}$ replaced by ${\boldsymbol{s}}$, a binary indicator for the received treatment (using the notation of the previous section). In this case, the exogenous selection assumption (conditional on ${\boldsymbol{w}}$) does not suffice for point identification of the objects of interest. The authors derive, however, sharp bounds on these quantities using monotone rearrangement inequalities. pac17 provides partial identification results for the coefficients in the linear projection of ${\boldsymbol{y}}$ on $({\boldsymbol{x}},{\boldsymbol{w}})$.
In order to discuss the partial identification approach to learning features of probability distributions in some level of detail while keeping this chapter to a manageable length, I have focused on a selection of papers. In this section I briefly mention several other excellent theoretical contributions that could be discussed more closely, as well as several papers that have applied partial identification analysis to answer important empirical questions.
While selectively observed data are commonplace in observational studies, in randomized experiments subjects are randomly placed in designated treatment groups conditional on ${\boldsymbol{x}}$, so that the assumption of exogenous selection is satisfied with respect to the assigned treatment. Yet, identification of some highly policy relevant parameters can remain elusive in the absence of strong assumptions. One challenge results from noncompliance, where individuals' received treatments differs from the randomly assigned ones. bal:pea97 derive sharp bounds on the ATE in this context, when $\mathcal{Y}=\mathbb{T}=\{0,1\}$. Even if one is interested in the intention-to-treat parameter, selectively observed data may continue to be a problem. For example, lee09 studies the wage effects of the Job Corps training program, which randomly assigns eligibility to participate in the program. Individuals randomized to be eligible were not compelled to receive treatment, hence lee09 focuses on the intention-to-treat effect. Because wages are only observable when individuals are employed, a selection problem persists despite the random assignment of eligibility to treatment, as employment status may be affected by the training program. lee09 obtains sharp bounds on the intention-to-treat effect, through a trimming procedure that leverages results in hor:man95. mol08MT analyzes the problem of identification of the ATE and other treatment effects, when the received treatment is unobserved for a subset of the population. Missing treatment data may be due to item or survey nonresponse in observational studies, or noncompliance with randomly assigned treatments that are not directly monitored. She derives sharp worst case bounds leveraging results in hor:man95, and she shows that these are a function of the available prior information on the distribution of missing treatments. If the response function is assumed monotone as in (ref), she obtains informative bounds without restrictions on the distribution of missing treatments.
Even randomly assigned treatments and perfect compliance with no missing data may not suffice for point identification of all policy relevant parameters. Important examples are given by hec:smi:cle97 and man97:mixing. hec:smi:cle97 show that features of the joint distribution of the potential outcomes of treatment and control, including the distribution of treatment effects impacts, cannot be point identified in the absence of strong restrictions. This is because although subjects are randomized to treatment and control, nobody's outcome is observed under both states. Nonetheless, the authors obtain bounds for the functionals of interest. mul18 derives related bounds on the probability that the potential outcome of one treatment is larger than that of the other treatment, and applies these results to health economics problems. man97:mixing shows that features of outcome distributions under treatment rules in which treatment may vary within groups cannot be point identified in the absence of strong restrictions. This is because data resulting from randomized experiments with perfect compliance allow for point identification of the outcome distributions under treatment rules that assign all persons with the same ${\boldsymbol{x}}$ to the same treatment group. However, such data only allow for partial identification of outcome distributions under rules in which treatment may vary within groups. man97:mixing derives sharp bounds for functionals of these distributions.
Analyses of data resulting from natural experiments also face identification challenges. hot:mul:san97 study what can be learned about treatment effects when one uses a contaminated instrumental variable, i.e. when a mean-independence assumption holds in a population of interest, but the observed population is a mixture of the population of interest and one in which the assumption doesn't hold. They extend the results of hor:man95 to learn about the causal effect of teenage childbearing on a teen mother's subsequent outcomes, using the natural experiment of miscarriages to form an instrumental variable for teen births. This instrument is contaminated because miscarriges may not occur randomly for a subset of the population (e.g., higher miscarriage rates are associated with smoking and drinking, and these behaviors may be correlated with the outcomes of interest).
Of course, analyses of selectively observed data present many challenges, including but not limited to the ones described in Section (ref). ath:imb06 generalize the difference-in-difference (DID) design to a changes-in-changes (CIC) model, where the distribution of the unobservables is allowed to vary across groups, but not overtime within groups, and the additivity and linearity assumptions of the DID are dispensed with. For the case that the outcomes have a continuous distribution, ath:imb06 provide conditions for point identification of the entire counterfactual distribution of effects of the treatment on the treatment group as well as the distribution of effects of the treatment on the control group, without restricting how these distributions differ from each other. For the case that the outcome variables are discrete, they provide partial identification results, as well as additional conditions compared to their baseline model under which point identification attains.
Motivated by the question of whether the age-adjusted mortality rate from cancer in 2000 was lower than that in the early 1970s, hon:lle06 study partial identification of competing risk models pet76. To answer this question, they need to contend with the fact that mortality rate from cardiovascular disease declined substantially over the same period of time, so that individuals that in the early 1970s might have died from cardiovascular disease before being diagnosed with cancer, do not in 2000. In this context, it is important to carry out the analysis without assuming that the underlying risks are independent. hon:lle06 show that bounds for the parameters of interest can be obtained as the solution to linear programming problems. The estimated bounds suggest much larger improvements in cancer mortality rates than previously estimated.
blu:gos:ich:meg07 use UK data to study changes over time in the distribution of male and female wages, and in wage inequality. Because the composition of the workforce changes over time, it is difficult to disentangle that effect from changes in the distribution of wages, given that the latter are observed only for people in the workforce. blu:gos:ich:meg07 begin their empirical analysis by reporting worst case bounds man94 on the CDF of wages conditional on covariates. They then consider various restrictions on treatment selection, e.g., a first order stochastic dominance assumption according to which people with higher wages are more likely to work, and derive tighter bounds under this assumption (and under weaker ones). Finally, they bring to bear shape restrictions. At each step of the analysis, they report the resulting bounds, thereby illuminating the role played by each assumption in shaping the inference. cha:che:mol:sch18 provide best linear approximations to the identification region for the quantile gender wage gap using Current Population Survey repeated cross-sections data from 1975-2001, using treatment selection assumptions in the spirit of blu:gos:ich:meg07 as well as exclusion restrictions.
bha:sha:vyt12 study the effect of Swan-Ganz catheterization on subsequent mortality.\footnote{The Swan-Ganz catheter is a device placed in patients in the intensive care unit to guide therapy.} Previous research had shown, using propensity score matching (assuming that there are no unobserved differences between catheterized and non catheterized patients) that Swan-Ganz catheterization increases the probability that patients die within 180 days from admission to the intensive care unit. bha:sha:vyt12 re-analyze the data using (and extending) bounds results obtained by sha:vyt11. These results are based on exclusion restrictions combined with a threshold crossing structure for both the treatment and the outcome variables in problems where $\mathcal{Y}=\mathcal{T}=\{0,1\}$. bha:sha:vyt12 use as instrument for Swan-Ganz catheterization the day of the week that the patient was admitted to the intensive care unit. The reasoning is that patients are less likely to be catheterized on the weekend, but the admission day to the intensive care unit is plausibly uncorrelated with subsequent mortality. Their results confirm that for some diagnoses, Swan-Ganz catheterization increases mortality at 30 days after catheterization and beyond.
man:pep18 use data from Maryland, Virginia and Illinois to learn about the impact of laws allowing individuals to carry concealed handguns (right-to-carry laws) on violent and property crimes. Point identification of these treatment effects is possible under invariance assumptions that certain features of treatment response are constant across states and years. man:pep18 propose the use of weaker but more credible restrictions according to which these features exhibit bounded variation -- the invariance case being the limit where the bound on variation equals zero. They carry out their analysis under different combinations of the bounded variation assumptions, and at each step they report the resulting bounds, thereby illuminating the role played by each assumption in shaping the inference.
mou:hen:mea18 provide sharp bounds on the joint distribution of potential (binary) outcomes in a Roy model with sector specific unobserved heterogeneity and self selection based on potential outcomes. The key maintained assumption is that the researcher has access to data that includes a stochastically monotone instrumental variable. This is a selection shifter that is restricted to affect potential outcomes monotonically. An example is parental education, which may not be independent from potential wages, but plausibly does not negatively affect future wages. Under this assumption, mou:hen:mea18 show that all observable implications of the model are embodied in the stochastic monotonicity of observed outcomes in the instrument, hence Roy selection behavior can be tested by checking this stochastic monotonicity. They apply the method to estimate a Roy model of college major choice in Canada and Germany, with special interest in the under-representation of women in STEM.
mog:san:tor18 provide a general method to obtain sharp bounds on a certain class of treatment effects parameters. This class is comprised of parameters that can be expressed as weighted averages of marginal treatment effects hec:vyt99,hec:vyt01,hec:vyt05. tor19pies provides a general method, based on copulas, to obtain sharp bounds on treatment effect parameters in semiparametric binary models. A notable feature of both mog:san:tor18 and tor19pies is that the bounds are obtained as solutions to convex (even linear) optimization problems, rendering them computationally attractive. fre:hor14 provide partial identification results and inference methods for a linear functional $\ell(g)$ when $g:\mathcal{X}\mapsto\mathbb{R}$ is such that ${\boldsymbol{y}}=g({\boldsymbol{x}})+\epsilon$ and $\mathbb{E}({\boldsymbol{y}}|{\boldsymbol{z}})=0$. The instrumental variable ${\boldsymbol{z}}$ and regressor ${\boldsymbol{x}}$ have discrete distributions, and ${\boldsymbol{z}}$ has fewer points of support than ${\boldsymbol{x}}$, so that $\ell(g)$ can only be partially identified. They impose shape restrictions on $g$ (e.g., monotonicity or convexity) to achieve interval identification of $\ell(g)$, and they show that the lower and upper points of the interval can be obtained by solving linear programming problems. They also show that the bootstrap can be used to carry out inference.
In this section I focus on the literature concerned with learning features of structural econometric models. These are models where economic theory is used to postulate relationships among observable outcomes ${\boldsymbol{y}}$, observable covariates ${\boldsymbol{x}}$, and unobservable variables $\nu$. For example, economic theory may guide assumptions on economic behavior (e.g., utility maximization) and equilibrium that yield a mapping from $({\boldsymbol{x}},\nu)$ to ${\boldsymbol{y}}$. The researcher is interested in learning features of these relationships (e.g., utility function, distribution of preferences), and to this end may supplement the data and economic theory with functional form assumptions on the mapping of interest and distributional assumptions on the observable and unobservable variables.
The earlier literature on partial identification of features of structural models includes important examples of nonparametric analysis of random utility models and revealed preference extrapolation, e.g. blo:mar60, mar60, hal73, mcf75, fal78, mcf:ric91, and others. The earlier literature also addresses semiparametric analysis, where the underlying models are specified up to parameters that are finite dimensional (e.g., preference parameters) and parameters that are infinite dimensional (e.g., distribution functions); important examples include mar:and44, mar52, fis66, har:kre79, kre81, lea81, man88, jov89, phi89, han:jag91, han:hea:lut95, lut96, and others. Contrary to the nonparametric bounds results discussed in Section (ref), and especially in the case of semiparametric models, structural partial identification often yields an identification region that is not constructive.\footnote{Of course, this is not always the case, as exemplified by the bounds in han:jag91.} Indeed, the boundary of the set is not obtained in closed form as a functional of the distribution of the observable data. Rather, the identification region can often be characterized as a level set of a properly specified criterion function.
The recent spark of interest in partial identification of structural microeconometric models was fueled by the work of man:tam02, tam03 and cil:tam09, and hai:tam03. Each of these papers has advanced the literature in fundamental ways, studying conceptually very distinct problems. man:tam02 are concerned with partial identification of the decision process yielding binary outcomes in a semiparametric model, when one of the explanatory variables is interval valued. Hence, the root cause of the identification problem they study is that the data is incomplete.\footnote{man:tam02 study also partial identification (and estimation) of nonparametric, semiparametric, and parametric conditional expectation functions that are well defined in the absence of a structural model, when one of the conditioning variables is interval valued. I refer to Section (ref) for a discussion.}
tam03 and cil:tam09 are concerned with identification (and estimation) of simultaneous equation models with dummy endogeneous variables which are representations of two-player entry games with multiple equilibria.\footnote{cil:tam09 consider more general multi-player entry games.} hai:tam03 are concerned with nonparametric identification and estimation of the distribution of valuations in a model of English auctions under weak assumptions on bidders' behavior. In both cases, the root cause of the identification problem is that the structural model is incomplete. This is because the model makes multiple predictions for the observed outcome variables (respectively: the players' actions; and the bidders' bids), but does not specify how one of them is selected to yield the observed data.
Set-valued predictions for the observable outcome (endogenous variables) are a key feature of partially identified structural models. The goal of this section is to explain how they result in a wide array of theoretical frameworks, and how sharp identification regions can be characterized using a unified approach based on random set theory. Although the work of man:tam02, tam03 and cil:tam09, and hai:tam03 has spurred many of the developments discussed in this section, for pedagogical reasons I organize the presentation based on application topic rather than chronologically. The work of pak10 and pak:por:ho:ish15 further stimulated a large empirical literature that applies partial identification methods to a wide array of questions of substantive economic importance, to which I return in Section (ref).
Let $\mathcal{I}$ denote a population of decision makers and $\mathcal{Y}=\{c_1,\dots,c_{|\mathcal{Y}|}\}$ a finite universe of potential alternatives (feasible set henceforth). Let $\mathfrak{U}$ be a family of real valued functions defined over the elements of $\mathcal{Y}$. Let $\in^* $ denote “is chosen from." Then observed choice is consistent with a random utility model if there exists a function $\mathfrak{u}_i$ drawn from $\mathfrak{U}$ according to some probability distribution, such that $\mathbb{P}(c \in^* C)=\mathbb{P}(\mathfrak{u}_i(c) \ge \mathfrak{u}_i(b)~\forall b \in C)$ for all $c\in C$, all non empty sets $C \subset \mathcal{Y}$, and all $i\in\mathcal{I}$ blo:mar60. See man07a for a textbook presentation of this class of models, and mat07 for a review of sufficient conditions for point identification of nonparametric and semiparametric limited dependent variables models.
As in the seminal work of mcf73, assume that the decision makers and alternatives are characterized by observable and unobservable vectors of real valued attributes. Denote the observable attributes by ${\boldsymbol{x}}_i \equiv \{{\boldsymbol{x}}_i^1,({\boldsymbol{x}}_{ic}^2,c\in\mathcal{Y})\},i\in\mathcal{I}$. These include attribute vectors ${\boldsymbol{x}}_i^1$ that are specific to the decision maker, as well as attribute vectors ${\boldsymbol{x}}_{ic}^2$ that include components that are specific to the alternative and components that are indexed by both. Denote the unobservable attributes (preferences) by $\nu_i\equiv(\zeta_i,\{\epsilon_{ic},~c\in\mathcal{Y}\}),i\in\mathcal{I}$. These are idiosyncratic to the decision maker and similarly may include alternative and decision maker specific terms. Denote $\mathcal{X},\mathcal{V}$ the supports of ${\boldsymbol{x}},\nu$, respectively.
In what follows, I label “standard" a random utility model that maintains some form of exogeneity for ${\boldsymbol{x}}_i$ (e.g., mean or quantile or statistical independence with $\nu_i$) and presupposes observation of data that include $\{({\boldsymbol{C}}_i,{\boldsymbol{y}}_i,{\boldsymbol{x}}_i):{\boldsymbol{y}}_i \in^* {\boldsymbol{C}}_i\}, i=1,\dots,n$, with ${\boldsymbol{C}}_i$ the choice set faced by decision maker $i$ and $|{\boldsymbol{C}}_i|\ge 2$ man75. Often it is also assumed that all members of the population face the same choice set, ${\boldsymbol{C}}_i=D$ for all $i\in\mathcal{I}$ and some known $D\subseteq\mathcal{Y}$, although this requirement is not critical to identification analysis.
man:tam02 provide inference methods for nonparametric, semiparametric, and parametric conditional expectation functions when one of the conditioning variables is interval valued. I have discussed their nonparametric and parametric sharp bounds on conditional expectations with interval valued covariates in Identification Problems (ref) and (ref), and Theorems SIR-(ref) and SIR-(ref), respectively. Here I focus on their analysis of semiparametric binary choice models. Compared to the generic notation set forth at the beginning of Section (ref), I let ${\boldsymbol{C}}_i=\mathcal{Y}=\{0,1\}$ for all $i\in\mathcal{I}$, and with some abuse of notation I denote the vector of observed covariates $({\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}},{\boldsymbol{w}})$.
Compared to Identification Problem (ref) (see p. (ref)), here one continues to impose ${\boldsymbol{x}}\in[{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}}]$ a.s. The sign restriction on $\delta$ replaces the monotonicity restriction (M) in Identification Problem (ref), but does not imply it unless the distribution of $\epsilon$ is independent of ${\boldsymbol{x}}$ conditional on ${\boldsymbol{w}}$. The quantile independence restriction is inspired by man85.
For given $\theta\in\Theta$, this model yields set valued predictions because ${\boldsymbol{y}}=1$ can occur whenever $\epsilon> -{\boldsymbol{w}}\theta-{\boldsymbol{x}}_{\mathrm{U}}$, whereas ${\boldsymbol{y}}=0$ can occur whenever $\epsilon\le -{\boldsymbol{w}}\theta-{\boldsymbol{x}}_{\mathrm{L}}$, and $-{\boldsymbol{w}}\theta-{\boldsymbol{x}}_{\mathrm{U}} \le -{\boldsymbol{w}}\theta-{\boldsymbol{x}}_{\mathrm{L}}$. Conversely, observation of ${\boldsymbol{y}}=1$ allows one to conclude that $\epsilon\in(-{\boldsymbol{w}}\theta-{\boldsymbol{x}}_{\mathrm{U}},+\infty)$, whereas observation of ${\boldsymbol{y}}=0$ allows one to conclude that $\epsilon\in(-\infty,-{\boldsymbol{w}}\theta-{\boldsymbol{x}}_{\mathrm{L}}]$, and these regions of possible realizations of $\epsilon$ overlap. In contrast, when ${\boldsymbol{x}}$ is observed the prediction is unique because the value $-{\boldsymbol{w}}\theta-{\boldsymbol{x}}$ partitions the space of realizations of $\epsilon$ in two disjoint sets, one associated with ${\boldsymbol{y}}=1$ and the other with ${\boldsymbol{y}}=0$. Figure (ref) depicts the model's set-valued predictions for ${\boldsymbol{y}}$ given $({\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})$ as a function of $\epsilon$, and the model's set valued predictions for $\epsilon$ given $({\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})$ as a function of ${\boldsymbol{y}}$.\footnote{Figure (ref) is based on Figure 1 in man:tam02. See che:ros19 for an extensive discussion of the duality between the model's set valued predictions for ${\boldsymbol{y}}$ as a function of $\epsilon$ and for $\epsilon$ as a function of ${\boldsymbol{y}}$, in both cases given the observed covariates.}
Why does this set-valued prediction hinder point identification? The reason is that the distribution of the observable data relates to the model structure in an incomplete manner. The model predicts $\mathsf{M}({\boldsymbol{y}}=1|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})=\int \mathsf{R}({\boldsymbol{y}}=1|{\boldsymbol{w}},{\boldsymbol{x}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})d\mathsf{R}({\boldsymbol{x}}|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})=\int \mathsf{R}(\epsilon>-{\boldsymbol{w}}\theta-{\boldsymbol{x}}|{\boldsymbol{w}},{\boldsymbol{x}})d\mathsf{R}({\boldsymbol{x}}|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}}),~({\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})$-a.s. Because the distribution $\mathsf{R}({\boldsymbol{x}}|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})$ is left completely unspecified, one can find multiple values for $(\theta,\mathsf{R}({\boldsymbol{x}}|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}}),\mathsf{R}(\epsilon|{\boldsymbol{w}},{\boldsymbol{x}}))$, satisfying the assumptions in Identification Problem (ref), such that $\mathsf{M}({\boldsymbol{y}}=1|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})=\mathsf{P}({\boldsymbol{y}}=1|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}}),~({\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})$-a.s. Nonetheless, in general, not all values of $\theta\in\Theta$ can be paired with some $\mathsf{R}({\boldsymbol{x}}|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})$ and $\mathsf{R}(\epsilon|{\boldsymbol{w}},{\boldsymbol{x}})$ so that they are compatible with $\mathsf{P}({\boldsymbol{y}}=1|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}}),~({\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})$-a.s. and with the maintained assumptions. Hence, $\theta$ can be partially identified using the information in the model and observed data.
Revisiting man:tam02's man:tam02 study of Identification Problem (ref) nearly 20 years later yields important insights on the differences between point and partial identification analysis. It is instructive to take as a point of departure the analysis of man85, which under the additional assumption that $({\boldsymbol{y}},{\boldsymbol{w}},{\boldsymbol{x}})$ is observed yields
In this case, $\theta$ is identified relative to $\vartheta\in\Theta$ if
man:tam02 extend this reasoning to the case that ${\boldsymbol{x}}$ is unobserved, but known to satisfy ${\boldsymbol{x}}\in [{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}}]$ a.s. The first part of their analysis, collected in their Proposition 2, characterizes the collection of values that cannot be distinguished from $\theta$ on the basis of $\mathsf{P}({\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})$ alone, through a clear generalization of (ref):
It is worth emphasizing that the characterization in (ref) depends on $\theta$, and makes no use of the information in $\mathsf{P}({\boldsymbol{y}}|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})$. The Corollary to Proposition 2 yields conditions on $\mathsf{P}({\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})$ under which either the sign of components of $\theta$, or $\theta$ itself, can be identified, regardless of the distribution of ${\boldsymbol{y}}|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}}$.
man:tam02 provide a second characterization, which presupposes knowledge of $\mathsf{P}({\boldsymbol{y}},{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})$, yields a set smaller than the one in (ref), and coincides with the result in Theorem SIR-(ref). man:tam02 use the same notation for the two sets, although the sets are conceptually and mathematically distinct.\footnote{This was confirmed in personal communication with Chuck Manski and Elie Tamer.} The result in Theorem SIR-(ref) is due to man:tam02, but the proof provided here is new, as is the use of random set theory in this application.\footnote{The proof closes a gap in the argument in man:tam02 connecting their Proposition 2 and Lemma 1, due to the fact that for a given $\vartheta$ the sets $\{({\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}}):\, \{{\boldsymbol{w}}\theta+{\boldsymbol{x}}_{\mathrm{U}}\le 0<{\boldsymbol{w}}\vartheta+{\boldsymbol{x}}_{\mathrm{L}}\} \cup \{{\boldsymbol{w}}\vartheta+{\boldsymbol{x}}_{\mathrm{U}}\le 0<{\boldsymbol{w}}\theta+{\boldsymbol{x}}_{\mathrm{L}}\}\}$ and $\{({\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}}):\, \{0<{\boldsymbol{w}}\vartheta+{\boldsymbol{x}}_{\mathrm{L}}\cap \mathsf{P}({\boldsymbol{y}}=1|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})\le 1-\alpha\}\cup \{{\boldsymbol{w}}\vartheta+{\boldsymbol{x}}_{\mathrm{U}}\le 0\cap \mathsf{P}({\boldsymbol{y}}=1|{\boldsymbol{w}},{\boldsymbol{x}}_{\mathrm{L}},{\boldsymbol{x}}_{\mathrm{U}})> 1-\alpha\}\}$ need not coincide, with the former being a subset of the latter due to part (c) of the proof of Proposition 2 in man:tam02.}
man10 studies random expected utility models, where agents choose the alternative that maximizes their expected utility. The core difference with standard models is that man10 does not fully specify the subjective beliefs that agents use to form their expectations, but only a set of such beliefs. man10 shows that the resulting, partially identified, discrete choice model can be formulated similarly to how man:tam02 treat interval valued covariates, and leverages their results to obtain bounds on preference parameters.\footnote{ber:mol:mol11 extend the analysis of man:tam02 to multinomial choice models with interval covariates.}
mag:mau08 consider a different but closely related model to the semiparametric binary response model studied by man:tam02. They assume that an instrumental variable ${\boldsymbol{z}}$ is available, that $\epsilon$ is independent of ${\boldsymbol{x}}$ conditional on $({\boldsymbol{w}},{\boldsymbol{z}})$, and that $Corr({\boldsymbol{z}},\epsilon)=0$. They assume that the distribution of ${\boldsymbol{x}}$ is absolutely continuous with support $[v_1,v_k]$, and that ${\boldsymbol{x}}$ is not a deterministic linear function of $({\boldsymbol{w}},{\boldsymbol{z}})$. They consider the case that ${\boldsymbol{x}}$ is unobserved but known to belong to one of the fixed (and known) intervals $[v_i,v_{i+1})$, $i=1,\dots,k-1$, with $\mathsf{R}[{\boldsymbol{x}}\in[v_i,v_{i+1})|{\boldsymbol{w}},{\boldsymbol{z}}]>0$ almost surely for all $i$. Finally, they assume that $(-{\boldsymbol{w}}\theta-\epsilon)\in [v_1,v_k]$ with probability one. They do not, however, make quantile independence assumptions.
Their point of departure is the fact that under these conditions, if ${\boldsymbol{x}}$ were observed, one could employ a transformation proposed by lew00 for the binary outcome ${\boldsymbol{y}}$, such that $\theta$ can be identified through a simple linear moment condition. Specifically, let
where $f_{\boldsymbol{x}}(\cdot|{\boldsymbol{w}},{\boldsymbol{z}})$ is the conditional density function of ${\boldsymbol{x}}$. Then, using the assumption that ${\boldsymbol{z}}$ and $\epsilon$ are uncorrelated, one has
With interval valued ${\boldsymbol{x}}$, mag:mau08 denote by ${\boldsymbol{x}}^*$ the random variable that takes value $i\in\{1,\dots,k-1\}$ if ${\boldsymbol{x}}\in[v_i,v_{i+1})$, so that the observed data are draws from the joint distribution of $({\boldsymbol{y}},{\boldsymbol{w}},{\boldsymbol{z}},{\boldsymbol{x}}^*)$. They let $\delta({\boldsymbol{x}}^*)=v_{{\boldsymbol{x}}^*+1}-v_{{\boldsymbol{x}}^*}$ denote the length of the ${\boldsymbol{x}}^*$-th interval, and define the transformed outcome variable:
The assumptions on ${\boldsymbol{x}}$ yield that, given ${\boldsymbol{z}}$ and ${\boldsymbol{w}}$, $\epsilon$ does not depend on ${\boldsymbol{x}}^*$. Moreover, $\mathsf{P}({\boldsymbol{y}}=1|{\boldsymbol{x}}^*,{\boldsymbol{w}},{\boldsymbol{z}})$ is non-decreasing in ${\boldsymbol{x}}^*$ and $\mathsf{F}_\epsilon(\cdot|{\boldsymbol{z}},{\boldsymbol{w}},{\boldsymbol{x}},{\boldsymbol{x}}^*)=\mathsf{F}_\epsilon(\cdot|{\boldsymbol{z}},{\boldsymbol{w}})$. mag:mau08 show that the sharp identification region for $\theta$ is
where $\mathbb{E}_\mathsf{P}({\boldsymbol{z}} {\boldsymbol{y}}^* + {\boldsymbol{z}} {\boldsymbol{U}})$ is the Aumann (or selection) expectation of the random interval ${\boldsymbol{z}} {\boldsymbol{y}}^* + {\boldsymbol{z}} {\boldsymbol{U}}$, see Definition (ref), with
In this expression, $r_{{\boldsymbol{x}}^*}({\boldsymbol{w}},{\boldsymbol{z}})\equiv\mathsf{P}({\boldsymbol{y}}=1|{\boldsymbol{x}}^*,{\boldsymbol{w}},{\boldsymbol{z}})$ and by convention $r_0({\boldsymbol{w}},{\boldsymbol{z}})=0$ and $r_K({\boldsymbol{w}},{\boldsymbol{z}})=1$, see mag:mau08. If $r_i({\boldsymbol{w}},{\boldsymbol{z}}),i=0,\dots,k$, were observed, this characterization would be very similar to the one provided by ber:mol08 for Identification Problem (ref), see equation (ref). However, these random functions need to be estimated. While the first-stage estimation of $r_i({\boldsymbol{w}},{\boldsymbol{z}}),i=0,\dots,k$, does not affect the identification arguments, it does complicate inference, see cha:che:mol:sch18 and the discussion in Section (ref).
Whereas the standard random utility model presumes some form of exogeneity for ${\boldsymbol{x}}$, in practice often some explanatory variables are endogenous. This problem has been addressed in the literature to obtain point identification of the model through a combination of several assumptions, including large support conditions, special regressors, control function restrictions, and more mat93,ber:lev:pak95,lew00,pet:tra10. hon:tam03 analyze the distinct but related problem of identification in a censored regression model with endogeneous explanatory variables, and provide sufficient conditions for point identification.\footnote{The estimator that they propose extends the minimum distance estimator put forward by man:tam02, see Section (ref), so that if the conditions required for point identification do not hold, it estimates the parameter's identification region (under regularity conditions). hon:tam03letters carry out a similar analysis for the binary choice model with endogenous explanatory variables.}
Here I discuss how to carry out identification analysis in the absence of such assumptions when instrumental variables ${\boldsymbol{z}}$ are available, as proposed by che:ros:smo13. They consider a more general case than I do here, with utility function that is not parametrically specified and not restricted to be separable in the unobservables. Even in that more general case, the identification analysis follows through similar steps as reported here.
The key challenge to identification here results because the distribution of $\nu$ can vary across different values of ${\boldsymbol{x}}$, both conditional and unconditional on ${\boldsymbol{z}}$. Why does this fact hinder point identification? For a given $\vartheta\in\Theta$ and for any $c\in\mathcal{Y}$ and $x\in\mathcal{X}$, the model yields that $c$ is optimal, and hence chosen, if and only if $\nu$ realizes in the set
Figure (ref) plots the set $\mathcal{E}_\vartheta({\boldsymbol{y}},{\boldsymbol{x}})$ in a stylized example with $\mathcal{Y}=\{1,2,3\}$ and $\mathcal{X}=\{x^1,x^2\}$, as a function of $(\epsilon_1-\epsilon_3,\epsilon_2-\epsilon_3)$.\footnote{This figure is based on Figures 1-3 in che:ros:smo13.} Consider the model implied distribution, denoted $\mathsf{M}$ below, of the optimal choice. Then, recalling the restriction ${\boldsymbol{z}}\protect\mathpalette{\protect\independenT}{\perp}\nu$, we have
Because the joint distribution of $({\boldsymbol{x}},\nu)$ conditional on ${\boldsymbol{z}}$ is left completely unrestricted (other than (ref)), one can find multiple triplets $(\vartheta,\mathsf{Q},\mathsf{R}(\nu|{\boldsymbol{x}},{\boldsymbol{z}}))$ satisfying the maintained assumptions and with $\mathsf{M}(c|{\boldsymbol{x}}\in R_x,{\boldsymbol{z}};\vartheta)=\mathsf{P}(c|{\boldsymbol{x}}\in R_x,{\boldsymbol{z}})$ for all $c\in\mathcal{Y}$ and $R_x\subseteq\mathcal{X}$, ${\boldsymbol{z}}$-a.s.
It is instructive to compare (ref)-(ref) with mcf73's mcf73 conditional logit. Under the standard assumptions, ${\boldsymbol{x}}\protect\mathpalette{\protect\independenT}{\perp}\nu$ so that no instrumental variables are needed. This yields $\mathsf{Q}(\nu)=\mathsf{R}(\nu|{\boldsymbol{x}})$ ${\boldsymbol{x}}$-a.s., and in addition $\mathsf{Q}$ is typically known, with corresponding simplifications in (ref). The resulting system of equalities can be inverted under standard order and rank conditions to yield point identification of $\theta$.
Further insights can be gained by looking at Figure (ref). As the value of ${\boldsymbol{x}}$ changes from $x^1$ to $x^2$, the region of values where, say, alternative 1 is optimal changes. When ${\boldsymbol{x}}$ is exogenous, say independent of $\nu$, this yields a system of equalities relating $(\theta,\mathsf{Q})$ to the observed distribution $\mathsf{P}({\boldsymbol{y}},{\boldsymbol{x}})$ which, as stated above, can be inverted to obtain point identification. When ${\boldsymbol{x}}$ is endogenous, this reasoning breaks down because the conditional distribution $\mathsf{R}(\nu|{\boldsymbol{x}},{\boldsymbol{z}})$ may change across realizations of ${\boldsymbol{x}}$. Figure (ref) also offers an instructive way to connect Identification Problem (ref) with the identification problem studied in Section (ref) (as well as with those in Sections (ref)-(ref) below). In the latter, the model has set-valued predictions for the outcome variable given realizations of the covariates and unobserved heterogeneity terms, which overlap across realizations of the unobserved heterogeneity terms. In the problem studied here, the model has singleton-valued predictions for the outcome variable of interest ${\boldsymbol{y}}$ as a function of the observable explanatory variables ${\boldsymbol{x}}$ and unobservables $\nu$. However, for given realization of $\nu$, the model admits sets of values for the endogenous variables $({\boldsymbol{y}},{\boldsymbol{x}})$, which overlap across realizations of $\nu$. Because the model is silent on the joint distribution of $({\boldsymbol{x}},\nu)$ (except for requiring that the marginal distribution of $\nu$ does not depend on ${\boldsymbol{z}}$), partial identification results.
It is possible to couple the maintained assumptions with the observed data to learn features of $(\theta,\mathsf{Q})$. Because the observed choice ${\boldsymbol{y}}$ is assumed to maximize utility, for the data generating $(\theta,\mathsf{Q})$ the model yields
with $\mathcal{E}_\theta({\boldsymbol{y}},{\boldsymbol{x}})$ a random closed set as per Definition (ref). Equation (ref) exhausts the modeling content of Identification Problem (ref). Theorem (ref) (as expressed in (ref)) can then be leveraged to extract its empirical content from the observed distribution $\mathsf{P}({\boldsymbol{y}},{\boldsymbol{x}},{\boldsymbol{z}})$. As a preparation for doing so, note that for given $F\in\mathcal{F}$ (with $\mathcal{F}$ the collection of closed subsets of $\mathcal{V}$) and $\vartheta\in\Theta$, we have
so that this probability can be learned from the observed data.
While Theorem SIR-(ref) relies on checking inequality (ref) for all $F\in\mathcal{F}$, the results in che:ros:smo13 and mol:mol18 can be used to obtain a smaller collection of sets over which to verify it. In particular, if ${\boldsymbol{x}}$ has a discrete distribution, it suffices to use a finite collection of sets. For example, in the case depicted in Figure (ref) with $\mathcal{X}=\{x^1,x^2\}$, che:ros:smo13 show that $\mathcal{H}_\mathsf{P}[\theta,\mathsf{Q}]$ is obtained by checking at most twelve inequalities in (ref). The left hand side of these inequalities is a linear function of six values that the distribution $\tilde\mathsf{Q}$ assigns to each of the component regions depicted in Figure (ref) (the one where $\mathcal{E}_\vartheta(1,x^1)\cap\mathcal{E}_\vartheta(1,x^2)$ realizes; the one where $\mathcal{E}_\vartheta(1,x^1)\cap\mathcal{E}_\vartheta(3,x^2)$ realizes; etc.) Hence, in this example, $(\vartheta,\tilde\mathsf{Q})\in\mathcal{H}_\mathsf{P}[\theta,\mathsf{Q}]$ if and only if $\tilde\mathsf{Q}$ assigns to these six regions a probability mass such that for $\vartheta$ the twelve inequalities characterized by che:ros:smo13 hold.
che:ros19 discuss related generalized instrumental variables models where random set methods are used to obtain characterizations of sharp identification regions in the presence of endogenous explanatory variables.
Compared to the general framework set forth at the beginning of Section (ref), as pointed out in man77, often the researcher observes $({\boldsymbol{y}}_i,{\boldsymbol{x}}_i)$ but not ${\boldsymbol{C}}_i$, $i=1,\dots,n$. Even when ${\boldsymbol{C}}_i$ is observable, the researcher may be unaware of which of its elements the decision maker actually evaluates before selecting one. In what follows, to shorten expressions, I refer to both the measurement problem of unobserved choice sets and the (cognitive) problem of limited consideration as “unobserved heterogeneity in choice sets."
Learning features of preferences using discrete choice data in the presence of unobserved heterogeneity in choice sets is a formidable task. When a decision maker chooses an alternative, this may be because her choice set equals the feasible set and the chosen alternative is the one yielding the highest utility. Then observed choice reveals preferences. But it can also be that the decision maker has access to/considers only the chosen alternative blo:mar60. Then observed choice is driven entirely by choice set composition, and is silent about preferences. A plethora of scenarios between these extremes is possible, but the researcher does not know which has generated the observed data. This fundamental identification problem calls either for restrictions on the random utility model and consideration set formation process, or for collection of richer data that eliminates unobserved heterogeneity in ${\boldsymbol{C}}_i$ or allows for enhanced modeling of it cap16.
A sizable literature spanning behavioral economics, econometrics, experimental economics, marketing, microeconomics, and psychology, has put forward different models to formalize the complex process that leads to the formation of the set of alternatives that the agent considers or can choose from sim59, how63, tve72. man77 proposes both a general econometric model where decision makers draw choice sets from an unknown distribution, as well as a specific model of choice set formation, independent from preferences, and studies their implications for the distributional structure of random utility models.\footnote{The specific model in man77 is often used in applications. It posits that each alternative $c\in\mathcal{Y}$ enters the decision maker's choice set with probability $\phi_c$, independently of the other alternatives. The probability $\phi_c$ may depend on observable individual characteristics, and $\phi_c=1$ for at least one option $c\in\mathcal{Y}$ (the “default" good).}
However, assumptions about the choice set formation process are often rooted in a desire to achieve point identification rather than in information contained in the model or observed data.\footnote{These assumptions are akin to assumptions about selection mechanisms in models with multiple equilibria. The latter are discussed further below in Section (ref), along with their criticisms.} It is then important to ask what can be learned about decision maker's preferences under minimal assumptions on the choice set formation process. Allowing for unrestricted dependence between choice sets and preferences, while challenging for identification analysis, is especially relevant. Indeed, decision makers' unobserved attributes may determine both their preferences and which items in the feasible set they pay attention to or are available to them (e.g., through unobserved liquidity constraints, unobserved characteristics such as religious preferences in the context of school choice, or behavioral phenomena such as aversion to extremes, salience, etc.). Here I use the framework put forward by bar:cou:mol:tei18 to study identification of discrete choice models with unobserved heterogeneity in choice sets and preferences.
The model just laid out has set valued predictions for the decision maker's optimal choice, because different alternatives might be optimal depending on which choice set the decision maker draws. Figure (ref), which is based on the analysis in bar:cou:mol:tei18, illustrates the set valued predictions in a stylized example. In the figure $\nu$ is assumed to be a scalar; $\bar{\nu}_{j,m}$ denotes the threshold value of $\nu$ above which $c_j$ yields higher utility than $c_m$ and below which $c_m$ yields higher utility than $c_j$ (the threshold's dependence on $({\boldsymbol{x}};\delta)$ is suppressed for notational convenience). Consider the case that $\nu\in[\bar{\nu}_{2,3},\bar{\nu}_{1,2}]$, so that $c_2$ is the option yielding the highest utility among all options in $\mathcal{Y}$. When $\kappa=|\mathcal{Y}|-1$, the agent may draw a choice set that does not include one of the alternatives in $\mathcal{Y}$. If the excluded alternative is not $c_2$ (or if ${\boldsymbol{C}}$ realizes equal to $\mathcal{Y}$), the model predicts that the decision maker chooses $c_2$. If ${\boldsymbol{C}}$ realizes equal to $\mathcal{Y}\setminus\{c_2\}$, the model predicts that the decision maker chooses the second best: $c_1$ if $\nu\in[\bar{\nu}_{1,3},\bar{\nu}_{1,2}]$, and $c_3$ if $\nu\in[\bar{\nu}_{2,3},\bar{\nu}_{1,3}]$. Conversely, observation of ${\boldsymbol{y}}=c_1$ allows one to conclude that $\nu\ge\bar\nu_{1,3}$, and ${\boldsymbol{y}}=c_2$ that $\nu\ge\bar\nu_{2,4}$, with $\bar\nu_{2,4}\le\bar\nu_{1,3}$, and these regions of possible realizations of $\nu$ overlap.
Why does this set valued prediction hinder point identification? The reason is similar to the explanation given for Identification Problem (ref): the distribution of the observable data relates to the model structure in an incomplete manner, because the distribution of the (unobserved) choice sets is left completely unspecified. bar:cou:mol:tei18 show that one can find multiple candidate distributions for ${\boldsymbol{C}}$ and parameter vectors $\vartheta$, such that together they yield a model implied distribution for ${\boldsymbol{y}}|{\boldsymbol{x}}$ that matches $\mathsf{P}({\boldsymbol{y}}|{\boldsymbol{x}})$, ${\boldsymbol{x}}$-a.s.
bar:cou:mol:tei18 propose to work directly with the set of model implied optimal choices given $({\boldsymbol{x}},\nu)$ associated with each possible realization of ${\boldsymbol{C}}$, which is depicted in Figure (ref) for a specific example. The key idea is that, according to the model, the observed choice maximizes utility among the alternatives in ${\boldsymbol{C}}$. Hence, for the data generating value of $\theta$, it belongs to the set of model implied optimal choices. With this, the authors are able to characterize $\mathcal{H}_\mathsf{P}[\theta]$ through Theorem (ref) as the collection of parameter vectors that satisfy a finite number of conditional moment inequalities.
Identification Problem (ref) sets up a structure where preferences include idiosyncratic components $\nu$ that are decision maker specific and can depend on ${\boldsymbol{C}}$, and where heterogeneity in ${\boldsymbol{C}}$ can be driven either by a measurement problem, or by the decision maker's limited attention to the options available to her. However, for computational and finite sample inference reasons, it restricts the family of utility functions to be known up to a finite dimensional parameter vector $\delta$.
A rich literature in decision theory has analyzed a different framework, where the decision maker's choice set is observable to the researcher, but the decision maker does not consider all alternatives in it mas:nak:ozb12,man:mar14. In this literature, the utility function is left completely unspecified, so that interest focuses on identification of preference orderings of the available options. Unobserved heterogeneity in preferences is assumed away, so that heterogeneous choice is driven by randomness in consideration sets. If the consideration set formation process is left unspecified or is subject only to weak restrictions, point identification of the preference orderings is not possible even if preferences are homogeneous and the researcher observes a representative agent facing multiple distinct choice problems with varying choice sets. cat:ma:mas:sul17 propose a general model for the consideration set formation process where the only restriction is a weak and intuitive monotonicity condition: the probability that any particular consideration set is drawn does not decrease when the number of possible consideration sets decreases. Within this framework, they provide revealed preference theory and testable implications for observable choice probabilities.
cat:ma:mas:sul17 posit that an observed distribution of choice $\mathsf{P}({\boldsymbol{y}}|{\boldsymbol{C}})$ has a random attention representation, and hence they name it a random attention model, if there exists a preference ordering $\succ$ over $\mathcal{Y}$ and a monotonic attention rule $\mu$ such that
The sharp identification region for the preference ordering, denoted $\mathcal{H}_\mathsf{P}[\succ]$ henceforth, is given by the collection of preference orderings for which one can find a monotonic attention rule to pair it with, so that (ref) holds.
Of course, an observed distribution of choice can be represented by multiple preference orderings and attention rules. The authors, however, show in their Lemma 1 that if for some $G\in\mathfrak{D}$ with $\{b,c\}\in G$,
then $c \succ b$ for any $\succ$ for which one can find a monotonic attention rule $\mu$ such that (ref) holds. Because of preference transitivity, one can also learn $a\succ b$ if in addition to the above condition one has $\mathsf{p}(a|G^\prime)>\mathsf{p}(a|G^\prime\setminus \{c\})$ for some $c\in G^\prime$ and $G^\prime\in\mathfrak{D}$. The authors further show in their Theorem 1 that the collection of preference relations associated with all possible instances of (ref) for all $c\in G$ and $G\in\mathfrak{D}$ yield all information about preferences given the observed choice probabilities. This yields a system of linear inequalities in $\mathsf{p}(c|G)$ that fully characterize $\mathcal{H}_\mathsf{P}[\succ]$. Let $\vec{\mathsf{p}}$ denote the vector with elements $[\mathsf{p}(c|G):c\in G,G\in\mathfrak{D}]$ and $\Pi_\succ$ denote a conformable matrix collecting the constraints on $\mathsf{P}({\boldsymbol{y}}|{\boldsymbol{C}})$ embodied in (ref) and its generalizations based on transitive closure. Then
The authors show that for any given preference ordering $\succ$, the matrix $\Pi_\succ$ characterizing whether $\succ \in \mathcal{H}_\mathsf{P}[\succ]$ through the system of linear inequalities in (ref) is unique, and they provide a simple algorithm to compute it. They also show that mild additional assumptions, such as, for example, that decision makers facing binary choice sets pay attention to both alternatives frequently enough, can substantially increase the informational content of the data (i.e., substantially tighten $\mathcal{H}_\mathsf{P}[\succ]$).
aba:ada18 and bar:mol:thi19 provide different sets of sufficient conditions for point identification of models of limited consideration. In both cases, the authors posit specific models of consideration set formation and provide sufficient conditions for point identification under exclusion and large support assumptions. aba:ada18 assume that unobserved heterogeneity in preferences and in consideration sets are independent. They exploit violations of Slutsky symmetry that result from inattention, assuming that for each alternative there is an observable characteristic with large support that does not affect the consideration probability of the other options. bar:mol:thi19 provide a thorough analysis of the extent of dependency between consideration and preferences under which semi-nonparametric point identification of the distribution of preferences and consideration attains. They exploit a requirement of standard economic theory --the Spence-Mirrlees single crossing property of utility functions-- coupled with a mild strengthening of the classic conditions for semi-nonparametric identification of discrete choice models with full consideration and identical choice sets mat07, assuming that there is at least one decision maker-specific characteristic with large support that affects utility but not consideration.
Building on mar60, man07b studies a question related but distinct from those in Identification Problems (ref)-(ref). He is concerned with prediction of choice behavior when decision makers face counterfactual choice sets. man07b frames this question as one of predicting treatment response (see Section (ref)). Here the collection of potential treatments is given by $\mathfrak{D}$, the nonempty subsets of the universe of feasible alternatives $\mathcal{Y}$, and the response function specifies the alternative chosen by a decision maker when facing choice set $G\in\mathfrak{D}$. man07b assumes that the researcher observes realized choice sets and chosen alternatives, $({\boldsymbol{y}},{\boldsymbol{C}})\sim\mathsf{P}$.\footnote{Here I suppress covariates for simplicity.} Under the standard assumptions laid out at the beginning of Section (ref), specifically if utility functions are (say) linear in $\epsilon_{ic}$ and the distribution of $\epsilon_{ic}$ is (say) Type I extreme value or multivariate normal, prediction of choice behavior with counterfactual choice sets is immediate (and point identified). man07b, however, leaves utility functions completely unspecified, and in fact works directly with preference orderings, which he labels decision maker's types. He places no restriction on the distribution of preference types, except requiring that they are independent of the observed choice sets. man07b shows that under these rather weak assumptions, the distribution of predicted choices from counterfactual choice sets can be partially identified, and characterized as the solution to linear programs.
Specifically, let ${\boldsymbol{y}}^*(G)$ denote the decision maker's optimal choice when facing choice set $G\in\mathfrak{D}$. Assume ${\boldsymbol{y}}^*(\cdot)\protect\mathpalette{\protect\independenT}{\perp}{\boldsymbol{C}}$, and let $y_k$ denote the choice function for a decision maker of type $k$ --that is, a decision maker with a specific preference ordering labeled $k$. One example of such preference ordering might be $c_1\succ c_2\succ\dots\succ c_{|\mathcal{Y}|}$. If a decision maker of this type faces, say, choice set $G=\{c_2,c_3,c_4\}$, then she chooses alternative $c_2$. Let $K$ denote the set of logically possible types, and $\theta_k$ the probability that a decision maker in the population is of type $k$. Suppose that the researcher posits a behavioral model specifying $K$, $\{y_k,k=1,\dots,K\}$, and restrictions that constrain $\theta$ to lie in some specified set of distributions. Let $\Theta$ denote the values of $\vartheta$ that satisfy these requirements plus the conditions $\vartheta_k\ge 0$ for all $k\in K$ and $\sum_{k\in K}\vartheta_k=1$. Then for any $c\in\mathcal{Y}$ and $\vartheta\in\Theta$, the model predicts
How can one partially identify this probability based on the observed data? Suppose ${\boldsymbol{C}}$ is observed to take realizations $D_1,\dots,D_m$. Then the data reveal
This yields that the sharp identification region for $\theta$ is
If the behavioral model is correctly specified, $\mathcal{H}_\mathsf{P}[\theta]$ is non-empty. In turn, the sharp identification region for each choice probability is
and its extreme points can be obtained by solving linear programs.
kit:sto19 provide closely related sharp bounds on features of counterfactual choices in the nonparametric random utility model of demand, where observable choices are repeated cross-sections and one allows for unrestricted, unobserved heterogeneity. Their approach builds on the work of kit:sto18, who test weather agents' behavior is consistent with the Axiom of Revealed Stochastic Preference (SARP) in a random utility model in which the utility function of each consumer over commodity bundles is assumed to satisfy only the basic restriction that “more is better" with no satiation. Because the testing exercise is to be carried out using repeated cross-sections data, the authors maintain the assumption that multiple populations of consumers who face distinct choice sets have the same distribution of preferences. With this structure in place, de facto the task is to test the full implications of rationality without functional form restrictions. kit:sto18's approach is based on several novel ideas. As a first step, they leverage an earlier insight of mcf05 to discretize the data without loss of information, so that they can define a large but finite set of rational preferences types. As a second step, they show that this implies that rationality can be tested by checking whether observed behavior lies in a cone corresponding to positive linear combinations of preference types. While the problem is discrete, its dimension is at first sight prohibitive. Nonetheless, Kitamura and Stoye are able to develop novel computational methods that render the problem tractable. They apply their method to the U.K. Household Expenditure Survey, adapting to their framework results on nonparametric instrumental variable analysis by imb:new09 so that they can handle price endogeneity.
kam18 builds on man07b to learn program effects when agents are randomly assigned to control or treatment. The treatment group is provided access to the program, while the control group is not. However, members of the control group may receive access to the program from outside the experiment, leading to noncompliance with the randomly assigned treatment. The researcher wants to learn about the average effect of program access on the decision to participate in the program and on the subsequent outcome. While sufficiently rich data may allow the researcher to learn these effects, kam18 is concerned with the identification problem that arises when the researcher only observes the treatment assignment status, the program participation decision, and the outcome, but not the receipt of program access for every agent. kam18 formalizes this problem as one where the received treatment is selected from a choice set that depends on the assigned treatment and is unobservable to the researcher, and the agents optimally choose whether to participate in the program by maximizing their utility function over their choice set. Importantly, the utility functions are not subject to parametric restrictions, similarly to man07b. But while man07b assumed independence of choice sets and preference types, kam18 allows them to be arbitrarily dependent on each other, as in bar:cou:mol:tei18. kam18's kam18 approach leverages specific assumptions on random assignment of treatments and on compliance (or lack thereof) of participants to obtain nonparametric bounds on the treatment effects of interest that can be characterized using tractable linear programs.
tam03 and cil:tam09 substantially enlarge the scope of partial identification analysis of structural models by showing how to apply it to learn features of payoff functions in static, simultaneous-move finite games of complete information with multiple equilibria. ber:tam06 extend the approach and considerations that follow to games of incomplete information. To start, here I focus on two-player entry games with complete information.\footnote{Completeness of information is motivated by the idea that firms in the industry have settled in a long-run equilibrium, and have detailed knowledge of both their own and their rivals' profit functions.}
From the econometric perspective, this is a generalization of a standard discrete choice model to a bivariate simultaneous response model which yields a stochastic representation of equilibria in a two player, two action game. Generically, for a given value of $\theta$ and realization of the payoff shifters, the model just laid out admits multiple equilibria (existence of PSNE is guaranteed because the interaction parameters are non-positive). In other words, it yields set valued predictions as depicted in Figure (ref).\footnote{This figure is based on Figure 1 in tam03.}
Why does this set valued prediction hinder point identification? Intuitively, the challenge can be traced back to the fact that for different values of $\theta\in\Theta$, one may find different ways to assign the probability mass in $[-{\boldsymbol{x}}_1\beta_1,-{\boldsymbol{x}}_1\beta_1-\delta_1)\times [-{\boldsymbol{x}}_2\beta_2,-{\boldsymbol{x}}_2\beta_2-\delta_2)$ to $(0,1)$ and $(1,0)$, so as to match the observed distribution $\mathsf{P}({\boldsymbol{y}}_1,{\boldsymbol{y}}_2|{\boldsymbol{x}}_1,{\boldsymbol{x}}_2)$. More formally, for fixed $\vartheta\in\Theta$ and given $({\boldsymbol{x}},\varepsilon)$ and $(y_1,y_2)\in\{0,1\}\times\{0,1\}$, let
so that in Figure (ref) $\mathcal{E}_\vartheta[(1,0),(0,1);{\boldsymbol{x}}]$ is the gray region, $\mathcal{E}_\vartheta[(0,1);{\boldsymbol{x}}]$ is the dotted region, etc. Let $\mathsf{R}(y_1,y_2|{\boldsymbol{x}},\varepsilon)$ be a selection mechanism that assigns to each possible outcome of the game $(y_1,y_2)\in\{0,1\}\times\{0,1\}$ the probability that it is played conditional on observable and unobservable payoff shifters. In order to be admissible, $\mathsf{R}(y_1,y_2|{\boldsymbol{x}},\varepsilon)$ must be such that $\mathsf{R}(y_1,y_2|{\boldsymbol{x}},\varepsilon)\ge 0$ for all $(y_1,y_2)\in\{0,1\}\times\{0,1\}$, $\sum_{(y_1,y_2)\in\{0,1\}\times\{0,1\}}\mathsf{R}(y_1,y_2|{\boldsymbol{x}},\varepsilon)=1$, and
Let $\Phi_r$ denote the probability distribution of a bivariate Normal random variable with zero means, unit variances, and correlation $r\in[-1,1]$. Let $\mathsf{M}(y_1,y_2|{\boldsymbol{x}})$ denote the model predicted probability that the outcome of the game realizes equal to $(y_1,y_2)$. Then the model yields
Because $\mathsf{R}(\cdot|{\boldsymbol{x}},\varepsilon)$ is left completely unspecified, other than the basic restrictions listed above that render it an admissible selection mechanism, one can find multiple values for $(\vartheta,\mathsf{R}(\cdot|{\boldsymbol{x}},\varepsilon))$ such that $\mathsf{M}(y_1,y_2|{\boldsymbol{x}})=\mathsf{P}(y_1,y_2|{\boldsymbol{x}})$ for all $(y_1,y_2)\in\{0,1\}\times\{0,1\}$ ${\boldsymbol{x}}$-a.s.
Multiplicity of equilibria implies that the mapping from the model's exogenous variables $({\boldsymbol{x}}_1,{\boldsymbol{x}}_2,\varepsilon_1,\varepsilon_2)$ to outcomes $({\boldsymbol{y}}_1,{\boldsymbol{y}}_2)$ is a correspondence rather than a function. This violates the classical \textquotedblleft principal assumptions\textquotedblright\ or \textquotedblleft coherency conditions\textquotedblright\ for simultaneous discrete response models discussed extensively in the econometrics literature hec78,gou80,sch81,mad83,blu:smi94. Such coherency conditions require the existence of a unique reduced form, mapping the model's exogenous variables and parameters to a unique realization of the endogenous variable; hence, they constrain the model to be recursive or triangular in nature. As pointed out by bjo:vuo84, however, the coherency conditions shut down exactly the social interaction effect of interest by requiring, e.g., that $\delta_1\delta_2=0$, so that at least one player's action has no impact on the other player's payoff.
The desire to learn about interaction effects coupled with the difficulties generated by multiplicity of equilibria prompted the earlier literature to provide at least two different ways to achieve point identification. The first one relies on imposing simplifying assumptions that shift focus to outcome features that are common across equilibria. For example, bre:rei88,bre:rei90,bre:rei91 and ber92 study entry games where the number, though not the identities, of entrants is uniquely predicted by the model in equilibrium. Unfortunately, however, these simplifying assumptions substantially constrain the amount of heterogeneity in player's payoffs that the model allows for. The second approach relies on explicitly modeling a selection mechanism which specifies the equilibrium played in the regions of multiplicity. For example, bjo:vuo84 assume it to be a constant; baj:hon:rya10 assume a more flexible, covariate dependent parametrization; and ber92 considers two possible selection mechanism specifications, one where the incumbent moves first, and the other where the most profitable player moves first. Unfortunately, however, the chosen selection mechanism can have non-trivial effects on inference, and the data and theory might be silent on which is more appropriate. A nice example of this appears in ber92. ber:tam06 review and extend a number of results on the identification of entry models extensively used in the empirical literature. jov89 discusses the observable implications of models with multiple equilibria, and within the analysis of a model with homogeneous preferences shows that partial identification is possible jov89. I refer to pau13 for a review of the literature on econometric analysis of games with multiple equilibria.
cil:tam09 show, on the other hand, that it is possible to partially identify entry models that allow for rich heterogeneity in payoffs and for any possible selection mechanism (even ones that are arbitrarily dependent on the unobservable payoff shifters after conditioning on the observed payoff shifters). In addition, tam03 provides sufficient conditions for point identification based on exclusion restrictions and large support assumptions. kli:tam12 analyze partial identification of nonparametric models of entry in a two-player model, drawing connections with the program evaluation literature.
cil:tam09 propose to use simple and tractable implications of the model to learn features of the structural parameters of interest. Specifically, they point out that the probability of observing any outcome of the game cannot be smaller than the model's implied probability that such outcome is the unique equilibrium of the game, and cannot be larger than the model's implied probability that such outcome is one of the possible equilibria of the game. Looking at Figure (ref) this means, for example, that the observed $\mathsf{P}(({\boldsymbol{y}}_1,{\boldsymbol{y}}_2)=(0,1)|{\boldsymbol{x}}_1,{\boldsymbol{x}}_2)$ cannot be smaller than the probability that $(\varepsilon_1,\varepsilon_2)$ realizes in the dotted region, and cannot be larger than the probability that it realizes either in the dotted region or in the gray region. Compared to the model predicted distribution in (ref), this means that $\mathsf{P}(({\boldsymbol{y}}_1,{\boldsymbol{y}}_2)=(0,1)|{\boldsymbol{x}}_1,{\boldsymbol{x}}_2)$ cannot be smaller than the expression obtained setting, for $\varepsilon\in\mathcal{E}_\vartheta[(1,0);(0,1);{\boldsymbol{x}}]$, $\mathsf{R}(0,1|{\boldsymbol{x}},\varepsilon)=0$, and cannot be larger than that obtained with $\mathsf{R}(0,1|{\boldsymbol{x}},\varepsilon)=1$. Denote by $\Phi(A_1,A_2;\rho)$ the probability that the bivariate normal with mean vector zero, variances equal to one, and correlation $\rho$ assigns to the event $\{\varepsilon_1\in A_1,\varepsilon_2\in A_2\}$. Then cil:tam09 show that any $\vartheta=[d_1,d_2,b_1,b_2,r]$ that is observationally equivalent to the data generating value $\theta$ satisfies, $({\boldsymbol{x}}_1,{\boldsymbol{x}}_2)$-a.s.,
While the approach of cil:tam09 is summarized here for a two player entry game, it extends without difficulty to any finite number of players and actions and to solution concepts other than pure strategy Nash equilibrium.
ara:tam08 build on the insights of cil:tam09 to study what is the identification power of equilibrium in games. To do so, they compare the set-valued model predictions and what can be learned about $\theta$ when one assumes only level-$k$ rationality as opposed to Nash play. In static entry games of complete information, they find that the model's predictions when $k\ge 2$ are similar to those obtained with Nash behavior and allowing for multiple equilibria and mixed strategies. mol:ros08 extend the analysis of ara:tam08 to the class of supermodular games.
The collections of parameter vectors satisfying (in)equalities (ref)-(ref) yields the sharp identification region $\mathcal{H}_\mathsf{P}[\theta]$ in the case of two player entry games with pure strategy Nash equilibrium as solution concept, as shown by ber:mol:mol11. When there are more than two players or more than two actions ara:tam08, the characterization in cil:tam09 obtained by extending the reasoning just laid out yields an outer region. ber:mol:mol11 use elements of random set theory to provide a general and computationally tractable characterization of the identification region that is sharp, regardless of the number of players and actions, or the solution concept adopted. For the case of PSNE with any finite number of players or actions, gal:hen11 provide a computationally tractable sharp characterization of the identification region using elements of optimal transportation theory.
ber:mol:mol11 provide a general approach based on random set theory that delivers sharp identification regions on parameters of structural semiparametric models with set valued predictions. Here I summarize it for the case of static, simultaneous move finite games of complete information, first with PSNE as solution concept and then with mixed strategy Nash equilibrium. Then I discuss games of incomplete information.
For a given $\vartheta\in\Theta$, denote the set of pure strategy Nash equilibria (depicted in Figure (ref)) as ${\boldsymbol{Y}}_\vartheta({\boldsymbol{x}},\varepsilon)$. It is easy to show that ${\boldsymbol{Y}}_\vartheta({\boldsymbol{x}},\varepsilon)$ is a random closed set as in Definition (ref). Under the assumption in Identification Problem (ref) that ${\boldsymbol{y}}$ results from simultaneous move, pure strategy Nash play, at the true DGP value of $\theta\in\Theta$, one has
Equation (ref) exhausts the modeling content of Identification Problem (ref). Theorem (ref) can be leveraged to extract its empirical content from the observed distribution $\mathsf{P}({\boldsymbol{y}},{\boldsymbol{x}})$. For a given $\vartheta\in\Theta$ and $K\subset\mathcal{Y}$, let $\mathsf{T}_{{\boldsymbol{Y}}_{\vartheta}({\boldsymbol{x}},\varepsilon)}(K;\Phi_r)$ denote the probability of the event $\{{\boldsymbol{Y}}_\vartheta({\boldsymbol{x}},\varepsilon)\cap K\neq \emptyset\}$ implied when $\varepsilon\sim\Phi_r$, ${\boldsymbol{x}}$-a.s.
The characterization provided in Theorem SIR-(ref) for games with multiple PSNE, taken from ber:mol:mol11, is equivalent to the one in gal:hen11. When $J=2$ and $\mathcal{Y}=\{0,1\}\times\{0,1\}$, the inequalities in (ref) reduce to (ref)-(ref). With more players and/or more actions, the inequalities in (ref) are a superset of those in (ref)-(ref), with the latter comprised of the ones in (ref) for $K=\{k\}$ and $k=\mathcal{Y}\setminus\{k\}$, for all $k\in\mathcal{Y}$. Hence, the inequalities in (ref) are more informative. Of course, the computational cost incurred to characterize $\mathcal{H}_\mathsf{P}[\theta]$ may grow with the number of inequalities involved. I discuss computational challenges in partial identification in Section (ref).
Next, I discuss the case that the outcome of the game results from simultaneous move, mixed strategy Nash play.\footnote{The same reasoning given here applies if instead of mixed strategy Nash the solution concept is correlated equilibrium, by replacing the set of MSNE below with the set of correlated equilibria.} When mixed strategies are allowed for, the model predicts multiple mixed strategy Nash equilibria (MSNE). But whereas when only pure strategies are allowed for, if the model is correctly specified, the observed outcome of the game is one of the predicted PSNE, with mixed strategy it is only the result of a random mixing draw from one of the predicted MSNE. Hence, the identification problem is more complex, and in order to obtain a tractable characterization of $\theta$'s sharp identification region one needs to use different tools from random set theory.
To keep the treatment simple here I continue to consider the case of two players with two strategies, as in Identification Problem (ref), with mixed strategies allowed for, and refer to mol:mol18 for the general case. Fix $\vartheta\in\Theta$. Let $\sigma_j:\{0,1\}\to [0,1]$ denote the probability that player $j$ enters the market, with $1-\sigma_j$ the probability that she stays out. With some abuse of notation, let $\mathfrak{u}_j(\sigma_j,\sigma_{-j},{\boldsymbol{x}}_j,\varepsilon_j,\vartheta)$ denote the expected payoff associated with the mixed strategy profile $\sigma=(\sigma_1,\sigma_2)$. For a given realization $(x,e)$ of $({\boldsymbol{x}},\varepsilon)$ and a given value of $\vartheta\in\Theta$, the set of mixed strategy Nash equilibria is
ber:mol:mol11 show that ${\boldsymbol{S}}_\vartheta\equiv S_\vartheta({\boldsymbol{x}},\varepsilon)$ is a random closed set in $[0,1]^2$. Its realizations are illustrated in Panel (a) of Figure (ref) as a function of $(\varepsilon_1,\varepsilon_2)$.\footnote{This figure is based on Figure 1 in ber:mol:mol11.}
Define the set of possible multinomial distributions over outcomes of the game associated with the selections $\sigma$ of each possible realization of ${\boldsymbol{S}}_{\vartheta}$ as
As ${\boldsymbol{Q}}_\vartheta$ is the image of a continuous map applied to the random compact set ${\boldsymbol{S}}_\vartheta$, it is a random compact set. Its realizations are plotted in Panel (b) of Figure (ref) as a function of $(\varepsilon_1,\varepsilon_2)$.
The multinomial distribution over outcomes of the game determined by a given $\sigma\in{\boldsymbol{S}}_\vartheta$ is a function of $\varepsilon$. To obtain the predicted distribution over outcomes of the game conditional on observed payoff shifters only, one needs to integrate out the unobservable payoff shifters $\varepsilon$. Doing so requires care, as it needs to be done for each ${\boldsymbol{q}}(\sigma)\in{\boldsymbol{Q}}_\vartheta$. First, observe that all the ${\boldsymbol{q}}(\sigma)\in{\boldsymbol{Q}}_\vartheta$ are contained in the $3$ dimensional unit simplex, and are therefore integrable. Next, define the conditional selection expectation (see Definition (ref)) of ${\boldsymbol{Q}}_\vartheta$ as
where $\operatorname{Sel}({\boldsymbol{S}}_\vartheta)$ is the set of all measurable selections from ${\boldsymbol{S}}_\vartheta$, see Definition (ref). By construction, $\mathbb{E}_{\Phi_r}({\boldsymbol{Q}}_\vartheta|{\boldsymbol{x}})$ is the set of probability distributions over action profiles conditional on ${\boldsymbol{x}}$ which are consistent with the maintained modeling assumptions, i.e., with all the model's implications (including the assumption that $\varepsilon\sim\Phi_r$). If the model is correctly specified, there exists at least one vector $\theta\in\Theta$ such that the observed conditional distribution $\mathsf{p}({\boldsymbol{x}})\equiv[\mathsf{P}({\boldsymbol{y}}=y^1|{\boldsymbol{x}}),\dots,\mathsf{P}({\boldsymbol{y}}=y^4|{\boldsymbol{x}})]^\top$ almost surely belongs to the set $\mathbb{E}_{\Phi_\rho}({\boldsymbol{Q}}_\theta|{\boldsymbol{x}})$. Indeed, by the definition of $\mathbb{E}_{\Phi_\rho}({\boldsymbol{Q}}_\theta|{\boldsymbol{x}})$, $\mathsf{p}({\boldsymbol{x}})\in \mathbb{E}_{\Phi_\rho}({\boldsymbol{Q}}_\theta|{\boldsymbol{x}})$ almost surely if and only if there exists ${\boldsymbol{q}}\in \operatorname{Sel}({\boldsymbol{Q}}_\theta)$ such that $\mathbb{E}_{\Phi_\rho}({\boldsymbol{q}}|{\boldsymbol{x}})=\mathsf{p}({\boldsymbol{x}})$ almost surely, with $\operatorname{Sel}({\boldsymbol{Q}}_\theta)$ the set of all measurable selections from ${\boldsymbol{Q}}_\theta$. Hence, the collection of parameter vectors $\vartheta\in\Theta$ that are observationally equivalent to the data generating value $\theta$ is given by the ones that satisfy $\mathsf{p}({\boldsymbol{x}})\in \mathbb{E}_{\Phi_r}({\boldsymbol{Q}}_\vartheta|{\boldsymbol{x}})$ almost surely. In turn, observing that by Theorem (ref) the set $\mathbb{E}_{\Phi_r}({\boldsymbol{Q}}_\vartheta|{\boldsymbol{x}})$ is convex, we have that $\mathsf{p}({\boldsymbol{x}})\in \mathbb{E}_{\Phi_r}({\boldsymbol{Q}}_\vartheta|{\boldsymbol{x}})$ if and only if $u^\top \mathsf{p}({\boldsymbol{x}})\leq h_{\mathbb{E}_{\Phi_r}({\boldsymbol{Q}}_\vartheta|{\boldsymbol{x}})}(u)$ for all $u$ in the unit ball roc70, where $h_{\mathbb{E}_{\Phi_r}({\boldsymbol{Q}}_\vartheta|{\boldsymbol{x}})}(u)$ is the support function of $\mathbb{E}_{\Phi_r}({\boldsymbol{Q}}_\vartheta|{\boldsymbol{x}})$, see Definition (ref).
For a fixed $u\in\mathbb{B}^4$, the possible realizations of $h_{{\boldsymbol{Q}}_\vartheta}(u)$ are plotted in Panel (c) of Figure (ref) as a function of $(\varepsilon_1,\varepsilon_2)$. The expectation of $h_{{\boldsymbol{Q}}_\vartheta}(u)$ is quite straightforward to compute, whereas calculating the set $\mathbb{E}_{\Phi_r}({\boldsymbol{Q}}_\vartheta|{\boldsymbol{x}})$ is computationally prohibitive in many cases. Hence, the characterization in (ref) is computationally attractive, because for each $\vartheta\in\Theta$ it requires to maximize an easy-to-compute superlinear, hence concave, function over a convex set, and check if the resulting objective value vanishes. Several efficient algorithms in convex programming are available to solve this problem, see for example the MatLab software for disciplined convex programming CVX gra:boy10. Nonetheless, $\mathcal{H}_\mathsf{P}[\theta]$ itself is not necessarily convex, hence tracing out its boundary is non-trivial. I return to computational challenges in partial identification in Section (ref).
I conclude this section discussing the case of static, simultaneous move finite games of incomplete information, using the results in ber:mol:mol11.\footnote{See ber:tam06 and pau13 for a thorough discussion of the literature on identification problems in games of incomplete information with multiple Bayesian Nash equilibria (BNE). ber:tam06 explain how to extend the approach proposed by cil:tam09 to obtain outer regions on $\theta$ when no restrictions are imposed on the equilibrium selection mechanism that chooses among the multiple BNE.} For clarity, I formalize the maintained assumptions.
With incomplete information, players' strategies are decision rules that map the support of $(\varepsilon,{\boldsymbol{x}})$ into $\{0,1\}$. The non-negativity condition on expected payoffs that determines each player's decision to enter the market results in equilibrium mappings (decision rules) that are step functions determined by a threshold: $y_j(\varepsilon_j) =\mathbf{1}(\varepsilon_j\geq t_j), j=1,2$. As a result, player $j$'s beliefs about player $3-j$'s probability of entry under the common prior assumption is $\int y_{3-j}(\varepsilon_{3-j}) d\mathsf{F}_\gamma(\varepsilon_{3-j}|{\boldsymbol{x}}) =1-\mathsf{F}_\gamma(t_{3-j}|{\boldsymbol{x}})$, and therefore player $j$'s best response cutoff is
Hence, the set of equilibria can be defined as the set of cutoff rules:
The equilibrium thresholds are functions of ${\boldsymbol{x}}$ and $\theta$ only. The set ${\boldsymbol{T}}_{\theta}({\boldsymbol{x}})$ might contain a finite number of equilibria (e.g., if the common prior is the Normal distribution), or a continuum of equilibria. For ease of notation I suppress its dependence on ${\boldsymbol{x}}$ in what follows.
Given the equilibrium decision rules (the selections of the set ${\boldsymbol{T}}_\theta$), it is possible to determine their associated action profiles. Because in the simple two-player entry game that I consider actions and outcomes coincide, I denote the set of admissible action profiles by ${\boldsymbol{Y}}_\theta$:
with $\operatorname{Sel}({\boldsymbol{T}}_\theta)$ the set of all measurable selections from ${\boldsymbol{T}}_\theta$, see Definition (ref). To obtain the predicted set of multinomial distributions for the outcomes of the game, one needs to integrate out $\varepsilon$ conditional on ${\boldsymbol{x}}$. Again this can be done by using the conditional Aumann expectation:
This set is closed and convex. Regardless of whether ${\boldsymbol{T}}_\theta$ contains a finite number of equilibria or a continuum, ${\boldsymbol{Y}}_\theta$ can take on only a finite number of realizations corresponding to each of the vertices of the three dimensional simplex, because the vectors ${\boldsymbol{y}}({\boldsymbol{t}})$ in (ref) collect threshold decision rules. This implies that $\mathbb{E}_{\mathsf{F}_\gamma}({\boldsymbol{Y}}_\theta|{\boldsymbol{x}})$ is a closed convex polytope ${\boldsymbol{x}}$-a.s., fully characterized by a finite number of supporting hyperplanes. Hence, it is possible to determine whether $\vartheta\in\mathcal{H}_\mathsf{P}[\theta]$ using efficient algorithms in linear programming.
One can use the same argument as in the proof of Theorem SIR-(ref), to show that the Aumann expectation/support function characterization of the sharp identification region in Theorem SIR-(ref) coincides with the characterization based on the capacity functional in Theorem SIR-(ref), when only pure strategies are allowed for. This shows that in this class of models, the capacity functional based characterization is a special case of the Aumann expectation/support function based one.
ara:tam08 study what is the identification power of equilibrium also in the case of static entry games with incomplete information. They show that in the presence of multiple equilibria, assuming Bayesian Nash behavior yields more informative regions for the parameter vector $\theta$ than assuming only rational behavior, but at the price of a higher computational cost.
pau:tan12 propose a procedure to test for the sign of the interaction effects (which here I have assumed to be non-positive) in discrete simultaneous games with incomplete information and (possibly) multiple equilibria. As a by-product of this procedure, they also provide a test for the presence of multiple equilibria in the DGP. The test does not require parametric specifications of players' payoffs, the distributions of their private signals, or the equilibrium selection mechanism. Rather, the test builds on the commonly invoked assumption that players' private signals are independent conditional on observed states.
gri14 introduces an important class of models with flexible information structure. Each player is assumed to have a vector of payoff shifters unobservable by the researcher composed of elements that are private information to the player, and elements that are known to all players. The results of ber:mol:mol11 reported in this section apply to this set-up as well.
hai:tam03 study what can be learned about the distribution of valuations in an open outcry English auction where symmetric bidders have independent private values for the object being auctioned. The standard theoretical model mil:web82, called “button auction" model, posits that each bidder holds down a button while the object's price rises continuously and exogenously, releasing it (in the dominant strategy equilibrium) when it reaches her valuation or all her opponents have left. In this case, the distribution of bidder's valuation can be learned exactly. hai:tam03 show that much can be learned about the distribution of valuations, even allowing for the fact that real-life auctions may depart from this stylized framework, as in the following identification problem.\footnote{Examples of departures from the standard model include the case where active bidding by a player's opponents may eliminate her incentives to bid close to her valuation or at all; the econometrician does not precisely observe the point at which each bidder drops out; there are discrete bid increments; etc. }
The model in Identification Problem (ref) delivers set valued predictions because given valuations $({\boldsymbol{v}}_1,\dots,{\boldsymbol{v}}_n)$, the two fundamental assumptions about bidder's behavior yield
where $\vec{{\boldsymbol{v}}}_n\equiv({\boldsymbol{v}}_{1:n},\dots,{\boldsymbol{v}}_{n:n})$ denotes the vector of order statistics of the valuations, and $V_n=\{v\in\mathbb{R}^n:\underline{v}\le v_1\le v_2\le\dots\le v_n\le \bar{v}\}$.\footnote{Using the same convention as for the bids, ${\boldsymbol{v}}_{i:n}$ denotes the $i$-th lowest of the $n$ valuations.} Figure (ref) provides a stylized depiction of a realization of this set for $\vec{{\boldsymbol{v}}}_n=v^0$ when there are three bidders ($n=3$), $\underline{v}=0$, and $\delta=0$. In words, ${\boldsymbol{B}}(\vec{{\boldsymbol{v}}}_n)$ collects the model predicted values of ordered bids. The fact that ${\boldsymbol{b}}_{i:n}\le {\boldsymbol{v}}_{i:n}$ for all $i$ results from assumption (1): since each bidder bids at most an amount equal to her valuation, the $i$-th highest bid cannot exceed the $i$-th highest valuation hai:tam03.\footnote{Note that ${\boldsymbol{b}}_{i:n}$ needs not be the bid made by the bidder with valuation ${\boldsymbol{v}}_{i:n}$.} The fact that ${\boldsymbol{b}}_{n:n}\ge {\boldsymbol{v}}_{n-1,n}-\delta$ follows immediately from assumption (2) hai:tam03. The fact that $\vec{{\boldsymbol{b}}}_n$ has to lie in $V_n$ follows because it is a vector of ordered bids.
Why does this set-valued prediction hinder point identification? The reason is that the distribution of the observable data relates to the model structure in an incomplete manner.\footnote{hai:tam03 provide the discussion summarized here. Additionally, in their Appendix B, they give a simple example of a two-bidder auction satisfying all assumptions in Identification Problem (ref), where two different distributions $\mathsf{Q}$ and $\tilde{\mathsf{Q}}$ yield the same distribution of ordered bids.} Define a bidding rule $\mathsf{B}({\boldsymbol{b}}_{1:n},\dots,{\boldsymbol{b}}_{n:n}|{\boldsymbol{v}}_{1:n},\dots,{\boldsymbol{v}}_{n:n})$ to be a conditional joint distribution for the order statistics of the bids conditional on the order statistics of the valuations. Then, for a given realization of the valuations ${\boldsymbol{v}}_{1:n}=v_1,\dots,{\boldsymbol{v}}_{n:n}=v_n$, the model requires that the support of $\mathsf{B}(\cdot|v_1,\dots,v_n)$ is in $B(\vec{v})$ as defined in (ref) with ${\boldsymbol{v}}_{1:n}=v_1,\dots,{\boldsymbol{v}}_{n:n}=v_n$, but imposes no other restriction on it. Hence, the model implied joint distribution of ordered bids is
where $\mathsf{Q}_{1,\dots,n:n}$ is the joint distribution of order statistics of the valuations implied by $\mathsf{Q}$. Since the bidding rule $\mathsf{B}$ is left completely unspecified (other than requiring it to be a valid joint conditional probability distribution with support in ${\boldsymbol{B}}$), one can find multiple pairs $(\mathsf{B} ,\mathsf{Q})$ satisfying the assumptions of Identification Problem (ref), such that $\mathsf{M}_{1,\dots,n:n}(\cdot;\mathsf{B},\mathsf{Q})=\mathsf{G}_{1,\dots,n:n}(\cdot)$, with $\mathsf{G}_{1,\dots,n:n}$ the observed joint CDF of the order statistics of the bids associated with $\mathsf{P}$.
hai:tam03 propose to use simple and tractable implications of the model to learn features of $\mathsf{Q}$. Recall that with i.i.d. valuations, the distribution of each order statistic uniquely determines $\mathsf{Q}(v)$, with $\mathsf{Q}(v)\equiv\mathsf{Q}({\boldsymbol{v}}\le v)$ for any $v\ge\underline{v}$, through:
where $\mathsf{Q}_{i:n}$ is the CDF of ${\boldsymbol{v}}_{i:n}$ and $\mathsf{q}_{\mathcal{B}}(\cdot;i,n-i+1)$ is the quantile function of a Beta-distributed random variable with parameters $i$ and $n-i+1$. Using this, their Lemmas 1 and 3 yield, respectively,
where, for any $v\ge\underline{v}$, $\mathsf{G}_{i:n}(v)\equiv\mathsf{P}({\boldsymbol{b}}_{i:n}\le v)$ denotes the observed CDF of ${\boldsymbol{b}}_{i:n}$ for $i=1,\dots,n$.
hai:tam03 also provide sharp bounds on the optimal reserve price, which I do not discuss here. However, they leave open the question of whether the collection of CDFs satisfying (ref)-(ref) yields the sharp identification region for $\mathsf{Q}$. As discussed in Sections (ref)-(ref), pointwise bounds on the CDF deliver tubes of admissible CDFs that in general yield outer regions on the CDF of interest. But in this identification problem, the issue of sharpness is even more subtle, and therefore addressed in the following subsection.
Before moving on to that discussion, I note that the work of hai:tam03 spurred a rich literature applying partial identification analysis to the study of auction models. tan11 studies first price sealed bid auctions with equilibrium behavior, where affiliated valuations prevent --in the absence of parametric restrictions on the distribution of the model primitives-- point identification of the model. He derives bounds on seller revenue under various counterfactual scenarios on reserve prices and auction formats. arm13 also studies first price sealed bid auctions with equilibrium behavior, but relaxes the independence assumptions on symmetric valuations by requiring it to hold only conditional on unobserved heterogeneity. He derives bounds on various functionals of the distributions of interest, including the mean bid and mean valuation. ara:gan:qui13 analyze second price auctions with correlated private values. In this case, the distribution of valuations is not point identified even under the assumptions of the button auction model ath:hai02. Nonetheless, ara:gan:qui13 show that interesting functionals of it (seller profits and bidder surplus) can be bounded, if one assumes that transaction prices are determined by the second highest valuation and imposes some restrictions on the joint distribution of the number of bidders and distribution of the valuations. kom13 studies a related model of second-price ascending auctions with arbitrary dependence in bidders' private values. She provides partial identification results for the joint distribution of values for any subset of bidders under various assumptions about what data the researcher observes. While in her framework the highest bid is never observed, she considers the case where only the winner's identity and the winning price are observed, and the case where all the identities and all the bids except for the highest bid are known. She also investigates the informational content of assuming positive dependence in bidders' values. gen:li14 are concerned with nonparametric identification of a two-stage entry and bidding game. Potential bidders are assumed to have private valuations and observe private signals before deciding whether to enter the auction. The dependence between signals and valuations is only minimally restricted. Hence, even with some excluded instruments that affect selection into the auction, the model primitives are only partially identified. The authors derive bounds on these primitives, and provide conditions under which point identification is restored. syr:tam:zia18 provide partial identification results in private value and common value auctions under weak restrictions on the information available to the bidders. Their approach leverages a result in ber:mor16 yielding an equivalence between distributions of valuations that obey the restrictions imposed by a Bayesian Correlated Equilibrium and those that obey the restrictions imposed by Bayesian Nash Equilibrium under some information structure. Such equivalence is particularly helpful because the set of Bayesian Correlated Equilibria can be characterized through linear programming, so that the sharp identification region provided by syr:tam:zia18 is given by the collection of parameter vectors $\vartheta$ for which a linear program is feasible. Related results leveraging the linear structure of correlated equilibria in the context of entry games include yan06, ber:mol:mol11, and mag:ron17.
hai:tam03's hai:tam03 bounds exploit the information contained in the marginal CDFs $\mathsf{G}_{i:n}$ for each $i$ and $n$. However, in Identification Problem (ref) additional information can be extracted from the joint distribution of ordered bids. che:ros17 obtain the sharp identification region $\mathcal{H}_\mathsf{P}[\mathsf{Q}]$ using random set methods (Artstein's characterization in Theorem (ref)) applied to a quantile function representation of the order statistics. Here I provide an equivalent characterization that uses equation (ref) directly, and which has not appeared in the literature before. Let $\mathcal{T}$ denote the space of probability distributions with support on $[\underline{v},\bar{v}]$, so that $\mathsf{Q}\in\mathcal{T}$. For a candidate distribution $\tilde{\mathsf{Q}}\in\mathcal{T}$, let $\tilde{\mathsf{Q}}_{1,\dots,n:n}$ denote the implied distribution of order statistics of $n$ i.i.d. random variables distributed $\tilde{\mathsf{Q}}$. Let $\tilde{{\boldsymbol{B}}}$ be a random closed set defined as in (ref) with respect to order statistics of i.i.d. random variables with distribution $\tilde{\mathsf{Q}}$. For a given set $K\in\mathcal{K}$, with $\mathcal{K}$ the collection of compact subsets of $\mathbb{R}^n$, let $\mathsf{T}_{\tilde{\boldsymbol{B}}}(K;\tilde{\mathsf{Q}})$ denote the probability of the event $\{\tilde{\boldsymbol{B}}\cap K\neq \emptyset\}$ implied by $\tilde{\mathsf{Q}}$.
In (ref), $\mathsf{P}(\vec{{\boldsymbol{b}}}_n\in K)$ is determined by the joint distribution of the ordered bids and hence can be learned from the data. On the other side, $\mathsf{T}_{\tilde{\boldsymbol{B}}}(K;\tilde{\mathsf{Q}})$ is a function of the model and $\tilde{\mathsf{Q}}\in\mathcal{T}$. Hence, it can be computed using (ref), with $\tilde{\boldsymbol{B}}$ defined with respect to order statistics of i.i.d. random variables with distribution $\tilde{\mathsf{Q}}\in\mathsf{T}$. To gain insights in the characterization of $\mathcal{H}_\mathsf{P}[\mathsf{Q}]$, consider for example the set $K=\{\prod_{i=1}^{n-1}(-\infty,+\infty)\}\times(-\infty,v]$. Plugging it in the inequalities in (ref), one obtains
which, using (ref), yields (ref). Similarly, plugging in the sets $K_j=\{\prod_{i=1}^{j-1}(-\infty,+\infty)\}\times[v,\infty)\times\{\prod_{j+1}^n(-\infty,+\infty)\}$, $j=1,\dots,n$, yields (ref). So the inequalities proposed by hai:tam03 are a subset of the inequalities yielding the sharp identification region in Theorem SIR-(ref). More information can be obtained by using additional sets $K$. For instance, the set $K=[v_1,\infty)\times[v_2,\infty)\times\{\prod_{i=1}^{n}(-\infty,+\infty)\}$, $v_2\ge v_1$, yields $\mathsf{P}({\boldsymbol{b}}_{1:n}\ge v_1,{\boldsymbol{b}}_{2:n}\ge v_2)\le \mathsf{Q}_{1,2:n}([v_1,\infty)\times[v_2,\infty))$, which further restricts $\mathsf{Q}$. Numerous examples can be given.
Characterization (ref) is stated using inequality (ref) for the collection of compact subsets of $\mathbb{R}^n$. One can instead use the (equivalent) inequality (ref), and show that in fact it suffices to check it for a much smaller collection of sets, as shown by che:ros17 mol:mol18. Nonetheless, this collection remains extremely large.
che:ros17auction further generalize the analysis in this section by dropping the requirement of independent private values. This allows them, for example, to consider affiliated private values. They show that even in this significantly more complex context, the key behavioral restrictions imposed by hai:tam03 to relate bids to valuations can be coupled with the use of random set theory, to characterize sharp identification regions.
Strategic models of network formation generalize the frameworks of single agents and multiple agents discrete choice models reviewed in Sections (ref) and (ref). They posit that pairs of agents (nodes) form, maintain, or sever connections (links) according to an explicit equilibrium notion and utility structure. Each individual's utility depends on the links formed by others (the network) and on utility shifters that may be pair-specific.
One may conjecture that the results reported in Sections (ref)-(ref) apply in this more general context too. While of course lessons can be carried over, network formation models present challenges that combined cannot be overcome without the development of new tools. These include the issue of equilibrium existence and the possibility of multiple equilibria when they exist, due to the interdependence in agents' choices (this problem was already discussed in Section (ref)). Another challenge is the degree of correlation between linking decisions, which interacts with how the observable data is generated: one may observe a growing number of independent networks, or a growing number of agents on a single network. Yet another challenge, which substantially increases the difficulties associated with the previous two, is the combinatoric complexity of network formation problems. The purpose of this section is exclusively to discuss some recent papers that have made important progress to address these specific challenges and carry out partial identification analysis. For a thorough treatment of the literature on network formation, I refer to the reviews in gra15, cha16, pau17, and gra19.\footnote{For a review of the literature on peer group effect analysis, see, e.g., bro:dur01hoe, blu:bro:dur:ioa11, pau17, and gra19.}
Depending on whether the researcher observes data from a single network or multiple independent networks, the underlying population of agents may be represented as a continuum or as a countably infinite set in the first case, or as a finite set in the second case. Henceforth, I denote generic agents as $i$, $j$, $k$, and $m$. I consider static models of undirected network formation with non-transferable utility.\footnote{Undirected means that if a link from node $i$ to node $j$ exists, then the link from $j$ to $i$ exists. The discussion that follows can be generalized to the case of models with transferable utility.} The collection of all links among nodes forms the network, denoted ${\boldsymbol{y}}$. For any pair $(i,j)$ with $i\neq j$, ${\boldsymbol{y}}_{ij}=1$ if they are linked, and ${\boldsymbol{y}}_{ij}=0$ otherwise (${\boldsymbol{y}}_{ii}=0$ for all $i$ by convention). The notation ${\boldsymbol{y}}-\{ij\}$ denotes the network that results if a link present between nodes $i$ and $j$ is deleted, while ${\boldsymbol{y}}+\{ij\}$ denotes the network that results if a link absent between nodes $i$ and $j$ is added. Denote agent $i$'s payoff by $\mathfrak{u}_i({\boldsymbol{y}},{\boldsymbol{x}},\epsilon)$. This payoff depends on the network ${\boldsymbol{y}}$ and the payoff shifters $({\boldsymbol{x}},\epsilon)$, with ${\boldsymbol{x}}$ observable both to the agents and to the researcher, $\epsilon$ only to the agents, and $({\boldsymbol{x}},\epsilon)$ collecting $({\boldsymbol{x}}_{ij},\epsilon_{ij})$ for all $i$ and $j$.\footnote{Here I consider a framework where the agents have complete information.}
Following much of the literature, I employ pairwise stability jac:wol96 as equilibrium notion: ${\boldsymbol{y}}$ is a pairwise stable network if all linked agents prefer not to sever their links, and all non-existing links are damaging to at least one agent. Formally,
Under this equilibrium notion, if equilibria exist multiplicity is likely; see, among others, the examples in gra15, pau17, and she18. The model is therefore incomplete, because it does not specify how an equilibrium is selected in the region of multiplicity. For the same reasons as discussed in the context of finite games in Section (ref), partial identification results (unless one is willing to impose restrictions on the equilibrium selection mechanism). However, as I explain below, an immediate application of the identification analysis carried out there presents enormous practical challenges because there are $2^{n(n-1)/2}$ possible network configurations to be checked for stability (and the dimensionality of the space of unobservables is also very large).
In what follows I consider two distinct frameworks that make different assumptions about the utility function and how the data is generated, and discuss what can be learned about the parameters of interest in these cases.
I first consider the case that the researcher observes data from multiple independent networks. I follow the set-up put forward by she18.
she18 analyzes this problem. She establishes equilibrium existence provided that $\delta_2\ge 0$ and $\delta_3\ge 0$ she18.\footnote{With transferable utility, she18 establishes existence for any $\delta_2,\delta_3\in\mathbb{R}$. See hel13 for an earlier analysis of existence and uniqueness of pairwise stable networks.} Given payoff shifters $({\boldsymbol{x}},\epsilon)$ and parameters $\vartheta\equiv[\tilde\delta_1~\tilde\delta_2~\tilde\delta_3~\tilde\gamma]\in\Theta$, let ${\boldsymbol{Y}}_\vartheta({\boldsymbol{x}},\epsilon)$ denote the collection of pairwise stable networks implied by the model. It is easy to show that ${\boldsymbol{Y}}_\vartheta({\boldsymbol{x}},\epsilon)$ is a random closed set as in Definition (ref). The networks in ${\boldsymbol{Y}}_\vartheta({\boldsymbol{x}},\epsilon)$ are $n\times n$ symmetric adjacency matrices with diagonal elements equal to zero and off diagonal elements in $\{0,1\}$. To ease notation, I omit ${\boldsymbol{Y}}_\vartheta$'s dependence on $({\boldsymbol{x}},\epsilon)$ in what follows. Under the assumption that ${\boldsymbol{y}}$ is a pairwise stable network, at the true data generating value of $\theta\in\Theta$, one has
Equation (ref) exhausts the modeling content of Identification Problem (ref). Theorem (ref) can be leveraged to extract its empirical content from the observed distribution $\mathsf{P}({\boldsymbol{y}},{\boldsymbol{x}})$. Let $\mathcal{Y}$ be the collection of $n\times n$ symmetric matrices with diagonal elements equal to zero and all other entries in $\{0,1\}$, so that $|\mathcal{Y}|=2^{n(n-1)/2}$. For a given set $K\subset\mathcal{Y}$, let $\mathsf{T}_{{\boldsymbol{Y}}_{\vartheta}}(K;\mathsf{F}_\gamma)$ denote the probability of the event $\{{\boldsymbol{Y}}_\vartheta\cap K\neq \emptyset\}$ implied when $\epsilon\sim\mathsf{F}_\gamma$, ${\boldsymbol{x}}$-a.s.
The characterization of $\mathcal{H}_\mathsf{P}[\theta]$ in Theorem SIR-(ref) is new to this chapter.\footnote{ gua19 has previously used Theorem D.1 in ber:mol:mol11, as I do here, to characterize sharp identification regions in unilateral and bilateral directed network formation games.} While technically it entails a finite number of conditional moment inequalities, in practice their number can be prohibitive as it can be as large as $2^{2^{n(n-1)/2}}-2$.\footnote{This number may be reduced drastically using the notion of core determining class of sets, see Definition (ref) and the discussion on p. (ref). Nonetheless, even with relatively few agents, the number of inequalities in (ref) may remain overwhelming.} Even using only a subset of the inequalities in (ref) to obtain an outer region, for example applying the insights in cil:tam09, may not be practical (with $n=20$, $|\mathcal{Y}|\approx 10^{57}$). Moreover, computation of $\mathsf{T}_{{\boldsymbol{Y}}_{\vartheta}}(K;\mathsf{F}_\gamma)$ may require (depending on the set $K$) evaluation of rather complex integrals.
To circumvent these challenges, she18 proposes to analyze network formation through subnetworks. A subnetwork is the restriction of a network to a subset of the agents (i.e., a subset of nodes and the links between them). For given $A\subseteq\{1,2,\dots,n\}$, let ${\boldsymbol{y}}^A=\{{\boldsymbol{y}}_{ij}\}_{i,j\in A, i\neq j}$ be the submatrix in ${\boldsymbol{y}}$ with rows and columns in $A$, and let ${\boldsymbol{y}}^{-A}$ be the remaining elements of ${\boldsymbol{y}}$ after ${\boldsymbol{y}}^A$ is deleted. With some abuse of notation, let $({\boldsymbol{y}}^A,{\boldsymbol{y}}^{-A})$ denote the composition of ${\boldsymbol{y}}^A$ and ${\boldsymbol{y}}^{-A}$ that returns ${\boldsymbol{y}}$. Recall that ${\boldsymbol{Y}}_\vartheta\equiv{\boldsymbol{Y}}_\vartheta({\boldsymbol{x}},\epsilon)$, and let
be the collection of subnetworks with rows and columns in $A$ that can be part of a pairwise stable network in ${\boldsymbol{Y}}_\vartheta$. Let ${\boldsymbol{x}}^A$ denote the subset of ${\boldsymbol{x}}$ collecting ${\boldsymbol{x}}_{ij}$ for $i,j\in A$. For a given $y^A\in\{0,1\}^{|A|}$, let $\mathsf{C}_{{\boldsymbol{Y}}_{\vartheta}^A}(y^A;\mathsf{F}_\gamma)$ and $\mathsf{T}_{{\boldsymbol{Y}}_{\vartheta}^A}(y^A;\mathsf{F}_\gamma)$ denote, respectively, the probability of the events $\{{\boldsymbol{Y}}_\vartheta^A=\{y^A\}\}$ and $\{\{y^A\}\in{\boldsymbol{Y}}_\vartheta^A\}$ implied when $\epsilon\sim\mathsf{F}_\gamma$, ${\boldsymbol{x}}$-a.s. The first event means that only the subnetwork $y^A$ is part of a pairwise stable network, while the second event means that $y^A$ is a possible subnetwork that is part of a pairwise stable network but other subnetworks may be part of it too. she18 provides the following outer region for $\theta$ by adapting the insight in cil:tam09 to subnetworks. In the theorem I abuse notation compared to Table (ref) by introducing a superscript, $A$, to make explicit the dependence of the outer region on it.
she18 further assumes that the selection mechanism ${\boldsymbol{u}}(\tilde{\boldsymbol{y}}|{\boldsymbol{Y}}_\vartheta)$ is invariant to permutations of the labels of the players. Under this condition and the maintained assumptions on $\epsilon$, she shows that the inequalities in (ref) are invariant under permutations of labels, so subnetworks in any two subsets $A,A'\subseteq\{1,2,\dots,n\}$ with $|A|=|A'|$ and ${\boldsymbol{x}}^A={\boldsymbol{x}}^{A'}$ yield the same inequalities for all $y^A=y^{A'}$. It is therefore sufficient to consider subnetwork $A$ and the inequalities in (ref) associated with it. Leveraging this result, she18 proposes an outer region obtained by looking at unlabeled subnetworks of size $|A|\le\bar{a}$ and given by
As long as the subnetworks are chosen to be small, e.g., $|A|\le 2,3,4$, the inequalities in (ref) can be computed even if the network is large. she18 shows that the inequalities in (ref) remain informative as $n$ grows. This fact highlights the importance of working with subnetworks. One could have applied the insight of cil:tam09 directly to the full network by setting ${\boldsymbol{u}}$ equal to zero and to one in (ref). The resulting bounds, however, would vanish to zero as $n$ grows and become uninformative for $\theta$. The characterization in Theorem OR-(ref) can be refined to obtain a smaller region, adapting the results in ber:mol:mol11 to subnetworks. The size of this refined region is weakly decreasing in $|A|$.\footnote{The idea of using random set methods on subnetworks to obtain the refined region was put forward in an earlier version of she18. She provided a proof that the refined region's size decreases weakly in $|A|$.} However, the refinement does not yield $\mathcal{H}_\mathsf{P}[\theta]$ because it is applied only to subnetworks.
miy16 considers a framework similar to the one laid out in Identification Problem (ref). He assumes non-negative externalities, and shows that in this case the set of pairwise stable equilibria is a complete lattice with a smallest and a largest equilibrium.\footnote{This approach exploits supermodularity, and is related to jia08 and ech05.} He then uses moment functions that are monotone in the pairwise stable network (so that they take their extreme values at the smallest and largest equilibria), to obtain moment conditions that restrict $\theta$. Examples of the moment functions used include the proportion of pairs with a link, the proportion of links belonging to traingles, and many more miy16.
gua19 considers unilateral and bilateral directed network formation games, still under a sampling framework where the researcher observes many independent networks. The equilibrium notion that she uses is pure strategy Nash. She assumes that the payoff that player $i$ receives from forming link $ij$ is allowed to depend on the number of additional players forming a link pointing to $j$, but rules out other spillover effects. Under this assumption and some regularity conditions, gua19 shows that the network formation game can be decomposed into local games (i.e., games whose sets of players and strategy profiles are subsets of the network formation game's ones), so that the network formation game is in equilibrium if and only if each local game is in equilibrium. She then obtains a characterization of $\mathcal{H}_\mathsf{P}[\theta]$ using elements of random set theory.
When the researcher observes data from a single network, extra care has to be taken to restrict the dependence among linking decisions. This can be done in various ways cha16. Here I consider a framework proposed by pau:shu:tam18.
Identification Problem (ref) enforces dimension reduction through the restrictions on depth and degree (the bounds $\bar{d}$ and $\bar{l}$), so that it is applicable to frameworks with networks that have limited degree distribution (e.g., close friendships network, but not Facebook network). It also requires that individual identities are irrelevant. This substantially reduces the richness of unobserved heterogeneity allowed for and the dimensionality of the space of unobservables. While the latter feature narrows the domain of applicability of the model, it is very beneficial to obtain a tractable characterization of what can be learned about $\theta$, and yields equilibria that may include isolated nodes, a feature often encountered in networks data.
pau:shu:tam18 study Identification Problem (ref) focusing on the payoff-relevant local subnetworks that result from the maintained assumptions. These are distinct from the subnetworks used by she18: whereas she18 looks at subnetworks formed by arbitrary individuals and whose size is chosen by the researcher on the base of computational tractability, pau:shu:tam18 look at subnetworks among individuals that are within a certain distance of each other, as determined by the structure of the preferences. On the other hand, she18's she18 analysis does not require that agents have a finite number of types nor bounds the number of links that they may form.
To characterize the local subnetworks relevant for identification analysis in their framework, pau:shu:tam18 propose the concepts of network type and preference class. A network type $t=(a,v)$ describes the local network up to distance $\bar{d}$ from the reference node. Here $a$ is a square matrix of size $1+\bar{l}\sum_{d=1}^{\bar{d}}(\bar{l}-1)^{d-1}$ that describes the local subnetwork that is utility relevant for an agent of type $t$. It consists of the reference node, its direct potential neighbors ($\bar{l}$ elements), its second order neighbors ($\bar{l}(\bar{l}-1)$ elements), through its $\bar{d}$-th order neighbors ($\bar{l}(\bar{l}-1)^{\bar{d}-1}$ elements). The other component of the type, $v$, is a vector of length equal to the size of $a$ that contains the observable characteristics of the reference node and her alters. The bounds $\bar{d}$ and $\bar{l}$ enforce dimension reduction by bounding the number of network types. The partial identification approach of pau:shu:tam18 depends on this number, rather than on the number of agents. For example, the number of moment inequalities is determined by the number of network types, not by the number of agents. As such, the approach yields its highest dividends for dimension reduction in large networks.
Let $\mathcal{T}$ denote the collection of network types generated from a preference structure $\mathfrak{u}$ and set of characteristics $\mathcal{X}$. For given realization $(x,e)$ of the observable characteristics and preference shocks of a reference agent, and for given $\vartheta\in\Theta$, define the collection of network types for which no agent wants to drop a link by
where $a_{-\ell}$ is equal to the local adjacency matrix $a$ but with the $\ell$-th link removed (that is, it sets the $(1,\ell+1)$ and $(\ell+1,1)$ elements of $a$ equal to zero). Because $({\boldsymbol{x}},\epsilon)$ are random vectors, ${\boldsymbol{H}}_\vartheta\equiv H_\vartheta({\boldsymbol{x}},\epsilon)$ is a random closed set as per Definition (ref). This random set takes on a finite number of realizations (equal to the possible subsets of $\mathcal{T}$), so that its distribution is completely determined by the probability with which it takes on each of these realizations. A preference class $H\subset\mathcal{T}$ is one of the possible realizations of ${\boldsymbol{H}}_\vartheta$ for some $\vartheta\in\Theta$. The model implied probability that ${\boldsymbol{H}}_\vartheta=H$ is given by
Observation of data from one network allows the researcher, under suitable restrictions on the sampling process, to learn the distribution of network types in the data (type shares), denoted $\mathsf{P}(t)$.\footnote{Full observation of the network is not required (and in practice it often does not occur). Sampling uncertainty results from it because in this model there is a continuum of agents.} For example, in a network of best friends with $\bar{l}=1$ and $\bar{d}=2$, and $\mathcal{X}=\{x^1,x^2\}$ (e.g., a simplified framework with only two possible races), agents are either isolated or in a pair. Network types are pairs for the agents' race and the best friend's race (with second element equal zero if the agent is isolated). Type shares are the fraction of isolated blacks, the fraction of isolated whites, the fraction of blacks with a black best friend, the fraction of whites with a black best friend, and the fraction of whites with a white best friend. The preference classes for a black agent are $H^1(b,e)=\{(b,0)\}$, $H^2(b,e)=\{(b,0),(b,b)\}$, $H^3(b,e)=\{(b,0),(b,w)\}$, $H^4(b,e)=\{(b,0),(b,w),(b,b)\}$ (and similarly for whites). In each case, being alone is part of the preference class, as there are no links to sever. In the second class the agent has a preference for having a black friend, in the third class for a white friend, and in the last class for a friend of either race. It is easy to see that the model is incomplete, as for a given realization of $\epsilon$ it makes multiple predictions on the agent's preference type.
pau:shu:tam18 propose to map the distribution of preference classes into the observed distribution of preference types in the data through the use of allocation parameters, denoted $\alpha_H(t)\in[0,1]$. These are distinct from but play the same role as a selection mechanism, and they represent a candidate distribution for $t$ given ${\boldsymbol{H}}_\vartheta=H$. The model, augmented with them, implies a probability that an agent is of network type $t$:
where $\mu_{v_1(t)}$ is the measure of reference agents with characteristics equal to the second component of the preference type $t$, ${\boldsymbol{x}}=v_1(t)$, and $\alpha\equiv\{\alpha_H(t):t\in \mathcal{T}, H\subset\mathcal{T}\}$.
pau:shu:tam18 provide a characterization of an outer region for $\theta$ based on two key implications of pairwise stability that deliver restrictions on $\alpha$. They also show that under some additional assumptions, this characterization yields $\mathcal{H}_\mathsf{P}[\theta]$ pau:shu:tam18. Here I focus on their more general result.
The first implication that they use is that existing links should not be dropped:
The condition in (ref) is embodied in $\bar\alpha\equiv\{\alpha_H(t):t\in H, H\subset\mathcal{T}\}$.
The second implication is that it should not be possible to establish mutually beneficial links among nodes that are far from each other. Let $t^\prime$ and $s^\prime$ denote the network types that are generated if one adds a link in networks of types $t$ and $s$ among two nodes that are at distance at least $2\bar{d}$ from each other and each have less than $\bar{l}$ links. Then the requirement is
In words, if a positive measure of agents of type $t$ prefer $t^\prime$ (i.e., $\alpha_H(t)>0$ for some $H$ such that $t^\prime\in H$), there must be zero measure of type $s$ individuals who prefer $s^\prime$, because otherwise the network is unstable. pau:shu:tam18 show that the conditions in (ref) can be embodied in a square matrix $q$ of size equal to the length of $\bar{\alpha}$. The entries of $q$ are constructed as follows. Let $H$ and $\tilde{H}$ be two preference classes with $t\in H$ and $s\in\tilde{H}$. With some abuse of notation, let $q_{\alpha_H(t),\alpha_{\tilde{H}}(s)}$ denote the element of $q$ corresponding to the index of the entry in $\bar\alpha$ equal to $\alpha_H(t)$ for the row, and to $\alpha_{\tilde{H}}(s)$ for the column. Then set $q_{\alpha_H(t),\alpha_{\tilde{H}}(s)}(\vartheta)=\mathbf{1}(t^\prime\in H)\mathbf{1}(s^\prime\in\tilde{H})$. It follows that this element yields the term $\big(\alpha_H(t)\mathbf{1}(t^\prime\in H)\big)\big(\alpha_{\tilde{H}}(s)\mathbf{1}(s^\prime\in \tilde{H})\big)$ in the quadratic form $\bar{\alpha}^\top q \bar{\alpha}$. As long as $\mu_{v_1(\cdot)}$ and $\mathsf{M}(\cdot|{\boldsymbol{x}};\vartheta)$ in (ref) are strictly positive, this term is equal to zero if and only if condition (ref) holds for types $t$ and $s$.\footnote{The possibility that $\mu_{v_1(\cdot)}$ or $\mathsf{M}(\cdot|{\boldsymbol{x}};\vartheta)$ are equal to zero can be accommodated by setting $q_{\alpha_H(t),\alpha_{\tilde{H}}(s)}(\vartheta)=(\mu_{v_1(t)}\mathsf{M}(H|v_1(t);\vartheta)\mathbf{1}(t^\prime\in H))(\mu_{v_1(s)}\mathsf{M}(H|v_1(s);\vartheta)\mathbf{1}(s^\prime\in\tilde{H}))$. However, in that case $q$ depends on $\vartheta$ and its computational cost increases.}
With this background, Theorem OR-(ref) below provides an outer region for $\theta$. The proof of this result follows from the arguments laid out above pau:shu:tam18.
The set in (ref) does not equal $\mathcal{H}_\mathsf{P}[\theta]$ in all models allowed for in Identification Problem (ref) because condition (ref) does not embody all implications of pairwise stability on non-existing links. While the optimization problem in (ref) is quadratic, it is not necessarily convex because $q$ may not be positive definite. Nonetheless, the simulations reported by pau:shu:tam18 suggest that $\mathcal{O}_\mathsf{P}[\theta]$ can be computed rapidly, as least for the examples they considered.
In order to discuss the partial identification approach to learning structural parameters of economic models in some level of detail while keeping this chapter to a manageable length, I have focused on a selection of papers. In this section I briefly mention several other excellent theoretical contributions that could be discussed more closely, as well as several empirical papers that have applied partial identification analysis of structural models to answer a wide array of questions of substantive economic importance.
pak10 and pak:por:ho:ish15 propose to embed revealed preference-based inequalities into structural models of both demand and supply in markets where firms face discrete choices of product configuration or of location. Revealed preference arguments are a trademark of the literature on discrete choice analysis. pak10 and pak:por:ho:ish15 use these arguments to leverage a subset of the model's implications to obtain easy-to-compute moment inequalities. For example, in the context of entry games such as the ones discussed in Section (ref), they propose to base inference on the implication that a player enters the market if and only if (s)he expects to make non-negative profits. This condition can be exploited even when players have heterogeneous (unobserved to the researcher) information sets, and it implies that the expected profits for entrants should be non-negative. Nonetheless, the condition does not suffice to obtain moment inequalities that include only observed payoff shifters and preference parameters. This is because the expected value of unobserved payoff shifters for entrants is not equal to zero, as the group of entrants is selected. The authors require the availability of valid (monotone) instrumental variables to solve this problem (see Section (ref) for uses of instrumental variables and monotone instrumental variables in partial identification analysis of treatment effects). Interesting features of their approach include that the researcher does not need to solve for the set of equilibria, nor to require that the distribution of unobservable payoff shifters is known up to finite dimensional parameter vector. Moreover, the same basic ideas can be applied to single agent models (with or without heterogeneous information sets). A shortcoming of the method is that the set of parameter vectors satisfying the moment inequalities may be wider than the sharp identification region under the maintained assumptions.
The breadth of applications of the approach proposed by pak10 and pak:por:ho:ish15 is vast.\footnote{Statistical inference in these papers is often carried out using the methods proposed by che:hon:tam07, ber:mol08, and and:soa10. Model specification tests, if carried out, are based on the method proposed by bug:can:shi15. See Sections (ref) and (ref), respectively, for a discussion of confidence sets and specification tests.} For example, ho09 uses it to model the formation of the hospital networks offered by US health insurers, and ho:ho:mor12 and lee13 use it to obtain bounds on firm fixed costs as an input to modeling product choices in the movie industry and in the US video game industry, respectively. hol11 estimates the effects of Wal-Mart's strategy of creating a high density network of stores. While the close proximity of stores implies cannibalization in sales, Wal-Mart is willing to bear it to achieve density economies, which in turn yield savings in distribution costs. His results suggest that Wal-Mart substantially benefits from high store density. ell:hou:tim13 measure the effects of chain economies, business stealing, and heterogeneous firms' comparative advantages in the discount retail industry. kaw:wat13 estimate a model of strategic voting and quantify the impact it has on election outcomes. As in other models analyzed in this section, the one they study yields multiple predicted outcomes, so that partial identification methods are required to carry out the empirical analysis if one does not assume a specific selection mechanism to resolve the multiplicity. They estimate their model on Japanese general-election data, and uncover a sizable fraction of strategic voters. They also estimate that only a small fraction of voters are misaligned (voting for a candidate other than their most preferred one). eiz14 studies whether the rapid removal from the market for personal computers of existing central processing units upon creation of new ones through innovation reduces surplus. He finds that a limited group of price-insensitive consumers enjoys the largest share of the welfare gains from innovation. A policy that kept older technologies on the shelf would allow for the benefits from innovation to reach price-sensitive consumers thanks to improved access to mobile computing, but total welfare would not increase because consumer welfare gains would be largely offset by producer losses. ho:pak14 analyze hospital referrals for labor and birth episodes in California in 2003, for patients enrolled with six health insurers that use, to a different extent, incentives to referring physicians groups to reduce hospital costs (capitation contracts). The aim is to learn whether enrollees with high-capitation insurers tend to be referred to lower-priced hospitals (ceteris paribus) compared to other patients with same-severity conditions, and whether quality of care was affected. Their model allows for an insurer-specific preference function that is additively separable in the hospital price paid by the insurer (which is allowed to be measured with error), the distance traveled, and plan and severity-specific hospital fixed effects. Importantly, unobserved heterogeneity entering the preference function is not assumed to be drawn from a distribution known up to finite dimensional parameter vector. The results of the empirical analysis indicate that the price paid by insurers to hospitals has an impact on referrals, with higher elasticity to price for insurers whose physicians groups are more highly capitated. dic:mor18 study how the information that potential exporters have to predict the profits they will earn when serving a foreign market influences their decisions to export. They propose a model where the researcher specifies and observes a subset of the variables that agents use to form their expectations, but may not observe other variables that affect firms' expectations heterogeneously (across firms and markets, and over time). Because only a subset of the variables entering the firms' information set is observed, partial identification results. They show that, under rational expectations, they can test whether potential exporters know and use specific variables to predict their export profits. They also use their model's estimates to quantify the value of information. wol18 studies the implications of the \$85 billion automotive industry bailout in 2009 on the commercial vehicle segment. He finds that had Chrysler and GM been liquidated (or aquired by a major competitor) rather than bailed out, the surviving firms would have experienced a rise in profits high enough to induce them to introduce new products.
A different use of revealed preference arguments appears in the contributions of blu:bro:cra08, blu:kri:mat14, hod:sto14,hod:sto15, man14, bar:mol:tei16, hau:new16, ada19, and many others. For example, man14 proposes a method to partially identify income-leisure preferences and to evaluate the associated effects of tax policies. He starts from basic revealed-preference analysis performed under the assumption that individuals prefer more income and leisure, and no other restriction. The analysis shows that observing an individual's time allocation under a status quo tax policy yields bounds on his allocation that may or may not be informative, depending on how the person allocates his time under the status quo policy and on the tax schedules. He then explores what more can be learned if one additionally imposes restrictions on the distribution of income-leisure preferences, using the method put forward by man07b. One assumption restricts groups of individuals facing different choice sets to have the same distribution of preferences. The other assumption restricts this distribution to a parametric family. kli:tar16 build on and expand man14's framework to evaluate the effect of Connecticut's Jobs First welfare reform experiment on women' labor supply and welfare participation decisions.
bar:mol:tei16 propose a method to learn features of households' risk preferences in a random utility model that nests expected utility theory plus a range of non-expected utility models.\footnote{Their model is based on the one put forward by bar:mol:odo:tei13. See bar:mol:odo:tei18 for a review of these and other non-expected utility models in the context of estimation of risk preferences.} They allow for unobserved heterogeneity in preferences (that may enter the utility function non-separably) and leave completely unspecified their distribution. The authors use revealed preference arguments to infer, for each household, a set of values for its unobserved heterogeneity terms that are consistent with the household's choices in the three lines of insurance coverage. As their core restriction, they assume that each household's preferences are stable across contexts: the household's utility function is the same when facing distinct but closely related choice problems. This allows them to use the inferred set valued data to partially identify features of the distribution of preferences, and to classify households into preference types. They apply their proposed method to analyze data on households' deductible choices across three lines of insurance coverage (home all perils, auto collision, and auto comprehensive).\footnote{Auto collision coverage pays for damage to the insured vehicle caused by a collision with another vehicle or object, without regard to fault. Auto comprehensive coverage pays for damage to the insured vehicle from all other causes, without regard to fault. Home all perils (or simply home) coverage pays for damage to the insured home from all causes, except those that are specifically excluded (e.g., flood, earthquake, or war).} Their results show that between 70 and 80 percent of the households make choices that can be rationalized by a model with linear utility and monotone, quadratic, or even linear probability distortions. These probability distortions substantially overweight small probabilities. By contrast, fewer than 40 percent can be rationalized by a model with concave utility but no probability distortions.
hau:new16 propose a method to carry out demand analysis while allowing for general forms of unobserved heterogeneity. Preferences and linear budget sets are assumed to be statistically independent (conditional on covariates and control functions). hau:new16 show that for continuous demand, average surplus is generally not identified from the distribution of demand for a given price and income, and therefore propose a partial identification approach. They use bounds on income effects to derive bounds on average surplus. They apply the bounds to gasoline demand, using data from the 2001 U.S. National Household Transportation Survey.
Another strand of empirical applications pertains to the analysis of discrete games. cil:tam09 use the method they develop, described in Section (ref), to study market structure in the US airline industry and the role that firm heterogeneity plays in shaping it. Their findings suggest that the competitive effects of each carrier increase in that carrier's airport presence, but also that the competitive effects of large carriers (American, Delta, United) are different from those of low cost ones (Southwest). They also evaluate the effect of a counterfactual policy repealing the Wright Amendment, and find that doing so would see an increase in the number of markets served out of Dallas Love.
gri14 proposes a model of static entry that extends the one in Section (ref) by allowing individuals to have flexible information structures, where players's payoffs depend on both a common-knowledge unobservable payoff shifter, and a private-information one. His characterization of $\mathcal{H}_\mathsf{P}[\theta]$ is based on using an unrestricted selection mechanism, as in ber:tam06 and cil:tam09. He applies the model to study the impact of supercenters such as Wal-Mart, that sell both food and groceries, on the profitability of rural grocery stores. He finds that entry by a supercenter outside, but within 20 miles, of a local monopolist's market has a smaller impact on firm profits than entry by a local grocer. Their entrance has a small negative effect on the number of grocery stores in surrounding markets as well as on their profits. The results suggest that location and format-based differentiation partially insulate rural stores from competition with supercenters.
A larger class of information structures is considered in the analysis of static discrete games carried out by mag:ron17. They allow for all information structures consistent with the players knowing their own payoffs and the distribution of opponents' payoffs. As solution concept they adopt the Bayes Correlated Equilibrium recently developed by ber:mor16. Also with this solution concept multiple equilibria are possible. The authors leave completely unspecified the selection mechanism picking the equilibrium played in the regions of multiplicity, so that partial identification attains. mag:ron17 use the random sets approach to characterize $\mathcal{H}_\mathsf{P}[\theta]$. They apply the method to estimate a model of entry in the Italian supermarket industry and quantify the effect of large malls on local grocery stores. nor:tan14 provide partial identification results (and Bayesian inference methods) for semiparametric dynamic binary choice models without imposing distributional assumptions on the unobserved state variables. They carry out an empirical application using rus87's model of bus engine replacement. Their results suggest that parametric assumptions about the distribution of the unobserved states can have a considerable effect on the estimates of per-period payoffs, but not a noticeable one on the counterfactual conditional choice probabilities. ber:com19 use the random sets approach to partially identify and estimate dynamic discrete choice models with serially correlated unobservables, under instrumental variables restrictions. They extend two-step dynamic estimation methods to characterize a set of structural parameters that are consistent with the dynamic model, the instrumental variables restrictions, and the data.\footnote{Statistical inference on $\theta$ is carried out using che:che:kat18's method.} gua19 uses the random sets approach and a network formation model, to learn about Italian firms' incentives for having their executive directors sitting on the board of their competitors.
bar:cou:mol:tei18 use the method described in Section (ref) to partially identify the distribution of risk preferences using data on deductible choices in auto collision insurance.\footnote{Statistical inference on projections of $\theta$ is carried out using kai:mol:sto19's method.} They posit an expected utility theory model and allow for unobserved heterogeneity in households' risk aversion and choice sets, with unrestricted dependence between them. Motivation for why unobserved heterogeneity in choice sets might be an important factor in this empirical framework comes from the earlier analysis of bar:mol:tei16 and novel findings that are part of bar:cou:mol:tei18's bar:cou:mol:tei18 contribution. They show that commonly used models that make strong assumptions about choice sets (e.g., the mixed logit model with each individual's choice set assumed equal to the feasible set, and various models of choice set formation) can be rejected in their data. With regard to risk aversion, their key finding is that their estimated lower bounds are significantly smaller than the point estimates obtained in the related literature. This suggests that the data can be explained by expected utility theory with lower and more homogeneous levels of risk aversion than it had been uncovered before. This provides new evidence on the importance of developing models that differ in their specification of which alternatives agents evaluate (rather than or in addition to models focusing on how they evaluate them), and to data collection efforts that seek to directly measure agents' heterogeneous choice sets cap16.
iar:shi:shu18 study the effect of pre-vote deliberation on the decisions of US appellate courts. The question of interest is weather deliberation increases or reduces the probability of an incorrect decision. They use a model where communication equilibrium is the solution concept, and only observed heterogeneity in payoffs is allowed for. In the model, multiple equilibria are again possible, and the authors leave the selection mechanism completely unspecified. They characterize $\mathcal{H}_\mathsf{P}[\theta]$ through an optimization problem, and structurally estimate the model on US Courts of Appeal data. iar:shi:shu18 compare the probability of making incorrect decisions under the pre-vote deliberation mechanism, to that in a counterfactual environment where no deliberation occurs. The results suggest that there is a range of parameters in $\mathcal{H}_\mathsf{P}[\theta]$, for which judges have ex-ante disagreement of imprecise prior information, for which deliberation is beneficial. Otherwise deliberation leads to lower effectiveness for the court.
dha:gai:mau18 propose a test for the hypothesis of rational expectations for the case that one observes only the marginal distributions of realizations and subjective beliefs, but not their joint distribution (e.g., when subjective beliefs are observed in one dataset, and realizations in a different one, and the two cannot be matched). They establish that the hypothesis of rational expectations can be expressed as testing that a continuum of moment inequalities is satisfied, and they leverage the results in and:shi17 to provide a simple-to-compute test for this hypothesis. They apply their method to test for and quantify deviations from rational expectations about future earnings, and examine the consequences of such departures in the context of a life-cycle model of consumption.
teb:tor:yan19 estimate the demand for health insurance under the Affordable Care Act using data from California. Methodologically, they use a discrete choice model that allows for endogeneity in insurance premiums (which enter as explanatory variables in the model) and dispenses with parametric assumptions about the unobserved components of utility leveraging the availability of instrumental variables, similarly to the framework presented in Section (ref). The authors provide a characterization of sharp bounds on the effects of changing premium subsidies on coverage choices, consumer surplus, and government spending, as solutions to linear programming problems, rendering their method computationally attractive.
Another important strand of theoretical literature is concerned with partial identification of panel data models. hon:tam06 consider a dynamic random effects probit model, and use partial identification analysis to obtain bounds on the model parameters that circumvent the initial conditions problem. ros12 considers a fixed effect panel data model where he imposes a conditional quantile restriction on time varying unobserved heterogeneity. Differencing out inequalities resulting from the conditional quantile restriction delivers inequalities that depend only on observable variables and parameters to be estimated, but not on the fixed effects, so that they can be used for estimation. che:fer:hah:new13 obtain bounds on average and quantile treatment effects in nonparametric and semiparametric nonseparable panel data models. kha:pon:tam16 provide partial identification results in linear panel data models when censored outcomes, with unrestricted dependence between censoring and observable and unobservable variables. Their results are derived for two classes of models, one where the unobserved heterogeneity terms satisfy a stationarity restriction, and one where they are nonstationary but satisfy a conditional independence restriction. tor19 provides a method to partially identify state dependence in panel data models where individual unobserved heterogeneity needs not be time invariant. pak:por16 study semiparametric multinomial choice panel models with fixed effects where the random utility function is assumed additively separable in unobserved heterogeneity, fixed effects, and a linear covariate index. The key semiparametric assumption is a group stationarity condition on the disturbances which places no restrictions on either the joint distribution of the disturbances across choices or the correlation of disturbances across time. pak:por16 propose a within-group comparison that delivers a collection of conditional moment inequalities that they use to provide point and partial identification results. ari19 proposes a related method, where partial identification relies on the observation of individuals whose outcome changes in two consecutive time periods, and leverages shape restrictions to reduce the number of between alternatives comparisons needed to determine the optimal choice.
The identification analysis carried out in Sections (ref)-(ref) presumes knowledge of the joint distribution $\mathsf{P}$ of the observable variables. That is, it presumes that $\mathsf{P}$ can be learned with certainty from observation of the entire population. In practice, one observes a sample of size $n$ drawn from $\mathsf{P}$. For simplicity I assume it to be a random sample.\footnote{This assumption is often maintained in the literature. See, e.g., and:soa10 for a treatment of inference with dependent observations. eps:kai:seo16 study inference in games of complete information as in Identification Problem (ref), imposing the i.i.d. assumption on the unobserved payoff shifters $\{\varepsilon_{i1},\varepsilon_{i2}\}_{i=1}^n$. The authors note that because the selection mechanism picking the equilibrium played in the regions of multiplicity (see Section (ref)) is left completely unspecified and may be arbitrarily correlated across markets, the resulting observed variables $\{{\boldsymbol{w}}_i\}_{i=1}^n$ may not be independent and identically distributed, and they propose an inference method to address this issue.}
Statistical inference on $\mathcal{H}_\mathsf{P}[\theta]$ needs to be conducted using knowledge of $\mathsf{P}_n$, the empirical distribution of the observable outcomes and covariates. Because $\mathcal{H}_\mathsf{P}[\theta]$ is not a singleton, this task is particularly delicate. To start, care is required to choose a proper notion of consistency for a set estimator $\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]$ and to obtain palatable conditions under which such consistency attains. Next, the asymptotic behavior of statistics designed to test hypothesis or build confidence sets for $\mathcal{H}_\mathsf{P}[\theta]$ or for $\vartheta\in\mathcal{H}_\mathsf{P}[\theta]$ might change with $\vartheta$, creating technical challenges for the construction of confidence sets that are not encountered when $\theta$ is point identified. Many of the sharp identification regions derived in Sections (ref)-(ref) can be written as collections of vectors $\vartheta\in\Theta$ that satisfy conditional or unconditional moment (in)equalities. For simplicity, I assume that $\Theta$ is a compact and convex subset of $\mathbb{R}^d$, and I use the formalization for the case of a finite number of unconditional moment (in)equalities:
In (ref), ${\boldsymbol{w}}_i\in\mathcal{W}\subseteq\mathbb{R}^{d_\mathcal{W}}$ is a random vector collecting all observable variables, with ${\boldsymbol{w}}\sim\mathsf{P}$; $m_j:\mathcal{W}\times\Theta\to\mathbb{R}$, $j\in\mathcal{J}\equiv\mathcal{J}_1\cup\mathcal{J}_2$, are known measurable functions characterizing the model; and $\mathcal{J}$ is a finite set equal to $\{1,\dots,|\mathcal{J}|\}$.\footnote{Examples where the set $\mathcal{J}$ is a compact set (e.g., a unit ball) rather than a finite set include the case of best linear prediction with interval outcome and covariate data, see characterization (ref) on p. (ref), and the case of entry games with multiple mixed strategy Nash equilibria, see characterization (ref) on p. (ref). A more general continuum of inequalities is also possible, as in the case of discrete choice with endogenous explanatory variables, see characterization (ref) on p. (ref). I refer to and:shi17 and ber:mol:mol11 for inference methods in the presence of a continuum of conditional moment (in)equalities.} Instances where $\mathcal{H}_\mathsf{P}[\theta]$ is characterized through a finite number of conditional moment (in)equalities and the conditioning variables have finite support can easily be recast as in (ref).\footnote{I refer to kha:tam09, and:shi13, che:lee:ros13, lee:son:wha13, arm14b,arm15, arm:cha16, che:che:kat18, and che18, for inference methods in the case that the conditioning variables have a continuous distribution.} Consider, for example, the two player entry game model in Identification Problem (ref) on p. (ref), where ${\boldsymbol{w}}=({\boldsymbol{y}}_1,{\boldsymbol{y}}_2,{\boldsymbol{x}}_1,{\boldsymbol{x}}_2)$. Using (in)equalities (ref)-(ref) and assuming that the distribution of $({\boldsymbol{x}}_1,{\boldsymbol{x}}_2)$ has $\bar{k}$ points of support, denoted $(x_{1,k},x_{2,k}),k=1,\dots,\bar{k}$, we have $|\mathcal{J}|=4\bar{k}$ and for $k=1,\dots,\bar{k}$,\footnote{In these expressions an index of the form $jk$ not separated by a comma equals the product of $j$ with $k$.}
In point identified moment equality models it has been common to conduct estimation and inference using a criterion function that aggregates moment violations han82. man:tam02 adapt this idea to the partially identified case, through a criterion function $q_\mathsf{P}:\Theta\to\mathbb{R}_+$ such that $q_\mathsf{P}(\vartheta)=0$ if and only if $\vartheta\in\mathcal{H}_\mathsf{P}[\theta]$. Many criterion functions can be used man:tam02,che:hon:tam07,rom:sha08,ros08,gal:hen09,and:gug09b,and:soa10,can10,rom:sha10. Some simple and commonly employed ones include
where $[x]_+=\max\{x,0\}$ and $\sigma_{\mathsf{P},j}(\vartheta)$ is the population standard deviation of $m_j({\boldsymbol{w}}_i;\vartheta)$. In (ref)-(ref) the moment functions are standardized, as doing so is important for statistical power and:soa10. To simplify notation, I omit the label and simply use $q_\mathsf{P}(\vartheta)$. Given the criterion function, one can rewrite (ref) as
To keep this chapter to a manageable length, I focus my discussion of statistical inference exclusively on consistent estimation and on different notions of coverage that a confidence set may be required to satisfy and that have proven useful in the literature.\footnote{Using the well known duality between tests of hypotheses and confidence sets, the discussion could be re-framed in terms of size of the test.} The topics of test of hypotheses and construction of confidence sets in partially identified models are covered in can:sha17, who provide a comprehensive survey devoted entirely to them in the context of moment inequality models. mol:mol18 provide a thorough discussion of related methods based on the use of random set theory.
When the identified object is a set, it is natural that its estimator is also a set. In order to discuss statistical properties of a set-valued estimator $\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]$ (to be defined below), and in particular its consistency, one needs to specify how to measure the distance between $\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]$ and $\mathcal{H}_\mathsf{P}[\theta]$. Several distance measures among sets exist mo1. A natural generalization of the commonly used Euclidean distance is the Hausdorff distance, see Definition (ref), which for given $A,B\subset\mathbb{R}^d$ can be written as
with $\mathbf{d}(a,B)\equiv\inf_{b\in B}\Vert a-b\Vert$.\footnote{The definition of the Hausdorff distance can be generalized to an arbitrary metric space by replacing the Euclidean metric by the metric specified on that space.} In words, the Hausdorff distance between two sets measures the furthest distance from an arbitrary point in one of the sets to its closest neighbor in the other set. It is easy to verify that $\mathbf{d}_H$ metrizes the family of non-empty compact sets; in particular, given non-empty compact sets $A,B\subset\mathbb{R}^d$, $\mathbf{d}_H(A,B) =0$ if and only if $A=B$. If either $A$ or $B$ is empty, $\mathbf{d}_H(A,B) =\infty$.
The use of the Hausdorff distance to conceptualize consistency of set valued estimators in econometrics was proposed by han:hea:lut95 and man:tam02.\footnote{It was previously used in the mathematical literature on random set theory, for example to formalize laws of large numbers and central limit theorems for random sets such as the ones in Theorems (ref) and (ref) art:vit75,gin:hah:zin83.}
mol98 establishes Hausdorff consistency of a plug-in estimator of the set $\{\vartheta\in\Theta:g_\mathsf{P}(\vartheta)\le 0\}$, with $g_\mathsf{P}:\mathcal{W}\times\Theta \to \mathbb{R}$ a lower semicontinuous function of $\vartheta\in\Theta$ that can be consistently estimated by a lower semicontinuous function $g_n$ uniformly over $\Theta$. The set estimator is $\{\vartheta\in\Theta:g_n(\vartheta)\le 0\}$. The fundamental assumption in mol98 is that $\{\vartheta\in\Theta:g_\mathsf{P}(\vartheta)\le 0\}\subseteq\operatorname{cl}(\{\vartheta\in\Theta:g_\mathsf{P}(\vartheta)< 0\})$, see mol:mol18 for a discussion. There are important applications where this condition holds. che:koc:men15 provide results related to mol98, as well as important extensions for the construction of confidence sets, and show that these can be applied to carry out statistical inference on the Hansen–Jagannathan sets of admissible stochastic discount factors han:jag91, the Markowitz–Fama mean–variance sets for asset portfolio returns mar52, and the set of structural elasticities in che12b's analysis of demand with optimization frictions. However, these methods are not broadly applicable in the general moment (in)equalities framework of this section, as mol98's key condition generally fails for the set $\mathcal{H}_\mathsf{P}[\theta]$ in (ref).
man:tam02 extend the standard theory of extremum estimation of point identified parameters to partial identification, and propose to estimate $\mathcal{H}_\mathsf{P}[\theta]$ using the collection of values $\vartheta\in\Theta$ that approximately minimize a sample analog of $q_\mathsf{P}$:
with $\tau_n$ a sequence of non-negative random variables such that $\tau_n\stackrel{p}{\rightarrow} 0$. In (ref), $q_n(\vartheta)$ is a sample analog of $q_\mathsf{P}(\vartheta)$ that replaces $\mathbb{E}_\mathsf{P}(m_j({\boldsymbol{w}}_i;\vartheta))$ and $\sigma_{\mathsf{P},j}(\vartheta)$ in (ref)-(ref) with properly chosen estimators, e.g.,
It can be shown that as long as $\tau_n=o_p(1)$, under the same assumptions used to prove consistency of extremum estimators of point identified parameters (e.g., with uniform convergence of $q_n$ to $q_\mathsf{P}$ and continuity of $q_\mathsf{P}$ on $\Theta$),
This yields that asymptotically each point in $\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]$ is arbitrarily close to a point in $\mathcal{H}_\mathsf{P}[\theta]$, or more formally that $\mathsf{P}(\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]\subseteq\mathcal{H}_\mathsf{P}[\theta])\to 1$. I refer to (ref) as inner consistency henceforth.\footnote{See ble15 for a pedagogically helpful proof for a semiparametric binary model.} red81 provides an early contribution establishing this type of inner consistency for maximum likelihood estimators when the true parameter is not point identified.
However, Hausdorff consistency requires also that
i.e., that each point in $\mathcal{H}_\mathsf{P}[\theta]$ is arbitrarily close to a point in $\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]$, or more formally that $\mathsf{P}(\mathcal{H}_\mathsf{P}[\theta]\subseteq\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta])\to 1$. To establish this result for the sharp identification regions in Theorem SIR-(ref) (parametric regression with interval covariate) and Theorem SIR-(ref) (semiparametric binary model with interval covariate), man:tam02 require the rate at which $\tau_n\stackrel{p}{\rightarrow} 0$ to be slower than the rate at which $q_n$ converges uniformly to $q_\mathsf{P}$ over $\Theta$.
What might go wrong in the absence of such a restriction? A simple example can help understand the issue. Consider a model with linear inequalities of the form
Suppose ${\boldsymbol{w}}\equiv({\boldsymbol{w}}_1,\dots,{\boldsymbol{w}}_6)$ is distributed multivariate normal, with $\mathbb{E}_\mathsf{P}({\boldsymbol{w}})=[6~0~2~0~{-2}~0]^\top$ and $\operatorname{Cov}_\mathsf{P}({\boldsymbol{w}})$ equal to the identity matrix. Then $\mathcal{H}_\mathsf{P}[\theta]=\{\vartheta=[\vartheta_1~\vartheta_2]^\top\in\Theta:\vartheta_1\in[0,6]~\text{and}~\vartheta_2=2\}$. However, with positive probability in any finite sample $q_n(\vartheta)=0$ for $\vartheta$ in a random region (e.g., a triangle if $q_n$ is the sample analog of (ref)) that only includes points that are close to a subset of the points in $\mathcal{H}_\mathsf{P}[\theta]$. Hence, with positive probability the minimizer of $q_n$ cycles between consistent estimators of subsets of $\mathcal{H}_\mathsf{P}[\theta]$, but does not estimate the entire set. Enlarging the estimator to include all points that are close to minimizing $q_n$ up to a tolerance that converges to zero sufficiently slowly removes this problem.
che:hon:tam07 significantly generalize the consistency results in man:tam02. They work with a normalized criterion function equal to $q_n(\vartheta)-\inf_{\tilde\vartheta\in\Theta}q_n(\tilde\vartheta)$, but to keep notation light I simply refer to it as $q_n$.\footnote{Using this normalized criterion function is especially important in light of possible model misspecification, see Section (ref).} Under suitable regularity conditions, they establish consistency of an estimator that can be a smaller set than the one proposed by man:tam02, and derive its convergence rate. Some of the key conditions required by che:hon:tam07 to study convergence rates include that $q_n$ is lower semicontinuous in $\vartheta$, satisfies various convergence properties among which $\sup_{\vartheta\in\mathcal{H}_\mathsf{P}[\theta]}q_n=O_p(1/a_n)$ for a sequence of normalizing constants $a_n\to\infty$, that $\tau_n\ge \sup_{\vartheta\in\mathcal{H}_\mathsf{P}[\theta]}q_n(\vartheta)$ with probability approaching one, and that $\tau_n\to 0$. They also require that there exist positive constants $(\delta,\kappa,\gamma)$ such that for any $\epsilon\in(0,1)$ there are $(d_\epsilon,n_\epsilon)$ such that
uniformly on $\{\vartheta\in\Theta:\mathbf{d}(\vartheta,\mathcal{H}_\mathsf{P}[\theta])\ge(d_\epsilon/a_n)^{1/\gamma}\}$ with probability at least $1-\epsilon$. In words, the assumption, referred to as polynomial minorant condition, rules out that $q_n$ can be arbitrarily close to zero outside $\mathcal{H}_\mathsf{P}[\theta]$. It posits that $q_n$ changes as at least a polynomial of degree $\gamma$ in the distance of $\vartheta$ from $\mathcal{H}_\mathsf{P}[\theta]$. Under some additional regularity conditions, che:hon:tam07 establish that
What is the role played by the polynomial minorant condition for the result in (ref)? Under the maintained assumptions $\tau_n\ge \sup_{\vartheta\in\mathcal{H}_\mathsf{P}[\theta]}q_n(\vartheta)\ge\kappa[\min\{\delta,\mathbf{d}(\vartheta,\mathcal{H}_\mathsf{P}[\theta])\}]^\gamma$, and the latter part of the inequality is used to obtain (ref). When could the polynomial minorant condition be violated? In moment (in)equalities models, che:hon:tam07 require $\gamma=2$.\footnote{che:hon:tam07 set $\gamma=1$ because they report the assumption for a criterion function that does not square the moment violations.} Consider a simple stylized example with (in)equalities of the form
with $\mathbb{E}_\mathsf{P}({\boldsymbol{w}}_1)=\mathbb{E}_\mathsf{P}({\boldsymbol{w}}_2)=\mathbb{E}_\mathsf{P}({\boldsymbol{w}}_3)=0$, and note that the sample means $(\bar{{\boldsymbol{w}}}_1,\bar{{\boldsymbol{w}}}_2,\bar{{\boldsymbol{w}}}_3)$ are $\sqrt{n}$-consistent estimators of $(\mathbb{E}_\mathsf{P}({\boldsymbol{w}}_1),\mathbb{E}_\mathsf{P}({\boldsymbol{w}}_2),\mathbb{E}_\mathsf{P}({\boldsymbol{w}}_3))$. Suppose $({\boldsymbol{w}}_1,{\boldsymbol{w}}_2,{\boldsymbol{w}}_3)$ are distributed multivariate standard normal. Consider a sequence $\vartheta_n=[\vartheta_{1n}~\vartheta_{2n}]^\top=[n^{-1/4}~n^{-1/4}]^\top$. Then $[\mathbf{d}(\vartheta_n,\mathcal{H}_\mathsf{P}[\theta])]^\gamma=O_p(n^{-1/2})$. On the other hand, with positive probability $q_n(\vartheta_n)=(\bar{{\boldsymbol{w}}}_3-\vartheta_{1n}\vartheta_{2n})^2=O_p\left(n^{-1}\right)$, so that for $n$ large enough $q_n(\vartheta_n)<[\mathbf{d}(\vartheta_n,\mathcal{H}_\mathsf{P}[\theta])]^\gamma$, violating the assumption. This occurs because the gradient of the moment equality vanishes as $\vartheta$ approaches zero, rendering the criterion function flat in a neighborhood of $\mathcal{H}_\mathsf{P}[\theta]$. As intuition would suggest, rates of convergence are slower the flatter $q_n$ is outside $\mathcal{H}_\mathsf{P}[\theta]$.
kai:mol:sto19CQ show that in moment inequality models with smooth moment conditions, the polynomial minorant assumption with $\gamma=2$ implies the Abadie constraint qualification (ACQ); see, e.g., baz:she:she06 for a definition and discussion of ACQ. The example just given to discuss failures of the polynomial minorant condition is in fact a known example where ACQ fails at $\vartheta=[0~0]^\top$.
che:hon:tam07 also consider the case that $q_n$ vanishes on subsets of $\Theta$ that converge in Hausdorff distance to $\mathcal{H}_\mathsf{P}[\theta]$ at rate $a_n^{-1/\gamma}$. While degeneracy might be difficult to verify in practice, che:hon:tam07 show that if it holds, $\tau_n$ can be set to zero. yil12 provides conditions on the moment functions, which are closely related to constraint qualifications kai:mol:sto19CQ under which it is possible to set $\tau_n=0$.
men14 studies estimation of $\mathcal{H}_\mathsf{P}[\theta]$ when the number of moment inequalities is large relative to sample size (possibly infinite). He provides a consistency result for criterion-based estimators that use a number of unconditional moment inequalities that grows with sample size. He also considers estimators based on conditional moment inequalities, and derives the fastest possible rate for estimating $\mathcal{H}_\mathsf{P}[\theta]$ under smoothness conditions on the conditional moment functions. He shows that the rates achieved by the procedures in arm14b,arm15 are (minimax) optimal, and cannot be improved upon.
ber:mol08 introduce to the econometrics literature inference methods for set valued estimators based on random set theory. They study the class of models where $\mathcal{H}_\mathsf{P}[\theta]$ is convex and can be written as the Aumann (or selection) expectation of a properly defined random closed set.\footnote{By Theorem (ref), the Aumann expectation of a random closed set defined on a nonatomic probability space is convex. In this chapter I am assuming nonatomicity of the probability space. Even if I did not make this assumption, however, when working with a random sample the relevant probability space is the product space with $n\to\infty$, hence nonatomic art:vit75. If $\mathcal{H}_\mathsf{P}[\theta]$ is not convex, ber:mol08's analysis applies to its convex hull.} They propose to carry out estimation and inference leveraging the representation of convex sets through their support function (given in Definition (ref)), as it is done in random set theory; see mo1 and mol:mol18. Because the support function fully characterizes the boundary of $\mathcal{H}_\mathsf{P}[\theta]$, it allows for a simple sample analog estimator, and for inference procedures with desirable properties.
An example of a framework where the approach of ber:mol08 can be applied is that of best linear prediction with interval outcome data in Identification Problem (ref).\footnote{kai:mol:sto19 establish that if ${\boldsymbol{x}}$ has finite support, $\mathcal{H}_\mathsf{P}[\theta]$ in Theorem SIR-(ref) can be written as the collection of $\vartheta\in\Theta$ that satisfy a finite number of moment inequalities, as posited in this section.} Recall that in that case, the researcher observes random variables $({\boldsymbol{y}}_{\mathrm{L}},{\boldsymbol{y}}_{\mathrm{U}},{\boldsymbol{x}})$ and wishes to learn the best linear predictor of ${\boldsymbol{y}}|{\boldsymbol{x}}$, with ${\boldsymbol{y}}$ unobserved and $\mathsf{R}({\boldsymbol{y}}_{\mathrm{L}}\le{\boldsymbol{y}}\le{\boldsymbol{y}}_{\mathrm{U}})=1$. For simplicity let ${\boldsymbol{x}}$ be a scalar. Given a random sample $\{{\boldsymbol{y}}_{\mathrm{L}i},{\boldsymbol{y}}_{\mathrm{U}i},{\boldsymbol{x}}_i\}_{i=1}^n$ from $\mathsf{P}$, the researcher can construct a random segment ${\boldsymbol{G}}_i$ for each $i$ and a consistent estimator $\hat{\Sigma}_n$ of the random matrix $\Sigma_\mathsf{P}$ in (ref) as
where ${\boldsymbol{Y}}_i=[{\boldsymbol{y}}_{\mathrm{L}i},{\boldsymbol{y}}_{\mathrm{U}i}]$ and $\overline{\boldsymbol{x}},\overline{{\boldsymbol{x}}^2}$ are the sample means of ${\boldsymbol{x}}_i$ and ${\boldsymbol{x}}^2_i$ respectively. Because in this problem $\mathcal{H}_\mathsf{P}[\theta]=\Sigma_\mathsf{P}^{-1}\mathbb{E}_\mathsf{P}{\boldsymbol{G}}$ (see Theorem SIR-(ref) on p. (ref)), a natural sample analog estimator replaces $\Sigma_\mathsf{P}$ with $\hat{\Sigma}_n$, and $\mathbb{E}_\mathsf{P}{\boldsymbol{G}}$ with a Minkowski average of ${\boldsymbol{G}}_i$ (see Appendix (ref), p. (ref) for a formal definition), yielding
The support function of $\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]$ is the sample analog of that of $\mathcal{H}_\mathsf{P}[\theta]$ provided in (ref):
where $f({\boldsymbol{x}}_i,u)=[1~{\boldsymbol{x}}_i]\hat\Sigma_n^{-1}u$. ber:mol08 use the Law of Large Numbers for random sets reported in Theorem (ref) to show that $\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]$ in (ref) is $\sqrt{n}$-consistent under standard conditions on the moments of $({\boldsymbol{y}}_{\mathrm{L}i},{\boldsymbol{y}}_{\mathrm{U}i},{\boldsymbol{x}}_i)$.
bon:mag:mau12 and cha:che:mol:sch18 significantly expand the applicability of ber:mol08's ber:mol08 estimator. bon:mag:mau12 show that it can be used in a large class of partially identified linear models, including ones that allow for the availability of instrumental variables. cha:che:mol:sch18 show that it can be used for best linear approximation of any function $f(x)$ that is known to lie within two identified bounding functions. The lower and upper functions defining the band are allowed to be any functions, including ones carrying an index, and can be estimated parametrically or nonparametrically. The method allows for estimation of the parameters of the best linear approximations to the set identified functions in many of the identification problems described in Section (ref). It can also be used to estimate the sharp identification region for the parameters of a binary choice model with interval or discrete regressors under the assumptions of mag:mau08, characterized in (ref) in Section (ref).
kai:san14 develop a theory of efficiency for estimators of sets $\mathcal{H}_\mathsf{P}[\theta]$ as in (ref) under the additional requirements that the inequalities $\mathbb{E}_\mathsf{P}(m_j({\boldsymbol{w}},\vartheta))$ are convex in $\vartheta\in\Theta$ and smooth as functionals of the distribution of the data. Because of the convexity of the moment inequalities, $\mathcal{H}_\mathsf{P}[\theta]$ is convex and can be represented through its support function. Using the classic results in bic:kla:rit:wel93, kai:san14 show that under suitable regularity conditions, the support function admits for $\sqrt{n}$-consistent regular estimation. They also show that a simple plug-in estimator based on the support function attains the semiparametric efficiency bound, and the corresponding estimator of $\mathcal{H}_\mathsf{P}[\theta]$ minimizes a wide class of asymptotic loss functions based on the Hausdorff distance. As they establish, this efficiency result applies to the estimators proposed by ber:mol08, including that in (ref), and by bon:mag:mau12.
kai16 further enlarges the applicability of the support function approach by establishing its duality with the criterion function approach, for the case that $q_\mathsf{P}$ is a convex function and $q_n$ is a convex function almost surely. This allows one to use the support function approach also when a representation of $\mathcal{H}_\mathsf{P}[\theta]$ as the Aumann expectation of a random closed set is not readily available. kai16 considers $\mathcal{H}_\mathsf{P}[\theta]$ and its level set estimator $\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]$ as defined, respectively, in (ref) and (ref), with $\Theta$ a convex subset of $\mathbb{R}^d$. Because $q_\mathsf{P}$ and $q_n$ are convex functions, $\mathcal{H}_\mathsf{P}[\theta]$ and $\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]$ are convex sets. Under the same assumptions as in che:hon:tam07, including the polynomial minorant and the degeneracy conditions, one can set $\tau_n=0$ and have $\mathbf{d}_H(\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta],\mathcal{H}_\mathsf{P}[\theta])=O_p(a_n^{-1/\gamma})$. Moreover, due to its convexity, $\mathcal{H}_\mathsf{P}[\theta]$ is fully characterized by its support function, which in turn can be consistently estimated (at the same rate as $\mathcal{H}_\mathsf{P}[\theta]$) using sample analogs as $h_{\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]}(u)=\max_{a_nq_n(\vartheta)\le 0}u^\top\vartheta$. The latter can be computed via convex programming.
kit:gia18 consider consistent estimation of $\mathcal{H}_\mathsf{P}[\theta]$ in the context of Bayesian inference. They focus on partially identified models where $\mathcal{H}_\mathsf{P}[\theta]$ depends on a “reduced form" parameter $\phi$ (e.g., a vector of moments of observable random variables). They recognize that while a prior on $\phi$ can be revised in light of the data, a prior on $\theta$ cannot, due to the lack of point identification. As such they propose to choose a single prior for the revisable parameters, and a set of priors for the unrevisable ones. The latter is the collection of priors such that the distribution of $\theta|\phi$ places probability one on $\mathcal{H}_\mathsf{P}[\theta]$. A crucial observation in kit:gia18 is that once $\phi$ is viewed as a random vector, as in the Bayesian paradigm, under mild regularity conditions $\mathcal{H}_\mathsf{P}[\theta]$ is a random closed set, and Bayesian inference on it can be carried out using elements of random set theory. In particular, they show that the set of posterior means of $\theta|{\boldsymbol{w}}$ equals the Aumann expectation of $\mathcal{H}_\mathsf{P}[\theta]$ (with the underlying probability measure of $\phi|{\boldsymbol{w}}$). They also show that this Aumann expectation converges in Hausdorff distance to the “true" identified set if the latter is convex, or otherwise to its convex hull. They apply their method to analyze impulse-response in set-identified Structural Vector Autoregressions, where standard Bayesian inference is otherwise sensitive to the choice of an unrevisable prior.\footnote{There is a large literature in macro-econometrics, pioneered by fau98, can:den02, and uhl05, concerned with Bayesian inference with a non-informative prior for non-identified parameters. I refer to kil:lut17 for a thorough review. Frequentist inference for impulse response functions in Structural Vector Autoregression models is carried out, e.g., in gra:moo:sch18 and gaf:mei:mon18. }
che:lee:ros13 propose an alternative to the notion of consistent estimator. Rather than asking that $\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]$ satisfies the requirement in Definition (ref), they propose the notion of half-median-unbiased estimator. This notion is easiest to explain in the case of interval identified scalar parameters. Take, e.g., the bound in Theorem SIR-(ref) for the conditional expectation of selectively observed data. Then an estimator of that interval is half-median-unbiased if the estimated upper bound exceeds the true upper bound, and the estimated lower bound falls below the true lower bound, each with probability at least $1/2$ asymptotically. More generally, one can obtain a half-median-unbiased estimator as
where $c_{1/2}(\vartheta)$ is a critical value chosen so that $\hat{\mathcal{H}}_{\mathsf{P}_n}[\theta]$ asymptotically contains $\mathcal{H}_\mathsf{P}[\theta]$ (or any fixed element in $\mathcal{H}_\mathsf{P}[\theta]$; see the discussion in Section (ref) below) with at least probability $1/2$. As discussed in the next section, $c_{1/2}(\vartheta)$ can be further chosen so that this probability is uniform over $\mathsf{P}\in\mathcal{P}$.
The requirement of half-median unbiasedness has the virtue that, by construction, an estimator such as (ref) is a subset of a $1-\alpha$ confidence set as defined in (ref) below for any $\alpha<1/2$, provided $c_{1-\alpha}(\vartheta)$ is chosen using the same criterion for all $\alpha\in(0,1)$. In contrast, a consistent estimator satisfying the requirement in Definition (ref) needs not be a subset of a confidence set. This is because the sequence $\tau_n$ in (ref) may be larger than the critical value used to obtain the confidence set, see equation (ref) below, unless regularity conditions such as degeneracy or others allow one to set $\tau_n$ equal to zero. Moreover, choice of the sequence $\tau_n$ is not data driven, and hence can be viewed as arbitrary. This raises a concern for the scope of consistent estimation in general settings.
However, reporting a set estimator together with a confidence set is arguably important to shed light on how much of the volume of the confidence set is due to statistical uncertainty and how much is due to a large identified set. One can do so by either using a half-median unbiased estimator as in (ref), or the set of minimizers of the criterion function in (ref) with $\tau_n=0$ (which, as previously discussed, satisfies the inner consistency requirement in (ref) under weak conditions, and is Hausdorff consistent in some well behaved cases).
I first discuss confidence sets $CS_n\subset\mathbb{R}^d$ defined as level sets of a criterion function. To simplify notation, henceforth I assume $a_n=n$.
In (ref), $c_{1-\alpha}(\vartheta)$ may be constant or vary in $\vartheta\in\Theta$. It is chosen to that $CS_n$ satisfies (asymptotically) a certain coverage property with respect to either $\mathcal{H}_\mathsf{P}[\theta]$ or each $\vartheta\in\mathcal{H}_\mathsf{P}[\theta]$. Correspondingly, different appearances of $c_{1-\alpha}(\vartheta)$ may refer to different critical values associated with different coverage notions. The challenging theoretical aspect of inference in partial identification is the determination of $c_{1-\alpha}$ and of methods to approximate it.
A first classification of coverage notions pertains to whether the confidence set should cover $\mathcal{H}_\mathsf{P}[\theta]$ or each of its elements with a prespecified asymptotic probability. Early on, within the study of interval-identified parameters, hor:man98,hor:man00 put forward a confidence interval that expands each of the sample analogs of the extreme points of the population bounds by an amount designed so that the confidence interval asymptotically covers the population bounds with prespecified probability.
che:hon:tam07 study the general problem of inference for a set $\mathcal{H}_\mathsf{P}[\theta]$ defined as the zero-level set of a criterion function. The coverage notion that they propose is pointwise coverage of the set, whereby $c_{1-\alpha}$ is chosen so that:
che:hon:tam07 provide conditions under which $CS_n$ satisfies (ref) with $c_{1-\alpha}$ constant in $\vartheta$, yielding the so called criterion function approach to statistical inference in partial identification. Under the same coverage requirement, bug10 and gal:hen13 introduce novel bootstrap methods for inference in moment inequality models. hen:mea:que15 propose an inference method for finite games of complete information that exploits the structure of these models.
ber:mol08 propose a method to test hypotheses and build confidence sets satisfying (ref) based on random set theory, the so called support function approach, which yields simple to compute confidence sets with asymptotic coverage equal to $1-\alpha$ when $\mathcal{H}_\mathsf{P}[\theta]$ is strictly convex. The reason for the strict convexity requirement is that in its absence, the support function of $\mathcal{H}_\mathsf{P}[\theta]$ is not fully differentiable, but only directionally differentiable, complicating inference. Indeed, fan:san18 show that standard bootstrap methods are consistent if and only if full differentiability holds, and they provide modified bootstrap methods that remain valid when only directional differentiability holds. cha:che:mol:sch18 propose a data jittering method that enforces full differentiability at the price of a small conservative distortion. kai:san14 extend the applicability of the support function approach to other moment inequality models and establish efficiency results. che:koc:men15 show that an Hausdorff distance-based test statistic can be weighted to enforce either exact or first-order equivariance to transformations of parameters. adu:ots16 provide empirical likelihood based inference methods for the support function approach. The test statistics employed in the criterion function approach and in the support function approach are asymptotically equivalent in specific moment inequality models ber:mol08,kai16, but the criterion function approach is more broadly applicable.
The field's interest changed to a different notion of coverage when imb:man04 pointed out that often there is one “true" data generating $\theta$, even if it is only partially identified. Hence, they proposed confidence sets that cover each $\vartheta\in\mathcal{H}_\mathsf{P}[\theta]$ with a prespecified probability. For pointwise coverage, this leads to choosing $c_{1-\alpha}$ so that:
If $\mathcal{H}_\mathsf{P}[\theta]$ is a singleton then (ref) and (ref) both coincide with the pointwise coverage requirement employed for point identified parameters. However, as shown in imb:man04, if $\mathcal{H}_\mathsf{P}[\theta]$ contains more than one element, the two notions differ, with confidence sets satisfying (ref) being weakly smaller than ones satisfying (ref). ros08 provides confidence sets for general moment (in)equalities models that satisfy (ref) and are easy to compute.
Although confidence sets that take each $\vartheta\in\mathcal{H}_\mathsf{P}[\theta]$ as the object of interest (and which satisfy the uniform coverage requirements described in Section (ref) below) have received the most attention in the literature on inference in partially identified models, this choice merits some words of caution. First, hen:ona12 point out that if confidence sets are to be used for decision making, a policymaker concerned with robust decisions might prefer ones satisfying (ref) (respectively, (ref) below once uniformity is taken into account) to ones satisfying (ref) (respectively, (ref) below with uniformity). Second, while in many applications a “true" data generating $\theta$ exists, in others it does not. For example, man:mol10 and giu:man:mol19 query survey respondents (in the American Life Panel and in the Health and Retirement Study, respectively) about their subjective beliefs on the probability chance of future events. A large fraction of these respondents, when given the possibility to do so, report imprecise beliefs in the form of intervals. In this case, there is no “true" point-valued belief: the “truth" is interval-valued. If one is interested in (say) average beliefs, the sharp identification region is the (Aumann) expectation of the reported intervals, and the appropriate coverage requirement for a confidence set is that in (ref) (respectively, (ref) below with uniformity).
In the context of interval identified parameters, such as, e.g., the mean with missing data in Theorem SIR-(ref) with $\theta\in\mathbb{R}$, imb:man04 pointed out that extra care should be taken in the construction of confidence sets for partially identified parameters, as otherwise they may be asymptotically valid only pointwise (in the distribution of the observed data) over relevant classes of distributions.\footnote{This discussion draws on many conversations with J\"{o}rg Stoye, as well as on notes that he shared with me, for which I thank him.} For example, consider a confidence interval that expands each of the sample analogs of the extreme points of the population bounds by a one-sided critical value. This confidence interval controls the asymptotic coverage probability pointwise for any DGP at which the width of the population bounds is positive. This is because the sampling variation becomes asymptotically negligible relative to the (fixed) width of the bounds, making the inference problem essentially one-sided. However, for every $n$ one can find a distribution $\mathsf{P}\in\mathcal{P}$ and a parameter $\vartheta\in\mathcal{H}_\mathsf{P}[\theta]$ such that the width of the population bounds (under $\mathsf{P}$) is small relative to $n$ and the coverage probability for $\vartheta$ is below $1-\alpha$. This happens because the proposed confidence interval does not take into account the fact that for some $\mathsf{P}\in\mathcal{P}$ the problem has a two-sided nature.
This observation naturally leads to a more stringent requirement of uniform coverage, whereby (ref)-(ref) are replaced, respectively, by
and $c_{1-\alpha}$ is chosen accordingly, to obtain either (ref) or (ref). Sets satisfying (ref) are referred to as confidence regions for $\mathcal{H}_\mathsf{P}[\theta]$ that are uniformly consistent in level (over $\mathsf{P}\in\mathcal{P}$). rom:sha10 propose such confidence regions, study their properties, and provide a step-down procedure to obtain them.
che:chr:tam18 propose confidence sets that are contour sets of criterion functions using cutoffs that are computed via Monte Carlo simulations from the quasi‐posterior distribution of the criterion and satisfy the coverage requirement in (ref). They recommend the use of a Sequential Monte Carlo algorithm that works well also when the quasi-posterior is irregular and multi-modal. They establish exact asymptotic coverage, non-trivial local power, and validity of their procedure in point identified and partially identified regular models, and validity in irregular models (e.g., in models where the reduced form parameters are on the boundary of the parameter space). They also establish efficiency of their procedure in regular models that happen to be point identified.
Sets satisfying (ref) are referred to as confidence regions for points in $\mathcal{H}_\mathsf{P}[\theta]$ that are uniformly consistent in level (over $\mathsf{P}\in\mathcal{P}$). Within the framework of imb:man04, sto09 shows that one can obtain a confidence interval satisfying (ref) by pre-testing whether the lower and upper population bounds are sufficiently close to each other. If so, the confidence interval expands each of the sample analogs of the extreme points of the population bounds by a two-sided critical value; otherwise, by a one-sided. sto09 provides important insights clarifying the connection between superefficient (i.e., faster than $O_p(1/\sqrt{n})$) estimation of the width of the population bounds when it equals zero, and certain challenges in imb:man04's proposed method.\footnote{Indeed, the confidence interval proposed by sto09 can be thought of as using a Hodges-type shrinkage estimator van97 for the width of the population bounds.} bon:mag:mau12 leverage sto09's results to obtain confidence sets satisfying (ref) using the support function approach for set identified linear models.
Obtaining confidence sets that satisfy the requirement in (ref) becomes substantially more complex in the context of general moment (in)equalities models. One of the key challenges to uniform inference stems from the fact that the behavior of the limit distribution of the test statistic depends on $\sqrt{n}\mathbb{E}_\mathsf{P}(m_j({\boldsymbol{w}}_i;\vartheta)),~j=1,\dots,|\mathcal{J}|$, which cannot be consistently estimated. rom:sha08,and:gug09b,and:soa10,can10,and:bar12,rom:sha:wol14, among others, make significant contributions to circumvent these difficulties in the context of a finite number of unconditional moment (in)equalities. and:shi13,che:lee:ros13,lee:son:wha13,arm14b,arm15,arm:cha16,che18, among others, make significant contributions to circumvent these difficulties in the context of a finite number of conditional moment (in)equalities (with continuously distributed conditioning variables). che:che:kat18 and and:shi17 study, respectively, the challenging frameworks where the number of moment inequalities grows with sample size and where there is a continuum of conditional moment inequalities.
I refer to can:sha17 for a thorough discussion of these methods and a comparison of their relative (de)merits bug:can:gug12,bug16.
The coverage requirements in (ref)-(ref) refer to confidence sets in $\mathbb{R}^d$ for the entire $\theta$ or $\mathcal{H}_\mathsf{P}[\theta]$. Often empirical researchers are interested in inference on a specific component or (smooth) function of $\theta$ (e.g., the returns to education; the effect of market size on the probability of entry; the elasticity of demand for insurance to price, etc.). For simplicity, here I focus on the case of a component of $\theta$, which I represent as $u^\top\theta$, with $u$ a standard basis vector in $\mathbb{R}^d$. In this case, the (sharp) identification region of interest is
One could report as confidence interval for $u^\top\theta$ the projection of $CS_n$ in direction $\pm u$. The resulting confidence interval is asymptotically valid but typically conservative. The extent of the conservatism increases with the dimension of $\theta$ and is easily appreciated in the case of a point identified parameter. Consider, for example, a linear regression in $\mathbb{R}^{10}$, and suppose for simplicity that the limiting covariance matrix of the estimator is the identity matrix. Then a 95% confidence interval for $u^\top\theta$ is obtained by adding and subtracting $1.96$ to that component's estimate. In contrast, projection of a 95% confidence ellipsoid for $\theta$ on each component amounts to adding and subtracting $4.28$ to that component's estimate.
It is therefore desirable to provide confidence intervals $CI_n$ specifically designed to cover $u^\top\theta$ rather then the entire $\theta$. Natural counterparts to (ref)-(ref) are
As shown in ber:mol08 and kai16 for the case of pointwise coverage, obtaining asymptotically valid confidence intervals is simple if the identified set is convex and one uses the support function approach. This is because it suffices to base the test statistic on the support function in direction $u$, and it is often possible to easily characterize the limiting distribution of this test statistic. See mol:mol18 for details.
The task is significantly more complex in general moment inequality models when $\mathcal{H}_\mathsf{P}[\theta]$ is non-convex and one wants to satisfy the criterion in (ref) or that in (ref). rom:sha08 and bug:can:shi17 propose confidence intervals of the form
where $\Theta(s)=\{\vartheta\in\Theta:u^\top\vartheta=s\}$ and $c_{1-\alpha}$ is such that (ref) holds. An important idea in this proposal is that of profiling the test statistic $nq_n(\vartheta)$ by minimizing it over all $\vartheta$s such that $u^\top\vartheta=s$. One then includes in the confidence interval all values $s$ for which the profiled test statistic's value is not too large. rom:sha08 propose the use of subsampling to obtain the critical value $c_{1-\alpha}(s)$ and provide high-level conditions ensuring that (ref) holds. bug:can:shi17 substantially extend and improve the profiling approach by providing a bootstrap-based method to obtain $c_{1-\alpha}$ so that (ref) holds. Their method is more powerful than subsampling (for reasonable choices of subsample size). bel:bug:che18 further enlarge the domain of applicability of the profiling approach by proposing a method based on this approach that is asymptotically uniformly valid when the number of moment conditions is large, and can grow with the sample size, possibly at exponential rates.
kai:mol:sto19 propose a bootstrap-based calibrated projection approach where
with
and $c_{1-\alpha}$ a critical level function calibrated so that (ref) holds. Compared to the simple projection of $CS_n$ mentioned at the beginning of Section (ref), calibrated projection (weakly) reduces the value of $c_{1-\alpha}$ so that the projection of $\theta$, rather than $\theta$ itself, is asymptotically covered with the desired probability uniformly.
che:chr:tam18 provide methods to build confidence intervals and confidence sets on projections of $\mathcal{H}_\mathsf{P}[\theta]$ as contour sets of criterion functions using cutoffs that are computed via Monte Carlo simulations from the quasi‐posterior distribution of the criterion, and that satisfy the coverage requirement in (ref). One of their procedures, designed specifically for scalar projections, delivers a confidence interval as the contour set of a profiled quasi-likelihood ratio with critical value equal to a quantile of the Chi-squared distribution with one degree of freedom.
The confidence sets discussed in this section are based on the frequentist approach to inference. It is natural to ask whether in partially identified models, as in well behaved point identified models, one can build Bayesian credible sets that at least asymptotically coincide with frequentist confidence sets. This question was first addressed by moo:sch12, with a negative answer for the case that the coverage in (ref) is sought out. In particular, they showed that the resulting Bayesian credible sets are a subset of $\mathcal{H}_\mathsf{P}[\theta]$, and hence too narrow from the frequentist perspective.
This discrepancy can be ameliorated when inference is sought out for $\mathcal{H}_\mathsf{P}[\theta]$ rather than for each $\vartheta\in\mathcal{H}_\mathsf{P}[\theta]$. nor:tan14, kli:tam16, kit:gia18, and lia:sim19 propose Bayesian credible regions that are valid for frequentist inference in the sense of (ref), where the first two build on the criterion function approach and the second two on the support function approach. All these contributions rely on the model being separable, in the sense that it yields moment inequalities that can be written as the sum of a function of the data only, and a function of the model parameters only (as in, e.g., (ref)-(ref)). In these models, the function of the data only (the reduced form parameter) is point identified, it is related to the structural parameters $\theta$ through a known mapping, and under standard regularity conditions it can be $\sqrt{n}$-consistently estimated. The resulting estimator has an asymptotically Normal distribution. The various approaches place a prior on the reduced form parameter, and standard tools in Bayesian analysis are used to obtain a posterior. The known mapping from reduced form to structural parameters is then applied to this posterior to obtain a credible set for $\mathcal{H}_\mathsf{P}[\theta]$.
Although partial identification often results from reducing the number of assumptions maintained in counterpart point identified models, care still needs to be taken in assessing the possible consequences of misspecification. This section's goal is to discuss the existing literature on the topic, and to provide some additional observations. To keep the notation light, I refer to the functional of interest as $\theta$ throughout, without explicitly distinguishing whether it belongs to an infinite dimensional parameter space (as in the nonparametric analysis in Section (ref)), or to a finite dimensional one (as in the semiparametric analysis in Section (ref)).
The original nonparametric “worst-case" bounds proposed by man89 for the analysis of selectively observed data and discussed in Section (ref) are not subject to the risk of misspecification, because they are based on the empirical evidence alone. However, often researchers are willing and eager to maintain additional assumptions that can help shrink the bounds, so that one can learn more from the available data. Indeed, early on man90 proposed the use of exclusion restrictions in the form of mean independence assumptions. Section (ref) discusses related ideas within the context of nonparametric bounds on treatment effects, and man03 provides a thorough treatment of other types of exclusion restriction. The literature reviewed throughout this chapter provides many more examples of assumptions that have proven useful for empirical research.
Broadly speaking, assumptions can be classified in two types man03. The first type is non-refutable: it may reduce the size of $\mathcal{H}_\mathsf{P}[\theta]$, but cannot lead to it being empty. An example in the context of selectively observed data is that of exogenous selection, or data missing at random conditional on covariates and instruments (see Section (ref), p. (ref)): under this assumption $\mathcal{H}_\mathsf{P}[\theta]$ is a singleton, but the assumption cannot be refuted because it poses a distributional (independence) assumption on unobservables.
The second type is refutable: it may reduce the size of $\mathcal{H}_\mathsf{P}[\theta]$, and it may result in $\mathcal{H}_\mathsf{P}[\theta]=\emptyset$ if it does not hold in the DGP. An example in the context of treatment effects is the assumption of mean independence between response function at treatment $t$ and instrumental variable ${\boldsymbol{z}}$, see (ref) in Section (ref). There the sharp bounds on $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}(t)|{\boldsymbol{x}}=x)$ are intersection bounds as in (ref). If the instrument is invalid, the bounds can be empty.
pon:tam11 consider the impact of misspecification on semiparametric partially identified models. One of their examples concerns a linear regression model of the form $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}})=\theta^\top{\boldsymbol{x}}$ when only interval data is available for ${\boldsymbol{y}}$ (as in Section (ref)). In this context, $\mathcal{H}_\mathsf{P}[\theta]=\{\vartheta\in\Theta:\mathbb{E}_\mathsf{P}({\boldsymbol{y}}_{\mathrm{L}}|{\boldsymbol{x}})\le \vartheta^\top{\boldsymbol{x}} \le\mathbb{E}_\mathsf{P}({\boldsymbol{y}}_{\mathrm{U}}|{\boldsymbol{x}}),~{\boldsymbol{x}}\text{-a.s.}\}$. The concern is that the conditional expectation might not be linear. pon:tam11 make two important observations. First, they argue that the set $\mathcal{H}_\mathsf{P}[\theta]$ is of difficult interpretation when the model is misspecified. When ${\boldsymbol{y}}$ is perfectly observed, if the conditional expectation is not linear, the output of ordinary least squares can be readily interpreted as the best linear approximation to $\mathbb{E}_\mathsf{Q}({\boldsymbol{y}}|{\boldsymbol{x}})$. This is not the case for $\mathcal{H}_\mathsf{P}[\theta]$ when only the interval data $[{\boldsymbol{y}}_{\mathrm{L}},{\boldsymbol{y}}_{\mathrm{U}}]$ is observed. They therefore propose to work with the set of best linear predictors for ${\boldsymbol{y}}|{\boldsymbol{x}}$ even in the partially identified case (rather than fully exploit the linearity assumption). The resulting set is the one derived by ber:mol08 and reported in Theorem SIR-(ref). pon:tam11 work with projections of this set, which coincide with the bounds in sto07.
pon:tam11 also point out that depending on the DGP, misspecification can cause $\mathcal{H}_\mathsf{P}[\theta]$ to be spuriously tight. This can happen, for example, if $\mathbb{E}_\mathsf{P}({\boldsymbol{y}}_{\mathrm{L}}|{\boldsymbol{x}})$ and $\mathbb{E}_\mathsf{P}({\boldsymbol{y}}_{\mathrm{U}}|{\boldsymbol{x}})$ are sufficiently nonlinear, even if they are relatively far from each other pon:tam11. Hence, caution should be taken when interpreting very tight partial identification results as indicative of a highly informative model and empirical evidence, as the possibility of model misspecification has to be taken into account. These observations naturally lead to the questions of how to test for model misspecification in the presence of partial identification, and of what are the consequences of misspecification for the confidence sets discussed in Section (ref).
With partial identification, a null hypothesis of correct model specification (and its alternative) can be expressed as
Tests for this hypothesis have been proposed both for the case of nonparametric as well as semiparametric partially identified models. I refer to san12 for specification tests in a partially identified nonparametric instrumental variable model; to kit:sto18 for a nonparametric test in random utility models that checks whether a repeated cross section of demand data might have been generated by a population of rational consumers (thereby testing for the Axiom of Revealed Stochastic Preference); and to gug:hah:kim08 and bon:mag:mau12 for specification tests in linear moment (in)equality models.
For the general class of moment inequality models discussed in Section (ref), rom:sha08, and:gug09b, gal:hen09, and and:soa10 propose a specification test that rejects the model if $CS_n$ in (ref) is empty, where $CS_n$ is defined with $c_{1-\alpha}(\vartheta)$ determined so as to satisfy (ref) and approximated according to the methods proposed in the respective papers. The resulting test, commonly referred to as by-product test because obtained as a by-product to the construction of a confidence set, takes the form
Denoting by $\mathcal{P}_0$ the collection of $\mathsf{P}\in\mathcal{P}$ such that $\mathcal{H}_\mathsf{P}[\theta]\neq\emptyset$, one has that the by-product test achieves uniform size control bug:can:shi15:
An important feature of the by-product test is that the critical value $c_{1-\alpha}(\vartheta)$ is not obtained to test for model misspecification, but it is obtained to insure the coverage requirement in (ref); hence, it is obtained by working with the asymptotic distribution of $nq_n(\vartheta)$. bug:can:shi15 propose more powerful model specification tests, using a critical value $c_{1-\alpha}$ that they obtain to ensure that (ref), rather than (ref), holds. In particular, they show that their tests dominate the by-product test in terms of power in any finite sample and in the asymptotic limit. Their critical value is obtained by working with the asymptotic distribution of $\inf_{\vartheta\in\Theta}nq_n(\vartheta)$. As such, their proposal resembles the classic approach to model specification testing ($J$-test) in point identified generalized method of moments models.
While it is possible to test for misspecification also in partially identified models, a word of caution is due on what might be the effects of misspecification on confidence sets constructed as in (ref) with $c_{1-\alpha}$ determined to insure (ref), as it is often done in empirical work. bug:can:gug12 show that in the presence of local misspecification, confidence sets $CS_n$ designed to satisfy (ref) fail to do so. In practice, the concern is that when the model is misspecified $CS_n$ might be spuriously small. Indeed, we have seen that it can be empty if the misspecification is sufficiently severe. If it is less severe but still present, it may lead to inference that is erroneously interpreted as precise.
It is natur