Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
115,150 characters · 28 sections · 63 citation commands
Identification of Incomplete Preferences
\thispagestyle{empty}
{.26in}
\setcounter{page}{1}
Since McFadden's 1974 paper, discrete choice models have been a cornerstone of applied economics analysis. These models, like much of economics, assume that individuals' preferences are complete: individuals can always rank all alternatives. However, there are important reasons why individuals' preference may not be complete; for example, one may have to choose without having enough information about each of the alternatives, or one may have to choose without having enough information about one's own preferences. Assuming this issue away in identification problems is convenient because completeness induces a definite choice, and therefore makes identification of rational behavior easier.\footnote{Indifference implies alternatives are ranked as equal, and is typically deemed a zero probability occurrence.} In this paper we tackle identification when individuals' preferences are allowed to be incomplete even if only aggregate behavior is observable (or, equivalently, individuals make only one observable choice). This is a classic discrete choice setting modified so that completeness does not necessarily hold. The objective is to identify properties of the probability distribution of preferences across heterogeneous individuals using either aggregated market level or individual choice data.
Our main results contradict two hypotheses one could make about the relationship between data and theory when preferences are not complete. The first is that without completeness anything can happen; data cannot tell us anything about the underlying preferences unless one makes assumptions about how choices might result when alternatives are not comparable. The second is that completeness cannot be ruled out by observable behavior. Pairing insights from econometric theory and decision theory, we show that parameters of the probability distribution of preferences across individuals are partially identified, and thus data restricts the possible underlying preferences even if nothing is assumed about how choices among non comparable alternatives are made. We then illustrate how data could show that some of the individuals in the population must have incomplete preferences, thus dispelling the second hypothesis. We also show that if preferences are not complete, a policy aimed at improving the frequency with which a particular alternative is chosen may actually have the opposite effect –a result that can only happen if preferences are not complete.
When preferences are not complete an individual may not be able to rank some pairs of alternatives. Therefore, observers cannot infer that a chosen alternative must have been at least as good as those not chosen. They can only infer that the alternatives not chosen could not have been strictly preferred to the chosen one. When using data to learn about the distribution of preferences in the population, allowing for incompleteness poses what may seem like an insurmountable problem. Because the theory says nothing about how choices between incomparable alternatives are made, data cannot help distinguish choices made because of a definite preference from choices made randomly when alternatives are not comparable. In standard discrete choice models this source of randomness is ruled out by assuming everyone's preferences are complete. One may suspect that when completeness is not imposed data cannot say anything about preferences. To the contrary, we show that also in this case data can be used to learn about the distribution of preference parameters. Obviously, identification has weaker properties than in the case of complete preferences.
We consider a setting in which the analyst has data on the choices of a population of heterogeneous rational agents who must choose one alternative from a given feasible set. With complete preferences, each alternative is associated with a utility value that is known to the decision maker but not known to the analyst. The analyst, on the other hand, knows from data the frequency with which each alternative is chosen. These observed choices are used to make inferences about the probability with which one option is preferred to the others in the population. Since completeness implies each alternative is associated with a utility value, rationality implies that the proportion of individuals who chose one alternative must be equal to the proportion of individuals who assign higher utility to that alternative.
We model incompleteness by allowing for the possibility that each alternative is associated with many utility values. An example of this framework is given by the interval order of Fishburn1970book, where each alternative is associated with an interval of utilities; other possible examples, in stochastic environments, are the multi-utility models of Aumann62, Ok02, and DubraMaccheroniOk04, or the multi-probability model of Bewley02; more recent examples combine Fishburn's and Bewley's models like Echenique-Pomatto-Vinson and Miyashita-Nakamura. We focus on the simplest of these settings, where alternatives are compared looking at the extremes of their utility intervals: their upper and lower utility. An alternative is better if its lower utility is larger than the upper utility of the other, worse if its upper utility is smaller the the lower utility of the other, and not comparable when neither of these conditions is satisfied.
Rationality implies that when two utility intervals are disjoint, the corresponding alternatives can be ranked and choice follows this ranking; when two intervals overlap, the corresponding alternatives are not comparable and choice is indeterminate. With more than two alternatives, rationality implies an alternative can be chosen as long as its upper utility is larger than the lower utility of all other alternatives. In other words, rationality means that decision makers choose an alternative that is not strictly dominated by any other in the feasible set. The choice set is thus the set of non-dominated alternatives; this is a singleton when preferences are complete, while it can contain many elements when they are not. Without completeness, the link between observed choice data and preferences is therefore weakened.
We assume there is a population of decision makers with potentially incomplete preferences from which a random sample is drawn. Although we relax completeness, we make no behavioral assumptions other than rationality. The choice set, the set of non-dominated alternatives, is a subset of the feasible set that is not necessarily a singleton. Because the sample is random and choice between non comparable alternatives is also random, the choice set is a random set and rationality implies that the observed choice is a random selection from this random set. Using data on choice probabilities, the parameters one would like to identify are the probabilities that the set of non-dominated alternatives is equal to each of the possible subsets of the feasible set. These parameters describe the distribution of (incomplete) preferences in the population.
Our first main result shows that although point identification is not possible in this setting, partial identification is. In particular, we characterize the identification region for the parameters of interest using a finite number of inequalities relating these parameters to the observed choice probabilities. When the feasible set contains only two options, the identification region for the probability an alternative is the preferred one is easy to describe with an interval. At one extreme, all individuals who choose an option do so because it is not comparable to the other, and thus the lower bound of the interval is zero. At the other extreme, rationality implies that the upper bound of the probability an alternative is the preferred one cannot exceed the fraction of consumers who chooses it. Since all individuals who prefer an alternative choose it (and some individuals might choose it even though they cannot compare it the other), the probability that a randomly selected individual prefers this alternative cannot exceed the observed fraction of individuals who chooses it. When the feasible set contains more than two options, the sharp identification region is harder to describe because it depends on more than these two inequalities as it can be a strict subset of the interval we just described.
Without additional assumptions one cannot rule out the possibility that non-comparability never occurs because all individuals who choose an alternative do so because it is ranked better than all others. This implies incompleteness is either absent or it is present but irrelevant to behavior. Our second main result uses the notion of an instrumental variable to rule out this possibility. We define an instrument as a random variable that is independent of preferences but correlated with choice; in particular, the fraction of individuals who choose an alternative changes for each realization of the instrument while preference characteristics do not. If such an instrumental variable exists there must be some individuals who have incomplete preferences. Intuitively, if the fraction of individuals who choose a certain alternative changes with the realization of the instrument, it cannot be that all individuals are always able to rank all alternatives. Existence of an instrument thus establishes that there must be some incomparable alternatives, making the identification region smaller.
We analyze the impact of a policy intervention aimed at increasing the chances a particular alternative is chosen by increasing the utility individuals give to that alternative. Examples of such a policy could be campaign advertising, or other forms of information provision, that target a particular alternative. When preferences are complete, one can easily show that increasing the utility of an alternative necessarily increases the frequency with which that alternative is chosen. Without completeness, this is no longer the case. Intuitively, unless the intervention is particularly strong, because of incompleteness the behavior of some individuals cannot be predicted even after the policy is enacted. There could be enough individuals that change behavior in favor of the non-targeted alternative to make the policy completely ineffective. Therefore, if one observes aggregate behavior go in the opposite direction than the one the intervention tried to achieve one cannot necessarily conclude the policy was ill-designed.
We study several extensions of the basic model that can also make the identification region smaller. First, we examine a model in which a known fraction of individuals pays attention only to a subset of the feasible set. Second, we show how minimizing ex-post regret could bring point identification. Third, in an extension that is particularly relevant for our application, we consider the situation in which the choices of a known fraction of the population are not observed.\footnote{In our application, voting, a known fraction of individuals go to the polls but abstain in some of the races on the ballot.} In this case, in addition to partial identification resulting from incomplete preferences, researchers can face non response or unobserved choices. We show that the identification region can then be written as a convex combination of an identification region resulting from incomplete preferences and an identification region resulting from non response. Finally, while our results are cast in a world without uncertainty, we show that one can allow for ambiguity by extending our framework to study identification of the preferences described in Bewley86.\footnote{Although we do not pursue it, a similar exercise could be performed in the multiple utility framework of Aumann62.} This extension shows that our results are not limited to interval orders and the notions of upper and lower utilities, but apply to any situation in which rational behavior is driven by two distinct thresholds.
We apply our methods to precinct-level data from Lorain county in Ohio, focusing on the two races for the Ohio Supreme Court that occurred in the 2018 midterm elections, and demonstrate how several behavioral assumptions can be combined to reduce the size of the identification set. This empirical application fits our setting in several ways. First, individual choices are not observable. Second, the phenomenon of roll-off voting can be thought of as a channel through which incompleteness can directly manifest itself: more than 20% of the turnout voters did not vote in these races even though they went to the polls.\footnote{A possible reason is that some voters cannot use party affiliation to rank candidates since it is not on the ballot even though candidates were selected by each party in their primary elections.} Third, voter registration data help us illustrate the role of limited consideration sets by assuming that partisan voters only consider candidates of their own party. Finally, Ohio election rules provide an example of an instrumental variable because the order in which the candidates are presented on the ballot changes from precinct to precinct. Order on the ballot is unlikely to be correlated with voters' preferences, while it can have an effect on vote shares, particularly when each candidate's party affiliation is not on the ballot as is the case for Judicial elections in Ohio. We show that the order has a noticeable, yet small, effect on the identification region in our races.
Literature Review
Our paper brings together two strands of literature: the one focused on theories of individual decision making that yield indeterminacy in behavior, and the one focused on the econometric identification of economic models that have non-unique predictions. The former group dates back to Luce56, who speaks about intransitive indifference, and has focused mostly on identifying individual preference characteristics from individuals choices (see Dziewulski-2021 for a recent example). To the best of our knowledge, we are the first to tackle econometric identification of incomplete preferences from choice data, including aggregate data, and we are the first to connect preference incompleteness in a deterministic setting with partial identification.
Our approach to the econometric identification of choice probabilities when preferences are incomplete follows the vast literature on partial identification. Manski2003 coins the term The Law of Decreasing Credibility saying that “The credibility of inference decreases with the strength of the assumptions maintained”. Seeking more credible results, a researcher can either weaken assumptions on the behavior of decision makers in the model or weaken assumptions on the data generating process. Our main focus is the first of these two - the behavioral aspect, and the methods developed in this paper apply to any situation where decision makers cannot fully rank the alternatives they face.
Ambiguity (or Knightian uncertainty) has generated much interest in the econometric literature on partial identification in the last two decades starting from Manski2000. Ambiguity requires that (a) the analyst and the decision makers agree on the set of possible states of nature and (b) that the analyst knows the set of probabilities used by each decision maker (behavioral ambiguity) or knows the set that contains each decision maker's expectations (observational ambiguity). ManskiMolinari2008 show that questioner design can lead to imprecise knowledge of agents preferences. GiustinelliManskiMolinari2021 elicit expectations from decision makers who may have imprecise subjective probabilities regarding late-onset dementia, and find that about half of the individual hold imprecise probabilities regarding their future health. Manski2018 looks at random utility models with linear utilities when some alternatives' characteristics are state dependent and the distribution function of the state is known to belong to a subset of the simplex. Section 2 of Manski2018 discusses observational ambiguity where agents have a unique subjective expectation for future states but the analyst only knows that this distribution belongs to a certain set of distributions. Section 3 of Manski2018 discusses behavioral ambiguity when decision makers do not have a unique subjective distribution on the states of nature. In both cases Manski assumes that a decision maker does not chose an alternative which is dominated by another. He shows that in a binary choice model when choices are observed and the degree of ambiguity is known, this non-dominance condition yields inequalities on the parameters of the model. The results described in our paper generalize his findings to non-binary choice models with general behavioral incomplete preferences. We also combine the behavioral source of partial identification with observational source of partial identification.
Limited consideration sets are related to theories introduced in EliazSpiegler2011, Masatlioglu2012, and Lleras2017. ManziniMariotti2014) introduce assumptions on the way individuals restrict their attention to a subset of the alternatives such that the parameters of the model are point identified. When consideration sets are unobserved, the model parameters are partially identified. BarseghyanCoughlinMolinariTeitelbaum2021 consider decision makers with complete preferences but with unobserved consideration sets. Their model leads to partial identification of model parameters such as risk aversion. BarseghyanMolinariThirkettle2021 consider decision makers with unobserved heterogeneous consideration sets and use exclusion restriction on alternatives characteristics to point identify the parameters of the model. To a large extent the models of consideration sets constrain the behavior of the individuals rather than relax the assumptions of the model. In the application we consider, voting, the set of feasible alternatives is revealed on the ballot and is similar to all decision makers for statewide races. We analyze the impact of an assumption that certain voters consider only a subset of the candidates that are affiliated with a certain party.
Finally, this paper is also related to behavioral models where rationality is weakened. Some papers focus on finding tests for rationality in the general sense as in KitamuraStoye2018 and by Hoderlein2011. These papers look for violations of utility maximizing behavior in data on individual choices. Violations, if found, are not attributed to failure of any specific assumption but rather to a general lack of rationality by some individuals. In our case, decision makers are perfectly rational even though their preferences are not complete.
Structure
The remainder of the paper is organized as follows. Section (ref) introduces incomplete preferences and describes rational behavior. Section (ref) discusses random decision makers and non-parametric identification of preferences distribution. Section (ref) focuses on binary choice. Section (ref) illustrates several extension of our basic framework. In section (ref) we apply our findings to voting data, and show how this data can be used to illustrate our findings. Section (ref) concludes.
Our objective is to analyze the standard discrete choice model without imposing the assumption that preferences are complete. In the standard model, completeness is reflected by the idea that choice between alternatives is driven by the utility assigned to each. This follows from the well known result that, with a finite number of alternatives, any complete and transitive preference relation has a utility function representing it. The utility function assigns a number to each alternative (its utility), and the ordering of these numbers is used to rank alternatives and thus to model choices. Without completeness this is no longer the case: one cannot necessarily assign a utility value to each alternative.
We focus on a model in which each alternative is associated with two numbers, and these two numbers are used to determine the ranking between alternatives, if such a ranking exists. Let $X$ be the set of alternatives over which an individual's preference order $\succ$ is defined. For each $x \in X$ we assume there exist two real numbers $\bar{u}(x)$ and $\underbar{u}(x)$, with $\bar{u}(x) \geq \underbar{u}(x)$. We refer to these two numbers as an alternative's upper and lower utility respectively, and we use the word vagueness to describe the length of the utility interval $[\underbar{u}(x),\bar{u}(x) ]$ associated to each $x$. Upper and lower utility can be used to describe the individual's preference ordering between any two alternatives, and this preference can then be used to describe her behavior when faced with a particular subset of feasible alternatives. First, for any $x,y \in X$ we say that $x$ is (strictly) preferred to $y$ if and only if $\bar{u}(y) < \underbar{u}(x)$. In words, one alternative is preferred if its lower utility exceeds the upper utility of the other. Two alternatives are not comparable when neither this inequality nor its opposite are satisfied (their utility intervals overlap). We say that two alternatives are indifferent if their upper and lower utilities are the same (their utility intervals are the same).
Next, we describe how choices are made from a given set of feasible alternatives $ \mathcal{A} \subseteq X $ given the preferences described above. The main idea is that choice must allow for the possibility that some alternatives are not ranked. In particular, an alternative can be chosen from a set provided there is no other option in that set that is strictly preferred to it.\footnote{This weakens the standard revealed preferences conclusions as noted in Eliaz-Ok-2006.} Suppose only two alternatives are feasible, then one of them can be chosen if it is either preferred to the other, or if it is not comparable to it. In other words, an alternative can be chosen as long as it is not dominated. In general, when there are more than two feasible possibilities in $\mathcal{A}$, an alternative can be chosen provided there is no other option in $\mathcal{A}$ that is strictly preferred to it. We formalize these ideas using the following definition.
$M(\mathcal{A})$ is the set of all alternatives in $\mathcal{A}$ that are not (strictly) dominated by any element of $ \mathcal{A} $. When preferences are complete, the set $M(\mathcal{A})$ includes only alternatives that are indifferent to each other. Without completeness, the set $M(\mathcal{A})$ may include several incomparable alternatives.
Let $y \in \mathcal{A}$ denote an observable choice from the feasible set $\mathcal{A}$; when this set is not a singleton the decision maker chooses an alternative from it randomly. Rationality means that dominated alternatives cannot be chosen even if behavior can be random because some alternatives are not comparable to each other.
In the case of incomplete preferences, rationality does not make unique predictions about behavior. When preferences are complete, typical assumptions imply that indifference is a zero probability event, and thus the analyst can treat $M(\mathcal{A})$ as a singleton.\footnote{This is usually done by assuming that utility functions are continuous.} Without completeness, however, assuming that indifference is a zero probability event does not necessarily imply that $M(\mathcal{A})$ is a singleton because incompleteness is not as knife-edge as indifference. One can rule out indifference between $x$ and $y$ by asking that the intervals $[\underbar{u}(x),\bar{u}(x)]$ and $[\underbar{u}(y),\bar{u}(y)$] are different, but this is not enough to rule out any overlap between them. Rationality also reflects the idea that no selection rule is a-priory imposed among incomparable alternative because it only says that an alternative that is strictly dominated is never chosen.
Although we cast the paper in the language of upper and lower utilities for ease of exposition, our identification results apply to any incomplete preference relation that yields a set of undominated alternatives constructed using only two numbers as stated in Definition (ref). There are many examples of such preferences, and we end this section by describing one of them as an illustration of how upper and lower utilities can obtain; in Section (ref) we show that our approach is also consistent with the Knightian decision theory of Bewley86 where preferences are not complete because of ambiguity.\footnote{Bewley86 was later published as Bewley02. Other examples could be Aumann62 and DubraMaccheroniOk04, or the more recent Echenique-Pomatto-Vinson and Miyashita-Nakamura.}
Interval orders are presented in Fishburn1970book and are well suited to illustrate our setup. They can be obtained under simple assumptions on preferences. Let $(X,\succ)$ be a partially ordered set such that $X$\ is a finite set and $\succ$ is a strict preference relation over $X$. The first property is irreflexivity: $\lnot ( x\prec x)$ (where $\lnot $ denotes logical negation). The second property is a special form of transitivity: $x\prec y$ and $z\prec w$ $\Rightarrow $ $x\prec w$ or $z\prec y$.\footnote{One can easily verify that these two properties imply the usual transitivity.} Fishburn calls a preference that satisfies these two properties an interval order, and shows that $\succ$ can be represented using two functions as the following theorem illustrates.
Thus, $y$ is preferred to $x$ if and only if the utility of $y$ exceeds the utility of $x$ by some strictly positive amount. One can think of each alternative as being associated with an interval on the real line. The lower bound of that interval represents its utility, while the width of that interval represents imprecision in that utility. In this spirit, Fishburn calls $\sigma $ the vagueness function since it measures the amount of imprecision in the utility associated with each alternative. The name interval order follows from the observation that alternatives are compared using intervals: when $y$ is preferred to $x$ the interval $\left[\underbar{u}(y),\underbar{u}(y)+\sigma(y)\right]$ lies to the right of the interval $\left[\underbar{u}(x),\underbar{u}(x)+\sigma(x)\right]$. Interval orders are clearly not necessarily complete. When the two intervals overlap, $x$ and $y$ are not comparable. In this setting, although strict preference is transitive, non-comparability is not. In other words, $x$ may be not comparable to $y$, and $y$ maybe not comparable to $z$, but $z$ is strictly preferred to $x$. This idea is sometimes referred to as intransitive indifference.\footnote{Quoting from Fishburn: \textquotedblleft For example, if you prefer your coffee black it seems fair to assume that your preference will not decrease as $x$, the number of grains of sugar in your coffee, increases. You might well be indifferent between $x=0$ and $x=1$, between $x=1$ and $x=2$, ... , but of course will prefer $x=0$ to $x=1000$.\textquotedblright}
Interval orders easily map to our framework by letting $\bar{u}(x)= \underbar{u}(x) +\sigma(x)$ so that each alternative is associated with a pair of numbers measuring the lower and upper bound of the alternative's `utility interval'. Using these two values, one can then talk about behavior when faced with a particular set of possibilities.\footnote{Using Fishburn's example, only cups of coffee which contains low amounts of sugar can be chosen. As soon as sugar content is high enough to make a cup strictly worse than a cup with no sugar at all, this cup is a dominated alternative. It, as well as any cup that contains more sugar, will not be chosen.} Inspired by Fishburn, we use the term vagueness to describe the difference between upper and lower utility of an alternative.
Notation. Throughout the paper we use capital Latin letters to denote sets and random sets. We use lower case Latin letters for random vectors and random functions. We use lower case Greek letters for parameter vectors and capital Greek letters for parameter sets. For a set $A \subset \Re^k$, $A^c$ denotes its complement. For a finite set $A$, $|A|$ denotes its cardinality. Scripted Latin letters are used for spaces or collections of similar objects.
We use $i \in I$ to denote a random individual from the population $I$. For each $i\in I$ the set of feasible choices is $\mathcal{A}$, and for each $a\in \mathcal{A}$ decision maker $i$ has a utility interval $\left[ \underbar{u}_i(a) ,\bar{u}_i(a) \right]$ as described in Section (ref).\footnote{We assume the set of alternatives is the same for all decision makers for simplicity. We explore the role of limited consideration sets in Section (ref) } As usual, utility values are known to the individual but not to the analyst.
We treat $\underbar{u}:\mathcal{A} \rightarrow \mathbb{R}$ and $\bar{u}:\mathcal{A} \rightarrow \mathbb{R}$ as two random functions such that for every $a\in \mathcal{A}$, $\mathbf{P}( \underbar{u}(a)\leq \bar{u}(a)) =1$. As in standard discrete choice models, we rule out indifference by assuming that $\underbar{u}(a)$ and $\bar{u}(a)$ are continuous random variables for every $a \in \mathcal{A}$. Our objective is to learn about choice probabilities using data on choices. For example, similarly to the case of complete preferences, we would like to estimate the probability that a random individual prefers alternative $a$ to all other alternatives, $\mathbf{P} \left( \underbar{u}(a) \geq \bar{u}(b) \, \forall b \neq a \right)$. Denoting this probability as $\theta_a$, complete preferences and zero probability of ties imply that $\sum_{a \in \mathcal{A}}\theta_a=1$. Moreover, rationality (choosing the alternative that yield the highest utility) implies that the preference probabilities, $\theta_a$, and the observed choice probabilities, $P_a$, are synonymous. When preferences are not complete, however, the probability that alternatives $a$ and $b$ are incomparable but both are preferred to alternative $c$ can be positive. In what follows we formalize the difference between preference probabilities and choice probabilities in the context of incomplete preferences.
For an individual $i \in I$, let $M_i=M_i(\mathcal{A})$ be the set of alternatives that are not dominated,
We can describe $M$ as a mapping $M:I \rightarrow \mathcal{K}(\mathcal{A})$, where $\mathcal{K}(\mathcal{A}) $ is the set of all non-empty subsets of $\mathcal{A}$. Since $\mathcal{A}$ is finite, $M_i$ is non empty (a maximum exists) and $\mathcal{K}(\mathcal{A})$ contains compact sets. Since $\underbar{u}$ and $\bar{u}$ are random variables, for all $A \in \mathcal{K}(\mathcal{A}) $, $\left\{ i: M_i \cap A \neq \emptyset \right\} \in \mathcal{F}$. Therefore, $M$ is a random set (see Appendix (ref) for a summary of the definitions and tools of random set theory used in the body of the paper).
For every (non-empty) $A \in \mathcal{K}(\mathcal{A})$ define the probability that a decision maker's set of non-dominated alternatives equals $A$ as $$\theta_A = \mathbf{P}(M = A).$$ This definition extends the standard concept of choice probabilities in models with complete preferences. Complete preferences and no ties mean that $\theta_A>0$ if and only if $|A|=1$ and $\theta_A=0$ otherwise. When preferences are not complete, however, there is a set $A$ of cardinality bigger than $1$ such that $\theta_A>0$. The collection $\mathbf{\theta} = \{ \theta_A \}_{A \in \mathcal{K}(\mathcal{A}) }$ describes all choice relevant parameters of the joint distribution of $\mathcal{U}=\{\underbar{u}(a),\bar{u}(a) : a \in \mathcal{A} \}$.
The vector of preference parameters $\theta$ satisfies the following properties:
Therefore, after ordering these parameters in some way, the vector of choice parameters is an element of the simplex $\Theta =\Delta(\mathcal{K}(\mathcal{A}))$.\footnote{In section (ref) we discuss abstaining where decision makers do not have to choose any alternative in $\mathcal{A}$. Here we assume that the choice probabilities sum to $1$ and therefore the choice parameters sum to $1$ as well.}
Before observing any data, we can only say that the vector of preference parameters lies in the simplex and the sum of any subset of choice parameters lies between $0$ and $1$ (See Figure (ref) in the next Section for an example).
Rationality (see Definition (ref)) implies that individual $i$'s choice, denoted $y_i$, is an element of the random set $M_i$. Without completeness, rationality allows for $y_i$ to be chosen from $M_i$ randomly. In the context of random set theory, this behavioral assumption translates to the following measurability assumption.
A random sample of decision makers from the population amounts to drawing a random sample of their sets of non-dominated alternatives. When preferences are not complete there is an additional layer of randomness: a decision maker chooses an alternative from the set $M_i$ in a way that can be random.\footnote{Although we do not pursue that route formally, the econometric model could be augmented with a selection mechanism that completes the model. See Definition 2.4 in BMM2011.}
A well known result from random set theory, Artstein's Lemma, connects the containment functional of the random set $M$ to the distribution function of a selection from that set; this connection is described by the Artstein's Inequalities. Theorem (ref) shows there are restrictions that these inequalities impose on the preferences parameter $\theta$ even if preferences are not complete.
Artstein's inequalities imply that the set $\Theta^I$ is the sharp identification region for the parameter $\theta$. These inequalities are both sufficient and necessary for a parameter to be included in the identification set and hence $\Theta^I$ is the sharp identification set (see BMM2011 and BMM2012). Therefore, even without completeness, the data imposes restriction on parameters of the joint distribution of preferences. The connection between the choice parameters and the order of the upper and lower utilities can be understood from the following relationship:
Therefore, restrictions imposed on the choice parameters can be translated into restrictions on the relative order of the upper and lower utilities.
If there exists $a \in A$ such that $0<\mathbf{P} (y=a) <1$, then $\Theta^{I}$ is a strict subset of $\Theta$ (this is trivially true when there are at least two alternatives). Theorem (ref) shows that $\theta_a \leq \mathbf{P} (y=a)$ which is strictly less than $1$.\footnote{To simplify notation, for $A=\{a\}$, a singleton, we denote $\theta_a=\theta_{\{a\}}$.} In defining the set $\Theta$ this restriction on $\theta_a$ is not imposed and therefore $\Theta^I \subsetneq \Theta$.
Next, we show that all the inequalities in the definition of the identification region are potentially binding. Let,
be the set of choice parameters that satisfy Artstein's inequalities for subsets $A$ with $|A|=1$. The following result gives conditions for this set to be larger than the identification set defined in Theorem (ref).
Proposition (ref) illustrates that inequalities involving sets with cardinality greater than 1 are potentially binding (in addition to inequalities related to subsets with cardinality 1). In terms of the capacity functional, the condition in Theorem (ref) can be replaced with $\exists A \subset \mathcal{A}$ such that $|A|>2$ and $C_M(A)>\sum_{a \in A} \theta_a$.
So far, we have obtained a general characterization of the identification region, and explored some of its properties. Next, we limit attention to a binary choice; this enables us to illustrate what we have learned so far, as well as present our next main results, using simple pictures and focusing on intuition.
In this section we focus on binary choice models; we assume that the set of alternatives is $\mathcal{A}_i=\left\{ a_0,a_1\right\} $ for all $i$. As an illustration of Theorem (ref), we first derive the identification region for parameters of interest similar to those identified in models with complete preferences, and show that these parameters are only partially identified. Since there are only two alternatives, we can illustrate our results with diagrams. Then, we present the second main result of the paper: an instrumental variable can shrink the identification region and potentially rule out the hypothesis that all individuals have complete preferences. Finally, we study how the identification region is affected by abstention and limited consideration sets.
We next derive the identification region implied by Theorem (ref), and show how it differs from the point identified case of complete preferences. We start with the case in which data identifies only the fraction of individuals who chose each alternative. We want to find out what these fractions can tell us about the distribution of utility values in the population. Formally, let $u_{i}(a_j) = [\underbar{u}_i(a_j) , \bar{u}_i(a_j) ]$ for $j=0,1$; this is the utility interval agent $i$ assigns to alternative $a_j$. Let $\mathcal{U}=(\underbar{u}(a_0),\bar{u}(a_0),\underbar{u}(a_1),\bar{u}(a_1))$ be the corresponding random vector of utilities. We seek to identify, or partially identify, features (parameters) of the joint distribution of $\mathcal{U}$. For binary choice in particular, one is interested in the relative position of the utility intervals $[\underbar{u}(a_0),\bar{u}(a_0)]$ and $[\underbar{u}(a_1),\bar{u}(a_1)]$.
When preferences are complete, each interval is a singleton, and what matters for choice is the order of the utilities resulting from alternatives $a_0$ and $a_1$. Without completeness, an individual chooses from the random set of non-dominated alternatives, $M$, that is defined as
Rationality means that $a_1$ is chosen with certainty if $\underbar{u}(a_1)>\bar{u}(a_0)$, while $a_0$ is chosen with certainty if $\underbar{u}(a_0)>\bar{u}(a_1)$.\footnote{Since we assume $\underbar{u}(a_i)$ and $\bar{u}(a_i)$ are continuous random variables equality can be ignored.} Figure (ref) illustrates the decision rule in terms of the random set $M$ and thus displays the decision maker's behavior. The axes are given by the difference between the lower utility of one alternative and the upper utility of the other, and the regions illustrating $M$ lie below the $-45^{\circ} $ line.\footnote{By definition, $\bar{u}(a_j)-\underbar{u}(a_j)\geq 0$ for $j=0,1$; adding over $j$ and rearranging one gets $[\underbar{u}(a_1)-\bar{u}(a_0)] \leq -[\underbar{u}(a_0)-\bar{u}(a_1)]$. When upper and lower utilities coincide because preferences are complete, then $M$ coincides with the $-45^{\circ} $ line}
Let
be the probabilities that a random decision maker prefers alternative $a_0$ over alternative $a_1$ and the probability that a random decision maker prefers alternative $a_1$ over alternative $a_0$, respectively. If preferences are complete and $\underbar{u}(a_j)=\bar{u}(a_j)$ for $j=0,1$, then $(\theta_0+\theta_1)=1$. Due to incompleteness we can only say that $0 \leq \theta_0+\theta_1 \leq 1$. Without any data, there are no further restrictions. The identification region of $(\theta_0,\theta_1)$ before observing data is depicted in Figure (ref) as the triangle between the axes and the negative $45^{\circ}$ line through the points $(0,1)$ and $(1,0)$.
In our setting, aggregate choices are observed by the analyst and this data helps narrow down possible values of $(\theta_0,\theta_1)$. Let $y_{i}$ denote the choice made by individual $i$, and denote the choice probability of alternative $a_1$ by $p_1 = \mathbf{P} ( y_{i}=a_1)$ and the choice probability of $a_0$ by $p_0=\mathbf{P} ( y_{i}=a_0)$. These are the fractions of individuals that choose $a_1$ and $a_0$ respectively, and are identified from the data generating process.
If preferences are complete, one has a point identified model where $(\theta_0,\theta_1)=(p_0,p_1)$. Without completeness, Theorem (ref) implies that even if point identification is not possible partial identification is. An agent's choice, $y$, is a selection from the random set $M$ because of Assumption (ref). Artstein inequalities imply that $y\in Sel(M) $ if and only if $\mathbf{P}(y \in K) \geq C_{M}(K) $ for every closed set $K$ where $C_{M}(K)$ is the containment functional. Substituting $K=\{a_0\}$ and $K=\{a_1\}$, this implies
and thus the identification region is
Since all individuals who prefer $a_0$ choose it (and some individuals might choose it even though they cannot compare it with $a_1$), the probability that a randomly selected individual prefers $a_0$ cannot exceed the fraction $p_0$ of individuals who chooses it. If, for example, $p_{0}=0.37$ then the identification region cannot admit the case where $\theta_0 = 0.4$ and $\theta_1 =0.6$.
The size of the group of people who choose an outcome even though they cannot compare it to the other could go from zero up to all those who made that choice. For this reason, the discrete choice model that allows for incomplete preferences is partially identified. The dark area in Figure (ref) represents the identification region for $(\theta_0,\theta_1)$ given a pair of $p_0$ and $p_1$ values. The combinations of $( \theta_0, \theta_1 )$ which are included in the two light colored triangles in Figure (ref) are eliminated from the identification region after observing data. Observing choices informs the researchers about the possible values of the probability that $a_0$ is strictly preferred to $a_1$ and the probability that $a_1$ is strictly preferred to $a_0$.
The identification region in equation ((ref)) can be useful in understanding which features of the theoretical model can be deduced from data. For example, denote with $\theta_{01}$ the fraction of decision makers whose preferences are incomplete. Clearly,
where $\theta_0 +\theta_1$ is the proportion of decision makers who can compare the two alternatives. From Figure (ref) one notices that the possibility that $\theta_0 + \theta_1=1$ (and thus $\theta_{01}=0$) is included in the identification region: this is the point where the identification region touches the line connecting $\theta_0=1$ to $\theta_1=1$. Thus, one cannot rule out the possibility that all decision makers have complete preferences. Similarly, one cannot rule out the possibility that all decision makers cannot compare the two alternatives, so that $\theta_0 = \theta_1=0$ and $\theta_{01} = 1$ (this is the origin). One cannot state anything sharper than $0 \leq \theta_{01} \leq 1$ without additional information about preferences.
Suppose one knows that at least a proportion $\nu>0$ of decision makers cannot rank the two alternatives ($\theta_{01} \geq \nu$).\footnote{In Section (ref) data suggests that at least a certain proportion of voters may not have been able to rank the candidates.} Then $\theta_0 + \theta_1 = 1-\theta_{01} \leq 1- \nu$. The corresponding identification region is shown in Figure (ref). The dark region includes all points consistent with two statements: at least $\nu$ proportion of decision makers cannot rank the two alternatives, and proportion $p_{0}$ of them chose alternative $a_0$.
Our results so far show that under reasonably few standard assumptions identification, albeit partial, is possible when preferences are not complete. Lack of completeness does not imply that “anything goes”. Next, we introduce the concept of an instrumental variable, and show how such a variable could shrink the identification region in an interesting way.
In this section we show how instrumental variables can refine the identification region established in Section (ref). In particular, we define an instrument as a random variable that influences choices while having (almost) no effect on preferences. When these instruments exist, they can be used to rule out the possibility that all decision makers have complete preferences. We present this result for the case of binary choice, but it extends to the general case.\footnote{One complication with more than two alternatives is that one may have partial identification even when preferences are complete.} Intuitively, if observed behavior changes with the realization of this random variable, then it must be the case that some individuals' behavior was not dictated by an actual ranking between the alternatives. In our voting application, a possible instrumental variable is represented by the order in which two candidates are presented on the ballot.
Think of a random variable that is independent of utilities but is correlated with choice. Intuitively, whenever a decision maker can rank the alternatives her choice cannot depend on the realized values of this random variable. However, when a decision maker cannot compare the alternatives her choice could depend on the random variable realizations (this dependence need not be deterministic). If observed aggregate choices change with the realized value of the instrument it must be that some decision makers were not able to rank alternatives. We formalize the idea of an instrument as follows.
Independence of $(\underbar{u}(a_j),\bar{u}(a_j)_{j=0,1})$ and $Z$ corresponds to independence of the instrumental variable and the unobservable element of the model - the underlying utilities in our case. The validity of this condition, however, is driven by behavioral assumptions which are case specific and are not testable. $\Delta_{0}>0$ corresponds to the relevance of the instrumental variable - the outcome variable (i.e. the choice) is affected by $Z$. The validity of this condition in Definition (ref) is testable. Given both assumptions in Definition (ref), an instrumental variable can then be used to establish the result that some individuals' preferences must not be complete.
Theorem (ref) says that the binding constraint on the probability that one outcome is preferred to the other is imposed by the lowest conditional probability that outcome is chosen, where conditioning is upon the values of the instrument. Figure (ref) illustrates the identification power of having an instrumental variable. The identification region does not contain any point on the $-45^{\circ}$ line. Therefore, we can rule out the possibility that all individuals had complete preferences. The distance of the identification region from the $-45^{\circ}$ line depends on $\Delta_0$ which measures the extent to which choices are influenced by $Z$.
One can, to a certain extent, relax the assumption that $Z$ is independent of the distribution of the utilities in $\mathcal{U}$ as explained in what follows. NevoRosen introduced the notion of imperfect instrumental variables in a linear regression model with endogenous regressors. In their context, an imperfect instrumental variable is a variable $Z$ correlated with the error term of the regression but to a much lesser degree than it is correlated with the endogenous regressor. NevoRosen show that, under some conditions, imperfect instruments can lead to partial identification of the regression parameters. We adapt the idea of imperfect instrumental variables to our model. Define the following quantities:
and
When Z is independent of preference parameters, as in Definition (ref), $\delta_0 = \delta_1 = 0$ and one has a perfect instrument. If $Z$ is not a perfect instrument, $\delta_0$ and $\delta_1$ measure the sensitivity of the utilities to changes in the value of the instrument $Z$.
The next result shows that if utilities depend on $Z$ to a lesser extent than choices depend on $Z$, the fraction of decision makers whose preferences are not complete is bounded away from $0$ by a strictly positive quantity.
Next, we consider the possibility that choices are unobserved for a subset of decision makers. This situation can occur for several reasons. First, as is the case with voting (see Section (ref)), individuals can refrain from making a decision (abstention). Second, the data collected is incomplete due to non-response or non-observability. Let $p_0$, $p_1$, and $\gamma$ represent the proportion of decision makers who chose alternative $a_0$, alternative $a_1$ and abstained, respectively, such that $p_0+p_1+\gamma=1$. Let $V \in \{0,1\}$ be a binary random variable indicating whether an individual's choice is observed, $V=1$, or unobserved, $V=0$.
To identify the probability that a random individual strictly prefers $a_0$ over $a_1$ given that this decision maker's choice is observed one can use the methods described in Section (ref). For example, one could assume that voters that went to the polls and then abstained must have done so because their preferences were incomplete. However, we are interested in identification when such an assumption is not made. In particular, we want to identify $\theta_0$ and $\theta_1$ in the entire population of decision makers, not only among those who made a choice. Using the total law of probability, we can write
Rationality implies that
These inequalities are described by the rectangular identification region in Figure (ref). We denote this set as $\Theta^{O}=\{(\theta_0,\theta_1): 0 \leq \theta_0 \leq p_0 \, , \, 0 \leq \theta_1 \leq p_1 \}$.
The quantities $\mathbf{P}(\underbar{u}_0>\bar{u}_1|V=0)$ and $\mathbf{P}(\underbar{u}_1>\bar{u}_0|V=0)$ are unidentified and satisfy the following inequality $$0 \leq Pr(\underbar{u}_0>\bar{u}_1|V=0)+Pr(\underbar{u}_1>\bar{u}_0|V=0) \leq 1.$$
Finally, we can combine both identification regions $\Theta^O$ and $\Theta^U$ using equation ((ref)),
where $\oplus$ is the Minkowski sum.
Since choices made by decision makers who abstain are unobserved, several assumptions can be made about $\Theta^U$. One can assume that all unobserved decision makers were not able to compare the two alternatives; in this case $\Theta^U=\{(0,0)\}$ and $\Theta^I = (1-\gamma) \Theta^O$ is the black rectangle in Figure (ref). Alternatively, one can be agnostic about the preferences of the unobserved decision makers and assume that $\Theta^U=\{ (\theta_0,\theta_1): 0 \leq \theta_0+\theta_1 \leq 1 \}$; these inequalities are represented by the triangular identification region in Figure (ref). In this case the joint identification region is the gray polygon in Figure (ref). Another common assumption is missing at random which implies that $\Theta^U=\Theta^O$ and yields $\Theta^I =\Theta^O$ which is the black rectangle in Figure (ref). Other assumptions on the behavior of the unobserved decision makers can be made depending on the application at hand.
One can combine the ideas discussed above with the notion of instrumental variables. Let $Z$ be an instrumental variable with a finite support $\mathcal{Z}$ satisfying Definition (ref). Equation ((ref)) holds for all values of the instrument $Z$. Furthermore, we assume that $\gamma$, the probability of abstaining remains constant across different values of the instrument $Z$. Equation ((ref)) holds for all values of the instruments which affect only the right hand-side. Therefore,
Combining equation ((ref)) and equation ((ref)), we have
Combining this information with the fact that $0 \leq \theta_{0|V=0} +\theta_{1|V=0} \leq 1$, we can write the joint identification region similarly to equation ((ref)),
where $\Theta^{O|Z}=\{ (\theta_0,\theta_1): \theta_0 \leq \inf_{z \in \mathcal{Z}} p_{0|Z=z}. \, \theta_1 \leq \inf_{z \in \mathcal{Z}} p_{1|Z=z} \}$. Theorem (ref) implies that if $\sup_{z\in \mathcal{Z}}p_{0|z} - \inf_{z\in \mathcal{Z}}p_{0|z}>0$, then $\Theta^{O|Z} \subsetneq \Theta^O$ and therefore $\Theta^{I|Z} \subsetneq \Theta^I$.
Partial identification in this section is a combination of both behavioral assumptions (incomplete preferences) and observational challenges (abstention). The identification regions reported in equations ((ref)) and ((ref)) combine both sources of partial identification: $\Theta^U$ is the result of decisions makers whose choice is unobserved (e.g. voters abstaining); $\Theta^O$ and $\Theta^{O|Z}$ are a result of an incomplete model. The resulting identification regions $\Theta^I$ and $\Theta^{I|Z}$ are a convex combination of both sources of partial identification using the weight $\gamma$. Manski2018 discusses behavioral and observational sources of partial identification separately. Overall, combining both behavioral and observational reasons for partial identification have not received as much attention in the literature. Here, we showed how they can be combined and the application in Section (ref) illustrates this combination in practice.
In Section (ref) we show that instrumental variables allow the analyst to provide evidence that at least some decision makers have preferences that are not complete. The origin is included in all identification regions presented so far (see Figure (ref)). Thus, one cannot rule out the possibility that all decision makers are unable to rank the two alternatives. In this section we explore a simple assumption that allows the analyst to exclude this possibility.
Consider the case where a known proportion of the decision makers considers, or pays attention to, only a subset of alternatives $\mathcal{A}' \subsetneq \mathcal{A}$.\footnote{See Masatlioglu2012 and Lleras2017 for formal models of decision making in the presence of limited consideration sets.} This could happen because some agents are unaware an alternative exist, or because some agents would never consider an alternative even when they are aware of it (for example, members of a political party would never vote for candidates who belong to a different party). When this happens, the set of non-dominated alternatives in Definition ((ref)) is applied to the decision maker with consideration set $\mathcal{A}'$ and gives the random set of non-dominated alternatives $M(\mathcal{A}')$.
In a binary choice context $\mathcal{A}=\{a_0,a_1\}$, and there are only three potential consideration sets: $\mathcal{A}=\{a_0,a_1\}$, $\mathcal{A}_0=\{a_0\}$, and $\mathcal{A}_1=\{a_1\}$. Let $\pi_0$ be the fraction of decision makers who have consideration set $\mathcal{A}_0$ and $\pi_1$ the fraction that have consideration set $\mathcal{A}_1$, and assume that $\pi_0$ and $\pi_1$ are known to the analyst. Since decision makers with consideration set $\mathcal{A}_0$ do not consider, or are unaware of, alternative ${a_1}$ their choice of $a_0$ is deterministic from the analyst's standpoint. Similar reasoning applies to decision makers with consideration set $\mathcal{A}_1$. This model is coherent if $\pi_0 \leq p_0$ and $\pi_1 \leq p_1$. In this case
If either $\pi_0$ or $\pi_1$ are strictly positive the identification set does not contain the origin. The case where $\pi_0,\pi_1>0$ is illustrated in Figure (ref) where the identification set of $(\theta_0,\theta_1)$ is the darker rectangle. The combination of consideration sets and instrumental variables assumptions is demonstrated using our empirical application in Section (ref).
Here we consider a simple parametric version of the binary choice model presented above as this is commonly used in applications. This version of the model helps us illustrate some features of the general model and is used in Section (ref) to study the effect of policy interventions.
We focus on the case in which only the utility of one of the alternatives is interval-valued. In particular, utilities are as follows:
While the utility from alternative $a_0$ is assumed to be the singleton $0$ (location normalization), the utility from alternative $a_1$ is the random interval $[\beta_i-\sigma,\beta_i+\sigma]$. Note that $\sigma=0$ represents the case of complete preferences. Let $\beta_i =\Bar{\beta} - \varepsilon_i$ where $E(\varepsilon_i)=0$ and $Var(\varepsilon_i)=1$ (scale normalization). Assume that $\varepsilon \sim F_{\varepsilon}$ is continuously distributed and is independent across decision makers. For example, assuming that $\varepsilon_i$s are independent with a standard normal distribution yields the Probit model with incomplete preferences, while assuming a logistic distribution for $\varepsilon_i$ yields the Logit model with incomplete preferences. In most applications one assumes that $\beta_i=X_i \beta - \varepsilon_i$ where $X_i$ is the vector of observed characteristics of alternative $a_1$ as perceived by individual $i$. For simplicity we restrict our attention to case where no covariates are present.\footnote{Manski2018 discusses a similar model of linear utility functions in the context of Knightian uncertainty (see also our Section (ref) below). Moreover, estimation of this parametric model can be handled using modified minimum distance estimator developed in ManskiTamer2002.}
Following equation ((ref)), the identification region can be written as
Combining these inequalities gives, $$\mathbf{P}(\varepsilon_i< \beta -\sigma) \leq p_1 \leq \mathbf{P}(\varepsilon_i \leq \beta +\sigma).$$ If one assumes that $\varepsilon$ has a continuous everywhere monotone CDF, $F_{\varepsilon}$, the identification set is
Note that the above identification region includes the point where $\beta-\sigma=\beta+\sigma$ implying $\beta=F_{\varepsilon}^{-1}(p_1)$ and $\sigma=0$. In other words, the absence of vagueness cannot be excluded.
In this simple setting, an instrumental variable is a random variable $Z$ such that the distribution of $\varepsilon|Z$ changes over the support of $Z$ while the determinants of utility, $\beta$ and $\sigma$, remain constant across the values of $Z$. If such a variable exists, the corresponding identification region is
If an instrumental variable exists, one can reject the hypothesis that $\sigma=0$.
In this section, we analyze the impact on the choice probabilities of a policy intervention aimed at changing the utility individuals give to one of the alternatives. Examples of such a policy could be campaign advertising, or other form of information provision, that targets a particular alternative. The objective of the policy is to increase the fraction of individuals who choose the targeted alternative. We show how this objective is harder to reach once incomplete preferences are allowed. We focus on binary choice and we first look at a model which is similar to the one of the previous Sections. We then specialize the model further and consider a simple linear utility parametric version of the general model.
Formally, policy intervention is modeled by assuming one can add a positive quantity, $\Delta>0$, to the utility, both upper and lower, of alternative $a_1$. We denote by $p_1$ the choice probability of $a_1$ before the policy and by $p^{\Delta}_1$ the choice probability after adding $\Delta$ to the utilities from $a_1$. The effect of the policy is measured by $p^{\Delta}_1-p_1$. We first describe the policy effect when preferences are complete, and then move to the case in which completeness is not assumed.
Let $\mathcal{U}=(\underbar{u}_0,\bar{u}_0,\underbar{u}_1,\bar{u}_1)$ be the lower and upper utilities for alternatives $a_0$ and $a_1$, respectively. Except for assuming that $P(\underbar{u}_j \leq \bar{u}_j)=1$ for $j=0,1$ we leave these utilities to be unspecified.
If individuals' preferences are complete, the utility of each alternative is a unique number, and therefore $\underbar{u}_0=\bar{u}_0=u_0$ and $\underbar{u}_1=\bar{u}_1=u_1$. In this case, $\theta_1$ and $\theta_0$ are point-identified by the corresponding choice probabilities. In particular, before the policy is introduced we have
After the policy intervention, the utility from choosing alternative $a_1$ is increased by $\Delta>0$, and becomes $u_1+\Delta$. In this case, the probability of choosing alternative $a_1$ must satisfy
Clearly, $p^{\Delta}_1 \geq p_1$ and therefore one can predict that the policy impact, defined as $p^{\Delta}_1-p_1$, when preferences are complete is necessarily positive under some mild assumptions on the cumulative distribution function of $u_0-u_1$.
Assume now that decision makers hold incomplete preferences. As before, let $p_1$ be the observed choice probability of $a_1$ before the policy intervention. We know from Section (ref) that this probability must satisfy the following inequality
After adding $\Delta$ to both $\underbar{u}_1$ and $\bar{u}_1$, the choice probability for $a_1$ will satisfy
Again, we measure the policy impact by looking at $p^{\Delta}_1 - p_1$ after $p_1$ is observed.
Given the pre-policy choice probability and the fact that $p^{\Delta}_1 \geq F(\Delta)$, we can say that $p^{\Delta}_1 -p_1 \geq F(\Delta)-p_1$. The last difference can be negative if the policy intervention $\Delta$ is small enough to have $F(\Delta) < p_1$. Hence adding $\Delta>0$ to the utility from $a_1$ can lead to a decrease in the probability in which it is chosen - a result we could not get with complete preferences.
Next, we add parametric assumptions as in Section (ref). That is, utilities are:
and $\beta_i =\Bar{\beta} - \varepsilon_i$ where $\varepsilon$ has a standard normal distribution and therefore $\beta_i \sim N(\Bar{\beta},1)$. Consider a policy intervention that has the ability to change the utility interval individual $i$ assigns to alternative $a_1$ by a fixed amount $\Delta>0$. After the policy intervention, the utility interval becomes
When preferences are complete we have $\sigma=0$, and the choice probability for alternative $a_1$ before the policy is
Since $p_1 = \Pr(y_i=a_1)$ is identified from the data generating process, $\Bar{\beta}$ is identified as $\Bar{\beta}= \Phi^{-1}(p_1)$. After the policy is enacted, a quantity $\Delta>0$ is added to the utility of alternative $a_1$, and the choice probability for alternative $a_1$ is
Since $\Bar{\beta}$ is identified, we can write,
Therefore, the predicted effect of this policy is
because $\Delta>0$. Therefore, when preferences are complete the policy effect on the observed frequency with which alternative $a_1$ is chosen is always positive.
If preferences are not complete, $\sigma>0$ and the identification region is $$\Theta^I = \{ (\Bar{\beta},\sigma) : \Bar{\beta} -\sigma \leq \Phi^{-1}(p_1) \leq \Bar{\beta} +\sigma \}.$$ The corresponding identification region for $\Bar{\beta}$ is $$B^I = \left[\Phi^{-1}(p_1)-\sigma, \Phi^{-1}(p_1)+\sigma \right].$$ Therefore, observing $p_1$ does not point identify $\Bar{\beta}$ but rather gives a bound for it. A policy that adds $\Delta$ to the utility of alternative $a_1$ effectively changes $\Bar{\beta}$ to $\Bar{\beta}+\Delta$. Therefore, the predicted choice probability after the policy is $$p_1^{\Delta} \in \{ p_1 : \Phi(\beta+\Delta-\sigma) \leq p_1 \leq \Phi(\beta+\Delta-\sigma), \ \beta \in B^I \}.$$ Here the effect of adding $\Delta$ to the utility of alternative $a_1$ is no longer guaranteed to be positive. The sign of the effect depends on the relationship between $\Delta$ and $\sigma$ as one can see in the diagrams in Appendix (ref).
In this section we illustrate how some of our identification results can be adapted to additional assumptions about individuals' behavior or the data generating process. In particular, we discuss incompleteness due to ambiguity (Section (ref)) and minmax regret behavior (Section (ref)). For ease of exposition, we focus on binary choice throughout.
Our identification framework can be adapted to different models of behavior when preferences are not complete as long as behavior is described by two thresholds. In the following, we illustrate how this could be done for Knightian uncertainty as described in Bewley02. Bewley shows that a strict preference relation that is not necessarily complete, but satisfies all other axioms of a standard expected utility framework, can be represented by a family of expected utility functions generated by a unique utility index and a set of probability distributions. Lack of completeness is thus reflected in multiplicity of beliefs: the unique subjective probability distribution of the standard expected utility framework is replaced by a set of probability distributions. When the preference relation is complete this set becomes a singleton.
Let $S$ denote the state space and, with abuse of notation, its cardinality. $\Delta (S)$ is the set of all probability distributions over $S$. Given an alternative $x \in X\subset \boldsymbol{R^S}$, $u(x(s))$ denotes the utility that alternative yields in state $s$. If $\pi \in \Delta (S)$, the expected utility of individual $i$ according to $\pi$ is given by
Let $\Pi \subset \Delta(S)$ be a closed and convex set of probability distributions on $S$. According to Bewley's Knightian Decision Theory, decision maker $i$'s preferences $\succ $ are described by the following result:
If the inequality in ((ref)) changes direction for different probability distributions in $\Pi$, the two alternatives are not comparable.
This model does not compare alternatives using intervals. Even though the values of $E_{\pi }[u(\cdot)]$ form an interval, comparisons are not made looking at the extremes of that interval but one probability distribution at a time. Two alternatives could be ranked even if the corresponding utility intervals overlap. Despite this difference, the results of the previous section can be applied by taking advantage of some simple algebra. Rearranging equation ((ref)) one gets
and therefore
We can then adapt the definition of the set of non-dominated alternatives as follows
As before, we let $\theta=(\theta_0,\theta_1)$ be the probabilities that alternative $a_0$ is preferred to $a_1$ and the probability that alternative $a_1$ ir preferred to $a_0$, respectively. By definition,
From here on, the analysis can proceed along lines similar to the ones provided in Section (ref), and we thus leave the details to the reader. A parametric version of this model is discussed in Manski2010 and Manski2018. Manski presents the identification region for the parameters using two inequalities that corresponds to the Artstein's inequalities we use in this paper. For the binary choice, the inequalities in Manski2010 and Manski2018 are both necessary and sufficient to achieve sharpness. Generalizing this result to a multi-nomial choice model requires defining the random set $M$ of non-dominated alternatives similarly to equation ((ref)) and use Artstein's inequalities to form the sharp identification set (see BMM2011).
So far, we have been silent about the way individuals make their choices when two or more alternatives are not comparable. This approach resulted in partial identification of the parameters of interest. Here we focus on the idea of minimizing maximal regret as a way to break the consumers' indecision.
Assume that after a decision is made the individual learns which of the possible utility values represent her 'true' utility for each alternative. Regret occurs if the realized utility of the chosen option turns out to be lower than the realized utility of an alternative that was not chosen. For a binary choice situation, recall that
and thus when the two utility intervals overlap both choices can be rational. If the individual chooses $a_0$ her maximal regret is $\bar{u}(a_1)-\underbar{u}(a_0)$, while if she chooses $a_1$ her maximal regret is $\bar{u}(a_0)-\underbar{u}(a_1)$. Therefore, if the decision maker minimizes her maximal regret when undecided, the corresponding rational choice region is as follows:
The minmax regret rule selects one alternative from the random set $M$. Figure (ref) describes the choice rule in equation ((ref)). The region in Figure (ref) where choice was indeterminate is now split between alternatives so that points above the 45 degree line mean alternative $a_1$ is chosen and points below it mean alternative $a_0$ is chosen. Since choice is no longer indeterminate, the corresponding model is point identified.
We implement the methods described in previous sections using data on voting.\footnote{Other papers estimated partially identified parameters in the context of voting. For example, KawaiWtanabe2013 estimate a model of strategic voting and quantify the impact it has on election outcomes. IaryczowerShiShum2018 estimate the effect of deliberation on collective choices in the context of criminal cases decided in the US courts of appeals. Both papers adopt set estimation as a result of multiple equilibria.} Specifically, we use precinct-level data obtained from the Ohio Board of Elections for the 2018 midterm elections in Lorain county, and focus on the two races for Justice of the Supreme Court of Ohio.\footnote{Since this paper deals with identification, we treat estimators as population level quantities and leave the statistical issues for future research.} We focus on one county to illustrate our results because it consists of a relatively homogeneous population. Election data fits our framework in many respects, the main one being that individuals' choices are not observable and one can only observe the share received by each candidate. The data from Ohio is also interesting because the candidates' order on the ballot varies from precinct to precinct. This change of order provides a potential instrumental variable. A full description of the data on Ohio 2018 midterm elections is in Appendix (ref) including explanation of how candidates' order on the ballot rotates.
In the 2018 midterm elections voters in Ohio participated in eight statewide races. In six of these races candidates' party affiliation was listed on the ballot, while in the remaining two it was not. The races where no party affiliation was noted on the ballot were the two races for the position of Justice of the Supreme Court of Ohio. The candidates had been selected in party-run primaries, and thus were affiliated with a party, but their affiliation was not actually printed on the ballot itself. As Table (ref) shows, in Lorain county the percentage of voters who refrained from expressing a preference is above 20% when party affiliation is not listed. About 40,000 voters who went to the polls, voted in most state-wide races, did not vote in the two state-wide races where a candidate's party affiliation was not printed on the ballot.\footnote{Similar numbers hold statewide, with more than 800,000 voters not voting in the races for Justice of the Supreme Court after having gone to the polls.} This is an example of what is sometimes called roll-off voting: fewer votes are cast in down-ballot races. It represents a particular form of abstention because the voter has already incurred the costs associated with going to the polls.
The recent literature on ballot roll-off in judicial elections (see Hall-Bonneau-2008 or Marble-2017), is mostly focused on empirically measuring it and its possible determinants. Similarly to what we find, this literature shows that affiliation on the ballot decreases ballot roll-off. There is also a theoretical literature explaining that abstention could stem from asymmetric information (Feddersen-Pesendorfer-1999), or context-dependent voting (Callander-Wilson-2006). In our framework, there could be an alternative reason for roll-off voting: incomplete preferences. When unable to rank alternatives, voters may decide not to make a choice and therefore do not vote in the corresponding race. Notice that a choice will be made for them anyhow, because the winner of the election will be selected as judge. In the context of voting, it is plausible to assume that those who came to the polls and did not vote behave that way because they could not compare the candidates. In the notation of equation (ref) in Section (ref), it means we can assume that $\Theta^U = \{(0,0)\}$. In other situations when a portion of the individuals do not report their choices (missing observations), a more plausible assumption is that $\Theta^U =\{(\theta_0,\theta_1):0 \leq \theta_0+\theta_1 \leq 1\}$.
The two races for a seat on the Ohio Supreme Court were Baldwin versus Donnelly and DeGenaro versus Stewart. As mentioned above, for these races the party affiliation of the candidates was not indicated on the ballot. We know, however, that Donnely, in the first race, and Stewart, in the second race, were affiliated with the Democratic Party. Table (ref) describes the results for these two races in Lorain county.
The no assumptions identification region for $(\theta_0,\theta_1)$ appears in Figures (ref) and (ref). These bounds are based on the choice probabilities of those who cast a vote in these races.
The political science literature suggests that the order in which candidates are presented on the ballot may affect the chances of these candidates to be elected (see, for example, Krosnick-Miller-Tichy-04 and Meredith-Salant-2013). As a result, ordering the candidates alphabetically may favor candidates with certain family names. For this reason, the Ohio Board of Elections rotates the order in which candidates appear on the ballot among the precincts (see Appendix (ref)). As a result, a candidate may appear first on the ballot in one precinct and last in another nearby precinct. The order of a candidate on the ballot, by construction, presents an example of an instrumental variable. Focusing on Lorain County to achieve a relatively homogeneous population given the covariates, we use an instrumental variable indicating whether a candidate appeared first on the ballot at a certain precinct.
As one can see in equation ((ref)), in both races for Ohio's supreme court judgeship, being first on the ballot gives a certain advantage over being second.\footnote{The advantage of appearing first on the ballot, when averaged over the whole state, is rather small. At the county level, the effect of the order is either large, rather small, or even negative. This uncounted difference between counties may be due to omitted covariates. We solve this issue by focusing on one county. Further empirical investigation is left for future work.} In the Baldwin versus Donnelly contest, appearing first on the ballot gives Donnelly a 2.6% advantage on average. In the race of DeGenaro versus Stewart, the advantage of being first is estimated to be 3.3% on average. The predicted choice probabilities calculated at the mean values of the covariates are
and the corresponding identification regions are illustrated in Figure (ref).
Next, we take into account that many voters have not expressed a preference in these races even though they have done so in other races. Following Equation ((ref)), we take a weighted average of the identification regions depicted in Figure (ref) and the identification region for those who did not vote. The weights, denoted as $\gamma$, are the conditional abstaining probabilities computed for each race. Specifically, using Table (ref), in the first race $\gamma=0.2470$ and $\gamma=0.2646$ in the second race. In this voting application, it is plausible to assume that all voters who abstained from choosing a supreme court judge could not compare the two candidates. This assumption corresponds to assuming that $\Theta^U=\{(0,0)\}$ as discussed in Section (ref). With this assumption the identification region corresponds to the orange rectangles in Figures (ref) and (ref). In other applications, however, it may be more plausible to assume that $\Theta^U=\{(\theta_0,\theta_1): 0 \leq \theta_0+\theta_1 \leq 1\}$. The identification regions in this case are the gray areas in Figures (ref) and (ref).
Finally, we introduce consideration sets by assuming that registered democrats and republicans only considered voting for a candidate of their own party. Figures (ref) and (ref) combine instrumental variables, consideration sets, and the presence of abstention. The dark identification region uses as lower bound the fraction of voters who are registered for the candidate's party and as upper bound the conditional probabilities in equation ((ref)) times the corresponding $1 - \gamma$.
The identification sets described in Figures (ref) and (ref) demonstrate how all the tools developed in this paper can be combined together. In practice, however, a researcher could start from the no-assumptions bounds described in Section (ref) and then add the assumptions she deems plausible for the application at hand - either all of them or only a subset.
In this paper we provided the sharp identification region for a discrete choice model in which individuals' preference may not be complete and only aggregate choice data is available. The identification region is a strict subset of the parameter space, and thus disproves the idea that everything goes when preferences are not complete. The identification region provides intuitive bounds on parameters of interest of the probability distribution of preferences across the population. Since our assumption do not rule out complete preferences, the identification region admits the two extreme possibilities of maximal incompleteness in which nobody can rank alternatives and no choice-relevant incompleteness in which everyone can rank alternatives. We illustrate how the existence of instrumental variables can rule out this last possibility.
Although we use interval orders as a way to describe preferences and behavior when incompleteness is allowed, our results extend beyond that model. In particular, any theory where behavior depends on two thresholds (instead of just one) can be accommodated into our framework. We have illustrated this using the Knightian decision theory of Bewley02, but we believe our results would also extend to the multi-utility framework of Aumann62 and DubraMaccheroniOk04, or to the more recent twofold conservatism model of Echenique-Pomatto-Vinson, as well as individual decision making models with “thick” indifference curves. All that one needs is to have a set of non-dominated alternatives that depends on two numbers per each alternative.