Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
179,557 characters · 16 sections · 0 citation commands
Identification in discrete choice models with imperfect information
\thispagestyle{empty} \setcounter{page}{1}
A fundamental issue in the empirical analysis of decision problems is the presence of frictions that prevent agents from learning the payoffs associated with the available alternatives. When estimating parameters and making counterfactual predictions, it is common to make strong assumptions about agents' beliefs, which can weaken the credibility of the results. In particular, most of the applied literature either imposes perfect information, or incorporates information frictions by fully specifying agents' beliefs, as in search models (\hyperlink{Mehta}{Mehta, Rajiv, and Srinivasan, 2003}; \hyperlink{Honka}{Honka and Chintagunta, 2016}; \hyperlink{Ursu}{Ursu, 2018}), models with rational inattention (\hyperlink{Caplin_Dean}{Caplin and Dean, 2015}; \hyperlink{Matejka_McKay}{Mat\u{e}jka and McKay, 2015}; \hyperlink{Fosgerau}{Fosgerau, Melo, de Palma, and Shum, 2020}; \hyperlink{Csaba}{Csaba, 2018}; \hyperlink{Caplin_Dean_Leahy}{Caplin, Dean, and Leahy, 2019}; \hyperlink{Brown}{Brown and Jeon, 2020}), and models with preferences for risk (for a review, see \hyperlink{Barseghyan}{Barseghyan, Molinari, O'Donoghue, and Teitelbaum, 2018}). Instead, this paper develops a methodology to identify preferences and counterfactual outcomes from cross-sectional choice data that imposes weak restrictions on the agents' beliefs and, hence, is robust to whether agents are perfectly or partially informed.
We consider a large class of static single-agent discrete choice models, where the decision maker (DM) has to choose an alternative from a finite set. The payoff generated by each alternative depends on the state of the world, which is randomly determined by nature. The DM has a prior on the state of the world. Moreover, the DM can refine their prior upon reception of a private signal representing the DM's information structure. This information structure can range from full revelation of the state of the world to no information whatsoever, depending on the latent frictions encountered by the DM in the learning process. The DM uses the acquired information structure to update their prior and obtain a posterior through the Bayes' rule. Lastly, the DM chooses an alternative maximising their expected payoff, where the expectation is computed via the posterior. Under additional assumptions on information structures, this framework accommodates additive random utility discrete choice models, discrete choice models with risk aversion, discrete choice models with rational inattention, and {some} discrete choice models with search. Our objective is (partially) identifying preferences and counterfactual outcomes while remaining agnostic about information structures.
The model just described is a game against nature (\hyperlink{Milnor}{Milnor, 1951}). That is, it is a 1-player game in which a single self-interested player must choose a strategy. The player's payoff depends on their own strategy and the realization of the state of the world, which is decided at random by a totally disinterested nature. Thus, we can use results from the theoretical and empirical literature on $N$-player games with $N\geq 1$ and weak assumptions on information structures to characterize the sharp identified set for the payoff parameters. In particular, we revisit our framework through the lens of 1-player Bayes Correlated Equilibrium (\hyperlink{Kamenica}{Kamenica and Gentzkow, 2011}; \hyperlink{BM_1}{Bergemann and Morris, 2013}; \hyperlink{BM_2}{2016}). A fundamental result in the theoretical literature on robust predictions (Theorem 1, \hyperlink{BM_2}{Bergemann and Morris, 2016}) is that the set of optimal strategies predicted by our model under a large range of possible information structures is {equivalent} to the collection of model-implied choice probabilities under the notion of 1-player Bayes Correlated Equilibrium. Further, the latter collection is a convex set defined by linear equalities and inequalities. Therefore, as shown by \hyperlink{Tamer_auctions}{Syrgkanis, Tamer, and Ziani (2021)} and \hyperlink{Magnolfi_Roncoroni}{Magnolfi and Roncoroni (2023)}, determining whether a given parameter value belongs to the sharp identified set consists of solving a linear program, which is a well-understood and computationally tractable problem.\footnote{Observe that the collection of model-implied choice probabilities under the notion of 1-player Bayes Correlated Equilibrium can also be written as the Aumann expectation of the random set of 1-player Bayes Correlated Equilibria. Therefore, the above characterization of the sharp identified set is equivalent to the one provided by \hyperlink{BMM}{Beresteanu, Molchanov, and Molinari (2011)}. A distinctive feature of our framework is that this Aumann expectation is defined by linear equalities and inequalities and, therefore, can be computed by solving a linear program. }
We make two methodological contributions. First, we develop a formal procedure to practically construct the sharp identified set for the payoff parameters when the state of the world is continuous. In such a case, a 1-player Bayes Correlated Equilibrium is an infinite-dimensional object, and, thus, the linear program to solve for each candidate parameter value is also infinite-dimensional. Previous papers simplify the analysis by assuming that the state of the world is discrete, or allowing for a continuous state of the world but, in practice, discretising its support in some arbitrary bins to operationalise the linear programming procedure. Here, instead, we propose a sieve approximation of the program along the lines of \hyperlink{Han}{Han and Yang (2023)} in their study of treatment effects. We test this procedure in simulations and implement it in the empirical application.
Second, we characterize sharp bounds on the counterfactual choice probabilities when agents receive information about the state of the world via a policy program. This is an important question in the empirical literature on single-agent decision problems across different fields and is largely absent in the literature on many-player games. To cite just a few examples, see \hyperlink{Hastings}{Hastings and Tejeda-Ashton (2008)} on retirement fund options in Mexico, \hyperlink{Hastings2}{Hastings and Weinstein (2008)} and \hyperlink{bettinger}{Bettinger, et al. (2012)} on school choice, and \hyperlink{Kling}{Kling, et al. (2012)} on Medicare Part D prescription drug plans. This question is typically answered using field experiments which can be very costly. Our characterization represents a powerful result because it allows the analyst to assess the effect of programs of information provision before conducting such interventions. We also provide sharp bounds on the maximum potential welfare cost of limited information which may keep agents from choosing their first best.
Lastly, for readers keen to dive deeper into the identification nuances of our framework, Section (ref) explores how it differs from traditional 2-player games commonly studied in the empirical literature.
We use our methodology to study voting behaviour in the UK. We consider the spatial model of voting, which is an important framework in political economy to explain individual preferences for parties (\hyperlink{Downs}{Downs 1957}; \hyperlink{Black}{Black, 1958}). This model postulates that an agent has a most preferred policy and votes for the party whose position is closest to their ideal. In empirical analysis, it is typically implemented by estimating a classical parametric discrete choice model with perfect information (\hyperlink{Alvarez1995}{Alvarez and Nagler, 1995}; \hyperlink{Alvarez1998}{1998}; \hyperlink{Alvarez2000}{2000}; \hyperlink{Alvarez2000_2}{Alvarez, Nagler, and Bowler, 2000}). However, in reality, uncertainty pervades voting (\hyperlink{Shepsle}{Shepsle, 1972}; \hyperlink{Weisberg}{Weisberg and Fiorina, 1980}; \hyperlink{Enelow}{Enelow and Hinich, 1981}; \hyperlink{Baron}{Baron, 1994}; \hyperlink{Matsusaka}{Matsusaka 1995}; \hyperlink{Carpini}{Carpini and Keeter, 1996}; \hyperlink{Lupia}{Lupia and McCubbins, 1998}; \hyperlink{Feddersen_Pesendorfer}{Feddersen and Pesendorfer, 1999}; \hyperlink{Tabellini2}{Mat\u{e}jka and Tabellini, 2019}). That is, voters may be aware of their own and the parties' attitudes towards some popular issues, but they might be less informed on how they themselves and the parties stand towards more technical or less debated topics, and on the traits of the candidates other than those publicly advertised. Further, their competence on these matters is likely to be arbitrarily different depending on, for example, political sentiment, civic sense, attentional limits, media exposure, and candidates' candor.
Despite the acknowledgement of the central role played by the sophistication of voters in determining voting patterns, only a few empirical works have attempted to take it into account while estimating a spatial voting framework (\hyperlink{Aldrich}{Aldrich and McKelvey, 1977}; \hyperlink{Bartels}{Bartels, 1986}; \hyperlink{Palfrey_Poole}{Palfrey and Poole, 1987}; \hyperlink{Franklin}{Franklin, 1991}; \hyperlink{Alvarez_solo}{Alvarez, 1998}; \hyperlink{Degan_Merlo2}{Degan and Merlo, 2011}). This has been done by exogenously and parametrically modelling how information frictions affect the perceptions of DMs about the returns to voting (for instance, via an additive, exogenous, and parametrically distributed evaluation error in the payoffs), or by parametrically specifying the probability of being informed versus uninformed when voting. Instead, our methodology permits us to incorporate voter uncertainty under weak assumptions on the latent, heterogeneous, and potentially endogenous process followed by voters to gather and evaluate information.
In particular, we focus on a setting where the state of the world consists of distances between the voters and the parties' ideological positions on a few popular policy issues, and of voter-party-specific taste variables capturing voter perception on candidates' qualities and parties' positions on obscure topics. We assume that each voter observes the realization of the former, but may be uncertain about the realization of the latter. We estimate the model using data from the British Election Study, 2017: Face-to-Face Post-Election Survey (\hyperlink{Fieldhouse}{Fieldhouse, et al., 2018}) on the UK general election held on 8 June 2017. We compare our findings with the results one gets under the standard assumption that all agents are perfectly informed about the returns to voting. Several conclusions on the payoff parameters achieved under the complete information assumption are not unambiguously corroborated when we remain agnostic about information structures.
We use our characterization of counterfactual bounds to robustly assess to what extent imperfect information affects the well-being of voters and parties. We imagine an omniscient mediator implementing a policy that gives voters perfect information about the state of the world. We simulate the counterfactual vote shares and study how they change compared to the factual scenario. This question has been debated at length in the literature. Political scientists have often answered it by arguing that a large population composed of possibly uninformed citizens act as if it was perfectly informed (for a review, see \hyperlink{Bartels2}{Bartels, 1996}). \hyperlink{Carpini}{Carpini and Keeter (1996)}, \hyperlink{Bartels2}{Bartels (1996)}, and \hyperlink{Degan_Merlo2}{Degan and Merlo (2011)} provide quantitative evidence to disconfirm such claims; the first two by using auxiliary data on the level of information of the survey respondents as rated by the interviewers or assessed by test items, and the latter by parametrically specifying the probability that a voter is informed. We contribute to this literature by providing a way to construct counterfactual vote shares under perfect information, which neither requires the difficult task of measuring voters' knowledge level in the factual scenario, nor imposes parametric assumptions on the probability that a voter is informed. We find that voters benefit from full information as it leads to a considerable drop in the abstention rate. We also find that transparency harms the two historically dominant parties, i.e., the Conservative Party and the Labour Party, and favours the other minor parties, i.e., the Liberal Democrats and the Green Party. This suggests that some payoff-relevant information is unobserved by voters, and the biggest parties in the British political scene benefit from such uncertainty. Moreover, we quantify the maximum voters' welfare cost of limited information and find that it is comparable in magnitude to reducing the left-right ideological distance from a given party by around three points.
In addition to studying the impact of information provision, we investigate how the parties' welfare changes when they modify their ideological positions on popular policy issues. We adapt to our setting Theorem 1 of \hyperlink{Bergemann_Brooks_Morris}{Bergemann, Brooks, and Morris (2022)}, which permits us to answer such a question while holding fixed the voters' information structures in the counterfactual scenario. In agreement with several post-election studies, we find that, by holding a strong left ideological position about tax and social care, the Labour Party gained numerous votes during the election campaign.
The remainder of the paper is organised as follows. Section (ref) describes the model. Section (ref) discusses identification. Section (ref) presents some simulations. Section (ref) illustrates the empirical application. Finally, section (ref) concludes. The proofs are in Appendix (ref).
In what follows, we shorten \hyperlink{BM_2}{Bergemann and Morris (2016)} as BM16, \hyperlink{Tamer_auctions}{Syrgkanis, Tamer, and Ziani (2021)} as STZ21, and \hyperlink{Magnolfi_Roncoroni}{Magnolfi and Roncoroni (2023)} as MR23. We refer to 1-player Bayes Correlated Equilibrium as 1BCE. Given a random variable $Z$ with support $\mathcal{Z}$, $P_Z$ denotes its probability mass function (if $Z$ is discrete) or density (if $Z$ is continuous). $\Delta(\mathcal{Z})$ denotes the set of all probability mass functions or densities with support contained in $\mathcal{Z}$.
We consider a DM who faces the problem of choosing an alternative from a {finite} set, $\mathcal{Y}$. There is an unknown state of the world, $V$, with support $\mathcal{V}$, that enters directly in the DM's utility, $u: \mathcal{Y}\times \mathcal{V} \rightarrow \mathbb{R}$. $u\in \mathcal{U}$, where $\mathcal{U}$ is the set of all functions mapping $\mathcal{Y}\times \mathcal{V}$ to $\mathbb{R}$. The DM does not observe the realization of $V$ and has a prior belief on it, $P_V$. Before making a choice, the DM has an opportunity to learn more about the state of the world and resolve some uncertainty about the payoffs generated by the alternatives. Formally, the DM can refine their prior on $V$ upon reception of a private signal, $T$, with support $\mathcal{T}$ and distribution $P_{T|V}(\cdot| v)$ conditional on $V=v$. We denote by $\mathcal{P}_{T| V}$ the family of the signal's conditional distributions for each $v\in \mathcal{V}$, i.e., $\mathcal{P}_{T| V}\coloneqq \{P_{T|V}(\cdot| v): v \in \mathcal{V}\}$. The DM uses $\mathcal{P}_{T| V}$ and the received signal realization, $t$, to update their prior on $V$ via Bayes' rule and obtains a posterior, $P_{V|T}(\cdot|t)$. The DM chooses alternative $y\in \mathcal{Y}$ maximising their expected utility computed under the posterior, $ \int_{\mathcal{V}} u(y, v) P_{V| T}(v|t)dv$. If there is more than one maximising alternative, the DM applies some tie-breaking rule.
The informativeness of $T$ about $V$ (in the Blackwell sense) is inherently related to the frictions potentially encountered by the DM while investigating the state of the world.\footnote{\hyperlink{blackwell51}{Blackwell (1951;} \hyperlink{blackwell53}{1953)} provides a rank-ordering of information structures in terms of their informativeness.} These frictions can stem from various sources, such as attentional and cognitive limits, financial constraints, spatial and temporal boundaries, and cultural and personal biases. If the DM faces no information frictions, they may process a signal revealing the realization of $V$ and discover the payoffs with certainty. Instead, if the DM experiences considerable information frictions, they may process a signal adding nothing to their prior on $V$. A signal whose informativeness is between such two extremes is plausible as well. In a typical empirical application, the information frictions the DM encounters are not observed by the researcher. Hence, we proceed without assumptions on $\mathcal{T}$ and $\mathcal{P}_{T| V}$.
We now provide a more compact representation of our framework. Following the terminology of \hyperlink{BM_2}{BM16}, we define the {\it baseline decision problem} faced by the DM as $G\coloneqq \{\mathcal{Y}, \mathcal{V}, u, P_V\}$. We also define the {\it information structure} processed by the DM as $S\coloneqq\{\mathcal{T}, \mathcal{P}_{T|V}\}$. $G$ represents what the DM knows before processing any signal. $S$ consists of the additional information the DM learns about $V$, together with the received realization of the signal. $S$ belongs to $\mathcal{S}$, which is the set of all possible information structures processed by the DM. $\mathcal{S}$ contains the information structure giving complete information ({\it complete information structure}), the information structure giving no information in addition to the prior ({\it null information structure}), and any information structure whose informativeness is between those two extremes. The pair $\{G, S\}$ constitutes the {\it augmented decision problem} faced by the DM.
We denote by $Y$ the DM's choice. An optimal (mixed) strategy in the augmented decision problem $\{G, S\}$ is a distribution of $Y$ conditional on $T$, $\mathcal{P}_{Y|T}\coloneqq \{P_{Y|T}(\cdot| t): t \in \mathcal{T}\}$, such that, for each $t\in \mathcal{T}$, the DM maximises their expected utility by choosing any alternative $y\in \mathcal{Y}$ featuring $P_{Y|T}(y| t)>0$.\footnote{A mixed strategy arises if there are ties.}
Observe that the model just described is a {\it game against nature} (\hyperlink{Milnor}{Milnor, 1951}). That is, it is a 1-player game in which a single self-interested player must choose a strategy. Nature is indifferent among outcomes, has no payoff, and chooses $V$ through randomisation. The player's payoff depends on their own strategy and the realization of $V$. One can categorize our model as a specific example of 2-player games, where nature serves as the second player. However, our framework diverges from traditional 2-player games commonly studied in the empirical literature. In those games, the researcher models the payoffs of both players, both players can affect each other's payoffs through their choices, and such choices are observed by the researcher. In contrast, in our model, nature's payoff is not specified, nature selects a value for $V$ randomly from $P_V$, and this value remains unobserved by the researcher. These differences have implications for identification power, as discussed in Section (ref).
Our framework encompasses several settings of empirical interest, primitives of which are typically estimated under strong assumptions on $\mathcal{T}$ and $\mathcal{P}_{T| V}$. For example, it includes additive random utility discrete choice models (Logit, Nested Logit, Mixed Logit, Probit, etc.), discrete choice models with preferences for risk (\hyperlink{Barseghyan_AER}{Barseghyan, Molinari, O' Donoghue, and Teitelbaum, 2013}; \hyperlink{Barseghyan_QE}{Barseghyan, Molinari, and Teitelbaum, 2016}), models with rational inattention (\hyperlink{Caplin_Dean}{Caplin and Dean, 2015}; \hyperlink{Matejka_McKay}{Mat\u{e}jka and McKay, 2015}; \hyperlink{Csaba}{Csaba, 2018}; \hyperlink{Caplin_Dean_Leahy}{Caplin, Dean, and Leahy, 2019}; \hyperlink{Brown}{Brown and Jeon, 2020}; \hyperlink{Fosgerau}{Fosgerau, Melo, de Palma, and Shum, 2020}), and some search models (\hyperlink{Hebert}{H\'ebert and Woodford, 2018}; \hyperlink{Morris}{Morris and Strack, 2019}). See Appendix (ref) for more details. Also, see Appendix (ref) for a connection with the consideration set literature.
We assume that the utility function, $u$, and prior, $P_V$, belong to parametric classes, $\{u(\cdot; \theta_u)\}_{\theta_u\in \Theta_U}$ and $\{P_V(\cdot; \theta_V)\}_{\theta_V\in \Theta_V}$, indexed by the finite-dimensional structural parameters $\theta_u\in \Theta_U$ and $\theta_V\in \Theta_V$, respectively. Let $\theta\coloneqq (\theta_u, \theta_V)\in \Theta\coloneqq\Theta_U\times \Theta_V $ denote a generic parameter vector and $\theta_0\coloneqq (\theta_{0,u}, \theta_{0,V})$ denote the true parameter vector. Let $Y_1,\dots, Y_n$ be an i.i.d. sample of choices, where each choice is the outcome of the augmented decision problem, $\{G(\theta_0), S_1\},\dots, \{G(\theta_0), S_n\}$, respectively.
We do not know the exact information structures, $S_1,\dots, S_n$, that were processed in each of these decision problems and remain agnostic about those. In particular, we allow $S_i$ to be different from $S_j$ for each $i\neq j$, implying that the empirical distribution of choices, $\mathbb{P}_Y$, is a mixture of optimal strategies over various information structures. This heterogeneity embeds the fact that different agents could encounter different information frictions and hence process more or less informative signals. We treat the information structures processed by DMs as nuisance parameters and study the question of identifying $\theta_0$ and counterfactuals of interest from $\mathbb{P}_Y$.\footnote{It is implicit in our discussion that we also remain agnostic about tie-breaking rules and allow them to be heterogenous across DMs.}
All DMs are assumed to rely on a common prior, $P_V(\cdot; \theta_{0,V})$. Some heterogeneity of priors across DMs can be introduced by including discrete payoff-relevant variables, $(X,\epsilon)$, that are observed by DMs together with the signal and correlated with $V$. In that case, each DM has a prior $P_{V|X,\epsilon}(\cdot;| x,e; \theta_{0,V})$ conditional on $(X,\epsilon)=(x,e)$. See Section (ref) and Appendix (ref) on how to add $(X,\epsilon)$ to our framework.
In certain settings, some or all the components of $V$ are observed by the researcher. For example, in models of insurance plans, the researcher often has data on the ex-post claim experience of the agents in the sample (see Example (ref) in Appendix (ref)). In those cases, $\theta_{0,V}$ could be identified directly from such additional data. In our general discussion below, we focus on the scenario where $V$ is unobserved to the researcher. This is the case considered in our empirical application on voting behaviour.
The identified set of $\theta_0$ can be characterised according to Result 3 of \hyperlink{Tamer_auctions}{STZ21} which encompasses any $N$-player games with $N\geq 1$. We explain how this result can be adapted to our specific case for the sake of completeness in our exposition. Intuitively, the identified set of $\theta_0$ is the set of $\theta$s for which the model predicts a distribution of $Y$ that matches $\mathbb{P}_Y$. Let $\mathcal{R}(\theta,S)$ be the set of optimal strategies of $\{G(\theta), S\}$ and $\mathcal{R}(\theta)$ be the set of model-implied choice probabilities while remaining agnostic about information structures. That is,
where we have used the fact that $Y$ is independent of $V$ conditional on $T$. Convexification (via the convex hull operator, $\text{Conv}\{\cdot\}$) allows us to include in $\mathcal{R}(\theta)$ distributions of $Y$ that are mixtures of optimal strategies over various information structures. This ensures the heterogeneity of information structures in the cross-section, as discussed above.\footnote{To understand the convexification step better, note that the information structures in our framework are econometrically similar to the equilibrium selection mechanisms in incomplete many-player games (\hyperlink{Tamer}{Tamer, 2003}; \hyperlink{CT}{Ciliberto and Tamer, 2009}). In the former, convexification allows the information structures to differ across DMs. In the latter, convexification allows the equilibrium selection mechanisms to differ across markets. See also \hyperlink{Tamer_auctions}{STZ21} and \hyperlink{Magnolfi_Roncoroni}{MR23} about convexification.} The identified set of $\theta_0$ is defined as, $$ \Theta^*\coloneqq \{\theta\in \Theta: \mathbb{P}_Y\in \mathcal{R}(\theta) \}. $$
The above definition of $\Theta^*$ is not helpful in practice. This is because constructing $\mathcal{R}(\theta)$ following ((ref)) is infeasible due to the necessity of exploring the large class $\mathcal{S}$, which contains infinite-dimensional objects. We overcome this issue by recalling that our decision problem is a 1-player game (game against nature), as discussed in Section (ref). Hence, we can use results from the theoretical literature on $N$-player games with $N\geq 1$ and weak assumptions on information to give a simpler characterization of $\mathcal{R}(\theta)$ and, in turn, $\Theta^*$.
In particular, we consider the notion of 1BCE (\hyperlink{Kamenica}{Kamenica and Gentzkow, 2011}; \hyperlink{BM_1}{Bergemann and Morris, 2013}; \hyperlink{BM_2}{BM16}). This notion refers to a theoretical setting where an omniscient mediator makes incentive-compatible recommendations to the DM as a function of the state of the world and consistent with the DM's prior. If the DM follows such recommendations, the resulting distribution of choices is a 1BCE. A fundamental result in the theoretical literature on robust predictions (Theorem 1, \hyperlink{BM_2}{BM16}) is that the set of optimal strategies that arise from adding an arbitrary information structure to the baseline decision problem $G(\theta)$ is equivalent to the set of 1BCE of $G(\theta)$. Therefore, to obtain $\mathcal{R}(\theta)$, we do not need to explore the collection of all possible information structures, but instead, we can calculate the set of 1BCE of $G(\theta)$. In the remainder of the section, we formalise the equivalent characterization of $\mathcal{R}(\theta)$ and $\Theta^*$ based on 1BCE. In Section (ref), we zoom into the computational part. In Section (ref), we discuss the identification power of our model. In Section (ref), we characterize bounds on counterfactuals of interest.
First, we give the definition of 1BCE of $G(\theta)$. In what follows, $P_{Y,V}$ denotes the joint distribution of $Y$ and $V$.
We now state Theorem 1 of \hyperlink{BM_2}{BM16}.
We use Theorem (ref) to equivalently rewrite $\mathcal{R}(\theta)$ and $\Theta^*$. For each $\theta\in \Theta$, let $\mathcal{W}(\theta)$ be the set of 1BCEs of $G(\theta)$. Let $\mathcal{Q}(\theta)$ be the set of distributions of $Y$ that arise from the 1BCEs of $G(\theta)$:
Theorem (ref) implies that $ \mathcal{R}(\theta)=\mathcal{Q}(\theta)$. This can be seen in three steps. First, observe that the {\it Consistency} and {\it Obedience} requirements defining 1BCE are linear in $P_{Y,V}$. Therefore, $\mathcal{Q}(\theta)$ is a convex set. Second, $\mathcal{R}(\theta)$ contains distributions of $Y$ that are mixture of optimal strategies over various information structures. Hence, by Theorem (ref), each $P_Y\in \mathcal{R}(\theta)$ maps into a mixture of 1BCEs. Third, since $\mathcal{Q}(\theta)$ is convex, any mixture of elements from the set is itself an element of $\mathcal{Q}(\theta)$. Therefore, $ \mathcal{R}(\theta)=\mathcal{Q}(\theta)$ and we can use $\mathcal{Q}(\theta)$ in place of $\mathcal{R}(\theta)$ to characterize $\Theta^*$, as shown by Result 3 of \hyperlink{Tamer_auctions}{STZ21} and stated in our Proposition (ref).
To see how to use Proposition (ref) in practice, we first rewrite it in a more explicit way using Definition (ref). By Proposition (ref) and Definition (ref), $\theta$ belongs to $ \Theta^*$ if and only if there exists a function $P_{Y,V}: \mathcal{Y}\times \mathcal{V}\rightarrow \mathbb{R}$ which satisfies the following constraints:
Observe that the above constraints are linear in $P_{Y,V}$. Therefore, when $\mathcal{V}$ is a finite set, ((ref)) reduces to a {\it finite-dimensional} linear program:
In turn, we can construct $\Theta^*$ following this procedure: first, generate a grid of points covering $\Theta$ as precisely as possible, depending on the available computational resources; second, for each $\theta$ in such a grid, check if ((ref)) has a solution with respect to $P_{Y,V}$; third, any $\theta$ for which ((ref)) has a solution belongs to the identified set.\footnote{Note that there is no need to recover the entire set of solutions of ((ref)) for a given $\theta$. Existence of at least one solution of ((ref)) is sufficient to include such a $\theta$ in the identified set.}
Observe that, when $\mathcal{V}$ is finite, we can dispense with the parameterisation of $u$ and $P_V$ via $\theta$, as these functions can be fully and flexibly characterised by a finite number of parameters, one for each combination of values of $(Y,V)$. In this case, we can also add nonparametric restrictions on $P_V$ to the linear program, such as monotonicity, concavity/convexity, and Lipschitz restriction, which can be written as linear constraints.
Further, we remark that ((ref)) is linear in $P_{Y,V}$ for a given $\theta$, but it is {\it not} linear in {\it both} $P_{Y,V}$ and $\theta$. This is why we grid over $\Theta$ to construct $\Theta^*$. Linearity of ((ref)) in $P_{Y,V}$ {\it and} the parameters is achieved when $u$ is {\it known} and $P_V$ is treated as a finite-dimensional vector of parameters without indexing it by $\theta_V$. In this case, we can obtain the identified set of moments of $P_V$ by solving a {\it unique} linear program, without gridding, as proposed by \hyperlink{Tamer_auctions}{STZ21} for an auction framework. In single-agent decision problems and other types of games, $u$ is typically unknown and, therefore, we have linearity {\it only} in $P_{Y,V}$.
The setting with finite $\mathcal{V}$ is considered by \hyperlink{Tamer_auctions}{STZ21} and \hyperlink{Magnolfi_Roncoroni}{MR23}. We refer to those papers for a discussion on the computational burden of solving ((ref)) as the cardinalities of $\mathcal{Y}$ and $\mathcal{V}$ increase.
When $\mathcal{V}$ is not a finite set, the simple finite-dimensional linear programming approach is no longer applicable as $P_{Y,V }$ is an {\it infinite-dimensional} object. This case has not been addressed in the econometric literature on games, where assuming a finite $\mathcal{V}$ or discretising a non-finite $\mathcal{V}$ in a few arbitrary bins to operationalise ((ref)) are standard practices. Here we propose a formal procedure to approximate $\Theta^*$ when $\mathcal{V}$ is not finite, along the lines of \hyperlink{Han}{Han and Yang (2023)}. As a preliminary step, it is useful to equivalently rewrite ((ref)) using the distribution of $Y$ conditional on $V$ as unknown, $\mathcal{P}_{Y|V}\coloneqq \{P_{Y|V}(\cdot |v): v\in \mathcal{V}\}$:\footnote{Note that the {\it Consistency} constraint is redundant in ((ref)) as $\sum_{y\in \mathcal{Y}} P_{Y|V}(y|v)P_{V}(v; \theta_V)=P_{V}(v; \theta_V)$ becomes $\sum_{y\in \mathcal{Y}} P_{Y|V}(y|v)=1$, which is one of the probability requirements.} \nobreak {
}
We start from the simple case where $V$ is a random variable and $\mathcal{V}\coloneqq [v_\ell, v_u]$. Consider the following sieve approximation of $P_{Y|V}(y|v)$ using Bernstein polynomials of order $K$:
where $a_{k,K}(v)\coloneqq {K \choose k} \frac{(v-v_\ell)^k (v_u-v)^{K-k}}{(v_u-v_\ell)^K}$ is a univariate Bernstein basis, $\lambda_{k,K}^y \coloneqq P_{Y|V}(y| v_\ell+(v_u-v_\ell)\frac{k}{K})$ is its coefficient, $\mathcal{K}\coloneqq \{0,1,\dots,K\}$, and $K$ is finite. It can be shown that this approximation tends to $P_{Y|V}(y|v)$ uniformly over $\mathcal{V}$ as $K$ goes to infinity. Observe that $a_{k,K}(v)\geq 0$. Hence, for each $y$, $P_{Y|V}(y|v)\geq 0$ for each $v$ if and only if $\lambda_{k,K}^y\geq 0$ for each $k$. Further, $\sum_{y\in \mathcal{Y}} P_{Y|V}(y|v) =1$ is approximately equal to $\sum_{y\in \mathcal{Y}} \lambda_{k,K}^y= 1$ for each $k$ (\hyperlink{Coolidge}{Coolidge, 1949}).\footnote{See also \hyperlink{Mogstad}{Mogstad, Santos, and Torgovitsky (2018)} for using Bernstein polynomials in their study of treatment effects to approximate the marginal treatment effect function.} Motivated by this result, a finite-dimensional linear program approximating ((ref)) can be obtained with respect to $\lambda\coloneqq (\lambda_{k,K}^y: (k,y)\in \mathcal{K}\times \mathcal{Y})$:
where $$
$$ We can then approximate $\Theta^*$ by verifying if (\ref{lin_pr_3}) has a solution with respect to $\lambda$ for each $\theta\in \Theta$.
The same logic applies when $V$ is a $D\times 1$ random vector and $\mathcal{V}\coloneqq \times _{d=1}^D [v_{\ell_d}, v_{u_d}] $. In this case, we can use ((ref)), where $\mathcal{K}\coloneqq \times_{d=1}^D\{0,1,\dots, K_d\}$, $K_d$ is finite for each $d=1,\dots, D$, $K$ is the cardinality of $\mathcal{K}$, $a_{k,K}(v)\coloneqq \prod_{d=1}^D{K_d \choose k_d} \frac{(v_d-v_{\ell_d})^{k_d} (v_{u_d}-v_d)^{K_d-k_d}}{(v_{u_d}-v_{\ell_d})^{K_d}}$ is a $D$-variate Bernstein basis, and $\lambda_{k,K}^y \coloneqq P_{Y|V}(y| v_{\ell_1}+(v_{u_1}-v_{\ell_1})\frac{k_1}{K_1},\dots ,v_{\ell_D}+(v_{u_D}-v_{\ell_D})\frac{k_D}{K_D} )$ is its coefficient. We can further generalise Bernstein polynomials to the case where $\mathcal{V}$ is the real line (\hyperlink{Szasz}{Szasz, 1950}; \hyperlink{Butzer}{Butzer, 1954}).
Suppose there are other {\it discrete} payoff-relevant variables which enter the DM's information set together with the signal. In this case, one should solve ((ref)) for each value of such variables. For instance, $u$ could depend on covariates $X$ observed by the DM and the researcher. $u$ could also depend on some variable $\epsilon$ observed by the DM but unobserved by the researcher. Our procedure allows $(X,V,\epsilon)$ to be correlated, so that DMs can have heterogenous priors, $P_{V|X,\epsilon}(\cdot;| x,e; \theta_{V})$ conditional on $(X,\epsilon)=(x,e)$. See Appendix (ref) on how to include $(X,\epsilon)$ in ((ref)).
Verifying if ((ref)) has a solution for a given $\theta$ is computationally easy using standard algorithms, such as the simplex optimizer or the interior-point optimizer. In Section (ref), we discuss how to choose $K$ in practice and the computing time.
Our identification procedure imposes weak assumptions on the DMs' behavior by allowing for any information structures through the 1BCE characterization. One might wonder whether being agnostic about information structures strips discrete choice models of any empirical content. Proposition (ref) shows that, {\it even under the restrictive assumption of complete information}, discrete choice models do not retain identification power with respect to $(u,P_V)$, unless further assumptions are made, such as exogenous covariates and (non)parametric restrictions on $(u,P_V)$. Therefore, to generate informative bounds, our framework necessitates additional assumptions on $(u,P_V)$, mirroring the requirements of traditional “complete-information” discrete choice models. In particular, in our simulations and empirical application, we show that our model can generate relatively tight bounds by introducing exogenous covariates and parameterising $u$ and $P_V$.
Nevertheless, even in the presence of exogenous covariates and parametric restrictions, we expect our framework to generally provide less informative bounds on the primitives than those in 2-player games and no assumptions on information structures. To see why, consider a 2-player entry game with $\mathcal{Y}\coloneqq \{0,1\}$. Let the payoff of player $i\in \{1,2\}$ be $Y_i(X_i \beta+ \delta Y_j+\epsilon_i)$, where $Y_j$ is the competitor's action, $X_i$ represents $i$'s covariates (scalar, for simplicity), and $\epsilon_i$ captures $i$'s characteristics unobserved by the researcher. Each player $i$ is assumed to observe $(X_i, X_j, \epsilon_i)$. The researcher remains agnostic about $i$'s knowledge of $\epsilon_j$ and, hence, their ability to accurately predict $Y_j$. Note that this framework resembles our model (a game against nature) by letting player $i$ be the DM, assigning player $j$'s role to nature, and setting $Y_j\coloneqq V$. However, while in the aforementioned 2-player game player $j$ chooses $Y_j$ to maximise their own payoff, in our setting nature is indifferent to the outcomes and selects $V$ randomly from $P_V$. Additionally, this value is unobserved by the researcher. This leads to a generic reduction in the identification power of our framework compared to 2-player games, given the fewer data and model restrictions available for analysis. For example, a common way to achieve point identification of $\beta$ and the parameters governing the distribution of $(\epsilon_1,\epsilon_2)$ in the above 2-player game is to use “at infinity” arguments (Theorem 1, \hyperlink{Tamer}{Tamer, 2003}; Proposition 3, \hyperlink{Magnolfi_Roncoroni}{MR23}). These involve finding extreme values of $X_j$ that induce player $j$ always to choose one action, so that player $i$'s problem turns to a single agent parametric discrete choice problem with complete information that we know to be point identified (\hyperlink{Manski}{Manski, 1988}). Such a strategy is clearly not implementable in our setting because the model lacks assumptions regulating nature's behaviour, let alone covariates influencing nature's latent choice of $V$, which needs to be integrated out.
A key objective in the empirical analysis of single-agent decision problems is predicting how individual choices and social welfare change in counterfactual decision settings deployed in the same environment. Our framework allows us to consider two types of policy experiments, illustrated in what follows.
In the first policy experiment, we study how the choice probabilities change in response to changes in the availability of information to agents about the state of the world. Suppose the policy maker implements some intervention that urges all agents to process information structure $S$. Some agents may settle for this information structure, while others might prefer to collect {\it further} information. In this new environment, the DM selects their favourite alternative from $\mathcal{Y}$ by solving the augmented decision problem $\{G(\theta_0), S^\dagger\}$, where $S^\dagger$ is an unknown information structure that is {\it at least as} informative as $S$. $S^\dagger$ is also called an {\it expansion} of $S$. Formally, $S^\dagger\coloneqq \{\mathcal{T}^\dagger, \mathcal{P}_{T^\dagger|V}\}$ is an expansion of ${S}\coloneqq \{{\mathcal{T}}, {\mathcal{P}}_{T|V}\}$ if there exists ${S}^\diamond \coloneqq \{{\mathcal{T}}^\diamond, {\mathcal{P}}_{T^\diamond |V}\} $ such that $S^\dagger$ is the {\it combination} of $S$ and ${S}^\diamond$. That is, $\mathcal{T}^\dagger \coloneqq \mathcal{T}\times \mathcal{T}^\diamond$, $\int_{\mathcal{T}^\diamond} P_{T^\dagger|V}( t, t^\diamond |v) d t^\diamond= P_{T|V}(t|v)$ for each $t\in \mathcal{T}$ and $v\in \mathcal{V}$, and $\int_{\mathcal{T}} P_{T^\dagger|V}( t, t^\diamond |v) d t = P_{T^\diamond |V}(t^\diamond |v)$ for each $t^\diamond \in \mathcal{T}^\diamond $ and $v\in \mathcal{V}$ (Definition 5, \hyperlink{BM_2}{BM16}).
We can characterize sharp bounds on the counterfactual choice probabilities while fixing the policy-implemented information structure $S$ and remaining agnostic about $S^\diamond$ and, hence, $S^\dagger$. To do this, we use once again Theorem 1 of \hyperlink{BM_2}{BM16}, which generically applies to any 1-player game whose {\it minimal} information structure $S$ has been set by the researcher. Namely, by Theorem 1 of \hyperlink{BM_2}{BM16}, the set of optimal strategies that arise from arbitrarily expanding $S$ is equivalent to the set of 1BCE of $\{G(\theta_0), S\}$. Therefore, the collection of counterfactual choice probabilities is simply the set of 1BCE of $\{G(\theta_0), S\}$. Proposition (ref) formalises these arguments.
In Proposition (ref), ((ref))-((ref)) define a 1BCE of $\{G(\theta), S\}$. In particular, ((ref)) is the {\it Consistency} constraint, ((ref)) is the {\it Obedience} constraint, and ((ref)) ensures that $P_{Y^\dagger,V,T}$ is a proper distribution. Observe that computing $\Psi^{\ell}(h;\theta)$ and $\Psi^{u}(h;\theta)$ for a given $\theta$ requires us to solve two linear programs with respect to $P_{Y^\dagger,V,T}$. Further, Proposition (ref) presumes that the identified set, $\Theta^*$, has been constructed in a pre-step. One could also append the program of Proposition (ref) to ((ref)), solve a unique program, and thus find $ \mathcal{P}^*_{Y^\dagger}(h)\times \Theta^*$ in one step. However, this strategy would not bring notable computational advantages because the programs of Proposition (ref) and ((ref)) are not linear in $\theta$.
Understanding how information changes behaviour is a key question in the literature on single-agent decision problems (see, for example, \hyperlink{Athey}{Athey, 2002}) and is largely absent in the literature on many-player games. In particular, as discussed in Section (ref), several empirical papers are concerned with assessing the impact on choices of sending agents information about payoff-relevant variables. These papers typically exploit field experiments which can be very costly. Proposition (ref) represents a powerful result because it allows the analyst to assess the effect of programs of information provision {\it prior} to conducting such interventions. In the empirical application of Section (ref), we apply Proposition (ref) to the benchmark case where the policy intervention fully reveals the state of the world to agents, i.e., $S$ is the complete information structure and, therefore, $S^\dagger=S$ for each DM.\footnote{Any expansion $S^\dagger$ of the complete information structure $S$ is equal to $S$.}
The complete information benchmark is also useful to evaluate the welfare cost of limited information which may keep agents from choosing their first best. Specifically, given the true parameter vector $\theta_0\in \Theta$, we consider the average gains in the model-implied {\it ex-post} utility from expanding every DM's information structure from the null information structure to the complete information structure:
where the first term is the average ex-post utility when all agents process the complete information structure and the second term is the average ex-post utility when all agents process the null information structure. Observe that, in the complete information scenario, each DM earns an ex-post payoff that is, on average, greater or equal than the ex-post payoff under the null information structure. Therefore, $\Delta \mathbb{E}_{\theta_0}$ captures the maximum welfare cost of limited information. The identified set of $\Delta \mathbb{E}_{\theta_0}$ is $ \cup_{\theta\in \Theta^*} \Delta \mathbb{E}_{\theta}$.
The second policy experiment considers a more traditional, yet important, question that interests the literature on both single-agent decision problems and games. In particular, we study how the choice probabilities change in response to changes in covariates $X$ entering the utility function and observed by the researcher and the DMs. Suppose the policy maker implements some intervention which shifts the realization of $X$ assigned to agents. In this new environment, the DM processes the same information structure as in the factual scenario (i.e., differently from Proposition (ref), information structures are now held fixed), but has to account for the new realization of $X$ in evaluating their payoffs. Proposition (ref) characterises sharp bounds on the counterfactual choice probabilities.
Before presenting Proposition (ref), we introduce some useful notation and provide the intuition behind the result. Let $\mathcal{X}$ be the finite support of $X$. Let $X^\dagger$ be the vector of covariates after the intervention, with finite support $\mathcal{X}^\dagger$. To compute the counterfactual choice probabilities, imagine a hypothetical scenario where the DM simultaneously chooses an alternative from $ \mathcal{Y}$ for the augmented decision problem $\{G(\theta_0, X), S\}$ and an alternative from $ \mathcal{Y}$ for the augmented decision problem $\{G(\theta_0, X^\dagger), S\}$. The choices for $\{G(\theta_0, X), S\}$ and $\{G(\theta_0, X^\dagger), S\}$ are denoted by $Y$ and $Y^\dagger$, respectively. $Y^\dagger$ and $S$ are not observed by the researcher. There is no interaction between the two choices except for the common information structure. $\{G(\theta_0, X), G(\theta_0, X^\dagger), S\}$ is called the {\it linked augmented decision problem}. $\{G(\theta_0, X), G(\theta_0, X^\dagger)\}$ is called the {\it linked baseline decision problem}. Let $\mathcal{P}_{Y,Y^\dagger,V| X, X^\dagger}$ be a 1BCE of the linked baseline decision problem and consider the set of such 1BCEs. By Theorem (ref) above, the marginal of these distributions on $(Y^\dagger, V)$ is precisely the set of counterfactual choice probabilities that we are looking for. These steps are formalised by Theorem 1 in \hyperlink{Bergemann_Brooks_Morris}{Bergemann, Brooks, and Morris (2022)} and readapted to our case by Proposition (ref).
In proposition (ref), ((ref))-((ref)) define a 1BCE of the linked baseline decision problem $\{G(\theta, X), G(\theta, X^\dagger)\}$. In particular, ((ref)) is the {\it Consistency} constraint, ((ref)) and ((ref)) are the {\it Obedience} constraints for each decision problem, ((ref)) ensures that $P_{Y, Y^\dagger,V|x, X^\dagger}$ is a proper distribution, and ((ref)) imposes that the factual choice probabilities are equal to the empirical ones.
In the empirical application of Section (ref), we apply Proposition (ref) to study how the well-being of UK parties changes when they modify their ideological positions on popular policy issues.
We consider a model specification close to the one used in the empirical application of Section (ref). The payoff function is
where $i$ indexes a generic DM, $\mathcal{Y}\coloneqq \{0,1,\dots, D\}$, $0$ denotes the outside option and its utility is normalised to zero, $X_{iy}$ and $V_{iy}$ are DM-alternative specific features. DM $i$ observes the realization of $X_i\coloneqq (X_{i1},..., X_{iD})$ but may be uncertain about the realization of $V_i\coloneqq (V_{i1},..., V_{iD})$. $X_i$ and $V_i$ are assumed to be independent. DM $i$ has a prior on $V_i$, which is assumed to be standard Normal. The researcher observes the choice made by DM $i$ and the realization of $X_i$ for a large sample of DMs, without knowing their information structures. Note that this framework reduces to a standard multinomial Probit model under the additional assumption that each DM processes the complete information structure.
In the first simulation exercise, we illustrate how to choose the order of the Bernstein polynomials, $K$. As developing data-driven procedures to choose $K$ is still an open question in nonparametric frameworks with point identification, here we follow the heuristic approach developed by \hyperlink{Han}{Han and Yang (2023)} for partially identified settings. We simulate the data from ((ref)) with $|\mathcal{Y}|=3$, $D=2$, and $\beta=1.3$. Each $X_{iy}$ is randomly drawn from a probability mass function constructed by taking a bivariate normal and then discretising it to have support $\mathcal{X}\coloneqq \{-2.4, -0.4, 0.3\}$.\footnote{That is, $\Pr(X=x)=\frac{\exp(-(x-\mu)^2/\sigma^2)}{\sum_{x\in \mathcal{X}} \exp(-(x-\mu)^2/\sigma^2)}$ for each $x\in \mathcal{X}$.} The data is generated assuming all DMs process the complete information structure. The first column of Table (ref) reports the identified set as $K$ increases. To obtain such identified set, we explored a grid of candidate values of $\beta$ between -30 and 30 equally distanced at 0.001. The second and third columns report $K_d$, which is taken to be constant across $d=1,\dots, D$, and $K\coloneqq(K_d+1)^{D}$, respectively. For a given covariate realization $x$ and parameter value $\beta$, the fourth column computes the number of unknowns of the linear program ((ref)), which is $(D+1) K$. Observe that the width of the bounds tends to increase weakly with $K$. When $K_d=3$ ($K =64$), the bounds are narrow, but this may be due to our misspecification of the smoothness of the family of functions to be approximated, $\mathcal{P}_{Y|V}$. When $K_d$ is sufficiently large, the width increases at a slower rate, and the bounds start to converge. In particular, the bounds become stable from $K_d=10$ ($K =1,331$) onwards. Therefore, we set $K_d=10$ for $d=1,\dots, D$ in the next simulations.\footnote{When the data are generated from ((ref)) with $X_i$ exogenous and $V_i$ distributed as a standard Normal, then $\beta=0$ is expected to be part of the identified set. To see why, suppose $\beta=0$ and each DM processes the null information structure, so that the posterior is equal to the prior. Then, given the standard Normal prior, each DM gets the same expected payoff, $\mathbb{E}(V_{iy})=0$, from every alternative $y\in \mathcal{Y}$ and chooses according to some tie-breaking rule. Clearly, there will be a tie-breaking rule that reproduces the data. Hence, $\beta=0$, coupled with the null information structure for all DMs and some tie-breaking rule, generates the data. Nevertheless, the model maintains enough identification power to exclude negative values of $\beta$, as shown in Table (ref).}
The fifth column of Table (ref) shows the average CPU time to assess if ((ref)) has a solution for a given $(x,\beta)$, using the MOSEK solver for Matlab. The CPU time includes the calculation of the integral $\gamma^{y,y',x}_{1,k,K}(\beta)\coloneqq \int_{\mathcal{V}} a_{k,K}(v)P_V(v)(u(y,x, v;\beta)-u(y', x, v; \beta)) dv$ for each $k\in \mathcal{K}\coloneqq \times_{d=1}^D \{0,1,\dots, K_d\}$ and $y,y'\in \mathcal{Y}$.\footnote{Recall that $P_V$ is assumed to be standard normal, with $V_i$ independent of $X_i$. Hence, $\gamma^{y,y'}_{2,k,K}\coloneqq \int_{\mathcal{V}} a_{k,K}(v) P_v(v)dv$ does not vary across values of the covariates and parameters and can be computed once.} These integrals are computed by Monte Carlo integration taking $10^4$ random draws from the standard Normal distribution. The number of unknowns of ((ref)) increases linearly in $K$ and exponentially in $K_d$. Hence, the CPU time increases approximately linearly in $K$ and exponentially in $K_d$. The total time to construct the identified set depends on the possibility of parallelising across $(x,\beta)$. Using a computing cluster, we could exploit 600 parallel workers by coding our procedure as an SGE array job. Based on those, the rough total CPU time is listed in the last column of Table (ref).
In the second simulation exercise, we investigate how the identified set varies as the cardinality of the support of $X_i$ increases. We generate data from ((ref)) with $|\mathcal{Y}|=3$, $D=2$, $\beta=1.3$, and all DMs processing the complete information structure. We distinguish two scenarios regarding the dependence between $X_i$ and $V_i$. In the first scenario, $X_i$ and $V_i$ are independent and the prior on $V_i$ is a standard normal, as in the first simulation exercise. In the second scenario, $X_i$ and $V_i$ are allowed to be correlated and the prior on $V_i$ conditional on $X_i=x$ is a normal distribution with mean and variance that vary with $x$. For each scenario, we study three cases. In the first case, each $X_{iy}$ is randomly drawn from a probability mass function constructed by taking a bivariate normal and then discretising it to have support $\mathcal{X}\coloneqq\{-2.4, -0.4, 0.3\}$, as in the first simulation exercise. In the second case, $\mathcal{X}\coloneqq\{-2.4, -0.5, -0.4, 0.1, 0.3\}$. In the third case, $\mathcal{X}\coloneqq\{-2.4, -0.7, -0.5, -0.4, 0.1, 0.2, 0.3\}$. We set $K_d=10$ for $d=1,\dots, D$ to construct the identified set. For each $\mathcal{X}$, the second and third columns of Table (ref) show the identified set in the first and second scenarios, respectively. When $X_i$ is exogenous, we find that the identified set shrinks as the cardinality of $\mathcal{X}$ increases. This is because variation in exogenous covariates induces variation in agents' choices, which helps the identification of $\beta$, as in standard parametric analysis. Conversely, when $X_i$ and $V_i$ are allowed to be correlated, the identified set becomes larger as the cardinality of $\mathcal{X}$ increases. This is because, as the cardinality of $\mathcal{X}$ increases, there are smaller groups of DMs with the same prior, which worsens the identification of $\beta$ by “weakening” the {\it Consistency} requirement of 1BCE.\footnote{When the data are generated from ((ref)) and the prior varies across realizations of $X_i$, $\beta=0$ may not be part of the identified set. To see why the logic of Footnote (ref) does not apply, suppose $\beta=0$ and each DM processes the null information structure, so that the posterior is equal to the prior. If $X_i=x$, then DM $i$ gets an expected payoff equal to $\mathbb{E}(V_{iy}|X_i=x)\coloneqq \mu_{y,x}$ from every alternative $y\in \mathcal{Y}\setminus \{0\}$, and a payoff equal to 0 from the outside option. Hence, the DMs' choices will differ depending on the realization of the covariates and the tie-breaking rule adopted. In turn, the resulting distribution of choices may not coincide with the empirical one and, so, $\beta=0$, may not belong to the identified set.}
In the third simulation exercise, we investigate how the identified set changes as the information structures processed by DMs in the underlying data generating process vary. We generate data from ((ref)) with $|\mathcal{Y}|=3$, $D=2$, $\beta=1.3$, $X_i$ independent of $V_i$, each $X_{iy}$ randomly drawn from a probability mass function constructed by taking a bivariate normal and then discretising it to have support $\mathcal{X}\coloneqq\{-2.4, -0.4, 0.3\}$, and $V_i$ distributed as a standard Normal. We set $K_d=10$ for $d=1,2$ to construct the identified set. We consider three scenarios. In the first scenario, all DMs process the complete information structure, as in the first simulation exercise. In the second scenario, 5/6 of the DMs process the complete information structure, and 1/6 process the null information structure. In the third scenario, all DMs process the null information structure. For each scenario, Table (ref) shows the identified set (second column) and the value of $\beta$ that would be identified if the researcher assumed that all DMs process the complete information structure (third column). Assuming that all DMs process the complete information structure, as in a standard multinomial Probit model, leads to recovering one parameter value contained in the identified set. When this assumption is misspecified (second and third scenario), the recovered parameter value differs from the truth. The model has the least identifying power in the third scenario, when all DMs process the null information structure. As soon as a significant proportion of DMs process the complete information structure (first and second scenarios), the identifying power of the model improves. Note, in fact, that, under the first scenario, the DMs take their decisions based on the actual payoffs. Instead, under the third scenario, the DMs choose based on the expected payoffs which are relatively homogenous in the population because computed using the same posterior. Hence, under the third scenario, there is less variation in the DMs' choices, which leads to wider bounds.
In this section, we use our methodology to study the determinants of voting behaviour during the UK general election held on 8 June 2017 and perform some counterfactual exercises aiming to evaluate the impact of uncertainty on the well-being of voters and parties.
The spatial model of voting is a dominant framework in political economy to explain individual preferences for parties and, in turn, how such preferences shape the policies implemented by democratic societies (\hyperlink{Downs}{Downs, 1957}; \hyperlink{Black}{Black, 1958}). This model posits that an agent has a most preferred policy (also called “bliss point”) and casts their vote in favour of the party whose position is closest to their ideal (i.e., she votes “ideologically”). In empirical analysis, it is typically implemented by estimating a classical parametric discrete choice model with perfect information. That is, it is assumed that each DM $i$ processes the complete information structure and votes for party $y\in \mathcal{Y}$ maximising their utility,
where $Z_{iy}\coloneqq|Z_{i}-Z_{y}|$ is an $M\times 1$ vector observed by the researcher, representing the distance between DM $i$'s opinion ($Z_i$) and party $y$'s opinion ($Z_y$) on $M$ issues, as measured in some common $M$-dimensional ideological metric space. $W_i$ is a vector of individual-specific covariates observed by the researcher. $X_i\coloneqq(W_i, Z_{iy}: y \in \mathcal{Y})$ collects the ideological distances and the individual-specific covariates. $V_i \coloneqq (V_{iy}: y \in \mathcal{Y})$ is a vector of tastes of DM $i$ for each party/candidate that is unknown to the researcher and independent of $X_i$, whose distribution belongs to a parametric family, thereby outlining a (Multinomial) Logit model, Probit model, Nested Logit model, etc. If voters vote ideologically, then each $\beta_y$ is expected to be negative so that DM $i$'s utility declines with increasing distance between $Z_i$ and $Z_y$. $\theta_u\coloneqq (\beta_y, \gamma_y: y \in \mathcal{Y})$ is the vector of payoff parameters.
The above framework is scientifically appealing because of its elegance and simplicity but it has limitations. Importantly, uncertainty affects voting (\hyperlink{Tabellini2}{Mat\u{e}jka and Tabellini, 2021}, and other references in Section (ref)). That is, voters may be unsure about their own and the parties' ideological positions and, more generally, about the qualities of the candidates. This is because of the inevitable difficulty of making precise political judgments and understanding associated returns, or because the parties deliberately obfuscate information to attract voters with different preferences and expand electoral support. More plausibly, in the wake of election campaigns, voters are conscious of their own and the parties' attitudes towards some popular issues, but might be uncertain about how they themselves and the parties stand towards more technical or less debated topics, and about the traits of the candidates other than those publicly advertised. Further, they may attempt to fill such gaps in information with various degrees of success and in different ways, depending on a priori inclination for certain parties, political sentiments, interest in specific issues, civic sense, attentional limits, participation in townhall debates, candidates' transparency, opinion makers, and media exposure. In turn, some individuals might become much more informed, others less, giving rise to heterogeneity in the public understanding of politics.
Despite the acknowledgement of the central role played by the sophistication of voters in determining voting patterns, only a few empirical works have attempted to take it into account while estimating a spatial voting framework. This has been done, for instance, through an additive, exogenous, and parametrically distributed error in the payoffs representing the evaluation mistakes made by voters, a parametric specification of the variance of the perceived party position across voters, or a parametric specification of the probability of being informed versus uninformed when voting (\hyperlink{Degan_Merlo2}{Degan and Merlo, 2011}, and other references in Section (ref)). By contrast, our methodology permits us to incorporate uncertainty under weak assumptions on the latent, heterogeneous, and potentially endogenous process followed by voters to gather and evaluate information.
In particular, we assume that, when assessing the returns to voting for party $y\in \mathcal{Y}$, DM $i$ is:
We follow the literature on voting under uncertainty which typically models priors as normal distributions (\hyperlink{Knight}{Knight and Schiff, 2010}; \hyperlink{Tabellini2}{Mat\u{e}jka and Tabellini, 2021}; \hyperlink{Yuksel}{Yuksel, 2022}). In particular, we assume that $V_i$ is distributed as a standard normal, independent of $X_i$. See also \hyperlink{Feddersen}{Feddersen and Pesendorfer (1997)} and \hyperlink{McMurray}{McMurray (2013)} on the use of the common prior assumption and Bayesian rationality in models of voting. Lastly, we consider abstention as the base category and normalise its payoff to zero as, for example, in \hyperlink{Knight}{Knight and Schiff (2010)}.
We estimate our model by using data on the UK general election held on 8 June 2017. Specifically, we use data from the British Election Study, 2017: Face-to-Face Post-Election Survey (\hyperlink{Fieldhouse}{Fieldhouse, et al., 2018}). The survey took place immediately after the election. It asks questions concerning key contemporary problems about political representation, accountability, and engagement, and aims to explain changes in party support. The interviewees constitutes an address-based random probability sample of eligible voters living in 468 wards in 234 Parliamentary Constituencies across England, Scotland, and Wales.
We believe that such data fit the framework described in Section (ref) for four main reasons. First, the UK parties were clearly focused on the topic of Brexit, along with issues of public health and austerity, thus inducing potential uncertainty among voters with respect to many other factors (\hyperlink{Hutton}{Hutton, 2017}; \hyperlink{Snowdon}{Snowdon and Demianyk, 2017}). Second, the UK political scene is dominated by historical parties. Hence, past election outcomes and consequent behaviour of parties can justify the common prior assumption on $V_i$. Third, the survey reports the positions of the respondents on topics that were debated at length before the election, which we discuss more precisely below. The survey also asks respondents to state the parties's positions with respect to those topics and the answers provided are substantially aligned. This suggests that there was no uncertainty among voters on those topics. Hence, they can be used to construct the vector $Z_{iy}$ for each party $y \in \mathcal{Y}$, whose realization is assumed to be in the information set of voters. The survey does not contain data on other relevant factors that might have induced uncertainty among voters. Hence, it is natural to treat the realization of $V_i$ as unobserved by the researcher. Fourth, the survey asks respondents to declare if they voted tactically. Only $2.16\%$ of the respondents answer affirmatively. We drop them from our final sample, in order for the assumption that voters vote ideologically to apply.
To limit the impact of Scottish and Welsh independentist fronts on our results, we focus on the respondents who reside in England. We consider the answers of respondents on which party they have voted for among the Conservative Party, Labour Party, Liberal Democrats, United Kingdom Independence Party (UKIP), Green Party, and none.
We collect in $Z_{iy}$ the distances between DM $i$'s position and party $y$'s position on four dimensions: EU integration, taxation and social care, income inequality, and left-right political orientation. More precisely, we select the answers of the respondents to the following questions (summarised with respect to the original version, for brevity):
Following the literature (\hyperlink{Alvarez1995}{Alvarez and Nagler, 1995}; \hyperlink{Alvarez1998}{1998}; \hyperlink{Alvarez2000}{2000}; \hyperlink{Alvarez2000_2}{Alvarez, Nagler, and Bowler, 2000}), we set party $y$'s position on dimensions 1-4 equal to the median placement of the party on each dimension across the sample, although as noticed above there is substantial alignment among the respondents' answers.
We collect in $W_i$ some demographic characteristics of respondents. In particular, we focus on gender, socio-economic class, and total income before tax. Recall that $\gamma_y$ captures the impact of $W_i$ on the vote shares. We allow this impact to be heterogeneous across the parties. To be parsimonious on the number of parameters to estimate, we further parameterise $\gamma_y$ by requiring that $\gamma_y\coloneqq\gamma Z^{\text{LR}}_y$ for every party $y$, where $Z^{\text{LR}}_y$ is the position of party $y$ with regards to left-right orientation. In other words, we assume that the aforementioned heterogeneity is driven by the position of each party in the left-right political spectrum. Similarly, to reduce dimensionality, we impose $\beta_y\coloneqq \beta$ for each party $y$.
In our final sample, $36.48\%$ of people have voted for the Labour Party, $36.65\%$ for the Conservative Party, $6.41\%$ for the Liberal Democrats, $1.73\%$ for UKIP, $1.56\%$ for the Green Party, and $17.17\%$ did not vote. Table (ref) presents some descriptive statistics. The second column refers to the positions of the respondents on dimensions 1-4 and reports the mean (rounded to the nearest integer), median, and standard deviation across the sample. The remaining columns reports $Z_y$ for each party $y$. As expected, the Conservative Party and UKIP are more right-wing, less concerned with income inequality, more Eurosceptic, and stronger supporters of low taxes and a minimal welfare state, than the Labour Party and the Green Party. The Liberal Democrats are more centrist.
The sample is gender balanced, with $48.97\%$ of males and $51.03\%$ of females. We assign label $1$ to females and $0$ to males. In the original data, the socio-economic class is divided into seven categories, following the Standard Occupation Classification 2010: professional occupations; managerial and technical occupations; skilled occupations - non-manual; skilled occupations - manual; partly skilled occupations; unskilled occupations; armed forces. To lessen the computational burden, we reorganise these categories into three groups. The first group is assigned label $0$ and collects professional occupations, managerial and technical occupations, skilled occupations - non-manual, and armed forces ($68.04\%$ of the sample). The second group is assigned label $1$ and collects skilled occupations - manual and partly skilled occupations ($29.42\%$ of the sample). The third group is assigned label $2$ and collects unskilled occupations ($2.54\%$ of the sample). Similarly, in the original data, the total income before tax is bracketed into $14$ categories. We reorganise these categories into four groups, which we construct by approximately following the UK income tax rates. The first group is for income between $\textsterling 0$ and $\textsterling15,599$ ($21.78\%$ of the sample). The second group is for income between $\textsterling15,600$ and $\textsterling49,999$ ($51.68\%$ of the sample). The third group is for income between $\textsterling50,000$ and $ \textsterling99,999$ ($21.45\%$ of the sample). The fourth group is for income above $\textsterling 100,000$ ($5.09\%$ of the sample). To each of the four groups, we assign as value the logarithm of the median income across the respondents belonging to that group ($9.4727$, $10.4282$, $11.1199$, and $12.6115$, respectively). We summarise these numbers in Table (ref).\\
We estimate the identified set of the $1\times 7$ vector $\theta_0 \coloneqq (\beta_0, \gamma_0)$ by evaluating a grid of 100,000 parameter values. The grid is constructed by exploring the parameter space, $\Theta\subseteq \mathbb{R}^7$, via the simulated annealing algorithm.\footnote{In the case of a high-dimensional vector of parameters, it is common in the partial identification literature to construct the grid of parameter values by using the simulated annealing algorithm (\hyperlink{CT}{Ciliberto and Tamer, 2009}). In particular, we proceed in four steps. First, we estimate $\theta_0$ by maximum likelihood under the assumption that all DMs process the complete information structure and obtain the estimate $\hat{\theta}^\text{com}$. Second, we construct an Halton set of $10^6$ points around $\hat{\theta}^\text{com}$. We draw 100 points at random from this set. We stack these points, together with $\hat{\theta}^\text{com}$, in an $101\times 7$ matrix, $A$. Third, we minimise the test statistic $\text{TS}_n(\theta)$ defined in Appendix (ref) with respect to $\theta$ by running the simulated annealing algorithm from each row of $A$ as starting point and experimenting at different temperatures. We save every parameter value encountered in the course of the algorithm. We stack all the saved parameter values in a matrix $G$. Fourth, we draw 100,000 rows at random from $G$. Such 100,000 rows constitute our final grid of candidate parameter values.}
To decide the order of the Bernstein polynomials, we take $K_d$ constant across $d=1,\dots, D$ and consider $K_d=3,5,7,10$. $D=5$ because there are five parties and the payoff from abstention is normalised to zero. For each of these four values of $K_d$, we construct an estimate of $\Theta^*$, $\widehat{\Theta}^*$, by replacing the empirical choice probabilities in ((ref)) with their sample analogues and solving ((ref)) for every realization $x$ of $X_i$ and parameter value $\theta$ in the grid. Table (ref) reports the projections of $\widehat{\Theta}^*$. As seen in Section (ref), the width of the bounds tends to increase weakly with $K_d$. The bounds become approximately stable from $K_d=5$ onwards. Therefore, in our next computations, we set $K_d=5$ for $d=1,\dots, D$.
Table (ref) provides some computational details, as in Table (ref) of Section (ref). In particular, the fourth column shows the average CPU time to assess if ((ref)) has a solution for a given $(x,\theta)$, using the MOSEK solver for Matlab. The CPU time includes the calculation of the integrals $\{\gamma^{y,y',x}_{1,k,K}(\theta)\}$ by Monte Carlo integration, taking $10^2$ random draws from the standard Normal distribution. The last column of Table (ref) reports the rough total CPU time to construct $\widehat{\Theta}^*$ based on exploiting 800 parallel workers from a computing cluster.
Table (ref) presents the inference results. In particular, the second and third columns report the maximum likelihood estimate of $\theta_0$ ($\hat{\theta}^\text{com}$) and 95% confidence intervals ($C^\text{com}_{0.95}$), respectively, under the assumption that all DMs process the complete information structure. The third and fourth columns report the projections of $\widehat{\Theta}^*$ and of the 95% confidence region for any $\theta\in \Theta^*$ ($C_{0.95}$), respectively. $C_{0.95}$ is constructed following \hyperlink{Andrews_Shi}{Andrews and Shi (2013)}, as outlined in Appendix (ref), based on 50 bootstrap samples.
Under the assumption that voters are fully informed, all the $\beta$ coefficients, except $\beta_2$, are statistically different from zero at $5\%$. This suggests that DMs vote ideologically on the EU, inequality, and left-right dimensions. That is, the smaller the distance between DM $i$ and party $y$'s ideological positions on those dimensions, the more likely DM $i$ votes for party $y$, ceteris paribus. Further, $\beta_4$ has the highest absolute value magnitude among the $\beta$ coefficients. That is, voters particularly disvalue casting their votes in favour of a party ideologically distant on the left-right axis. A one-unit increase in the ideological distance on the left-right axis produces a payoff decrease that is roughly 2, 23, and 4 times bigger than the payoff decrease produced by a one-unit increase in the ideological distance on the EU, social care, and inequality dimensions, respectively.
When we remain agnostic about voter sophistication, all the projections of $\widehat{\Theta}^*$ for the $\beta$ coefficients include zero. Therefore, differently from above, we cannot reject the possibility that the election outcomes have been generated under some combinations of information structures under which the ideological distances on the EU, social care, inequality, and left-right dimensions are irrelevant for voter preferences. Nevertheless, the model maintains enough identification power to completely exclude positive values of $\beta_3$ and almost entirely positive values of $\beta_4$. This is roughly confirmed by the projections of $C_{0.95}$. In particular, $\beta_4$ can have the highest absolute value magnitude among the $\beta$ coefficients, in agreement with the maximum likelihood results. The fact that this finding on $\beta_4$ is robust to the restrictions on the information environment reflects several post-election descriptive studies run by political experts, which emphasise that the traditional left-right values, rather than specific policy issues, have been the main driver of the British electoral behaviour in 2017 (\hyperlink{Hobolt}{Hobolt, 2018}).
The upper bounds of the projections for $\beta_1$ and $\beta_2$ are non-negligibly positive. While this aligns with the non-significance of $\beta_2$ under the complete information assumption ($C^\text{com}_{0.95}$ includes positive values), it is in contrast with the significance of $\beta_1$ under the complete information assumption and the 2017 election being often referred to as the “Brexit election” (\hyperlink{Mellon}{Mellon, et al., 2018}). Upon closer inspection, however, the inconclusive results on $\beta_1$ reflect that the 2017 pre-election period saw a substantial increase in the relationship between EU referendum choice and Labour versus Conservative vote choice, with a sort of alignment of the remain-leave axis with the traditional left-right axis. The parties with the clearest positions against Brexit (the Liberal Democrats) and in favour of Brexit (UKIP) lost many supporters. These switched {\it en masse} to the Labour Party, offering a “soft Brexit” and the Conservative Party, offering a “hard Brexit”, respectively (\hyperlink{Mellon}{Mellon, et al., 2018}; \hyperlink{Heath}{Heath and Goodwin, 2017}). Such a tendency may have dampened the role of the distance in the Brexit sentiment in determining the preferences of voters, an insight that is not picked up under the complete information assumption.
\paragraph{Information provision.} The uncertainty about the payoffs resulting from voting can occur due to deliberate strategies of the candidates who “{\it becloud}” their characteristics and opinions “{\it in a fog of ambiguity}” (\hyperlink{Downs}{Downs, 1957}, p.136), in order to expand the electoral support by attracting groups of voters with different political preferences (\hyperlink{Campbell}{Campbell, 1983}; \hyperlink{Dahlberg}{Dahlberg, 2009}; \hyperlink{Tomz}{Tomz and van Houweling, 2009}; \hyperlink{Somer}{Somer-Topcu, 2015}). It remains unclear, however, to what extent such uncertainty affects the vote shares and, in turn, the election results. A better understanding is important for designing transparency policies that can improve citizens' welfare and parties' well-being. We investigate this question by imagining an omniscient mediator who implements a policy that gives voters complete information. This can be achieved, for instance, by organising school campaigns that develops political literacy; forcing candidates to publicly disclose their assets, liabilities, and criminal records; and enforcing a strict regulation regarding campaign spending and airtime.\footnote{See, for example, \hyperlink{Niemi}{Niemi and Junn (1998)}, \hyperlink{Hooghe}{Hooghe and Wilkenfeld (2007)}, and \hyperlink{Pontes}{Pontes, Henn, and Griffiths (2019)} on the impact of civic education on political engagement.} We simulate the counterfactual vote shares under complete information and study how they change compared to the factual scenario.
This question has been largely debated in the literature. As explained by \hyperlink{Bartels2}{Bartels (1996)}, political scientists have often answered it by arguing that a large population composed of possibly uninformed citizens acts as if it was fully informed, either because each voter uses cues and information shortcuts helping them to figure out what she needs to know about the political world; or because individual deviations from fully informed voting cancel out in a large election, producing the same aggregate election outcome as if voters were fully informed. \hyperlink{Carpini}{Carpini and Keeter (1996)} and \hyperlink{Bartels2}{Bartels (1996)} are the first studies to use quantitative evidence to disconfirm such claims. They simulate counterfactual vote shares under complete information using data on the level of information of the survey respondents as rated by the interviewers or assessed by test items. \hyperlink{Degan_Merlo2}{Degan and Merlo (2011)} propose an alternative approach, which is closer to ours. They consider a spatial model of voting with latent uncertainty. Differently from us, they estimate such a model by parametrically and exogenously specifying the probability that a voter is informed. They use their estimates to obtain counterfactual vote shares under complete information and find that making citizens more informed about electoral candidates decreases abstention. We contribute to this thread of the literature by providing a way to construct counterfactual vote shares under complete information, which neither requires the difficult task of measuring voters' level of information in the factual scenario, nor imposes parametric assumptions on the probability that a voter is informed.
Recall from Proposition (ref) that, given a realization $x$ of $X_i$ and a parameter value $\theta$, $\Psi^{\ell}(y|x;\theta)$ and $\Psi^{\ell}(y|x;\theta)$ are the minimum and maximum counterfactual probabilities of choosing alternative $y\in \mathcal{Y}$ when the DMs face the augmented decision problem $\{G(\theta,x), S^\dagger\}$, where $S^\dagger$ is some unknown expansion of the information structure $S$ implemented by the policy program, possibly heterogeneous across DMs. In the case analysed, $S$ is the complete information structure and, hence, $S^\dagger=S$ for each DM. Moreover, observe that, when all DMs process the complete information structure, there is a unique model-implied choice distribution. Therefore, $\Psi^{\ell}(y|x;\theta)=\Psi^{u}(y|x;\theta)\coloneqq \Psi(y|x;\theta)$. In light of all this, we compute $$ \widehat{\Delta}^*_y\coloneqq \cup_{\theta \in \widehat{\Theta}^*} \sum_{x}(\Psi(y|x;\theta)-\widehat{\mathbb{P}}_Y(y|x))\widehat{\mathbb{P}}_X(x) \quad \text{and} \quad \widehat{\Delta}^*_{y,0.95} \coloneqq \cup_{\theta \in C_{0.95}} \sum_{x}(\Psi(y|x;\theta)-\widehat{\mathbb{P}}_Y(y|x))\widehat{\mathbb{P}}_X(x), $$ for each $y\in \mathcal{Y}$, where $\widehat{\mathbb{P}}_Y(y|x)$ and $\widehat{\mathbb{P}}_X(x)$ are the sample probability of choosing $y$ conditional on $x$ and the sample probability of $x$, respectively. $\widehat{\Delta}^*_y$ and $\widehat{\Delta}^*_{y,0.95}$ are the estimated gains/losses in vote shares under complete information compared to the factual scenario.
Table (ref) reveals that, when voters are fully informed, abstention drops with respect to the factual scenario. This shows that voters are more confident in choosing a party and aligns with the empirical results in \hyperlink{Degan_Merlo2}{Degan and Merlo (2011)}. We also find that the “losers” from the policy intervention are the two biggest parties, i.e., the Conservative Party and the Labour Party. Conversely, the “winners” from the policy intervention are the other minor parties, i.e., the Liberal Democrats and the Green Party. This suggests that there exists some payoff-relevant information unobserved by voters, and the historically dominant parties in the British political scene benefit the most from such uncertainty.\footnote{The observed drop in abstention is not a mechanical effect of the normalisation to zero of the payoff from not voting. This is because the observed realization of $V_{iy}$ could add to or subtract from the component of the payoff “observed pre-signal”, $\beta_y^\top Z_{iy} + \gamma_y^\top W_i$. Further, note that it could be that many/all voters in the population already observe the realization of $V_i$, in which case we should expect no significant change in the abstention share, regardless of the normalisation adopted.}
Lastly, we quantify the voters' maximum welfare cost of limited information, based on ((ref)). In particular, we compute $$ \widehat{W}^*\coloneqq \cup_{\theta \in \widehat{\Theta}^*} \Delta \mathbb{E}_{\theta} \quad \text{and} \quad \widehat{W}^*_{0.95}\coloneqq \cup_{\theta \in C_{0.95}} \Delta \mathbb{E}_{\theta}, $$ where $$
$$ Table \ref{count1_table2} highlights that up to 1.1697 utility points would be gained, on average, if all voters were perfectly informed. We interpret this utility increase by translating it into the corresponding change in a specific covariate. For example, consider $Z_{i4}$ and take the estimated lower bound of its coefficient, $\beta_4$, which is $ -0.3948$, as shown in Table \ref{res2}. A utility increase of 1.1697 is equivalent to the utility gain achieved by reducing the ideological distance from a given party on the left-right dimension by approximately $1.1697/0.3948\approx 3$ points.
\paragraph{Changes in covariates.} Various political experts sustain that, while at the beginning of the 2017 election campaign the Conservative Party had a sizeable lead in the opinion polls over the Labour Party, as the campaign progressed the Labour Party recovered ground because it strengthened its left ideological position on social spending and nationalization of key public services (for example, \hyperlink{Heath}{Heath and Goodwin, 2017}; \hyperlink{Mellon}{Mellon, at al., 2018}). To evaluate this, we reset the Labour Party's placement on dimension 2 (social care) to be two points less (i.e., $5$ instead of $7$) and study if the Labour Party's well-being worsens by looking at the change in its vote shares.
Recall from Proposition (ref) that, given a realization $x$ of $X_i$ and a parameter value $\theta$, $\Phi^{\ell}(y|x, x^\dagger;\theta)$ and $\Phi^{u}(y|x, x^\dagger;\theta)$ are the minimum and maximum counterfactual probabilities of choosing alternative $y\in \mathcal{Y}$ when the DMs face the augmented decision problem $\{G(\theta,x^\dagger), S\}$, where $S$ is the unknown information structure $S$ processed in the factual scenario, possibly heterogeneous across DMs, and $x$ and $x^\dagger$ are the covariate realizations before and after the intervention, respectively. Given $y$ denoting the Labour Party, we compute $$ \widehat{\Lambda}_y\coloneqq \cup_{\theta \in \widehat{\Theta}^*} \Big[ \sum_{(x, x^\dagger)} (\Phi^{\ell}(y| x, x^\dagger; \theta)-\widehat{\mathbb{P}}_Y(y|x))\widehat{\mathbb{P}}_X(x), \sum_{(x, x^\dagger)} (\Phi^{u}(y| x, x^\dagger; \theta)-\widehat{\mathbb{P}}_Y(y|x))\widehat{\mathbb{P}}_X(x)\Big], $$ and $$ \widehat{\Lambda}_{y,0.95}\coloneqq \cup_{\theta \in C_{0.95}} \Big[ \sum_{(x, x^\dagger)} (\Phi^{\ell}(y| x, x^\dagger; \theta)-\widehat{\mathbb{P}}_Y(y|x))\widehat{\mathbb{P}}_X(x), \sum_{(x, x^\dagger)} (\Phi^{u}(y| x, x^\dagger; \theta)-\widehat{\mathbb{P}}_Y(y|x))\widehat{\mathbb{P}}_X(x)\Big]. $$ $\widehat{\Lambda}^*_y$ and $\widehat{\Lambda}^*_{y,0.95}$ are the estimated gains/losses in the Labour Party's vote share compared to the factual scenario. We also calculate the difference between the counterfactual and factual choice probabilities under the assumption that all voters process the complete information structure: $$ \widehat{\Lambda}^{\text{com}}_y\coloneqq \sum_{x} (P_Y(y| x^\dagger; \hat{\theta}^{\text{com}}) -\widehat{\mathbb{P}}_Y(y|x))\widehat{\mathbb{P}}_X(x), $$ where $P_Y(y| x^\dagger; \hat{\theta}^{\text{com}})$ is the counterfactual probability of choosing $y$ conditional on $x^\dagger$ under complete information and $\theta=\hat{\theta}^{\text{com}}$, obtained using the standard Multivariate Probit formulas.
$\widehat{\Lambda}^{\text{com}}_y$ in Table (ref) reveals that, when voters are assumed to be fully informed, weakening the social care position to 5 leads to a decrease in the Labour Party's vote share. Our results partly confirm this finding, as $\widehat{\Lambda}_y$ and $\widehat{\Lambda}_{y,0.95}$ mostly lie on the negative real line. This supports the claim that, by strengthening its left ideological position on the social care dimension, the Labour Party gained some votes during the election campaign.
In this paper, we study identification of preferences in static single-agent discrete choice models where decision makers may be imperfectly informed about the state of the world. We leverage the notion of 1BCE by \hyperlink{BM_2}{BM16} to provide a tractable characterization of the sharp identified set. We make three main methodological contributions. First, by reinterpreting our framework as a 1-player game against nature, we provide insightful comparisons between the identification power of our framework and several many-player games. Second, we develop a formal procedure to practically construct the sharp identified set for the payoff parameters when the state of the world is continuous. Third, we characterize sharp bounds on the counterfactual choice probabilities when agents receive information about the state of the world via a policy program, which is an important question in the empirical literature on single-agent decision problems across different fields. The method developed in this paper is used to estimate a spatial voting model under weak assumptions on agents' information about the returns to voting for the 2017 UK general election. We show the usefulness of our methodology to quantify the consequences of imperfect information on the welfare of voters and parties.