Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
195,729 characters · 10 sections · 64 citation commands
Revealed Information
\pagenumbering{gobble}
JEL codes: D44, D82, D83\\ Keywords: revealed preference, revealed information, distributions with given marginals, support function, stochastic choice out of menus, state-dependent stochastic choice, information design, Bayesian persuasion, flows in networks, optimal transport \pagenumbering{arabic}
When economic agents make decisions under uncertainty, they rely on information about unknown factors. Yet, researchers rarely observe this information directly, creating a fundamental challenge: How can we infer what agents know from their observed choices? This challenge has significant implications for economic modeling, because assumptions about information can dramatically impact model predictions and parameter estimates.
In this paper, we provide a framework to determine when observed choice patterns can be rationalized by some information structure, without requiring the researcher to observe the relationship between choices and underlying states. Specifically, given a decision maker's (DM) utility function, prior beliefs, and an observed distribution of actions, we characterize when this action distribution can be explained as the result of the DM\ optimally responding to some information about the state. Our main contribution is a support-function characterization that translates this question into a finite system of inequalities involving the utility function, prior, and action distribution.
Consider a concrete example: an analyst studying whether a judge's bail decisions are informed by recidivism risk rambachan2022identifying. The analyst observes only the frequency with which the judge grants bail, not the frequency with which the judge grants bail conditional on whether the defendant will recidivate. Our framework allows the analyst to determine which combinations of the judge's utility function and prior beliefs about recidivism would make the observed bail decisions consistent with the judge having some information about recidivism risk.
Many recent empirical studies seek preference estimates robust to informational assumptions, recognizing how strongly these assumptions affect outcomes. For instance, dickstein2018exporters, dickstein2024patient, gualdani2019identification, and rambachan2022identifying develop methods to understand the role of information in firms' export decisions, physicians' treatment recommendations, voter choices, and prediction mistakes, respectively. Similar approaches appear in multi-agent settings such as auctions \citep*{syrgkanis2017inference} and entry games magnolfi2019estimation.
Some of these approaches rely on Bayes correlated equilibrium (\ensuremath{\mathrm{BCE}}), developed by bergemann2016bayes for games and kamenica2011bayesian for single-agent settings. Given a payoff structure\textemdash players' utility functions and their common prior over states\textemdash\ensuremath{\mathrm{BCE}}\ provides conditions under which an outcome distribution can be rationalized as if players had access to information before playing. Importantly, checking whether an outcome is a \ensuremath{\mathrm{BCE}}\ requires verifying only a finite system of linear inequalities.
However, a critical gap exists between what \ensuremath{\mathrm{BCE}}\ requires and the data empirical researchers typically have available. \ensuremath{\mathrm{BCE}}\ presumes the analyst observes the joint distribution over payoff-relevant states and action profiles. In our judge example, \ensuremath{\mathrm{BCE}}\ would require observing bail decisions conditional on whether defendants would recidivate\textemdash data that are rarely available. In practice, analysts typically observe only the marginal distribution of actions (the frequency of bail grants), not the action distribution conditional on the state (the frequency of bail grants conditional on recidivism risk). For a given payoff structure, rather than checking finitely many linear inequalities, the analyst checks whether a \ensuremath{\mathrm{BCE}}\ exists whose marginals over the actions matches the observed choices.
Our paper bridges this gap in single-agent settings by characterizing when marginal distributions over states and actions are consistent with a (single-agent) \ensuremath{\mathrm{BCE}}\ given the DM's utility function. Formally, given the triple of a utility function, prior beliefs, and observed action distribution, we study when a \ensuremath{\mathrm{BCE}}\ exists whose marginals over the states and actions coincide with the DM's prior and action distribution. When such a \ensuremath{\mathrm{BCE}}\ exists, we say the marginals\textemdash the DM's prior and the observed action distribution\textemdash are \ensuremath{\mathrm{BCE}}-consistent given the utility function.
Our main contributions are in Sections (ref) and (ref). In (ref), we provide a characterization in terms of a system of finitely many inequalities, linear in both marginal distributions, such that the marginals are \ensuremath{\mathrm{BCE}}-consistent if and only if these inequalities hold. For a given action distribution and utility function, these inequalities characterize the support function of the set of priors that make the DM's choices consistent with information. Although (ref) precisely identifies the finitely many inequalities that must hold for the marginals to be \ensuremath{\mathrm{BCE}}-consistent, the characterization is rather implicit in that it does not describe them in closed form. Our remaining results consider assumptions on the cardinality of the set of states or on the utility function under which we provide closed-form characterizations of the system of finitely many inequalities in (ref).
(ref) provides a closed-form characterization when the state space has at most three elements. As we explain in the main text, the inequalities in (ref) are always a subset of those in (ref), and they can be readily expressed in terms of primitives even when the state space has more than three elements. Thus, they can be used to rule out pairs of prior beliefs and action distributions that are not \ensuremath{\mathrm{BCE}}-consistent given the utility function.
Theorems (ref) and (ref) offer parallel characterizations for utility functions with affine and two-step differences (Definitions (ref) and (ref)), respectively. For instance, all binary decision problems and the discrete analog of quadratic loss have affine differences, whereas the discrete analog of absolute error loss has two-step differences. Both results rest on (ref), which narrows the search for halfspaces defining the set of \ensuremath{\mathrm{BCE}}-consistent distributions for utility functions satisfying increasing differences and concavity in actions (the discrete counterpart to first-order approach conditions). In (ref), we further generalize this characterization to compact Polish action and state spaces under the first‑order approach (cf. \citealp*{kolotilin2023persuasion}).
In (ref), we apply our results to study comparative statics and \ensuremath{\mathrm{BCE}}-consistency across decision problems. (ref) applies (ref) to study comparative statics in the set of \ensuremath{\mathrm{BCE}}-consistent marginals when considering changes to the DM's prior or utility function. Building on results in bergemann2022counterfactuals, we show in (ref) how our results can be used to study whether a single information structure exists that rationalizes a DM's choices across different decision problems. In (ref), we apply our results to study \ensuremath{\mathrm{BCE}}-consistency in simple games.
Because our characterization results are not constructive, we study in (ref) which information structures make the marginals \ensuremath{\mathrm{BCE}}-consistent. In (ref), we characterize the Bayes plausible distributions over posteriors that implement a given action distribution, interpreting \ensuremath{\mathrm{BCE}}-consistency as a market-clearing condition in a persuasion economy and building on gale1957theorem.
Our results have implications for both empirical work and theoretical analysis. Empirically, for a given action distribution and utility function, our characterization results non-parametrically identify the set of priors such that the prior and action distribution are \ensuremath{\mathrm{BCE}}-consistent given the utility function, which is useful whenever the analyst has no information on what the prior should be, but may have auxiliary data on the DM’s payoffs.\footnote{By contrast, studies of risk aversion across domains assume the DM's beliefs about expected claim rates coincide with the frequencies in the data and estimate the curvature of the utility function \citep*{cohen2007estimating, barseghyan2011risk,barseghyan2013nature}.} Furthermore, for a given action distribution, our results characterize joint restrictions on the prior and the utility function for the triple of the utility function, prior, and action distribution to be consistent with information.
Our framework also has applications in behavioral economics, particularly for cognitive uncertainty models where $\mathrm{DM}$s exhibit random behavior across instances of the same problem (e.g., \citealp*{khaw2021cognitive}; enke2023cognitive). In these models, the state often represents the correct action and the DM\ has a noisy perception of this state. Whereas laboratory experiments may generate state-dependent choice data, outside the lab, analysts typically only observe average choices. Our results can test whether behavior is consistent with Bayesian cognitive uncertainty\textemdash whether noisy perception of states can be rationalized via an information structure.
Theoretically, our results open up the study of marginal information design\textemdash akin to reduced-form implementation in mechanism design\textemdash where an information designer cares only about the DM's actions, and not the state of the world. From this perspective, results such as (ref) reveal the structure of the binding constraints in information design problems, and we expect it can be used to further the study of Bayesian persuasion.
\paragraph{Related Literature} Our analysis relates to several strands of literature. lu2016random, rehbeck2023revealed, de2022rationalizing, and azrieli2022marginal study related rationalization problems but differ in either the available data or characterization approach. In lu2016random, the analyst observes the DM's stochastic choice out of every possible menu. As in our paper, the analyst in rehbeck2023revealed and de2022rationalizing observes the DM's stochastic choice from a single menu in a static and dynamic decision problem, respectively. Unlike our paper, their characterization results are in terms of the non-existence of a (possibly randomized) deviation, akin to Pearce's lemma. In azrieli2022marginal, the analyst observes the distribution of menus the DM\ faces and the DM's distribution of choices, but not the distribution of choices from each menu. Despite the difference in the settings, we discuss how (ref), which we obtain relying on gale1957theorem, can be obtained using their results.
A literature in decision theory and experimental economics studies when state-dependent stochastic choice data can be rationalized via costly information acquisition and whether the data identify the information acquisition costs (e.g., \citealp*{caplin2015revealed,caplin2017rationally,chambers2020costly,dewan2020estimating,denti2022posterior,caplin2023rationalizable}). Like we do, many of these papers provide results for a given utility function.\footnote{To be sure, caplin2023rationalizable consider recovering the utility function.} Unlike our paper, the DM's prior is observed in the data. Relatedly, ergin2010unique and dillenberger2014theory,dillenberger2023subjective study when menu choice data is consistent with costly information acquisition.
A literature in information design studies problems with (given) marginals. arieli2021feasible and morris2020no characterize joint distributions over posterior beliefs that are consistent with some information structure with binary and finitely many states, respectively. Assuming the sender and receiver care only about the posterior mean of the states, toikka2022bayesian show the sender's problem is a linear programming problem that only depends on the marginal distribution over actions. kolotilin2023persuasion characterize properties of optimal information structures assuming the first-order approach applies in a large class of persuasion problems with nonlinear sender preferences. strack2024privacy show the optimization over privacy-preserving signals can be cast as an optimal transport problem.
Methodologically, our paper relates to the econometrics literature on random sets for partial identification, where support functions are used to study the Aumann expectation of a random set (e.g., \citealp*{galichon2011set,beresteanu2011sharp,molchanov2018random}). Our extension to compact Polish action and state spaces relies on strassen1965existence, which also appears in galichon2011set.
In lieu of an organizational paragraph, we collect here mathematical notation and definitions used throughout the paper: \paragraph{Mathematical conventions and definitions} For a finite set \ensuremath{X}, we denote by $\ensuremath{\mathbb{R}}^\ensuremath{X}$ the set of vectors of length $|\ensuremath{X}|$. Depending on context, we refer to elements of $\ensuremath{\mathbb{R}}^\ensuremath{X}$ either as functions from $\ensuremath{X}$ to the reals or as vectors in $\ensuremath{\mathbb{R}}^\ensuremath{X}$. We reserve $\ensuremath{v}$ (serif) for the function $\ensuremath{v}:\ensuremath{X}\rightarrow\ensuremath{\mathbb{R}}$ and $\ensuremath{\boldsymbol{\ensuremath{v}}}$ (boldface) for the vector in $\ensuremath{\mathbb{R}}^\ensuremath{X}$. When we wish to emphasize the length of \ensuremath{\boldsymbol{\ensuremath{v}}}\ we write $\ensuremath{\boldsymbol{\ensuremath{v}}}\in\ensuremath{\mathbb{R}}^{|\ensuremath{X}|}$. If $\ensuremath{\boldsymbol{\ensuremath{v}}},\ensuremath{\tilde{\ensuremath{\boldsymbol{\ensuremath{v}}}}}$ are two vectors in $\ensuremath{\mathbb{R}}^\ensuremath{X}$, we denote by $\ensuremath{\boldsymbol{\ensuremath{v}}}\ensuremath{\tilde{\ensuremath{\boldsymbol{\ensuremath{v}}}}}$ their inner product, $\sum_{i=1}^{|\ensuremath{X}|}\ensuremath{\boldsymbol{\ensuremath{v}}}_\ensuremath{i}\ensuremath{\tilde{\ensuremath{\boldsymbol{\ensuremath{v}}}}}_\ensuremath{i}$.
Given two nonempty subsets $\ensuremath{V},\ensuremath{\ensuremath{V}^\prime}\subset\ensuremath{\mathbb{R}}^\ensuremath{X}$, their Minkowski sum is the set $\ensuremath{V}+\ensuremath{\ensuremath{V}^\prime}=\{\ensuremath{\boldsymbol{\ensuremath{v}}}+\ensuremath{\tilde{\ensuremath{\boldsymbol{\ensuremath{v}}}}}:\ensuremath{\boldsymbol{\ensuremath{v}}}\in\ensuremath{V},\ensuremath{\tilde{\ensuremath{\boldsymbol{\ensuremath{v}}}}}\in\ensuremath{\ensuremath{V}^\prime}\}$. Given a set $\ensuremath{V}\subset\ensuremath{\mathbb{R}}^\ensuremath{X}$, the support function of \ensuremath{V}\ is the mapping $\ensuremath{\boldsymbol{\ensuremath{p}}}\mapsto\sup\{\ensuremath{\boldsymbol{\ensuremath{p}}}\ensuremath{\boldsymbol{\ensuremath{v}}}:\ensuremath{\boldsymbol{\ensuremath{v}}}\in V\}$. A (convex) cone \ensuremath{C}\ is a subset of $\ensuremath{\mathbb{R}}^\ensuremath{X}$ that is closed under addition and non-negative scalar multiplication. A vector $\boldsymbol{w}\in\ensuremath{C}$ is an extreme ray if no linearly independent $\ensuremath{\boldsymbol{\ensuremath{v}}},\ensuremath{\tilde{\ensuremath{\boldsymbol{\ensuremath{v}}}}}\in\ensuremath{C}$ and positive scalars $\lambda,\gamma$ exist such that $\boldsymbol{w}=\lambda\ensuremath{\boldsymbol{\ensuremath{v}}}+\gamma\ensuremath{\tilde{\ensuremath{\boldsymbol{\ensuremath{v}}}}}$. Note that if $\ensuremath{\boldsymbol{\ensuremath{v}}}\in\ensuremath{C}$ is an extreme ray, so is $\lambda\ensuremath{\boldsymbol{\ensuremath{v}}}$ for $\lambda>0$. When we refer to an extreme ray, we refer to one representative of this equivalence class. Finally, a polytope is a bounded subset of $\ensuremath{\mathbb{R}}^\ensuremath{X}$ defined as the intersection of finitely many halfspaces of the form $\ensuremath{\boldsymbol{\ensuremath{p}}}_l\ensuremath{\boldsymbol{\ensuremath{v}}}\leq\ensuremath{b}_l$ for some $(\ensuremath{\boldsymbol{\ensuremath{p}}}_l,\ensuremath{b}_l)\in\ensuremath{\mathbb{R}}^{\ensuremath{X}}\times\ensuremath{\mathbb{R}}$, $l\in\{1,\dots,L\}$.
\paragraph{A decision problem with given marginals} Our model considers a DM\ taking an action under uncertainty about a state of the world. We denote by $\ensuremath{\Omega}=\{\ensuremath{\omega}_1,\dots,\ensuremath{\omega}_\ensuremath{I}\}$ the finite set of states of the world and by $\ensuremath{A}=\{\ensuremath{a}_1,\dots,\ensuremath{a}_{\ensuremath{J}}\}$ the finite set of actions. The DM's utility function $\ensuremath{u}:\ensuremath{A}\times\ensuremath{\Omega}\rightarrow\ensuremath{\mathbb{R}}$ describes the DM's payoff as a function of the action she takes and the state of the world. The tuple $\langle\ensuremath{\Omega},\ensuremath{A},\ensuremath{u}\rangle$ defines the decision problem.
Letting \ensuremath{\Delta(\ensuremath{\Omega})}\ denote the set of distributions over \ensuremath{\Omega}, the DM's utility defines the set of beliefs for which a given action $\ensuremath{a}\in\ensuremath{A}$ is optimal, \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}, as follows:
Below, we regard \ensuremath{\Delta(\ensuremath{\Omega})}\ as a full-dimensional set in $\ensuremath{\mathbb{R}}^{\ensuremath{I}-1}$. Furthermore, to streamline the presentation, we implicitly assume each nonempty \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ is full dimensional in $\ensuremath{\mathbb{R}}^{\ensuremath{I}-1}$.\footnote{Because $\cup_{\ensuremath{a}\in\ensuremath{A}}\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}=\ensuremath{\Delta(\ensuremath{\Omega})}$, at least one \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ is full dimensional in $\ensuremath{\mathbb{R}}^{\ensuremath{I}-1}$. Assuming all nonempty \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ are full-dimensional sets in $\ensuremath{\mathbb{R}}^{\ensuremath{I}-1}$ simplifies exposition.} (ref) deals with the general case.\footnote{Our results remain the same, but properly stating them requires defining the embedding of a lower-dimensional subset into $\ensuremath{\mathbb{R}}^{\ensuremath{I}-1}$.} Thus, whenever we refer to the dimension of a set, we mean its dimension in $\ensuremath{\mathbb{R}}^{\ensuremath{I}-1}$, even if our notation does not make it explicit.
We take the viewpoint of an analyst who knows the decision problem, but not whether the DM\ has access to further information before taking her action. The analyst also observes the DM's distribution over actions, $\ensuremath{\ensuremath{\nu}_0}\in\Delta(\ensuremath{A})$. The analyst's goal is to determine for which prior distributions over the states, $\ensuremath{\ensuremath{\mu}_0}\in\ensuremath{\Delta(\ensuremath{\Omega})}$, \ensuremath{\ensuremath{\nu}_0}\ can be rationalized as the result of the DM\ optimally choosing her actions after observing the outcome of an information structure.
For a given prior distribution $\ensuremath{\ensuremath{\mu}_0}\in\ensuremath{\Delta(\ensuremath{\Omega})}$, the results in myerson1982optimal, kamenica2011bayesian, and bergemann2016bayes imply an information structure exists that rationalizes \ensuremath{\ensuremath{\nu}_0}\ given the utility function \ensuremath{u}\ if and only if the pair of distributions \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ satisfy the following:
In words, an information structure exists that rationalizes the DM's choices \ensuremath{\ensuremath{\nu}_0}\ given her utility function \ensuremath{u}\ and prior belief \ensuremath{\ensuremath{\mu}_0}\ if and only if a joint distribution $\ensuremath{\pi}\in\Delta(\ensuremath{A}\times\ensuremath{\Omega})$ exists that satisfies Equations (ref), (ref), and (ref). (ref) states that if the DM\ knows \ensuremath{a}\ has been drawn according to \ensuremath{\pi}\textemdash but not the state\textemdash the DM\ finds \ensuremath{a}\ optimal. Equations (ref) and (ref) state the joint distribution is consistent with the pair \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}: the generated information averages out to the prior (ref), and the DM's average choices coincide with \ensuremath{\ensuremath{\nu}_0}\ (ref).
With (ref) at hand, we can now state the analyst's problem formally: given \ensuremath{\ensuremath{\nu}_0}\ and \ensuremath{u}, the analyst seeks to characterize the set of prior distributions \ensuremath{\ensuremath{\mu}_0}\ such that $\ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\in\ensuremath{\ensuremath{\mathrm{BCE}}\left(\ensuremath{u}\right)}$. We denote the set of all such priors by \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}; that is,
Taking \ensuremath{\ensuremath{\nu}_0}\ as given, the question of whether \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}\ is equivalent to whether \ensuremath{\ensuremath{\mu}_0}\ belongs in \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}.
\paragraph{The action marginal as a distribution over posteriors} A joint distribution $\ensuremath{\pi}\in\Delta(\ensuremath{A}\times\ensuremath{\Omega})$ with marginals $\ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}$ induces a belief system, $\{\ensuremath{\mu}(\cdot|\ensuremath{a})\in\ensuremath{\Delta(\ensuremath{\Omega})}:\ensuremath{a}\in\ensuremath{A}\}$, describing the DM's beliefs conditional on action \ensuremath{a}, which satisfies that for all actions $\ensuremath{a}\in\ensuremath{A}$, \[\ensuremath{\ensuremath{\nu}_0}(\ensuremath{a})\ensuremath{\mu}(\ensuremath{\omega}|\ensuremath{a})=\ensuremath{\pi}(\ensuremath{a},\ensuremath{\omega}).\] In this case, one can view \ensuremath{\ensuremath{\nu}_0}\ as a distribution over posteriors and the belief system $\left(\ensuremath{\mu}(\cdot|\ensuremath{a})\right)_{\ensuremath{a}\in\ensuremath{A}}$ as its support. Consequently, whether \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}\ is equivalent to whether a belief system $\left(\ensuremath{\mu}(\cdot|\ensuremath{a})\right)_{\ensuremath{a}\in\ensuremath{A}}$ exists that satisfies the following. First, for all states $\ensuremath{\omega}\in\ensuremath{\Omega}$,
and for all $\ensuremath{a},\ensuremath{\ensuremath{a}^\prime}\in\ensuremath{A}$,
Then, Equations (ref) and (ref) require that (i) \ensuremath{\ensuremath{\nu}_0}\ induces a Bayes plausible distribution over posteriors and (ii) for all actions \ensuremath{a}, the posterior belief $\ensuremath{\mu}(\cdot|\ensuremath{a})$ is an element of $\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}$. Under this interpretation, the action distribution \ensuremath{\ensuremath{\nu}_0}\ describes the frequency with which inducing beliefs in $\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}$ is necessary. The results in (ref) use the representation of \ensuremath{\mathrm{BCE}}-consistency given the utility function \ensuremath{u}\ implied by Equations (ref) and (ref).
We close this section with (ref), which compares our approach with that in the empirical work discussed in the introduction. Readers interested in the characterization results can jump to (ref), with little loss of continuity.
In this section, we introduce our basic characterization result, (ref).
\paragraph{A Minkowski-sum representation}
For a fixed utility function \ensuremath{u}, Equations (ref) and (ref) allow us to immediately characterize the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ as the \ensuremath{\ensuremath{\nu}_0}-weighted Minkowski sum of the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}. That is, we claim
Hence, $\ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\in\ensuremath{\ensuremath{\mathrm{BCE}}\left(\ensuremath{u}\right)}$ if and only if \ensuremath{\ensuremath{\mu}_0}\ is in the \ensuremath{\ensuremath{\nu}_0}-weighted Minkowski sum of the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}. That any prior in \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ is an element of $\sum_{\ensuremath{a}\in\ensuremath{A}}\ensuremath{\ensuremath{\nu}_0}(\ensuremath{a})\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}$ follows from Equations (ref) and (ref). Conversely, consider a collection of beliefs $\{\ensuremath{\mu}(\cdot|\ensuremath{a})\in\ensuremath{\Delta(\ensuremath{\Omega})}:\ensuremath{a}\in\ensuremath{A}\}$ such that for all $\ensuremath{a}\in\ensuremath{A}$, $\ensuremath{\mu}(\cdot|\ensuremath{a})\in\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}$. This collection together with the prior $\ensuremath{\ensuremath{\ensuremath{\mu}_0}^\prime}=\sum_{\ensuremath{a}\in\ensuremath{A}}\ensuremath{\ensuremath{\nu}_0}(\ensuremath{a})\ensuremath{\mu}(\cdot|\ensuremath{a})$ satisfy Equations (ref) and (ref); hence, the pair $(\ensuremath{\ensuremath{\ensuremath{\mu}_0}^\prime},\ensuremath{\ensuremath{\nu}_0})$ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}. We illustrate this and other results in this section with (ref):
The Minkowski-sum representation of the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ in (ref) has several implications. First, a prior exists such that \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}\ if and only if the support of \ensuremath{\ensuremath{\nu}_0}\ does not include strictly dominated actions, that is, actions $\ensuremath{a}\in\ensuremath{A}$ for which \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ is empty. Letting $\ensuremath{A}_+$ denote the support of \ensuremath{\ensuremath{\nu}_0}, we assume in what follows that $\ensuremath{A}_+$ contains no strictly dominated actions. Second, because the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ are full-dimensional polytopes, \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ is itself a full-dimensional polytope (cf. (ref)).\footnote{To be fully precise, \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ is full-dimensional so long as an $\ensuremath{a}\in\ensuremath{A}_+$ exists such that \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ is full-dimensional.} As such, it can be represented either as the convex hull of its extreme points (the so-called $V$-representation) or as the intersection of finitely many halfspaces (the so-called $H$-representation). In fact, letting $\ensuremath{\mathrm{ext}}(\ensuremath{X})$ denote the set of extreme points of a subset \ensuremath{X}\ of $\ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{\Omega}}$, we have that
In words, any extreme point of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ is a \ensuremath{\ensuremath{\nu}_0}-weighted combination of extreme points in the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}, but the opposite may not hold.
Although (ref) characterizes the set of \ensuremath{\mathrm{BCE}}-consistent distributions for a fixed utility function \ensuremath{u}\ and action distribution \ensuremath{\ensuremath{\nu}_0}, it does not capture the joint restrictions on the prior and utility that arise from requiring \ensuremath{\ensuremath{\nu}_0}\ to be rationalizable via information. Moreover, even for fixed \ensuremath{u}, computing \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ is nontrivial\textemdash let alone doing so for every \ensuremath{u}\ the analyst may wish to consider. Most algorithms for Minkowski sums require the $V$-representation of the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}, which is generally unavailable, with the notable exception of bergemann2015limits.\footnote{See das2024worst. In fact, obtaining the $V$-representation of the set \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ from its $H$-representation in (ref) is known to be a computationally complex problem weibel2007minkowski.}
In the remainder of the paper, we focus on studying the $H$-representation of the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}. First, as our results below highlight, the $H$-representation allows us to understand which pairs of prior and utility function can jointly rationalize the \ensuremath{\ensuremath{\nu}_0}\ via information (cf. (ref)). Second, our characterization of the $H$-representation of the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ relies on the support function of this set, an object that has received increasing attention in the econometrics literature on random sets for partial identification molchanov2018random.\footnote{That literature uses the support function to study the Aumann expectation of a random set, with the Minkowski sum serving as its empirical analog. Because \ensuremath{A}\ is finite, the Minkowski sum exactly represents \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}. The key difference is that in our setting, the data\textemdash the action distribution\textemdash determines the weights of the Minkowski sum, not its summands, whereas in econometrics, data informs the summands, with weights given by empirical frequencies.}
\paragraph{The support function of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}} Because the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ is convex, it can be characterized via its support function. Moreover, the support function of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ is the \ensuremath{\ensuremath{\nu}_0}-weighted sum of the support function of the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ rockafellar1970convex. Together, these arguments lead to the following statement, which we record for future reference:
In words, \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}\ if and only if for all vectors $\ensuremath{\boldsymbol{\ensuremath{p}}}\in\ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{\Omega}}$, \ensuremath{\ensuremath{\mu}_0}\ lies below the support function of the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}. The left-hand side of (ref) represents the support function of the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ via the support functions of the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}.\footnote{Whereas (ref) is an immediate consequence of the Minkowski-sum representation of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}, it can alternatively be obtained from the dual representation of the system of equations that define the set \ensuremath{\ensuremath{\mathrm{BCE}}\left(\ensuremath{u}\right)}\ in (ref) (see (ref)) and from strassen1965existence.} An immediate consequence of (ref) is that to define the $H$-representation of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\textemdash defined by a collection of normal vector-height pairs $(\ensuremath{\boldsymbol{\ensuremath{p}}}_i,b_i)_{i\in I}$\textemdash the normal vectors alone suffice. By (ref), the height of the halfspace with normal vector \ensuremath{\boldsymbol{\ensuremath{p}}}\ is the \ensuremath{\ensuremath{\nu}_0}-weighted sum of the values of the support functions of the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ in direction \ensuremath{\boldsymbol{\ensuremath{p}}}. Finally, we note that (ref) uses our ongoing assumption that the support of \ensuremath{\ensuremath{\nu}_0}\ contains no strictly dominated actions. However, by changing $\max$ for $\sup$ in (ref), the support function of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ also accounts for whether \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ is empty by following the convention that the supremum over an empty set is $-\infty$.
\paragraph{Test functions} (ref) provides an $H$-representation of the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}, albeit not the most useful one, because checking whether $\ensuremath{\ensuremath{\mu}_0}\in\ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}$ requires verifying (ref) holds for infinitely many vectors $\ensuremath{\boldsymbol{\ensuremath{p}}}\in\ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{\Omega}}$. The results that follow refine the result in (ref) by describing sets of test functions $\ensuremath{\pazocal{P}}\subsetneq\ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{\Omega}}$ such that if (ref) holds for all vectors in $\ensuremath{\boldsymbol{\ensuremath{p}}}\in\ensuremath{\pazocal{P}}$, it holds for all vectors $\ensuremath{\boldsymbol{\ensuremath{p}}}\in\ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{\Omega}}$.\footnote{Ideally, one would like to obtain the minimal $H$-representation of the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}. Whereas the results in Theorems (ref)--(ref) provide the minimal-$H$ representation of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ in some environments, we use the term test functions because these vectors are often a superset of the minimal-$H$ representation.}
The Minkowski-sum representation of the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ immediately implies the face structure of the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ determine that of the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ and hence its $H$-representation. The concepts of the normal fan of a polytope and the common refinement of the normal fans of a collection of polytopes, to which we turn next, allow us to describe how the face structure of the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ determine the $H$-representation of the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}:
Intuitively, the normal fan of a polytope “partitions” normal vectors \ensuremath{\boldsymbol{\ensuremath{p}}}\ according to the faces at which the objective, $\ensuremath{\boldsymbol{\ensuremath{p}}}\ensuremath{\boldsymbol{\ensuremath{v}}}$, attains its maximum value on \ensuremath{P}.\footnote{Technically speaking, it is not a partition, which can be easily seen in (ref). The direction $(1,1)$ belongs in the normal cone of the face defined by $\ensuremath{\mu}_1+\ensuremath{\mu}_2=1$, and also in the normal cone of the face defined by $(\ensuremath{\mu}_1,\ensuremath{\mu}_2)=(\nicefrac{1}{2},\nicefrac{1}{2})$.} In line with the partition interpretation, the common refinement of a collection of normal fans is the meet of these partitions.
We illustrate (ref) in the context of (ref) in (ref). In (ref), the blue arrows denote the normal vectors defining the polytope $\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a}_1)$. Clockwise from the left, these vectors are given by the utility differences $\ensuremath{u}(\ensuremath{a}_\ensuremath{j},\cdot)-\ensuremath{u}(\ensuremath{a}_1,\cdot)$ for $j\in\{2,3\}$\textemdash viewed as vectors in $\ensuremath{\mathbb{R}}^2$\textemdash the direction $(1,1)$\textemdash arising from the constraint that $\ensuremath{\mu}_1+\ensuremath{\mu}_2\leq1$\textemdash and the (negative of the) canonical direction $e_{\ensuremath{\omega}_2}=(0,1)$\textemdash arising from the nonnegativity constraint on $\ensuremath{\mu}_2$. The normal fan of $\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a}_1)$ is given by the collection of cones depicted in dotted red, including the one-dimensional ones. (ref) depicts the vectors that generate the cones in the normal fan. (ref) illustrates how the normal fan partitions $\ensuremath{\mathbb{R}}^2$, with vectors in the same cell representing linear functions in $\ensuremath{\mathbb{R}}^\ensuremath{\Omega}$ that attain their maximum on the same face of $\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a}_1)$. (ref) illustrates the normal fan of $\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a}_2)$. (ref) illustrates the common refinement of the normal fans of $\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a}_1)$ and $\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a}_2)$, which refines the partitions of $\ensuremath{\mathbb{R}}^2$ induced by each of the normal fans.
The normal vectors defining the minimal $H$-representation of a full-dimensional polytope correspond to the extreme rays of the one-dimensional cones in the normal fan of that polytope ziegler2012lectures. Figures (ref) and (ref) illustrate this observation. For ease of reference, below, we denote the extreme rays of the one-dimensional cones in the normal fan of a polytope \ensuremath{P}\ by $\ensuremath{\pazocal{P}}_{\ensuremath{\mathrm{ext}}}(\ensuremath{N}(\ensuremath{P}))$. Consequently, the minimal $H$-representation of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ obtains from the extreme rays of the one-dimensional cones in the normal fan of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}.\footnote{Recall that we mean one-dimensional in $\ensuremath{\mathbb{R}}^{\ensuremath{I}-1}$.} The normal fan of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ is in fact the common refinement of the normal fans of the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ ziegler2012lectures. For instance, the normal fan in (ref) is the normal fan of $\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a}_1)+\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a}_2)$. It follows that the one-dimensional cones in the normal fan of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ are the one-dimensional cones in the common refinement of the normal fans of the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}.
(ref) summarizes this discussion and presents our basic characterization of the set \ensuremath{\ensuremath{\mathrm{BCE}}\left(\ensuremath{u}\right)}.
The proof of (ref) and other results can be found in (ref).
(ref) characterizes the set of priors \ensuremath{\ensuremath{\mu}_0}\ such that \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}\ via a finite system of inequalities \ensuremath{\ensuremath{\mu}_0}\ must satisfy. In practice, the analyst knows neither the prior \ensuremath{\ensuremath{\mu}_0}\ nor the DM's utility \ensuremath{u}. From this perspective, (ref) describes the joint restrictions on the pairs $(\ensuremath{\ensuremath{\mu}_0},\ensuremath{u})$ for which an information structure exists that rationalizes the given action marginal \ensuremath{\ensuremath{\nu}_0}. Thus, the approach in this paper also offers an alternative perspective on how to carry out the identification exercise. Typically, the analyst specifies the prior and the utility function up to a finite-dimensional parameter. Instead, for a given utility function \ensuremath{u}, the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ provides a nonparametric representation of all priors for which \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}, which is useful whenever the analyst has no information on what the prior should be, but may have auxiliary data on the DM's payoffs.
The characterization in (ref) leaves open the question of how to obtain the vectors that generate the one-dimensional cones in the common refinement of the normal fans $\{\ensuremath{N}(\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}):\ensuremath{a}\in\ensuremath{A}_+\}$. The answer to this question is the focus of the rest of the paper. The main challenge is that the one-dimensional cones of the normal fan of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ may not be determined solely from the one-dimensional cones of the normal fan of the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\textemdash the extreme rays of which we know\textemdash but from intersections of higher-dimensional cones in these normal fans\textemdash the extreme rays of which a priori we do not know. This issue is exacerbated by the cardinality of the set of states and actions, or by how complex the utility differences $\ensuremath{u}(\ensuremath{a},\cdot)-\ensuremath{u}(\ensuremath{\ensuremath{a}^\prime},\cdot)$ are. For that reason, our results place restrictions either on the cardinality of the states or the utility function. When \ensuremath{\Omega}\ contains at most three states, (ref) shows the one-dimensional cones of the normal fan of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ obtain directly from those of the normal fans $\{\ensuremath{N}(\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}):\ensuremath{a}\in\ensuremath{A}_+\}$. (ref) characterizes the extreme rays of the one-dimensional cones of $\ensuremath{N}(\ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}})$ for canonical classes of utility functions.
\paragraph{Connecting the $H$-representations of \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ and of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}} Suppose the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ is nonempty, and let \ensuremath{\boldsymbol{\ensuremath{p}}}\ denote an extreme ray of a one-dimensional cone of \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}, for some action $\ensuremath{a}\in\ensuremath{A}_+$. Then, \ensuremath{\boldsymbol{\ensuremath{p}}}\ is an extreme ray of a one-dimensional cone of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}.\footnote{When \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ has dimension less than $\ensuremath{I}-1$, this statement applies to those \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ with dimension equal to that of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}.} Hence, the minimal set of test functions always includes the extreme rays of the one-dimensional cones of the normal fan of \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ whenever $\ensuremath{a}\in\ensuremath{A}_+$, that is,
Thus, asking when this inclusion is an equality is natural. (ref) below shows this is the case whenever $\ensuremath{|\ensuremath{\Omega}|}\leq3$. To state (ref), define
where $\ensuremath{u}(\ensuremath{\ensuremath{a}^{\prime\prime}},\cdot)-\ensuremath{u}(\ensuremath{\ensuremath{a}^\prime},\cdot)\in\ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{\Omega}}$ is the vector that collects the payoff differences between \ensuremath{\ensuremath{a}^{\prime\prime}}\ and \ensuremath{\ensuremath{a}^\prime}\ as a function of the state, and $e_\ensuremath{\omega}$ is the vector in \ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{\Omega}}\ with a 1 in the \ensuremath{\omega}-coordinate and $0$ elsewhere. For a given $\ensuremath{\ensuremath{a}^\prime}\in\ensuremath{A}$, these vectors define the $H$-representation of \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{\ensuremath{a}^\prime})}\ and hence contain $\ensuremath{\pazocal{P}}_{\ensuremath{\mathrm{ext}}}(\ensuremath{N}(\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{\ensuremath{a}^\prime})}))$. Taking union over the different actions $\ensuremath{\ensuremath{a}^\prime}$ yields the set $\ensuremath{\pazocal{P}}_{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}}$.\footnote{We could refine the set $\ensuremath{\pazocal{P}}_{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}}$ by eliminating normal directions to \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ that do not define a facet of \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}.}
(ref) summarizes the above discussion:
In (ref), (ref) corresponds to (ref) evaluated at the canonical vectors $-e_\ensuremath{\omega}$, whereas (ref) corresponds to (ref) evaluated at the vectors $\ensuremath{u}(\ensuremath{\ensuremath{a}^{\prime\prime}},\cdot)-\ensuremath{u}(\ensuremath{\ensuremath{a}^\prime},\cdot)$ for different action pairs $(\ensuremath{\ensuremath{a}^\prime},\ensuremath{\ensuremath{a}^{\prime\prime}})$.
We remark that as long as a direction $\ensuremath{\boldsymbol{\ensuremath{p}}}\in\ensuremath{\pazocal{P}}_{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}}$ is an extreme ray in the one-dimensional cone of the normal fan of some \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ such that $\ensuremath{a}\in\ensuremath{A}_+$, the corresponding equation in the statement of (ref) is part of the $H$-representation of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ independently of the cardinality of the states. Thus, whereas any $\ensuremath{\boldsymbol{\ensuremath{p}}}\in\ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{\Omega}}$ defines through (ref) a necessary condition for \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ to be \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}, (ref) gives us a precise sense in which equations (ref) and (ref) are the necessary conditions and can be used to rule out pairs \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ that are not \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}. In fact, we can provide intuition for Equations (ref) and (ref) by reasoning about their necessity, starting from (ref). If \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}, then we can find a belief system that satisfies for each $\ensuremath{a}\in\ensuremath{A}$,
The above implication holds because \ensuremath{\mathrm{BCE}}-consistency of \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ implies that for each $\ensuremath{a}\in\ensuremath{A}_+$, $\ensuremath{\mu}(\cdot|\ensuremath{a})$ is an element of \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}. Thus, if \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent, the left-hand side of (ref) is a lower bound on the prior. In particular, (ref) implies the non-negativity constraints on \ensuremath{\ensuremath{\mu}_0}. Whenever the left-hand side of (ref) is positive, rationalizing \ensuremath{\ensuremath{\nu}_0}\ via information requires the DM\ to assign strictly positive probability to some states.
Intuitively, (ref) verifies whether finding a belief system $\{\ensuremath{\mu}(\cdot|\ensuremath{a}):\ensuremath{a}\in\ensuremath{A}\}$ that satisfies the belief martingale property relative to \ensuremath{\ensuremath{\mu}_0}\ is possible. Indeed, if the inequality in (ref) failed for some state, either the frequency with which the DM\ is taking a given action, or the minimum probability the DM\ needs to assign to this state so that taking a given action is optimal, reflects that the DM\ is more optimistic than at the prior. In this case, \ensuremath{\ensuremath{\nu}_0}\ cannot be rationalized via information.
To provide intuition for (ref), considering the binary-action case is useful. Assume $\ensuremath{A}=\{\ensuremath{a}_1,\ensuremath{a}_2\}$: \ensuremath{\mathrm{BCE}}-consistency of \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ implies a belief system exists that satisfies at most two obedience constraints, which we can write as the following chain of inequalities:
That is, information resulting in beliefs $\ensuremath{\mu}(\cdot|\ensuremath{a}_1)$ and $\ensuremath{\mu}(\cdot|\ensuremath{a}_2)$ alters the relative ranking of $\ensuremath{a}_1$ and $\ensuremath{a}_2$. Equation (ref) implies that if we add up both sides of the above chain, we obtain the following:
In other words, although information can alter the relative ranking between the two actions, it cannot systematically do so: on average, the ranking between $\ensuremath{a}_1$ and $\ensuremath{a}_2$ must coincide with how the DM\ ranks these two actions at the prior \ensuremath{\ensuremath{\mu}_0}. This observation is the analogue to the belief martingale condition, albeit in terms of the DM's payoffs; hence we refer to it as a payoff martingale condition. Using once again the property that $\ensuremath{\mu}(\cdot|\ensuremath{a})\in\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}$, the above equality implies (ref).
Consider now the case in which the DM\ has three actions, $\ensuremath{A}=\{\ensuremath{a}_1,\ensuremath{a}_2,\ensuremath{a}_3\}$, and once again, (ref) for the pair $\ensuremath{a}_1,\ensuremath{a}_2$. Whereas the obedience constraints feature the comparison between these two actions when $\ensuremath{a}_1$ or $\ensuremath{a}_2$ is recommended, no such inequality arises when $\ensuremath{a}_3$ is recommended. Still, (ref) adds up over all actions, including $\ensuremath{a}_3$. The reason is that when ensuring $\ensuremath{a}_3$ is optimal, the obedience constraints place no restrictions on the relative ranking of $\ensuremath{a}_1$ and $\ensuremath{a}_2$. Yet, the ranking of these two actions has to average to their ranking at the prior over all beliefs the DM\ has. (ref) checks that the payoff martingale condition can be satisfied by placing bounds on the induced relative rankings for each action pair $(\ensuremath{\ensuremath{a}^\prime},\ensuremath{\ensuremath{a}^{\prime\prime}})$ across all action recommendations.
We illustrate (ref) using (ref): \setcounter{example}{0}
Whereas (ref) provides an instance in which all equations in (ref) are needed to define \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}, this is not always the case. Indeed, when $\ensuremath{\Omega}=\{\ensuremath{\omega}_1,\ensuremath{\omega}_2\}$, the belief martingale equations (ref) alone define this set. In the case of two states, the prior is summarized by the probability of $\ensuremath{\omega}_2$, $\ensuremath{\ensuremath{\mu}_0}(\ensuremath{\omega}_2)$. It is immediate that the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ is an interval; hence, it is defined by two inequalities. The same is true of the sets $\{\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}:\ensuremath{a}\in\ensuremath{A}\}$, which in a slight abuse of notation, we define as $\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}=[\underline{\ensuremath{\mu}}(\ensuremath{\omega}_2|\ensuremath{a}),\overline{\ensuremath{\mu}}(\ensuremath{\omega}_2|\ensuremath{a})]$. (ref) shows the lower and upper bounds of \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}\ define the lower and upper bounds of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}:
We make two remarks. First, the right-hand side of the expression in (ref) obtains from (ref) at $\ensuremath{\omega}_1$. Second, the representation of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ via its extreme points provides one way of understanding (ref) (cf. (ref)). With binary states, the beliefs $\{\underline{\ensuremath{\mu}}(\ensuremath{\omega}_2|\ensuremath{a}),\overline{\ensuremath{\mu}}(\ensuremath{\omega}_2|\ensuremath{a})\}$ are the extreme points of \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}. Furthermore, the extreme points of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ are \ensuremath{\ensuremath{\nu}_0}-weighted convex combinations of the extreme points of the sets \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}, but not the reverse. (ref) states that only the \ensuremath{\ensuremath{\nu}_0}-weighted convex combination of the minimal extreme points and of the maximal extreme points can be extreme in \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}.
\paragraph{Beyond simple state spaces} As anticipated above, the result in (ref) does not extend to larger state spaces: when the cardinality of \ensuremath{\Omega}\ is at least four, an extreme ray in a one-dimensional cone of the normal fan of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ may not be an extreme ray in a one-dimensional cone of any of the normal fans $\{\ensuremath{N}(\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}):\ensuremath{a}\in\ensuremath{A}\}$.\footnote{Recall the normal fan of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ is the common refinement of the normal fans $\ensuremath{N}(\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})})$. Hence, a normal cone of \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\ obtains by intersecting normal cones in $\{\ensuremath{N}(\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}):\ensuremath{a}\in\ensuremath{A}\}$. In three or more dimensions, the extreme rays of the intersection of two cones need not be extreme rays in any of the cones in the intersection.} We illustrate this possibility with a simple binary-action, four-state example:
In this section, we identify assumptions on the utility function \ensuremath{u}, which allow us to refine the basic characterization in (ref) without imposing assumptions on the cardinality of the state space. Concretely, we focus on monotone and concave decision problems in which the utility function \ensuremath{u}\ satisfies concavity and increasing differences assumptions ((ref)). (ref) identifies a set of test functions under this assumption. We next consider two special cases: affine utility differences ((ref)) and two-step utility differences ((ref)). In each case, we characterize the set of \ensuremath{\mathrm{BCE}}-consistent marginals via a system of finitely many inequalities. In contrast to the results in (ref), for which we relied on the properties of the Minkowski sum, the characterization in this section relies on the dual of the program induced by checking the feasibility of Equations (ref), (ref), and (ref). This dual approach allows us to identify both the test functions and the value of the support function of the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}\textemdash the left-hand side of (ref)\textemdash and thus provide a more succinct characterization of the set \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}. (ref) extends the results in this section to the case in which the sets of actions and states are compact Polish spaces.
\paragraph{Monotone and concave decision problems} We focus on decision problems in which the utility function \ensuremath{u}\ satisfies a combination of concavity and increasing-differences assumptions. To state these assumptions, the ordering of the states and the actions matters. Recall we are indexing the actions with $\ensuremath{j}\in\{1,\dots,\ensuremath{J}\}$ and the states with $\ensuremath{i}\in\{1,\dots,\ensuremath{I}\}$. Definitions (ref) and (ref) below place restrictions on the utility difference across adjacent actions, which we denote by
In words, under increasing differences, the DM\ finds higher index actions more attractive than lower index ones in higher states.
In words, a utility function is concave$^*$ if, for a given state, there are decreasing returns to increasing the actions\textemdash that is, $\ensuremath{u}(\cdot,\ensuremath{\omega})$ is concave\textemdash and, for any belief, the DM\ can be indifferent between at most two actions.
When the DM's utility function satisfies the above definitions, we say the decision problem is monotone and concave. We record this in (ref) for ease of reference:
(ref) affords the following simplification in determining whether a pair of distributions is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}. Recall that \ensuremath{\mathrm{BCE}}-consistency of \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is equivalent to the feasibility of the system defined by Equations (ref), (ref), and (ref). (ref) implies we can ignore all obedience constraints not involving adjacent action pairs, $(\ensuremath{a}_\ensuremath{j},\ensuremath{a}_{\ensuremath{j}+1})$ and $(\ensuremath{a}_\ensuremath{j},\ensuremath{a}_{j-1})$, thereby reducing the number of obedience constraints to $2\ensuremath{J}-2$. Furthermore, (ref) also implies that, given an action $\ensuremath{a}_\ensuremath{j}$, the adjacent obedience constraints, $(\ensuremath{a}_\ensuremath{j},\ensuremath{a}_{\ensuremath{j}+1})$ and $(\ensuremath{a}_\ensuremath{j},\ensuremath{a}_{j-1})$, cannot simultaneously bind whenever $\ensuremath{j}\in\{2,\dots,\ensuremath{J}-1\}$.
\paragraph{Dual formulation of \ensuremath{\mathrm{BCE}}-consistency} Recall our goal is to identify a set of test functions for (ref). Whereas the results in (ref) identified such functions by relying on properties of the Minkowski sum, the results in this section rely on the analysis of a problem dual to determining the feasibility of the equations that define \ensuremath{\mathrm{BCE}}-consistency of \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}, Equations (ref), (ref), and (ref).
Consider the problem of choosing $\ensuremath{\pi}\in\Delta(\ensuremath{A}\times\ensuremath{\Omega})$ to maximize $0$ subject to Equations (ref), (ref), and (ref). This problem has value $0$ and hence, the system of equations in (ref) is feasible, if and only if program (ref) below has nonnegative value:
In the above program, the vectors \ensuremath{\boldsymbol{\ensuremath{p}}}\ and \ensuremath{\boldsymbol{\ensuremath{q}}}\ are the Lagrange multipliers on the constraints (ref) and (ref), respectively. That the notation for the multiplier on (ref) coincides with that of the vectors in (ref) is not a coincidence: the analysis that follows shows the dual variables \ensuremath{\boldsymbol{\ensuremath{p}}}\ that solve this program are intimately related to the test functions in (ref). The vectors \ensuremath{\ensuremath{\boldsymbol{\ensuremath{\lambda}}}^\uparrow}\ and \ensuremath{\ensuremath{\boldsymbol{\ensuremath{\lambda}}}^\downarrow}\ are the multipliers on the adjacent obedience constraints: conditional on a recommendation to take $\ensuremath{a}_\ensuremath{j}$, $\ensuremath{\ensuremath{\lambda}^\uparrow}_\ensuremath{j}$ is the multiplier on the upward-looking constraint that $\ensuremath{a}_\ensuremath{j}$ is better than $\ensuremath{a}_{\ensuremath{j}+1}$, and $\ensuremath{\ensuremath{\lambda}^\downarrow}_\ensuremath{j}$ is the multiplier on the downward-looking constraint that $\ensuremath{a}_{j}$ is better than $\ensuremath{a}_{j-1}$. Finally, because these constraints cannot simultaneously bind, complementary slackness implies the condition $\ensuremath{\ensuremath{\lambda}^\uparrow}_\ensuremath{j}\ensuremath{\ensuremath{\lambda}^\downarrow}_\ensuremath{j}=0$ must hold at a solution. Below, we follow the convention that $\ensuremath{\ensuremath{\lambda}^\downarrow}_1=\ensuremath{\ensuremath{\lambda}^\uparrow}_{\ensuremath{|\ensuremath{A}|}}=0$.
We make two observations.\footnote{Recall that we are assuming no action in $\ensuremath{A}_+$ is strictly dominated. Absent this assumption, the value of the dual is $-\infty$ and hence, the primal is unfeasible.} First, program (ref) is always feasible because we can always set to $0$ the coordinates of each of the vectors \ensuremath{\boldsymbol{\ensuremath{p}}}, \ensuremath{\boldsymbol{\ensuremath{q}}}, \ensuremath{\ensuremath{\boldsymbol{\ensuremath{\lambda}}}^\uparrow}, and \ensuremath{\ensuremath{\boldsymbol{\ensuremath{\lambda}}}^\downarrow}\ and still satisfy the conditions of the program. Therefore, a solution exists. Second, in any solution to program (ref), the dual variable \ensuremath{\boldsymbol{\ensuremath{p}}}\ must be of the form
regardless of the choice of \ensuremath{\boldsymbol{\ensuremath{q}}}, \ensuremath{\ensuremath{\boldsymbol{\ensuremath{\lambda}}}^\uparrow}, and \ensuremath{\ensuremath{\boldsymbol{\ensuremath{\lambda}}}^\downarrow}. (ref) below provides a building block for the rest of this section. It shows that in analyzing the value of program (ref), we need only consider solutions in which either \ensuremath{\ensuremath{\boldsymbol{\ensuremath{\lambda}}}^\uparrow}\ or \ensuremath{\ensuremath{\boldsymbol{\ensuremath{\lambda}}}^\downarrow}\ is zero. In other words, we need only consider solutions in which either all adjacent upward-looking or all adjacent downward-looking constraints are non-binding. By (ref), this in turn has implications for the directions \ensuremath{\boldsymbol{\ensuremath{p}}}\ at which testing (ref) is sufficient.
To state (ref), define
In words, directions $\ensuremath{\boldsymbol{\ensuremath{p}}}\in\ensuremath{\pazocal{P}}^\uparrow$ correspond to directions for which the downward-looking constraints are nonbinding, whereas directions $\ensuremath{\boldsymbol{\ensuremath{p}}}\in\ensuremath{\pazocal{P}}^\downarrow$ correspond to directions for which the upward-looking constraints are nonbinding. Increasing differences implies that every vector in $\ensuremath{\pazocal{P}}^\uparrow$ is an increasing function of \ensuremath{\omega}, whereas every vector in $\ensuremath{\pazocal{P}}^\downarrow$ is a decreasing function of $\ensuremath{\omega}$.
(ref) identifies the vectors in $\ensuremath{\ensuremath{\pazocal{P}}^\uparrow}\cup\ensuremath{\ensuremath{\pazocal{P}}^\downarrow}$ as test functions for the \ensuremath{\mathrm{BCE}}-consistency of \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ given \ensuremath{u}. From an information design perspective, this result uncovers an interesting property of the extremal information structures that implement a given action distribution in monotone and concave decision problems. From the point of view of characterizing the extreme points of the set of joint distributions $\ensuremath{\pi}(\ensuremath{a},\ensuremath{\omega})$ that obediently implement a given \ensuremath{\ensuremath{\nu}_0}, (ref) says we can restrict attention to those in which either no downward-looking obedience constraint binds, or no upward-looking obedience constraint binds.
In the following sections, we use (ref) together with additional assumptions on the decision problem to identify a finite subset of the test functions in $\ensuremath{\pazocal{P}}^\uparrow\cup\ensuremath{\pazocal{P}}^\downarrow$.
Throughout this section, we assume the utility function satisfies the following condition:
We make three observations. First, in the case of binary actions, affine utility difference entails no loss of generality. Second, by reordering the states, it is without loss of generality to assume $\ensuremath{d}$ is increasing in \ensuremath{\omega}. Hence, decision problems with affine utility differences satisfy increasing differences. Thus, decision problems with affine utility differences are monotone and concave decision problems. Finally, we note the connection between affine utility differences when \ensuremath{A}\ and \ensuremath{\Omega}\ are finite sets, and quadratic loss in the general case in which they are compact, convex subsets of \ensuremath{\mathbb{R}}. The analogue to the payoff difference across adjacent actions in quadratic loss is the derivative of the utility function $-(\ensuremath{a}-\ensuremath{\omega})^2$ with respect to \ensuremath{a}:
which is analogous to condition (ref) when \ensuremath{d}\ is the identity, $\ensuremath{\gamma}=2$, and $\ensuremath{\kappa}=-2\ensuremath{a}$. Below, this connection becomes apparent once we note the analogy between (ref) and the characterization of feasible action distributions under quadratic loss as mean-preserving contractions of the prior (cf. strassen1965existence).
\paragraph{Test functions for affine utility differences} (ref) characterizes the test functions under the assumption of affine utility differences. We first state the result and then provide intuition for it. To state the result, we first define a family of vectors $\ensuremath{\boldsymbol{\ensuremath{p}}}\in\ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{\Omega}}$ and $\ensuremath{\boldsymbol{\ensuremath{q}}}\in\ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{A}}$, which below play the role of the test functions and the value of the support function at those test functions, respectively.
For any $\ensuremath{\ensuremath{\omega}^\star}\in\ensuremath{\Omega}$, define the vectors $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\uparrow}_{\ensuremath{\ensuremath{\omega}^\star}},\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\downarrow}_{\ensuremath{\ensuremath{\omega}^\star}}\in\ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{\Omega}}$ as
and let $\ensuremath{\pazocal{P}}_{\ensuremath{\mathrm{AUD}}}=\{\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\uparrow}_{\ensuremath{\ensuremath{\omega}^\star}},\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\downarrow}_{\ensuremath{\ensuremath{\omega}^\star}}:\ensuremath{\ensuremath{\omega}^\star}\in\ensuremath{\Omega}\}$. The notation highlights that $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\uparrow}_{\ensuremath{\ensuremath{\omega}^\star}}$ and $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\downarrow}_{\ensuremath{\ensuremath{\omega}^\star}}$ are elements of the sets $\ensuremath{\pazocal{P}}^\uparrow$ and $\ensuremath{\pazocal{P}}^\downarrow$, respectively (cf. (ref)). Note the vectors $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\uparrow}_{\ensuremath{\omega}_\ensuremath{I}}$ and $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\downarrow}_{\ensuremath{\omega}_1}$ correspond to the (normalized) utility differences, while the vectors $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\uparrow}_{\ensuremath{\omega}_1}$ and $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\downarrow}_{\ensuremath{\omega}_\ensuremath{I}}$ are proportional to the vectors $\{\mathbf{1},-\mathbf{1}\}$, where $\mathbf{1}$ is the vector with $1$ in every coordinate.
To each vector in $\ensuremath{\pazocal{P}}_{\ensuremath{\mathrm{AUD}}}$, associate the vectors $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{q}}}^\uparrow}_{\ensuremath{\ensuremath{\omega}^\star}},\ensuremath{\ensuremath{\boldsymbol{\ensuremath{q}}}^\downarrow}_{\ensuremath{\ensuremath{\omega}^\star}}\in\ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{A}}$, defined as follows:
(ref) reduces the question of whether \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}\ to checking at most $2\ensuremath{|\ensuremath{\Omega}|}$ linear inequalities, together with $\ensuremath{|\ensuremath{\Omega}|}+1$ constraints implied by $\ensuremath{\ensuremath{\mu}_0}\in\ensuremath{\Delta(\ensuremath{\Omega})}$.\footnote{To be precise, checking that the coordinates of the prior add up to one is implied by (ref) evaluated at $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\uparrow}_{\ensuremath{\omega}_1}$ and $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\downarrow}_{\ensuremath{\omega}_\ensuremath{I}}$.} Interestingly, the slopes, $\ensuremath{\gamma}_\ensuremath{j}$, and constants, $\ensuremath{\kappa}_\ensuremath{j}$, determine the heights of the halfspaces that define \ensuremath{\ensuremath{\mathrm{M}}\ensuremath{(\ensuremath{u},\ensuremath{\ensuremath{\nu}_0})}}, but not their normal vectors.
To provide some intuition for why $\ensuremath{\pazocal{P}}_{\ensuremath{\mathrm{AUD}}}$ is the minimal set of test functions, consider the case of quadratic utility. In this case, we know \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}\ if and only if \ensuremath{\ensuremath{\mu}_0}\ dominates \ensuremath{\ensuremath{\nu}_0}\ in the convex order. In other words, if and only if, for all concave functions $f$, the expected value of $f$ under \ensuremath{\ensuremath{\nu}_0}\ dominates that under \ensuremath{\ensuremath{\mu}_0}. Any concave function can be obtained as affine combinations of the functions $\min\{\ensuremath{\omega}-c,0\}$ and $\min\{c-\ensuremath{\omega},0\}$ for different values of $c\in\ensuremath{\mathbb{R}}$ border1991functional. In fact, testing that the expected value of these functions under \ensuremath{\ensuremath{\nu}_0}\ is greater than that under \ensuremath{\ensuremath{\mu}_0}\ is enough to determine whether \ensuremath{\ensuremath{\mu}_0}\ dominates \ensuremath{\ensuremath{\nu}_0}\ in the convex order. (The representation of the convex order via the comparison of the cumulative distributions of \ensuremath{\ensuremath{\nu}_0}\ and \ensuremath{\ensuremath{\mu}_0}\ comes from this finding.) The test functions $\ensuremath{\ensuremath{p}^\uparrow}_{\ensuremath{\ensuremath{\omega}^\star}}$ and $\ensuremath{\ensuremath{p}^\downarrow}_{\ensuremath{\ensuremath{\omega}^\star}}$ are the analogue of the $\min\{\ensuremath{\omega}-c,0\}$ and $\min\{c-\ensuremath{\omega},0\}$ functions in the case in which \ensuremath{d}\ is not the identity.
We illustrate (ref) in the context of (ref): \setcounter{example}{1}
\paragraph{Binary actions} When the DM\ only has two actions, assuming affine utility differences, together with $\ensuremath{\gamma}=1$ and $\ensuremath{\kappa}=0$, is without loss. Thus, (ref) characterizes the set of \ensuremath{\mathrm{BCE}}-consistent distributions for all decision problems with binary actions. We record the corresponding characterization in (ref) below, where we take advantage of the binary action assumption to provide more explicit expressions for the conditions in (ref).
When $\ensuremath{A}=\{\ensuremath{a}_1,\ensuremath{a}_2\}$, (ref) implies \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}\ if and only if for all $\ensuremath{\ensuremath{\omega}^\star}\in\ensuremath{\Omega}_+\equiv\{\ensuremath{\omega}\in\ensuremath{\Omega}:\ensuremath{d}(\ensuremath{\omega})>0\}$,
where in the above expressions, $\ensuremath{\omega}\geq\ensuremath{\ensuremath{\omega}^\star}$ and $\ensuremath{\omega}\leq\ensuremath{\ensuremath{\omega}^\star}$ signify states with higher and lower indices than \ensuremath{\ensuremath{\omega}^\star}, respectively.
We can refine the above expressions by figuring out the states \ensuremath{\ensuremath{\omega}^\star}\ that minimize the right-hand side of these inequalities. To this end, define
whenever these sets are nonempty. Otherwise, let $\ensuremath{\omega}_{\ensuremath{a}_1}(\ensuremath{\ensuremath{\mu}_0})=\ensuremath{\omega}_\ensuremath{I}$ when the first set is empty, and let $\ensuremath{\omega}_{\ensuremath{a}_2}(\ensuremath{\ensuremath{\mu}_0})=\ensuremath{\omega}_1$ when the second is empty. We note the following: First, in the above expressions, $\min$ and $\max$ are over the state indices. Second, unless the DM\ is indifferent between both actions at \ensuremath{\ensuremath{\mu}_0}, at least one of the two sets is nonempty. To see this, suppose that $\ensuremath{a}_1$ is uniquely optimal at the prior; hence, the first set is empty. Then, $\ensuremath{\omega}_1$ satisfies the conditions defining the set on the right-hand side.
To understand the definition of $\ensuremath{\omega}_{\ensuremath{a}_1}(\ensuremath{\ensuremath{\mu}_0})$, consider the information structures that maximize the probability the DM\ takes $\ensuremath{a}_1$. Intuitively, such an information structure should pool states below a threshold state, \ensuremath{\ensuremath{\omega}^\star}, and this threshold state is an element of $\ensuremath{\Omega}_+$. Then, $\ensuremath{\omega}_{\ensuremath{a}_1}(\ensuremath{\ensuremath{\mu}_0})$ is the smallest state such that the recommendation to take $\ensuremath{a}_1$ when states below $\ensuremath{\omega}_{\ensuremath{a}_1}(\ensuremath{\ensuremath{\mu}_0})$ are pooled is disobedient. In other words, the information structure that maximizes the probability of taking $\ensuremath{a}_1$ pools states strictly below $\ensuremath{\omega}_{\ensuremath{a}_1}(\ensuremath{\ensuremath{\mu}_0})$ with probability $1$, and $\ensuremath{\omega}_{\ensuremath{a}_1}(\ensuremath{\ensuremath{\mu}_0})$ with a probability determined by the DM's binding obedience constraints. The intuition for $\ensuremath{\omega}_{\ensuremath{a}_2}(\ensuremath{\ensuremath{\mu}_0})$ is similar. With these definitions, the characterization for the case of binary actions is as follows:
Whenever $\ensuremath{a}_1$ is optimal at the prior, the right-hand side of (ref) is $0$. Instead, when $\ensuremath{a}_2$ is optimal at the prior, the right-hand side of (ref) is 1. Consequently, when the DM\ is indifferent between both actions at the prior, all action distributions can be rationalized.
By reordering the states, decision problems with two-step utility differences satisfy increasing differences. Thus, decision problems with two-step utility differences are monotone and concave decision problems. Decision problems with two-step utility differences are pinned down by the values the vector $\ensuremath{d}(\ensuremath{a}_{\ensuremath{j}+1},\ensuremath{a}_\ensuremath{j},\cdot)$ takes on $\ensuremath{\omega}_1$ and $\ensuremath{\omega}_\ensuremath{I}$, and the state at which $\ensuremath{d}(\ensuremath{a}_{\ensuremath{j}+1},\ensuremath{a}_\ensuremath{j},\cdot)$ switches between those values. In what follows, we denote by $\ensuremath{\ensuremath{i}^\star}(\ensuremath{j})$ the highest index state at which $\ensuremath{d}(\ensuremath{a}_{\ensuremath{j}+1},\ensuremath{a}_\ensuremath{j},\cdot)$ coincides with $\ensuremath{\ensuremath{\underline{\ensuremath{d}}}_{j+1,j}}$. Concavity implies $\ensuremath{\ensuremath{i}^\star}(\ensuremath{j})$ is increasing in \ensuremath{j}.
The discrete analog of the absolute loss function is a limiting case of two-step utility differences.\footnote{To be sure, this special case does not satisfy the second requirement of concavity*. We can instead consider a perturbed decision problem with $\tilde{d}(\ensuremath{a}_{\ensuremath{j}+1},\ensuremath{a}_\ensuremath{j},\ensuremath{\omega})={d}(\ensuremath{a}_{\ensuremath{j}+1},\ensuremath{a}_\ensuremath{j},\ensuremath{\omega})-j\varepsilon$ for some very small $\varepsilon>0$. This perturbed decision problem has two-step utility differences, so our characterization below applies.}\footnote{yang2024monotone and kolotilin2024distributions characterize the distributions over posterior quantiles consistent with a prior distribution over the states.} To see this, let $\ensuremath{\Omega}_\ensuremath{j}$ denote the set of states in which action \ensuremath{j}\ is optimal, $\ensuremath{\Omega}_\ensuremath{j}=\{\ensuremath{\omega}:(\forall\ensuremath{a}\in\ensuremath{A})\ensuremath{u}(\ensuremath{a}_\ensuremath{j},\ensuremath{\omega})\geq\ensuremath{u}(\ensuremath{a},\ensuremath{\omega})\}$. Under (ref), these sets are ordered in that $\ensuremath{\Omega}_\ensuremath{j}$ lies to the left of $\ensuremath{\Omega}_{\ensuremath{j}+1}$\textemdash in terms of the state indices. We can then define
or $\ensuremath{u}(\ensuremath{a}_\ensuremath{j},\ensuremath{\omega})=-c|k-j|$ for $\ensuremath{\omega}\in\ensuremath{\Omega}_k$.
\paragraph{Test functions for two-step utility differences} (ref) characterizes the test functions under the assumption of two-step utility differences. In particular, we show the test functions for two-step utility differences are basically those in $\ensuremath{\pazocal{P}}_{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}}$.
For each $\ensuremath{j}\in\ensuremath{A}$, we define $\ensuremath{\ensuremath{p}^\uparrow}_\ensuremath{j}(\ensuremath{\omega})=\ensuremath{d}(\ensuremath{a}_{\ensuremath{j}+1},\ensuremath{a}_j,\ensuremath{\omega})$ and $\ensuremath{\ensuremath{p}^\downarrow}_\ensuremath{j}=-\ensuremath{d}(\ensuremath{a}_{\ensuremath{j}+1},\ensuremath{a}_{j},\ensuremath{\omega})$. Denote by $\ensuremath{\pazocal{P}}_2$ the set of vectors $\{\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\uparrow}_j,\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\downarrow}_j:j\in\{1,\dots,\ensuremath{J}-1\}\}$ and note it is a subset of $\ensuremath{\pazocal{P}}_{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}}$\textemdash it is the left-most set in (ref). To each of these vectors, associate the vectors $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{q}}}^\uparrow}_\ensuremath{j},\ensuremath{\ensuremath{\boldsymbol{\ensuremath{q}}}^\downarrow}_\ensuremath{j}\in\ensuremath{\ensuremath{\mathbb{R}}^\ensuremath{A}}$, defined as follows:
For instance, in the case of absolute loss, $\ensuremath{\ensuremath{q}^\uparrow}_\ensuremath{j}(\ensuremath{a}_k)=c\mathbbm{1}_{i^\star(k)>i^\star(j)}$ and $\ensuremath{\ensuremath{q}^\downarrow}_\ensuremath{j}(\ensuremath{a}_k)=-c\mathbbm{1}_{i^\star(k-1)<i^\star(j)}$.
(ref) reduces the question of whether \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}\ to checking $2(\ensuremath{|\ensuremath{A}|}-1)$ linear inequalities, together with $\ensuremath{|\ensuremath{\Omega}|}+1$ constraints implied by $\ensuremath{\ensuremath{\mu}_0}\in\ensuremath{\Delta(\ensuremath{\Omega})}$.
Two-step utility differences is a case in which (i) the vectors in $\ensuremath{\pazocal{P}}_{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}}$ are test functions, without restricting the cardinality of the state space, and (ii) the belief martingale equations are implied by either the payoff martingale equations or the non-negativity condition that \ensuremath{\boldsymbol{\ensuremath{\ensuremath{\mu}_0}}}\ is a distribution. We obtain (i) because of the simple structure of the utility differences. Implicitly, (ref) shows that under two-step utility differences no new extreme rays are generated when we intersect the normal cones of \ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}. Regarding (ii), fix a state $\ensuremath{\omega}_i$ and consider the belief martingale condition for that state:
Recall that $\ensuremath{\Omega}_j=\{\ensuremath{\omega}_{\ensuremath{\ensuremath{i}^\star}(j-1)+1},\dots,\ensuremath{\omega}_{\ensuremath{\ensuremath{i}^\star}(j)}\}$. If $i\neq\ensuremath{\ensuremath{i}^\star}(j)$ for any $j$, then the left-hand side of the above expression is $0$, so that the above equation reduces to the non-negativity constraint on $\ensuremath{\ensuremath{\mu}_0}(\ensuremath{\omega}_i)$. The same is true when $i=\ensuremath{\ensuremath{i}^\star}(j)$ for some $j$, but $\ensuremath{\Omega}_j$ is not a singleton. Instead, when $i=\ensuremath{\ensuremath{i}^\star}(j)$ for some $j$ and $\ensuremath{\Omega}_j$ is a singleton, then one can verify the above equation is implied by (ref).
We conclude this section with (ref), which summarizes the extension of our results to the case in which \ensuremath{\Omega}\ and \ensuremath{A}\ are compact, convex, Polish spaces and the first-order approach applies (cf. kolotilin2023persuasion). Readers interested in applications can jump to (ref), with little loss of continuity.
In this section, we apply our results (i) to elucidate comparative statics of the set of \ensuremath{\mathrm{BCE}}-consistent marginals ((ref)) and (ii) to study \ensuremath{\mathrm{BCE}}-consistency across decision problems ((ref)). (ref) further illustrates our results in the context of simple multi-agent settings.
Under the assumption of affine utility differences, we consider in this section changes to the prior or the utility function that preserve the rationalization of a given marginal distribution over actions.
\paragraph{Changes to the prior} Suppose \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}\ and let \ensuremath{\ensuremath{\ensuremath{\mu}_0}^\prime}\ denote another prior belief. When can we say $(\ensuremath{\ensuremath{\ensuremath{\mu}_0}^\prime},\ensuremath{\ensuremath{\nu}_0})$ are \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}? The following definition is key:
Note $\ensuremath{\ensuremath{\mu}_0}\circ\ensuremath{d}^{-1}$ is the distribution of payoffs\textemdash measured by the utility difference \ensuremath{d}\textemdash faced by the DM\ under the prior distribution \ensuremath{\ensuremath{\mu}_0}. When \ensuremath{\ensuremath{\ensuremath{\mu}_0}^\prime}\ is a \ensuremath{d}-mean-preserving spread of \ensuremath{\ensuremath{\mu}_0}, the DM\ faces a riskier payoff distribution under \ensuremath{\ensuremath{\ensuremath{\mu}_0}^\prime}\ than under \ensuremath{\ensuremath{\mu}_0}. Conversely, we can obtain the payoff distribution under \ensuremath{\ensuremath{\mu}_0}\ by a garbling of the payoff distribution under \ensuremath{\ensuremath{\ensuremath{\mu}_0}^\prime}. A fortiori, if \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}, so is $(\ensuremath{\ensuremath{\ensuremath{\mu}_0}^\prime},\ensuremath{\ensuremath{\nu}_0})$.
(ref) summarizes this discussion:
The result follows from the shape of the test functions $\ensuremath{\pazocal{P}}_{\ensuremath{\mathrm{AUD}}}$ and noting that if \ensuremath{\ensuremath{\ensuremath{\mu}_0}^\prime}\ is a \ensuremath{d}-mean preserving spread of \ensuremath{\ensuremath{\mu}_0}, then for all \ensuremath{\ensuremath{\omega}^\star}, $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\uparrow}_{\ensuremath{\ensuremath{\omega}^\star}}\ensuremath{\ensuremath{\ensuremath{\mu}_0}^\prime}\leq\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\uparrow}_{\ensuremath{\ensuremath{\omega}^\star}}\ensuremath{\ensuremath{\mu}_0}$ and $\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\downarrow}_{\ensuremath{\ensuremath{\omega}^\star}}\ensuremath{\ensuremath{\ensuremath{\mu}_0}^\prime}\leq\ensuremath{\ensuremath{\boldsymbol{\ensuremath{p}}}^\downarrow}_{\ensuremath{\ensuremath{\omega}^\star}}\ensuremath{\ensuremath{\mu}_0}$.
We illustrate (ref) with the following example:
\paragraph{Binary actions and payoff shifters} Consider now the case of binary actions, $\ensuremath{A}=\{\ensuremath{a}_1,\ensuremath{a}_2\}$, so that without loss, we can take $\ensuremath{u}(\ensuremath{a}_2,\ensuremath{\omega})-\ensuremath{u}(\ensuremath{a}_1,\ensuremath{\omega})=\ensuremath{d}(\ensuremath{\omega})$. Suppose we parameterize the utility differences by $\ensuremath{d}(\cdot,\ensuremath{\theta})$ such that $\ensuremath{\theta}<\ensuremath{\ensuremath{\theta}^\prime}$ implies $\ensuremath{d}(\cdot,\ensuremath{\theta})\leq\ensuremath{d}(\cdot,\ensuremath{\ensuremath{\theta}^\prime})$. That is, as we move from \ensuremath{\theta}\ to \ensuremath{\ensuremath{\theta}^\prime}, action $\ensuremath{a}_2$ becomes more attractive.
(ref) characterizes how the lower and upper bounds on $\ensuremath{\ensuremath{\nu}_0}(\ensuremath{a}_2)$ change as the utility differences shift in favor of the higher action.
The result is intuitive: as $\ensuremath{a}_2$ becomes more attractive, both the minimal probability with which the DM\ must take $\ensuremath{a}_2$ and the maximal probability with which the DM\ may take $\ensuremath{a}_2$ so that \ensuremath{\ensuremath{\nu}_0}\ is consistent with information are higher.
The result has the following implication: suppose \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given $\ensuremath{d}(\cdot,\ensuremath{\theta})$ for some real-valued parameter \ensuremath{\theta}. Then, \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent given $\ensuremath{d}(\cdot,\ensuremath{\ensuremath{\theta}^\prime})$ for all $\ensuremath{\ensuremath{\theta}^\prime}\in[\underline{\ensuremath{\theta}}_{\ensuremath{\ensuremath{\mu}_0}},\overline{\ensuremath{\theta}}_{\ensuremath{\ensuremath{\mu}_0}}]$, where $\mathrm{UB}(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\boldsymbol{\ensuremath{d}}}(\underline{\ensuremath{\theta}}_{\ensuremath{\ensuremath{\mu}_0}}))=\ensuremath{\ensuremath{\nu}_0}(\ensuremath{a}_2)$ and $\mathrm{LB}(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\boldsymbol{\ensuremath{d}}}(\overline{\ensuremath{\theta}}_{\ensuremath{\ensuremath{\mu}_0}}))=\ensuremath{\ensuremath{\nu}_0}(\ensuremath{a}_2)$. That is, for each \ensuremath{\ensuremath{\mu}_0}\ such that \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent for some parameter, we can identify an interval of parameter values for which \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ is \ensuremath{\mathrm{BCE}}-consistent.
We illustrate (ref) with (ref), based on bergemann2022counterfactuals:
\paragraph{Binary actions and ratio-ordered payoffs} Say the parameterized utility difference $\ensuremath{d}(\cdot,\ensuremath{\theta})$ is ratio ordered if for $\ensuremath{\omega}<\ensuremath{\ensuremath{\omega}^\prime}$, the ratio $\ensuremath{d}(\ensuremath{\omega},\ensuremath{\theta})/\ensuremath{d}(\ensuremath{\ensuremath{\omega}^\prime},\ensuremath{\theta})$ is increasing in \ensuremath{\theta}. We have the following result:
We illustrate (ref) with (ref):
Suppose the analyst observes the DM's choices across $N$ different decision problems, each indexed by a set of actions $\ensuremath{\ensuremath{A}_\ensuremath{n}}$ and utility function $\ensuremath{\ensuremath{u}_\ensuremath{n}}:\ensuremath{\ensuremath{A}_\ensuremath{n}}\times\ensuremath{\Omega}\to\ensuremath{\mathbb{R}}$. The analyst's data are now the joint distribution over actions across decision problems, which we denote by $\ensuremath{\bar{\ensuremath{\nu}}_0}\in\Delta(\ensuremath{A}_1\times\dots\ensuremath{A}_N)$. Our results so far allow us to understand whether $(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\bar{\ensuremath{\nu}}_0}|_{\ensuremath{\ensuremath{A}_\ensuremath{n}}})$ is \ensuremath{\mathrm{BCE}}-consistent in decision problem $\ensuremath{D}_\ensuremath{n}=\langle\ensuremath{\Omega},\ensuremath{\ensuremath{A}_\ensuremath{n}},\ensuremath{\ensuremath{u}_\ensuremath{n}}\rangle$, where $\ensuremath{\bar{\ensuremath{\nu}}_0}|_{\ensuremath{\ensuremath{A}_\ensuremath{n}}}$ is the marginal of \ensuremath{\bar{\ensuremath{\nu}}_0}\ over \ensuremath{\ensuremath{A}_\ensuremath{n}}. Instead, we now study when a single information structure exists that rationalizes the choices \ensuremath{\ensuremath{\nu}_0}\ made by the DM\ across all decision problems.
Consider now an auxiliary decision problem $\bar\ensuremath{D}=\langle\ensuremath{\Omega}, \bar\ensuremath{A},\bar u\rangle$, where $\bar\ensuremath{A}=\times_{i=1}^N\ensuremath{\ensuremath{A}_\ensuremath{n}}$. In this decision problem, choices are action profiles, $\ensuremath{a}\in\times_{\ensuremath{n}\in\ensuremath{N}}\ensuremath{\ensuremath{A}_\ensuremath{n}}$, and payoffs are $\bar u(\ensuremath{a},\ensuremath{\omega})=\sum_{\ensuremath{n}=1}^\ensuremath{N}\ensuremath{\ensuremath{u}_\ensuremath{n}}(\ensuremath{\ensuremath{a}_\ensuremath{n}},\ensuremath{\omega})$. The results in bergemann2022counterfactuals imply that for a given prior distribution \ensuremath{\ensuremath{\mu}_0}, an information structure exists that rationalizes the DM's choices \ensuremath{\bar{\ensuremath{\nu}}_0}\ across decision problems $\left(\ensuremath{D}_\ensuremath{n}\right)_{\ensuremath{n}\leq N}$ if and only if $(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\bar{\ensuremath{\nu}}_0})$ is \ensuremath{\mathrm{BCE}}-consistent in the auxiliary decision problem $\bar\ensuremath{D}$:
Consequently, the results in Sections (ref) and (ref) can be used to study \ensuremath{\mathrm{BCE}}-consistency across decision problems. We illustrate this result using a variation of (ref):
Note we can interpret the joint distribution over action profiles in this section as coming from a game between $N$ players, where player \ensuremath{n}\ has action set \ensuremath{\ensuremath{A}_\ensuremath{n}}\ and payoffs $\ensuremath{\ensuremath{u}_\ensuremath{n}}$. Consider now the question of whether we can find a public information structure that rationalizes \ensuremath{\bar{\ensuremath{\nu}}_0}\ as if the players observe the realization of the public information structure prior to non-cooperatively playing the game. As we show in (ref), (ref) characterizes the pairs $(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\bar{\ensuremath{\nu}}_0})$ that admit such rationalization. (ref) also applies our results to study ring-network games kneeland2015identifying.
Whereas the results in the previous sections characterize the set of \ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}\ that are \ensuremath{\mathrm{BCE}}-consistent given \ensuremath{u}, they remain silent as to the set of information structures that rationalize the marginal \ensuremath{\ensuremath{\nu}_0}\ given the DM's prior \ensuremath{\ensuremath{\mu}_0}\ and utility function \ensuremath{u}.\footnote{In a sense, this observation is consistent with the paper's motivation: the applied literature treats information as a nuisance parameter and hence, is not necessarily interested in estimating the (parameters of the) information structure.} In this section, we tackle this problem in the single-agent setting, where an information structure can be identified with the distribution over posteriors $\ensuremath{\tau}\in\Delta(\ensuremath{\Delta(\ensuremath{\Omega})})$ it induces kamenica2011bayesian. Given $\ensuremath{(\ensuremath{\ensuremath{\mu}_0},\ensuremath{\ensuremath{\nu}_0})}$ and \ensuremath{u}, we characterize in this section the set of distributions over posteriors (if any) that rationalize the DM's distribution over actions. Without loss of generality, the analysis that follows assumes the distribution over posteriors \ensuremath{\tau}\ has finite support myerson1982optimal,kamenica2011bayesian; we denote the support of \ensuremath{\tau}\ by $\mathrm{supp}\,\ensuremath{\tau}$.
For a distribution over posteriors \ensuremath{\tau}\ to rationalize \ensuremath{\ensuremath{\nu}_0}, two conditions must be satisfied. First, the mean of \ensuremath{\tau}\ must equal the prior, \ensuremath{\ensuremath{\mu}_0}; that is, \ensuremath{\tau}\ must be Bayes plausible. Second, the distribution over posteriors \ensuremath{\tau}\ must induce the distribution over actions \ensuremath{\ensuremath{\nu}_0}. Formally, let $\ensuremath{a}^*(\ensuremath{\mu})$ denote the DM's optimal set of actions when their belief is \ensuremath{\mu}. Then, \ensuremath{\tau}\ induces \ensuremath{\ensuremath{\nu}_0}\ if a decision rule $\ensuremath{\alpha}:\ensuremath{\Delta(\ensuremath{\Omega})}\to\Delta(\ensuremath{A})$ exists such that for every $\ensuremath{a}\in\ensuremath{A}$,
where $\ensuremath{a}\in\text{ supp }\ensuremath{\alpha}(\ensuremath{\mu})(\cdot)$ only if $\ensuremath{a}\in\ensuremath{a}^*(\ensuremath{\mu})$.
Whereas in the previous section we interpreted the marginal distribution \ensuremath{\ensuremath{\nu}_0}\ as a Bayes plausible distribution over posteriors, (ref) shows this analogy is perhaps incomplete. Indeed, whenever the distribution over posteriors \ensuremath{\tau}\ induces beliefs such that $\ensuremath{a}^*(\ensuremath{\mu})$ is not a singleton, specifying the DM's tie-breaking rule \ensuremath{\alpha}\ is necessary to determine whether the frequency with which the DM\ takes actions under \ensuremath{\tau}\ matches that under \ensuremath{\ensuremath{\nu}_0}.
\paragraph{A demand and supply problem} The problem of determining whether a decision rule \ensuremath{\alpha}\ exists satisfying (ref) admits the following interpretation (cf. gale1957theorem): beliefs $\ensuremath{\mu}\in\mathrm{supp}\,\ensuremath{\tau}$ are supplied in quantities $\ensuremath{\tau}(\ensuremath{\mu})$ and demanded by actions $\ensuremath{a}\in\ensuremath{A}$ in quantities $\ensuremath{\ensuremath{\nu}_0}(\ensuremath{a})$. The demand $\ensuremath{\ensuremath{\nu}_0}(\ensuremath{a})$ can only be satisfied by certain beliefs\textemdash those that satisfy $\ensuremath{\mu}\in\ensuremath{\ensuremath{\ensuremath{\Delta}_\ensuremath{u}^*}(\ensuremath{a})}$. The decision rule \ensuremath{\alpha}\ describes how much of a given belief $\ensuremath{\mu}\in\mathrm{supp}\,\ensuremath{\tau}$ is allocated to action \ensuremath{a}. That \ensuremath{\tau}\ implements \ensuremath{\ensuremath{\nu}_0}\ is equivalent to being able to allocate the supply of beliefs to satisfy the action demands in a market-clearing way.
As we argue in (ref), the above is an instance of the supply and demand problem studied in gale1957theorem, with the distributions \ensuremath{\ensuremath{\nu}_0}\ and \ensuremath{\tau}\ determining the demanded and supplied quantities, respectively. Building on the main theorem in that paper, (ref) below characterizes when \ensuremath{\tau}\ implements \ensuremath{\ensuremath{\nu}_0}. To state (ref), one final piece of notation is needed. Given a Bayes plausible $\ensuremath{\tau}\in\Delta(\ensuremath{\Delta(\ensuremath{\Omega})})$, we construct a measure over subsets \ensuremath{B}\ of the set of actions \ensuremath{A}\ as follows. For each $\ensuremath{B}\subseteq\ensuremath{A}$, define the push-forward measure $\ensuremath{\tau}_\ensuremath{A}(\ensuremath{B})$ as
In words, each action subset \ensuremath{B}\ has mass equal to the probability that \ensuremath{\tau}\ induces a belief under which \ensuremath{B}\ is the set of optimal actions. (ref) characterizes the distribution over posteriors that implement \ensuremath{\ensuremath{\nu}_0}\ when the DM's prior is \ensuremath{\ensuremath{\mu}_0}\ and their utility function is \ensuremath{u}:
To interpret (ref), note the following. The left-hand side is the probability under which the agent takes some action \ensuremath{a}\ in the set \ensuremath{B}. Instead, the right-hand side is the probability under which the agent finds some action in the set \ensuremath{B}\ optimal, but no action that is not in \ensuremath{B}. (ref) then states the frequency with which the agent takes actions in \ensuremath{B}\ has to be at least the frequency with which an action in \ensuremath{B}\ is optimal.\footnote{(ref) is intimately connected to the Border-Matthews-Maskin-Riley characterization of reduced-form implementation in auctions. The latter states that a reduced-form auction (a collection of interim probabilities of trade for each buyer) has an auction implementation if and only if for all subsets of buyer-type profiles, the probability the reduced form auction allocates the good to types in that set is no larger than the prior probability of buyer types in that set. Letting $\ensuremath{\tau}_\ensuremath{A}$ and \ensuremath{\ensuremath{\nu}_0}\ play the role of the reduced-form auction and of the type distribution, respectively, (ref) is morally related to the Border-Matthews-Maskin-Riley inequalities.}
The proof of (ref) is in (ref) and follows from three steps. First, we show how to map our problem into that in gale1957theorem. Second, whereas gale1957theorem allows for the supply-demand equations to hold as weak inequalities, we show that any solution to Gale's problem clears the market exactly and, hence, satisfies (ref). This result follows from \ensuremath{\ensuremath{\nu}_0}\ and \ensuremath{\tau}\ being probability distributions. gale1957theorem refers to such solutions as maximal flows. Finally, we show the necessary and sufficient condition for the existence of a feasible and maximal flow in gale1957theorem is equivalent to (ref).
\paragraph{Connection with stochastic choice} We now draw a connection with stochastic choice from menus, which, among other things, allows us to illustrate why (ref) implies \ensuremath{\tau}\ implements \ensuremath{\ensuremath{\nu}_0}.
It follows from the results in gale1957theorem that (ref) implies a system of conditional probabilities $\{\ensuremath{\sigma}(\cdot|\ensuremath{B})\in\Delta(\ensuremath{A}):\ensuremath{B}\subset\ensuremath{A}\}$ exists such that (i) for all $\ensuremath{B}\subseteq\ensuremath{A}$, $\ensuremath{\sigma}(\ensuremath{B}|\ensuremath{B})=1$ and (ii)
The conditional probabilities $\left(\ensuremath{\sigma}(\cdot|\ensuremath{B})\right)_{\ensuremath{B}\subseteq\ensuremath{A}}$ can be interpreted as the DM's stochastic choice from the menus $\{\ensuremath{B}:\ensuremath{B}\subseteq\ensuremath{A}\}$. (ref) states that the probability that the DM\ chooses action \ensuremath{a}\ under marginal is the probability that the agent faces a menu \ensuremath{B}\ that has \ensuremath{a}\ available and the DM\ chooses \ensuremath{a}\ out of \ensuremath{B}.\footnote{ (ref) is another demand and supply problem, where $\ensuremath{\tau}_\ensuremath{A}$ is the supply of action subsets. The results in gale1957theorem imply (ref) is equivalent to the existence of $\ensuremath{\sigma}$. The existence of $\ensuremath{\sigma}$ can also be established using azrieli2022marginal's extension of Hall's marriage theorem (see their Proposition 9).}
Consider now the following “experiment”: we first draw a menu $\ensuremath{B}$ using $\ensuremath{\tau}_\ensuremath{A}$ and then draw an action $\ensuremath{a}\in\ensuremath{B}$ according to $\ensuremath{\sigma}(\cdot|\ensuremath{B})$. We only inform the DM\ of the drawn action, and not the menu from which it was drawn. Because we draw menu \ensuremath{B}\ only when it is the optimal set of actions, we only recommend \ensuremath{a}\ when following the recommendation is optimal for the DM. As we explain below, (ref) then implies this “experiment” induces the DM\ to take actions with the desired frequency.
Technically, we have not described an experiment\textemdash a collection of signal distributions conditional on the state of the world\textemdash but one can do so immediately as follows: for each $\ensuremath{\omega}\in\ensuremath{\Omega}$ and $\ensuremath{a}\in\ensuremath{A}$,
In this experiment, the DM\ receives an action recommendation conditional on the state of the world, so that (ref) describes the DM's state-dependent stochastic choice.
The previous discussion connects two sets of conditional distributions over choices that arise in the stochastic choice literature: stochastic choices conditional on a state of the world ((ref)) and stochastic choices out of a menu\textemdash$\ensuremath{\sigma}$ in (ref). Indeed, the measure $\ensuremath{\tau}_\ensuremath{A}$ can be interpreted as the frequency with which the agent faces different menus\textemdash action subsets in this case\textemdash whereas the measure \ensuremath{\ensuremath{\nu}_0}\ represents the frequency with which the agent makes different choices. In other words, the pair $(\ensuremath{\tau}_\ensuremath{A},\ensuremath{\ensuremath{\nu}_0})$ is analogous to the dataset in azrieli2022marginal. Whereas they show the core condition in (ref) characterizes the existence of $\left(\ensuremath{\sigma}(\cdot|\ensuremath{B})\right)_{\ensuremath{B}\subseteq\ensuremath{A}}$ in (ref) given $(\ensuremath{\tau}_\ensuremath{A},\ensuremath{\ensuremath{\nu}_0})$, we show the core condition characterizes the set of Bayes plausible distribution over posteriors that induce \ensuremath{\ensuremath{\nu}_0}.
{\singlespacing}