EconBase
← Back to paper

Improving Robust Decisions with Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

84,389 characters · 15 sections · 46 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Improving Robust Decisions with Data

abstractA decision-maker faces uncertainty governed by a data-generating process (DGP), which is only known to belong to a set of sequences of independent but possibly non-identical distributions. A robust decision maximizes the expected payoff against the worst possible DGP in this set. This paper characterizes when and how such robust decisions can be objectively improved with data --- that is, yield higher expected payoffs under the true DGP regardless of which DGP is the truth. It further develops simple and novel inference procedures that achieve such improvement, while common methods (e.g., maximum likelihood) may fail to do so. JEL: D81, D83, C44 Keywords: Maxmin expected utility, $\Gamma$-minimax, learning under ambiguity, statistical decisions

Introduction

When a decision-maker (DM) lacks sufficient knowledge about the probability law governing an uncertain environment, they may want their decisions to be robust. This concern for robustness is often modeled by ranking choices by their worst-case expected payoffs among a set of possible laws 10.1214/aoms/1177698602, GILBOA1989141, HansenSargent+2007, Carroll2019. With insufficient knowledge, gathering data is a classic means of learning about the environment to help guide decisions. Here, learning involves using data to revise the set of possible laws. A revision with data is said to provide an objective improvement if the data-revised robust decision yields a higher expected payoff under the true law than the robust decision without any data-revision, the benchmark robust decision. This paper studies when and how data may be used to achieve such an objective improvement regardless of which possible law is the true one.

The analysis is conducted in a decision environment where it is a priori unclear whether and how data can generate objective improvement --- one in which the true law cannot be uniquely identified even asymptotically. Specifically, the DM faces a countable sequence of random experiments (e.g., coin flips) that share the same set of outcomes. The probability law governing these realizations, referred to as the data-generating process (DGP), is a sequence of independent but possibly non-identical distributions over the outcomes. The DM initially knows there is a set of possible DGPs that contains the true one, observes sample data given by realizations of some experiments, and then makes a decision whose payoff depends only on future outcomes.

This non-identical decision environment captures settings with unobserved heterogeneity. For example, consider an online platform, say Netflix, making content recommendations to users based on feedback from other users with similar profiles. Users' preferences, however, are usually also determined by their various offline activities that cannot be observed, which often further affect how they interact with the platform. As a result, the sequential user feedback data collected by Netflix can be viewed as generated by a sequence of independent but possibly non-identical distributions.\footnote{See Cao2014 for a discussion on the issue of non-IIDness in online recommender systems.} In this case, Netflix is motivated to use its data in a robust manner to guard against the concern that the data could be more from one type of user, whereas future users that determine payoffs could be more of a different type.

Regardless of what data Netflix/the DM observes, they can always rely on their initial belief to make the benchmark decision, which is robust against all initially possible DGPs. Notice, because the experiments are independent, the benchmark decision does not depend on the data at all.\footnote{With independence, the set of conditional distributions over future outcomes is solely determined by the set of DGPs regardless of the realized outcomes. Independence is assumed because this fact eases the exposition. In more general environments, the benchmark should be understood as the robust decision under full Bayesian or prior-by-prior updating pires2002rule. Namely, the DM applies Bayes' rule to update every DGP in the initial set conditioning on the sample data to obtain a set of posteriors over future outcomes. In this case, all results either directly apply or generalize under standard conditions, see Appendix (ref).} As a result, the benchmark decision ensures robustness, but at the cost of ignoring any information that might lead to better decisions.

The goal of this paper is to develop a strategy for using data that addresses both the concern for robustness and the desire for improvement. To preview the final output, the proposed strategy ensures that, in a broad class of decision problems, no matter which DGP in the initial set is true, a sufficient amount of data will almost surely lead the data-revised decision to objectively improve upon the benchmark decision, and sometimes strictly so. With finite samples, the same holds with a pre-specified probability. Hence, this strategy enables the DM to achieve higher expected payoffs under the truth while preserving robustness, i.e., failing to do so only with either zero or a small pre-specified probability.

What must be true for such a strategy? The following example illustrates a key observation of this paper.

example[Introductory Example] Suppose Netflix is deciding how to recommend a movie to a population of users who might either like it (thumbs-Up) or dislike it (thumbs-Down). Let $\{U, D\}$ denote these two possible outcomes observed after a recommendation. Let full state space be $\{U, D\}^{\infty}$, i.e., the infinite Cartesian product of these outcomes. Denote any probability distribution over $\{U, D\}$ by the probability of $U$. A data-generating process, $P$, can be written as $P = P_{1} \times P_{2} \times \cdots$ with $P_{i} \in [0,1]$. Suppose Netflix's initial knowledge about the users is given by the set $\{(1/3)^{\infty}\} \cup \{3/5, 1\}^{\infty}$, where $(1/3)^{\infty}$ denotes an i.i.d. sequence with marginal probability $1/3$ and the other set is defined as \begin{equation*} \{3/5, 1\}^{\infty} \equiv \{P: P_{i} \in \{3/5, 1\}\}. \end{equation*} Intuitively, users share some similarities: they either all like the movie with the same low probability $1/3$, or all like it with high but possibly heterogeneous probabilities, $3/5$ or $1$. Suppose Netflix observes feedback from having recommended the movie to $N$ users and must decide how to recommend it to future users. Decision Problem \expandafter\@slowromancap\romannumeral 1@. Suppose Netflix can recommend the movie either aggressively (a) or mildly (m).\footnote{For instance, the movie can be recommended either on top of the front page or further down the list. All payoffs can be interpreted as cardinal measures of how much the recommendation affects users' likelihood of interacting with Netflix.} The aggressive recommendation yields a higher payoff than a mild one if users like the movie, but is worse otherwise. Specifically, let $a(U) = 2$, $a(D) = -1$, $m(U) = 1/2$, and $m(D) = 0$ be the state-contingent payoffs. Within the initial set of DGPs, the worst-case DGP for both types of recommendation is $(1/3)^{\infty}$, since it gives the lowest probability of getting a $U$. Thus, the benchmark decision is to choose $m$. With the sample data, Netflix can make a data-revised decision by revising the initial set of DGPs. A commonly used method is maximum likelihood updating, which revises the initial set to the subset of DGPs that maximize the likelihood of observing the data.\footnote{Proposed and axiomatized in GILBOA199333 and CHENG2022102587, this updating rule is applied in Epstein2007 for studying problems in a similar setting.} When the true DGP is $(1/3)^{\infty}$, with $N$ sufficiently large, Netflix will almost surely observe sample data with an empirical frequency of $U$ close to $1/3$. The likelihood of any such sample, however, is never maximized by the true DGP in the initial set. Indeed, given any such sequence of outcomes, one can always find a DGP in the set $\{3/5, 1\}^{\infty}$ that assigns probability 1 to outcome $U$ whenever $U$ is observed and probability $3/5$ to outcome $U$ whenever $D$ is observed. In this case, notice that for all $N$, \begin{equation*} (1/3)^{N/3} \times (2/3)^{2N/3} < 1^{N/3} \times (2/5)^{2N/3}, \end{equation*} i.e., the likelihood of observing the given sample under this heterogeneous DGP is always strictly higher than under the true one. Thus, with maximum likelihood updating, Netflix will asymptotically almost surely rule out the true DGP from their data-revised set.\footnote{While the present paper, to my knowledge, is the first to make such an observation in the context of maximum likelihood updating, it is intrinsically related to the infamous incidental parameter problem in making estimates using maximum likelihood, discovered by Neyman1948 (see LANCASTER2000391 for a review), and more broadly, the problem of overfitting. The same observation continues to hold for likelihood-based updating rules, such as those proposed in Epstein2007 and CHENG2022102587.} As the worst-case probability of $U$ in the data-revised set increases to $3/5$, the aggressive recommendation becomes Netflix's data-revised decision. However, this decision gives Netflix an expected payoff of $0$ under the true DGP, which is strictly lower than the benchmark $(1/6)$. Therefore, with maximum likelihood updating, Netflix will almost surely be worse off than under their benchmark decision when the true DGP is $(1/3)^{\infty}$.

The preceding example shows that ruling out the true DGP can lead to objectively worse decisions. Theorem (ref) formalizes this observation by showing that accommodating (i.e., containing, with the necessary technical generalizations) the true DGP when revising the initial set is necessary for guaranteeing objective improvements. This necessity holds even when restricting attention to basic decision problems, those consist of two alternatives, one of which yields a constant payoff. Since likelihood-based rules fail this criterion, this paper develops a simple empirical distribution method that accommodates the truth, as illustrated below. In the introductory example, this also proves sufficient for objective improvement.

example[Introductory Example Continued] Under the empirical distribution method, Netflix revises the initial set to include DGPs whose average mixture of sample marginals, i.e., $1/N\sum_{i=1}^{N} P_{i}$ in this example, is close enough to the empirical distribution of observed outcomes. Kolmogorov's strong law of large numbers ensures that the true DGP will be retained in such data-revised sets asymptotically almost surely. More importantly, whenever the true DGP is retained, the data-revised decision is $m$ if the true DGP is $(1/3)^{\infty}$, and $a$ if the true DGP belongs to the set $\{3/5, 1\}^{\infty}$. In all cases, it is either the same or strictly better than the benchmark decision, thereby achieving objective improvement.

Notice that in Decision Problem \expandafter\@slowromancap\romannumeral 1@, both alternatives rank the two states $U$ and $D$ the same in terms of payoffs. Hence, a higher probability of $U$ leads to a higher expected payoff under each alternative. I define a decision problem (i.e., a set of alternatives) that possesses this property as a monotone decision problem.\footnote{This terminology stems from the definition of a “monotone decision problem" in ATHEY2018101.} More specifically, all alternatives in a monotone decision problem are positive affine transformations of a common payoff function. In this class of problems, accommodating the truth is both necessary and sufficient for achieving objective improvement (Theorem (ref)). Monotone decision problems arise naturally in economic contexts; Section (ref) further illustrates how this result may be applied to canonical principal-agent models.

However, this sufficiency does not extend beyond monotone decision problems (Theorem (ref)). In other words, such monotonicity is the property that characterizes when accommodating the truth suffices for objective improvement.

In searching for objective improvements in all decision problems, Theorem (ref) and Corollary (ref) together provide an impossibility result: It is achieved if and only if the true DGP is uniquely identified from the data. Such identification, however, is generally infeasible in non-identical environments.\footnote{As in the example, for any sample size $N$, there always exist multiple DGPs in the set $\{3/5, 1\}^{\infty}$ such that they are the same up to the $N$-th experiment, but have different marginals over future outcomes.} Nevertheless, accommodating the truth still provides a weaker yet meaningful guarantee of objective improvement in all decision problems (Proposition (ref)). In particular, it ensures that the decision maker would never prefer to forgo the opportunity to revise decisions using data in exchange for the certainty-equivalent payoff of the benchmark decision.

To summarize, all these results jointly provide three decision-theoretic justifications for accommodating the truth in data revision: (1) it is necessary for objective improvements, (2) it is sufficient in monotone decision problems; and (3) it provides a weaker but still useful guarantee in all decision problems. The paper then proceeds to develop practical methods that accommodate the truth in non-identical environments.

First, I show that the simple empirical distribution method accommodates the truth asymptotically almost surely (Theorem (ref)).

Additional concerns arise when considering data samples of bounded size. The standard finite-sample approach is to develop methods that accommodate the truth with a pre-specified probability. This entails making statistical inferences from independent but non-identical samples, which can be computationally challenging. To address this difficulty, I develop an easy-to-implement method, the augmented i.i.d. test, which operates as a simple augmentation to the standard practice of constructing confidence regions from i.i.d. data. Theorem (ref) establishes that this method accommodates the truth with at least the required pre-specified probability. For a concrete illustration, Section (ref) presents a Bernoulli model with ambiguous nuisance parameters and shows that the data-revised sets given by the augmented i.i.d. test are simple extensions of the Wilson Confidence Intervals Wilson1927.

Similar non-identical decision environments have been adopted in the literature to study, for instance, social learning with heterogeneous individuals reshidi2020information, Chen2019, dynamic portfolio choice under ambiguous idiosyncratic shocks Epstein2007, and estimating parameters of incomplete models Epstein2016. In all these models, the revision rules developed in this paper can provide new perspectives and often distinct predictions. As an illustration, in Section (ref), I consider the model in reshidi2020information. They assume the DM applies prior-by-prior updating and show there can be non-vanishing ambiguity. Proposition (ref) here shows all that ambiguity vanishes asymptotically if the DM instead applies the revision rules proposed here. In fact, learning under the new revision rules is more effective compared to commonly used updating rules and is guaranteed to be correct.

Finally, I briefly summarize the main contributions of this paper to the literature on robust statistical decisions wald1950statistical, Watson2016, Hansen2016, manski2021econometrics. A more detailed discussion is provided in Section (ref). Most papers in this literature focus on only the “data-revised decision” in their decision context. The main innovation here is to explicitly compare the data-revised decision with the benchmark decision that could have been made without using data. This comparison turns out to be not always obvious. Indeed, I show that it is impossible to guarantee the data-revised decision to be always better whenever the true DGP is not uniquely identified. In light of this impossibility, the present paper contributes to the literature by (1) characterizing when the simple criterion of accommodating the truth suffices for objective improvement and (2) providing practical inference methods that achieve it in both asymptotic and finite-sample settings.

Outline. Section (ref) introduces the decision environment. Section (ref) characterizes when and how objective improvements can be achieved. Section (ref) defines and studies the proposed revision rules. Section (ref) presents two applications with parametric models. Section (ref) provides a further discussion on related literature. All proofs are collected in the Appendix.

The Decision Environment

There is a countably infinite sequence of random experiments that all have the same finite set of outcomes $S$ with generic element $s$.\footnote{The finiteness assumption is made for simplicity. Section (ref) and Appendix (ref) provide discussions on how to extend the results in this paper to infinite state spaces.} The experiments are ordered and indexed by the set $\mathbb{N} = \{1,2,\cdots\}$. The full state space is denoted by $\Omega = S^{\infty} \equiv \prod_{i} S_{i}$ with generic element $\omega$. Let $\Sigma$ denote the discrete $\sigma$-algebra on $S$ and $\Sigma^{\infty}$ the product $\sigma$-algebra on $\Omega$.

For any given sample size $N \in \mathbb{N}$, let $S^{N} \equiv \prod_{i=1}^{N}S_{i}$ denote the sample states. A decision-maker (DM) observes sample data $\omega^{N} \in S^{N}$ and makes a decision whose payoff depends only on future unrealized experiments. Let $\Omega_{N} \equiv \prod_{i=N+1}^{\infty} S_{i} \equiv S_{N}^{\infty}$ denote the future states with generic element $\tilde{\omega}_{N}$. Define $\Sigma_{N}$ and $\Sigma_{N}^{\infty}$ similarly.

The DM's decision is the choice from a set of alternatives, named acts. An act is defined as a bounded $\Sigma_{N}^{\infty}$-measurable function, $f: \Omega_{N} \rightarrow \mathbb{R}$, that maps future states to payoffs (measured in utilities). Let $\mathcal{F}$ be the space of all acts endowed with the product topology.\footnote{For simplicity, I suppress the dependence of $\mathcal{F}$ on $N$. This is without loss of generality since any act on $\Omega_{N}$ can be isomorphically defined on $\Omega_{M}$ for all $M \in \mathbb{N}$.} An act is finitely-based if it depends only on finitely many experiments (i.e., not on tail events). Let $x \in \mathbb{R}$ denote a constant act that gives the same payoff $x$ in all future states. A decision problem $D$ is a nonempty and compact subset of $\mathcal{F}$. Let $\mathcal{D}$ denote the collection of all decision problems. Say that a decision problem $D$ is binary if $|D| = 2$.

I call a sequence of independent but possibly non-identical distributions over outcomes a data-generating process (DGP). Formally, let $\Delta(\Omega)$ denote the set of all countably additive probability measures on $(\Omega,\Sigma^\infty)$, endowed with the topology of setwise convergence. The collection of all DGPs is the subset $\Delta_{\text{indep}}(\Omega)= \prod_{i=1}^\infty \Delta(S_i)\subseteq\Delta(\Omega)$. On $\Delta_{\text{indep}}(\Omega)$, the subspace topology inherited from $\Delta(\Omega)$ under setwise convergence coincides with the product topology on $\prod_{i=1}^\infty \Delta(S_i)$. For each $P \in \Delta_{\text{indep}}(\Omega)$, write $P_i$ for its marginal distribution on $S_i$, $P^N$ for its joint marginal on $S^N$, and $P_N^\infty$ for its marginal on $\Omega_N = S_N^{\infty}$.

Suppose the DM knows there is an initial set $\mathcal{P}$ of possible DGPs. When fixing some sample data $\omega^{N}$, let $P^{*}$ denote the true DGP that generates the data and governs the future states. The DM's initial knowledge is “correctly specified”, meaning that it always contains the true DGP. Using the terminology from Bayesian learning literature Kalai1993, the DM's initial knowledge contains a “grain of truth”:

assumption$P^{*} \in \mathcal{P}$.

What this assumption also means is that every DGP in the initial set could be the true one governing the experiments. In addition, assume the set $\mathcal{P}$ is compact and all $P \in \mathcal{P}$ have full support.\footnote{This is a standard assumption for studying the update of a set of probability measures. It is also used when establishing a Central Limit Theorem in this environment. See the proof of Lemma (ref) for details.} Let $\mathcal{P}_{i}$, $\mathcal{P}^{N}$ and $\mathcal{P}_{N}^{\infty}$ denote the sets of their corresponding marginals.

Given a decision problem $D$, the DM can make a robust decision by using the initial set of DGPs. Formally, let the benchmark decision, denoted by $c(D) \in D$, be given by

equation*[equation* omitted — 339 chars of source]

where $\overline{co}(\mathcal{P}_{N}^{\infty})$ denotes the closed and convex hull of $\mathcal{P}_{N}^{\infty}$. The “min” is well-defined as $\mathcal{P}$ is compact. The equality follows since the minimum can be achieved at an extreme point and thus belongs to $\mathcal{P}$. Moreover, $c(D)$ is defined as a singleton by imposing an arbitrary (but fixed) tie-breaking rule.

With sample data $\omega^{N}$, the DM can also choose to revise their initial set to a data-revised set of DGPs. Let $\mathcal{P}(\omega^{N}) \subseteq \Delta_{indep}(\Omega)$ denote the revised set. Given Assumption (ref), further let $\mathcal{P}(\omega^{N}) \subseteq \mathcal{P}$ as there is no need to consider DGPs outside $\mathcal{P}$. Importantly, the data-revised set is assumed to depend only on the initial set $\mathcal{P}$ and sample data $\omega^{N}$ but not on $D$, i.e., the specific decision problem considered. In other words, the DM's revision rule is purely “data-based”. In the literature on belief updating, this is known as consequentialism, see hanany2007updating, hanany2009updating and Siniscalchi2009-SINTOO-5 for discussions.

Let $c(D, \omega^{N}) \in D$ denote the DM's data-revised decision given (the closed and convex hull of) their data-revised set $\mathcal{P}(\omega^{N})$, i.e.,

equation*[equation* omitted — 227 chars of source]

where the “min” is well defined as $\overline{\mathrm{co}}\big(\mathcal{P}(\omega^{N})_{N}^{\infty}\big)$ is compact and $c(D,\omega^{N})$ is also defined to be a singleton by imposing a tie-breaking rule, subject to the following consistency requirement:

assumptionIf $c(D)$ is among the maximizers in the data-revised program, then $c(D,\omega^{N})=c(D)$. Otherwise, the tie-breaking can be arbitrary.

Assumption (ref) makes sure that if the DM's data-revised set coincides with their initial set, then their data-revised decision is the same as the benchmark. This rules out uninteresting complications when the DM's decisions are different only because of different tie-breakings.\footnote{Also notice both decisions are defined as deterministic choices from $D$, which seems to rule out randomizations from the DM's decisions. This is, in fact, more general as one can explicitly add those randomizations as acts and it will just be another well-defined decision problem. The current definition allows for richer decision patterns: By imposing different tie-breaking rules, it allows the DM to have different preferences in terms of whether randomization hedges against ambiguity. See Saito2015 and Ke2020 for discussions and characterizations of such preferences.}

Importantly, the DM's objective payoff from an act is determined by the true DGP $P^{*}$ that actually governs the future. To make this explicit, let

equation*[equation* omitted — 108 chars of source]

denote the expected payoff of act $f$ evaluated under the true DGP $P^{*}$. For any decision problem $D$, the DM's objective payoffs from their benchmark and data-revised decisions are then $W(c(D), P^{*}) $ and $W(c(D,\omega^{N}), P^{*})$, respectively.

Because the DM's decisions depend only on future states and are made using the robust criterion, it is useful to define the following notions that suitably generalize set inclusions:

definitionThe data-revised set $\mathcal{P}(\omega^{N})$ accommodates a DGP $P$ if \begin{equation*} P_{N}^{\infty} \in \overline{co}(\mathcal{P}(\omega^{N})_{N}^{\infty}). \end{equation*} The data-revised set $\mathcal{P}(\omega^{N})$ refines the initial set $\mathcal{P}$ if \begin{equation*} \overline{co}(\mathcal{P}(\omega^{N})_{N}^{\infty}) \subsetneqq \overline{co}(\mathcal{P}_{N}^{\infty}). \end{equation*} Say that $\mathcal{P}(\omega^{N})$ is a truth-accommodating refinement if it accommodates the true DGP and refines the initial set.

Notice that if $P^{*} \in \mathcal{P}(\omega^{N})$, then $\mathcal{P}(\omega^{N})$ accommodates $P^{*}$. Moreover, $\mathcal{P}(\omega^{N})$ refines $\mathcal{P}$ only if $\mathcal{P}(\omega^{N})$ is a proper subset of $\mathcal{P}$. In practice, since a convex combination of DGPs in $\Delta_{indep}(\Omega)$ remains in $\Delta_{indep}(\Omega)$ only when the DGPs share identical marginals in all but one experiment, a truth-accommodating refinement is thus essentially $(P^{*})_{N}^{\infty} \in \mathcal{P}(\omega^{N})_{N}^{\infty} \subsetneqq \mathcal{P}_{N}^{\infty}$.

Objective Improvements

Fix an initial set $\mathcal{P}$, some sample data $\omega^{N}$ generated by the true DGP $P^{*} \in \mathcal{P}$. This section investigates what guarantees objective improvements across different decision problems. To this end, consider the following definition.

definitionLet $\mathcal{C} \subseteq \mathcal{D}$ denote a class of decision problems. A data revision is said to provide objective improvement in $\mathcal{C}$ if, for all $D \in \mathcal{C}$, \begin{equation*} W(c(D, \omega^{N}), P^{*}) \geq W(c(D), P^{*}), \end{equation*} and the inequality is strict for some $D \in \mathcal{C}$.

Among all decision problems, the simplest possible form is the choice between an uncertain and a constant act. Such a canonical form is often used to model, for example, the decision of whether or not to approve a new drug, implement a new policy, convict a defendant, and invest in an asset. I call such decision problems, i.e., binary with a constant act, basic decision problems. Arguably, basic decision problems are the building blocks of more complex decision problems, thus any general enough class of decision problems should include them as special cases. Guaranteeing objective improvements at least in all basic decision problems is a reasonable minimal requirement for any data revision. The following result identifies a necessary condition for this guarantee: The data-revised set must accommodate the true DGP.

theoremIf the data-revised set $\mathcal{P}(\omega^{N})$ does not accommodate the true DGP $P^{*}$, then there exists a basic decision problem for which the data-revised decision is objectively worse than the benchmark decision, i.e., $W(c(D, \omega^{N}), P^{*}) < W(c(D), P^{*})$.

The proof of Theorem (ref) relies on a standard separating hyperplane argument by noticing that $P^{*} \in \mathcal{P}$ (Assumption (ref)), but $(P^{*})_{N}^{\infty} \notin \overline{co}(\mathcal{P}(\omega^{N})_{N}^{\infty})$ (Definition (ref)). The key takeaway is that if the data-revised set fails to accommodate the true DGP, then one can easily construct a basic decision problem for which the data-revised decision is objectively worse. In this sense, Theorem (ref) highlights that accommodating the true DGP is a necessary condition for achieving objective improvement in any class of decision problems including basic decision problems as special cases.

A truth-accommodating refinement, a slight strengthening of this necessary condition, can be shown to be also sufficient for objective improvement in basic decision problems. The next section establishes this result by characterizing the essentially largest class of decision problems for which this sufficiency holds, a class that includes all basic decision problems.

Monotone Decision Problems

What structural feature of a decision problem makes a truth-accommodating refinement sufficient for objective improvement? To build intuition, consider a betting decision problem, where every available act is a bet on the same event. Formally, in a betting decision problem, there exists $A \subseteq \Omega_{N}$ such that for each $f$, for all $\tilde{\omega} \in A$ and $\tilde{\omega}' \in A^{c}$,

equation*[equation* omitted — 78 chars of source]

In words, all acts rank outcomes in the same way across the two events, $A$ and $A^{c}$, but differ in their payoff tradeoffs. As a result, a higher belief in $A$ (i.e., a greater probability assigned to $A$) leads to a higher expected payoff of every act. Moreover, all acts can be ordered so that higher acts are optimal under higher beliefs. This is in line with the definition of monotone decision problems in ATHEY2018101,\footnote{Their definition says that the optimal act is monotone in signal $x$, which corresponds to a posterior belief over states. The additional property here is that it also leads to a higher expected payoff.} which highlights that both payoffs and choices move in the same direction as beliefs.

This form of monotonicity is precisely what makes a truth-accommodating refinement sufficient for objective improvement. When choosing how much to bet on the event $A$, the worst-case DGP is always the one assigning the lowest probability to $A$. A truth-accommodating refinement raises this worst-case probability while keeping it below the true probability. Monotonicity then guarantees that the data-revised decision moves closer to the exact optimal act under $P^{*}$, thereby delivering a higher objective payoff than the benchmark.

Motivated by this intuition, I now generalize the betting problem to higher dimensions by defining a class of monotone decision problems.

definitionA decision problem $D$ is called a monotone decision problem if there exists an act $g \in \mathcal{F}$ (not necessarily contained in $D$) such that, for every $f \in D$, \begin{equation*} f = \lambda_{f} g + c_{f}, \end{equation*} for some $\lambda_{f} \geq 0$ and $c_{f} \in \mathbb{R}$. Let $\mathcal{D}_{m}$ denote the class of monotone decision problems.

Intuitively, all acts in a monotone decision problem are non-negative affine transformations of a common reference act $g$. Hence they induce the same ranking over future states and satisfy

equation*[equation* omitted — 82 chars of source]

Thus the DM's evaluation reduces to the single statistic $W(g,P)$: a higher belief in $g$ (i.e., a greater expected payoff of $g$) raises the expected payoff of every act. This generalizes the monotonicity observed in betting problems. Importantly, the “monotone” label encodes a one-dimensional structure: although the underlying state space may be high-dimensional, all payoff-relevant variation collapses to the single expectation of $g$. From a geometric perspective, after subtracting the constant, all acts are aligned in the same direction. Thus, choosing among these acts is essentially trading off between their sensitivity to beliefs $(\lambda_{f})$ against their guaranteed payoffs $(c_{f})$, just as in betting decision problems one trades off payoffs across two events.

Basic decision problems are trivially monotone, since the only non-constant act can be taken as the reference act. It thus follows from Theorem (ref) that accommodating the true DGP is necessary for objective improvement in monotone decision problems. The next result establishes sufficiency.

theoremA data revision provides objective improvement in monotone decision problems if and only if the data-revised set is a truth-accommodating refinement.

Theorem (ref) delivers a clear and powerful message: a truth-accommodating refinement is sufficient to guarantee objective improvement in monotone decision problems, regardless of the true DGP. Its proof builds on the intuition developed in betting decision problems, but the extension is substantial: monotone decision problems constitute a broader class that encompasses a wide range of economically relevant environments. To illustrate this concretely, I next show that the problem of choosing among linear contracts in a canonical principal-agent model is a monotone decision problem. Consequently, Theorem (ref) applies directly and yields new insights.

example[Improving Linear Contracts with Data] A principal hires an agent to work on a project whose output is given by $y = \theta e + \epsilon$, where $\theta \in \mathbb{R}_{+}$ denotes the agent's productivity, $e \in \mathbb{R}_{+}$ the agent's effort, and $\epsilon$ a random noise with $\mathbb{E}[\epsilon]=0$. The agent's effort cost is $c(e)=k e^{2}/2$ for some $k>0$. The principal faces uncertainty about the agent's productivity. Let $\Theta \subseteq \mathbb{R}_{+}$ denote the set of possible productivity levels and $\mathcal{P}\subseteq \Delta(\Theta)$ the principal's initial belief set. The principal's objective is to choose a robustly optimal linear contract to offer the agent, given $\mathcal{P}$ and a possible revision using data from past projects. A linear contract specifies a base wage $\alpha \in \mathbb{R}$ and a share $\beta \in [0,1]$ of output. Given $(\alpha,\beta)$, the agent chooses effort $e$ to maximize $\mathbb{E}_{\epsilon} \big[\alpha + \beta(\theta e + \epsilon) - k e^{2}/2\big]$, yielding the optimal effort $e^{*}(\theta,\alpha,\beta)=\beta\theta/k \geq 0$. Thus, given $(\alpha,\beta)$ and $\theta$, the agent's expected payoff is $U_{A}(\theta;\alpha,\beta) = \alpha+\beta^{2}\theta^{2}/2k$ and the principal's expected payoff is \begin{align*} \pi_{(\alpha,\beta)}(\theta) = \mathbb{E}_{\epsilon} \big[(1-\beta)(\theta e^{*}(\theta,\alpha,\beta)+\epsilon)-\alpha\big] =\frac{\beta(1-\beta)}{k}\theta^{2}-\alpha. \end{align*} Let $u_{0}:\Theta\to\mathbb{R}$ denote the agent's outside-option payoff. Under robustness considerations, the principal restricts attention to contracts that satisfy individual rationality (IR) uniformly across all types.\footnote{The uniform IR condition provides one natural way to ensure that the linear-contract problem is monotone. Nevertheless, the monotone structure can also be preserved under alternative IR formulations, provided that the set of participating types remains invariant across contracts.} This pins down contracts to those satisfying \begin{equation*} \alpha\ge \sup_{\theta\in\Theta} \left( u_{0}(\theta)-\frac{\beta^{2}\theta^{2}}{2k} \right)\equiv \alpha_{\min}(\beta). \end{equation*} Because the principal's payoff is decreasing in $\alpha$, it is without loss to focus on $\alpha=\alpha_{\min}(\beta)$. Thus, the principal's expected payoff from a linear contract with share $\beta$ becomes \begin{equation*} \pi_{\beta}(\theta) =\frac{\beta(1-\beta)}{k}\theta^{2}-\alpha_{\min}(\beta) \equiv \lambda_{\beta}\,g(\theta)+c_{\beta}. \end{equation*} This implies that choosing among feasible contracts (i.e., among $\beta\in[0,1]$) is to choose among acts $\pi_{\beta}$ that are non-negative affine transformations of the common act $g(\theta) = \theta^{2}$, i.e., a monotone decision problem. By Theorem (ref), if the principal revises the initial belief set $\mathcal{P}$ about the agent's productivity to a truth-accommodating refinement using data, the resulting data-revised contract is guaranteed to yield a higher expected payoff than the benchmark contract regardless of the true productivity distribution.

The linear-contract example illustrates how monotone decision problems can arise in economic settings. The key is to identify a decision problem where all available options can be represented as non-negative affine transformations of a common act. In this example, the assumption of linear contracts is essential for representing the principal's decision problem as monotone. This assumption may be partially justified by the prominent role of linear contracts under robustness concerns Carroll2015, carroll2016. Nevertheless, to extend the observations here to more contracting environments, the generalizations identified in the next section are useful.

Monotone-Like Decision Problems

Theorem (ref) identifies monotonicity as a sufficient condition for a truth-accommodating refinement to guarantee objective improvement when $\mathcal{P}$ is arbitrary and acts are functions from states to payoffs. Its core intuition, in fact, extends more broadly as additional structure is imposed on either the initial set or the acts. Notice the key force behind Theorem (ref) is two-fold: (i) all acts share a common worst-case DGP, and (ii) a truth-accommodating refinement shifts the robust decision towards higher payoffs under every DGP in the revised set, thereby ensuring an objective improvement regardless of which DGP is the truth. In this section, I illustrate that these two forces can be found in several alternative formulations that are relevant in applications.

Monotone Decision Problems Under FOSD

Fix an order on $\Omega$. Say that $\mathcal{P}$ is FOSD-comparable if all DGPs in $\mathcal{P}$ are totally ordered by first-order stochastic dominance with respect to this order. A decision problem $D$ is monotone under FOSD if, under the same order on $\Omega$, every act in $D$ is increasing and for any $f, g \in D$, the difference $f - g$ is monotone (either increasing or decreasing). This requirement is weaker than Definition (ref) as it does not restrict acts to be affine transformation of one another. Nevertheless, it preserves the two forces identified above and thus guarantees objective improvement under truth-accommodating refinements.

corollaryFix an order on $\Omega$ and suppose $\mathcal{P}$ is FOSD-comparable. If the data-revised set is a truth-accommodating refinement, then the data revision provides objective improvement in all decision problems that are monotone under FOSD.

To see why, notice when $\mathcal{P}$ is FOSD-comparable, all increasing acts share the same worst-case DGP. If $\mathcal{P}(\omega^{N})$ is a truth-accommodating refinement, then the true DGP $P^{*}$ dominates the common worst-case DGP $P_{1}$ in $\mathcal{P}(\omega^{N})$, which in turn dominates the common worst-case DGP $P_{2}$ in $\mathcal{P}$. Let $f = c(D)$ and $g = c(D, \omega^{N})$. It follows that $W(f, P_{2}) \leq W(g, P_{2})$ and $W(g, P_{1}) > W(f, P_{1})$. Let $h = g - f$ and by monotonicity, $h$ must be increasing. Hence $W(h, P^{*}) \geq W(h, P_{1}) > 0$, which implies the desired objective improvement.

example[Improving (Non-Linear) Contracts with Data] Continuing the previous example, but additionally suppose the principal's initial knowledge $\mathcal{P}$ is FOSD-comparable with respect to $\Theta \subseteq \mathbb{R}_{+}$. This additional structure allows the principal to consider more general contracts beyond linear ones while still being able to guarantee objective improvement under truth-accommodating refinements. Concretely, let $\Theta = [\underline{\theta}, \overline{\theta}] \subseteq \mathbb{R}_{+}$ and suppose the principal expands the contract space to include the following quadratic form of contracts: \begin{align*} w_{(\beta, \gamma)}(y) = \alpha_{\min}(\beta,\gamma) + \beta y + \frac{\gamma}{2} y^{2}, \end{align*} with $\beta \in [0,1]$, $\gamma < (1-\beta)k/\overline{\theta}^{2}$, and $\alpha_{\min}(\beta,\gamma)$ the minimal base wage satisfying uniform IR. For simplicity, let $\epsilon \equiv 0$. The upper bound on $\gamma$ ensures that $e^{*}(\theta, \beta, \gamma) = \beta\theta/(k - \gamma \theta^{2})$ is always well-defined and the principal's payoff, \begin{align*} \pi_{(\beta, \gamma)}(\theta) = \frac{(1-\beta)\beta\theta^{2}}{k - \gamma \theta^{2}} - \frac{\gamma \beta^{2} \theta^{4}}{2(k - \gamma \theta^{2})^{2}} - \alpha_{\min}(\beta,\gamma), \end{align*} is increasing in $\theta$. However, it is not necessarily true that for any two feasible contracts $(\beta, \gamma)$ and $(\beta', \gamma')$, the difference $\pi_{(\beta, \gamma)}(\theta) - \pi_{(\beta', \gamma')}(\theta)$ is monotone in $\theta$. Nevertheless, if the principal's optimal contracts before and after data revision are such that this difference is monotone in $\theta$, then objective improvement can be established by examining only these two contracts: by Corollary (ref), objective improvement holds in the decision problem restricted to this pair. Since the principal in fact selects these same two contracts in the full problem, enlarging the contract space to include additional contracts that are not chosen does not affect the objective-improvement conclusion.

Monotone Moment Decision Problems

In fields such as information design gentzkow2016, Kolotilin2017,dworczak2019a, among others, decision problems are sometimes modeled as choosing among options whose payoffs depend on the distribution over states only through a scalar moment (e.g., the mean). Formally, fix a bounded measurable map $m: \Omega_{N} \rightarrow \mathbb{R}$ and write $m(P) \equiv \mathbb{E}_{P}[m(\omega)]$. A moment act is a function $\overline{f}:\mathbb{R} \to \mathbb{R}$ that assigns payoff $\overline{f}(m(P))$ under distribution $P$. Note that while every (state–contingent) act induces an expectation under each $P$, not every moment act corresponds to such a state–contingent act, especially when $\overline f$ is nonlinear in the moment (e.g., $\overline f(x)=x^{2}$).

A decision problem $\overline D$ is a monotone moment decision problem if all acts are increasing moment acts and single crossing holds: for all $\overline f,\overline g\in\overline D$, if $\overline f(x)\ge \overline g(x)$ at some $x$, then $\overline f(x')\ge \overline g(x')$ for all $x'\ge x$. As before, these conditions ensure the same two forces identified above. The following corollary summarizes the conclusion.

corollaryIf the data-revised set is a truth-accommodating refinement, then the data revision provides objective improvement in all monotone moment decision problems.
example[Improving Linear Contracts with Data Under Generalized Preferences] Continuing the linear-contract example, another key assumption that makes the principal's problem monotone is the principal's payoff form. Allowing the principal's preference to depend non-linearly on the expected output but still linear in the expected payment would generally break the monotonicity as in Definition (ref). Specifically, let $u: \mathbb{R} \to \mathbb{R}$ be an increasing function such that the principal's payoff under a linear contract $(\alpha, \beta) \in \mathbb{R} \times [0,1]$ and a distribution $P \in \Delta(\Theta)$ is \begin{equation*} \overline{\pi}_{(\alpha,\beta)}(P) = u\left(\frac{\beta}{k} \mathbb{E}_{P}[\theta^{2}]\right) - \frac{\beta^{2}}{k} \mathbb{E}_{P}[\theta^{2}] - \alpha. \end{equation*} Let $m(P) = \mathbb{E}_{P} [\theta^{2}]$, then $\overline{\pi}_{(\alpha,\beta)}$ is in the form of a moment act. When $u(\cdot)$ is twice differentiable, $u'(z) \geq 1$ and $u'(z) + z u''(z) \geq 2$ on the relevant range of moments is sufficient for all $\overline{\pi}_{(\alpha,\beta)}$ to be increasing and pairwise single-crossing in $m(P)$. Then by Corollary (ref), again, a truth-accommodating refinement guarantees objective improvement in choosing among linear contracts.

Monotone decision problems under FOSD and monotone moment decision problems illustrate two distinct yet complementary directions for generalizing the notion of monotonicity while preserving the two key forces underlying Theorem (ref). As emphasized at the beginning, within any specific decision context, whenever these two forces are present, a truth-accommodating refinement guarantees objective improvement.

Beyond Monotonicity

Outside the monotone and monotone-like classes identified above, where a common worst case and a directional monotonicity of payoff differences obtain, the guarantee that truth-accommodating refinements yield objective improvement can break down. The next result shows that this failure is robust: it could arise for arbitrary non-singleton refinements, or arbitrary non-monotone binary decision problems.

theoremThe following statements are true. \begin{enumerate}[(i)] • For any non-singleton data-revised set $\mathcal{P}(\omega^{N})$ that refines an initial set $\mathcal{P}$, there exists some $P^{*} \in \mathcal{P}(\omega^{N})$ and a non-monotone decision problem $D$ such that \begin{equation*} W(c(D, \omega^{N}), P^{*}) < W(c(D), P^{*}). \end{equation*} • For all non-monotone binary decision problem $D$ with finitely-based acts $f_{1}$ and $f_{2}$, if there exists $P \neq P' \in \Delta_{indep}(\Omega)$ such that $W(f_{1}, P) > W(f_{2}, P)$, and $W(f_{1}, P') < W(f_{2}, P')$, then there exists an initial set $\mathcal{P}$, a data-revised set $\mathcal{P}(\omega^{N})$, and a true DGP $P^{*}$ such that $\mathcal{P}(\omega^{N})$ is a truth-accommodating refinement, yet \begin{equation*} W(c(D, \omega^{N}), P^{*}) < W(c(D), P^{*}). \end{equation*} \end{enumerate}

Theorem (ref) identifies two different impossibility directions for extending sufficiency beyond monotone decision problems. Part (i) rules out any guarantee based solely on the refinement itself: for every non–singleton refinement $\mathcal{P}(\omega^{N})$ of $\mathcal{P}$ there is a true DGP and a non-monotone problem for which the data-revised choice underperforms the benchmark. Part (ii) emphasizes that there is no particular way of deviating from monotonicity that maintains sufficiency.\footnote{The statement restricts to finitely-based acts to avoid a small caveat involving tail events. See Remark (ref) in the proof for details. In addition, similar negative results can be stated for non-binary decision problems as the argument involves only two acts, those that are chosen by the benchmark and data-revised decisions. The presence of other acts only complicates the construction.} In other words, monotonicity is not merely sufficient; it is essentially the boundary for when a truth-accommodating refinement guarantees objective improvement.

To gain some intuition of why without monotonicity, a truth-accommodating refinement could lead to a strictly worse decision, consider the following example.

example[Introductory Example continued] Decision Problem \expandafter\@slowromancap\romannumeral 2@. Netflix decides whether to include (i) or remove (r) the movie from their recommendations. The key difference from Decision Problem \expandafter\@slowromancap\romannumeral 1@ is that the two actions now rank the states in opposite ways: removing the movie yields a higher payoff when users dislike it, whereas including it yields a higher payoff when users like it. Numerically, let the payoffs be $i(U) = 1$, $i(D) = -1$, $r(U) = 0$, and $r(D) = 1$. Because the two alternatives rank the two states differently, this decision problem is non-monotone. In this case, Netflix's benchmark decision is to choose $r$. If Netflix revises the initial belief using the empirical distribution method described in the introduction, then their data-revised decision will be $r$ when the true DGP is $(1/3)^{\infty}$ and $i$ when the true DGP belongs to the set $\{3/5, 1\}^{\infty}$. However, if the true DGP is $(3/5)^{\infty} \in \{3/5, 1\}^{\infty}$, the expected payoff of $i$ is $1/5$, strictly lower than the expected payoff of $r$, equal to $2/5$.

In Decision Problem \expandafter\@slowromancap\romannumeral 2@, the two acts rank the two states differently. While a higher probability of $U$ implies a “higher” optimal act, it does not necessarily lead to a higher expected payoff. This non-monotonicity invalidates the previous intuition. In this case, the worst-case DGPs for the two acts are those assigning the lowest probability to $U$ and $D$, respectively. While a truth-accommodating refinement still ensures that the lowest probabilities of $U$ and $D$ in the data-revised set are greater than those in the initial set but less than the true probabilities, the different levels of probability increase may lead the data-revised decision to be further away from the exact optimal act.

Decision Problem \expandafter\@slowromancap\romannumeral 2@ illustrates one direction where a decision problem can deviate from monotonicity. For all the other directions, see the example in the proof of Theorem (ref) for an illustration.

Impossibility of Objective Improvement in All Decisions

Are there stronger conditions than truth-accommodating refinements that could guarantee objective improvement beyond monotone decision problems? This section provides a negative answer when considering the set of all decision problems, i.e., when $\mathcal{C} = \mathcal{D}$.

Obviously, a sufficient condition is when $P^{*}$ is uniquely identified from $\omega^{N}$ and $\mathcal{P}$ is not a singleton. In this case, by letting $\mathcal{P}(\omega^{N}) = \{P^{*}\}$, the DM's data-revised decision is exactly optimal against $P^{*}$, thus always improves. However, revising a non-singleton initial set to a singleton set containing the true DGP is not always feasible, especially when the possible DGPs can be non-identical. When the data-revised set is not a singleton, the following theorem shows that objective improvement in all decision problems requires the true DGP to be effectively uniquely identified.

theoremA data revision provides objective improvement in all decision problems if and only if there exists $\alpha \in (0,1]$ such that \begin{equation*} \overline{co}(\mathcal{P}(\omega^{N})_{N}^{\infty}) = \alpha P^{*\infty}_{N} + (1-\alpha)\overline{co}(\mathcal{P}_{N}^{\infty}). \end{equation*}

The condition $\overline{co}(\mathcal{P}(\omega^{N})_{N}^{\infty}) = \alpha P^{*\infty}_{N} + (1-\alpha)\overline{co}(\mathcal{P}_{N}^{\infty})$ says that, after taking the closed convex hull of future marginals, the data-revised set needs to be a convex combination of the initial set and the true DGP. Observe that the only way to form such a data-revised set requires knowing exactly what $P^{*}$ is. But if it is the case, the DM should just let $\{P^{*}\}$ be their data-revised set. On the other hand, if $\overline{co}(\mathcal{P}(\omega^{N})_{N}^{\infty}) = \alpha P^{*\infty}_{N} + (1-\alpha)\overline{co}(\mathcal{P}_{N}^{\infty})$ for some $\alpha \in (0,1]$, then the same cannot hold for any other $P$ and $\alpha$ whenever $P_{N}^{\infty} \neq P_{N}^{*\infty}$. This can be seen by considering the probability of any arbitrary event. This impossibility thus leads to the following corollary.

corollaryAny given data-revised set can provide objective improvement in all decision problems under at most one DGP (up to the same marginal over future states).

Given Corollary (ref), if there are multiple DGPs the DM can by no means distinguish using the sample data (like the ones in the introductory example), then no data-revised set would be able to guarantee objective improvement in all decision problems under all of them.

Crucially, a truth-accommodating refinement does not have this issue: A data-revised set continues to be a truth-accommodating refinement no matter which DGP accommodated by it turns out to be the truth. Therefore, Theorem (ref) indeed further implies that the objective improvement can be guaranteed simultaneously under multiple DGPs, contrasting to the conclusion in Corollary (ref).

In addition, a truth-accommodating refinement can also provide a weaker improvement guarantee in all decision problems:

propositionIf the data-revised set $\mathcal{P}(\omega^{N})$ is a truth-accommodating refinement, then for all $D \in \mathcal{D}$, \begin{equation*} W(c(D, \omega^{N}), P^{*}) \geq \min\limits_{P \in \mathcal{P}} W(c(D), P), \end{equation*} and the inequality is strict for some $D \in \mathcal{D}$.

In words, the expected payoff from the data-revised decision is never lower than the guaranteed payoff of the benchmark decision. Hence, whenever the revision rule constitutes a truth-accommodating refinement, the DM would never prefer to forgo the opportunity to revise their decision using data in exchange for the benchmark's certainty-equivalent payoff. This improvement guarantee is relatively weak but not trivial: a misled data-revised decision could be objectively worse than this certainty equivalent. Truth accommodation ensures that it cannot happen. In other words, a truth-accommodating refinement guarantees that learning from data is always valuable relative to receiving the ex-ante certainty equivalent payoff, providing an additional rationale for adopting revision rules that accommodate the truth.

As a final remark, all results in this section also apply when comparing two data-revised sets, say $\mathcal{P}_{1}(\omega^{N})$ and $\mathcal{P}_{2}(\omega^{N})$. When both sets accommodate the truth and $\mathcal{P}_{1}(\omega^{N})$ is a refinement of $\mathcal{P}_{2}(\omega^{N})$, all conclusions about objective improvement holds by viewing $\mathcal{P}_{1}(\omega^{N})$ as a truth-accommodating refinement of $\mathcal{P}_{2}(\omega^{N})$.

Revision Rules that Accommodate the Truth

As established in the previous section, for improving robust decisions, it is necessary and sometimes sufficient to accommodate the true DGP in revising the initial set (the refinement part only guarantees it to be sometimes strict). Define a revision rule as the mapping from an initial set $\mathcal{P}$ and sample data $\omega^{N}$ to a data-revised set $\mathcal{P}(\omega^{N})$. When the sample size is unbounded, the following definition formalizes a notion of accommodating the truth in the asymptotic sense.

definitionA revision rule accommodates the truth asymptotically almost surely if, for any initial set $\mathcal{P}$, for every $P^{*} \in \mathcal{P}$, and for $P^{*}$-almost every $\omega \in \Omega$, there exists $\bar{N}(\omega)$ such that, for all $N \geq \bar{N}(\omega)$, the data-revised set $\mathcal{P}(\omega^{N})$ accommodates the DGP $P^{*}$.

This definition says that, regardless of which possible DGP governs the uncertainty, the revision rule ensures that, with a sufficient amount of data, the data-revised set accommodates the true DGP almost surely. Therefore, whenever accommodating the truth is sufficient for objective improvement, such improvements are also achieved asymptotically almost surely.

Likelihood-based rules have been shown in the introductory example to violate this property. Here, I propose a revision rule based on empirical distributions: For any sample data $\omega^{N}$, let $\boldsymbol\Phi(\omega^{N}) \in \Delta(S)$ denote the empirical distribution, i.e., for any outcome $s \in S$, $\boldsymbol\Phi(\omega^{N})(s) \equiv N^{-1}\sum_{i=1}^{N} I\{\omega_{i} = s\}.$ For any $P \in \Delta_{indep}(\Omega)$ and for any $N$, the average of sample marginals, $\bar{P}^{N} \in \Delta(S)$, is defined to be the distribution over $S$ given by the average mixture of the marginal distributions over each component of the sample states, i.e., $\bar{P}^{N} \equiv N^{-1}\sum_{i=1}^{N}P_{i}$. For any $p,q\in \Delta(S)$, let $\rho(p,q)$ denote the sup-norm distance.

definitionThe data-revised sets are obtained by the empirical distribution method if, for some pre-specified $\epsilon > 0$ and for every $\omega^{N}$, \begin{equation*} \mathcal{P}(\omega^{N}) = \left\{P\in \mathcal{P}: \rho (\bar{P}^{N}, \boldsymbol\Phi(\omega^{N}) ) < \epsilon \right\}. \end{equation*}

The empirical distribution method is a formalization of the simple heuristic of retaining a DGP if its average of sample marginals is close enough to the empirical distribution. Such a heuristic can be used when DGPs are i.i.d., for it happens to coincide with retaining DGPs that maximize the likelihood in this special case. With possible non-identical DGPs, while maximum likelihood is no longer useful, this heuristic remains valid.

theoremFor all $\epsilon > 0$, the empirical distribution method accommodates the truth asymptotically almost surely.

The proof of Theorem (ref) is to verify that Kolmogorov's strong law of large numbers holds. This also suggests that the independence assumption is not crucial for this result. As long as there is a version of the strong law of large numbers, one can obtain the same conclusion. Notable cases include when the DGPs satisfy Markov property or are weakly dependent dejong_1995.

In addition, the empirical distribution method is not the only revision rule that accommodates the truth asymptotically almost surely. For instance, when the odd and even experiments are known to have different characteristics, one may apply the empirical distribution method separately to the odd and even experiments to obtain a potentially more refined data-revised set. Such revision rules, however, typically need to be tailored to the specific structure in a case-by-case manner. In contrast, the empirical distribution method stands out for its simplicity and general applicability.

Accommodating the Truth with Finite Sample

As Theorem (ref) holds for all $\epsilon > 0$, letting $\epsilon \rightarrow 0$ obtains the theoretic limit of the data-revised sets under the empirical distribution method. For applications with finite samples, the standard approach is to derive the $\epsilon$-bound as a function of the sample size that ensures a pre-specified asymptotic probability of accommodating the truth.\footnote{With a finite sample, the only revision rule that accommodates the truth almost surely is to keep using the initial set, because of the full-support assumption. Thus, the only meaningful notion of accommodating the truth with a finite sample is the probabilistic approach, which is also a standard practice in statistical inferences.} This section develops a novel and simple method to achieve this. First, define the following finite-sample notion of accommodating the truth.

definitionA revision rule accommodates the truth with an asymptotic level $1-\alpha$ if, for any initial set $\mathcal{P}$ and for every $P^{*} \in \mathcal{P}$, \begin{equation*} \liminf_{N \rightarrow \infty} P^{*} (\{\omega^{N}: \mathcal{P}(\omega^{N}) accommodates P^{*}\}) \geq 1-\alpha. \end{equation*}

A revision rule satisfying Definition (ref) ensures that, regardless of which DGP governs the data, the data-revised set accommodates the true DGP with asymptotic probability at least $1-\alpha$. In particular, the level $1-\alpha$ is understood uniformly --- the probability bound holds uniformly over all possible DGPs in $\mathcal{P}$. Consequently, objective improvement is guaranteed with at least the same asymptotic probability whenever accommodating the truth is sufficient.

Definition (ref) effectively requires the data-revised sets to serve as consistent confidence regions for the true DGP.\footnote{The coverage is in the weaker sense of accommodating but the difference is unimportant.} One way to ensure this property is to construct the confidence regions directly and use them as the data-revised sets.

Constructing confidence regions is theoretically straightforward exploiting the well-known duality with hypothesis tests. For any $P \in \mathcal{P}$, consider testing the null hypothesis $P^{*} = P$, against the unrestricted alternative $P^{*} \neq P$. Let $A_{N, \alpha}(P) \subseteq S^{N}$ denote the region of acceptance: the set of sample data under which the null cannot be rejected. Then find regions that satisfy

equation*[equation* omitted — 82 chars of source]

that is, the probability of accepting the null when it is true is at least $1-\alpha$ asymptotically. Given any data $\omega^{N}$, construct the data-revised set as

equation*[equation* omitted — 99 chars of source]

Such data-revised sets contain the true DGP with asymptotic probability $1-\alpha$, uniformly across all $P$ in $\mathcal{P}$.

When all possible DGPs are i.i.d., this construction is standard and tractable for two convenient features. First, by the central limit theorem, the regions of acceptance can be obtained from probability contours of the corresponding multivariate Gaussian approximations. The mean vectors and covariance matrices depend only on the unique marginal distribution and thus remain fixed as the sample size grows. Second, because each i.i.d. distribution is uniquely determined by its marginal, the number of tests remains fixed regardless of the sample size.

Both features do not carry over when DGPs may be non-identical. First, in this case, both the mean vectors and covariance matrices depend on all marginals across sample states. As a result, for every sample size, both need to be recomputed even for the same DGP. Second, a non-identical DGP is determined by all its marginals, so each additional observation increases the number of tests.

To overcome these difficulties, I present in the following a novel method for constructing confidence regions for non-identical DGPs. This method retains the simplicity of the i.i.d. case while guaranteeing that the resulting confidence regions cover the true non-identical DGP with at least the required asymptotic level.

Formally, for any $p \in \Delta(S)$, let $p^{\infty}$ denote the i.i.d. distribution over $\Omega$ with marginal $p$. Let $A^{*}_{N,\alpha} (p^{\infty})$ denote its region of acceptance with asymptotic level $1-\alpha$, constructed using the corresponding Gaussian approximation.\footnote{The exact form is standard and is given by equation (ref) in the appendix, with additional notations.} For any $P \in \Delta_{indep}(\Omega)$, recall $\bar{P}^{N} \in \Delta(S)$ denotes the average of sample marginals. Let $A^{*}_{N, \alpha}((\bar{P}^{N})^{\infty})$ denote the region of acceptance constructed using the Gaussian approximation of the i.i.d. distribution $(\bar{P}^{N})^{\infty}$. Consider the following revision rule:

definitionThe data-revised sets are obtained using the augmented i.i.d. test with asymptotic level $1-\alpha$ if, for every $\omega^{N}$, \begin{equation*} \mathcal{P}(\omega^{N}) = \left\{P \in \mathcal{P}: \omega^{N} \in A^{*}_{N,\alpha}((\bar{P}^{N})^{\infty}) \right\}. \end{equation*}

Intuitively, the augmented i.i.d. test follows a two-step procedure:

enumerate[(i)] • For any sample data $\omega^{N}$, construct a confidence region as if the initial set consists of all i.i.d. DGPs. • For each DGP in the initial set, retain it in the data-revised set if its average of sample marginals coincides with the marginal of some i.i.d. DGP in the previous confidence region.

Notice the first step is the standard procedure for constructing confidence regions from i.i.d. sample data. The essential departure is the second step, which augments the i.i.d. confidence region by also including the possible non-identical DGPs. Importantly, this augmentation is achieved through a straightforward comparison, adding virtually no computational difficulty. Therefore, implementing the augmented i.i.d. test is as tractable as conventional statistical inferences based on i.i.d. samples.

theoremThe augmented i.i.d. test with asymptotic level $1-\alpha$ accommodates the truth with the same asymptotic level.

The proof of this theorem relies on a key observation: For all $P \in \Delta_{indep}(\Omega)$, for all $N$ and $\alpha$,

equation*[equation* omitted — 90 chars of source]

which further implies that

equation*[equation* omitted — 92 chars of source]

Therefore, when testing the null hypothesis $P^{*} = P$, using the region of acceptance $A^{*}_{N,\alpha}((\bar{P}^{N})^{\infty})$, the probability of accepting $P$ when it is true is at least greater than the probability when using $A^{*}_{N,\alpha}(P)$. The latter probability is, by construction, asymptotically greater than $1-\alpha$.

The key relation, $A^{*}_{N, \alpha}(P) \subseteq A^{*}_{N,\alpha}((\bar{P}^{N})^{\infty})$, is shown in Lemma (ref) by deriving a result relating the average covariance matrices of the two distributions, $P$ and $(\bar{P}^{N})^{\infty}$. Specifically, subtracting the average covariance matrix of $P_{N}$ from that of $((\bar{P}^{N})^{\infty})_{N}$ yields a positive semi-definite matrix. This result generalizes a well-known variance inequality for mixtures of binomial distributions wang1993 to the multinomial case.\footnote{Specifically, the average variance of an i.i.d. binomial distribution is always weakly greater than the average variance of a non-identical binomial distribution whose average mean is the same as the i.i.d. distribution.}

remarkOne potential concern is that computing the probability contours of multivariate Gaussian distributions may be difficult when $|S|$ is large. A practical alternative is to construct a Bonferroni-type confidence region by forming confidence intervals for the probability of each outcome $s$, each with confidence level $1-\alpha/(|S|-1)$. The intersection of all such confidence intervals then yields a confidence region with overall confidence level $1-\alpha$. However, this region is generally more conservative than that obtained directly from the multivariate Gaussian distribution. Moreover, using the corresponding result in wang1993 for binomial distributions, one can show that the Bonferroni-type confidence region constructed from i.i.d. distributions guarantees at least the same coverage probability for non-identical distributions.

Applications with Parametric Models

This section illustrates how the theoretical results translate into familiar statistical and economic settings. In particular, it studies two applications where the initial sets are given by specific parametric models. The first application showcases a setting where the data-revised sets obtained under the augmented i.i.d. test have a closed-form solution, and are closely related to the standard Wilson confidence interval. The second application highlights new findings in a commonly studied model of learning under ambiguity. These applications confirm the practical relevance of the proposed revision rules and the associated theoretical results.

Bernoulli Model with Ambiguous Nuisance Parameters

This model is a generalization of the one studied in Walley1991-WALSRW. Suppose the DM faces a sequence of coin flips, with outcome space $\Omega = \{H, T\}^{\infty}$, representing heads and tails. The probability of getting a head from the $i$-th coin flip is determined by both a structural parameter $\theta \in [0,1]$ and a nuisance parameter $\psi_{i} \in [0,1]$:

equation*[equation* omitted — 64 chars of source]

for some fixed $\delta \in [0,1]$. The structural parameter is common across all flips and can be interpreted as a systematic characteristic of the coins. But each coin flip is also affected by its idiosyncratic feature captured by $\psi_{i}$. Throughout, I use the probability of $H$ to represent a probability distribution over $\{H, T\}$. Because each $\psi_{i}$ is only known to lie in $[0,1]$, each structural parameter $\theta$ corresponds to a set of possible DGPs:

equation*[equation* omitted — 142 chars of source]

Let $\Theta = [0,1]$ denote the set of structural parameters. The initial set of DGPs is therefore $\mathcal{P} = \cup_{\theta \in \Theta} \mathcal{P}_{\theta}$. The DM observes outcomes of $N$ coin flips. Their goal is to estimate the true structural parameter and make a set-valued prediction for the probability of getting a head in the next flip. The benchmark estimate is simply the initial set of parameters $\Theta = [0,1]$, so the benchmark prediction is $\mathcal{P}_{N+1} = [0,1]$.

Consider the DM's asymptotic prediction using the empirical distribution method. For simplicity, I ignore the pre-specified $\epsilon$ by taking it to be arbitrarily small. Then the data-revised set is given by

equation*[equation* omitted — 132 chars of source]

Let the DM's data-revised estimate of the structural parameter be given by

equation*[equation* omitted — 133 chars of source]

i.e., a structural parameter is considered possible whenever there is a corresponding DGP retained in $\mathcal{P}(\omega^{N})$. For any $\theta \in \Theta$, there exists $P \in \mathcal{P}_{\theta}$ satisfying the above condition if and only if $\Phi(\omega^{N}) \in \left[(1-\delta)\theta, (1-\delta)\theta + \delta \right]$.

As a result, it follows that

equation*[equation* omitted — 185 chars of source]

Notice the data-revised prediction is also completely determined by the data-revised estimate of the structural parameter and is given by

equation*[equation* omitted — 169 chars of source]

Intuitively, the asymptotic prediction under the empirical distribution method is a “$\delta$-fattening” of the observed empirical frequency of heads.

Next, consider the finite-sample estimate and prediction at asymptotic level $1-\alpha$ using the augmented i.i.d. test. For any sample data $\omega^{N}$, the first step is to construct the confidence interval as if the underlying DGPs were i.i.d. binomial distributions. Specifically, the corresponding confidence interval is the Wilson Interval.\footnote{Different from the probably more famous Wald Interval which uses the sample variance, Wilson Interval is constructed by directly inverting the statistical tests, thus using the null variance. Wilson Interval has considerably better asymptotic performance than the Wald Interval. See Brown2001 for a discussion.} Let $z_{\alpha/2}$ denote the upper $100(\alpha/2) \%$ quantile of the standard normal distribution. Let $[\underline{W}(\omega^{N}), \overline{W}(\omega^{N})]$ denote the Wilson Interval which has the following closed-form expressions:

align*[align* omitted — 486 chars of source]

Given the i.i.d. confidence interval, the second step is to consider non-identical DGPs whose average of sample marginals falls into this confidence interval: A structural parameter $\theta$ is retained in the data-revised estimate if and only if $[(1-\delta)\theta, (1-\delta)\theta + \delta] \cap [\underline{W}(\omega^{N}), \overline{W}(\omega^{N})] \neq \emptyset$. Hence, the data-revised estimate is

equation*[equation* omitted — 201 chars of source]

Similarly, the data-revised prediction in this case is

equation*[equation* omitted — 185 chars of source]

Notice the prediction is again a $\delta$-fattening of the Wilson Interval. Therefore, The resulting expressions are not only analytically tractable but also intuitively interpretable.

Gaussian Signals with Ambiguous Variances

Prior-by-prior or full-Bayesian updating is the most commonly used update rule in models of learning under ambiguity in the literature. However, its asymptotic result is often hard to derive and is known only in specific parametric models. This section revisits one such model from reshidi2020information and shows that applying the empirical distribution method yields a simpler analysis and entirely different conclusions.

In this model, the DM aims to learn the state of the world $\theta \in \Theta \equiv \mathbb{R}$ by observing a countably infinite sequence of signals denoted by $\{x_{i}\}_{i=1}^{\infty}$. Each $x_{i}$ is a Gaussian random variable with mean $\theta$ and variance $\sigma_{i}^{2}$, and let $g_{i}(\theta, \sigma_{i})$ denote its probability density function. The signals are mutually independent, but each $\sigma_{i}$ is only known to lie in $[\underline{\sigma}, \overline{\sigma}]$. Thus, the DM observes a sequence of independent but possibly heterogeneous Gaussian random variables.

Formally, for each state $\theta$, let

equation*[equation* omitted — 172 chars of source]

denote the set of possible DGPs over the signal sequence. The initial set is $\mathcal{P} = \cup_{\theta \in \Theta} \mathcal{P}_{\theta}.$ For every $N \in \mathbb{N}$, let $\hat{x}^{N} \equiv (\hat{x}_{1}, \hat{x}_{2}, \cdots, \hat{x}_{N})$ be a sequence of signal realizations and let $\mathcal{P}(\hat{x}^{N})$ denote the data-revised set of DGPs. Because the DM's goal is to learn the true state, define

equation*[equation* omitted — 136 chars of source]

as the set of states consistent with the data-revised set of DGPs.

Directly applying full Bayesian updating here would amount to retain all possible DGPs, implying $\mathcal{P}(\hat{x}^{N} ) \equiv \mathcal{P}$ and hence $\Theta(\hat{x}^{N}) \equiv \Theta$ for all $\hat{x}^{N}$. In reshidi2020information, they apply full Bayesian updating differently by assuming a prior $\mu$ over $\Theta$. Then applying full Bayesian updating is to apply Bayes' rule to update $\mu$ under each possible DGP. Specifically, for each $\theta$, let $P_{\theta} \in \mathcal{P}_{\theta}$ denote a specific DGP, the posterior probability

equation*[equation* omitted — 156 chars of source]

This yields a set of posterior beliefs over $\Theta$.\footnote{This way of applying full Bayesian updating can be incorporated into the present framework by letting the state space be $\Theta \times S^{\infty}$ and allowing dependence in DGPs. Applying (ref) in Appendix (ref) yields exactly the same posterior distribution over $\Theta$.} Their main result (Theorem 1) shows that in any state $\theta$ and some possible DGP, the set of posteriors converges almost surely to a set of degenerate distributions over a non-singleton set of states. That is, ambiguity does not vanish asymptotically.

Consider revising the initial set using the empirical distribution method. Because only the mean matters, it suffices to consider their sample mean.

definitionThe revision of states is obtained by the sample mean method if, for some pre-specified $\epsilon > 0$, and for every $\hat{x}^{N}$, \begin{equation*} \Theta(\hat{x}^{N} ) = \left\{\theta \in \Theta: \left|N^{-1}\sum\limits_{i=1}^{N}\hat{x}_{i} - \theta \right| < \epsilon \right\}. \end{equation*}

As in the empirical distribution method, $\Theta(\hat{x}^{N} )$ retains a state if it is close enough to the sample mean. Because the mean of each marginal distribution equal to $\theta$, applying Kolmogorov's strong law of large numbers yields the following result.

propositionFor any $\theta^{*} \in \Theta$ and any $P^{*} \in \mathcal{P}_{\theta^{*}}$, the revision of states obtained by the sample mean method with any $\epsilon > 0$ contains the true state asymptotically almost surely, i.e., for any $\epsilon > 0$, \begin{equation*} P^{*}\left(\hat{x}: \lim\limits_{N \rightarrow \infty} \theta^{*} \in \Theta(\hat{x}^{N})\right) = 1. \end{equation*}

As the conclusion holds for any $\epsilon > 0$, taking $\epsilon$ arbitrarily small makes the revised set of states arbitrarily precise. Thus, even with Gaussian signals that have unknown and possibly heterogeneous variances, the true state can still be identified asymptotically almost surely. This stands in sharp contrast to full Bayesian updating, under which ambiguity persists.

The key factor enabling asymptotic identification here is that all Gaussian signals share the same mean. When the means themselves are also ambiguous, the sample-mean method may still yield asymptotic ambiguity over a non-vanishing set of states. Nonetheless, the key here is that this asymptotic prediction is obtained straightforwardly with the sample-mean method, whereas the corresponding asymptotic result for full-Bayesian updating in this setting remains unknown.

Related Literature

The decision environment formulated in this paper is closely related to some in the literature on decisions under ambiguity. In particular, it directly generalizes the setting studied by Epstein2007. They assume the DM applies maximum likelihood updating to revise the initial set. The present paper highlights possible concerns with this approach. Epstein2016 develop robust confidence regions when the possible data-generating processes are belief functions. Belief functions impose restrictions on the possible marginal distributions. In contrast, the environment studied here allows for arbitrary marginals. But the main conceptual difference from their paper and other papers on asymptotic learning under ambiguity, such as Marinacci2002 and MARINACCI2019144, is that the present paper emphasizes implications for decision making in addition to asymptotic learning.

This paper also contributes to the literature on dynamic decisions under ambiguity by proposing new rules for revising or updating sets of distributions. See gilboa_marinacci_2013 and CHENG2022102587 for recent developments. The essential departure of the present paper from this literature is that it evaluates decisions using an objective criterion. The objective criterion leads to a characterization of accommodating-the-truth property, conceptually analogous to statistical consistency. In this way, this paper draws a connection between a classical concept from statistics and the theory of decisions under ambiguity, following the line of research by cerreia-vioglio2013 and denti2022. This objective approach also resonates with some recent studies of misspecified learning that evaluate performance according to an objective measure, such as Frick2021 and he2020evolutionarily.

Finally, this paper develops a useful augmenting technique for making inferences in the presence of independent but non-identical distributions. In essence, this technique can be applied on top of standard statistical procedures developed under the i.i.d. assumption. The data-revised set naturally serves as a set-valued identification object and is therefore related to the partial identification literature (see surveys by Tamer2010, canay/shaikh:2017, and MOLINARI2020355). In that literature, the decision environment is typically formulated so that the data-generating distribution is point identified, while the payoff-relevant parameters are only partially identified, see, for example, christensen2023optimal. By contrast, the present paper focuses on situations where the distribution itself is partially identified and develops relevant inference methods for such settings.