Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
68,444 characters · 18 sections · 42 citation commands
19pt19ptA Two-Ball Ellsberg Paradox: An Experiment
\tikzstyle{level 1}=[level distance=30mm, sibling distance=30mm] \tikzstyle{level 2}=[level distance=30mm, sibling distance=15mm] \tikzstyle{level 3}=[level distance=20mm]
\onehalfspacing
We present in this paper a novel and simple experiment that documents a new channel through which decision-makers might shy away from ambiguous choice situations and prefer risky situations that are “objectively" worse for them. Risk refers to situations where various outcomes can occur with precise probabilities. Ambiguity (or uncertainty) refers to situations where agents cannot attach precise probabilities to the various possible outcomes. Ambiguity aversion refers to the preference for risky situations over similar ambiguous situations.
In 1961, Daniel Ellsberg (Ellsberg) showed in his famous ‘Ellsberg Paradox’ experiment that people generally exhibit ambiguity aversion, falsifying the popular Subjective Expected Utility theory of decision-making under uncertainty (savage1972foundations). Since then, theorists have devised many other models of rational decision-making that account for the Ellsberg Paradox. Amid the early axiomatic work, Schmeidler proposed Choquet expected utility (CEU), GilboaSchmeidler designed Maximin expected utility (MEU), GMM suggested the alpha-maximin expected utility ($\alpha$-MEU). The second generation of models includes Klibanoff, MMR, and Strzalecki. All these models can accommodate typical Ellsberg choices, but they cannot explain the results of our experiment.
Ambiguity and ambiguity aversion are relevant to economists and policymakers since, in most real-world situations, agents cannot attach precise probabilities to the possible outcomes. Relying on the standard models cited above, researchers have explored the important implications of ambiguity in diverse economic fields. To name a few: in environmental economics, millner2013scientific demonstrate the effects of ambiguity on the social cost of carbon by integrating ambiguity within Nordhaus' famous integrated assessment model of climate economy. lange2008uncertainty investigate the learning effects of climate policy under ambiguity. In health economics, treich2010value shows under which conditions ambiguity aversion increases the value of a statistical life. In macro-finance, ju2012ambiguity show how ambiguity aversion can account for the equity premium puzzle.
Our paper focuses on the following thought experiment proposed by Jabarian1. There are two urns, R and A, and two types of balls: red and blue. In urn R, there are 100 balls: 50 red and 50 blue balls. In urn A, there are 100 balls in unknown proportion. Subjects draw two balls successively (with replacement) from one of the two urns and win a prize if the balls match in color. We focus on the following gambles. $RR$: drawing twice from urn R, versus $AA$: drawing twice from urn A. Which gamble would subjects prefer: $RR$ or $AA$?
Although urn R may seem more attractive than urn A, since its content is merely risky, gamble $AA$ has an unambiguously larger win probability than gamble $RR$. Indeed, the more unevenly distributed urn A's contents, the higher the win probability of gamble $AA$. Urn \textbf{A} having 50% red balls and 50% blue balls is the \textit{worst case}: it yields a probability of winning equal to $\frac{1}{2}$. Any other distribution gives a probability of winning that is higher. If the distribution were, say, $(\frac{1}{3},\frac{2}{3})$, the probability of winning would be $\frac{5}{9}>\frac{1}{2}$. We call the preference for gamble $RR$ over $AA$ the \textit{Two-Ball Ellsberg Paradox}, and among a nationally representative sample in an incentivized online experiment, we find that it is widespread.
Unless subjects have beliefs over the two draws that are not consistent with the information given to them (say, subjects somehow believe the draws are not independent), the choice of RR over AA cannot be reconciled with existing models of ambiguity aversion in a straightforward manner. For instance, in the Maxmin Expected utility model of GilboaSchmeidler, if subjects entertained the entire simplex over $\{R,B\}$ for the composition of an urn U and form beliefs over the two draws by composing each prior with itself, this would lead to indifference between gambles $RR$ and $AA$.\footnote{Dominance, as in gajdos2008attitude, implies that $RR$ cannot be strictly preferred to $AA$.} Actually, if one assumes that the set of priors is a subset of $\{q \in \Delta(\{R,B\}^2) | q=p\times p; p\in\Delta(\{R,B\})\}$, then the choice RR over AA is incompatible with $\alpha$-maxmin expected utility for any $\alpha$. This choice is also incompatible with savage1972foundations's Subjective Expected Utility model if beliefs are a product measure of the type $p\times p$.
\paragraph{Aims of the paper.} We report the results of an incentivized experiment on a nationally representative US sample to test whether people show Two-Ball Ellsberg Paradox preferences and if so, to understand why this may be the case. Our experiment aims to answer two main questions. First, do subjects occasionally avoid an ambiguous act by choosing an unambiguous lottery worse than any lottery the ambiguous act could produce? Second, is this preference entirely due to failing to understand Two-Ball gambles? We also investigate how subjects' Two-Ball Ellsberg Paradox preferences relate to other “paradoxical” behaviors, such as ambiguity aversion (as in the Classical Ellsberg Paradox) and complexity aversion (as measured by two-ball versus one-ball gambles, by compound versus simple gambles by the combination of the two).
\paragraph{Experimental design.} We used a multiple price list (MPL) to elicit subjects' Certainty Equivalents (CEs), allowing us to determine the strength of their preferences for various gambles. Laboratory, and even more online, experiments eliciting subjects' CEs for gambles are often subject to significant measurement error, i.e., biases in estimating coefficients and correlations. We rely on the Obviously Related Instrumental Variables (ORIV) approach of GSY to correct these errors.
Due to time constraints, our experiment is divided into four treatments and subjects are randomly assigned to one of them. Treatment nudging tests whether subjects' preference for urn R over urn A is due to their failing to understand the gambles. Treatment complexity tests whether subjects respond in the same way to complexity, in the form of compound lotteries, as they do to ambiguity. Treatment robustness tests whether subjects' preference for urn \textbf{R} over urn \textbf{A} is due to a false belief that the contents of urn \textbf{U} are changed between draws from that urn, and it also explores how the \textit{number or proportion of draws that are from urn} \textbf{U} affects subjects' preferences. Lastly, treatment \textsc{ order} tests whether the order of draws from various urns matters and explores whether the \textit{mere presence of ambiguity} affects subjects' decisions.
\paragraph{Main results.} Overall, we determined that 55% of subjects exhibit a distaste for ambiguity that cannot be explained by most existing models or by the fact that subjects failed to understand the situations presented to them. Our results suggest that subjects avoid ambiguity per se as opposed to avoiding it simply because it may yield a worse outcome. The two most important findings supporting this idea of ambiguity aversion per se originate from our core treatment, nudging.
This treatment revealed that a lack of understanding does not entirely explain the preference for avoiding ambiguity. Subjects “correctly” select the act with a larger win probability when comparing two similar two-ball acts that are both ambiguous; that is, they correctly identify that more unevenly distributed urns are preferable in two-ball gambles. Furthermore, those same subjects went on to exhibit as strong a preference for $RR$ over $AA$ as subjects who were not exposed to such an opportunity to be “nudged” toward identifying more uneven urns as better. This suggests that subjects' preference for avoiding ambiguity -- even when it can only improve their odds of receiving a given prize -- cannot be due to failing to understand the gambles.
\paragraph{Contribution to the literature.}
The recent literature has produced several experimental tests and thought experiments that challenge existing decision-making models under risk and ambiguity. The following are the most resembling the one in this paper, either because they present some scenarios to decision-makers that resemble our two-ball gamble or because they reach similar conclusions to ours.
epstein2019ambiguous employ a similar two-ball gamble in one of their supplemental treatments from a 2014 experiment: subjects win the gamble if and only if two balls drawn with replacement from a single ambiguous urn (containing at most two different colors of balls) have the same color. However, they do not elicit subjects' Certainty Equivalents for this gamble and, crucially, do not observe the choice over a risky bet; they merely ask whether subjects prefer this gamble or a gamble that wins if and only if the first of the balls drawn is red (essentially, a 1-ball ambiguous gamble). Hence, from their results, one cannot directly estimate the proportion of people who prefer a 50-50 risky gamble to a two-ball ambiguous gamble. However, among subjects in their experiment who exhibited monotone and transitive choices, 21.6% exhibited a strict preference for the 1-ball ambiguous gamble over the 2-ball ambiguous gamble - a proportion that is consistent with our results when one takes into account the possible preference for a 50-50 risky gamble over a 1-ball ambiguous gamble.
fleurbaeyambiguity2019 constructs a thought experiment in which there is a risky urn $R$ containing 50 red and 50 blue balls and an ambiguous urn $A$ containing only red and blue balls; the decision maker draws two balls sequentially from some combination of these urns and wins if the two balls have the same color.\footnote{We thank Marc Fleurbaey for having provided access to his unpublished manuscript.} The decision maker must choose to either (1) draw a ball from urn $R$, observe its color, then choose to draw the second ball from either (a) urn $R$ (with replacement) or (b) urn $A$; or (2) draw a ball from urn $A$ and then a ball from urn $R$. The payoffs from winning vary slightly between options (1a), (1b), and (2) in such a way as to generate time-inconsistent or otherwise “unreasonable" behavior from certain theoretical types of decision-makers. In contrast to this thought experiment, our paper's central two-ball gamble compares two draws from urn $A$ to two draws from urn $R$, and the decision maker does not have any choices between the drawing of the first and second balls. Although both papers consider situations in which individuals may pay to avoid the mere presence of ambiguity, only our Two-Ball Ellsberg Paradox is an example in which individuals choose a dominated gamble to avoid such ambiguity.
yang2017testing designed an experiment in which two balls are drawn with replacement from a single urn containing only red and white balls; the subject wins 50 yuan if the first ball is red and independently wins an additional 50 yuan if the second ball is white. This gamble has the properties that (i) its expected payout is 50 yuan regardless of the composition of the urn and (ii) the variance of this payout decreases as the urn's contents become more dispersed (e.g., it has zero variance when the urn is all red balls or all white balls). Thus, when choosing whether to play this gamble drawing with replacement from a risky urn $R$ whose contents are known to be 50% each color versus an ambiguous urn $A$ whose distribution between red/white is unknown, a decision maker with risk-averse decision maker should think urn $A$ is better than urn $R$, assuming her preferences satisfy the monotonicity axiom. After eliciting subjects' risk attitudes, the authors estimate that as many as 45% of risk-averse subjects choose urn $A$ over urn $R$, violating all theories that include a monotonicity axiom. These results are qualitatively similar to the ones we find in our present paper, except that our central two-ball gamble's payoff has not only variance that decreases with the dispersion of the urn's contents but also has mean that increases with this dispersion. Hence, urn $A$ may be attractive to those who are risk-averse and those who are risk-seeking.
More broadly, our paper belongs to the literature on experimental tests and thought experiments questioning models of decision-making under risk and ambiguity. Among the most famous and relevant experiments, we can cite the following. Machina2009 presents thought experiments demonstrating plausible violations of the CEU model. MachinaParadoxes employ variations on Machina2009 to propose plausible violations of several of the decision-making models mentioned above. ReflectionExperiment test one of these examples, the so-called "reflection example" of Machina2009, in an experimental setting and reject the MEU and variational preferences models and the smooth ambiguity model of Klibanoff. Blavatskyy proposes a variant of Machina's reflection example, casting doubt on more recent decision-making models under ambiguity, such as VectorEU's Vector Expected Utility.
Halevy demonstrates that even in relatively simple settings, individuals' choices may be inconsistent with models that assume that individuals correctly reduce compound lotteries - an assumption that is implicitly part of models built on the framework of AnscombeAumann. He further shows that ambiguity aversion in the classic experiment of Ellsberg is closely correlated with an adverse reaction to compound lotteries. GSY replicate Halevy's experiment with a correction for measurement error, revealing that this correlation may be close to 1. Schneider examine whether subjects' preferences satisfy weak separability, a weak form of the monotonicity axiom of AnscombeAumann, and find that nearly half of the subjects violate it. Furthermore, subjects violating the monotonicity axiom generally make choices consistent with first-order stochastic dominance when choosing under risk, demonstrating that these violations are likely unrelated to a lack of understanding of their choices.
Our experiment distinguishes itself from these by proposing a new class of decision problems: Two-Ball Ellsberg drawings versus non-ambiguous drawings. As the Two-Ball Ellsberg drawings feature ambiguity but guarantee a minimum win probability at least as large as a related non-ambiguous gamble, they enable us to test whether a subject avoids ambiguity per se instead of avoiding ambiguity because it may yield a worse outcome.
\paragraph{Structure of the paper.} Section (ref) provides an overview of our experimental design and methodology. Section (ref) presents the results of our core gambles: the Two-Ball Ellsberg Paradox. Sections (ref), (ref), (ref) and (ref) describe and summarize the results of our experimental treatments meant to test various hypotheses about what may be generating the Two-Ball Ellsberg Paradox. Section (ref) discusses some remaining plausible explanations of subjects' behavior, and section (ref) concludes.
Our experiment was designed to answer two primary questions. First, to what extent do subjects prefer urn R over urn A in our two-ball gamble? Second, what possible explanations of this “paradoxical” preference can be falsified? Answering the first question only requires asking subjects about a few different gambles. However, since many possible explanations exist for a preference for urn R over urn A, our experiment includes many gambles designed to address the second question.
We used Prolific to run our experiment and collect our data. Prolific is an online survey platform that, due to its participant pool's quality, is increasingly used in economics to conduct surveys and incentivized experiments. Our sample comprised 880 participants, selected to be nationally representative in age and gender. Of these initial 880 participants, 708 passed the basic attention-screening questions and criteria described at the end of this section.
Due to the constraints on subjects' time and attention inherent in an online experiment, our various gambles were divided across four treatments, with each subject completing exactly one treatment. All treatments ask subjects about our central two-ball gamble (playing with urn A versus urn R), and all treatments elicit subjects' ambiguity attitudes via the classic two-urns Ellsberg paradox. Beyond this, each treatment contains some gambles specific to that treatment. Similar gambles were grouped into blocks, and gambles within a block were presented in random order.\footnote{The order of the blocks was also randomized; we detail the particular randomization for each treatment as elaborated in the sections (ref), (ref), (ref) and (ref).}
In each gamble, the subject can either "win" (gain \$3) or "lose" (gain nothing). After viewing instructions explaining the conditions under which the current gamble will win or lose, the subject must report her certainty equivalent (CE) for that gamble from a multiple price list (MPL) containing dollar amounts between \$0 and \$3 in increments of 10 cents. Compared to eliciting choices, the MPL allows us to measure the intensity of subjects' preferences.
Laboratory and online experiments eliciting subjects' CEs for gambles are often prone to significant measurement error. To correct this, we rely on the Obviously Related Instrumental Variables (ORIV) method of GSY. Compared to other methods to correct measurement errors, such as using the first elicitation as an instrument for the second, the ORIV approach generally results in lower standard errors. We, therefore, elicit subjects' CEs twice for most of our gambles.
Including all duplicate questions, each treatment contains 11 or 12 gambles in total. In each treatment, three\footnote{Although incentivizing only one gamble would allow us to raise the monetary stakes of each question, doing so would create too large a variance in different subjects' payoffs, which was undesirable for this online experiment.} of these gambles were selected at random for incentivization: if a gamble was selected, then a random row of the MPL for that gamble was chosen, and subjects were given what they reported they preferred from that row.\footnote{For example, if Gamble X was selected for incentivization, and then the row “\$1.20” was selected at random for this gamble, the following happens. (A) If the subject reported she preferred a fixed \$1.20 payment to play Gamble X, then she received \$1.20. (B) If the subject reported she preferred playing Gamble X to receiving \$1.20, then we simulated Gamble X and gave her \$3 if it won and \$0 if it lost.} Subjects received an average payment of \$3.50 from the incentivized questions, plus a fixed \$2 payment for completing the experiment.
Since the monetary stakes of the experiment were not very high, there is a reason for concern that subjects may answer at random to quickly finish the experiment. We employed three screening criteria to address this concern: (1) After the experiment instructions, but before the gambles, subjects were given a 3-question basic comprehension quiz about the instructions. Any subject who failed at least one of these questions was given a small payment and forced to leave the experiment. (2) Subjects were given a standard attention-screening question between each of the experiment's major sections. Subjects failing at least one such question were removed from our analysis. (3) If, across our two elicitations of a subject's CE for the same gamble, the subject reported two CEs that differed by more than \$1 (that is, one-third the size of the \$3 MPL table), that subject was removed from our analysis.\footnote{Other reasonable thresholds for exclusion, such as “differed by more than \$1.50,” yield qualitatively similar results in our analysis as detailed in the Appendix.} Out of an initial pool of 880 subjects, 172 were removed due to violating at least one of the criteria (1)-(3).
The block 2Ball contains this experiment's central gambles and is present in all four of our treatments. It contains two gambles, named $RR$ and $AA$. The block Ellsberg replicates the classic Ellsberg paradox to elicit subjects' attitudes towards risk and ambiguity and is also present in all four of our treatments; it contains two gambles named $R$ and $A$. Table (ref) describes these gambles.
The blocks 2BallD and EllsbergD contain duplicate gambles of those in blocks 2Ball and Ellsberg. The standard practice when double-eliciting CEs requires the two “duplicate” gambles measuring the same CE to have slightly different wordings so as two constitute two independent measurements of that CE. To accomplish this, whenever we duplicate a block of gambles, we slightly change the specified total number of balls in a given urn without changing the \textit{proportion} of balls of each color. For example, in block \textbf{2BallD}, urn \textbf{R} contains 40 red and 40 blue balls rather than 50 red and 50 blue.
For each gamble $X$ that is double-elicited, we use the notation $X_i^j$ to represent the $j$-th elicitation of subject $i$'s CE for gamble $X$, and we use the notation \[ X_i = \frac{X_i^1 + X_i^2}{2} \] to denote the average CE of subject $i$ for gamble $X$. So, for example, $RR_{36}^2$ represents the 2nd elicitation of subject 36's CE for gamble $RR$, and $A_{15}$ denotes subject 15's average CE for gamble $A$. Figure (ref) shows the CDFs of the empirical distributions of the CEs for $RR$, $AA$, $R$, and $A$; Table (ref) in Appendix (ref) contains summary statistics for each elicitation of these CEs.
Other than at a few extreme CE values that were reported by a total of less than 10% of subjects, these empirical CDFs lie in the same vertical order everywhere. This suggests that on average, subjects prefer the gambles in the order \[ R \succ RR \succ A \succ AA. \]
Nearly all widely-used models of decision making under risk and ambiguity cannot explain a preference for $R$ over $AA$ or a preference for $RR$ over $AA$, since gamble $AA$ has a win probability of at least 50% while gambles $R$ and $RR$ have a win probability of exactly 50%. Throughout this paper, we use the variable $R-AA$ to measure the extent to which individuals exhibit the “2-Ball Ellsberg Paradox.” Although it may be slightly more natural to compare gamble $AA$ to gamble $RR$, we choose to compare it to gamble $R$ since gamble $R$ provides a more standard basis of comparison for the other gambles in our experiment, such as the compound lottery $C$ in treatment complexity.\footnote{See Section (ref) below.} In any case, among the standard models of decision making, $R-AA$ taking a statistically significant positive value is sufficient to falsify all the same models would be falsified by a statistically significant positive value of $RR-AA$. Figure (ref) shows the distribution of individuals' reported CE differences $R^j-AA^j$ in each of the two elicitations $j$.
Averaging across both elicitations, a majority (54.9%) of subjects exhibited 2-Ball Ellsberg Paradox preferences by reporting a value $R-AA$ that was greater than zero. The average CE for gamble $R$ is 118.13 cents, while the average CE for gamble $AA$ is only 101.03 cents. The 17.1 cent difference between these averages is statistically significant $(t = 11.7)$; individuals are willing to pay about 17% more for gamble $R$ than they are for the higher-win-probability gamble $AA$.
One natural explanation for subjects' preference for gamble $R$ over gamble $AA$ is that subjects fail to understand the fact that gamble $AA$ must have at least as large of a win probability than gamble $R$. Our treatment nudging was designed to test this hypothesis in two ways. First, we present subjects with a variation of the 2-ball ambiguous gamble wherein we place bounds on the contents of urn A to test if subjects can successfully identify that more unevenly distributed urns yield higher win probabilities in 2-ball gambles. We find that subjects \textit{do} correctly identify this -- or at least, they make choices consistent with such an understanding. Second, we check if those subjects who were exposed to Treatment \textsc{ nudging} -- who "correctly" identified more unevenly distributed urns as more preferable -- tended to exhibit less of a preference for gamble $R$ over gamble $AA$ than those who were not exposed to this treatment. We find that subjects in Treatment \textsc{ nudging} do not show a statistically significant difference in their values $R-AA$ from subjects not exposed to this treatment ($p=.63$).
The gambles unique to treatment nudging are those in block BoundedA. In this block, subjects play a 2-ball gamble: two balls are drawn from an urn A containing 100 balls, all red or blue, but whose exact contents are unknown. The subject wins \$3 if the two balls have the same color. In each gamble in block BoundedA, some further information is given about the contents of urn A, as described in Table (ref).
In treatment nudging, subjects complete the blocks BoundedA, Ellsberg and 2Ball as well as the duplicate blocks EllsbergD and 2BallD. The order in which these blocks were presented was determined at random, independently for each subject assigned to this treatment, according to Figure (ref).
In this figure, the initial split between line 1 (with BoundedA at the beginning) and line 2 (with BoundedA at the end) indicates that subjects were randomized uniformly between doing block BoundedA either before or after all the other blocks in the treatment. Furthermore, the fact that the boxes containing “Ellsberg, 2Ball” and “EllsbergD, 2BallD” are adjacent and shaded, in the same way, indicates that, within each of these two randomized groups, there is further randomization as to whether the blocks Ellsberg and 2Ball are both completed before blocks \textbf{EllsbergD} and \textbf{2BallD} or are both completed after these two blocks. Finally, in any box containing multiple blocks, those blocks were completed in a random order (e.g., block \textbf{Ellsberg} is either completed before or after block \textbf{2Ball}). Hence, Figure (ref) indicates that there are 16 possible orders in which subjects could complete the blocks in treatment \textsc{ nudging}.
The first main result of this treatment is that participants prefer more unevenly distributed urns when all involve ambiguity. Table (ref) below summarizes subjects' CEs for the gambles unique to block nudging. Subjects prefer $BB^{60-100}$ to $BB^{40-60}$ by an average of 33.6 cents ($t > 9$). Similarly, they prefer $BB^{95-100}$ to $BB^{60-100}$ by an average of 75.4 cents ($t > 13$). Of the 179 subjects in treatment nudging, only 22 subjects reported a larger CE for $BB^{40-60}$ than $BB^{60-100}$; and even among those 22 subjects, the average CE for $BB^{95-100}$ was massively larger than the average CE for $BB^{60-100}$ (mean of difference = 68.64, $t=3.36$).
These data indicate that in 2-ball gambles wherein all options involve ambiguity, subjects “correctly” report significantly larger CEs for gambles drawing from more unevenly distributed urns. This suggests that the preference for gamble $RR$ over gamble $AA$ may be due to a preference to avoid the mere presence of ambiguity and not due to a lack of understanding that more unevenly distributed urns yield higher win probabilities in 2-ball gambles.
The second main result of this treatment is that participants are not “nudged” by bounded ambiguity. One might argue that, while subjects' choices in block BoundedA indicate a significant understanding of the fact that more unevenly distributed urns yield higher win probabilities in 2-ball gambles, this understanding does not exist among subjects who were not exposed to the “leading questions” found in block BoundedA. Indeed, perhaps subjects only come to understand this fact when confronted with stark examples such as an urn containing at least 95% red balls.
Suppose the hypothesis in the previous paragraph was correct. In that case, we should expect to find that the CE difference $R-AA$ is significantly smaller (or, more negative) among subjects who completed block BoundedA before completing blocks 2Ball and 2BallD than it is among subjects who did not complete BoundedA before 2Ball and \textbf{2BallD}. Completing block \textbf{BoundedA} should “nudge” subjects into being less susceptible to the 2-ball Ellsberg paradox.
Half of the 179 subjects randomly assigned to treatment nudging completed BoundedA before the blocks 2Ball and 2BallD, whereas none of the subjects randomly assigned to other treatments did so. So if a nudging effect exists, then it should manifest as a statistically significant (negative) difference between the $R-AA$ values in the treatment \textsc{ nudging} versus those in the other treatments.
If we let $I^{TN}$ be the indicator variable for assignment to Treatment nudging, then in a regression of $Z := R - AA$ on $I^{TN}$, the slope coefficient represents the causal effect of being in Treatment nudging on the preference for $R$ over $AA$. A statistically significant negative slope coefficient would indicate that Treatment nudging has a nudging effect, causing subjects to manifest less preference for $R$ over $AA$.
Table (ref) shows the results of such a regression, first using individual elicitations and then the averages across elicitations. As shown, the slope coefficient is not statistically significant, and it is positive in the case using averages. Thus, we fail to reject the hypothesis that there is no nudging effect ($p=.63$). The results of Treatment nudging therefore provide strong evidence that subjects' preference for gamble $R$ over gamble $AA$ has little to do with a lack of understanding that gamble $AA$ has a larger win probability.
One might argue that the preference for gamble $R$ over gamble $AA$ is not due to an aversion to the ambiguity present in gamble $AA$ but instead to the complexity present in gamble $AA$. "Complexity" is a concept difficult to define precisely, and it is not the aim of this paper to do so. However, experiments like Halevy's have established the potential relevance of specific types of complexity, such as the compoundness of lotteries. With this in mind, we test whether the preferences for gamble $R$ over gamble $AA$ is indistinguishable from the preference for a simple 50-50 gamble like $R$ over a compound 50-50 gamble, call it $C$ as described in Table (ref). We designed Treatment complexity to test whether these specific types of complexity may be the primary factors generating the Two-Ball Ellsberg paradox.
The ORIV-corrected correlation between the “Two-Ball Ellsberg paradox” preference $R-AA$ and the “aversion to compound lotteries” preference $R-C$ is extremely high, but the strength of preference $R-AA$ is greater than that of preference $R-C$. We find similar results when we compare $R-AA$ to other “paradoxical” preferences, such as the classic Ellsberg paradox preference $R-A$.
The gambles unique to treatment complexity are those in block Compound. In this block, subjects play two gambles involving an urn C that contains 100 balls, all either red or blue. Subjects are informed that before each gamble begins, the contents of urn C are determined uniformly at random (i.e., each of its 101 possible balls compositions are equally likely to be realized). Table (ref) summarizes these gambles.
In other words, block Compound consists of two gambles: a compound lottery $C$ and a “Two-Ball Compound” gamble $CC$. Gamble $CC$ is the same as the ambiguous gamble $AA$, except its urn's contents are determined by a known lottery rather than an unknown, ambiguous procedure.
Block CompoundD contains duplicate questions of those in block Compound. In treatment complexity, subjects complete the blocks Compound, Ellsberg and 2Ball as well as the duplicate blocks \textbf{CompoundD}, \textbf{EllsbergD} and \textbf{2BallD}. The order in which these blocks were presented was determined at random, independently for each subject assigned to this treatment, according to Figure (ref). Its interpretation is analogous to that of Figure (ref); there are 12 different orders in which the six blocks comprising Treatment \textsc{ complexity} could be completed. Table (ref) in Appendix (ref) contains summary statistics for each elicitation of CEs for gambles $C$ and $CC$.
The variable $R-A$ measures subjects' ambiguity aversion in the classic Ellsberg paradox, while $R-RR$ measures their preference for a simple 50-50 gamble to a Two-Ball 50-50 gamble. $R-C$ measures subjects' preference for a simple 50-50 gamble over a Compound 50-50 gamble, and $R-CC$ measures their preference for a simple 50-50 gamble over a Two-Ball Compound 50-50 gamble. Table (ref) in Appendix (ref) contains summary statistics for each elicitation of these CE differences.
Table (ref) computes the ORIV-adjusted correlations between our central variable $R-AA$ and these other variables.
As the table shows, the preference for $R$ over $AA$ is extremely tightly correlated with each of the preferences mentioned in the previous paragraph. Thus, from the analyst's point of view, a subject exhibiting one of these “paradoxical” preferences to a certain degree of strength (as measured by standard deviations above the population mean) makes it exceedingly likely that she will exhibit these other “paradoxical” preferences to a similar degree of strength. In particular, this finding replicates Halevy's and GSY's conclusions that ambiguity aversion in the classic Ellsberg paradox is tightly linked to failure to reduce compound lotteries.
Besides correlations, it is worthwhile to examine the differences between the variables in the table above. $R-AA$ is larger than all of $R-A$, $R-RR$, and $R-C$ ($t > 4$ in all cases) and is larger than $R-CC$ by a statistically insignificant amount ($t = 1.05$). This suggests that, according to most subjects, gamble $AA$ is likely the “worst” of gambles $AA$, $A$, $RR$, $C$, and $CC$ - perhaps because gamble $AA$ combines ambiguity and Two-Ball complexity. The only possible competitor for being the “worst” is gamble $CC$, which is identical to gamble $AA$ except that its urn's contents are determined randomly rather than in an ambiguous manner.
Another hypothesis that might explain a preference for gamble $R$ over gamble $AA$ is that subjects incorrectly believe that the contents of urn A can change between the two draws (with replacement) made from it in gamble $AA$. Let us call this the False Independence hypothesis. If this hypothesis were true, then gamble $AA$ need not have a higher win probability than gamble $R$, and may be less desirable since its win probability is ambiguous while that of gamble $R$ is not. We designed Treatment robustness to test the False Independence hypothesis, as well as to contain exploratory gambles meant to motivate further research that will be discussed in Section (ref).\footnote{Due to constraints on the maximum number of gambles we could fit in a given treatment, we could not fit all such exploratory gambles in a single treatment. Thus, we split them across two treatments.}
The gambles unique to treatment robustness are those in blocks Independent and 3Ball. In block Independent, subjects draw a ball from each of two ambiguous urns (containing only red and blue balls) whose contents were determined independently; they win \$3 if the two balls have the same color. In block 3Ball, subjects draw 3 balls in total, with replacement, from some combination of a single ambiguous urn A and a single risky urn \textbf{R}, in a certain order. They win \$3 if \textit{all three} balls have the same color. The gambles in these blocks are summarized in Table (ref).
In treatment robustness, subjects complete the blocks Independent, 3Ball, Ellsberg and 2Ball as well as the duplicate blocks EllsbergD and \textbf{2BallD}. The order in which these blocks were presented was determined at random, independently for each subject assigned to this treatment, according to Figure (ref). Its interpretation is analogous to that of Figure (ref); there are 48 different orders in which the six blocks comprising Treatment \textsc{ robustness} could be completed.
In gamble $IA$, the only one in block Independent, the mean of the subjects' CEs was 107.839, and its standard deviation was 68.733. We discuss the gambles in block 3Ball in Section (ref).
On average, the 192 subjects in Treatment robustness slightly preferred $AA$ to $IA$, but the difference is not statistically significant (mean = 1.73, $t=.60$). An indifference between $AA$ and $IA$ would be consistent with the False Independence hypothesis, while a strict preference for $AA$ over $IA$ would falsify this hypothesis. We, therefore, fail to reject the False Independence hypothesis. A future experiment that replicates block Independent with a larger sample size or larger payments may be able to falsify it; or, see Section (ref) for discussion of a variation on block BoundedA that may be able to shed further light on this hypothesis.
A further hypothesis that could explain a preference for gamble $R$ over gamble $AA$ is that the mere presence of ambiguity in gamble $AA$ makes it undesirable to subjects or that, more generally, gambles containing a larger “amount” of ambiguity are less preferable (all else equal). The purpose of this paper is not to precisely articulate what a “distaste for the mere presence of ambiguity” means or to define the “amount of ambiguity” present in a gamble, and it is difficult to separate concepts like these from a distaste for the presence or amount of complexity.
Hence, it is difficult to test the hypothesis that the Two-Ball Ellsberg paradox is due to a distaste for the amount of ambiguity present in gamble $AA$. Nonetheless, we expected that eliciting subjects' CEs for certain gambles closely related to $AA$ - namely gambles $AR$ and $RA$ described below and gambles $AAA$, $RRR$ and $RAA$ from Section (ref) - may shed light on this hypothesis. We thus designed Treatment order, and block 3Ball from Treatment robustness, to provide additional information that may motivate future research.
The gambles unique to treatment order are found in block 2BallMixed - an expanded version of block 2Ball that contains gambles not only $RR$ and $AA$ as before but also gambles $AR$ and $RA$. In each gamble, subjects draw two balls - either from the same urn and with replacement or from distinct urns - in a certain order, and they win \$3 if the two balls have the same color. Table (ref) summarizes these gambles.
Block 2BallMixedD contains duplicate questions of those in block 2BallMixed. In treatment order, subjects complete the blocks 2BallMixed and Ellsberg as well as the duplicate blocks 2BallMixedD and \textbf{EllsbergD}. The order in which these blocks were presented was determined at random, independently for each subject assigned to this treatment, according to Figure (ref). Its interpretation is analogous to that of Figure (ref); there are 8 different orders in which the 4 blocks comprising Treatment \textsc{ order} could be completed.
Table (ref) in Appendix (ref) contains summary statistics for each elicitation of CEs for gambles $RR$, $AA$, $AR$, and $RA$. Besides, table (ref) in Appendix (ref) contains summary statistics for each elicitation of the differences between the CEs $RR$, $AA$, $AR$, and $RA$.
The average CEs for the four gambles in block 2BallMixed are ranked in the order \[ RR > AA > AR > RA, \] but the only statistically significant differences between these variables are those between $RR$ and each of the other three. Hence, we cannot rule out the possibility that subjects' average preferences are of the form \[ RR \succ AA \sim AR \sim RA. \]
It is worth noting that gambles $RA$ and $AR$ each have a win probability of exactly 50%: whatever ball is drawn from urn A, the ball from urn R has a 50% chance of matching it. Hence, if subjects care only about win probabilities and exhibit no aversion to the mere presence of ambiguity, then we should expect a preference ordering like $AA \succsim RR \sim RA \sim AR$. Meanwhile, if subjects do not care at all about win probabilities and respond solely based on a distaste for the mere presence of ambiguity, then we should expect a preference ordering like $RR \succ AR \sim RA \succ AA.$\footnote{Here, we have informally formulated a “distaste for the mere presence of ambiguity” to be such that each additional instance of a draw from an ambiguous urn makes the overall gamble more distasteful.} A strict preference for RR over RA may also be explainable by subjects failing to realize that the preference of urn R automatically hedges against the ambiguity in urn A and causes gamble RA to have an unambiguous probability of winning of 50%. Lastly, strict preference between $AR$ and $RA$ would indicate order sensitivity and that a more intricate explanation is required.
Although we do not find significant evidence of order sensitivity, we find that neither a total indifference to the mere presence of ambiguity nor a preference based solely on avoiding the mere presence of ambiguity, is sufficient to explain subjects' behavior. Indeed, gamble $AA$ is neither at least as good as all three other gambles (as the first simplistic theory would predict) nor worse than all three other gambles (as the second would predict). We speculate that subjects may find gamble $AA$ to be at least as good as gambles $AR$ and $RA$ because its larger probability of winning (than the 50% given by $AR$ or $RA$) offsets its increased presence of ambiguity, but find $AA$ to be worse than $RR$ because $RR$'s complete lack of ambiguity makes it significantly more attractive.
Recall the gambles in block 3Ball summarized in Table (ref) above. Table (ref) presents summary statistics of subjects' CEs for these gambles.
These reported CEs are too large for a classical risk-averse agent who correctly calculates the probabilities of winning.\footnote{Our results from the simple 50-50 gamble in block Ellsberg suggest that subjects are on average slightly risk averse.} Notice that $RRR$ has a win probability of exactly $\frac{1}{4}$, but subjects report an average CE of 97.7 cents for this gamble - a value significantly larger than the risk-neutral CE of 75 cents ($t=4.67$). Similarly, subjects on average value gamble $RAA$ significantly at more than half as much as gamble $AA$ (difference of means = 37.77, $t=11.03$). Thus, subjects seemingly overweight the win probabilities of 3-Ball gambles.
Despite the general overweighting of win probabilities, comparisons between CEs for these 3-Ball gambles remain qualitatively similar to the comparisons between the CEs for Two-Ball gambles. Similarly to how subjects on average preferred $RR$ to $AA$, we find that subjects on average prefer $RRR$ to $AAA$ (mean = 6.59, $t=2.16$), even though $AAA$ must have at least as large of a win probability as $RRR$. Also, just as we did not find a statistically significant difference between $RA$ and $AA$ in Section (ref), we find no statistically significant difference between $RAA$ and $AAA$ (mean = 1.43, $t=.56$).
Likewise, these results suggest that perhaps the “amount” of ambiguity present in a gamble, measured in terms of the total number of (or proportion of) draws that come from ambiguous urns, does not matter as much as the mere presence of ambiguity at all.
\paragraph{Distaste for Ambiguity or Complexity?} Our treatment nudging suggests that individuals' choice of gamble $RR$ over $AA$ is not due entirely to a lack of cognition or a failure to reduce compound lotteries. Nonetheless, one might argue that the preference for $RR$ over $AA$ is due to a distaste for complexity rather than ambiguity. Even in this case, we have identified the mere presence of ambiguity as a driver of change in people's behavior, perhaps through the complexity, it introduces or perhaps through other means.\footnote{If ambiguity is simply a type of complexity, then models of contingent reasoning such as martinez2019failures may be good candidates to explain our results.}
One might further suggest that the preference for $RR$ over $AA$ comes from ”inappropriate” beliefs, i.e., beliefs that are not the product of the same distribution over the contents of the urn. Our treatment robustness was designed to address this concern but its results lacked the statistical significance to rule out this explanation.
Whether explained or not as an instance of complexity, people harboring a distaste for the mere presence of ambiguity has potentially widespread implications for economics. Subjects may prefer to gamble $R$ to $A$ in the classic Ellsberg paradox primarily because they dislike the mere presence of ambiguity and not, for instance, entirely because they hold concern for worst-case scenarios, as GilboaSchmeidler would suggest. Models ignoring a distaste for ambiguity per se would incorrectly predict individuals' behavior in a variety of situations. Hence, new models may be required.
\paragraph{Raiffa Critique.} Unlike in the original Ellsberg paradox, a subject cannot eliminate the ambiguity present in gamble $AA$ by introducing randomization in her choice of color (as in raiffa1961risk). Indeed, gamble $AA$ does not ask subjects to choose a color. Even if we presented subjects with a modified version of gamble $AA$ wherein they choose either red or blue and win if and only if both balls drawn were of the chosen color (and compared this to a similarly modified version of gamble $RR$), it is still the case that randomizing one's color choice does not eliminate the ambiguity in the payoff of gamble $AA$. If $p$ is the (ambiguous) proportion of red balls in urn A, then this modified version of gamble $AA$ has win probability $p^2$ when you bet on red and win probability $(1-p)^2$ when you bet on blue.
Randomizing your choice of color 50-50 would thus mean that the gamble's win probability is $.5p^2 + .5(1-p)^2 \geq .25$. In contrast, the modified version of gamble $RR$ has a probability of winning $.25$ regardless of the color on which you bet (or whether you randomized your choice of color). It is still the case that gamble $AA$ has an ambiguous win probability and that it is at least as large as (and in all but one case, strictly larger than) that of $RR$.
\paragraph{A different experiment to reject the False Independence hypothesis.} Recall the "False Independence" hypothesis mentioned in Section (ref): Do subjects imagine that our "two draws with replacement from the same ambiguous urn" are actually "two draws from two ambiguous urns whose contents were determined independently"? Our experiment can't rule out the False Independence assumption as a driving factor in the 2-Ball Ellsberg paradox, but here we suggest how a further experiment might do so.
A variation on block BoundedA may be sufficient to show that the False Independence assumption cannot fully explain our results. Consider a version of gamble $BB^{95-100}$ wherein instead of the gamble specifying that the urn contains between 95 and 100 red balls, it merely specifies that at least 95 of the 100 balls in the urn are of the same color. Suppose subjects imagined the two draws from the specified urn as “one draw from each of two distinct urns, whose contents were determined in a specified manner but were determined independently.” Then we should not find a strong preference for this version of gamble $BB^{95-100}$ over gamble $AA$.
Indeed, suppose subjects believe in False Independence. In that case, they might easily imagine this new version of gamble $BB^{95-100}$ to have a win probability close to 50%. For although it is possible in their minds that "both urns" contain at least 95 red balls (or that both contain at least 95 blue balls), it is equally possible to them that "one urn contains at least 95 red balls while the other contains at least 95 blue balls". In other words, their CEs for this version of gamble $BB^{95-100}$ should certainly not be radically larger than their CEs for gamble $AA$. If such a radical difference in CEs as we found between the original version of gamble $BB^{95-100}$ and gamble $AA$ were still found under this modified version of $BB^{95-100}$, this would suggest that False Independence is not the primary factor generating our results.
Two-Ball gambles are a rich class of decision problems. Because they can involve ambiguity but guarantee a minimum win probability that is at least as large as that of some other gamble, they allow us to test whether subjects avoid ambiguity per se as opposed to avoiding ambiguity because it may yield a worse outcome.
The most striking case of preferring a gamble with lower win probability is that subjects preferred the 50-50 gamble, $R$, to the Two-Ball ambiguous gamble, $AA$. This preference is closely correlated with the traditional Ellsberg preference for $R$ over a 1-Ball ambiguous gamble $A$, and also with the preference for $R$ over the compound 50-50 gamble $C$, as well as the preference for $R$ over the Two-Ball 50-50 gamble $RR$. These close relationships suggest that it may be difficult to separate an aversion to ambiguity per se from an aversion to complexity.
It is implausible that subjects prefer $R$ to $AA$ simply due to a poor understanding of Two-Ball gambles. In the block BoundedA, subjects correctly and strongly identified that more unevenly distributed urns are more likely to win. Moreover, the lack of a "nudging" effect from being in the treatment containing BoundedA suggests that subjects' preference for $R$ over $AA$ is deliberate.
Although the presence of an ambiguous draw within a gamble is associated with a significantly lower CE, it remains unclear whether having more ambiguity, as measured perhaps by the number or proportion of ambiguous draws present in a gamble, has an additional negative effect on the CE. This presents an interesting question for further research.
\singlespacing