Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
103,803 characters · 19 sections · 63 citation commands
Decision Conflict, Power Logit, and the Deferral Outside Option
\\ University of Glasgow\ \\ \today }
{ Keywords:\\ Decision difficulty; outside option; choice deferral; quality/price competition; discrete choice; estimation.}
\thispagestyle{empty}
\setcounter{page}{1}
It is a well-established fact that people often “choose not to choose” when they find it hard to compare the active-choice alternatives available to them, even when all these alternatives are individually considered “good enough” and are paid attention to. Real-world examples of such behavior include: (i) employees who operated within an “active decision” pension-savings environment and did not sign up for one of the plans that were available to them within, say, a day, week or month of first notice, possibly even opting for indefinite non-enrolment;\footnote{Such behavior is documented in carroll-choi-laibson-madrian-metrick, for example.} (ii) patients who, instead of choosing “immediately” one of the active treatments that were recommended to them against a medical condition, delayed making such a choice---often at a health cost---due to “facing a treatment dilemma”;\footnote{See knops-goossens-ubbink-legemate-stalpers-bossuyt13. See also oconnor95,DecConfScale19 for additional references and overview of the use of a “decisional conflict scale” in medical decision making that was developed “to measure a person's perceptions of their uncertainty in making a choice about health care options, the modifiable factors contributing to uncertainty, and the quality of the decision made”.} (iii) doctors who were willing to prescribe the single available drug to treat a medical condition but were not prepared to prescribe anything when they had to decide from the expanded set that contained one more drug, because “the difficulty in deciding between the two medications led some physicians to recommend not starting either” redelmeier-shafir95. \footnote{Other works that find evidence associating decision difficulty with such choice paralysis in different environments include tversky-shafir92,dhar97,dhar-simonson03,danan&ziegelmeyer06, bhatia-mullett16,CCGT22.} Even though in many such cases people do end up making an active choice eventually, the following question arises for the decision analyst: \textit{Should the observed individual's initial avoidance/delay to make such a choice be ignored as non-relevant?} Put more concretely, should the decision of a medical doctor who eventually prescribes one of the two possible medications after several reminders be thought of as similar to their decision to prescribe that medication without delay in a different situation? In this paper we regard the decision to avoid/delay making an active choice that is caused by such context-specific “decision conflict” as providing meaningful information about both the individual and the alternatives in question, regardless of which active choice---if any---was eventually made at that menu at a later time.
In their influential monograph, janis-mann77 defined “decision conflicts” as the “simultaneous opposing tendencies to accept and reject a given course of action” and identified “hesitation, vacillation, [and] feelings of uncertainty” to be among their most prominent symptoms “whenever the decision comes within the focus of attention”.\footnote{pochonetal08 is a targeted study in the neuroscience literature on the brain regions that are activated when subjects face decision conflict.} Motivated by the relevance of hesitation-driven opt-out decisions for understanding preferences and explaining behavior, our specific goal in this paper is to model choice in the presence of a choice-avoidance/deferral outside option within a stochastic choice framework in ways that deviate as little as possible from existing well-understood modelling practices and, at the same time, make predictions that are in line with some findings from the empirical/experimental literature and evade existing models. We pursue this by extending in disciplined ways the foundational luce59/logit model and its econometric specification pioneered by mcfadden73. Specifically, we propose the class of decision-conflict logit models which, in their most general form, are a straightforward but so far unexplored extension of the logit with an outside option that assign a menu-dependent value to that option while retaining the menu-invariance assumption on all active-choice alternatives. The relative value of the outside option at a menu in turn determines the probability of avoiding/deferring choice and can be interpreted as proxying decision difficulty.
We show that, despite its simplicity, this general model can be the starting point for richly structured special cases. In particular, we introduce and focus on the broad class of power logit models that are examples of of such cases where decision difficulty depends in intuitive ways on the logit values of all active-choice alternatives. In these models, decision difficulty could be thought of as driven by the agent's noisy resampling of the menu's elements. In the quadratic logit special case of this class of models, for example, such resampling takes the form of the choice probability of a market alternative emerging as the product of two logit probabilities according to a single value function/criterion. Intuitively, the agent is more likely to choose an active-choice alternative if and only if its value realizations according to this criterion are much larger than those of everything else feasible across both rounds of sampling. Conversely, the agent is more likely to avoid/defer choice when no alternative achieves such unanimous clear dominance. This model could therefore be thought of as capturing a hesitant decision maker who behaves as if they used an objective criterion to compare alternatives (e.g. sum or multiply each option's values across all relevant attributes) but is aware that their subjective evaluation according to this objective criterion may be imperfect, possibly due to cognitive limitations, thereby leading them to performing this task twice. This model and its power-logit generalization appear to be the first to provide a theory where the no-choice outside option is feasible and has an endogenously determined, menu-dependent value.
We further show that these structured models explain the following empirical phenomena that various studies in cognitive and consumer psychology have documented about decisions that allow agents to avoid/delay making an active choice:
We illustrate the applicability of our analysis both in theoretical and empirical settings. In our theoretical application we analyse a duopolistic market where firms compete simultaneously in price and quality under the common-knowledge assumption that consumer demand is determined by the power-logit model where a product's value is defined, intuitively, by its quality/price ratio. We derive simple and economically interpretable intuitive closed-form solutions for all equilibrium variables in the model: price, quality, profits and a notion of consumer welfare that appears suitable in environments where consumers opt out due to indecisiveness or overload. A key feature of the (symmetric) equilibrium is that, as the power parameter capturing consumers' decision difficulty increases, firms increase their products' quality/price ratio and see their profits decreased, both because of the reduced profit margins and because of the lower share of consumers who buy any product. Intuitively, this is driven by each firm increasing its quality/price ratio in an effort to reduce the consumer's decision difficulty and mitigate the risk of losing them to the rival firm or of driving them out of the market altogether.
In our empirical application we first show how the classic assumptions and argument that underpin the discrete-choice formulation of the logit without an outside option mcfadden73 must be modified and extended in order for both the quadratic logit and the more general power logit models to admit a similar discrete-choice formulation and be taken to the data for maximum-likelihood estimation of their respective parameters. We then show the potential fruitfulness of such analyses by estimating both the quadratic and power logit models on the deferral-permitting discrete-choice data with film decisions from the survey experiment of bhatia-mullett16, using the participants' subjective ratings of the different films as the explanatory variable. To assess the added value of the hereby proposed models on these data, we use standard criteria to evaluate their goodness of fit and compare them to those of baseline logit models with a fixed and inferior or a random outside option. Our analysis suggests that both the power and quadratic logit often perform better compared to either version of the baseline logit under these performance criteria, particularly in those situations where theory suggests they would do so. Hence, the proposed power-logit framework could be considered in the analysis of similar datasets whenever the researcher suspects that the observed opting-out/deferring behavior might be due to decision difficulty rather than the relative unattractiveness of the available active-choice alternatives.
Let $X$ be the grand choice set of finitely many active-choice alternatives with generic elements $a,b\in X$. Let $\mathcal{M}:=\{A: \emptyset\neq A\subseteq X\}$ be the collection of all menus of such alternatives, and $\mathcal{B}$ its sub-collection that comprises all binary menus. The outside option is denoted by $o\not\in X$. We clarify that this is not a status quo option (e.g. a tenant's current rental agreement) which can, in principle, be compared to the other feasible alternatives (e.g. other housing options) on the same or a similar set of relevant attributes. Instead, this option is devoid of attributes and its value to the decision maker is unobservable to the analyst.\footnote{hensher-rose-greene15 , for example, describe this distinction as follows: “At this point, it is worthwhile considering choice situations in which there exists the possibility to `choose not to choose', or to remain with some status quo alternative. Many choice situations present decision makers with examples of both types of alternatives. For example, a person can elect to stay at home and not see a movie if three potential movie alternatives showing at a local cinema at some preferred time do not appeal to them. Likewise, a decision maker facing the expiration of their rental agreement may elect to simply renew their current rental contract or move apartments, hence signing a new lease. In the case of a no choice alternative, the alternative labelled \textnormal{`none'} will be devoid of any attribute levels (e.g., there is no movie ticket price, no time spent at the cinema, etc., associated with going to the movies). The absence of attributes, however, does not mean that the decision maker is indifferent to that alternative. In the movie ex- ample, if the three movies on offer are romantic comedies, then staying at home and not attending any of them might be the most preferred option.” A formal distinction in the treatment of status-quo and choice-deferral outside options in a deterministic choice-theoretic framework is provided in gerasimou16a.} A \textit{random free-choice model} on $X$ is a function $\rho:X\times \mathcal{M}\rightarrow\mathbb{R}_{+}$ such that $\rho(a,A)\in [0,1]$ for all $A\in\mathcal{M}$ and all $a\in A$; $\rho(a,A)=0$ for all $A\in{\mathcal M}$ and all $a\not\in A$; and $\sum_{a\in A} \rho(a,A)\leq 1$, where $\rho(o,A):=1-\sum_{a\in A} \rho(a,A)\leq 1$ is the probability of choosing the---always feasible---outside option at menu $A$. To simplify notation, for $A,B\in{\mathcal M}$ with $B\subseteq A$ we write $\rho(B,A):=\sum_{b\in B}\rho(b,A)$.
We start by introducing the logit with a general outside option as the model that comprises value functions $u:X\rightarrow\mathbb{R}_{++}$ and $D:\mathcal{M}\rightarrow\mathbb{R}_{+}$ such that, for every menu $A\in{\mathcal M}$ and active-choice alternative $a\in A$,
where the pair $(u,D)$ is unique up to a common positive linear transformation. In this model, $u$ captures the menu-independent values of active-choice alternatives and $D(A)$ the menu-dependent value of the outside option. Like the baseline Luce model [see (ref) below], all active-choice alternatives in (ref) are assigned menu-independent values that determine their relative likelihood of being chosen. Unlike the baseline model, where this property also extends to the outside option, here the probability of making an active choice in the first place (equivalently, of avoiding/deferring this decision) is determined by the menu-dependent value of $D$.
Two properties characterize the class of models that can be represented in this way:\\
\noindentA1 (Positivity).\\ For all $A\in{\mathcal M}$ and all $a\in A$: $\rho(a,A)>0$.\\
\noindentA2 (The Active-Choice Luce Axiom).\\ For all $A,B\in{\mathcal M}$ and all $a,b\in A\cap B$:
A1 is standard and allows for a crisp illustration of the main ideas that we put forward in this paper. A2 imposes the familiar kind of IIA-consistency only in the odds of pairs of active-choice alternatives, while allowing odds that involve such an alternative and the outside option to deviate from it. That is, $\frac{\rho(o,A)}{\rho(b,A)}\neq \frac{\rho(o,B)}{\rho(b,B)}$ is allowed by A2.
Indeed, adapting the arguments in luce59 yields an equivalence between A1-A2 and the existence of a function $u:X\rightarrow\mathbb{R}_{++}$ such that, for every $A\in{\mathcal M}$ and $a\in A$,
where
for arbitrary and fixed $\alpha>0$ and $z\in X$. It follows then that for every $A\in{\mathcal M}$ there is a unique $D(A)\geq 0$ that makes (ref) true, with
Finally, it is immediate that $(u,D)$ and $(u',D')$ represent the same $\rho$ if and only if $u=\alpha u'$ and $D=\alpha D'$ for some $\alpha>0$.
We now compare (ref) to the baseline logit with an outside option anderson_etal,hensher-rose-greene15 and without such an option. We recall that a random choice model $\rho$ on $X\cup\{o\}$ admits the former representation if there is a function $u:X\cup\{o\}\rightarrow\mathbb{R}_{++}$ such that, for all $A\in{\mathcal M}$ and $a\in A$,
On the other hand, $\rho$ admits a logit representation without an outside option if there exists some $u:X\rightarrow\mathbb{R}_{++}$ such that
The latter obviously implies $\sum_{a\in A}\rho(a,A)=1$ for all $A\in{\mathcal M}$, so that the opportunity to defer is either infeasible in this model or feasible but never acted upon. Thus, (ref) includes (ref) as a special case when $o\not\in X$. In addition, (ref) extends (ref) but without nesting it unless A1-A2 and $\rho$ operate on the enriched domain $X\cup \{o\}$.
We define the power logit model by the existence of a menu-independent stimulus intensity value function $\widehat{u}:X\rightarrow\mathbb{R}_{++}$ and a parameter $p\geq 1$ such that, for every menu $A$ and alternative $a$ in $A$,
Clearly, this model predicts $\rho(o,A)>0$ at every menu $A$ if and only if $p>1$, and reduces to (ref) at $p=1$.
The agent portrayed in (ref) could be thought of as behaving according to the standard logit with a single valuation criterion but, possibly aware of their decision difficulty, also as if they sampled all alternatives more than once before making a decision. For example, in the quadratic logit case of special interest where $p=2$, the agent might be thought of as sampling the same menu twice. Because the resulting value realizations generally differ across these two rounds of sampling due to the postulated randomness, this individual would be more likely to choose an active-choice alternative if its perceived signal/stimulus intensity from both inspections, captured by the two value realizations of $\widehat{u}$, is relatively high, and as being more likely to avoid/defer choice when this is not true for any such alternative. When deciding which insurance plan to buy, for example, an agent whose behavior is approximated by the quadratic logit may review the top-rated plans from a service comparison website in the morning, receive some value stimuli/signals from each of them, and then go back and repeat this process in the evening. Assuming that the two sampling rounds are independent (admittedly, a demanding assumption), an insurance plan is more likely to be chosen at the end of this two-stage process if its relative stimulus/signal intensity is sufficiently high to make the product stand out despite the agent's hesitation.
The intuition in the more general case where $p\neq 2$ in (ref) is analogous and admits a probabilistic explanation. Specifically, if the analyst a priori restricts $p$ to lie between 1 and 2, then $p-1$ might be interpreted as the (exogenous) probability that the agent will engage in two rounds of sampling, equalling 1 in the limit where the quadratic logit decision process emerges with certainty. Similarly, if $p$ is assumed to lie between 2 and 3, then $p-2$ could be thought of as the probability that the agent will perform three rounds of sampling, conditional on the analyst expecting them to do at least two. More generally, the power parameter $p$ in this model could be viewed as reflecting the agent's propensity to engage in possibly multiple rounds of sampling.
That this model is a logit with a general outside option may not be obvious at first glance but quickly becomes so upon noticing that one can write
At $p=2$ these expressions admit the simpler and more easily interpretable form
where the last step makes use of the notational convention
This clarifies that the quadratic logit $\rho\sim(\widehat{u})^2$ is an additive $(u,D)$ model in the sense that the value of the outside option at every menu depends additively on the value of that option at each of its binary submenus. It also clarifies that the latter value takes a symmetric Cobb-Douglas form with respect to $\widehat{u}$. We will return to additivity later in this section but note here that the quadratic case where $p=2$ is the only one where the $(u,D)$ representation of (ref) has this property.
We proceed by noting the following direct implication of the power-logit model:\\
A3 (Desirability & Complexity)\\ For all $A\in{\mathcal M}$: $\rho(A,A)=1$ $\Longleftrightarrow$ $|A|=1$.\\
To motivate the intuition behind (and the label assigned to) A3 we first recall that, as was clarified early on, our aim here is to model decision difficulty that is rooted in a fully attentive individual's potential inability to make some preference comparisons between otherwise desirable options. If a single such option was feasible to such an individual, therefore, one might expect that person to immediately choose that one option. If on the other hand there are at least two available options and the individual is not forced to make a choice immediately, then the experimental/empirical evidence suggests that there is at least some probability that this person's attempt to find a most preferred option and choose that option will not be fruitful reasonably quickly. To the extent that this is so, a legitimate approach from the analyst's perspective would be to portray that decision maker as deferring choice with positive probability whenever at least one non-trivial comparison is required.
Imagine, for example, a patient like those reported on in knops-goossens-ubbink-legemate-stalpers-bossuyt13 who has been diagnosed with a life-threatening disease. Suppose that their doctor informs them that there is only one available treatment that can cure this disease, and asks whether they would like to sign up for this treatment. One would expect the patient to sign up immediately because there would be no benefit from delaying their only chance for a cure. Now suppose instead that the doctor tells the patient that there are two possible treatments: one with high efficacy but severe side effects, and another with milder side effects but lower cure rates. Even though either one of these treatments would have been chosen immediately if it was the only feasible one (see tversky-shafir92,redelmeier-shafir95,dhar97,CCGT22, for example), here one might expect the patient to delay making such an active choice, perhaps until they think about the conflicting pros and cons and then ultimately determine which treatment would be best for them. Situations of this kind are compatible with and, in fact, motivate our modelling framework in this paper.
In the spirit of these examples, A3 postulates that an active choice is made with certainty only at singleton menus and, as such, formalises the behavioral mechanisms outlined above. Of course, one can easily think of situations where this axiom is descriptively invalid. Yet for analytical purposes it is a useful property because it allows for completely isolating the decision-difficulty channel to deferrals from other potential channels such as undesirability of the available alternatives or limited attention, which have quite distinct behavioral origins.
The following statement readily follows from the preceding analysis.
We will refer to this special class of generalized logit models with a context-dependent outside option as the class of decision-conflict logit models, and to the menu function $D$ that captures the varying appeal of opting out at different menus as the decision cost or decision complexity function. Justifying such a name for the function $D$ given the requirement that it be zero-valued only at singletons may benefit from some additional explanation that supplements the preceding discussion. When the decision environment is such that avoidance/deferral is caused solely by decision difficulty instead of other factors (e.g. none of the active-choice alternatives is good enough, or none is considered due to limited-attention constraints), our decision maker is portrayed as not having any problem deciding between deferring or choosing the only available active-choice option: they do the latter. By contrast, the decision between deferring or choosing from two or more such options is at least somewhat costly because of the effort that is necessary to make the relevant preference comparisons.
We will refer to both a decision-conflict logit $\rho=(u,D)$ and $D$ as monotonic if
If $D(A)>D(B)$ is always true when $A\supset B$, then $D$ and $\rho=(u,D)$ will be called strictly monotonic. In line with our intended interpretation of $D$ as a complexity/cost function, the total number of pairs of distinct alternatives increases as a menu expands, hence so does the expected number of comparisons between alternatives that a fully-attentive individual needs to make. In expectation, therefore, decision difficulty also goes up in absolute terms when more alternatives are added to a menu. Importantly, however, this does not imply that deferring always becomes more likely once a menu is expanded when $D$ is monotonic (we will return to this point soon). But Monotonicity does have a familiar general implication for active-choice alternatives, which in the standard random forced-choice environments was originally stated in block-marschak60:
In particular, monotonic models satisfy what we will refer to as active-choice regularity, whereby the probability of such alternatives cannot increase when more options are added to a menu. Crucially, however, as we discuss and illustrate by example later, this property does not hold for the outside option.
When it comes to using a decision-conflict logit in suitable applications, the analyst must first decide whether to employ a special case where function $D$ is set exogenously or one where it is determined endogenously. In the first case the choice might be dictated by the analyst's a priori assessment of the specific environment in question and could include, for example, defining $D$ as the menu-cardinality function iyengar-lepper00,iyengaretal04 or, if the alternatives have clearly identifiable attributes, some measure of similarity in attribute space spektor-gluth-fontanesi-rieskamp19. The analyst's choice in the second case might instead be dictated by an agnosticism towards what is the most appropriate functional form for $D$, and by resorting instead to a general decision process where $D$ is a function of the feasible options' $u$-values. The power-logit class of models is clearly of this kind.
We proceed with an illustration of how a general $(u,D)$ model, or even the more structured power logit, make predictions that help explain intuitively the three empirically documented choice-deferral phenomena mentioned in the Introduction. We start by noting that findings and arguments from the consumer-psychology literature reported in dhar97,sela-berger-liu09,scheibehenne-greifeneder-todd10, among others, suggest that decision makers are sometimes more likely to avoid/delay choice when the feasible alternatives are perceived to be of similar value. The last authors noted, for example, that as the most attractive feasible options become more similar when new items are added to a menu, it can become more difficult for the decision maker to justify the choice of any particular option, which in turn would increase the likelihood of choice deferral. This is what we earlier referred to as “similarity-driven deferral”. In the same direction, but focusing on response times rather than deferral decisions, bhatia-mullett18 recently reported evidence to suggest that choice between similarly attractive options is significantly correlated with longer response times.
Our next result shows how the power logit predicts such an effect. More specifically, an interesting feature of this model is that its predicted probability of opting out at a menu as a function of the number of active-choice alternatives at that menu is bounded above in a simple way, and that upper bound is attained precisely when all feasible alternatives are of the same value.
In this model, therefore, an agent's decision difficulty at a menu, as revealed by the deferral probability at that menu, is maximized when all feasible active-choice alternatives are equally desirable, and this maximum difficulty is increasing in proportion to the total number of such alternatives at a decreasing rate (Figure (ref)).\footnote{A clarifying remark may be due at this point. Equal values (specifically, utilities) between two or more alternatives are---in most of economic theory---associated with positive indifference, which in turn is interpreted as suggesting that the individual in question would be equally happy with any of these alternatives. By contrast, the influential drift diffusion model in neuroeconomics krajbich-armel-rangel10,baldassi&etal, fudenberg&newey&strack&strzalecki, which originates in the psychology literature ratcliff&mckoon08, and related experimental evidence that have been of increasing visibility and interest in the economics literature lately make the opposite predictions/observations. The latter in turn are broadly in line with the general predictions of the power-logit model that we focus on in this paper. Considering the different motivations, methodological frameworks and intended interpretations in the two literatures, however, the seeming discrepancy is in our view more an issue of semantics than it is one of substance. In any case, our use of the term “value” rather than “utility” in reference to the terms appearing in logit formulae is partly motivated by this issue.}
We now turn to the power-logit model's comparative statics in the important class of binary menus. Figure (ref) illustrates, with a quadratic-logit example, the general pattern in the behavior of $\rho(a,\{a,b\})$ and $\rho(o,\{a,b\})$ as the stimulus intensity of $a$ changes while that of $b$ is held fixed. Interestingly, the monotonic increase of $\rho(a,\{a,b\})$ in $\widehat{u}(a)$ occurs at an increasing rate as this value approaches the $\frac{\widehat{u}(b)}{2}$ stimulus-intensity threshold from below than when $\widehat{u}(a)$ increases monotonically beyond $\frac{\widehat{u}(b)}{2}$. Intuitively, the inflection-point stimulus intensity value $\frac{\widehat{u}(b)}{2}$ that dissects $\rho(a,\{a,b\})$---viewed as a function of $\widehat{u}(a)$---into convex and concave regions suggests that marginal improvements in the appeal of $a$ lead to more rapid market share increases when this alternative is still “catching up” with $b$ than when it has become sufficiently close to (or surpassed) it in attractiveness. On the other hand, $\rho(o,\{a,b\})$ is a strictly concave function of $\widehat{u}(a)$ and, consistent with Proposition (ref), attains its maximum value of $\frac{1}{2}$ when $\widehat{u}(a)=\widehat{u}(b)$. Thus, the model's novel prediction here, of relevance both from a consumer-welfare and a seller-profit perspective, is that minimal re-designing of a menu that consists of equally attractive alternatives is more likely to be effective at reducing opt-out behavior if one of the original alternatives becomes less rather than more appealing, other things equal.
Now, since, as was discussed previously, it is not generally true that avoiding/deferring becomes more likely as menus expand even for monotonic decision-conflict logit models, it is naturally of interest to understand when, exactly, such behavior is to be expected in this environment. The general idea in answering this question is that, even if decision difficulty increases in absolute terms when new alternatives are introduced, when these new alternatives are sufficiently better than the pre-existing ones their added value will offset the elevated decision cost and will ultimately result in a higher probability of making an active choice at the larger menu. To state this more formally we will abuse notation slightly by letting
stand for the total Luce value at menu $S\in{\mathcal M}$.
This eloquent equivalence clarifies that the choice probability of opting out will decrease following menu expansion if and only if the marginal benefit of this expansion, as measured by the percentage increase in total value, exceeds its marginal cost, as measured by the percentage increase in decision complexity. This is a distinctive property of decision-conflict logit models. It clarifies that they do not belong to the random-utility class\footnote{See apesteguia-ballester-lu, stoye19, strzalecki24 and references therein.} with an outside option, and enables them to explain simply the non-monotonic and dominance-driven effect that menu expansion has been known to exert on the probability of deferring scheibehenne-greifeneder-todd10,chernev-bockenholt-goodman15, which we earlier referred to as the “roller-coaster choice overload” effect.
Indeed, citing several studies in consumer psychology, the meta-analysis in chernev-bockenholt-goodman15 notes that “it has been shown that consumers are more likely to make a purchase from an assortment when it contains a dominant option than when such an option is absent” (p. 338). This finding is important for the interpretation and policy responses to choice-overload phenomena of the kind that were first reported in iyengar-lepper00. To our knowledge, the decision-conflict logit is the first random-choice model that predicts this dominance-driven emergence and disappearance of choice-overload effects, and it does so without imposing any undesirability or inattention constraints. Table (ref) illustrates an example such effect that is predicted by the quadratic logit model.
Finally, as the next result establishes, the power-logit also predicts another important choice-deferral phenomenon, known as the “relative-desirability” effect dhar97,white-hoffrage-reisen15,bhatia-mullett16. This refers to situations where, choosing the outside option becomes more likely in binary menus as the available options become more equally desirable, other things equal.
This result, illustrated in Figure (ref), is distinct from the similarity-driven deferral effect that was discussed in relation to Proposition (ref) because it compares the probabilities of opting out at two distinct binary menus as a function of the absolute value/stimulus-intensity differences between the two active-choice alternatives, rather than focusing on when this probability is maximized within the same menu of any size. It clarifies, indeed, that the model predicts relative-desirability effects irrespective of whether the stimulus-intensity absolute difference of the two alternatives or that between their power-logit values---which emerge from the stimulus-intensity values via the (convex) power transformation---is used to assess relative desirability.
We proceed with an illustration of the potential usefulness of the power-logit functional form in the analysis of oligopolistic markets when consumers potentially face comparison difficulties and may avoid/delay making an active choice.\footnote{piccione-spiegler12, spiegler15, bachi&spiegler and gerasimou&papi have recently suggested distinct approaches to study such markets.} To this end, we consider a market where two profit-maximizing firms compete for a single consumer (equivalently, a unit mass of consumers) by offering a product that is differentiated in quality, $q_i$, and price, $p_i$. Producing a product of quality $q_i$ costs $q_i$ to firm $i=1,2$, while $0\leq q_i\leq p_i \leq I$ and $I>0$ denotes consumer income. Furthermore, a consumer's value from product $(q_i,p_i)$ coincides with that product's quality-price ratio:
This assumption further implies
for all $(q_i,p_i)$. Such a “value-for-money” specification imposes intuitive positive and negative dependences of $u$ on quality and price, respectively, with the former being linear and the latter strictly convex. Moreover, while identifying value with quality-price ratios as in (ref) rather than with quality-price differences $q_i-p_i$ appears to be a novel modelling assumption, it is consistent with some central implications of the behavioral choice model by bordalloetal13 concerning consumer preferences for high quality-price ratio products, even though that model starts from very different primitives and features a quality-price difference value function instead.
The two firms choose their products' quality and price levels simultaneously and under complete information. The market share of product $(q_i,p_i)$ at menu/strategy profile $\big((q_1,p_1),(q_2,p_2)\bigr)$ is determined by the power logit model
where $s\geq 1$ and $s=1$ in the baseline special case where there is no decision difficulty. Under the above assumptions, each firm $i=1,2$ solves
The strategic trade-off in this model, which applies both when $s=1$ and $s>1$, is that each firm wishes to increase its quality/price ratio in order to expand its market share, while at the same time also wishing to decrease it in order to enlarge its profit markup.
Turning to consumer welfare, taking into account that decision conflict can potentially drive the consumer out of the market altogether, and that -by A3- this would be undesirable, we consider a utilitarian-like welfare measure that weighs the possible value levels at a given strategy profile by the probabilities that these values will actually be realized at that profile. We formalize this with the consumer welfare function $W:\mathbb{R}^4_{++}\rightarrow [0,1]$ defined by
This welfare indicator may be particularly relevant in cases where consumer surplus is equilibrium-invariant, as will turn out to be the case in the present environment.\footnote{A related measure that identifies welfare with the proportion of consumers who make an active choice was studied in spiegler15, while gerasimou&papi introduced an index that is similar to $W$ but features instead the probability-weighted product variety that is associated with a strategy profile.}
Perhaps surprisingly, this duopolistic model leads to the following simple and intuitive equilibrium predictions:
Thus, although the equilibrium pricing strategy features full surplus extraction irrespective of the value of the hesitation/resampling parameter $s$, the equilibrium quality level increases in $s$ at the rate $\frac{s}{2+s}$. starting at the low of $\frac{I}{2}$ in the baseline case of logit market shares and no consumer hesitation ($s=1$), and approaching $I$ as $s$ becomes large. An intuitive interpretation of this fact is that decision conflict inevitably introduces a third “competitor” into the market, the outside option, that becomes more “powerful” as $s$ grows. The power logit predicts that the choice probability of the outside option goes down as the value of one of the two products is unilaterally increased, while the choice probability of the comparatively more appealing product simultaneously goes up during the process. This in turn creates incentives for each firm to unilaterally increase its quality level relative to the baseline logit case. But since increasing quality is costly, the above-mentioned strategic trade-off that is embedded in each firm's profit function eventually kicks in and halts this increase at the above symmetric-equilibrium level.
Notably, while consumer surplus is zero in equilibrium because each firm's profits turn out to be strictly increasing in its product's price, consumer welfare changes in an interesting way as $s$ varies. In particular, despite the increase in the attainable value level in equilibrium once firms best-respond to consumers' hesitation and resampling, welfare decreases in $s$. This decrease is caused by the fact that in the power logit with two equally attractive products the consumer is more/equally/less likely to defer than to make an active choice when $s>2/s=2/s<2$ and, conditional on doing the latter, equally likely to choose either of the two available products (Proposition (ref)). The implication of this in the present environment is that the higher value that the consumer receives in expectation under the equilibrium with some decision conflict ($s>1$) is not sufficiently high to offset the lower value that they receive with certainty under the equilibrium with no conflict ($s=1$). The firms' profits, finally, also decrease when consumers are hesitant relative to the case where there they are not. This large decrease is intuitive and contributed by the reduced probability of the consumer choosing either product, as well as by the reduction in the firms' profit margins that is brought about by the improvement in quality. Figure (ref) illustrates these facts graphically when $I$ is normalized to 1.
It is often the case in empirical applications that the choice frequencies available to the analyst are obtained from the choices made by a cross section of individuals who are presented with the same menu, rather than from a single decision maker's repeated choices at that menu. Random-utility based discrete choice estimation in those cases is often carried out under the assumption that the observable component of every individual's utility coincides, and that the error term in that model's formulation captures all individual heterogeneity that is unobserved to the analyst. Adopting and adapting this assumption to our non-random-utility environment, in this section (and in Appendix A) we show how the other assumptions and formal argument that underpin the discrete-choice formulation of the logit model without an outside option that was pioneered by mcfadden73 can be modified to arrive at a similar discrete-choice version of the quadratic- and power-logit models. It is worth remarking that, as we show in Section 5, the use of otherwise standard discrete-choice datasets is sufficient towards estimating these models, as long as they are obtained from a “free choice” decision environment, i.e. one where individuals could choose the no-choice outside option, where the analyst observes both the active choices and those of the latter option.
We start by denoting the set of all quadratic-logit decision makers by $\{1,\ldots,n,\ldots,N\}$. Keeping the menu $A:=\{a_1,\ldots,a_i,\ldots,a_k\}\subseteq X$ fixed throughout this and the next subsection, we proceed by recalling and breaking down the baseline assumptions of the discrete-choice formulation of the baseline logit in (ref) as follows:\\ 1. Random utility [structural assumption]: there is some function $u_n:X\rightarrow\mathbb{R}$ such that
where $v_n(a_i)$ and $\epsilon_{ni}$ are, respectively, the deterministic (observable to the analyst) and stochastic (unobservable) contributions to agent $n$'s utility from good $a_i$. Denoting by $x_{ni}$ and $\beta$, respectively, two $m$-vectors of observable product/consumer characteristics and estimable coefficients that capture their relative importance via the relationship specified by some function $g:\mathbb{R}^m\times \mathbb{R}^m\rightarrow\mathbb{R}$, in applications it is assumed that
and, often, that the dependence of $v_n$ on $\beta$ and $x_{ni}$ via $g$ is linear-additive:
where $\cdot$ denotes the inner product.\\ 2. Random utility maximization [behavioral assumption]: for all $a_i\in A$,
3. Gumbel noise [distributional assumption]: the error term $\epsilon_{ni}$ is independently and identically distributed across $i$ according to the standard Gumbel density
As has been widely known since the seminal contribution of mcfadden73,\footnote{luce&suppes and, indeed, mcfadden73 also credit Eric W. Holman and Anthony A. J. Marley with this discovery.} these assumptions jointly imply the analytically convenient and famous form
We proceed by examining how the premises and conclusion of this classic discrete-choice logit model are affected and can be modified when we assume that decision maker $n$ uses the single but noisy value criterion captured by $u_n$ to sample the values of the alternatives in $A$ twice, as per the quadratic special case of the power logit (focusing on the quadratic case here is done for simplicity of the exposition; we deal with the general case later). To this end, and recalling the interpretation that was put forward at the beginning of Section 2.3, we first note that maintaining the additivity and linearity assumption implies that at the end of the second round of sampling the individual has perceived two values for each alternative $a_i\in A$,
These generally distinct values across the two rounds will vary according to the distribution of $\epsilon_{ni}$. Such multiplicity of value realizations in turn implies that each alternative $a_i\in A$ is ultimately associated with a vector of values $\big(u^1_n(a_i),u^2_n(a_i)\bigr)$. With utility now being vector-valued, however, the utility-maximization behavioral assumption that underpins (ref) is no longer applicable in an obvious way. To break this impasse we assume that the random utility maximization behavioral assumption is replaced by a dominance assumption whereby
Turning, finally, to the modification of the distributional assumption (ref), to make it operational in the quadratic-logit framework we assume that the random errors $\epsilon_{ni}^1$ and $\epsilon_{ni}^2$ are independent across all alternatives $i\leq k$ and across the two sampling rounds $l\leq 2$. As was also anticipated in the discussion of Section 3.1, this is indeed a demanding simplifying assumption that we hope future studies will be able to relax.
With these assumptions in place we can now write
where each integral is $k$-dimensional, the first step makes use of the above behavioral, distributional and independence assumptions on $\epsilon_{ni}^l$, while the last step follows from the derivation of the discrete-choice logit [see, for example, train09].
An important difference between the discrete-choice version of the logit with an outside option in (ref) and its quadratic-logit counterpart is that in the former case the modeller specifies the value of that option exogenously (see anderson_etal,hensher-rose-greene15), whereas in the latter case this value emerges endogenously as a function of the observable characteristics of all active-choice alternatives. Indeed, assuming now---and in the remainder of this section and the next---both (ref) and (ref), upon rewriting (ref) as
one observes that
By contrast, in the baseline model we have
where $x_{no}$ is set by the analyst.
In Appendix A we consider the general case of the power logit model, formulate its likelihood function, and identify the first-order conditions on $p$ and $\beta$ for its maximization.
For our application we use the survey-experiment data with film choices that were collected by bhatia-mullett16. In that study, 58 subjects were initially asked to rate from 1 (least desirable) to 9 (most desirable)\footnote{A typo in bhatia-mullett16 erroneously suggests that the highest rating was 7 instead.} the 100 most voted-on (hence most popular) films on the IMDB online platform (https://www.imdb.com) at the time. Following that, subjects were presented with 100 distinct binary menus with films that were drawn from that list, with the respective images presented side by side. In the free-choice treatment, subjects were asked to choose either the film positioned on the left or on the right of each menu, or to defer the decision (these choices were entered by clicking on the left, right and up keys, respectively).
Regarding the instructions that subjects received, the authors highlighted (p. 136) that these “were created to avoid any suggestion of an explicit time limit (e.g. to suggest that participants should defer if they cannot decide quickly enough) or that deferral was a third comparable option (e.g. in the form of a status quo or default movie). More specifically, the instructions stated that if participants preferred the movie on the left/right then they should press the left/right arrow. If they could not make a decision about which of the two movies they preferred then they should press the up arrow instead.” In the forced-choice treatment, the same 100 menus were presented but deferral was not feasible.
The study featured a within-subject design and subjects were randomly assigned to start the experiment in either of the two treatments. There was no limit in the time subjects had available to make their $2\times100$ decisions.
Although bhatia-mullett16 focused mainly on the relationship between choice deferral and response times, they also reported on the relationship between ratings and active-choice probabilities conditional on an active choice being made. Specifically, they found that the film with a higher rating, where relevant, is chosen 83% of the time (p. 137). Enabled in this way by the theoretical analysis of the previous sections, our focus here instead is on the unconditional analysis of the explanatory value of the subjects' own ratings on their subsequent active-choice and deferral decisions, and on comparing the results from this analysis when it builds either on the baseline logit with an outside option or on the hereby proposed power and quadratic logit.\footnote{We recall that, as was clarified in Sections 2 and 3, the power logit and the baseline logit with an outside option are non-nested models.} In particular, on each of the 100 binary menus in this dataset (for which, we recall, 58 observations are available) we estimate and compare the goodness of fit of the models that we lay out below.
In line with existing practices (see, for example, pp. 411-414 in hensher-rose-greene15), to estimate this model we treat the outside option as an explicit alternative with a fixed value that is common to all subjects.{\footnote{Under these two conditions the exact value of the outside option's “rating” is unimportant for this model's maximized log-likelihood and estimate of $\beta_1^A$, mattering only for the estimates of $\beta_0^{l,A}$ and $\beta_0^{r,A}$.}} Doing so leads to the following three-parameter multinomial logit specification:
The left-hand-side terms denote the estimated probabilities of subject $n$ choosing “left”, “right” or “defer” at binary menu $A$. On the right hand side, $\beta_1^A$ and $\beta_0^{l,A}$, $\beta_0^{r,A}$ are, respectively, the estimated slope and intercept coefficients at menu $A$. The former captures the effect that a unitary increase in subject $n$'s rating of the left (right) film--denoted here by $rat.Left$ ($rat.Right$)--has on the log-odds of choosing that film over deferring when the latter option's value is fixed. The option-specific intercepts $\beta_0^{l,A}$ and $\beta_0^{r,A}$ on the other hand capture the log-odds of choosing, respectively, the left and right film over deferring when the relevant film's rating is zero. Hence, including these terms in the estimation is essential for otherwise the prediction would be equal choice probabilities for “left”, “right” and “defer” if both films had a zero rating. This, in turn, would go against the model's treatment of the outside option as any other alternative that is more likely to be chosen as the other feasible options become worse.
We also consider the variant of the preceding model where, instead of assuming a fixed common value (“rating”) for the outside option, we allow it to vary across subjects and menus by randomizing uniformly over the permissible rating values.\\
As discussed in the previous subsection, estimating the quadratic logit amounts to estimating the parameter $\gamma^A$ in
There are some important differences between this model and the multinomial logit with an outside option laid out above. First, unlike that model, the quadratic logit does not include any intercept terms. This is in line with the theoretical predictions of the general version of this model (Proposition (ref)), according to which all active-choice options are equally likely to be chosen when they have the same value. Including alternative-specific intercept terms here would go against this prediction as it would lead to generally distinct predicted probabilities for the left and right film when their ratings are identically equal to zero. Second, unlike $\beta^A$, the slope coefficient $\gamma^{A}$ here captures the log-odds of choosing one film over the other (in particular, not of choosing one film over deferring) following a unitary change in the former film's rating. More specifically, given (ref), (ref), (ref) and (ref), a more appropriate interpretation of this coefficient is that it captures the relevant change in the log-odds of choosing one film over the other following a unitary increase in the former's rating conditional on an active choice having been made, while the unconditional change in these log-odds is obtained by multiplying them by $1-\rho(o,A)$. By contrast, (ref) clarifies that the log-odds of choosing a film over deferring following a unitary increase in that option's rating is captured by $2\gamma^A$ instead.
Estimating this more general model now involves finding simultaneously optimal values for the slope coefficient $\theta^A$ and the power parameter $p_A$ in
The parameter $\theta^A$ here admits an analogous interpretation to $\gamma^A$ in the quadratic logit, while the term $p_A\theta^{A}$ is interpretable as the effect that a unitary change in a film's rating has on the log-odds of choosing that film over deferring.
We perform a goodness-of-fit analysis and comparison of the four models that aim to assess their explanatory and predictive performance separately on each of the 100 menus.\footnote{The results presented in this subsection were obtained with code written in the R programming language (baseR, v4.5.2) with RStudio Rstudio, and utilising the “mlogit” mlogit, “optimx” optimx, “plyr” wickham11 and “tidyverse” tidyverse packages/libraries.} To this end, we focus on the maximized log-likelihood value, the Akaike (AIC) and Bayesian (BIC) information criteria, and each model's proportion of correct predictions. In particular, denoting by $\widehat{L}_A$, $k$ and $N_A$, respectively, a model's maximized log-likelihood value at menu $A$, the number of its parameters and its sample size, recall that $AIC = 2k-2\log(\widehat{L}_A)$ and $BIC = k\log(N_A)-2\log(\widehat{L}_A)$. The value of $k$ is 3 for the two multinomial logit models with a fixed and random outside option, 2 for the power logit and 1 for the quadratic logit. The sample size is $N_A=58$ in all four models and for each one of the 100 menus. In the prediction analysis we used the models' 100 menu-specific predicted choices per subject ($5800=100 \times 58$ in total) to subjects' actual choices at each menu. A model was taken to make a correct prediction for a given subject at a given menu if it predicted a weakly highest choice probability for the option that was actually chosen by that subject in that menu. The same principle was applied in a comprehensive cross-validation analysis where data from all but one of the 100 menus were pooled together and the resulting model estimates were used to make out-of-sample predictions at the excluded menu.
Figure (ref) plots the 100 pairs of power- and slope-parameter estimates that emerge from the power-logit model. The mean, median and standard deviation of the $p$ estimates in those regressions are 1.51, 1.47 and 0.27, respectively. The slope-parameter estimates on the other hand have a mean, median and standard deviation of 0.43, 0.40 and 0.15, suggesting that the effect of a one-unit increase in a film's rating is an approximately 53% increase in the odds of choosing that film over the alternative. For comparison, the mean/median and standard deviation in the slope estimates corresponding to the baseline logit with a fixed outside option are 0.58 and 0.15, respectively, pointing to an approximately 78% increase in the above-mentioned odds.
Interestingly, there is a negative correlation (Spearman $\rho=-0.33$) between the $\widehat{p}$ and $\widehat{\theta}$ estimates in these data. The fact that $\widehat{p}$ tends to be lower at menus where $\widehat{\theta}$ is higher, however, can be interpreted intuitively through the lens of this model. Specifically, when $\widehat{p}$ is high, the deferral frequency also tends to be high. When deferrals are primarily caused by the relative undesirability of the two films, as per the logit with an outside option, a higher value of the slope parameter would be expected, in line with the $\overline{\beta}_1>\overline{\theta}$ finding. This is so because, in this model, the marginal effect of a unitary change in a film's rating is more likely to be high when both films have a low rating. But when deferrals are not primarily due to undesirability but, instead, are mainly caused by decision difficulty, then relatively low values of $\widehat{\theta}$ could be observed not because of low but because of similar ratings, and by the harder comparison that such similarity entails.
The potential presence of such a channel is further supported by the negative correlation (Spearman $\rho=-0.17$) between the $\widehat{p}$ estimates and average (across subjects) absolute differences in ratings at the respective menus. The mean, median and standard deviation of this variable at the 100 menus are 2.25, 2.21 and 0.45, respectively. The bottom-right quarter of the scatter plot in Figure (ref) reveals the presence of 31 menus with an estimated $p$ in excess of its median value of 1.47 and an average absolute difference in ratings between the two films at each of these menus below its median of 2.25. The mean and median estimates of the power-logit slope parameter $\widehat{\theta}$ at these 31 menus are 0.39, while the corresponding statistics in the remaining 69 menus are 0.46 and 0.42. The difference in the distribution of $\widehat{\theta}$ between these two groups is statistically significant ($p=0.044$; two-sided Mann-Whitney test) and corroborates this intuition and theoretical prediction.
We now turn to the results of the goodness-of-fit comparison that is summarized in Table (ref). The logit with an inferior outside option performs better than the other three models in most menus under each of the log-likelihood (85), AIC (75) and BIC (60) criteria. Together, the power and quadratic logit provide the best fits under AIC and BIC in 18 and 34 menus, respectively, followed by the logit with a random outside option (7 and 6 menus). Under the proportion-of-correct-predictions criterion, on the other hand, the power logit performs better (41.6%), followed by the baseline logit with a fixed or random outside option (35.6% and 36.4%, respectively) and by the quadratic logit (31.1%).\footnote{The 5- and 4-percentage point differences in the rates of correct predictions between the power logit and the baseline logit with either an inferior or a random outside option are significant ($p<0.001$ in both cases; $p$-values from 2-sided Fisher's exact tests).} Our cross-validation analysis (Table (ref)) focuses on the two models that the preceding analysis suggests are the descriptively leading ones. It reveals that the baseline logit---but not the power logit---makes correct out-of-sample predictions regarding the option that is most likely to be chosen in nearly two thirds of all cases. Conversely, the power logit---but not the baseline logit---makes correct out-of-sample predictions in nearly one fifth of all cases. Finally, both models make correct out-of-sample predictions 15% of the time.
Further light on the relevance of the behavioral channel that was discussed earlier can now be shed by comparing the models' fit in those menus where the average film ratings are high and low. This is relevant because the mechanism underpinning the logit with a fixed outside option suggests that choosing that option is more likely when the average rating is low. Intuitively, therefore, we would expect this model to provide a better fit in the latter group of menus than the power logit does. To this end, we compare the two models' AIC and BIC scores in the two groups of 50 menus with above- and below-median average total rating (the median value of this statistic is 11.44). In line with this intuition, the baseline logit performs better than the power logit in a higher proportion of the 50 menus with a low average rating than those with a high such rating, under both criteria (AIC: 92% vs 74%; BIC: 84% vs 60%), with the difference in proportions being significant in both cases.\footnote{ The respective $p$-values from two-sided Fisher's exact tests are $p=0.03$ and $p=0.01$.}
These results suggest that the hereby proposed class of power-logit discrete-choice models with an endogenously determined menu-dependent value of the outside option can indeed provide meaningful explanatory gains relative to the baseline logit model with a fixed or random outside option. Moreover, these explanatory gains often occur in those decision environments where intuition and the theoretical analysis of previous sections would suggest that the proposed model should indeed perform better. We hope that this illustration will be helpful to the experimenter or empirical researcher who is interested in creating and analyzing similar free-choice datasets.
As was illustrated in the proof-of-concept empirical application of the previous section, and as was also explained in Section 2, standard discrete choice models with an outside option that are based on random-utility maximization treat this option just like any other alternative and predict that it is more likely to be chosen when its utility is higher than that of all feasible active-choice alternatives. anderson_etal and hensher-rose-greene15, for example, are textbook references that discuss this approach in detail. The models that we study in this paper differ radically from this (un-)desirability approach to modelling choice of the outside option, as they predict that every active-choice alternative is always chosen when it is the only feasible one (A3, Section 2). In addition, the more structured power-logit special case of this model predicts that the probability of opting out at larger menus increases as the feasible such alternatives become more equally appealing, in line with relevant empirical evidence.
Starting with man&mar14, several random choice models of limited attention---also logically distinct from the modelling framework that this paper proposes---have included an outside option as a model-closing assumption that requires this option to be chosen when no attention is paid to any of the feasible active-choice alternatives brady-rehbeck16,aguiar17,aguiar-boccardi-kashaev-kim23. Because of this assumption, deferring/opting out becomes less likely in these models as menus become bigger. horan19 recently clarified how the deferral option can be removed from these models without affecting their general features and primary purpose, which is to explain active-choice decision making subject to cognitive/attention constraints.
We proceed with a more detailed comparison of the logit with a general outside option that is formulated in (ref) and the intuitive generalization of the classic nested logit model ben-akiva73,mcfadden78 that was recently proposed and analysed in kovach-tserejigmid19. The latter assumes that the set of alternatives $X$ can be partitioned into nests $X_1,\ldots,X_K$, and that there exist a non-negative function $v^*$ on the collection $\bigcup_{i=1}^K2^{X_i}$ and a strictly positive function $u^*$ on $X$ such that, for all $A\in{\mathcal M}$ and $a\in A\cap X_i$,
That paper did not consider an outside option as it focused on explaining the kinds of canonical violations of A2 that motivated the original development of nested logit as a generalization of baseline logit. Yet the version of (ref) that is closest to (ref) emerges when: (i) $X$ is expanded to $X\cup\{o\}$ and partitioned into the nests\footnote{See, for example, the top branch of the auto-mobile choice model in Figure 1 of goldberg95.} $\{X_1=\{o\}, X_2=X\}$; (ii) the collection of menus is ${\mathcal M}^*:=\{A\cup\{o\}: \emptyset\neq A\subseteq X\}$. In this case the choice probability of active-choice alternative $a$ at decision problem $A^*\in{\mathcal M}^*$ reduces to
Using the notational convention $A^*\equiv A\cup\{o\}$ for $A\in{\mathcal M}$, a little algebra shows that (ref) and (ref) become equivalent if and only if $u(a)\equiv u^*(a)$ for all $a\in X$ and $D(A)\equiv \dfrac{v^*(\{o\})u^*(o)}{v(A)}$ for all $A\in{\mathcal M}$. Thus, unless $v^*(\{o\})=0$ or $u^*(o)=0$, equivalence between the logit with a general outside option and the generalized nested logit with a fixed outside option is possible only if $D(A)>0$ for every $A\in{\mathcal M}$. The models that are representable as in (ref) which do not impose this restriction at singleton menus. Moreover, (ref) has two additional degrees of freedom compared to (ref): one because $u^*$ takes $|X\cup \{o\}|=|X|+1$ values; and another because $v^*$ takes $|\mathcal{M^*}\cup\{o\}|=|\mathcal{M}|+1$ values. Therefore, even when $D(A)>0$ for all $A\in{\mathcal M}$, (ref) is not uniquely recoverable from (ref). Thus, despite the structural similarity between (ref) and (ref)---which is perhaps best seen by contrasting the two multiplicative terms in generalized nested logit with the corresponding ones in (ref)---the two models differ in some essential ways.
We note, finally, that also related to (ref) but logically and interpretively distinct from it is the focal logit model of kovach-tserenjigmid21 when a fixed outside option is introduced into the latter. That model's components comprise: (i) a menu-independent value function over alternatives $u^{**}$ on $X\cup\{o\}$; (ii) a menu-dependent focus function $F$ that assigns a consideration set $F(A^*)$ to every problem $A^*\in{\mathcal M}^*$; (iii) a menu-dependent focality bias function $\delta$ that gives a `value boost' to alternatives in $F(A^*)$. Formally, the choice probability of active-choice alternative $a\in A^*$ in this model is given by
where $\mathds{1}\{\cdot\}$ is the indicator function. Although (ref) and (ref) are distinct, they intersect in the special case where $u(a)\equiv u^{**}(a)$ for all $a\in X$; $u^{**}(o)\equiv 1$; $F(A^*)\equiv \{o\}$ for all $A^*\in{\mathcal M}^*$; and hence $\delta(A^*)\equiv D(A)-1$ for all $A^*\in{\mathcal M}$ (recall that $A\equiv A^*\setminus\{o\}$).\footnote{The author is grateful to Levent \"{U}lk\"{u} for alerting him to this connection between (ref) and (ref).} The last restriction implies $D(\{a\})>0$ for all $a\in X$, hence $\rho(a,\{a\})<1$. The special cases of (ref) that we studied here do not impose this restriction.
Understanding the “easy” and “hard” parts of people's preference comparisons, as these are revealed by their active-choice or choice-avoidance/delay decisions, is important methodologically and also for practical applications such as effective choice architecture. The present paper contributes in this respect by introducing the tractable power logit model and its quadratic-logit special cases. These models---both of which belong to a general class of logit models with a context-dependent outside option---assume that people can avoid/delay making an active choice and are more likely to opt out in free-choice problems when it is harder for them to identify a best alternative from those available to them. This prediction is supported empirically and differs from the predictions of existing models where the outside option is chosen due to the undesirability of all feasible alternatives, limited attention, or other sources of bounded-rational behavior. In conjunction with the insights from the relevant decision-making literature, our analysis suggests that decision-conflict logit models can help theoretical and applied empirical economists think formally and perhaps more realistically about non-strategic as well as strategic situations where decision makers: (i) are presented sufficiently small menus, so that limited-attention considerations are not pertinent; (ii) consider all feasible active-choice alternatives to be desirable/good enough, so that any one of them would be expected to be chosen if it were the only feasible item; (iii) find it difficult to compare these alternatives due to their complexity or due to potentially non-trivial trade-offs these generate; and (iv) are not forced to make an active choice.
Proof of Proposition (ref). In the main text.\\
Proof of Proposition (ref).
For the second claim, suppose $\widehat{u}(a)=\widehat{u}(b):=c$ for all $a,b\in A$. By (ref), $\rho(o,A) = 1- \sum\limits_{a\in A}\left( \frac{\widehat{u}(a)}{\sum\limits_{b\in A}\widehat{u}(b)} \right)^p = 1-|A|\left(\frac{c}{|A|c}\right)^p = 1-|A|^{1-p}$. Thus, $\rho(o,A)$ is independent of the specific $\widehat{u}$ values at $A$ whenever these values coincide. This readily implies that, viewed as the function
$\rho(o,A)$ has any $|A|$-vector of $\widehat{u}$ values $(c,\ldots,c)$ as a critical point that trivially satisfies both the first- and second-order conditions of local optimality. Yet, because the determinant of the Hessian matrix at any such point is zero, it is not immediately clear if this point is a local maximizer. To show that this is indeed so, by symmetry it suffices to consider marginal deviations in a single direction; say, an $\epsilon$ increase or decrease in $\widehat{u}(a_1)$. Since $\widehat{u}(a_i)=\widehat{u}(a_j)\equiv c>0$, by assumption, this and (ref) yield
Suppose to the contrary that this weakly exceeds $1-|A|^{1-p}$. Without loss of generality, write $\epsilon:=mc$ for some small $m>0$ or $m<0$. We have
To ease notation, write $n:=|A|$. Rearranging, observe that the above is true if and only if $n^{1-p}(n+m)^pc^p \geq (n-1)c^p+(1+m)^pc^p$, which in turn is true if and only if $n^{1-p}(n+m)^p \geq n-1 +(1+m)^p$. Rearranging further, we get $n\left(\dfrac{n+m}{n}\right)^p \geq n+1 + (m+1)^p$ from which we finally obtain $\left(1+\dfrac{m}{n}\right)^p - \dfrac{n+1+(m+1)^p}{n} \geq 0$ Taking the limit as $m\rightarrow 0$ and rearranging leads to $n\geq n+2$, which is impossible. We have therefore established that the above critical point is indeed a local maximizer of $\rho(o,A)$.
We proceed toward showing that it is in fact a global maximizer, thereby concluding the proof. To this end, notice first that $\rho(o,A)<1$, by strict positivity of $\widehat{u}$. Suppose to the contrary that there is a non-constant $|A|$-vector $(\widehat{u}(a_1),\ldots,\widehat{u}(a_{|A|})$ that satisfies the first-order conditions of optimality that are derived from (ref). Differentiating and rearranging pins down these conditions to
Solving this system leads to $\widehat{u}(a_1^*)=\widehat{u}(a_2^*)=\ldots=\widehat{u}(a^*_{|A|})$, contradicting the supposed non-constancy of the postulated alternative local maximizer. It follows that $\rho(o,A)$ is maximized at any constant $|A|$-vector only. From this and the second claim that was established earlier it now follows that this maximum is indeed given by $1-|A|^{1-p}$, as per the first claim. $\blacksquare$\\
Proof of Proposition (ref).
To dispense with the absolute value sign, assume without loss of generality that $\widehat{u}(a)>\widehat{u}(b)$ and $\widehat{u}(c)>\widehat{u}(d)$. We will first show that (ref) holds under either of the postulated conditions. Following that, we will show that (ref) $\Leftrightarrow$ (ref), also under either condition.
Starting with (ref), consider first the case where $\widehat{u}(a)+\widehat{u}(b)=\widehat{u}(c)+\widehat{u}(d)$. Denote this common sum by $s$. We have $\rho(o,\{a,b\})>\rho(o,\{c,d\}) \Leftrightarrow \frac{\widehat{u}(c)^p+\widehat{u}(d)^p}{s^p} >\frac{\widehat{u}(a)^p+\widehat{u}(b)^p}{s^p}$. This is equivalent to
Suppose to the contrary that
This and the postulated equality yield $\widehat{u}(a)\geq \widehat{u}(c)$. Furthermore, this and (ref) jointly imply $\widehat{u}(a)>\widehat{u}(c)$ and $\widehat{u}(d)>\widehat{u}(b)$. Thus,
In view of (ref), observe that the terms $\frac{\widehat{u}(a)-\widehat{u}(c)}{\widehat{u}(a)-\widehat{u}(b)}$ and $\frac{\widehat{u}(c)-\widehat{u}(b)}{\widehat{u}(a)-\widehat{u}(b)}$ are convex weights. Hence, since $\widehat{u}(\cdot)\mapsto\widehat{u}(\cdot)^p$ is a strictly convex function, we have
Adding (ref) to (ref) and recalling that $\widehat{u}(a)+\widehat{u}(b)=\widehat{u}(c)+\widehat{u}(d)=s$ yields
which contradicts (ref). Thus,
holds. Conversely, suppose (ref) is true. This and the postulated equality together imply
Applying the preceding convexity argument using (ref) yields (ref), thereby completing the proof that (ref) holds under the first postulate.
We now show that (ref) is true when $u(a)+u(b)=u(c)+u(d)$ or, equivalently,
holds instead. Let $t$ denote this common sum. We have $\rho(o,\{a,b\})>\rho(o,\{c,d\}) \Leftrightarrow \frac{\widehat{u}(c)^p+\widehat{u}(d)^p}{\big(\widehat{u}(c)+\widehat{u}(d)\bigr)^p} >\frac{\widehat{u}(a)^p+\widehat{u}(b)^p}{\big(\widehat{u}(a)+\widehat{u}(b)\bigr)^p} \Leftrightarrow \frac{t}{\big(\widehat{u}(c)+\widehat{u}(d)\bigr)^p} >\frac{t}{\big(\widehat{u}(a)+\widehat{u}(b)\bigr)^p} \Leftrightarrow \big(\widehat{u}(a)+\widehat{u}(b)\bigr)^p > \big(\widehat{u}(c)+\widehat{u}(d)\bigr)^p $. This is true if and only if $\widehat{u}(a)+\widehat{u}(b) > \widehat{u}(c) + \widehat{u}(d)$, which is equivalent to
Suppose to the contrary that
From (ref) and (ref) we get $\widehat{u}(a)>\widehat{u}(c)$ and $\widehat{u}(b)<\widehat{u}(d)$. Thus,
By (ref), (ref) and convexity of $\widehat{u}(\cdot)\mapsto\widehat{u}(\cdot)^p$ we have
which contradicts (ref). Hence, (ref) holds. Conversely, suppose (ref) is true and assume to the contrary that (ref) is violated, i.e.
Rearranging (ref),
By (ref) + (ref) we obtain $\widehat{u}(a)<\widehat{u}(c)$. This and (ref) in turn imply $\widehat{u}(b)<\widehat{u}(d)$. Hence,
By (ref) we have
Finally, (ref), (ref) and convexity of $\widehat{u}(\cdot)\mapsto\widehat{u}(\cdot)^p$ jointly lead to the same contradiction as above. This completes the proof that (ref) holds under the second postulate as well.
We now show that (ref) holds under either of the postulated conditions. That is, we verify that $\widehat{u}(a)-\widehat{u}(b) < \widehat{u}(c)-\widehat{u}(d)$ $\Leftrightarrow$ $\widehat{u}(a)^p-\widehat{u}(b)^p<\widehat{u}(c)^p-\widehat{u}(d)^p$. Suppose first that $\widehat{u}(a)+\widehat{u}(b)=\widehat{u}(c)+\widehat{u}(d)$. Let $\widehat{u}(a)-\widehat{u}(b) < \widehat{u}(c)-\widehat{u}(d)$ be true and assume to the contrary that
The former two assumptions imply $\widehat{u}(a)<\widehat{u}(c)$, $\widehat{u}(b)>\widehat{u}(d)$ and therefore
Using again the convexity argument that revolved around (ref) and (ref) we get
By (ref) and (ref) we now obtain $\widehat{u}(a)>\widehat{u}(c)$, which is a contradiction. Conversely, suppose $\widehat{u}(a)^p-\widehat{u}(b)^p<\widehat{u}(c)^p-\widehat{u}(d)^p$ and assume to the contrary that $\widehat{u}(a)-\widehat{u}(b) \geq \widehat{u}(c)-\widehat{u}(d)$. This and $\widehat{u}(a)+\widehat{u}(b) = \widehat{u}(c)+\widehat{u}(d)$ jointly imply $\widehat{u}(a)>\widehat{u}(c)$ and $\widehat{u}(b)<\widehat{u}(d)$. Thus, we have $\widehat{u}(a)>\widehat{u}(c)>\widehat{u}(d)>\widehat{u}(b)$. Using the above convexity argument once again we obtain $\widehat{u}(c)^p+\widehat{u}(d)^p<\widehat{u}(a)^p+\widehat{u}(b)^p$. Subtracting $\widehat{u}(a)^p-\widehat{u}(b)^p<\widehat{u}(c)^p-\widehat{u}(d)^p$ from this inequality yields $\widehat{u}(b)>\widehat{u}(d)$, a contradiction.
Finally, we establish (ref) under the postulate
Let
and again assume to the contrary that (ref) is true. By (ref) + (ref) we get $\widehat{u}(a) \geq \widehat{u}(c)$. This and (ref) implies $\widehat{u}(b) > \widehat{u}(d)$. But $\widehat{u}(a)\geq \widehat{u}(c)$ and (ref) also implies $\widehat{u}(b) \leq \widehat{u}(d)$. This is impossible. Conversely, suppose $\widehat{u}(a)^p-\widehat{u}(b)^p<\widehat{u}(c)^p-\widehat{u}(d)^p$. This and the postulated $\widehat{u}(a)^p+\widehat{u}(b)^p = \widehat{u}(c)^p+\widehat{u}(d)^p$ jointly imply $\widehat{u}(c)>\widehat{u}(a)$ and $\widehat{u}(b)<\widehat{u}(d)$. Together with the without-loss initial assumption whereby $\widehat{u}(a)>\widehat{u}(b)$ and $\widehat{u}(c)>\widehat{u}(d)$, this in turn implies $\widehat{u}(c)>\widehat{u}(a)>\widehat{u}(d)>\widehat{u}(b)$. Assume to the contrary that $\widehat{u}(a)-\widehat{u}(b) \geq \widehat{u}(c)-\widehat{u}(d)$. This is equivalent to $\widehat{u}(d)-\widehat{u}(b)\geq \widehat{u}(c)-\widehat{u}(a)>0$. Rearranging (ref), we also have $\widehat{u}(b)^p-\widehat{u}(d)^p = \widehat{u}(c)^p-\widehat{u}(a)^p$. Since $\widehat{u}(\cdot)\mapsto\widehat{u}(\cdot)^p$ is a strictly increasing function, it follows from the above that the left hand side of this equation is negative while the right hand positive. This is a contradiction. Thus, (ref) holds in this case too. $\blacksquare$\\
Proof of Proposition (ref).
Firm $i=1,2$ maximizes $\pi_i$ with respect to $q_i$ and $p_i$ taking the choices of the other firm $j\neq i$ as given. Differentiating $\pi_i$ with respect to $p_i$, $q_i$ and simplifying we get
Setting the two equations equal to zero yields the first-order conditions
It can be checked upon rearranging these conditions in $\frac{q_i}{p_i}$ form (which, in particular, is a non-negative term) and simplifying that they cannot be satisfied simultaneously under the assumption that $p_i,q_i,s\geq 0$ and $I>0$. This implies that there is no equilibrium where firms choose interior strategies. Since $q_i^{*}\leq p_i^{*}$ must hold, this fact and (ref), (ref) together imply either $p_i^{*}=0$ or $p_i^{*}=I$. Because the latter (former) case is associated with a strictly positive (zero) profit, it follows that
for $i=1,2$. Since the problem is symmetric, by (ref) and $p_i^{*}=I$ we get
for $i=1,2$. Solving this system yields
as claimed. The remaining assertions are verifiable by simple substitution.$\blacksquare$\\
{
\linespread{1}
\printbibliography[notkeyword=software,heading=bibliography,title={Core References}] \printbibliography[keyword=software,heading=bibliography,title={Software References}]
}