EconBase
← Back to paper

Exogenous Consideration and Extended Random Utility

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

55,358 characters · 14 sections · 72 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Exogenous Consideration and Extended Random Utility

abstractIn a consideration set model, an individual maximizes utility among the considered alternatives. I relate a consideration set additive random utility model to classic discrete choice and the extended additive random utility model, in which utility can be $-\infty$ for infeasible alternatives. When observable utility shifters are bounded, all three models are observationally equivalent. Moreover, they have the same counterfactual bounds and welfare formulas for changes in utility shifters like price. For attention interventions, welfare cannot change in the full consideration model but is completely unbounded in the limited consideration model. The identified set for consideration set probabilities has a minimal width for any bounded support of shifters, but with unbounded support it is a point: identification “towards” infinity does not resemble identification “at” infinity.

Introduction

To many economists, “discrete choice” is synonymous with logit or the general nonparametric additive random utility model (ARUM) studied in mcfadden1981econometric. ARUM is widely used, but has many flaws. One criticism is that the representation assumes individuals consider all alternatives, which has motivated analysis of models that do not assume full consideration. Questions that arise include: When is it possible to falsify the hypothesis of full consideration? When can we identify the probability an alternative is considered? How does limited consideration alter counterfactual or welfare analysis?

This paper studies these questions for two models that “update” the classic additive random utility model to allow limited consideration. The first is the extended additive random utility model, in which unobservable utility shocks can equal $-\infty$ for some alternatives. Such alternatives are unavailable, not considered, or arbitrarily undesirable. The second is the additive random utility model with consideration sets. In this model, an individual first forms a consideration set, and then maximizes utility among alternatives in this set. Consideration set formation can arbitrarily depend on unobservable taste shocks, but does not depend on utility indices or covariates and so I term it exogenous.

Overall, this paper finds a divide between results that are robust across all ARUM family models, and results that sharply differ. Classical analysis concerning empirical content (over bounded sets), welfare, and counterfactuals is robust to consideration sets. The robust results establish that the procedural concern that individuals do not consider all alternatives is not automatically an empirical concern, at least for the models studied here. For an attention intervention, which is outside of classical analysis, an extreme contrast emerges: the analyst cannot meaningfully bound welfare increases for an attention intervention for the consideration set model. While the nonparametric consideration set model is more general and allows new questions to be tackled, it is not always capable of answering those questions. This is relevant because a major motivation of the recent literature on consideration sets is to measure welfare changes given attention interventions. More structure is necessary or richer observables, such as menu variation manzini2014stochastic, covariates that only shift attention goeree2008limited, or the extra assumption of independence between consideration and preference heterogeneity abaluck2021consumers,barseghyan2021discrete. I now elaborate on the results.

For the empirically important case of bounded regressors, all ARUM-type models are observationally equivalent. Thus, any behavior that can be described by an ARUM consideration set model can be described by a full consideration ARUM model. This spills over to counterfactual analysis involving a change in utility indices (e.g. a price change): all models reach the same conclusions concerning bounds on choice probabilities. To an outside observer, it is “as if” each individual considers all alternatives. This does not mean ARUM is a “good” consideration set model. Instead, it formally establishes that the empirical limitations of the hypothesis of full consideration cannot be separated from general empirical failures of ARUM. When ARUM is falsified, it is not immediate whether this is due to endogenous consideration, classic endogeneity concerns, or other concerns entirely such as failures of individual optimization.

I next study identification of consideration set probabilities. I find a severe discontinuity. The identified set of a consideration probability has a minimal width for any bounded set of utility index variation, yet with unbounded variation it is a point. Thus, large bounded support does not resemble full support. The intuition for this is we cannot distinguish $\varepsilon_k$ being “very negative” from $\varepsilon_k = -\infty$, interpreted as not considered. This discontinuity also manifests in counterfactual analysis with an attention intervention that makes an alternative always be considered.

Welfare analysis is subtle in consideration set models. masatlioglu2012revealed highlight that the principle of revealed preference breaks. When an alternative is chosen, it is not obvious which alternatives it is preferred to. This makes direct utility comparisons challenging, but I show indirect utility comparisons are straightforward. Specifically, for a change in utility indices such as prices, the change in average indirect utility (or social surplus mcfadden1981econometric) is identified. The reason for this is indirect utility calculations operate over a maximizing envelope and are unaffected by points below the envelope. Here, for all models, choice probabilities are the derivative of average indirect utility and we can integrate to identify changes in welfare. Thus, classic “consumer surplus” formulas are robust to exogenous limited consideration. Such formulas -- or local versions as in “sufficient statistics” chetty2009sufficient -- are widely used empirically.

Next I analyze an attention intervention that makes a specific alternative always be considered. In the classic model there is no scope for a change. In contrast, for the limited consideration model, welfare bounds are the trivial bounds $[0,\infty)$. The model cannot rule out no welfare change, allows an arbitrarily high welfare change, and the bounds do not depend on choice probabilities. The reason for this is simple: individuals who do not consider an alternative may have an arbitrarily high latent utility shock for it. This makes analysis similar to the question of assessing the welfare impact of a new good, which requires enough structure. The benchmark model of exogenous consideration is useful because this result formally establishes that a model that generalizes this setup in one way -- such as by relaxing the exogenous consideration assumption -- must specialize in another to guarantee a finite welfare bound.

The primary analysis assumes utility indices for each alternative are known in order to focus on the role of unobservables. I show observational equivalence extends to the case of unknown utility indices, provided they are continuous and covariates vary over a bounded set. I show that in fact, the identified set for utility indices is the same for each model under these assumptions. This means that identification results for utility indices for one ARUM-type model can be ported to the others.\footnote{Results can always go from the more general consideration set models to classic ARUM, e.g. allen2019identification covers a generalization of the extended ARUM. Porting an identification result from classic ARUM to the other models requires allowing choice probabilities to be separated from 1 to allow limited consideration. This rules out certain classic nonparametric identification results as in Theorem 4 in matzkin1993nonparametric.}

The specific exogenous consideration set model studied here appears to be new, but I mention some setups related to exogenous consideration. The consideration set models of manzini2014stochastic, aguiar2017random, and jagabathula2021demand are formally random utility models in the sense of block1960random, when using menu variation. Consideration set models that can be written as random utility models can be interpreted as exogenous consideration set models: alternatives in the consideration set have a random utility that is well separated above the other alternatives, and the random utilities are independent of the menu and thus exogenous. In a setting of decisions under uncertainty, barseghyan2021discrete assumes covariate shifters are independent of consideration to identify preferences and consideration sets. Likewise, abaluck2021consumers has unobservables independent of prices that shift desirability. Though these models may be general random utility models due to the independence assumption, none of these models are additive random utility models.

This paper is part of a large literature studying choice sets that are latent. A classic paper is manski1977structure. A choice set could be latent for many reasons, including a good going out of stock conlon2013demand. The term “consideration set” is not standardized in economics, but often refers to a latent mental choice set that is not modelled as an endogenous solution to an optimization problem (e.g. masatlioglu2012revealed).\footnote{Outside of economics, more emphasis is placed on the formation of the consideration set itself, e.g. campbell1969existence, narayana1975consumer, hauser1990evaluation, and roberts1991development. See honka2019empirical and crawford2021survey for further references spanning economics, transportation, and marketing.} This distinguishes the terminology from explicit multi-stage choice in which a first stage is optimally chosen,\footnote{Exceptions include caplin2019rational and parts of aguiar2023random.} as in classic nested logit and certain models of search weitzman1978optimal. The theoretical literature has studied empirical content and identification focusing on variation in the set of available alternatives, termed menu variation (manzini2014stochastic, brady2016menu, aguiar2017random, kashaev2022random). Econometric and empirical analysis has focused on covariate variation (chiang1998markov, barseghyan2021discrete, abaluck2021consumers, lu2022estimating, crawford2021survey), including the use of exclusion restrictions that alter consideration but not preferences bacsar2004parameterized,goeree2008limited. See Chapter 11 in strzalecki2023 for additional references and explicit links between menu and covariate variation.\footnote{See also fosgerau2013choice, kawaguchi2021designing, and allen2022revealed in discrete choice, and he2023identification and agarwal2022demand in matching.} In contrast with most of this literature, the goal here is not to separately identify consideration and preferences. This paper instead shows that classic ARUM can be interpreted as a consideration set model. In this sense the paper is closer to matvejka2015rational and fosgerau2020discrete, who interpret classic discrete choice through the lens of rational inattention.

The rest of this paper is organized as follows. Section (ref) defines the models. Section (ref) presents observational equivalence results, and then studies identification of consideration sets and the related question of distiguishing between the models. Section (ref) studies counterfactual analysis with a change in covariates or attention. Section (ref) studies welfare analysis. Section (ref) generalizes the baseline setup to cover unknown utility indices. Section (ref) presents concluding remarks.

The Models

This paper formalizes models in terms of structural choice probabilities $p : U \rightarrow \Delta^K$. The number of alternatives is $K \geq 2$, $\Delta^K$ is the probability simplex consisting of $K$-dimensional vectors that sum to $1$ and are non-negative. The argument of $p$ is a $K$-dimensional vector $u \in U$, interpreted as a vector of utility indices. Thus, $p(u)$ is the vector of choice probabilities given utility indices, $p_k(u)$ is the probability of choosing alternative $k$, and $u_k$ is the utility index for alternative $k$. I treat the utility indices as known in the core analysis. I study observational equivalence and identification questions using variation in utility indices.\footnote{An alternative interpretation of our core analysis is that $u_k$ is negative price, in which we study price variation. In Section (ref) we study unknown utility indices so that $u_k$ is replaced by $v_k(x_k)$ for a vector $x_k$ of observable covariates and unknown function $v_k$. This covers variation in covariates when the utility index is not known.} The set $U \subseteq \mathbb{R}^K$ is the set of utility indices for which $p(u)$ is identified.

I first define the extended additive random utility model and classic additive random utility model. In these models, the individual chooses the alternative $k$ that maximizes $u_k + \varepsilon_k$. Here, $u = (u_1, \ldots, u_K)$ is observable to each agent and the econometrician, and $\varepsilon = (\varepsilon_1, \ldots, \varepsilon_K)$ is unobservable to the econometrician. I formalize consistency of choice probabilities with models below.

defnChoice probabilities $p$ are consistent with the extended additive random utility model (ARUM-E) if there is a distribution $\mu$ over $\varepsilon$ such that: \begin{enumerate}[(i)] • $\varepsilon_k < \infty$ $\mu$-a.s. for each $k$. • $p_k(u) = \Pr_{\mu}(u_k + \varepsilon_k > \max_{j \neq k} u_j + \varepsilon_j)$ for every $u \in U$ and $k$. \end{enumerate} If in addition $\mu$ can be chosen such that $\varepsilon_k > -\infty$ $\mu$-a.s. for each $k$, then $p$ is consistent with the classic additive random utility model (ARUM).

Any such distribution $\mu$ over $\varepsilon$ is said to rationalize the corresponding model. This paper rules out endogeneity between covariates and unobservables because it treats utility indices as parameters that change while the unobservables have the same distributions. Part (ii) assumes a unique maximizer with probability $1$, for any value of utility indices $u \in U$. It rules out utility ties to avoid discussing tie-breaking rules for each model.

Classic ARUM is the model studied in mcfadden1981econometric among many others, and in econometrics is often just referred to as the random utility model. ARUM-E is a generalization of ARUM. The definition of ARUM-E as a stand-alone model appears to be new, but it is a special case of the perturbed utility model studied in allen2019identification. When $\varepsilon_k = -\infty$, alternative $k$ is either unavailable, not considered, or arbitrarily bad.\footnote{Page 4 in fosgerau2013choice mentions allowing $u_k = -\infty$ to accommodate menu variation, i.e. $k$ is unavailable. This paper makes $u$ be finite and thus does not study menu variation.} Thus in ARUM-E, each individual chooses among alternatives with finite $\varepsilon_k$ so as to maximize utility. Note that since a maximizer exists with probability $1$, some $\varepsilon_k$ is finite with probability $1$.

I now formalize the additive random utility model with consideration sets. Let $\mathcal{S}$ denote the set of nonempty subsets of $\{1, \ldots, K\}$. I introduce $\eta \in H$ as an unobservable that alters consideration. The attention function $S : H \rightarrow \mathcal{S}$ describes which sets are considered given $\eta$. Each individual chooses to maximize $u_k + \varepsilon_k$ among alternatives $k$ in the consideration set $S(\eta)$.

defnChoice probabilities $p$ are consistent with the consideration set additive random utility model (ARUM-CS) if there is a distribution $\nu$ over $(\varepsilon,\eta)$ such that: \begin{enumerate}[(i)] • $-\infty < \varepsilon_k < \infty$ $\nu$-a.s. for each $k$. • $p_k(u) = \Pr_{\nu}(\{ u_k + \varepsilon_k > \max_{j \in S(\eta) : j \neq k} u_j + \varepsilon_j \} \cap \{ k \in S(\eta)\})$ for every $u \in U$ and $k$. \end{enumerate}

Part (i) restricts each $\varepsilon_k$ to be finite. Part (ii) models that $k$ is chosen when it maximizes utility among alternatives in the consideration set. Note that this definition allows variables determining consideration sets ($\eta$) to be arbitrarily correlated with the other shocks ($\varepsilon$).

To close the setup, we collect the definitions of distributions that rationalize $p$ according to the respective models.

defnLet $\mathcal{M}^{ARUM}$ denote the set of distributions over real-valued $\varepsilon$ that rationalize $p$ according to ARUM. Let $\mathcal{M}^{ARUM-E}$ denote the set of distributions over extended real-valued $\varepsilon$ that rationalize $p$ according to ARUM-E. Let $\mathcal{M}^{ARUM-CS}$ denote the set of distributions over $(\varepsilon,\eta)$ that rationalize $p$ according to ARUM-CS.
remark[Extended Reals] ARUM-E requires working with addition in the extended reals. This paper defines $a + -\infty = -\infty$ for any $a \neq \infty$. Thus, when $\varepsilon_k = -\infty$ we have $u_k + \varepsilon_k = -\infty$ because the utility index $u_k$ is finite. Note that because $a > b$ only when $a > -\infty$, part (ii) of the definition of ARUM-E implies that for each $u \in U$, with probability $1$, $\varepsilon_k$ is finite for some $k$.
remark[Normalizations] For ARUM and ARUM-CS, it is a normalization to set $u_1 + \varepsilon_1 = 0$ for alternative $1$, so that we can interpret the remaining $u_k + \varepsilon_k$ numbers as differences relative to alternative $1$. By normalization I mean the empirical content of the model is the same with this extra assumption. The equality is not a normalization in ARUM-E because setting $u_1 + \varepsilon_1 = 0$ implies that alternative $1$ is always considered.\footnote{This would imply that $\sup_{u \in U} p_1(u) = 1$ if $U = \mathbb{R}^K$ for example. I note that it is a normalization to set $u_1 = 0$ because utility here is always finite.} In particular, ARUM-E should not generally be interpreted as a model in differences. If $\varepsilon_{1} = - \infty$ for example, then $\varepsilon_k - \varepsilon_1$ is not defined. Thus for ARUM-E, the analyst should not make an assumption that one alternative has a set utility of $0$, unless the analyst also wishes to assume the alternative is always considered.

Observational Equivalence

I first present conditions under which the models are observationally equivalent, and then in Section (ref) study how to distinguish between the models and the closely related question of identification of features of the distribution of consideration sets.

Say that $U$ has bounded utility differences if $\sup_{u \in U} |u_k - u_j| < \infty$, for any alternatives $j$ and $k$.

prop\begin{enumerate}[(i)] • For any $U \subseteq \mathbb{R}^K$, ARUM-E and ARUM-CS are observationally equivalent, i.e. $p$ is consistent with ARUM-E if and only if it is consistent with ARUM-CS. • If $U$ has bounded utility differences, then ARUM, ARUM-E, and ARUM-CS are observationally equivalent. \end{enumerate}

Part (i) states that we can never distinguish between ARUM-E and ARUM-CS.\footnote{See Proposition 5 in barseghyan2021discrete for a related random utility model.} Working with the extended reals thus provides an alternative representation to a two-step choice model.\footnote{Encoding constraints in this manner is used in convex analysis, cf. rockafellar2015convex.} The more noteworthy result is (ii), which shows that when utility differences are bounded, it is impossible to distinguish such models from the classic additive random utility model.\footnote{Proposition 2 in jagabathula2021demand shows observational equivalence of a general random utility model (with nonseparable shocks) and a general consideration set model. We differ by considering different models and by using variation in latent utility indices and not variation in the set of available alternatives.} Thus, while it may be intuitively unappealing to assume individuals consider all alternatives, this concern can be addressed within ARUM by allowing a sufficiently flexible distribution of disturbances. Note that ARUM can be interpreted as a special case of ARUM-CS in which $\eta$ is constant and thus independent of $\varepsilon$. Thus, when $U$ has bounded utility differences, adding independence between $\eta$ and $\varepsilon$ is without loss of generality for studying empirical content.

Observational equivalence over bounded sets means that empirical content results that use local structure have equivalent translations between the models. Thus, local empirical content characterizations of ARUM in koning1994compatibility and koning2003discrete are also characterizations of ARUM-E and ARUM-CS for bounded sets.\footnote{Conditions in koning1994compatibility are the classical Williams-Daly-Zachary-McFadden conditions proven to characterize ARUM in mcfadden1981econometric, except dropping the requirement that probabilities limit to $0$ and $1$ at extreme values of covariates.} Or stated differently, their ARUM results are automatically robust to consideration sets.

I next study when it is possible to distinguish between ARUM and the other models, and the closely related question of identification of the distribution of latent feasibility sets.

Identification and Distinguishability

I begin by presenting sharp identification results for certain consideration set probabilities. I then use these to establish that ARUM can be distinguished from ARUM-E (or ARUM-CS) if and only if certain marginal consideration set probabilities are identified. With Proposition (ref) this establishes that in general, unbounded utility differences are necessary for point identification of these marginal probabilities.

The identification results concern probabilities of the form $\Pr(\varepsilon_k > -\infty)$ for ARUM-E, and $\Pr(k \in S(\eta))$ for ARUM-CS, i.e. the probability $k$ is considered. These probabilities are relevant because they must be 1 for ARUM. I define identified sets for these marginal probabilities, which reflect that there are typically multiple distributions that can rationalize the structural choice probabilities. I define the identified set for $\Pr(\varepsilon_k > -\infty)$ as the union over all possible distributions that rationalize $p$ according to ARUM-E: \[ \Theta^{k,E} = \cup_{\mu \in \mathcal{M}^{ARUM-E}} \Pr_{\mu} (\varepsilon_k > -\infty). \] Similarly for $\Pr(k \in S(\eta))$ I define \[ \Theta^{k,CS} = \cup_{\nu \in \mathcal{M}^{ARUM-CS}} \Pr_{\nu}(k \in S(\eta)). \]

lemmaThe identified sets for marginal consideration probabilities are convex and satisfy \[ \Theta^{k,E} = \Theta^{k,CS}. \] Moreover, $\sup_{u \in U} p_k(u)$ is a lower bound for each set. That is, for any $c \in \Theta^{k,E}$, it follows that $\sup_{u \in U} p_k(u) \leq c$.

Page 1980 in barseghyan2021discrete notes a lower bound in a related context.

Distinguishability and identification questions hinge on whether $k$ can be made extremely attractive, i.e. when \[ \sup_{u \in U} \min_{j \neq k} \{ u_k - u_j \} = \infty. \] Let $\mathcal{K}^{EA}$ be the set of alternatives that can be made extremely attractive. This set $\mathcal{K}^{EA}$ is empty if and only if $U$ has bounded utility differences. As a final definition, say $u \in U$ is $k$-maximal if for any $w \in U$ and $j$, $u_k - u_j \geq w_k - w_j$.

propSuppose $p$ is consistent with ARUM-E or equivalently ARUM-CS. \begin{enumerate}[(i)] • If $k \in \mathcal{K}^{EA}$, then \[ \Theta^{k,E} = \Theta^{k,CS} = \sup_{u \in U} p_k(u). \] • If $U$ has bounded utility differences (equivalently, $\mathcal{K}^{EA}$ is empty) and $\Theta^{k,E}$ is a singleton, then \[ \Theta^{k,E} = \Theta^{k,CS} = \sup_{u \in U} p_k(u) = 1. \] • If $U$ has bounded utility differences (equivalently, $\mathcal{K}^{EA}$ is empty) and contains a $k$-maximal point $u^*$, then \[ \Theta^{k,E} = \Theta^{k,CS} = \left[p_k(u^*),1 \right] = \left[ \sup_{u \in U} p_k(u), 1 \right]. \] \end{enumerate}

Part (i) states that point identification of marginal consideration set probabilities is possible for any alternative that can be made extremely attractive. Part (ii) states that with bounded utility differences, point identification of these marginal consideration probabilities is equivalent to $\sup_{u \in U} p_k(u) = 1$. Part (iii) characterizes the identified set for marginal consideration probabilities for domains that contain a $k$-maximal point. An example of a set that contains $k$-maximal points for any $k$ is U a Cartesian product of compact intervals or more generally, a product of compact subsets of $\mathbb{R}$. One such set is $\{ 0, 1\}^K$, which has a $k$-maximal element for each $k$. For example, the $1$-maximal element is $(1,0, \ldots, 0)$.

Unbounded regressors have been used previously to identify the distribution of consideration sets in a variety of settings abaluck2021consumers,barseghyan2021discrete,hyung2021,kashaev2023peer.\footnote{abaluck2021consumers identify changes in consideration set probabilities as prices change, for a different generalization of ARUM. Loosely, $S(\eta)$ can depend on $u$ (negative prices) in their setup in certain ways. That paper requires prices to limit to infinity to identify the levels of consideration set probabilities.} The primary novelty here is to highlight that for the ARUM family studied here, large support is typically necessary for point identification of the marginal distributions of consideration sets.

I present three implications of Proposition (ref). To state the first result I generalize the notation a bit to accommodate a growing domain over which $p$ is identified. To that end, for any $U \subseteq \mathbb{R}^K$ let $\Theta^{k,E}(U) = \cup_{\mu \in \mathcal{M}^{ARUM-E}(U)} \Pr_{\mu} (\varepsilon_k > -\infty)$, where $\mathcal{M}^{ARUM-E}(U)$ denotes distributions that rationalize $p$ according to ARUM-E when $p$ is identified over $U$. Define $\Theta^{k,CS}(U)$ analogously.

cor[Discontinuous Identification] Let $\{ U^s \}_{s = 1}^{\infty}$ be an increasing collection of subsets of $\mathbb{R}^K$, i.e. $U^s \subseteq U^m$ for $s \leq m$. Assume further that each set $U^s$ is a compact rectangle, i.e. $U^s = \times_{j = 1}^K [\underline{u}^s_j, \overline{u}_j^s]$. Let $U^{\infty} := \cup_{s = 1}^{\infty} U^s$. \begin{enumerate}[(i)] • For any $k$, \[ \left[\sup_{u \in U^{\infty}} p_k(u),1 \right] \subseteq \cap_{s = 1}^{\infty} \Theta^{k,E}(U^s) = \cap_{s = 1}^{\infty} \Theta^{k,CS}(U^s). \] • For any $k$ that can be made extremely attractive with domain $U^{\infty}$, \[ \sup_{u \in U^{\infty}} p_k(u) = \Theta^{k,E}(U^{\infty}) = \Theta^{k,CS}(U^{\infty}). \] \end{enumerate}

Putting these results together with Proposition (ref)(iii), we conclude that if $U^{\infty}$ is rich enough so that $k$ can be made extremely attractive, the left hand side of Corollary (ref)(i) is a point if and only if $k$ is considered with probability $1$. In general, there is a severe discontinuity of identification contrasting the case of bounded variation in utility indices with the limit of full variation. For any set $U^s$, no matter how big, the minimal width of the identified set is $1 - \sup_{u \in U^{\infty}} p_k(u)$, but in the limit it is a single point. Thus, it is not the case that large bounded variation resembles the case of unbounded variation. See magnac2007identification and khan2010irregular for other contexts with such a discontinuity, and how this makes estimation challenging.

The rest of the paper keeps $U$ fixed.

cor[Distinguishability] Assume $p$ is consistent with ARUM-E or equivalently ARUM-CS. Assume each alternative can be made extremely attractive, i.e. $\mathcal{K}^{EA} = \{1, \ldots, K\}$. The following are equivalent: \begin{enumerate}[(i)] • $p$ is consistent with ARUM. • $\sup_{u \in U} p_k(u) = 1$ for each $k$. \end{enumerate}

This characterizes the extra empirical restrictions of ARUM.

cor[Nontrivial Rationalizations] Assume $p$ is consistent with either ARUM-E or ARUM-CS, and that for some $k$, $\sup_{u \in U} p_k(u) < 1$. Assume $U$ has bounded utility differences and contains a $k$-maximal point. It follows that: \begin{enumerate}[(i)] • $p$ can be rationalized according to ARUM-E with a distribution $\mu \in \mathcal{M}^{ARUM-E}$ that satisfies \[ \Pr_{\mu}(\varepsilon_k > -\infty) < 1. \]$p$ can be rationalized according to ARUM-CS with a distribution $\nu \in \mathcal{M}^{ARUM-CS}$ that satisfies \[ \Pr_{\nu}(k \in S(\eta)) < 1. \] \end{enumerate}

This establishes that if the choice probability for some alternative $k$ is separated from $1$, then we we can rationalize $p$ with a distribution in which some alternative is not considered with positive probability. Thus, it is “as if” some individuals do not consider all alternatives.

Counterfactuals

This section builds on the previous results to study implications for counterfactuals. I analyze two types of counterfactual interventions. The first is a utility index intervention, which involves setting the utility index $u$ to a new value, holding the distribution of unobservables fixed. If the utility index for each alternative is its (negative) price, this amounts to a counterfactual price change. The second is an attention intervention, which changes consideration set formation, holding everything else fixed. When we identify probabilities over a region with bounded utility differences, the models all deliver the same counterfactual bounds for utility-interventions, but deliver contrasting bounds for attention interventions.

Utility Index Interventions

Suppose the analyst knows $p$ over $U$, and is interested in theory-consistent counterfactuals at a new value $u^C \not\in U$. I formulate counterfactual restrictions in terms of an extension $p^C : U \cup \{u^c\} \rightarrow \Delta^K$ that agrees with $p$ on the set $U$. That is, $p^C(u) = p(u)$ for $u \in U$. The interpretation is $p^C(u^C)$ represents choice probabilities at the new value of utility indices. If $p^C$ is consistent with ARUM, say it is ARUM-consistent, and similarly for the other models. The set of ARUM-consistent functions is given by $\mathcal{P}_{ARUM}^C$. For ARUM, define the identified set for counterfactuals at the new value $u^C$ as \[ \Theta^{C}_{ARUM} = \cup_{p^C \in \mathcal{P}^C_{ARUM}} p^C(u^C). \] For any element $a \in \Theta^{C}_{ARUM}$, there is a distribution over $\varepsilon$ that matches the known probabilities over $U$, and that yields $a$ as the choice probability at the new value $u^C$. For ARUM-E and ARUM-CS, define the identified sets analogously.

prop\begin{enumerate}[(i)] • If $U \subseteq \mathbb{R}^K$ has bounded utility differences, then for any $u^C \in \mathbb{R}^K$, \[ \Theta^C_{ARUM} = \Theta^{C}_{ARUM-E} = \Theta^{C}_{ARUM-CS}. \] • For any $U \subseteq \mathbb{R}^K$ and $u^C \in \mathbb{R}^K$, \[ \Theta^{C}_{ARUM-E} = \Theta^{C}_{ARUM-CS}. \] \end{enumerate}

Part (i) states that for utility interventions, the identified sets for counterfactuals are the same for the three models when utility differences are bounded. Note that for general unbounded $U \subseteq \mathbb{R}^K$, it is possible that \[ \Theta^C_{ARUM} \neq \Theta^{C}_{ARUM-E} \] because the left hand side can be empty when the right hand side is not. This can happen when $p$ is consistent with ARUM-E but not ARUM. In this case, there is no extension of $p$ to the counterfactual point $p^C(u^C)$ using the model ARUM, and so $\Theta^C_{ARUM}$ is empty.

Attention Interventions

An attention intervention changes consideration set formation without changing preferences. I record a first fact.

factAttention interventions in ARUM do not change structural choice probabilities.

This is no longer true in consideration set models and so I study the scope for such interventions. In ARUM-CS, I model an intervention by changing the mapping $S(\cdot)$, which maps latent attention factors to consideration sets. A $k$-attention intervention makes alternative $k$ always feasible. That is, $S$ changes to $\tilde{S}^k$, which is $S$ with $k$ added to the consideration set, i.e. $\tilde{S}^k(\eta) = S(\eta) \cup \{k\}$. Given an attention intervention, an admissible counterfactual is given by a counterfactual probability function $p^{\tilde{S}^k} : U \rightarrow \Delta^K$. This $p^{\tilde{S}^k}$ must satisfy the property that there is a distribution $\nu \in \mathcal{M}^{ARUM-CS}$ over $(\varepsilon,\eta)$ such that for any $u \in U$ and $j$, \[ p^{\tilde{S}^k}_j(u) = \Pr_{\nu} \left( \left\{ u_j + \varepsilon_j > \max_{\ell \in \tilde{S}^k (\eta) : \ell \neq k} u_{\ell} + \varepsilon_{\ell} \right\} \right), \] and that \[ p_j(u) = \Pr_{\nu} \left( \left\{ u_j + \varepsilon_j > \max_{\ell \in S(\eta) : \ell \neq k} u_{\ell} + \varepsilon_{\ell} \right\} \cap \{ j \in S(\eta)\} \right). \] In words, there must be a distribution of tastes ($\varepsilon$) and latent attention factors $(\eta)$ that generates the counterfactual probability $p^{\tilde{S}^k}$ while matching the primitive structural choice probability $p$. Let $\mathcal{P}^{Attention}$ be the set of such $p^{\tilde{S}^k}$ functions.

I study bounds on structural choices probabilities before and after a $k$-attention intervention. I analyze bounds on the quantity \[ p^{\tilde{S}^k}_k(u) - p_k(u), \] where $p^{\tilde{S}^k}$ is an admissible counterfactual mapping. An obvious feature is that this difference is nonnegative, because adding $k$ to the consideration set cannot make its choice probability go down. I characterize the highest it can be below.

propAssume $U$ has bounded utility differences and $U$ contains a $k$-maximal point. The identified set for the maximal change given a $k$-attention intervention is given by \[ \cup_{p^{\tilde{S}^k} \in \mathcal{P}^{Attention}} \sup_{u \in U} \left\{ p^{\tilde{S}^k}_k(u) - p_k(u) \right\} = \left[0, 1 - \sup_{u \in U} p_k(u) \right]. \]

The left and side is the definition of the identified set. The lower bound $0$ states that we cannot rule out that an attention intervention may have no change in choice probabilies. The upper bound states that the attention intervention may shift choice probability for alternative $k$ up by the amount $1 - \sup_{u \in U} p_k(u)$. Thus, the maximum scope for a $k$-attention intervention is higher when $\sup_{u \in U} p_k(u)$ is lower. Indeed, when this supremum is $1$, there is no scope for $k$-attention interventions to change the probability of choosing alternative $k$, because the data indicate $k$ is always considered.

When $U$ has unbounded utility differences, I show a qualitatively different result for any alternative that can be made extremely attractive.

propAssume $U \subseteq \mathbb{R}^K$ and that $k$ can be made extremely attractive, i.e. $k \in \mathcal{K}^{EA}$. For every admissible counterfactual $p^{\tilde{S}^k} \in \mathcal{P}^{Attention}$ we have \[ \sup_{u \in U} \left\{ p^{\tilde{S}^k}_k(u) - p_k(u) \right\} = 1 - \sup_{u \in U} p_k(u). \]

This states that the maximum scope for $k$-attention interventions is uniquely identified for any alternative that can be made extremely attractive. Moreover, as long as $\sup_{u \in U} p_k(u) < 1$, we identify that there must be scope for $k$-attention interventions. This contrasts with the bounded utility difference case of Proposition (ref), which states that it is always possible that an attention intervention does not change choice probabilities. Thus, there is a severe discontinuity between the unbounded and bounded cases, as in Corollary (ref).

I omit formal analysis of ARUM-E and mention a key consideration that this paper is agnostic about. In ARUM-E, $\varepsilon_k = -\infty$ may be interpreted as alternative $k$ either not being available, not considered, or arbitrarily unattractive. The specific interpretation determines the scope for the specific intervention. For example, if $k$ is always considered but sometimes not available because it is sold out, then an attention intervention may do nothing. Likewise, if $k$ is always considered but arbitrarily unattractive to some people (e.g. $k$ is a bag of peanuts and some people have a severe allergy), then an attention intervention will do nothing.

Welfare

I now study welfare analysis using the different ARUM-family models. Paralleling the counterfactual analysis of Section (ref), I analyze two welfare questions. The first is how welfare changes when utility indices change from $u$ to $\tilde{u}$, e.g. $u = -p$ is negative price and some prices change. I show that welfare analysis does not depend on which ARUM family model the analyst uses, at least when $U$ is convex. The second question is how welfare changes given an attention intervention that makes each individual always consider an alternative. I establish the striking result that for ARUM-CS, it is not possible to obtain meaningful bounds on welfare changes for interventions that make an alternative always be considered.

Throughout, I work with an average indirect utility notion for welfare changes. For $\mu \in \mathcal{M}^{ARUM}$, define the average indirect utility $V^{ARUM}_{\mu} : \mathbb{R}^K \rightarrow \mathbb{R}$ as \[ V_{\mu}^{ARUM}(u) = \mathbb{E}_{\mu} \left[ \max_{k} \{ u_k + \varepsilon_k \} - \max_{k} \{ \varepsilon_k \} \right]. \] Note this function is defined over all of $\mathbb{R}^K$ not just $U$. This $V$ is sometimes called the social surplus function. Subtraction is standard and ensures that this expectation exists and is finite (sorensen2022mcfadden). Define ARUM-E identically, except the distribution $\mu^E$ can allow $\varepsilon_k = -\infty$ for some $k$.\footnote{Recall for ARUM-E, a unique maximizer exists $\mu^E$-a.s., so $\max_k \{\varepsilon_k \}$ is finite $\mu^E$-a.s.} For ARUM-CS, define \[ V_{\nu}^{ARUM-CS}(u) = \mathbb{E}_{\nu} \left[ \max_{k \in S(\eta)} \{ u_k + \varepsilon_k \} - \max_{k} \{ \varepsilon_k \} \right], \] where $\nu$ is a distribution over $(\varepsilon,\eta)$.

Utility Index Changes

The following envelope theorem is very useful.

lemmaAssume $p$ is consistent with ARUM, ARUM-E, and ARUM-CS. For any $u \in U$, $\mu \in \mathcal{M}^{ARUM}$, $\mu^E \in \mathcal{M}^{ARUM-E}$, and $\nu \in \mathcal{M}^{ARUM-CS}$, it follows that: \[ p(u) = \nabla_u V_{\mu}^{ARUM}(u) = \nabla_u V_{\mu^{E}}^{ARUM-E}(u) = \nabla_u V_{\nu}^{ARUM-CS}(u). \]

Differentiability of each $V$ is established in the proof using the fact that the maximizer is unique with probability $1$. I note that Lemma (ref) holds for $U$ discrete or even a singleton $\{ u \}$. The equality for ARUM is well-known under different conditions, e.g. mcfadden1981econometric.\footnote{For ARUM, mcfadden1981econometric and shi2018estimating impose a density assumption that rules out utility ties for any $u \in \mathbb{R}^K$. Here we assume ties occur with probability $0$ only for $u \in \mathcal{U}$. See sorensen2022mcfadden for a version allowing utility ties (which formally is more general than ARUM defined here), in which case the value function may no longer be differentiable.} Lemma 1 in allen2019identification covers ARUM-E under alternative high-level conditions. The ARUM-CS case does not have clear precedent.

I consider the welfare change in moving from $u$ to $\tilde{u}$. For ARUM this is given by \[ \Delta^{ARUM}(\tilde{u},u,\mu) = V_{\mu}^{ARUM}(\tilde{u}) - V_{\mu}^{ARUM}(u), \] and I similarly define $\Delta^{ARUM-E}$ and $\Delta^{ARUM-CS}$. Recall that as defined in Section (ref), $\mathcal{M}^{ARUM}$ is the set of distributions that rationalize $p$ according to ARUM, and similarly for $\mathcal{M}^{ARUM-E}$ and $\mathcal{M}^{ARUM-CS}$. The identified set for welfare changes in ARUM is given by \[ \cup_{\mu \in \mathcal{M}^{ARUM}} \Delta^{ARUM}(\tilde{u},u,\mu), \] and defined analogously for ARUM-E and ARUM-CS.

I use the envelope theorem above to integrate probabilities and identify differences in average indirect utility.

propAssume $p$ is consistent with ARUM, ARUM-E, and ARUM-CS. Assume $U \subseteq \mathbb{R}^K$. Let $u, \tilde{u} \in U$, and assume that for any $\alpha \in [0,1]$, $\alpha u + (1 - \alpha) \tilde{u} \in U$. It follows that \begin{align*} \cup_{\mu \in \mathcal{M}^{ARUM}} \Delta^{ARUM}(\tilde{u},u,\mu) & = \cup_{\mu \in \mathcal{M}^{ARUM-E}} \Delta^{ARUM-E}(\tilde{u},u,\mu) \\ & = \cup_{\nu \in \mathcal{M}^{ARUM-CS}} \Delta^{ARUM-CS}(\tilde{u},u,\nu) \\ & = \int_0^1 p(t \tilde{u} + (1 - t) u) \cdot (\tilde{u} - u) dt. \end{align*}

This result provides a constructive point identification result for welfare changes that applies to each model. Point identification for ARUM and ARUM-E is known under slightly different conditions, e.g. results in small1981applied and Theorem 4 in allen2019identification. The novelty here is that (i) we can also identify welfare changes for ARUM-CS, and (ii) all models deliver the same welfare implications for utility changes.

Attention Changes

I now analyze how welfare changes with an attention intervention like in Section (ref). In ARUM, attention increases do not change welfare because individuals pay attention to all alternatives. In ARUM-E, the role of attention interventions depends on the interpretation of the event $\varepsilon_k = -\infty$. It could be due to stock out, limited consideration, or just the alternative being arbitrarily undesirable. The interpretation determines the role of attention interventions. I thus focus on ARUM-CS, where this ambiguity does not arise because $\varepsilon_k$ must be finite and this paper interprets $S(\eta)$ as the set of items that are considered.

Define now \[ V_{\nu}^{ARUM-CS}(u,S) = \mathbb{E}_{\nu} \left[ \max_{k \in S(\eta)} \{ u_k + \varepsilon_k \} - \max_{k} \{ \varepsilon_k \} \right]. \] Here, $S$ is added as an argument because I consider changes that make an alternative available. That is, I consider a $k$-attention intervention that appends $k$ to the consideration set via $\tilde{S}^k(\eta) = S(\eta) \cup \{ k \}$. The welfare change given a $k$-attention intervention is \[ V_{\nu}^{ARUM-CS}(u,\tilde{S}^k) - V_{\nu}^{ARUM-CS}(u,S). \] The identified set for this difference is defined across distributions $\nu$ that rationalize $p$ with the model ARUM-CS. I characterize it as follows.

propAssume $p$ is consistent with ARUM-CS. Let $U \subseteq \mathbb{R}^K$, $u \in U$, and let $k$ be a specific alternative. \begin{enumerate}[(i)] • If $\sup_{u \in U} p_k(u) = 1$, then the identified set for a $k$-attention intervention is \[ \cup_{\nu \in \mathcal{M}^{ARUM-CS}} \left\{V_{\nu}^{ARUM-CS}(u,\tilde{S}^k) - V_{\nu}^{ARUM-CS}(u,S) \right\} = 0. \] • If $\sup_{u \in U} p_k(u) < 1$ and $U$ is bounded and contains a $k$-maximal point, or alternative $U$ is unbounded and $k$ can be made extremely attractive, then the identified set for a $k$-attention intervention is \[ \cup_{\nu \in \mathcal{M}^{ARUM-CS}} \left\{V_{\nu}^{ARUM-CS}(u,\tilde{S}^k) - V_{\nu}^{ARUM-CS}(u,S) \right\} = [0,\infty). \] \end{enumerate}

Part (i) states that when alternative $k$ takes probability arbitrarily close to $1$, then an attention intervention does not alter welfare. Part (ii) is striking, since it states that the only restriction is that higher attention cannot decrease utility. The precise values of $p$ are irrelevant when $\sup_{u \in U} p_k(u) < 1$. This indicates a fundamental limit of welfare analysis in this setting. This limit arises because when an alternative is not considered, its utility shock may be arbitrarily high. Thus, making it be considered may lead to an arbitrarily high increase in indirect utility.

Unknown Utility Indices

The previous analysis focused on the case in which utility indices are known. In practice, it is common to consider covariates that vary and alter utility in a way that is not known a priori. This section extends a snapshot of the previous results to this setting.

I describe the modified setup. Instead of the utility index $u_k$ for alternative $k$, replace it with $v_k(x_k)$, where $v_k$ is an unknown function and $x_k$ is a vector of observable regressors that alter the desirability of alternative $k$. Such regressors could be different across alternatives (such as alternative characteristics) or common (such as demographic variables). This setup allows $v_k$ to be linear but does not require this. All covariates are collected in $x = (x'_1, \ldots, x'_K)'$. Each $x_k \in \mathbb{R}^{d_k}$ and the overall vector satisfies $x \in \mathbb{R}^d$ where $d = \sum_{k = 1}^K d_k$. Utility indices are collected as $v(x) = (v_1(x_1), \ldots, v_K(x_K))'$. The structural probabilities $p : U \rightarrow \Delta^K$ are not directly identified here. Instead, I consider the mapping $\tilde{p}(x) = p(v_1(x_1), \ldots, v_K(x_K))$. Then $\tilde{p}_k(x)$ is the probability of choosing alternative $k$ given covariates $x$. I assume $\tilde{p} : \mathcal{X} \rightarrow \Delta^K$ is known over $\mathcal{X}$. One can interpret $\mathcal{X}$ as the observable range over which covariates vary.

Distinguishability and Identification of Utility Indices

I study when it is possible to distinguish between models when utility indices are not known in advance, and the related question of identification of $v$ for each model.

defnChoice probabilities $\tilde{p}$ are consistent with ARUM with utility function $v : \mathcal{X} \rightarrow \mathbb{R}^K$ if there exists a function $p : v(\mathcal{X}) \rightarrow \Delta^K$ such that $p$ is consistent with ARUM and $\tilde{p}(x) = p(v(x))$ for each $x \in \mathcal{X}$.

The set $v(\mathcal{X})$ is the range of $v$. Consistency with ARUM-E and ARUM-CS is defined analogously.

propAssume $\mathcal{X} \subseteq \mathbb{R}^d$ and that $\tilde{p}$ is consistent with either ARUM, ARUM-E, or ARUM-CS with a utility $v$ that is bounded over $\mathcal{X}$. It follows that $\tilde{p}$ is consistent with ARUM, ARUM-E, and ARUM-CS with the same utility $v$.

An important case is when $\mathcal{X}$ is compact and we restrict attention to continuous utility functions $v$. This means $v$ is bounded over $\mathcal{X}$ and so any such $v$ that rationalizes one model can rationalize any other. I formalize this with some additional notation. To that end, let $\mathcal{U}$ denote a set of candidate utility functions from $\mathcal{X}$ to $\mathbb{R}^K$, which forms the parameter space. For the model ARUM, define \[ \Theta^{ARUM}_{\mathcal{U}} = \{ v \in \mathcal{U} \mid \tilde{p} \text{ is consistent with ARUM with utility } v \}. \] Define this identified set similarly for each model.

corAssume $\mathcal{X} \subseteq \mathbb{R}^d$ is compact and $\mathcal{U}$ consists of continuous functions. \begin{enumerate}[(i)] • (Identified Sets are Equal.) \[ \Theta^{ARUM}_{\mathcal{U}} = \Theta^{ARUM-E}_{\mathcal{U}} = \Theta^{ARUM-CS}_{\mathcal{U}}. \] • (Point Identification.) If one of the sets $\Theta^{ARUM}_{\mathcal{U}}$, $\Theta^{ARUM-E}_{\mathcal{U}}$ and $\Theta^{ARUM-CS}_{\mathcal{U}}$ is a singleton, then all are singletons. • (Observational Equivalence.) Assume $\tilde{p}$ is consistent with one of the models ARUM, ARUM-E, or ARUM-CS for some $v \in \mathcal{U}$. It follows that $\tilde{p}$ is consistent with all of the models ARUM, ARUM-E, and ARUM-CS for the same $v \in \mathcal{U}$. \end{enumerate}

Part (i) states the identified sets are the same here. Part (ii) states point identification in one model is equivalent to point identification in all. For example, allen2019identification identify utility indices in a generalization of ARUM-E allowing compact support for regressors and requiring continuous utility indices. Part (iii) states the models are observationally equivalent here. This indicates that counterfactual analysis involving covariate varition also coincides.

Without compactness of $\mathcal{X}$, ARUM is no longer observationally equivalent to the other models, and the identified sets can differ. Lack of observational equivalence is true for the simple example $\mathcal{U}$ equal to the identity mapping (so that covariates are scalar for each alternative), and $\mathcal{X} = \mathbb{R}^K$. Then from Corollary (ref) we conclude ARUM is not observationally equivalent to ARUM-E or ARUM-CS. To see why identified sets can differ when $\mathcal{X}$ is not compact, consider a setting in which $\sup_{x \in \mathcal{X}} \tilde{p}_k(x) < 1$ for each $k$. ARUM requires $v$ with bounded utility differences, whereas ARUM-E and ARUM-CS do not.

Identification of Marginal Consideration Probabilities

I now turn to identification of marginal consideration probabilities $\Pr(\varepsilon_{k} > -\infty)$ and $\Pr(k \in S(\eta))$ when utility indices are not known in advance.

propAssume $\mathcal{X} \subseteq \mathbb{R}^{d}$ is the Cartesian product of $d$ compact sets, each in $\mathbb{R}$, and that $\mathcal{U}$ consists of continuous functions. Assume $\tilde{p}$ is known over $\mathcal{X}$ and is consistent with ARUM-E for some $v \in \mathcal{U}$. The identified sets for $\Pr(\varepsilon_{k} > -\infty)$ and $\Pr(k \in S(\eta))$ are equal, and are given by \[ \left[\sup_{x \in \mathcal{X}} \tilde{p}_k(x) , 1 \right]. \]

This result provides conditions under which identification of utility indices is separate from identification of these marginal consideration set probabilities: sharp bounds depend only on $\sup_{x \in \mathcal{X}} \tilde{p}_k(x)$.

Discussion

How does limited consideration alter analysis relative to full-consideration models? The answer depends on the baseline model of choice and the extension to allow limited consideration. This paper asks the question with the baseline additive random utility models and two extensions that allow exogenous limited consideration. The main motivation for studying these models is that despite its limitations, the additive random utility model is widely used, and the extensions provide a natural “update” to the classic model.

The broad finding of this paper is that with the exception of attention interventions -- which clearly require a model of limited attention -- the classic and updated models deliver identical answers to certain questions in empirically relevant settings. The classic nonparametric model can thus be interpreted as a consideration set model. The sharpest contrast is measuring welfare with an attention intervention that makes one alternative be considered. The consideration set model only concludes the welfare contrast is somewhere in the trivial set $[0,\infty)$, while the full attention model states it must be zero.

I close with further remarks and qualifications.

This paper adds ARUM to the list of consideration set models. The union of the empirical content of models that have been used to study consideration sets is now large. I illustrate by contrasting two other models. For general random utility (with nonseparable unobservables as in block1960random), regularity states that if an alternative is added to the choice set, the probability of choosing an existing alternative cannot go up. The model studied in manzini2014stochastic is a random utility model and requires regularity, while the model of cattaneo2020random requires violations of regularity to identify limited consideration. Comparing these two models shows that whether regularity is or is not a hallmark of limited consideration is thus model specific.

Overall, it seems that the “right” model is setting-question specific. ARUM is a reasonable candidate if the analyst wishes to do classic analysis that is robust to exogenous consideration sets. An important caveat is that this paper shows this only for the nonparametric model studied here. This paper is silent on comparison of a parsimonious ARUM model with a parsimonious consideration set model. In order for ARUM to mimic a consideration set model, it must have fat tails for the distribution of shocks. Loosely, if $\varepsilon_k = -\infty$ with positive probability, ARUM can only mimic this with a distribution of shocks that allows very negative values. Traditional econometric analysis rules this out. For example, common parametric choices like logit and normal shocks have very thin tails, and may be ill-suited when consideration sets may be relevant.