Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
98,404 characters · 11 sections · 41 citation commands
Attention Overload
\onehalfspacing
Keywords: attention frequency, limited and random attention, revealed preference, partial identification, high-dimensional inference.
\thispagestyle{empty}
\doublespacing
\setcounter{page}{1}
This paper studies decision making in settings where the decision makers confront an abundance of options, their consideration sets are random, and their attention span is limited. We assume that the attention any alternative receives will (weakly) decrease as the number of rivals increases, a nonparametric restriction on the attention rule of decision makers, which we called Attention Overload. If attention was deterministic, our proposed behavioral assumption would simply say that if a product grabs the consumer's consideration in a large supermarket, then it will grab her attention in a small convenience store as there are fewer options \citep*{Reutskaja2009SatisfactionSatiate,Visschers-Hess-Siegrist_2010_PHN,Reutskaja_et_al_2011_AER,Geng_2016_EI}.
Our baseline choice model has two components: a random attention rule and a homogeneous preference ordering, but we later enhance our model to allow for heterogeneous preferences. The random attention rule is the probability distribution on all possible consideration sets. To introduce our attention overload assumption formally, we define the amount of attention a product receives as the frequency it enters the consideration set, termed Attention Frequency. Attention overload then implies that the attention frequency should not increase as the choice set expands. For preferences, we assume that the decision makers have a complete and transitive (initially homogeneous, later heterogeneous) preferences over the alternatives, and that they pick the best alternative in their consideration sets. In this general setting where attention is random and limited, and products compete for attention, we aim to elicit compatible preference orderings and attention frequencies solely from observed choices.
Existing random attention models cannot capture, or are incompatible with, attention overload. For example, Manzini-Mariotti_2014_ECMA consider a parametric attention model with independent consideration where each alternative has a constant attention frequency even when there are more alternatives, and therefore their model does not allow decision makers to be more attentive in smaller decision problems. \citet*{Aguiar_2017_EL} also share the same feature of constant attention frequency. On the other hand, recent research has tried to incorporate menu-dependent attention frequency (Demirkan-Kimya_2020_JME) under the framework of independent consideration. This model is so general that it allows for the opposite behavior than attention overload (i.e., being more attentive in larger choice sets). The recent models of Brady-Rehbeck_2016_ECMA and Cattaneo-Ma-Masatlioglu-Suleymanov_2020_JPE also allow for the possibility that an alternative receives less attention even when the choice set gets smaller. See Section (ref) and the supplemental appendix (Section SA.1) for more discussion on related literature.
We contribute to the decision theory literature by introducing the attention overload nonparametric restriction on the attention frequency, which is the key building block to achieve both preference ordering and attention frequency (point or partial) identification from observed choice data. Our results do not require the attention rule to be observed, nor to satisfy other restrictions beyond those implicitly imposed by attention overload. Since our revealed preference and attention elicitation results are derived from nonparametric restrictions on the consideration set formation, without committing to any particular parametric attention rule, they are more robust to misspecification biases Matzkin_2007_Handbook,Matzkin_2013_ARE,Molinari_2020_HandbookCh.
The fact that attention is not observed poses unique challenges to both identification of and statistical inference on the decision maker's preference. This is because one can only identify (and consistently estimate) the choice probabilities from typical choice data, while our main restriction is imposed on the attention rule. Furthermore, as our attention overload assumption does not require a parametric model of consideration set formation, the set of compatible attention rules is usually quite large. In other words, the attention rule is almost never uniquely identified in our model. We nonetheless show that our attention overload assumption, despite being very general, still delivers nontrivial empirical content: we prove in Section (ref) that a preference ordering is compatible with our attention overload model if and only if the choice probability satisfies a system of inequality constraints, which corresponds to a form of regularity violation. Furthermore, to improve computation and practical implementation, we discuss how to leverage binary choice problems for identification, estimation, and inference.
Besides revealed preference, information about attention frequency is also an object of interest. For example, it enables marketers to gauge the effectiveness of their marketing strategies, or policy-markers to assess whether consumers allocate their attention to better products. Despite the fact that the underlying attention rule may not be identifiable, we show in Section (ref) that our nonparametric attention overload behavioral assumption allows for (point or partial) identification of the attention frequency using standard choice data. This result appears to be the first nonparametric identification result of a relevant feature of an attention rule in the random limited attention literature: revealed attention analysis has not been possible under nonparametric identifying restrictions in prior work.
In Section (ref), we enhance our attention overload model to accommodate multiple decision makers with heterogeneous preference orderings. We begin by providing a general characterization of the partially identified set of distributions on preference orderings and deterministic attention rules. Due to the nonparametric nature of the identifying assumptions, however, the resulting identified set is arguably too large to be useful in practice. Thus, we then further discipline the amount of heterogeneity in the model to (almost) point identify the distribution of preference orderings: we propose the idea of List-Based Attention Overload, where alternatives are presented to the customers as a list that correlates with both heterogeneous preferences and random limited attention. Many real-life situations involve consumers encountering alternatives in the form of a list simon1955behavioral,rubinstein2006model. Our restricted heterogeneous preference model is motivated by the observation that an item's placement on a list has a profound impact on its recollection and evaluation by subjects ellison2009search,augenblick2016ballot,biswas2010order,Levavetal2010. For example, a ranked list of search results provided by a web platform can affect both the search behavior and the perception of individuals about the quality of products Reutskaja_et_al_2011_AER.
Our underlying idea is that a common set of characteristics among the decision makers is taken into account to construct the list. For example, a list emerging from search results can be a good proxy for their preferences if decision makers perceive that the search result reflects the true quality of the listed items westerwick2013effects. Indeed, many commercial websites collect individual consumers’ behavioral data and try to match each consumer with specific products. The list can be thought of as the outcome of personalized recommendations, and therefore individuals facing the same list would share similar tastes. However, individuals might tend to favor their status quo and assign a relatively higher rank to their reference point compared to the rest of the items in the original list. Thus, the existence of the list allows for both heterogeneous preferences and random attention, but restricts the total number of potential preference orderings and attention rules allowed in the choice model.
Attention overload implies that it is often impractical for decision makers to conduct exhaustive searches when many products are on the list. We thus assume that a decision maker investigates alternatives to construct her limited attention consideration set through the list: she might consider only a subset of the alternatives available. Our list-based attention overload model imposes three basic behavioral restrictions on the consideration set formation for a given list: (i) whenever an alternative is considered, all alternatives in the list before it are also taken into account; (ii) if an alternative is not recognized in a smaller set, then it cannot be recognized in a larger set; and (iii) in binary problems, both options are always considered. These assumptions are, for example, supported by eye-tracking studies showing that people tend to scan search engine results in order of appearance, and then fixate on the top-ranked results even if lower-ranked results are more relevant pernice2018people. To capture heterogeneity in cognitive ability, the model accommodates individuals with different consideration sets as long as they satisfy the above behavioral restrictions, thereby allowing for list-based heterogeneity in random limited attention.
To discipline the amount of preference heterogeneity, we also introduce three behavioral axioms characterizing our proposed heterogeneous (preference) attention overload model for a given list. The first axiom captures a restricted form of regularity violation: removing alternatives will not decrease the choice probabilities of a product as long as there is another product listed before it in both decision problems. The second axiom states that binary choice probabilities decrease as the opponent is ranked higher in the list. The last axioms requires that the total binary choice probabilities against the immediate predecessor in the list must be less than or equal to one. We then show that preference and attention frequencies are (point or partially) identifiable under nonparametric assumptions on the list and attention formation mechanisms, even when the true underlying list is unknown to the researcher.
Based on our identification results, covering both homogeneous and heterogeneous preferences settings, we develop econometric methods for revealed preference and attention analysis in both homogeneous and heterogeneous preference settings, which are directly applicable to standard choice data. We only assume that a random sample of choice problems and choices selections is observed, and then provide methods for estimation of and inference on the preference ordering (homogeneous case), or the preferences frequency (heterogeneous case), and the attention frequency of the decision makers. For example, our methods allow for (i) test whether a specific preference ordering is compatible with our attention overload model, (ii) construct (asymptotically) valid confidence sets, (iii) conduct overall model specification testing, and (iv) estimate preference frequencies in heterogeneous settings. For revealed attention, we obtain (point or partial) identification estimates for attention frequencies in both homogeneous and heterogeneous attention overload models. To establish the validity of our econometric methods, we employ the latest results on high-dimensional normal approximation \citep*{Chernozhukov-Chetverikov-Kato-Koike_2022_AOS}. This is crucial because the number of inequality constraints involved in our statistical inference procedures may not be small relative to the sample size. While allowing the dimension (complexity) of the inference problems to be much larger than the sample size, we explicitly characterize the error from a normal approximation for the estimated choice probabilities, thereby shedding light on the finite-sample performance of our proposed econometric methods.
Econometric methods based on revealed preference theory have a long tradition in economics and many other social and behavioral sciences. See \citet*{Matzkin_2007_Handbook,Matzkin_2013_ARE}, \citet*{Molinari_2020_HandbookCh}, and references therein. There is only a handful of recent studies bridging decision theory and econometric methods by connecting discrete choice and limited consideration. Contributions to this new research area include \citet*{Abaluck_Adams_2021}, \citet*{Barseghyan-Coughlin-Molinari-Teitelbaum_2021_ECMA}, \citet*{Barseghyan-Molinari-Thirkettle_2021_AER}, \citet*{Cattaneo-Ma-Masatlioglu-Suleymanov_2020_JPE}, and \citet*{Dardanoni-Manzini-Mariotti-Tyson_2020_ECMA}, among others. Each of these papers imposes different (parametric) identification assumptions on the random consideration and the preferences, producing different levels of identification of preference orderings and consideration set rules. We contribute to this emerging literature with new nonparametric results on the identification of and inference for the preference and attention frequencies when decision makers only pay attention to a subset of possibly too many alternatives at random.
The rest of the paper proceeds as follows. Section (ref) introduces the setup and our key attention overload assumption under homogeneous preferences, and then proves our main characterization result. That section also presents computationally attractive identification results based on binary comparisons, discusses partial identification of the attention frequency, and outlines valid econometric methods for revealed preference and revealed attention analyses with homogeneous preferences. Section (ref) introduces the idea of list-based attention overload to allow for heterogeneous preferences, and presents (point or partial) identification of preference and attention frequencies. That section considers settings where the underlying true list may or may not be known, and also discusses principled econometric methods for estimation and inference using only observed choice data. The appendix contains the proofs of our results, while the supplemental appendix collects (i) in-depth related literature discussion, (ii) simulation evidence, and (iii) omitted technical lemmas and proof details. We also provide a software package and replication code in R implementing our empirical methods.
The theoretical analysis in this section revolves around the assumption that attention frequency is monotonic and preferences are homogeneous. Then, in Section (ref), we enhance our choice model to allow for heterogeneous preferences. We denote the grand alternative set as $X$, and its cardinality by $|X|$. A typical element of $X$ is denoted by $a$. We let $\mathcal{D}$ be a collection of non-empty subsets of $X$ representing the collection of choice problems. In this section, we allow incomplete data where $\mathcal{D}$ is a strict subset of all non-empty subsets of $X$, which makes the model still applicable when there is missing data. A choice rule is a map $\pi: X \times \mathcal{D} \to [0,1]$ such that $\pi(a | S)= 0$ for all $a \notin S$ and $\sum_{a \in S}\pi(a|S)=1$ for all $ S \in \mathcal{D}$. $\pi(a |S)$ represents the probability that the decision maker chooses alternative $a$ from the choice problem $S$. We assume that the choice rule is identifiable from data; that is, it is known or estimable for the purpose of learning about features of the underlying data generating process in econometrics language (Section (ref)).
An important feature of our model is that consideration sets can be random. An attention rule is a map $\mu: 2^X \times \mathcal{D} \to [0,1]$ such that $\mu(T | S)= 0$ for all $T \not \subseteq S$ and $\sum_{T \subseteq S} \mu(T|S)=1$ for all $ S \in \mathcal{D}$. $\mu(T|S)$ represents the probability of paying attention to the consideration set $T \subseteq S $ when the choice problem is $S$. This formulation also allows for deterministic attention rules (e.g., $\mu(S|S)=1$ represents full attention). The choice rule and attention rule are standard features of (rational) choice models with random attention. In this paper, we consider a novel feature of these models that is related to the amount of attention each alternative captures for a given $\mu$. We can extract this information from the attention rule by simply summing up the frequencies of consideration sets containing the alternative.
$\phi_\mu(a|S)$ represents the total probability that $a$ attracts attention in $S$. Whenever $\mu$ is clear from the context, we will omit the subscript $\mu$ to reduce notation. In deterministic attention models, the attention that one alternative receives is either zero or one (i.e., whether it is being considered or not). However, in stochastic environments, attention is probabilistic: this means that the attention one alternative receives may not be binary.
When decision makers are overwhelmed by an abundance of options, every choice alternative competes for attention. This implies that as the number of alternatives increases, the competition gets more fierce: the attention frequency to a product should decrease weakly when the set of available alternatives is expanded by adding more options. We call this property Attention Overload, the novel nonparametric identifying restriction in this paper.
If we allow the consideration set to be empty, then we should also require that the frequency of paying attention to nothing increases when the choice set expands. This is related to the choice overload behavioral phenomenon. At this point, we exclude the possibility of paying attention to nothing for simplicity (i.e., $\mu(\emptyset|S)=0$ for all $ S \in \mathcal{D}$). An attention rule $\mu$ satisfies attention overload if its corresponding attention frequency is monotonic in the sense of Assumption (ref). Section (ref) compares and contrasts Assumption (ref) with other choice models in the literature (see also Section SA.1 in the supplemental appendix).
Given the nonparametric attention overload restriction in Assumption (ref), the choice rule can be defined accordingly. A (rational) decision maker who follows the attention overload choice model maximizes her utility according to a preference ordering $\succ$ under each realized consideration set.
To summarize, the unknown model primitives are the attention rule $\mu$ and the preference ordering $\succ$. We only assume that the choice rule $\pi$ is observable (i.e., point identifiable and estimable from data), and we do not require additional information beyond standard observable choice data. We next investigate the behavioral implications of AOM. Section (ref) shows that it is possible to (point or partially) identify the underlying preference ordering by exploiting the attention overload Assumption (ref), and Section (ref) presents (point or partial) identification results for the attention frequency. Section (ref) builds on those identification results and develops feasible econometric methods. Section (ref) further demonstrates how to incorporate and study heterogeneous preferences in the presence of attention overload.
We first investigate characterization and preference elicitation, as they have a close relationship. We aim to determine whether a data generating process possesses an AOM representation, and if it is feasible to identify preference orderings from observed choice data. To accomplish this, we investigate whether a specific preference ordering can accurately represent the data. This is a challenging task due to several potential issues. First, the data may not have an AOM representation at all. Second, even if an AOM representation exists, the actual preference may differ from the proposed preference ordering. Third, even if the proposed preference aligns with the underlying preference ordering, it is still necessary to construct an attention rule that satisfies attention overload and accurately represents the data. Our first main result addresses these challenges by providing a tight representation without the requirement of constructing an attention rule, which can be a laborious task when there are many alternatives.
AOM has several behavioral implications. Assume that $(\succ,\mu)$ represents $\pi$. Since attention is a requirement for a choice, any choice probability is always bounded above by attention frequency, i.e., $\phi(a|S) \geq \pi(a|S)$. Then, by attention overload, we must have $\phi(a|T)\geq \phi(a|S) \geq \pi(a|S)$ for $T \subseteq S$. In addition, the difference $\phi(a|T) - \pi(a|T)$ captures the probability that $a$ receives attention but is not chosen in $T$. As a consequence, in these cases, a better option must be chosen in $T$, which implies $\phi(a|T) - \pi(a|T) \leq \pi( U_\succ (a)|T)$, where $U_\succ (a)$ denotes the strict upper contour set of $a$ with respect to $\succ$. (With a slight abuse of notation, we set $\pi( U_\succ (a)|T) = \pi( U_\succ (a)\cap T|T) = \sum_{b\in T: b\succ a} \pi(b|T)$.) Combining these observations, we get $ \pi(a|S) \leq \phi(a|S) \leq \phi(a|T) \leq \pi(U_\succeq (a)|T)$, where $U_\succeq (a)$ denotes the upper contour set of $a$ with respect to $\succ$. It follows that $\pi(a|S) \leq \pi(U_\succeq (a)|T)$ whenever $\succ$ represents the data. This condition only refers to preferences, not to the attention rule. Therefore, the following axiom must be satisfied whenever $\succ$ represents the data.
Axiom (ref) applies to the choice rule $\pi$, which is point identifiable and estimable from standard choice data and is stated in terms of a preference ordering $\succ$, a key unobservable primitive of our model. Given $\succ$, it is routine to check whether $\pi$ satisfies $\succ$-Regularity. This axiom is closely related to, but different from, the classical regularity condition. Axiom (ref) trivially implies the regularity condition for the best alternative $a^*$ in $T$, as $U_\succeq (a^*)=\{a^*\}$ and $\pi(U_\succeq (a^*)|T)=\pi(a^*|T) \geq \pi(a^*|S)$. Hence, the full power of regularity is assumed. For other alternatives, the regularity condition is partially relaxed. At the other extreme, $\succ$-Regularity does not restrict the choice probabilities for the worst alternative, $a_*$, since $U_\succeq (a_*)=X$, and hence $\pi(U_\succeq (a_*)|T)=1 \geq \pi(a_*|S)$ for all $T$, implying that $\succ$-Regularity holds trivially. The following result shows that $\succ$-Regularity is not only necessary but also a sufficient condition for $\succ$ to represent the data.
An immediate corollary is that $\pi$ is AOM if and only if there exists $\succ$ such that $\succ$-Regularity is satisfied. We provided above the proof of the necessity of $\succ$-Regularity. The proof of sufficiency, which relies on Farkas's Lemma, is given in the appendix. $\succ$-Regularity informs us whether $\pi$ has an AOM representation with $\succ$. Of course, it is possible that $\succ$-Regularity can be violated for $\succ$ but is satisfied for another preference $\succ'$. Hence, $\succ$-Regularity allows us to identify all possible preference orderings without constructing the underlying attention rule $\mu$.
We now turn to the discussion of revealed preference. Option $b$ is revealed to be preferred to option $a$ if $b \succ a$ for all $\succ$ representing $\pi$. Theorem (ref) suggests that $\succ$-Regularity could be a handy method to identify the underlying preference. We first define a binary relation based on $\succ$-Regularity property:
Again, $\text{P}_\pi$ does not require the construction of all AOM representations. Given a candidate preference, finding the corresponding attention rule satisfying attention overload could be a daunting task. On the other hand, checking whether $\pi$ satisfying $\succ$-Regularity is straightforward because $\succ$-Regularity does not require finding the underlying attention rule. Indeed, we utilize this fact in Section (ref) to develop econometric methods. The next result states the revealed preference of this model.
Although Theorem (ref) and Corollary (ref) bypass the need of constructing the underlying attention rule for revealed preference analysis, checking all possible $|X|!$ preference orderings can be computationally expensive. In addition, if the analyst is interested in learning the preference between two alternatives, $a$ and $b$, then some constraints suggested by $\succ$-Regularity may provide little to no relevant information. Fortunately, a key observation is that regularity violations at binary choice problems can reveal the decision maker's preference. Although this result does not exhaust the nonparametric identification power of Assumption (ref), it can be handy and computationally more attractive.
More specifically, if $a,b\in S$ and $\pi(a|S) > \pi(a|\{a,b\})$, then any $(\succ, \mu)$ representing $\pi$ must rank $b$ above $a$, hence $b$ must be preferred to $a$. To reach such a conclusion, assume the contrary: there exists $(\succ, \mu)$ representing $\pi$ such that $a \succ b $ and $\mu$ satisfying attention overload. First, the attention frequency is always greater (or equal) than the choice probability for any alternative and in any choice set: $\pi(a|S) \leq \phi (a| S)$. In addition, they are equal for the best alternative in any choice set: $a$ is $\succ$-best in $S$ implies $\pi(a|S) = \phi (a| S)$. Given $a\succ b$, we have $\phi (a| \{a,b\}) =\pi(a|\{a,b\}) < \pi(a|S) \leq \phi(a|S)$. This contradicts our attention overload assumption. The next proposition formalizes this observation.
Proposition (ref) provides a guideline to easy-to-implement revealed preference analysis without knowledge about each particular representation. A natural question is whether we can generalize the implication of Proposition (ref) for an arbitrary set $T\subseteq S$ instead of only for binary sets. The answer is not straightforward because identification from regularity violation may not be as sharp when there are more than two alternatives in the smaller set. From Proposition (ref), we are able to claim revealed preference between two alternatives, but when the smaller set contains more than two alternatives, we only know there are some alternatives better than $a$ in the smaller set. To see this, suppose not, and $a$ is the best alternative in the smaller set. We must have $\pi(a|T)=\phi(a|T)$. Given that $\phi(a|T)=\pi(a|T)<\pi(a|S)\leq \phi(a,S)$, it contradicts the attention overload assumption. We put this observation in the following proposition.
Both Proposition (ref) and (ref) are based on regularity violations, and Proposition (ref) implies Proposition (ref) when $T$ is a binary set. The following example demonstrates how these propositions can be used to limit possible preferences orderings first in order to then apply Theorem (ref): regularity violation alone does not exhaust the nonparametric identification power of attention overload in this example.
In the above example, we show that (i) the data has multiple AOM representations, (ii) $a \succ_1 b\succ c \succ_1 d$ and $a \succ_2 c\succ_2 b \succ_2 d$ represent the data, and (iii) since only $\succ_1$ and $\succ_2$ satisfy Axiom (ref), the revealed preference P$_\pi$ is only missing information on $b$ and $c$ (otherwise it is complete). While we had $24$ possible candidates, Proposition (ref) implied only $12$ of them were viable candidates. Then, Proposition (ref) eliminated $8$ of the remaining ones. Finally, only two satisfied Axiom (ref). Hence, Example (ref) demonstrates how our main results can be used constructively to identify the set of plausible preferences, while also substantially reducing the computational burden.
We established how revealed preferences analysis can be done with our nonparametric attention overload assumption. Due to the nature of attention overload, one might suspect that the decision makers are more likely to pay full attention when there are only two alternatives. We now assume that extreme limited attention (i.e., considering only a single option) at binaries cannot exceed a preset probability level, and investigate its implications for revealed preference.
Assume that for any $\eta \geq 0.5$,
This condition puts an upper limit on the magnitude of extreme limited attention. As $\eta$ increases, the probability of limited attention can increase: $\eta=1$ imposes no constraint on attention behavior. This assumption does not impose a lower bound for full attention. Even in the extreme case, where $\eta$ is equal to $0.5$, it is still possible that there is no full attention ($\mu(\{a,b\}|\{a,b\})=0$). Hence, it is not a demanding condition when compared to full attention at binaries.
Condition (ref) can generate additional revealed preferences: $a \text{P}_B b$ if $\pi(a|\{a,b\})>\eta$, for any $\eta \geq 0.5$. Whenever we observe $\pi(a|\{a,b\})>\eta$, the choice probability of $a$ cannot be entirely attributed to the attention on the singleton set, $\mu(\{a\}|\{a,b\}) \leq \eta$. Then, we must have $\mu(\{a,b\}|\{a,b\})>0$ and the decision maker chooses $a$ over $b$ when she pays attention to both alternatives. Hence, it implies that $a$ must be better than $b$.
We could also interpret the parameter $\eta$ as a measure of how cautious the policy-maker is when making a welfare judgment. If $\eta=1$, the policy-maker would not draw any conclusion from binary comparisons only ($\text{P}_B=\emptyset$). The choice $\eta=0.5$ is commonly used in the literature marschak1959binary,fishburn1998stochastic, which would refer to the largest $\text{P}_B$ in our setup---almost uniquely identified.
Condition (ref) provides additional revealed preference information if the data on binary comparisons are available. Under (ref), the revealed preferences of our model must include $\text{P}_B$. We can then extend our characterization theorem: if $\pi$ satisfies $\succ$-Regularity where $\succ$ includes $\text{P}_B$, then the data has an AOM representation with $\mu$ satisfying (ref). More importantly, $\text{P}_B$ improves the result of Proposition (ref) (and hence Proposition (ref)) by restricting the set of plausible preferences. We revisit Example (ref) to illustrate this point.
Our attention overload model builds on the simple nonparametric requirement that each alternative gets weakly less attention in bigger choice problems, which is captured by monotonicity in attention frequency (Assumption (ref)). Given a dataset, one might want to learn how the attention frequency changes across different alternatives and choice problems. For example, marketers might want to gauge the effectiveness of their marketing strategies, or policy markers could be interested in assessing whether consumers allocate their attention to better products. Since we do not put any restriction on the attention rule, the attention frequency can vary depending on the actual attention rule that the decision maker has. This section shows that it is possible to develop upper and lower bounds for the attention frequency and thus achieve partial identification of $\phi$.
Consider bounding $\phi$ from below first. For any superset $R\supseteq S$, the attention overload assumption implies that $\pi(a|R)\leq \phi(a|R)\leq \phi(a|S)$. Therefore, for any $S$, $\phi(a|S) \geq \max_{R\supseteq S} \pi(a|R)$. This lower bound on the attention frequency only uses information from the choice rule, which is estimable from standard choice data. Importantly, this lower bound does not require a particular AOM representation, that is, it does not require knowledge of the underlying attention rule. It is also possible to derive an upper bound for $\phi$, although in this case the bound will depend on the preference ordering. Consider a preference $\succ$ and an attention rule $\mu$ satisfying attention overload, so that $\pi$ is an AOM with $(\succ,\mu)$. Then, for any subset $T\subseteq S$, $\phi(a|S)\leq \phi(a|T) \leq \pi(U_\succeq (a)|T)$, which implies that $\phi(a|S) \leq \min_{T\subseteq S} \pi(U_\succeq (a)|T)$. These observations give the following theorem.
We now consider three extreme cases of Theorem (ref). If both the lower bound and the upper bound are $1$, then $a$ attracts full attention at $S$ (Revealed Full Attention). If both bounds are zero, then $a$ does not attract any attention at $S$ (Revealed Inattention). The third case happens when the lower bound is zero and the upper bound is one (No Revealed Attention). Indeed, these three cases are the only possibilities when the data is deterministic, which was studied by \citet*{Lleras_et_al_2017_JET}. However, they did not provide any characterization result for revealed attention. Theorem (ref) provides such characterization not only for stochastic choice but also for its deterministic counterpart, and hence our theorem is also a novel contribution in the competing attention framework for deterministic choice theory.
Since the stochastic data is richer, Theorem (ref) covers another interesting case, which we call Partial Revealed Attention: the upper bound is strictly below one and/or the lower bound is strictly above zero. To illustrate revealed attention, we revisit Example (ref).
In some cases, the attention frequency will be uniquely identified for certain alternatives. For instance, in addition to Example 1, we have $\pi(c|R)=0.25$ for some $R\supseteq \{a,b,c,d\}$. Then, the lower bound for $\phi(c|\{a,b,c,d\})$, while being free of the underlying preference, must be $0.25$. Hence, the attention frequency is point-identified to be $0.25$ since the upper bound is also $0.25$.
Theorem (ref) is useful in real world applications to inform a firm/government how much attention each product/policy receives among other options. While the lower bound can be interpreted as the pessimistic evaluation for attention, the upper bound captures optimistic evaluation. The question is whether these local pessimistic (optimistic) evaluations hold globally, that is, we ask whether there is an underlying attention rule $\mu$ satisfying attention overload such that the attention frequencies agree with the pessimistic (optimistic) evaluations for every set. Due to the richness in attention rule allowed by our Assumption (ref), it turns out that the answer is affirmative.
This theorem concerns the pessimistic evaluation case, but an analogous result can be established for the optimistic evaluation. As a consequence, in econometrics language, Theorem (ref) delivers the sharp identified set for $\phi$.
We compare AOM to other existing (random) attention models: tversky1972elimination, Manzini-Mariotti_2014_ECMA, Brady-Rehbeck_2016_ECMA, Aguiar_2017_EL, Cattaneo-Ma-Masatlioglu-Suleymanov_2020_JPE, and Demirkan-Kimya_2020_JME. With the exception of Cattaneo-Ma-Masatlioglu-Suleymanov_2020_JPE, which imposes a nonparametric restriction, all the other models introduce the idea of random limited attention with a parametric restriction on the attention rule. We show that none of these models can capture the attention overload assumption by comparing their underlying attention rules. Section SA.1 in the supplemental appendix provides further comparisons between AOM and the related literature.
Consider two individuals, Ann and Ben. In a larger decision problem, $S$, Ann pays attention to all alternatives with probability one (full attention, $\mu_{\text{Ann}}(S|S)=1$), while Ben experiences attention overload and focuses only on a single alternative $a$ in $S$ while ignoring the rest (limited attention, $\mu_{\text{Ben}}(\{a\}|S)=1$). We chose these two extreme cases to make our point clear. Existing evidence suggests that, as the size of the available options decreases, the phenomenon of choice overload becomes less evident, leading decision makers to overlook less alternatives. Hence, assuming $|T|<|S|$, Ann continues to exhibit full attention in $T$ but Ben considers more alternatives in $T$. Table (ref) summarizes the comparisons with the literature using these two decision makers.
The Independence Attention Model (IAM) of Manzini-Mariotti_2014_ECMA and “Elimination by Aspects” Model (EAM) tversky1972elimination,Aguiar_2017_EL make the same predictions for Ann and Ben. Manzini-Mariotti_2014_ECMA considers a parametric model of limited attention where each alternative has a constant attention frequency. Full attention on the larger selection implies that the attention frequency for each alternative in $S$ is one. (Some of the parametric models require the attention parameters be strictly between zero and one, but we can capture these examples either by allowing the parameters to be equal to zero and one, or by taking a limit.) “Elimination by Aspects” attention is an adaptation of tversky1972elimination into limited attention, where alternatives are exogeneously bundled into categories, and the decision maker considers each of these categories with certain probabilities. Aguiar_2017_EL characterizes a special case of this model with the default option where each category includes the default option. For Ann, both models make the same prediction, which is consistent with attention overload: Ann should pay full attention in $T$. However, these models do not allow Ben to be more attentive for smaller sets: Ben must pay attention to the singleton $\{a\}$ with probability one. In IAM, this is because the attention frequencies for other alternatives in $S$ are zero, while in EAM, $a$ is the only alternative belonging to the most popular category. Hence, Manzini-Mariotti_2014_ECMA and tversky1972elimination,Aguiar_2017_EL are too restrictive to accommodate attention overload. Furthermore, Demirkan-Kimya_2020_JME drops the menu-independence assumption in IAM, which leads to no restriction on the attention rule for $T$, and hence cannot accommodate ateention overload either.
The Logit Attention Model (LAM) of Brady-Rehbeck_2016_ECMA and the Random Attention Model (RAM) of Cattaneo-Ma-Masatlioglu-Suleymanov_2020_JPE make the same prediction for our individuals. LAM could be interpreted as a parametric limited attention model where each subset could be the consideration set with some probability. Since Ann exhibits full attention in the larger set, $S$ must be the most probable consideration set, and since $S$ is not a subset of $T$, Ann is allowed to focus on any subset of $T$, including a single alternative. For Ben, this model implies Ben must continue to pay attention only to $a$. RAM imposes a (nonparametric) monotonicity assumption on the attention rule, so that $\mu(T|S) \leq \mu(T| S\setminus b)$ for every $b \notin T\subseteq S$. In RAM, the monotonic attention does not impose any restriction on Ann's behavior (for example, it could be $\mu_{\text{Ann}}(\{a\}|T)=1$), but Ben must pay attention to the same single alternative ($\mu_{\text{Ben}}(\{a\}|T)=1$). In other words, Ann can exhibit the opposite of choice overload, while Ben must exhibit limited attention even in a smaller problem. Therefore, these two models stand in contrast to the concept of attention overload.
In contrast to the existing literature, our AOM introduces a novel nonparametric assumption on attention frequency, which measures how much attention each alternative (rather than each consideration set) receives. That is, $\phi(b|S) \leq \phi(b|S\setminus c)$ for every $b \in S$. If Ann behaves according to AOM, $\mu_{\text{Ann}}(S|S)=1$ implies $\phi_{\text{Ann}}(b|S)=1$ for all $b$, which dictates $\phi_{\text{Ann}}(b|T)=1$ for all $b\in T$ (full attention). For Ben, AOM imposes that he must consider $a$ for sure ($\phi_{\text{Ben}}(a|T)=1$), but attention frequency for other alternatives can increase. In particular, AOM allows that $\phi_{\text{Ben}}(b|T)=1$ for all $b\in T$ (full attention). This discussion makes it clear that our attention overload property is distinct from all other (parametric or nonparametric) random limited attention models. As a consequence, our AOM captures novel empirical findings and describes novel attention allocation behaviors compared to the existing models in the literature.
We obtained several testable implications and related results for the AOM: Theorem (ref), Propositions (ref) and (ref), and Theorem (ref). Our next goal is to develop econometric methods to implement these findings using real data, which can help elicit preferences, conduct empirical testing of our AOM, and provide confidence sets for attention frequencies. To this end, we rely on a random sample of observations consisting of choice data for $n$ units indexed by $i=1,2,\dots,n$. Each unit faces a choice problem $Y_i$, and her choice is denoted by $y_i\in Y_i$. This is formally stated in the assumption below. We recall that $\mathcal{D}\subseteq 2^X\setminus \emptyset$ is a collection of choice problems.
A given preference ordering $\succ$ is compatible with our AOM if and only if $\succ$-Regularity holds. In other words, each preference ordering corresponds to a collection of inequality constraints, which we collect in the following hypothesis:
To construct a test statistic for testing the hypotheses in (ref), we can replace the unknown choice probabilities by their estimates, $\widehat{\pi}(a|S)$ and $\widehat{\pi}(U_\succeq (a)|T)$, leading to $\widehat{D}(a|S,T) = \widehat{\pi}(a|S) - \widehat{\pi}(U_\succeq (a)|T)$. For example, $\widehat{\pi}(a|S) = \frac{1}{N_S}\sum_{i=1}^n \mathbbm{1}(Y_i=S,y_i=a)$, with $N_S=\sum_{i=1}^n \mathbbm{1}(Y_i=S)$. We define the following test statistic \[\mathsf{T}(\succ) = \max\left\{\max_{{a\in T\subset S;\ T,S\in\mathcal{D}}}{\widehat{D}(a|S,T)}/{\widehat{\sigma}(a|S,T)}\ ,\ 0\right\}, \] where $\widehat\sigma(a|S,T)$ is the standard error of $\widehat{D}(a|S,T)$, that is, $\widehat{\sigma}^2(a|S,T) = \frac{1}{N_S}\widehat\pi(a|S)\big(1-\widehat\pi(a|S)\big) + \frac{1}{N_T}\widehat\pi(U_\succeq (a)|T)\big(1-\widehat\pi(U_\succeq (a)|T)\big). $ The outer $\max$ operation in $\mathsf{T}(\succ)$ guarantees that we will never reject the null hypothesis if none of the estimated differences $\widehat{D}(a|S,T)$ are strictly positive. In other words, a preference is not ruled out by our analysis if $\succ$-Regularity holds in the sample. The statistic depends on a specific preference ordering which we would like to test against for: such dependence is explicitly reflected by the notation $\mathsf{T}(\succ)$.
We investigate the statistical properties of the test statistic in order to construct valid inference procedures, building on the recent work of \citet*{Chernozhukov-Chetverikov-Kato_2019_RESTUD} and \citet*{Chernozhukov-Chetverikov-Kato-Koike_2022_AOS} for many moment inequality testing. (See also \citet*{Molinari_2020_HandbookCh} for an overview and further references.) We seek for a critical value, denoted by $\mathrm{cv}(\alpha,\succ)$, such that under the null hypothesis (i.e., when the preference is compatible with our AOM), $\mathbb{P}\left[ \mathsf{T}(\succ) > \mathrm{cv}(\alpha,\succ) \right] \leq \alpha + \mathfrak{r}_{\succ}$, where $\alpha\in(0,1)$ denotes the desired significance level of the test, and $\mathfrak{r}_{\succ}$ denotes a quantifiable error of approximation (which should vanish in large samples with possibly many inequalities).
To provide some intuition on the critical value construction, the Studentized test statistic, ${\widehat{D}(a|S,T)}/{\widehat{\sigma}(a|S,T)}$, is approximately normally distributed with mean ${D(a|S,T)}/{\sigma(a|S,T)}$. Since $D(a|S,T)\leq 0$ under the null hypothesis, the above normal distribution will be first-order stochastically dominated by the standard normal distribution. Letting $\widehat{\mathbf{D}}$ be the column vector collecting all $\widehat{D}(a|S,T)$, and $\boldsymbol{\Omega}$ be its correlation matrix, then our test statistic $\mathsf{T}(\succ)$ will be dominated by the maximum of a normal vector with a zero mean and a variance of $\boldsymbol{\Omega}$, up to the error from normal approximation. Using properties of Bernoulli random variables, an estimate of $\boldsymbol{\Omega}$ can be constructed with the estimated choice probabilities and the effective sample sizes. We denote the estimated correlation matrix by $\widehat{\boldsymbol{\Omega}}$. The critical value is then the $(1-\alpha)$-quantile of the maximum of a Gaussian vector: $\mathrm{cv}(\alpha,\succ) = \inf\{ t : \mathbb{P}[\mathsf{T}^{\mathtt{G}}(\succ) \leq t|\text{Data}]\geq 1-\alpha \}$ with $\mathsf{T}^{\mathtt{G}}(\succ) = \max\{\max(\widehat{\boldsymbol{\Omega}}^{1/2}\mathbf{z} )\ ,\ 0\}$, where the inner $\max$ operation computes the maximum over the elements of $\widehat{\boldsymbol{\Omega}}^{1/2}\mathbf{z}$, and $\mathbf{z}$ denotes a standard normal random vector of suitable dimension. Precise definitions and omitted details are given in the supplemental appendix to conserve space. The theorem below offers formal statistical guarantee on the validity of our proposed test.
This theorem shows that the error in distributional approximation, $\mathfrak{r}_{\succ}$, only depends on the dimension of the problem (i.e., $\mathfrak{c}_1$) logarithmically, and therefore our estimation and inference procedures remain valid even if the test statistic $\mathsf{T}(\succ)$ involves comparing “many” pairs of estimated choice probabilities. By providing non-asymptotic statistical guarantees, our procedures can accommodate situations where both the number of alternatives and the number of choice problems are large, and hence they are expected to perform well in finite samples, leading to more robust welfare analysis results and policy recommendations.
It is routine to incorporate condition ((ref)) into our econometric implementation. Specifically, the test statistic and the critical value we introduced above are based on moment inequality testing. To accommodate the new assumption on attentive at binaries, one only needs to include additional probability comparisons corresponding to the $\eta$-constrained revealed preference.
Given the testing procedures we developed, it is easy to construct valid confidence sets by test inversion. To be precise, a dual asymptotically valid $100(1-\alpha)\%$ level confidence set is $\mathsf{CS}(1-\alpha) = \{\succ:\ \mathsf{T}(\succ) \leq \mathrm{cv}(\alpha,\succ)\}$. Therefore, for any preference $\succ$ that is compatible with our AOM, we have the statistical guarantee on coverage: $ \mathbb{P}\left[ \succ\ \in \mathsf{CS}(1-\alpha) \right] \geq 1 - \alpha - \mathfrak{r}_{\succ}.$
Revealed preference between two alternatives, $a$ and $b$, can be analyzed based on Theorem (ref) by checking, say, if $a\succ b$ for all identified preferences $\succ$ in $ \mathsf{CS}(1-\alpha)$. This approach to revealed preference has the advantage that it exhausts the identification power of our AOM. On the other hand, it comes with a nontrivial computational cost as discussed in Section (ref). Therefore, we also discuss how Proposition (ref) can be implemented in practice. The econometric methods stemming from Theorem (ref) and Proposition (ref) are complementary: while the former provides a more systematic framework for preference revelation, the latter can be handy if binary comparisons are available in the data or if the analyst is particularly interested in inferring preference ordering among pairs of alternatives. (Also see Example (ref), which demonstrates that applying Proposition (ref) first to a choice data may greatly reduce the number of preference orderings to be tested against $\succ$-Regularity.)
We fix two alternatives, say $a$ and $b$, and let $\mathcal{D}_{ab}$ be the collection of choice problems containing both $a$ and $b$, excluding the binary comparison; that is, $\mathcal{D}_{ab} = \{ S\in \mathcal{D}:\ S \supsetneq \{a,b\}\}$. We also recall the simplified notation $\pi(a|b) = \pi(a|\{a,b\})$ for binary comparisons. Then, we may deduce the preference ordering $b\succ a$ if we are able to reject $\mathsf{H}_0: \max_{S\in\mathcal{D}_{ab}} D(a|S,b) \leq 0$, where we define $D(a|S,b) = \pi(a|S) - \pi(a|b)$. Constructing a test statistic is straightforward:
where $\widehat{D}(a|S,b) = \widehat{\pi}(a|S) - \widehat{\pi}(a|b)$, and $\widehat{\sigma}(a|S,b)$ is its standard error. We employ the same technique to construct a critical value. Letting $\widehat{\mathbf{D}}$ be the column vector collecting all $\widehat{D}(a|S,b)$, and $\widehat{\boldsymbol{\Omega}}$ be its estimated correlation matrix. The critical value is $\mathrm{cv}(\alpha,ab) = \inf\{ t : \mathbb{P}[\mathsf{T}^{\mathtt{G}}(ab) \leq t|\text{Data}]\geq 1-\alpha \}$ with $\mathsf{T}^{\mathtt{G}}(ab) = \max\{\max(\widehat{\boldsymbol{\Omega}}^{1/2}\mathbf{z} )\ ,\ 0\}$. The next result follows from Theorem (ref).
One of our main econometric contributions is a careful study of the properties of the estimated correlation matrix, $\widehat{\boldsymbol{\Omega}}$, where we provide an explicit bound on the supremum of the entry-wise estimation error $\Vert \widehat{\boldsymbol{\Omega}} - {\boldsymbol{\Omega}}\Vert_{\infty}$. This is further combined with the results in \citet*{Chernozhukov-Chetverikov-Kato-Koike_2022_AOS} to establish a normal approximation for the centered and normalized inequality constraints, say $(\widehat{D}(a|S,T) - D(a|S,T))/\widehat{\sigma}(a|S,T)$. Cattaneo-Ma-Masatlioglu-Suleymanov_2020_JPE considered preference elicitation under a monotonic attention assumption, and proposed estimation and inference procedures based on pairwise comparison of choice probabilities as in (ref). However, their econometric analysis assumed “fixed dimension,” and hence did not allow the complexity of the problem to be “large” relative to the sample size, which is required in this paper.
We now discuss how to operationalize the partial identification result in Theorem (ref) on attention frequency. We will illustrate with the lower bound, $\phi(a|S)\geq \max_{R\supseteq S} \pi(a|R)$, since the upper bound follows analogously. A na\"{i}ve implementation would replace the unknown choice probabilities by their estimates. Unfortunately, the uncertainty in the estimated choice probabilities will be amplified by the maximum operator, leading to over-estimated lower bounds. Our aim is to provide a construction of the lower bound, denoted as $\underline{\widehat{\phi}}(a|S)$ such that $\mathbb{P}[ \phi(a|S) \geq \underline{\widehat{\phi}}(a|S) ] \geq 1 - \alpha + \mathfrak{r}_{\underline{\phi}(a|S)}$, with $\alpha\in(0,1)$ denoting the desired significance level, and $\mathfrak{r}_{{\underline{\phi}}(a|S)}$ denoting the error in approximation, which should become smaller as the sample size increases.
Our construction is based on computing the maximum over a collection of adjusted empirical choice probabilities. To be precise, we define
where $\widehat{\sigma}(a|R) = \sqrt{\widehat{\pi}(a|R)(1-\widehat{\pi}(a|R))/N_R}$ is the standard error of the estimated probability $\widehat{\pi}(a|R)$, and $\mathrm{cv}(\alpha,\underline{{\phi}}(a|S)) = \inf\{ t:\ \mathbb{P}[ \max(\mathbf{z}) \leq t ] \geq 1-\alpha \}$, with $\mathbf{z}$ is a standard normal random vector of dimension $|\{R\in\mathcal{D}:R\supseteq S\}|$, the number of supersets of $S$.
To provide some intuition for the construction, we begin with the normal approximation: $\widehat{\pi}(a|R) \overset{\mathrm{a}}{\thicksim} \text{Normal}(\pi(a|R),\ \sigma(a|R))$. Then, the estimated choice probabilities are mutually independent since they are constructed from different subsamples, and
where $\approx$ denotes an approximation in large samples. Heuristically, the above demonstrates that with high probability (approximately $1-\alpha$) the true choice probabilities, $\pi(a|R)$, are bounded from below by $\widehat{\pi}(a|R) - \mathrm{cv}(\alpha,\underline{{\phi}}(a|S))\cdot\widehat{\sigma}(a|R)$. This is made possible by the adjustment term we added to the estimated choices probabilities. This idea is formalized in the following theorem, which offers precise probability guarantees.
Due to space limitations, the supplemental appendix reports simulation evidence showcasing the empirical performance of our theoretical and methodological econometric results. To be more precise, we test against four preference orderings using simulated data, and in this process, we vary the number of choice problems available ($\mathcal{D}$) in the data and the effective sample size of each choice problem ($N_S$). Overall, our procedure performs well: it is able to reject preference orderings that are not compatible with $\succ$-Regularity with nontrivial power, while at the same time maintaining size.
The analysis so far considered homogeneous preferences, and hence our choice model assumed that every decision maker has the same taste but different levels of attentiveness. In this section, we present a model that describes the choice behavior at individual and population levels allowing for preference heterogeneity in the underlying data-generating process as well as attention heterogeneity. Our heterogeneous preference choice model can be used empirically with aggregate data on a group of distinct decision makers where each may differ not only on what they pay attention to but also on what they prefer. The model also allows heterogeneous preferences that may correlate with attention, full independence being a special case.
We consider a group of individuals who differ not only in how they pay attention but also in their preferences. We focus on individuals with deterministic choices. To do so, we first define a deterministic attention rule satisfying the attention overload property. An attention rule $\mu$ is deterministic if $\mu(T|S)$ is either $0$ or $1$. Since $\mu(\cdot|S)$ is a probability distribution, there exists a unique subset of $S$, say $T$, such that $\mu(T|S)=1$. Hence, we can define a mapping $\Gamma$ from $2^X$ to $2^X \setminus \emptyset$ so that $\Gamma(S)=T$ where $\mu(T|S)=1$. $\Gamma(S)$ represents the consideration set under $S$. Our Assumption (ref) implies the following condition for the consideration set mapping $\Gamma$: If an alternative is recognized in a larger set, it is also recognized in a smaller one; this condition first appeared in Lleras_et_al_2017_JET. Formally, if $a \in \Gamma(S)$ and $a \in T \subset S$, then $a \in \Gamma(T)$. Indeed, this property is equivalent to Assumption (ref) in the class of the deterministic attention rules.\footnote{When $\mu$ is deterministic, then $\phi_{\mu}(a|S)$ is $1$ if $a \in \Gamma(S)$, otherwise $0$. If $\phi_{\mu}(a|S)=0$, there is no need to check. If $\phi_{\mu}(a|S)=1$, then $a$ must be in $\Gamma(S)$, which implies that $a$ must be in $\Gamma(T)$. Hence $\phi_{\mu}(a|T)=1$ satisfying Assumption (ref).} We denote all consideration sets satisfying Assumption (ref) by $\mathcal{AO}$.
We are now ready to define our model allowing heterogeneous preferences. Let $\mathcal{P}$ denote the collection of all preferences. Consider a population of individuals where each individual is endowed with two primitives; a preference ordering $\succ \in \mathcal{P} $ and a deterministic consideration set mapping $\Gamma \in \mathcal{AO}$. Each pair $(\Gamma,\succ)$ represents a choice type in the population, whose choices are deterministic. Let $\tau$ be a probability distribution on $\mathcal{AO} \times \mathcal{P}$, so $\tau(\Gamma,\succ)$ is the probability of $(\Gamma,\succ)$ being the choice type.\footnote{One could imagine a more general model where each type is a pair of $(\mu, \succ)$ where $\mu$ is a stochastic attention rule satisfying Assumption (ref). Then $\tau$ is a probability distribution over $(\mu, \succ)$. As we elaborate below, this general model would suffer from non-uniqueness even more.} Each $\tau$ naturally induces a probabilistic choice function:
The “rational” types are allowed to be on the support of $\tau$ (take $\Gamma(S)=S$). On the other hand, there are many other “non-rational” choice types. Hence, one might wonder whether this model has any prediction power. It is routine to show this model has empirical content when the number of alternatives is greater or equal to $4$. Using an idea from mcfadden1990stochastic, we are able to provide a full characterization for the model.
This theorem provides a method for falsifying the model: a single violated inequality suffices to indicate that the data cannot be accurately represented by this model.\footnote{Demonstrating sufficiency, unfortunately, involves an infinite number of linear inequalities. This is also true for the Axiom of Revealed Stochastic Preference (ARSP) of mcfadden1990stochastic, for which kitamura2018nonparametric develop empirically tests thereof.} However, identifying $\tau$ within this model is highly unsatisfactory. For instance, it is possible to construct multiple representations $\{\tau_i\}$ such that (i) they all represent the same choice behavior, (ii) they all differ from each other, and (iii) their supports are disjoint. Hence, there is no hope for point identification, and little hope for informative partial identification results; even partial identification does not help constrain the parameters to lie in a strict subset of the parameter space. This complexity arises from two sources of non-uniqueness. Firstly, even under full attention, the classical Random Utility Model (RUM) is known to suffer from non-uniqueness due to varying tastes fishburn1998stochastic,turansick2022identification. Secondly, the issue is exacerbated by the consideration set mapping, which can lead to different consideration sets representing the same deterministic choice, even with fixed preferences. Consequently, without additional structure, developing useful identification results for $\tau$ at the present level of generality is arguably a hopeless exercise.
To achieve point identification with preference heterogenity, we impose further nonparametric restriction on the heterogeneity in preference types and possible consideration sets to discipline our proposed choice model: our key idea is that alternatives are presented to the decision maker as a list that correlates with both heterogeneous preferences and random (limited) attention.\footnote{In decision theory, rubinstein2006model is the seminal paper introducing the idea that decision makers encounter the alternatives in the form of a list. This idea has been influential since then (see for example, horan2010sequential,guney2014theory,Yildiz2016TE,aguiar2016satisficing,kovach2020satisficing,Ishiietal_2021_JME,tserenjigmid2021order,YEGANE2022,KOSHEVOY2023,manzini2024_approval).} Our modeling strategy that individuals come across various options presented as a list is motivated by real-world scenarios. For example, customers often browse Amazon's ordered search results or receive a ranked list of advertisements, not to mention decision makers employing web search engines.
We first assume that the list is observable. This assumption may be reasonable in situations where we can observe Amazon's product list for each product category, a ballot for a specific election, Google's search results for a keyword, etc. In Appendix (ref), we relax this assumption and endogenize the list, allowing for the identification of heterogeneous preferences when the true underlying list is unknown to the researcher. The supplementary appendix (Section SA-1) gives a review of the related literature on limited attention and heterogeneous preferences, and of choice theory over a list.
In this subsection, we propose a model that accommodates varying tastes and attention mechanisms with point identification. Formally, we assume there is a list of items represented by the linear order $\triangleright$. Let $\langle a_1,a_2,\dots,a_{|X|}\rangle$ be the enumeration of the elements in $X$ with respect to $\triangleright$, where $a_j$ denotes the item in the $j$th position, and $|X|$ is the size of the grand set $X$. In other words, $j < k$ implies that $a_j$ appears earlier in the list than $a_k$, which is equivalent to saying $a_j \triangleright a_k$. We will use both notation, $j < k$ and $a_j \triangleright a_k$, interchangeably. For a choice problem $S\subseteq X$, we also enumerate its elements as $\langle a_{s_1},a_{s_2},\dots,a_{s_{|S|}}\rangle$.
It is often impractical for consumers to conduct exhaustive searches because their attention is limited. Through the list, a decision maker investigates alternatives to construct her consideration set but she might consider only a subset of the alternatives available to her due to limited attention. Our proposed model will impose three basic behavioral restrictions on the consideration set formation for a given list. Let $\Gamma(S)$ be the consideration set when the choice problem is $S$. First, we assume that the consideration set obeys the underlying order: if $a_k$ belongs to the consideration set, so does every feasible alternative that appeared before $a_k$, i.e., $a_k \in \Gamma(S)$ and $a_j \in S$ such that $j <k$ imply $a_j \in \Gamma(S)$. In other words, whenever an alternative is considered, all alternatives in the list before it are also taken into account. Second, we assume that alternatives are the deterministic counterpart of AOM as discussed before. Finally, we assume that each individual always considers both items in binary problems, i.e, $\Gamma(S)=S$ whenever $|S|= 2$.
We denote $\mathcal{AO}_{\triangleright}$ as the set of all consideration set rules satisfying list-based attention overload with respect to $\triangleright$.
Individuals are also heterogeneous in terms of their preferences. Unlike RUM, we assume that the set of preferences is related to the underlying list. First, our model recognizes that some individuals perceive search results as reflecting the true quality of listed items. Indeed, many commercial websites collect individual consumers’ behavioral data and try to match each consumer with personally relevant products. The list can be thought of as the outcome of personalized recommendations. Individuals facing the same list share similar tastes. However, our model also captures the idea that individuals might favor their status quo, meaning that they assign a relatively higher rank to their reference point compared to other items in the original list. This assumption restricts potential preferences exhibited in the model. For all $j<k$, define $\succ_{k j}$ as a linear order where the $k$th alternative in $\triangleright$ is moved to the $j$th position. To give some examples, $\succ_{21}$ corresponds to the ordering $\langle a_2,a_1,a_3,a_4,\dots,a_{|X|}\rangle$, and $\succ_{42}$ is $\langle a_1,a_4,a_2,a_3,\dots,a_{|X|}\rangle$. We call $\succ_{k j}$ a single improvement of $\triangleright$, and we use $\mathcal{P}_{\triangleright}$ to denote the set of all single improvements of $\triangleright$ including $\triangleright$ itself. For notational convenience, we let $\succ_{kk}=\triangleright$. If $\succ \in \mathcal{P}_{\triangleright} $, this implies that there exists a unique alternative $a_k $ such that its relative ranking improved with respect to $\triangleright$. Orderings involving multiple changes to the original list order $\triangleright$, such as $\langle a_2,a_1,a_4,a_3,\dots,a_{|X|}\rangle$, are not allowed.
Equipped with $\mathcal{AO}_{\triangleright} \times \mathcal{P}_{\triangleright}$, we state the precise definition of the model.
HAOM$_{\triangleright}$ introduces heterogeneity both in terms of preferences and attention. This feature makes this model independent of RUM. (If the support of $\tau$ consists of only choice types with $\Gamma(S)=S$, then the model becomes a special case of RUM.) Due to limited attention, HAOM$_{\triangleright}$ allows choice types outside of the preference maximization paradigm, and therefore it captures behaviors outside of RUM. On the other hand, since RUM allows more preference types, some choice behaviors can be only captured by RUM but not by HAOM$_{\triangleright}$. Having said that, HAOM$_{\triangleright}$ still encompasses more choice types than RUM, even though we restrict the set of possible preferences. This is because HAOM$_{\triangleright}$ considers two types of heterogeneity (attention and preferences), while RUM only allows for heterogeneity in preference ordering. Our AOM can be regarded as another extreme of HAOM$_{\triangleright}$, as it requires that every choice type has the same preferences but it imposes a mild nonparametric restriction on attention. In the supplemental appendix we discuss further the relationship between AOM and HAOM, and their connections with prior literature.
We now provide a list of behavioral postulates describing the implications of HAOM$_{\triangleright}$ for a given list $\triangleright$. Since the model is more involved in terms of types, the following results require that $\mathcal{D}$ includes all non-empty subsets of $X$, which is a common assumption on choice models that involve random utility.
We first highlight that HAOM$_{\triangleright}$ allows violations of the regularity condition, while regularity always holds in RUM. However, HAOM$_{\triangleright}$ limits the types of regularity violations that are permissible. For example, there will be no regularity violation as long as the first alternative in the list is always present. This is because only choice types $\{\Gamma:a_k\in \Gamma(S)\} \times \{\succ_{k1}\}$ will pick $a_k$ in the presence of $a_1$ (assuming $k> 1$), and they will continue choosing $a_k$ in a smaller choice problem $T\subset S$. Indeed, we can generalize this intuition: removing alternatives will not decrease the choice probabilities of a product as long as there is another product listed before it in both decision problems. Consider two products $a_k$ and $a_j$ such that $j<k$, then the choice probability of $a_k$ obeys regularity conditions, i.e., $\pi(a_k|T) \geq \pi(a_k|S)$ for $T\subset S$ provided that $a_j \in T$. This is our first behavioral postulate for HAOM$_{\triangleright}$.
The following property imposes a structure on binary choices. It simply says everything else equal, being listed earlier increases choice probabilities on binary comparisons. For example, compared to the third product in the list, the first product is chosen more frequently than the second one. In other words, binary choice probabilities decrease as the opponent is ranked higher in the list. For example, $a_3$ is going to be chosen against $a_2$ more often than against $a_1$. This is because the individual considers both alternatives in every binary comparison. Hence, being listed earlier is reflected in choice probabilities. We adopt the following notation for binary comparisons: $\pi(a_k|a_{\ell}):=\pi(a_k|\{a_k,a_{\ell}\})$.
Again consider binary comparisons. Assume we have $\pi(a_2|a_1)=0.6$. This implies that the frequency of $\succ_{21}$ must be $0.6$, and those types must prefer $a_2$ over $a_3$ as well (our model only allows preferences that are single improvements over the original list order). Hence, $\pi(a_3|a_2)$ must be smaller than $0.4 =1-0.6$. The next axiom generalizes this idea: the total binary choice probabilities against the immediate predecessor in the list must be less than or equal to $1$.
HAOM$_{\triangleright}$ introduces some compatibility among all the preference types in the population because each preference type is a single improvement of $\triangleright$. However, it allows for a significant level of heterogeneity in attention. Our next theorem indicates that HAOM$_{\triangleright}$ can still make predictions because it states that the three axioms completely characterize HAOM$_{\triangleright}$.
Theorem (ref) establishes both a necessary and sufficient condition for HAOM$_{\triangleright}$. The importance of this theorem lies in its applicability to any dataset, even when RUM is not applicable. This makes it a powerful tool for studying choice behaviors beyond utility maximization. In contrast to RUM, our choice model enjoys strong identification for preferences in $\mathcal{P}_{\triangleright}$: the frequency of each preference type in $\mathcal{P}_{\triangleright}$ is uniquely (point) identified. Formally, we define the preference type frequency for each $\succ$, $\tau(\succ):=\tau(\{(\Gamma,\succ):\Gamma \in \mathcal{AO}_{\triangleright} \})$; that is, $\tau(\succ)$ represents the total probability of $\succ$ being the underlying preferences. If both $\tau_1$ and $\tau_2$ are HAOM$_{\triangleright}$ representations of $\pi$, then $\tau_1(\succ)=\tau_2(\succ)$. The uniqueness of HAOM$_{\triangleright}$ is in sharp contrast to the non-uniqueness of RUM.
Given the uniqueness result, we now ask whether we can reveal the frequency of specific preference types. Our revealed preference serves (at least) two purposes. First, it identifies the specific form of heterogeneity in the population in terms of preferences. Second, it provides a unique weight for each preference type within this heterogeneity. For example, by using these weights, a policymaker can evaluate how a particular policy affects each agent in the heterogeneous population and then combine these effects with precisely identified weights. Furthermore, when we interpret $\succ_{k j}$ as the status quo bias preferences, $\tau(\succ_{k j})$ is the frequency of people whose default option is $a_k$ and the strength of bias is $j-k$ (the difference between the original and final ranking of the default $a_k$). Hence, our identification result can be used to measure the status quo bias in the data. For the theorem below, we recall the adopted convention that $\succ_{11}:= \triangleright$, which corresponds to choice types whose preference coincides with the original list order.
To provide some intuition, recall from Definition (ref) that decision makers pay full attention in binary comparisons. This implies that for $j < k$, the choice probability $\pi(a_k|a_j)$ is just the frequency of choice types who rank $a_k$ at or higher than the $j$th position; that is, $\pi(a_k|a_j) = \sum_{\ell \leq j} \tau(\succ_{k\ell})$. This observation justifies the first part of the theorem. Now take $j = k-1$. Then $\pi(a_k|a_{k-1})$ is the frequency of choice types who do not agree with $\triangleright$ on the ranking of $a_k$, which leads to the second conclusion of the theorem.
The identification result provided by Theorem (ref) is based on binary choices. While it provides point identification, it can only be used when all binary comparisons $\{a_k,a_j\}$ such that $k>j$ are available in the data. When the data is incomplete, in the sense that not all possible comparisons are observed (or identifiable and estimable in econometrics language), we can provide bounds for the frequency of preference types: if a non-top alternative is chosen, it must be attributed to the types who like that alternative better than the top alternative. In this sense, the choice probability is the lower bound of those types since some of these decision maker types might not have looked far enough down the list (i.e., random and limited attention).
To close this subsection, we demonstrate how to perform revealed attention analysis on our choice model with heterogeneous preference and random attention; c.f., Section (ref). Given $\tau$, we define the attention frequency for an alternative $\phi_\tau(a|S):=\tau(\{(\Gamma,\succ):a \in \Gamma(S)\})$; that is, $\phi_\tau(a|S)$ represents the total probability of $a$ being considered in the population when the choice set is $S$. Under RUM, this is assumed to be equal to one for all alternatives in $S$ (full attention). In our model, while full attention holds for the first alternative on the list, it might not hold for other alternatives. Even though we cannot point identify the attention frequency for other alternatives, we provide upper and lower bounds. For notation convenience, we will drop the subscript and use $\phi(a|S) = \phi_\tau(a|S)$.
We first discuss how to empirically bound the frequency of different preference types. To implement the partial identification result in Proposition (ref), we fix two positions, $1\leq j < k$. We are interested in bounding the fraction of decision maker who rank $a_k$ at or above the $j$th position. In other words, we consider the following parameter of interest: $ \theta_{k j} = \sum_{\ell \leq j}\tau(\succ_{k\ell}).$ We have $\theta_{k j} \geq \theta_{k s_1} \geq \pi(a_k|S)$ for any choice problem $\langle a_{s_1},a_{s_2},\dots,a_{s_{|S|}} \rangle$ satisfying (i) $a_k \in S$, and (ii) $s_1\leq j$, where the first inequality follows from the definition of $\theta_{k j}$, and the second inequality follows by Proposition (ref).
We construct a lower bound for $\theta_{k j}$ by computing the maximum over a collection of adjusted empirical choice probabilities. (The same idea has been used to empirically bound the attention frequency in Section (ref) and Theorem (ref).) Let $\mathcal{D}_{kj} = \{ S: a_k\in S,\ s_1\leq j \}$. We define
As before, $\widehat{\sigma}(a_k|S) = \sqrt{\widehat{\pi}(a_k|S)(1-\widehat{\pi}(a_k|S))/N_S}$ is the standard error of the estimated probability $\widehat{\pi}(a_k|S)$, and the critical value is $\mathrm{cv}(\alpha,{\underline{\theta}}_{k j}) = \inf\{ t:\ \mathbb{P}\big[ \max(\mathbf{z}) \leq t \big] \geq 1-\alpha \}$, where $\mathbf{z}$ is a standard normal random vector of dimension $|\mathcal{D}_{k j}|$. The statistical validity of the proposed lower bound is guaranteed by the following theorem.
We now discuss how to bound the attention frequency using the results of Theorem (ref). We will illustrate with the lower bound since the upper bound follows analogously. Fix some choice problem $S$, and some option $a_k\in S$ which differs from the top-listed one. As before, $\mathcal{D}$ is the collection of choice problems available in the data. We form the lower bound as
where $\widehat{\pi}(U_{\triangleright}(a_k)|R)$ is the empirical probability of choosing an option listed before $a_k$ in $R$, and $\widehat{\sigma}(U_{\triangleright}(a_k)|R) = \sqrt{\widehat{\pi}(U_{\triangleright}(a_k)|R)(1-\widehat{\pi}(U_{\triangleright}(a_k)|R))/N_R}$ is its standard error. The critical value is constructed as $\mathrm{cv}(\alpha,\underline{{\phi}}(a_k|S)) = \inf\left\{ t:\ \mathbb{P}\big[ \max(\mathbf{z}) \leq t \big] \geq 1-\alpha \right\}$, where $\mathbf{z}$ is a standard normal random vector of dimension $|\{R\in\mathcal{D}:R\supseteq S\}|$, the number of supersets of $S$.
To showcase the performance of our proposed econometric methods, we report the results of a simulation study on preference elicitation in the supplemental appendix. Our numerical findings illustrate how the lower bound on preference distribution (i.e., $\theta_{kj}$ above) changes as we vary the effective sample size and the number of choice problems available in the data.
\setstretch{1.4}